Information processing device, information processing method, computer program, and system
The system addresses complex camera settings and integration of image editing behaviors by using a learning model to automatically adjust camera settings based on user data, ensuring high-quality, preference-matching images with reduced complexity and costs.
Patent Information
- Application Number
- PCT/JP2025/003293
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-29
- Filing Date
- 2025-01-31
- Publication Date
- 2025-10-02
AI Technical Summary
Existing camera systems struggle with complex and time-consuming setting operations for capturing images that reflect photographers' intentions and preferences, and there is a lack of efficient methods to integrate image editing behaviors into the shooting process, leading to inconsistent results and high computational and storage costs.
A system that utilizes a learning model trained on metadata and shooting/editing settings to estimate appropriate camera settings, allowing for automatic adjustment based on past user tendencies, reducing the need for manual input and simplifying the setting process.
Enables cameras to generate images that reflect users' intentions and preferences without complex manual settings, reducing operational complexity and costs while ensuring high reproducibility of edited images.
Smart Images

Figure JP2025003293_02102025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, computer program, and system
[0001] The technology disclosed in this specification (hereinafter referred to as "the present disclosure") relates to an information processing device and information processing method, a computer program, an imaging device, and a system that perform processing related to camera imaging.
[0002] So-called "digital data," which uses semiconductor image sensors to convert light received from the outside world into digital data and record it, has become widespread. When taking pictures with a digital camera, camera settings such as aperture value and shutter speed are typically set using, for example, a user interface on the camera body. It is also common to adjust brightness, color temperature, color correction, contrast, and other aspects of images taken with a digital camera using an editing application. Recently, many digital cameras have an auto-shooting function that automatically controls camera settings based on the detection values of a sensor built into the camera. For example, numerous technologies have been proposed for focus control, which automatically focuses on a subject (see, for example, Patent Document 1).
[0003] WO2023 / 286301 JP2020-39123A, paragraphs 0020-0030
[0004] An object of the present disclosure is to provide an information processing device and information processing method for performing processing related to camera settings, a computer program, an imaging device, a system, a camera system, and a camera device that simplify setting operations.
[0005] The present disclosure has been made in consideration of the above-mentioned problems, and a first aspect thereof is an information processing device including: a collection unit that collects metadata and shooting setting values at the time of shooting and editing setting values at the time of image editing of the shot image, linked to an image shot by a camera, and stores the collected metadata in storage; a generation unit that trains a model to learn the correlation between the metadata, shooting setting values, and editing information stored in the storage, and generates a learning model that estimates the shooting setting values and editing information from the metadata; and a transfer unit that transfers the generated learning model to the camera.
[0006] The metadata includes at least one of first information that analyzes the characteristics of the subject scene, second information that digitizes the situation at the time of shooting, and third information that is a reduced image of the subject scene.
[0007] The generation unit trains a model to learn the correlation between metadata, shooting settings, and editing information based on a specific user's shooting and editing operations, and generates a learning model that reflects the user's past shooting and image editing tendencies.
[0008] A second aspect of the present disclosure is an information processing method including: a step of collecting metadata and shooting settings at the time of shooting and editing information at the time of image editing of the shot image, linked to an image shot by a camera, and storing the collected metadata in a storage; a generation step of training a model to learn the correlation between the metadata, shooting settings, and editing information stored in the storage, and generating a learning model that estimates the shooting settings and editing information from the metadata; and a transfer step of transferring the generated learning model to the camera.
[0009] Furthermore, a third aspect of the present disclosure is a computer program written in a computer-readable format to cause a computer to function as: a collection unit that collects metadata and shooting settings at the time of shooting and editing information at the time of image editing of the shot image, linked to the image shot by the camera, and stores them in storage; a generation unit that trains a model to learn the correlation between the metadata, shooting settings, and editing information stored in the storage, and generates a learning model that estimates the shooting settings and editing information from the metadata; and a transfer unit that transfers the generated learning model to the camera.
[0010] A computer program according to a third aspect of the present disclosure defines a computer program written in a computer-readable format to perform predetermined processing on a computer. The computer program can be provided to a computer capable of executing various program codes in a computer-readable format via a storage medium or communication medium, such as an optical disk, a magnetic disk, or a semiconductor memory, or a communication medium such as a network. By installing the computer program according to the third aspect of the present disclosure on a computer via any of these media, a cooperative effect is exerted on the computer, and the same effects as those of the information processing device according to the first aspect of the present disclosure can be obtained.
[0011] A fourth aspect of the present disclosure is an imaging device having a photographing function unit, the imaging device comprising: a generation unit that analyzes a subject scene to generate metadata; and an estimation unit that estimates setting values from the metadata using a learning model that has learned the correlation between the metadata and setting values; and the imaging device controls the photographing function unit using the setting values to perform a photographing operation.
[0012] The estimation unit estimates shooting setting values and editing information from the metadata using the learning model. Also, the imaging device according to a fourth aspect further includes a conversion unit that converts the editing information into parameter values that can be handled within the imaging device, and controls the imaging function unit using the shooting setting values and the parameter values to perform a shooting operation.
[0013] In addition, the imaging device according to the fourth aspect further includes a monitoring function unit that generates a monitoring image of the shooting function unit, and controls the monitoring function unit using the shooting setting values and the parameter values to display a monitoring screen.
[0014] Furthermore, a fifth aspect of the present disclosure is a system including a camera, an editing device that edits images captured by the camera, and a model providing device that provides a learning model to the camera, wherein the camera outputs metadata at the time of shooting and shooting setting values of the camera, linked to the captured image, and the editing device outputs editing information at the time of editing the captured image, and the model providing device includes a collection unit that collects the metadata, shooting setting values, and editing information and stores them in storage, a generation unit that trains a model to learn the correlation between the metadata, shooting setting values, and editing information stored in the storage, and generates a learning model that estimates the shooting setting values and editing information from the metadata, and a transfer unit that transfers the generated learning model to the camera.
[0015] However, the term "system" used here refers to a logical collection of multiple devices (or functional modules that realize specific functions), regardless of whether each device or functional module is contained within a single housing. In other words, both a single device consisting of multiple parts or functional modules and a collection of multiple devices are considered "systems."
[0016] Furthermore, a sixth aspect of the present disclosure is a camera system comprising: a camera unit having a first learning model unit; and a generation AI model unit that generates a second image based on a first image captured by the camera unit; and the first learning model is changed using a learning model learned based on the first image and the second image.
[0017] Furthermore, a seventh aspect of the present disclosure is a camera device having a machine learning model, comprising: a control unit; and a first learning model, wherein the control unit changes the first learning model based on a second learning model learned by comparing a first image captured by the camera device with a second image generated from the first image by a generation AI model unit.
[0018] FIG. 1 is a diagram showing the configuration of a system 10 to which the present disclosure is applied. FIG. 2 is a diagram showing how metadata, shooting setting values, and editing setting values for each captured image are linked and accumulated for each user. FIG. 3 is a diagram showing how a learning model is generated using a large amount of data accumulated for each user. FIG. 4 is a diagram showing the operation of the system 10 up to the time when the model providing device 13 generates a learning model. FIG. 5 is a diagram showing example shooting setting values set in the camera device 11 during shooting. FIG. 6 is a diagram showing an example of "information analyzing the characteristics of a subject scene" included in the metadata. FIG. 7 is a diagram showing an example of "information digitizing the situation at the time of shooting" included in the metadata. FIG. 8 is a diagram showing example editing setting values set in an image editing application. FIG. 9 is a diagram showing the operation of the system 10 after the camera device 11 acquires a learning model. FIG. 10 is a diagram showing the operation of the system 10 for the camera device 11 to display an image for monitoring using the learning model. FIG. 11 is a diagram showing a modified example of system operation. FIG. 12 is a diagram showing the configuration of a camera device 11 to which the present disclosure is applicable. FIG. 13 is a diagram showing the internal configuration of the image processing unit 1203. FIG. 14 is a diagram showing specific operations of the system 10 to which the camera device 11 shown in FIG. 12 is applied. FIG. 15 is a flowchart showing a processing procedure by which the camera device 11 performs a photographing operation. FIG. 16 is a flowchart showing a processing procedure by which the editing device 12 performs an editing operation on a photographed image output from the camera device 11. FIG. 17 is a flowchart showing a processing procedure by which the model providing device 13 generates a learning model. FIG. 18 is a diagram showing an example of system operation when monitoring a subject scene. FIG. 19 is a flowchart showing a processing procedure by which the camera device 11 monitors a subject scene. FIG. 20 is a diagram showing a first application example of the system 10. FIG. 21 is a diagram showing a second application example of the system 10. FIG. 22 is a diagram showing a third application example of the system 10. FIG. 23 is a diagram showing a third application example of the system 10. FIG. 24 is a diagram showing an example of the hardware configuration of an information processing device.
[0019] Hereinafter, embodiments of the present disclosure will be described in the following order with reference to the drawings.
[0020] A. Overview A-1. System configuration A-2. Example of system operation A-3. Other example of system operation A-4. Modifications A-5. Effects B. Specific examples B-1. Configuration of camera device B-2. Example of system operation B-3. Other example of system operation C. Application examples C-1. First application example C-2. Second application example C-3. Third application example D. Configuration of information processing device
[0021] A. Overview Photographers need to configure various settings on their camera equipment to capture images that reflect their intentions and preferences. For some photographers, operating these camera settings can be complex and difficult. Camera equipment that offers an automatic shooting function is configured to analyze the subject scene using various sensors, including an image sensor, and calculate and automatically configure the camera's shooting settings based on the results. However, images generated using the automatic shooting function are based on logical conditional judgments. In other words, the camera's shooting settings calculated for the same subject scene using the automatic shooting function are uniform, which can deviate from the diverse intentions and preferences of each photographer.
[0022] To begin with, there are an infinite number of types of subject scenes, and automatically calculating the camera's shooting settings based on the results of analyzing the scene requires a complex algorithm that includes numerous conditional branching decisions, which carries the risk of compromising operating quality due to program bugs, etc.
[0023] Recently, image processing technology for editing captured images has been developing. Photographers can use editing applications to process captured images to reflect their own intentions and preferences. However, as the number of captured images increases with the evolution of camera equipment, the editing process can become extremely time-consuming.
[0024] If an image could be generated at the time of capture, taking into account the behavior of image processing through editing, the above-described editing work would be unnecessary. However, setting up camera equipment to take into account the behavior of image processing through editing is complex and requires specialized knowledge. Furthermore, even if there is a means of processing an image to reflect one's intentions and preferences, the captured image and the edited image are different, and it is not possible to immediately check what the final image will look like at the time of capture. Furthermore, it is difficult for an inexperienced user to easily imitate the images captured by a user skilled in shooting and editing with camera equipment.
[0025] Editing can be facilitated by providing users with editing settings corresponding to specific pre-edited images based on past trends obtained from correlations between a large number of pre-edited images and the editing settings corresponding to each image. However, setting camera equipment in consideration of the behavior of image processing through editing remains complex and difficult. Furthermore, calculating such correlations requires a large number of pre-edited images with large data sizes, which results in problems such as high storage costs for accumulating the data and high computational costs for handling image data with a large amount of information.
[0026] Therefore, the present disclosure proposes a technology for generating images that reflect the diverse intentions and preferences of various photographers based on the photographers' past shooting and image editing tendencies. Specifically, the present disclosure trains an AI (Artificial Intelligence) model that estimates appropriate shooting settings and editing settings from metadata at the time of image capture based on the correlation between metadata from past image captures and camera shooting settings and editing settings for image editing (hereinafter, the "AI model" will also be referred to as a "learning model"). When a subsequent shooting instruction is issued, the learning model is used to set the shooting settings and editing settings estimated from the metadata, thereby performing shooting processing. Therefore, according to the present disclosure, the camera device uses the learning model to perform settings operations on the camera device taking into account the behavior of image processing through editing, thereby enabling the generation of captured images equivalent to those after editing. Note that the present disclosure may also use a "dictionary" that derives appropriate shooting settings and editing settings from metadata instead of a "learning model" that learns the correlation between metadata, shooting settings, and editing settings for image editing. In this specification, the term "learning model" is intended to include both models currently being trained and models that have already been trained.
[0027] The metadata may include, for example, the characteristics of the subject scene and the circumstances at the time of shooting, and may be acquired from image analysis results, hardware configuration information of the camera device, detection results of sensors mounted on the camera device, etc. The metadata may include reduced images obtained by reducing the size of the images captured by the camera device. In the present disclosure, the model is trained using compressed metadata, which can reduce storage costs and computational costs compared to training a model using captured images with large data sizes as they are.
[0028] A-1. System Configuration Fig. 1 shows a schematic configuration of a system 10 to which the present disclosure is applied. The illustrated system 10 is composed of a camera device 11 through which a user takes photos, an editing device 12 through which the user edits images taken with the camera device 11, and a model providing device 13 that provides a learning model that has learned the past shooting and image editing tendencies of each user. Note that, while it is generally assumed that the user of the camera device 11 and the user of the editing device 12 are the same, the users may be different.
[0029] Each time a user takes a photograph, the camera device 11 outputs the captured image along with metadata and the photographing setting values set by the user at the time of photographing. The captured image output from the camera device 11 is, for example, a RAW image before editing, but may of course be an image in another format. The metadata includes, for example, the characteristics of the subject scene and the circumstances at the time of photographing, but the specific information that constitutes this will be described in detail later. The specific parameters that constitute the photographing setting values will also be described in detail later.
[0030] When an image captured by the camera device 11 is input, the editing device 12 performs editing processing, such as image processing, on the captured image in accordance with a user's editing operations, using, for example, an image editing application. The editing device 12 then outputs the edited image as well as the editing setting values set by the user during image editing. The edited image output from the editing device 12 is, for example, in YUV format, but may be in other formats such as JPEG (Joint Photographic Experts Group). Specific parameters that constitute the editing setting values will be described in detail later. The editing device 12 may also output some metadata.
[0031] The model providing device 13 collects metadata and shooting setting values for each captured image from the camera device 11 and also collects editing setting values for the captured image from the editing device 12. As shown in FIG. 2, the model providing device 13 associates and stores the metadata, shooting setting values, and editing setting values for each captured image for each user. Then, as shown in FIG. 3, the model providing device 13 uses a large amount of data stored for each user to train a model on the correlation between the metadata and the shooting setting values and editing setting values, thereby generating a learning model that estimates shooting setting values and editing setting values from the metadata. The learning model generated by the model providing device 13 is a personalized learning model that estimates appropriate shooting setting values and editing setting values based on the target user's past shooting and image editing tendencies. The learning model generated by the model providing device 13 is provided to the target user's camera device 11. The model providing device 13 may, for example, retrain the model each time new data is accumulated and provide the retrained learning model to the camera device 11.
[0032] After acquiring the learning model, the camera device 11 can input the metadata at that time into the learning model each time a photograph is taken and control the photographing operation using the photographing setting values and editing setting values output from the learning model. In other words, even if the user does not input the photographing setting values and editing setting values one by one, the camera device 11 can generate photographed images with high reproducibility equivalent to those after editing by the editing device 12.
[0033] Furthermore, even if a user does not have specialized knowledge regarding the setting operations of the camera device 11 or image editing, by introducing a learning model generated using accumulated data of experienced users into the camera device 11, the user can imitate images equivalent to those taken and edited by an experienced user.
[0034] The camera device 11 may be a general digital camera. The editing device 12 may be, for example, an information processing device such as a personal computer (PC) with an image editing application installed, or an information terminal such as a smartphone or tablet. The camera device 11 and the editing device 12 may be physically independent devices as shown in FIG. 1 , or may be physically integrated devices (for example, an information terminal equipped with a camera function).
[0035] The model providing device 13 is, for example, an information processing device such as a PC on which a correlation learning application is installed. The correlation learning application here is an application that has a function of generating a learning model by having a model learn the correlation between information on shooting conditions and shooting setting values and editing setting values.
[0036] The model providing device 13 is, for example, a server installed on the cloud. For simplicity of explanation, only one camera device 11, one editing device 12, and one model providing device 13 are depicted in Fig. 1 , but in reality, one model providing device 13 may collect data from a large number of camera devices 11 and editing devices 12, perform correlation learning based on a huge amount of data accumulated for each user, generate a learning model for each user, and provide the learning model for each user to each of the large number of camera devices 11.
[0037] As a variant, the model providing device 13 may be an information processing device integrated with the editing device 12, or the camera equipment 11, editing device 12, and model providing device 13 may all be incorporated into a single device. The model providing device 13 integrated with the camera equipment 11 and editing device 12 generates a learning model for a specific user.
[0038] A-2. Example of System Operation Figure 4 shows the operation of the system 10 up to the point where the model providing device 13 generates a learning model.
[0039] The camera device 11 includes a photographing function unit 401. The photographing function unit 401 includes a photographing execution unit having a lens and an image sensor, and a signal processing unit that performs signal processing such as development on an image signal (RAW image) output from the image sensor to generate a YUV image (or an RGB image, or a JPEG image). To simplify the drawing, the photographing execution unit and the signal processing unit are not shown in FIG.
[0040] The user sets shooting setting values via a user interface (not shown in FIG. 4 ) of the camera device 11 each time a photograph is taken. The photographing function unit 401 then takes a photograph based on the set shooting setting values. FIG. 5 shows examples of shooting setting values that are set in the camera device 11 when a photograph is taken. The shooting setting values include an aperture value, a shutter speed, a sensitivity value, an exposure compensation value, a color temperature, and a focus area type, but it is not necessary to include all of these; conversely, the shooting setting values may include other shooting parameters. The photographing function unit 401 may take either still images or videos.
[0041] The photographing function unit 401 also generates metadata when photographing based on the photographing setting values that have been set. The metadata includes, for example, the characteristics of the subject scene and the circumstances at the time of photographing, and can be acquired from the results of image analysis, hardware configuration information of the camera device, and detection results of sensors mounted on the camera device. The metadata may include a reduced image of the photographed image. The photographing function unit generates metadata every time it receives a photographing instruction from the user.
[0042] 6 shows an example of "information analyzing the characteristics of the subject scene" included in the metadata. The "information analyzing the characteristics of the subject scene" includes information such as "frequency values obtained by classifying the brightness values of each pixel in the image into five levels," "ratio of the values of the brightest pixel and the darkest pixel," "level of the brightness of the pixel at the focus position among the five levels," "frequency values obtained by classifying the colors of each pixel in the image into five colors," "level of the color of the pixel at the focus position among the five levels," "whether or not a moving object is present in the subject scene," and "distance information about the subject." However, it is not necessary to include all of these information, and other information may also be included. Although not shown in FIG. 6, the metadata may also include a reduced image obtained by reducing the image captured by the camera device.
[0043] 7 shows an example of "information that digitizes the situation at the time of shooting" included in the metadata. "Name of the photographing device," "Name of the lens used," "Date and time of shooting," "Ambient illuminance value," "Camera tilt (or camera shake)," etc. are included in the "information that digitizes the situation at the time of shooting," but it is not necessary to include all of these, and other information may also be included. The camera device 11 is assumed to have a metadata acquisition unit (not shown) that acquires this metadata and stores it in a file every time the photographing function unit performs a photographing operation.
[0044] When the camera device 11 creates a file of the captured image, it also stores the metadata and shooting setting values used in capturing the image in the file. The metadata and shooting setting values may be stored in the same file as the captured image, or may be stored in a separate file from the captured image if they are linked to the captured image. In Figure 4, the file group of captured images output from the camera device 11 is indicated by reference numeral 411, and the file groups of metadata and shooting setting values linked to each captured image are indicated by reference numerals 412 and 413, respectively.
[0045] 4 , the editing device 12 inputs a captured image 411 from the camera device 11 and performs editing processing, such as image processing, on the captured image in accordance with a user's editing operation using an image editing application 402. However, the image editing application 402 may operate within the camera device 11 rather than the editing device 12. The image editing application 402 processes the captured image input from the camera device 11 based on editing setting values set by the user. FIG. 8 shows examples of editing setting values set in the image editing application 402. The editing setting values include brightness, color temperature adjustment, color correction, contrast, white level, black level, and saturation, but it is not necessary to include all of these; conversely, other editing parameters may be included in the editing setting values.
[0046] When the image edited by the image editing application 402 is converted into a file and output, the editing setting values are also stored in a file linked to the original captured image. In Fig. 4, the group of files of the edited image output from the image editing application 402 is indicated by reference numeral 414, and the group of files of the editing setting values linked to the original captured image is indicated by reference numeral 415. Note that the editing device 12 may store some of the metadata used in image editing in the above-mentioned metadata file.
[0047] The model providing device 13 collects file groups 412, 413, and 415 of metadata, shooting setting values, and editing setting values that are respectively linked to the original captured images and stored, and stores the files in the storage 403 by linking them to the captured images for each user (see FIG. 2). The model providing device 13 may collect the file groups 412, 413, and 415 periodically, or may perform the operation in response to an instruction from the user.
[0048] Then, within the model providing device 13, the model learning unit 404 trains a model to learn the correlation between the metadata, shooting settings, and editing settings associated with each captured image for a specific user, and generates a learning model 405 that estimates shooting settings and editing settings from the metadata, and saves the model as a file. The learning model 405 generated by the model providing device 13 is a model for a specific user that estimates appropriate shooting settings and editing settings for each subject scene based on the specific user's past shooting and image editing tendencies. The model providing device 13 transfers the learning model 405 to the camera device 11 of the specific user. The model providing device 13 may transfer the file of the learning model 405 in response to a download request from the camera device 11, for example, or may push the file of the learning model 405.
[0049] The model providing device 13 performs the process of generating a learning model using compressed metadata, thereby reducing storage costs and computational costs compared to training a model using captured images with large data sizes as they are.
[0050] Fig. 9 shows the operation of the system 10 after the camera device 11 acquires the learning model. However, Fig. 9 assumes that the learning model generated for a specific user has already been transferred to the camera device 11 in accordance with the operation of the system 10 shown in Fig. 4, and that the learning model is available on the camera device 11 side.
[0051] When the camera device 11 receives a shooting instruction from a user, the shooting function unit 401 analyzes the subject scene and generates metadata as indicated by reference numeral 417. This metadata 417 is input into the learning model 405 provided by the model providing device 13, which estimates shooting setting values and editing setting values as indicated by reference numerals 418 and 419, respectively. The editing setting values 419 are parameter values for the image editing application 402, and are converted into parameter values in a format that can be handled within the camera device 11 as indicated by reference numeral 420. The shooting function unit 401 then sets the aperture value, shutter speed, sensitivity value, exposure compensation value, color temperature, and focus area type in the shooting execution unit (described above) based on the shooting setting values 418 estimated using the learning model 405, and performs a shooting operation, and sets the converted parameter values 420 in the signal processing unit (described above) to perform signal processing of the image signal (such as developing a RAW image).
[0052] In this way, by using the learning model 405 provided by the model providing device 13, the camera equipment 11 can generate, according to the subject scene, an image 421 having a visual effect that reflects the visual effect trends of a group of images 414 that the same user has previously generated by taking and editing images, simply by the user performing a shooting operation (in other words, without performing image editing).
[0053] Note that even after acquiring the learning model 405 from the model providing device 13, the camera device 11 stores the metadata and shooting setting values in a file linked to the captured image each time a photograph is taken using the shooting setting values set by the user. Also, each time a captured image is edited by the user using the editing device 12, the editing setting values are stored in a file linked to the original captured image. Also, even after transferring the learning model 405 to the camera device 11, the model providing device 13 re-trains the model using the newly collected file groups of metadata, shooting setting values, and editing setting values, and re-transfers the file of the re-trained learning model to the camera device 11.
[0054] A-3. Other Examples of System Operation This section A-3 describes the system operation when monitoring the subject scene through the viewfinder before shooting.
[0055] Fig. 10 shows the operation of the system 10 for the camera device 11 to display an image for monitoring using a learning model. However, Fig. 10 assumes that a learning model generated for a specific user has already been transferred to the camera device 11 in accordance with the operation of the system 10 shown in Fig. 4, and that the learning model is available on the camera device 11 side.
[0056] In addition to the photographing function unit 401 , the camera device 11 further includes a viewfinder 408 and a monitoring function unit 407 that generates a monitoring image to be displayed on the viewfinder 408 .
[0057] The monitoring function unit 407 constantly analyzes the subject scene captured by the shooting function unit 401 in real time to generate metadata 417. This metadata 417 is input to the learning model 405 to estimate shooting setting values 418 and editing setting values 419. The editing setting values 419 are parameter values for the image editing application 402, and are therefore converted into parameter values 420 in a format that can be handled within the camera device 11.
[0058] When the monitoring function unit 407 acquires a live view image from the shooting function unit 401 based on the shooting setting values 418 estimated using the learning model 405, it generates a monitoring image equivalent to the live view image after editing by the image editing application 402 based on the parameter values 420 converted from the editing setting values, and displays the monitoring image on the viewfinder 408.
[0059] In this way, by using the learning model 405 provided by the model providing device 13, the camera equipment 11 can display on the viewfinder 408 a monitoring image having a visual effect that reflects the visual effect trends of a group of images previously taken and generated by the same user through image editing, depending on the subject scene at the time of monitoring.
[0060] A-4. Modifications Section A-2 above described the operation of the system 10, which generates a learning model that learns the correlation between shooting setting values and editing setting values and metadata, and utilizes the learning model on the camera device 11. Section A-4 describes, as a modification, a modification in which shooting setting values are not used in model learning, but a learning model that learns the correlation between only editing setting values and metadata is generated, and the learning model is utilized on the camera device 11, with reference to FIG.
[0061] Each time a user takes a photograph, the user sets photographing setting values via a user interface (not shown in FIG. 11 ) of the camera device 11. Then, the photographing function unit 401 takes a photograph based on the set photographing setting values. Note that the photographing function unit 401 may take either a still image or a video.
[0062] When the camera device 11 converts a captured image into a file, it also stores metadata about the image capture in the file. The metadata may be stored in the same file as the captured image, or may be stored in a separate file from the captured image if it is linked to the captured image. The group of captured image files output from the camera device 11 is indicated by reference numeral 411, and the metadata linked to each captured image is indicated by reference numeral 412. Unlike the system operation example shown in FIG. 4, the shooting setting values for each shooting operation are not saved in a file.
[0063] The editing device 12 inputs a captured image 411 from the camera device 11 and performs editing processing, such as image processing, on the captured image in accordance with a user's editing operations using an image editing application 402. When the editing device 12 files the edited image and outputs it, it also stores editing setting values in a file linked to the original captured image. The group of files of the edited images output from the image editing application 402 is indicated by reference numeral 414, and the group of files of the editing setting values linked to the original captured image is indicated by reference numeral 415. Note that the editing device 12 may store some of the metadata used in image editing in the metadata file.
[0064] The model providing device 13 collects the metadata file groups 412 and 415 that are linked to the original photographed images and stored, and stores the files 412 and 415 linked to the photographed images for each user in the storage 403. The model providing device 13 may collect the file groups 412 and 415 periodically, or may perform the collection operation in response to an instruction from the user.
[0065] Then, within the model providing device 13, a model learning unit 1101 trains a model to learn the correlation between metadata and editing setting values associated with each captured image for a specific user, generates a learning model 1103 that estimates editing setting values from the metadata, and saves the model as a file. The learning model generated by the model providing device 13 is a model for a specific user that estimates appropriate editing setting values for each subject scene based on the specific user's past image editing tendencies. The model providing device 13 transfers a file 1104 of the generated learning model to the camera device 11 of the specific user. The model providing device 13 may transfer the learning model file 1104 in response to a download request from the camera device 11, for example, or may push-distribute the learning model file 1104.
[0066] On the camera device 11 side, after acquiring the learning model 1103, upon receiving a shooting instruction from the user, the shooting function unit 401 analyzes the subject scene and generates metadata 417. This metadata 417 is input into the learning model 1103 to estimate editing setting values 1105. Since the editing setting values 1105 are parameter values for the image editing application 402, they are converted into parameter values 1106 in a format that can be handled within the camera device 11. Then, the shooting function unit 401 sets shooting setting values such as aperture value, shutter speed, sensitivity value, exposure compensation value, color temperature, and focus area type based on the shooting instruction from the user, performs shooting operations, and also performs signal processing of the image signal (such as developing a RAW image) using the converted parameter values 1106.
[0067] In the system operation shown in FIG. 11, the user needs to set shooting settings each time a shooting instruction is issued, but even without performing image editing, it is possible to generate an image with visual effects that reflect the visual effect trends of a group of images previously generated by the same user through shooting and image editing, depending on the subject scene.
[0068] A-5. Effects This section A-5 summarizes the effects provided by the system 10 to which the present disclosure is applied. However, the effects described here are merely examples, and the effects provided by the present disclosure are not limited to these. Furthermore, the present disclosure may provide additional effects in addition to the effects described above. Further objects, features, and advantages of the present disclosure will be apparent from the embodiments described in this specification and a more detailed description based on the accompanying drawings.
[0069] In the system 10, once the camera device 11 has acquired the learning model, the user can take images that reflect the user's intentions and preferences according to the subject scene, without having to perform complex setting operations for shooting settings each time a photo is taken, and without having to process the captured image using the image editing application 402.
[0070] Furthermore, in the system 10, when the user monitors the subject scene before shooting using the viewfinder of the camera device 11, the monitoring image can be displayed that reflects the user's intentions and preferences according to the subject scene, without the user having to perform complex operations to set shooting settings. Therefore, the user can shoot while checking an image equivalent to the final image processed using the image editing application 402 before shooting.
[0071] Furthermore, the system 10 allows an inexperienced user to imitate images equivalent to those taken and edited by an experienced user by applying the learning model generated through the shooting and image editing of an experienced user to his or her own camera device 11.
[0072] Furthermore, the system 10 utilizes a learning model generated through the user's photography and image editing, and controls the camera equipment 11 using a simple algorithm that does not include numerous conditional branching decisions, thereby reducing the risk of impairing operational quality due to program bugs.
[0073] B. Specific Example B-1. Configuration of Camera Device Fig. 12 schematically shows the configuration of a camera device 11 to which the present disclosure can be applied. The illustrated camera device 11 includes an imaging optical system 1201, an imaging unit 1202, an image processing unit 1203, a display unit 1204, a recording unit 1205, an operation unit 1206, a control unit 1207, a sensor unit 1208, and a bus 1209.
[0074] The imaging optical system 1201 is configured by combining multiple optical lenses (neither of which is shown), such as a focus lens and a zoom lens. The imaging optical system 1201 drives the focus lens and the zoom lens based on control signals from the image processing unit 1203 and the control unit 1207, and forms an optical image of a subject on the imaging surface of the imaging unit 1202. The imaging optical system 1201 also includes an iris mechanism, a shutter mechanism, and the like.
[0075] The imaging unit 1202 is configured using an image sensor such as a CMOS (Complementary Metal Oxide Semiconductor) or a CCD (Charge Coupled Device). The imaging unit 1202 performs photoelectric conversion to generate an analog imaging signal corresponding to an optical image of a subject. The imaging unit 1202 also performs CDS (Correlated Double Sampling) processing, AGC (Auto Gain Control) processing, and AD (Analog to Digital) conversion processing on the analog imaging signal, and outputs the captured image as a digital image signal.
[0076] The image processing unit 1203 generates a captured image by performing development processing on the digital image signal and further processing on the developed digital image signal. A detailed description of the image processing unit 1203 will be given later.
[0077] The display unit 1204 is a display device configured, for example, with an LCD (Liquid Crystal Display), a PDP (Plasma Display Panel), an organic EL (Electro Luminescence) panel, etc. The display unit 1204 is mainly used as a viewfinder, and displays live view images during shooting, as well as still images and videos recorded in the recording unit 1205. The display unit 1204 also functions as a user interface, with a touch panel superimposed on the screen, and displays a menu screen and accepts touch operations on the screen.
[0078] The recording unit 1205 is configured using a recording medium such as a hard disc drive (HDD) or a solid state drive (SSD). The recording medium may be fixedly provided within the camera device 11 or may be detachably attached to the camera device 11. The recording unit 1205 records the captured image processed by the image processing unit 1203 on the recording medium. For example, the recording unit 1205 records an image signal of a still image in a compressed file format based on a predetermined standard such as JPEG, and also records information related to the recorded image (e.g., EXIF (Exchangeable Image File Format) data including additional information such as the date and time of shooting) in association with the image. For example, at least a portion of metadata (see FIGS. 6 and 7 ) may be recorded in the recording unit 1205 as Exif related to the captured image. Specifically, a reduced image of the captured image may be recorded as an Exif file. The recording unit 1205 records the moving image signal in a compressed file format based on a predetermined standard, such as MPEG2 (Moving Picture Experts Group 2). The compression and decompression of the image signal may be performed by the recording unit 1205 or the image processing unit 1203. The recording unit 1205 also stores the learning model transferred from the model providing device 13.
[0079] The operation unit 1206 includes, for example, a power button for turning the power of the camera device 11 body on / off, a release button for issuing an instruction to start recording a captured image, an operator for zoom adjustment, and a touch panel integrated with the screen of the display unit 1204. The operation unit 1206 accepts operations from the user (photographer), generates operation signals according to the user operations, and outputs the signals to the control unit 107. Shooting setting values and some editing setting values can be input via the operation unit 1206.
[0080] The control unit 1207 is composed of a CPU (Central Processing Unit), RAM (Random Access Memory), ROM (Read Only Memory), etc. The ROM stores information about the configuration of the camera device 11 itself, such as the name of the camera device 11 and the name of the lens used in the imaging optical system 1201, information about part of the metadata, and programs that are read and executed by the CPU in a non-volatile manner. The RAM is used as a work memory when the CPU executes programs. The CPU executes various processes and issues commands according to the programs stored in the ROM, thereby comprehensively controlling the operation of each unit so that the camera device 111 operates in accordance with user operations accepted by the operation unit 1206.
[0081] In this embodiment, the ROM stores a learning model (described above) that estimates appropriate shooting settings and editing settings from metadata such as the characteristics of a subject scene and the situation at the time of shooting, a control program for estimating appropriate shooting settings and editing settings from the characteristics of a subject scene and the situation at the time of shooting when shooting is instructed using the learning model and performing shooting processing, and a control program for converting and saving data (learning data) such as the metadata and shooting settings required to generate the learning model into a file. Then, when shooting is instructed via the operation unit 1206, the CPU comprehensively controls the shooting operation of the camera device 11 based on the shooting settings and editing settings estimated using the learning model.
[0082] The programs executed by the control unit 1207 may be pre-installed at the time of product shipment, or may be updated as needed after shipment by download or via a recording medium. The ROM may be an electrically erasable programmable ROM (EEPROM), which can be electrically rewritten, so that the stored programs and other data (such as learning models) can be updated.
[0083] The sensor unit 1208 includes various sensors for obtaining information necessary for controlling the shooting operation and other operations within the camera device 11. For example, the sensor unit 1208 includes a clock for measuring the date and time, a GPS (Global Positioning System) sensor for receiving GPS signals, an illuminance sensor, an IMU (Inertial Measurement Unit) for detecting the tilt and camera shake of the camera device 11 body, and the like. Any type of sensor element may be incorporated into the sensor unit 1208.
[0084] The bus 1209 electrically interconnects the above-mentioned components, enabling transmission and reception of image signals, control signals, etc. The bus 1209 may include multiple types of buses.
[0085] Note that a stacked CMOS sensor with a multi-layer structure may be used for the imaging unit 1202 (see, for example, Patent Document 2). When a stacked CMOS is used, a pixel unit including a pixel array is formed on a first layer, and logic units such as the image processing unit 1203 and memories usable as at least a part of the storage area of the recording unit 1205 are formed on the second and subsequent layers.
[0086] 13 is a schematic diagram showing the internal configuration of the image processing unit 1203. The image processing unit 1203 shown in the figure includes a development processing unit 1301, a post-processing unit 1302, a detection processing unit 1303, and a focus control unit 1304.
[0087] The development processing unit 1301 performs clamping, defect correction, demosaic, and other processes on the digital image signal supplied from the imaging unit 1202. The development processing unit 1301 performs clamping to set the signal level of a pixel in a black state where no light is present to a predetermined level in the image signal generated by the imaging unit 1202. The development processing unit 1301 also performs defect correction to correct the pixel signal of a defective pixel using, for example, pixel signals of surrounding pixels. Furthermore, the development processing unit 1301 performs demosaic processing to generate an image signal in which each pixel represents each YUV component (Y is luminance, U is the difference between luminance and blue, and V is the difference between luminance and red) from the pixel signal generated by the imaging unit 1202, in which each pixel represents one color component.
[0088] The post-processing unit 1302 performs noise removal, white balance adjustment, gamma adjustment, distortion adjustment, image stabilization, image quality improvement processing, image enlargement / reduction processing, encoding processing, and the like as necessary on the image signal after development processing by the development processing unit 1301. For example, the post-processing unit 1302 performs white balance processing using a white balance coefficient calculated by the detection processing unit 1303. The post-processing unit 1302 may also perform some of the post-processing on the image signal based on shooting setting values and editing setting values input via the operation unit 1206. The post-processing unit 1302 then outputs an image signal used to display an image (live view image) to the display unit 1204, and outputs an image signal used to record the image to the recording unit 1205.
[0089] The detection processing unit 1303 calculates a white balance coefficient based on the luminance value and integral value using a detection area portion set automatically or by a user out of the image signal after development processing by the development processing unit 1301, and outputs the white balance coefficient to the post-processing unit 1302. The detection processing unit 1303 may also determine the exposure state based on the image signal of the detection area, generate an exposure control signal based on this determination result, and output the signal to the imaging optical system 1201, thereby driving the aperture or the like so that the subject in the detection area is appropriately bright.
[0090] The focus control unit 1304 generates a control signal to focus on a predetermined subject based on the image signal after development processing by the development processing unit 1301, and outputs the control signal to the imaging optical system 1201. For example, the focus control unit 1304 performs focus control by generating a depth map of the imaging area from the image signal after development processing by the development processing unit 1301, or by calculating a depth value of the predetermined subject. For example, the depth map may be generated using a deep learning DNN model, or may be generated using other means.
[0091] B-2. Example of System Operation Fig. 14 shows a specific operation of the system 10 to which the camera device 11 shown in Fig. 12 is applied.
[0092] The control unit 1207 of the camera equipment 11 operates, in addition to control software (SW) for shooting operations based on shooting setting values, indicated by reference numeral 1401, a metadata generation algorithm, indicated by reference numeral 1402, that generates metadata that digitizes the characteristics of a subject scene and the shooting situation; a filing algorithm, indicated by reference numeral 1403, that files the metadata and shooting setting values and stores them linked to the captured image; an estimation algorithm, indicated by reference numeral 1404, that estimates shooting setting values and editing setting values from metadata using a learning model (or a dictionary lookup algorithm that derives shooting setting values and editing setting values from metadata using a dictionary); a conversion algorithm, indicated by reference numeral 1405, that converts editing setting values into parameter values that can be handled by the camera equipment 11; and a control algorithm, indicated by reference numeral 1406, that controls shooting operations via the control software 1401 based on the shooting setting values and the converted parameter values.
[0093] An editing application is executed on the CPU of the editing device 12. The editing application processes the captured image (RAW image) generated by the camera device 11 to generate an edited image (e.g., a YUV image, an RGB image, a JPEG image, etc.), further processes the edited image after image processing by the camera device 11, or reprocesses an edited image that has already been processed, thereby outputting the edited image. The editing application also includes a filing algorithm, indicated by the reference numeral 1411, that files the editing setting values used to edit the captured image and links the file to the original captured image.
[0094] The CPU of the model providing device 13 also operates a file collection algorithm denoted by reference numeral 1421 and a model learning algorithm denoted by reference numeral 1422. The file collection algorithm 1421 collects file groups of metadata, shooting setting values, and editing setting values that have been filed by the filing algorithms 1403 and 1411, links them to the original captured images for each user, and stores them in the storage 403. The model learning algorithm 1422 utilizes the file groups stored in the storage 403 to learn a model for a specific user using the correlation between the metadata, shooting setting values, and editing setting values that are linked to each captured image, and generates a learning model that estimates the shooting setting values and editing setting values from the metadata.
[0095] FIG. 15 shows in the form of a flowchart the processing procedure for the camera device 11 to perform a photographing operation in the system shown in FIG.
[0096] First, the user inputs shooting setting values using the operation unit 1206 or the like (step S1501).
[0097] The control software 1401 controls the imaging optical system 1201 and the imaging unit 1202 based on the imaging setting values input in step S1501, and generates a captured image (RAW image) (step S1502).
[0098] Next, the metadata generation algorithm 1402 controls the image processing unit 1203 to calculate information obtained by analyzing the characteristics of the subject scene from the captured image, such as a RAW image, and converts information about the situation at the time of shooting into data to generate metadata (step S1503).
[0099] Then, the filing algorithm 1403 files the shooting setting values acquired in step S1501 and the metadata generated in step S1503, and stores the files in association with the currently shot image (step S1504).
[0100] Next, the camera device 11 checks whether or not it has already acquired a learning model, that is, whether or not a learning model exists in the recording unit 1205 (step S1505).
[0101] If the camera device 11 has already acquired a learning model (Yes in step S1505), the estimation algorithm 1404 uses the learning model to estimate the shooting setting values and editing setting values from the metadata generated in step S1503 (step S1506).
[0102] Next, the conversion algorithm 1405 converts the edit setting values estimated in step S1506 into parameter values that can be handled by the camera device 11 (step S1507).
[0103] Then, the control algorithm 1406 controls the imaging optical system 1201 and the imaging unit 1202 via the control software 1401 based on the shooting setting values estimated in step S1506 to perform shooting and generate a captured image (such as a RAW image) (step S1508).
[0104] Next, the control algorithm 1406 controls the image processing unit 1203 via the control software 1401 based on the parameter values converted in step S1507 to develop the captured image such as the RAW image generated in step S1508, and generates a processed image (for example, a YUV image, an RGB image, a JPEG image, etc.) having a visual effect that reflects the tendency of the visual effect of a group of images generated by the same user through previous capture and image editing (step S1509), thereby completing the capture operation.
[0105] On the other hand, if the camera equipment 11 has not yet acquired a learning model (No in step S1505), the control software 1401 controls the image processing unit 1203 based on the shooting setting values input in step S1501, develops the captured image such as the RAW image generated in step S1502, generates a processed image (e.g., a YUV image, an RGB image, a JPEG image, etc.) (step S1510), and terminates the shooting operation.
[0106] FIG. 16 shows, in the form of a flowchart, the processing procedure in which the editing device 12 performs an editing operation on the photographed images output from the camera device 11 in the system shown in FIG.
[0107] The image editing application running on the CPU of the editing device 12 performs editing processing using an image (RAW image) captured by the camera device 11 or a processed image (e.g., a YUV image, an RGB image, a JPEG image, etc.) and outputs an edited image (e.g., a YUV image, an RGB image, a JPEG image, etc.) (step S1601). Here, the image editing application processes the RAW image or YUV image captured by the camera device 11 based on editing setting values (see FIG. 8 ) such as brightness, color temperature adjustment, color correction, contrast, white level, black level, and saturation input by the user.
[0108] Then, the filing algorithm 11 files the edit setting values used by the image editing application to edit the captured image in step S1601, associates the file with the original captured image, and saves the file (step S1602), thereby completing the editing operation.
[0109] FIG. 17 shows in the form of a flowchart the processing procedure for the model providing device 13 to generate a model in the system shown in FIG.
[0110] The file collection algorithm 1421 collects files of metadata and shooting setting values that have been filed by the filing algorithm 1403 in the camera equipment 11, and files of editing setting values that have been filed by the filing algorithm 1411 in the editing device 12, and links them to the captured images for each user (see Figure 2) and stores them in the storage 403 (step S1701).
[0111] The model learning algorithm 1422 does not perform model learning for a user until the number of files stored in the storage 403 for that user reaches a certain number (No in step S1702), and returns to step S1701 to repeat the collection of file groups.
[0112] On the other hand, when the number of files stored in storage 403 for a certain user reaches a certain number (Yes in step S1702), the model learning algorithm 1422 utilizes the group of files stored in storage 403 to learn a model for that user using the correlation between the metadata linked to each captured image and the shooting setting values and editing setting values, and generates a learning model that estimates appropriate shooting setting values and editing setting values from the metadata that reflect the user's past shooting and image editing tendencies, and saves the model as a file (step S1703).
[0113] The model providing device 13 then transfers the learning model file to the camera device 11 of the corresponding user (step S1704), and ends this process. The model providing device 13 may transfer the learning model file 416 in response to a download request from the camera device 11, for example, or may push-distribute the learning model file.
[0114] B-3. Other System Operations In this section B-3, the system operations when monitoring a subject scene on the display unit 1201 (viewfinder) in the camera device 11 will be described.
[0115] Fig. 18 shows an example of system operation when monitoring a subject scene. However, it is assumed that the camera device 11 has already acquired a learning model from the model providing device 13, and the editing device 12 and the model providing device 13 are not shown in Fig. 18.
[0116] The control unit 1207 of the camera equipment 11 operates, in addition to control software (SW) for shooting operations based on shooting setting values, indicated by reference numeral 1801, a metadata generation algorithm for generating metadata digitizing the characteristics of the subject scene and the shooting situation, indicated by reference numeral 1802, an estimation algorithm for estimating shooting setting values and editing setting values from metadata using a learning model, indicated by reference numeral 1803 (or a dictionary lookup algorithm for deriving shooting setting values and editing setting values from metadata using a dictionary), a conversion algorithm for converting editing setting values into parameter values that can be handled by the camera equipment 11, indicated by reference numeral 1804, and a control algorithm for controlling shooting operations and display operations on a monitor screen via the control software 1801 based on the shooting setting values and the converted parameter values, indicated by reference numeral 1805.
[0117] FIG. 19 shows in the form of a flowchart the processing procedure for the camera device 11 to monitor a subject scene.
[0118] First, the user inputs shooting setting values using the operation unit 1206 or the like (step S1901). The control software 1801 controls the imaging optical system 1201 and the imaging unit 1202 based on the shooting setting values input in step S1501 to generate a RAW image (step S1902).
[0119] Next, the metadata generation algorithm 1802 controls the image processing unit 1203 to calculate information obtained by analyzing the characteristics of the subject scene from the captured image, such as a RAW image, and further generates information on a reduced version of the captured image, while converting information about the situation at the time of shooting into data to generate metadata (step S1903).
[0120] If the display of the monitoring image on the display unit (viewfinder) 1207 is to continue (Yes in step S1904), the camera equipment 11 then checks whether or not the learning model has already been acquired, i.e., whether or not a learning model exists in the recording unit 1205 (step S1905).
[0121] If the camera device 11 has already acquired a learning model (Yes in step S1905), the estimation algorithm 1803 uses the learning model to estimate the shooting setting values and editing setting values from the metadata generated in step S1903 (step S1906).
[0122] Next, the conversion algorithm 1804 converts the edit setting values estimated in step S1906 into parameter values that can be handled by the camera device 11 (step S1907).
[0123] Then, the control algorithm 1805 controls the imaging optical system 1201 and the imaging unit 1202 via the control software 1801 based on the imaging setting values estimated in step S1906 to perform imaging and generate a captured image such as a RAW image (step S1908).
[0124] Next, the control algorithm 1805 controls the image processing unit 1203 via the control software 1801 based on the parameter values converted in step S1907 to develop the captured image such as the RAW image generated in step S1908, and generates a processed image (e.g., a YUV image, an RGB image, a JPEG image, etc.) having a visual effect that reflects the visual effect trends of a group of images previously captured and generated by the same user through image editing (step S1909).
[0125] Next, the control algorithm 1805 controls the display unit 1207 via the control software 1801 based on the shooting setting values estimated in step S1906 to display the processed image (e.g., YUV image, RGB image, JPEG image, etc.) generated in step S1909 (step S1910). Thereafter, the process returns to step S1904, and the same processes as above are repeatedly executed.
[0126] On the other hand, if the camera equipment 11 has not yet acquired a learning model (No in step S1905), the control software 1801 controls the image processing unit 1203 based on the shooting setting values input in step S1901, develops the captured image such as the RAW image generated in step S1902, and generates a processed image (e.g., a YUV image, an RGB image, a JPEG image, etc.) (step S1911).
[0127] Next, the control software 1801 controls the display unit 1207 based on the shooting setting values input in step S1901 to display the processed image (e.g., YUV image, RGB image, JPEG image, etc.) generated in step S1910 (step S1911). Thereafter, the process returns to step S1904, and the same processes as above are repeatedly executed.
[0128] Then, when the display of the monitoring image on the display unit (viewfinder) 1207 is completed (No in step S1904), this processing also ends.
[0129] C. Application Examples C-1. First Application Example Fig. 20 shows, as an application example of the operation of the system 10 shown in Figs. 4 and 9, the generation of a learning model in the case where difference information between an original captured image and an edited image is used as learning data instead of the editing setting values of an image editing application, and the operation of the system 10 after the camera device 11 acquires the learning model.
[0130] First, the operation up to generating a learning model in the system 10 shown in FIG. 20 will be described.
[0131] Each time a user takes a photograph, the user sets photographing setting values (aperture value, shutter speed, sensitivity value, exposure compensation value, color temperature, focus area type, etc.) via a user interface (not shown in FIG. 4 ) of the camera device 11. Then, the photographing function unit 401 takes a photograph based on the set photographing setting values. In addition, when taking a photograph based on the set photographing setting values, the photographing function unit 401 generates metadata (see FIGS. 6 and 7 ).
[0132] When the camera device 11 creates a file of the captured image, it also stores the metadata and shooting setting values used in capturing the image in the file. The metadata and shooting setting values may be stored in the same file as the captured image, or may be stored in a separate file from the captured image if they are linked to the captured image. In Fig. 20, the file group of captured images output from the camera device 11 is indicated by reference numeral 411, and the file groups of metadata and shooting setting values linked to each captured image are indicated by reference numerals 412 and 413, respectively.
[0133] Next, when the editing device 12 receives the captured image 411 from the camera device 11, it performs editing processing such as image processing on the captured image in accordance with the user's editing operations using the image editing application 402. The image editing application 402 processes the captured image input from the camera device 11 based on the editing setting values (see FIG. 8 ) set by the user.
[0134] Then, when the edited image created by the image editing application 402 is converted into a file and output, difference information between the original captured image and the edited image is stored in a file linked to the original captured image. Note that the editing device 12 may store some of the metadata used in image editing in the metadata file. In Fig. 20, the group of edited image files output from the image editing application 402 is indicated by reference numeral 414, and difference information between the original captured image and the edited image is indicated by reference numeral 2201. The "difference information" between images referred to here includes, for example, the following information:
[0135] - The difference in brightness value of each corresponding pixel between the original captured image and the edited image. - The difference in color of each corresponding pixel between the original captured image and the edited image.
[0136] The model providing device 13 collects file groups 412, 413, 2201 of metadata, shooting setting values, and difference information between the shot image and the edited image that are each linked to the original shot image, and stores them in storage 403 linked to the shot image for each user.
[0137] Then, within the model providing device 13, the model learning unit 2202 trains a model on the correlation between metadata, shooting settings, and difference information between the captured image and the edited image associated with each captured image for a specific user, and generates a learning model 2203 that estimates shooting settings and difference information between the captured image and the edited image from the metadata. The learning model 2203 is a model for a specific user that estimates appropriate shooting settings and editing settings for each subject scene based on the specific user's past shooting and image editing tendencies. The model providing device 13 performs the learning model generation process using compressed metadata and difference information between images, thereby reducing storage costs and computational costs compared to training a model using captured images with large data sizes as they are. The model providing device 13 transfers the learning model 2203 to the specific user's camera device 11.
[0138] Next, the operation of the system 10 shown in FIG. 20 after the camera device 11 acquires a learning model will be described.
[0139] When the camera device 11 receives a shooting instruction from a user, the shooting function unit 401 analyzes the subject scene and generates metadata as indicated by reference numeral 417. This metadata 417 is input into the learning model 2203 provided by the model providing device 13, which estimates shooting setting values and difference information between the captured image and the edited image as indicated by reference numerals 418 and 2204, respectively. Furthermore, as indicated by reference numeral 2205, the difference information between the captured image and the edited image is converted into parameter values in a format that can be handled within the camera device 11. The shooting function unit 401 then sets the aperture value, shutter speed, sensitivity value, exposure compensation value, color temperature, and focus area type in the shooting execution unit (described above) based on the shooting setting values 418 estimated using the learning model 2203, and performs a shooting operation, and sets the converted parameter values 2205 in the signal processing unit (described above) to perform signal processing of the image signal (such as developing a RAW image).
[0140] In this way, by using the learning model 2203 provided by the model providing device 13, the camera equipment 11 can generate, according to the subject scene, an image 2206 having a visual effect that reflects the visual effect trends of a group of images 414 previously generated by the same user through shooting and image editing, simply by the user performing a shooting operation (in other words, without performing image editing).
[0141] C-2. Second Application Example Figure 21 shows, as a further application example of the operation of the system 10 shown in Figure 20, the generation of a learning model when a captured image is automatically edited using a generative AI model in the editing device 12 instead of an image editing application, and the operation of the system 10 after the camera device 11 acquires the generative AI model.
[0142] First, the operation up to generating a learning model in the system 10 shown in FIG. 21 will be described.
[0143] The operation of the camera device 11 before acquiring the learning model is the same as that of the application example shown in FIG. 20 above, and therefore will not be described here.
[0144] When the editing device 12 receives the captured image 411 from the camera device 11, it performs automatic editing processing on the captured image using the generative AI model 2101. The generative AI model 2101 may generate an edited image of the captured image. The generative AI model 2101 may be an image editing model that has been trained to edit the captured image or to generate an edited image of the captured image. Model parameters and the like are input to the generative AI model 2101 via an external input device 2110.
[0145] Then, when the edited image created by the generative AI model 2101 is converted into a file and output, difference information between the original captured image and the edited image is stored in a file linked to the original captured image. Note that the editing device 12 may store some of the metadata used in image editing in the above-mentioned metadata file. In FIG. 21 , the group of files of edited images output from the generative AI model 2101 is indicated by reference numeral 2102, and difference information between the original captured image and the edited image is indicated by reference numeral 2103. The "difference information" between images referred to here is the same as described above.
[0146] The model providing device 13 collects file groups 412, 413, 2103 of metadata, shooting setting values, and difference information between the shot image and the edited image that are each linked to the original shot image, and stores them in storage 403 linked to the shot image for each user.
[0147] Then, within the model providing device 13, the model learning unit 2104 trains a model to learn the correlation between metadata, shooting settings, and difference information between the captured image and the edited image associated with each captured image for a specific user, and generates a learning model 2105 that estimates the shooting settings and difference information between the captured image and the edited image from the metadata, and saves the model as a file. The learning model 2105 generated by the model providing device 13 is a model for a specific user that estimates appropriate shooting settings and editing settings for each subject scene based on the specific user's past shooting and image editing tendencies. The model providing device 13 transfers the learning model 2105 to the camera device 11 of the specific user. If the camera device 11 already has the learning model 2100, it updates or modifies the learning model 2100 based on the learning model 2105 newly provided by the model providing device 13.
[0148] Next, the operation of the system 10 shown in FIG. 21 after the camera device 11 acquires a learning model will be described.
[0149] When the camera device 11 receives a shooting instruction from a user, the shooting function unit 401 analyzes the subject scene and generates metadata as indicated by reference numeral 417. This metadata 417 is input into the learning model 2100, which estimates shooting setting values and difference information between the captured image and the edited image as indicated by reference numerals 2106 and 2107, respectively. Furthermore, as indicated by reference numeral 2108, the difference information between the captured image and the edited image is converted into parameter values in a format that can be handled within the camera device 11. The shooting function unit 401 then sets the aperture value, shutter speed, sensitivity value, exposure compensation value, color temperature, and focus area type in the shooting execution unit (described above) based on the shooting setting values 2106 estimated using the learning model 2100 to perform a shooting operation, and sets the converted parameter values 2108 in the signal processing unit (described above) to perform signal processing of the image signal (such as developing a RAW image).
[0150] In this way, by using the learning model 2100 provided by the model providing device 13, the camera equipment 11 can generate an image 2109 having a visual effect that reflects the visual effect trends of the image group 2102 generated by automatic editing using the generation AI model 2101 according to the subject scene, simply by the user performing the shooting operation (in other words, without performing image editing).
[0151] C-3. Third Application Example Figure 22 shows an application example of the system 10 shown in Figure 1. In the system 10 shown in Figure 1, only the model providing device 13 is placed in the cloud, and a general user is configured to take images and edit the taken images using a camera device 11 and an editing device 12. In contrast, in the system 10 shown in Figure 22, both the model providing device 13 and the editing device 12 are placed in the cloud, and a general user is configured to take images using the camera device 11.
[0152] The operation of the system 10 shown in FIG. 22 is described below.
[0153] A user uploads an image (e.g., a RAW image) captured by the camera device 11 to the editing device 12 on the cloud, and the editing device 12 then performs editing processing on the uploaded captured image. The camera device 11 can then download the edited image (e.g., a YUV image, an RGB image, a JPEG image, etc.) from the editing device 12. The camera device 11 also associates the metadata and shooting setting values used when the image was captured with the captured image and uploads the resulting image to the model providing device 13. The metadata is as already described with reference to FIGS. 6 and 7 , and a reduced image of the captured image may also be included in the metadata.
[0154] The editing device 12 performs editing processing on the captured image uploaded from the camera device 11, for example, based on an editing operation from the user via the camera device 11 (or another information terminal owned by the user). The editing device 12 may perform image editing using an image editing application, or may generate an edited image using a generative AI model. The editing device 12 also links editing information generated during image editing to the captured image and transfers it to the model providing device 13. The editing information includes at least one of editing setting values used in image editing and difference information between the original captured image and the edited image.
[0155] The model providing device 13 collects metadata, shooting setting values, and editing information linked to captured images from the camera equipment 11 and the editing device 12. Then, using the large amount of collected metadata, shooting setting values, and editing information, the model providing device 13 trains a model to learn the correlation between the metadata, shooting setting values, and editing information, and generates a learning model that estimates the shooting setting values and editing information from the metadata. The model providing device 13 then transfers the learning model to the camera equipment 11.
[0156] By using the learning model provided by the model providing device 13, the camera equipment 11 can generate an image with visual effects that reflect the visual effect trends of the image edited by the editing device 12, depending on the subject scene, simply by the user performing the shooting operation (in other words, without performing image editing).
[0157] The advantage of a system configuration in which both the model providing device 13 and the editing device 12 are placed in the cloud as shown in Figure 22 is that a service consisting of a set of the model providing device 13 and the editing device 12 can be provided to a large number of camera devices 11, as shown in Figure 23.
[0158] The model providing device 13 can provide each camera device 11 with a deep learning model using a huge amount of metadata, shooting setting values, and editing information linked to images captured by the many camera devices 11. In this case, the learning model should be said to be intended for general users rather than for a specific user.
[0159] In addition, the model providing device 13 can provide a professional-grade learning model to a large number of camera devices 11, which has been trained using metadata and shooting setting values linked to images taken by experts such as professional photographers, and editing information from professional image editors.
[0160] The editing device 12 may perform image editing using an image editing application, or may generate edited images using a generative AI model. Since the image editing application is shared by many users, an expensive deep learning generative AI model may be used to provide an image editing service or an edited image generation service.
[0161] D. Configuration of Information Processing Device Fig. 24 shows an example of the hardware configuration of an information processing device 2000 that can operate as an editing device or a model providing device. This information processing device 2000 includes a CPU (Central Processing Unit) 2001, a ROM (Read Only Memory) 2002, a RAM (Random Access Memory) 2003, a host bus 2004, a bridge 2005, an expansion bus 2006, an interface unit 2007, an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013. The information processing device 2000 is configured, for example, by a personal computer.
[0162] The CPU 2001 controls the overall operation of the information processing device 2000 in accordance with various programs. When performing computationally intensive processing such as model learning on the information processing device 2000, it is desirable that the CPU 2001 be a multi-core CPU (e.g., Apple M1 Max, etc.), or that the information processing device 2000 further be equipped with a multi-core processor such as a GPU (Graphics Processing Unit) or a GPGPU (General-purpose computing on graphics processing unit) (e.g., NVIDIA's "Quadro A6000"). However, for convenience, these will be collectively referred to as the CPU 2001 below.
[0163] The ROM 2002 stores in a nonvolatile manner programs (such as a basic input / output system) and calculation parameters used by the CPU 2001. The RAM 2003 is used to load programs to be executed by the CPU 2001 and to temporarily store parameters such as working data that change as appropriate during program execution. Programs loaded into the RAM 2003 and executed by the CPU 2001 include, for example, various application programs and an operating system (OS).
[0164] The CPU 2001, ROM 2002, and RAM 2003 are interconnected by a host bus 2004, which includes a CPU bus and other components. The CPU 2001 executes various application programs in an execution environment provided by an OS through the cooperative operation of the ROM 2002 and RAM 2003, thereby enabling various functions and services to be realized. If the information processing device 2000 is a personal computer, the OS may be, for example, Microsoft Windows (registered trademark), Unix (registered trademark), or a successor OS. Furthermore, the application program or some of the modules in the application program may use an existing library that is stored, shared, or made public through, for example, a source code management service.
[0165] When the information processing device 2000 operates as an editing device, an editing application that performs editing, such as image processing, on images captured by a camera device is executed on the CPU 2001. This editing application includes a filing algorithm that files the editing setting values used to edit the captured image and links the files to the original captured image. Furthermore, when the information processing device 2000 operates as a model providing device, the CPU 2001 runs a file collection algorithm that collects file groups of filed metadata, shooting setting values, and editing setting values, and a model learning algorithm that generates a learning model that estimates shooting setting values and editing setting values that reflect the user's past shooting and image editing tendencies from metadata, based on the correlation between the metadata, shooting setting values, and editing setting values linked to each captured image.
[0166] The host bus 2004 is connected to an expansion bus 2006 via a bridge 2005. The expansion bus 2006 is, for example, a PCI (Peripheral Component Interconnect) bus or PCI Express, and the bridge 2005 is based on the PCI standard. However, the information processing device 2000 does not need to be configured so that the circuit components are separated by the host bus 2004, bridge 2005, and expansion bus 2006, and may be implemented so that almost all circuit components are interconnected by a single bus (not shown).
[0167] The interface unit 2007 connects peripheral devices such as an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013 in accordance with the standards of the expansion bus 2006. However, not all of the peripheral devices shown in Fig. 24 are necessarily required, and the information processing device 2000 may further include peripheral devices not shown. Furthermore, the peripheral devices may be built into the main body of the information processing device 2000, or some of the peripheral devices may be externally connected to the main body of the information processing device 2000.
[0168] The input unit 2008 is composed of an input control circuit that generates an input signal based on an input from a user and outputs the signal to the CPU 2001. If the information processing device 2000 is a personal computer, the input unit 2008 may include a keyboard, a mouse, and a touch panel. The output unit 2009 includes, for example, a display device such as a liquid crystal display (LCD) device, an organic electroluminescence (EL) display device, and an LED (light emitting diode), and an audio output device such as a speaker.
[0169] The storage unit 2010 stores files such as programs (applications, OS, etc.) executed by the CPU 2001 and various data. The storage unit 2010 is configured with a large-capacity storage device such as an SSD (Solid State Drive) or an HDD (Hard Disk Drive), but may also include an external storage device. When the information processing device 2000 operates as a model providing device, the storage unit 2010 accumulates, for each user, a group of files associated with the original captured image, including metadata, shooting setting values, and editing setting values.
[0170] The removable storage medium 2012 is a storage medium configured as a cartridge, such as a microSD card. The drive 2011 performs read and write operations on the loaded removable storage medium 2012. The drive 2011 outputs data read from the removable storage medium 2012 to the RAM 2003 or the storage unit 2010, and writes data on the RAM 2003 or the storage unit 2010 to the removable storage medium 2012.
[0171] The communication unit 2013 is a device that performs wireless communication such as Wi-Fi (registered trademark), Bluetooth (registered trademark), or cellular communication networks such as 4G and 5G. The communication unit 2013 may also include terminals such as a Universal Serial Bus (USB) or a High-Definition Multimedia Interface (HDMI) (registered trademark), and may further include a function for performing HDMI (registered trademark) communication with USB devices such as scanners and printers, displays, and the like. When the information processing device 2000 operates as an editing device, it acquires captured images from a camera device and transfers files of editing setting values to a model providing device through the communication unit 2013. When the information processing device 2000 operates as a model providing device, it acquires file groups of metadata, shooting setting values, and editing setting values from a camera device or an editing device through the communication unit 2013, and transfers a generated learning model to the camera device.
[0172] The present disclosure has been described in detail above with reference to specific embodiments. However, the present disclosure should not be construed as being limited to the above-described embodiments, and it is obvious that those skilled in the art can modify or substitute the embodiments without departing from the spirit of the present disclosure. Furthermore, the effects described in this specification are merely examples, and the effects brought about by the present disclosure are not limited thereto, and additional effects not described in this specification may exist.
[0173] Although the present specification has mainly described an embodiment in which the present disclosure is applied to a system consisting of a physically independent camera device, editing device, and model providing device, the gist of the present disclosure is not limited thereto. For example, the present disclosure can also be applied to a system consisting of a camera device with editing functionality and a model providing device. Furthermore, the model providing device may be, for example, a server deployed in the cloud, which generates and provides models reflecting each user's past shooting and image editing tendencies for the camera devices of multiple users. Alternatively, the camera device may be equipped with a model generation function and internally generate a learning model reflecting the camera device user's past shooting and image editing tendencies.
[0174] In short, the present disclosure has been described in the form of examples, and the contents of the specification should not be interpreted as limiting. To determine the gist of the present disclosure, the claims should be taken into consideration.
[0175] The series of processes described in this specification can be executed by hardware, software, or a configuration that combines hardware and software. When executing processes by software, a program recording a processing sequence related to realizing the present disclosure is installed in memory in a computer incorporated in dedicated hardware and executed. It is also possible to install the program in a general-purpose computer capable of executing various processes and execute the processes related to realizing the present disclosure.
[0176] The program can be stored in advance on a recording medium installed in the computer, such as a HDD, SSD, or ROM. Alternatively, the program can be temporarily or permanently stored on a removable recording medium such as a flexible disk, CD-ROM (Compact Disc Read Only Memory), MO (Magneto Optical) disk, DVD (Digital Versatile Disc), BD (Blu-Ray Disc (registered trademark)), magnetic disk, or USB (Universal Serial Bus) memory. Using such a removable recording medium, a program related to the realization of the present disclosure can be provided as so-called package software.
[0177] The program may also be transferred wirelessly or via a wire from a download site to a computer via a network such as a wide area network (WAN) typified by cellular, a local area network (LAN), the Internet, etc. The computer can receive the program transferred in this manner and install it in a large-capacity storage device such as an HDD or SSD within the computer.
[0178] The present disclosure may also be configured as follows.
[0179] (1) An information processing device comprising: a collection unit that collects metadata and shooting settings at the time of shooting and editing information at the time of image editing of the captured image, linked to an image captured by a camera, and stores the collected information in storage; a generation unit that trains a model to learn the correlation between the metadata, shooting settings, and editing information stored in the storage, and generates a learning model that estimates the shooting settings and editing information from the metadata; and a transfer unit that transfers the generated learning model to the camera.
[0180] (2) The information processing device described in (1) above, wherein the metadata includes at least one of first information that analyzes the characteristics of the subject scene, second information that digitizes the situation at the time of shooting, and third information that reduces the image of the subject scene.
[0181] (3) The information processing device described in (2) above, wherein the first information includes at least one of the following: a frequency value obtained by classifying the brightness value of each pixel in the image into five levels; a ratio of the values of the brightest pixel and the darkest pixel; a level in the five-level classification of the brightness of the focus position pixel; a frequency value obtained by classifying the color of each pixel in the image into five colors; a level in the five-level classification of the color of the focus position pixel; whether or not there is a moving object in the subject scene; and distance information of the subject; and the second information includes at least one of the name of the photographing equipment; the name of the lens used; the date and time of the photograph; an ambient illuminance value; and a tilt of the camera (or camera shake).
[0182] (4) The information processing device according to any one of (1) to (3), wherein the shooting setting values include at least one of an aperture value, a shutter speed, a sensitivity value, an exposure compensation value, a color temperature, and a focus area type.
[0183] (5) The information processing device according to any one of (1) to (4), wherein the editing information includes editing setting values consisting of at least one of brightness, color temperature adjustment, color correction, contrast, white level, black level, and saturation.
[0184] (6) The information processing device according to (1), wherein the editing information includes difference information between the captured image and the edited image.
[0185] (6-1) The information processing device described in (6) above, wherein the difference information between the captured image and the edited image includes at least one of the difference in brightness value of each corresponding pixel between the original captured image and the edited image, and the difference in color of each corresponding pixel between the original captured image and the edited image.
[0186] (7) The information processing device described in any one of (1) to (6) above, wherein the generation unit trains a model to learn correlations between metadata, shooting settings, and editing settings based on a specific user's shooting and editing operations, and generates a learning model that reflects the user's past shooting and image editing tendencies.
[0187] (8) An information processing method comprising: a step of collecting metadata and shooting settings at the time of shooting and editing information at the time of image editing of the shot image, linked to the image shot by the camera, and storing them in storage; a generation step of training a model to learn the correlation between the metadata, shooting settings, and editing information stored in the storage, and generating a learning model that estimates the shooting settings and editing information from the metadata; and a transfer step of transferring the generated learning model to the camera.
[0188] (9) A computer program written in a computer-readable format to cause a computer to function as: a collection unit that collects metadata and shooting settings at the time of shooting and editing information at the time of image editing of the captured image by linking them to images captured by a camera and stores them in storage; a generation unit that trains a model to learn the correlation between the metadata, shooting settings, and editing information stored in the storage, and generates a learning model that estimates the shooting settings and editing information from the metadata; and a transfer unit that transfers the generated learning model to the camera.
[0189] (10) An imaging device having a photographing function unit, comprising: a generation unit that analyzes a subject scene to generate metadata; and an estimation unit that estimates setting values from the metadata using a learning model that has learned the correlation between the metadata and setting values, and the imaging device controls the photographing function unit using the setting values to perform a photographing operation.
[0190] (11) The imaging device described in (10) above, wherein the estimation unit uses the learning model to estimate shooting setting values and editing information from the metadata, and further includes a conversion unit that converts the editing information into parameter values that can be handled within the imaging device, and controls the imaging function unit using the shooting setting values and the parameter values to perform shooting operations.
[0191] (12) The imaging device according to (11) above, further comprising a monitoring function unit that generates a monitoring image of the imaging function unit, and controls the monitoring function unit using the imaging setting values and the parameter values to display a monitoring screen.
[0192] (13) The imaging device according to any one of (10) to (12) above, wherein metadata and setting values are saved during shooting.
[0193] (14) The imaging device according to any one of (10) to (13), wherein the estimation unit uses the learning model generated using metadata and shooting setting values at the time of shooting and editing information at the time of editing the captured image.
[0194] (15) A system comprising a camera, an editing device that edits images captured by the camera, and a model providing device that provides a learning model to the camera, wherein the camera outputs metadata at the time of shooting and shooting setting values of the camera linked to the captured image, and the editing device outputs editing information at the time of image editing of the captured image, and the model providing device comprises a collection unit that collects the metadata, shooting setting values, and editing information and stores them in storage, a generation unit that generates a learning model that estimates the shooting setting values and editing information from the metadata by having a model learn the correlation between the metadata, shooting setting values, and editing information stored in the storage, and a transfer unit that transfers the generated learning model to the camera.
[0195] (16) A camera system comprising: a camera unit having a first learning model unit; and a generation AI model unit that generates a second image based on a first image captured by the camera unit; and the first learning model is changed using a learning model learned based on the first image and the second image.
[0196] (17) A camera device having a machine learning model, comprising: a control unit; and a first learning model, wherein the control unit changes the first learning model based on a second learning model learned by comparing a first image captured by the camera device with a second image generated from the first image by a generation AI model unit.
[0197] 10...system, 11...camera equipment, 12...editing device 13...model providing device 401...photographing function unit, 402...image editing application 403...storage, 404...model learning unit 407...monitoring function unit, 408...viewfinder 1101...model learning unit 1201...imaging optical system, 1202...imaging unit, 1203...image processing unit 1204...display unit, 1205...recording unit, 1206...operation unit 1207...control unit, 1208...sensor unit, 1209...bus 1301...development processing unit, 1302...post processing unit 1303...detection processing unit, 1304...focus control unit 2000...information processing device, 2001...CPU, 2002...ROM 2003...RAM, 2004...host bus, 2005...bridge 2006...expansion bus, 2007...interface unit 2008...input unit, 2009...output unit, 2010...storage unit, 2011...drive, 2012...removable recording medium, 2013...communication unit
Claims
1. An information processing device comprising: a collection unit that collects metadata and shooting settings at the time of shooting and editing information at the time of image editing of the captured image, linked to images captured by a camera, and stores the collected information in storage; a generation unit that trains a model to learn the correlation between the metadata, shooting settings, and editing information stored in the storage, and generates a learning model that estimates the shooting settings and editing information from the metadata; and a transfer unit that transfers the generated learning model to the camera.
2. The information processing device according to claim 1, wherein the metadata includes at least one of first information obtained by analyzing the characteristics of the subject scene and second information obtained by converting the situation at the time of shooting into digital form.
3. The information processing device of claim 2, wherein the first information includes at least one of the following: a frequency value obtained by classifying the brightness value of each pixel in the image into five levels; a ratio of the values of the brightest pixel and the darkest pixel; a level in the five-level classification of the brightness of the focus position pixel; a frequency value obtained by classifying the color of each pixel in the image into five colors; a level in the five-level classification of the color of the focus position pixel; whether or not there is a moving object in the subject scene; and distance information of the subject; and the second information includes at least one of the name of the photographing equipment; the name of the lens used; the date and time of the photograph; an ambient illuminance value; and a tilt of the camera (or camera shake).
4. The information processing device according to claim 1, wherein the shooting setting values include at least one of an aperture value, a shutter speed, a sensitivity value, an exposure compensation value, a color temperature, and a focus area type.
5. The information processing device according to claim 1, wherein the editing information includes editing setting values consisting of at least one of brightness, color temperature adjustment, color correction, contrast, white level, black level, and saturation.
6. The information processing device according to claim 1, wherein the editing information includes difference information between the captured image and the edited image.
7. The information processing device of claim 1, wherein the generation unit trains a model to learn the correlation between metadata, shooting settings, and editing information based on a specific user's shooting and editing operations, and generates a learning model that reflects the user's past shooting and image editing tendencies.
8. An information processing method comprising: a step of collecting metadata and shooting settings at the time of shooting and editing information at the time of image editing of the shot image, linked to the image shot by the camera, and storing them in storage; a generation step of training a model to learn the correlation between the metadata, shooting settings, and editing information stored in the storage, and generating a learning model that estimates the shooting settings and editing information from the metadata; and a transfer step of transferring the generated learning model to the camera.
9. A computer program written in a computer-readable format to cause a computer to function as: a collection unit that collects metadata and shooting settings at the time of shooting and editing information at the time of image editing of the captured image, linked to images captured by a camera, and stores them in storage; a generation unit that trains a model to learn the correlation between the metadata, shooting settings, and editing information stored in the storage, and generates a learning model that estimates the shooting settings and editing information from the metadata; and a transfer unit that transfers the generated learning model to the camera.
10. An imaging device having a photographing function unit, comprising: a generation unit that analyzes a subject scene and generates metadata; and an estimation unit that estimates setting values from the metadata using a learning model that has learned the correlation between metadata and setting values; and an imaging device that controls the photographing function unit using the setting values to perform a photographing operation.
11. The imaging device described in claim 10, wherein the estimation unit uses the learning model to estimate shooting setting values and editing information from the metadata, and further includes a conversion unit that converts the editing information into parameter values that can be handled by the imaging device, and controls the shooting function unit using the shooting setting values and parameter values to perform shooting operations.
12. The imaging device according to claim 11, further comprising a monitoring function unit that generates a monitoring image of the imaging function unit, and controls the monitoring function unit using the imaging setting values and the parameter values to display a monitoring screen.
13. The imaging device according to claim 10, wherein metadata and setting values are saved when capturing an image.
14. The imaging device according to claim 10, wherein the estimation unit uses the learning model generated using metadata and shooting setting values at the time of shooting and editing information at the time of editing the captured image.
15. A system comprising: a camera; an editing device that edits images captured by the camera; and a model providing device that provides a learning model to the camera, wherein the camera outputs metadata at the time of shooting and shooting setting values of the camera linked to the captured image; the editing device outputs editing information at the time of editing the captured image; the model providing device comprises a collection unit that collects the metadata, shooting setting values, and editing information and stores them in storage; a generation unit that generates a learning model that estimates the shooting setting values and editing information from the metadata by having a model learn the correlation between the metadata, shooting setting values, and editing information stored in the storage; and a transfer unit that transfers the generated learning model to the camera.
16. A camera system comprising: a camera unit having a first learning model unit; and a generation AI model unit that generates a second image based on a first image captured by the camera unit; and the first learning model is changed using a learning model learned based on the first image and the second image.
17. A camera device having a machine learning model, comprising: a control unit; and a first learning model, wherein the control unit changes the first learning model based on a second learning model learned by comparing a first image captured by the camera device with a second image generated from the first image by a generation AI model unit.
Citation Information
Patent Citations
Imaging device and imaging method
JP2019146022A
Image processing device and control method thereof
JP2022111133A