Image processing device, imaging device, image processing method, and program
The image processing device addresses redundancy in composite images by associating subject information based on similarity criteria, improving image usability through selective recording.
Patent Information
- Application Number
- JP2022028383
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-25
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2042-02-25
AI Technical Summary
Conventional image processing technologies fail to address redundancy issues in subject information when combining multiple raw images to generate a composite image, leading to reduced usability due to redundant subject information association.
An image processing device that acquires and synthesizes images, associating subject information only if the similarity between images meets a predetermined criterion, thereby reducing redundancy in the composite image.
Reduces redundancy of subject information in composite images, enhancing their usability by selectively recording relevant subject information.
Smart Images

Figure 0007797245000001 
Figure 0007797245000002 
Figure 0007797245000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device, an imaging device, an image processing method, and a program. [Background technology]
[0002] In recent years, artificial intelligence (AI) technologies such as deep learning have been utilized in various technical fields. For example, a function for detecting human faces from captured images has been known in digital still cameras. Patent Document 1 also discloses a technology for accurately detecting and recognizing animals such as dogs and cats, in addition to detecting humans.
[0003] Furthermore, there are known techniques for creating a composite image by compositing a plurality of raw images, such as HDR compositing and multiple compositing. In relation to this technique, Patent Document 2 discloses that only the shooting information of an image (raw image) including a main subject is added to the composite image and recorded. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-099559 [Patent Document 2] Japanese Patent Application Publication No. 2019-009577 Summary of the Invention [Problem to be solved by the invention]
[0005] Consider a case where subject information of raw images is estimated and recorded using AI technology, and multiple raw images are combined to generate a composite image. In this case, if all of the subject information of each raw image is unconditionally associated with the composite image, depending on the similarity of the subject information between the raw images, the subject information associated with the composite image may become redundant, reducing the usability of the subject information. However, conventional technology cannot address this issue.
[0006] The present invention has been made in consideration of this situation, and aims to provide a technology for associating subject information of each material image with a composite image so as to reduce redundancy of the subject information associated with the composite image. [Means for solving the problem]
[0007] In order to solve the above problem, the present invention provides an image processing device comprising: an acquisition means for acquiring a first image, first subject information representing a first subject detected in the first image, a second image, and second subject information representing a second subject detected in the second image; a synthesis means for generating a synthetic image by synthesizing the first image and the second image; and a recording means for recording either the first subject information or the second subject information in association with the synthetic image if the similarity between the first subject information and the second subject information satisfies a predetermined criterion, and for recording both the first subject information and the second subject information in association with the synthetic image if the similarity between the first subject information and the second subject information does not satisfy the predetermined criterion. [Effects of the Invention]
[0008] According to the present invention, it is possible to reduce redundancy of subject information associated with a composite image.
[0009] Other features and advantages of the present invention will become more apparent from the accompanying drawings and the following detailed description of the preferred embodiment of the present invention. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of a digital camera 100. [Figure 2] 1 is a flowchart of HDR shooting processing executed by the digital camera 100. [Figure 3] FIG. 4 is a diagram showing an example of the configuration of a material image file. [Figure 4] 10 is a flowchart of HDR composition processing executed by the digital camera 100. [Figure 5] 1A and 1B are diagrams showing examples of material images and an HDR composite image. [Figure 6] 10A and 10B are diagrams showing examples of annotation information recorded in a material image file and an HDR composite image file. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.
[0012] In the following description, a digital camera (image capture device) is used as an example of an image processing device that performs object classification using an inference model. However, in the following embodiments, the image processing device is not limited to a digital camera. The image processing device in the following embodiments may be any device that has the functions of a digital camera described below, and may be, for example, a smartphone or a tablet PC.
[0013] [First embodiment] ●Configuration of digital camera 100 FIG. 1 is a block diagram showing an example configuration of a digital camera 100. Barrier 10 is a protective member that covers the imaging unit, including the photographing lens 11, of digital camera 100, to prevent the imaging unit from getting dirty or damaged. The operation of barrier 10 is controlled by barrier control unit 43. The photographing lens 11 forms an optical image on the imaging surface of imaging element 13. Shutter 12 has an aperture function. The imaging element 13 is composed of, for example, a CCD or CMOS sensor, and converts the optical image formed on the imaging surface by the photographing lens 11 via shutter 12 into an electrical signal.
[0014] The A / D converter 15 converts the analog image signal output from the image sensor 13 into a digital image signal. The digital image signal converted by the A / D converter 15 is written to memory 25 as so-called RAW image data. At the same time, development parameters corresponding to each piece of RAW image data are generated based on information at the time of shooting and written to memory 25. The development parameters consist of various parameters used in image processing for recording images in the JPEG format or the like, such as exposure settings, white balance, color space, and contrast.
[0015] The timing generating unit 14 is controlled by the memory control unit 22 and the system control unit 50 and supplies clock signals and control signals to the image sensor 13, the A / D converter 15, and the D / A converter 21.
[0016] The image processing unit 20 performs various image processing such as predetermined pixel interpolation processing, color conversion processing, correction processing, resizing processing, and image synthesis processing on data from the A / D converter 15 or data from the memory control unit 22. The image processing unit 20 also performs predetermined image processing and arithmetic processing using image data obtained by capturing an image, and provides the obtained arithmetic results to the system control unit 50. The system control unit 50 controls the exposure control unit 40 and focus control unit 41 based on the provided arithmetic results, thereby realizing AF (autofocus) processing, AE (autoexposure) processing, and EF (flash pre-flash) processing.
[0017] In this embodiment, the system control unit 50 can perform shooting with two different exposure settings. The first exposure setting is an "optimal exposure setting," in which the system control unit 50 feeds back the results of AE (auto exposure) processing to the exposure control unit 40, thereby obtaining an image with appropriate exposure. The second exposure setting is an "underexposure setting," in which the system control unit 50 provides an offset that lowers the exposure with respect to the results of AE (auto exposure) processing and feeds this back to the exposure control unit 40, thereby obtaining an image with darker exposure.
[0018] The image processing unit 20 also performs predetermined calculations using the captured image data and performs AWB (auto white balance) processing based on the calculation results. Furthermore, the image processing unit 20 reads image data stored in the memory 25 and performs compression or decompression processing using a method such as the JPEG method, the MPEG-4 AVC method, the HEVC (High Efficiency Video Coding) method, or a lossless compression method for uncompressed RAW data. The image processing unit 20 then writes the processed image data to the memory 25.
[0019] The image processing unit 20 also performs predetermined arithmetic processing using captured image data to edit various types of image data. For example, the image processing unit 20 can perform a trimming process to adjust the display range and size of an image by hiding unnecessary parts around the image data, and a resizing process to enlarge or reduce the size of image data and screen display elements. Furthermore, the image processing unit 20 can perform RAW development, which applies image processing such as color conversion to data that has been compressed or expanded using a lossless compression method on uncompressed RAW data, converts it to JPEG format, and creates image data. The image processing unit 20 can also perform video clipping, which clips out specified frames from a video format such as MPEG-4, converts them to JPEG format, and saves them.
[0020] The image processing unit 20 also includes a compositing processing circuit that combines multiple pieces of image data. The image processing unit 20 is capable of performing additive compositing processing, weighted additive compositing processing, and area-designated compositing processing. The area-designated compositing processing is a process in which an area to be used for compositing is designated for each material image, and the designated area of each material image is composited.
[0021] The image processing unit 20 also performs processing such as superimposing an OSD (On-Screen Display) such as a menu or arbitrary characters to be displayed on the display unit 23 together with the image data for display.
[0022] Furthermore, the image processing unit 20 performs subject detection processing to detect the subject present in the image data and the subject area by using the input image data and information on the distance to the subject obtained from the image sensor 13 at the time of shooting. Detectable information (subject detection information) includes information on the position, size, and tilt of the subject area in the image, as well as information on the likelihood.
[0023] The memory control unit 22 controls the A / D converter 15, the timing generation unit 14, the image processing unit 20, the image display memory 24, the D / A converter 21, and the memory 25. The RAW image data generated by the A / D converter 15 is written to the image display memory 24 or the memory 25 via the image processing unit 20 and the memory control unit 22, or directly via the memory control unit 22.
[0024] The image data for display written in the image display memory 24 is displayed on a display unit 23 configured by a TFT LCD or the like via a D / A converter 21. By using the display unit 23 to sequentially display image data obtained by capturing an image, it is possible to realize an electronic finder function that displays a live image.
[0025] The memory 25 has a storage capacity sufficient to store a predetermined number of still images and a predetermined period of moving images, and stores the captured still images and moving images. The memory 25 can also be used as a working area for the system control unit 50.
[0026] The exposure control unit 40 controls the shutter 12, which has an aperture function. The exposure control unit 40 also has a flash dimming function by working in conjunction with a flash 44. The focus control unit 41 adjusts focus by driving a focus lens (not shown) included in the photographing lens 11 based on instructions from the system control unit 50. The zoom control unit 42 controls zooming by driving a zoom lens (not shown) included in the photographing lens 11. The flash 44 has an AF assist light projection function and a flash dimming function.
[0027] The system control unit 50 controls the entire digital camera 100. The nonvolatile memory 51 is an electrically erasable and recordable nonvolatile memory, such as an EEPROM. The nonvolatile memory 51 stores not only programs but also map information and the like.
[0028] The shutter switch 61 (SW1) is turned on during operation of the shutter button 60, and instructs the start of operations such as AF processing, AE processing, AWB processing, and EF processing. The shutter switch 62 (SW2) is turned on when the shutter button 60 is completely operated, and instructs the start of a series of shooting operations including exposure processing, development processing, and recording processing. In the exposure processing, a signal read from the image sensor 13 is written to the memory 25 as RAW image data via the A / D converter 15 and the memory control unit 22. In the development processing, the RAW image data written to the memory 25 is developed through calculations in the image processing unit 20 and the memory control unit 22, and then written to the memory 25 as image data. In the recording processing, the image data is read from the memory 25, compressed by the image processing unit 20, stored in the memory 25, and then written to the external recording medium 91 via the card controller 90.
[0029] The operation unit 63 includes various operation members such as buttons and a touch panel. For example, the operation unit 63 includes a power button, a menu button, a mode switch for switching between shooting mode, playback mode, and other special shooting modes, a cross key, a set button, a macro button, and a multi-screen playback page break button. Furthermore, for example, the operation unit 63 includes a flash setting button, a single-shot / continuous-shot / self-timer switching button, a menu movement + (plus) button, a menu movement - (minus) button, a shooting quality selection button, an exposure compensation button, a date / time setting button, etc.
[0030] When recording image data on the external recording medium 91, the metadata generation and analysis unit 70 generates various metadata, such as information conforming to the Exchangeable Image File Format (Exif) standard to be attached to the image data, based on information at the time of shooting. Furthermore, when reading image data recorded on the external recording medium 91, the metadata generation and analysis unit 70 analyzes the metadata attached to the image data. Examples of metadata include shooting setting information at the time of shooting, image data information related to the image data, and feature information of the subject included in the image data. Furthermore, when recording moving image data, the metadata generation and analysis unit 70 can also generate and attach metadata for each frame.
[0031] The power supply 80 includes a primary battery such as an alkaline battery or a lithium battery, a secondary battery such as a NiCd battery, a NiMH battery, or a Li battery, or an AC adapter, etc. The power supply control unit 81 supplies the power from the power supply 80 to each unit of the digital camera 100.
[0032] The card controller 90 transmits and receives data to and from an external recording medium 91 such as a memory card. The external recording medium 91 is configured, for example, as a memory card, and records images (still images and videos) captured by the digital camera 100.
[0033] The inference engine 73 uses the inference model recorded in the inference model recording unit 72 to perform inference on image data input via the system control unit 50. The system control unit 50 can record inference models input from an external device (not shown) via the communication unit 71 in the inference model recording unit 72. The system control unit 50 can also record inference models obtained by re-learning the inference model using the learning unit 74 in the inference model recording unit 72. Note that the inference models recorded in the inference model recording unit 72 may be updated by inputting an inference model from an external device or by re-learning the inference model using the learning unit 74. Therefore, the inference model recording unit 72 holds version information so that the version of the inference model can be identified.
[0034] The inference engine 73 also has a neural network design 73a. The neural network design 73a has a configuration in which an intermediate layer (neurons) is arranged between an input layer and an output layer. Image data is input to the input layer from the system control unit 50. Several layers of neurons are arranged in the intermediate layer. The number of neuron layers is determined appropriately in the design. The number of neurons in each layer is also determined appropriately in the design. In the intermediate layer, weighting is performed based on the inference model recorded in the inference model recording unit 72. In the output layer, an inference result according to the image data input to the input layer is output.
[0035] In this embodiment, the inference model recorded in the inference model recording unit 72 is assumed to be an inference model that infers the classification of the subject included in the image. An inference model generated by deep learning is used, using as training data image data of various subjects and the results of their classification (for example, classification of animals such as dogs and cats, or classification of subject types such as people, animals, plants, and buildings). Therefore, when an image and information indicating the area of the subject detected in this image are input to the inference engine 73 that uses the inference model, an inference result indicating the classification (type) of this subject is output.
[0036] The learning unit 74 re-learns the inference model upon receiving a request from the system control unit 50 or the like. The learning unit 74 has a teacher data recording unit 74a. The teacher data recording unit 74a records information related to teacher data for the inference engine 73. The learning unit 74 re-learns the inference engine 73 using the teacher data recorded in the teacher data recording unit 74a, and can update the inference engine 73 using the inference model recording unit 72.
[0037] The communication unit 71 has a communication circuit for transmitting and receiving. The communication performed by the communication circuit may be wireless communication such as Wi-Fi or Bluetooth (registered trademark), or wired communication such as Ethernet or USB.
[0038] ●HDR shooting processing and HDR compositing processing Next, the HDR shooting process and HDR compositing process executed by the digital camera 100 will be described with reference to Figures 2 to 6. Figure 2 is a flowchart of the HDR shooting process executed by the digital camera 100. Unless otherwise specified, the process of each step in this flowchart is realized by the system control unit 50 of the digital camera 100 controlling each component of the digital camera 100 in accordance with a program. When the operation mode of the digital camera 100 is set to HDR shooting mode, the HDR shooting process in this flowchart starts. Note that the user can set the operation mode of the digital camera 100 to HDR shooting mode by operating the operation unit 63 to display a menu screen on the display unit 23 and selecting HDR shooting mode on the menu screen.
[0039] In S202, the system control unit 50 determines whether or not a shooting instruction has been issued by the user. The user can issue a shooting instruction by pressing the shutter button 60 to turn on the shutter switches 61 (SW1) and 62 (SW2). The system control unit 50 repeats the determination process in S202 until a shooting instruction has been issued by the user. When a shooting instruction has been issued by the user, the processing proceeds to S203.
[0040] In S203, the system control unit 50 performs shooting processing with an appropriate exposure setting (appropriate shooting processing). In the appropriate shooting processing, the system control unit 50 performs AF (autofocus) processing and AE (autoexposure) processing using the focus control unit 41 and the exposure control unit 40, and then stores in the memory 25 the image signal output from the image sensor 13 via the A / D converter 15. At this time, the system control unit 50 controls so that an image with appropriate exposure is obtained by feeding back the result of the AE (autoexposure) processing to the exposure control unit 40 with the appropriate exposure setting. In addition, the image processing unit 20 performs compression processing on the image signal stored in the memory 25 in accordance with the user's settings, thereby generating image data in a format in accordance with the user's settings (for example, JPEG format) and storing the image data in the memory 25.
[0041] In S204, the image processing unit 20 performs subject detection processing on the image signal stored in the memory 25, and acquires information about the subject included in the image (subject detection information).
[0042] In S205, the system control unit 50 uses the inference engine 73 to perform inference processing on the subject detected in the image signal (material image) stored in the memory 25. The system control unit 50 identifies the subject area in the image based on the image signal stored in the memory 25 and the subject detection information acquired in S204. The system control unit 50 inputs the image signal (material image) and information indicating the subject area in the material image to the inference engine 73. As a result of the inference processing performed by the inference engine 73 for each subject area, an inference result indicating the classification (type) of the subject included in the subject area is output. Note that in addition to the inference result, the inference engine 73 may output information related to the inference processing, such as debug information and logs related to the operation of the inference processing.
[0043] In S206, the system control unit 50 records a file including the image data generated in S203, the subject detection information acquired in S204, and the inference result acquired in S205 on the external recording medium 91 as a material image file for HDR synthesis.
[0044] FIG. 3 is a diagram showing an example of the structure of a material image file. As shown in FIG. 3, the material image file 300 is divided into multiple storage areas, including an Exif area 301 that stores metadata in accordance with the Exif standard and an image data area 308 that records compressed image data. The material image file 300 also includes an annotation information area 310 that records annotation information. If the material image file 300 is a JPEG format file, each of the multiple storage areas is defined by a marker. For example, if a user instructs image recording in JPEG format, the material image file 300 is recorded in JPEG format. In this case, the image data generated in S203 is recorded in the image data area 308 in JPEG format, and the information in the Exif area 301 is recorded in an area defined by, for example, an APP1 marker. Furthermore, the information in the annotation information area 310 is recorded in an area defined by, for example, an APP11 marker. If a user instructs image recording in HEIF (High Efficiency Image File Format) format, the material image file 300 is recorded in the HEIF file format. In this case, the information in the Exif area 301 and the annotation information area 310 is recorded in a Metadata Box, etc. Similarly, when the user instructs to record an image in RAW format, the information in the Exif area 301 and the annotation information area 310 is recorded in a predetermined area such as a Metadata Box.
[0045] The subject detection information acquired in S204 is recorded by the metadata generation and analysis unit 70 in a subject detection information tag 306 in a MakerNote 305 (an area where manufacturer-specific metadata can be written in a confidential format in principle) included in the Exif area 301. In addition, if there is version information of the current inference model recorded in the inference model recording unit 72 or debug information output by the inference engine 73 in S205, this information is recorded in the MakerNote 305 as inference model management information 307.
[0046] The inference result obtained in S205 is recorded as annotation information in the annotation information area 310. The location of the annotation information area 310 is indicated by the annotation information link 303 included in the annotation link information storage tag 302. In this embodiment, it is assumed that the annotation information is written in a text format such as XML or JSON.
[0047] 2, in S207, the system control unit 50 performs shooting processing with an underexposure setting (underexposure shooting processing). In the underexposure shooting processing, the system control unit 50 performs the same processing as in S203, but controls the system control unit 50 to obtain a darkly exposed image by feeding back the results of the AE (auto exposure) processing to the exposure control unit 40 with an underexposure setting. In addition, as in S203, the system control unit 50 generates image data in a format (for example, JPEG format) according to the user's settings and stores it in the memory 25.
[0048] In S208, the image processing unit 20 acquires information about the subject (subject detection information) included in the image obtained by the underexposure processing through the same processing as in S204.
[0049] In S209, the system control unit 50 acquires an inference result regarding the subject included in the image signal (material image) obtained by the underexposure processing, through the same processing as in S205.
[0050] In S210, similar to S206, the system control unit 50 records a file including the image data generated in S207, the subject detection information acquired in S208, and the inference result acquired in S209 on the external recording medium 91 as a material image file for HDR synthesis.
[0051] Thereafter, the processing returns to step S202, and when the next shooting instruction is given, the system control unit 50 again executes the processing from step S203 onwards.
[0052] 5, an example of a material image obtained by the HDR shooting process of FIG. 2 will be described. Material image 501 is an image generated by appropriate shooting processing. Inference processing on material image 501 has resulted in inference results corresponding to mountain 504, sky 507, slope 510, and cloud 513. Material image 502 is an image generated by underexposed shooting processing. Inference processing on material image 502 has resulted in inference results corresponding to mountain 505, sky 508, slope 511, and cloud 514. The inference result for material image 501 is recorded in the material image file as annotation information 601 shown in FIG. 6, and the inference result for material image 502 is recorded in the material image file as annotation information 602 shown in FIG. 6. As shown in FIG. 6, the annotation information includes, for each subject, an inference result (subject information representing the subject) including information indicating the area and type of the subject.
[0053] Next, the HDR compositing process will be described with reference to Figure 4. Unless otherwise specified, the processing of each step in this flowchart is realized by the system control unit 50 of the digital camera 100 controlling each component of the digital camera 100 in accordance with a program. The HDR compositing process in this flowchart starts when the operating mode of the digital camera 100 is set to HDR compositing mode. Note that the user can set the operating mode of the digital camera 100 to HDR compositing mode by operating the operation unit 63 to display a menu screen on the display unit 23 and selecting HDR compositing mode on the menu screen.
[0054] In S401, the system control unit 50 displays a user interface (image selection UI) on the display unit 23 for the user to select material images for HDR compositing processing. The image selection UI displays, for example, thumbnails of material images generated by appropriate shooting processing. The user can select a material image displayed as a thumbnail by operating the operation unit 63. When a material image is selected in the image selection UI, a material image generated by under-shooting processing corresponding to the selected material image is also selected as a material image for HDR compositing processing.
[0055] It should be noted that the two source image files generated by the appropriate shooting process and the underexposed shooting process included in one HDR shooting process are related to each other in some way. For example, the two source image files can be related to each other by including a unique character string common to the file names of the two source image files as identification information.
[0056] In S402, the system control unit 50 determines whether or not the user has completed the selection of material images. The system control unit 50 repeats the determination process in S402 until the user has completed the selection of material images. When the user has completed the selection of material images, the processing proceeds to S403.
[0057] In S403, the system control unit 50 reads out from the external recording medium 91 the material image file corresponding to the material image selected in S402, and stores it in the memory 25.
[0058] In S404, the system control unit 50 parses (analyzes) the material image file stored in the memory 25 in S403, and extracts the image data (material image), subject detection information, and inference results.
[0059] In S405, the image processing unit 20 generates an HDR composite image with an expanded dynamic range of brightness by combining the two material images obtained in S404. Any existing method in the technical field of generating HDR composite images by image combination can be used as the combining method here. As an example, the image processing unit 20 determines, on a pixel-by-pixel basis, which of the material images has a properly exposed pixel based on a comparison of brightness at corresponding positions in the two material images. The image processing unit 20 then combines the properly exposed pixels contained in the two material images to generate an HDR composite image. The generated HDR composite image is stored in the memory 25.
[0060] 5 is an example of an HDR composite image generated in S405. In this example, the slope 510 is synthesized with the pixels of the slope 510 included in the material image 501, and the mountains 505, sky 508, and clouds 514 are synthesized with the pixels of the mountains 505, sky 508, and clouds 514 included in the material image 502.
[0061] In S406, the system control unit 50 determines inference results to be excluded from the recording target. Specifically, the system control unit 50 identifies inference results that are similar between two material images. For example, in the case of material images 501 and 502 shown in FIG. 5, combinations of mountains 504 and 505, skies 507 and 508, slopes 510 and 511, and clouds 513 and 514 are identified as inference results that are similar between material images 501 and 502. When two inference results (e.g., mountains 504 and 505) are similar between two material images, if both of them are recorded in association with the HDR composite image, the HDR composite image will contain redundant inference results, which will reduce usability for the user. Therefore, when two inference results (e.g., mountains 504 and 505) are similar between two material images, the system control unit 50 excludes one of them from the recording target.
[0062] Here, an example of a method for identifying similar inference results between two material images will be described. The system control unit 50 determines whether the similarity between the inference result of one material image (for example, mountain 504) and the inference result of the other material image (mountain 505) satisfies a predetermined criterion. If the predetermined criterion regarding similarity is satisfied, the system control unit 50 can identify these two inference results as similar inference results between the two material images. The predetermined criterion regarding similarity is not particularly limited, but as an example, a criterion based on the degree of overlap of the subject regions can be adopted. For example, if the degree of overlap between the area of the subject identified by the inference result of one material image (for example, the rectangular area of mountain 504 identified by the coordinates (x11, y11, w11, h11) included in annotation information 601) and the area of the subject identified by the inference result of the other material image (for example, the rectangular area of mountain 505 identified by the coordinates (x21, y21, w21, h21) included in annotation information 602) is equal to or greater than a predetermined degree, the system control unit 50 determines that the similarity between these two inference results satisfies a predetermined criterion.
[0063] Next, two examples of a method for determining inference results to be excluded from recording will be described. For convenience of explanation, attention will be focused on mountains 504 and 505 as inference results that are similar between two material images.
[0064] As a first example, the system control unit 50 determines which of the components derived from mountain 504 and the components derived from mountain 505 is more prevalent in the HDR blended image. If the components derived from mountain 504 are more prevalent than the components derived from mountain 505, it is considered that the exposure of mountain 504 is more appropriate than that of mountain 505, and the reliability or importance of the inference result for mountain 504 is higher than that of mountain 505. Therefore, the system control unit 50 excludes mountain 505 (i.e., mountain 504 is recorded). Conversely, if the components derived from mountain 505 are more prevalent than the components derived from mountain 504, the system control unit 50 excludes mountain 504 (i.e., mountain 505 is recorded).
[0065] As a second example, the system control unit 50 determines which inference results to exclude from recording based on the level of detail of the type of subject indicated by the inference results. As shown in annotation information 601 in FIG. 6, the inference results for each subject include information indicating the type of subject. The information indicating the type of subject may also include more detailed information about the type. For "Subject 1" corresponding to mountain 504 in annotation information 601, there is information indicating that the type is mountain, more specifically, Mount Fuji. On the other hand, for "Subject 1" corresponding to mountain 505 in annotation information 602, there is information indicating that the type is mountain, but no more detailed information. Therefore, the system control unit 50 excludes inference results with relatively less detailed information indicating the type of subject, so that inference results with more detailed information indicating the type of subject are recorded. Therefore, for mountains 504 and 505, the inference result corresponding to mountain 505 is excluded.
[0066] Note that a configuration may be adopted in which similar inference results are not excluded if certain conditions are met. For example, when two material images are captured in a predetermined shooting mode, the system control unit 50 does not exclude the inference results regardless of whether the inference results are similar (thus, for example, both mountains 504 and 505 are recorded). The predetermined shooting mode is, for example, night scene mode. As described in the first example above regarding the method for determining inference results to be excluded from recording, the reliability or importance of inference results for subjects detected in inappropriately exposed areas is generally considered to be low. However, in night scene mode, the inference results for subjects detected in underexposed areas may be important to the user. Therefore, when two material images are captured in night scene mode, the system control unit 50 does not exclude the inference results regardless of whether the inference results are similar.
[0067] In S407, the system control unit 50 records the HDR composite image generated in S405 and the inference results of each material image (excluding the inference results excluded in S406) as an HDR composite image file on the external recording medium 91. The system control unit 50 also includes subject detection information (obtained in S204 and S209 in FIG. 2) corresponding to the recorded inference results in the HDR composite image file. The HDR composite image file may have the same structure as the material image file described with reference to FIG. 3.
[0068] 6 is an example of annotation information (i.e., inference results recorded in association with the HDR composite image) included in the HDR composite image file generated in S407. As can be understood from a comparison with the annotation information 601 and 602 corresponding to the material images 501 and 502, the annotation information 603 corresponds to the case where the above-described first example is used as a method for determining inference results to be excluded from recording.
[0069] Modifications regarding the shooting process and composition process In the above description, the two material images to be combined are assumed to be images captured by a shooting process using different exposure settings (optimal exposure setting and underexposure setting). However, the different exposure settings in the shooting process of this embodiment are not limited to optimal exposure setting and underexposure setting. For example, an overexposure setting may be used instead of the underexposure setting. In the overexposure setting, the system control unit 50 applies an offset to increase the exposure to the result of the AE (auto exposure) processing and feeds it back to the exposure control unit 40, thereby obtaining a brightly exposed image.
[0070] Furthermore, the number of material images to be combined is not limited to 2. For example, a configuration may be adopted in which three material images are taken with different exposure settings (for example, an optimum exposure setting, an underexposure setting, and an overexposure setting), and these three material images are combined.
[0071] In addition, in the above description, the compositing process of the material images is HDR compositing process (compositing process that expands the dynamic range of brightness), but the compositing process of this embodiment is not limited to HDR compositing process and may be, for example, depth compositing process (compositing process that expands the depth of field). In this case, the system control unit 50 photographs two (or more) material images with different focus distance settings instead of photographing two (or more) material images with different exposure settings.
[0072] As described above, according to the first embodiment, the digital camera 100 acquires a plurality of material images (e.g., material image 501 and material image 502) and subject information representing subjects detected in each material image (e.g., information about each subject in annotation information 601 and 602). The digital camera 100 also generates a composite image (e.g., HDR composite image 503) by combining the plurality of material images. The digital camera 100 then records the subject information in association with the composite image. During this recording, if the similarity between the subject information of one material image (first subject information) and the subject information of another material image (second subject information) satisfies a predetermined criterion, the digital camera 100 records either the first subject information or the second subject information. This reduces the redundancy of the subject information (inference result) associated with the composite image, improving the usability of the subject information (inference result) for the user.
[0073] [Other embodiments] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0074] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0075] 11... photographing lens, 13... imaging element, 20... image processing unit, 25... memory, 50... system control unit, 51... non-volatile memory, 72... inference model recording unit, 73... inference engine, 74... learning unit, 100... digital camera
Claims
1. an acquisition means for acquiring a first image, first object information representing a first object detected in the first image, a second image, and second object information representing a second object detected in the second image; a synthesis means for generating a synthesized image by synthesizing the first image and the second image; a recording means for recording one of the first subject information and the second subject information in association with the composite image when the similarity between the first subject information and the second subject information satisfies a predetermined criterion, and for recording both of the first subject information and the second subject information in association with the composite image when the similarity between the first subject information and the second subject information does not satisfy the predetermined criterion; and An image processing device comprising:
2. When the degree of overlap between the area of the first subject specified by the first subject information and the area of the second subject specified by the second subject information is equal to or greater than a predetermined degree, the similarity between the first subject information and the second subject information satisfies the predetermined criterion.
2. The image processing device according to claim 1, wherein:
3. When the similarity between the first object information and the second object information satisfies the predetermined criterion, When the component derived from the first object is more than the component derived from the second object in the composite image, the recording means records the first object information in association with the composite image; When the component derived from the second object is more numerous in the composite image than the component derived from the first object, the recording means records the second object information in association with the composite image.
3. The image processing device according to claim 1, wherein the image processing device is a computer.
4. the first object information includes information indicating a type of the first object, the second object information includes information indicating a type of the second object, When the similarity between the first object information and the second object information satisfies the predetermined criterion, When the information indicating the type of the first subject is more detailed than the information indicating the type of the second subject, the recording means records the first subject information in association with the composite image; When the information indicating the type of the second subject is more detailed than the information indicating the type of the first subject, the recording means records the second subject information in association with the composite image.
3. The image processing device according to claim 1, wherein the image processing device is a computer.
5. the first image and the second image are images taken with different exposure settings, The combining means combines the first image and the second image so as to expand the dynamic range of luminance.
5. The image processing device according to claim 1, wherein the image processing device is a computer.
6. When the first image and the second image are photographed in a predetermined photographing mode, the recording means records both the first object information and the second object information in association with the composite image, regardless of whether the similarity between the first object information and the second object information satisfies the predetermined criterion.
6. The image processing device according to claim 5,
7. The predetermined photographing mode is a night view mode.
7. The image processing device according to claim 6,
8. the first image and the second image are images taken at different focus distance settings, The combining means combines the first image and the second image so as to extend the depth of field.
5. The image processing device according to claim 1, wherein the image processing device is a computer.
9. a detection means for detecting the first subject in the first image and the second subject in the second image; a generating means for generating the first object information representing the first object detected in the first image and generating the second object information representing the second object detected in the second image; Further provided with The acquisition means acquires the first object information and the second object information generated by the generation means.
9. The image processing device according to claim 1, wherein the image processing device is a computer.
10. The generating means generates the first object information and the second object information by performing an inference process using an inference model on the first object detected in the first image and the second object detected in the second image.
10. The image processing device according to claim 9,
11. An image processing device according to any one of claims 1 to 10; an imaging means for generating the first image and the second image; Equipped with The acquisition means acquires the first image and the second image generated by the imaging means. An imaging device characterized by:
12. An image processing method executed by an image processing device, an acquisition step of acquiring a first image, first object information representing a first object detected in the first image, a second image, and second object information representing a second object detected in the second image; a combining step of combining the first image and the second image to generate a combined image; a recording step of recording either the first subject information or the second subject information in association with the composite image when the similarity between the first subject information and the second subject information satisfies a predetermined criterion, and recording both the first subject information and the second subject information in association with the composite image when the similarity between the first subject information and the second subject information does not satisfy the predetermined criterion; An image processing method comprising:
13. A program for causing a computer to function as each of the means of the image processing apparatus according to any one of claims 1 to 10.
Citation Information
Patent Citations
Image reproduction device
JP2013017141A
Information apparatus, server, image file, image file generation method, image file management method, and program
JP2015002441A
Image processing apparatus, image processing method, and program
JP2015099559A
Imaging apparatus
JP2019009577A
Image processing device and image processing method
JP2019041152A