Image processing device, imaging device, image processing method, and program

The image processing apparatus addresses the challenge of overlapping subjects in composite images by synthesizing and recording subject information from individual images, facilitating complete subject recognition.

JP7799475B2Active Publication Date: 2026-01-15CANON KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021206266
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-20
Publication Date
2026-01-15
Estimated Expiration
2041-12-20

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately detect and recognize subjects in composite images where subjects from multiple element images overlap, as they fail to effectively handle overlapping subjects with different postures.

Method used

An image processing apparatus that acquires subject information from individual images and synthesizes them, recording this information alongside the composite image, even when subjects in different postures are superimposed.

Benefits of technology

Enables the retrieval of subject information for subjects that may not be detectable in the composite image, ensuring comprehensive subject recognition across multiple images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007799475000001
    Figure 0007799475000001
  • Figure 0007799475000002
    Figure 0007799475000002
  • Figure 0007799475000003
    Figure 0007799475000003
Patent Text Reader

Abstract

To provide a technology that enables subject information representing a subject to be acquired together with a synthesized image even if the subject detected in a material image cannot be detected in the synthesized image generated from a plurality of material images.SOLUTION: The image processing device includes: acquiring means for acquiring a first image, first subject information representing a first subject detected in the first image, a second image, and second subject information representing a second subject detected in the second image; synthesizing means for generating a synthesized image by synthesizing the first image and the second image; and recording means for recording the first subject information and the second subject information in association with the synthesized image.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device, an imaging device, an image processing method, and a program. [Background technology]

[0002] In recent years, artificial intelligence (AI) technologies such as deep learning have been utilized in various technical fields. For example, a function for detecting human faces from captured images has been known in digital still cameras. Patent Document 1 also discloses a technology for accurately detecting and recognizing animals such as dogs and cats, in addition to detecting humans.

[0003] Furthermore, there are known techniques for creating a composite image by combining a plurality of raw images, such as multiple composition and trajectory composition. In relation to this technique, Patent Document 2 discloses that only the shooting information of an image (raw image) including a main subject is added to the composite image and recorded. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-099559 [Patent Document 2] Japanese Patent Application Publication No. 2019-009577 Summary of the Invention [Problem to be solved by the invention]

[0005] Consider a case where AI technology is used to detect and recognize subjects in a composite image created by combining multiple element images (multiple composition, trajectory composition, etc.). In the composite image, there is a possibility that the subjects in each element image are overlapping in the same location. In such a case, there is a problem that it is difficult to correctly detect and recognize all the subjects included in the composite image. However, the technologies of Patent Documents 1 and 2 cannot address this problem.

[0006] The present invention has been made in view of the above circumstances, and aims to provide a technique that makes it possible to acquire subject information representing a subject along with a composite image generated from a plurality of material images, even when the subject detected in the material image cannot be detected in the composite image generated from the material images. [Means for solving the problem]

[0007] In order to solve the above-mentioned problems, the present invention provides an image processing apparatus including: an acquisition unit that acquires a first image, first subject information representing a first subject detected in the first image, a second image, and second subject information representing a second subject detected in the second image; and a synthesis unit that generates a synthesized image by synthesizing the first image and the second image. ,before The above composite image as an image file and a recording means for recording the The recording means records both the first subject information and the second subject information, which may be impossible to generate from the composite image when the first subject and the second subject are the same subject but with different postures and are superimposed and composited in the composite image, in the image file in which the composite image is stored. The present invention provides an image processing device characterized by the above. [Effects of the Invention]

[0008] According to the present invention, even if a subject detected in a material image cannot be detected in a composite image generated from multiple material images, it is possible to obtain subject information representing this subject together with the composite image.

[0009] Other features and advantages of the present invention will become more apparent from the accompanying drawings and the following detailed description of the preferred embodiment of the present invention. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of a digital camera 100. [Figure 2] 10 is a flowchart of a multiple composite photographing process executed by the digital camera 100. [Figure 3] FIG. 1A is a diagram showing an example of the configuration of a material image file, and FIGS. 1B to 1C are diagrams showing examples of the configuration of a composite image file. [Figure 4] 4 shows material images 401 to 411 and a composite image 412 as examples of material images and composite images obtained as a result of the processing of S203 to S208. FIG. [Figure 5] 10A to 10B are diagrams showing an example of annotation information including inference results for a material image, and FIG. 10C is a diagram showing an example of annotation information including inference results for a composite image. [Figure 6] FIG. 1A shows an example of the configuration of main annotation information, and FIG. 1B shows an example of the configuration of sub-annotation information. [Figure 7] FIG. 10 is a diagram showing an example of the configuration of sub-annotation information. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0012] In the following description, a digital camera (image capture device) is used as an example of an image processing device that performs object classification using an inference model. However, in the following embodiments, the image processing device is not limited to a digital camera. The image processing device in the following embodiments may be any device that has the functions of a digital camera described below, and may be, for example, a smartphone or a tablet PC.

[0013] [First embodiment] ●Configuration of digital camera 100 FIG. 1 is a block diagram showing an example configuration of a digital camera 100. Barrier 10 is a protective member that covers the imaging unit, including the photographing lens 11, of digital camera 100, to prevent the imaging unit from getting dirty or damaged. The operation of barrier 10 is controlled by barrier control unit 43. The photographing lens 11 forms an optical image on the imaging surface of imaging element 13. Shutter 12 has an aperture function. The imaging element 13 is composed of, for example, a CCD or CMOS sensor, and converts the optical image formed on the imaging surface by the photographing lens 11 via shutter 12 into an electrical signal.

[0014] The A / D converter 15 converts the analog image signal output from the image sensor 13 into a digital image signal. The digital image signal converted by the A / D converter 15 is written to memory 25 as so-called RAW image data. At the same time, development parameters corresponding to each piece of RAW image data are generated based on information at the time of shooting and written to memory 25. The development parameters consist of various parameters used in image processing for recording images in the JPEG format or the like, such as exposure settings, white balance, color space, and contrast.

[0015] The timing generating unit 14 is controlled by the memory control unit 22 and the system control unit 50 and supplies clock signals and control signals to the image sensor 13, the A / D converter 15, and the D / A converter 21.

[0016] The image processing unit 20 performs various image processing such as predetermined pixel interpolation processing, color conversion processing, correction processing, resizing processing, and image synthesis processing on data from the A / D converter 15 or data from the memory control unit 22. The image processing unit 20 also performs predetermined image processing and arithmetic processing using image data obtained by capturing an image, and provides the obtained arithmetic results to the system control unit 50. The system control unit 50 controls the exposure control unit 40 and focus control unit 41 based on the provided arithmetic results, thereby realizing AF (autofocus) processing, AE (autoexposure) processing, and EF (flash pre-flash) processing.

[0017] The image processing unit 20 also performs predetermined calculations using the captured image data and performs AWB (auto white balance) processing based on the calculation results. Furthermore, the image processing unit 20 reads image data stored in the memory 25 and performs compression or decompression processing using a method such as the JPEG method, the MPEG-4 AVC method, the HEVC (High Efficiency Video Coding) method, or a lossless compression method for uncompressed RAW data. The image processing unit 20 then writes the processed image data to the memory 25.

[0018] The image processing unit 20 also performs predetermined arithmetic processing using captured image data to edit various types of image data. For example, the image processing unit 20 can perform a trimming process to adjust the display range and size of an image by hiding unnecessary parts around the image data, and a resizing process to enlarge or reduce the size of image data and screen display elements. Furthermore, the image processing unit 20 can perform RAW development, which applies image processing such as color conversion to data that has been compressed or expanded using a lossless compression method on uncompressed RAW data, converts it to JPEG format, and creates image data. The image processing unit 20 can also perform video clipping, which clips out specified frames from a video format such as MPEG-4, converts them to JPEG format, and saves them.

[0019] The image processing unit 20 also includes a compositing processing circuit that combines multiple pieces of image data. In this embodiment, the image processing unit 20 is capable of executing additive compositing processing, weighted additive compositing processing, comparatively bright compositing processing, and comparatively dark compositing processing. The comparatively bright compositing processing is processing that generates one composite image from multiple material images by selecting the brightest pixel value among the multiple material images as the pixel value of each pixel of the composite image. The comparatively dark compositing processing is processing that generates one composite image from multiple material images by selecting the darkest pixel value among the multiple material images as the pixel value of each pixel of the composite image.

[0020] The image processing unit 20 also performs processing such as superimposing an OSD (On-Screen Display) such as a menu or arbitrary characters to be displayed on the display unit 23 together with the image data for display.

[0021] Furthermore, the image processing unit 20 performs subject detection processing to detect the subject present in the image data and the subject area by using the input image data and information on the distance to the subject obtained from the image sensor 13 at the time of shooting. Detectable information (subject detection information) includes information on the position, size, and tilt of the subject area in the image, as well as information on the likelihood.

[0022] The memory control unit 22 controls the A / D converter 15, the timing generation unit 14, the image processing unit 20, the image display memory 24, the D / A converter 21, and the memory 25. The RAW image data generated by the A / D converter 15 is written to the image display memory 24 or the memory 25 via the image processing unit 20 and the memory control unit 22, or directly via the memory control unit 22.

[0023] The image data for display written in the image display memory 24 is displayed on a display unit 23 configured by a TFT LCD or the like via a D / A converter 21. By using the display unit 23 to sequentially display image data obtained by capturing an image, it is possible to realize an electronic finder function that displays a live image.

[0024] The memory 25 has a storage capacity sufficient to store a predetermined number of still images and a predetermined period of moving images, and stores the captured still images and moving images. The memory 25 can also be used as a working area for the system control unit 50.

[0025] The exposure control unit 40 controls the shutter 12, which has an aperture function. The exposure control unit 40 also has a flash dimming function by working in conjunction with a flash 44. The focus control unit 41 adjusts focus by driving a focus lens (not shown) included in the photographing lens 11 based on instructions from the system control unit 50. The zoom control unit 42 controls zooming by driving a zoom lens (not shown) included in the photographing lens 11. The flash 44 has an AF assist light projection function and a flash dimming function.

[0026] The system control unit 50 controls the entire digital camera 100. The nonvolatile memory 51 is an electrically erasable and recordable nonvolatile memory, such as an EEPROM. The nonvolatile memory 51 stores not only programs but also map information and the like.

[0027] The shutter switch 61 (SW1) is turned on during operation of the shutter button 60, and instructs the start of operations such as AF processing, AE processing, AWB processing, and EF processing. The shutter switch 62 (SW2) is turned on when the shutter button 60 is completely operated, and instructs the start of a series of shooting operations including exposure processing, development processing, and recording processing. In the exposure processing, a signal read from the image sensor 13 is written to the memory 25 as RAW image data via the A / D converter 15 and the memory control unit 22. In the development processing, the RAW image data written to the memory 25 is developed through calculations in the image processing unit 20 and the memory control unit 22, and then written to the memory 25 as image data. In the recording processing, the image data is read from the memory 25, compressed by the image processing unit 20, stored in the memory 25, and then written to the external recording medium 91 via the card controller 90.

[0028] The operation unit 63 includes various operation members such as buttons and a touch panel. For example, the operation unit 63 includes a power button, a menu button, a mode switch for switching between shooting mode, playback mode, and other special shooting modes, a cross key, a set button, a macro button, and a multi-screen playback page break button. Furthermore, for example, the operation unit 63 includes a flash setting button, a single-shot / continuous-shot / self-timer switching button, a menu movement + (plus) button, a menu movement - (minus) button, a shooting quality selection button, an exposure compensation button, a date / time setting button, etc.

[0029] When recording image data on the external recording medium 91, the metadata generation and analysis unit 70 generates various metadata, such as information conforming to the Exchangeable Image File Format (Exif) standard to be attached to the image data, based on information at the time of shooting. Furthermore, when reading image data recorded on the external recording medium 91, the metadata generation and analysis unit 70 analyzes the metadata attached to the image data. Examples of metadata include shooting setting information at the time of shooting, image data information related to the image data, and feature information of the subject included in the image data. Furthermore, when recording moving image data, the metadata generation and analysis unit 70 can also generate and attach metadata for each frame.

[0030] The power supply 80 includes a primary battery such as an alkaline battery or a lithium battery, a secondary battery such as a NiCd battery, a NiMH battery, or a Li battery, or an AC adapter, etc. The power supply control unit 81 supplies the power from the power supply 80 to each unit of the digital camera 100.

[0031] The card controller 90 transmits and receives data to and from an external recording medium 91 such as a memory card. The external recording medium 91 is configured, for example, as a memory card, and records images (still images and videos) captured by the digital camera 100.

[0032] The inference engine 73 uses the inference model recorded in the inference model recording unit 72 to perform inference on image data input via the system control unit 50. The system control unit 50 can record inference models input from an external device (not shown) via the communication unit 71 in the inference model recording unit 72. The system control unit 50 can also record inference models obtained by re-learning the inference model using the learning unit 74 in the inference model recording unit 72. Note that the inference models recorded in the inference model recording unit 72 may be updated by inputting an inference model from an external device or by re-learning the inference model using the learning unit 74. Therefore, the inference model recording unit 72 holds version information so that the version of the inference model can be identified.

[0033] The inference engine 73 also has a neural network design 73a. The neural network design 73a has a configuration in which an intermediate layer (neurons) is arranged between an input layer and an output layer. Image data is input to the input layer from the system control unit 50. Several layers of neurons are arranged in the intermediate layer. The number of neuron layers is determined appropriately in the design. The number of neurons in each layer is also determined appropriately in the design. In the intermediate layer, weighting is performed based on the inference model recorded in the inference model recording unit 72. In the output layer, an inference result according to the image data input to the input layer is output.

[0034] In this embodiment, the inference model recorded in the inference model recording unit 72 is assumed to be an inference model that infers the classification of the subject included in the image. An inference model generated by deep learning is used, using as training data image data of various subjects and the results of their classification (for example, classification of animals such as dogs and cats, or classification of subject types such as people, animals, plants, and buildings). Therefore, when an image and information indicating the area of ​​the subject detected in this image are input to the inference engine 73 that uses the inference model, an inference result indicating the classification of this subject is output.

[0035] The learning unit 74 re-learns the inference model upon receiving a request from the system control unit 50 or the like. The learning unit 74 has a teacher data recording unit 74a. The teacher data recording unit 74a records information related to teacher data for the inference engine 73. The learning unit 74 re-learns the inference engine 73 using the teacher data recorded in the teacher data recording unit 74a, and can update the inference engine 73 using the inference model recording unit 72.

[0036] The communication unit 71 has a communication circuit for transmitting and receiving. Specifically, the communication performed by the communication circuit may be wireless communication such as Wi-Fi or Bluetooth (registered trademark), or wired communication such as Ethernet or USB.

[0037] Composition processing by the image processing unit 20 The following describes the compositing process performed by the image processing unit 20 to combine multiple image data (multiple material images). The image processing unit 20 can perform four compositing processes: additive compositing, weighted additive compositing, comparatively bright compositing, and comparatively dark compositing. Let I_i(x, y) be the pixel value of image i (i = 1 to N) before compositing, where x and y represent coordinates within the screen, and let I(x, y) be the pixel value of the composite image. The pixel values ​​may be the values ​​of the R, G1, G2, and B signals in the Bayer array, or the value of a luminance signal (luminance value) obtained from a group of R, G1, G2, and B signals. In this case, the Bayer array signals may be interpolated so that R, G, and B signals are present for each pixel, and then the luminance value for each pixel may be calculated. For example, the luminance value may be calculated by weighted addition of the R, G, and B signals, such as Y = 0.3 × R + 0.59 × G + 0.11 × B, where Y is the luminance value. The synthesis process is performed based on pixel values ​​that have been aligned by performing processes such as alignment between multiple images as necessary.

[0038] The additive synthesis process is performed in accordance with the following equation: That is, the image processing unit 20 performs an additive process on the pixel values ​​of N images for each pixel to generate a synthesized image. I(x,y)=I_1(x,y)+I_2(x,y)+···+I_N(x,y)

[0039] The weighted addition synthesis process is performed according to the following formula, where ai (i = 1 to N) is a weighting coefficient. That is, the image processing unit 20 generates a synthesized image by performing weighted addition process on the pixel values ​​of N images for each pixel. When a1 + a2 + + aN = 1, the following formula corresponds to weighted average process. I(x,y)=a1×I_1(x,y)+a2×I_2(x,y)+···+aN×I_N(x,y)

[0040] The lightening combination process is performed according to the following formula: That is, the image processing unit 20 generates a combined image by selecting the maximum pixel value of the N images for each pixel. I(x,y)=max(I_1(x,y),I_2(x,y),...,I_N(x,y))

[0041] The comparatively dark combination process is performed according to the following formula: That is, the image processing unit 20 generates a combined image by selecting the minimum pixel value of the N images for each pixel. I(x,y)=min(I_1(x,y),I_2(x,y),...,I_N(x,y))

[0042] ●Multiple composite shooting processing Next, the multiple composite shooting process executed by the digital camera 100 will be described with reference to Figures 2 to 7. Figure 2 is a flowchart of the multiple composite shooting process executed by the digital camera 100. Unless otherwise specified, the processing of each step in this flowchart is realized by the system control unit 50 of the digital camera 100 controlling each component of the digital camera 100 in accordance with a program. When the operation mode of the digital camera 100 is set to multiple shooting mode, the multiple composite shooting process of this flowchart begins. Note that the user can set the operation mode of the digital camera 100 to multiple shooting mode by operating the operation unit 63 to display a menu screen on the display unit 23 and selecting multiple shooting mode on the menu screen.

[0043] In S202, the system control unit 50 determines whether or not a shooting instruction has been issued by the user. The user can issue a shooting instruction by pressing the shutter button 60 to turn on the shutter switches 61 (SW1) and 62 (SW2). The system control unit 50 repeats the determination process in S202 until a shooting instruction has been issued by the user. When a shooting instruction has been issued by the user, the processing proceeds to S203.

[0044] The processes of S203 to S208 are repeatedly executed until it is determined in S209, which will be described later, that the shooting instruction is no longer continuing. In the following explanation, it is assumed that the processes of S203 to S208 have been executed 11 times (thus, 11 material images have been generated). Fig. 4 shows material images 401 to 411 and composite image 412 as examples of material images and composite images obtained as a result of the processes of S203 to S208.

[0045] In S203, the system control unit 50 performs an image capturing process. In the image capturing process, the system control unit 50 performs an AF (autofocus) process and an AE (autoexposure) process using the focus control unit 41 and the exposure control unit 40, and then stores the image signal output from the image sensor 13 via the A / D converter 15 in the memory 25. The image processing unit 20 also performs a compression process on the image signal stored in the memory 25 in accordance with the user's settings, thereby generating image data in a format (for example, JPEG format) in accordance with the user's settings.

[0046] In S204, the image processing unit 20 performs subject detection processing on the image signal stored in the memory 25, and acquires information about the subject included in the image (subject detection information).

[0047] In S205, the system control unit 50 uses the inference engine 73 to perform inference processing on the subject detected in the image signal (material image) stored in the memory 25. The system control unit 50 identifies the subject area in the image based on the image signal stored in the memory 25 and the subject detection information acquired in S204. The system control unit 50 inputs the image signal (material image) and information indicating the subject area in the material image to the inference engine 73. As a result of the inference processing performed by the inference engine 73 for each subject area, an inference result indicating the classification of the subject included in the subject area is output. Note that in addition to the inference result, the inference engine 73 may output information related to the inference processing, such as debug information and logs related to the operation of the inference processing.

[0048] In S206, the system control unit 50 records a file including the image data generated in S203, the subject detection information acquired in S204, and the inference results acquired in S205 on the external recording medium 91 as a multiple composite material image file.

[0049] FIG. 3A shows an example of the structure of a material image file. As shown in FIG. 3A, the material image file 300 is divided into multiple storage areas, including an Exif area 301 that stores metadata in accordance with the Exif standard and an image data area 308 that records compressed image data. The material image file 300 also includes an annotation information area 310 that records annotation information. If the material image file 300 is a JPEG format file, each of the multiple storage areas is defined by a marker. For example, if a user instructs image recording in JPEG format, the material image file 300 is recorded in JPEG format. In this case, the image data generated in S203 is recorded in the image data area 308 in JPEG format, and the information in the Exif area 301 is recorded in an area defined by, for example, an APP1 marker. The information in the annotation information area 310 is recorded in an area defined by, for example, an APP11 marker. When a user instructs that an image be recorded in the High Efficiency Image File Format (HEIF), the material image file 300 is recorded in the HEIF file format. In this case, the information in the Exif area 301 and the annotation information area 310 is recorded in a Metadata Box or the like. Similarly, when a user instructs that an image be recorded in the RAW format, the information in the Exif area 301 and the annotation information area 310 is recorded in a predetermined area such as a Metadata Box.

[0050] The subject detection information acquired in S204 is recorded by the metadata generation and analysis unit 70 in a subject detection information tag 306 in a MakerNote 305 (an area where manufacturer-specific metadata can be written in a confidential format in principle) included in the Exif area 301. In addition, if there is version information of the current inference model recorded in the inference model recording unit 72 or debug information output by the inference engine 73 in S205, this information is recorded in the MakerNote 305 as inference model management information 307.

[0051] The inference result obtained in S205 is recorded as annotation information in the annotation information area 310. The location of the annotation information area 310 is indicated by the annotation information link 303 included in the annotation link information storage tag 302. In this embodiment, it is assumed that the annotation information is written in a text format such as XML or JSON.

[0052] 5(a) and 5(b) are diagrams showing examples of annotation information including inference results for material images. The system control unit 50 manages the same subject included in multiple material images captured consecutively using the same subject number (subject identification information that identifies the subject). For example, since subject 502 in material images 401 and 411 does not move, the same inference result is recorded for subject 502 as "subject 1" for both material images 401 and 411. Furthermore, subject 503 in material image 401 and subject 504 in material image 411 are the same subject, although their postures are different. Therefore, both subject 503 and subject 504 are recorded as "subject 2." Of the inference results for "subject 2," information on the subject's position (coordinates of head position, eye position, etc.) varies between material images, but the same information is recorded for other information (gender, age, name, etc.) for each material image.

[0053] 2, in S207, the image processing unit 20 performs a synthesis process on the material images. In the first process of S207 (i.e., when processing the material image 401), the image processing unit 20 saves the image data generated in S202 as a synthetic image in a synthetic image area of ​​the memory 25. In the second or subsequent process of S207 (i.e., when processing any of the material images 402 to 411), the image processing unit 20 synthesizes the synthetic image saved in the synthetic image area of ​​the memory 25 with the image data created in S202, and saves the synthesized image in the synthetic image area of ​​the memory 25 as a new synthetic image.

[0054] In S208, the system control unit 50 performs a process of generating sub-annotation information for the composite image based on the inference result obtained in S205 (i.e., the inference result of the material image). Specifically, in the first process of S208 (i.e., when processing the material image 401), the system control unit 50 generates sub-annotation information including the inference result obtained in S205 in the memory 25. In the second or subsequent process of S207 (i.e., when processing any of the material images 402 to 411), the system control unit 50 adds information about the inference result obtained in S205 to the sub-annotation information stored in the memory 25. This makes it possible to carry over the inference result of the material image to the composite image.

[0055] 6(b) and 7(a) are diagrams showing configuration examples of sub-annotation information. As shown in FIG. 6(b), the system control unit 50 may simply add the inference results obtained in S205 for each material image to the sub-annotation information. In this case, the finally obtained sub-annotation information includes all inference results corresponding to all material images. Alternatively, as shown in FIG. 7(a), the system control unit 50 may add, to the sub-annotation information, difference information between the inference results obtained in S205 and existing inference results included in the sub-annotation information.

[0056] In S209, the system control unit 50 determines whether the user is continuing to issue a shooting instruction. The user can continue to issue a shooting instruction by continuing to press the shutter button 60 and keeping the shutter switches 61 (SW1) and 62 (SW2) ON. If the shooting instruction is continuing, the processing returns to S203; if the shooting instruction is not continuing, the processing proceeds to S210.

[0057] In S210, the image processing unit 20 performs subject detection processing on the composite image generated by the processing of S207, and acquires information on the subjects included in the composite image (subject detection information). The processing of S210 is similar to the processing of S204, except that the processing target is the composite image rather than the material image.

[0058] In S211, the system control unit 50 performs inference processing on the composite image using the inference engine 73. The processing in S211 is similar to the processing in S205, except that the processing target is the composite image rather than the material images. FIG. 5(c) is a diagram showing an example of annotation information including the inference result for the composite image. Note that the system control unit 50 manages the same subject included in one or more material images and the composite image using the same subject number (subject identification information for identifying the subject). For example, as can be seen from FIGS. 5(a) to 5(c), the subject 502 included in the composite image 412 is the same subject as the subject 502 included in the material images 401 and 411, and therefore these subjects are all recorded as "subject 1." Furthermore, at the positions of subjects 503 and 504 included in the material images 401 and 411, the subjects move in each material image, so multiple subjects overlap in the composite image. Since the subject is not detected from the overlapping subjects and it cannot be inferred that the subject is a person, no subject corresponding to a person is recorded in the inference result for the composite image.

[0059] In S212, the system control unit 50 records a file including the composite image generated in S207, the sub-annotation information generated in S207, the subject detection information obtained in S210, and the inference result obtained in S211 as a composite image file on the external recording medium 91.

[0060] 3(b) and 3(c) are diagrams showing examples of the structure of a composite image file. As shown in Fig. 3(b) and 3(c), the composite image generated in S207 is saved in the image data area 308 of the composite image file 320 or 330. In addition, the subject detection information acquired in S210 is recorded in the subject detection information tag 306 in the MakerNote 305 of the composite image file 320 or 330.

[0061] In the case of the composite image file 320 shown in FIG. 3(b), the inference result obtained from the composite image in S211 is recorded in a main annotation information area 323. Furthermore, the sub-annotation information generated in S208 is recorded in a sub-annotation information area 324. In the case of FIG. 3(b), the main annotation information area 323 and the sub-annotation information area 324 are storage areas defined by separate APP11 markers or separate Meta data boxes, etc. The location of the main annotation information area 323 is indicated by a main annotation information link 321 included in the annotation link information storage tag 302. The sub-annotation information area 324 is indicated by a sub-annotation information link 322 included in the annotation link information storage tag 302.

[0062] In the case of the composite image file 330 shown in Fig. 3(c), the main annotation information and sub-annotation information are recorded in the same storage area (annotation information area 310), such as an area specified by an APP11 marker or a Metadata Box. In the annotation information area 310, the main annotation information and sub-annotation information are stored separately in separate tags (main annotation information tag 331 and sub-annotation information tag 332). The location of the annotation information area 310 is indicated by the annotation information link 303 included in the annotation link information storage tag 302.

[0063] FIG. 6(a) is a diagram showing an example of the configuration of main annotation information including an inference result, which is recorded in the main annotation information area 323 or the main annotation information tag 331. As shown in FIG. 6(a), the main annotation information may include information identifying an image (image identification information), such as the file number of a composite image file, recorded in association with the inference result of a subject detected in the composite image. Similarly, as shown in FIGS. 6(b) and 7(a), the sub-annotation information may include information identifying a material image (image identification information), such as the number of a material image file, recorded in association with the inference result of a subject detected in the material image. Alternatively, as shown in FIG. 7(b), the sub-annotation information may not include information identifying a material image (image identification information), such as the number of the material image file. For example, if the material image file is not saved (if the material image is discarded after the composite image is generated), information identifying the material image is unnecessary. In such a case, the configuration of FIG. 7(b) may be adopted.

[0064] As described above, according to the first embodiment, the digital camera 100 acquires a plurality of material images (for example, material image 401 and material image 402) and subject information (for example, information including the inference results by the inference engine 73) that represents the subject detected in each material image. The digital camera 100 also generates a composite image by combining the plurality of material images. Then, the digital camera 100 records the subject information of each material image in association with the composite image, for example, by generating and recording a composite image file that includes the subject information of each material image and the composite image.

[0065] In this way, according to the first embodiment, the subject information of each material image is recorded in association with the composite image. Therefore, even if a subject detected in a material image cannot be detected in a composite image generated from multiple material images, it is possible to obtain subject information representing this subject together with the composite image.

[0066] [Other embodiments] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0067] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0068] 11... photographing lens, 13... imaging element, 20... image processing unit, 25... memory, 50... system control unit, 51... non-volatile memory, 72... inference model recording unit, 73... inference engine, 74... learning unit, 100... digital camera

Claims

1. an acquisition means for acquiring a first image, first object information representing a first object detected in the first image, a second image, and second object information representing a second object detected in the second image; a synthesis means for generating a synthesized image by synthesizing the first image and the second image; a recording means for recording the composite image as an image file; Equipped with The recording means records both the first subject information and the second subject information, which may be impossible to generate from the composite image when the first subject and the second subject are the same subject but have different postures and are superimposed and composited in the composite image, in the image file in which the composite image is stored.

1. An image processing device comprising:

2. When the first subject and the second subject are the same subject, the recording means records the second subject information as difference information with respect to the first subject information.

2. The image processing device according to claim 1, wherein:

3. The recording means records first image identification information for identifying the first image in the image file in association with the first subject information, and records second image identification information for identifying the second image in the image file in association with the second subject information.

3. The image processing device according to claim 1, wherein the image processing device is a computer.

4. a generating unit configured to detect a subject from the composite image and generate subject information representing the subject; If the first subject and the second subject are the same third subject that does not move, the generating means generates third object information representing the third object detected from the composite image; the recording means records the third subject information in the image file in which the composite image is stored; When the first subject and the second subject are the same fourth subject with different postures and are superimposed and combined in the combined image, and therefore cannot be detected from the combined image, the generating means does not generate fourth object information representing the fourth object from the composite image, and the recording means does not record the fourth object information in the image file; 3. The image processing device according to claim 1, wherein the image processing device is a computer.

5. The image file is divided into a plurality of storage areas, the first object information and the second object information are stored in a first storage area among the plurality of storage areas; The third object information is stored in a second storage area different from the first storage area among the plurality of storage areas.

5. The image processing device according to claim 4.

6. The image file is divided into a plurality of storage areas, The first object information, the second object information, and the third object information are stored in the same storage area among the plurality of storage areas.

5. The image processing device according to claim 4.

7. the image file is a JPEG format file, Each of the plurality of storage areas is defined by a marker.

7. The image processing device according to claim 5, wherein the image processing device is a computer.

8. the first object information includes first object identification information that identifies the first object, the second object information includes second object identification information that identifies the second object, When the first subject and the second subject are the same subject, the first subject identification information is equal to the second subject identification information.

8. The image processing device according to claim 4, wherein the image processing device is a computer.

9. The generating means generates the third object information by performing an inference process on the third object detected in the composite image using an inference model.

9. The image processing device according to claim 4, wherein the image processing device is a computer.

10. The inference model is configured to infer a classification of an object.

10. The image processing device according to claim 9,

11. The image processing device according to any one of claims 1 to 3; an imaging means for generating the first image and the second image; a generation means for detecting the first subject in the first image, detecting the second subject in the second image, generating the first subject information representing the first subject detected in the first image, and generating the second subject information representing the second subject detected in the second image; Equipped with The acquisition means acquires the first image and the second image generated by the imaging means, and the first subject information and the second subject information generated by the generation means. An imaging device characterized by:

12. An image processing device according to any one of claims 4 to 10; an imaging means for generating the first image and the second image; Equipped with the generating means detects the first subject in the first image, detects the second subject in the second image, generates the first subject information representing the first subject detected in the first image, and generates the second subject information representing the second subject detected in the second image; The acquisition means acquires the first image and the second image generated by the imaging means, and the first subject information and the second subject information generated by the generation means. An imaging device characterized by:

13. An image processing method executed by an image processing device, an acquisition step of acquiring a first image, first object information representing a first object detected in the first image, a second image, and second object information representing a second object detected in the second image; a combining step of combining the first image and the second image to generate a combined image; a recording step of recording the composite image as an image file; Equipped with The recording step records both the first subject information and the second subject information, which may be impossible to generate from the composite image when the first subject and the second subject are the same subject but have different postures and are superimposed and composited in the composite image, in the image file in which the composite image is stored. An image processing method comprising:

14. A program for causing a computer to function as each of the means of the image processing apparatus according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Image processing apparatus and program

    JP2011211636A

  • Control device and storage medium

    JP2015001609A

  • Image processing apparatus, image processing method, and program

    JP2015099559A

  • Imaging apparatus

    JP2019009577A