Image processing device, image processing method, and program

The image processing system improves image selection by analyzing captions and features to prioritize images based on user-defined preferences, enhancing the accuracy of album creation.

JP7814924B2Active Publication Date: 2026-02-17CANON KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021212731
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-27
Publication Date
2026-02-17
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

Existing image selection techniques fail to accurately prioritize images based on user-defined preferences and captions, leading to undesirable selections.

Method used

An image processing system that analyzes captions and image features to determine a priority subject, generates captions for uncaptioned images, and scores images based on these factors to improve selection accuracy.

Benefits of technology

Enhances the selection process by prioritizing images that align with user preferences and captions, ensuring more appropriate image choices for album creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007814924000001
    Figure 0007814924000001
  • Figure 0007814924000002
    Figure 0007814924000002
  • Figure 0007814924000003
    Figure 0007814924000003
Patent Text Reader

Abstract

To preferably select images.SOLUTION: A program for selecting an image from a candidate image group causes a computer to function as: obtaining means for obtaining the candidate image group including a plurality of images; determining means for determining a specific condition for preferentially selecting an image from the candidate image group: image analyzing means for analyzing the images in the candidate image group; caption analyzing means for analyzing captions attached to the images in the candidate image group; and selecting means for selecting a specific image from the candidate image group based on results of the determining means, the image analyzing means, and the caption analyzing means.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technique for selecting an image. [Background technology]

[0002] There is an automatic layout technology that automatically selects images for creating an album from among a plurality of images, automatically determines an album template, and automatically assigns images to the template.

[0003] Patent Document 1 discloses a technology for recognizing a subject to be preferentially laid out (hereinafter referred to as a priority subject) and at least one sub-subject, estimating the state of the priority subject based on the relationship between the recognized subjects, and selecting an image based on the state of the priority subject.

[0004] Patent Document 2 discloses a technique for selecting a template or stamp image based on comments attached to the laid-out images when creating an album from images posted on a social networking service (SNS). This method calculates a score based on the relevance of the image to predetermined keywords, making it possible to select a highly relevant template or stamp image. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Publication No. 2018-097492 [Patent Document 2] Patent Publication No. 2021-071870 [Non-patent literature]

[0006] [Non-Patent Document 1] Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. “Show and Tell: A Neural Image Caption Generator”, 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 3156-3164 [Non-patent document 2] Tomas Mikolov, Kai Chen, Greg S. Corrado, and Jeffrey Dean “Efficient Estimation of Word Representations in Vector Space”, International Conference on Learning Representations (ICLR), 2013 Summary of the Invention [Problem to be solved by the invention]

[0007] A technique for selecting an appropriate image is required.

[0008] Therefore, an object of the present invention is to suitably select an image. [Means for solving the problem]

[0009] A program according to one aspect of the present invention is a program for selecting an image from a group of candidate images, the program comprising: an acquisition unit for acquiring the group of candidate images including a plurality of images; a determination unit for determining specific conditions for preferentially selecting images from the group of candidate images; and an image analysis unit for analyzing images in the group of candidate images. a caption generating means for generating a caption when a caption is not added to an image in the candidate image group; Captions attached to the images in the candidate image group or a caption generated by the caption generating means a caption analysis means for analyzing the image; and a selection means for selecting a specific image from the group of candidate images based on the results of the determination means, the image analysis means, and the caption analysis means. The caption analysis means analyzes the caption while the caption generation means is generating the caption. It is characterized by: [Effects of the Invention]

[0010] According to the present invention, an image can be suitably selected. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 10 is a diagram illustrating a problem in a comparative example. [Figure 2] FIG. 2 is a block diagram showing the hardware configuration of the image processing apparatus. [Figure 3] FIG. 2 is a software block diagram of an album creation application. [Figure 4] FIG. 10 is a diagram illustrating an example of a UI provided by an album creation application. [Figure 5] 10 is a flowchart illustrating an automatic layout process. [Figure 6] FIG. 10 is a diagram illustrating image feature amounts. [Figure 7] FIG. 1 is a diagram illustrating an automatic caption generation model. [Figure 8] FIG. 10 is a diagram showing caption analysis information. [Figure 9] 10 is a flowchart showing a scoring process. [Figure 10] FIG. 10 is a diagram showing a group of templates used for layout of image data. [Figure 11] 10A and 10B are diagrams illustrating the effects of the embodiment. [Figure 12] 10 is a flowchart illustrating an automatic layout process. [Figure 13] 10 is a flowchart showing a scoring process. [Figure 14] 10 is a flowchart illustrating an automatic layout process. [Figure 15] 10 is a flowchart showing a caption generation and analysis process. [Figure 16] 10 is a flowchart showing a caption generation and analysis process. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, preferred embodiments of the image processing apparatus according to the present invention will be described in detail with reference to the accompanying drawings, however, the scope of the invention is not limited to the illustrated examples.

[0013] Before describing this case, as a comparative example, image selection using information on a priority subject when image caption analysis, which will be described later, is not used will be described with reference to FIG. 1. FIG. 1(a) is an image primarily capturing a train, and FIG. 1(b) is an image in which a person who crossed in front of the camera while capturing the image of FIG. 1(a) is captured. When the priority subject is set to "train," it is assumed that the user expects the image of FIG. 1(a) to be automatically selected, and the image of FIG. 1(b) should not be selected. However, in the comparative example, the subjects recognized from these images are "train" and "person" in both FIG. 1(a) and FIG. 1(b). Therefore, both FIG. 1(a) and FIG. 1(b) are determined to contain the priority subject "train," and control is performed to select them preferentially. Therefore, conventional methods may result in the selection of an undesirable image of the priority subject.

[0014] In the following embodiments, a method for improving the accuracy of image selection by setting a priority subject, acquiring a caption associated with an image, and using information obtained by analyzing the acquired caption will be described. In the following embodiments, a caption specifically refers to text associated with an image. The caption is added or set to an image by an application other than an application for creating an album (hereinafter also referred to as an "app"), which will be described later. Specifically, the other application may be, for example, a social networking service (SNS) application that allows users to post images to SNS, or an image management application that allows users to manage multiple images and view them. A user can enter any text in these applications to add or set a caption to an image. Text automatically generated by the application may be added or set to an image as a caption. In this case, the other application may be an application that analyzes an image and automatically adds or sets text appropriate to the analysis results as a caption to the image. The album creation application, which will be described later, realizes the following embodiments by, for example, acquiring and analyzing captions set by the other application as described above. Note that the caption is not limited to the above-mentioned form, and may be information other than information added by the app, such as EXIF ​​information added or set by the camera to a photographic image.

[0015] <<First Embodiment>> <System Description> In this embodiment, a method for automatically generating a layout by running an application for creating an album on the image processing device 200 will be described as an example. In the following description, unless otherwise specified, the term "image" includes still images, videos, and frame images extracted from videos. Furthermore, the term "image" may also include still images, videos, and frame images from videos that are stored on a network, such as a network service or network storage, and that can be obtained via a network.

[0016] 2 is a block diagram showing the hardware configuration of the image processing device 200. The image processing device 200 may be, for example, a personal computer (hereinafter referred to as a PC) or a smartphone. In this embodiment, the image processing device 200 will be described as a PC. The image processing device 200 includes a CPU 201, a ROM 202, a RAM 203, a HDD 204, a display 205, a keyboard 206, a pointing device 207, and a data communication unit 208.

[0017] The CPU (Central Processing Unit or Processor) 201 comprehensively controls the image processing device 200 and, for example, reads out a program stored in the ROM 202 into the RAM 203 and executes it, thereby realizing the operation of this embodiment. Although FIG. 2 shows one CPU, the system may be configured with multiple CPUs. The ROM 202 is a general-purpose ROM and stores, for example, a program executed by the CPU 201. The RAM 203 is a general-purpose RAM and is used, for example, as a working memory for temporarily storing various information when the CPU 201 executes a program. The HDD (Hard Disk) 204 is a storage medium (storage unit) for storing image files, a database for holding processing results such as image analysis, and templates used by an album creation application.

[0018] The display 205 displays to the user the user interface (UI) of this embodiment and an electronic album as a layout result of image data (hereinafter also referred to as "images"). The keyboard 206 and pointing device 207 accept instructions and operations from the user. The display 205 may have a touch sensor function. The keyboard 206 is used, for example, when the user inputs the number of spreads of the album he or she wants to create on the UI displayed on the display 205. Note that in this specification, the term "spread" corresponds to one display window in a display and is an area that typically corresponds to two pages in a printout, and refers to a pair of adjacent pages printed on a sheet that the user can view at a glance. The pointing device 207 is used, for example, when the user clicks a button on the UI displayed on the display 205.

[0019] The data communication unit 208 communicates with external devices such as SNS or cloud services via a wired or wireless network. The data communication unit 208 transmits, for example, data laid out by the automatic layout function to a printer or server that can communicate with the image processing device 200. The data communication unit 208 also transmits data related to the automatic layout processing to an external cloud computer so that part or all of the automatic layout processing, which will be described later, can be performed by the cloud computer. A data bus 209 connects the blocks in FIG. 2 so that they can communicate with each other.

[0020] 2 is merely an example, and is not limited to this. For example, the image processing device 200 may not have the display 205, and the UI may be displayed on an external display.

[0021] The album creation application in this embodiment is stored in the HDD 204. As will be described later, the application is started by the user selecting the application icon displayed on the display 205 with the pointing device 207 and clicking or double-clicking the icon.

[0022] <Software block description> Fig. 3 is a diagram showing the software blocks of the album creation application. The album creation application described above includes program modules corresponding to the components shown in Fig. 3. The CPU 201 executes the program modules, causing the CPU 201 to function as each component shown in Fig. 3. Hereinafter, each component shown in Fig. 3 will be described assuming that each component executes various processes. Fig. 3 also shows a software block diagram particularly relating to an automatic layout processing unit 318 that executes the automatic layout function.

[0023] The album creation condition specification unit 301 specifies album creation conditions to the automatic layout processing unit 318 in response to UI operations using the pointing device 207. In this embodiment, the album creation conditions can include a group of album candidate images including candidate images to be used in the album, the number of spreads, the type of template, and whether the subjects of the images used in the album should be prioritized as people or pets. The album creation conditions can also include the theme of the album to be created, conditions such as whether image correction should be performed on the album, a photo number adjustment amount for adjusting the number of photos to be placed in the album, and the commercial materials for creating the album. The group of album candidate images can be specified based on attribute information of individual images, such as the date and time of capture, or based on the structure of a file system containing images, such as a device and directory. Alternatively, two arbitrary images can be specified, and all images captured between the dates and times when the image data of the two images were captured can be used as the target image group.

[0024] The image acquisition unit 302 acquires a group of album candidate images designated by the album creation condition designation unit 301 from the HDD 204. The image acquisition unit 302 outputs, as meta information (additional data accompanying the image), information such as image width or height information contained in the acquired image, shooting date and time information contained in the Exif information at the time of shooting, or information indicating whether the image is included in the user image group, to the image analysis unit 304. The image acquisition unit 302 also outputs the acquired image data to the image conversion unit 303. Identification information is assigned to each image, and the meta information output to the image analysis unit 304 and the image data output to the image analysis unit 304 via the image conversion unit 303 (described later) can be associated by the image analysis unit 304.

[0025] The images stored in the HDD 204 include still images and frame images extracted from videos. The still images and frame images are acquired from an imaging device such as a digital camera or a smart device. The imaging device may be included in the image processing device 200 or an external device. If the imaging device is an external device, the images are acquired via the data communication unit 208. The still images and extracted images may be acquired from a network or a server via the data communication unit 208. Examples of images acquired from a network or a server include SNS images. The program executed by the CPU 201 analyzes data attached to each image to determine the storage source. The SNS images may be acquired from an SNS via an application, and the acquisition destination may be managed within the application. The images are not limited to the above-mentioned images, and other types of images may be used.

[0026] The image conversion unit 303 converts the image data input from the image acquisition unit 302 into pixel count and color information to be used by the image analysis unit 304, and outputs the converted information to the image analysis unit 304. In this embodiment, the image is converted to a predetermined pixel count (for example, 420 pixels on the short side) and the long side is converted to a size that maintains the ratio of the original sides. Furthermore, the image is converted to a unified color space such as sRGB for color analysis. In this way, the image conversion unit 303 converts the image into an analysis image with a unified pixel count and color space. The image conversion unit 303 outputs the converted image to the image analysis unit 304. The image conversion unit 303 also outputs the image to the layout information output unit 315 and the image correction unit 317.

[0027] The image analysis unit 304 analyzes the image data of the analysis image input from the image conversion unit 303 using a method described below to acquire image features. Image features refer to, for example, meta-information stored in the image or features that can be acquired by analyzing the image. The analysis process executes processes such as focus estimation, face detection, personal recognition, and object determination to acquire these image features. Other image features include color, brightness, resolution, data volume, and degree of blur / shake, but other image features may also be acquired. The image analysis unit 304 extracts and combines necessary information from the meta-information input from the image acquisition unit 302 along with these image features, and outputs the combined information as features to the image scoring unit 307. The image analysis unit 304 also outputs shooting date and time information to the spread allocation unit 312.

[0028] The caption acquisition unit 305 acquires captions attached to the acquired images and outputs them to the caption analysis unit 306. The caption generation unit 319 generates captions for images without captions by applying a known caption generation model, and outputs them to the caption analysis unit 306.

[0029] The caption analysis unit 306 analyzes the caption input from the caption acquisition unit 305 using a method described later, acquires caption analysis information, and outputs it to the image scoring unit 307 .

[0030] The image scoring unit 307 assigns a score to each image in the group of album candidate images using the feature amount acquired from the image analysis unit 304 and the caption analysis information acquired from the caption analysis unit 306. The score here is an index showing the suitability of each image for the layout, with a higher score indicating a more suitable layout. The scoring results are output to the image selection unit 311 and the image layout unit 314.

[0031] The photo number adjustment amount input unit 308 inputs the adjustment amount for adjusting the number of photos to be arranged in the album, which is specified by the album creation condition specification unit 301, to the photo number determination unit 310. The spread number input unit 309 inputs the number of spreads in the album, which is specified by the album creation condition specification unit 301, to the photo number determination unit 310 and the spread allocation unit 312. The number of spreads in the album corresponds to the number of templates in which multiple images are arranged.

[0032] The photo number determination unit 310 determines the total number of photos that will make up the album based on the adjustment amount specified from the photo number adjustment amount input unit 308 and the number of spreads specified from the spread number input unit 309, and inputs the determined number to the image selection unit 311.

[0033] The image selection unit 311 selects images based on the number of photos input from the photo number determination unit 310 and the score calculated by the image score unit 307, creates a list of layout images to be used in the album, and provides it to the spread allocation unit 312.

[0034] The spread allocation unit 312 allocates each image to a spread using the shooting date information for the group of images selected by the image selection unit 311. Here, an example of allocation in units of spreads will be described, but allocation in units of pages may also be performed.

[0035] The template input section 313 reads from the HDD 204 a plurality of templates corresponding to the template information designated by the album creation condition designation section 301 , and inputs them to the image layout section 314 .

[0036] The image layout unit 314 performs layout processing for images on each spread. Specifically, for the spread to be processed, the image layout unit 314 determines a template suitable for the images selected by the image selection unit 311 from the multiple templates input by the template input unit 313, and determines the layout of each image.

[0037] The layout information output unit 315 outputs layout information to be displayed on the display 205 in accordance with the layout determined by the image layout unit 314. The layout information is, for example, bitmap data in which the data of the selected image selected by the image selection unit 311 is laid out in the determined template.

[0038] The image correction condition input section 316 provides the image correction section 317 with ON / OFF information for the image correction specified by the album creation condition specification section 301. Examples of types of correction include brightness correction, dodging correction, red-eye correction, and contrast correction. ON or OFF of image correction may be specified for each type of correction, or may be specified for all types collectively.

[0039] The image correction unit 317 corrects the layout information held by the layout information output unit 315 based on the image correction conditions received from the image correction condition input unit 316. The number of pixels of the images processed by the image conversion unit 303 to the image correction unit 317 can be changed according to the size of the layout image determined by the image layout unit 314. In this embodiment, image correction is performed on each image after the layout image is generated, but the present invention is not limited to this, and each image may be corrected before being laid out on a spread or a page.

[0040] When the album creation application is installed in the image processing apparatus 200, a startup icon is displayed on the top screen (desktop) of the OS (operating system) that operates on the image processing apparatus 200. When the user double-clicks the startup icon displayed on the display 205 with the pointing device 207, the application program stored in the HDD 204 is loaded into the RAM 203 and is started by being executed by the CPU 201.

[0041] Note that some or all of the functions of the components of the software block may be realized by using a dedicated circuit. Also, some or all of the functions of the components of the software block may be realized by using a cloud computer.

[0042] <Example of UI screen> FIG. 4 is a diagram showing an example of an application startup screen 401 provided by the album creation application. The application startup screen 401 is displayed on the display 205. The user sets the album creation conditions described later via the application startup screen 401. The album creation condition specifying unit 301 acquires the setting contents from the user through this UI screen.

[0043] The path box 402 on the application startup screen 401 displays the storage location (path) in the HDD 204 of a plurality of images (for example, a plurality of image files) that are the targets of album creation. When the folder selection button 403 is instructed by a click operation with the pointing device 207 from the user, a selection screen of the folders standardly installed in the OS is displayed. On the folder selection screen, the folders set in the HDD 204 are displayed in a tree structure, and the user can select the folder containing the images to be the targets of album creation with the pointing device 207. The path of the folder in which the album candidate image group selected by the user is stored is displayed in the path box 402.

[0044] The theme selection drop-down list 404 accepts a theme setting from the user. A theme is an indicator for providing a kind of uniformity to the images to be laid out, such as "travel," "ceremony," or "everyday life." The template designation area 405 is an area where the user can designate template information, and the template information is displayed as an icon. In the template designation area 405, icons of multiple template information are displayed side by side, and the user can select template information by clicking with the pointing device 207.

[0045] The number of spreads box 406 accepts the user's setting of the number of spreads in the album. The user inputs a number directly into the number of spreads box 406 via the keyboard 206, or inputs a number from a list into the number of spreads box 406 using the pointing device 207.

[0046] A check box 407 accepts ON / OFF designation of image correction from the user. When the check box is checked, image correction ON is designated, and when the check box is not checked, image correction OFF is designated. In this embodiment, all image corrections are turned ON / OFF with one button, but this is not limited to this, and a check box may be provided for each type of image correction.

[0047] The priority mode selection button 408 accepts a priority mode setting from the user, indicating whether to prioritize the selection of portrait images or pet images in the album to be created. In this embodiment, the priority mode is selected from two modes, portrait and pet, but is not limited to these. Other modes, such as landscape, vehicle, or food, may also be used. Based on the priority mode set here, the image scoring unit 307 determines a priority subject to be used as a criterion for corrections, etc., when scoring images.

[0048] The number of photos adjustment 409 is for adjusting the number of images to be arranged on each spread of the album using a slider bar. The user can adjust the number of images to be arranged on each spread of the album by moving the slider bar left or right. The number of photos adjustment 409 allows the number of images that can be arranged on a spread to be adjusted by assigning an appropriate number, for example, -5 for few and +5 for many. Note that there may also be a form in which the user can input the number of photos without using a slider bar.

[0049] The product material specification unit 410 sets the product material of the album to be created. The product material can be set to the size of the album and the type of paper for the album. In addition, the cover type and binding type can be set individually.

[0050] When the user presses the OK button 411, the album creation condition designation unit 301 outputs the content set on the application startup screen 401 to the automatic layout processing unit 318 of the album creation application.

[0051] At this time, the path entered in path box 402 is transmitted to image acquisition unit 302. In addition, the number of spreads entered in number of spreads box 406 is transmitted to number of spreads input unit 309. Template information selected in template designation area 405 is transmitted to template input unit 313. ON / OFF of image correction in the image correction check box is transmitted to image correction condition input unit 316. Reset button 412 on application launch screen 401 is a button for resetting each setting information on application launch screen 401.

[0052] <Processing flow> Fig. 5 is a flowchart showing the processing of the automatic layout processing unit 318 of the album creation application. The flowchart shown in Fig. 5 is realized, for example, by the CPU 201 reading a program stored in the HDD 204 into the RAM 203 and executing it. In the explanation of Fig. 5, it will be explained that the components shown in Fig. 3, which function when the CPU 201 executes the album creation application, execute the processing. The automatic layout processing will be explained with reference to Fig. 5. Note that the symbol "S" in the explanation of each process indicates a step in the flowchart (this also applies to the present embodiment and subsequent embodiments).

[0053] In S501, the image scoring unit 307 determines a priority subject based on the priority mode information specified by the album creation condition specifying unit 301. For example, if a people priority mode that prioritizes the selection of portrait images is specified, subjects related to people, such as "people," "men," "women," and "children," are determined as priority subjects. On the other hand, if a pet priority mode that prioritizes the selection of pet images is specified, subjects related to pets, such as "pets," "dogs," "cats," and "hamsters," are determined as priority subjects. In this way, in S501, at least one priority subject linked to the specified priority mode is determined.

[0054] In this embodiment, the priority subject is determined based on the priority mode specified using the priority mode selection button 408, but this is not limiting, and the user may, for example, specify an arbitrary priority subject via a priority subject box (not shown). Also, the priority subject may be determined based on a theme specified using the theme selection dropdown list 404.

[0055] In S502, the image conversion unit 303 converts the image to generate an analysis image. The image used for analysis here is an image from the group of album candidate images stored in the folder in the HDD 204 specified by the album creation condition specification unit 301. Therefore, at the time of S502, it is assumed that various settings have been completed via the UI screen of the application launch screen 401, and the album creation conditions and the group of album candidate images have already been set. The image conversion unit 303 reads the group of album candidate images from the HDD 204 to the RAM 203. Then, the image conversion unit 303 converts the image in the read image file into an analysis image having a predetermined number of pixels and color information, as described above. In this embodiment, the image is converted into an analysis image having a short side of 420 pixels and color information converted to sRGB.

[0056] In S503, the image analysis unit 304 executes an analysis process on the analysis image generated in S502 to acquire image feature quantities. In this embodiment, the analysis process includes acquisition of the degree of focus, face detection, personal recognition, and object determination, but is not limited to these, and other analysis processes may also be executed. Details of the process executed by the image analysis unit 304 in S503 will be described below.

[0057] The image analysis unit 304 extracts necessary meta information from the meta information received from the image acquisition unit 302. For example, the image analysis unit 304 acquires the shooting date and time as time information of the image in the image file from the Exif information attached to the image file read from the HDD 204. Note that the acquired meta information may include, for example, image position information or F-number. Furthermore, information other than that attached to the image file may also be acquired as meta information. For example, schedule information linked to the shooting date and time of the image may be acquired.

[0058] As mentioned above, the image analysis unit 304 acquires image features from the analysis image generated in S502. An example of an image feature is the degree of focus. Edge detection is used to determine the degree of focus. A Sobel filter is a commonly known edge detection method. The Sobel filter detects edges, and the gradient of the edge is calculated by dividing the difference in brightness between the start and end points of the edge by the distance between the start and end points. Based on the average gradient of the edges in an image, an image with a larger average gradient can be considered to be more in focus than an image with a smaller average gradient. By setting multiple thresholds with different values ​​for the gradient, it is possible to determine which threshold the gradient is equal to or greater than, and output an evaluation value of the focus level. In this embodiment, two different thresholds are set in advance, and the focus level is determined using three levels: "○", "△", and "×". For example, threshold values ​​are set in advance so that "◯" is determined as the focus tilt desired to be used in the album, "Δ" is determined as an acceptable focus tilt, and "×" is determined as an unacceptable tilt. The threshold settings may be provided by, for example, the creator of the album creation application, or may be set on a user interface. Note that, for example, image brightness, color tone, saturation, or resolution may be acquired as image feature amounts.

[0059] Furthermore, the image analysis unit 304 performs face detection on the analysis image generated in S502. Here, a known method can be used for the face detection process. For example, Adaboost, which creates a strong classifier from multiple weak classifiers, is used for the face detection process. In this embodiment, a face image of a person (object) is detected by the strong classifier created by Adaboost. The image analysis unit 304 extracts a face image and acquires the upper left coordinate value and the lower right coordinate value of the position of the detected face image. By having these two types of coordinates, the image analysis unit 304 can acquire the position and size of the face image.

[0060] The image analysis unit 304 performs personal recognition by comparing a facial image in the processing target image, which is detected by face detection based on the analysis image, with a representative facial image stored for each personal ID in the face dictionary database. The image analysis unit 304 obtains the similarity between each of the representative facial images and the facial image in the processing target image. The image analysis unit 304 also identifies the representative facial image whose similarity is equal to or greater than a threshold and has the highest similarity. The personal ID corresponding to the identified representative facial image is then set as the ID of the facial image in the processing target image. If the similarity between all of the representative facial images and the facial image in the processing target image is less than the threshold, the image analysis unit 304 registers the facial image in the processing target image as a new representative facial image in the face dictionary database, associated with a new personal ID.

[0061] The image analysis unit 304 also performs object recognition on the analysis image generated in S502. Here, a known method can be used for the object recognition process. In this embodiment, objects are recognized by a classifier created by deep learning. The classifier outputs a likelihood of 0 to 1 for each object, and recognizes an object that exceeds a certain threshold as being present in the image. By recognizing the object image, the image analysis unit 304 can acquire the type of object, such as a pet such as a dog or cat, a flower, food, a building, a figurine, or a landmark. While this embodiment distinguishes objects, this is not limited thereto, and each type may be acquired by recognizing facial expressions, photographic composition, or scenes such as travel or wedding scenes. Alternatively, the likelihood output from the classifier itself before the discrimination is performed may be used.

[0062] FIG. 6 is a diagram showing image feature amounts. The image analysis unit 304 stores the image feature amounts acquired in S502 in a storage area such as the ROM 202, distinguishing them by ID for identifying each image (analysis image) as shown in FIG. 6. For example, as shown in FIG. 6, the shooting date and time information acquired in S502, the focus determination result, the number of detected faces and their position information and similarity, and the type of recognized object are stored in a table format. Note that the position information of the face image is stored distinguishing them by the personal ID acquired in S502. Furthermore, when multiple types of objects are recognized from one image, in the table shown in FIG. 6, all of the multiple types of objects are stored in the row corresponding to that one image.

[0063] In S504, the caption acquisition unit 305 determines whether or not a caption is attached to the image. If it is determined that a caption is attached, the process proceeds to S505, and if it is determined that a caption is not attached, the process proceeds to S506.

[0064] In S505, the caption acquisition unit 305 acquires the caption attached to the image. If there is a caption added by the user, or if there is a history of adding a caption when creating a past album, etc., the caption data is acquired.

[0065] In S506, the caption generation unit 319 automatically generates a caption for the image using a known caption generation model. Although the method for automatically generating a caption is not particularly limited, in this embodiment, the caption is automatically generated using the Show and Tell model described in Non-Patent Document 1.

[0066] Figure 7 is a diagram explaining a caption generation model using the Show and Tell model as an example. The caption generation model is broadly composed of three networks. The three networks are a Convolutional Neural Network (CNN), a Word Embedding (We) word embedding, and a Long Short Term Memory (LSTM). The CNN converts images into feature vectors. The Word Embedding (WE) word embeddings convert words into feature vectors. The LSTM outputs the probability of the next word appearing. When generating a caption, first an image is input into the CNN. The feature vector obtained from this input is then input into the LSTM, which calculates the word appearance probability sequentially from the beginning of the sentence, and outputs the word string with the highest product of the word appearance probabilities as the caption sentence.

[0067] In S507, the caption analysis unit 306 analyzes the caption acquired in S505 and the caption generated in S506 to acquire caption analysis information. In this embodiment, syntactic analysis is performed as the analysis process. Syntactic analysis is a process of breaking down language into morphemes and further clarifying the syntactic relationships between them. Known methods such as operator precedence, top-down syntactic analysis, or bottom-up syntactic analysis may be used to realize syntactic analysis. The caption analysis unit 306 performs syntactic analysis on the caption acquired in S505 and the caption generated in S506 to acquire elements of the caption, such as the subject, verb, object, or complement.

[0068] In this embodiment, the caption analysis process ends here, but further analysis may be performed on the element words obtained by the syntactic analysis. For example, the element words may be converted into distributed representations using a known technique. Distributed representation is a representation method in which characters or words are embedded in a vector space and are represented as a single point in that space. An example of a known technique is Word2Vec, described in Non-Patent Document 2.

[0069] Fig. 8 is a diagram showing caption analysis information. The caption analysis unit 306 stores the caption analysis information acquired in S507 in a storage area such as the ROM 202, sorting it by ID that identifies each image as shown in Fig. 8. For example, as shown in Fig. 8, each element of the subject, verb, object, or complement acquired in S507 is stored in a table format. Note that Fig. 8 also shows an example in which words are converted into embedded representations.

[0070] In S508, the image scoring unit 307 assigns a score to each image in the group of album candidate images. The score referred to here is an index indicating the appropriateness of each image for the layout. Scoring means assigning a score to each image (scoring). The assigned scores are provided to the image selection unit 311 and are referenced when selecting images to be used in the layout, which will be described later.

[0071] 9 is a flowchart showing the details of the scoring process in S508. The scoring process performed in S508 will be described below with reference to FIG.

[0072] First, in S901, the image scoring section 307 calculates the average value and standard deviation of the album candidate image group for each image feature acquired in S502. In S902, the image scoring section 307 determines whether the processing of S901 has been completed for all image feature items. If it is determined that the processing has not been completed, the processing from S901 is repeated. If it is determined that the processing has been completed, the processing proceeds to S903.

[0073] In S903, the image scoring unit 307 calculates the score for each image to be scored (referred to as "image of interest") using the following formula (1). Note that the images to be scored are images in the album candidate image group. Sji=50-|10×(μi−fji) / σi|...Equation (1)

[0074] Here, j is the index of the image of interest, i is the index of the image feature, fji is the image feature of the image of interest, and Sji is the score corresponding to the image feature fji. Also, μi and σi respectively represent the average value and standard deviation for each image feature of the album candidate images. Then, the image scoring unit 307 calculates the score for each image of interest using the score Sji for each image feature calculated by equation (1) and the following equation (2). Pj=Σi(Sji) / Ni Equation (2)

[0075] Here, Pj indicates the score of each target image, and Ni indicates the number of image feature items. In other words, the score of each target image is calculated as the average of the scores of each image feature. Note that, since it is preferable that images used in an album are in focus, a predetermined score may be added to target images with a focus feature value of "0" shown in FIG. 7.

[0076] In S904, the image scoring unit 307 corrects the score calculated in S903 based on the caption analysis information acquired in S507. One correction method is to increase the score calculated in S903 when the subject information acquired in S507 matches the priority subject information set in S501. In a caption added by a user, the subject is important information that determines the important subject in the image, and the subject subject is considered to be the main subject. Therefore, this method allows scoring to be performed in a way that makes it easier to select an image of a desirable priority subject, such as one in which the subject desired for layout priority is the important subject in the image. In this embodiment, the score of an image in which the subject information and the priority subject information match is increased by, for example, 20 points, but other increments may also be used.

[0077] Another correction method may be to decrease the score calculated in S903 when the subject information and the priority object information do not match. This method makes it possible to control the selection of images of priority objects that are not desirable for layout because the object that is desired to be preferentially laid out is not an important object in the image.

[0078] In the above correction method, the score is corrected based on whether the subject information and the priority object information match. However, a perfect match is not required when using a priority mode for the object, such as that selected by the priority mode selection button 408 in FIG. 4, and the score may be corrected based on whether they are synonymous. For example, when the "pet priority mode" is selected by the priority mode selection button 408 in FIG. 4, the score may be corrected based on whether the subject matches information synonymous with "pet," such as "dog" or "cat." This method enables more flexible score correction. A specific example of a method is to use the well-known WordNet to determine whether the priority object and the subject are synonymous, and increase the score if they are synonymous. Alternatively, synonym relationships between words may be stored in ROM 202 in advance, and a search may be performed to determine whether the priority object and the subject are synonymous.

[0079] Furthermore, if it is determined in S504 that a caption is not attached, the score calculated in S903 may be reduced rather than proceeding to the caption generation step (S506). One could also think of images to which the user has not added a caption as being less important to the user than images to which captions have been added. Therefore, according to this method, by reducing the score of images to which no captions have been added, images that are important to the user and to which captions have been added are relatively more likely to be selected.

[0080] Conversely, the score of an image determined to have a caption attached may be increased in S504. By using this method, the score of an image with a caption attached increases, making it more likely that an image that is important to the user who added the caption will be selected.

[0081] Furthermore, in S505, when a word is represented by a distributed representation, the score may be corrected based on the relationship between the priority object and the acquired subject on the distributed representation spatial vector. For example, the score may be increased when the distance between the priority object and the acquired subject on the spatial vector is equal to or less than a certain threshold. This method makes it possible to correct the score based on whether the priority object and the acquired subject are semantically similar, rather than whether they completely match. In this case, it is desirable to convert the word of the priority object into a distributed representation in advance before S904.

[0082] In S905, the image scoring unit 307 determines whether the processes of S903 and S904 have been completed for all images in the album candidate image group in the user-specified folder. If it is determined that the processes have not been completed, the processes from S903 onward are repeated. If it is determined that the processes have been completed, the scoring process in FIG. 9 ends.

[0083] Returning to the explanation of Fig. 5, following S508, in S509 the image scoring unit 307 determines whether or not image scoring in S508 has been completed for all images in the album candidate image group in the user-specified folder. If it is determined that the process has not been completed, the process is repeated from S502. If it is determined that the process has been completed, the process proceeds to S510.

[0084] In S510, the photo number determination unit 310 determines the number of photos to be arranged in the album. In this embodiment, the number of photos to be arranged in the album is determined by equation (3) using the adjustment amount for adjusting the number of photos in a spread input from the photo number adjustment amount input unit 308 and the number of spreads input from the number of spreads input unit 309. Number of photos = [Number of spreads × (Number of basic photos + adjustment amount)] Formula (3)

[0085] Here, [·] indicates a floor function that discards the decimal portion, and the basic number of photos indicates the number of images to be placed on a spread when no adjustment is made. In this embodiment, the basic number of photos is set to six, taking into consideration the appearance of the layout, and is pre-programmed into the album creation application program.

[0086] In S511, the image selection unit 311 selects images to be laid out based on the scores for each image calculated by the image scoring unit 307 and the number of photos determined by the photo number determination unit 310. Hereinafter, the selected group of images will be referred to as a layout image group. In this embodiment, the image selection unit 311 selects images from the group of images specified by the album creation condition specification unit 301 in descending order of the scores assigned by the image scoring unit 307, the total number of images to be laid out. Note that, as a method of image selection, a higher selection probability may be set for images with higher scores, and images may be selected based on probability. In this way, by selecting based on probability, layout images can be changed each time the automatic layout function by the automatic layout processing unit 318 is executed. For example, if the user is not satisfied with the automatic layout result, the user may be able to obtain a different layout result from the previous one by pressing a reselect button (not shown) in the UI.

[0087] Alternatively, image selection unit 311 may select as layout images images whose scores calculated by image score unit 307 are equal to or greater than a certain threshold. In this case, the number of photos does not need to be determined by photo number determination unit 310. In this case, a value that allows the number of selected images to be the number of double-page spreads becomes the upper limit that can be set as the threshold.

[0088] In S512, the spread allocation unit 312 divides and allocates the layout image group acquired in S511 into image groups equal to the number of spreads input from the spread number input unit 309. In this embodiment, the layout images are arranged in the order of shooting time acquired in S503, and are divided at locations where the time difference between shooting times of adjacent images is large. This process is repeated until the layout images are divided into the number of spreads input from the spread number input unit 309. In other words, division is performed (number of spreads - 1) times. This makes it possible to create an album in which images are arranged in order of shooting time. Note that the process of S512 may be performed on a page-by-page basis rather than on a spread-by-spread basis.

[0089] In S513, the image layout unit 314 determines the image layout. An example will be described below in which the template input unit 313 inputs (a) to (p) of Fig. 10 for a certain spread in accordance with the specified template information.

[0090] 10 is a diagram showing a group of templates used for laying out image data. Each of the multiple templates included in the group of templates corresponds to a double-page spread. Template 1001 is a single template. Template 1001 includes a main slot 1002, a sub-slot 1003, and a sub-slot 1004. Main slot 1002 is the main slot (a frame for laying out images) within template 1001, and is larger in size than sub-slot 1003 and sub-slot 1004.

[0091] Here, the number of slots for the input template is specified as 3, for example. Figure 10(q) shows three images selected according to the specified number of templates, arranged in order of the date and time of their capture. The three images are also arranged with their orientation (portrait or landscape) distinguished.

[0092] Here, in each group of images allocated to a spread, the image with the highest score calculated by the image score unit 307 is set as the main slot, and the other images are set as sub-slots. Note that whether an image is for the main slot or sub-slot may be set based on a certain image feature acquired by the image analysis unit 304, or may be set randomly.

[0093] Here, it is assumed that image data 1005 is for the main slot, and image data 1006 and 1007 are for the sub-slots. In this embodiment, image data with an older capture date and time is laid out in the upper left of the template (main slot 1002 in template 1001), and images with a newer capture date and time are laid out in the lower right (sub-slot 1004 in template 1001). In FIG. 10(q), image data 1005 for the main slot is portrait-oriented and has the most recent capture date and time, so it is laid out so that the lower right of the template becomes the main slot. Therefore, the templates of FIGS. 10(i) to 10(l) are candidates. Furthermore, the older image data 1006 for the sub-slot is a portrait image, and the newer image data 1007 is a landscape image. As a result, the template of FIG. 10(j) is determined as the template most suitable for the selected three image data, and the layout is determined. In S513, it is determined which image is to be laid out in which slot of which template.

[0094] In S514, the image correction unit 317 executes image correction. When information indicating that image correction is ON is input from the image correction condition input unit 316, the image correction unit 317 executes image correction. As image correction, for example, dodging correction (brightness correction), red-eye correction, or contrast correction is executed. When information indicating that image correction is OFF is input from the image correction condition input unit 316, the image correction unit 317 does not execute image correction. Image correction can also be executed on image data whose short side is 1200 pixels and whose size has been converted to the sRGB color space, for example.

[0095] In S515, layout information output unit 315 creates layout information. Image layout unit 314 lays out the image data that has been subjected to image correction in S514 in each slot of the template determined in S513. At this time, image layout unit 314 resizes and lays out the image data to be laid out in accordance with the size information of the slot. Then, layout information output unit 315 generates bitmap data in which the image data is laid out in the template as an output image.

[0096] In S516, the image layout unit 314 determines whether the processes from S513 to S515 have been completed for all spreads. If it is determined that the processes have not been completed, the process from S513 is repeated. If it is determined that the processes have been completed, the automatic layout process in FIG. 5 ends.

[0097] <Effects of the first embodiment> As described above, according to this embodiment, it is possible to suitably select an image. The difference in the effect of image selection between the comparative example and this embodiment will be described below with reference to the drawings.

[0098] FIG. 11 is a diagram illustrating the effects of this embodiment. Images 1101 and 1102 are the same as those in FIGS. 1(a) and 1(b), respectively. When the priority subject is a "train," image 1101 is an image that should be selected, and image 1102 is an image that should not be selected. Conventionally, setting the priority subject to "train" made both image 1101 and image 1102, which contain trains, more likely to be selected. In this embodiment, in addition to setting the priority subject, captions associated with the images are acquired and subjected to syntactic analysis to identify subjects that could be the main subject, thereby increasing the score for images whose subjects match the priority subject. According to this method, the processes from S505 to S508 determine that the subject of image 1101 is a train 1103, and the subject of image 1102 is a person 1104. Therefore, image 1101, whose subject matches the priority subject and can be the main subject, receives an increased score and is more likely to be selected, while image 1102, whose score is not corrected, is relatively less likely to be selected. In other words, the user can select an image of a priority subject that is more desirable.

[0099] <<Second embodiment>> In the second embodiment, the image analysis process of S503 described in the first embodiment is not performed, and the caption analysis result is used to score the image.

[0100] The software block diagram of the album creation application in this embodiment is basically the same as that in FIG. 3 of the first embodiment, but since image analysis processing is not performed, the image analysis unit 304 may be omitted.

[0101] <Processing flow> Fig. 12 is a flowchart showing the processing of the automatic layout processing unit 318 of the album creation application in the second embodiment. The automatic layout processing in the second embodiment will be described with reference to Fig. 12. Note that the basic processing of the automatic layout processing is similar to the example described in the first embodiment, and the following description will focus on the differences.

[0102] In S1201, the caption analysis unit 306 analyzes the caption acquired in S505 or the caption generated in S506 to acquire caption analysis information. In this embodiment, the caption analysis unit 306 also performs syntactic analysis on the caption acquired in S505 or the caption generated in S506 to acquire elements such as the subject, verb, object, or complement of the caption. Then, in this embodiment, the words that constitute each element are represented as a distributed representation using a known technique. In this embodiment, the distributed representation of words is realized using Word2Vec.

[0103] In S1202, the image scoring unit 307 determines whether or not the caption analysis in S1201 has been completed for all images in the album candidate image group in the user-specified folder. If it is determined that the caption analysis has not been completed, the processing from S502 is repeated. If it is determined that the caption analysis has been completed, the processing proceeds to S1203.

[0104] In S1203, the image scoring unit 307 assigns a score to each image in the album candidate image group. In the first embodiment, the scoring is performed using image feature amounts obtained by analyzing the image and caption analysis information obtained by analyzing the caption. In this embodiment, the scoring is performed using only the caption analysis information.

[0105] Fig. 13 is a flowchart showing the details of the scoring process in S1203. The scoring process performed in S1203 will be explained below with reference to Fig. 13. First, in S1301, the image scoring unit 307 selects one element from each element (subject, verb, object, or complement) of the syntactic analysis result in the caption analysis information acquired in S1201.

[0106] In S1302, the image scoring unit 307 performs clustering on the elements selected in S1301 and divides the images into clusters. In this embodiment, the Ward method is used as the clustering method. Of course, the clustering method is not limited to this, and for example, the longest distance method or the k-means method may be used. In S1303, the image scoring unit 307 determines whether the processing of S1302 has been completed for each element of the parsing result. If it is determined that it has not been completed, the processing from S1301 is repeated. If it is determined that it has been completed, the processing proceeds to S1304.

[0107] In S1304, the image scoring unit 307 calculates the score of the image of interest for each element of the parsing result using equation (4). Skj=50×(Nji / Nk) Equation (4)

[0108] Here, k is the index of the image of interest, j is the index of the element in the parsing result, i is the index of the cluster related to element j, and Skj is the score corresponding to element j in image of interest k. Furthermore, Nk is the number of images included in the group of candidate images for the album, and Nji is the number of images of interest included in cluster i in element j. According to formula (4), images that contain words that appear frequently in the group of captions linked to the group of candidate images for the album will have a higher score and will be more likely to be selected. In other words, it is possible to select images that have a consistent feel across each element.

[0109] Then, the image scoring unit 307 calculates the score of each image of interest using the score Skj for each element of each image of interest calculated by equation (4) and equation (5). Pk=Σj(Sjk) / Nj (5)

[0110] Here, Pk represents the score of each target image, and Nj represents the number of elements. In other words, the score of each target image is calculated as the average of the scores of each element.

[0111] Below, we will explain how to calculate scores using Equations (4) and (5) using two target images as examples. Target image 1 is an image for which the result of syntactic analysis is "a train running through the mountains." In target image 1, the subject is "train," the verb is "running," and the complement indicating the scene is "mountain." Assume that the number of images (Nk) included in the album candidate image group is 100. Syntactic analysis of the 100 images reveals that 25 images have the subject "train," 10 images have the verb "running," and 5 images have the scene "mountain." In this case, when the score for each element of target image 1 is calculated using Equation (4), the results are: subject = 12.5 points, verb = 5 points, object = 0 point, and scene = 2.5 points. Applying Equation (5) to this result to calculate the average score for each element yields Pk = 5 points, which becomes the score for target image 1.

[0112] Similarly, assume that image of interest 2 is an image whose syntactic analysis results in "a train running along the seashore." In image of interest 2, the subject is "train," the verb is "running," and the complement indicating the scene is "sea." Assume that 100 candidate images for the album contain 10 images whose scene is "sea." In this case, calculating the scores for each element of image of interest 2 using formula (4) yields the following: subject = 12.5 points, verb = 5 points, object = 0 point, and scene = 5 points. Applying formula (5) to this result to calculate the average score for each element yields Pk = 5.6 points, which becomes the score for image of interest 2. Therefore, comparing image of interest 1 and image of interest 2, image of interest 2 has a higher score, making it more likely to be selected at this point. In practice, the score is adjusted based on the relationship between the priority subject and the subject, as explained below, and the scoring process is completed.

[0113] In S1305, the image scoring unit 307 corrects the score calculated in S1304 based on the caption analysis information acquired in S1202. In this embodiment, as in the first embodiment, the score is corrected based on the relationship between the priority object and the subject. The specific correction method is the same as in S904 in the first embodiment, and therefore will not be described here.

[0114] In S1306, the image scoring unit 307 determines whether the processes of S1304 and S1305 have been completed for all images in the album candidate image group in the user-specified folder. If it is determined that the processes have not been completed, the processes from S1304 onward are repeated. If it is determined that the processes have been completed, the scoring process in FIG. 13 ends.

[0115] Returning to the description of Fig. 12, following S1203, in S1204 the image scoring unit 307 determines whether or not the image scoring of S1203 has been completed for all images in the album candidate image group in the user-specified folder. If it is determined that the scoring has not been completed, the processing of S1203 is repeated. If it is determined that the scoring has been completed, the processing proceeds to S510. The processing thereafter is the same as in the first embodiment, and therefore description thereof will be omitted. With the processing of S516, the automatic layout processing of Fig. 12 ends.

[0116] <Effects of the second embodiment> As described above, according to this embodiment, automatic layout processing is possible using only caption analysis information, without performing the image analysis processing of S503 in the first embodiment. Therefore, the processing load due to image analysis processing can be eliminated, and processing speed can be increased.

[0117] <Modification of the second embodiment> In the above embodiment, in S1303, the image scoring unit 307 calculates the score of the image of interest for each element of the parsing result using formula (4), thereby enabling consistent image selection. However, the image scoring unit 307 may calculate the score of the image of interest for each element of the parsing result using the following formula (6) instead of formula (4). Skj=50×(1-Nji / Nk)...Equation (6)

[0118] According to formula (6), images with words that appear less frequently in the captions linked to the album candidate images, i.e., words that appear more sporadically, will have a higher score and will be more likely to be selected. In other words, it is possible to select images with a wide variety of elements.

[0119] <<Third Embodiment>> In the third embodiment, in the caption generation process of S506 described in the first embodiment, the caption generation is not completed, but information during the generation is extracted and used for image scoring.

[0120] <Processing flow> Fig. 14 is a flowchart showing the processing of the automatic layout processing unit 318 of the album creation application in the third embodiment. The automatic layout processing in the third embodiment will be described with reference to Fig. 14. Note that the basic processing of the automatic layout processing is similar to the example described in the first embodiment, and the following description will focus on the differences.

[0121] In S1401, the caption generation unit 319 automatically generates and analyzes a caption for an image. In this embodiment, too, the caption is automatically generated using the Show and Tell model described in Non-Patent Document 1.

[0122] In the Show and Tell model, a word string with a high probability of occurrence can be obtained in the process of completing caption generation up to the end of the sentence. In this embodiment, based on the above characteristics of the caption generation model, the subject is estimated before caption generation is completed using information obtained in the caption generation process.

[0123] Fig. 15 is a diagram showing the caption generation and analysis process. The caption generation and analysis process performed in S1401 will be described below with reference to Fig. 15. In S1501, the caption generation unit 319 estimates the i-th word using a Show and Tell model. In the Show and Tell model, the top multiple words are candidates based on the word appearance probability known from the output of the LSTM, and multiple word string candidates are estimated together with the word candidates estimated up to the i-1th word.

[0124] In S1502, the caption generation unit 319 determines, from among the multiple word sequence candidates estimated up to the (i-1)th word sequence, the word sequence with the highest product of the occurrence probabilities of the words included in the word sequence as the representative word sequence. In other words, from the multiple word sequence candidates, the word sequence estimated to be most suitable as a caption is determined. Up to this point, no word estimation has been performed for the i-th word and the process proceeds to S1503.

[0125] In S1503, the caption generation unit 319 acquires the part of speech of the ith word in the representative word string. A part of speech is a group into which words are classified according to grammatical criteria, such as nouns or verbs. In this embodiment, the correspondence between words and parts of speech is stored in advance in the ROM 202, and the part of speech is acquired based on the estimated word. As another method for acquiring the part of speech, a morphological analysis may be performed on the representative word string to acquire the estimated part of speech for the ith word. Morphological analysis is the process of dividing a sentence in a natural language written in characters into the smallest meaningful linguistic units (morphemes).

[0126] In S1504, the caption generation unit 319 determines whether the part of speech of the i-th word acquired in S1503 is a noun. A word determined to be a noun may be the subject of the word string. If it is determined to be a noun, the process proceeds to S1505. If it is determined not to be a noun, the process from S1501 is repeated. In S1505, the caption generation unit 319 outputs the i-th word of the representative string determined to be a noun in S1504. When the process of S1505 ends, the caption generation and analysis process of FIG. 15 ends. According to the above-mentioned method, it is possible to output the noun (subject) of the representative word string by estimating the words up to the i-th word.

[0127] In this embodiment, the part-of-speech acquisition process of S1503 and the noun determination process of S1504 are performed after the representative word string determination process of S1502, but the representative word string determination process of S1502 may also be performed after the noun determination process of S1504. In other words, the part-of-speech acquisition process of S1503 and the noun determination process of S1504 may be performed for each of a plurality of word string candidates, and the representative word string determination process of S1502 may be performed for one or more word string candidates determined to be nouns. According to this method, the noun determination process of S1504 can be performed for more word strings, and therefore nouns may be output in fewer steps.

[0128] Here, Fig. 16 is a diagram showing a caption generation and analysis process using a method different from that of Fig. 15. As will be explained below, the caption generation and analysis shown in Fig. 15 can also be performed using the processing flow shown in Fig. 16. Note that some of the processing is the same as the example explained in Fig. 15, and the following explanation will focus on the differences.

[0129] In S1601, the caption generation unit 319 determines whether the i-th word in the representative word string acquired in S1502 matches the priority subject acquired in S501. If it is determined that they match, even if the representative word string is a word string in which estimation of words has only been completed up to the i-th word, it can be estimated that the i-th word when parsed is the subject. If it is determined that they match, the process proceeds to S1602. If it is determined that they do not match, the process repeats from S1501.

[0130] In S1602, the caption generation unit 319 performs syntactic analysis on the representative word string. In S1603, the caption generation unit 319 outputs the word determined to be the subject as a result of the syntactic analysis in S1602. When the processing of S1603 ends, the caption generation and analysis processing of Fig. 16 ends. As with Fig. 15, the method of Fig. 16 also makes it possible to output the subject of the representative word string even if word estimation has only been performed up to the i-th word, thereby shortening the processing time.

[0131] Returning to the explanation of Fig. 14, the subject information estimated in S1401 is added to the caption analysis information acquired in S507 and is used for image scoring in S508. Then, with the processing of S516, the automatic layout processing of Fig. 14 ends. Note that, although the Show and Tell model is used as the caption generation model in this embodiment, this is not limiting, and other caption generation models may be used as long as information during caption generation, such as the state in which the subject has been estimated, can be acquired.

[0132] Furthermore, the subject estimation performed in this embodiment during caption generation can also be used when generating captions in embodiment 1 or 2, allowing the process to proceed to caption analysis more quickly.

[0133] <Effects of the third embodiment> As described above, according to this embodiment, information during caption generation can be extracted and used for image scoring without waiting for the completion of the caption generation process, thereby reducing the processing load associated with the caption generation process.

[0134] <<Other embodiments>> The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

Claims

1. 1. A program for selecting an image from a group of candidate images, comprising: Computer, an acquisition means for acquiring the candidate image group including a plurality of images; a determining means for determining specific conditions for preferentially selecting an image from the group of candidate images; image analysis means for analyzing the images of the candidate image group; a caption generating means for generating a caption when a caption is not added to an image of the candidate image group; a caption analysis means for analyzing captions attached to images in the candidate image group or captions generated by the caption generation means; a selection means for selecting a specific image from the group of candidate images based on the results of the determination means, the image analysis means, and the caption analysis means; It functions as The caption analysis means analyzes the caption while the caption generation means is generating the caption.

2. 2. The program according to claim 1, wherein the specific conditions include a setting of a priority subject for preferentially selecting the specific image.

3. 3. The program according to claim 2, wherein the caption analyzing means breaks down the caption into words and determines the subject of an image in the candidate image group.

4. 4. The program according to claim 3, wherein the selection means preferentially selects an image in the candidate image group when the subject of the image determined by the caption analysis means matches the priority subject.

5. 2. The program according to claim 1, wherein the caption generating means generates captions for images using a Show and Tell model.

6. 6. The program according to claim 5, wherein the caption analysis means further analyzes the caption generated by the caption generation means.

7. 2. The program according to claim 1, wherein the image analysis means estimates the degree of focus of the image, detects a face, recognizes a person, or determines an object.

8. 8. The program according to claim 1, wherein the specific conditions include a degree of focus of the image, a number of faces, or an object.

9. 1. A program for selecting an image from a group of candidate images, comprising: Computer, an acquisition means for acquiring a candidate image group including a plurality of images; a determining means for determining specific conditions for preferentially selecting an image from the group of candidate images; a caption generating means for generating a caption when a caption is not added to an image of the candidate image group; a caption analysis means for analyzing captions attached to images in the candidate image group or captions generated by the caption generation means; a selection means for selecting a specific image from the group of candidate images based on the results of the determination means and the caption analysis means; It functions as The caption analysis means analyzes the caption while the caption generation means is generating the caption.

10. 10. The program according to claim 9, wherein the specific conditions include setting a priority subject for preferentially selecting the specific image.

11. 11. The program according to claim 10, wherein the caption analyzing means breaks down the caption into words and determines the subject of an image in the candidate image group.

12. The program according to claim 11, wherein the selection means preferentially selects an image in the candidate image group when the subject of the image determined by the caption analysis means matches the priority subject.

13. 10. The program according to claim 9, wherein the caption generating means generates captions for images using a Show and Tell model.

14. 14. The program according to claim 13, wherein the caption analysis means further analyzes the caption generated by the caption generation means.

15. an acquisition means for acquiring a candidate image group including a plurality of images; a determining means for determining specific conditions for preferentially selecting an image from the group of candidate images; image analysis means for analyzing the images of the candidate image group; a caption generating means for generating a caption when a caption is not added to an image of the candidate image group; a caption analysis means for analyzing captions attached to images in the candidate image group or captions generated by the caption generation means; a selection means for selecting a specific image from the group of candidate images based on the results of the determination means, the image analysis means, and the caption analysis means; Equipped with The image processing device is characterized in that the caption analysis means analyzes the caption while the caption generation means is generating the caption.

16. an acquisition step of acquiring a candidate image group including a plurality of images; a determining step of determining specific conditions for preferentially selecting images from the group of candidate images; an image analysis step of analyzing the images of the candidate image group; a caption generating step of generating a caption when a caption is not assigned to an image of the candidate image group; a caption analysis step of analyzing captions attached to images in the candidate image group or captions generated by the caption generation step; a selection step of selecting a particular image from the group of candidate images based on results of the determination step, the image analysis step, and the caption analysis step; Equipped with The method for controlling an image processing device, wherein the caption analysis step analyzes the caption during the generation of the caption by the caption generation step.

17. an acquisition means for acquiring a candidate image group including a plurality of images; a determining means for determining specific conditions for preferentially selecting an image from the group of candidate images; a caption generating means for generating a caption when a caption is not added to an image of the candidate image group; a caption analysis means for analyzing captions attached to images in the candidate image group or captions generated by the caption generation means; a selection means for selecting a specific image from the group of candidate images based on the results of the determination means and the caption analysis means; Equipped with The image processing device is characterized in that the caption analysis means analyzes the caption while the caption generation means is generating the caption.

18. an acquisition step of acquiring a candidate image group including a plurality of images; a determining step of determining specific conditions for preferentially selecting images from the group of candidate images; a caption generating step of generating a caption when a caption is not assigned to an image of the candidate image group; a caption analysis step of analyzing captions attached to images in the candidate image group or captions generated by the caption generation step; a selection step of selecting a particular image from the group of candidate images based on the results of the determination step and the caption analysis step; Equipped with The control method for an image processing device, characterized in that the caption analysis step functions to analyze the caption during the generation of the caption by the caption generation step.

Citation Information

Patent Citations

  • Synthesis image creation assist device, method and program, and record medium therefor

    JP2015069431A

  • Image processing method, image adjustment method, image adjustment program, and image adjustment device

    JP2017028412A

  • Image processing device, system, method and program

    JP2018097492A

  • Information processing device, information processing method and program

    JP2021071870A