Image content determination device, image content determination method, and storage medium

The image content determination device performs facial and character recognition on the image and automatically assigns labels using similar facial image information, which solves the problem of tedious manual label input by users and improves information reliability and efficiency.

CN115315695BActive Publication Date: 2026-03-27FUJIFILM CORP
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-21
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing technologies, users need to manually input label information for a large number of images, which leads to cumbersome operations, and the reliability of information estimation based on image analysis is low.

Method used

The image content determination device performs face and character recognition on the image, obtains relevant information by using image information of similar faces, and automatically assigns labels to the image.

Benefits of technology

It enables automatic tag assignment without requiring users to manually input tags, thus improving the reliability and efficiency of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115315695B_ABST
    Figure CN115315695B_ABST
Patent Text Reader

Abstract

The image content determination device includes at least one processor that performs a first recognition process of recognizing a character and a face of a first person from a first image including the character and the face of the first person, a first acquisition process of acquiring first person-related information about the first person included in the first image based on the recognized character and the face of the first person, a second recognition process of recognizing a face of a second person from a second image including the face of the second person, and a second acquisition process of acquiring second person-related information about the second person included in the second image, wherein the second person-related information is acquired using the first person-related information corresponding to the first image including the face of the first person similar to the face of the second person in the second acquisition process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image content determination device, an image content determination method, and a storage medium. Background Technology

[0002] In recent years, online storage services have emerged that can transmit and store user-held image data, such as photos, via the internet. Users can download and view images stored in the memory using mobile devices and / or PCs (Personal Computers).

[0003] In such online storage services, in order to easily search for the images that users want to view from a large number of images stored in memory, the images stored in memory are given tag information that allows for keyword search (Japanese Patent Application Publication No. 2009-526302 and Japanese Patent Application Publication No. 2010-067014).

[0004] Japanese Patent Application Publication No. 2009-526302 and Japanese Patent Application Publication No. 2010-067014 disclose, for example, the following technology: when two images each contain the face of a person and one image is assigned a label such as the name of the person by user input, the label information assigned to one image is copied to the other image based on the similarity of the faces contained in the two images. Summary of the Invention

[0005] However, the technologies described in Japanese Patent Application Publication No. 2009-526302 and Japanese Patent Application Publication No. 2010-067014 require users to pre-input label information for the source images of the image being copied, which causes inconvenience for users. For example, if there are many images, it is tedious for users to view each image one by one, confirm the image content, and assign corresponding label information.

[0006] Therefore, as a method to assign labeling information to images without causing trouble for users, one could consider performing image analysis to determine the image content and assigning labeling information based on the determination results. For example, methods for determining image content using image analysis could include estimating the ages of people in the image, or, when the image contains multiple people, estimating the relationships (family relationships, etc.) between the people based on their estimated ages.

[0007] However, the accuracy of image content determination based on image analysis is also limited. Therefore, when estimating information related to people contained in an image, performing image analysis using only the data of the image that is the object of image content determination can result in low reliability of the information obtained through estimation.

[0008] In view of the above problems, one embodiment of the present invention provides an image content determination device, an image content determination method, and an image content determination program that can obtain highly reliable information as information related to people contained in an image without causing trouble to the user.

[0009] means for solving technical problems

[0010] The image content determination apparatus of the present invention includes at least one processor, which performs the following processes: performing a first recognition process to identify characters and the face of a first person from a first image containing characters and the face of a first person; performing a first acquisition process to obtain first person-related information related to the first person contained in the first image based on the identified characters and the face of the first person; performing a second recognition process to identify the face of a second person from a second image containing the face of a second person; and performing a second acquisition process to obtain second person-related information related to the second person contained in the second image. In the second acquisition process, the second person-related information is obtained using first person-related information corresponding to the first image containing the face of a first person similar to the face of the second person.

[0011] The image content determination apparatus of the present invention operates in a method comprising at least one processor, wherein the processor performs the following processes: performing a first recognition process for recognizing characters and the face of a first person from a first image containing characters and the face of a first person; performing a first acquisition process for acquiring first person-related information related to the first person contained in the first image based on the recognized characters and the face of the first person; performing a second recognition process for recognizing the face of a second person from a second image containing the face of a second person; performing a second acquisition process for acquiring second person-related information related to the second person contained in the second image; wherein, in the second acquisition process, the second person-related information is acquired using first person-related information corresponding to the first image containing the face of a first person similar to the face of the second person.

[0012] The operating procedure of the image content determination device of the present invention is an operating procedure for enabling a computer including at least one processor to function as an image content determination device. The processor performs a first recognition process to identify characters and the face of a first person from a first image containing characters and the face of a first person; performs a first acquisition process to obtain first person-related information related to the first person contained in the first image based on the identified characters and the face of the first person; performs a second recognition process to identify the face of a second person from a second image containing the face of a second person; and performs a second acquisition process to obtain second person-related information related to the second person contained in the second image. In the second acquisition process, the second person-related information is obtained using first person-related information corresponding to the first image containing the face of a first person similar to the face of the second person. Attached Figure Description

[0013] Figure 1 This is an explanatory diagram showing an overview of online storage services.

[0014] Figure 2 This is a block diagram of an image content determination device.

[0015] Figure 3 This is a functional block diagram of the CPU in the image content determination device.

[0016] Figure 4 This is an explanatory diagram of the classification process performed by the classification department.

[0017] Figure 5 This is an explanatory diagram of the first identification process performed by the first identification unit and the first acquisition process performed by the first acquisition unit.

[0018] Figure 6 This is an explanatory diagram illustrating an example of the first acquisition process.

[0019] Figure 7 This is a table representing an example of the list of information for the first image.

[0020] Figure 8 This is an explanatory diagram of the second identification process performed by the second identification unit.

[0021] Figure 9 This is a table representing an example of the second image information list.

[0022] Figure 10 This is an explanatory diagram of the second acquisition process performed by the second acquisition unit.

[0023] Figure 11 This is an illustration of the labeling process performed by the labeling department.

[0024] Figure 12This is an example of a table representing a second image information list with added information and tags for a second person.

[0025] Figure 13 This is a flowchart of the image content determination and processing.

[0026] Figure 14 This is a schematic diagram showing the outline of the first embodiment.

[0027] Figure 15 This is a schematic diagram showing the outline of the second embodiment.

[0028] Figure 16 This is an explanatory diagram illustrating an example of the second acquisition process in the second embodiment.

[0029] Figure 17 This is an explanatory diagram illustrating an example of the first acquisition process in the third embodiment.

[0030] Figure 18 This is an explanatory diagram illustrating an example of the second acquisition process in the third embodiment.

[0031] Figure 19 This is an explanatory diagram illustrating an example of the first acquisition process in the fourth embodiment.

[0032] Figure 20 This is an explanatory diagram illustrating an example of the second acquisition process in the fourth embodiment.

[0033] Figure 21 This is an explanatory diagram illustrating an example of the first acquisition process in the fifth embodiment.

[0034] Figure 22 This is an explanatory diagram illustrating an example of the second acquisition process in the fifth embodiment.

[0035] Figure 23 This is an explanatory diagram based on the classification process of whether or not a specific word is present.

[0036] Figure 24 This is an explanatory diagram illustrating an example of installing a program stored in a storage medium onto an image content determination device. Detailed Implementation

[0037] [First Implementation]

[0038] exist Figure 1In this invention, the image content determination device 2, as an example of the technology, constitutes part of an image transmission system. The image transmission system is a system that stores images P of multiple users, such as user A and user B, in a memory 4 and transmits the stored images P via a communication network N according to requests from each user. Images P are digital data such as photographs held by each user. From the user's perspective, the service provided by the image transmission system is a service that stores images in the memory 4 via the communication network N; therefore, it is also referred to as online storage service, etc. When using the image transmission system, each user signs a usage contract with the operator operating the image transmission system. For users who have signed a usage contract, for example, an account is created for each user and a storage area for storing each user's images P is allocated in the memory 4. When signing the usage contract, the operator receives personal information such as the user's name and date of birth and registers the obtained personal information as the user's account information.

[0039] The memory 4 is a data storage device such as a hard disk drive or a solid-state drive. The memory 4 is communicatively connected to the image content determination device 2 and also functions as an external memory for the image content determination device 2. Furthermore, the memory 4 can be connected to the image content determination device 2 via a network, such as a WAN (Wide Area Network) like the Internet, or a LAN (Local Area Network) like Wi-Fi. The connection between the network and the image content determination device 2 can be wired or wireless. Additionally, the memory 4 can be a recording medium directly connected to the image content determination device 2 using a USB (Universal Serial Bus) or can be built into the image content determination device 2. Moreover, the memory 4 is not limited to a single device; it can be composed of multiple devices per data unit and / or per capacity unit.

[0040] Users, including User A and User B, can, for example, launch an online storage service application installed on Smart Device 6 to upload image data of photos taken with Smart Device 6 to Storage 4 via Communication Network N. Furthermore, each user can access the online storage service via PC. Each user can also upload image data of photos taken with Digital Camera 8 to Storage 4 via PC. Additionally, each user can use Scanner 10 to read printed photos PA and upload the digitized image data to Storage 4 via PC or Smart Device 6. Alternatively, printed photos PA can be digitized using the photography function of Smart Device 6 or Digital Camera 8, instead of digitizing them using Scanner 10.

[0041] The printed photo PA also includes greeting cards created by individual users. These greeting cards include New Year's cards, Christmas cards, summer greeting cards, and winter greeting cards. Additionally, the digitization of the printed photo PA can be outsourced to an online storage service provider, replacing the user's manual digitization and upload to storage device 4.

[0042] Image data uploaded by each user is stored as image P in memory 4. Then, the image P uploaded to memory 4 is tagged using image content determination device 2. Memory 4 is provided with, for example, an unprocessed folder 12 for storing images P that have not been tagged and a processed folder 14 for storing images P that have been tagged.

[0043] In the unprocessed folder 12, a dedicated folder is set up for each user, such as user A's dedicated folder 12A and user B's dedicated folder 12B. The images P held by each user are stored in each user's dedicated folder. Image data uploaded by user A is stored in user A's dedicated folder 12A, which is set up in the unprocessed folder 12. Image data uploaded by user B is stored in user B's dedicated folder 12B, which is set up in the unprocessed folder 12.

[0044] Image content determination device 2 is a device that uses image analysis techniques such as face recognition, character recognition, and scene discrimination to determine the content of image P. In this example, image content determination device 2 also assigns the determination result of the content of image P as tag information for keyword searching of image P to add tags to image P.

[0045] Furthermore, the tag information assigned to image P can be information other than the determination result of image P's content, such as accompanying information like image P's Exif (Exchangeable Image File Format) information. In addition to the camera equipment manufacturer and model name, the Exif information also includes the date and time of the photograph, GPS (Global Positioning System) information indicating the location of the photograph. The Exif information is recorded as metadata within the image P file and can be used as a tag for searching.

[0046] The image content determination device 2 has the function of determining the content of the image P in addition to the Exif information and assigning information related to the people contained in the image P as tag information.

[0047] For example, when image P is a first image P1 such as a New Year's card, the New Year's card often includes a family photo containing the faces of multiple people who make up a family. Furthermore, the New Year's card contains characters such as the names of the multiple people who make up the family, and the date. If it can be determined that the first image P1 is a New Year's card, it can be estimated that the relationship between the multiple people in the photo within the first image P1 is a family, and the names contained in the first image P1 are the names of the members of that family. Thus, in greeting cards such as New Year's cards, image P contains not only the faces of the people, but also character information related to the people, such as their names.

[0048] Furthermore, in image P, besides images like the first image P1 that contain both a person's face and characters, there are also images like the second image P2 that contain a person's face but not characters. Regarding this second image P2 that does not contain characters, the image content determination device 2 also analyzes the image content, thereby estimating information related to the person contained in the second image P2.

[0049] The image content determination device 2 has the function of analyzing the content of a second image P2, which contains a person's face and characters but does not contain characters, using information about the person obtained from an image P, such as the first image P1, that contains a person's face. This function will be explained in detail below.

[0050] Image content determination device 2 determines the image content for each image group of each user, for example. For example, when determining the image content of user A's image group, image content determination device 2 performs image content determination processing on user A's image P stored in user A's dedicated folder 12A in the unprocessed folder 12.

[0051] Image P includes a first image P1 and a second image P2. The first image P1 is an image containing characters and a person's face. The person included in the first image P1 corresponds to the first person involved in the technology of this invention. As an example of the first image P1, it has an image with a character area. An image with a character area refers to an image that includes a photographic area AP containing the face of the first person and a character area AC containing characters arranged in a blank area outside the outline of the photographic area AP. The blank area can be a plain color without patterns, or it can have patterns, etc. Greeting cards such as New Year's cards often have images with character areas.

[0052] In this example, the first image P1 is a New Year's card and is an image with a character area. Therefore, the first image P1 is an image that includes a photo area AP showing the faces of multiple first persons who make up a family and a character area AC in the blank area of ​​the photo area AP containing New Year's greetings such as "Happy New Year", the names of family members, addresses, etc.

[0053] The second image P2 is an image containing a person's face. The person contained in the second image P2 corresponds to the second person involved in the technology of this invention. As an example of the second image P2, it is an image without a character area. An image without a character area refers to an image that only contains the photo area AP containing the face of the second person. The second image P2 is an image that does not include a character area AC outside the photo area AP, except for the characters in the background, etc., of the second person reflected in the photo area AP containing the face of the second person.

[0054] The image content determination device 2 obtains first person-related information R1 from the first image P1, which relates to the first person. Furthermore, when determining the image content of the second image P2, it determines that the first image P1 contains a first person similar to the second person contained in the second image P2. Moreover, the image content determination device 2 obtains second person-related information R2 related to the second person in the second image P2 based on the determined first person-related information R1 of the first image P1.

[0055] Furthermore, the image content determination device 2 adds tags to the second image P2 based on the acquired information R2 related to the second person. The tagged second image P2 is stored in the processed folder 14. As an example, each user also has a dedicated folder for the processed folder 14; user A's second image P2 is stored in user A's dedicated folder 14A, and user B's second image P2 is stored in user B's dedicated folder 14B.

[0056] In addition, Figure 1 In the already processed folder 14, only the second image P2 is stored. However, when the first image P1 is also newly tagged with the result of obtaining the first person's related information R1 from the first image P1, the first image P1 is also stored in the already processed folder 14.

[0057] Thus, the first image P1 and the second image P2 of each user, which have been tagged, are stored in a folder that can be sent to each user for viewing. At this time, each user can use the tag information to perform keyword searches, etc.

[0058] like Figure 2 As an example, the computer constituting the image content determination device 2 includes a CPU (Central Processing Unit) 18, internal memory 20, program internal memory 22, communication I / F 24, and external device I / F 26. They are interconnected via a bus 28.

[0059] The aforementioned memory 4 is communicatively connected to the image content determination device 2 via external device I / F 26. The computer constituting the image content determination device 2 and the memory 4 are configured, for example, at the base of an operator providing online storage services, along with other devices constituting the image transmission system. Furthermore, the communication I / F 24 serves as an interface for controlling the transmission of various information with external devices.

[0060] The internal program memory 22 stores a classification program 30, an identification program 31, a first acquisition program 32, a second acquisition program 34, and a tagging program 35. Among these programs, the identification program 31, the first acquisition program 32, and the second acquisition program 34 are programs for enabling the computer constituting the image content determination device 2 to operate as the "image content determination device" according to the technology of this invention. These programs are an example of the "image content determination program" according to the technology of this invention.

[0061] Internal memory 20 serves as internal memory for processing by CPU 18 and for storing data such as dictionary data (described later) required for processing by CPU 18, as well as the first image information list 48 and the second image information list 50 (described later). CPU 18 loads the classification program 30, recognition program 31, first acquisition program 32, second acquisition program 34, and tagging program 35 stored in program internal memory 22 into internal memory 20.

[0062] like Figure 3 As an example, the CPU 18 functions as a classification unit 36, an identification unit 38, a first acquisition unit 40, a second acquisition unit 42, and a tagging unit 44 by executing a classification program 30, an identification program 31, a first acquisition program 32, a second acquisition program 34, and a tagging program 35 on the internal memory 20. The CPU 18 is an example of a "processor" according to the technology of the present invention.

[0063] Regarding the processing of the image content determination device 2, this example will use the determination of the content of image P by user A as an example for explanation. Figure 3 In the process of processing image P of user A, the classification unit 36 ​​reads image P from user A's dedicated folder 12A. The classification unit 36 ​​classifies the read image P into image 1 P1 and image 2 P2.

[0064] The recognition unit 38 includes a first recognition unit 38-1 and a second recognition unit 38-2. The first recognition unit 38-1 performs a first recognition process to recognize characters and the face of a first person from a first image P1 containing characters and the face of a first person. Specifically, the first recognition unit 38-1 recognizes the face of the first person contained in the first image P1 from the photo area AP of the first image P1, and recognizes characters from the character area AC. The second recognition unit 38-2 performs a second recognition process to recognize the face of a second person contained in the second image P2 from the photo area AP of the second image P2.

[0065] The first acquisition unit 40 performs a first acquisition process to acquire the first person-related information R1 contained in the first image P1 based on the characters and the face of the first person recognized by the first recognition unit 38-1.

[0066] The second acquisition unit 42 performs a second acquisition process to acquire second person-related information R2 related to the second person contained in the second image P2. In this second acquisition process, the second person-related information R2 is acquired using the first person-related information R1 corresponding to the first image P1 containing the face of the first person similar to the second person. The tagging unit 44 assigns tag information to the second image P2 based on the second person-related information R2.

[0067] refer to Figure 4 An example of the classification process performed by the classification unit 36 ​​will be described. The classification unit 36 ​​determines whether an image P includes a photographic region AP and a character region AC. For example, the classification unit 36 ​​performs contour extraction on the image P using methods such as edge detection and detects the photographic region AP and the character region AC from the extracted contours. Furthermore, the photographic region AP and the character region AC have features that can distinguish them from other regions, such as features related to the pixel values ​​of each pixel and the arrangement of pixel values. The classification unit 36 ​​detects the photographic region AP and the character region AC from the image P by investigating such features contained in the image P. When the characters contained in the image P are printed or written with the same pen, it is considered that the pixel values ​​of the pixels corresponding to the characters contained in the image P are similar within a certain range. Therefore, for example, the pixels constituting the image P can be analyzed using two-dimensional coordinates. When a pixel column with pixel values ​​representing a predetermined range of similarity is arranged with a predetermined width or more on the first axis (X-axis) and the pixel column is continuously arranged with a predetermined width or more on the second axis (Y-axis), it is determined to be a character, and the region containing the character is determined to be the character region AC.

[0068] The character area AC contains not only kanji, hiragana, katakana, and letters, but also numbers and symbols. Characters are not limited to those defined by a specific font; handwritten characters are also included. Character recognition within the character area AC can be performed using OCR (Optical Character Recognition / Reader) and other character recognition technologies. Alternatively, machine learning-based character recognition techniques can also be employed.

[0069] Furthermore, the classification unit 36 ​​uses face recognition techniques such as contour extraction and pattern matching to identify the faces of people from the photo region AP. Of course, face recognition techniques using machine learning can also be employed. As an example, the classification unit 36 ​​detects the face image PF representing the identified face in the photo region AP and classifies the image P based on the presence or absence of the face image PF.

[0070] exist Figure 3 The text explains that the classification is into two types: image 1 P1 and image 2 P2. More specifically, as... Figure 4 As shown, the classification unit 36 ​​classifies image P into three types—first image P1, second image P2, and third image P3—based on the presence or absence of a face image PF and a character region AC. Specifically, firstly, image P that includes both the photo region AP and the character region AC, and where the photo region AP includes the face image PF, is classified as first image P1. Then, image P that includes the photo region AP but not the character region AC, and where the photo region AP includes the face image PF, is classified as second image P2. Furthermore, image P that includes the photo region AP but not the character region AC, and where the photo region AP does not include the face image PF, is classified as third image P3. Additionally, in Figure 4 In the example, the third image P3 is illustrated by excluding the character region AC, but the necessary condition for the third image P3 is that the photo region AP does not include the face image PF, but it may or may not include the character region AC.

[0071] The memory 4 contains categorized folders 13 for storing the first image P1, the second image P2, and the third image P3 after categorization. Within these categorized folders 13, each user has a separate folder for storing the first image P1 (folder 13-1), the second image P2 (folder 13-2), and the third image P3 (folder 13-3). Figure 4 In the example, the three image folders, namely the first image folder 13-1, the second image folder 13-2, and the third image folder 13-3, are user A's private folders.

[0072] Next, refer to Figures 5-7An example of the first recognition process and the first acquisition process performed on the first image will be explained.

[0073] like Figure 5 As shown, the first recognition unit 38-1 reads the first image P1 one by one from the first image folder 13-1 of the classified folder 13 and performs the first recognition process. In the following example, when distinguishing multiple first images P1, like first image P1-1 and first image P1-2, the symbol P1 is labeled with subdivision symbols "-1", "-2", and "-3". Figure 5 The diagram shows an example of performing the first recognition process on the first images P1-4. The first recognition process includes first face recognition processing, character recognition processing, and photographic scene discrimination processing.

[0074] In the first face recognition process, the first recognition unit 38-1 recognizes the face of the first person M1 contained in the photo region AP of the first image P1-4. As the face recognition technology, the same technology used in the classification unit 36 ​​is employed. For example, the first recognition unit 38-1 extracts a rectangular region containing the recognized face within the photo region AP as a first face image PF1. When multiple faces of the first person M1 are contained within the photo region AP, as in the first image P1-4, recognition of all faces of the first person M1 is performed, and a first face image PF1 is extracted for all recognized faces. Figure 5 In the example, the photo region AP contains three first persons M1, so three first face images PF1 are extracted. Furthermore, when it is necessary to distinguish multiple first persons M1, the symbol M1 is labeled with subdivision symbols A, B, and C, just like first persons M1A, M1B, and M1C.

[0075] Additionally, there are situations where, within the photo area AP, a face of a person who is difficult to identify as the main subject is reflected in the background of the first person M1, which is the main subject. As a countermeasure, for example, when a relatively small face is included within the photo area AP, the small face can be determined as not being the main subject and excluded from the extraction objects. Furthermore, for example, if the size of the area containing the first face image PF1 within the photo area AP is less than a predetermined area, it can be excluded.

[0076] In character recognition processing, the first recognition unit 38-1 recognizes the string CH from the character region AC included in the first image P1-4. The string CH consists of multiple characters and is an example of a character. In character recognition processing, the string CH recognized within the character region AC is converted into text data using character recognition technology.

[0077] In the photographic scene discrimination process, the first recognition unit 38-1 determines the photographic scene of the photograph shown in the photograph region AP of the first image P1-4. Examples of photographic scenes include portraits and landscapes. Landscapes include mountains, seas, cities, night scenes, indoor scenes, outdoor scenes, festivals, ceremonies, and watching sporting events. Photographic scenes are determined, for example, through image analysis using pattern matching and machine learning. Figure 5 In the example, the photographic scene of images P1-4 in the first image is classified as "portrait" and "outdoor". Thus, there can be multiple classification results for the photographic scene.

[0078] As an example, the first acquisition unit 40 performs a first acquisition process based on a first face image PF1 representing the face of the first person M1, the string CH, and the photographic scene. The first acquisition process includes primary processing and secondary processing.

[0079] The first processing step involves using the dictionary data 46 to determine the meaning of the string CH and obtaining the determined meaning as primary information. This primary information serves as the basis for various determinations in the second processing step. The second processing step involves obtaining the relevant information R1 of the first person based on the obtained primary information and the first face image PF1, etc.

[0080] The results of the first recognition process and the first acquisition process are recorded in the first image information list 48. The first image information list 48 is a file that records the first image information, including the first face image PF1 acquired in the first acquisition process for each first image P1, the information obtained according to the string CH, and the relevant information R1 of the photographic scene and the first person. In addition to the information acquired in the first acquisition process, the first image information also includes supplementary information (such as Exif information) when the first image P1 has supplementary information. Furthermore, the first image information also includes the string CH recognized by the first recognition process. The supplementary information and the string CH are also recorded in the first image information list 48. The image information of multiple first images P1 is listed by recording the individual image information of multiple first images P1 in the first image information list 48.

[0081] refer to Figure 6 Specific examples of the first and second processing steps in the first acquisition process are explained. For example... Figure 6As shown, the first acquisition unit 40 refers to the dictionary data 46 in one processing step to determine the meaning of the string CH. The dictionary data 46 stores data that establishes a correspondence between multiple patterns of strings and the meaning of those strings. For example, the dictionary data 46 registers various typical patterns of strings representing "New Year's greetings". If the string CH matches the pattern of "New Year's greetings", then the meaning of the string CH is determined to be "New Year's greetings". Furthermore, the dictionary data 46 registers various typical patterns of strings representing "name" and "address", etc. If the string CH matches the pattern of "name" and "address", then the meaning of the string CH is determined to be "name" and "address". In addition to name and address, the meaning of the string CH also includes telephone number, nationality, workplace, school name, age, date of birth, and hobbies. The dictionary data 46 also registers typical patterns of these strings, enabling the determination of various meanings of the string CH. In addition, although the dictionary data 46 is set to be recorded in the internal memory 22, it is not limited to this and can also be recorded in the memory 4.

[0082] exist Figure 6 In the examples, the string CH of "Happy New Year" was identified as "New Year's greetings". The string CH of "2020 New Year's Day" was identified as "date". The string CH of "1-1, ××-cho, 00-ku, Tokyo" was identified as "address". The string CH of "Yamada Taro, Hanako, Ichiro" was identified as "name".

[0083] Furthermore, in one processing step, for example, the category of the content of the first image P1 is estimated based on the discriminative meaning of the string CH. The category of the content of the first image P1 refers to information such as whether the first image P1 represents a New Year's card or a Christmas card. Thus, in one processing step, the discriminative meaning of the string CH, such as "New Year's greetings," "date," "name," and "address," and the category of the content of the first image P1 estimated based on the meaning of the string CH (New Year's card in this example) are obtained as primary information. Primary information is information obtained solely from the string CH, and the meaning of the string CH discriminated through one processing step is also a general meaning.

[0084] In the secondary processing, the first acquisition unit 40 uses the primary information as the base information to acquire first person-related information R1 related to the first person contained in the first image P1. In this example, the first images P1-4 are New Year's cards, and the category of the content of the first images P1-4 contained in the primary information is New Year's cards. In the case of New Year's cards, it is likely that the "name" and "address" contained in the character area AC are the "address" and "name" of the first person M1 contained in the photo area AP. Since the primary information of the first images P1-4 contains "New Year's cards", the first acquisition unit 40 estimates that the "address" and "name" contained in the primary information are the "name" and "address" of the first person M1 in the photo area AP.

[0085] That is, at the point of first processing, the meaning of the string CH for "address" and "name" is only recognized as a general meaning unrelated to a specific person. However, in the second processing, the meaning of the string CH, like the "address" and "name" of the first person M1 detected by recognizing a face from the first image P1, becomes a specific meaning determined based on its relationship with the first person M1. The "name" and "address" contained in the first image P1 are information about the "name" and "address" of the first person M1 contained in the photo area AP contained in the first image P1, obtained based on the characters recognized from the first image P1 and the face of the first person M1, and are an example of the first person-related information R1.

[0086] Furthermore, in the case of New Year's cards, when the photo area AP contains the faces of multiple first persons M1, the relationship between these multiple first persons M1 is often a family relationship such as spouses or parents and children. Therefore, since "New Year's card" is included in the information of the first images P1-4, the first acquisition unit 40 estimates that the multiple first persons M1 in the photo area AP are family members. Since the first images P1-4 contain three first persons M1, the relationship of the three first persons M1 is estimated to be a family of three. The information that the relationship of the three first persons M1 is a parent-child relationship and a family of three is obtained based on the characters identified from the first image P1 and the faces of the first persons M1, and is an example of the first person-related information R1.

[0087] Furthermore, as an example, the first acquisition unit 40 analyzes the first face images PF1 of the three individuals M1A, M1B, and M1C contained in the first images P1-4 to estimate the gender and age of the three individuals M1A, M1B, and M1C. In this example, it is estimated that the first individual M1A is a male aged 30-39, the first individual M1B is a female aged 30-39, and the first individual M1C is a child under 10 years old. Based on this estimation result and the information about the three individuals' families, the first acquisition unit 40 acquires the first-person related information R1, which states that the first individual M1A is a "husband" and "father," the first individual M1B is a "wife" and "mother," and the first individual M1C is the child of the first individuals M1A and M1B.

[0088] Thus, the first acquisition unit 40 acquires first person-related information R1 based on the characters identified from the first image P1 and the face of the first person M1. When there are multiple first images P1, the first acquisition unit 40 performs the first recognition process and the first acquisition process on each first image P1 to acquire information and first person-related information R1 once. The first person-related information R1 acquired in this way is recorded in the first image information list 48. In addition, although the first image information list 48 is set to be recorded in the internal memory 22, it is not limited to this and may also be recorded in the memory 4.

[0089] exist Figure 7 As an example, the first image information list 48 contains first image information obtained from multiple first images P1 held by user A, including the first face image PF1, the scene, the string CH, primary information, and first person-related information R1, and stores them in a corresponding relationship with each of the first images P1-1, P1-2, P1-3, ... . The first image information list 48 is stored, for example, together with each user's image P in the storage area allocated to each user within the memory 4.

[0090] exist Figure 7 In the first image information list 48 shown, Exif information is recorded as supplementary information in first images P1-2 and P1-3, but no Exif information is recorded in first images P1-1 and P1-4. This indicates that, for example, first images P1-2 and P1-3 are images taken using a smart device 6 or digital camera 8 that has the function of attaching Exif information at the time of shooting. On the other hand, it indicates that first images P1-1 and P1-4, which do not record Exif information, are images digitized by reading a printed photo PA using a scanner 10 or the like.

[0091] Furthermore, in the first image P1-1, as information R1 related to the first person, it includes information that the pet of the first person M1 is a dog. This information is obtained, for example, by estimating that the dog is the pet of the first person M1 when the dog is reflected together with the first person M1 in the first image P1-1.

[0092] and, Figure 7 The first images P1-1 to P1-4 shown are examples of New Year's cards with "Yamada Taro" as the sender. For example, it is an example where user A "Yamada Taro" saves the first image P1 of a New Year's card with himself as the sender in memory 4.

[0093] In images P1-1 to P1-4, the sending years are arranged chronologically. Image P1-1 is dated "2010," making it the earliest, while image P1-4 is dated "2020," making it the most recent. During the first acquisition process, the name "Yamada Taro" is common to all images P1-1 to P1-4. Therefore, it is possible to estimate that the name of the first person M1A, who is common to all images P1-1 to P1-4, is "Yamada Taro." Furthermore, the first image information list 48 records the first facial image PF1 of the first person M1A included in each of images P1-1 to P1-4, along with the date, thus enabling the tracking of changes in the first person M1A's face. These changes in the first person M1A's face in each decade are also included in the first person-related information R1. In other words, the information related to the first person, R1, also includes information obtained from multiple first images, P1.

[0094] The first image information, which includes the first person-related information R1 recorded in the first image information list 48, is used not only as tag information for the first image P1, but also as a prerequisite for determining the image content for tagging the second image P2.

[0095] Next, refer to Figures 8-11 This document explains the second recognition process, the second acquisition process, and the labeling process performed on the second image P2.

[0096] As in Figure 8 As an example, the second recognition unit 38-2 sequentially reads the second image P2 one by one from the second image folder 13-2 of the classified folder 13 and performs the second recognition process. In the following examples, similar to the first image P1, when distinguishing multiple second images P2, the symbol P2 is labeled with a subdivision symbol, just like the second images P2-1 and P2-2. Figure 8The diagram illustrates an example of performing the second recognition process on the second image P2-1. The second recognition process includes a second face recognition process and a scene discrimination process.

[0097] In the second face recognition process, the second recognition unit 38-2 uses the same face recognition technology as the first recognition unit 38-1 to recognize the face of the second person M2 contained in the photo region AP of the second image P2-1. For example, the second recognition unit 38-2 extracts a rectangular region containing the recognized face within the photo region AP as a second face image PF2. When multiple second persons M2 are contained within the photo region AP, as in the second image P2-1, the faces of all second persons M2 are recognized, and a second face image PF2 is extracted for all recognized faces. Figure 8 In the example, the photo region AP of the second image P2-1 contains three faces of the second person M2, therefore three second face images PF2 are extracted. Similar to the first person M1, regarding the second person M2, when it is necessary to distinguish multiple second persons M2, the symbol M2 is labeled with subdivision symbols A, B, and C, just like with second persons M2A, M2B, and M2C. When a relatively small face is included as background within the photo region AP, the process of determining the small face as not a primary subject and excluding it from the extracted objects is the same as the first recognition process.

[0098] In the scene discrimination process, the second recognition unit 38-2 determines the scene of the photograph shown in the photograph region AP of the second image P2-1. The scene discrimination method is the same as that for the first image P1. Figure 8 In the example, the photographic scene was identified as "portrait" and "indoor". The result of the second recognition process is recorded in the second image information list 50. Although the second image information list 50 is set to be recorded in internal memory 22, it is not limited to this and may also be recorded in memory 4.

[0099] like Figure 9 As an example, the second image information list 50 is a file that records second face images PF2, representing the face of the second person M2 identified from the second image P2 in the second recognition process, and second image information of the photographic scene. Second images P2-1 and P2-3 contain the faces of the second person M2 of three people, therefore three second face images PF2 are recorded as second image information. Second image P2-2 contains the faces of the second person M2 of four people, therefore four second face images PF2 are recorded as second image information. Second image P2-4 contains the faces of the second person M2 of two people, therefore two second face images PF2 are recorded as second image information.

[0100] Furthermore, in the second image information list 50, in addition to "portrait" and "outdoors," "shrine" is also recorded as the photographic scene for the second image P2-3. This is determined, for example, based on the presence of a shrine house or torii gate in the background of the photographic area AP of the second image P2-3. Furthermore, in addition to "portrait," "sea" is also recorded as the photographic scene for the second image P2-4. This is determined based on the presence of the sea and a ship in the background of the photographic area AP of the second image P2-4.

[0101] Furthermore, in addition to the information identified in the second recognition process, the second image information list 50 also includes supplementary information (such as Exif information) when the second image P2 has such information. The second image information list 50 records the individual image information of multiple second images P2. In the second image information list 50, Exif information is recorded in second images P2-1, P2-3, and P2-4 (from P2-1 to P2-4), but not in second image P2-2.

[0102] The accompanying information includes GPS information. The GPS information for image P2-1 indicates that the location was Hawaii. The GPS information for image P2-3 indicates that the location was Tokyo. Furthermore, the GPS information for image P2-4 indicates that the location was over Tokyo Bay.

[0103] like Figure 10 As an example, the second acquisition unit 42 performs a second acquisition process to acquire second person-related information R2 related to the second person M2 contained in the second image P2. The second acquisition process includes similar image search processing and formal processing. Figure 10 The example is an example of performing the second acquisition process on the second image P2-1.

[0104] In the similar image search processing, the second acquisition unit 42 reads the second face image PF2 of the second image P2-1 of the processing object from the second image information list 50. Then, the second acquisition unit 42 compares the second face image PF2 with the first face image PF1 contained in the first image P1 of the same user A. Then, it searches from multiple first images P1 for a first image P1 that contains a first face image PF1 similar to the second face image PF2 contained in the second image P2-1. Figure 10 In the example, the first face image PF1 and the second face image PF2 for verification are read from the first image information list 48 and the second image information list 50, respectively.

[0105] The second acquisition unit 42 checks each of the second face images PF2 of the second person M2 contained in the second image P2-1 against the first face image PF1. Since the second image P2-1 contains three second persons M2 and three second face images PF2, the second acquisition unit 42 checks each of the three second face images PF2 against the first face image PF1. However, there is a case where the first image P1 also contains multiple first persons M1, and the first face images PF1 also contain a number corresponding to the number of persons. In this case, each first face image PF1 is checked.

[0106] In this example, the second face image PF2 of the three people in image P2-1 is compared with the first face image PF1 of the one person in image P1-1. In this case, the combinations for comparison are 3×1, resulting in three possibilities. Next, the second face image PF2 of the three people in image P2-1 is compared with the first face image PF1 of the two people in image P1-2. In this case, the combinations for comparison are 3×2, resulting in six possibilities. Next, the second face image PF2 of the three people in image P2-1 is compared with the first face image PF1 of the three people in image P1-3. In this case, the combinations for comparison are 3×3, resulting in nine possibilities. Finally, the second face image PF2 of the three people in image P2-1 is compared with the first face image PF1 of the three people in image P1-4. Similar to images P1-3, images P1-4 also contain three faces (PF1). Therefore, in the case of images P1-4, the possible combinations for verification are 3×3 (9 in total). This verification is performed the number of times corresponding to the number of images P1. Furthermore, this embodiment describes a cyclical verification of images of people contained in images P1 with images of people contained in images P2, but it is not limited to this. For example, it could be as follows: Analyzing the second person M2A contained in the second image P2, if the image is similar to the first person M1A in images P1-4 to a predetermined level or higher, priority is given to verifying the first person M1 (e.g., first person M1B and first person M1C) contained in images P1-4 other than the first person M1A.

[0107] By comparing multiple second face images PF2 contained in the second image P2 of the processing object with multiple first face images PF1 contained in the first image P1, a first image P1 containing a face similar to that of the second person M2 is searched. Regarding the determination of facial similarity, for example, if the similarity evaluation value is above a pre-set threshold, it is determined to be similar. The similarity evaluation value is calculated using image analysis techniques such as pattern matching based on feature quantities representing facial morphological features and machine learning.

[0108] exist Figure 10 In the example, image P1, which contains the face of the first person M1 that is similar to the face of the second person M2 (one of the three people in image P2-1), yields four images: P1-1, P1-2, P1-3, and P1-4. When there are many images being searched, a pre-set number of images with high similarity scores can be extracted, while images with low similarity scores are excluded.

[0109] The second acquisition unit 42 reads the first image information from the first image information list 48, which contains the first person information R1 corresponding to each of the first images P1-1 to P1-4 that were searched.

[0110] In the formal processing, the second acquisition unit 42 uses image information containing information R1 related to the first person to acquire information R2 related to the second person. First, the second acquisition unit 42 estimates that the second person M1 in the second image P2-1 is a family of three, based on the similarity between the faces of the second persons M2A, M2B, and M2C in the second image P2-1 and the first persons M1A, M1B, and M1C in the family of three in the first image P1-4. Furthermore, the GPS information included in the accompanying information of the second image P2-1 is "Hawaii," meaning that the location of the second image P2-1 is "Hawaii." In contrast, the address of the first person M1 included in the information R1 related to the first person is "Tokyo." Based on this comparison between the location and address, the second acquisition unit 42 estimates that "the second image P2-1 is a family photo taken during a trip to Hawaii." The second acquisition unit 42 acquires the estimation results of "the second person M2 of the three people is a family" and "the second image P2-1 is a family photo taken during a trip to Hawaii" as the second person-related information R2 related to the second person M2.

[0111] In addition, R2, which is related to the second person in the second image P2-1, besides Figure 10In addition to the information exemplified herein, such as the first person-related information R1 obtained from the first images P1-4, it may also include gender and age obtained through image analysis of the facial features of the second person M2 included in the second images P2-1. Furthermore, as described later, the accuracy of the estimation results such as gender and age estimated through image analysis can be verified using the first person-related information R1.

[0112] The second person-related information R2 obtained through the second acquisition process is recorded in the second image information list 50 (see reference). Figure 12 The information related to the second person, R2, is used for labeling the second image, P2.

[0113] like Figure 11 As an example, the tagging unit 44 performs tagging processing on the second image P2-1 of the processing object based on the second person-related information R2 obtained by the second acquisition unit 42. In the tagging processing, the tagging unit 44 extracts keywords used in the tag information from the second person-related information R2. For example, when the second person-related information R2 is "the second image P2-1 is a family photo taken during a trip to Hawaii," the tagging unit 44 extracts "family," "trip," and "Hawaii" as keywords used in the tag information from the second person-related information R2. Furthermore, the keywords used in the tag information can be the words themselves contained in the second person-related information R2, or they can be different words that share a common substantive meaning. Examples of different words that share a common substantive meaning include, for example, "overseas" and "United States," which geographically include "Hawaii." If we consider Japan as a base point, all three words can be included in the broader concept of "overseas," therefore, it can be said that they share a common substantive meaning.

[0114] The tagging unit 44 assigns these keywords as tag information to the second image P2-1. The tagging unit 44 stores the second image P2-1 with the assigned tag information in the user A's dedicated folder 14A set in the processed folder 14.

[0115] like Figure 12 As an example, the second person-related information R2 acquired by the second acquisition unit 42 and the tag information assigned by the tagging unit 44 are associated with the second image P2-1 and recorded in the second image information list 50. In the second image information list 50, the second person-related information R2 and tag information are recorded for each second image P2.

[0116] Next, refer to Figure 13 The flowchart illustrates the function of the structure described above. As an example, the image content determination process for the second image P2 in the image content determination device 2 is as follows: Figure 13 Proceed in the order shown.

[0117] In this example, the image content determination device 2 performs image content determination processing on each image P of each user at a preset time interval. The preset time interval could be, for example, monitoring the number of unprocessed images P uploaded by users to the memory 4 and determining when the number of unprocessed images P reaches a preset number. For example, when the number of unprocessed images P uploaded by user A to the memory 4 reaches the preset number, the image content determination device 2 performs image content determination processing on user A's image P. Alternatively, the preset time interval could also be the time interval for newly uploaded images P from new users. The following explanation uses the case of performing image content determination processing on user A's image P as an example.

[0118] In the image content determination process, firstly, the classification unit 36... Figure 13 The classification process is performed in step ST10. During the classification process, as follows... Figure 4 As an example, the classification unit 36 ​​reads the unprocessed image P of user A from the unprocessed folder 12. Then, it classifies the image P into one of three types: a first image P1, a second image P2, or a third image P3, based on whether the photo area AP contains a face image PF and whether the image P contains a character area AC. When the image P includes the photo area AP and the character area AC, and the photo area AP contains the face image PF, the classification unit 36 ​​classifies the image P as the first image P1. When the image P includes the photo area AP containing the face image PF but does not include the character area AC, the classification unit 36 ​​classifies the image P as the second image P2. When the image P includes the photo area AP that does not contain the face image PF, or when it does not include the photo area AP, the classification unit 36 ​​classifies the image P as the third image P3.

[0119] The classification unit 36 ​​performs classification processing on all unprocessed images P of each user. The classified first image P1, second image P2, and third image P3 are stored in the classified folder 13 respectively.

[0120] Next, the first identification unit 38-1 in Figure 13 The first identification process is performed in step ST20. In the first identification process, such as... Figure 5 As an example, the first recognition unit 38-1 performs a first recognition process on the first image P1 within the classified folder 13. In this first recognition process, the first recognition unit 38-1 first performs a first face recognition process to identify the face of the first person M1 contained in the photo region AP of the first image P1. Figure 5In the case of the first image P1-4 as an example, since the face of the first person M1 of the three people is contained in the photo area AP, the face of the first person M1 of the three people is identified from the first image P1-4. The first recognition unit 38-1 extracts the face of the first person M1 of the three people identified from the first image P1-4 as three first face images PF1.

[0121] Subsequently, the first recognition unit 38-1 performs character recognition processing on the first image P1. The first recognition unit 38-1 extracts the string "CH" from the character region AC included in the first image P1. Figure 5 In the case of the first image P1-4 shown, the strings CH such as "Tokyo 00 Ward..." and "Yamada Taro" are identified.

[0122] Subsequently, the first recognition unit 38-1 performs photographic scene discrimination processing on the first image P1. In the photographic scene discrimination processing, the first recognition unit 38-1 discriminates photographic scenes such as "portrait" and "outdoor".

[0123] Next, the first acquisition unit 40 in Figure 13 In step ST30, the first acquisition process is executed. In the first acquisition process, such as... Figure 5 As an example, the first acquisition unit 40 performs a first acquisition process based on the string CH, which is an example of a recognized character, and a first face image PF1 representing the face of the first person M1. The first acquisition process includes a first processing and a second processing.

[0124] In one processing step, the first acquisition unit 40 determines the general meaning of the string CH while referring to the dictionary data 46. For example, as... Figure 6 As shown, the string CH “Tokyo 00 Ward…” is generally interpreted as an address. Furthermore, the string CH “Yamada Taro” is generally interpreted as a name. And the string CH “Happy New Year” is interpreted as “New Year’s greetings”. Additionally, in one processing step, since the string CH contains “New Year’s greetings”, the category of the content of image P1 is estimated to be “New Year’s card”. This information is obtained as primary information and is used as the basis for secondary processing.

[0125] In secondary processing, such as Figure 6 As shown, the first acquisition unit 40 uses primary information as the basis to acquire first person-related information R1 related to the first person M1 contained in the first image P1. For example... Figure 6As shown, the information in the first image P1-4 includes "New Year's card", therefore the first acquisition unit 40 estimates that the "name" and "address" included in the first information are the "name" and "address" of the first person M1 in the photo area AP. Furthermore, since the first image P1-4 is a "New Year's card", the first acquisition unit 40 estimates that the three first persons M1 in the photo area AP are a family of three.

[0126] The first acquisition unit 40 acquires the estimated information as the first person-related information R1. The first acquisition unit 40 records the primary information obtained in one processing step and the first person-related information R1 obtained in a secondary processing step in the first image information list 48. For example... Figure 7 As an example, in addition to primary information and information related to the first person R1, the first image information 48 also records secondary information including supplementary information and the first face image PF1.

[0127] The unprocessed first image P1 is subjected to the classification processing of step ST10 up to the first acquisition processing of step ST30. As a result, image information of multiple first images P1 is recorded in the first image information list 48.

[0128] Next, the second identification unit 38-2 in Figure 13 The second identification process is performed in step ST40. In the second identification process, as... Figure 8 As an example, the second recognition unit 38-2 performs a second recognition process on the second image P2 within the classified folder 13. In this second recognition process, the second recognition unit 38-2 first recognizes the face of the second person M2 contained within the photographic region AP of the second image P2. Figure 8 In the case of the second image P2-1 as an example, since the face of the second person M2 of the three people is contained in the photo area AP, the second recognition unit 38-2 recognizes the face of the second person M2 of the three people in the second image P2-1 and extracts the region containing the recognized face as three second face images PF2. Subsequently, the second recognition unit 38-2 performs scene discrimination processing on the second image P2 to determine the photographic scene. Figure 8 In the example, the photographic scene of image P2-1 in image 2 was classified as "portrait" and "indoor".

[0129] The second recognition unit 38-2 performs a second recognition process on the second image P2 of the processing object. For example... Figure 9 As an example, the second face image PF2, representing the face of the second person M2 identified from the second image P2, and the photographic scene are recorded in the second image information list 50.

[0130] Next, the second acquisition unit 42 performs a second acquisition process in step ST50. This second acquisition process includes similar image search processing and formal processing. For example... Figure 10 As an example, the second acquisition unit 42 first compares the second face image PF2 with the first face image PF1, thereby performing a similar image search process on the first image P1 that contains a face similar to the face of the first person M1 contained in the second image P2-1. Figure 10 In the example, similar image search processing is used to search for four first images P1 (P1-1 to P1-4) as first images P1 containing the face of a first person M1 that is similar to any one of the faces of the three people included in the second image P2-1. The second acquisition unit 42 reads the first image information containing the first person-related information R1 corresponding to the searched first image P1 from the first image information list 48. When like Figure 10 When there are multiple searched first images P1 as in the example, the second acquisition unit 42 reads the first image information containing the first person information R1 corresponding to each first image P1 from the first image information list 48.

[0131] In the formal processing, the second acquisition unit 42 acquires the second person's related information R2 based on the first image information containing the first person's related information R1. Figure 10 In the example, the first person information R1 contains information that the first person M1 in the three people in the first image P1-4 is a family of three. Furthermore, the faces of the second person M2 in the three people in the second image P2-1 are all similar to the faces of the first person M1 in the three people in the first image P1-4. Based on this information, the second acquisition unit 42 estimates that the second person M2 in the three people in the second image P2 is a family. Furthermore, the GPS information for the second image P2-1 indicates the shooting location as "Hawaii," while the address of the three-person family recorded in the first person information R1 of the first images P1-4 is "Tokyo." By checking this information, the second acquisition unit 42 estimates that "the second image P2-1 is a family photo taken during a trip to Hawaii." The second acquisition unit 42 records this estimation result in the second image information list 50 (see reference). Figure 12 ).

[0132] Second acquisition unit 42 in Figure 13 In step ST60, tagging processing is performed. In this tagging process, the second acquisition unit 42 adds tag information to the second image P2 based on the acquired second person-related information R2. Figure 11 In the example, according to Figure 10In the example, the relevant information R2 of the second person is obtained, and the second image P2-1 is labeled with information such as "family, travel, Hawaii, ...".

[0133] The second acquisition unit 42 performs processing on the multiple unprocessed second images P2. Figure 13 The process proceeds from the second identification step in step ST40 to the tagging process in step ST60. The result is as follows: Figure 12 As shown in the second image information list 50 as an example, multiple second images P2 are labeled with tags. The tags are used as keywords for searching the second image P2.

[0134] If the above content is presented in summary, then as follows: Figure 14 As shown. That is, in the image content determination device 2 of this example, the first recognition unit 38-1 performs a first recognition process to recognize the character shown as the string CH and the face of the first person M1 from a first image P1 containing characters and the face of the first person, such as a New Year's card like the first image P1-4. Then, the first acquisition unit 40 performs a first acquisition process to acquire first person-related information R1 related to the first person M1 contained in the first image P1 based on the recognized string CH and the face of the first person M1. When the first image P1-4 is a New Year's card, the first person-related information R1 includes the "name" and "address" of the first person M1, so the "name" and "address" of the first person M1, as well as information that multiple first persons M1 belong to one family, are acquired.

[0135] Then, the second recognition unit 38-2 performs a second recognition process to identify the face of the second person M2 from the second image P2 containing the face of the second person M2. In the case that the second image P2 is the second image P2-1, the face of the second person M2 of the three people is identified. Then, the second acquisition unit 42 performs a second acquisition process to acquire second person-related information R2 related to the second person M2 contained in the second image P2-1. The second acquisition process is a process of acquiring the second person-related information R2 using the first person-related information R1 corresponding to the first image P1 containing the face of the first person M1, which is similar to the face of the second person M2. Figure 14 In the example, in the second acquisition process, the first person-related information R1 corresponding to the first image P1-4, which contains the first person M1 of the three people who are similar to the second person M2 of the three people contained in the second image P2-1, is acquired. Then, the first person-related information R1, which is the three-person family, is used to acquire the second person-related information R2, such as "the three second people M2 are a family" and "the second image P2-1 is a family photo taken during a trip to Hawaii".

[0136] The string CH contained in the first image P1, such as the greeting card, often contains accurate personal information about the first person M1, such as address and name. This information is highly reliable as the basis for obtaining the first person-related information R1 associated with the first person M1. Therefore, the first person-related information R1 obtained using the string CH contained in the first image P1 is also highly reliable. Then, in this example, the image content determination device 2, when determining the image content of the second image P2, determines the first image P1 related to the second image P2 based on the similarity between the face of the second person M2 and the face of the first person M1, and obtains the first person-related information R1 corresponding to the determined first image P1. Then, in obtaining the second person-related information R2 for the second person M2, the first person-related information R1 of the first person M1, which is highly likely to be the same person as the second person M2, is used.

[0137] Therefore, according to the image content determination device 2 in this example, compared to the conventional method that does not utilize the first person-related information R1 corresponding to the first image P1, it is possible to obtain highly reliable second person-related information R2 as information related to the person M2 contained in the second image P2. Furthermore, in the image content determination device 2 of this example, the CPU 18 performs a series of processes to obtain the first person-related information R1 and the second person-related information R2. Therefore, it avoids the inconvenience caused to the user as in the past.

[0138] As an example, the information related to the second person, R2, is used as the tag information for the second image, P2. This tag information is generated based on the information related to the second person, R2, and therefore has high reliability as information representing the image content of the second image, P2. Thus, the likelihood of assigning appropriate tag information representing the image content to the second image, P2, is high, and the probability of finding the second image, P2, as desired by the user when performing a keyword search on P2 is also increased.

[0139] In this example, the first image P1 illustrates an image with a character region, including a photographic region AP containing the face of the first person M1 and a character region AC containing characters as a blank area outside the outline of the photographic region AP. The second image P2 illustrates an image without characters, consisting only of a photographic region AP containing the face of the second person M2.

[0140] Greeting cards and identification documents often use images with character regions. When the first image P is such an image with a character region, the characters contained in the character region AC are highly likely to represent information related to the first person M1 contained in the photo region AP. Therefore, the first person-related information R1 obtained from the characters in the character region AC is also meaningful and reliable. By utilizing such first person-related information R1, for example, compared to using information obtained from an image consisting only of the photo region AP containing the face of the first person M1 as first person-related information R1, it is easier to obtain meaningful and reliable information as second person-related information R2.

[0141] Furthermore, when the second image P2 is an image without character regions, it contains less information compared to the case of an image with character regions, thus lacking clues for determining the image content. Therefore, the amount of information related to the second person R2 that can be obtained solely from the second image P2, which is an image without character regions, is limited. Therefore, when obtaining the information related to the second person R2 from such a second image P2, utilizing the information related to the first person R1 from the first image P1, which is an image with character regions, is particularly effective.

[0142] Furthermore, the first image P1 contains an image representing at least one of a greeting card and identification. Greeting cards include seasonal greeting cards such as New Year's cards and Christmas cards, as well as summer greeting cards. In addition, greeting cards may include postcards announcing a child's birth, children's ceremonies such as Shichi-Go-San (celebrating the growth of a 7-year-old girl, a 5-year-old boy, and a 3-year-old boy and girl), school and graduation notices, and relocation notices. Identification documents include driver's licenses, passports, employee ID cards, and student ID cards. The information recorded in such greeting cards and identification documents is particularly accurate; therefore, compared to cases where the first image P1 only contains an image representing a commercially available art postcard, the first image P1 is especially effective as a source of highly reliable first-person related information R1. Furthermore, greeting cards may contain diverse information about the person, such as hobbies; therefore, compared to cases where the first image P1 is an image representing a direct mail advertisement, the first-person related information R1 has a higher probability of providing diverse information.

[0143] In the image content determination device 2 of this example, the classification unit 36 ​​performs a classification process to classify multiple images P into a first image P1 and a second image P2 before performing the first recognition process and the second recognition process. In this way, by pre-classifying multiple images P into a first image P1 and a second image P2, the process of obtaining information related to the first person R1 and information related to the second person R2 can be performed more efficiently compared to the case where no classification process is performed before each recognition process.

[0144] In this example, the first person's related information R1 is obtained from the first image P1, which has the same owner as the second image P2. "Same owner" means that both the first image P1 and the second image P2 are stored in the same user's account's storage area within memory 4. When the owners of the first image P1 and the second image P2 are the same, the commonality between the first person M1 contained in the first image P1 and the second person M2 contained in the second image P2 is high compared to the case where the owners of the first image P1 and the second image P2 are different. When obtaining the second person's related information R2 from the second image P2, meaningful first person's related information R1, which has a high correlation with the second person M2, can be utilized. Therefore, compared to the case where the owners of the first image P1 and the second image P2 are different, the reliability of the obtained second person's related information R2 is improved. Furthermore, meaningful first person's related information R1 is easily obtained; in other words, there is less noise. Therefore, when the holder of the first image P1 and the holder of the second image P2 are the same, the processing efficiency for obtaining highly reliable information about the second person R2 is also improved compared to the case where the holders are different.

[0145] Alternatively, the first person-related information R1 corresponding to the first image P1 held by a person different from the holder of the second image P2 can also be used. The reason is as follows. For example, there may be a relationship between the two holders, such as the holder of the first image P1 and the holder of the second image P2 being a family member, friends, or people who participated in the same event. In this case, when obtaining the second person-related information R2 for the second image P2, it is possible to obtain meaningful information by using the first person-related information R1 corresponding to the first image P1 held by a different person. Furthermore, the users who can utilize the first person-related information R1 based on user A's image group are limited to users who meet predefined conditions. These predefined conditions may be specified by user A, or they may be images with a predefined number or proportion of images similar to those included in user A's image group.

[0146] In this example, the first person's information R1 includes, for example, at least one of the following: the first person M1's name, address, phone number, age, date of birth, and hobbies. The first person's information R1 containing this information is effective as a clue to obtain the second person's information R2. For example, the first person M1's name is highly valuable for determining the second person M2's name, and the first person M1's phone number is highly valuable for determining the second person M2's address. Furthermore, the address may not be an accurate address; it can be merely a postal code or simply the name of a prefecture. In addition to the above, the first person's information R1 may also include any of the following: nationality or the name of an organization. Examples of organizations include workplace names, school names, and circle names. This information is also effective as a clue to obtain both the first person's information R1 and the second person's information R2.

[0147] In this example, the first person-related information R1 and the second person-related information R2 include family relationships. Family relationships are an example of information representing the relationships between multiple first persons M1 included in the first image P1 or information representing the relationships between multiple second persons M2 included in the second image P2. Thus, when multiple first persons M1 are included in the first image P1 or when multiple second persons M2 are included in the second image P2, the first person-related information R1 or the second person-related information R2 can contain information representing the relationships between multiple persons. As shown in the example above, information representing the relationships between multiple first persons M1 is effectively used to estimate the relationships between multiple second persons M2 identified as first persons M1. Furthermore, by including information representing the relationships between multiple second persons M2 in the second person-related information R2, more diverse labeling information can be assigned compared to the case where the second person-related information R2 only contains information related to each of the multiple second persons M2 individually.

[0148] Information representing relationships among multiple first persons M1 or multiple second persons M2 may include, in addition to family relationships such as spouses, parents and children, and siblings, at least one of the following: kinship relationships including grandparents, friendship relationships, and teacher-student relationships. Furthermore, the "relationship among multiple first persons M1" is not limited to family relationships and kinship relationships; it may also include interpersonal relationships such as friendship or teacher-student relationships. Therefore, according to this structure, compared to the case where the information representing relationships among multiple first persons M1 or multiple second persons M2 is only information representing family relationships, more reliable first-person related information R1 or second-person related information R2 can be obtained.

[0149] According to this example, the second acquisition unit 42 utilizes GPS information, such as Exif information, which is an example of supplementary information attached to the second image P2. Among the supplementary information of the second image P2, such as Exif information, there is more information useful for acquiring GPS information and other information related to the second person R2. By utilizing the supplementary information, the information related to the second person R2 with higher reliability can be acquired compared to not utilizing it. Furthermore, in this example, the second acquisition unit 42 utilizes the supplementary information attached to the second image P2 in the second acquisition process for acquiring the information related to the second person R2, but of course, the first acquisition unit 40 could also utilize the supplementary information attached to the first image P1 in the first acquisition process for acquiring the information related to the first person R1.

[0150] In the above example, the example of obtaining first person-related information R1 from the first images P1-4 of a New Year's card and second person-related information R2 ("This is a family photo taken during a trip to Hawaii") from the second image P2-1 of a family photo was used as an illustration. Besides the above example, there are various first images P1 and second images P2, and various methods can be considered regarding which first person-related information R1 is obtained from which first image P1. Furthermore, various methods can also be considered regarding which second person-related information R2 is obtained from which second image P2 based on which first person-related information R1. Such various methods are shown in the following embodiments.

[0151] In the following embodiments, the structure of the image content determination device 2 is the same as that of the first embodiment described above, and the basic processing sequence until the acquisition of the second person-related information R2 is also the same. Figure 13 The processing order shown is the same. The only difference lies in the content of information such as the type of at least one of the first image P1 and the second image P2, the content of the first person-related information R1, and the content of the second person-related information R2. Therefore, in the following embodiments, the differences from the first embodiment will be emphasized.

[0152] [Second Implementation]

[0153] exist Figure 15 and Figure 16 In the second embodiment shown as an example, the second person-related information R2 of the second image P2-2 is obtained using the first image P1-4. The first image P1-4 is the same as that described in the first embodiment above, so the processing performed on the first image P1-4 will be omitted.

[0154] The second image P2-2 contains the faces of the second person M2 of the four people. Therefore, in the second recognition process, the second recognition unit 38-2 recognizes the faces of the second person M2 of the four people from the second image P2-2 and extracts the four recognized second face images PF2.

[0155] exist Figure 16 In the second acquisition process shown, during the similar image search process, the second acquisition unit 42 searches for a first image P1 containing a face of a first person M1 that is similar to the face of the second person M2 by comparing the second face images PF2 of each of the four second persons M2 included in the second image P2-2 with the first face images PF1 of each of the four first persons M1 included in the first image P1. Then, the second acquisition unit 42 retrieves the first person-related information R1 from the first image information list 48. Figure 16 In the example, compared with the first embodiment Figure 10 Similarly, for the first image P1, images P1-1 to P1-4 are also searched. Then, their first person-related information R1 is obtained.

[0156] exist Figure 16 In the example, the faces of three of the four people in image P2-2 (M2A-M2D) whose second persons M2A-M2C are similar to the faces of the three first persons M1A-M1C in image P1-4. In the formal processing, the second acquisition unit 42 estimates, based on this verification result, that three of the four people in image P2-2 (M2A-M2D) whose second persons M2A-M2C are a family. The second acquisition unit 42 also estimates who the remaining second person M2D in image P2-2 is. The second acquisition unit 42 estimates the age and gender of the second persons M2A-M2D by performing image analysis on image P2-2. In this example, it is estimated that second persons M2A and M2B are a man and a woman aged 30-39, second person M2C is a child under 10 years old, and second person M2D is a woman aged 60-69. The ages of the second person M2A and the second person M2B differ from the age of the second person M2D by approximately 20 years or more. Furthermore, the second person M2D is different from the first person M1A through M1C identified as having a parent-child relationship in the relevant information R1 for the first person. Therefore, the second acquisition unit 42 estimates that the second person M2D is the child of the second person M2A and the second person M2B, i.e., the grandmother of the second person M2C. Based on this estimation, in the second image P2-2, the second person M2D is estimated to be "a woman aged 60-69 and the grandmother of the second person M2C, i.e., the child." The second acquisition unit 42 acquires this estimation result as the relevant information R2 for the second person.

[0157] Image analysis of the second image P2-2 allows estimation of the age and gender of the second person M2 among the four individuals. In this example, the information R2 of the second person is obtained using the information R1 of the first person (M1A-M1C), who are similar to the second persons M2A-M2C and belong to a three-person family. Thus, by using the information R1 of the first person, misidentification of the second person M2D as the child (i.e., the mother of the second person M2C) can be prevented. That is, in this case, during the second acquisition process, the second acquisition unit 42 derives the information R2 of the second person based on the second image P2 and determines the accuracy of the derived information R2 of the second person based on the information R1 of the first person. Therefore, compared to the case where the information R1 of the first person is highly reliable and the derived information R2 of the second person is directly obtained using the information R1 of the first person in determining the accuracy of the information R2 of the second person, a more reliable information R2 of the second person can be obtained.

[0158] In this example, we can assign a label like "grandmother" to image P2-2. Having such a label makes searching for photos of "grandmother" much easier.

[0159] [Third Implementation]

[0160] exist Figure 17 and Figure 18 In the third embodiment shown as an example, in addition to the first person-related information R1 obtained from the first image P1, the account information of the user who owns the first image P1 is also used to estimate the age of the second person M2A reflected in the second image P2.

[0161] In the third embodiment, such as Figure 18 As shown, the second image P2 of the processing object is the same as the second image P2-1 in the first embodiment. Similarly, the first image P1 searched through the similar image search process is also the same as the first images P1-1 to P1-4 in the first embodiment (see reference). Figure 10 ).

[0162] As described in the first embodiment, the holder of the first images P1-1 to P1-4 is user A, and user A's account information was registered when signing the online storage service contract. The account information is stored, for example, in the storage area allocated to user A within the memory 4. The account information includes user A's name "Yamada Taro" and his date of birth, for example, April 1, 1980. The account information is registered for each user, and as a storage format, it can be assigned to each first image P1 like Exif information. Furthermore, account information can be assigned to only one of multiple first images P1 for the same user. In any case, within the memory 4, the account information is associated with each user and each of that user's multiple first images P1. The account information is an example of accompanying information in each first image P1, with the meaning of establishing an association, along with Exif information.

[0163] And, as Figure 7 As shown, images P1-1 to P1-4 are all New Year's cards. All of these images contain the face of the first person, M1A, and the string "CH" ("Yamada Taro"). Furthermore, the face of the first person, M1, contained in image P1-1 is only the face of the first person, M1A, and the name contained in image P1-1 is only "Yamada Taro". Moreover, in images P1-1 to P1-4, the string "CH" contains dates such as "New Year's Day 2010" and "New Year's Day 2014," which can be estimated as the approximate year of photography for the photographic area AP.

[0164] like Figure 17 As shown, in the first acquisition process performed to obtain the first person's related information R1, the first acquisition unit 40 acquires the user A's account information in addition to the string CH contained in the character area AC of the first images P1-1 to P1-4. Since the face of the first person M1, which is shared by all of the first images P1-1 to P1-4, is only the face of the first person M1A, and the shared string CH is only "Yamada Taro", the first acquisition unit 40 estimates that the first person M1A is "Yamada Taro". Moreover, since the string CH "Yamada Taro" is the same as the name "Yamada Taro" in the account information, the first person M1A is estimated to be user A, and the first person M1A's date of birth is "April 1, 1980" contained in the account information.

[0165] And, as Figure 14As shown, the first image P1-4 is a "New Year's card," and contains the string CH representing the date "New Year's Day 2020." Based on this date, the first acquisition unit 40 estimates the year of photography of the photo area AP of the first image P1-4 to be around 2020. Then, assuming the year of photography of the photo area AP of the first image P1-4 is 2020, the birth year of the first person M1A included in the first image P1-4 is 1980, so the age of the first person M1A is estimated to be about 40 years old. The first acquisition unit 40 estimates the age of the first person M1A at the time of photography of the first image P1-4 to be about 40 years old through such estimation. Although this estimated age of the first person M1A uses account information, it is information obtained based on the face of the first person M1A identified from the first image P1-4 and the string CH "New Year's Day 2020," and is therefore an example of the first person-related information R1. Furthermore, the first acquisition unit 40 also performs the same estimation on the first images P1-1 to P1-3, estimating the age of the first person M1A at the time of photography of each of the first images P1-1 to P1-3.

[0166] like Figure 18 As shown, the second acquisition unit 42, in the formal processing, acquires the determination result that the face of the second person M2A in the second image P2-1 is most similar to the face of the first person M1A in the first image P1-4 as the processing result of the similarity image search processing. Furthermore, the second acquisition unit 42 acquires information from the first person-related information R1 in the first image P1-4 that the estimated age of the first person M1A is 40 years old. Based on this information, the second acquisition unit 42 acquires the estimated age of the second person M2A in the second image P2-1 as 40 years old as the second person-related information R2.

[0167] Thus, in the third embodiment, when estimating the age of the second person M2A in the second image P2-1, the second acquisition unit 42 searches for a first image P1-4 containing a face similar to that of the second person M2A in the second image P2-1, and uses the first person-related information R1 from the searched first image P1-4 to acquire the estimated age of the second person M2A, which is second person-related information R2. Therefore, compared to estimating the age of the second person M2A based on the second face image PF2, the reliability of the age estimation is improved by utilizing the highly reliable first person-related information R1.

[0168] In addition, the estimated age can be within a certain range, such as 40-49 years old, 40-45 years old, or 38-42 years old.

[0169] Furthermore, this example illustrates the use of account information during the first acquisition process performed by the first acquisition unit 40, but account information can also be used during the second acquisition process performed by the second acquisition unit 42. For example, in Figure 18 In the formal processing, as described above, the second acquisition unit 42 acquires the result of determining the most similarity between the face of the second person M2A in the second image P2-1 and the face of the first person M1A in the first image P1-4. After acquiring this determination result, the second acquisition unit 42 can use account information to estimate the age of the first person M1A in the first image P1-4 and the age of the second person M2A in the second image P2-1.

[0170] [Fourth Implementation]

[0171] exist Figure 19 and Figure 20 In the fourth embodiment shown, the first acquisition unit 40 determines from a plurality of first images P1 the year in which the number of family members changes, and acquires first person-related information R1 related to the change in the number of family members. The second acquisition unit 42 uses the first person-related information R1 to acquire second person-related information R2 related to second images P2 taken after that year.

[0172] In the fourth embodiment, for example, Figure 19 As shown, the first acquisition unit 40 acquires the changes in the number of family members of the first person M1A from multiple first images P1 of user A as first person-related information R1. The first images P1-1 to P1-3 are the same as the first images P1-1 to P1-3 shown in the above embodiments (see reference). Figure 7 and Figure 17 In the photo area AP of the first image P1-1, the first person M1A "Yamada Taro" is displayed alone, and the date 2010 is included as the string CH. Based on this information, the first acquisition unit 40 retrieves the relevant information R1 from the first image P1-1 that the first person M1A was single in 2010. Figure 7 As shown, information R1 related to the first person in a two-person family (M1A and M1B) in 2014 is obtained from image P1-2. Information R1 related to the first person in a three-person family (M1A, M1B, and M1C) in 2015 is obtained from image P1-3. Furthermore, referring to the information R1 related to the first person in image P1-2, it can be seen that in January 2014, the first person M1A and M1B were a two-person family, and the first person M1C did not exist. Therefore, the first person M1C included in image P1-3 is a child born in 2014. This information is also obtained as information R1 related to the first person.

[0173] like Figure 20 As shown, when determining the image content of the second image P2-3, the second acquisition unit 42 searches for the first image P1-3 as the first image P1 containing the face of the first person M1, which is similar to the face of the second person M2 contained in the second image P2-3. The search is performed based on the similarity between the faces of the second persons M2A and M2B contained in the second image P2-3 and the faces of the first persons M1A and M1B contained in the first image P1-3. Then, the second acquisition unit 42 reads the first person-related information R1 corresponding to the first image P1-3 from the first image information list 48.

[0174] In its formal processing, the second acquisition unit 42 estimates the age of the child (second person M2C) when the second image P2-3 was taken to be 5 years old based on the information "the child (first person M1C) was born in 2014" contained in the first person information R1 of the first image P1-3 and the information "2019" contained in the accompanying information of the second image P2-3. Furthermore, the second acquisition unit 42 estimates that "the second image P2-3 was taken at a shrine" based on the photography date "November 15th" and the photography scene of the second image P2-3 being a "shrine." This information is acquired as the second person information R2. Based on this second person information R2, the second image P2-3 is, for example, labeled with "Shichi-Go-San."

[0175] As explained above, according to the fourth embodiment, the first acquisition unit 40 determines the year in which a family member was added based on a plurality of first images P1-1 to P1-3, and acquires the year in which the family member was added as first person-related information R1. The second acquisition unit 42 estimates the age of the child in the second images P2-3 acquired after that year by using the year in which the family member was added as the year the child was born, and acquires the estimated age of the child as second person-related information R2. Therefore, according to this structure, compared with the case of acquiring first person-related information R1 from only one first image P1, it is possible to acquire a variety of information as first person-related information R1, and further, it is possible to acquire highly reliable information as second person-related information R2.

[0176] Furthermore, based on the child's estimated age and the date of the photograph, it is estimated that the second image P2-3 is a commemorative photograph of an event corresponding to the child's age, such as Shichigosan. According to the technology of the present invention, since the first person-related information R1 is utilized, a more reliable second person-related information R2 can be obtained compared to the case where the event related to the second person M2 is estimated solely from the image content of the second image P2.

[0177] In addition, the types of events related to the second character M2, who is tagged with the second image P2, include traditional events celebrating a child's healthy growth, such as the Shichi-Go-San festival and shrine visits (events celebrating a baby's growth), as well as events celebrating longevity, such as the 60th birthday and the 88th birthday, life events such as weddings and school ceremonies, and planned events such as festivals and concerts. Furthermore, events also include school festivals such as sports meets and academic achievement presentations.

[0178] [Fifth Implementation]

[0179] In the above embodiments, the first image P1 held by user A is described as a New Year's card sent by user A. However, the first image P1 held by user A can also be an image such as a New Year's card received by user A.

[0180] Figure 21 and Figure 22 The fifth embodiment shown is an example of obtaining second person information R2 from a second image P2 using first person information R1 of a first image P1 where user A is the recipient.

[0181] like Figure 21 As shown, the first image P1-5 is the image of the New Year's card for which user A is the recipient. It is stored together with the first image P1-4, which is the first image folder 13-1, which is the folder for user A.

[0182] In the first image P1-4 where user A is the sender, the sender's name "Yamada Taro" is included. However, in the first image P1-5 where user A is the recipient, the sender's name includes "Sato Saburo" instead of "Yamada Taro". Furthermore, the character area AC in the first image P1-5 contains the strings "Happy New Year" and "Let's go fishing next time" CH.

[0183] In the first acquisition process of the first image P1-5, the first acquisition unit 40 estimates that the first image P1-5 is a "New Year's card" because it contains the string "Happy New Year" as CH. Furthermore, since the sender's name is "Sato Saburo," the first person M1F contained in the photo area AP is estimated to be named "Sato Saburo." Also, based on the message "Let's go fishing next time" contained as the string CH, the first acquisition unit 40 estimates that the first person M1F's hobby in the first image P1-5 is "fishing." Furthermore, since the first image P1-5 is stored as the first image P1-5 of user A "Yamada Taro," the first acquisition unit 40 estimates that the sender "Sato Saburo" is a friend of user A "Yamada Taro." The first acquisition unit 40 acquires this information as the first person-related information R1 for the first image P1-5.

[0184] like Figure 22 As shown, in the similar image search processing, the second acquisition unit 42 searches for the first image P1-5 based on the similarity of the face of the second person M2F contained in the second image P2-4 and the first person M1F contained in the first image P1-5. The second acquisition unit 42 acquires the relevant information R1 of the first person in the searched first image P1-5.

[0185] In the formal processing, the second acquisition unit 42, based on the information R1 related to the first person in the first image P1-5, which states "hobby is fishing," and the fact that the shooting scene of the second image P2-4 is "sea," and the GPS information is also Tokyo Bay (see also reference),... Figure 9 The second acquisition unit 42 obtains the second person-related information R2, which states that "the second image P2-4 is a photo taken while sea fishing." Since the scene depicted in the second image P2-4 is the sea, the GPS information is Tokyo Bay, and fish are visible, image analysis of the second image P2-4 can estimate that the second image P2-4 shows two people, M2A and M2F, fishing. The second acquisition unit 42 can then use this estimation result to derive the second person-related information R2, indicating that fishing is the hobby of the two people, M2A and M2F. The second acquisition unit 42 can then use the first person-related information R1 to determine the accuracy of the derived second person-related information R2. Therefore, the reliability of the second person-related information R2 is improved.

[0186] Based on the relevant information R2 of the second person, for example, the second image P2-4 is labeled with the tag "Sea: Fishing".

[0187] As explained above, in the fifth embodiment, a New Year's card addressed to user A is used as the first image P1. When a friend of user A is displayed as the first person M1 in the first image P1, the second acquisition unit 42 acquires second person-related information R2 from the second images P2-4 containing the friend's face using first person-related information R1. Therefore, according to this structure, for example, compared to the case where second person-related information R2 is acquired based on user A's first person-related information R1, more reliable second person-related information R2 can be acquired.

[0188] In the above embodiments, the first image P1 is described as an image with a character region, having a photo region AP and a character region AC. However, the first image P1 is not limited to an image with a character region; for example, it could be an image with a character region such as... Figure 23The character-mapped image 52 is an image in which a specific word, pre-registered as a character, is mapped onto a photo region AP containing only the face of the first person.

[0189] In the technology of this invention, in order to obtain highly reliable information R1 related to a first person, it is preferable that the first image P1 contains characters representing highly reliable information related to the first person M1. As explained in the above embodiments, it is considered that images with character regions AC contain more characters representing highly reliable information than images without character regions. However, even if there is no character region AC, if characters representing highly reliable information are contained in the photo region AP, it is preferable to actively use that image P as the first image P1.

[0190] Specific words that indicate highly reliable information related to the first person M1 include, for example, entrance ceremonies, graduation ceremonies, coming-of-age ceremonies, weddings, and birthday celebrations. These specific words are highly likely to be used to express various information about events related to the first person. For example, like... Figure 23 As with character-mapped image 52, when character-mapped image 52 containing the string "○○ University Graduation Ceremony" is used as the first image P1, the first acquisition unit 40 can obtain from character-mapped image 52 the university that the first person M1 of the first image P1 graduated from. This information can be said to be highly reliable as information related to the first person R1. Furthermore, as Figure 23 As shown, when the graduation year, such as "2020", is included as a character, the first acquisition unit 40 is able to acquire the graduation year and use the acquired graduation year in the estimation of the approximate age of the first character M1.

[0191] In the first embodiment, such as Figure 4 As shown, Figure 23 The character-mapped image 52 does not have a character region AC, and therefore is classified as image P2 by the classification unit 36. To classify the character-mapped image 52 as image P1, the following condition needs to be added to the classification unit 36's criteria for classifying image P as image P1: Even without a character region AC, for image P that contains both a person's face and a specific word in the photo region AP, the condition for classification as image P1 is added. Thus, the character-mapped image 52 is classified as image P1.

[0192] For example, such as Figure 23As shown, an example will be described where the character input image 52 has characters including the specific word "graduation ceremony" and the character input image 53 has characters including "under construction" but not including the specific word. Although there is no character area AC in the character input image 52, the photo area AP includes both the face of a person and the specific word, so the classification unit 36 classifies the character input image 52 as the first image P1. On the other hand, since there is no character area AC in the character input image 53 and only the face of a person is included in the photo area AP without including the specific word, the classification unit 36 classifies the character input image 53 as the second image P2. The classification unit 36 determines the presence or absence of the specific word by referring to the thesaurus data 46 in which the specific word is pre-registered.

[0193] Moreover, the specific word can be, for example, the date printed on the printed photo PA (refer to Figure 1 ). In addition, the first image is not limited to the character input image 52, and can also be an image including a specific word handwritten on the printed photo PA. The printed photo PA including the handwritten specific word means, for example, a printed photo PA on which a user has written a date such as "October 〇, 2010" and a specific word such as "〇〇' s graduation ceremony" with a ballpoint pen or the like when organizing the printed photo PA. Such handwritten information often contains information related to the person shown in the photo. Compared with the case of classifying the image P with the handwritten specific word as the second image P2 by using the image P with the handwritten specific word as the first image P1, diverse and highly reliable first person-related information R1 can be obtained, and further highly reliable second person-related information R2 can be obtained.

[0194] Moreover, as the specific word, it can also include words such as "Happy New Year" and "Merry Christmas" that can be judged as greeting cards. For example, even for greeting cards such as New Year cards or Christmas cards, there are cases where there is no character area AC distinguished from the photo area AP. In this case, the specific word is often included in the photo area AP. If words that can be judged as greeting cards are registered as specific words, greeting cards with only the photo area AP can be used as the first image P1.

[0195] Moreover, in Figure 23The example described illustrates how the presence or absence of a specific word in the character-free regions of images 52 and 53 can classify images 52 containing a specific word as the first image P1. However, not only images without character regions, but also images with character regions, such as the first images P1-4 of the greeting cards shown in the above embodiments, can be classified as the first image P1 by identifying a specific word. Furthermore, not all images with character regions contain meaningful character information; therefore, the presence or absence of a specific word can exclude images P1 that do not contain meaningful character information.

[0196] Furthermore, in the above embodiment, an example was described with a classification unit 36 ​​provided in the image content determination device 2, but the classification unit 36 ​​may also be omitted. For example, the image content determination device 2 may process the first image P1 and the second image P2, which have already been classified by other devices.

[0197] In the above embodiments, for example, the hardware structure of the computer that performs various processes of the classification unit 36, recognition unit 38, first acquisition unit 40, second acquisition unit 42, and tagging unit 44 of the image content determination device 2 can utilize various processors as shown below. Among these processors, besides the CPU 18, a general-purpose processor that performs the functions of various processing units by executing software (e.g., classification program 30, recognition program 31, first acquisition program 32, second acquisition program 34, and tagging program 35), it also includes FPGAs (Field Programmable Gate Arrays) as processors whose circuit structure can be changed after manufacturing, PLDs (Programmable Logic Devices) and / or ASICs (Application Specific Integrated Circuits) as processors with dedicated circuit structures designed specifically for performing specific processes. A GPU (Graphics Processing Unit) can also be used instead of an FPGA.

[0198] A processing unit can consist of one of these various processors, or it can consist of a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs and / or a combination of CPU and FPGA or a combination of CPU and GPU). Furthermore, a single processor can also constitute multiple processing units.

[0199] As examples of a single processor comprising multiple processing units, there are two main approaches: First, as exemplified by computers such as client machines and servers, a single processor is constructed using a combination of one or more CPUs and software, and this processor functions as multiple processing units. Second, as exemplified by SoCs (System-on-Chips), a processor that implements the overall functionality of a system containing multiple processing units is used, implemented by a single IC (Integrated Circuit) chip. In these cases, various processing units are constructed using one or more of the aforementioned processors as part of the hardware architecture.

[0200] Furthermore, as the hardware structure of these various processors, more specifically, they can use circuits composed of circuit elements such as semiconductor elements.

[0201] Furthermore, in the first embodiment described above, various programs including a classification program 30, an identification program 31, a first acquisition program 32, a second acquisition program 34, and a tagging program 35 are stored in the internal program memory 22; however, the technology of the present invention is not limited thereto. Figure 2 Similarly, the memory 4 shown can store various programs on any portable storage medium, such as an SSD or USB (Universal Serial Bus) internal storage. In this case, as an example, such as... Figure 24 As shown, various programs stored in storage medium 60 are connected to and installed in image content determination device 2 in the same way as memory 4. CPU 18 performs classification processing, first recognition processing, first acquisition processing, second recognition processing, second acquisition processing, and tagging processing according to the various programs installed.

[0202] Furthermore, similar to memory 4, it can be accessed via communication network N (reference). Figure 1 Various programs are stored in the storage unit of other computers or server devices connected to the image content determination device 2, and various programs are downloaded to the image content determination device 2 according to the request of the image content determination device 2. In this case, the CPU 18 performs classification processing, first recognition processing, first acquisition processing, second recognition processing, second acquisition processing, and tagging processing according to the downloaded programs.

[0203] As described in the above embodiments, the image content determination device of the present invention may add the following appendix.

[0204] [Note 1]

[0205] The first image may include an image with a character area, which includes: a photographic area containing the face of the first person; and a character area containing characters as a blank area outside the outline of the photographic area. The second image may be an image without a character area that only contains the photographic area containing the face of the second person.

[0206] [Note 2]

[0207] The first image may be an image representing at least one of the greeting card and identification document.

[0208] [Note 3]

[0209] The first image may include an image of a photo area containing only the face of the first person, or an image of a characterless area where a specific word pre-registered as a character is projected onto the photo area.

[0210] [Note 4]

[0211] The first image may contain specific words that are pre-registered as characters.

[0212] [Note 5]

[0213] The processor can perform classification processing to classify multiple images into Image 1 and Image 2.

[0214] [Note 6]

[0215] Information about the first character can be obtained from the first image of the same holder as the second image.

[0216] [Note 7]

[0217] The information related to the first person may include at least one of the following: the first person's name, address, phone number, age, date of birth, and hobbies.

[0218] [Note 8]

[0219] In at least one of the first acquisition process and the second acquisition process, the processor may utilize the accompanying information attached to the first image or the second image.

[0220] [Note 9]

[0221] In the second acquisition process, the processor can derive information about the second person based on the second image and determine the accuracy of the derived information about the second person based on the information about the first person.

[0222] [Note 10]

[0223] The information related to the second person can be at least one of the events related to the second person and the estimated age of the second person.

[0224] [Note 11]

[0225] When the first image contains the faces of multiple first persons, the information related to the first persons may include information indicating the relationship between the multiple first persons, and / or, when the second image contains the faces of multiple second persons, the information related to the second persons may include information indicating the relationship between the multiple second persons.

[0226] [Note 12]

[0227] Information indicating relationships between multiple first persons or multiple second persons may include at least one of family relationships, kinship relationships, and friendship relationships.

[0228] [Note 13]

[0229] In the second acquisition process, the processor can utilize the first person information corresponding to multiple first images in the acquisition of the second person information.

[0230] The technology of the present invention can also be appropriately combined with the various embodiments and / or variations described above. Furthermore, it is not limited to the embodiments described above; various structures can certainly be adopted as long as they do not depart from the spirit of the invention. In addition to programs, the technology of the present invention also relates to storage media for non-transitory storage of programs.

[0231] The descriptions and illustrations above are detailed explanations of the parts related to the technology of this invention, and are merely one example of the technology of this invention. For example, the descriptions related to the above-described structure, function, effect, and effect are examples of the structure, function, effect, and effect of the parts related to the technology of this invention. Therefore, without departing from the spirit of this invention, unnecessary parts may be deleted from the descriptions and illustrations above, or new elements may be added or replaced. Furthermore, to avoid complications and to facilitate understanding of the parts related to the technology of this invention, descriptions related to common technical knowledge that are not particularly necessary to explain in terms of enabling the implementation of this invention have been omitted from the descriptions and illustrations above.

[0232] In this specification, "A and / or B" has the same meaning as "at least one of A and B". That is, "A and / or B" can mean only A, only B, or a combination of A and B. Furthermore, in this specification, the same approach applies to situations where three or more items are connected by "and / or".

[0233] All disclosures of Japanese Patent Application No. 2020-058617, filed on March 27, 2020, are incorporated herein by reference. Furthermore, all documents, patent applications, and technical standards described in this specification are incorporated herein by reference to the same extent as those specifically and separately described and incorporated herein by reference.

Claims

1. An image content determination device, comprising at least one processor and a memory, The memory includes an unprocessed folder storing untagged images and a processed folder storing tagged images. The images include: a first image containing characters and the face of a first person; and a second image consisting of a characterless area containing only the face of a second person. The processor performs the following processing: Perform a first recognition process to identify the characters and the face of the first person from the first image containing the characters; Perform a first acquisition process to obtain first person-related information related to the first person contained in the first image based on the identified characters and the face of the first person; Perform the first tagging process, in which the obtained information related to the first person is used to assign search tag information to the first image for keyword searching of the first image; Perform a second recognition process to identify the face of the second person from the second image containing the face of the second person; as well as A second acquisition process is performed to obtain second person-related information concerning the second person included in the second image. The second acquisition process includes a similar image search process and a formal process. In the similar image search process, a first image containing a face similar to the face of the second person included in the second image is searched by comparing a second face image representing the face of the second person identified from the second image and a first face image representing the face of the first person identified from the first image. In the formal process, the second person-related information is obtained using the first person-related information corresponding to the first image containing the face of the first person similar to the face of the second person. A second tagging process is performed, in which tag information is added to the second image based on the obtained information related to the second person. This tag information is used as keywords for searching the second image.

2. The image content determination device according to claim 1, wherein, The first image includes an image with a character region, which includes: a photographic area containing the face of the first person; and a character region containing the characters as a blank area outside the outline of the photographic area.

3. The image content determination device according to claim 1 or 2, wherein, The first image includes an image representing at least one of a greeting card and an identity document.

4. The image content determination device according to claim 1 or 2, wherein, The first image includes a character-mapped image, which is an image of a characterless area containing only a photographic region of the face of the first person, and a specific word pre-registered as the character is mapped into the photographic region.

5. The image content determination device according to claim 1 or 2, wherein, The first image contains a specific word that has been pre-registered as the character.

6. The image content determination device according to claim 1 or 2, wherein, The processor performs a classification process that categorizes multiple images into the first image and the second image.

7. The image content determination device according to claim 1 or 2, wherein, The information related to the first person was obtained from the first image held by the same holder as the second image.

8. The image content determination device according to claim 1 or 2, wherein, The information related to the first person includes at least one of the following: the first person's name, address, phone number, age, date of birth, and hobbies.

9. The image content determination device according to claim 1 or 2, wherein, In at least one of the first acquisition process and the second acquisition process, the processor utilizes additional information accompanying the first image or the second image.

10. The image content determination device according to claim 9, wherein, In the second acquisition process, the processor derives the relevant information of the second person based on the second image, and determines the accuracy of the derived relevant information of the second person based on the relevant information of the first person.

11. The image content determination device according to claim 1 or 2, wherein, The information related to the second person is at least one of events related to the second person and the estimated age of the second person.

12. The image content determination device according to claim 1 or 2, wherein, When the first image contains the faces of multiple first persons, the information related to the first persons includes information indicating the relationship between the multiple first persons, and / or, when the second image contains the faces of multiple second persons, the information related to the second persons includes information indicating the relationship between the multiple second persons.

13. The image content determination device according to claim 12, wherein, Information indicating the relationship between multiple first persons or multiple second persons includes at least one of family relationship, kinship relationship, and friendship relationship.

14. The image content determination device according to claim 1 or 2, wherein, In the second acquisition process, the processor utilizes the first person-related information corresponding to the plurality of the first images in the acquisition of the second person-related information.

15. A method for determining image content, comprising the following steps: In the memory, create an unprocessed folder to store images that have not been tagged and a processed folder to store images that have been tagged. The image comprises: a first image containing characters and the face of a first person; and a second image, which is a characterless region consisting only of a photographic area containing the face of a second person; Perform a first recognition process to identify the character and the face of the first person from a first image containing the character and the face of the first person; Perform a first acquisition process to obtain first person-related information related to the first person contained in the first image based on the identified characters and the face of the first person; The first tagging process is performed, in which the first person information is used to assign the first image with search tag information for keyword search of the first image; the second recognition process is performed to identify the face of the second person from the second image containing the face of the second person but without a character area. as well as A second acquisition process is performed to obtain second person-related information concerning the second person included in the second image. The second acquisition process includes a similar image search process and a formal process. In the similar image search process, a first image containing a face similar to the face of the second person included in the second image is searched by comparing a second face image representing the face of the second person identified from the second image and a first face image representing the face of the first person identified from the first image. In the formal process, the second person-related information is obtained using the first person-related information corresponding to the first image containing the face of the first person similar to the face of the second person. A second tagging process is performed, in which tag information is added to the second image based on the obtained information related to the second person. This tag information is used as keywords for searching the second image.

16. A storage medium storing an image content determination program for causing a computer including at least one processor to perform a process comprising the following steps: In the memory, create an unprocessed folder to store images that have not been tagged and a processed folder to store images that have been tagged. The image comprises: a first image containing characters and the face of a first person; and a second image, which is a characterless region consisting only of a photographic area containing the face of a second person; Perform a first recognition process to identify the character and the face of the first person from a first image containing the character and the face of the first person; Perform a first acquisition process to obtain first person-related information related to the first person contained in the first image based on the identified characters and the face of the first person; Perform the first tagging process, in which the obtained information related to the first person is used to assign search tag information to the first image for keyword searching of the first image; Perform a second recognition process to identify the face of the second person from a second image that contains the face of the second person but has no character area; as well as A second acquisition process is performed to obtain second person-related information concerning the second person included in the second image. The second acquisition process includes a similar image search process and a formal process. In the similar image search process, a first image containing a face similar to the face of the second person included in the second image is searched by comparing a second face image representing the face of the second person identified from the second image and a first face image representing the face of the first person identified from the first image. In the formal process, the second person-related information is obtained using the first person-related information corresponding to the first image containing the face of the first person similar to the face of the second person. A second tagging process is performed, in which tag information is added to the second image based on the obtained information related to the second person. This tag information is used as keywords for searching the second image.

Citation Information

Patent Citations

  • Digital data tagging method and system

    JP2009526302A

  • Image classification device and image classification method

    JP2010067014A

  • Game machine

    JP2020058617A

  • System and method for providing objectified image renderings using recognition information from images

    US20060251338A1

  • Automatic Creation Of A Scalable Relevance Ordered Representation Of An Image Collection

    US20110038550A1