Information processing device, control method thereof and program

JP2023163733A5Pending Publication Date: 2025-05-12CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022074828
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-04-28
Publication Date
2025-05-12

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To select an image suitable for a user's purpose from a plurality of images.SOLUTION: An information processing device according to the present invention comprises: face detection means which detects a face region from an image; processing means which processes the image on the basis of the detected face region; attribute detection means which detects a subject with a prescribed attribute on the basis of the processed image; region detection means which detects an unnecessary region included in the processed image in which the subject with the prescribed attribute is detected; and selection means which selects the image as an inappropriate image or an appropriate image on the basis of the detected unnecessary region.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, a control method thereof, a program, and a recording medium. [Background technology]

[0002] When a large number of images are taken with a camera, it is a heavy burden for the user to select images suitable for the user's purpose. Patent Document 1 discloses an imaging device that detects faces from captured images and deletes images in which no faces are detected or images that do not meet the shooting conditions based on information such as facial expressions, facial orientation, eye opening / closing, and line of sight. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2007-20105 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the imaging device of Patent Document 1 can only classify images based on the facial area of ​​the captured image itself, so it may not select an image that is suitable for the user's purpose, or may select an image that is not suitable for the user's purpose. The present invention has been made in view of the above-mentioned problems, and has as its object to select an image suitable for a user's purpose from a plurality of images. [Means for solving the problem]

[0005] The information processing device of the present invention is characterized by having a face detection means for detecting a face area from an image, a processing means for processing the image based on the detected face area, an attribute detection means for detecting a subject with a predetermined attribute based on the processed image, an area detection means for detecting an unnecessary area included in the processed image in which the subject with the predetermined attribute has been detected, and a selection means for selecting the image as an inappropriate image or an appropriate image based on the detected unnecessary area. [Effects of the Invention]

[0006] According to the present invention, it is possible to select an image suitable for a user's purpose from a plurality of images. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 illustrates an example of a configuration of an information processing system. [Figure 2] FIG. 2 is a diagram illustrating an example of a functional configuration of an image selection device. [Figure 3] 10 is a flowchart illustrating an example of processing performed by the image selection device. [Figure 4] 10 is a flowchart illustrating an example of an inappropriate image selection process. [Figure 5] 10 is a flowchart illustrating an example of a suitable image selection process. [Figure 6] FIG. 10 is a diagram showing an example of a folder in which images are stored according to the results of an image selection process. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. FIG. 1 is a diagram illustrating an example of the configuration of an information processing system 10. As shown in FIG. The information processing system 10 includes an imaging device 100, an image selection device 110, and an external device 120. The imaging device 100, the image selection device 110, and the external device 120 are communicably connected to one another via a network 130. The network 130 may be wireless or may be wired via a USB cable or the like.

[0009] The imaging device 100 is a digital camera that automatically captures images according to predetermined conditions. For example, the imaging device 100 automatically captures images at predetermined time intervals or automatically captures images based on information about a subject (e.g., the subject's face) detected by the imaging device 100. The imaging device 100 also has a storage unit that stores captured images. The imaging device 100 transmits a plurality of stored images to the image selection device 110 via the network 130. The imaging device 100 may also have an automatic pan-tilt function.

[0010] The image selection device 110 automatically selects multiple images received from the imaging device 100. The image selection device 110 is an example of an information processing device. The multiple images received from the imaging device 100 may include many images that are not suitable for the user's purpose because they were automatically captured by the imaging device 100. Therefore, by having the image selection device 110 select images suitable for the user's purpose, the user's work of visually selecting images suitable for the purpose can be reduced. The specific configuration and processing of the image selection device 110 will be described later. The image selection device 110 transmits the automatically selected images to the external device 120 via the network 130. The image selection device 110 may be integrated with the imaging device 100.

[0011] The external device 120 stores images automatically selected and transmitted by the image selection device 110. The external device 120 is an external PC or an external server. The external device 120 is, for example, a device managed by a business that sells images to users. The images stored in the external device 120 can be viewed or selected by a user via a smartphone, tablet terminal, PC, or the like owned by the user, for example, through a website managed by the business. Note that the external device 120 may be an information processing system that allows the user to directly view or select images without going through a website.

[0012] <Hardware configuration> An example of the hardware configuration of the image selection device 110 will be described with reference to FIG. The image selection device 110 includes a CPU 101 , a RAM 102 , a ROM 103 , a storage device 104 , and a network interface I / F 105 .

[0013] The CPU 101 executes programs stored in the ROM 103 and programs such as an OS (operating system) and applications loaded from the storage device 104 to the RAM 102. By executing the programs stored in the ROM 103, the CPU 101 performs the function of an image selection device that reads in a plurality of images and selects from each of the read images an image suitable for the user's purpose. However, this is not limited to this case, and the image selection device 110 may also include an image selection processing unit realized by a dedicated circuit such as a GPU or ASIC.

[0014] The RAM 102 is the main memory of the CPU 101 and functions as a work area and the like. The ROM 103 stores various programs. The storage device 104 stores the images received from the imaging device 100. The storage device 104 may be built into the image selection device 110 in advance, or may be a removable recording medium. The network I / F 105 functions as an interface for connecting the image selection device 110 to the network 130. The network 130 may be, for example, a LAN or a public switched telephone network (PSTN). The system bus 111 is a communication path for communication between the hardware elements.

[0015] <Functional configuration> An example of the functional configuration of the image selection device 110 will be described with reference to FIG. FIG. 2 is a diagram showing an example of the functional configuration of the image selection device 110. Note that FIG. 2 shows only functions related to the process of selecting images suitable for a user's purpose. The functional configuration shown in FIG. 2 is realized by the CPU 101 executing a program recorded in the ROM 103 or a program such as an application loaded from the storage device 104 to the RAM 102. Note that the execution results of each process are stored in the RAM 102.

[0016] The image selection device 110 is configured to include an inappropriate image selection unit 210 and an appropriate image selection unit 220. The inappropriate image selection unit 210 selects images that are not suitable for the user's purpose (inappropriate images) from a plurality of images. The appropriate image selection unit 220 selects images that are suitable for the user's purpose (appropriate images) to be finally presented to the user from the images that have not been selected as inappropriate images. However, the image selection device 110 may be configured not to include the appropriate image selection unit 220, and to present all images that the inappropriate image selection unit 210 has not selected as inappropriate images to the user.

[0017] Here, an image suitable for a user's purpose is an image desired by the user. In this embodiment, for example, a case is assumed in which daily photos at a kindergarten or nursery school are sold to parents. The purpose in this embodiment is a purpose in which a user purchases images of children. Therefore, an image suitable for the purpose in this embodiment (appropriate image) is an image that the user wishes to purchase. On the other hand, an image that is not suitable for the purpose in this embodiment (inappropriate image) is an image that the user does not wish to purchase. For example, an inappropriate image is an image in which an adult, rather than a child, is the main subject, or an image in which part of another person's arm or body is captured in front of the child, the subject, obstructing the subject. Note that an image suitable for a purpose is not limited to an image that the user wishes to purchase, but may also be an image that the user wishes to keep without deleting.

[0018] The inappropriate image selection unit 210 includes a face detection unit 211 , an image processing unit 212 , an attribute detection unit 213 , an area detection unit 214 , and an inappropriate image determination unit 215 . The face detection unit 211 detects a face area from an image using image processing technology. The image processing unit 212 processes the image based on the face region detected by the face detection unit 211. The image processing unit 212 performs processing such as trimming the image around the face region. The attribute detection unit 213 detects a subject with a predetermined attribute based on the image processed by the image processing unit 212. The attribute detection unit 213 of this embodiment detects the attribute of the subject, such as whether the subject is an adult or a child. The area detection unit 214 detects unnecessary areas that were captured during shooting and are included in an image in which a subject with a predetermined attribute has been detected by the attribute detection unit 213. The inappropriate image determination unit 215 selects images that are not suitable for the user's purpose (inappropriate images) from among the processed images based on the determination results of the face detection unit 211, the attribute detection unit 213, and the area detection unit 214.

[0019] The suitable image selection unit 220 includes a grouping unit 221 and a suitable image determination unit 222 . A grouping section 221 sets images that have not been selected as inappropriate images by the inappropriate image selection section 210 as suitable candidate images, and groups the suitable candidate images by images similar in composition, color tone, etc., or by individuals appearing in the images. The appropriate image determination unit 222 selects a designated number of images suitable for the user's purpose from each group.

[0020] <Flowchart> Fig. 3 is a flowchart relating to the process of selecting images by the image selection device 110. The flowchart shown in Fig. 3 is implemented by the CPU 101 executing a program recorded in the ROM 103 or a program such as an application loaded from the storage device 104 to the RAM 102.

[0021] In S301, the inappropriate image selection unit 210 acquires an image dataset to be selected as images suitable for a user's purpose. In this embodiment, the inappropriate image selection unit 210 acquires the image dataset by receiving, via the network I / F 105, a plurality of images automatically captured by the imaging device 100. However, the image dataset is not limited to images received from the imaging device 100, and may be images stored in the ROM 103 or the storage device 104, or images received via the network I / F 105 from the external device 120. Furthermore, the image dataset is not limited to images automatically captured by the imaging device 100, and may be images intentionally captured by a photographer.

[0022] In S302, the inappropriate image selection unit 210 acquires one image from which inappropriate images are to be selected from the acquired image data set. At this time, the inappropriate image selection unit 210 may perform processing so that an image intentionally taken by the photographer is not selected as an inappropriate image. Specifically, for an image intentionally taken by the photographer, the photographer or the imaging device associates evaluation information with the image and records it. Here, the evaluation information is a value indicating a favorite or an evaluation rank value, etc. Therefore, the inappropriate image selection unit 210 may read the evaluation information when acquiring an image, and if evaluation information equal to or greater than a predetermined value is associated with the image, the acquired image may not be treated as a target image, and the process may proceed to S307, which will be described later, by skipping S302 to S306.

[0023] In S303, the inappropriate image selection unit 210 executes an inappropriate image selection process to determine whether the image acquired in S302 is an inappropriate image. In the case of the application of this embodiment, images that do not show a child or images in which the main subject is not a child are unnecessary images for the guardian. Also, images that show unnecessary areas such as body parts of a person other than the main subject, such as the arm or back, or tableware placed on a table in front of the main subject, or images in which the main subject is obscured by unnecessary areas are also unnecessary images for the guardian.

[0024] FIG. 4 is a flowchart showing the inappropriate image selection process. <Face area detection> In S401, the face detection unit 211 detects a face area from a target image using image processing technology and obtains the coordinates and size of the face area. Here, the face detection unit 211 uses machine learning technology as the image processing technology for detecting the face area, and outputs the coordinates and size of the face area in response to inputting the image into a neural network model. The face detection unit 211 stores the output coordinates and size of the face area. Note that face detection technology using known image processing technology may also be used to detect the face area.

[0025] In S402, the inappropriate image determination unit 215 determines whether or not the face detection unit 211 detected a face area from the target image in S401. If a face area is detected, the process proceeds to S403, and if a face area is not detected, the process proceeds to S409. Here, if a face area is not detected, the image does not include a human subject, and corresponds to an image that is not suitable for the user's use in this embodiment.

[0026] <Image processing> In S403, the image processing unit 212 processes the target image based on the face area detected in S401. Specifically, the image processing unit 212 trims the target image so as to increase the size of the subject's face area based on information about the coordinates and size of the face area and parameters related to photography used when the target image was captured. Here, the parameters related to photography include ISO sensitivity, shutter speed, etc., and are recorded in association with the image.

[0027] Here, the image processing unit 212 determines the trimming area so that a certain margin is formed around the coordinates of the face area, with the coordinates of the face area as the center. The size of the margin may be changed according to the size of the face area, such as by making the left and right margins the same width as the face area, or may be adjusted according to the image size after trimming, such as by adjusting the image to a width of 2000 pixels. Furthermore, if the trimming area is too small, problems with image quality, such as noticeable noise and reduced resolution, may occur. Therefore, for images with high ISO sensitivity, the size of the trimming area may be changed based on shooting parameters, such as by setting a larger margin around the face when trimming.

[0028] The method for determining the trimming area is not limited to the above-described method. For example, a machine learning model may be used to estimate the orientation of the face or body from an image, and if the person is facing forward, the trimming area may be set so that the left and right margins are equal. On the other hand, if the person is facing right, the trimming area may be set so that the margin ratio on the right side of the subject is larger. Alternatively, the trimming area may be calculated directly from the target image using a machine learning model.

[0029] Furthermore, there are cases where multiple face areas are detected in S401. In this case, the image processing unit 212 may determine the trimming areas so that all of the detected face areas are included in the trimmed image and generate one trimmed image, or may determine the trimming area for each detected face area and generate multiple trimmed images.

[0030] The image processing unit 212 stores the generated trimmed image in the ROM 103 or the storage device 104. At this time, the image processing unit 212 may overwrite the trimmed image on the image to be trimmed, or may leave the image to be trimmed as it is. Note that the image processing unit 212 is not limited to processing for generating a trimmed image, but may also perform image processing for improving image quality, for example, by converting pixel values ​​of the trimmed image or the target image through contrast adjustment and color correction.

[0031] <Detection of specified attributes> In S404, the attribute detection unit 213 detects a subject with a predetermined attribute based on the image (trimmed image) processed in S403. The predetermined attribute in this embodiment refers to a child when subjects are classified as adults or children. In this embodiment, the attribute detection unit 213 uses image processing technology to detect whether the attribute of the main subject is a child. Here, the image for detecting whether it is a child is a trimmed image with a facial region as its center. For example, in an image taken at a distance from the child and showing the child at a small size, or in an image showing multiple children and adults, the detection accuracy decreases when detecting whether the attribute of the main subject is a child because the image region is small. On the other hand, in this embodiment, the detection accuracy can be improved by detecting whether the attribute of the main subject is a child based on the trimmed image. Furthermore, the size of the image data of the trimmed image is smaller than that of the image before trimming, which reduces processing time.

[0032] In addition, when face detection or trimming is performed after detecting the attributes of the main subject, if the face detection or attribute detection is incorrect, a trimmed image that is not suitable for the user's purpose will ultimately be generated and presented to the user. For example, if an image containing both an adult and a child is detected as the main subject to be a child, but the child's facial area cannot be detected, a trimmed image that focuses only on the adult will be generated, and an image in which the main subject is an adult will be presented to the user. By detecting the attributes of the main subject based on a trimmed image centered on the facial area as in this embodiment, it is possible to prevent the presentation of an image that is not suitable for the user's purpose.

[0033] To detect whether the attribute of the main subject is a child, the attribute detection unit 213 first determines, as the main subject, the subject whose face region is closest to the center coordinates of the larger face region based on the coordinates and size information of the face region detected in S401. Next, by estimating the age of the main subject using a known age estimation technique or machine learning model for the determined face region of the main subject, it is possible to detect whether the attribute of the main subject is a child. Alternatively, machine learning technique may be used to detect the attribute of the subject by inputting an image into a neural network model. Alternatively, machine learning technique may be used to directly detect whether the main subject is a child using an image recognition model that performs binary judgment using image data as input.

[0034] In S405, the inappropriate image determination unit 215 determines whether or not the attribute detection unit 213 detected a subject with a predetermined attribute in S404. In this embodiment, the attribute detection unit 213 determines whether or not the attribute of the main subject of the trimmed image is a child. If the attribute of the main subject is a child, the process proceeds to S406, and if not, the process proceeds to S409. Here, a trimmed image in which the attribute of the main subject is not a child corresponds to an image that is not suitable for the user's use in this embodiment. In the present embodiment, the predetermined attribute is a child when the subject is classified into adults and children. However, the predetermined attribute may also be an adult. Furthermore, the predetermined attribute may be, for example, a male or female when the subject is classified by gender, or a person wearing any clothes or not wearing any clothes when the subject is classified into a person wearing any uniform. In this way, the predetermined attribute, the method for determining the attribute, the design of the machine learning model, and the combination of models can be changed as appropriate depending on the user's purpose.

[0035] <Detection of unwanted areas> In S406, the region detection unit 214 detects unnecessary regions (unnecessary regions) included in the trimmed image in which a subject with a predetermined attribute has been detected. In this embodiment, the unnecessary regions are regions that are unnecessary as elements constituting the image, such as parts of the body other than the main subject, such as the arms or back, or tableware placed on a desk in front of the main subject. In this embodiment, the region detection unit 214 detects unnecessary regions included in the trimmed image in which the attribute of the main subject is a child. Here, the image for detecting the unnecessary region is a pre-trimmed image with the face region at its center. Therefore, depending on the positional relationship between the unnecessary region and the main subject, all or part of the unnecessary region can be deleted from the image by trimming in S403. Furthermore, by detecting the unnecessary region in the trimmed image with the face region at its center, it is possible to prevent an image that is suitable for the user's purpose from being excessively determined to be inappropriate.

[0036] In S407, the inappropriate image determination unit 215 sorts out inappropriate images based on the unnecessary region detected in S406. In this embodiment, if an unnecessary region exists in the trimmed image, more specifically, if the unnecessary region exists in the trimmed image over a certain area (above a threshold) or if the unnecessary region hides the main subject, the inappropriate image determination unit 215 proceeds to S409. On the other hand, in this embodiment, if an unnecessary region does not exist in the trimmed image, more specifically, if the unnecessary region does not exist over a certain area or above in the trimmed image and the unnecessary region does not hide the main subject, the inappropriate image determination unit 215 proceeds to S408. In S408, the inappropriate image determination unit 215 determines that the trimmed image is not an inappropriate image.

[0037] A method for determining whether a trimmed image is inappropriate can be to use segmentation technology using a deep learning model or the like to separate and recognize the main subject area from unnecessary areas such as the arms and back, and determine whether the unnecessary areas are detected as being larger than a threshold. Alternatively, an image recognition model that performs binary judgment using the trimmed image as input can directly determine whether the image is inappropriate. The image recognition model that performs binary judgment can perform judgment separately: a model that judges body parts other than the subject, such as the arms and back, as unnecessary areas, and a model that judges tableware placed on a table in front of the main subject as unnecessary areas. Alternatively, a model that judges two unnecessary areas simultaneously can be created. Alternatively, machine learning technology can be used to input the trimmed image into a neural network model to determine unnecessary areas.

[0038] <Inappropriate image judgment> In S409, the inappropriate image determination unit 215 determines, as an inappropriate image, the target image acquired in S302 or the image processed in S403 (trimmed image). Here, images determined as inappropriate images include images in which a face area was not detected in S402, images in which it was determined in S405 that the main subject does not have a predetermined attribute, and images in which it was determined in S407 that an unnecessary area of ​​a certain size or more exists. By the above-described inappropriate image selection process, the target image acquired in S302 or the trimmed image processed in S403 can be sorted into inappropriate images and non-inappropriate images.

[0039] Returning to the flowchart of FIG. 3, the explanation will be continued. In S304, the inappropriate image selection unit 210 determines whether the target image or the trimmed image is determined to be an inappropriate image. If the target image or the trimmed image is determined to be an inappropriate image, the process proceeds to S305, and if it is determined not to be an inappropriate image, the process proceeds to S306. In S305, the inappropriate image selection unit 210 classifies the target image or the cropped image as an inappropriate image. In this embodiment, the inappropriate image selection unit 210 adds the target image or the cropped image to the inappropriate image list by storing the file name or file path of the target image or the cropped image in RAM 102. Note that if the cropped image has not been overwritten on the target image in S403, the file names or file paths of both the target image and the cropped image are stored. In S306, the inappropriate image selection unit 210 classifies the target image or the cropped image as a suitable candidate image. In this embodiment, the inappropriate image selection unit 210 adds the target image or the cropped image to the suitable candidate image list by storing the file name or file path of the target image or the cropped image in RAM 102. Note that if the target image is not overwritten with the cropped image in S403, only the file name or file path of the cropped image can be stored. The reason why an image that is not determined to be an inappropriate image is not stored as a suitable image will be explained in S308 below.

[0040] In S307, the inappropriate image selection unit 210 determines whether the above-described processing has been performed on all images in the image data set acquired in S301. If there are images that have not been processed, the process returns to S302 and continues processing on the unprocessed images. On the other hand, if the above-described processing has been performed on all images in the image data set, the process proceeds to S308.

[0041] <Selecting appropriate images> In S308, the suitable image selection unit 220 executes a suitable image selection process to select suitable images to be finally presented to the user from the image data set classified as suitable candidate images in S306. Here, all images not determined as inappropriate by the inappropriate image selection unit 210 may be presented to the user, but this may result in a large number of images being presented to the user, or conversely, in a case where almost no images can be presented to the user. For example, there may be a large number of images to be selected, and even after removing the inappropriate images, there may still be a large number of images, or there may be almost no images determined as inappropriate. In this case, it may be difficult to present images suitable for the user's purpose based solely on the selection results by the inappropriate image selection unit 210. Therefore, in this embodiment, the appropriate image selection unit 220 selects images to be finally presented to the user from images not classified as inappropriate images, thereby making it possible to present images suitable for the user's purpose.

[0042] In the case of an application such as this embodiment in which everyday photographs of a kindergarten or the like are sold to parents, it is preferable to select images presented to the user without bias in composition, scene, or person as subject. Therefore, by grouping suitable candidate images by composition, scene, or person as subject, and presenting an appropriate number of images from the group to the user as suitable images, the effectiveness of suitable image selection can be improved.

[0043] FIG. 5 is a flowchart showing the suitable image selection process. In S501, the suitable image selection unit 220 acquires the image data set classified as suitable candidate images in S306. The grouping unit 221 performs the processes of S502 to S509 on the acquired image data set, thereby classifying similar images into the same group.

[0044] <Similar image determination> In S502, the grouping unit 221 acquires, from the image data set acquired in S501, image A having the oldest shooting time and image B, which was shot next to image A among images shot with the same imaging device as image A. Information on the shooting time and information on the imaging device are recorded in association with the images.

[0045] In S503, the grouping unit 221 determines whether the two images acquired in S502 are similar in composition, color tone, etc. As a method for determining whether two images are similar, feature points may be calculated from the frequency components of each image, and images may be determined to be similar if the sum of the distances between the feature points is equal to or less than a threshold, or similarity may be determined from the spectrogram of each hue. Alternatively, feature vectors of each image may be calculated using a machine learning model, and similarity may be determined from the distance between each feature vector.

[0046] In S504, the grouping unit 221 determines whether the two images are similar images based on the result of the determination in S503. If the images are determined to be similar, the process proceeds to S505, and if the images are not similar, the process proceeds to S506. In S505, the grouping unit 221 stores image B as an image in the similar image group to which image A belongs. In S506, the grouping unit 221 creates a new similar image group separate from the similar image group to which image A belongs, and stores image B as an image in the created similar image group.

[0047] In S507, the grouping unit 221 determines whether or not there is an image in the image data set acquired in S501 that was captured by the same imaging device as image B and at a later capture time than image B. If there is a corresponding image, the process proceeds to S508, and if there is no corresponding image, the process proceeds to S509.

[0048] In S508, the grouping unit 221 treats the image treated as image B as image A' for the next processing, acquires image C captured after image B as image B' for the next processing, and returns to S503. Therefore, the grouping unit 221 creates similar image groups by comparing images captured by the same imaging device that are adjacent when arranged in the order in which they were captured. For example, if images A and B are determined to be similar images and then images B and C are determined to be similar images, images A, B, and C belong to the same similar image group. On the other hand, if images B and C are determined not to be similar images, images A and B are stored as one similar image group, and image C is stored as a separate similar image group. Thereafter, similarity between image D, which was captured after image C, is determined.

[0049] In S509, the grouping unit 221 determines whether the above-described process has been performed on all of the image data sets of suitable candidate images acquired in S501. If the result of the determination is that unprocessed images remain, the process returns to S502, and the image with the oldest shooting time is acquired from among the unprocessed images as image A, and the process continues. On the other hand, if the process has been performed on all suitable candidate images, the process proceeds to S510.

[0050] In the present embodiment, a method has been described in which adjacent images captured by the same imaging device are compared when arranged in the order in which they were captured, but the present invention is not limited to this. For example, if it is determined in S503 that images A and B are similar images, when the process returns to S502, image C, which was captured after image B, is treated as image B', and it is determined whether images A and B' (image C) are similar images. If images A and C are determined to be similar images, image D, which was captured after image C, is similarly treated as image B' and it is determined whether they are similar images. On the other hand, if it is determined that images A and C are not similar images, image C may be treated as image A' and image D as image B', and similar images may be determined, and the grouping process may proceed.

[0051] In addition, in this embodiment, a method for creating an image group for each image capture device has been described, but this is not limited to this, and similar image groups may be created by determining the similarity of all images captured by different image capture devices in the order in which they were taken as described above. Alternatively, similar image groups may be created by determining the similarity in the character string order of the file names.

[0052] <Selecting an appropriate image> In S510, the suitable image determination unit 222 selects a specified number of images for each similar image group created by the grouping unit 221 and stores them as suitable images. Here, the number of images selected as suitable images may be a predetermined number from the similar image group (for example, 1 image), or a predetermined proportion from the similar image group (for example, 30%). The number or proportion to be selected can be set in advance by an administrator of the image selection device 110, or can be set by the suitable image determination unit 222 itself.

[0053] A method for selecting a specified number of images from an image group involves estimating the facial expression of the main subject using a machine learning model that determines facial expressions such as smiling and crying for the face area stored in RAM 102, and selecting images that are not expressionless. Furthermore, if values ​​for the degree of smile and eye openness at the time of capture are recorded in association with the images, a selection score may be calculated so that images with a high degree of smile and open eyes receive a higher score, and images with high scores may be selected. Alternatively, the image quality may be estimated from the ISO sensitivity and shutter speed recorded in the images, and images with good image quality may be selected as suitable images. Images determined to be suitable images are added to a suitable image list and stored in RAM 102.

[0054] In S511, the suitable image determination unit 222 determines whether or not the processing of S510 has been performed on all of the similar image groups created in S506. If there are similar image groups that have not been processed, the process returns to S510 and continues processing on the unprocessed similar image groups. On the other hand, if the processing of S510 has been performed on all similar image groups, the suitable image selection processing of S308 ends.

[0055] Finally, the suitable image selection unit 220 may move the images in the suitable image list stored as suitable images in the suitable image selection process of S308 to a folder designated by the user and present them to the user. Alternatively, the images may be automatically transferred from the network I / F 105 to the external device 120 via the network 130. Alternatively, the images stored as inappropriate images in S305 may be automatically deleted from the ROM 103 or the storage device 104.

[0056] As described above, according to this embodiment, an image is processed based on a face area detected from the image, a subject with a predetermined attribute is detected based on the processed image, and an image is selected based on an unnecessary area detected from the processed image in which the subject with the predetermined attribute was detected. In this way, by detecting a subject with a predetermined attribute based on the processed image, it is possible to improve detection accuracy. Furthermore, by detecting an unnecessary area in the processed image, all or part of the unnecessary area can be deleted from the image in advance, so that images suitable for the user's purpose can be efficiently selected.

[0057] (Variation 1) In the flowchart of Figure 5 described above, steps S502 to S509 are used to create similar image groups based on similarities in composition, color tone, etc. However, instead of this process, image groups may be created for each individual subject using personal authentication technology. The suitable image selection unit 220 receives, for example, representative images of individuals to be photographed from the external device 120, and stores them in the RAM 102, the ROM 103, or the storage device 104. At this time, there is no limit to the number of people to be stored, as long as there is one or more representative images.

[0058] The grouping unit 221 uses personal authentication technology to determine whether the face area data detected by the face detection unit 211 is the same person as the representative image. If it is determined that the person is the same as the representative image, the grouping unit 221 adds the image to the personal image group of that person. If the image is not the same as the representative image, a new personal image group is created and stored as an image in that personal image group. At this time, the face area and its surrounding image are stored as the representative image of the new person. Here, the personal authentication technology may be of a type that uses a machine learning model that takes the image as input to calculate a facial feature vector and determines whether the person is the same based on the feature vector. Alternatively, it may be of a type that determines whether the face area data and the representative image data are the same person, and performs the same determination for all registered people.

[0059] Another method for creating image groups for individual subjects is to use a machine learning model that uses the image as input to calculate feature vectors for the face region and record them in association with the image. Similar processing is performed on other images, and once feature vectors have been obtained from all face regions, images in which the distance between the feature vectors of the face region is less than a threshold may be considered to be images in which the same person appears, creating a personal image group.

[0060] Furthermore, for each personal image group created by the above-described processing, similar images may be determined by the processing from S502 to S509, and a similar image group may be created for each individual subject. After that, the appropriate image determination unit 222 selects a specified number of images from the image group created by the above-described processing through the processing from S510 onwards, and stores them as appropriate images.

[0061] (Variation 2) 4, if the main subject has a predetermined attribute in S405, the process proceeds to S406, and an unnecessary region is detected from the cropped image in which the main subject with the predetermined attribute has been detected. However, this is not the only case. If the image data set contains few images that include a subject with the predetermined attribute, it is preferable not to determine the image as inappropriate even if it contains a portion of an unnecessary region.

[0062] Therefore, the area detection unit 214 may detect unnecessary areas later according to the number of trimmed images in which a subject with a predetermined attribute has been detected. Specifically, the inappropriate image selection unit 210 counts the number of trimmed images in which a subject with a predetermined attribute has been detected, and if the counted number is less than a predetermined number, the process proceeds to S408 without proceeding to S406. Therefore, if the number of trimmed images is small, the area detection unit 214 will not detect unnecessary areas, and it is possible to reduce the number of trimmed images determined to be inappropriate images. In this way, by detecting unnecessary areas according to the number of trimmed images in which a subject with a predetermined attribute has been detected, it is possible to prevent an excessive increase in inappropriate images. It is preferable that the area detection unit 214 performs image processing such as trimming or masking on an image that includes an unnecessary area in part.

[0063] Furthermore, in the flowchart of FIG. 5 described above, a case where a predetermined number or a predetermined ratio of images is selected from a similar image group in S510 has been described, but this is not limited to this case. If the number of images included in a similar image group is equal to or less than a predetermined number, the suitable image determination unit 222 can select all of the images included in the similar image group. Furthermore, the area detection unit 214 may detect unnecessary areas later depending on the number of images included in the similar image group. Specifically, the area detection unit 214 detects unnecessary areas when the number of images included in the similar image group is equal to or greater than a predetermined number, and does not detect unnecessary areas when the number of images included in the similar image group is less than the predetermined number. Therefore, if the number of images included in a similar image group is small, the area detection unit 214 will not detect unnecessary areas, and the number of images included in the similar image group can be increased.

[0064] An example in which images imported from the imaging device 100 are stored in association with each folder according to the results of the sorting process, as described above, will be described with reference to FIG. 6 . Images imported from the imaging device 100 are stored in the imported folder 601. When the above-described sorting process is executed, each image is stored in a sorted folder 602 according to the results. Folders 603 and 604 for each imaging device are included in a lower layer of the sorted folder 602. Folders 605 and 607 for each shooting date are included in a lower layer of the folder 603 for each imaging device. Folders 609 and 607 for each shooting date are included in a lower layer of the folder 605 for the shooting date. An inappropriate image folder 609 and a similar image folder 611 are included in the inappropriate image folder 609. Images classified as inappropriate images in step S305 are stored in the similar image folder 611. Images that were added to the similar image group in step S505 but were not selected as suitable images in step S510 are stored in the similar image folder 611. Images selected as suitable image candidates are stored directly under the shooting date folder 605. By storing images in each folder according to the results of automatic selection in this way, the user can easily select the desired image from the group of images presented as suitable image candidates. It also makes it easier to check what kind of image an image is, even if it has been automatically selected as inappropriate and stored in a folder.

[0065] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a recording medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0066] The various controls described above as being performed by CPU 101 may be performed by a single piece of hardware, or the entire device may be controlled by multiple pieces of hardware (e.g., multiple processors or circuits) sharing the processing.

[0067] The present invention has been described in detail above based on its preferred embodiments, but the present invention is not limited to these specific embodiments, and various forms within the scope of the invention that do not deviate from the gist of the invention are also included in the present invention.

[0068] The disclosure of this embodiment also includes the following configuration. (Configuration 1) An information processing device, a face detection means for detecting a face area from an image; a processing means for processing the image based on the detected face area; attribute detection means for detecting a subject having a predetermined attribute based on the processed image; an area detection means for detecting an unnecessary area included in the processed image in which the subject having the predetermined attribute has been detected; a selection means for selecting the image as an inappropriate image or an appropriate image based on the detected unnecessary region; An information processing device comprising: (Configuration 2) The selecting means is 2. The information processing apparatus according to configuration 1, wherein the image is selected based on a positional relationship between the detected unnecessary region and the subject having the predetermined attribute. (Configuration 3) The selecting means is 3. The information processing device according to configuration 1 or 2, characterized in that an image in which a face area is not detected by the face detection means, and an image in which a face area is detected by the face detection means and the predetermined attribute is not detected by the attribute detection means, are sorted out as inappropriate images. (Configuration 4) The selecting means is The information processing device according to any one of configurations 1 to 3, characterized in that at least one of an image in which an unnecessary area detected by the area detection means exists, an image in which an unnecessary area detected by the area detection means exists at a level equal to or greater than a threshold, and an image in which an unnecessary area detected by the area detection means hides a subject having the predetermined attribute detected by the attribute detection means is selected as an inappropriate image. (Configuration 5) 5. The information processing apparatus according to any one of configurations 1 to 4, further comprising a grouping means for classifying the images selected by the selecting means into groups. (Configuration 6) The grouping means 6. The information processing device according to configuration 5, wherein the images are grouped by at least one of classifications based on similar images and classifications based on individual subjects included in the images. (Configuration 7) The grouping means 6. The information processing device according to configuration 5, wherein images are taken by the same imaging device and are grouped according to classification of similar images. (Configuration 8) 8. The information processing apparatus according to any one of configurations 5 to 7, further comprising a selection means for selecting a predetermined number or a predetermined ratio of images from the images grouped by the grouping means. (Configuration 9) 9. The information processing device according to configuration 8, wherein the image selected by the selection means is selected as an appropriate image. (Configuration 10) the face detection means includes a neural network model; 10. The information processing device according to any one of configurations 1 to 9, wherein the coordinates and size of a face area are output in response to inputting an image to the neural network model. (Configuration 11) The processing means is a trimming means for trimming the image around the face area based on the coordinates and size of the face area detected by the face detection means; 11. The information processing device according to any one of configurations 1 to 10, further comprising: image processing means for converting an image by contrast adjustment and color correction. (Configuration 12) the attribute detection means includes a neural network model; 12. The information processing device according to any one of configurations 1 to 11, wherein an attribute of a subject is detected in response to inputting an image into the neural network model. (Configuration 13) the region detection means includes a neural network model; 13. The information processing device according to any one of configurations 1 to 12, wherein an unnecessary region is determined in response to inputting an image into the neural network model. (Configuration 14) 14. The information processing device according to any one of configurations 1 to 13, wherein the images include images automatically captured in accordance with predetermined conditions. (Method 1) A control method for an information processing device, comprising: a face detection step of detecting a face region from the image; a processing step of processing the image based on the detected face area; an attribute detection step of detecting a subject having a predetermined attribute based on the processed image; an area detection step of detecting an unnecessary area included in the processed image in which the subject having the predetermined attribute has been detected; a sorting step of sorting the image as an unsuitable image or a suitable image based on the detected unnecessary region; 1. A method for controlling an information processing device, comprising: (Program 1) A program for causing a computer to function as each means of the information processing device described in any one of configurations 1 to 14. (Recording Medium 1) A computer-readable recording medium having recorded thereon a program for causing a computer to function as each means of the information processing device described in any one of configurations 1 to 14. [Explanation of symbols]

[0069] 100: Imaging device 110: Image selection device (information processing device) 210: Inappropriate image selection unit 211: Face detection unit 212: Image processing unit 213: Attribute detection unit 214: Area detection unit 215: Inappropriate image determination unit 220: Appropriate image selection unit 221: Grouping unit 222: Appropriate image determination unit

Claims

1. An information processing device, a detection means for detecting a region of a first object from the image; a processing means for processing the image based on the detected region of the first object; a subject detection means for detecting a first subject having a first attribute from the processed image; a selection means for detecting a second subject having a second attribute different from the first attribute from the processed image when the first subject having the first attribute is detected from the processed image, and selecting the processed image based on a detection result of the second subject having the second attribute; 13. An information processing device comprising:

2. The information processing device according to claim 1, characterized in that the processed image is selected based on the presence or absence of the second subject having the second attribute in the processed image.

3. The information processing device of claim 1, characterized in that when a second subject having the second attribute is detected from the processed image, the second subject is selected based on the positional relationship between the area of ​​the second subject having the second attribute and the area of ​​the first subject having the first attribute, and the size of the area of ​​the second subject.

4. The information processing device described in claim 1, characterized in that the first object is a person's face.

5. The information processing device described in claim 1, characterized in that the first attribute is a child.

6. The information processing device described in claim 1, characterized in that the second attribute is an obstruction that hides the subject.

7. An information processing device as described in claim 1, characterized in that the processing by said processing means is trimming of the image.

8. The information processing device according to claim 1, characterized in that the processed image is sorted as a suitable image or an inappropriate image.

9. The information processing device according to claim 8, characterized in that the processed image is stored in a folder corresponding to the result of being selected as the appropriate image or the inappropriate image.

10. The information processing device described in Claim 8, characterized in that among the processed images, images in which the first object area is not detected, and images in which the first object area is detected and the first subject having the first attribute is not detected are selected as inappropriate images.

11. The information processing device of claim 8, characterized in that at least any of the following images are selected as inappropriate images: an image in which the second subject having the second attribute is present; an image in which the area of ​​the second subject having the second attribute is present above a threshold; and an image in which the area of ​​the second subject having the second attribute hides part of the area of ​​the first subject having the first attribute.

12. The information processing apparatus according to claim 8, further comprising a grouping means for classifying the processed images selected as the suitable images into groups.

13. In response to inputting the image into a neural network model, coordinates and a size of the first object region are output; 2. The information processing apparatus according to claim 1, wherein said processing means processes the image by cropping the image based on the output coordinates and size.

14. The information processing device according to claim 13, characterized in that the processing means processes the cropped image by converting it through contrast adjustment and color correction.

15. The information processing device of claim 1, characterized in that the first subject having the first attribute or the second subject having the second attribute is detected by inputting the processed image into a neural network model.

16. 16. The information processing apparatus according to claim 1, wherein the image includes an image that is automatically captured in accordance with a predetermined condition.

17. A method for controlling an information processing device, comprising: a detection step of detecting a region of a first object from the image; a processing step of processing the image based on the detected region of the first object; a subject detection step of detecting a first subject having a first attribute from the processed image; a sorting step of detecting a second subject having a second attribute different from the first attribute from the processed image when the first subject having the first attribute is detected from the processed image, and sorting the processed image based on a detection result of the second subject having the second attribute; 13. A method for controlling an information processing apparatus comprising the steps of:

18. A program for causing a computer to function as each of the means of the information processing device according to claim 1.

19. 2. A computer-readable recording medium having recorded thereon a program for causing a computer to function as each of the means of the information processing apparatus according to claim 1.