Information processing device, information processing method, and information processing program
The information processing device addresses mosaic processing challenges by setting size limits and adjusting intensity for each face region, ensuring effective deletion of personal information and maintaining face recognition in images with multiple individuals.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK CORPORATION
- Filing Date
- 2025-01-16
- Publication Date
- 2026-07-29
AI Technical Summary
Existing image processing techniques struggle to perform mosaic processing on images with multiple individuals, often resulting in either overly strong or weak mosaic intensities that either obscure or fail to delete personal information effectively.
An information processing device that detects face regions, sets lower and upper limits for face image resizing, and adjusts mosaic intensity based on these limits to ensure appropriate processing for each face region, using a constant reduction ratio and size constraints to generate mosaic images.
The device achieves balanced mosaic processing that effectively deletes personal information while maintaining recognizable human faces, preventing stretching or blurring issues, thus enhancing privacy protection.
Smart Images

Figure 2026122735000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program.
Background Art
[0002] Conventionally, a technique for deleting personal information reflected in an image has been known. For example, a person area, which is an area in an image captured by a camera device where a person appears, is detected, and privacy processing with different intensities is performed on the person area according to the depth or a predetermined index related to the depth associated with the coordinates of the person area.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Means for Solving the Problems
[0004] The information processing apparatus according to the present application includes: a detection unit that detects each of a plurality of face regions including the face of each of a plurality of persons from an image; a determination unit that determines whether or not the size of a first reduced face image, which is an image of the face region reduced at a predetermined reduction rate, is equal to or greater than a lower limit size and equal to or less than an upper limit size; and a generation unit that, when the determination unit determines that the size of the first reduced face image is less than the lower limit size, resizes the size of the first reduced face image to the lower limit size, and when the determination unit determines that the size of the first reduced face image is greater than the upper limit size, resizes the size of the first reduced face image to the upper limit size, and generates a first enlarged face image obtained by enlarging the resized first reduced face image to the size of the image of the original face region.
[0005] Furthermore, the information processing device according to the present application includes: a detection unit that detects each of a plurality of face regions, each containing the face of a plurality of persons, from an image; a generation unit that generates a plurality of enlarged face images by enlarging each of the plurality of reduced face images, which are images of the face regions reduced to each of a plurality of predetermined sizes, to the size of the original face region image; a determination unit that determines whether the similarity between the image feature quantities of each of the plurality of enlarged face images and the image feature quantities of the original face region image exceeds a predetermined similarity; and, if the determination unit determines that the similarity exceeds the predetermined similarity, an acquisition unit that acquires the enlarged face image corresponding to the predetermined size for which the similarity has been determined to exceed the predetermined similarity as a mosaic image. [Brief explanation of the drawing]
[0006] [Figure 1] Figure 1 shows an image after applying mosaic processing according to the first comparative technique. [Figure 2] Figure 2 shows an image after applying mosaic processing according to the second comparative technique. [Figure 3] Figure 3 shows an image after performing mosaic processing according to the first embodiment. [Figure 4] Figure 4 is a diagram illustrating the overview of the mosaic processing. [Figure 5] Figure 5 is a diagram illustrating the reduction process of mosaic processing (Pattern 1). [Figure 6] Figure 6 is a diagram illustrating the reduction process of mosaic processing (pattern 2). [Figure 7] Figure 7 illustrates the case where a lower limit for the size of the reduced face image is not set during the mosaic reduction process (pattern 2). [Figure 8] Figure 8 is a diagram illustrating the setting of the lower limit value for the size of the reduced face image in the reduction process (pattern 2) of the mosaic processing according to the first embodiment. [Figure 9] Figure 9 illustrates the case where no upper limit is set for the size of the reduced face image during the mosaic reduction process (pattern 2). [Figure 10] Figure 10 is a diagram illustrating the setting of the upper limit value for the size of the reduced face image in the reduction process (pattern 2) of the mosaic processing according to the first embodiment. [Figure 11] Figure 11 shows an example of the configuration of an information processing system according to the first embodiment. [Figure 12] Figure 12 shows an example of the configuration of an information processing device according to the first embodiment. [Figure 13] Figure 13 is a flowchart showing an example of the processing procedure of the information processing apparatus according to the first embodiment. [Figure 14] Figure 14 shows an example of the configuration of an information processing device according to the second embodiment. [Figure 15] Figure 15 is a diagram illustrating the overview of the processing of the information processing apparatus according to the second embodiment. [Figure 16] Figure 16 is a diagram illustrating the overview of the processing of the information processing apparatus according to the second embodiment. [Figure 17] Figure 17 is a flowchart showing an example of the processing procedure of the information processing device according to the second embodiment. [Figure 18] Figure 18 is a flowchart showing an example of the processing procedure of the information processing device according to the second embodiment. [Figure 19] Figure 19 shows an example of the configuration of an information processing device according to the third embodiment. [Figure 20] Figure 20 is a flowchart showing an example of the processing procedure of the information processing apparatus according to the third embodiment. [Figure 21] Figure 21 shows an example of a hardware configuration. [Modes for carrying out the invention]
[0007] Hereinafter, embodiments for implementing the information processing apparatus, information processing method, and information processing program according to the present application (hereinafter referred to as "embodiments") will be described in detail while referring to the drawings. Note that the information processing apparatus, information processing method, and information processing program according to the present application are not limited by these embodiments. Also, in the following embodiments, the same parts are denoted by the same reference numerals, and redundant explanations are omitted.
[0008] [1. Introduction] First, referring to FIGS. 1 to 10, the background for the inventors to create the first embodiment will be described.
[0009] Generally, for mosaic processing of an image in which a person is shown, it is desirable that the person shown in the image is not identified as a specific person, and that mosaic processing with an appropriate intensity is performed so that the face of the person shown in the image can be recognized as a human face. On the other hand, if mosaic processing is performed on an image in which a plurality of persons are shown with a uniform mosaic intensity, there may be cases where the mosaic intensity is too strong and the face is crushed, or the mosaic intensity is too weak and personal information cannot be deleted. This point will be described using FIGS. 1 to 3.
[0010] FIGS. 1 to 3 show images G11 to G13 in which three persons U1 to U3 are shown. In FIGS. 1 to 3, person U1 is located at the forefront of the space shown in the image. In other words, person U1 is shown at the forefront of the image. Also, person U3 is located at the deepest part of the space shown in the image. In other words, person U3 is shown at the deepest part of the image. Also, person U2 is located in the space between person U1 and person U3. In other words, person U2 is shown behind person U1 (in front of person U3).
[0011] FIG. 1 is a diagram showing an image on which mosaic processing according to a first comparative technique is executed. FIG. 1 shows an image G11 on which uniform and strong mosaic processing is executed on the image in order to delete personal information of a person U1 shown at the forefront of the image. In FIG. 1, for the face region of the person U1 shown at the forefront of the image, mosaic processing with an appropriate intensity according to the size of the face region is executed. For this reason, in FIG. 1, it is possible to recognize the face of the person U1 as a human face, and personal information of the person U1 has been successfully deleted so that the face of the person U1 is not recognized as the face of a specific person. However, in FIG. 1, for each of the face regions M12 and M13 of the persons U2 and U3 shown in the back of the image, strong intensity (i.e., coarse) mosaic processing that is not suitable for the sizes of the face regions M12 and M13 is executed. For this reason, in FIG. 1, the faces of the persons U2 and U3 shown in the back of the image are blurred. In other words, in FIG. 1, the faces of the persons U2 and U3 shown in the back of the image cannot be recognized as human faces.
[0012] FIG. 2 is a diagram showing an image on which mosaic processing according to a second comparative technique is executed. FIG. 2 shows an image G12 on which uniform and weak mosaic processing is executed on the image so that the faces of the persons U2 and U3 shown in the back of the image can be recognized as human faces and the faces of the persons U2 and U3 are not recognized as the faces of specific persons. In FIG. 2, for the face regions of the persons U2 and U3 shown in the back of the image, mosaic processing with an appropriate intensity according to the size of the face region is executed. For this reason, in FIG. 2, it is possible to recognize the faces of the persons U2 and U3 as human faces, and personal information of the persons U2 and U3 has been successfully deleted so that the faces of the persons U2 and U3 are not recognized as the faces of specific persons. However, in FIG. 2, for the face region M21 of the person U1 shown in the front of the image, weak intensity (i.e., smooth) mosaic processing that is not suitable for the size of the face region M21 is executed. For this reason, in FIG. 2, the face of the person U1 shown in the front of the image can be recognized as the face of a specific person. In other words, in FIG. 2, the personal information of the person U1 shown in the front cannot be deleted.
[0013] Figure 3 shows an image after mosaic processing according to the first embodiment. The mosaic processing according to the first embodiment is performed by the information processing device 100 according to the first embodiment. Figure 3 shows image G13 in which the information processing device 100 has performed mosaic processing of appropriate intensity according to the respective sizes of the face regions M31 to M33 of each person U1 to U3. In Figure 3, the information processing device 100 performs stronger mosaic processing on the face region M31 of person U1, which is in the foreground of the image. As a result, the information processing device 100 is able to recognize the face of person U1 in the foreground as a human face and successfully removes the personal information of person U1. Furthermore, the information processing device 100 performs weaker mosaic processing on the face regions M32 and M33 of people U2 and U3, which are in the background of the image, than on the face region M31 of person U1. As a result, the information processing device 100 is able to recognize the faces of people U2 and U3 in the background as human faces and successfully removes the personal information of people U2 and U3.
[0014] Figure 4 is a diagram illustrating the overview of mosaic processing. Figure 4 illustrates the general overview of mosaic processing applied to the face region of a person detected in an image. Generally, mosaic processing applied to the face region of a person is performed by reducing the size of the face region image (Step 1) and then enlarging the reduced image to the size of the original face region image (Step 2). For example, the information processing device 100 acquires an image G21 of the face region of a person contained in the image. Next, the information processing device 100 generates a reduced face image G22 by reducing the image G21 of the face region of the person (Step 1: reduction processing). Next, the information processing device 100 generates an enlarged face image G23 by enlarging the reduced face image G22 to the size of the original face region image G21 (Step 2: enlargement processing). In this way, the information processing device 100 acquires the enlarged face image G23 as an image (hereinafter referred to as a mosaic image) on which mosaic processing has been applied to the face region image G21.
[0015] Figure 5 is a diagram illustrating the mosaic reduction process (pattern 1). Figure 5 illustrates an example of the mosaic reduction process described in Figure 4. In Figure 5, the information processing device 100 reduces the image G31 of a person's face to a predetermined size. In the following, the original image of a person's face may be referred to as the input face image. In other words, the information processing device 100 generates a reduced face image G32 by reducing the image G31 of a person's face to a predetermined size. In other words, the information processing device 100 generates a reduced face image G32 by reducing the image G31 of a person's face to a predetermined width and height. For example, the information processing device 100 generates a reduced face image G32 by reducing the image G31 of a person's face to a predetermined width of 30 x height of 40 pixels. Note that 30 x 40 pixels is just one example of a predetermined width and height, and the predetermined width and height are not limited to 30 x 40 pixels. For example, the predetermined width and height may be larger than 30 x 40 pixels, or smaller than 30 x 40 pixels.
[0016] Figure 6 is a diagram illustrating the mosaic reduction process (pattern 2). Figure 6 illustrates a different example of the mosaic reduction process described in Figure 4, as shown in Figure 5. In Figure 6, the information processing device 100 reduces the image G33 of the person's face region by a constant reduction ratio relative to the size of the input face image. In other words, the information processing device 100 generates a reduced face image G34 by reducing the image G33 of the person's face region by a constant reduction ratio relative to the size of the input face image. For example, the information processing device 100 generates a reduced face image G34 with a width of 130 x height of 150 pixels by reducing the image G33 of the person's face region, which is 260 x 300 pixels wide, to 0.5 times its original size. Note that 0.5 is just one example of a constant reduction ratio, and the constant reduction ratio is not limited to 0.5. For example, the constant reduction ratio may be greater than 0.5 or less than 0.5.
[0017] As explained in Figures 5 and 6, there are two types of reduction processing in mosaic processing: reduction processing that reduces the input face image to a predetermined size (Pattern 1), and reduction processing that reduces the input face image at a constant reduction ratio (Pattern 2). In the case of reduction processing in Pattern 1, if the size of the input face image is large, the size of the reduced face image will be relatively quite small compared to the size of the input face image. Therefore, when the reduced face image reduced by the reduction processing in Pattern 1 is enlarged to the size of the input face image, the reduced face image is stretched considerably relative to the size of the reduced face image. Furthermore, if an enlarged face image is generated in which the reduced face image is stretched considerably relative to the size of the reduced face image, a mosaic processing of a strong intensity (i.e., coarse) that is not suitable for the size of the input face image will be performed. For this reason, in the case of reduction processing in Pattern 1, depending on the size of the input face image, a mosaic processing of a strong intensity (i.e., coarse) that is not suitable for the size of the input face image may be performed. In contrast, in the case of the reduction process of Pattern 2, compared to the reduction process of Pattern 1, the image is reduced at a constant reduction ratio relative to the size of the input face image, making it less likely to perform a strong (i.e., coarse) mosaic process that is unsuitable for the size of the input face image. For this reason, in the mosaic process according to the first embodiment, the reduction process (Pattern 2) that reduces the image at a constant reduction ratio relative to the size of the input face image is used as the basis of the logic.
[0018] As described above, the mosaic processing according to the first embodiment is based on a reduction process (pattern 2) that reduces the input face image by a constant reduction ratio. However, in the reduction process (pattern 2) that reduces the input face image by a constant reduction ratio, if a lower limit is not set for the size of the reduced face image obtained by reducing the input face image, a mosaic processing of a strong intensity (i.e., coarse) that is unsuitable for the size of the input face image may be performed. Furthermore, as a result of performing a mosaic processing of a strong intensity (i.e., coarse) that is unsuitable for the size of the input face image, it may become impossible to recognize the faces in the mosaic-processed face image (enlarged face image) as human faces.
[0019] Figure 7 illustrates the case where a lower limit for the size of the reduced face image is not set in the mosaic reduction process (pattern 2). Figure 7 is a diagram illustrating the case where a lower limit for the size of the reduced face image is not set in the mosaic reduction process (pattern 2). Figure 7 shows an input face image G41 with a width of 138 × a height of 153 pixels. It also shows a reduced face image G42 with a width of 4 × a height of 4 pixels, obtained by reducing the input face image G41 by a constant reduction ratio of 0.03. Furthermore, it shows an enlarged face image G43, obtained by enlarging the reduced face image G42 to the size of the input face image G41 with a width of 138 × a height of 153 pixels.
[0020] In Figure 7, since no lower limit is set for the size of the reduced face image, the size of the reduced face image G42 is relatively quite small (for example, less than 10 pixels in width or height). Therefore, when the reduced face image G42 is enlarged to the size of the input face image G41, an enlarged face image G43 is generated, which is a greatly stretched version of the reduced face image G42. Furthermore, when an enlarged face image G43 is generated, a strong (i.e., coarse) mosaic processing that is unsuitable for the size of the input face image is performed. For this reason, in the reduction process (Pattern 2) which reduces the input face image by a constant reduction ratio, it is desirable to set a lower limit for the size of the reduced face image obtained by reducing the input face image.
[0021] Using Figure 8, the setting of the lower limit value for the size of the reduced face image in the mosaic reduction process (pattern 2) according to the first embodiment will be explained. Figure 8 is a diagram for explaining the setting of the lower limit value for the size of the reduced face image in the mosaic reduction process (pattern 2) according to the first embodiment. In Figure 8, the information processing device 100 acquires an input face image G51 with a width of 138 × a height of 153 pixels. The information processing device 100 also determines whether the size of the reduced face image obtained by reducing the input face image G51 by a certain reduction ratio of 0.03 times is greater than or equal to the lower limit size. In Figure 8, the lower limit size is 20 × 20 pixels. Note that 20 × 20 pixels is just one example of a lower limit size, and the lower limit size is not limited to 20 × 20 pixels. For example, the lower limit size may be larger than 20 × 20 pixels, or smaller than 20 × 20 pixels.
[0022] In Figure 8, the information processing device 100 determines that the size of the reduced face image obtained by reducing the input face image G51 by a constant reduction ratio of 0.03 times falls below the lower limit size. Furthermore, if the information processing device 100 determines that the size of the reduced face image falls below the lower limit size, it resizes the reduced face image to the lower limit size. Specifically, if the information processing device 100 determines that the size of the reduced face image falls below the lower limit size, it reduces the size of the input face image G51 to the lower limit size. In other words, the information processing device 100 generates a reduced face image G52 by reducing the size of the input face image G51 to the lower limit size. Subsequently, the information processing device 100 generates an enlarged face image G53 by enlarging the reduced face image G52 to the size of the input face image G51.
[0023] In Figure 8, the information processing device 100 has a lower limit set for the size of the reduced face image, which prevents the size of the reduced face image G52 from becoming relatively too small (for example, less than 10 pixels in width or less than 10 pixels in height). In other words, because the information processing device 100 has a lower limit set for the size of the reduced face image, it can reduce the size of the reduced face image G52 relative to the size of the input face image G51 to an appropriate size (for example, 10 pixels or more in width and 10 pixels or more in height). Therefore, when the information processing device 100 enlarges the reduced face image G52 to the size of the input face image G51, it can prevent the generation of an enlarged face image G53 in which the reduced face image G52 is greatly stretched relative to the size of the reduced face image G52. In other words, the information processing device 100 can generate an enlarged face image G53 in which the reduced face image G52 is enlarged to an appropriate size relative to the size of the reduced face image G52. Furthermore, the information processing device 100 can prevent the generation of an enlarged face image G53 that is significantly stretched relative to the size of the reduced face image G52, thus preventing the execution of a mosaic process with a strong intensity (i.e., coarse) that is unsuitable for the size of the input face image. In other words, the information processing device 100 can generate an enlarged face image G53 that is appropriately enlarged relative to the size of the reduced face image G52, thus enabling the execution of a mosaic process with an appropriate intensity suitable for the size of the input face image.
[0024] Furthermore, as described above, the mosaic processing according to the first embodiment is based on a reduction process (pattern 2) that reduces the input face image by a certain reduction ratio. Also, as described above, the mosaic processing according to the first embodiment is based on a lower limit for the size of the reduced face image. However, if an upper limit is not set for the size of the reduced face image obtained by reducing the input face image, a mosaic processing of a weak intensity (i.e., smooth) that is unsuitable for the size of the input face image will be performed. As a result of performing a mosaic processing of a weak intensity (i.e., smooth) that is unsuitable for the size of the input face image, the face in the mosaic-processed face image (enlarged face image) may be recognized as the face of a specific person (i.e., personal information cannot be deleted).
[0025] Figure 9 illustrates the case where the upper limit for the size of the reduced face image is not set in the mosaic reduction process (pattern 2). Figure 9 is a diagram illustrating the case where the upper limit for the size of the reduced face image is not set in the mosaic reduction process (pattern 2). Figure 9 shows an input face image G61 with a width of 2295 × a height of 3236 pixels. It also shows a reduced face image G62 with a width of 69 × a height of 97 pixels, obtained by reducing the input face image G61 by a constant reduction ratio of 0.03. Furthermore, it shows an enlarged face image G63, obtained by enlarging the reduced face image G62 to the size of the input face image G61 with a width of 2295 × a height of 3236 pixels.
[0026] In Figure 9, since there is no upper limit set for the size of the reduced face image, the size of the reduced face image G62 is relatively quite large (for example, more than 50 pixels in width or more than 50 pixels in height). Therefore, when the reduced face image G62 is enlarged to the size of the input face image G61, an enlarged face image G63 is generated, which is a smoothly enlarged version of the reduced face image G62. Furthermore, when an enlarged face image G63 is generated, a mosaic process with a weak intensity (i.e., smooth) that is not suitable for the size of the input face image is performed. For this reason, in the reduction process (pattern 2) which reduces the input face image by a constant reduction ratio, it is desirable to set an upper limit on the size of the reduced face image obtained by reducing the input face image.
[0027] Using Figure 10, the setting of the upper limit value for the size of the reduced face image in the mosaic reduction process (pattern 2) according to the first embodiment will be explained. Figure 10 is a diagram for explaining the setting of the upper limit value for the size of the reduced face image in the mosaic reduction process (pattern 2) according to the first embodiment. In Figure 10, the information processing device 100 acquires an input face image G71 with a width of 2295 × a height of 3236 pixels. The information processing device 100 also determines whether the size of the reduced face image obtained by reducing the input face image G71 by a certain reduction ratio of 0.03 times is less than or equal to the upper limit size. In Figure 10, the upper limit size is 30 × 30 pixels. Note that 30 × 30 pixels is just one example of an upper limit size, and the upper limit size is not limited to 30 × 30 pixels. For example, the upper limit size may be larger than 30 × 30 pixels, or smaller than 30 × 30 pixels.
[0028] In Figure 10, the information processing device 100 determines that the size of the reduced face image obtained by reducing the input face image G71 by a constant reduction ratio of 0.03 times exceeds the upper limit size. Furthermore, if the information processing device 100 determines that the size of the reduced face image exceeds the upper limit size, it resizes the reduced face image to the upper limit size. Specifically, if the information processing device 100 determines that the size of the reduced face image exceeds the upper limit size, it reduces the size of the input face image G71 to the upper limit size. In other words, the information processing device 100 generates a reduced face image G72 by reducing the size of the input face image G71 to the upper limit size. Subsequently, the information processing device 100 generates an enlarged face image G73 by enlarging the reduced face image G72 to the size of the input face image G71.
[0029] In Figure 10, the information processing device 100 has an upper limit on the size of the reduced face image, which prevents the size of the reduced face image G72 from becoming relatively large (for example, more than 50 pixels in width or more than 50 pixels in height). In other words, because the information processing device 100 has an upper limit on the size of the reduced face image, it can reduce the size of the reduced face image G72, which is reduced relative to the size of the input face image G71, to an appropriate size (for example, less than 50 pixels in width and less than 50 pixels in height). Therefore, when the information processing device 100 enlarges the reduced face image G72 to the size of the input face image G71, it can prevent the generation of an enlarged face image G73 that is smoothly enlarged relative to the size of the reduced face image G72. In other words, the information processing device 100 can generate an enlarged face image G73 that is appropriately enlarged relative to the size of the reduced face image G72. Furthermore, the information processing device 100 can prevent the generation of an enlarged face image G73 that is smoothly enlarged relative to the size of the reduced face image G72, thus preventing the execution of a weak (i.e., smooth) mosaic process that is unsuitable for the size of the input face image. In other words, the information processing device 100 can generate an enlarged face image G73 that is appropriately enlarged relative to the size of the reduced face image G72, thus enabling the execution of a mosaic process of appropriate intensity suitable for the size of the input face image.
[0030] [2. First Embodiment] [2-1. Configuration of the Information Processing System] An example of the configuration of the information processing system 1 according to the first embodiment will be described using Figure 11. Figure 11 is a diagram showing an example of the configuration of the information processing system 1 according to the first embodiment. As shown in Figure 11, the information processing system 1 includes a user terminal 10 and an information processing device 100. The user terminal 10 and the information processing device 100 are connected to each other via a predetermined communication network (network N) by wired or wireless means.
[0031] The user terminal 10 is an information processing device used by the user. For example, the user terminal 10 may be an information processing device such as a desktop PC (Personal Computer) or a notebook PC. Alternatively, the user terminal 10 may be a smart device such as a smartphone or tablet.
[0032] The information processing device 100 is a device that performs information processing according to the first embodiment. The information processing device 100 performs information processing according to the first embodiment by executing an information processing program according to the first embodiment. Specifically, the information processing device 100 detects each of a plurality of face regions, each containing the face of a plurality of persons, from an image. The information processing device 100 also determines whether the size of the first reduced face image, which is an image of a face region reduced by a predetermined reduction ratio, is greater than or equal to a lower limit size and less than or equal to an upper limit size. If the information processing device 100 determines that the size of the first reduced face image is less than the lower limit size, it resizes the first reduced face image to the lower limit size, and if the information processing device 100 determines that the size of the first reduced face image is greater than the upper limit size, it resizes the first reduced face image to the upper limit size. The information processing device 100 also generates a first enlarged face image by enlarging the resized first reduced face image to the size of the original face region image.
[0033] [2-2. Configuration of Information Processing Device] An example of the configuration of the information processing device 100 according to the first embodiment will be described using Figure 12. Figure 12 is a diagram showing an example of the configuration of the information processing device according to the first embodiment. The information processing device 100 includes a communication unit 110, a storage unit 120, and a control unit 130.
[0034] (Communications Department 110) The communication unit 110 is implemented using a NIC (Network Interface Card), an antenna, etc. The communication unit 110 is connected to various networks by wired or wireless means, and performs information transmission and reception, for example, with the user terminal 10.
[0035] (Storage unit 120) The storage unit 120 is implemented by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or by storage devices such as hard disks and optical discs. Specifically, the storage unit 120 stores the information processing program according to the first embodiment.
[0036] (Control unit 130) The control unit 130 is a controller, and is realized, for example, by executing various programs stored in the memory device inside the information processing device 100 using RAM as the working area, using a CPU (Central Processing Unit) or MPU (Micro Processing Unit), etc. Alternatively, the control unit 130 is a controller and can be realized, for example, by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array).
[0037] The control unit 130 has a detection unit 131, a determination unit 132, and a generation unit 133 as functional units, and may realize or execute the information processing operations described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in Figure 12, and other configurations are also possible as long as they perform the information processing described later. Also, each functional unit represents the function of the control unit 130 and does not necessarily have to be physically separated.
[0038] (Detection unit 131) The detection unit 131 acquires an image. For example, the detection unit 131 acquires an image from the user terminal 10.
[0039] Furthermore, the detection unit 131 detects each of multiple face regions containing the faces of multiple people from the image. For example, when an image is input, the detection unit 131 uses a face detection model, which is a machine learning model trained to output each of multiple face regions containing the faces of multiple people in the image, to detect each of multiple face regions containing the faces of multiple people from the image. For example, the detection unit 131 inputs the acquired image into the face detection model to detect each of multiple face regions containing the faces of multiple people from the image. By detecting each of the multiple face regions, the detection unit 131 acquires each of the multiple face regions.
[0040] (Judgment unit 132) The determination unit 132 generates a first reduced face image, which is an image of the face region reduced by a predetermined reduction ratio. For example, the determination unit 132 generates a first reduced face image in which the width and height of the original face region image are each reduced by a predetermined reduction ratio. Here, the predetermined reduction ratio may be, for example, 0.5 times or 0.03 times. Note that 0.5 times or 0.03 times are just examples of predetermined reduction ratios, and the predetermined reduction ratio is not limited to 0.5 times or 0.03 times. For example, the predetermined reduction ratio may be less than 0.03 times or greater than 0.03 times. Also, the predetermined reduction ratio may be less than 0.5 times or greater than 0.5 times.
[0041] Furthermore, the determination unit 132 determines whether the size of the first reduced face image, which is an image of the face region reduced by a predetermined reduction ratio, is greater than or equal to the lower limit size and less than or equal to the upper limit size. For example, the determination unit 132 determines whether the size of the first reduced face image is greater than or equal to the lower limit size. For example, the determination unit 132 determines whether the width and height of the first reduced face image are greater than or equal to the width and height corresponding to the lower limit size. For example, if the determination unit 132 determines that the width and height of the first reduced face image are greater than or equal to the width and height corresponding to the lower limit size, it determines that the size of the first reduced face image is greater than or equal to the lower limit size. On the other hand, if the determination unit 132 determines that either the width or height of the first reduced face image is less than the width or height corresponding to the lower limit size, it determines that the size of the first reduced face image is not greater than or equal to the lower limit size (it is less than the lower limit size). Furthermore, if the determination unit 132 determines that the size of the first reduced face image is not greater than or equal to the lower limit size (i.e., less than or equal to the lower limit size), it determines that the size of the first reduced face image is greater than or equal to the lower limit size and not less than or equal to the upper limit size.
[0042] Furthermore, the determination unit 132 determines whether the size of the first reduced face image is less than or equal to the upper limit size. For example, the determination unit 132 determines whether the width and height of the first reduced face image are less than or equal to the width and height corresponding to the upper limit size. For example, if the determination unit 132 determines that the width and height of the first reduced face image are less than or equal to the width and height corresponding to the upper limit size, it determines that the size of the first reduced face image is less than or equal to the upper limit size. On the other hand, if the determination unit 132 determines that either the width or height of the first reduced face image exceeds the width or height corresponding to the upper limit size, it determines that the size of the first reduced face image is not less than or equal to the upper limit size (it exceeds the upper limit size). If the determination unit 132 determines that the size of the first reduced face image is not less than or equal to the upper limit size (it exceeds the upper limit size), it determines that the size of the first reduced face image is greater than or equal to the lower limit size and not less than or equal to the upper limit size.
[0043] (Generation unit 133) If the determination unit 132 determines that the size of the first reduced face image is greater than or equal to the lower limit size and less than or equal to the upper limit size, the generation unit 133 resizes the first reduced face image to either the lower limit size or the upper limit size. Specifically, if the determination unit 132 determines that the size of the first reduced face image is less than the lower limit size, the generation unit 133 resizes the first reduced face image to the lower limit size and generates a first enlarged face image by enlarging the resized first reduced face image to the size of the original face region image. More specifically, if the determination unit 132 determines that the size of the first reduced face image is less than the lower limit size, the generation unit 133 reduces the size of the original face region image to the lower limit size. In other words, the generation unit 133 generates a second reduced face image by reducing the size of the original face region image to the lower limit size. Subsequently, the generation unit 133 generates a first enlarged face image by enlarging the second reduced face image to the size of the original face region image.
[0044] Furthermore, if the determination unit 132 determines that the size of the first reduced face image exceeds the upper limit size, the generation unit 133 resizes the first reduced face image to the upper limit size and generates a first enlarged face image by enlarging the resized first reduced face image to the size of the original face region image. More specifically, if the determination unit 132 determines that the size of the first reduced face image exceeds the upper limit size, the generation unit 133 reduces the size of the original face region image to the upper limit size. In other words, the generation unit 133 generates a third reduced face image by reducing the size of the original face region image to the upper limit size. Subsequently, the generation unit 133 generates a first enlarged face image by enlarging the third reduced face image to the size of the original face region image.
[0045] Furthermore, if the determination unit 132 determines that the size of the first reduced face image is greater than or equal to the lower limit size and less than or equal to the upper limit size, the generation unit 133 generates a second enlarged face image by enlarging the first reduced face image to the size of the original face region image.
[0046] [2-3. Processing Procedure] An example of the processing procedure of the information processing device 100 according to the first embodiment will be explained using Figure 13. Figure 13 is a flowchart of an example of the processing procedure of the information processing device according to the first embodiment. In Figure 13, the detection unit 131 acquires an image (step S101). Subsequently, the detection unit 131 detects face regions from the image (step S102). For example, the detection unit 131 detects each of a plurality of face regions, each containing the face of a plurality of people, from the image. By detecting each of the plurality of face regions, the detection unit 131 acquires each of the plurality of face regions.
[0047] Furthermore, the determination unit 132 generates a first reduced face image, which is an image of the face region reduced by a predetermined reduction ratio (step S103). Subsequently, the determination unit 132 determines whether the size of the first reduced face image is greater than or equal to the lower limit size and less than or equal to the upper limit size (step S104).
[0048] If the determination unit 132 determines that the size of the first reduced face image is greater than or equal to the lower limit size and less than or equal to the upper limit size (step S104; No), the generation unit 133 resizes the first reduced face image to the lower limit size or the upper limit size (step S105). For example, if the determination unit 132 determines that the size of the first reduced face image is less than the lower limit size, the generation unit 133 resizes the first reduced face image to the lower limit size and generates a first enlarged face image by enlarging the resized first reduced face image to the size of the original face region image. Also, if the determination unit 132 determines that the size of the first reduced face image is greater than the upper limit size, the generation unit 133 resizes the first reduced face image to the upper limit size and generates a first enlarged face image by enlarging the resized first reduced face image to the size of the original face region image. Next, the generation unit 133 generates a first enlarged face image by enlarging the resized first reduced face image to the size of the original face region image (step S106).
[0049] On the other hand, if the determination unit 132 determines that the size of the first reduced face image is greater than or equal to the lower limit size and less than or equal to the upper limit size (step S104; Yes), the generation unit 133 generates a second enlarged face image by enlarging the first reduced face image to the size of the original face region image (step S107).
[0050] [2-4. Effects] As described above, the information processing device 100 according to the first embodiment includes a detection unit 131, a determination unit 132, and a generation unit 133. The detection unit 131 detects each of a plurality of face regions, each containing the face of a plurality of persons, from an image. The determination unit 132 determines whether the size of a first reduced face image, which is an image of a face region reduced by a predetermined reduction ratio, is greater than or equal to a lower limit size and less than or equal to an upper limit size. If the determination unit 132 determines that the size of the first reduced face image is less than the lower limit size, the generation unit 133 resizes the first reduced face image to the lower limit size. If the determination unit 132 determines that the size of the first reduced face image is greater than the upper limit size, the generation unit 133 resizes the first reduced face image to the upper limit size and generates a first enlarged face image by enlarging the resized first reduced face image to the size of the original face region image.
[0051] This allows the information processing device 100 to perform appropriate mosaic processing according to the size of each face area of multiple people included in the image. Furthermore, because the information processing device 100 can perform appropriate mosaic processing according to the size of each face area of multiple people included in the image, it can contribute to achieving Sustainable Development Goal (SDG) 9, "Build resilient infrastructure, promote inclusive and sustainable industrialization and foster innovation."
[0052] Furthermore, if the determination unit 132 determines that the size of the first reduced face image is greater than or equal to the lower limit size and less than or equal to the upper limit size, the generation unit 133 generates a second enlarged face image by enlarging the first reduced face image to the size of the original face region image.
[0053] This allows the information processing device 100 to perform appropriate mosaic processing according to the size of each face area of the multiple people included in the image.
[0054] [3. Second Embodiment] In the second embodiment, a method for automatically determining the appropriate values for the predetermined reduction ratio, lower limit size, and upper limit size according to the first embodiment will be described.
[0055] [3-1. Configuration of the Information Processing System] Using Figure 11, an example of the configuration of the information processing system 1A according to the second embodiment will be described. In this embodiment, the configuration of the information processing system 1A is the same as in the first embodiment, so you can refer to Figure 11, which shows an example of the configuration of the information processing system 1 according to the first embodiment. In the second embodiment, the information processing device 100 shown in Figure 11 is replaced by the information processing device 100A according to the second embodiment. As the user terminal 10 is the same as in the first embodiment, the description of the user terminal 10 will be omitted here.
[0056] The information processing device 100A is a device that performs information processing according to the second embodiment. The information processing device 100A performs information processing according to the second embodiment by executing an information processing program according to the second embodiment. Specifically, the information processing device 100A determines a predetermined reduction ratio based on a reference reduction ratio and the size of the original face region image. The information processing device 100A also determines a lower limit size based on the face detection inference result. The information processing device 100A also determines an upper limit size based on the similarity of image features extracted from the face region image.
[0057] [3-2. Configuration of Information Processing Device] Using Figure 14, an example of the configuration of the information processing device 100A according to the second embodiment will be described. Figure 14 is a diagram showing an example of the configuration of the information processing device 100A according to the second embodiment. The information processing device 100A includes a communication unit 110, a storage unit 120, and a control unit 130A. Note that, except for the control unit 130A, the functional units are common to those included in the information processing device 100 according to the first embodiment, so here only the control unit 130A, which is not common to the first embodiment, will be described, and the description of the other functional units will be omitted.
[0058] (Control Unit 130A) The control unit 130A is a controller, and is implemented, for example, by a CPU or MPU executing various programs stored in the memory device inside the information processing device 100A using RAM as the working area. Alternatively, the control unit 130A is a controller and can be implemented, for example, by an integrated circuit such as an ASIC or FPGA.
[0059] The control unit 130A has a detection unit 131, a determination unit 132, a generation unit 133, and a decision unit 134 as functional units, and may realize or execute the information processing operations described below. Note that the internal configuration of the control unit 130A is not limited to the configuration shown in Figure 14, and other configurations are also acceptable as long as they perform the information processing described later. Furthermore, each functional unit represents the function of the control unit 130A and does not necessarily have to be physically distinct.
[0060] The following describes each element of the control unit 130A in order, but since all elements except the determination unit 134 are the same as those in the first embodiment, their descriptions are omitted here.
[0061] (Decision Section 134) The determination unit 134 determines a predetermined reduction ratio based on a reference reduction ratio and the size of the original face region image. Specifically, the determination unit 134 determines a predetermined reduction ratio based on the size of the reference image, the reference reduction ratio, and the size of the original face region image. More specifically, the determination unit 134 determines a predetermined reduction ratio using the following formula (1). In the following formula (1), the reduction ratio基準 This corresponds to the reference reduction ratio. Also, in equation (1) below, the input image corresponds to the original face region image.
[0062]
number
[0063] The processing overview of the information processing device 100A according to the second embodiment will be explained using Figure 15. Figure 15 is a diagram for explaining the processing overview of the information processing device 100A according to the second embodiment. In Figure 15, the determination unit 134 generates multiple reduced face images by reducing the input face image to a predetermined size (pattern 1), thereby reducing the original face region image to each of a plurality of predetermined sizes. For example, the determination unit 134 generates multiple reduced face images, which are images of the face region reduced to each of a plurality of predetermined sizes, starting from a size smaller than the size of the original face region image and gradually increasing. The determination unit 134 also generates multiple enlarged face images by enlarging each of the multiple reduced face images, each reduced to each of a plurality of predetermined sizes, to the size of the original face image.
[0064] Figure 15 shows how multiple enlarged face images G81 to G84 are arranged from top to bottom in order of their predetermined size (reduction size) in the reduction process, from smallest to largest. Here, a small predetermined size in the reduction process corresponds to a strong mosaic effect. Conversely, a large predetermined size in the reduction process corresponds to a weak mosaic effect. For example, the determination unit 134 performs face detection on each of the multiple enlarged face images G81 to G84 and determines whether or not a face can be detected from each of the multiple enlarged face images G81 to G84. Here, determining that a face can be detected from an enlarged face image corresponds to determining that a face that can be recognized as a human face can be detected from the mosaic-processed face image. On the other hand, determining that a face cannot be detected from an enlarged face image corresponds to determining that a face that can be recognized as a human face cannot be detected from the mosaic-processed face image. The determination unit 134 also determines a lower limit size that is larger than the predetermined size corresponding to the enlarged face image for which it was determined that a face could not be detected. In other words, the determination unit 134 determines a lower limit size that is larger than a predetermined size corresponding to the enlarged face image in which it has been determined that it is not possible to detect a face that can be recognized as a human face.
[0065] In Figure 15, the determination unit 134 determines that faces cannot be detected in the enlarged face images G81 and G82. On the other hand, the determination unit 134 determines that faces can be detected in the enlarged face images G83 and G84. For example, the determination unit 134 performs face detection on each of the multiple enlarged face images G81 to G84 in order from the smallest predetermined size, determines whether or not faces can be detected in order from the smallest predetermined size, and determines a size larger than the predetermined size corresponding to the enlarged face image G82, for which it was determined that no face could be detected, as the lower limit size. For example, the determination unit 134 determines whether or not faces can be detected in order from the smallest predetermined size, and if faces are detected a predetermined number of times consecutively, it determines the smallest predetermined size included in the predetermined number of times as the lower limit size. Here, since the face detection model may output incorrect detection results, it may not be possible to say that there is a high probability that the results of only one detection are accurate and that there are faces that can be recognized as human faces. For this reason, if faces are detected a predetermined number of times consecutively, it can be said that there is a higher probability that there are faces that can be recognized as human faces.
[0066] As described above, the determination unit 134 generates a plurality of third enlarged face images by enlarging each of the plurality of second reduced face images, which are images of face regions reduced to a plurality of predetermined sizes, back to the size of the original face region image. The determination unit 134 also performs face detection on each of the plurality of third enlarged face images, determines whether or not a face can be detected from each of the plurality of third enlarged face images, and determines a lower limit size that is larger than the predetermined size corresponding to the third enlarged face image for which it is determined that a face cannot be detected.
[0067] Furthermore, the determination unit 134 generates a plurality of third enlarged face images by enlarging each of the plurality of second reduced face images, which are face region images reduced to a plurality of predetermined sizes that gradually increase from a size smaller than the original face region image, back to the size of the original face region image. The determination unit 134 also performs face detection on each of the plurality of third enlarged face images in order from the smallest predetermined size, and determines whether or not faces can be detected in order from the smallest predetermined size. If faces are detected a predetermined number of times consecutively, the smallest predetermined size included in the predetermined number of times is determined as the lower limit size. For example, the predetermined number of times can be any number as long as it is two or more.
[0068] The processing overview of the information processing device 100A according to the second embodiment will be explained using Figure 16. Figure 16 is a diagram illustrating the processing overview of the information processing device according to the second embodiment. The determination unit 134 generates a plurality of reduced face images by reducing the original face region image to each of a plurality of predetermined sizes through a reduction process (pattern 1) that reduces the input face image to a predetermined size. For example, the determination unit 134 generates a plurality of reduced face images, which are images of the face region reduced to each of a plurality of predetermined sizes that gradually decrease from the size of the original face region image. The determination unit 134 also generates a plurality of enlarged face images by enlarging each of the plurality of reduced face images, each of which has been reduced to each of a plurality of predetermined sizes, to the size of the original face image.
[0069] Figure 16 shows how multiple enlarged face images G91 to G93 are arranged from top to bottom in order of their predetermined size (reduction size) during the reduction process, from largest to smallest. Here, a larger predetermined size during the reduction process corresponds to a weaker mosaic effect. Conversely, a smaller predetermined size during the reduction process corresponds to a stronger mosaic effect. For example, the decision unit 134 extracts image features F91 to F93 from each of the multiple enlarged face images G91 to G93. For example, the decision unit 134 uses a feature extraction model, which is a machine learning model trained to output image features that represent the characteristics of an image when an image is input, to extract image features F91 to F93 from each of the multiple enlarged face images G91 to G93. For example, the image features may be vectors. For example, the decision unit 134 inputs each of the multiple enlarged face images G91 to G93 into a feature extraction model and extracts the image features F91 to F93 from each of the multiple enlarged face images G91 to G93. The decision unit 134 also extracts the image feature F9 of the original face region image G9 from the original face region image G9. For example, the decision unit 134 uses a feature extraction model to extract the image feature F9 of the original face region image G9 from the original face region image G9. For example, the decision unit 134 inputs the original face region image G9 into a feature extraction model and extracts the image feature F9 of the original face region image G9.
[0070] Furthermore, the determination unit 134 calculates the similarity between each of the image feature quantities F91 to F93 of the multiple enlarged face images G91 to G93 and the image feature quantity F9 of the original face region image G9. Here, a high similarity between the image feature quantities of the enlarged face image and the image feature quantities of the original face region image corresponds to a high probability that the image contains a face that can be recognized as the face of a specific individual. Conversely, a low similarity between the image feature quantities of the enlarged face image and the image feature quantities of the original face region image corresponds to a low probability that the image contains a face that can be recognized as the face of a specific individual. For example, the determination unit 134 calculates cosine similarity as an example of similarity. The determination unit 134 also determines whether the similarity is below a predetermined similarity and sets a size smaller than a predetermined size corresponding to the enlarged face image for which the similarity is determined to be below the predetermined similarity as the upper limit size. In Figure 16, the determination unit 134 determines that the cosine similarity between image feature quantity F91 and image feature quantity F9 exceeds the predetermined cosine similarity. On the other hand, the determination unit 134 determines that the cosine similarity between each of the image features F92 to F93 and the image feature F9 is less than or equal to a predetermined cosine similarity. For example, the determination unit 134 calculates the cosine similarity for each of the predetermined sizes in descending order, determines whether the cosine similarity for each of the predetermined sizes is less than or equal to a predetermined cosine similarity, and determines a size smaller than the predetermined size corresponding to the enlarged face image G92 for which the cosine similarity has been determined to be less than or equal to a predetermined cosine similarity as the upper limit size. For example, the determination unit 134 determines the upper limit size to be the largest predetermined size that is less than or equal to a predetermined cosine similarity. For example, the determination unit 134 determines the upper limit size to be the predetermined size of the enlarged face image G92 that corresponds to the largest predetermined size among the enlarged face images G92 and G93 for which the cosine similarity is less than or equal to a predetermined cosine similarity.
[0071] As described above, the determination unit 134 generates multiple third enlarged face images by enlarging each of the multiple second reduced face images, which are images of face regions reduced to each of multiple predetermined sizes, to the size of the original face region image. It calculates the similarity between the image features of each of the multiple third enlarged face images and the image features of the original face region image, determines whether the similarity is less than or equal to a predetermined similarity, and determines an upper limit size that is smaller than the predetermined size corresponding to the third enlarged face image for which the similarity is determined to be less than or equal to the predetermined similarity.
[0072] Furthermore, the determination unit 134 generates multiple third enlarged face images by enlarging each of the multiple second reduced face images, which are face region images reduced to multiple predetermined sizes that gradually decrease from the size of the original face region image, back to the size of the original face region image. The determination unit calculates the similarity of each image in order from the largest predetermined size, and determines whether the similarity is less than or equal to the predetermined similarity, in order from the largest predetermined size. If it is determined that the similarity is less than or equal to the predetermined similarity, the determination unit sets the largest predetermined size at which the similarity is less than or equal to the predetermined similarity as the upper limit size.
[0073] [3-3. Processing Procedure] An example of the processing procedure of the information processing device 100A according to the second embodiment will be explained using Figure 17. Figure 17 is a flowchart of an example of the processing procedure of the information processing device 100A according to the second embodiment. In Figure 17, the detection unit 131 acquires an image of a face region (step S201). For example, the detection unit 131 acquires an image. Subsequently, the detection unit 131 acquires an image of a face region by detecting the face region from the image. For example, the detection unit 131 acquires each of a plurality of face regions by detecting each of a plurality of face regions containing the faces of each of a plurality of people from the image.
[0074] Furthermore, the determination unit 134 generates a plurality of second reduced face images, each of which is an image of a face region reduced to a plurality of predetermined sizes (step S202). Subsequently, the determination unit 134 generates a plurality of third enlarged face images by enlarging each of the plurality of second reduced face images to the size of the original face region image (step S203). Subsequently, the determination unit 134 performs face detection on each of the plurality of third enlarged face images (step S204). For example, the determination unit 134 performs face detection on each of the plurality of third enlarged face images in order from the smallest predetermined size.
[0075] Furthermore, the determination unit 134 determines whether or not a face can be detected from each of the multiple third enlarged face images. For example, the determination unit 134 determines whether or not a face can be detected in order from the smallest predetermined size. Subsequently, the determination unit 134 determines whether or not a face has been detected a predetermined number of times consecutively (step S205).
[0076] If the determination unit 134 determines that it has not detected a predetermined number of faces consecutively (step S205; No), it performs face detection on each of the multiple third enlarged face images (step S204). On the other hand, if the determination unit 134 determines that it has detected a predetermined number of faces consecutively (step S205; Yes), it determines the smallest predetermined size included in the predetermined number of detections as the lower limit size (step S206).
[0077] An example of the processing procedure of the information processing device 100A according to the second embodiment will be explained using Figure 18. Figure 18 is a flowchart of an example of the processing procedure of the information processing device 100A according to the second embodiment. In Figure 18, the detection unit 131 acquires an image of a face region (step S301). For example, the detection unit 131 acquires an image. Subsequently, the detection unit 131 acquires an image of a face region by detecting the face region from the image. For example, the detection unit 131 acquires each of a plurality of face regions by detecting each of a plurality of face regions containing the faces of each of a plurality of people from the image.
[0078] Next, the determination unit 134 generates a plurality of second reduced face images, each of which is an image of a face region reduced to a plurality of predetermined sizes (step S302). Subsequently, the determination unit 134 generates a plurality of third enlarged face images by enlarging each of the plurality of second reduced face images to the size of the original face region image (step S303). Subsequently, the determination unit 134 calculates the similarity between the image features of each of the plurality of third enlarged face images and the image features of the original face region image (step S304). The determination unit 134 determines whether the similarity is less than or equal to a predetermined similarity (step S305). For example, the determination unit 134 determines whether the similarity is less than or equal to a predetermined similarity for each of the predetermined sizes, starting from the largest.
[0079] If the determination unit 134 determines that the similarity is not below a predetermined similarity (step S305; No), it determines whether the similarity is below a predetermined similarity (step S305). For example, the determination unit 134 determines whether the similarity is below a predetermined similarity for each predetermined size, starting from the largest. On the other hand, if the determination unit 134 determines that the similarity is below a predetermined similarity (step S305; Yes), it determines the largest predetermined size that is below a predetermined threshold as the upper limit size (step S306).
[0080] [3-4. Effects] As described above, the information processing device 100A according to the second embodiment includes a determination unit 134. The determination unit 134 determines a predetermined reduction ratio based on a reference reduction ratio and the size of the original face region image.
[0081] This allows the information processing device 100A to automatically determine an appropriate value for a predetermined reduction ratio.
[0082] Furthermore, the determination unit 134 generates a plurality of third enlarged face images by enlarging each of the plurality of second reduced face images, which are images of face regions reduced to a plurality of predetermined sizes, back to the size of the original face region image. It then performs face detection on each of the plurality of third enlarged face images and determines whether or not a face can be detected from each of the plurality of third enlarged face images. For third enlarged face images for which it is determined that a face cannot be detected, it determines a size larger than the predetermined size corresponding to that image as the lower limit size.
[0083] As a result, the information processing device 100A can determine a lower limit size that is larger than a predetermined size corresponding to an enlarged face image in which it has been determined that it cannot detect a face that can be recognized as a human face. Therefore, the information processing device 100A can automatically determine an appropriate value for the lower limit size.
[0084] Furthermore, the determination unit 134 generates a plurality of third enlarged face images by enlarging each of the plurality of second reduced face images, which are face region images that have been reduced to a plurality of predetermined sizes that gradually increase from a size smaller than the original face region image, to the size of the original face region image. Face detection is performed on each of the plurality of third enlarged face images in order from the smallest predetermined size, and it is determined whether or not faces can be detected in order from the smallest predetermined size. If faces are detected a predetermined number of times consecutively, the smallest predetermined size included in the predetermined number of times is determined as the lower limit size.
[0085] As a result, the information processing device 100A can determine the smallest size among predetermined sizes corresponding to an enlarged face image with a higher probability of containing a face that can be recognized as a human face as the lower limit size. Therefore, the information processing device 100A can automatically determine an appropriate value for the lower limit size.
[0086] Furthermore, the determination unit 134 generates multiple third enlarged face images by enlarging each of the multiple second reduced face images, which are images of face regions reduced to multiple predetermined sizes, to the size of the original face region image. It calculates the similarity between the image feature quantities of each of the multiple third enlarged face images and the image feature quantities of the original face region image, determines whether the similarity is less than or equal to a predetermined similarity, and determines an upper limit size that is smaller than the predetermined size corresponding to the third enlarged face image for which the similarity is determined to be less than or equal to the predetermined similarity.
[0087] As a result, the information processing device 100A can determine an upper limit size smaller than a predetermined size corresponding to an enlarged face image that has a low probability of containing a face that can be recognized as the face of a specific individual. Therefore, the information processing device 100A can automatically determine an appropriate value for the upper limit size.
[0088] Furthermore, the determination unit 134 generates multiple third enlarged face images by enlarging each of the multiple second reduced face images, which are face region images reduced to multiple predetermined sizes that gradually decrease from the size of the original face region image, back to the size of the original face region image. The determination unit calculates the similarity of each image in order from the largest predetermined size, and determines whether the similarity is less than or equal to the predetermined similarity, in order from the largest predetermined size. If it is determined that the similarity is less than or equal to the predetermined similarity, the determination unit sets the largest predetermined size at which the similarity is less than or equal to the predetermined similarity as the upper limit size.
[0089] As a result, the information processing device 100A can determine the largest size among those smaller than a predetermined size corresponding to an enlarged face image with a low probability of containing a face that can be recognized as the face of a specific individual as the upper limit size. Therefore, the information processing device 100A can automatically determine an appropriate value for the upper limit size.
[0090] [4. Third Embodiment] In the third embodiment, a method for performing mosaic processing of appropriate intensity without using the predetermined reduction ratio, lower limit size, and upper limit size of the first embodiment will be described.
[0091] [4-1. Configuration of the Information Processing System] Using Figure 11, an example of the configuration of the information processing system 1B according to the third embodiment will be described. In this embodiment, the configuration of the information processing system 1B is the same as that of the first embodiment, so you can refer to Figure 11, which shows an example of the configuration of the information processing system 1 according to the first embodiment. In the third embodiment, the information processing device 100 shown in Figure 11 is replaced by the information processing device 100B according to the third embodiment. As the user terminal 10 is the same as in the first embodiment, the description of the user terminal 10 will be omitted here.
[0092] The information processing device 100B is a device that performs information processing according to the third embodiment. The information processing device 100B performs information processing according to the third embodiment by executing an information processing program according to the third embodiment. Specifically, the information processing device 100B detects each of a plurality of face regions, each containing the face of a plurality of persons, from an image. The information processing device 100B also generates a plurality of enlarged face images by enlarging each of the plurality of reduced face images, which are images of face regions reduced to each of a plurality of predetermined sizes, to the size of the original face region image. The information processing device 100B also determines whether the similarity between the image feature quantities of each of the plurality of enlarged face images and the image feature quantities of the original face region image exceeds a predetermined similarity. If the information processing device 100B determines that the similarity exceeds the predetermined similarity, it acquires the enlarged face image corresponding to the predetermined size for which the similarity exceeded the predetermined similarity as a mosaic image.
[0093] [4-2. Configuration of Information Processing Device] Using Figure 19, an example of the configuration of the information processing device 100B according to the third embodiment will be described. Figure 19 is a diagram showing an example of the configuration of the information processing device 100B according to the third embodiment. The information processing device 100B includes a communication unit 110, a storage unit 120, and a control unit 130B. Since all parts except the control unit 130B are common to the functional units included in the information processing device 100 according to the first embodiment, only the control unit 130B, which is not common to the first embodiment, will be described here, and the description of the other functional units will be omitted.
[0094] (Control Unit 130B) The control unit 130B is a controller, and is implemented, for example, by a CPU or MPU executing various programs stored in the memory device inside the information processing device 100B using RAM as the working area. Alternatively, the control unit 130B is a controller and can be implemented, for example, by an integrated circuit such as an ASIC or FPGA.
[0095] The control unit 130B has a detection unit 131, a generation unit 132B, a determination unit 133B, and an acquisition unit 134B as functional units, and may realize or execute the information processing operations described below. Note that the internal configuration of the control unit 130B is not limited to the configuration shown in Figure 19, and other configurations are also acceptable as long as they perform the information processing described later. Furthermore, each functional unit represents the function of the control unit 130B and does not necessarily have to be physically distinct.
[0096] The following describes each element of the control unit 130B in order, but the detection unit 131 is the same as in the first embodiment, so its description is omitted here.
[0097] (Generation unit 132B) The generation unit 132B generates multiple enlarged face images by enlarging each of the multiple reduced face images, which are images of face regions reduced to each of a plurality of predetermined sizes, back to the size of the original face region image. For example, the generation unit 132B generates multiple reduced face images, which are images of face regions reduced to each of a plurality of predetermined sizes. Subsequently, the generation unit 132B generates multiple enlarged face images by enlarging each of the multiple reduced face images back to the size of the original face region image. In addition, the generation unit 132B generates multiple enlarged face images by enlarging each of the multiple reduced face images, which are images of face regions reduced to each of a plurality of predetermined sizes, starting from an image size smaller than the size of the original face region image and gradually increasing in size. For example, the generation unit 132B generates multiple reduced face images, which are images of face regions reduced to each of a plurality of predetermined sizes, starting from an image size smaller than the size of the original face region image and gradually increasing in size. Subsequently, the generation unit 132B generates multiple enlarged face images by enlarging each of the multiple reduced face images back to the size of the original face region image.
[0098] (Judgment unit 133B) The determination unit 133B determines whether the similarity between the image features of each of the multiple enlarged face images and the image features of the original face region image exceeds a predetermined similarity. Specifically, the determination unit 133B uses a feature extraction model to extract the image features of each of the multiple enlarged face images generated by the generation unit 132B. For example, the determination unit 133B inputs each of the multiple enlarged face images generated by the generation unit 132B into the feature extraction model to obtain the image features of each of the multiple enlarged face images. The determination unit 133B also uses the feature extraction model to extract the image features of the original face region image. For example, the determination unit 133B inputs the image of the original face region obtained by the detection unit 131 into the feature extraction model to obtain the image features of the original face region image. The determination unit 133B also calculates the similarity between the image features of each of the multiple enlarged face images and the image features of the original face region image. For example, the determination unit 133B calculates cosine similarity as an example of similarity. Next, the determination unit 133B determines whether the similarity exceeds a predetermined similarity, starting with the smallest predetermined size. For example, the determination unit 133B determines whether the cosine similarity exceeds a predetermined cosine similarity, starting with the smallest predetermined size.
[0099] (Acquisition part 134B) If the acquisition unit 134B determines that the similarity exceeds a predetermined similarity, it acquires the enlarged face image corresponding to the predetermined size for which the similarity exceeds the predetermined similarity as a mosaic image. For example, if the acquisition unit 134B determines that the similarity exceeds a predetermined similarity, it acquires the enlarged face image corresponding to the smallest predetermined size for which the similarity exceeds the predetermined similarity as a mosaic image. In other words, the acquisition unit 134B acquires the enlarged face image corresponding to the similarity for which the similarity exceeds a predetermined similarity as a mosaic image. For example, if the acquisition unit 134B determines that the cosine similarity exceeds a predetermined cosine similarity, it acquires the enlarged face image corresponding to the smallest predetermined size for which the cosine similarity exceeds a predetermined cosine similarity as a mosaic image. In other words, the acquisition unit 134B acquires the enlarged face image corresponding to the smallest predetermined size from among a plurality of enlarged face images for which the cosine similarity exceeds a predetermined cosine similarity as a mosaic image.
[0100] [4-3. Processing Procedure] An example of the processing procedure of the information processing device 100B according to the third embodiment will be explained using Figure 20. Figure 20 is a flowchart of an example of the processing procedure of the information processing device 100B according to the third embodiment. In Figure 20, the detection unit 131 acquires an image (step S401). Subsequently, the detection unit 131 detects face regions from the image (step S402). For example, the detection unit 131 detects each of a plurality of face regions, each containing the face of a plurality of people, from the image. By detecting each of the plurality of face regions, the detection unit 131 acquires each of the plurality of face regions.
[0101] Furthermore, the generation unit 132B generates a plurality of reduced face images, each of which is an image of a face region reduced to a plurality of predetermined sizes (step S403). For example, the generation unit 132B generates a plurality of reduced face images, each of which is an image of a face region reduced to a plurality of predetermined sizes, starting from an image size smaller than the size of the original face region image and gradually increasing. Subsequently, the generation unit 132B generates a plurality of enlarged face images by enlarging each of the plurality of reduced face images to the size of the original face region image (step S404).
[0102] Furthermore, the determination unit 133B calculates the similarity between the image features of each of the multiple enlarged face images and the image features of the original face region image (step S405). Subsequently, the determination unit 133B determines whether the similarity exceeds a predetermined similarity (step S406). For example, the determination unit 133B determines whether the similarity exceeds the predetermined similarity for each image in order from the smallest predetermined size.
[0103] If the determination unit 133B determines that the similarity does not exceed a predetermined similarity (step S406; No), it determines whether the similarity exceeds a predetermined similarity (step S406). For example, the determination unit 133B determines whether the similarity exceeds a predetermined similarity for each predetermined size in order from the smallest to the largest. On the other hand, if the determination unit 133B determines that the similarity exceeds a predetermined similarity (step S406; Yes), it obtains an enlarged face image corresponding to the predetermined size for which the similarity exceeded the predetermined similarity as a mosaic image (step S407). For example, the determination unit 133B obtains an enlarged face image corresponding to the smallest predetermined size for which the similarity exceeded the predetermined similarity as a mosaic image.
[0104] [4-4. Variations] The determination unit 133B performs face detection on each of the multiple enlarged face images and determines whether or not a face can be detected in each of the multiple enlarged face images. If the determination unit 133B determines that a face can be detected, the acquisition unit 134B acquires the enlarged face image corresponding to the predetermined size in which a face was determined to be detectable as a mosaic image.
[0105] Furthermore, the generation unit 132B generates multiple enlarged face images by enlarging each of the multiple reduced face images, which are images of face regions reduced to multiple predetermined sizes that are gradually smaller than the size of the original face region image, back to the size of the original face region image. The determination unit 133B determines whether the similarity is less than or equal to a predetermined similarity, starting with the largest predetermined size. If the determination unit 133B determines that the similarity is less than or equal to a predetermined similarity, the acquisition unit 134B acquires the enlarged face image corresponding to the largest predetermined size for which the similarity was determined to be less than or equal to a predetermined similarity as a mosaic image.
[0106] [4-5. Effects] As described above, the information processing device 100B according to the third embodiment includes a detection unit 131, a generation unit 132B, a determination unit 133B, and an acquisition unit 134B. The detection unit 131 detects each of a plurality of face regions, each containing the face of a plurality of persons, from an image. The generation unit 132B generates a plurality of enlarged face images by enlarging each of the plurality of reduced face images, which are images of face regions reduced to each of a plurality of predetermined sizes, to the size of the original face region image. The determination unit 133B determines whether the similarity between the image feature quantities of each of the plurality of enlarged face images and the image feature quantities of the original face region image exceeds a predetermined similarity. If the determination unit 133B determines that the similarity exceeds the predetermined similarity, the acquisition unit 134B acquires the enlarged face image corresponding to the predetermined size for which the similarity exceeded the predetermined similarity as a mosaic image.
[0107] As a result, the information processing device 100B can perform appropriate mosaic processing according to the size of each face area of multiple people included in the image. Furthermore, because the information processing device 100 can perform appropriate mosaic processing according to the size of each face area of multiple people included in the image, it can contribute to achieving Sustainable Development Goal (SDG) 9, "Build resilient infrastructure, promote inclusive and sustainable industrialization and foster innovation." In addition, the information processing device 100B can perform mosaic processing of appropriate intensity for each of the multiple face areas included in the image without using predetermined reduction ratios, lower size limits, and upper size limits.
[0108] Furthermore, the generation unit 132B generates multiple enlarged face images by enlarging each of the multiple reduced face images, which are images of the face region reduced to a number of predetermined sizes that gradually increase from an image size smaller than the original face region image, back to the size of the original face region image. The determination unit 133B determines whether the similarity exceeds a predetermined similarity threshold, starting from the smallest predetermined size. If the determination unit 133B determines that the similarity exceeds a predetermined similarity threshold, the acquisition unit 134B acquires the enlarged face image corresponding to the smallest predetermined size for which the similarity exceeded the predetermined similarity threshold as a mosaic image.
[0109] As a result, the information processing device 100B can perform mosaic processing of appropriate intensity on each of the multiple face regions included in the image without using predetermined reduction ratios, lower size limits, and upper size limits.
[0110] Furthermore, the determination unit 133B performs face detection on each of the multiple enlarged face images and determines whether or not a face can be detected in each of the multiple enlarged face images. If the determination unit 133B determines that a face can be detected, the acquisition unit 134B acquires the enlarged face image corresponding to the predetermined size in which the face was determined to be detectable as a mosaic image.
[0111] As a result, the information processing device 100B can perform mosaic processing of appropriate intensity on each of the multiple face regions included in the image without using predetermined reduction ratios, lower size limits, and upper size limits.
[0112] Furthermore, the generation unit 132B generates multiple enlarged face images by enlarging each of the multiple reduced face images, which are images of face regions reduced to multiple predetermined sizes that are gradually smaller than the size of the original face region image, back to the size of the original face region image. The determination unit 133B determines whether the similarity is less than or equal to a predetermined similarity, starting with the largest predetermined size. If the determination unit 133B determines that the similarity is less than or equal to a predetermined similarity, the acquisition unit 134B acquires the enlarged face image corresponding to the largest predetermined size for which the similarity was determined to be less than or equal to a predetermined similarity as a mosaic image.
[0113] As a result, the information processing device 100B can perform mosaic processing of appropriate intensity on each of the multiple face regions included in the image without using predetermined reduction ratios, lower size limits, and upper size limits.
[0114] [5. Hardware Configuration] Furthermore, the information processing device 100 according to the above-described embodiment is realized by a computer 1000 having a configuration such as that shown in Figure 21. The following explanation will use the information processing device 100 as an example. Figure 21 is a diagram showing an example of the hardware configuration. The computer 1000 is connected to an output device 1010 and an input device 1020, and has a configuration in which an arithmetic unit 1030, a primary storage device 1040, a secondary storage device 1050, an output interface 1060, an input interface 1070, and a network interface 1080 are connected by a bus 1090.
[0115] The arithmetic unit 1030 operates based on programs stored in the primary storage device 1040 and the secondary storage device 1050, as well as programs read from the input device 1020, and executes various processes. The arithmetic unit 1030 can be implemented using, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), an ASIC (Application Specific Integrated Circuit), or an FPGA (Field Programmable Gate Array).
[0116] The primary storage device 1040 is a memory device, such as RAM (Random Access Memory), that temporarily stores data used by the arithmetic unit 1030 for various calculations. The secondary storage device 1050 is a storage device where data used by the arithmetic unit 1030 for various calculations and various databases are registered, and can be implemented using ROM (Read Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), flash memory, etc. The secondary storage device 1050 may be internal storage or external storage. The secondary storage device 1050 may also be a removable storage medium such as USB (Universal Serial Bus) memory or SD (Secure Digital) memory card. The secondary storage device 1050 may also be cloud storage (online storage), NAS (Network Attached Storage), file server, etc.
[0117] The output I / F 1060 is an interface for transmitting information to be output to output devices 1010, such as displays, projectors, and printers, and is implemented using connectors of standards such as USB (Universal Serial Bus), DVI (Digital Visual Interface), and HDMI (High Definition Multimedia Interface). The input I / F 1070 is an interface for receiving information from various input devices 1020, such as mice, keyboards, keypads, buttons, and scanners, and is implemented using, for example, USB.
[0118] Furthermore, the output interface 1060 and input interface 1070 may be wirelessly connected to the output device 1010 and input device 1020, respectively. In other words, the output device 1010 and input device 1020 may be wireless devices.
[0119] Furthermore, the output device 1010 and the input device 1020 may be integrated as a touch panel. In this case, the output I / F 1060 and the input I / F 1070 may also be integrated as an input / output I / F.
[0120] The input device 1020 may also be a device that reads information from, for example, an optical recording medium such as a CD (Compact Disc), DVD (Digital Versatile Disc), or PD (Phase Change Rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.
[0121] The network interface 1080 receives data from other devices via network N and sends it to the computing unit 1030, and also transmits data generated by the computing unit 1030 to other devices via network N.
[0122] The arithmetic unit 1030 controls the output device 1010 and the input device 1020 via the output interface 1060 and the input interface 1070. For example, the arithmetic unit 1030 loads a program from the input device 1020 or the secondary storage device 1050 onto the primary storage device 1040 and executes the loaded program.
[0123] For example, when computer 1000 functions as an information processing device 100, the arithmetic unit 1030 of computer 1000 realizes the functions of the control unit 130 by executing a program loaded onto the primary storage device 1040. Alternatively, the arithmetic unit 1030 of computer 1000 may load a program obtained from another device via the network interface 1080 onto the primary storage device 1040 and execute the loaded program. Furthermore, the arithmetic unit 1030 of computer 1000 may cooperate with other devices via the network interface 1080 and call and use program functions, data, etc., from other programs on other devices.
[0124] [6. Other] Although embodiments of the present invention have been described above, the present invention is not limited by the content of these embodiments. Furthermore, the aforementioned components include those that can be easily conceived by those skilled in the art, those that are substantially the same, and those that fall within the so-called equivalent range. Moreover, the aforementioned components can be combined as appropriate. Furthermore, various omissions, substitutions, or modifications of the components can be made without departing from the gist of the embodiments described above.
[0125] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above document and drawings can be arbitrarily changed unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.
[0126] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.
[0127] For example, the information processing devices 100, 100A, and 100B described above may be implemented using multiple server computers, and the configuration can be flexibly changed, such as by calling external platforms via APIs (Application Programming Interfaces) or network computing depending on the function.
[0128] Furthermore, the embodiments and modifications described above can be combined as appropriate, provided that the processing content is not inconsistent. [Explanation of Symbols]
[0129] 100, 100A, 100B Information Processing Devices 110 Communications Department 120 Storage section 130, 130A, 130B control unit 131 Detection unit 132, 133B Judgment section 133, 132B generation section 134 Decision Section 134B Acquisition Department
Claims
1. A detection unit that detects each of multiple face regions, including each of the faces of multiple people, from an image, A determination unit that determines whether the size of the first reduced face image, which is an image of the face region reduced by a predetermined reduction ratio, is greater than or equal to a lower limit size and less than or equal to an upper limit size, If the determination unit determines that the size of the first reduced face image is below the lower limit size, the size of the first reduced face image is resized to the lower limit size. If the determination unit determines that the size of the first reduced face image exceeds the upper limit size, the size of the first reduced face image is resized to the upper limit size. A generation unit generates a first enlarged face image by enlarging the resized first reduced face image to the size of the original face region image, An information processing device equipped with the following features.
2. The generating unit is If the determination unit determines that the size of the first reduced face image is greater than or equal to the lower limit size and less than or equal to the upper limit size, a second enlarged face image is generated by enlarging the first reduced face image to the size of the original face region image. The information processing apparatus according to claim 1.
3. A determination unit that determines the predetermined reduction ratio based on a reference reduction ratio and the size of the original image of the face region. The information processing apparatus according to claim 1, further comprising:
4. A determination unit generates a plurality of third enlarged face images by enlarging each of a plurality of second reduced face images, which are images of the face region reduced to each of a plurality of predetermined sizes, to the size of the original face region image; performs face detection on each of the plurality of third enlarged face images; determines whether or not a face can be detected from each of the plurality of third enlarged face images; and determines the lower limit size to be a size larger than the predetermined size corresponding to the third enlarged face image for which it was determined that a face could not be detected. The information processing apparatus according to claim 1, further comprising:
5. The aforementioned determination unit, The process involves generating a plurality of third enlarged face images by enlarging each of the plurality of second reduced face images, which are images of the face region reduced to each of the plurality of predetermined sizes, starting from a size smaller than the original face region image and gradually increasing in size, to the size of the original face region image; performing face detection on each of the plurality of third enlarged face images in order from the smallest predetermined size; determining whether or not the face can be detected in order from the smallest predetermined size; and if the face is detected a predetermined number of times consecutively, determining the smallest predetermined size included in the predetermined number of times as the lower limit size. The information processing apparatus according to claim 4.
6. A determination unit generates a plurality of third enlarged face images by enlarging each of a plurality of second reduced face images, which are images of the face region reduced to each of a plurality of predetermined sizes, to the size of the original face region image; calculates the similarity between the image feature quantities of each of the plurality of third enlarged face images and the image feature quantities of the original face region image; determines whether the similarity is less than or equal to a predetermined similarity; and determines the upper limit size to be a size smaller than the predetermined size corresponding to the third enlarged face image for which the similarity has been determined to be less than or equal to the predetermined similarity; The information processing apparatus according to claim 1, further comprising:
7. The aforementioned determination unit, The process generates a plurality of third enlarged face images by enlarging each of the plurality of second reduced face images, which are images of the face region reduced to each of the plurality of predetermined sizes that are gradually smaller than the original face region image, to the size of the original face region image, calculating the similarity for each of the predetermined sizes in descending order, determining whether the similarity is less than or equal to the predetermined similarity for each of the predetermined sizes in descending order, and if it is determined that the similarity is less than or equal to the predetermined similarity, determining the largest predetermined size that results in a similarity less than or equal to the predetermined similarity as the upper limit size. The information processing apparatus according to claim 6.
8. A detection unit that detects each of multiple face regions, including each of the faces of multiple people, from an image, A generation unit generates multiple enlarged face images by enlarging each of the multiple reduced face images, which are images of the face region reduced to each of a plurality of predetermined sizes, back to the size of the original face region image. A determination unit that determines whether the similarity between the image feature quantities of each of the multiple enlarged face images and the image feature quantities of the original face region image exceeds a predetermined similarity, If the determination unit determines that the similarity exceeds a predetermined similarity, the acquisition unit acquires an enlarged face image corresponding to the predetermined size for which the similarity has been determined to exceed the predetermined similarity as a mosaic image. An information processing device equipped with the following features.
9. The generating unit is Each of the multiple reduced face images, which are images of the face region that have been reduced to each of the multiple predetermined sizes, starting from an image size smaller than the original face region image and gradually increasing in size, is enlarged to the original face region image size to generate the multiple enlarged face images. The determination unit, The similarity is determined to exceed the predetermined similarity, starting with the smallest predetermined size. The acquisition unit is, If the determination unit determines that the similarity exceeds the predetermined similarity, the enlarged face image corresponding to the smallest predetermined size for which the similarity exceeds the predetermined similarity is acquired as the mosaic image. The information processing apparatus according to claim 8.
10. The determination unit, Face detection is performed on each of the multiple enlarged face images, and it is determined whether or not a face can be detected from each of the multiple enlarged face images. The acquisition unit is, If the determination unit determines that the face can be detected, the enlarged face image corresponding to the predetermined size for which the face was determined to be detectable is acquired as the mosaic image. The information processing apparatus according to claim 8.
11. The generating unit is Multiple enlarged face images are generated by enlarging each of the multiple reduced face images, which are images of the face region that have been reduced to each of the multiple predetermined sizes that are gradually smaller than the size of the original face region image, back to the size of the original face region image. The determination unit, The similarity is determined to be less than or equal to the predetermined similarity for each of the predetermined sizes, starting with the largest. The acquisition unit is, If the determination unit determines that the similarity is less than or equal to the predetermined similarity, the enlarged face image corresponding to the largest predetermined size for which the similarity is determined to be less than or equal to the predetermined similarity is acquired as the mosaic image. The information processing apparatus according to claim 8.
12. An information processing method implemented by a program executed by an information processing device, A detection process that detects each of multiple face regions, each containing the face of each of multiple people, from an image, A determination step of determining whether the size of the first reduced face image, which is an image of the face region reduced by a predetermined reduction ratio, is greater than or equal to a lower limit size and less than or equal to an upper limit size, If the determination step determines that the size of the first reduced face image is below the lower limit size, the size of the first reduced face image is resized to the lower limit size. If the determination step determines that the size of the first reduced face image exceeds the upper limit size, the size of the first reduced face image is resized to the upper limit size. A generation step of generating a first enlarged face image by enlarging the resized first reduced face image to the size of the original face region image, Information processing methods including
13. An information processing method implemented by a program executed by an information processing device, A detection process that detects each of multiple face regions, each containing the face of each of multiple people, from an image, A generation step of generating multiple enlarged face images by enlarging each of the multiple reduced face images, which are images of the face region reduced to each of a plurality of predetermined sizes, to the size of the original face region image, A determination step of determining whether the similarity between the image feature quantities of each of the multiple enlarged face images and the image feature quantities of the original face region image exceeds a predetermined similarity, If the determination step determines that the similarity exceeds a predetermined similarity, the acquisition step involves acquiring an enlarged face image corresponding to the predetermined size for which the similarity was determined to exceed the predetermined similarity, as a mosaic image. Information processing methods including
14. A detection procedure for detecting each of multiple face regions, each containing the face of each of multiple people, from an image, A determination procedure for determining whether the size of a first reduced face image, which is an image of the face region reduced by a predetermined reduction ratio, is greater than or equal to a lower limit size and less than or equal to an upper limit size, If the determination procedure determines that the size of the first reduced face image is below the lower limit size, the size of the first reduced face image is resized to the lower limit size. If the determination procedure determines that the size of the first reduced face image exceeds the upper limit size, the size of the first reduced face image is resized to the upper limit size. A generation procedure for generating a first enlarged face image by enlarging the resized first reduced face image to the size of the original face region image, An information processing program that causes a computer to execute something.
15. A detection procedure for detecting each of multiple face regions, each containing the face of each of multiple people, from an image, A generation procedure for generating multiple enlarged face images by enlarging each of the multiple reduced face images, which are images of the face region reduced to each of a plurality of predetermined sizes, back to the size of the original face region image, A determination procedure for determining whether the similarity between the image feature quantities of each of the multiple enlarged face images and the image feature quantities of the original face region image exceeds a predetermined similarity, If the determination procedure determines that the similarity exceeds a predetermined similarity, the acquisition procedure includes obtaining an enlarged face image corresponding to the predetermined size for which the similarity was determined to exceed the predetermined similarity as a mosaic image. An information processing program that causes a computer to execute something.