Image processing apparatus, image processing method, and image processing program

The image processing apparatus uses multiple learning models to generate and specify region groups, addressing the challenge of inaccurate region specification, resulting in natural and personalized image processing outcomes.

JP7705025B2Active Publication Date: 2025-07-09FURYU KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021083336
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-17
Publication Date
2025-07-09
Estimated Expiration
2041-05-17

AI Technical Summary

Technical Problem

Existing image processing technologies struggle to accurately specify desired regions in images, leading to unnatural image data outcomes.

Method used

An image processing apparatus and method that utilizes multiple learning models to generate and specify region groups, allowing precise identification of target regions for processing, including skin, hair, and specific parts of a person, using techniques like semantic and instance segmentation.

Benefits of technology

Enables accurate and natural image processing by distinguishing between different regions of a person, ensuring consistent and personalized image enhancements for each individual, even in varied poses and compositions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007705025000001
    Figure 0007705025000001
  • Figure 0007705025000002
    Figure 0007705025000002
  • Figure 0007705025000003
    Figure 0007705025000003
Patent Text Reader

Abstract

To provide favorable image processing for a user by specifying a region being the image processing object.SOLUTION: An image processing device comprises: a reception unit (303) which receives image data including a person; a generation processing unit (304) which executes first generation processing for generating a first region group from the image data by using a first learning model that has learned a relation between image data for learning and the first region group being at least one or more regions for the person included in the image data for learning, and second generation processing for generating a second region group from the image data by using a second learning model that has learned a relation between the image data for learning and the second region group being at least one or more regions different from the first region group for the person included in the image data for learning; a specification unit (305) which specifies the object region being the processing object according to the first and second region groups from the image data; and an image processing unit (306) which performs prescribed image processing on the object region of the image data.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an image processing method, and an image processing program for processing an image of a person.

Background Art

[0002] Conventionally, a photo sticker creation apparatus that performs predetermined image processing on a captured image and provides it to a user is known. Specifically, there is one that generates a plurality of mask images with different shapes for a predetermined part and specifies it as an object of image processing. (For example, see Patent Document 1).

[0003] Patent Document 1 describes, for example, recognizing a region of a predetermined color existing in the lower region of a face image as a lip, generating a mask image corresponding to the recognized lip portion, and generating a lip image.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In Patent Document 1, there is a possibility that a desired region cannot be accurately specified. In such a case, the image data obtained by image processing may become unnatural.

[0006] An object of the present invention is to solve the above problems, and to provide an image processing apparatus, an image processing method, and an image processing program that specify a specific region of a user to be an object of image processing and perform preferable image processing for the user, and a region specification model and a model generation method used for these image processings.

Means for Solving the Problems

[0007] A first aspect of the image processing apparatus according to the present invention includes a reception unit that receives image data including a person, a first learning model that has learned the relationship between the learning image data including the person and a first region group that is at least one or more regions of the person included in the learning image data, a first generation processing unit that generates the first region group from the image data using the first learning model, a second learning model that has learned the relationship between the learning image data including the person and a second region group that is at least one or more regions different from the first region group of the person included in the learning image data, a second generation processing unit that generates the second region group from the image data using the second learning model, a specifying unit that specifies a target region to be processed according to the first and second region groups from the image data, and an image processing unit that performs predetermined image processing on the target region of the image data.

[0008] A first aspect of the image processing method according to the present invention includes a reception step of receiving image data including a person, a first generation step of generating a first region group from the image data using a first learning model that has learned the relationship between the learning image data including the person and a first region group that is at least one or more regions of the person included in the learning image data, a second generation step of generating a second region group from the image data using a second learning model that has learned the relationship between the learning image data including the person and a second region group that is at least one or more regions different from the first region group of the person included in the learning image data, a specifying step of specifying a target region to be processed according to the first and second region groups from the image data, and an image processing step of performing predetermined image processing on the target region of the image data.

[0009] A first aspect of the image processing program according to the present invention causes a computer to function as a reception unit that receives image data including a person, a first generation processing unit that generates a first region group from the image data using a first learning model that has learned the relationship between the learning image data including the person and a first region group that is at least one or more regions of the person included in the learning image data, a second generation processing unit that generates a second region group from the image data using a second learning model that has learned the relationship between the learning image data including the person and a second region group that is at least one or more regions different from the first region group of the person included in the learning image data, a specifying unit that specifies a target region to be processed according to the first and second region groups from the image data, and an image processing unit that performs predetermined image processing on the target region of the image data.

Advantages of the Invention

[0010] According to the present invention, it is possible to specify a region to be subjected to image processing and provide preferable image processing for a user.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Embodiments for Carrying Out the Invention

[0012] Hereinafter, specific embodiments of the present invention will be described with reference to the accompanying drawings. Note that members having the same function are denoted by the same reference numerals, and the description thereof will be omitted as appropriate. Furthermore, the shapes of the configurations shown in each drawing, and dimensions such as length, depth, and width are not intended to reflect the actual shapes and dimensions, and are appropriately changed for the sake of clarity and simplification of the drawings.

[0013] The image processing apparatus according to the present invention will be described as a photo sticker creation apparatus as an example. The photo sticker creation apparatus to which the present invention is applied is a game apparatus that allows a user to perform shooting, editing, etc. as a game, and provides the user with the captured and edited images as photo stickers or data. The photo sticker creation apparatus 1 is installed, for example, in game centers, shopping malls, and stores in tourist destinations.

[0014] In the game provided by the photo sticker creation device, the user uses the camera provided in the photo sticker creation device to take pictures of themselves or others. The user edits the captured image by synthesizing a foreground image and / or a background image, or by using editing functions (scribbling editing functions) such as pen image input or stamp image input as editing and synthesizing images, thereby designing the captured image to be rich in colors. After the game ends, the user receives a photo sticker or the like on which the edited image is printed as a result. Alternatively, the photo sticker creation device provides the edited image to the user's mobile terminal, and the user can also receive the result via the mobile terminal.

[0015] (Configuration of Photo Sticker Creation Device) Figures 1 and 2 are perspective views showing a configuration example of the appearance of the photo sticker creation device 1. The photo sticker creation device 1 is a game machine that provides captured images and edited images. The photo sticker creation device 1 provides images to the user by printing the images on sticker paper or by transmitting the images to a server so that the images can be viewed on the user's mobile terminal. The users of the photo sticker creation device 1 mainly consist of young women such as junior high school girls and high school girls. In the photo sticker creation device 1, a plurality of users, mainly two or three people per group, can enjoy the game. Of course, in the photo sticker creation device 1, a single user can also enjoy the game.

[0016] In the photo sticker creation device 1, the user performs a shooting operation with themselves as the subject. The user synthesizes a synthetic image such as handwritten characters or a stamp image with the image selected from the captured images obtained by shooting through an editing operation. As a result, the captured image is edited into a magnificent image. The user receives the sticker paper on which the edited image, that is, the edited image, is printed and ends a series of games.

[0017] As shown in Figure 1, the photo sticker creation device 1 is basically configured by installing the shooting unit 11 and the editing unit 12 in contact with each other.

[0018] The photographing unit 11 is composed of a pre-selection operation unit 20, a photographing operation unit 21, and a background unit 22. The pre-selection operation unit 20 is installed on the side surface of the photographing operation unit 21. The space in front of the pre-selection operation unit 20 becomes a pre-selection space where pre-selection processing is performed. Also, the photographing operation unit 21 and the background unit 22 are installed at a predetermined distance apart. The space formed between the photographing operation unit 21 and the background unit 22 becomes a photographing space where photographing processing is performed.

[0019] As pre-selection processing, the pre-selection operation unit 20 performs guidance for introducing a game provided by the photo sticker creating device 1, or performs processing for allowing the user to select various settings in the photographing processing performed in the photographing space. The pre-selection operation unit 20 is provided with a coin insertion slot into which the user inserts money, a touch panel monitor used for various operations, and the like. The pre-selection operation unit 20 appropriately guides the user in the pre-selection space to the photographing space according to the availability of the photographing space.

[0020] The photographing operation unit 21 is a device for photographing the user as a subject. The photographing operation unit 21 is located in front of the user who has entered the photographing space. On the front surface of the photographing operation unit 21 facing the photographing space, a camera, a touch panel monitor used for various operations, and the like are provided. When the left side surface is defined as the left side surface and the right side surface is defined as the right side surface as viewed from the user facing forward in the photographing space, the left side surface of the photographing operation unit 21 is constituted by the side panel 41A, and the right side surface is constituted by the side panel 41B. Further, the front surface of the photographing operation unit 21 is constituted by the front panel 42. It is assumed that the above-described pre-selection operation unit 20 is installed on the side panel 41A. Note that the pre-selection operation unit 20 may be installed on the side panel 41B, or may be installed on both of the side panels 41A and 41B.

[0021] The background part 22 is composed of a rear panel 51, a side panel 52A, and a side panel 52B. The rear panel 51 is a plate-shaped member located on the back side of the user facing the front. The side panel 52A is a plate-shaped member attached to the left end of the rear panel 51 and narrower in width than the side panel 41A. The side panel 52B is a plate-shaped member attached to the right end of the rear panel 51 and narrower in width than the side panel 41B.

[0022] The side panel 41A and the side panel 52A are provided in substantially the same plane. The upper parts of the side panel 41A and the side panel 52A are connected by a connecting part 23A which is a plate-shaped member. The lower parts of the side panel 41A and the side panel 52A are connected by a connecting part 23A' which is a member made of, for example, metal provided on the floor surface. Similarly, the side panel 41B and the side panel 52B are provided in substantially the same plane. The upper parts of the side panel 41B and the side panel 52B are connected by a connecting part 23B. The lower parts of the side panel 41B and the side panel 52B are connected by a connecting part 23B'.

[0023] Note that, for example, a green chroma key sheet is attached to the surface of the rear panel 51 on the shooting space side. The photo sticker creating device 1 performs chroma key compositing in the shooting process and the editing process by shooting with the chroma key sheet as the background. Thereby, the background image desired by the user is composited onto the sheet part.

[0024] The opening formed by being surrounded by the side panel 41A, the connecting part 23A, and the side panel 52A serves as the entrance and exit of the shooting space. Also, the opening formed by being surrounded by the side panel 41B, the connecting part 23B, and the side panel 52B serves as the entrance and exit of the shooting space.

[0025] Above the imaging space, a ceiling surrounded by the front of the imaging operation unit 21, the connecting part 23A, and the connecting part 23B is formed. A ceiling strobe unit 24 is provided in a part of the ceiling. One end of the ceiling strobe unit 24 is fixed to the connecting part 23A, and the other end is fixed to the connecting part 23B. The ceiling strobe unit 24 incorporates a strobe that irradiates light into the imaging space in accordance with imaging. Inside the ceiling strobe unit 24, in addition to the strobe, a fluorescent lamp is provided. Thus, the ceiling strobe unit 24 also functions as illumination for the imaging space.

[0026] The editing unit 12 is a device for performing editing on the captured image. The editing unit 12 is connected to the imaging unit 11 such that one side surface is in contact with the front panel 42 of the imaging operation unit 21.

[0027] Assuming the configuration of the editing unit 12 shown in FIG. 1 as the front side configuration, configurations used in the editing work are provided on each of the front side and the back side of the editing unit 12. With this configuration, two sets of users can perform the editing work simultaneously.

[0028] The front side of the editing unit 12 is composed of a surface 61 and an inclined surface 62 formed above the surface 61. The surface 61 is perpendicular to the floor surface and is substantially parallel to the side panel 41A of the imaging operation unit 21. On the inclined surface 62, as a configuration used in the editing work, a tablet-integrated monitor and a touch pen are provided. On the left side of the inclined surface 62, a columnar support portion 63A that supports one end of the lighting device 64 is provided. On the right side of the inclined surface 62, a columnar support portion 63B that supports the other end of the lighting device 64 is provided. On the upper surface of the support portion 63A, a support portion 65 that supports the curtain rail 26 is provided.

[0029] Above the editing unit 12, a curtain rail 26 is attached. The curtain rail 26 is composed of a combination of three rails 26A, 26B, and 26C. The three rails 26A, 26B, and 26C are combined so that their shape when viewed from above is U-shaped. One end of the rails 26A and 26B provided in parallel is fixed to the connecting part 23A and the connecting part 23B respectively, and the other ends of the rails 26A and 26B are joined to both ends of the rail 26C respectively.

[0030] A curtain is attached to the curtain rail 26 so that the space in front of the front of the editing unit 12 and the space in front of the back cannot be seen from the outside. The space in front of the front of the editing unit 12 and the space behind the back surrounded by the curtain become the editing space where the user performs the editing work.

[0031] Also, as will be described later, a discharge port for discharging the printed sticker paper is provided on the left side surface of the editing unit 12. The space in front of the left side surface of the editing unit 12 becomes a printing waiting space where the user waits for the printed sticker paper to be discharged.

[0032] (Movement of the user) Here, the flow of the image creation game and the movement of the user accompanying it will be described. FIG. 3 is a diagram for explaining the spatial movement of the user during the photo sticker creation game.

[0033] First, the user inserts money into the coin slot in the pre-selection space A0, which is the space in front of the pre-selection operation unit 20. Next, the user makes various settings according to the screen displayed on the touch panel monitor. The user performs, for example, pre-selection operations such as selecting a course related to shooting performed in the shooting space.

[0034] After finishing the pre-selection work, the user enters the shooting space A1 formed between the shooting operation unit 21 and the background unit 22 through the entrance G1 between the side panel 41A and the side panel 52A as indicated by the white arrow #1. Then the user performs the shooting work using the camera, touch panel monitor, etc. provided on the shooting operation unit 21.

[0035] After the user finishes the shooting operation, as shown by the white arrow #2, the user exits the shooting space A1 from the entrance G1 and moves to the editing space A2-1, or as shown by the white arrow #3, exits the shooting space A1 from the entrance G2 and moves to the editing space A2-2.

[0036] The editing space A2-1 is the editing space on the front side of the editing unit 12. On the other hand, the editing space A2-2 is the editing space on the back side of the editing unit 12. Which space, the editing space A2-1 or the editing space A2-2, the user moves to is guided by the screen display of the touch panel monitor of the shooting operation unit 21 or the like. For example, the user is guided to move to the empty one of the two editing spaces. The user who has moved to the editing space A2-1 or the editing space A2-2 starts the editing work. The user in the editing space A2-1 and the user in the editing space A2-2 can perform the editing work simultaneously.

[0037] After the editing work is completed, the printing of the edited image is started. When the printing starts, the user who has finished the editing work in the editing space A2-1 moves from the editing space A2-1 to the printing waiting space A3 as shown by the white arrow #4. Also, the user who has finished the editing work in the editing space A2-2 moves from the editing space A2-2 to the printing waiting space A3 as shown by the white arrow #5.

[0038] The user who has moved to the printing waiting space A3 waits for the completion of the image printing. When the printing is completed, the user receives the sticker paper discharged from the discharge port provided on the right side surface of the editing unit 12 and finishes the series of image creation games.

[0039] (Configuration of the pre-selection unit) Next, the configuration of each device will be described. FIG. 4 is a front view of the pre-selection operation unit 20.

[0040] A touch panel monitor 71 is provided above the pre-selection operation unit 20. The touch panel monitor 71 is composed of a monitor such as an LCD (Liquid Crystal Display) and a touch panel laminated thereon. The touch panel monitor 71 has a function of displaying various GUIs (Graphical User Interfaces) and receiving selection operations of users. On the touch panel monitor 71, a screen used for pre-selection processing for performing selection of a course related to shooting, selection of an image serving as a background in an edit target image to be edited, layout of a created image, selection of at least one of BGM (Back Ground Music), sound, and voice flowing during an image creation game, and input of a user's name is displayed.

[0041] As for the course related to shooting, a two-person course for two users to perform shooting and a large-group course for three or more users to perform shooting are prepared. Also, a couple course for a male and female couple to perform shooting may be prepared.

[0042] A speaker 72 is provided below the touch panel monitor 71. The speaker 72 outputs guidance voices, BGM, sound effects, etc. of the pre-selection process. Also, a coin insertion and return slot 73 for users to insert coins is provided adjacent to the speaker 72.

[0043] (Configuration of the shooting unit) FIG. 5 is a front view of a shooting operation unit 21 as a shooting unit. The shooting operation unit 21 is configured to be surrounded by a side panel 41A, a side panel 41B, and a front panel 42.

[0044] A camera unit 81 is provided at the center of the front panel 42. The camera unit 81 is composed of a camera 91 and a touch panel monitor 92 as a display unit.

[0045] The camera 91 is, for example, a single-lens reflex camera and is attached inside the camera unit 81 so that the lens is exposed. The camera 91 has an imaging device such as a CCD (Charge Coupled Device) image sensor or a CMOS (Complementary Metal Oxide Semiconductor) image sensor, and photographs a user in the photographing space A1. The moving image captured by the camera 91 (hereinafter also referred to as a live view display image) is displayed on the touch panel monitor 92 in real time. The still image captured by the camera 91 at a predetermined timing such as when shooting is instructed is saved as a photographed image.

[0046] The touch panel monitor 92 is provided below the camera 91. The touch panel monitor 92 is composed of a monitor such as an LCD and a touch panel laminated thereon. The touch panel monitor 92 has a function as a live view monitor for displaying the moving image captured by the camera 91 and a function of displaying various GUIs and receiving the selection operations of the user. Specific examples of the selection operation include selection of a shooting course, instructions for starting and ending shooting (control instructions for shooting), selection of the degree of eye distortion and color, and the degree of correction of skin color, selection of an image to be a background image in the created image, and selection of BGM (sound / audio) during an image creation game. The touch panel monitor 92 displays the moving image (live view image) and the still image (photographed image) captured by the camera 91.

[0047] Above the camera unit 81, an upper strobe 82 with a curved light emitting surface facing the user is installed. The upper strobe 82 irradiates light onto the user's face and upper body from directly above the front.

[0048] Also, below the camera unit 81, a foot strobe 85 that irradiates light onto the user's lower body and feet is provided. Note that the upper strobe 82 and the foot strobe 85 include strobes and fluorescent lamps.

[0049] Although not shown in FIGS. 1 and 5, a speaker 93 is provided, for example, near the ceiling of the front panel 42. The speaker 93 outputs guidance voices for shooting processing, BGM, sound effects, etc. according to the voice signals output from the control unit 201.

[0050] (Configuration of the background part) FIG. 6 is a front view of the shooting space A1 side of the background part 22.

[0051] Above the rear panel 51, a rear top strobe 101 is installed. The rear top strobe 101 irradiates the user with light from above the rear.

[0052] In the figure, a rear left strobe 102 is installed to the left of the rear panel 51. The rear left strobe 102 irradiates the user from the right side of the rear. In the figure, a rear right strobe 103 is installed to the right of the rear panel 51. The rear right strobe 103 irradiates the user from the left side of the rear.

[0053] Also, a chroma key sheet 121 may be attached to the surface of the rear panel 51 on the shooting space A1 side (the front side in the figure). The color of the chroma key sheet 121 is, for example, green.

[0054] Although not shown, a chroma key sheet may also be attached to the surfaces of the side panels 52A and 52B on the shooting space A1 side, similar to the chroma key sheet 121.

[0055] (Configuration of the editing unit) FIG. 7 is a front view of the editing space A2-1 side of the editing unit 12.

[0056] Almost at the center of the inclined surface 62, a tablet-integrated monitor 131 is provided. On the left side of the tablet-integrated monitor 131, a touch pen 132A is provided. On the right side of the tablet-integrated monitor 131, a touch pen 132B is provided. The tablet-integrated monitor 131 is configured by being provided such that the tablet exposes the display. The tablet enables operation input using the touch pen 132A or the touch pen 132B. For example, an editing screen used for an editing operation is displayed on the tablet-integrated monitor 131. For example, when two users perform an editing operation simultaneously, the touch pen 132A is used by the user on the left side facing the tablet-integrated monitor 131, and the touch pen 132B is used by the user on the right side facing the tablet-integrated monitor 131.

[0057] FIG. 8 is a left side view of the editing unit 12.

[0058] A sticker paper discharge port 161 is provided on the lower side of the left side surface of the editing unit 12. A printer as an output unit is provided inside the editing unit 12. By that printer, an image of the user in the editing space A2-1 or an image of the user in the editing space A2-2 is printed on sticker paper in a predetermined layout and discharged from the sticker paper discharge port 161.

[0059] (Internal Configuration of Photo Sticker Creation Device) FIG. 9 is a block diagram showing a configuration example of the interior of the photo sticker creation device 1. In FIG. 9, the same components as those described above are denoted by the same reference numerals. Redundant descriptions will be omitted as appropriate.

[0060] The control unit 201 is composed of a CPU (Central Processing Unit) or the like. The control unit 201 executes programs stored in the ROM (Read Only Memory) 206 and the storage unit 202, and controls the overall operation of the photo seal creating apparatus 1. The storage unit 202, the communication unit 203, the drive 204, the ROM 206, and the RAM (Random Access Memory) 207 are connected to the control unit 201. Each component of the pre-selection operation unit 20, the shooting operation unit 21, the background unit 22, the editing operation units 27A and 27B, and the printing operation unit 28 is also connected to the control unit 201.

[0061] The storage unit 202 is a non-volatile recording medium such as a hard disk or a flash memory. The storage unit 202 stores various setting information and the like supplied from the control unit 201. The information stored in the storage unit 202 is appropriately read by the control unit 201.

[0062] The communication unit 203 is an interface for a network such as the Internet. The communication unit 203 communicates with an external device according to the control by the control unit 201. For example, the communication unit 203 transmits a captured image or an edited image selected by the user to the server. The image transmitted from the communication unit 203 is stored in a predetermined storage area in the server and is displayed or downloaded to a mobile terminal that has accessed the server.

[0063] A removable medium 205 composed of an optical disk, a semiconductor memory, or the like is appropriately mounted on the drive 204. Programs and data read from the removable medium 205 by the drive 204 are supplied to the control unit 201 and are stored or installed in the storage unit 202.

[0064] The ROM 206 stores programs and data executed in the control unit 201. The RAM 207 temporarily stores data and programs processed by the control unit 201.

[0065] The pre-selection operation unit 20 implements pre-selection processing for users in the pre-selection space A0. The pre-selection operation unit 20 is composed of a touch panel monitor 71, a speaker 72, and a coin processing unit 74.

[0066] The touch panel monitor 71 displays various selection screens according to the control by the control unit 201 and receives the operations of the user on the selection screens. The input signal representing the content of the user's operation is supplied to the control unit 201, and various settings are performed.

[0067] The coin processing unit 74 detects the insertion of coins into the coin insertion / return port 73. When the coin processing unit 74 detects that coins of a predetermined amount have been inserted, it outputs a start signal instructing the start of the game to the control unit 201.

[0068] The shooting operation unit 21 implements shooting processing for users in the shooting space A1. The shooting unit 220 is composed of an upper strobe 82, a left strobe 83, a right strobe 84, a foot strobe 85, a camera 91, a touch panel monitor 92, and a speaker 93.

[0069] The upper strobe 82, the left strobe 83, the right strobe 84, and the foot strobe 85 are arranged in the shooting space A1 and emit light according to the illumination control signal supplied from the control unit 201.

[0070] The camera 91 takes pictures according to the shutter control by the control unit 201 and outputs the captured images (image data) obtained by the shooting to the control unit 201.

[0071] The editing operation unit 27A implements editing processing for users in the editing space A2-1. The editing operation unit 27A is composed of a tablet built-in monitor 131, touch pens 132A, 132B, and a speaker 133. The editing operation unit 27B implements editing processing for users in the editing space A2-2 and has the same configuration as the editing operation unit 27A. In the following, when the editing operation units 27A and 27B are not particularly distinguished, they are simply referred to as the editing operation unit 27.

[0072] The built-in monitor 131 of the tablet displays an editing screen according to the control by the control unit 201 and receives the user's operations on the editing screen. An input signal representing the content of the user's operation is supplied to the control unit 201, and the captured image to be edited is edited.

[0073] The printing operation unit 28 realizes a printing process of providing the user in the printing waiting space A3 with the sticker paper on which the created image has been printed. The printing operation unit 28 is configured to include a printer 140. A sticker paper unit 141 is attached to the printer 140.

[0074] Based on the print data supplied from the control unit 201, the printer 140 prints the edited image on the sticker paper 142 stored in the sticker paper unit 141 and discharges it to the sticker paper discharge port 161.

[0075] (Functional blocks of the photo sticker creating device) FIG. 10 is a block diagram showing the functional blocks of the photo sticker creating device 1. The photo sticker creating device 1 functions as a pre-selection unit 210, a photographing unit 220, an editing unit 230, and a printing unit 240. Further, by executing the photo sticker creation program of the present invention, the control unit 201 functions as a pre-selection processing unit 301, a photographing processing unit 302, a reception unit 303, a generation processing unit 304, a specifying unit 305, an image processing unit 306, a composition unit 307, an editing processing unit 308, and a printing processing unit 309.

[0076] The pre-selection unit 210 includes the above-described pre-selection operation unit 20 and a pre-selection processing unit 301. The pre-selection processing unit 301 performs pre-selection processing by controlling the touch panel monitor 71, the speaker 72, and the coin processing unit 74 in the pre-selection operation unit 20. The pre-selection processing unit 301 causes the touch panel monitor 71 to display a selection screen or the like for selecting a course related to the shooting performed in the shooting space A1. The pre-selection processing unit 301 also accepts an operation input from the user to the touch panel monitor 71. Specifically, the pre-selection processing unit 301 accepts a selection operation input for the selection screen displayed on the touch panel monitor 71, an input of the user's name, and the like. The pre-selection processing unit 301 also controls the output of guidance for explaining various selection operations. The pre-selection processing unit 301 causes the touch panel monitor 71 to display a screen for explaining various selection operations, or causes the speaker 72 to output a voice for explaining various selection operations.

[0077] The shooting unit 220 includes the above-described shooting operation unit 21 and a shooting processing unit 302. The shooting processing unit 302 performs shooting processing by controlling the camera 91, the touch panel monitor 92, and the speaker 93 in the shooting operation unit 21.

[0078] The shooting processing unit 302 accepts an operation input from the user to the touch panel monitor 92. For example, the shooting processing in the shooting space A1 is started when the shooting processing unit 302 accepts, as an input, a contact operation on the touch panel monitor 92 by the user.

[0079] The shooting processing unit 302 controls the camera 91 and shoots the user as a subject. There are moving images and still images for shooting. The shooting processing unit 302 controls the display of the touch panel monitor 92 to cause the moving image captured by the camera 91 to be displayed as a live view on the touch panel monitor 92, or to cause the still image that is the shooting result to be displayed as a captured image.

[0080] In addition, the photographing processing unit 302 causes the touch panel monitor 92 to display an instruction screen for explaining the number of photographed images, the standing position of the user, sample poses, messages regarding the line of sight, and photographing timing, etc. Further, the voice of the narration corresponding to each instruction screen and the BGM are output from the speaker 93.

[0081] The reception unit 303 receives the image data obtained by photographing the user by the photographing unit 220.

[0082] The generation processing unit 304 generates a group of regions using the learning data including a person and the learned model in which the regions of the person have been learned. Specifically, the generation processing unit 304 generates at least two or more of the “skin and hair mask”, the “specific part mask”, and the “person region mask” which are groups of regions.

[0083] (Skin and hair mask) First, the generation process of the skin and hair mask will be described. The generation processing unit 304 uses the learned skin and hair learning model for the relationship between the learning image data including a person and the skin and hair region group which is at least one or more regions of the person included in the learning image data, and generates the skin and hair region group from the image data received by the reception unit 303. Specifically, the generation processing unit 304 generates a skin and hair mask in which at least the regions of the skin and hair of the person are separated from other regions. For example, the generation processing unit 304 generates a skin and hair mask including a skin region 500, a hair region 501, and other regions 502 as shown in FIG. 11(b) from the image data as shown in FIG. 11(a). The skin region 500 is the region of the skin of the person exposed from the clothes, specifically, the regions such as the face, neck, and limbs. Further, the hair region 501 is specifically the region of the hair of the person. The skin region 500 and the hair region 501 are combined here to be the skin and hair region group. Note that, for example, as shown in FIG. 11(b), the other regions in the skin and hair mask include the regions of the lips, eyes, and eyebrows. Also, in the example shown in FIG. 11(b), the other regions 502 and the background region 503 are distinguished, but the background may also be included in the other regions.

[0084] As shown in an example in FIG. 12(a), the skin and hair learning model M is a model that has been pre-trained by a learning device on the relationship between a plurality of sets of learning image data and correct answer data as learning data. Specifically, the learning image data includes people. Also, the correct answer data includes the skin area, hair area, and other areas of the people included in the learning image data. The skin and hair learning model M obtained by learning the relationship between a plurality of sets of learning image data and correct answer data by a learning device can output the skin area, hair area, and other areas of the people included in the image data for image data including people, as shown in FIG. 12(b).

[0085] The generation processing unit 304 generates a skin and hair mask, for example, using semantic segmentation. By using semantic segmentation, fine-grained area extraction at the pixel level becomes possible, and a group of areas can be generated at the pixel level. Thereby, in the image processing unit 306 described later, image processing at the pixel level can be realized.

[0086] (Specific part mask) Subsequently, the generation process of the specific part mask will be described. The generation processing unit 304 generates a specific part region group from the image data using a part learning model that has learned the relationship between the learning image data including people and the specific part region group that is at least one or more regions for the people included in the learning image data. The specific part region group is a region group different from the skin and hair region group. Specifically, the generation processing unit 304 generates a specific part mask that separates specific parts of a person, for example, the regions of the head and limbs from other regions. For example, the generation processing unit 304 generates a specific part mask including a head region 510, limb regions 511, 512, and other regions 513, 514 as shown in FIG. 11(c) from the image data as shown in FIG. 11(a). The head region 510 and the limb regions 511, 512 are combined here as the specific part region group.

[0087] In the example shown in FIG. 11(c), the arm portion 511, which is a limb region, is a portion that includes the sleeve portion, which is a part of the clothing, and the fingertips exposed from the sleeve. Also, the leg portion 512, which is a limb region, is the leg portion exposed from the clothing. Although not included in FIG. 11(c), the leg region can include the toes. At this time, it is assumed that shoes or sandals are worn on the feet, and the leg portion includes up to the toes of such shoes. Also, in the example shown in FIG. 11(c), in the other regions, as the torso, the part 513 of the body part of the user's upper garment and the part 514 of the skirt, which is the lower garment, are detected as separate regions, but the part of the body part of the upper garment and the lower garment may be integrated. The division of each region in the specific part mask is not limited to the example shown in FIG. 11(c), and it is sufficient if at least the head region and the limb regions to be processed by the image processing unit 306 described later can be specified. Also, if a person is wearing pants that fit the legs instead of a skirt as the lower garment, the region of the pants may be regarded as the leg part of the limb region. And, as shown in FIG. 11(c), it is preferable that the neck region is not included in the head region 510 but is included in the torso region 513. This is because it is assumed that image processing that is not to be performed is required for the face region in the neck region. Also, in the case of performing image processing on the region from the neck to the décolletage region exposed from the clothing, for the neck and the décolletage region, a specific region group may be generated as a region different from the torso of the clothing. Here, how the generation processing unit 304 specifies the regions from the image data depends on what correct data of the specific part region group is used for the learning process together with the learning image data. Also, what correct data to use can be determined by what image processing is to be performed by the image processing unit 306 described later. In the example shown in FIG. 11(c), the other region 515 is distinguished from the background region, but the background may also be included in the other regions 513 and 514. That is, it is not essential to distinguish the regions 513 to 515 shown in FIG. 11(c).

[0088] Although the description using the illustration is omitted, the part learning model is generated by learning, in a learning device, the relationship between learning image data including a person and specific part areas such as the face area and the limb areas of the person which are correct answer data. The part learning model thus obtained can output the areas of specific parts of the person included in the image data for the image data including the person. In the example shown in FIG. 11(c), it is an example of a part specific mask generated using a part learning model learned with correct answer data that divides the head area 510 into the areas of the face and hair, the body area 513 into the area from the neck to the hem of the upper garment, the arm area 511 into the area from the upper arm to the fingertips, the area 514 of the lower garment, and the leg area 512 below the hem of the lower garment.

[0089] The generation processing unit 304 generates a specific part mask, for example, using semantic segmentation. By using semantic segmentation, it becomes possible to extract fine areas in pixel units, and a group of areas can be generated in pixel units. Thereby, in the image processing unit 306 described later, image processing in pixel units can be realized.

[0090] (Person area mask) Subsequently, the generation process of the person area mask will be described. The generation processing unit 304 generates a person area group from the image data received by the reception unit 303 using a region learning model that has learned the relationship between the learning image data including a person and the person area group that is the area of the person included in the learning image data. Specifically, the generation processing unit 304 generates a person area mask including the area for each person. For example, as shown in FIG. 11(a), even when a plurality of people are included in the image data and the hair and arms of the plurality of people overlap, as shown in FIG. 11(d), a person area mask divided into areas 520 and 521 for each person is generated. The areas of all the people included in this image data are combined to form a person area group. Incidentally, if only one person is included in the image data, the area of that one person is used as the person area group.

[0091] Although the description using the drawings is omitted, the region learning model is generated by having a learning device learn the relationship between learning image data including a person and the person region which is correct answer data, as in the example described above with reference to Fig. 12(a). The region learning model thus obtained can output the person region included in the image data for the image data including a person. For example, when the image data includes a plurality of persons, the region learning model can specify and output a region for each person.

[0092] Note that the learning image data used for generating the skin and hair learning model, the part learning model, and the region learning model does not necessarily have to be the same, but by using the same image data and performing learning simultaneously, high accuracy can be obtained when generating each region group in each learning model.

[0093] The generation processing unit 304 generates a person region mask, for example, using instance segmentation. By using instance segmentation, the region of each object can be classified in pixel units. Thereby, in the image processing unit 306 described later, image processing in pixel units can be realized.

[0094] The specifying unit 305 specifies a target region that is the target of image processing from the image data according to two or more region groups generated by the generation processing unit 304.

[0095] (Specification of Target Region Using Skin and Hair Mask and Specific Part Mask) The specific part 305 specifies the target area 600 of the skin of a person's face separately from other areas 601 according to the group of skin and hair areas indicated by the skin and hair mask and the specific part area indicated by the specific part mask. Specifically, as shown in Fig. 13(a), the area 600 of the skin of the person's face is specified using the hair and skin area mask shown in Fig. 11(b) and the specific area mask shown in Fig. 11(c). For example, the example of the image data shown in Fig. 11(a) is an image taken with a hand in front of the face of the person on the left and the hand overlapping the face. As shown in Fig. 13(a), by specifying the area of the skin of the face separately from the area of the skin of the hand, it is possible to target only the face for image processing. Also, in the example shown in Fig. 13(a), the area 602 of the person's hair is also specified, and it is possible to target it separately from the area 600 of the skin of the person's face and other areas 601 for image processing.

[0096] For example, when performing image processing such as makeup, if a hand overlaps and appears at the position where lipstick, lip gloss, cheek, eyeshadow, eyeliner, eyebrow, etc. are applied, there is a risk of obtaining an unnatural image in which the color applied to the face is also applied to the hand by the image processing. Also, although it is not uncommon for the face and the hand to appear to have different skin colors, there is a risk of obtaining an unnatural image in which the face and the hand are subjected to the same processing (for example, whitening processing). On the other hand, as shown in Fig. 13(a), by specifying the area of the skin of the face, it is possible to perform image processing only on the face part and prevent performing the same image processing as the face on other parts, thereby obtaining a natural image.

[0097] Specifically, in the case of the photo sticker creating device 1, there are a variety of variations in the poses of the users when taking pictures. Therefore, even if only the area around the face is photographed, there may be image data in which a hand exists in front of the face. For such image data, if the area including the face area and the hand area is extracted as the skin area, as described above, even when processed on the hand, since the hand area and the face area are distinguished, it is possible to perform only the necessary processing on the face area.

[0098] (Specification of the target area using the skin and hair mask and the person area mask) The specific part 305 identifies the target skin areas 610A, 610B and the hair areas 611A, 611B for each person according to the skin and hair area group indicated by the skin and hair mask and the person area group indicated by the person area mask. Specifically, using the skin and hair area mask shown in Fig. 11(b) and the person area mask shown in Fig. 11(d), as shown in Fig. 13(b), the skin areas 610A, 610B and the hair areas 611A, 611B for each person are identified. For example, when a plurality of people are included in the image data, the skin color and hair color of each person may be different. In such a case, it is not preferable to perform image processing on the skin colors and hair colors of the plurality of people included in the image in the same way. If the hair colors of two people are clearly different, if the hair of the two people is processed in the same way for image processing, there is a risk of an unnatural image. Also, for example, when the skin colors of two people are different, if the skin of the two people is processed in the same way for image processing, there is a risk of an unnatural image. On the other hand, as shown in Fig. 13(b), by identifying the skin areas 610A, 610B and the hair areas 611A, 611B for each person, it is possible to realize image processing for the skin of each person and image processing for the hair of each person, so that a natural image with image processing suitable for each person can be obtained.

[0099] Specifically, in the case of the photo sticker creation device 1, the number of people to be photographed is generally not one person but a plurality of people. Also, in the case of the photo sticker photographing device 1, there are a rich variety of poses of the user when photographing. For a simple group photo, the parts such as the arms, hands and legs of each person can be identified as the arms, hands and legs etc. at a "position close to the body". However, with the photo sticker photographing device 1, image data can be photographed in various poses. Therefore, it is difficult to identify the arms, hands and legs of each person based on the criterion of "a position close to the body". Therefore, as described above, by identifying the skin area of each person's area based on the generated person area mask, the skin areas of the people can be distinguished and natural image processing suitable for each person can be performed.

[0100] Note that the identification of such skin and hair regions for each person is a so-called panoptic segmentation technique because it is a combination of a skin and hair mask obtained by semantic segmentation and a person region mask obtained by instance segmentation.

[0101] (Identification of Target Region Using Specific Part Mask and Person Region Mask) The specific part 305 identifies the region of the specific part for each person according to the group of specific part regions indicated by the specific part mask and the group of person regions indicated by the person region mask. For example, the specific part 305 uses the specific region mask shown in Fig. 11(c) and the person region mask shown in Fig. 11(d) to identify the regions 620A, 620B, 621A, 621B of the limbs of the person as shown in Fig. 11(c). Also, the specific part 305 identifies the regions 622A, 622B of the head of the person, the regions 623A, 623B of the body of the upper garment, and the regions 624A, 624B of the lower garment as shown in Fig. 11(c). As an example, assume that at the time of image data shooting, a specific part of a person protruded forward compared to the torso. In this case, due to the perspective relationship, in the image data, as shown in Fig. 14, a part of the person's body (the tip of the right person's foot in the case of Fig. 14(a), the lower body (especially the knee) of the right person in the case of Fig. 14(b)) may appear larger than its original size. Since such an image looks unbalanced and unnatural, it may be image-processed so that the part that appears larger than normal looks balanced and natural. At this time, as shown in Fig. 13(c), by identifying which part is the specific part of which person, it becomes possible to realize the image target. Thereby, for example, a natural image obtained by image-processing an unbalanced part can be obtained.

[0102] Specifically, in the case of the photo sticker creation device 1, it is common for the number of people to be photographed to be more than one. Also, in the case of the photo sticker photographing device 1, there is a wide variety of poses that users can take when taking photos. For example, in the case of a simple group photo, it is common to be photographed standing still upright. However, with the photo sticker photographing device 1, image data can be photographed in various poses such as with hands extended forward, arms or knees bent. Therefore, it is not uncommon for the image data photographed by the photo sticker photographing device 1 to be unbalanced image data in which parts of the body parts such as each person's arms, hands, legs, feet, and joints of the arms and legs are unnaturally emphasized and appear large. Also, in the case of the photo sticker photographing device 1, compared to general photography, since image data is photographed in a narrow space, it is likely to become image data in which some of the body parts of the person are emphasized as described above and look unnatural overall. Therefore, as described above, using the regions of specific parts of each person specified using the specific part mask and the person region mask, it is possible to identify unnatural parts and perform image processing so that the unnatural parts look natural.

[0103] Note that the specification of the regions of specific parts for each person is a so-called panoptic segmentation technique because it is a combination of the part specification mask obtained by semantic segmentation and the person region mask obtained by instance segmentation.

[0104] (Specification of the target region using the skin hair mask, specific part mask, and person region mask) The specific part 305 identifies the specific part to be image - processed for each person according to the group of skin - hair regions indicated by the skin - hair mask, the group of specific - part regions indicated by the specific - part mask, and the group of person regions indicated by the person - region mask. Specifically, using the skin - hair mask shown in Fig. 11(b), the specific - part mask shown in Fig. 11(c), and the person - region mask shown in Fig. 11(d), as shown in Fig. 13(d), as the skin regions of the specific parts for each person, the facial skin regions 630A, 630B and the limb skin regions 631A, 631B, 632A, 632B for each person are identified. Also, the hair regions 633A, 633B for each person are identified. According to this, in addition to being able to distinguish the skin - hair image processing for each person, it is also possible to distinguish the image processing of the facial skin and the limb skin even for the same person. Therefore, since the image processing can be distinguished for each person and for each part of the same person, a more natural image can be obtained. Specifically, for each person, the image processing of the facial skin and the image processing of the limb skin can be made different. Also, even when a hand exists in front of the face, since the facial region and the hand region can be distinguished, it is possible to prevent performing the same processing as the face on the hand region.

[0105] Specifically, in the case of the photo sticker creation device 1, the number of people to be photographed is generally not one person but a plurality of people. Also, in the case of the photo sticker photographing device 1, there are a variety of poses of the user when taking a photo. Therefore, even if only the area around the face is photographed, there may be image data in which a hand exists in front of the face. In the image data photographed by the photo sticker photographing device 1, poses in which each person overlaps complexly are also preferred. Also, in the image data photographed by the photo sticker photographing device 1, a hand may exist in front of one's own or another person's face. In addition, in the photo sticker photographing device 1, image data is photographed in various poses such as a hand being extended forward or an arm or knee being bent. Therefore, it is not uncommon for the image data photographed by the photo sticker photographing device 1 to be unbalanced image data in which a part of the parts such as the arms, hands, legs, toes, and joints of the arms and legs of each person is unnaturally emphasized and appears large. Also, in the case of the photo sticker photographing device 1, compared with general photography, image data is photographed in a narrow space, so that a part of the parts of the person as described above is emphasized, and it is likely to become image data that looks unnatural as a whole.

[0106] As described above, by using the hair and skin mask, the specific part mask, and the person area mask, the skin area of the face and the skin area of the hands and feet of each person can be specified respectively. Thereby, for each person, necessary processing can be performed on the face or other skin areas. Also, the area of each part of each person can be specified, the fine parts of the skin, the unnatural parts that can be subjected to appropriate image processing can be specified, and image processing can be performed so that the unnatural parts look natural.

[0107] Note that the specification of the specific parts to be the image processing target for each person is a combination of the hair and skin mask and the part specification mask obtained by semantic segmentation and the person area mask obtained by instance segmentation, so it is a so-called panoptic segmentation technique.

[0108] The image processing unit 306 performs predetermined image processing on the target area of the image data specified by the specifying unit 305. For example, the image processing unit 306 may perform image processing to make the skin appear fair. Also, the image processing unit 306 may perform image processing on the skin area of the image data as if makeup has been applied. Specifically, it may execute processing such as applying eyeshadow around the eyes, drawing eyeliner, or applying blush on the cheeks. Further, the image processing unit 306 may perform processing on the hair area of the image data so that shine appears.

[0109] When the image processing unit 306 can distinguish between the face and the skin area other than the face, the processing of the skin area of the face and the processing of the skin area other than the face may be made different. For example, the color of the skin of the face may be brighter than the color of the skin of the hands and feet. In such a case, the color of the skin of the face may be adjusted to be brighter than the color of the skin of the hands and feet. Also, image processing may be performed to make only the face area appear smaller.

[0110] When the image processing unit 306 can distinguish the skin area for each person, the processing of the skin area for each person may be made different for each person. For example, it is common for the skin colors of multiple people to be different depending on their original skin color and the degree of sunburn at that time. Therefore, when the skin area for each person is distinguished, adjustments can be made according to the situation. For example, the skin color of a sunburned person may not be made as bright as the skin color of a person with fair skin.

[0111] When the image processing unit 306 can distinguish specific parts for each person, as shown in FIG. 14(a), when a specific part of a person appears unnatural compared to other parts, image processing may be performed on the unnatural-looking part to make it appear natural. Specifically, in FIG. 14(a), since one foot tip of the person on the right is in front of the torso, the foot tip part appears very unnatural. Therefore, image processing is performed to shrink the foot tip part so that it appears natural.

[0112] Specifically, the image processing unit 306 compares the size of each identified part of each person with the size of the face of each person. When the size of each part is outside the specified ratio range compared to the size of the face, the image processing unit 306 corrects it to an ideal value, considering it unnatural. As a result, the image processing unit 306 can correct areas that appear unnatural, for example, making overly long areas shorter. As shown in FIG. 14(a), in order to enable image processing to correct the tip of the foot to be smaller (shorter) when the tip of the foot appears large (long), in the above-mentioned specific part mask, the area of the tip of the foot and the area of the leg are distinguished and specified, so that image processing of the tip of the foot becomes possible.

[0113] Also, in the example shown in FIG. 14(b), the person on the right has their upper body receded and their legs bent, and the knee part is in front of the torso, so the knee part appears very unnatural and disproportionate. In such a case, the image processing unit 306 performs image processing to shrink the knee part so that it looks natural.

[0114] Specifically, the image processing unit 306 compares the size of each identified part of each person with the size of the face of each person and calculates the positional relationship of predetermined coordinates (for example, joints, tip parts (such as fingertips), the top of the head, etc.) of each part. When the size of each part is outside the specified ratio range compared to the size of the face, the image processing unit 306 corrects it to an ideal value, considering it unnatural. Also, when the distance between predetermined coordinates is outside the specified range, the image processing unit 306 corrects it to an ideal value, considering it unnatural. As a result, the image processing unit 306 can correct areas that appear unnatural, for example, making overly long areas shorter. As shown in FIG. 14(b), when the knee part bends and the area around the knee appears unnaturally large (thick), the image processing unit 306 performs image processing to make the area around the knee smaller (thinner) so that it looks natural.

[0115] Such an unnatural-looking depiction is caused by factors such as perspective. As described above, particularly in the photo sticker creation device 1, in addition to shooting image data with a variety of poses, image data is shot in a narrow shooting space, making it easy for such an unnatural-looking depiction to occur.

[0116] Note that the image processing unit 306 can combine a plurality of the above-described processes. For example, a plurality of image processes such as image processing for each person, image processing for skin color adjustment, image processing for makeup, image processing for hair gloss, and image processing for unnatural portions can be combined.

[0117] The composition unit 307 uses the processed image as a composition image, composes the composition image into the moving image captured by the camera 91, and causes the composed image to be displayed as a live view display image on the touch panel monitor 92. Therefore, the user can perform shooting while confirming the finished image in real time.

[0118] The editing unit 230 includes the above-described editing operation units 27A and 27B and an editing processing unit 308. The editing processing unit 308 performs editing processing by controlling the built-in monitor 131 and the speaker 133 in the tablet in the editing operation units 27A and 27B. The editing processing unit 308 receives an operation input from the user using the touch pens 132A and 132B with respect to the built-in monitor 131 of the tablet.

[0119] In addition, the editing processing unit 308 performs predetermined image processing on the captured image as the image to be edited according to a selection operation on the selection screen displayed on the built-in monitor 131 of the tablet, and displays it on the built-in monitor 131 of the tablet. Alternatively, the editing processing unit 308 performs predetermined image processing on the composition image according to an input operation on the editing screen displayed on the built-in monitor 131 of the tablet, or generates a new composition image according to the input operation, composes it into the captured image, and displays it on the built-in monitor 131 of the tablet.

[0120] In addition, the editing processing unit 308 controls the output of guidance for explaining how to proceed with the editing. For example, the editing processing unit 308 causes a guidance screen for explaining how to proceed with the editing to be displayed on the built-in monitor 131 of the tablet, or causes guidance voice for explaining how to proceed with the editing to be output from the speaker 133. Further, the editing processing unit 308 controls the communication unit 203 and performs processing related to communication via a network such as the Internet. Also, the editing processing unit 308 may perform printing processing by controlling the printer 140 of the printing operation unit 28.

[0121] The printing unit 240 includes the above-described printing operation unit 28 and a printing processing unit 309. The printing processing unit 309 receives printing data from the editing processing unit 308 and performs printing processing by controlling the printer 140 of the printing operation unit 28. Here, although the printing unit 240 that outputs a photo sticker as an example of the output unit for outputting image data has been described, the method of outputting image data is not limited to this. For example, a transmission means for transmitting the image data itself to an external communication terminal using a network or the like may be used as the output unit.

[0122] As described above, in exchange for charging, the photo sticker creation device 1 provides a photo sticker creation game in which various contrivances (such as selection of shooting poses, BGM, narration, etc.) are provided to boost the user's mood. Therefore, the created image by the photo sticker creation device 1 becomes an image that draws out the user's happy expression or a gorgeous image with elaborate taste.

[0123] In addition, the photo sticker creation device 1 is well-equipped with facilities such as writing, and since highly advanced techniques can be used for image deformation processing (such as the size of the subject's eyes and the length of the legs) and color correction (such as whitening processing of the subject's skin), the image created by the photo sticker creation device 1 becomes an image in which the user looks good.

[0124] In addition, compared to image processing outside the photo sticker creation device 1 (such as processing in an application for photo processing), editing (scribbling on the image) can be easily performed, and the variations of such editing are also rich. From this aspect as well, it can be said that the created image by the photo sticker creation device 1 is finished relatively splendidly compared to the images photographed and image-processed outside the photo sticker creation device 1.

[0125] (Flow of the photo sticker creation game) Next, the processing flow in which the user plays the photo sticker creation game on the photo sticker creation device 1 will be described with reference to FIG. 15. FIG. 15 is a flowchart showing the processing flow from the start of the game on the photo sticker creation device 1 to creating a photo sticker in the game.

[0126] In the state before the game starts, the control unit 201 that functions as the pre-selection processing unit 301 of the photo sticker creation device 1 causes the touch panel monitor 71 of the pre-selection operation unit 20 to display a message prompting the insertion of a coin. Further, as shown in FIG. 15, the control unit 201 determines whether a coin has been inserted into the coin insertion return port 73 based on the presence or absence of an activation signal from the coin processing unit 74 (S1). When the control unit 201 determines that no coin has been inserted into the coin insertion return port 73 (S1: NO), the control unit 201 continues the determination process of whether a coin has been inserted.

[0127] The user who intends to start the game inserts a coin into the coin insertion return port 73 in the pre-selection space A0, which is the space in front of the pre-selection operation unit 20. When a coin is inserted into the coin insertion return port 73, an activation signal instructing the start of the game is output from the coin processing unit 74. When the control unit 201 inputs the activation signal from the coin processing unit 74, the control unit 201 determines that a coin has been inserted into the coin insertion return port 73 (S1: YES) and executes pre-customer service processing for the user (S2).

[0128] In the pre-reception process, the control unit 201 causes the touch panel monitor 71 to display messages and the like that prompt the selection of a course, the input of a name, the selection of a design, and the like. When the user makes various selections or inputs according to the messages and the like displayed on the touch panel monitor 71, the control unit 201 sets the shooting course, name, design, print layout, and the like. The control unit 201 causes the touch panel monitor 92 to display a plurality of types of background images for composition so that the user can select them.

[0129] When the pre-reception process ends, the control unit 201 causes the touch panel monitor 71 to display a message or the like that prompts the user to move to the shooting space A1 and perform shooting. The control unit 201 functioning as the shooting processing unit 302 causes the touch panel monitor 92 of the shooting operation unit 21 to display a message that prompts the user to touch the screen. In addition, a start button may be displayed together with this message or instead of this message. Further, the control unit 201 causes the speaker 93 to output a narration that prompts the user to touch the screen together with the BGM. When the user who has moved to the shooting space A1 touches the touch panel monitor 92, the control unit 201 reads that the touch panel monitor 92 has been touched and starts the shooting process (S3).

[0130] In the shooting process, the control unit 201 causes the touch panel monitor 92 to display guidance regarding writing, for example, and prompts the user to select the writing level. When the user selects the writing level, the control unit 201 sets the writing level to the selected level.

[0131] In addition, the control unit 201 causes the touch panel monitor 92 to display an instruction screen for explaining the number of shots and outputs the corresponding narration from the speaker 93. In this embodiment, as an example, the number of shots is set to 7.

[0132] Next, the control unit 201 causes the touch panel monitor 92 to display an instruction screen for guiding the user to the standing position, and causes the corresponding narration to be output from the speaker 93.

[0133] After displaying the instruction screen as described above, the control unit 201 causes the touch panel monitor 92 to display, as a live view display image, a synthesized live view display image obtained by synthesizing the background image for synthesis and the moving image acquired by the camera 91 according to the shooting course selected by the user (S4). Thereby, the user can pose while checking the finished image.

[0134] The control unit 201 performs the live view display until immediately before the end of the countdown for shooting. During that time, the control unit 201 causes the touch panel monitor 92 to display a sample pose together with or instead of the live view display. The control unit 201 causes the speaker 93 to output a narration corresponding to the sample pose.

[0135] The control unit 201 manages the time from the start to the end of the live view display, and when a preset predetermined time has elapsed, performs a countdown using the display on the touch panel monitor 92 and the voice from the speaker 93.

[0136] At the end timing of the countdown, the control unit 201 transmits an illumination control signal to the upper strobe 82, the left strobe 83, the right strobe 84, and the foot strobe 85, and transmits a shutter signal to the camera 91.

[0137] Thereby, the upper strobe 82, the left strobe 83, the right strobe 84, and the foot strobe 85 irradiate flashes, and the camera 91 acquires a photographed image in which the illuminated user is shown together with the background. In the present embodiment, as an example, the processes from step S3 to step S4 are repeated a plurality of times to acquire seven photographed images. Further, the control unit 201 causes the storage unit 202 to store an edited target image obtained by synthesizing the background image for synthesis with the photographed image.

[0138] Here, an example has been described in which a photographed image is acquired at preset time intervals. However, the acquisition timing of the photographed image is not limited to this. For example, when the photographing operation unit 21 has an operation button for photographing, the photographed image may be acquired at the timing when this operation button is operated.

[0139] After the photographing is completed, the control unit 201 displays a guidance screen on the touch panel monitor 92 that prompts the user to move to either the editing space A2-1 or the editing space A2-2, and outputs a voice guidance for movement to the speaker 93.

[0140] Then, the control unit 201 executes an editing process that allows the user to edit the image to be edited (S5). More specifically, the control unit 201 displays the image to be edited on the built-in monitor 131 of the tablet, and allows the user to draw a stamp image, a pen image, etc. on this image to be edited using the touch pens 132A and 132B, and creates an edited image.

[0141] After that, the control unit 201 displays a guidance screen on the built-in monitor 131 of the tablet that prompts the user to move to the printing waiting space A3 where the sticker outlet 161 is provided, and outputs a voice guidance for movement to the speaker 133.

[0142] Furthermore, the control unit 201 executes a printing process of arranging the edited image in the printing layout selected by the preselection operation unit 20 to create a printing image, and printing this printing image on the sticker 142 (S6).

[0143] When the printing process is completed, the control unit 201 executes a discharging process of the sticker 142 (S7), discharges the printed sticker 142 from the sticker outlet 161, provides it to the user as a photo sticker, and ends the game. In this way, the photographed image of the user created by the photo sticker creating device 1 can be output as a photo sticker. Although detailed description is omitted, the photo sticker creating device 1 of the present embodiment can also output the photographed image to a mobile terminal or the like by communication.

[0144] (Details of the synthesis process) Next, with reference to the flowchart shown in FIG. 16, the synthesis process of step S4 in the flowchart of FIG. 15 will be described. As shown in FIG. 16, in the synthesis process of step S4, the control unit 201 receives the image data captured by the imaging unit 220 (S41).

[0145] The control unit 201 generates a group of regions from the image data received in step S41 (S42). The control unit 201 generates at least two or more groups of regions among the hair and skin region group, the specific part region group, and the person region group.

[0146] The control unit 201 specifies the target region for image processing according to the plurality of groups of regions generated in step S42 (S43).

[0147] The control unit 201 executes predetermined image processing on the target region of the image data specified in step S43 (S44).

[0148] The control unit 201 synthesizes an image using the image processed in step S44 (S46). The image synthesized here is acquired while being confirmed by the user through live view display.

[0149] (Display example) FIG. 17 shows an example of the editing screen 400 displayed on the tablet built-in monitor 131. As shown in FIG. 17, the editing screen 400 includes a first editing unit 401 and a second editing unit 402.

[0150] The first editing unit 401 includes a first thumbnail unit 403, a first image to be edited display unit 405, a first operation button display unit 407, and a first palette 409. Similarly, the second editing unit 402 includes a second thumbnail unit 404, a second image to be edited display unit 406, a second operation button display unit 408, and a second palette 410.

[0151] Thumbnails of a plurality of images to be edited Im1 to Im5 that have been photographed are displayed in the first thumbnail section 403 and the second thumbnail section 404. The thumbnails displayed in the first thumbnail section 403 and the second thumbnail section 404 can be selected with the touch pen 132A or 132B. In the example shown in FIG. 17, the images to be edited Im1 and Im3 are selected respectively. Although not shown, characters such as "being selected" may be displayed on the thumbnails of the selected images to be edited. Further, when seven images are photographed in step S3 of the flowchart in FIG. 15, seven images to be edited are also displayed in the thumbnail sections 403 and 404 respectively. However, FIG. 17 will be described by way of example in which five images to be edited are displayed for convenience.

[0152] The first image-to-be-edited display section 405 displays the image to be edited selected in the first thumbnail section 403. Further, the second image-to-be-edited display section 406 displays the image to be edited selected in the second thumbnail section 404. When the image to be edited selected in the thumbnail section 403 or 404 is changed, the image to be edited displayed in the image-to-be-edited display section 405 or 406 is also changed.

[0153] The first palette 409 is used to select the type of content, etc. when decorating the image to be edited displayed in the first image-to-be-edited display section 405 with content such as characters, patterns, colors, etc. Further, the second palette 410 is used to select the type of content, etc. when decorating the image to be edited displayed in the second image-to-be-edited display section 406 with content. As will be described in detail later, in the example shown in FIG. 17, each of the palettes 409 and 410 includes tabs for "cosplay", "scribbling on the face", "items", and "makeup".

[0154] As described above, in the present embodiment, by extracting each area, it becomes possible to execute the processing assumed for each area, and it is possible to realize image processing that is preferable for the user.

[0155] <<Modification Example 1>> In the above example, in the generation processing unit 304, a hair and skin mask is generated using the first learning model, a specific part mask is generated using the second learning model, and a person area mask is generated using the third learning model. In the specific part 305, it has been described that the processing area is specified using a plurality of masks necessary for image processing. However, the generation processing unit 304 may generate a plurality of masks using a learning model that has learned a plurality of types of area groups in advance. For example, the learning model used by the generation processing unit 304 has learned the relationship between image data as shown in FIG. 11(a), a hair and skin area group as shown in FIG. 11(b), a specific part area group as shown in FIG. 11(c), and at least two or more area groups selected from a person area group as shown in FIG. 11(d). Therefore, the generation processing unit 304 can generate a plurality of types of masks, which are each area group, from the image data according to this learned model. Here, when the learned model has learned the relationship between the image data and the plurality of area groups, the generation processing unit 304 may generate only the masks of the area group of the requested type. For example, the generation processing unit 304 generates a hair and skin area mask and a person area mask from the image data. Then, the specific part 305 specifies each area to be processed using the plurality of masks that are the area groups generated by the generation processing unit 304.

[0156] <<Modification Example 2>> In addition, the generation processing unit 304 may generate regions using a learned learning model of at least any one of the relationship of regions that distinguish the skin region of a person's face and other regions as shown in Fig. 13(a), the relationship of the skin region, hair region, and other regions for each person as shown in Fig. 13(b), the relationship of regions of specific parts for each person as shown in Fig. 13(c), and the relationship of the face skin, skin other than the face, and regions of specific parts for each person as shown in Fig. 13(d). In this case, the region itself generated by the generation processing unit 304 using the learning model becomes the target region to be processed. For example, a learned learning model of "skin and hair region group" and "person region group" generates a region group including, for each person, the skin region and other regions as shown in Fig. 13(b) from the input image data. Therefore, the specifying unit 305 specifies such a generated region group as the target region to be processed. That is, since the region itself generated by the generation processing unit 304 is the target region, the specifying unit 305 may be integrated with the generation processing unit 304.

[0157] Note that the combination of the skin and hair region group and the specific part region group uses semantic segmentation. On the other hand, the person region group uses instance segmentation. Therefore, in each of the above examples, the combination of the skin and hair region group and the person region group, the combination of the part specifying region group and the person region group, and the combination of the skin and hair region group, the part specifying region group, and the person region group use panoptic segmentation.

[0158] (Example of software implementation) The functional blocks of the photo sticker creation device 1 may be realized by a logic circuit (hardware) formed in an integrated circuit (IC chip) or the like, or may be realized by software using a CPU (Central Processing Unit).

[0159] In the latter case, the photo sticker creating device 1 includes a CPU that executes instructions of a program which is software for realizing each function, a ROM (Read Only Memory) or a storage device (collectively referred to as "recording medium") in which the program and various data are recordable in a readable manner by a computer (or CPU), a RAM (Random Access Memory) for expanding the program, and the like. Then, when the computer (or CPU) reads the program from the recording medium and executes it, the object of the present invention is achieved. As the recording medium, a "non-transitory tangible medium", for example, a tape, a disk, a card, a semiconductor memory, a programmable logic circuit, or the like can be used. Further, the program may be supplied to the computer via any transmission medium (such as a communication network or a broadcast wave) capable of transmitting the program. Note that the present invention can also be realized in the form of a data signal embedded in a carrier wave, in which the program is embodied by electronic transmission.

[0160] In the above description, the image processing apparatus has been described as the photo sticker creating device 1. However, it is not limited thereto. Therefore, even if the image processing apparatus is, for example, another device that does not create photo stickers, the same effect can be obtained. As an example, the image processing apparatus may be used as a device for taking image data in a photo studio.

[0161] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope shown in the claims. Embodiments obtained by appropriately combining technical means disclosed in different embodiments are also included in the technical scope of the present invention.

Explanation of Reference Numerals

[0162] 1 Photo sticker creating device (image processing apparatus) 301 Preselection processing unit 302 Shooting processing unit 303 Reception unit 304 Generation processing unit 305 Identification unit 306 Image processing unit 307 Composition unit 308 Editing processing unit 309 Printing processing unit (output unit)

Claims

1. A receiving unit that receives image data including a person, A first generation process that generates a specific part region group from the image data using a first learned model that has learned the relationship between learning image data including a person and the specific part region group, which is at least one or more regions of the person included in the learning image data, and a second generation process that generates a person region group from the image data using a second learned model that has learned the relationship between learning image data including a person and a person region group, which is at least one or more regions different from the specific part region group of the person included in the learning image data, and a generation processing unit that executes the above, A specifying unit that specifies a target region to be processed according to the specific part region group and the person region group from the image data, An image processing unit that performs predetermined image processing on the target region of the image data, The image data includes a plurality of persons, The specific part region group is at least the regions of the face and limbs of a person, The person region group is the region of each person included in the image data, The first generation process uses semantic segmentation, The second generation process uses instance segmentation, The specifying unit specifies at least the regions of the face and limbs for each person, An image processing apparatus.

2. The generation processing unit further executes a third generation process that generates a skin and hair region group including the skin, hair, and other regions of a person from the image data using a third learned model that has learned the relationship between image data including a person and the contour of the person included in the image data, The third generation process uses semantic segmentation, The specifying unit specifies, for each person, the face and the region of the skin other than the face as the target region according to the specific part region group, the person region group, and the skin and hair region group, The image processing apparatus according to Claim 1.

3. A receiving step of receiving image data including a person, A first generation step of generating a specific part region group from the image data using a first learned model that has learned the relationship between learning image data including a person and the specific part region group, which is at least one or more regions of the person included in the learning image data, Using a second learning model that has learned the relationship between the learning image data including a person and a group of person regions, which are at least one or more regions different from the specific part region group for the person included in the learning image data, to generate a group of person regions from the image data, which is a second generation step; A specifying step of specifying a target region to be processed according to the specific part region group and the person region group from the image data; An image processing step of performing predetermined image processing on the target region of the image data; including the image data includes a plurality of persons; the specific part region group is at least the regions of the face and limbs of a person; the person region group is the region of each person included in the image data; the first generation step utilizes semantic segmentation; the second generation step utilizes instance segmentation; the specifying step specifies at least the regions of the face and limbs for each person image processing method.

4. A computer, a receiving unit that receives image data including a person; a first generation processing unit that uses a first learning model that has learned the relationship between the learning image data including a person and a specific part region group, which is at least one or more regions for the person included in the learning image data, to generate a specific part region group from the image data; a second generation processing unit that uses a second learning model that has learned the relationship between the learning image data including a person and a group of person regions, which are at least one or more regions different from the specific part region group for the person included in the learning image data, to generate a group of person regions from the image data; a specifying unit that specifies a target region to be processed according to the specific part region group and the person region group from the image data; an image processing unit that performs predetermined image processing on the target region of the image data; to function as, the image data includes a plurality of persons; the specific part region group is at least the regions of the face and limbs of a person; the person region group is the region of each person included in the image data; the first generation processing unit utilizes semantic segmentation; the second generation processing unit utilizes instance segmentation; the specifying unit specifies at least the regions of the face and limbs for each person image processing program.

Citation Information

Patent Citations

  • Pattern extraction device

    JP2005222304A

  • Face image photographing method and device

    JP2006215947A

  • Image processing system and image processing method

    JP2016148933A

  • Information processing equipment, image area selection method, computer program, and storage media

    JP2019086899A

  • Photo-creating game machine, image processing method, and program

    JP2020053831A