Image processing device, image processing method and image processing program

The image processing apparatus uses learning models to generate and specify region groups, addressing the challenge of inaccurate region specification, resulting in improved image processing outcomes.

JP2025109961AActive Publication Date: 2025-07-25FURYU KK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025085613
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-25
Estimated Expiration
2041-05-17

AI Technical Summary

Technical Problem

Existing image processing technologies struggle to accurately specify desired regions, leading to unnatural image data processing outcomes.

Method used

An image processing apparatus and method utilizing first and second learning models to generate and specify region groups, enabling precise target region identification and subsequent processing.

Benefits of technology

Enables accurate and user-friendly image processing by specifying regions for targeted processing, enhancing the quality of image data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025109961000001_ABST
    Figure 2025109961000001_ABST
Patent Text Reader

Abstract

To provide favorable image processing for a user by specifying a region being the image processing object.SOLUTION: An image processing device comprises: a reception unit (303) which receives image data including a person; a generation processing unit (304) which executes first generation processing for generating a first region group from the image data by using a first learning model that has learned a relation between image data for learning and the first region group being at least one or more regions for the person included in the image data for learning, and second generation processing for generating a second region group from the image data by using a second learning model that has learned a relation between the image data for learning and the second region group being at least one or more regions different from the first region group for the person included in the image data for learning; a specification unit (305) which specifies the object region being the processing object according to the first and second region groups from the image data; and an image processing unit (306) which performs prescribed image processing on the object region of the image data.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an image processing method, and an image processing program for processing an image of a person.

Background Art

[0002] Conventionally, there is known a photo sticker creation apparatus that performs predetermined image processing on a photographed image and provides it to a user. Specifically, there is one that generates a plurality of mask images having different shapes for a predetermined part and specifies it as an object of image processing. (For example, see Patent Document 1).

[0003] Patent Document 1 describes, for example, recognizing a region of a predetermined color existing in the lower region of a face image as a lip, generating a mask image corresponding to the recognized lip portion, and generating a lip image.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In Patent Document 1, there is a possibility that a desired region cannot be accurately specified. In such a case, the image data obtained by image processing may become unnatural.

[0006] An object of the present invention is to solve the above problems and provide an image processing apparatus, an image processing method, and an image processing program that specify a specific region of a user to be an object of image processing and perform preferable image processing for the user, and a region specifying model and a model generation method used for these image processings.

Means for Solving the Problems

[0007] A first aspect of the image processing apparatus according to the present invention includes a reception unit that receives image data including a person, a first learning model that has learned the relationship between the learning image data including the person and a first region group that is at least one or more regions of the person included in the learning image data, a first generation processing unit that generates the first region group from the image data using the first learning model, a second learning model that has learned the relationship between the learning image data including the person and a second region group that is at least one or more regions different from the first region group of the person included in the learning image data, a second generation processing unit that generates the second region group from the image data using the second learning model, a specifying unit that specifies a target region to be processed according to the first and second region groups from the image data, and an image processing unit that performs predetermined image processing on the target region of the image data.

[0008] A first aspect of the image processing method according to the present invention includes a reception step of receiving image data including a person, a first generation step of generating a first region group from the image data using a first learning model that has learned the relationship between the learning image data including the person and a first region group that is at least one or more regions of the person included in the learning image data, a second generation step of generating a second region group from the image data using a second learning model that has learned the relationship between the learning image data including the person and a second region group that is at least one or more regions different from the first region group of the person included in the learning image data, a specifying step of specifying a target region to be processed according to the first and second region groups from the image data, and an image processing step of performing predetermined image processing on the target region of the image data.

[0009] A first aspect of the image processing program according to the present invention causes a computer to function as a reception unit that receives image data including a person, a first generation processing unit that generates a first region group from the image data using a first learning model that has learned the relationship between the learning image data including the person and a first region group that is at least one or more regions of the person included in the learning image data, a second generation processing unit that generates a second region group from the image data using a second learning model that has learned the relationship between the learning image data including the person and a second region group that is at least one or more regions different from the first region group of the person included in the learning image data, a specifying unit that specifies a target region to be processed according to the first and second region groups from the image data, and an image processing unit that performs predetermined image processing on the target region of the image data.

Advantages of the Invention

[0010] According to the present invention, it is possible to specify a region to be subjected to image processing and provide image processing favorable for a user.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Embodiments for Carrying Out the Invention

[0012] Hereinafter, specific embodiments of the present invention will be described with reference to the accompanying drawings. Note that members having the same function are denoted by the same reference numerals, and the description thereof will be omitted as appropriate. Furthermore, the shapes of the configurations shown in each drawing, as well as dimensions such as length, depth, and width, do not reflect the actual shapes and dimensions, and are appropriately changed for the sake of clarity and simplification of the drawings.

[0013] The image processing apparatus according to the present invention will be described as a photo sticker creating device as an example. The photo sticker creating device to which the present invention is applied is a game device that allows a user to perform shooting, editing, etc. as a game and provides the user with the captured and edited images as photo stickers or data. The photo sticker creating device 1 is installed, for example, in game centers, shopping malls, and stores at tourist spots.

[0014] In the game provided by the photo sticker creation device, the user uses the camera provided in the photo sticker creation device to take pictures of themselves or others. The user edits the captured image by synthesizing a foreground image and / or a background image, or by using editing functions (scribbling editing functions) such as pen image input or stamp image input as editing and synthesizing images, thereby designing the captured image to be rich in color. After the game ends, the user receives a photo sticker or the like on which the edited image is printed as a result. Alternatively, the photo sticker creation device provides the edited image to the user's mobile terminal, and the user can also receive the result via the mobile terminal.

[0015] (Configuration of the photo sticker creation device) FIGS. 1 and 2 are perspective views showing a configuration example of the appearance of the photo sticker creation device 1. The photo sticker creation device 1 is a game machine that provides captured images and edited images. The photo sticker creation device 1 provides an image to the user by printing the image on sticker paper or by transmitting the image to a server so that the image can be viewed on the user's mobile terminal. The users of the photo sticker creation device 1 are mainly young women such as junior high school girls and high school girls. In the photo sticker creation device 1, a plurality of users, mainly two or three people per group, can enjoy the game. Of course, in the photo sticker creation device 1, a single user can also enjoy the game.

[0016] In the photo sticker creation device 1, the user performs a shooting operation with himself / herself as the subject. The user synthesizes a synthetic image such as handwritten characters or a stamp image with the image selected from the captured images obtained by shooting through an editing operation. As a result, the captured image is edited into a magnificent image. The user receives a sticker paper on which the edited image, which is the edited image, is printed and ends a series of games.

[0017] As shown in FIG. 1, the photo sticker creation device 1 is basically configured by installing the shooting unit 11 and the editing unit 12 in contact with each other.

[0018] The photographing unit 11 is composed of a pre-selection operation unit 20, a photographing operation unit 21, and a background unit 22. The pre-selection operation unit 20 is installed on the side surface of the photographing operation unit 21. The space in front of the pre-selection operation unit 20 serves as a pre-selection space where pre-selection processing is performed. Also, the photographing operation unit 21 and the background unit 22 are installed at a predetermined distance apart. The space formed between the photographing operation unit 21 and the background unit 22 serves as a photographing space where photographing processing is performed.

[0019] As pre-selection processing, the pre-selection operation unit 20 performs guidance for introducing a game provided by the photo sticker creating device 1, or performs processing for allowing the user to select various settings in the photographing processing performed in the photographing space. The pre-selection operation unit 20 is provided with a coin insertion slot into which the user inserts money, a touch panel monitor used for various operations, and the like. The pre-selection operation unit 20 appropriately guides the user in the pre-selection space to the photographing space according to the availability status of the photographing space.

[0020] The photographing operation unit 21 is a device for photographing the user as a subject. The photographing operation unit 21 is positioned in front of the user who has entered the photographing space. On the front surface of the photographing operation unit 21 facing the photographing space, a camera, a touch panel monitor used for various operations, and the like are provided. Assuming that the left side surface is the left side surface and the right side surface is the right side surface when viewed from the user facing forward in the photographing space, the left side surface of the photographing operation unit 21 is constituted by a side panel 41A, and the right side surface is constituted by a side panel 41B. Further, the front surface of the photographing operation unit 21 is constituted by a front panel 42. It is assumed that the above-described pre-selection operation unit 20 is installed on the side panel 41A. Note that the pre-selection operation unit 20 may be installed on the side panel 41B, or may be installed on both of the side panels 41A and 41B.

[0021] The background part 22 is composed of a rear panel 51, a side panel 52A, and a side panel 52B. The rear panel 51 is a plate-shaped member located on the back side of the user facing the front. The side panel 52A is a plate-shaped member attached to the left end of the rear panel 51 and narrower in width than the side panel 41A. The side panel 52B is a plate-shaped member attached to the right end of the rear panel 51 and narrower in width than the side panel 41B.

[0022] The side panel 41A and the side panel 52A are provided in substantially the same plane. The upper parts of the side panel 41A and the side panel 52A are connected by a connecting part 23A which is a plate-shaped member. The lower parts of the side panel 41A and the side panel 52A are connected by a connecting part 23A' which is a member provided on the floor surface, for example, made of metal. Similarly, the side panel 41B and the side panel 52B are provided in substantially the same plane. The upper parts of the side panel 41B and the side panel 52B are connected by a connecting part 23B. The lower parts of the side panel 41B and the side panel 52B are connected by a connecting part 23B'.

[0023] Note that, for example, a green chroma key sheet is attached to the surface of the rear panel 51 on the shooting space side. The photo sticker creating device 1 performs chroma key compositing in the shooting process and the editing process by shooting with the chroma key sheet as the background. Thereby, the background image desired by the user is composited on the sheet part.

[0024] The opening formed by being surrounded by the side panel 41A, the connecting part 23A, and the side panel 52A serves as the entrance and exit of the shooting space. Also, the opening formed by being surrounded by the side panel 41B, the connecting part 23B, and the side panel 52B serves as the entrance and exit of the shooting space.

[0025] Above the imaging space, a ceiling surrounded by the front of the imaging operation unit 21, the connecting part 23A, and the connecting part 23B is formed. A ceiling strobe unit 24 is provided in a part of the ceiling. One end of the ceiling strobe unit 24 is fixed to the connecting part 23A, and the other end is fixed to the connecting part 23B. The ceiling strobe unit 24 incorporates a strobe that irradiates light into the imaging space according to imaging. Inside the ceiling strobe unit 24, in addition to the strobe, a fluorescent lamp is provided. Thus, the ceiling strobe unit 24 also functions as illumination for the imaging space.

[0026] The editing unit 12 is a device for editing the captured image. The editing unit 12 is connected to the imaging unit 11 such that one side surface is in contact with the front panel 42 of the imaging operation unit 21.

[0027] Assuming the configuration of the editing unit 12 shown in FIG. 1 as the front-side configuration, configurations used in the editing work are provided on each of the front side and the back side of the editing unit 12. With this configuration, two sets of users can perform the editing work simultaneously.

[0028] The front side of the editing unit 12 is composed of a surface 61 and an inclined surface 62 formed above the surface 61. The surface 61 is perpendicular to the floor surface and is substantially parallel to the side panel 41A of the imaging operation unit 21. On the inclined surface 62, as a configuration used in the editing work, a tablet-integrated monitor and a touch pen are provided. On the left side of the inclined surface 62, a columnar support part 63A that supports one end of the lighting device 64 is provided. On the right side of the inclined surface 62, a columnar support part 63B that supports the other end of the lighting device 64 is provided. On the upper surface of the support part 63A, a support part 65 that supports the curtain rail 26 is provided.

[0029] Above the editing unit 12, a curtain rail 26 is attached. The curtain rail 26 is composed of a combination of three rails 26A, 26B, and 26C. The three rails 26A, 26B, and 26C are combined so that their shape when viewed from above is U-shaped. One end of the rails 26A and 26B provided in parallel is fixed to the connecting part 23A and the connecting part 23B respectively, and the other ends of the rails 26A and 26B are joined to both ends of the rail 26C respectively.

[0030] A curtain is attached to the curtain rail 26 so that the space in front of the front of the editing unit 12 and the space in front of the back cannot be seen from the outside. The space in front of the front of the editing unit 12 and the space behind the back surrounded by the curtain become the editing space where the user performs editing work.

[0031] Also, as will be described later, a discharge port for discharging the printed sticker paper is provided on the left side surface of the editing unit 12. The space in front of the left side surface of the editing unit 12 becomes a printing waiting space where the user waits for the printed sticker paper to be discharged.

[0032] (Movement of the user) Here, the flow of the image creation game and the movement of the user accompanying it will be described. Fig. 3 is a diagram for explaining the spatial movement of the user during the photo sticker creation game.

[0033] First, the user inserts money into the coin slot in the pre-selection space A0, which is the space in front of the pre-selection operation unit 20. Next, the user makes various settings according to the screen displayed on the touch panel monitor. The user performs pre-selection operations such as selecting a course related to shooting performed in the shooting space, for example.

[0034] After finishing the pre-selection work, the user enters the shooting space A1 formed between the shooting operation unit 21 and the background unit 22 through the entrance G1 between the side panel 41A and the side panel 52A as indicated by the white arrow #1. Then the user performs shooting work using the camera, touch panel monitor, etc. provided on the shooting operation unit 21.

[0035] After the user finishes the shooting operation, as indicated by the white arrow #2, the user exits the shooting space A1 from the entrance G1 and moves to the editing space A2-1, or as indicated by the white arrow #3, exits the shooting space A1 from the entrance G2 and moves to the editing space A2-2.

[0036] The editing space A2-1 is the editing space on the front side of the editing unit 12. On the other hand, the editing space A2-2 is the editing space on the back side of the editing unit 12. Whether the user moves to the editing space A2-1 or the editing space A2-2 is guided by the screen display of the touch panel monitor of the shooting operation unit 21, etc. For example, the user is guided to move to the empty one of the two editing spaces. The user who has moved to the editing space A2-1 or the editing space A2-2 starts the editing work. The user in the editing space A2-1 and the user in the editing space A2-2 can perform the editing work simultaneously.

[0037] After the editing work is completed, the printing of the edited image is started. When the printing is started, the user who has finished the editing work in the editing space A2-1 moves from the editing space A2-1 to the printing waiting space A3 as indicated by the white arrow #4. Also, the user who has finished the editing work in the editing space A2-2 moves from the editing space A2-2 to the printing waiting space A3 as indicated by the white arrow #5.

[0038] The user who has moved to the printing waiting space A3 waits for the completion of the image printing. When the printing is completed, the user receives the seal paper discharged from the discharge port provided on the right side surface of the editing unit 12 and finishes a series of image creation games.

[0039] (Configuration of the pre-selection unit) Next, the configuration of each device will be described. FIG. 4 is a front view of the pre-selection operation unit 20.

[0040] A touch panel monitor 71 is provided above the pre-selection operation unit 20. The touch panel monitor 71 is composed of a monitor such as an LCD (Liquid Crystal Display) and a touch panel laminated thereon. The touch panel monitor 71 has a function of displaying various GUIs (Graphical User Interfaces) and receiving selection operations of users. On the touch panel monitor 71, a screen used for pre-selection processing is displayed to perform selection of a course related to shooting, selection of an image serving as a background in an edit target image to be edited, layout of a created image, selection of at least one of BGM (Back Ground Music), sound, and voice flowing during an image creation game, and input of a user's name.

[0041] As for the course related to shooting, a two-person course for two users to perform shooting and a large-group course for three or more users to perform shooting are prepared. Also, a couple course for a male and female couple to perform shooting may be prepared.

[0042] A speaker 72 is provided below the touch panel monitor 71. The speaker 72 outputs guiding voices, BGM, sound effects, etc. of the pre-selection processing. Also, a coin insertion and return slot 73 for users to insert coins is provided adjacent to the speaker 72.

[0043] (Configuration of the shooting unit) FIG. 5 is a front view of a shooting operation unit 21 as a shooting unit. The shooting operation unit 21 is configured to be surrounded by a side panel 41A, a side panel 41B, and a front panel 42.

[0044] A camera unit 81 is provided at the center of the front panel 42. The camera unit 81 is composed of a camera 91 and a touch panel monitor 92 as a display unit.

[0045] The camera 91 is, for example, a single-lens reflex camera and is attached inside the camera unit 81 so that the lens is exposed. The camera 91 has an imaging device such as a CCD (Charge Coupled Device) image sensor or a CMOS (Complementary Metal Oxide Semiconductor) image sensor, and photographs a user in the photographing space A1. The moving image captured by the camera 91 (hereinafter also referred to as a live view display image) is displayed on the touch panel monitor 92 in real time. The still image captured by the camera 91 at a predetermined timing such as when shooting is instructed is saved as a photographed image.

[0046] The touch panel monitor 92 is provided below the camera 91. The touch panel monitor 92 is composed of a monitor such as an LCD and a touch panel laminated thereon. The touch panel monitor 92 has a function as a live view monitor for displaying the moving image captured by the camera 91 and a function for displaying various GUIs and accepting the user's selection operations. Specific examples of the selection operations include selection of a shooting course, instructions for starting and ending shooting (control instructions for shooting), selection of the degree of eye distortion and color, and the degree of correction of skin color, selection of an image to be a background image in the created image, and selection of BGM (sound / audio) during an image creation game. The moving image (live view image) and still image (photographed image) captured by the camera 91 are displayed on the touch panel monitor 92.

[0047] Above the camera unit 81, an upper strobe 82 with a curved light emitting surface facing the user is installed. The upper strobe 82 irradiates light onto the user's face and upper body from directly above the front.

[0048] Also, below the camera unit 81, a foot strobe 85 that irradiates light onto the user's lower body and feet is provided. Note that the upper strobe 82 and the foot strobe 85 include a strobe and a fluorescent lamp.

[0049] Although not shown in FIGS. 1 and 5, a speaker 93 is provided, for example, near the ceiling of the front panel 42. The speaker 93 outputs guidance voices for shooting processing, BGM, sound effects, etc. according to the voice signal output from the control unit 201.

[0050] (Configuration of the background part) FIG. 6 is a front view of the shooting space A1 side of the background part 22.

[0051] An upper-back flash 101 is installed above the rear panel 51. The upper-back flash 101 irradiates the user with light from above the back.

[0052] In the figure, a left-back flash 102 is installed to the left of the rear panel 51. The left-back flash 102 irradiates the user from the right side of the back. In the figure, a right-back flash 103 is installed to the right of the rear panel 51. The right-back flash 103 irradiates the user from the left side of the back.

[0053] Also, a chroma key sheet 121 may be attached to the surface of the rear panel 51 on the shooting space A1 side (the front side in the figure). The color of the chroma key sheet 121 is, for example, green.

[0054] Although not shown, a chroma key sheet may also be attached to the surfaces of the side panels 52A and 52B on the shooting space A1 side, similar to the chroma key sheet 121.

[0055] (Configuration of the editing unit) FIG. 7 is a front view of the editing space A2-1 side of the editing unit 12.

[0056] Almost at the center of the inclined surface 62, a tablet-built-in monitor 131 is provided. On the left side of the tablet-built-in monitor 131, a touch pen 132A is provided. On the right side of the tablet-built-in monitor 131, a touch pen 132B is provided. The tablet-built-in monitor 131 is configured by being provided so that the tablet exposes the display. The tablet enables operation input using the touch pen 132A or the touch pen 132B. For example, an editing screen used for an editing operation is displayed on the tablet-built-in monitor 131. For example, when two users perform an editing operation simultaneously, the touch pen 132A is used by the user on the left side facing the tablet-built-in monitor 131, and the touch pen 132B is used by the user on the right side facing the tablet-built-in monitor 131.

[0057] Figure 8 is a left side view of the editing unit 12.

[0058] A sticker sheet discharge port 161 is provided on the lower side of the left side surface of the editing unit 12. A printer as an output unit is provided inside the editing unit 12. By that printer, an image of the user in the editing space A2-1 or an image of the user in the editing space A2-2 is printed on a sticker sheet in a predetermined layout and discharged from the sticker sheet discharge port 161.

[0059] (Internal Configuration of the Photo Sticker Creation Device) Figure 9 is a block diagram showing a configuration example inside the photo sticker creation device 1. In Figure 9, the same components as those described above are labeled with the same reference numerals. Redundant descriptions will be omitted as appropriate.

[0060] The control unit 201 is composed of a CPU (Central Processing Unit) or the like. The control unit 201 executes programs stored in the ROM (Read Only Memory) 206 and the storage unit 202, and controls the overall operation of the photo seal creating device 1. The storage unit 202, the communication unit 203, the drive 204, the ROM 206, and the RAM (Random Access Memory) 207 are connected to the control unit 201. Each component of the pre-selection operation unit 20, the shooting operation unit 21, the background unit 22, the editing operation units 27A and 27B, and the printing operation unit 28 is also connected to the control unit 201.

[0061] The storage unit 202 is a non-volatile recording medium such as a hard disk or a flash memory. The storage unit 202 stores various setting information and the like supplied from the control unit 201. The information stored in the storage unit 202 is appropriately read by the control unit 201.

[0062] The communication unit 203 is an interface for a network such as the Internet. The communication unit 203 communicates with an external device according to the control by the control unit 201. For example, the communication unit 203 transmits a photographed image or an edited image selected by the user to the server. The image transmitted from the communication unit 203 is stored in a predetermined storage area in the server and is displayed or downloaded to a mobile terminal that has accessed the server.

[0063] A removable medium 205 made of an optical disk, a semiconductor memory, or the like is appropriately mounted on the drive 204. Programs and data read from the removable medium 205 by the drive 204 are supplied to the control unit 201 and are stored or installed in the storage unit 202.

[0064] The ROM 206 stores programs and data executed in the control unit 201. The RAM 207 temporarily stores data and programs processed by the control unit 201.

[0065] The pre-selection operation unit 20 implements pre-selection processing for users in the pre-selection space A0. The pre-selection operation unit 20 includes a touch panel monitor 71, a speaker 72, and a coin processing unit 74.

[0066] The touch panel monitor 71 displays various selection screens according to the control by the control unit 201 and receives the operations of the user on the selection screens. The input signal representing the content of the user's operation is supplied to the control unit 201, and various settings are performed.

[0067] The coin processing unit 74 detects the insertion of coins into the coin insertion / return port 73. When the coin processing unit 74 detects that coins of a predetermined amount have been inserted, it outputs a start signal instructing the start of the game to the control unit 201.

[0068] The shooting operation unit 21 implements shooting processing for users in the shooting space A1. The shooting unit 220 includes an upper strobe 82, a left strobe 83, a right strobe 84, a foot strobe 85, a camera 91, a touch panel monitor 92, and a speaker 93.

[0069] The upper strobe 82, the left strobe 83, the right strobe 84, and the foot strobe 85 are arranged in the shooting space A1 and emit light according to the illumination control signal supplied from the control unit 201.

[0070] The camera 91 performs shooting according to the shutter control by the control unit 201 and outputs the captured image (image data) obtained by the shooting to the control unit 201.

[0071] The editing operation unit 27A implements editing processing for users in the editing space A2-1. The editing operation unit 27A includes a tablet built-in monitor 131, touch pens 132A, 132B, and a speaker 133. The editing operation unit 27B implements editing processing for users in the editing space A2-2 and has the same configuration as the editing operation unit 27A. Hereinafter, when the editing operation units 27A and 27B are not particularly distinguished, they are simply referred to as the editing operation unit 27.

[0072] The tablet-integrated monitor 131 displays an editing screen according to the control by the control unit 201 and receives the user's operations on the editing screen. An input signal representing the content of the user's operations is supplied to the control unit 201, and the captured image to be edited is edited.

[0073] The printing operation unit 28 realizes a printing process of providing the user in the printing waiting space A3 with the sticker paper on which the created image has been printed. The printing operation unit 28 is configured to include a printer 140. A sticker paper unit 141 is attached to the printer 140.

[0074] Based on the print data supplied from the control unit 201, the printer 140 prints the edited image on the sticker paper 142 stored in the sticker paper unit 141 and discharges it to the sticker paper discharge port 161.

[0075] (Functional blocks of the photo sticker creation device) FIG. 10 is a block diagram showing the functional blocks of the photo sticker creation device 1. The photo sticker creation device 1 functions as a pre-selection unit 210, a photographing unit 220, an editing unit 230, and a printing unit 240. Further, by executing the photo sticker creation program of the present invention, the control unit 201 functions as a pre-selection processing unit 301, a photographing processing unit 302, a reception unit 303, a generation processing unit 304, a specifying unit 305, an image processing unit 306, a composition unit 307, an editing processing unit 308, and a printing processing unit 309.

[0076] The pre-selection unit 210 includes the above-described pre-selection operation unit 20 and a pre-selection processing unit 301. The pre-selection processing unit 301 performs pre-selection processing by controlling the touch panel monitor 71, the speaker 72, and the coin processing unit 74 in the pre-selection operation unit 20. The pre-selection processing unit 301 causes the touch panel monitor 71 to display a selection screen or the like for selecting a course related to shooting performed in the shooting space A1. Further, the pre-selection processing unit 301 receives an operation input from the user with respect to the touch panel monitor 71. Specifically, the pre-selection processing unit 301 receives a selection operation input with respect to the selection screen displayed on the touch panel monitor 71, an input of the user's name, and the like. Further, the pre-selection processing unit 301 controls the output of guidance for explaining various selection operations. The pre-selection processing unit 301 causes the touch panel monitor 71 to display a screen for explaining various selection operations, or causes the speaker 72 to output a voice for explaining various selection operations.

[0077] The shooting unit 220 includes the above-described shooting operation unit 21 and a shooting processing unit 302. The shooting processing unit 302 performs shooting processing by controlling the camera 91, the touch panel monitor 92, and the speaker 93 in the shooting operation unit 21.

[0078] The shooting processing unit 302 receives an operation input from the user with respect to the touch panel monitor 92. For example, the shooting processing in the shooting space A1 is started when the shooting processing unit 302 receives, as an input, a contact operation on the touch panel monitor 92 by the user.

[0079] The shooting processing unit 302 controls the camera 91 and shoots the user as a subject. There are moving images and still images for shooting. The shooting processing unit 302 controls the display of the touch panel monitor 92 to cause the moving image captured by the camera 91 to be displayed as a live view on the touch panel monitor 92, or to cause the still image that is the shooting result to be displayed as a shot image.

[0080] In addition, the photographing processing unit 302 causes the touch panel monitor 92 to display an instruction screen for explaining the number of photographed images, the standing position of the user, the sample pose, the message about the line of sight, the photographing timing, and the like. Further, the voice of the narration corresponding to each instruction screen and the BGM are output from the speaker 93.

[0081] The reception unit 303 receives the image data obtained by photographing the user by the photographing unit 220.

[0082] The generation processing unit 304 generates a group of regions using the learning data including a person and the learned model in which the regions about the person have been learned. Specifically, the generation processing unit 304 generates at least two or more of the “skin and hair mask”, the “specific part mask”, and the “person region mask” which are groups of regions.

[0083] (Skin and hair mask) First, the generation process of the skin and hair mask will be described. The generation processing unit 304 uses the learned skin and hair learning model for the relationship between the learning image data including a person and the skin and hair region group which is at least one or more regions about the person included in the learning image data, and generates the skin and hair region group from the image data received by the reception unit 303. Specifically, the generation processing unit 304 generates a skin and hair mask in which at least the regions of the skin and hair of the person are separated from other regions. For example, the generation processing unit 304 generates a skin and hair mask including a skin region 500, a hair region 501, and other regions 502 as shown in FIG. 11(b) from the image data as shown in FIG. 11(a). The skin region 500 is the region of the skin of the person exposed from the clothes, specifically, the regions such as the face, neck, and limbs. Further, the hair region 501 is specifically the region of the hair of the person. The skin region 500 and the hair region 501 are combined here as the skin and hair region group. Note that, for example, as shown in FIG. 11(b), the other regions in the skin and hair mask include the regions of the lips, eyes, and eyebrows. Also, in the example shown in FIG. 11(b), although the other regions 502 and the background region 503 are distinguished, the background may also be included in the other regions.

[0084] As shown in an example in FIG. 12(a), the skin and hair learning model M is a model that has been pre-trained by a learning device on the relationship between a plurality of sets of training image data and correct answer data as training data. Specifically, the training image data includes people. Also, the correct answer data includes the skin area, hair area, and other areas of the people included in the training image data. The skin and hair learning model M obtained by learning the relationship between a plurality of sets of training image data and correct answer data by a learning device can output the skin area, hair area, and other areas of the people included in the image data for image data including people, as shown in FIG. 12(b).

[0085] The generation processing unit 304 generates a skin and hair mask, for example, using semantic segmentation. By using semantic segmentation, fine-grained area extraction at the pixel level becomes possible, and a group of areas can be generated at the pixel level. Thereby, in the image processing unit 306 described later, image processing at the pixel level can be realized.

[0086] (Specific part mask) Subsequently, the generation process of the specific part mask will be described. The generation processing unit 304 generates a specific part area group from the image data using a part learning model that has learned the relationship between the training image data including people and the specific part area group that is at least one or more areas for the people included in the training image data. The specific part area group is a group of areas different from the skin and hair area group. Specifically, the generation processing unit 304 generates a specific part mask that separates specific parts of a person, for example, the areas of the head and limbs from other areas. For example, the generation processing unit 304 generates a specific part mask including a head area 510, limb areas 511, 512, and other areas 513, 514 as shown in FIG. 11(c) from the image data as shown in FIG. 11(a). The head area 510 and the limb areas 511, 512 together are regarded as the specific part area group here.

[0087] In the example shown in FIG. 11(c), the arm part 511, which is a limb region, is a part that includes the sleeve part, which is a part of the clothing, and the part from the sleeve to the fingertips that are exposed from the sleeve. Also, the leg part 512, which is a limb region, is the part of the leg that is exposed from the clothing. Although not included in FIG. 11(c), the leg region can include up to the toes. At this time, it is assumed that shoes or sandals are worn on the feet, but the leg part includes up to the toes of such shoes or the like. Also, in the example shown in FIG. 11(c), in the other regions, as the torso, an example is shown where the body part 513 of the user's upper garment and the skirt part 514, which is the lower garment, are detected as separate regions, but the body part of the upper garment and the lower garment may be integrated. The division of each region in the specific part mask is not limited to the example shown in FIG. 11(c), and it is sufficient if at least the head region and the limb regions to be processed by the image processing unit 306 described later can be specified. Also, if a person is wearing pants that fit the legs instead of a skirt as the lower garment, the region of the pants may be the leg part of the limb region. And, as shown in FIG. 11(c), the neck region is preferably not included in the head region 510 but is included in the torso region 513. This is because it is assumed that image processing that is not to be performed is required for the face region in the neck region. Also, in the case of performing image processing on the region from the neck to the décolletage region that is exposed from the clothing, a specific region group may be generated for the neck and the décolletage region as regions different from the torso of the clothing. Here, how the generation processing unit 304 specifies regions from the image data depends on what correct data of the specific part region group is used for the learning process together with the learning image data. Also, what correct data to use can be determined by what image processing is to be performed by the image processing unit 306 described later. Note that in the example shown in FIG. 11(c), the other region 515 is distinguished from the background region, but the background may also be included in the other regions 513, 514. That is, it is not essential to distinguish the regions 513 to 515 shown in FIG. 11(c).

[0088] Although the description using the illustration is omitted, the part learning model is generated by having the learning device learn the relationship between the learning image data including a person and specific part regions such as the face region and the limb regions of the person which are the correct answer data, as in the example described above using FIG. 12(a). The part learning model thus obtained can output the regions of specific parts of the person included in the image data for the image data including a person. In the example shown in FIG. 11(c), it is an example of a part specific mask generated using a part learning model learned with correct answer data that divides the head region 510 into the regions of the face and hair, the body region 513 into the region from the neck to the hem of the upper garment, the arm region 511 into the region from the upper arm to the fingertips, the lower garment region 514, and the leg region 512 below the hem of the lower garment.

[0089] The generation processing unit 304 generates a specific part mask using, for example, semantic segmentation. By using semantic segmentation, fine region extraction at the pixel unit becomes possible, and a group of regions can be generated at the pixel unit. Thereby, in the image processing unit 306 described later, image processing at the pixel unit can be realized.

[0090] (Person region mask) Subsequently, the generation process of the person region mask will be described. The generation processing unit 304 uses a region learning model that has learned the relationship between the learning image data including a person and the person region group which is the region of the person included in the learning image data, to generate a person region group from the image data received by the reception unit 303. Specifically, the generation processing unit 304 generates a person region mask including the region for each person. For example, as shown in FIG. 11(a), even when a plurality of persons are included in the image data and the hair and arms of the plurality of persons overlap, as shown in FIG. 11(d), a person region mask divided into regions 520, 521 for each person is generated. The regions of all the persons included in this image data are combined to form a person region group. Incidentally, if only one person is included in the image data, the region of that one person is taken as the person region group.

[0091] Although the description using the illustration is omitted, the region learning model is generated by having a learning device learn the relationship between learning image data including a person and the person region which is the correct answer data, as in the example described above with reference to Fig. 12(a). The region learning model thus obtained can output the person region included in the image data for the image data including a person. For example, when the image data includes a plurality of persons, the region learning model can specify and output a region for each person.

[0092] Note that the learning image data used for generating the skin and hair learning model, the part learning model, and the region learning model does not necessarily have to be the same, but by using the same image data for simultaneous learning, high accuracy can be obtained when generating each region group in each learning model.

[0093] The generation processing unit 304 generates a person region mask, for example, by using instance segmentation. By using instance segmentation, the region of each object can be classified in pixel units. As a result, pixel-level image processing can be realized in the image processing unit 306 described later.

[0094] The specifying unit 305 specifies a target region to be subjected to image processing from the image data according to two or more region groups generated by the generation processing unit 304.

[0095] (Specification of the target region using the skin and hair mask and the specific part mask) The specific part 305 specifies the target area 600 of the skin of a person's face separately from other areas 601 according to the group of skin and hair areas indicated by the skin and hair mask and the specific part area indicated by the specific part mask. Specifically, as shown in Fig. 13(a), the area 600 of the skin of the person's face is specified using the hair and skin area mask shown in Fig. 11(b) and the specific area mask shown in Fig. 11(c). For example, the example of the image data shown in Fig. 11(a) is an image taken with a hand in front of the face of the person on the left and the hand overlapping the face. As shown in Fig. 13(a), by specifying the area of the skin of the face separately from the area of the skin of the hand, it is possible to target only the face for image processing. Also, in the example shown in Fig. 13(a), the area 602 of the person's hair is also specified, and it is possible to target it separately from the area 600 of the skin of the person's face and other areas 601 for image processing.

[0096] For example, when performing image processing such as makeup, if a hand overlaps the position where lipstick, lip gloss, cheek, eyeshadow, eyeliner, eyebrow, etc. are applied and appears in the image, there is a risk of obtaining an unnatural image in which the color applied to the face is also applied to the hand due to image processing. Also, although it is not uncommon for the face and hands to appear to have different skin colors, there is a risk of obtaining an unnatural image in which the face and hands are subjected to the same processing (e.g., whitening processing). On the other hand, as shown in Fig. 13(a), by specifying the area of the skin of the face, it is possible to perform image processing only on the face part and prevent the same image processing as the face from being performed on other parts, thereby obtaining a natural image.

[0097] Specifically, in the case of the photo sticker creating device 1, there are a variety of variations in the poses of the users when taking pictures. Therefore, even if only the area around the face is photographed, there may be image data in which a hand is present in front of the face. For such image data, if an area including the face area and the hand area is extracted as the skin area, as described above, even when processed on the hand, since the hand area and the face area are distinguished, it is possible to perform only the necessary processing on the face area.

[0098] (Specification of the target area using the skin and hair mask and the person area mask) The specific part 305 identifies the target skin areas 610A, 610B and hair areas 611A, 611B for each person according to the skin and hair area group indicated by the skin and hair mask and the person area group indicated by the person area mask. Specifically, using the skin and hair area mask shown in FIG. 11(b) and the person area mask shown in FIG. 11(d), as shown in FIG. 13(b), the skin areas 610A, 610B and hair areas 611A, 611B for each person are identified. For example, when a plurality of people are included in the image data, the skin color and hair color of each person may be different. In such a case, it is not preferable to perform image processing on the skin colors and hair colors of the plurality of people included in the image in the same method. If the hair colors of two people are clearly different, if the hair of the two people is processed in the same way for image processing, there is a risk of obtaining an unnatural image. Also, for example, when the skin colors of two people are different, if the skin of the two people is processed in the same way for image processing, there is a risk of obtaining an unnatural image. On the other hand, as shown in FIG. 13(b), by identifying the skin areas 610A, 610B and hair areas 611A, 611B for each person, it becomes possible to realize image processing for the skin of each person and image processing for the hair of each person, so that a natural image with appropriate image processing for each person can be obtained.

[0099] Specifically, in the case of the photo sticker creating device 1, the number of people to be photographed is generally not one person but a plurality of people. Also, in the case of the photo sticker photographing device 1, there are a variety of variations in the poses of the users when photographing. For a simple group photo, the parts such as the arms, hands and legs of each person can be identified as the arms, hands and legs at a position "close to the body". However, with the photo sticker photographing device 1, image data can be photographed in various poses. Therefore, it is difficult to identify the arms, hands and legs of each person based on the criterion of "a position close to the body". Therefore, as described above, by identifying the skin area of each person's area based on the generated person area mask, the skin areas of the people can be distinguished and appropriate natural image processing can be performed for each person.

[0100] Note that the identification of the skin and hair regions for each person is based on the combination of the skin and hair mask obtained by semantic segmentation and the person region mask obtained by instance segmentation, so it is a so-called panoptic segmentation technique.

[0101] (Identification of the target region using the specific part mask and the person region mask) The specific part 305 identifies the region of the specific part for each person according to the group of specific part regions indicated by the specific part mask and the group of person regions indicated by the person region mask. For example, as shown in FIG. 11(c), the specific part 305 uses the specific region mask shown in FIG. 11(c) and the person region mask shown in FIG. 11(d) to identify the regions 620A, 620B, 621A, and 621B of the limbs of the person as shown in FIG. 11(c). Further, as shown in FIG. 11(c), the specific part 305 identifies the regions 622A, 622B of the head of the person, the regions 623A, 623B of the body of the upper garment, and the regions 624A, 624B of the lower garment. As an example, assume that at the time of shooting the image data, a specific part of the person protrudes forward compared to the torso. In this case, due to the perspective relationship, in the image data, as shown in FIG. 14, a part of the person's body (in the case of FIG. 14(a), the toes of the person on the right side; in the case of FIG. 14(b), the lower body (especially the knees) of the person on the right side) may appear larger than the original size. Since such an image looks unbalanced and unnatural, it may be necessary to perform image processing on the part that appears larger than normal so that it looks balanced and natural. At this time, as shown in FIG. 13(c), by identifying which part is the specific part of which person, it is possible to realize the image target. Thereby, for example, a natural image obtained by performing image processing on an unbalanced part can be obtained.

[0102] Specifically, in the case of the photo sticker creation device 1, the number of people to be photographed is generally not one person but a plurality of people. Also, in the case of the photo sticker photographing device 1, there are a rich variety of poses of the user when taking a photo. For example, in the case of a simple group photo, it is common to be photographed standing still, but with the photo sticker photographing device 1, image data is photographed in various poses such as hands being extended forward, arms and knees being bent. Therefore, it is not uncommon for the image data photographed by the photo sticker photographing device 1 to be unbalanced image data in which some parts of the parts of each person's arms, hands, legs, feet, joints of the arms and legs, etc. are unnaturally emphasized and appear large. Also, in the case of the photo sticker photographing device 1, compared with general photo shooting, since image data is photographed in a narrow space, it is likely to become image data in which some parts of the above-mentioned parts of the person are emphasized and appear unnatural as a whole. Therefore, as described above, it is possible to identify unnatural parts using the areas of specific parts of each person identified using the specific part mask and the person area mask, and perform image processing so that the unnatural parts look natural.

[0103] Note that the identification of the areas of specific parts for each person in this way is a so-called panoptic segmentation technique because it is a combination of a part identification mask obtained by semantic segmentation and a person area mask obtained by instance segmentation.

[0104] (Identification of target areas using a hair and skin mask, a specific part mask, and a person area mask) The specific part 305 specifies a specific part to be subjected to image processing for each person according to the group of skin and hair regions indicated by the skin and hair mask, the group of specific part regions indicated by the specific part mask, and the group of person regions indicated by the person region mask. Specifically, using the skin and hair mask shown in FIG. 11(b), the specific part mask shown in FIG. 11(c), and the person region mask shown in FIG. 11(d), as shown in FIG. 13(d), as the skin regions of the specific parts for each person, the facial skin regions 630A and 630B and the limb skin regions 631A, 631B, 632A, and 632B for each person are specified. Also, the hair regions 633A and 633B for each person are specified. According to this, in addition to being able to distinguish the image processing of the skin and hair for each person, it is also possible to distinguish the image processing of the facial skin and the limb skin even for the same person. Therefore, since the image processing can be distinguished for each person and for each part of the same person, a more natural image can be obtained. Specifically, for each person, the image processing of the facial skin and the image processing of the limb skin can be made different. Also, even when a hand exists in front of the face, since the facial region and the hand region can be distinguished, it is possible to prevent performing the same processing as the face on the hand region.

[0105] Specifically, in the case of the photo sticker creation device 1, the number of people to be photographed is generally not one person but a plurality of people. Also, in the case of the photo sticker photographing device 1, there are a variety of poses of the user when taking a photo. Therefore, even if only the area around the face is photographed, there may be image data in which a hand exists in front of the face. In the image data photographed by the photo sticker photographing device 1, poses in which each person overlaps complexly are also preferred. Also, in the image data photographed by the photo sticker photographing device 1, a hand may exist in front of one's own or another person's face. In addition, in the photo sticker photographing device 1, image data is photographed in various poses such as a hand being extended forward, an arm or a knee being bent. Therefore, it is not uncommon for the image data photographed by the photo sticker photographing device 1 to be unbalanced image data in which a part of the parts such as the arms, hands, legs, toes, and joints of the arms and legs of each person is unnaturally emphasized and appears large. Also, in the case of the photo sticker photographing device 1, compared with general photo shooting, image data is photographed in a narrow space, so it is likely to become image data in which a part of the parts of the person as described above is emphasized and looks unnatural as a whole.

[0106] As described above, by using the hair and skin mask, the specific part mask, and the person area mask, the skin area of the face and the skin area of the hands and feet of each person can be specified respectively. Thereby, for each person, necessary processing can be performed on the face or other skin areas. Also, the area of each part of each person can be specified, the fine parts of the skin can be identified, the unnatural parts that can be subjected to appropriate image processing can be identified, and image processing can be performed so that the unnatural parts look natural.

[0107] Note that the specification of the specific part to be the image processing target for each person is a combination of the hair and skin mask and the part specification mask obtained by semantic segmentation and the person area mask obtained by instance segmentation, so it is a so-called panoptic segmentation technique.

[0108] The image processing unit 306 performs predetermined image processing on the target area of the image data specified by the specifying unit 305. For example, the image processing unit 306 may perform image processing to make the skin appear fair. Also, the image processing unit 306 may perform image processing on the skin area of the image data as if makeup has been applied. Specifically, it may execute processing such as applying eyeshadow around the eyes, drawing an eyeliner, or applying blush on the cheeks. Further, the image processing unit 306 may perform processing on the hair area of the image data so that gloss appears.

[0109] When the image processing unit 306 can distinguish between the face and the skin areas other than the face, the processing of the facial skin area and the processing of the skin areas other than the face may be made different. For example, the color of the facial skin may be brighter than the color of the skin of the hands and feet. In such a case, the color of the facial skin may be adjusted to be brighter than the color of the skin of the hands and feet. Also, image processing may be performed to make only the face area appear smaller.

[0110] When the image processing unit 306 can distinguish the skin areas for each person, the processing of the skin areas for each person may be made different for each person. For example, it is common for the skin colors of multiple people to be different depending on their original skin color and the degree of sunburn at that time. Therefore, when the skin areas for each person are distinguished, adjustments can be made according to the situation. For example, the skin color of a sunburned person may not be made as bright as the skin color of a person with fair skin.

[0111] When the image processing unit 306 can distinguish specific parts for each person, as shown in FIG. 14(a), when a specific part of a person appears unnatural compared to other parts, image processing may be performed to make the unnatural-looking part appear natural. Specifically, in FIG. 14(a), since one of the feet of the person on the right is in front of the torso, the foot appears very unnatural. Therefore, image processing is performed to shrink the foot so that it appears natural.

[0112] Specifically, the image processing unit 306 compares the size of each identified part of each person with the size of the face of each person. When the size of each part is outside the specified ratio range compared to the size of the face, the image processing unit 306 corrects it to an ideal value, considering it unnatural. Thereby, the image processing unit 306 can correct the unnatural-looking parts, for example, making the overly long parts shorter. As shown in FIG. 14(a), in order to enable image processing to correct the overly large (long) looking toe to be smaller (shorter), in the above-described specific part mask, the toe part area and the leg part area are distinguished and specified, enabling image processing of the toe part.

[0113] Also, in the example shown in FIG. 14(b), the person on the right has their upper body receded and their legs bent, with the knee part in front of the torso, so the knee part appears very unnatural. In such a case, the image processing unit 306 performs image processing to shrink the knee part so that it looks natural.

[0114] Specifically, the image processing unit 306 compares the size of each identified part of each person with the size of the face of each person and calculates the positional relationship of the predetermined coordinates (for example, joints, tip parts (such as fingertips), the top of the head, etc.) of each part. When the size of each part is outside the specified ratio range compared to the size of the face, the image processing unit 306 corrects it to an ideal value, considering it unnatural. Also, when the distance between the predetermined coordinates is outside the specified range, the image processing unit 306 corrects it to an ideal value, considering it unnatural. Thereby, the image processing unit 306 can correct the unnatural-looking parts, for example, making the overly long parts shorter. As shown in FIG. 14(b), when the knee part bends and the area around the knee looks unnaturally large (thick), the image processing unit 306 performs image processing to make the area around the knee smaller (thinner) so that it looks natural.

[0115] Such an unnatural-looking portrayal is caused by factors such as perspective. As described above, particularly in the photo sticker creation device 1, in addition to taking image data with a variety of poses, the image data is taken in a narrow shooting space, making it easy for such an unnatural-looking portrayal to occur.

[0116] Note that the image processing unit 306 can combine a plurality of the above-described processes. For example, a plurality of image processes such as image processing for each person, image processing for skin color adjustment, image processing for makeup, image processing for hair gloss, and image processing for unnatural areas can be combined.

[0117] The synthesizing unit 307 uses the processed image as a synthesis image, synthesizes the synthesis image with the moving image captured by the camera 91, and causes the synthesized image to be displayed as a live view display image on the touch panel monitor 92. Therefore, the user can perform shooting while confirming the finished image in real time.

[0118] The editing unit 230 includes the above-described editing operation units 27A and 27B and an editing processing unit 308. The editing processing unit 308 performs editing processing by controlling the built-in monitor 131 and the speaker 133 in the tablet in the editing operation units 27A and 27B. The editing processing unit 308 receives user operation inputs using the touch pens 132A and 132B for the built-in monitor 131 of the tablet.

[0119] In addition, the editing processing unit 308 performs predetermined image processing on the captured image as the image to be edited according to a selection operation on the selection screen displayed on the built-in monitor 131 of the tablet, and displays it on the built-in monitor 131 of the tablet. Alternatively, the editing processing unit 308 performs predetermined image processing on the synthesis image according to an input operation on the editing screen displayed on the built-in monitor 131 of the tablet, or generates a new synthesis image according to the input operation, synthesizes it with the captured image, and displays it on the built-in monitor 131 of the tablet.

[0120] In addition, the editing processing unit 308 controls the output of guidance for explaining how to proceed with the editing. For example, the editing processing unit 308 causes a guidance screen for explaining how to proceed with the editing to be displayed on the built-in monitor 131 of the tablet, or causes guidance audio for explaining how to proceed with the editing to be output from the speaker 133. Also, the editing processing unit 308 controls the communication unit 203 to perform processing related to communication via a network such as the Internet. Further, the editing processing unit 308 may perform printing processing by controlling the printer 140 of the printing operation unit 28.

[0121] The printing unit 240 includes the above-described printing operation unit 28 and a printing processing unit 309. The printing processing unit 309 receives printing data from the editing processing unit 308 and performs printing processing by controlling the printer 140 of the printing operation unit 28. Here, although the printing unit 240 that outputs a photo sticker, which is an example of printing data, has been used for explanation as an output unit for outputting image data, the method of outputting image data is not limited to this. For example, a transmission means for transmitting the image data itself to an external communication terminal using a network or the like may be used as the output unit.

[0122] In this way, in exchange for charging, the photo sticker creation device 1 provides a photo sticker creation game in which various devices (selection of shooting poses, BGM, narration, etc.) are provided to boost the user's mood. Therefore, the created image by the photo sticker creation device 1 becomes an image that draws out the user's happy expression or a gorgeous image with elaborate taste.

[0123] Also, the photo sticker creation device 1 is well-equipped with facilities such as writing, and since highly advanced techniques can be used for image deformation processing (for example, the size of the subject's eyes or the length of the legs) and color correction (whitening processing of the subject's skin), etc., the image created by the photo sticker creation device 1 becomes an image in which the user looks good.

[0124] In addition, compared to image processing outside the photo sticker creation device 1 (for example, processing in an application for photo processing), it is easier to edit (scribble on the image), and the variations of such editing are also rich. From this aspect as well, it can be said that the created image by the photo sticker creation device 1 is finished in a relatively colorful manner compared to the images photographed and image-processed outside the photo sticker creation device 1.

[0125] (Flow of the photo sticker creation game) Next, the processing flow of the user playing the photo sticker creation game on the photo sticker creation device 1 will be described with reference to FIG. 15. FIG. 15 is a flowchart showing the processing flow from the start of the game on the photo sticker creation device 1 to creating a photo sticker in the game.

[0126] In the state before the game starts, the control unit 201 that functions as the pre-selection processing unit 301 of the photo sticker creation device 1 causes the touch panel monitor 71 of the pre-selection operation unit 20 to display a message prompting the insertion of a coin. Also, as shown in FIG. 15, the control unit 201 determines whether a coin has been inserted into the coin insertion return port 73 based on the presence or absence of a start signal from the coin processing unit 74 (S1). When the control unit 201 determines that no coin has been inserted into the coin insertion return port 73 (S1: NO), it continues the determination process of whether a coin has been inserted.

[0127] The user who wants to start the game inserts a coin into the coin insertion return port 73 in the pre-selection space A0, which is the space in front of the pre-selection operation unit 20. When a coin is inserted into the coin insertion return port 73, a start signal instructing the start of the game is output from the coin processing unit 74. When the control unit 201 receives the start signal from the coin processing unit 74, it determines that a coin has been inserted into the coin insertion return port 73 (S1: YES) and executes pre-reception processing for the user (S2).

[0128] In the pre-reception process, the control unit 201 causes the touch panel monitor 71 to display messages prompting the selection of a course, the input of a name, the selection of a design, etc. When the user makes various selections or inputs according to the messages displayed on the touch panel monitor 71, the control unit 201 sets the shooting course, name, design, print layout, etc. The control unit 201 causes the touch panel monitor 92 to display a plurality of types of background images for composition so that the user can select a background image for composition.

[0129] When the pre-reception process ends, the control unit 201 causes the touch panel monitor 71 to display a message or the like prompting the user to move to the shooting space A1 and perform shooting. The control unit 201 functioning as the shooting processing unit 302 causes the touch panel monitor 92 of the shooting operation unit 21 to display a message prompting the user to touch the screen. In addition, a start button may be displayed together with this message or instead of this message. Further, the control unit 201 causes the speaker 93 to output a narration prompting the user to touch the screen together with the BGM. When the user who has moved to the shooting space A1 touches the touch panel monitor 92, the control unit 201 reads that the touch panel monitor 92 has been touched and starts the shooting process (S3).

[0130] In the shooting process, the control unit 201 causes the touch panel monitor 92 to display guidance regarding writing, for example, and prompts the user to select the writing level. When the user selects the writing level, the control unit 201 sets the writing level to the selected level.

[0131] In addition, the control unit 201 causes the touch panel monitor 92 to display an instruction screen for explaining the number of shots and outputs the corresponding narration from the speaker 93. In this embodiment, as an example, the number of shots is set to 7.

[0132] Next, the control unit 201 causes the touch panel monitor 92 to display an instruction screen for guiding the user to the standing position, and causes the speaker 93 to output the corresponding narration.

[0133] After displaying the instruction screen as described above, the control unit 201 causes the touch panel monitor 92 to display, as a live view display image, a synthesized live view display image obtained by synthesizing the background image for synthesis and the moving image acquired by the camera 91 according to the shooting course selected by the user (S4). Thereby, the user can pose while checking the finished image.

[0134] The control unit 201 performs the live view display until immediately before the end of the countdown for shooting. During that time, the control unit 201 causes the touch panel monitor 92 to display sample poses together with or instead of the live view display. The control unit 201 causes the speaker 93 to output a narration corresponding to the sample pose.

[0135] The control unit 201 manages the time from the start to the end of the live view display. When a preset predetermined time has elapsed, a countdown is performed with the display on the touch panel monitor 92 and the voice from the speaker 93.

[0136] At the end timing of the countdown, the control unit 201 transmits an illumination control signal to the top strobe 82, the left strobe 83, the right strobe 84, and the foot strobe 85, and transmits a shutter signal to the camera 91.

[0137] Thereby, the top strobe 82, the left strobe 83, the right strobe 84, and the foot strobe 85 irradiate flashes, and the camera 91 acquires a photographed image in which the illuminated user is shown together with the background. In the present embodiment, as an example, the processes from step S3 to step S4 are repeated a plurality of times to acquire seven photographed images. Further, the control unit 201 causes the storage unit 202 to store an edited target image obtained by synthesizing the background image for synthesis with the photographed image.

[0138] Here, an example has been described in which captured images are acquired at preset time intervals. However, the acquisition timing of the captured images is not limited to this. For example, when the imaging operation unit 21 has an operation button for imaging, the captured image may be acquired at the timing when this operation button is operated.

[0139] After the imaging is completed, the control unit 201 displays, on the touch panel monitor 92, a guidance screen that prompts the user to move to either the editing space A2-1 or the editing space A2-2, and outputs a voice guidance for movement to the speaker 93.

[0140] Then, the control unit 201 executes an editing process that allows the user to edit the image to be edited (S5). Specifically, the control unit 201 displays the image to be edited on the built-in monitor 131 of the tablet, and allows the user to draw a stamp image, a pen image, etc. on this image to be edited using the touch pens 132A and 132B, thereby creating an edited image.

[0141] Thereafter, the control unit 201 displays, on the built-in monitor 131 of the tablet, a guidance screen that prompts the user to move to the printing waiting space A3 where the sticker outlet 161 is provided, and outputs a voice guidance for movement to the speaker 133.

[0142] Furthermore, the control unit 201 executes a printing process of arranging the edited image in the print layout selected by the pre-selection operation unit 20 to create a print image, and printing this print image on the sticker 142 (S6).

[0143] When the printing process is completed, the control unit 201 executes the discharge process of the sticker 142 (S7), discharges the printed sticker 142 from the sticker outlet 161, provides it to the user as a photo sticker, and ends the game. In this way, the captured image of the user created by the photo sticker creating device 1 can be output as a photo sticker. Although detailed description is omitted, the photo sticker creating device 1 of the present embodiment can also output the captured image to a mobile terminal or the like by communication.

[0144] (Details of the synthesis process) Next, with reference to the flowchart shown in FIG. 16, the synthesis process in step S4 of the flowchart in FIG. 15 will be described. As shown in FIG. 16, in the synthesis process of step S4, the control unit 201 receives the image data captured by the imaging unit 220 (S41).

[0145] The control unit 201 generates a region group from the image data received in step S41 (S42). The control unit 201 generates at least two or more of the hair and skin region group, the specific part region group, and the person region group.

[0146] The control unit 201 specifies the target region for image processing according to the plurality of region groups generated in step S42 (S43).

[0147] The control unit 201 executes predetermined image processing on the target region of the image data specified in step S43 (S44).

[0148] The control unit 201 synthesizes an image using the image processed in step S44 (S46). The image synthesized here is acquired while being confirmed by the user through live view display.

[0149] (Display example) FIG. 17 shows an example of the editing screen 400 displayed on the in-built monitor 131 of the tablet. As shown in FIG. 17, the editing screen 400 includes a first editing unit 401 and a second editing unit 402.

[0150] The first editing unit 401 includes a first thumbnail unit 403, a first image to be edited display unit 405, a first operation button display unit 407, and a first palette 409. Similarly, the second editing unit 402 includes a second thumbnail unit 404, a second image to be edited display unit 406, a second operation button display unit 408, and a second palette 410.

[0151] Thumbnails of a plurality of images to be edited Im1 to Im5 that have been photographed are displayed in the first thumbnail section 403 and the second thumbnail section 404. The thumbnails displayed in the first thumbnail section 403 and the second thumbnail section 404 can be selected with the touch pen 132A or 132B. In the example shown in FIG. 17, the images to be edited Im1 and Im3 are selected, respectively. Although illustration is omitted, characters such as "being selected" may be displayed on the thumbnails of the selected images to be edited. Further, when seven images are photographed in step S3 of the flowchart in FIG. 15, seven images to be edited are also displayed in the thumbnail sections 403 and 404, respectively. However, FIG. 17 will be described by way of example in which five images to be edited are displayed for convenience.

[0152] The first image-to-be-edited display section 405 displays the image to be edited selected in the first thumbnail section 403. Further, the second image-to-be-edited display section 406 displays the image to be edited selected in the second thumbnail section 404. When the image to be edited selected in the thumbnail section 403 or 404 is changed, the image to be edited displayed in the image-to-be-edited display section 405 or 406 is also changed.

[0153] The first palette 409 is used to select the type of content, etc. when decorating the image to be edited displayed in the first image-to-be-edited display section 405 with content such as characters, patterns, colors, etc. Further, the second palette 410 is used to select the type of content, etc. when decorating the image to be edited displayed in the second image-to-be-edited display section 406 with content. As will be described in detail later, in the example shown in FIG. 17, each of the palettes 409 and 410 includes tabs for "imitation", "scribbling on the face", "items", and "makeup".

[0154] As described above, in the present embodiment, by extracting each area, it becomes possible to execute the processing assumed for each area, and preferable image processing for the user can be realized.

[0155] <<Modification Example 1>> In the above example, in the generation processing unit 304, a hair and skin mask is generated using the first learning model, a specific part mask is generated using the second learning model, and a person area mask is generated using the third learning model. In the specific part 305, it has been described that the processing area is specified using a plurality of masks necessary for image processing. However, the generation processing unit 304 may generate a plurality of masks using a learning model that has already learned a plurality of types of area groups in advance. For example, the learning model used by the generation processing unit 304 has learned the relationship between image data as shown in FIG. 11(a), a hair and skin area group as shown in FIG. 11(b), a specific part area group as shown in FIG. 11(c), and at least two or more area groups selected from a person area group as shown in FIG. 11(d). Therefore, the generation processing unit 304 can generate a plurality of types of masks that are each area group from the image data according to this learned model. Here, when the learned model has learned the relationship between the image data and the plurality of area groups, the generation processing unit 304 may generate only the masks of the area group of the requested type. For example, the generation processing unit 304 generates a hair and skin area mask and a person area mask from the image data. Then, the specific part 305 specifies each area to be processed using the plurality of masks that are the area groups generated by the generation processing unit 304.

[0156] <<Modification Example 2>> In addition, the generation processing unit 304 may generate regions using a learned learning model for at least any one of the relationship of regions that distinguish the skin region of a person's face and other regions as shown in FIG. 13(a), the relationship of the skin region, hair region, and other regions for each person as shown in FIG. 13(b), the relationship of regions of specific parts for each person as shown in FIG. 13(c), and the relationship of the face skin, skin other than the face, and regions of specific parts for each person as shown in FIG. 13(d). In this case, the region itself generated by the generation processing unit 304 using the learning model becomes the target region to be processed. For example, a learned learning model for the "skin and hair region group" and the "person region group" generates a region group including, for example, the skin region and other regions for each person as shown in FIG. 13(b) from the input image data. Therefore, the specifying unit 305 specifies the region group generated in this way as the target region to be processed. That is, since the region itself generated by the generation processing unit 304 is the target region, the specifying unit 305 may be integrated with the generation processing unit 304.

[0157] Note that the combination of the skin and hair region group and the specific part region group uses semantic segmentation. On the other hand, the person region group uses instance segmentation. Therefore, in each of the above-described examples, the combination of the skin and hair region group and the person region group, the combination of the part specifying region group and the person region group, and the combination of the skin and hair region group, the part specifying region group, and the person region group use panoptic segmentation.

[0158] (Example of software implementation) The functional blocks of the photo sticker creation device 1 may be realized by a logic circuit (hardware) formed in an integrated circuit (IC chip) or the like, or may be realized by software using a CPU (Central Processing Unit).

[0159] In the latter case, the photo sticker creating device 1 includes a CPU that executes instructions of a program which is software for realizing each function, a ROM (Read Only Memory) or a storage device (collectively referred to as "recording medium") in which the program and various data are recordable in a computer-readable manner, a RAM (Random Access Memory) for expanding the program, and the like. Then, when the computer (or CPU) reads the program from the recording medium and executes it, the object of the present invention is achieved. As the recording medium, a "non-transitory tangible medium", for example, a tape, a disk, a card, a semiconductor memory, a programmable logic circuit, etc. can be used. Further, the program may be supplied to the computer via any transmission medium (such as a communication network or a broadcast wave) capable of transmitting the program. Note that the present invention can also be realized in the form of a data signal embedded in a carrier wave in which the program is embodied by electronic transmission.

[0160] In the above description, the image processing device has been described as the photo sticker creating device 1. However, it is not limited thereto. Therefore, even if the image processing device is, for example, another device that does not create photo stickers, the same effect can be obtained. As an example, the image processing device may be used as a device for taking image data in a photo studio.

[0161] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope shown in the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.

Explanation of Reference Numerals

[0162] 1 Photo sticker creating device (image processing device) 301 Preselection processing unit 302 Shooting processing unit 303 Reception unit 304 Generation processing unit 305 Identification unit 306 Image Processing Unit 307 Composition Unit 308 Editing Processing Unit 309 Printing Processing Unit (Output Unit)

Claims

1. A reception unit that receives image data including a person, A first generation process that generates a first region group from the image data using a first learning model that has learned the relationship between learning image data including a person and the first region group, which is at least one or more regions for the person included in the learning image data, and a second generation process that generates a second region group from the image data using a second learning model that has learned the relationship between learning image data including a person and the second region group, which is at least one or more regions different from the first region group for the person included in the learning image data, and a generation processing unit that executes the processes, A specifying unit that specifies a target region to be processed according to the first and second region groups from the image data, An image processing unit that performs predetermined image processing on the target region of the image data, An image processing apparatus comprising the above.

2. The first region group is a person's skin, hair, and other regions, The second region group is at least the regions of a person's face and limbs, The first generation process and the second generation process utilize semantic segmentation, The specifying unit specifies a target region of a person's facial skin. The image processing apparatus according to Claim 1.

3. The image data includes a plurality of persons, The first region group is a person's skin, hair, and other regions, The second region group is the region of each person included in the image data, The first generation process utilizes semantic segmentation, The second generation process utilizes instance segmentation, The specifying unit specifies the skin region for each person. The image processing apparatus according to Claim 1.

4. The image data includes a plurality of persons, The first region group is at least the regions of a person's face and limbs, The second region group is the region of each person included in the image data, The first generation process utilizes semantic segmentation, The second generation process utilizes instance segmentation, The specifying unit specifies at least the regions of a person's face and limbs for each person. The image processing apparatus according to Claim 1.

5. The image data includes a plurality of persons, The first region group is a person's skin, hair, and other regions, The second region group is at least the regions of a person's face and limbs, The generation processing unit further executes a third generation process of generating a third region group including regions of each person from the image data using a third learned model that has learned the relationship between the image data including a person and the contour of the person included in the image data. The first generation process uses semantic segmentation. The second generation process uses semantic segmentation. The third generation process uses instance segmentation. The specifying unit specifies, for each person, a face and a skin region other than the face as target regions according to the first, second, and third region groups. The image processing apparatus according to claim 1.

6. A receiving unit that receives image data including a person, A generation processing unit that generates a learned region group from the image data using a learned model that has learned the relationship between learning image data including a person and a first region group that is at least one or more regions of the person included in the learning image data, the relationship between a second region group that is at least one or more regions different from the first region group, and at least two or more of a third region group different from the first region group and the second region group, A specifying unit that specifies a target region to be processed from the image data according to the learned region group, An image processing unit that performs predetermined image processing on the target region of the image data, An image processing apparatus comprising:

7. A receiving step of receiving image data including a person, A first generation step of generating a first region group from the image data using a first learned model that has learned the relationship between learning image data including a person and a first region group that is at least one or more regions of the person included in the learning image data, A second generation step of generating a second region group from the image data using a second learned model that has learned the relationship between learning image data including a person and a second region group that is at least one or more regions different from the first region group of the person included in the learning image data, A specifying step of specifying a target region to be processed from the image data according to the first and second region groups, An image processing step of performing predetermined image processing on the target region of the image data, An image processing method including:

8. A computer, A receiving unit that receives image data including a person, A first generation processing unit that generates a first group of regions from the image data using a first trained learning model that has learned the relationship between learning image data including a person and a first group of regions that are at least one or more regions for the person included in the learning image data. A second generation processing unit that generates a second group of regions from the image data using a second trained learning model that has learned the relationship between learning image data including a person and a second group of regions that are at least one or more regions different from the first group of regions for the person included in the learning image data. A specifying unit that specifies a target region to be processed according to the first and second groups of regions from the image data. An image processing unit that performs predetermined image processing on the target region of the image data. An image processing program that causes it to function.

Citation Information

Patent Citations

  • Image processing device, image processing method, and computer program

    JP2018173885A

  • Information processing equipment, image area selection method, computer program, and storage media

    JP2019086899A

  • Photo-creating game machine, image processing method, and program

    JP2020053831A