System
The system addresses the challenge of taking ideal selfies and group photos by generating 3D facial models and adjusting angles, poses, and expressions to naturally combine with background images, enhancing usability and image quality.
Patent Information
- Application Number
- JP2024118171
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-02-04
AI Technical Summary
The difficulty in taking ideal selfies and group photos due to location restrictions or physical challenges, and the complexity of combining facial images with background images naturally, especially with varying light directions and color tones.
A system that registers a user's facial image, generates feature point data, creates a 3D facial model, and automatically adjusts angles, poses, and facial expressions to combine naturally with a background photo, analyzing light direction and color tone for natural synthesis.
Enables the generation of ideal photos in challenging locations and high-quality group photos without physical gathering, improving usability and natural image synthesis.
Smart Images

Figure 2026017389000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] When taking selfies, there are problems such as the difficulty of taking the ideal photo because selfies are prohibited or physically difficult depending on the location or situation. There is also the issue of it being difficult to take group photos because it is difficult for everyone to gather together due to teleworking, etc. [Means for solving the problem]
[0005] The system provides a means for a user to register their own facial image and generate feature point data by extracting feature points from the registered facial image, and a means for generating a 3D facial model based on the feature point data. The system also provides a means for automatically generating the angle, pose, and facial expression to be combined with a background photo, and a means for representing the facial image as a 3D model according to the generated angle, pose, and facial expression, and combining it naturally with the background photo. The system further provides a means for analyzing the light direction and color tone based on the background photo specified by the user and calculating the parameters of the combined image based on the analysis. The system also provides a means for the user to specify the desired pose and facial expression, and a means for generating corresponding movements for the facial portion of the 3D model based on the specified pose and facial expression. This allows for the generation of ideal photos even in places where selfies are difficult, and allows group photos to be taken without physically gathering together.
[0006] "User" refers to an individual who uses the system to register their own facial image and generate a photo.
[0007] "Facial image" refers to a photograph containing the user's own face that the user registers in the system.
[0008] "Feature points" refer to position information of the eyes, nose, mouth, etc. extracted from a facial image.
[0009] "3D model" refers to a digital representation of a face generated in three dimensions based on feature point data.
[0010] A "background photo" refers to a photo selected by the user that will be combined with a facial image.
[0011] "Angle" refers to information about the orientation of the face.
[0012] "Pose" refers to the posture or position of the whole body or upper body.
[0013] "Facial expression" refers to the expression of emotions formed by the movement of facial muscles.
[0014] "Synthesis parameters" refer to information necessary to naturally synthesize a background photo and a facial image.
[0015] "Movement" refers to the movement of a 3D model that corresponds to a specified pose or facial expression.
[0016] "Light direction" refers to information about the position and angle of the light source in the background photo.
[0017] "Hue" refers to information about the tone and saturation of the color of the background photo. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and creates images that are naturally composited against a background photograph. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos.
[0040] 1. The process by which users register their facial images
[0041] User
[0042] First, the user launches the application on their smartphone or PC and selects the facial image registration menu.
[0043] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[0044] Terminal
[0045] The device automatically extracts feature points (e.g., the positions of the eyes, nose, and mouth) from each captured facial image.
[0046] A 3D model of the face is generated based on the feature point data.
[0047] The generated 3D model and feature point data are sent to the server.
[0048] server
[0049] The server receives the data sent from the terminal and stores it in a database.
[0050] The saved feature point data and 3D model are used in subsequent synthesis processes.
[0051] 2. Angle, pose, and facial expression generation process
[0052] User
[0053] Users can select a background photo within the application or take a new one.
[0054] The user specifies the desired pose and facial expression.
[0055] server
[0056] The server analyzes the background photo selected by the user and calculates the direction and color of the light.
[0057] Based on the pose and facial expression specified by the user, the angle, pose, and facial expression are automatically generated for the registered 3D model.
[0058] Calculate synthesis parameters for synthesizing the background photo analysis data and the 3D model.
[0059] The facial image is represented as a 3D model based on the generated angle, pose, and expression, and is then naturally combined with the background photo.
[0060] Terminal
[0061] The terminal receives the composite image and the composite parameters from the server and performs the composite process.
[0062] The synthesized image is presented to the user for confirmation.
[0063] 3. Photo generation and storage process
[0064] User
[0065] The user checks the generated composite image and presses the save button if satisfied.
[0066] Use further editing functions as needed.
[0067] Terminal
[0068] The terminal stores the composite image in local storage.
[0069] It also provides users with the option to share images on social media or cloud services, depending on their choice.
[0070] Specific examples
[0071] 1. Example of face registration
[0072] User A starts the app and registers three facial images: front, left, and right.
[0073] The device extracts feature points from these images and creates a 3D model of the face.
[0074] The generated data is sent to the server, which stores it in a database.
[0075] 2. Examples of group photos
[0076] A family of four wants to take a photo together, but one of them is in a remote location and it is physically impossible to take the photo.
[0077] The three people took a group photo on site, and the facial image of one person in a remote location was registered through the application.
[0078] Users select a group photo as a background within the app, and the server generates the appropriate angle, pose, and facial expression based on a remote facial model of a single person, and combines it with the background photo.
[0079] The composite image is displayed on the smartphone and can be saved by the user.
[0080] As described above, by implementing this invention, it becomes possible to easily generate ideal photos even in places where it is difficult to take selfies, and it is also possible to create high-quality group photos without having to physically gather together.
[0081] The processing flow will be explained below.
[0082] Program processing details
[0083] The process by which users register their facial images
[0084] Step 1:
[0085] The user launches the application on their smartphone or PC.
[0086] Step 2:
[0087] The user selects a face image registration menu and follows the instructions to take photos of the face from the front, left side, and right side.
[0088] Step 3:
[0089] The terminal receives each captured face image.
[0090] Step 4:
[0091] The device extracts feature points (such as the positions of the eyes, nose, and mouth) from the facial image.
[0092] Step 5:
[0093] The device generates a 3D model of the user's face based on the extracted feature point data.
[0094] Step 6:
[0095] The device sends the generated 3D model and feature point data to the server.
[0096] Step 7:
[0097] The server stores the received data in a database.
[0098] The process of generating angles, poses, and facial expressions
[0099] Step 1:
[0100] The user selects the photo creation menu of the application.
[0101] Step 2:
[0102] The user selects a background photo or takes a new one.
[0103] Step 3:
[0104] The user specifies the desired pose and facial expression.
[0105] Step 4:
[0106] The server receives the background photo and runs an image analysis algorithm to analyze the light direction and color tone.
[0107] Step 5:
[0108] Based on the analysis results, the server calculates the synthesis parameters for combining the image with the background photo.
[0109] Step 6:
[0110] The server automatically generates angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[0111] Step 7:
[0112] The server uses the generated data and synthesis parameters to represent the facial image as a 3D model and naturally combines it with the background photo.
[0113] Step 8:
[0114] The server transmits the composite image to the terminal.
[0115] Photo generation and storage process
[0116] Step 1:
[0117] The terminal receives the composite image sent from the server.
[0118] Step 2:
[0119] The terminal displays the received composite image on the screen and asks the user for confirmation.
[0120] Step 3:
[0121] The user checks the composite image and, if satisfied, presses the save button.
[0122] Step 4:
[0123] The device stores the composite image in local storage.
[0124] Step 5:
[0125] The device displays a notification to the user that the save is complete.
[0126] Step 6:
[0127] If the user wishes to share the image on a social networking site or cloud service, the device will provide the option to do so and process it accordingly.
[0128] Example 1
[0129] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0130] In conventional facial image synthesis systems, it was difficult for users to naturally combine their own facial images with background images, and differences in light direction and color tone often resulted in unnatural results. It was also difficult to naturally reflect the user's desired pose and facial expression. Furthermore, the process of efficiently extracting feature points from multiple facial images and generating a 3D model was complicated, resulting in poor usability for the entire system.
[0131] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0132] In this invention, the server includes means for a user to register his or her own facial image, means for extracting feature points from the registered facial image and generating feature point data, means for generating a 3D facial model based on the feature point data, means for analyzing a background image and identifying the light direction and color tone, means for automatically generating an angle, pose, and facial expression based on the analysis results, means for representing the facial image as a 3D model according to the generated angle, pose, and facial expression and naturally combining it with the background image, and means for presenting the combined image to the user and saving it after confirmation. This allows the user to easily generate natural combined images and improves operability, enabling more satisfying facial image synthesis.
[0133] "User" refers to an individual or group that uses this system to register their own facial images and create composite images.
[0134] "Means for registering" refers to a method or device that allows users to upload and store their facial images in the system.
[0135] "Feature points" are data that indicate specific parts of a face image, such as the eyes, nose, and mouth.
[0136] "Feature point data" is a data set that includes information on feature points extracted from a face image.
[0137] A "3D facial model" is a three-dimensional shape of a face reconstructed in three dimensions based on feature point data of a facial image.
[0138] A "background image" is a photograph or image selected by the user as a composite object.
[0139] The "analyzing means" refers to a method or device for calculating and analyzing the light direction and color tone of the background image.
[0140] "Means for automatically generating angles, poses, and facial expressions" refers to a method or device that has the function of automatically creating facial angles, poses, and facial expressions based on user specifications and analysis results.
[0141] The "synthesis means" refers to a method or device for naturally synthesizing the generated 3D face model with a background image.
[0142] The "presenting means" refers to a method or device for displaying the synthesized image to the user and requesting confirmation.
[0143] "Storing means" refers to a method or device for storing the confirmed composite image within the system or in external storage.
[0144] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and creates images that are naturally composited against a background photograph. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos.
[0145] The process by which users register their facial images
[0146] User
[0147] The user launches the application on their smartphone or PC and selects the facial image registration menu.
[0148] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[0149] Terminal
[0150] The device automatically extracts feature points (e.g., the positions of the eyes, nose, and mouth) from each captured facial image using Python's OpenCV library.
[0151] Based on the extracted feature point data, a 3D model of the face is generated using 3D modeling software such as Blender.
[0152] The generated 3D model and feature point data are sent to the server via REST API.
[0153] server
[0154] The server receives the data sent from the device and stores it in a MySQL database using the Django framework.
[0155] The saved feature point data and 3D model are used in subsequent synthesis processes.
[0156] The process of generating angles, poses, and facial expressions
[0157] User
[0158] Users can select a background photo within the application or take a new one.
[0159] The user specifies the desired pose and facial expression.
[0160] server
[0161] The server analyzes the background photo selected by the user and calculates the light direction and color tone using PIL (Python Imaging Library) and OpenCV.
[0162] Based on the pose and facial expression specified by the user, the angle, pose, and facial expression are automatically generated for the registered 3D model. Here, generative AI models (such as StyleGAN and DeepFaceLab) are utilized.
[0163] Calculate synthesis parameters for synthesizing the background photo analysis data and the 3D model.
[0164] The facial image is represented based on a 3D model according to the generated angle, pose, and expression, and is then naturally combined with the background photo.
[0165] The process by which the device generates a composite image and presents it to the user
[0166] Terminal
[0167] The terminal receives the composite image and the composite parameters from the server.
[0168] The synthesis process is performed and the completed image is presented to the user.
[0169] User
[0170] The user checks the presented composite image and presses the save button if satisfied.
[0171] Further modifications can be made within the app if needed.
[0172] The generated images can also be shared on social media or cloud services.
[0173] Specific examples
[0174] Specific examples of face registration
[0175] The user launches the app and takes a series of photos of their face - front, left, and right - and registers them. The device uses Python's OpenCV library to extract feature points from these images and creates a 3D face model using Blender. The generated data is sent to the server via a REST API and stored in a MySQL database using the Django framework.
[0176] Examples of group photos
[0177] If a family wants to take a group photo together, but one person is remote and cannot physically be present, the remote person can register their face image through the app. The user then takes a group photo on-site and selects a background photo within the app. The server then analyzes the image, calculating the light direction and color tone, and generates the appropriate angle, pose, and facial expression based on a 3D facial model of the remote person, which is then combined with the background photo. The resulting composite image is then displayed on the user's smartphone.
[0178] Prompt Sentence Examples
[0179] "We will create a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and creates an image that is composited into a background photo. We will also analyze the background photo, such as the direction of light and color tone, to ensure that the composite image looks natural."
[0180] By implementing this invention, ideal selfies can be easily generated even in places where it is difficult to take selfies, and high-quality group photos can be created without physically gathering together.
[0181] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0182] The flow of this system's program processing
[0183] Step 1:
[0184] explanation
[0185] Users launch the application on their smartphone or PC, select the face image registration menu, and follow the app's guide to take and register photos of their face from the front, left side, and right side.
[0186] Input and Output
[0187] Input: Face image (front, left, right)
[0188] Output: Three registered face images
[0189] Specific actions
[0190] The user clicks the "Register Face Image" button on the app, activates the camera to take face images from various angles, and then presses the "Confirm / Register" button.
[0191] Step 2:
[0192] explanation
[0193] The device automatically extracts feature points (such as the positions of the eyes, nose, and mouth) from the captured facial image using Python's OpenCV library. Based on the extracted feature point data, a 3D model of the face is generated using 3D modeling software such as Blender.
[0194] Input and Output
[0195] Input: Three registered face images
[0196] Output: feature point data, 3D face model
[0197] Specific actions
[0198] The device uses the OpenCV library to detect feature points from each facial image, then uses Blender to generate a 3D model of the face based on these feature points. The generated model and feature point data are then sent to the server via a REST API.
[0199] Step 3:
[0200] explanation
[0201] The server receives the data sent from the device and stores it in a MySQL database using the Django framework. The stored feature point data and 3D model are then used in the synthesis process.
[0202] Input and Output
[0203] Input: feature point data, 3D face model
[0204] Output: Data stored in the database
[0205] Specific actions
[0206] The server receives the transmitted data via the Django framework and stores it in a MySQL database.
[0207] Step 4:
[0208] explanation
[0209] Users can select a background photo within the application or take a new one, and specify the desired pose and facial expression.
[0210] Input and Output
[0211] Input: Background photo, desired pose and facial expression
[0212] Output: User selections and specifications
[0213] Specific actions
[0214] Once the user selects (or takes) a background photo, a UI for specifying the pose and expression is displayed.
[0215] Step 5:
[0216] explanation
[0217] The server analyzes the background photo selected by the user and calculates the light direction and color tone using PIL (Python Imaging Library) and OpenCV.
[0218] Input and Output
[0219] Input: Background photo
[0220] Output: Analysis results of light direction and color tone
[0221] Specific actions
[0222] The server loads the background photo using the PIL or OpenCV library and analyzes the light direction and color tone.
[0223] Step 6:
[0224] explanation
[0225] The server uses a generative AI model (e.g., StyleGAN or DeepFaceLab) to automatically generate angles, poses, and expressions for the registered 3D model based on the poses and expressions specified by the user.
[0226] Input and Output
[0227] Input: feature point data, 3D face model, desired pose and expression
[0228] Output: Generated angles, poses, and expressions
[0229] Specific actions
[0230] The server uses a generative AI model to generate poses and facial expressions based on feature point data and a 3D model.
[0231] Step 7:
[0232] explanation
[0233] The server calculates synthesis parameters for combining the analysis data of the background photo with a 3D model, and then represents the facial image as a 3D model according to the generated angle, pose, and expression, which is then naturally combined with the background photo.
[0234] Input and Output
[0235] Input: background photo, analysis data, generated angles, poses, facial expressions
[0236] Output: Synthesized face image
[0237] Specific actions
[0238] The server calculates the synthesis parameters and uses OpenCV to naturally synthesize the face image and background photo.
[0239] Step 8:
[0240] explanation
[0241] The terminal receives the composite image and the composite parameters from the server, performs the composite process, and presents the completed composite image to the user.
[0242] Input and Output
[0243] Input: synthetic image, synthetic parameters
[0244] Output: The composite image presented to the user
[0245] Specific actions
[0246] The device retrieves the composite image from the server via an HTTP request and displays it to the user on the app.
[0247] Step 9:
[0248] explanation
[0249] The user can review the composite image presented to them and, if satisfied, press the save button. If necessary, they can make further edits within the app. The generated image can also be shared via social media or cloud services.
[0250] Input and Output
[0251] Input: Composite image
[0252] Output: Saved and shared images
[0253] Specific actions
[0254] When the user presses the "Save" button, the image is saved to the smartphone's local storage, and when the user presses the "Share" button, the image is uploaded to social media or a cloud service.
[0255] (Application example 1)
[0256] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0257] Modern electronic payment services require enhanced security. However, many current systems have limited user authentication methods, making it difficult to simultaneously improve convenience and security. Furthermore, many systems can only authenticate users at specific angles or poses, which can be inconvenient for users. There is a need to solve these issues and realize more accurate and convenient facial recognition. Furthermore, a system is needed that can accurately authenticate users even when they have different facial expressions or poses.
[0258] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0259] In this invention, the server includes: a means for a user to register his or her own facial image; a means for extracting feature points from the registered facial image and generating feature point data; a means for generating a 3D facial model based on the feature point data; a means for automatically generating an angle, pose, and facial expression to be combined with a background photograph; a means for representing the facial image as a 3D model according to the generated angle, pose, and facial expression and combining it naturally with the background photograph; a means for performing electronic authentication using analysis data of the background photograph and the user's 3D model; and a means for presenting the combined image to the user and saving it after confirmation. This enables advanced electronic authentication by generating a natural combined image based on the user's registered facial image and background photograph. A means for the user to specify a desired pose and facial expression; a means for generating corresponding movements for the facial portion of the 3D model based on the specified pose and facial expression; a means for performing electronic authentication based on the generated 3D model and a captured image of the user; and a means for reconstructing the facial image according to the generated movements, thereby providing an authentication system that combines convenience and security.
[0260] "Means for users to register their own facial images" refers to a function that allows users to register their own facial images using devices such as smartphones or computers.
[0261] "Means for extracting feature points and generating feature point data" refers to a function that uses AI technology to detect feature points such as the eyes, nose, and mouth from registered facial images and converts them into data.
[0262] The "means for generating a 3D model of a face" is a function for generating a 3D model of a face using feature point data.
[0263] "Means for automatically generating angles, poses, and facial expressions for compositing with background photos" is a function that uses AI to automatically generate the optimal facial angle, pose, and facial expression for compositing with a background photo.
[0264] "Means for representing a facial image as a 3D model according to the generated angle, pose, and expression, and for naturally combining it with a background photograph" refers to a function for reproducing a facial image as a 3D model based on the generated angle, pose, and expression, and for naturally combining it with a background photograph.
[0265] "Means for electronic authentication using background photo analysis data and a 3D model of the user" is a function for securely and accurately authenticating a user using analysis data of the light direction and color tone of the background photo and a 3D model of the user.
[0266] The "means for presenting the synthesized image to the user and saving it after confirmation" is a function for displaying the generated synthesized image to the user and saving the image after obtaining confirmation.
[0267] "Means for generating corresponding movements for the facial part of a 3D model based on a specified pose or expression" is a function that moves the facial part of a 3D model according to the pose or expression specified by the user.
[0268] The "means for reconstructing a facial image according to the generated movements" is a function for reconstructing a facial image based on the generated poses and facial movements.
[0269] "Means for electronic authentication" refers to an authentication function that uses a user's facial image or 3D model to allow access to the system.
[0270] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and creates images that are naturally composited with background photographs. Below, we will provide a detailed explanation of how this system is realized.
[0271] 1. Facial image registration process
[0272] Users register their facial images using a smartphone or computer. Specifically, they take photos of their face from the front, left side, and right side, and register them through the application.
[0273] The device receives a facial image and extracts its features using a library such as dlib. This generates key feature data such as the eyes, nose, and mouth. A 3D model of the face is then generated based on this feature data. This 3D model and feature data are then sent to a server and stored in a database.
[0274] 2. Angle, pose, and facial expression generation process
[0275] The server analyzes the background photo selected or taken by the user. Specifically, it calculates the direction of light and color tone. Based on this analysis data, it automatically generates the appropriate angle, pose, and expression for the registered 3D model according to the pose and expression specified by the user. Using AI technology such as DeepFace, it calculates synthesis parameters to naturally combine the user's 3D face model with the background photo.
[0276] 3. Electronic Authentication Process
[0277] The device performs electronic authentication using background photo analysis data and a 3D model of the user. Specifically, it uses DeepFace's Facenet model to match the captured image of the user with the 3D model for accurate authentication.
[0278] 4. Photo Generation and Storage Process
[0279] Users can review the generated composite image and save it if they are satisfied. Further editing is available if necessary. The composite image is saved to local storage, and options for sharing it to social media and cloud services are also provided.
[0280] The system that realizes this invention executes a series of processes, from registering a facial image, generating a 3D model, using that model for electronic authentication, and finally generating and saving a composite image. This enables users to be accurately and securely authenticated even with different poses and expressions, and to obtain a high-quality composite image.
[0281] Examples:
[0282] A user installs the FacePay app and registers images of their face, showing the front, left, and right sides. After making a purchase at a convenience store, they perform payment using facial recognition, and the payment is completed upon successful authentication.
[0283] Example prompt sentence:
[0284] "Design an application that generates a 3D face model from a facial image registered by a user and uses that model for facial authentication during electronic payments."
[0285] As described above, by implementing the present invention, it is possible to generate ideal photos even when selfies are difficult, and to create high-quality group photos without physically gathering together. Furthermore, by using this system, the security of electronic payments is improved, and user convenience is also enhanced.
[0286] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0287] Step 1:
[0288] The user registers a facial image. In this process, the user launches an application on a smartphone or PC and takes and registers facial images from the front, left side, and right side.
[0289] Input: Face image taken by the user (front, left, right)
[0290] Output: Facial feature point data and a 3D model of the face generated by the device
[0291] Step 2:
[0292] The device processes the facial image provided by the user. Specifically, it uses the dlib library to extract feature points such as the eyes, nose, and mouth from the captured facial image. Based on this feature point data, a 3D model of the face is generated. The generated data is sent to a server and stored in a database.
[0293] Input: A face image obtained from the user
[0294] Output: Feature point data and a 3D face model
[0295] Step 3:
[0296] When a user selects or takes a background photo, the server analyzes the background photo, detecting the direction and color of the light, and then uses this information to calculate the direction and color of the light in the background photo.
[0297] Input: A background photo selected or taken by the user
[0298] Output: Analysis data of background photo (light direction and color tone)
[0299] Step 4:
[0300] The server automatically generates the angle, pose, and facial expression for the registered 3D model according to the pose and facial expression specified by the user. Based on this information, the server expresses the facial image as a 3D model and calculates synthesis parameters for naturally combining it with the background photo.
[0301] Input: Analysis data of user-specified poses and expressions, and background photos
[0302] Output: angle, pose, facial expression, synthesis parameters
[0303] Step 5:
[0304] The device uses the synthesis parameters and 3D model received from the server to naturally synthesize the background photo and facial image. Specifically, the synthesis process is carried out using DeepFace's AI technology.
[0305] Input: Synthesis parameters and 3D model provided by the server
[0306] Output: Composite image
[0307] Step 6:
[0308] The device will then present the combined image to the user for review, and if the user is satisfied, the image will be saved to local storage, with the option to share it via social media or cloud services if desired.
[0309] Input: The composite image
[0310] Output: User confirmation and save
[0311] Step 7:
[0312] The device performs electronic authentication using background photo analysis data and a 3D model of the user, and uses DeepFace's Facenet model to match the captured image with the 3D model.
[0313] Input: Captured image, background photo analysis data, user 3D model
[0314] Output: Authentication result (success or failure)
[0315] Through a series of processes from registration to synthesis and authentication, highly accurate and convenient electronic authentication is achieved.
[0316] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0317] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by the user, creating images that blend naturally with the background photo, and further combines it with an emotion engine that recognizes the user's emotions. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos.
[0318] 1. The process by which users register their facial images
[0319] User
[0320] First, the user launches the application on their smartphone or PC and selects the facial image registration menu.
[0321] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[0322] Terminal
[0323] The device automatically extracts feature points (such as the position of the eyes, nose, and mouth) from each captured facial image.
[0324] A 3D model of the face is generated based on the feature point data.
[0325] The generated 3D model and feature point data are sent to the server.
[0326] server
[0327] The server receives the data sent from the terminal and stores it in a database.
[0328] The saved feature point data and 3D model are used in subsequent synthesis processes.
[0329] 2. Angle, pose, and facial expression generation process
[0330] User
[0331] Users can select a background photo within the application or take a new one.
[0332] The user specifies the desired pose and facial expression.
[0333] server
[0334] The server analyzes the background photo selected by the user and calculates the direction and color of the light.
[0335] Based on the pose and facial expression specified by the user, the angle, pose, and facial expression are automatically generated for the registered 3D model.
[0336] Calculate synthesis parameters for synthesizing the background photo analysis data and the 3D model.
[0337] Emotion Engine
[0338] The emotion engine recognizes the user's emotion from the registered facial image.
[0339] The facial expression is adjusted and generated based on the user's desired pose, facial expression, and recognized emotion.
[0340] The emotion engine can also analyze the user's voice input, recognize the user's emotion based on the analysis results, and generate facial expressions based on that.
[0341] server
[0342] The server uses the generated data and synthesis parameters to represent the facial image as a 3D model and naturally combines it with the background photo.
[0343] The server transmits the composite image to the terminal.
[0344] 3. Photo generation and storage process
[0345] Terminal
[0346] The terminal receives the composite image from the server.
[0347] The terminal displays the composite image on the screen and asks the user for confirmation.
[0348] User
[0349] The user checks the composite image and, if satisfied, presses the save button.
[0350] Use further editing functions as needed.
[0351] Terminal
[0352] The device stores the composite image in local storage.
[0353] It also provides the option to share to social media and cloud services.
[0354] Specific examples
[0355] 1. Example of face registration
[0356] User A starts the app and registers three facial images: front, left, and right.
[0357] The device extracts feature points from these images and generates a 3D model of the face.
[0358] The generated data is sent to the server, which stores it in a database.
[0359] 2. Specific examples of emotion engines
[0360] User B selects a landscape photo as a background in the app, and the app recognizes B's current emotion.
[0361] If the user desires a "smiling" facial expression, but the emotion engine recognizes the emotion of "surprise," it applies the surprised expression to the facial image and generates a 3D model.
[0362] The server calculates the synthesis parameters based on this data and synthesizes it with the landscape photo.
[0363] 3. Examples of group photos
[0364] A family of four wants to take a photo together, but one of them is in a remote location and it is physically impossible to take the photo.
[0365] The three people took a group photo on site, and the facial image of one person in a remote location was registered through the application.
[0366] Users select a group photo as a background within the app, and the server generates the appropriate angle, pose, and facial expression based on a remote facial model of a single person, and combines it with the background photo.
[0367] The emotion engine recognizes the current emotion of a person in a remote location and generates and adjusts facial expressions to match that emotion.
[0368] The composite image is displayed on the smartphone and can be saved by the user.
[0369] As described above, by implementing this invention, it becomes possible to easily create ideal photos even in places where it is difficult to take selfies, and it is also possible to create high-quality group photos without physically gathering together.Furthermore, the emotion engine can generate natural facial expressions that better match the user's emotions, improving image quality and satisfaction.
[0370] The processing flow will be explained below.
[0371] Program processing details
[0372] The process by which users register their facial images
[0373] Step 1:
[0374] The user launches the application on their smartphone or PC.
[0375] Step 2:
[0376] The user selects a face image registration menu and follows the instructions to take photos of the face from the front, left side, and right side.
[0377] Step 3:
[0378] The terminal receives each captured face image.
[0379] Step 4:
[0380] The device extracts feature points (such as the positions of the eyes, nose, and mouth) from the facial image.
[0381] Step 5:
[0382] The device generates a 3D model of the user's face based on the extracted feature point data.
[0383] Step 6:
[0384] The device sends the generated 3D model and feature point data to the server.
[0385] Step 7:
[0386] The server stores the received data in a database.
[0387] The process of generating angles, poses, and facial expressions
[0388] Step 1:
[0389] The user selects the photo creation menu of the application.
[0390] Step 2:
[0391] The user selects a background photo or takes a new one.
[0392] Step 3:
[0393] The user specifies the desired pose and facial expression.
[0394] Step 4:
[0395] The server receives the background photo and runs an image analysis algorithm to analyze the light direction and color tone.
[0396] Step 5:
[0397] Based on the analysis results, the server calculates the synthesis parameters for combining the image with the background photo.
[0398] Step 6:
[0399] The server automatically generates angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[0400] Step 7:
[0401] The emotion engine recognizes the user's emotions from registered facial images or voice input.
[0402] Step 8:
[0403] The emotion engine adjusts and generates poses and expressions based on the recognized emotions.
[0404] Step 9:
[0405] The server uses the generated data and synthesis parameters to represent the facial image as a 3D model and naturally combines it with the background photo.
[0406] Step 10:
[0407] The server transmits the composite image to the terminal.
[0408] Photo generation and storage process
[0409] Step 1:
[0410] The terminal receives the composite image sent from the server.
[0411] Step 2:
[0412] The terminal displays the received composite image on the screen and asks the user for confirmation.
[0413] Step 3:
[0414] The user checks the composite image and, if satisfied, presses the save button.
[0415] Step 4:
[0416] The device stores the composite image in local storage.
[0417] Step 5:
[0418] The device displays a notification to the user that the save is complete.
[0419] Step 6:
[0420] If the user wishes to share the image on a social networking site or cloud service, the device will provide the option to do so and process it accordingly.
[0421] Example 2
[0422] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0423] Conventional photo compositing technologies have difficulty in generating natural poses and facial expressions desired by users, and require manual setting of compositing parameters such as light direction and color tone. This has required a great deal of time and effort to generate a photo that satisfies the user. Furthermore, when taking a group photo and not everyone is present, it is difficult to composite people in remote locations naturally. Furthermore, there has been a lack of mechanisms for appropriately reflecting facial expressions that match the user's emotions. The objective of this invention is to solve these problems and realize easier and more natural photo compositing.
[0424] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0425] In this invention, the server includes: means for a user to register his or her own facial image; means for extracting feature points from the registered facial image and generating feature point data; means for generating a 3D facial model based on the feature point data; means for automatically generating an angle, pose, and facial expression to be composited with a background photograph; means for representing the facial image as a 3D model according to the generated angle, pose, and facial expression and naturally composited with the background photograph; means including an emotion recognition engine that recognizes the user's emotion and adjusts and generates the facial expression of the facial image based on the emotion; means for presenting the composite image to the user and saving it after confirmation; and means for saving facial image data and feature point data received from the terminal. This enables the user to easily specify a desired pose and facial expression and automatically generate a natural composite photograph, and further enables the realization of a photograph with a realistic expression that reflects the user's emotion.
[0426] "User" refers to an individual who uses this system to register a facial image and perform various operations to generate a composite photograph.
[0427] "Facial image" refers to image data of a user's face, and includes photographs of the front, left side, and right side.
[0428] "Feature points" refer to position data of the eyes, nose, mouth, etc. extracted from a facial image, and are used to represent the structure and shape of the face.
[0429] "Feature point data" is a representation of feature points as numerical data, and is the basic data for generating a 3D model of a face.
[0430] A "3D facial model" refers to a digital model of a face expressed in three-dimensional space based on feature point data.
[0431] "Background photo" refers to photo data that is designated or taken by the user and used as the background of the composite photo.
[0432] "Angle" refers to a parameter that indicates the direction in which the 3D face model is facing relative to the background photo.
[0433] "Pose" refers to parameters that indicate the posture and gesture of a 3D facial model.
[0434] "Facial expression" refers to parameters that indicate emotions or expressions applied to a 3D facial model.
[0435] An "emotion recognition engine" refers to a software module that recognizes emotions from a user's facial image or voice input, and adjusts and generates facial expressions based on those emotions.
[0436] "Synthesis parameters" refer to the settings such as light direction and color tone required to naturally composite a 3D facial model onto a background photo.
[0437] A "composite photo" refers to an image created by combining a background photo with a 3D model of a face.
[0438] "Terminal" refers to a device such as a smartphone or PC on which a user runs an application.
[0439] "Server" refers to a remote server that stores facial images and feature point data and performs various processing.
[0440] MODE FOR CARRYING OUT THE INVENTION
[0441] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on a facial image registered by the user, creating a naturally composite image against a background photograph. It also incorporates an emotion engine that recognizes the user's emotions. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos. Below, we explain each step for implementing this invention.
[0442] 1. The process by which users register their facial images
[0443] User
[0444] The user launches the application on their smartphone or PC and selects the facial image registration menu.
[0445] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[0446] Terminal
[0447] The device uses libraries such as OpenCV and dlib to automatically extract feature points (such as the position of the eyes, nose, and mouth) from each captured facial image.
[0448] The device uses tools like Blender or Three.js to generate a 3D model of the face based on the feature point data.
[0449] The generated 3D model and feature point data are sent to the server.
[0450] server
[0451] The server receives the data sent from the device and stores it in a MySQL or MongoDB database.
[0452] The saved feature point data and 3D model are used in subsequent synthesis processes.
[0453] 2. Angle, pose, and facial expression generation process
[0454] User
[0455] The user selects a background photo within the application or takes a new one.
[0456] The user uses the UI within the application to specify the desired pose and expression.
[0457] server
[0458] The server uses Python's PIL library and OpenCV to analyze the background photo selected by the user and calculate the light direction and color tone.
[0459] The server uses a deep learning model (using TensorFlow or PyTorch) to automatically generate angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[0460] The server calculates the synthesis parameters for synthesizing the analysis data of the background photo with the 3D model.
[0461] Emotion Engine
[0462] The emotion engine uses Microsoft Azure's Face API and Google Cloud Vision API to recognize the user's emotions from registered facial images.
[0463] The facial expressions are adjusted and generated based on the user's desired pose and facial expression and the recognized emotion.
[0464] The emotion engine uses the Google Cloud Speech-to-Text API to analyze the user's voice input to recognize emotions and generate facial expressions based on them.
[0465] server
[0466] The server uses an automated script in Adobe Photoshop or ImageMagick to use the generated data and synthesis parameters to naturally synthesize the 3D face model with the background photo.
[0467] The server transmits the composite image to the terminal.
[0468] 3. Photo generation and storage process
[0469] Terminal
[0470] The terminal receives the composite image from the server, displays it on the screen, and asks the user for confirmation.
[0471] User
[0472] The user checks the composite image and, if satisfied, presses the "Save" button. Editing functions are available as needed.
[0473] Terminal
[0474] The device saves the composite image to local storage using the Android or iOS local storage API.
[0475] The device provides the option to share to social media and cloud services, using the Facebook API and Google Drive API.
[0476] Specific examples
[0477] 1. Example of face registration
[0478] User A starts the app and registers three facial images: front, left, and right.
[0479] The device uses OpenCV to extract feature points from these images and uses Blender to generate a 3D model of the face.
[0480] The generated data is sent to the server, which stores it in a database.
[0481] 2. Specific examples of emotion engines
[0482] User B selects a landscape photo as a background in the app, and the app uses the Microsoft Azure Face API to recognize B's current emotion.
[0483] If the user desires a "smiling" facial expression, but the emotion engine recognizes the emotion of "surprise," it applies the surprised expression to the facial image and generates a 3D model.
[0484] The server calculates the synthesis parameters based on this data and synthesizes it with the landscape photo.
[0485] 3. Examples of group photos
[0486] A family of four wants to take a photo together, but one of them is in a remote location and it is physically impossible to take the photo.
[0487] The three people took a group photo on site, and the facial image of one person in a remote location was registered through the application.
[0488] The user selects a group photo as a background within the app, and the server generates the appropriate angle, pose, and expression based on a single face model located remotely, and combines it with the background photo.
[0489] The emotion engine recognizes the current emotion of a person in a remote location and generates and adjusts facial expressions to match that emotion.
[0490] The composite image is displayed on the smartphone and can be saved by the user.
[0491] As described above, by implementing this invention, it becomes possible to easily create ideal photos even in places where it is difficult to take selfies, and to create high-quality group photos without physically gathering together. Furthermore, the emotion engine can generate natural facial expressions that match the user's emotions, improving image quality and satisfaction.
[0492] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0493] Processing flow
[0494] 1. The process by which users register their facial images
[0495] Step 1:
[0496] The user starts the application and selects the face image registration menu.
[0497] Input: An action on the user interface.
[0498] Output: Camera launch and registration menu display.
[0499] Specific operation: Click the "Register face image" button on the application's home screen, and the following screen will appear.
[0500] Step 2:
[0501] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[0502] Input: Camera operation and photography.
[0503] Output: 3 face images (front, left, right).
[0504] Specific operation: The camera screen will appear and you will be prompted to "Take a photo of the front" and then take a photo, repeating this three times.
[0505] Step 3:
[0506] The device uses libraries such as OpenCV and dlib to automatically extract feature points (such as the position of the eyes, nose, and mouth) from each captured facial image.
[0507] Input: 3 face images.
[0508] Output: Feature point data for each face image.
[0509] Specific operation: Using the dlib facial feature detection model, an algorithm is run to extract 68 facial feature points from an image.
[0510] Step 4:
[0511] The device uses Blender or Three.js to generate a 3D model of the face based on feature point data.
[0512] Input: feature point data.
[0513] Output: 3D face model data.
[0514] Specific operation: Using Blender's Python API, run a script that generates a 3D model based on feature point data.
[0515] Step 5:
[0516] The device sends the generated 3D model and feature point data to the server.
[0517] Input: 3D face model data, feature point data.
[0518] Output: Sending data to the server.
[0519] Specific operation: 3D model data and feature point data are sent to the server using an HTTP POST request.
[0520] Step 6:
[0521] The server receives the data sent from the device and stores it in a MySQL or MongoDB database.
[0522] Input: 3D face model data, feature point data.
[0523] Output: Data stored in the database.
[0524] Specific operation: Analyzes the received data and executes SQL insert statements or MongoDB document inserts to store it in the database.
[0525] 2. Angle, pose, and facial expression generation process
[0526] Step 1:
[0527] The user selects a background photo within the application or takes a new one.
[0528] Input: Select or take a background photo.
[0529] Output: A background photo that you select or take.
[0530] Specific action: Select a photo from the gallery or launch the camera to take a new photo.
[0531] Step 2:
[0532] The user uses the UI within the application to specify the desired pose and expression.
[0533] Input: Specify pose and facial expression.
[0534] Output: Specified pose and facial expression.
[0535] Specific Actions: Use the drop-down menus and sliders to select the desired pose and expression.
[0536] Step 3:
[0537] The server uses Python's PIL library and OpenCV to analyze the background photo selected by the user and calculate the light direction and color tone.
[0538] Input: Background photo data.
[0539] Output: Light direction and color data.
[0540] Specific operation: OpenCV is used to perform image histogram and edge detection, and algorithms are run to identify the direction and color of the light source.
[0541] Step 4:
[0542] The server uses a deep learning model (using TensorFlow or PyTorch) to automatically generate angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[0543] Input: 3D model data, pose and facial expression specification.
[0544] Output: Newly generated angle, pose, and expression data.
[0545] Specific operation: Input data is fed to a pre-trained model using TensorFlow or PyTorch, and the resulting angle and facial expression data is obtained.
[0546] Step 5:
[0547] The emotion engine uses Microsoft Azure's Face API and Google Cloud Vision API to recognize the user's emotions from registered facial images.
[0548] Input: Facial image data.
[0549] Output: Emotion data.
[0550] Specific operation: Send facial image data to the API and obtain emotion recognition results.
[0551] Step 6:
[0552] The server calculates the synthesis parameters for synthesizing the analysis data of the background photo with the 3D model.
[0553] Input: Light direction and color data, 3D model data.
[0554] Output: Synthesis parameters.
[0555] Specific operation: Calculates parameters for naturally combining a 3D model with a background photo, taking into account light direction, color tone, and shadow position.
[0556] Step 7:
[0557] The server uses an automated script in Adobe Photoshop or ImageMagick to use the generated data and synthesis parameters to naturally synthesize the 3D face model with the background photo.
[0558] Input: 3D model data, background photo data, synthesis parameters.
[0559] Output: Composite image.
[0560] Specific operation: Using Adobe Photoshop scripts and ImageMagick, the 3D model is composited with a background photo based on the provided parameters.
[0561] Step 8:
[0562] The server transmits the composite image to the terminal.
[0563] Input: Synthetic image data.
[0564] Output: Send image to device.
[0565] Specific operation: The composite image is sent to the device via an HTTP response or WebSocket message.
[0566] 3. Photo generation and storage process
[0567] Step 1:
[0568] The terminal receives the composite image from the server, displays it on the screen, and asks the user for confirmation.
[0569] Input: Synthetic image data from the server.
[0570] Output: Image displayed to the user.
[0571] Specific operation: The received image data is set to a GUI element for display and displayed to the user.
[0572] Step 2:
[0573] The user checks the composite image and, if satisfied, presses the "Save" button. Editing functions are available as needed.
[0574] Input: User confirmation action.
[0575] Output: Save image or edit data.
[0576] Specific operation: Press the "Save" button to save the image. Press the edit button to display the edit screen.
[0577] Step 3:
[0578] The device saves the composite image to local storage using the Android or iOS local storage API.
[0579] Input: Synthetic image data.
[0580] Output: Image data saved to local storage.
[0581] What it does: It uses Android and iOS file storage APIs to save the composite image to the user's device.
[0582] Step 4:
[0583] The device provides the option to share to social media and cloud services, using the Facebook API and Google Drive API.
[0584] Input: User's share action.
[0585] Output: Upload images to social media or cloud services.
[0586] Specific operation: When you press the share button, the image will be uploaded via the API of the selected social media or cloud service.
[0587] (Application example 2)
[0588] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0589] In today's digital society, there is a growing demand for generating high-quality images with natural expressions and poses in situations where it is difficult for users to take selfies, where it is physically impossible to take group photos, or when virtual try-on simulations are required. However, conventional systems have struggled to generate natural images in these situations, particularly in automatically adjusting facial expressions to reflect the user's emotions. Furthermore, in virtual stores, there is a lack of technology to enhance the user's virtual try-on experience, and there is a need to improve user satisfaction.
[0590] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a means for a user to register their own facial image, a means for extracting feature points from the registered facial image and generating feature point data, and a means for generating a 3D facial model based on the feature point data. This enables high-quality images with natural angles, poses, and facial expressions to be automatically generated and naturally combined with background photos, even in situations where it is difficult for the user to take a selfie. The system also includes a means for recognizing the user's emotions and adjusting the facial expression based on the emotions, thereby enabling more natural image generation. Furthermore, the system also includes a means for simulating virtual trying-on of an item selected by the user in a virtual store, thereby improving the user's virtual trying-on experience.
[0591] A "user" is an entity that registers their own facial image in the system and uses the image generation and virtual try-on simulation.
[0592] "Means for registering face images" is a function that allows users to upload their own face images to the system.
[0593] The "means for extracting feature points" is a function that automatically extracts feature points such as the eyes, nose, and mouth from registered face images.
[0594] "Feature point data" is data including the position coordinates of the eyes, nose, mouth, etc. extracted from a face image.
[0595] The "means for generating a 3D face model" is a function that constructs a three-dimensional model of the user's face based on feature point data.
[0596] "Means for automatically generating angles, poses, and facial expressions" is a function that automatically generates specified angles, poses, and facial expressions for a 3D model.
[0597] A "background photo" is a photo of the background to be combined with the user's facial image.
[0598] "Natural synthesis means" is a function that seamlessly integrates a facial image with the generated angle, pose, and expression into a background photo.
[0599] The "means for presenting to the user" is a function for displaying the synthesized image on the user's screen.
[0600] The "means for saving" is a function that allows the user to save the composite image after checking it.
[0601] The "means for simulating virtual fitting" is a function that combines an object (for example, clothing) selected by the user with a three-dimensional model to generate the results of virtual fitting.
[0602] The "means for recognizing emotions" is a function that automatically determines the user's current emotions from their voice and image.
[0603] The "means for adjusting facial expressions" is a function that appropriately changes the facial expression of a 3D model based on the recognized emotion.
[0604] This system uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by the user, creating images that blend naturally with the background photograph. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to generate more natural facial expressions. This system is particularly intended for virtual try-on and image generation in virtual stores.
[0605] 1. The process by which users register their facial images
[0606] First, the user launches the application on their smartphone or PC and selects the facial image registration menu. For registration, they take photos of their face from the front, left side, and right side. The device extracts feature points from the captured facial images and generates a 3D model of the face based on these. This process uses an image processing library such as OpenCV. The generated 3D model and feature point data are then sent to the server, which then stores the data in a database.
[0607] 2. Virtual try-on simulation process
[0608] Within the application, users select the clothing and accessories they want to try on. The server then combines the selected product information with a 3D model of the user's face to generate a virtual try-on experience. This process uses 3D modeling libraries such as Three.js. Users can also specify their desired pose and angle.
[0609] 3. Emotion Recognition and Facial Expression Generation Process
[0610] The emotion engine analyzes the user's facial images and voice to recognize their current emotions. Based on these emotions, it then appropriately adjusts and generates the facial expressions of the 3D model. This process uses the Microsoft Azure Emotion API and other tools.
[0611] 4. The process of combining with the background photo
[0612] After the user selects a background photo, the server analyzes the photo, calculates the light direction and color tone, calculates the parameters necessary for synthesis, and expresses the facial image as a 3D model, which is then naturally synthesized with the background photo. At this stage, the pose and facial expression specified by the user are reflected.
[0613] 5. Synthesis and saving process
[0614] The server generates a composite image and sends it to the device, which then presents it to the user, who can save it after reviewing it. The option to share it on social media or cloud services is also provided.
[0615] Specific examples
[0616] For example, if a user tries on a new summer dress and generates a composite photo with a shopping mall background, the process would proceed as follows:
[0617] The user registers a face image and selects a new summer dress.
[0618] The selected dress is then composited with a 3D model of a face, a smiling expression is applied, and the dress is composited with the background of a shopping mall.
[0619] The generated composite image is displayed on a smartphone, allowing the user to view and save it.
[0620] Prompt Sentence Examples
[0621] "Generate a composite photo of you trying on a new summer dress and smiling with a shopping mall background"
[0622] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0623] Step 1: User registers a face image
[0624] The user launches the application on their smartphone or PC and selects the facial image registration menu. Next, they take photos of their face from the front, left side, and right side and upload them to the application. These input images serve as the base data for subsequent processing. The device receives these facial images and uses OpenCV to extract feature points such as the eyes, nose, and mouth. Feature point data is generated, and a 3D model of the face is created based on this. This 3D model is then used in subsequent synthesis processing.
[0625] Step 2: User selects an object for virtual try-on
[0626] Within the application, the user selects the clothing and accessories they want to try on. The data for this selection is provided as input to the system. The server receives this and prepares to combine a 3D model of the user's face with the information about the selected items. This process uses a 3D modeling library such as Three.js, and the output is a virtual try-on ready image.
[0627] Step 3: The emotion engine analyzes facial images and voice
[0628] The user's facial image and voice are provided as input to the emotion engine. The server analyzes this data using Microsoft Azure Emotion API or similar to recognize the user's current emotion. Emotion data is generated as a result of the analysis. This emotion data is then used to adjust facial expressions.
[0629] Step 4: Generate and adjust facial expressions based on emotion data
[0630] The server adjusts the facial expression of the user's 3D model based on the emotion data obtained in step 3. For example, if the user is excited, it applies a smiling expression. The individual feature point data of the registered facial image is used to generate the facial expression of the 3D model. The output is the adjusted 3D model.
[0631] Step 5: Select a background photo and analyze it
[0632] The user selects a background photo, which is sent as input to the server, which analyzes the background photo to calculate the light direction and color tone. Image processing algorithms are used for the analysis, and light and color parameters for compositing are generated as output.
[0633] Step 6: Generate the composite image
[0634] The server uses the 3D model of the virtual try-on garment generated in step 2, the 3D model of the user adjusted in step 4, and the light and color parameters obtained in step 5 to naturally combine the background photo and facial image. The synthesis algorithm outputs a natural-looking composite image.
[0635] Step 7: Send the composite image to your device
[0636] The server sends the composite image to the terminal, which then presents the received composite image to the user, who then checks the presented composite image on the screen.
[0637] Step 8: Save and share your composite image
[0638] The user can review the composite image and save it if they are satisfied. The device then saves the composite image in local storage and provides the option to share it on social media or cloud services if desired. Finally, the saved composite image can be used according to the user's purpose.
[0639] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0640] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0641] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0642] [Second embodiment]
[0643] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0644] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0645] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0646] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0647] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0648] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0649] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0650] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0651] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0652] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0653] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0654] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0655] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and creates images that are naturally composited against a background photograph. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos.
[0656] 1. The process by which users register their facial images
[0657] User
[0658] First, the user launches the application on their smartphone or PC and selects the facial image registration menu.
[0659] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[0660] Terminal
[0661] The device automatically extracts feature points (e.g., the positions of the eyes, nose, and mouth) from each captured facial image.
[0662] A 3D model of the face is generated based on the feature point data.
[0663] The generated 3D model and feature point data are sent to the server.
[0664] server
[0665] The server receives the data sent from the terminal and stores it in a database.
[0666] The saved feature point data and 3D model are used in subsequent synthesis processes.
[0667] 2. Angle, pose, and facial expression generation process
[0668] User
[0669] Users can select a background photo within the application or take a new one.
[0670] The user specifies the desired pose and facial expression.
[0671] server
[0672] The server analyzes the background photo selected by the user and calculates the direction and color of the light.
[0673] Based on the pose and facial expression specified by the user, the angle, pose, and facial expression are automatically generated for the registered 3D model.
[0674] Calculate synthesis parameters for synthesizing the background photo analysis data and the 3D model.
[0675] The facial image is represented as a 3D model based on the generated angle, pose, and expression, and is then naturally combined with the background photo.
[0676] Terminal
[0677] The terminal receives the composite image and the composite parameters from the server and performs the composite process.
[0678] The synthesized image is presented to the user for confirmation.
[0679] 3. Photo generation and storage process
[0680] User
[0681] The user checks the generated composite image and presses the save button if satisfied.
[0682] Use further editing functions as needed.
[0683] Terminal
[0684] The terminal stores the composite image in local storage.
[0685] It also provides users with the option to share images on social media or cloud services, depending on their choice.
[0686] Specific examples
[0687] 1. Example of face registration
[0688] User A starts the app and registers three facial images: front, left, and right.
[0689] The device extracts feature points from these images and creates a 3D model of the face.
[0690] The generated data is sent to the server, which stores it in a database.
[0691] 2. Examples of group photos
[0692] A family of four wants to take a photo together, but one of them is in a remote location and it is physically impossible to take the photo.
[0693] The three people took a group photo on site, and the facial image of one person in a remote location was registered through the application.
[0694] Users select a group photo as a background within the app, and the server generates the appropriate angle, pose, and facial expression based on a remote facial model of a single person, and combines it with the background photo.
[0695] The composite image is displayed on the smartphone and can be saved by the user.
[0696] As described above, by implementing this invention, it becomes possible to easily generate ideal photos even in places where it is difficult to take selfies, and it is also possible to create high-quality group photos without having to physically gather together.
[0697] The processing flow will be explained below.
[0698] Program processing details
[0699] The process by which users register their facial images
[0700] Step 1:
[0701] The user launches the application on their smartphone or PC.
[0702] Step 2:
[0703] The user selects a face image registration menu and follows the instructions to take photos of the face from the front, left side, and right side.
[0704] Step 3:
[0705] The terminal receives each captured face image.
[0706] Step 4:
[0707] The device extracts feature points (such as the positions of the eyes, nose, and mouth) from the facial image.
[0708] Step 5:
[0709] The device generates a 3D model of the user's face based on the extracted feature point data.
[0710] Step 6:
[0711] The device sends the generated 3D model and feature point data to the server.
[0712] Step 7:
[0713] The server stores the received data in a database.
[0714] The process of generating angles, poses, and facial expressions
[0715] Step 1:
[0716] The user selects the photo creation menu of the application.
[0717] Step 2:
[0718] The user selects a background photo or takes a new one.
[0719] Step 3:
[0720] The user specifies the desired pose and facial expression.
[0721] Step 4:
[0722] The server receives the background photo and runs an image analysis algorithm to analyze the light direction and color tone.
[0723] Step 5:
[0724] Based on the analysis results, the server calculates the synthesis parameters for combining the image with the background photo.
[0725] Step 6:
[0726] The server automatically generates angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[0727] Step 7:
[0728] The server uses the generated data and synthesis parameters to represent the facial image as a 3D model and naturally combines it with the background photo.
[0729] Step 8:
[0730] The server transmits the composite image to the terminal.
[0731] Photo generation and storage process
[0732] Step 1:
[0733] The terminal receives the composite image sent from the server.
[0734] Step 2:
[0735] The terminal displays the received composite image on the screen and asks the user for confirmation.
[0736] Step 3:
[0737] The user checks the composite image and, if satisfied, presses the save button.
[0738] Step 4:
[0739] The device stores the composite image in local storage.
[0740] Step 5:
[0741] The device displays a notification to the user that the save is complete.
[0742] Step 6:
[0743] If the user wishes to share the image on a social networking site or cloud service, the device will provide the option to do so and process it accordingly.
[0744] Example 1
[0745] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0746] In conventional facial image synthesis systems, it was difficult for users to naturally combine their own facial images with background images, and differences in light direction and color tone often resulted in unnatural results. It was also difficult to naturally reflect the user's desired pose and facial expression. Furthermore, the process of efficiently extracting feature points from multiple facial images and generating a 3D model was complicated, resulting in poor usability for the entire system.
[0747] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0748] In this invention, the server includes means for a user to register his or her own facial image, means for extracting feature points from the registered facial image and generating feature point data, means for generating a 3D facial model based on the feature point data, means for analyzing a background image and identifying the light direction and color tone, means for automatically generating an angle, pose, and facial expression based on the analysis results, means for representing the facial image as a 3D model according to the generated angle, pose, and facial expression and naturally combining it with the background image, and means for presenting the combined image to the user and saving it after confirmation. This allows the user to easily generate natural combined images and improves operability, enabling more satisfying facial image synthesis.
[0749] "User" refers to an individual or group that uses this system to register their own facial images and create composite images.
[0750] "Means for registering" refers to a method or device that allows users to upload and store their facial images in the system.
[0751] "Feature points" are data that indicate specific parts of a face image, such as the eyes, nose, and mouth.
[0752] "Feature point data" is a data set that includes information on feature points extracted from a face image.
[0753] A "3D facial model" is a three-dimensional shape of a face reconstructed in three dimensions based on feature point data of a facial image.
[0754] A "background image" is a photograph or image selected by the user as a composite object.
[0755] The "analyzing means" refers to a method or device for calculating and analyzing the light direction and color tone of the background image.
[0756] "Means for automatically generating angles, poses, and facial expressions" refers to a method or device that has the function of automatically creating facial angles, poses, and facial expressions based on user specifications and analysis results.
[0757] The "synthesis means" refers to a method or device for naturally synthesizing the generated 3D face model with a background image.
[0758] The "presenting means" refers to a method or device for displaying the synthesized image to the user and requesting confirmation.
[0759] "Storing means" refers to a method or device for storing the confirmed composite image within the system or in external storage.
[0760] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and creates images that are naturally composited against a background photograph. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos.
[0761] The process by which users register their facial images
[0762] User
[0763] The user launches the application on their smartphone or PC and selects the facial image registration menu.
[0764] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[0765] Terminal
[0766] The device automatically extracts feature points (e.g., the positions of the eyes, nose, and mouth) from each captured facial image using Python's OpenCV library.
[0767] Based on the extracted feature point data, a 3D model of the face is generated using 3D modeling software such as Blender.
[0768] The generated 3D model and feature point data are sent to the server via REST API.
[0769] server
[0770] The server receives the data sent from the device and stores it in a MySQL database using the Django framework.
[0771] The saved feature point data and 3D model are used in subsequent synthesis processes.
[0772] The process of generating angles, poses, and facial expressions
[0773] User
[0774] Users can select a background photo within the application or take a new one.
[0775] The user specifies the desired pose and facial expression.
[0776] server
[0777] The server analyzes the background photo selected by the user and calculates the light direction and color tone using PIL (Python Imaging Library) and OpenCV.
[0778] Based on the pose and facial expression specified by the user, the angle, pose, and facial expression are automatically generated for the registered 3D model. Here, generative AI models (such as StyleGAN and DeepFaceLab) are utilized.
[0779] Calculate synthesis parameters for synthesizing the background photo analysis data and the 3D model.
[0780] The facial image is represented based on a 3D model according to the generated angle, pose, and expression, and is then naturally combined with the background photo.
[0781] The process by which the device generates a composite image and presents it to the user
[0782] Terminal
[0783] The terminal receives the composite image and the composite parameters from the server.
[0784] The synthesis process is performed and the completed image is presented to the user.
[0785] User
[0786] The user checks the presented composite image and presses the save button if satisfied.
[0787] Further modifications can be made within the app if needed.
[0788] The generated images can also be shared on social media or cloud services.
[0789] Specific examples
[0790] Specific examples of face registration
[0791] The user launches the app and takes a series of photos of their face - front, left, and right - and registers them. The device uses Python's OpenCV library to extract feature points from these images and creates a 3D face model using Blender. The generated data is sent to the server via a REST API and stored in a MySQL database using the Django framework.
[0792] Examples of group photos
[0793] If a family wants to take a group photo together, but one person is remote and cannot physically be present, the remote person can register their face image through the app. The user then takes a group photo on-site and selects a background photo within the app. The server then analyzes the image, calculating the light direction and color tone, and generates the appropriate angle, pose, and facial expression based on a 3D facial model of the remote person, which is then combined with the background photo. The resulting composite image is then displayed on the user's smartphone.
[0794] Prompt Sentence Examples
[0795] "We will create a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and then creates an image that is composited into a background photo. We will also analyze the background photo, such as the direction of light and color tone, to ensure that the composite image looks natural."
[0796] By implementing this invention, ideal selfies can be easily generated even in places where it is difficult to take selfies, and high-quality group photos can be created without physically gathering together.
[0797] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0798] The flow of this system's program processing
[0799] Step 1:
[0800] explanation
[0801] Users launch the application on their smartphone or PC, select the face image registration menu, and follow the app's guide to take and register photos of their face from the front, left side, and right side.
[0802] Input and Output
[0803] Input: Face image (front, left, right)
[0804] Output: Three registered face images
[0805] Specific actions
[0806] The user clicks the "Register Face Image" button on the app, activates the camera to take face images from various angles, and then presses the "Confirm / Register" button.
[0807] Step 2:
[0808] explanation
[0809] The device automatically extracts feature points (such as the positions of the eyes, nose, and mouth) from the captured facial image using Python's OpenCV library. Based on the extracted feature point data, a 3D model of the face is generated using 3D modeling software such as Blender.
[0810] Input and Output
[0811] Input: Three registered face images
[0812] Output: feature point data, 3D face model
[0813] Specific actions
[0814] The device uses the OpenCV library to detect feature points from each facial image, then uses Blender to generate a 3D model of the face based on these feature points. The generated model and feature point data are then sent to the server via a REST API.
[0815] Step 3:
[0816] explanation
[0817] The server receives the data sent from the device and stores it in a MySQL database using the Django framework. The stored feature point data and 3D model are then used in the synthesis process.
[0818] Input and Output
[0819] Input: feature point data, 3D face model
[0820] Output: Data stored in the database
[0821] Specific actions
[0822] The server receives the transmitted data via the Django framework and stores it in a MySQL database.
[0823] Step 4:
[0824] explanation
[0825] Users can select a background photo within the application or take a new one, and specify the desired pose and facial expression.
[0826] Input and Output
[0827] Input: Background photo, desired pose and facial expression
[0828] Output: User selections and specifications
[0829] Specific actions
[0830] Once the user selects (or takes) a background photo, a UI for specifying the pose and expression is displayed.
[0831] Step 5:
[0832] explanation
[0833] The server analyzes the background photo selected by the user and calculates the light direction and color tone using PIL (Python Imaging Library) and OpenCV.
[0834] Input and Output
[0835] Input: Background photo
[0836] Output: Analysis results of light direction and color tone
[0837] Specific actions
[0838] The server loads the background photo using the PIL or OpenCV library and analyzes the light direction and color tone.
[0839] Step 6:
[0840] explanation
[0841] The server uses a generative AI model (e.g., StyleGAN or DeepFaceLab) to automatically generate angles, poses, and expressions for the registered 3D model based on the poses and expressions specified by the user.
[0842] Input and Output
[0843] Input: feature point data, 3D face model, desired pose and expression
[0844] Output: Generated angles, poses, and expressions
[0845] Specific actions
[0846] The server uses a generative AI model to generate poses and facial expressions based on feature point data and a 3D model.
[0847] Step 7:
[0848] explanation
[0849] The server calculates synthesis parameters for combining the analysis data of the background photo with a 3D model, and then represents the facial image as a 3D model according to the generated angle, pose, and expression, which is then naturally combined with the background photo.
[0850] Input and Output
[0851] Input: background photo, analysis data, generated angles, poses, facial expressions
[0852] Output: Synthesized face image
[0853] Specific actions
[0854] The server calculates the synthesis parameters and uses OpenCV to naturally synthesize the face image and background photo.
[0855] Step 8:
[0856] explanation
[0857] The terminal receives the composite image and the composite parameters from the server, performs the composite process, and presents the completed composite image to the user.
[0858] Input and Output
[0859] Input: synthetic image, synthetic parameters
[0860] Output: The composite image presented to the user
[0861] Specific actions
[0862] The device retrieves the composite image from the server via an HTTP request and displays it to the user on the app.
[0863] Step 9:
[0864] explanation
[0865] The user can review the composite image presented to them and, if satisfied, press the save button. If necessary, they can make further edits within the app. The generated image can also be shared via social media or cloud services.
[0866] Input and Output
[0867] Input: Composite image
[0868] Output: Saved and shared images
[0869] Specific actions
[0870] When the user presses the "Save" button, the image is saved to the smartphone's local storage, and when the user presses the "Share" button, the image is uploaded to social media or a cloud service.
[0871] (Application example 1)
[0872] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0873] Modern electronic payment services require enhanced security. However, many current systems have limited user authentication methods, making it difficult to simultaneously improve convenience and security. Furthermore, many systems can only authenticate users at specific angles or poses, which can be inconvenient for users. There is a need to solve these issues and realize more accurate and convenient facial recognition. Furthermore, a system is needed that can accurately authenticate users even when they have different facial expressions or poses.
[0874] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0875] In this invention, the server includes: a means for a user to register his or her own facial image; a means for extracting feature points from the registered facial image and generating feature point data; a means for generating a 3D facial model based on the feature point data; a means for automatically generating an angle, pose, and facial expression to be combined with a background photograph; a means for representing the facial image as a 3D model according to the generated angle, pose, and facial expression and combining it naturally with the background photograph; a means for performing electronic authentication using analysis data of the background photograph and the user's 3D model; and a means for presenting the combined image to the user and saving it after confirmation. This enables advanced electronic authentication by generating a natural combined image based on the user's registered facial image and background photograph. A means for the user to specify a desired pose and facial expression; a means for generating corresponding movements for the facial portion of the 3D model based on the specified pose and facial expression; a means for performing electronic authentication based on the generated 3D model and a captured image of the user; and a means for reconstructing the facial image according to the generated movements, thereby providing an authentication system that combines convenience and security.
[0876] "Means for users to register their own facial images" refers to a function that allows users to register their own facial images using devices such as smartphones or computers.
[0877] "Means for extracting feature points and generating feature point data" refers to a function that uses AI technology to detect feature points such as the eyes, nose, and mouth from registered facial images and converts them into data.
[0878] The "means for generating a 3D model of a face" is a function for generating a 3D model of a face using feature point data.
[0879] "Means for automatically generating angles, poses, and facial expressions for compositing with background photos" is a function that uses AI to automatically generate the optimal facial angle, pose, and facial expression for compositing with a background photo.
[0880] "Means for representing a facial image as a 3D model according to the generated angle, pose, and expression, and for naturally combining it with a background photograph" refers to a function for reproducing a facial image as a 3D model based on the generated angle, pose, and expression, and for naturally combining it with a background photograph.
[0881] "Means for electronic authentication using background photo analysis data and a 3D model of the user" is a function for securely and accurately authenticating a user using analysis data of the light direction and color tone of the background photo and a 3D model of the user.
[0882] The "means for presenting the synthesized image to the user and saving it after confirmation" is a function for displaying the generated synthesized image to the user and saving the image after obtaining confirmation.
[0883] "Means for generating corresponding movements for the facial part of a 3D model based on a specified pose or expression" is a function that moves the facial part of a 3D model according to the pose or expression specified by the user.
[0884] The "means for reconstructing a facial image according to the generated movements" is a function for reconstructing a facial image based on the generated poses and facial movements.
[0885] "Means for electronic authentication" refers to an authentication function that uses a user's facial image or 3D model to allow access to the system.
[0886] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and creates images that are naturally composited with background photographs. Below, we will provide a detailed explanation of how this system is realized.
[0887] 1. Facial image registration process
[0888] Users register their facial images using a smartphone or computer. Specifically, they take photos of their face from the front, left side, and right side, and register them through the application.
[0889] The device receives a facial image and extracts its features using a library such as dlib. This generates key feature data such as the eyes, nose, and mouth. A 3D model of the face is then generated based on this feature data. This 3D model and feature data are then sent to a server and stored in a database.
[0890] 2. Angle, pose, and facial expression generation process
[0891] The server analyzes the background photo selected or taken by the user. Specifically, it calculates the direction of light and color tone. Based on this analysis data, it automatically generates the appropriate angle, pose, and expression for the registered 3D model according to the pose and expression specified by the user. Using AI technology such as DeepFace, it calculates synthesis parameters to naturally combine the user's 3D face model with the background photo.
[0892] 3. Electronic Authentication Process
[0893] The device performs electronic authentication using background photo analysis data and a 3D model of the user. Specifically, it uses DeepFace's Facenet model to match the user's captured image with the 3D model for accurate authentication.
[0894] 4. Photo Generation and Storage Process
[0895] Users can review the generated composite image and save it if they are satisfied. Further editing is available if necessary. The composite image is saved to local storage, and options for sharing it to social media and cloud services are also provided.
[0896] The system that realizes this invention executes a series of processes, from registering a facial image, generating a 3D model, using that model for electronic authentication, and finally generating and saving a composite image. This enables users to be accurately and securely authenticated even with different poses and expressions, and to obtain a high-quality composite image.
[0897] Examples:
[0898] A user installs the FacePay app and registers images of their face, showing the front, left, and right sides. After making a purchase at a convenience store, they perform payment using facial recognition, and the payment is completed upon successful authentication.
[0899] Example prompt sentence:
[0900] "Design an application that generates a 3D face model from a facial image registered by a user and uses that model for facial authentication during electronic payments."
[0901] As described above, by implementing the present invention, it is possible to generate ideal photos even when selfies are difficult, and to create high-quality group photos without physically gathering together. Furthermore, by using this system, the security of electronic payments is improved, and user convenience is also enhanced.
[0902] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0903] Step 1:
[0904] The user registers a facial image. In this process, the user launches an application on a smartphone or PC and takes and registers facial images from the front, left side, and right side.
[0905] Input: Face image taken by the user (front, left, right)
[0906] Output: Facial feature point data and a 3D model of the face generated by the device
[0907] Step 2:
[0908] The device processes the facial image provided by the user. Specifically, it uses the dlib library to extract feature points such as the eyes, nose, and mouth from the captured facial image. Based on this feature point data, a 3D model of the face is generated. The generated data is sent to a server and stored in a database.
[0909] Input: A face image obtained from the user
[0910] Output: Feature point data and a 3D face model
[0911] Step 3:
[0912] When a user selects or takes a background photo, the server analyzes the background photo, detecting the direction and color of the light, and then uses this information to calculate the direction and color of the light in the background photo.
[0913] Input: A background photo selected or taken by the user
[0914] Output: Analysis data of background photo (light direction and color tone)
[0915] Step 4:
[0916] The server automatically generates the angle, pose, and facial expression for the registered 3D model according to the pose and facial expression specified by the user. Based on this information, the server expresses the facial image as a 3D model and calculates synthesis parameters for naturally combining it with the background photo.
[0917] Input: Analysis data of user-specified poses and expressions, and background photos
[0918] Output: angle, pose, facial expression, synthesis parameters
[0919] Step 5:
[0920] The device uses the synthesis parameters and 3D model received from the server to naturally synthesize the background photo and facial image. Specifically, the synthesis process is performed using DeepFace's AI technology.
[0921] Input: Synthesis parameters and 3D model provided by the server
[0922] Output: Composite image
[0923] Step 6:
[0924] The device will then present the combined image to the user for review, and if the user is satisfied, the image will be saved to local storage, with the option to share it via social media or cloud services if desired.
[0925] Input: The composite image
[0926] Output: User confirmation and save
[0927] Step 7:
[0928] The device performs electronic authentication using background photo analysis data and a 3D model of the user, and uses DeepFace's Facenet model to match the captured image with the 3D model.
[0929] Input: Captured image, background photo analysis data, user 3D model
[0930] Output: Authentication result (success or failure)
[0931] Through a series of processes from registration to synthesis and authentication, highly accurate and convenient electronic authentication is achieved.
[0932] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0933] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by the user, creating images that blend naturally with the background photo, and further combines it with an emotion engine that recognizes the user's emotions. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos.
[0934] 1. The process by which users register their facial images
[0935] User
[0936] First, the user launches the application on their smartphone or PC and selects the facial image registration menu.
[0937] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[0938] Terminal
[0939] The device automatically extracts feature points (such as the position of the eyes, nose, and mouth) from each captured facial image.
[0940] A 3D model of the face is generated based on the feature point data.
[0941] The generated 3D model and feature point data are sent to the server.
[0942] server
[0943] The server receives the data sent from the terminal and stores it in a database.
[0944] The saved feature point data and 3D model are used in subsequent synthesis processes.
[0945] 2. Angle, pose, and facial expression generation process
[0946] User
[0947] Users can select a background photo within the application or take a new one.
[0948] The user specifies the desired pose and facial expression.
[0949] server
[0950] The server analyzes the background photo selected by the user and calculates the direction and color of the light.
[0951] Based on the pose and facial expression specified by the user, the angle, pose, and facial expression are automatically generated for the registered 3D model.
[0952] Calculate synthesis parameters for synthesizing the background photo analysis data and the 3D model.
[0953] Emotion Engine
[0954] The emotion engine recognizes the user's emotion from the registered facial image.
[0955] The facial expression is adjusted and generated based on the user's desired pose, facial expression, and recognized emotion.
[0956] The emotion engine can also analyze the user's voice input, recognize the user's emotion based on the analysis results, and generate facial expressions based on that.
[0957] server
[0958] The server uses the generated data and synthesis parameters to represent the facial image as a 3D model and naturally combines it with the background photo.
[0959] The server transmits the composite image to the terminal.
[0960] 3. Photo generation and storage process
[0961] Terminal
[0962] The terminal receives the composite image from the server.
[0963] The terminal displays the composite image on the screen and asks the user for confirmation.
[0964] User
[0965] The user checks the composite image and, if satisfied, presses the save button.
[0966] Use further editing functions as needed.
[0967] Terminal
[0968] The device stores the composite image in local storage.
[0969] It also provides the option to share to social media and cloud services.
[0970] Specific examples
[0971] 1. Example of face registration
[0972] User A starts the app and registers three facial images: front, left, and right.
[0973] The device extracts feature points from these images and generates a 3D model of the face.
[0974] The generated data is sent to the server, which stores it in a database.
[0975] 2. Specific examples of emotion engines
[0976] User B selects a landscape photo as a background in the app, and the app recognizes B's current emotion.
[0977] If the user desires a "smiling" facial expression, but the emotion engine recognizes the emotion of "surprise," it applies the surprised expression to the facial image and generates a 3D model.
[0978] The server calculates the synthesis parameters based on this data and synthesizes it with the landscape photo.
[0979] 3. Examples of group photos
[0980] A family of four wants to take a photo together, but one of them is in a remote location and it is physically impossible to take the photo.
[0981] The three people took a group photo on site, and the facial image of one person in a remote location was registered through the application.
[0982] Users select a group photo as a background within the app, and the server generates the appropriate angle, pose, and facial expression based on a remote facial model of a single person, and combines it with the background photo.
[0983] The emotion engine recognizes the current emotion of a person in a remote location and generates and adjusts facial expressions to match that emotion.
[0984] The composite image is displayed on the smartphone and can be saved by the user.
[0985] As described above, by implementing this invention, it becomes possible to easily create ideal photos even in places where it is difficult to take selfies, and it is also possible to create high-quality group photos without physically gathering together.Furthermore, the emotion engine can generate natural facial expressions that better match the user's emotions, improving image quality and satisfaction.
[0986] The processing flow will be explained below.
[0987] Program processing details
[0988] The process by which users register their facial images
[0989] Step 1:
[0990] The user launches the application on their smartphone or PC.
[0991] Step 2:
[0992] The user selects a face image registration menu and follows the instructions to take photos of the face from the front, left side, and right side.
[0993] Step 3:
[0994] The terminal receives each captured face image.
[0995] Step 4:
[0996] The device extracts feature points (such as the positions of the eyes, nose, and mouth) from the facial image.
[0997] Step 5:
[0998] The device generates a 3D model of the user's face based on the extracted feature point data.
[0999] Step 6:
[1000] The device sends the generated 3D model and feature point data to the server.
[1001] Step 7:
[1002] The server stores the received data in a database.
[1003] The process of generating angles, poses, and facial expressions
[1004] Step 1:
[1005] The user selects the photo creation menu of the application.
[1006] Step 2:
[1007] The user selects a background photo or takes a new one.
[1008] Step 3:
[1009] The user specifies the desired pose and facial expression.
[1010] Step 4:
[1011] The server receives the background photo and runs an image analysis algorithm to analyze the light direction and color tone.
[1012] Step 5:
[1013] Based on the analysis results, the server calculates the synthesis parameters for combining the image with the background photo.
[1014] Step 6:
[1015] The server automatically generates angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[1016] Step 7:
[1017] The emotion engine recognizes the user's emotions from registered facial images or voice input.
[1018] Step 8:
[1019] The emotion engine adjusts and generates poses and expressions based on the recognized emotions.
[1020] Step 9:
[1021] The server uses the generated data and synthesis parameters to represent the facial image as a 3D model and naturally combines it with the background photo.
[1022] Step 10:
[1023] The server transmits the composite image to the terminal.
[1024] Photo generation and storage process
[1025] Step 1:
[1026] The terminal receives the composite image sent from the server.
[1027] Step 2:
[1028] The terminal displays the received composite image on the screen and asks the user for confirmation.
[1029] Step 3:
[1030] The user checks the composite image and, if satisfied, presses the save button.
[1031] Step 4:
[1032] The device stores the composite image in local storage.
[1033] Step 5:
[1034] The device displays a notification to the user that the save is complete.
[1035] Step 6:
[1036] If the user wishes to share the image on a social networking site or cloud service, the device will provide the option to do so and process it accordingly.
[1037] Example 2
[1038] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1039] Conventional photo compositing technologies have difficulty in generating natural poses and facial expressions desired by users, and require manual setting of compositing parameters such as light direction and color tone. This has required a great deal of time and effort to generate a photo that satisfies the user. Furthermore, when taking a group photo and not everyone is present, it is difficult to composite people in remote locations naturally. Furthermore, there has been a lack of mechanisms for appropriately reflecting facial expressions that match the user's emotions. The objective of this invention is to solve these problems and realize easier and more natural photo compositing.
[1040] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1041] In this invention, the server includes: means for a user to register his or her own facial image; means for extracting feature points from the registered facial image and generating feature point data; means for generating a 3D facial model based on the feature point data; means for automatically generating an angle, pose, and facial expression to be composited with a background photograph; means for representing the facial image as a 3D model according to the generated angle, pose, and facial expression and naturally composited with the background photograph; means including an emotion recognition engine that recognizes the user's emotion and adjusts and generates the facial expression of the facial image based on the emotion; means for presenting the composite image to the user and saving it after confirmation; and means for saving facial image data and feature point data received from the terminal. This enables the user to easily specify a desired pose and facial expression and automatically generate a natural composite photograph, and further enables the realization of a photograph with a realistic expression that reflects the user's emotion.
[1042] "User" refers to an individual who uses this system to register a facial image and perform various operations to generate a composite photograph.
[1043] "Facial image" refers to image data of a user's face, and includes photographs of the front, left side, and right side.
[1044] "Feature points" refer to position data of the eyes, nose, mouth, etc. extracted from a facial image, and are used to represent the structure and shape of the face.
[1045] "Feature point data" is a representation of feature points as numerical data, and is the basic data for generating a 3D model of a face.
[1046] A "3D facial model" refers to a digital model of a face expressed in three-dimensional space based on feature point data.
[1047] "Background photo" refers to photo data that is designated or taken by the user and used as the background of the composite photo.
[1048] "Angle" refers to a parameter that indicates the direction in which the 3D face model is facing relative to the background photo.
[1049] "Pose" refers to parameters that indicate the posture and gesture of a 3D facial model.
[1050] "Facial expression" refers to parameters that indicate emotions or expressions applied to a 3D facial model.
[1051] An "emotion recognition engine" refers to a software module that recognizes emotions from a user's facial image or voice input, and adjusts and generates facial expressions based on those emotions.
[1052] "Synthesis parameters" refer to the settings such as light direction and color tone required to naturally composite a 3D facial model onto a background photo.
[1053] A "composite photo" refers to an image created by combining a background photo with a 3D model of a face.
[1054] "Terminal" refers to a device such as a smartphone or PC on which a user runs an application.
[1055] "Server" refers to a remote server that stores facial images and feature point data and performs various processing.
[1056] MODE FOR CARRYING OUT THE INVENTION
[1057] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on a facial image registered by the user, creating a naturally composite image against a background photograph. It also incorporates an emotion engine that recognizes the user's emotions. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos. Below, we explain each step for implementing this invention.
[1058] 1. The process by which users register their facial images
[1059] User
[1060] The user launches the application on their smartphone or PC and selects the facial image registration menu.
[1061] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[1062] Terminal
[1063] The device uses libraries such as OpenCV and dlib to automatically extract feature points (such as the position of the eyes, nose, and mouth) from each captured facial image.
[1064] The device uses tools like Blender or Three.js to generate a 3D model of the face based on the feature point data.
[1065] The generated 3D model and feature point data are sent to the server.
[1066] server
[1067] The server receives the data sent from the device and stores it in a MySQL or MongoDB database.
[1068] The saved feature point data and 3D model are used in subsequent synthesis processes.
[1069] 2. Angle, pose, and facial expression generation process
[1070] User
[1071] The user selects a background photo within the application or takes a new one.
[1072] The user uses the UI within the application to specify the desired pose and expression.
[1073] server
[1074] The server uses Python's PIL library and OpenCV to analyze the background photo selected by the user and calculate the light direction and color tone.
[1075] The server uses a deep learning model (using TensorFlow or PyTorch) to automatically generate angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[1076] The server calculates the synthesis parameters for synthesizing the background photo analysis data with the 3D model.
[1077] Emotion Engine
[1078] The emotion engine uses Microsoft Azure's Face API and Google Cloud Vision API to recognize the user's emotions from registered facial images.
[1079] The facial expressions are adjusted and generated based on the user's desired pose and facial expression and the recognized emotion.
[1080] The emotion engine uses the Google Cloud Speech-to-Text API to analyze the user's voice input to recognize emotions and generate facial expressions based on them.
[1081] server
[1082] The server uses an automated Adobe Photoshop script and ImageMagick to use the generated data and compositing parameters to naturally composite the 3D face model with the background photo.
[1083] The server transmits the composite image to the terminal.
[1084] 3. Photo generation and storage process
[1085] Terminal
[1086] The terminal receives the composite image from the server, displays it on the screen, and asks the user for confirmation.
[1087] User
[1088] The user checks the composite image and, if satisfied, presses the "Save" button. Editing functions are available as needed.
[1089] Terminal
[1090] The device saves the composite image to local storage using the Android or iOS local storage API.
[1091] The device provides the option to share to social media and cloud services, using the Facebook API and Google Drive API.
[1092] Specific examples
[1093] 1. Example of face registration
[1094] User A starts the app and registers three facial images: front, left, and right.
[1095] The device uses OpenCV to extract feature points from these images and uses Blender to generate a 3D model of the face.
[1096] The generated data is sent to the server, which stores it in a database.
[1097] 2. Specific examples of emotion engines
[1098] User B selects a landscape photo as a background in the app, and the app uses the Microsoft Azure Face API to recognize B's current emotion.
[1099] If the user desires a "smiling" facial expression, but the emotion engine recognizes the emotion of "surprise," it applies the surprised expression to the facial image and generates a 3D model.
[1100] The server calculates the synthesis parameters based on this data and synthesizes it with the landscape photo.
[1101] 3. Examples of group photos
[1102] A family of four wants to take a photo together, but one of them is in a remote location and it is physically impossible to take the photo.
[1103] The three people took a group photo on site, and the facial image of one person in a remote location was registered through the application.
[1104] The user selects a group photo as a background within the app, and the server generates the appropriate angle, pose, and expression based on a single face model located remotely, and combines it with the background photo.
[1105] The emotion engine recognizes the current emotion of a person in a remote location and generates and adjusts facial expressions to match that emotion.
[1106] The composite image is displayed on the smartphone and can be saved by the user.
[1107] As described above, by implementing this invention, it becomes possible to easily create ideal photos even in places where it is difficult to take selfies, and to create high-quality group photos without physically gathering together. Furthermore, the emotion engine can generate natural facial expressions that match the user's emotions, improving image quality and satisfaction.
[1108] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1109] Processing flow
[1110] 1. The process by which users register their facial images
[1111] Step 1:
[1112] The user starts the application and selects the face image registration menu.
[1113] Input: An action on the user interface.
[1114] Output: Camera launch and registration menu display.
[1115] Specific operation: Click the "Register face image" button on the application's home screen, and the following screen will appear.
[1116] Step 2:
[1117] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[1118] Input: Camera operation and photography.
[1119] Output: 3 face images (front, left, right).
[1120] Specific operation: The camera screen will appear and you will be prompted to "Take a photo of the front" and then take a photo, repeating this three times.
[1121] Step 3:
[1122] The device uses libraries such as OpenCV and dlib to automatically extract feature points (such as the position of the eyes, nose, and mouth) from each captured facial image.
[1123] Input: 3 face images.
[1124] Output: Feature point data for each face image.
[1125] Specific operation: Using the dlib facial feature detection model, an algorithm is run to extract 68 facial feature points from an image.
[1126] Step 4:
[1127] The device uses Blender or Three.js to generate a 3D model of the face based on feature point data.
[1128] Input: feature point data.
[1129] Output: 3D face model data.
[1130] Specific operation: Using Blender's Python API, run a script that generates a 3D model based on feature point data.
[1131] Step 5:
[1132] The device sends the generated 3D model and feature point data to the server.
[1133] Input: 3D face model data, feature point data.
[1134] Output: Sending data to the server.
[1135] Specific operation: 3D model data and feature point data are sent to the server using an HTTP POST request.
[1136] Step 6:
[1137] The server receives the data sent from the device and stores it in a MySQL or MongoDB database.
[1138] Input: 3D face model data, feature point data.
[1139] Output: Data stored in the database.
[1140] Specific operation: Analyzes the received data and executes SQL insert statements or MongoDB document inserts to store it in the database.
[1141] 2. Angle, pose, and facial expression generation process
[1142] Step 1:
[1143] The user selects a background photo within the application or takes a new one.
[1144] Input: Select or take a background photo.
[1145] Output: A background photo that you select or take.
[1146] Specific action: Select a photo from the gallery or launch the camera to take a new photo.
[1147] Step 2:
[1148] The user uses the UI within the application to specify the desired pose and expression.
[1149] Input: Specify pose and facial expression.
[1150] Output: Specified pose and facial expression.
[1151] Specific Actions: Use the drop-down menus and sliders to select the desired pose and expression.
[1152] Step 3:
[1153] The server uses Python's PIL library and OpenCV to analyze the background photo selected by the user and calculate the light direction and color tone.
[1154] Input: Background photo data.
[1155] Output: Light direction and color data.
[1156] Specific operation: OpenCV is used to perform image histogram and edge detection, and algorithms are run to identify the direction and color of the light source.
[1157] Step 4:
[1158] The server uses a deep learning model (using TensorFlow or PyTorch) to automatically generate angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[1159] Input: 3D model data, pose and facial expression specification.
[1160] Output: Newly generated angle, pose, and expression data.
[1161] Specific operation: Input data is fed to a pre-trained model using TensorFlow or PyTorch, and the resulting angle and facial expression data is obtained.
[1162] Step 5:
[1163] The emotion engine uses Microsoft Azure's Face API and Google Cloud Vision API to recognize the user's emotions from registered facial images.
[1164] Input: Facial image data.
[1165] Output: Emotion data.
[1166] Specific operation: Send facial image data to the API and obtain emotion recognition results.
[1167] Step 6:
[1168] The server calculates the synthesis parameters for synthesizing the analysis data of the background photo with the 3D model.
[1169] Input: Light direction and color data, 3D model data.
[1170] Output: Synthesis parameters.
[1171] Specific operation: Calculates parameters for naturally combining a 3D model with a background photo, taking into account light direction, color tone, and shadow position.
[1172] Step 7:
[1173] The server uses an automated script in Adobe Photoshop or ImageMagick to use the generated data and synthesis parameters to naturally synthesize the 3D face model with the background photo.
[1174] Input: 3D model data, background photo data, synthesis parameters.
[1175] Output: Composite image.
[1176] Specific operation: Using Adobe Photoshop scripts and ImageMagick, the 3D model is composited with a background photo based on the provided parameters.
[1177] Step 8:
[1178] The server transmits the composite image to the terminal.
[1179] Input: Synthetic image data.
[1180] Output: Send image to device.
[1181] Specific operation: The composite image is sent to the device via an HTTP response or WebSocket message.
[1182] 3. Photo generation and storage process
[1183] Step 1:
[1184] The terminal receives the composite image from the server, displays it on the screen, and asks the user for confirmation.
[1185] Input: Synthetic image data from the server.
[1186] Output: Image displayed to the user.
[1187] Specific operation: The received image data is set to a GUI element for display and displayed to the user.
[1188] Step 2:
[1189] The user checks the composite image and, if satisfied, presses the "Save" button. Editing functions are available as needed.
[1190] Input: User confirmation action.
[1191] Output: Save image or edit data.
[1192] Specific operation: Press the "Save" button to save the image. Press the edit button to display the edit screen.
[1193] Step 3:
[1194] The device saves the composite image to local storage using the Android or iOS local storage API.
[1195] Input: Synthetic image data.
[1196] Output: Image data saved to local storage.
[1197] What it does: It uses Android and iOS file storage APIs to save the composite image to the user's device.
[1198] Step 4:
[1199] The device provides the option to share to social media and cloud services, using the Facebook API and Google Drive API.
[1200] Input: User's share action.
[1201] Output: Upload images to social media or cloud services.
[1202] Specific operation: When you press the share button, the image will be uploaded via the API of the selected social media or cloud service.
[1203] (Application example 2)
[1204] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1205] In today's digital society, there is a growing demand for generating high-quality images with natural expressions and poses in situations where it is difficult for users to take selfies, where it is physically impossible to take group photos, or when virtual try-on simulations are required. However, conventional systems have struggled to generate natural images in these situations, particularly in automatically adjusting facial expressions to reflect the user's emotions. Furthermore, in virtual stores, there is a lack of technology to enhance the user's virtual try-on experience, and there is a need to improve user satisfaction.
[1206] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a means for a user to register their own facial image, a means for extracting feature points from the registered facial image and generating feature point data, and a means for generating a 3D facial model based on the feature point data. This enables high-quality images with natural angles, poses, and facial expressions to be automatically generated and naturally combined with background photos, even in situations where it is difficult for the user to take a selfie. The system also includes a means for recognizing the user's emotions and adjusting the facial expression based on the emotions, thereby enabling more natural image generation. Furthermore, the system also includes a means for simulating virtual trying-on of an item selected by the user in a virtual store, thereby improving the user's virtual trying-on experience.
[1207] A "user" is an entity that registers their own facial image in the system and uses the image generation and virtual try-on simulation.
[1208] "Means for registering face images" is a function that allows users to upload their own face images to the system.
[1209] The "means for extracting feature points" is a function that automatically extracts feature points such as the eyes, nose, and mouth from registered face images.
[1210] "Feature point data" is data including the position coordinates of the eyes, nose, mouth, etc. extracted from a face image.
[1211] The "means for generating a 3D face model" is a function that constructs a three-dimensional model of the user's face based on feature point data.
[1212] "Means for automatically generating angles, poses, and facial expressions" is a function that automatically generates specified angles, poses, and facial expressions for a 3D model.
[1213] A "background photo" is a photo of the background to be combined with the user's facial image.
[1214] "Natural synthesis means" is a function that seamlessly integrates a facial image with the generated angle, pose, and expression into a background photo.
[1215] The "means for presenting to the user" is a function for displaying the synthesized image on the user's screen.
[1216] The "means for saving" is a function that allows the user to save the composite image after checking it.
[1217] The "means for simulating virtual fitting" is a function that combines an object (for example, clothing) selected by the user with a three-dimensional model to generate the results of virtual fitting.
[1218] The "means for recognizing emotions" is a function that automatically determines the user's current emotions from their voice and image.
[1219] The "means for adjusting facial expressions" is a function that appropriately changes the facial expression of a 3D model based on the recognized emotion.
[1220] This system uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by the user, creating images that blend naturally with the background photograph. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to generate more natural facial expressions. This system is particularly intended for virtual try-on and image generation in virtual stores.
[1221] 1. The process by which users register their facial images
[1222] First, the user launches the application on their smartphone or PC and selects the facial image registration menu. For registration, they take photos of their face from the front, left side, and right side. The device extracts feature points from the captured facial images and generates a 3D model of the face based on these. This process uses an image processing library such as OpenCV. The generated 3D model and feature point data are then sent to the server, which then stores the data in a database.
[1223] 2. Virtual try-on simulation process
[1224] Within the application, users select the clothing and accessories they want to try on. The server then combines the selected product information with a 3D model of the user's face to generate a virtual try-on experience. This process uses 3D modeling libraries such as Three.js. Users can also specify their desired pose and angle.
[1225] 3. Emotion Recognition and Facial Expression Generation Process
[1226] The emotion engine analyzes the user's facial images and voice to recognize their current emotions. Based on these emotions, it then appropriately adjusts and generates the facial expressions of the 3D model. This process uses the Microsoft Azure Emotion API and other tools.
[1227] 4. The process of combining with the background photo
[1228] After the user selects a background photo, the server analyzes the photo, calculates the light direction and color tone, calculates the parameters necessary for synthesis, and expresses the facial image as a 3D model, which is then naturally synthesized with the background photo. At this stage, the pose and facial expression specified by the user are reflected.
[1229] 5. Synthesis and saving process
[1230] The server generates a composite image and sends it to the device, which then presents it to the user, who can save it after reviewing it. The option to share it on social media or cloud services is also provided.
[1231] Specific examples
[1232] For example, if a user tries on a new summer dress and generates a composite photo with a shopping mall background, the process would proceed as follows:
[1233] The user registers a face image and selects a new summer dress.
[1234] The selected dress is then composited with a 3D model of a face, a smiling expression is applied, and the dress is composited with the background of a shopping mall.
[1235] The generated composite image is displayed on a smartphone, allowing the user to view and save it.
[1236] Prompt Sentence Examples
[1237] "Generate a composite photo of you trying on a new summer dress and smiling with a shopping mall background"
[1238] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1239] Step 1: User registers a face image
[1240] The user launches the application on their smartphone or PC and selects the facial image registration menu. Next, they take photos of their face from the front, left side, and right side and upload them to the application. These input images serve as the base data for subsequent processing. The device receives these facial images and uses OpenCV to extract feature points such as the eyes, nose, and mouth. Feature point data is generated, and a 3D model of the face is created based on this. This 3D model is then used in subsequent synthesis processing.
[1241] Step 2: User selects an object for virtual try-on
[1242] Within the application, the user selects the clothing and accessories they want to try on. The data for this selection is provided as input to the system. The server receives this and prepares to combine a 3D model of the user's face with the information about the selected items. This process uses a 3D modeling library such as Three.js, and the output is a virtual try-on ready image.
[1243] Step 3: The emotion engine analyzes facial images and voice
[1244] The user's facial image and voice are provided as input to the emotion engine. The server analyzes this data using Microsoft Azure Emotion API or similar to recognize the user's current emotion. Emotion data is generated as a result of the analysis. This emotion data is then used to adjust facial expressions.
[1245] Step 4: Generate and adjust facial expressions based on emotion data
[1246] The server adjusts the facial expression of the user's 3D model based on the emotion data obtained in step 3. For example, if the user is excited, it applies a smiling expression. The individual feature point data of the registered facial image is used to generate the facial expression of the 3D model. The output is the adjusted 3D model.
[1247] Step 5: Select a background photo and analyze it
[1248] The user selects a background photo, which is sent as input to the server, which analyzes the background photo to calculate the light direction and color tone. Image processing algorithms are used for the analysis, and light and color parameters for compositing are generated as output.
[1249] Step 6: Generate the composite image
[1250] The server uses the 3D model of the virtual try-on garment generated in step 2, the 3D model of the user adjusted in step 4, and the light and color parameters obtained in step 5 to naturally combine the background photo and facial image. The synthesis algorithm outputs a natural-looking composite image.
[1251] Step 7: Send the composite image to your device
[1252] The server sends the composite image to the terminal, which then presents the received composite image to the user, who then checks the presented composite image on the screen.
[1253] Step 8: Save and share your composite image
[1254] The user can review the composite image and save it if they are satisfied. The device then saves the composite image in local storage and provides the option to share it on social media or cloud services if desired. Finally, the saved composite image can be used according to the user's purpose.
[1255] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1256] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1257] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1258] [Third embodiment]
[1259] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1260] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1261] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1262] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1263] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1264] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1265] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1266] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1267] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1268] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1269] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1270] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1271] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and creates images that are naturally composited against a background photograph. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos.
[1272] 1. The process by which users register their facial images
[1273] User
[1274] First, the user launches the application on their smartphone or PC and selects the facial image registration menu.
[1275] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[1276] Terminal
[1277] The device automatically extracts feature points (e.g., the positions of the eyes, nose, and mouth) from each captured facial image.
[1278] A 3D model of the face is generated based on the feature point data.
[1279] The generated 3D model and feature point data are sent to the server.
[1280] server
[1281] The server receives the data sent from the terminal and stores it in a database.
[1282] The saved feature point data and 3D model are used in subsequent synthesis processes.
[1283] 2. Angle, pose, and facial expression generation process
[1284] User
[1285] Users can select a background photo within the application or take a new one.
[1286] The user specifies the desired pose and facial expression.
[1287] server
[1288] The server analyzes the background photo selected by the user and calculates the direction and color of the light.
[1289] Based on the pose and facial expression specified by the user, the angle, pose, and facial expression are automatically generated for the registered 3D model.
[1290] Calculate synthesis parameters for synthesizing the background photo analysis data and the 3D model.
[1291] The facial image is represented as a 3D model based on the generated angle, pose, and expression, and is then naturally combined with the background photo.
[1292] Terminal
[1293] The terminal receives the composite image and the composite parameters from the server and performs the composite process.
[1294] The synthesized image is presented to the user for confirmation.
[1295] 3. Photo generation and storage process
[1296] User
[1297] The user checks the generated composite image and presses the save button if satisfied.
[1298] Use further editing functions as needed.
[1299] Terminal
[1300] The terminal stores the composite image in local storage.
[1301] It also provides users with the option to share images on social media or cloud services, depending on their choice.
[1302] Specific examples
[1303] 1. Example of face registration
[1304] User A starts the app and registers three facial images: front, left, and right.
[1305] The device extracts feature points from these images and creates a 3D model of the face.
[1306] The generated data is sent to the server, which stores it in a database.
[1307] 2. Examples of group photos
[1308] A family of four wants to take a photo together, but one of them is in a remote location and it is physically impossible to take the photo.
[1309] The three people took a group photo on site, and the facial image of one person in a remote location was registered through the application.
[1310] Users select a group photo as a background within the app, and the server generates the appropriate angle, pose, and expression based on a remote facial model of a single person, and combines it with the background photo.
[1311] The composite image is displayed on the smartphone and can be saved by the user.
[1312] As described above, by implementing this invention, it becomes possible to easily generate ideal photos even in places where it is difficult to take selfies, and it is also possible to create high-quality group photos without having to physically gather together.
[1313] The processing flow will be explained below.
[1314] Program processing details
[1315] The process by which users register their facial images
[1316] Step 1:
[1317] The user launches the application on their smartphone or PC.
[1318] Step 2:
[1319] The user selects a face image registration menu and follows the instructions to take photos of the face from the front, left side, and right side.
[1320] Step 3:
[1321] The terminal receives each captured face image.
[1322] Step 4:
[1323] The device extracts feature points (such as the positions of the eyes, nose, and mouth) from the facial image.
[1324] Step 5:
[1325] The device generates a 3D model of the user's face based on the extracted feature point data.
[1326] Step 6:
[1327] The device sends the generated 3D model and feature point data to the server.
[1328] Step 7:
[1329] The server stores the received data in a database.
[1330] The process of generating angles, poses, and facial expressions
[1331] Step 1:
[1332] The user selects the photo creation menu of the application.
[1333] Step 2:
[1334] The user selects a background photo or takes a new one.
[1335] Step 3:
[1336] The user specifies the desired pose and facial expression.
[1337] Step 4:
[1338] The server receives the background photo and runs an image analysis algorithm to analyze the light direction and color tone.
[1339] Step 5:
[1340] Based on the analysis results, the server calculates the synthesis parameters for combining the image with the background photo.
[1341] Step 6:
[1342] The server automatically generates angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[1343] Step 7:
[1344] The server uses the generated data and synthesis parameters to represent the facial image as a 3D model and naturally combines it with the background photo.
[1345] Step 8:
[1346] The server transmits the composite image to the terminal.
[1347] Photo generation and storage process
[1348] Step 1:
[1349] The terminal receives the composite image sent from the server.
[1350] Step 2:
[1351] The terminal displays the received composite image on the screen and asks the user for confirmation.
[1352] Step 3:
[1353] The user checks the composite image and, if satisfied, presses the save button.
[1354] Step 4:
[1355] The device stores the composite image in local storage.
[1356] Step 5:
[1357] The device displays a notification to the user that the save is complete.
[1358] Step 6:
[1359] If the user wishes to share the image on a social networking site or cloud service, the device will provide the option to do so and process it accordingly.
[1360] Example 1
[1361] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1362] In conventional facial image synthesis systems, it was difficult for users to naturally combine their own facial images with background images, and differences in light direction and color tone often resulted in unnatural results. It was also difficult to naturally reflect the user's desired pose and facial expression. Furthermore, the process of efficiently extracting feature points from multiple facial images and generating a 3D model was complicated, resulting in poor usability for the entire system.
[1363] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1364] In this invention, the server includes means for a user to register his or her own facial image, means for extracting feature points from the registered facial image and generating feature point data, means for generating a 3D facial model based on the feature point data, means for analyzing a background image and identifying the light direction and color tone, means for automatically generating an angle, pose, and facial expression based on the analysis results, means for representing the facial image as a 3D model according to the generated angle, pose, and facial expression and naturally combining it with the background image, and means for presenting the combined image to the user and saving it after confirmation. This allows the user to easily generate natural combined images and improves operability, enabling more satisfying facial image synthesis.
[1365] "User" refers to an individual or group that uses this system to register their own facial images and create composite images.
[1366] "Means for registering" refers to a method or device that allows users to upload and store their facial images in the system.
[1367] "Feature points" are data that indicate specific parts of a face image, such as the eyes, nose, and mouth.
[1368] "Feature point data" is a data set that includes information on feature points extracted from a face image.
[1369] A "3D facial model" is a three-dimensional shape of a face reconstructed in three dimensions based on feature point data of a facial image.
[1370] A "background image" is a photograph or image selected by the user as a composite object.
[1371] The "analyzing means" refers to a method or device for calculating and analyzing the light direction and color tone of the background image.
[1372] "Means for automatically generating angles, poses, and facial expressions" refers to a method or device that has the function of automatically creating facial angles, poses, and facial expressions based on user specifications and analysis results.
[1373] The "synthesis means" refers to a method or device for naturally synthesizing the generated 3D face model with a background image.
[1374] The "presenting means" refers to a method or device for displaying the synthesized image to the user and requesting confirmation.
[1375] "Storing means" refers to a method or device for storing the confirmed composite image within the system or in external storage.
[1376] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and creates images that are naturally composited against a background photograph. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos.
[1377] The process by which users register their facial images
[1378] User
[1379] The user launches the application on their smartphone or PC and selects the facial image registration menu.
[1380] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[1381] Terminal
[1382] The device automatically extracts feature points (e.g., the positions of the eyes, nose, and mouth) from each captured facial image using Python's OpenCV library.
[1383] Based on the extracted feature point data, a 3D model of the face is generated using 3D modeling software such as Blender.
[1384] The generated 3D model and feature point data are sent to the server via REST API.
[1385] server
[1386] The server receives the data sent from the device and stores it in a MySQL database using the Django framework.
[1387] The saved feature point data and 3D model are used in subsequent synthesis processes.
[1388] The process of generating angles, poses, and facial expressions
[1389] User
[1390] Users can select a background photo within the application or take a new one.
[1391] The user specifies the desired pose and facial expression.
[1392] server
[1393] The server analyzes the background photo selected by the user and calculates the light direction and color tone using PIL (Python Imaging Library) and OpenCV.
[1394] Based on the pose and facial expression specified by the user, the angle, pose, and facial expression are automatically generated for the registered 3D model. Here, generative AI models (such as StyleGAN and DeepFaceLab) are utilized.
[1395] Calculate synthesis parameters for synthesizing the background photo analysis data and the 3D model.
[1396] The facial image is represented based on a 3D model according to the generated angle, pose, and expression, and is then naturally combined with the background photo.
[1397] The process by which the device generates a composite image and presents it to the user
[1398] Terminal
[1399] The terminal receives the composite image and the composite parameters from the server.
[1400] The synthesis process is performed and the completed image is presented to the user.
[1401] User
[1402] The user checks the presented composite image and presses the save button if satisfied.
[1403] Further modifications can be made within the app if needed.
[1404] The generated images can also be shared on social media or cloud services.
[1405] Specific examples
[1406] Specific examples of face registration
[1407] The user launches the app and takes a series of photos of their face - front, left, and right - and registers them. The device uses Python's OpenCV library to extract feature points from these images and creates a 3D face model using Blender. The generated data is sent to the server via a REST API and stored in a MySQL database using the Django framework.
[1408] Examples of group photos
[1409] If a family wants to take a group photo together, but one person is remote and cannot physically be present, the remote person can register their face image through the app. The user then takes a group photo on-site and selects a background photo within the app. The server then analyzes the image, calculating the light direction and color tone, and generates the appropriate angle, pose, and facial expression based on a 3D facial model of the remote person, which is then combined with the background photo. The resulting composite image is then displayed on the user's smartphone.
[1410] Prompt Sentence Examples
[1411] "We will create a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and then creates an image that is composited into a background photo. We will also analyze the background photo, such as the direction of light and color tone, to ensure that the composite image looks natural."
[1412] By implementing this invention, ideal selfies can be easily generated even in places where it is difficult to take selfies, and high-quality group photos can be created without physically gathering together.
[1413] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1414] The flow of this system's program processing
[1415] Step 1:
[1416] explanation
[1417] Users launch the application on their smartphone or PC, select the face image registration menu, and follow the app's guide to take and register photos of their face from the front, left side, and right side.
[1418] Input and Output
[1419] Input: Face image (front, left, right)
[1420] Output: Three registered face images
[1421] Specific actions
[1422] The user clicks the "Register Face Image" button on the app, activates the camera to take face images from various angles, and then presses the "Confirm / Register" button.
[1423] Step 2:
[1424] explanation
[1425] The device automatically extracts feature points (such as the positions of the eyes, nose, and mouth) from the captured facial image using Python's OpenCV library. Based on the extracted feature point data, a 3D model of the face is generated using 3D modeling software such as Blender.
[1426] Input and Output
[1427] Input: Three registered face images
[1428] Output: feature point data, 3D face model
[1429] Specific actions
[1430] The device uses the OpenCV library to detect feature points from each facial image, then uses Blender to generate a 3D model of the face based on these feature points. The generated model and feature point data are then sent to the server via a REST API.
[1431] Step 3:
[1432] explanation
[1433] The server receives the data sent from the device and stores it in a MySQL database using the Django framework. The stored feature point data and 3D model are then used in the synthesis process.
[1434] Input and Output
[1435] Input: feature point data, 3D face model
[1436] Output: Data stored in the database
[1437] Specific actions
[1438] The server receives the transmitted data via the Django framework and stores it in a MySQL database.
[1439] Step 4:
[1440] explanation
[1441] Users can select a background photo within the application or take a new one, and specify the desired pose and facial expression.
[1442] Input and Output
[1443] Input: Background photo, desired pose and facial expression
[1444] Output: User selections and specifications
[1445] Specific actions
[1446] Once the user selects (or takes) a background photo, a UI for specifying the pose and expression is displayed.
[1447] Step 5:
[1448] explanation
[1449] The server analyzes the background photo selected by the user and calculates the light direction and color tone using PIL (Python Imaging Library) and OpenCV.
[1450] Input and Output
[1451] Input: Background photo
[1452] Output: Analysis results of light direction and color tone
[1453] Specific actions
[1454] The server loads the background photo using the PIL or OpenCV library and analyzes the light direction and color tone.
[1455] Step 6:
[1456] explanation
[1457] The server uses a generative AI model (e.g., StyleGAN or DeepFaceLab) to automatically generate angles, poses, and expressions for the registered 3D model based on the poses and expressions specified by the user.
[1458] Input and Output
[1459] Input: feature point data, 3D face model, desired pose and expression
[1460] Output: Generated angles, poses, and expressions
[1461] Specific actions
[1462] The server uses a generative AI model to generate poses and facial expressions based on feature point data and a 3D model.
[1463] Step 7:
[1464] explanation
[1465] The server calculates synthesis parameters for combining the analysis data of the background photo with a 3D model, and then represents the facial image as a 3D model according to the generated angle, pose, and expression, which is then naturally combined with the background photo.
[1466] Input and Output
[1467] Input: background photo, analysis data, generated angles, poses, facial expressions
[1468] Output: Synthesized face image
[1469] Specific actions
[1470] The server calculates the synthesis parameters and uses OpenCV to naturally synthesize the face image and background photo.
[1471] Step 8:
[1472] explanation
[1473] The terminal receives the composite image and the composite parameters from the server, performs the composite process, and presents the completed composite image to the user.
[1474] Input and Output
[1475] Input: synthetic image, synthetic parameters
[1476] Output: The composite image presented to the user
[1477] Specific actions
[1478] The device retrieves the composite image from the server via an HTTP request and displays it to the user on the app.
[1479] Step 9:
[1480] explanation
[1481] The user can review the composite image presented to them and, if satisfied, press the save button. If necessary, they can make further edits within the app. The generated image can also be shared via social media or cloud services.
[1482] Input and Output
[1483] Input: Composite image
[1484] Output: Saved and shared images
[1485] Specific actions
[1486] When the user presses the "Save" button, the image is saved to the smartphone's local storage, and when the user presses the "Share" button, the image is uploaded to social media or a cloud service.
[1487] (Application example 1)
[1488] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1489] Modern electronic payment services require enhanced security. However, many current systems have limited user authentication methods, making it difficult to simultaneously improve convenience and security. Furthermore, many systems can only authenticate users at specific angles or poses, which can be inconvenient for users. There is a need to solve these issues and realize more accurate and convenient facial recognition. Furthermore, a system is needed that can accurately authenticate users even when they have different facial expressions or poses.
[1490] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1491] In this invention, the server includes: a means for a user to register his or her own facial image; a means for extracting feature points from the registered facial image and generating feature point data; a means for generating a 3D facial model based on the feature point data; a means for automatically generating an angle, pose, and facial expression to be combined with a background photograph; a means for representing the facial image as a 3D model according to the generated angle, pose, and facial expression and combining it naturally with the background photograph; a means for performing electronic authentication using analysis data of the background photograph and the user's 3D model; and a means for presenting the combined image to the user and saving it after confirmation. This enables advanced electronic authentication by generating a natural combined image based on the user's registered facial image and background photograph. A means for the user to specify a desired pose and facial expression; a means for generating corresponding movements for the facial portion of the 3D model based on the specified pose and facial expression; a means for performing electronic authentication based on the generated 3D model and a captured image of the user; and a means for reconstructing the facial image according to the generated movements, thereby providing an authentication system that combines convenience and security.
[1492] "Means for users to register their own facial images" refers to a function that allows users to register their own facial images using devices such as smartphones or computers.
[1493] "Means for extracting feature points and generating feature point data" refers to a function that uses AI technology to detect feature points such as the eyes, nose, and mouth from registered facial images and converts them into data.
[1494] The "means for generating a 3D model of a face" is a function for generating a 3D model of a face using feature point data.
[1495] "Means for automatically generating angles, poses, and facial expressions for compositing with background photos" is a function that uses AI to automatically generate the optimal facial angle, pose, and facial expression for compositing with a background photo.
[1496] "Means for representing a facial image as a 3D model according to the generated angle, pose, and expression, and for naturally combining it with a background photograph" refers to a function for reproducing a facial image as a 3D model based on the generated angle, pose, and expression, and for naturally combining it with a background photograph.
[1497] "Means for electronic authentication using background photo analysis data and a 3D model of the user" is a function for securely and accurately authenticating a user using analysis data of the light direction and color tone of the background photo and a 3D model of the user.
[1498] The "means for presenting the synthesized image to the user and saving it after confirmation" is a function for displaying the generated synthesized image to the user and saving the image after obtaining confirmation.
[1499] "Means for generating corresponding movements for the facial part of a 3D model based on a specified pose or expression" is a function that moves the facial part of a 3D model according to the pose or expression specified by the user.
[1500] The "means for reconstructing a facial image according to the generated movements" is a function for reconstructing a facial image based on the generated poses and facial movements.
[1501] "Means for electronic authentication" refers to an authentication function that uses a user's facial image or 3D model to allow access to the system.
[1502] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and creates images that are naturally composited with background photographs. Below, we will provide a detailed explanation of how this system is realized.
[1503] 1. Facial image registration process
[1504] Users register their facial images using a smartphone or computer. Specifically, they take photos of their face from the front, left side, and right side, and register them through the application.
[1505] The device receives a facial image and extracts its features using a library such as dlib. This generates key feature data such as the eyes, nose, and mouth. A 3D model of the face is then generated based on this feature data. This 3D model and feature data are then sent to a server and stored in a database.
[1506] 2. Angle, pose, and facial expression generation process
[1507] The server analyzes the background photo selected or taken by the user. Specifically, it calculates the direction of light and color tone. Based on this analysis data, it automatically generates the appropriate angle, pose, and expression for the registered 3D model according to the pose and expression specified by the user. Using AI technology such as DeepFace, it calculates synthesis parameters to naturally combine the user's 3D face model with the background photo.
[1508] 3. Electronic Authentication Process
[1509] The device performs electronic authentication using background photo analysis data and a 3D model of the user. Specifically, it uses DeepFace's Facenet model to match the user's captured image with the 3D model for accurate authentication.
[1510] 4. Photo Generation and Storage Process
[1511] Users can review the generated composite image and save it if they are satisfied. Further editing is available if necessary. The composite image is saved to local storage, and options for sharing it to social media and cloud services are also provided.
[1512] The system that realizes this invention executes a series of processes, from registering a facial image, generating a 3D model, using that model for electronic authentication, and finally generating and saving a composite image. This enables users to be accurately and securely authenticated even with different poses and expressions, and to obtain a high-quality composite image.
[1513] Examples:
[1514] A user installs the FacePay app and registers images of their face, showing the front, left, and right sides. After making a purchase at a convenience store, they perform payment using facial recognition, and the payment is completed upon successful authentication.
[1515] Example prompt sentence:
[1516] "Design an application that generates a 3D face model from a facial image registered by a user and uses that model for facial authentication during electronic payments."
[1517] As described above, by implementing the present invention, it is possible to generate ideal photos even when selfies are difficult, and to create high-quality group photos without physically gathering together. Furthermore, by using this system, the security of electronic payments is improved, and user convenience is also enhanced.
[1518] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1519] Step 1:
[1520] The user registers a facial image. In this process, the user launches an application on a smartphone or PC and takes and registers facial images from the front, left side, and right side.
[1521] Input: Face image taken by the user (front, left, right)
[1522] Output: Facial feature point data and a 3D model of the face generated by the device
[1523] Step 2:
[1524] The device processes the facial image provided by the user. Specifically, it uses the dlib library to extract feature points such as the eyes, nose, and mouth from the captured facial image. Based on this feature point data, a 3D model of the face is generated. The generated data is sent to a server and stored in a database.
[1525] Input: A face image obtained from the user
[1526] Output: Feature point data and a 3D face model
[1527] Step 3:
[1528] When a user selects or takes a background photo, the server analyzes the background photo, detecting the direction and color of the light, and then uses this information to calculate the direction and color of the light in the background photo.
[1529] Input: A background photo selected or taken by the user
[1530] Output: Analysis data of background photo (light direction and color tone)
[1531] Step 4:
[1532] The server automatically generates the angle, pose, and facial expression for the registered 3D model according to the pose and facial expression specified by the user. Based on this information, the server expresses the facial image as a 3D model and calculates synthesis parameters for naturally combining it with the background photo.
[1533] Input: Analysis data of user-specified poses and expressions, and background photos
[1534] Output: angle, pose, facial expression, synthesis parameters
[1535] Step 5:
[1536] The device uses the synthesis parameters and 3D model received from the server to naturally synthesize the background photo and facial image. Specifically, the synthesis process is performed using DeepFace's AI technology.
[1537] Input: Synthesis parameters and 3D model provided by the server
[1538] Output: Composite image
[1539] Step 6:
[1540] The device will then present the combined image to the user for review, and if the user is satisfied, the image will be saved to local storage, with the option to share it via social media or cloud services if desired.
[1541] Input: The composite image
[1542] Output: User confirmation and save
[1543] Step 7:
[1544] The device performs electronic authentication using background photo analysis data and a 3D model of the user, and uses DeepFace's Facenet model to match the captured image with the 3D model.
[1545] Input: Captured image, background photo analysis data, user 3D model
[1546] Output: Authentication result (success or failure)
[1547] Through a series of processes from registration to synthesis and authentication, highly accurate and convenient electronic authentication is achieved.
[1548] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1549] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by the user, creating images that blend naturally with the background photo, and further combines it with an emotion engine that recognizes the user's emotions. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos.
[1550] 1. The process by which users register their facial images
[1551] User
[1552] First, the user launches the application on their smartphone or PC and selects the facial image registration menu.
[1553] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[1554] Terminal
[1555] The device automatically extracts feature points (such as the position of the eyes, nose, and mouth) from each captured facial image.
[1556] A 3D model of the face is generated based on the feature point data.
[1557] The generated 3D model and feature point data are sent to the server.
[1558] server
[1559] The server receives the data sent from the terminal and stores it in a database.
[1560] The saved feature point data and 3D model are used in subsequent synthesis processes.
[1561] 2. Angle, pose, and facial expression generation process
[1562] User
[1563] Users can select a background photo within the application or take a new one.
[1564] The user specifies the desired pose and facial expression.
[1565] server
[1566] The server analyzes the background photo selected by the user and calculates the direction and color of the light.
[1567] Based on the pose and facial expression specified by the user, the angle, pose, and facial expression are automatically generated for the registered 3D model.
[1568] Calculate synthesis parameters for synthesizing the background photo analysis data and the 3D model.
[1569] Emotion Engine
[1570] The emotion engine recognizes the user's emotion from the registered facial image.
[1571] The facial expression is adjusted and generated based on the user's desired pose, facial expression, and recognized emotion.
[1572] The emotion engine can also analyze the user's voice input, recognize the user's emotion based on the analysis results, and generate facial expressions based on that.
[1573] server
[1574] The server uses the generated data and synthesis parameters to represent the facial image as a 3D model and naturally combines it with the background photo.
[1575] The server transmits the composite image to the terminal.
[1576] 3. Photo generation and storage process
[1577] Terminal
[1578] The terminal receives the composite image from the server.
[1579] The terminal displays the composite image on the screen and asks the user for confirmation.
[1580] User
[1581] The user checks the composite image and, if satisfied, presses the save button.
[1582] Use further editing functions as needed.
[1583] Terminal
[1584] The device stores the composite image in local storage.
[1585] It also provides the option to share to social media and cloud services.
[1586] Specific examples
[1587] 1. Example of face registration
[1588] User A starts the app and registers three facial images: front, left, and right.
[1589] The device extracts feature points from these images and generates a 3D model of the face.
[1590] The generated data is sent to the server, which stores it in a database.
[1591] 2. Specific examples of emotion engines
[1592] User B selects a landscape photo as a background in the app, and the app recognizes B's current emotion.
[1593] If the user desires a "smiling" facial expression, but the emotion engine recognizes the emotion of "surprise," it applies the surprised expression to the facial image and generates a 3D model.
[1594] The server calculates the synthesis parameters based on this data and synthesizes it with the landscape photo.
[1595] 3. Examples of group photos
[1596] A family of four wants to take a photo together, but one of them is in a remote location and it is physically impossible to take the photo.
[1597] The three people took a group photo on site, and the facial image of one person in a remote location was registered through the application.
[1598] Users select a group photo as a background within the app, and the server generates the appropriate angle, pose, and facial expression based on a remote facial model of a single person, and combines it with the background photo.
[1599] The emotion engine recognizes the current emotion of a person in a remote location and generates and adjusts facial expressions to match that emotion.
[1600] The composite image is displayed on the smartphone and can be saved by the user.
[1601] As described above, by implementing this invention, it becomes possible to easily create ideal photos even in places where it is difficult to take selfies, and it is also possible to create high-quality group photos without physically gathering together.Furthermore, the emotion engine can generate natural facial expressions that better match the user's emotions, improving image quality and satisfaction.
[1602] The processing flow will be explained below.
[1603] Program processing details
[1604] The process by which users register their facial images
[1605] Step 1:
[1606] The user launches the application on their smartphone or PC.
[1607] Step 2:
[1608] The user selects a face image registration menu and follows the instructions to take photos of the face from the front, left side, and right side.
[1609] Step 3:
[1610] The terminal receives each captured face image.
[1611] Step 4:
[1612] The device extracts feature points (such as the positions of the eyes, nose, and mouth) from the facial image.
[1613] Step 5:
[1614] The device generates a 3D model of the user's face based on the extracted feature point data.
[1615] Step 6:
[1616] The device sends the generated 3D model and feature point data to the server.
[1617] Step 7:
[1618] The server stores the received data in a database.
[1619] The process of generating angles, poses, and facial expressions
[1620] Step 1:
[1621] The user selects the photo creation menu of the application.
[1622] Step 2:
[1623] The user selects a background photo or takes a new one.
[1624] Step 3:
[1625] The user specifies the desired pose and facial expression.
[1626] Step 4:
[1627] The server receives the background photo and runs an image analysis algorithm to analyze the light direction and color tone.
[1628] Step 5:
[1629] Based on the analysis results, the server calculates the synthesis parameters for combining the image with the background photo.
[1630] Step 6:
[1631] The server automatically generates angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[1632] Step 7:
[1633] The emotion engine recognizes the user's emotions from registered facial images or voice input.
[1634] Step 8:
[1635] The emotion engine adjusts and generates poses and expressions based on the recognized emotions.
[1636] Step 9:
[1637] The server uses the generated data and synthesis parameters to represent the facial image as a 3D model and naturally combines it with the background photo.
[1638] Step 10:
[1639] The server transmits the composite image to the terminal.
[1640] Photo generation and storage process
[1641] Step 1:
[1642] The terminal receives the composite image sent from the server.
[1643] Step 2:
[1644] The terminal displays the received composite image on the screen and asks the user for confirmation.
[1645] Step 3:
[1646] The user checks the composite image and, if satisfied, presses the save button.
[1647] Step 4:
[1648] The device stores the composite image in local storage.
[1649] Step 5:
[1650] The device displays a notification to the user that the save is complete.
[1651] Step 6:
[1652] If the user wishes to share the image on a social networking site or cloud service, the device will provide the option to do so and process it accordingly.
[1653] Example 2
[1654] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1655] Conventional photo compositing technologies have difficulty in generating natural poses and facial expressions desired by users, and require manual setting of compositing parameters such as light direction and color tone. This has required a great deal of time and effort to generate a photo that satisfies the user. Furthermore, when taking a group photo and not everyone is present, it is difficult to composite people in remote locations naturally. Furthermore, there has been a lack of mechanisms for appropriately reflecting facial expressions that match the user's emotions. The objective of this invention is to solve these problems and realize easier and more natural photo compositing.
[1656] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1657] In this invention, the server includes: means for a user to register his or her own facial image; means for extracting feature points from the registered facial image and generating feature point data; means for generating a 3D facial model based on the feature point data; means for automatically generating an angle, pose, and facial expression to be composited with a background photograph; means for representing the facial image as a 3D model according to the generated angle, pose, and facial expression and naturally composited with the background photograph; means including an emotion recognition engine that recognizes the user's emotion and adjusts and generates the facial expression of the facial image based on the emotion; means for presenting the composite image to the user and saving it after confirmation; and means for saving facial image data and feature point data received from the terminal. This enables the user to easily specify a desired pose and facial expression and automatically generate a natural composite photograph, and further enables the realization of a photograph with a realistic expression that reflects the user's emotion.
[1658] "User" refers to an individual who uses this system to register a facial image and perform various operations to generate a composite photograph.
[1659] "Facial image" refers to image data of a user's face, and includes photographs of the front, left side, and right side.
[1660] "Feature points" refer to position data of the eyes, nose, mouth, etc. extracted from a facial image, and are used to represent the structure and shape of the face.
[1661] "Feature point data" is a representation of feature points as numerical data, and is the basic data for generating a 3D model of a face.
[1662] A "3D facial model" refers to a digital model of a face expressed in three-dimensional space based on feature point data.
[1663] "Background photo" refers to photo data that is designated or taken by the user and used as the background of the composite photo.
[1664] "Angle" refers to a parameter that indicates the direction in which the 3D face model is facing relative to the background photo.
[1665] "Pose" refers to parameters that indicate the posture and gesture of a 3D facial model.
[1666] "Facial expression" refers to parameters that indicate emotions or expressions applied to a 3D facial model.
[1667] An "emotion recognition engine" refers to a software module that recognizes emotions from a user's facial image or voice input, and adjusts and generates facial expressions based on those emotions.
[1668] "Synthesis parameters" refer to the settings such as light direction and color tone required to naturally composite a 3D facial model onto a background photo.
[1669] A "composite photo" refers to an image created by combining a background photo with a 3D model of a face.
[1670] "Terminal" refers to a device such as a smartphone or PC on which a user runs an application.
[1671] "Server" refers to a remote server that stores facial images and feature point data and performs various processing.
[1672] MODE FOR CARRYING OUT THE INVENTION
[1673] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on a facial image registered by the user, creating a naturally composite image against a background photograph. It also incorporates an emotion engine that recognizes the user's emotions. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos. Below, we explain each step for implementing this invention.
[1674] 1. The process by which users register their facial images
[1675] User
[1676] The user launches the application on their smartphone or PC and selects the facial image registration menu.
[1677] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[1678] Terminal
[1679] The device uses libraries such as OpenCV and dlib to automatically extract feature points (such as the position of the eyes, nose, and mouth) from each captured facial image.
[1680] The device uses tools like Blender or Three.js to generate a 3D model of the face based on the feature point data.
[1681] The generated 3D model and feature point data are sent to the server.
[1682] server
[1683] The server receives the data sent from the device and stores it in a MySQL or MongoDB database.
[1684] The saved feature point data and 3D model are used in subsequent synthesis processes.
[1685] 2. Angle, pose, and facial expression generation process
[1686] User
[1687] The user selects a background photo within the application or takes a new one.
[1688] The user uses the UI within the application to specify the desired pose and expression.
[1689] server
[1690] The server uses Python's PIL library and OpenCV to analyze the background photo selected by the user and calculate the light direction and color tone.
[1691] The server uses a deep learning model (using TensorFlow or PyTorch) to automatically generate angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[1692] The server calculates the synthesis parameters for synthesizing the background photo analysis data with the 3D model.
[1693] Emotion Engine
[1694] The emotion engine uses Microsoft Azure's Face API and Google Cloud Vision API to recognize the user's emotions from registered facial images.
[1695] The facial expressions are adjusted and generated based on the user's desired pose and facial expression and the recognized emotion.
[1696] The emotion engine uses the Google Cloud Speech-to-Text API to analyze the user's voice input to recognize emotions and generate facial expressions based on them.
[1697] server
[1698] The server uses an automated script in Adobe Photoshop or ImageMagick to use the generated data and synthesis parameters to naturally synthesize the 3D face model with the background photo.
[1699] The server transmits the composite image to the terminal.
[1700] 3. Photo generation and storage process
[1701] Terminal
[1702] The terminal receives the composite image from the server, displays it on the screen, and asks the user for confirmation.
[1703] User
[1704] The user checks the composite image and, if satisfied, presses the "Save" button. Editing functions are available as needed.
[1705] Terminal
[1706] The device saves the composite image to local storage using the Android or iOS local storage API.
[1707] The device provides the option to share to social media and cloud services, using the Facebook API and Google Drive API.
[1708] Specific examples
[1709] 1. Example of face registration
[1710] User A starts the app and registers three facial images: front, left, and right.
[1711] The device uses OpenCV to extract feature points from these images and uses Blender to generate a 3D model of the face.
[1712] The generated data is sent to the server, which stores it in a database.
[1713] 2. Specific examples of emotion engines
[1714] User B selects a landscape photo as a background in the app, and the app uses the Microsoft Azure Face API to recognize B's current emotion.
[1715] If the user desires a "smiling" facial expression, but the emotion engine recognizes the emotion of "surprise," it applies the surprised expression to the facial image and generates a 3D model.
[1716] The server calculates the synthesis parameters based on this data and synthesizes it with the landscape photo.
[1717] 3. Examples of group photos
[1718] A family of four wants to take a photo together, but one of them is in a remote location and it is physically impossible to take the photo.
[1719] The three people took a group photo on site, and the facial image of one person in a remote location was registered through the application.
[1720] The user selects a group photo as a background within the app, and the server generates the appropriate angle, pose, and expression based on a single face model located remotely, and combines it with the background photo.
[1721] The emotion engine recognizes the current emotion of a person in a remote location and generates and adjusts facial expressions to match that emotion.
[1722] The composite image is displayed on the smartphone and can be saved by the user.
[1723] As described above, by implementing this invention, it becomes possible to easily create ideal photos even in places where it is difficult to take selfies, and to create high-quality group photos without physically gathering together. Furthermore, the emotion engine can generate natural facial expressions that match the user's emotions, improving image quality and satisfaction.
[1724] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1725] Processing flow
[1726] 1. The process by which users register their facial images
[1727] Step 1:
[1728] The user starts the application and selects the face image registration menu.
[1729] Input: An action on the user interface.
[1730] Output: Camera launch and registration menu display.
[1731] Specific operation: Click the "Register face image" button on the application's home screen, and the following screen will appear.
[1732] Step 2:
[1733] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[1734] Input: Camera operation and photography.
[1735] Output: 3 face images (front, left, right).
[1736] Specific operation: The camera screen will appear and you will be prompted to "Take a photo of the front" and then take a photo, repeating this three times.
[1737] Step 3:
[1738] The device uses libraries such as OpenCV and dlib to automatically extract feature points (such as the position of the eyes, nose, and mouth) from each captured facial image.
[1739] Input: 3 face images.
[1740] Output: Feature point data for each face image.
[1741] Specific operation: Using the dlib facial feature detection model, an algorithm is run to extract 68 facial feature points from an image.
[1742] Step 4:
[1743] The device uses Blender or Three.js to generate a 3D model of the face based on feature point data.
[1744] Input: feature point data.
[1745] Output: 3D face model data.
[1746] Specific operation: Using Blender's Python API, run a script that generates a 3D model based on feature point data.
[1747] Step 5:
[1748] The device sends the generated 3D model and feature point data to the server.
[1749] Input: 3D face model data, feature point data.
[1750] Output: Sending data to the server.
[1751] Specific operation: 3D model data and feature point data are sent to the server using an HTTP POST request.
[1752] Step 6:
[1753] The server receives the data sent from the device and stores it in a MySQL or MongoDB database.
[1754] Input: 3D face model data, feature point data.
[1755] Output: Data stored in the database.
[1756] Specific operation: Analyzes the received data and executes SQL insert statements or MongoDB document inserts to store it in the database.
[1757] 2. Angle, pose, and facial expression generation process
[1758] Step 1:
[1759] The user selects a background photo within the application or takes a new one.
[1760] Input: Select or take a background photo.
[1761] Output: A background photo that you select or take.
[1762] Specific action: Select a photo from the gallery or launch the camera to take a new photo.
[1763] Step 2:
[1764] The user uses the UI within the application to specify the desired pose and expression.
[1765] Input: Specify pose and facial expression.
[1766] Output: Specified pose and facial expression.
[1767] Specific Actions: Use the drop-down menus and sliders to select the desired pose and expression.
[1768] Step 3:
[1769] The server uses Python's PIL library and OpenCV to analyze the background photo selected by the user and calculate the light direction and color tone.
[1770] Input: Background photo data.
[1771] Output: Light direction and color data.
[1772] Specific operation: OpenCV is used to perform image histogram and edge detection, and algorithms are run to identify the direction and color of the light source.
[1773] Step 4:
[1774] The server uses a deep learning model (using TensorFlow or PyTorch) to automatically generate angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[1775] Input: 3D model data, pose and facial expression specification.
[1776] Output: Newly generated angle, pose, and expression data.
[1777] Specific operation: Input data is fed to a pre-trained model using TensorFlow or PyTorch, and the resulting angle and facial expression data is obtained.
[1778] Step 5:
[1779] The emotion engine uses Microsoft Azure's Face API and Google Cloud Vision API to recognize the user's emotions from registered facial images.
[1780] Input: Facial image data.
[1781] Output: Emotion data.
[1782] Specific operation: Send facial image data to the API and obtain emotion recognition results.
[1783] Step 6:
[1784] The server calculates the synthesis parameters for synthesizing the analysis data of the background photo with the 3D model.
[1785] Input: Light direction and color data, 3D model data.
[1786] Output: Synthesis parameters.
[1787] Specific operation: Calculates parameters for naturally combining a 3D model with a background photo, taking into account light direction, color tone, and shadow position.
[1788] Step 7:
[1789] The server uses an automated Adobe Photoshop script and ImageMagick to use the generated data and compositing parameters to naturally composite the 3D face model with the background photo.
[1790] Input: 3D model data, background photo data, synthesis parameters.
[1791] Output: Composite image.
[1792] Specific operation: Using Adobe Photoshop scripts and ImageMagick, the 3D model is composited with a background photo based on the provided parameters.
[1793] Step 8:
[1794] The server transmits the composite image to the terminal.
[1795] Input: Synthetic image data.
[1796] Output: Send image to device.
[1797] Specific operation: The composite image is sent to the device via an HTTP response or WebSocket message.
[1798] 3. Photo generation and storage process
[1799] Step 1:
[1800] The terminal receives the composite image from the server, displays it on the screen, and asks the user for confirmation.
[1801] Input: Synthetic image data from the server.
[1802] Output: Image displayed to the user.
[1803] Specific operation: The received image data is set to a GUI element for display and displayed to the user.
[1804] Step 2:
[1805] The user checks the composite image and, if satisfied, presses the "Save" button. Editing functions are available as needed.
[1806] Input: User confirmation action.
[1807] Output: Save image or edit data.
[1808] Specific operation: Press the "Save" button to save the image. Press the edit button to display the edit screen.
[1809] Step 3:
[1810] The device saves the composite image to local storage using the Android or iOS local storage API.
[1811] Input: Synthetic image data.
[1812] Output: Image data saved to local storage.
[1813] What it does: It uses Android and iOS file storage APIs to save the composite image to the user's device.
[1814] Step 4:
[1815] The device provides the option to share to social media and cloud services, using the Facebook API and Google Drive API.
[1816] Input: User's share action.
[1817] Output: Upload images to social media or cloud services.
[1818] Specific operation: When you press the share button, the image will be uploaded via the API of the selected social media or cloud service.
[1819] (Application example 2)
[1820] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1821] In today's digital society, there is a growing demand for generating high-quality images with natural expressions and poses in situations where it is difficult for users to take selfies, where it is physically impossible to take group photos, or when virtual try-on simulations are required. However, conventional systems have struggled to generate natural images in these situations, particularly in automatically adjusting facial expressions to reflect the user's emotions. Furthermore, in virtual stores, there is a lack of technology to enhance the user's virtual try-on experience, and there is a need to improve user satisfaction.
[1822] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a means for a user to register their own facial image, a means for extracting feature points from the registered facial image and generating feature point data, and a means for generating a 3D facial model based on the feature point data. This enables high-quality images with natural angles, poses, and facial expressions to be automatically generated and naturally combined with background photos, even in situations where it is difficult for the user to take a selfie. The system also includes a means for recognizing the user's emotions and adjusting the facial expression based on the emotions, thereby enabling more natural image generation. Furthermore, the system also includes a means for simulating virtual trying-on of an item selected by the user in a virtual store, thereby improving the user's virtual trying-on experience.
[1823] A "user" is an entity that registers their own facial image in the system and uses the image generation and virtual try-on simulation.
[1824] "Means for registering face images" is a function that allows users to upload their own face images to the system.
[1825] The "means for extracting feature points" is a function that automatically extracts feature points such as the eyes, nose, and mouth from registered face images.
[1826] "Feature point data" is data including the position coordinates of the eyes, nose, mouth, etc. extracted from a face image.
[1827] The "means for generating a 3D face model" is a function that constructs a three-dimensional model of the user's face based on feature point data.
[1828] "Means for automatically generating angles, poses, and facial expressions" is a function that automatically generates specified angles, poses, and facial expressions for a 3D model.
[1829] A "background photo" is a photo of the background to be combined with the user's facial image.
[1830] "Natural synthesis means" is a function that seamlessly integrates a facial image with the generated angle, pose, and expression into a background photo.
[1831] The "means for presenting to the user" is a function for displaying the synthesized image on the user's screen.
[1832] The "means for saving" is a function that allows the user to save the composite image after checking it.
[1833] The "means for simulating virtual fitting" is a function that combines an object (for example, clothing) selected by the user with a three-dimensional model to generate the results of virtual fitting.
[1834] The "means for recognizing emotions" is a function that automatically determines the user's current emotions from their voice and image.
[1835] The "means for adjusting facial expressions" is a function that appropriately changes the facial expression of a 3D model based on the recognized emotion.
[1836] This system uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by the user, creating images that blend naturally with the background photograph. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to generate more natural facial expressions. This system is particularly intended for virtual try-on and image generation in virtual stores.
[1837] 1. The process by which users register their facial images
[1838] First, the user launches the application on their smartphone or PC and selects the facial image registration menu. For registration, they take photos of their face from the front, left side, and right side. The device extracts feature points from the captured facial images and generates a 3D model of the face based on these. This process uses an image processing library such as OpenCV. The generated 3D model and feature point data are then sent to the server, which then stores the data in a database.
[1839] 2. Virtual try-on simulation process
[1840] Within the application, users select the clothing and accessories they want to try on. The server then combines the selected product information with a 3D model of the user's face to generate a virtual try-on experience. This process uses 3D modeling libraries such as Three.js. Users can also specify their desired pose and angle.
[1841] 3. Emotion Recognition and Facial Expression Generation Process
[1842] The emotion engine analyzes the user's facial images and voice to recognize their current emotions. Based on these emotions, it then appropriately adjusts and generates the facial expressions of the 3D model. This process uses the Microsoft Azure Emotion API and other tools.
[1843] 4. Compositing process with background photo
[1844] After the user selects a background photo, the server analyzes the photo, calculates the light direction and color tone, calculates the parameters necessary for synthesis, and expresses the facial image as a 3D model, which is then naturally synthesized with the background photo. At this stage, the user's specified pose and facial expression are reflected.
[1845] 5. Synthesis and saving process
[1846] The server generates a composite image and sends it to the device, which then presents it to the user, who can save it after reviewing it. The option to share it on social media or cloud services is also provided.
[1847] Specific examples
[1848] For example, if a user tries on a new summer dress and generates a composite photo with a shopping mall background, the process would proceed as follows:
[1849] The user registers a face image and selects a new summer dress.
[1850] The selected dress is then composited with a 3D model of a face, a smiling expression is applied, and the dress is composited with the background of a shopping mall.
[1851] The generated composite image is displayed on a smartphone, allowing the user to view and save it.
[1852] Prompt Sentence Examples
[1853] "Generate a composite photo of you trying on a new summer dress and smiling with a shopping mall background"
[1854] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1855] Step 1: User registers a face image
[1856] The user launches the application on their smartphone or PC and selects the facial image registration menu. Next, they take photos of their face from the front, left side, and right side and upload them to the application. These input images serve as the base data for subsequent processing. The device receives these facial images and uses OpenCV to extract feature points such as the eyes, nose, and mouth. Feature point data is generated, and a 3D model of the face is created based on this. This 3D model is then used in subsequent synthesis processing.
[1857] Step 2: User selects an object for virtual try-on
[1858] Within the application, the user selects the clothing and accessories they want to try on. The data for this selection is provided as input to the system. The server receives this and prepares to combine a 3D model of the user's face with the information about the selected items. This process uses a 3D modeling library such as Three.js, and the output is a virtual try-on ready image.
[1859] Step 3: The emotion engine analyzes facial images and voice
[1860] The user's facial image and voice are provided as input to the emotion engine. The server analyzes this data using Microsoft Azure Emotion API or similar to recognize the user's current emotion. Emotion data is generated as a result of the analysis. This emotion data is then used to adjust facial expressions.
[1861] Step 4: Generate and adjust facial expressions based on emotion data
[1862] The server adjusts the facial expression of the user's 3D model based on the emotion data obtained in step 3. For example, if the user is excited, it applies a smiling expression. The individual feature point data of the registered facial image is used to generate the facial expression of the 3D model. The output is the adjusted 3D model.
[1863] Step 5: Select a background photo and analyze it
[1864] The user selects a background photo, which is sent as input to the server, which analyzes the background photo to calculate the light direction and color tone. Image processing algorithms are used for the analysis, and light and color parameters for compositing are generated as output.
[1865] Step 6: Generate the composite image
[1866] The server uses the 3D model of the virtual try-on garment generated in step 2, the 3D model of the user adjusted in step 4, and the light and color parameters obtained in step 5 to naturally combine the background photo and the facial image. The synthesis algorithm outputs a natural-looking composite image.
[1867] Step 7: Send the composite image to your device
[1868] The server sends the composite image to the terminal, which then presents the received composite image to the user, who then checks the presented composite image on the screen.
[1869] Step 8: Save and share your composite image
[1870] The user can review the composite image and save it if they are satisfied. The device then saves the composite image in local storage and provides the option to share it via social media or cloud services if desired. Finally, the saved composite image can be used according to the user's purpose.
[1871] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1872] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1873] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1874] [Fourth embodiment]
[1875] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1876] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1877] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1878] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1879] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1880] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1881] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1882] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1883] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1884] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1885] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1886] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1887] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1888] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and creates images that are naturally composited against a background photograph. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos.
[1889] 1. The process by which users register their facial images
[1890] User
[1891] First, the user launches the application on their smartphone or PC and selects the facial image registration menu.
[1892] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[1893] Terminal
[1894] The device automatically extracts feature points (e.g., the positions of the eyes, nose, and mouth) from each captured facial image.
[1895] A 3D model of the face is generated based on the feature point data.
[1896] The generated 3D model and feature point data are sent to the server.
[1897] server
[1898] The server receives the data sent from the terminal and stores it in a database.
[1899] The saved feature point data and 3D model are used in subsequent synthesis processes.
[1900] 2. Angle, pose, and facial expression generation process
[1901] User
[1902] Users can select a background photo within the application or take a new one.
[1903] The user specifies the desired pose and facial expression.
[1904] server
[1905] The server analyzes the background photo selected by the user and calculates the direction and color of the light.
[1906] Based on the pose and facial expression specified by the user, the angle, pose, and facial expression are automatically generated for the registered 3D model.
[1907] Calculate synthesis parameters for synthesizing the background photo analysis data and the 3D model.
[1908] The facial image is represented as a 3D model based on the generated angle, pose, and expression, and is then naturally combined with the background photo.
[1909] Terminal
[1910] The terminal receives the composite image and the composite parameters from the server and performs the composite process.
[1911] The synthesized image is presented to the user for confirmation.
[1912] 3. Photo generation and storage process
[1913] User
[1914] The user checks the generated composite image and presses the save button if satisfied.
[1915] Use further editing functions as needed.
[1916] Terminal
[1917] The terminal stores the composite image in local storage.
[1918] It also provides users with the option to share images on social media or cloud services, depending on their choice.
[1919] Specific examples
[1920] 1. Example of face registration
[1921] User A starts the app and registers three facial images: front, left, and right.
[1922] The device extracts feature points from these images and creates a 3D model of the face.
[1923] The generated data is sent to the server, which stores it in a database.
[1924] 2. Examples of group photos
[1925] A family of four wants to take a photo together, but one of them is in a remote location and it is physically impossible to take the photo.
[1926] The three people took a group photo on site, and the facial image of one person in a remote location was registered through the application.
[1927] Users select a group photo as a background within the app, and the server generates the appropriate angle, pose, and facial expression based on a remote facial model of a single person, and combines it with the background photo.
[1928] The composite image is displayed on the smartphone and can be saved by the user.
[1929] As described above, by implementing this invention, it becomes possible to easily generate ideal photos even in places where it is difficult to take selfies, and it is also possible to create high-quality group photos without having to physically gather together.
[1930] The processing flow will be explained below.
[1931] Program processing details
[1932] The process by which users register their facial images
[1933] Step 1:
[1934] The user launches the application on their smartphone or PC.
[1935] Step 2:
[1936] The user selects a face image registration menu and follows the instructions to take photos of the face from the front, left side, and right side.
[1937] Step 3:
[1938] The terminal receives each captured face image.
[1939] Step 4:
[1940] The device extracts feature points (such as the positions of the eyes, nose, and mouth) from the facial image.
[1941] Step 5:
[1942] The device generates a 3D model of the user's face based on the extracted feature point data.
[1943] Step 6:
[1944] The device sends the generated 3D model and feature point data to the server.
[1945] Step 7:
[1946] The server stores the received data in a database.
[1947] The process of generating angles, poses, and facial expressions
[1948] Step 1:
[1949] The user selects the photo creation menu of the application.
[1950] Step 2:
[1951] The user selects a background photo or takes a new one.
[1952] Step 3:
[1953] The user specifies the desired pose and facial expression.
[1954] Step 4:
[1955] The server receives the background photo and runs an image analysis algorithm to analyze the light direction and color tone.
[1956] Step 5:
[1957] Based on the analysis results, the server calculates the synthesis parameters for combining the image with the background photo.
[1958] Step 6:
[1959] The server automatically generates angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[1960] Step 7:
[1961] The server uses the generated data and synthesis parameters to represent the facial image as a 3D model and naturally combines it with the background photo.
[1962] Step 8:
[1963] The server transmits the composite image to the terminal.
[1964] Photo generation and storage process
[1965] Step 1:
[1966] The terminal receives the composite image sent from the server.
[1967] Step 2:
[1968] The terminal displays the received composite image on the screen and asks the user for confirmation.
[1969] Step 3:
[1970] The user checks the composite image and, if satisfied, presses the save button.
[1971] Step 4:
[1972] The device stores the composite image in local storage.
[1973] Step 5:
[1974] The device displays a notification to the user that the save is complete.
[1975] Step 6:
[1976] If the user wishes to share the image on a social networking site or cloud service, the device will provide the option to do so and process it accordingly.
[1977] Example 1
[1978] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1979] In conventional facial image synthesis systems, it was difficult for users to naturally combine their own facial images with background images, and differences in light direction and color tone often resulted in unnatural results. It was also difficult to naturally reflect the user's desired pose and facial expression. Furthermore, the process of efficiently extracting feature points from multiple facial images and generating a 3D model was complicated, resulting in poor usability for the entire system.
[1980] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1981] In this invention, the server includes means for a user to register his or her own facial image, means for extracting feature points from the registered facial image and generating feature point data, means for generating a 3D facial model based on the feature point data, means for analyzing a background image and identifying the light direction and color tone, means for automatically generating an angle, pose, and facial expression based on the analysis results, means for representing the facial image as a 3D model according to the generated angle, pose, and facial expression and naturally combining it with the background image, and means for presenting the combined image to the user and saving it after confirmation. This allows the user to easily generate natural combined images and improves operability, enabling more satisfying facial image synthesis.
[1982] "User" refers to an individual or group that uses this system to register their own facial images and create composite images.
[1983] "Means for registering" refers to a method or device that allows users to upload and store their facial images in the system.
[1984] "Feature points" are data that indicate specific parts of a face image, such as the eyes, nose, and mouth.
[1985] "Feature point data" is a data set that includes information on feature points extracted from a face image.
[1986] A "3D facial model" is a three-dimensional shape of a face reconstructed in three dimensions based on feature point data of a facial image.
[1987] A "background image" is a photograph or image selected by the user as a composite object.
[1988] The "analyzing means" refers to a method or device for calculating and analyzing the light direction and color tone of the background image.
[1989] "Means for automatically generating angles, poses, and facial expressions" refers to a method or device that has the function of automatically creating facial angles, poses, and facial expressions based on user specifications and analysis results.
[1990] The "synthesis means" refers to a method or device for naturally synthesizing the generated 3D face model with a background image.
[1991] The "presenting means" refers to a method or device for displaying the synthesized image to the user and requesting confirmation.
[1992] "Storing means" refers to a method or device for storing the confirmed composite image within the system or in external storage.
[1993] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and creates images that are naturally composited against a background photograph. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos.
[1994] The process by which users register their facial images
[1995] User
[1996] The user launches the application on their smartphone or PC and selects the facial image registration menu.
[1997] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[1998] Terminal
[1999] The device automatically extracts feature points (e.g., the positions of the eyes, nose, and mouth) from each captured facial image using Python's OpenCV library.
[2000] Based on the extracted feature point data, a 3D model of the face is generated using 3D modeling software such as Blender.
[2001] The generated 3D model and feature point data are sent to the server via REST API.
[2002] server
[2003] The server receives the data sent from the device and stores it in a MySQL database using the Django framework.
[2004] The saved feature point data and 3D model are used in subsequent synthesis processes.
[2005] The process of generating angles, poses, and facial expressions
[2006] User
[2007] Users can select a background photo within the application or take a new one.
[2008] The user specifies the desired pose and facial expression.
[2009] server
[2010] The server analyzes the background photo selected by the user and calculates the light direction and color tone using PIL (Python Imaging Library) and OpenCV.
[2011] Based on the pose and facial expression specified by the user, the angle, pose, and facial expression are automatically generated for the registered 3D model. Here, generative AI models (such as StyleGAN and DeepFaceLab) are utilized.
[2012] Calculate synthesis parameters for synthesizing the background photo analysis data and the 3D model.
[2013] The facial image is represented based on a 3D model according to the generated angle, pose, and expression, and is then naturally combined with the background photo.
[2014] The process by which the device generates a composite image and presents it to the user
[2015] Terminal
[2016] The terminal receives the composite image and the composite parameters from the server.
[2017] The synthesis process is performed and the completed image is presented to the user.
[2018] User
[2019] The user checks the presented composite image and presses the save button if satisfied.
[2020] Further modifications can be made within the app if needed.
[2021] The generated images can also be shared on social media or cloud services.
[2022] Specific examples
[2023] Specific examples of face registration
[2024] The user launches the app and takes a series of photos of their face - front, left, and right - and registers them. The device uses Python's OpenCV library to extract feature points from these images and creates a 3D face model using Blender. The generated data is sent to the server via a REST API and stored in a MySQL database using the Django framework.
[2025] Examples of group photos
[2026] If a family wants to take a group photo together, but one person is remote and cannot physically be present, the remote person can register their face image through the app. The user then takes a group photo on-site and selects a background photo within the app. The server then analyzes the image, calculating the light direction and color tone, and generates the appropriate angle, pose, and facial expression based on a 3D facial model of the remote person, which is then combined with the background photo. The resulting composite image is then displayed on the user's smartphone.
[2027] Prompt Sentence Examples
[2028] "We will create a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and then creates an image that is composited into a background photo. We will also analyze the background photo, such as the direction of light and color tone, to ensure that the composite image looks natural."
[2029] By implementing this invention, ideal selfies can be easily generated even in places where it is difficult to take selfies, and high-quality group photos can be created without physically gathering together.
[2030] The flow of the identification process in the first embodiment will be described with reference to FIG.
[2031] The flow of this system's program processing
[2032] Step 1:
[2033] explanation
[2034] Users launch the application on their smartphone or PC, select the face image registration menu, and follow the app's guide to take and register photos of their face from the front, left side, and right side.
[2035] Input and Output
[2036] Input: Face image (front, left, right)
[2037] Output: Three registered face images
[2038] Specific actions
[2039] The user clicks the "Register Face Image" button on the app, activates the camera to take face images from various angles, and then presses the "Confirm / Register" button.
[2040] Step 2:
[2041] explanation
[2042] The device automatically extracts feature points (such as the positions of the eyes, nose, and mouth) from the captured facial image using Python's OpenCV library. Based on the extracted feature point data, a 3D model of the face is generated using 3D modeling software such as Blender.
[2043] Input and Output
[2044] Input: Three registered face images
[2045] Output: feature point data, 3D face model
[2046] Specific actions
[2047] The device uses the OpenCV library to detect feature points from each facial image, then uses Blender to generate a 3D model of the face based on these feature points. The generated model and feature point data are then sent to the server via a REST API.
[2048] Step 3:
[2049] explanation
[2050] The server receives the data sent from the device and stores it in a MySQL database using the Django framework. The stored feature point data and 3D model are then used in the synthesis process.
[2051] Input and Output
[2052] Input: feature point data, 3D face model
[2053] Output: Data stored in the database
[2054] Specific actions
[2055] The server receives the transmitted data via the Django framework and stores it in a MySQL database.
[2056] Step 4:
[2057] explanation
[2058] Users can select a background photo within the application or take a new one, and specify the desired pose and facial expression.
[2059] Input and Output
[2060] Input: Background photo, desired pose and facial expression
[2061] Output: User selections and specifications
[2062] Specific actions
[2063] Once the user selects (or takes) a background photo, a UI for specifying the pose and expression is displayed.
[2064] Step 5:
[2065] explanation
[2066] The server analyzes the background photo selected by the user and calculates the light direction and color tone using PIL (Python Imaging Library) and OpenCV.
[2067] Input and Output
[2068] Input: Background photo
[2069] Output: Analysis results of light direction and color tone
[2070] Specific actions
[2071] The server loads the background photo using the PIL or OpenCV library and analyzes the light direction and color tone.
[2072] Step 6:
[2073] explanation
[2074] The server uses a generative AI model (e.g., StyleGAN or DeepFaceLab) to automatically generate angles, poses, and expressions for the registered 3D model based on the poses and expressions specified by the user.
[2075] Input and Output
[2076] Input: feature point data, 3D face model, desired pose and expression
[2077] Output: Generated angles, poses, and expressions
[2078] Specific actions
[2079] The server uses a generative AI model to generate poses and facial expressions based on feature point data and a 3D model.
[2080] Step 7:
[2081] explanation
[2082] The server calculates synthesis parameters for combining the analysis data of the background photo with a 3D model, and then represents the facial image as a 3D model according to the generated angle, pose, and expression, which is then naturally combined with the background photo.
[2083] Input and Output
[2084] Input: background photo, analysis data, generated angles, poses, facial expressions
[2085] Output: Synthesized face image
[2086] Specific actions
[2087] The server calculates the synthesis parameters and uses OpenCV to naturally synthesize the face image and background photo.
[2088] Step 8:
[2089] explanation
[2090] The terminal receives the composite image and the composite parameters from the server, performs the composite process, and presents the completed composite image to the user.
[2091] Input and Output
[2092] Input: synthetic image, synthetic parameters
[2093] Output: The composite image presented to the user
[2094] Specific actions
[2095] The device retrieves the composite image from the server via an HTTP request and displays it to the user on the app.
[2096] Step 9:
[2097] explanation
[2098] The user can review the composite image presented to them and, if satisfied, press the save button. If necessary, they can make further edits within the app. The generated image can also be shared via social media or cloud services.
[2099] Input and Output
[2100] Input: Composite image
[2101] Output: Saved and shared images
[2102] Specific actions
[2103] When the user presses the "Save" button, the image is saved to the smartphone's local storage, and when the user presses the "Share" button, the image is uploaded to social media or a cloud service.
[2104] (Application example 1)
[2105] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2106] Modern electronic payment services require enhanced security. However, many current systems have limited user authentication methods, making it difficult to simultaneously improve convenience and security. Furthermore, many systems can only authenticate users at specific angles or poses, which can be inconvenient for users. There is a need to solve these issues and realize more accurate and convenient facial recognition. Furthermore, a system is needed that can accurately authenticate users even when they have different facial expressions or poses.
[2107] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2108] In this invention, the server includes: a means for a user to register his or her own facial image; a means for extracting feature points from the registered facial image and generating feature point data; a means for generating a 3D facial model based on the feature point data; a means for automatically generating an angle, pose, and facial expression to be combined with a background photograph; a means for representing the facial image as a 3D model according to the generated angle, pose, and facial expression and combining it naturally with the background photograph; a means for performing electronic authentication using analysis data of the background photograph and the user's 3D model; and a means for presenting the combined image to the user and saving it after confirmation. This enables advanced electronic authentication by generating a natural combined image based on the user's registered facial image and background photograph. A means for the user to specify a desired pose and facial expression; a means for generating corresponding movements for the facial portion of the 3D model based on the specified pose and facial expression; a means for performing electronic authentication based on the generated 3D model and a captured image of the user; and a means for reconstructing the facial image according to the generated movements, thereby providing an authentication system that combines convenience and security.
[2109] "Means for users to register their own facial images" refers to a function that allows users to register their own facial images using devices such as smartphones or computers.
[2110] "Means for extracting feature points and generating feature point data" refers to a function that uses AI technology to detect feature points such as the eyes, nose, and mouth from registered facial images and converts them into data.
[2111] The "means for generating a 3D model of a face" is a function for generating a 3D model of a face using feature point data.
[2112] "Means for automatically generating angles, poses, and facial expressions for compositing with background photos" is a function that uses AI to automatically generate the optimal facial angle, pose, and facial expression for compositing with a background photo.
[2113] "Means for representing a facial image as a 3D model according to the generated angle, pose, and expression, and for naturally combining it with a background photograph" refers to a function for reproducing a facial image as a 3D model based on the generated angle, pose, and expression, and for naturally combining it with a background photograph.
[2114] "Means for electronic authentication using background photo analysis data and a 3D model of the user" is a function for securely and accurately authenticating a user using analysis data of the light direction and color tone of the background photo and a 3D model of the user.
[2115] The "means for presenting the synthesized image to the user and saving it after confirmation" is a function for displaying the generated synthesized image to the user and saving the image after obtaining confirmation.
[2116] "Means for generating corresponding movements for the facial part of a 3D model based on a specified pose or expression" is a function that moves the facial part of a 3D model according to a pose or expression specified by the user.
[2117] The "means for reconstructing a facial image according to the generated movements" is a function for reconstructing a facial image based on the generated poses and facial movements.
[2118] "Means for electronic authentication" refers to an authentication function that uses a user's facial image or 3D model to allow access to the system.
[2119] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by users, and creates images that are naturally composited with background photographs. Below, we will provide a detailed explanation of how this system is realized.
[2120] 1. Facial image registration process
[2121] Users register their facial images using a smartphone or computer. Specifically, they take photos of their face from the front, left side, and right side, and register them through the application.
[2122] The device receives a facial image and extracts its features using a library such as dlib. This generates key feature data such as the eyes, nose, and mouth. A 3D model of the face is then generated based on this feature data. This 3D model and feature data are then sent to a server and stored in a database.
[2123] 2. Angle, pose, and facial expression generation process
[2124] The server analyzes the background photo selected or taken by the user. Specifically, it calculates the direction of light and color tone. Based on this analysis data, it automatically generates the appropriate angle, pose, and expression for the registered 3D model according to the pose and expression specified by the user. Using AI technology such as DeepFace, it calculates synthesis parameters to naturally combine the user's 3D face model with the background photo.
[2125] 3. Electronic Authentication Process
[2126] The device performs electronic authentication using background photo analysis data and a 3D model of the user. Specifically, it uses DeepFace's Facenet model to match the captured image of the user with the 3D model for accurate authentication.
[2127] 4. Photo Generation and Storage Process
[2128] Users can review the resulting composite image and save it if they are satisfied. Further editing is available if needed. The composite image is saved to local storage, and users are also given the option to share it via social media or cloud services.
[2129] The system that realizes this invention executes a series of processes, from registering a facial image, generating a 3D model, using that model for electronic authentication, and finally generating and saving a composite image. This enables users to be accurately and securely authenticated even with different poses and expressions, and to obtain a high-quality composite image.
[2130] Examples:
[2131] A user installs the FacePay app and registers images of their face, showing the front, left, and right sides. After making a purchase at a convenience store, they perform payment using facial recognition, and the payment is completed upon successful authentication.
[2132] Example prompt sentence:
[2133] "Design an application that generates a 3D face model from a facial image registered by a user and uses that model for facial authentication during electronic payments."
[2134] As described above, by implementing the present invention, it is possible to generate ideal photos even when selfies are difficult, and to create high-quality group photos without physically gathering together. Furthermore, by using this system, the security of electronic payments is improved, and user convenience is also enhanced.
[2135] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2136] Step 1:
[2137] The user registers a facial image. In this process, the user launches an application on a smartphone or PC and takes and registers facial images from the front, left side, and right side.
[2138] Input: Face image taken by the user (front, left, right)
[2139] Output: Facial feature point data and a 3D model of the face generated by the device
[2140] Step 2:
[2141] The device processes the facial image provided by the user. Specifically, it uses the dlib library to extract feature points such as the eyes, nose, and mouth from the captured facial image. Based on this feature point data, a 3D model of the face is generated. The generated data is sent to a server and stored in a database.
[2142] Input: A face image obtained from the user
[2143] Output: Feature point data and a 3D face model
[2144] Step 3:
[2145] When a user selects or takes a background photo, the server analyzes the background photo, detecting the direction and color of the light, and then uses this information to calculate the direction and color of the light in the background photo.
[2146] Input: A background photo selected or taken by the user
[2147] Output: Analysis data of background photo (light direction and color tone)
[2148] Step 4:
[2149] The server automatically generates the angle, pose, and facial expression for the registered 3D model according to the pose and facial expression specified by the user. Based on this information, the server expresses the facial image as a 3D model and calculates synthesis parameters for naturally combining it with the background photo.
[2150] Input: Analysis data of user-specified poses and expressions, and background photos
[2151] Output: angle, pose, facial expression, synthesis parameters
[2152] Step 5:
[2153] The device uses the synthesis parameters and 3D model received from the server to naturally synthesize the background photo and facial image. Specifically, the synthesis process is carried out using DeepFace's AI technology.
[2154] Input: Synthesis parameters and 3D model provided by the server
[2155] Output: Composite image
[2156] Step 6:
[2157] The device will then present the combined image to the user for review, and if the user is satisfied, the image will be saved to local storage, with the option to share it via social media or cloud services if desired.
[2158] Input: The composite image
[2159] Output: User confirmation and save
[2160] Step 7:
[2161] The device performs electronic authentication using background photo analysis data and a 3D model of the user, and uses DeepFace's Facenet model to match the captured image with the 3D model.
[2162] Input: Captured image, background photo analysis data, user 3D model
[2163] Output: Authentication result (success or failure)
[2164] Through a series of processes from registration to synthesis and authentication, highly accurate and convenient electronic authentication is achieved.
[2165] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2166] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on facial images registered by the user, creating images that blend naturally with the background photo, and further combines it with an emotion engine that recognizes the user's emotions. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos.
[2167] 1. The process by which users register their facial images
[2168] User
[2169] First, the user launches the application on their smartphone or PC and selects the facial image registration menu.
[2170] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[2171] Terminal
[2172] The device automatically extracts feature points (such as the position of the eyes, nose, and mouth) from each captured facial image.
[2173] A 3D model of the face is generated based on the feature point data.
[2174] The generated 3D model and feature point data are sent to the server.
[2175] server
[2176] The server receives the data sent from the terminal and stores it in a database.
[2177] The saved feature point data and 3D model are used in subsequent synthesis processes.
[2178] 2. Angle, pose, and facial expression generation process
[2179] User
[2180] Users can select a background photo within the application or take a new one.
[2181] The user specifies the desired pose and facial expression.
[2182] server
[2183] The server analyzes the background photo selected by the user and calculates the direction and color of the light.
[2184] Based on the pose and facial expression specified by the user, the angle, pose, and facial expression are automatically generated for the registered 3D model.
[2185] Calculate synthesis parameters for synthesizing the background photo analysis data and the 3D model.
[2186] Emotion Engine
[2187] The emotion engine recognizes the user's emotion from the registered facial image.
[2188] The facial expression is adjusted and generated based on the user's desired pose, facial expression, and recognized emotion.
[2189] The emotion engine can also analyze the user's voice input, recognize the user's emotion based on the analysis results, and generate facial expressions based on that.
[2190] server
[2191] The server uses the generated data and synthesis parameters to represent the facial image as a 3D model and naturally combines it with the background photo.
[2192] The server transmits the composite image to the terminal.
[2193] 3. Photo generation and storage process
[2194] Terminal
[2195] The terminal receives the composite image from the server.
[2196] The terminal displays the composite image on the screen and asks the user for confirmation.
[2197] User
[2198] The user checks the composite image and, if satisfied, presses the save button.
[2199] Use further editing functions as needed.
[2200] Terminal
[2201] The device stores the composite image in local storage.
[2202] It also provides the option to share to social media and cloud services.
[2203] Specific examples
[2204] 1. Example of face registration
[2205] User A starts the app and registers three facial images: front, left, and right.
[2206] The device extracts feature points from these images and generates a 3D model of the face.
[2207] The generated data is sent to the server, which stores it in a database.
[2208] 2. Specific examples of emotion engines
[2209] User B selects a landscape photo as a background in the app, and the app recognizes B's current emotion.
[2210] If the user desires a "smiling" facial expression, but the emotion engine recognizes the emotion of "surprise," it applies the surprised expression to the facial image and generates a 3D model.
[2211] The server calculates the synthesis parameters based on this data and synthesizes it with the landscape photo.
[2212] 3. Examples of group photos
[2213] A family of four wants to take a photo together, but one of them is in a remote location and it is physically impossible to take the photo.
[2214] The three people took a group photo on site, and the facial image of one person in a remote location was registered through the application.
[2215] Users select a group photo as a background within the app, and the server generates the appropriate angle, pose, and facial expression based on a remote facial model of a single person, and combines it with the background photo.
[2216] The emotion engine recognizes the current emotion of a person in a remote location and generates and adjusts facial expressions to match that emotion.
[2217] The composite image is displayed on the smartphone and can be saved by the user.
[2218] As described above, by implementing this invention, it becomes possible to easily create ideal photos even in places where it is difficult to take selfies, and it is also possible to create high-quality group photos without physically gathering together.Furthermore, the emotion engine can generate natural facial expressions that better match the user's emotions, improving image quality and satisfaction.
[2219] The processing flow will be explained below.
[2220] Program processing details
[2221] The process by which users register their facial images
[2222] Step 1:
[2223] The user launches the application on their smartphone or PC.
[2224] Step 2:
[2225] The user selects a face image registration menu and follows the instructions to take photos of the face from the front, left side, and right side.
[2226] Step 3:
[2227] The terminal receives each captured face image.
[2228] Step 4:
[2229] The device extracts feature points (such as the positions of the eyes, nose, and mouth) from the facial image.
[2230] Step 5:
[2231] The device generates a 3D model of the user's face based on the extracted feature point data.
[2232] Step 6:
[2233] The device sends the generated 3D model and feature point data to the server.
[2234] Step 7:
[2235] The server stores the received data in a database.
[2236] The process of generating angles, poses, and facial expressions
[2237] Step 1:
[2238] The user selects the photo creation menu of the application.
[2239] Step 2:
[2240] The user selects a background photo or takes a new one.
[2241] Step 3:
[2242] The user specifies the desired pose and facial expression.
[2243] Step 4:
[2244] The server receives the background photo and runs an image analysis algorithm to analyze the light direction and color tone.
[2245] Step 5:
[2246] Based on the analysis results, the server calculates the synthesis parameters for combining the image with the background photo.
[2247] Step 6:
[2248] The server automatically generates angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[2249] Step 7:
[2250] The emotion engine recognizes the user's emotions from registered facial images or voice input.
[2251] Step 8:
[2252] The emotion engine adjusts and generates poses and expressions based on the recognized emotions.
[2253] Step 9:
[2254] The server uses the generated data and synthesis parameters to represent the facial image as a 3D model and naturally combines it with the background photo.
[2255] Step 10:
[2256] The server transmits the composite image to the terminal.
[2257] Photo generation and storage process
[2258] Step 1:
[2259] The terminal receives the composite image sent from the server.
[2260] Step 2:
[2261] The terminal displays the received composite image on the screen and asks the user for confirmation.
[2262] Step 3:
[2263] The user checks the composite image and, if satisfied, presses the save button.
[2264] Step 4:
[2265] The device stores the composite image in local storage.
[2266] Step 5:
[2267] The device displays a notification to the user that the save is complete.
[2268] Step 6:
[2269] If the user wishes to share the image on a social networking site or cloud service, the device will provide the option to do so and process it accordingly.
[2270] Example 2
[2271] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2272] Conventional photo compositing technologies have difficulty in generating natural poses and facial expressions desired by users, and require manual setting of compositing parameters such as light direction and color tone. This has required a great deal of time and effort to generate a photo that satisfies the user. Furthermore, when taking a group photo and not everyone is present, it is difficult to composite people in remote locations naturally. Furthermore, there has been a lack of mechanisms for appropriately reflecting facial expressions that match the user's emotions. The objective of this invention is to solve these problems and realize easier and more natural photo compositing.
[2273] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[2274] In this invention, the server includes: means for a user to register his or her own facial image; means for extracting feature points from the registered facial image and generating feature point data; means for generating a 3D facial model based on the feature point data; means for automatically generating an angle, pose, and facial expression to be composited with a background photograph; means for representing the facial image as a 3D model according to the generated angle, pose, and facial expression and naturally composited with the background photograph; means including an emotion recognition engine that recognizes the user's emotion and adjusts and generates the facial expression of the facial image based on the emotion; means for presenting the composite image to the user and saving it after confirmation; and means for saving facial image data and feature point data received from the terminal. This enables the user to easily specify a desired pose and facial expression and automatically generate a natural composite photograph, and further enables the realization of a photograph with a realistic expression that reflects the user's emotion.
[2275] "User" refers to an individual who uses this system to register a facial image and perform various operations to generate a composite photograph.
[2276] "Facial image" refers to image data of a user's face, and includes photographs of the front, left side, and right side.
[2277] "Feature points" refer to position data of the eyes, nose, mouth, etc. extracted from a facial image, and are used to represent the structure and shape of the face.
[2278] "Feature point data" is a representation of feature points as numerical data, and is the basic data for generating a 3D model of a face.
[2279] A "3D facial model" refers to a digital model of a face expressed in three-dimensional space based on feature point data.
[2280] "Background photo" refers to photo data that is designated or taken by the user and used as the background of the composite photo.
[2281] "Angle" refers to a parameter that indicates the direction in which the 3D face model is facing relative to the background photo.
[2282] "Pose" refers to parameters that indicate the posture and gesture of a 3D facial model.
[2283] "Facial expression" refers to parameters that indicate emotions or expressions applied to a 3D facial model.
[2284] An "emotion recognition engine" refers to a software module that recognizes emotions from a user's facial image or voice input, and adjusts and generates facial expressions based on those emotions.
[2285] "Synthesis parameters" refer to the settings such as light direction and color tone required to naturally composite a 3D facial model onto a background photo.
[2286] A "composite photo" refers to an image created by combining a background photo with a 3D model of a face.
[2287] "Terminal" refers to a device such as a smartphone or PC on which a user runs an application.
[2288] "Server" refers to a remote server that stores facial images and feature point data and performs various processing.
[2289] MODE FOR CARRYING OUT THE INVENTION
[2290] This invention is a system that uses AI to automatically generate angles, poses, and facial expressions based on a facial image registered by the user, creating a naturally composite image against a background photograph. It also incorporates an emotion engine that recognizes the user's emotions. This system is useful in situations where it is difficult for users to take selfies or when it is physically difficult to take group photos. Below, we explain each step for implementing this invention.
[2291] 1. The process by which users register their facial images
[2292] User
[2293] The user launches the application on their smartphone or PC and selects the facial image registration menu.
[2294] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[2295] Terminal
[2296] The device uses libraries such as OpenCV and dlib to automatically extract feature points (such as the position of the eyes, nose, and mouth) from each captured facial image.
[2297] The device uses tools like Blender or Three.js to generate a 3D model of the face based on the feature point data.
[2298] The generated 3D model and feature point data are sent to the server.
[2299] server
[2300] The server receives the data sent from the device and stores it in a MySQL or MongoDB database.
[2301] The saved feature point data and 3D model are used in subsequent synthesis processes.
[2302] 2. Angle, pose, and facial expression generation process
[2303] User
[2304] The user selects a background photo within the application or takes a new one.
[2305] The user uses the UI within the application to specify the desired pose and expression.
[2306] server
[2307] The server uses Python's PIL library and OpenCV to analyze the background photo selected by the user and calculate the light direction and color tone.
[2308] The server uses a deep learning model (using TensorFlow or PyTorch) to automatically generate angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[2309] The server calculates the synthesis parameters for synthesizing the analysis data of the background photo with the 3D model.
[2310] Emotion Engine
[2311] The emotion engine uses Microsoft Azure's Face API and Google Cloud Vision API to recognize the user's emotions from registered facial images.
[2312] The facial expressions are adjusted and generated based on the user's desired pose and facial expression and the recognized emotion.
[2313] The emotion engine uses the Google Cloud Speech-to-Text API to analyze the user's voice input to recognize emotions and generate facial expressions based on them.
[2314] server
[2315] The server uses an automated script in Adobe Photoshop or ImageMagick to use the generated data and synthesis parameters to naturally synthesize the 3D face model with the background photo.
[2316] The server transmits the composite image to the terminal.
[2317] 3. Photo generation and storage process
[2318] Terminal
[2319] The terminal receives the composite image from the server, displays it on the screen, and asks the user for confirmation.
[2320] User
[2321] The user checks the composite image and, if satisfied, presses the "Save" button. Editing functions are available as needed.
[2322] Terminal
[2323] The device saves the composite image to local storage using the Android or iOS local storage API.
[2324] The device provides the option to share to social media and cloud services, using the Facebook API and Google Drive API.
[2325] Specific examples
[2326] 1. Example of face registration
[2327] User A starts the app and registers three facial images: front, left, and right.
[2328] The device uses OpenCV to extract feature points from these images and uses Blender to generate a 3D model of the face.
[2329] The generated data is sent to the server, which stores it in a database.
[2330] 2. Specific examples of emotion engines
[2331] User B selects a landscape photo as a background in the app, and the app uses the Microsoft Azure Face API to recognize B's current emotion.
[2332] If the user desires a "smiling" facial expression, but the emotion engine recognizes the emotion of "surprise," it applies the surprised expression to the facial image and generates a 3D model.
[2333] The server calculates the synthesis parameters based on this data and synthesizes it with the landscape photo.
[2334] 3. Examples of group photos
[2335] A family of four wants to take a photo together, but one of them is in a remote location and it is physically impossible to take the photo.
[2336] The three people took a group photo on site, and the facial image of one person in a remote location was registered through the application.
[2337] The user selects a group photo as a background within the app, and the server generates the appropriate angle, pose, and expression based on a single face model located remotely, and combines it with the background photo.
[2338] The emotion engine recognizes the current emotion of a person in a remote location and generates and adjusts facial expressions to match that emotion.
[2339] The composite image is displayed on the smartphone and can be saved by the user.
[2340] As described above, by implementing this invention, it becomes possible to easily create ideal photos even in places where it is difficult to take selfies, and to create high-quality group photos without physically gathering together. Furthermore, the emotion engine can generate natural facial expressions that match the user's emotions, improving image quality and satisfaction.
[2341] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2342] Processing flow
[2343] 1. The process by which users register their facial images
[2344] Step 1:
[2345] The user starts the application and selects the face image registration menu.
[2346] Input: An action on the user interface.
[2347] Output: Camera launch and registration menu display.
[2348] Specific operation: Click the "Register face image" button on the application's home screen, and the following screen will appear.
[2349] Step 2:
[2350] The user follows the app's instructions to take and register photos of their face from the front, left side, and right side.
[2351] Input: Camera operation and photography.
[2352] Output: 3 face images (front, left, right).
[2353] Specific operation: The camera screen will appear and you will be prompted to "Take a photo of the front" and then take a photo, repeating this three times.
[2354] Step 3:
[2355] The device uses libraries such as OpenCV and dlib to automatically extract feature points (such as the position of the eyes, nose, and mouth) from each captured facial image.
[2356] Input: 3 face images.
[2357] Output: Feature point data for each face image.
[2358] Specific operation: Using the dlib facial feature detection model, an algorithm is run to extract 68 facial feature points from an image.
[2359] Step 4:
[2360] The device uses Blender or Three.js to generate a 3D model of the face based on feature point data.
[2361] Input: feature point data.
[2362] Output: 3D face model data.
[2363] Specific operation: Using Blender's Python API, run a script that generates a 3D model based on feature point data.
[2364] Step 5:
[2365] The device sends the generated 3D model and feature point data to the server.
[2366] Input: 3D face model data, feature point data.
[2367] Output: Sending data to the server.
[2368] Specific operation: 3D model data and feature point data are sent to the server using an HTTP POST request.
[2369] Step 6:
[2370] The server receives the data sent from the device and stores it in a MySQL or MongoDB database.
[2371] Input: 3D face model data, feature point data.
[2372] Output: Data stored in the database.
[2373] Specific operation: Analyzes the received data and executes SQL insert statements or MongoDB document inserts to store it in the database.
[2374] 2. Angle, pose, and facial expression generation process
[2375] Step 1:
[2376] The user selects a background photo within the application or takes a new one.
[2377] Input: Select or take a background photo.
[2378] Output: A background photo that you select or take.
[2379] Specific action: Select a photo from the gallery or launch the camera to take a new photo.
[2380] Step 2:
[2381] The user uses the UI within the application to specify the desired pose and expression.
[2382] Input: Specify pose and facial expression.
[2383] Output: Specified pose and facial expression.
[2384] Specific Actions: Use the drop-down menus and sliders to select the desired pose and expression.
[2385] Step 3:
[2386] The server uses Python's PIL library and OpenCV to analyze the background photo selected by the user and calculate the light direction and color tone.
[2387] Input: Background photo data.
[2388] Output: Light direction and color data.
[2389] Specific operation: OpenCV is used to perform image histogram and edge detection, and algorithms are run to identify the direction and color of the light source.
[2390] Step 4:
[2391] The server uses a deep learning model (using TensorFlow or PyTorch) to automatically generate angles, poses, and facial expressions for the registered 3D model based on the poses and facial expressions specified by the user.
[2392] Input: 3D model data, pose and facial expression specification.
[2393] Output: Newly generated angle, pose, and expression data.
[2394] Specific operation: Input data is fed to a pre-trained model using TensorFlow or PyTorch, and the resulting angle and facial expression data is obtained.
[2395] Step 5:
[2396] The emotion engine uses Microsoft Azure's Face API and Google Cloud Vision API to recognize the user's emotions from registered facial images.
[2397] Input: Facial image data.
[2398] Output: Emotion data.
[2399] Specific operation: Send facial image data to the API and obtain emotion recognition results.
[2400] Step 6:
[2401] The server calculates the synthesis parameters for synthesizing the background photo analysis data with the 3D model.
[2402] Input: Light direction and color data, 3D model data.
[2403] Output: Synthesis parameters.
[2404] Specific operation: Calculates parameters for naturally combining a 3D model with a background photo, taking into account light direction, color tone, and shadow position.
[2405] Step 7:
[2406] The server uses an automated script in Adobe Photoshop or ImageMagick to use the generated data and synthesis parameters to naturally synthesize the 3D face model with the background photo.
[2407] Input: 3D model data, background photo data, synthesis parameters.
[2408] Output: Composite image.
[2409] Specific operation: Using Adobe Photoshop scripts and ImageMagick, the 3D model is composited with a background photo based on the provided parameters.
[2410] Step 8:
[2411] The server transmits the composite image to the terminal.
[2412] Input: Synthetic image data.
[2413] Output: Send image to device.
[2414] Specific operation: The composite image is sent to the device via an HTTP response or WebSocket message.
[2415] 3. Photo generation and storage process
[2416] Step 1:
[2417] ...
Claims
1. A means for a user to register his / her own facial image; A means for extracting feature points from a registered face image and generating feature point data; A means for generating a 3D model of a face based on the feature point data; A method for automatically generating angles, poses, and facial expressions to be combined with background photos, A means for expressing a facial image as a 3D model according to the generated angle, pose, and expression, and for naturally combining the 3D model with a background photograph; A means for presenting the synthesized image to a user and saving it after confirmation; A system including:
2. means for analyzing the direction and color tone of light based on a background photograph designated by a user; means for calculating parameters of a composite image based on the analysis results; The system of claim 1 , wherein the system naturally combines the background photograph and the facial image according to the calculated parameters.
3. A means for the user to specify a desired pose or expression; means for generating a corresponding motion for a face part of a 3D model based on a specified pose and expression; The system of claim 1 , wherein the system reconstructs a facial image according to the generated motion.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A
Cited By
Systems and methods for charge state assignment in mass spectrometry
US12683139B2