system

The system generates and delivers realistic travel photos by combining user avatars with travel destination backgrounds, addressing the challenge of creating and sharing such images automatically and interactively.

JP2026035228APending Publication Date: 2026-03-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024138071
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Users face challenges in creating realistic travel photos that combine their own facial images with real-world scenery due to time and financial constraints, and existing systems lack features for automatic generation and delivery at specified times.

Method used

A system that receives a user's facial photograph, generates an avatar using facial recognition, combines it with a background image of a travel destination, and delivers the composite image in a specified manner, optionally including additional images and emotions.

Benefits of technology

Enables users to easily obtain and share realistic travel photos that simulate actual travel experiences, providing a richer and interactive experience on social media and other platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035228000001_ABST
    Figure 2026035228000001_ABST
Patent Text Reader

Abstract

Provide a system. [Solution] means for receiving a facial photograph of the user; means for generating an avatar of the user based on the received facial photograph using a facial recognition algorithm; means for acquiring a background image based on a travel destination designated by a user; A means for synthesizing the acquired background image and the generated avatar to create a realistic travel photo; means for transmitting the created travel photos in a manner designated by the user; A system including:
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, many people have expressed a desire to travel abroad, but time and financial constraints often make it difficult to realize this dream. Furthermore, even for users interested in social media and image generation technology, there are limited ways to easily obtain realistic images that combine real-world scenery with themselves. Given this background, there is a demand for a system that can provide users with photos that make them appear as if they are actually traveling. [Means for solving the problem]

[0005] This invention provides a system including means for receiving a user's facial photograph, means for generating an avatar for the user based on the received facial photograph using a facial recognition algorithm, means for acquiring a background image based on a travel destination designated by the user, means for creating a realistic travel photograph by combining the acquired background image with the generated avatar, and means for transmitting the created travel photograph in a manner designated by the user, thereby enabling users to easily receive realistic photographs that make them feel as if they are traveling the world.

[0006] Furthermore, the system can provide a richer experience by including a means for generating additional optional images, including images of specific local items and meals, based on the travel destination. Also, the system can provide continuous and consistent service by including a means for saving the travel destination selection information and transmission method information specified by the user and automatically generating travel photos based on the predetermined information at the set date and time.

[0007] "User" refers to any individual or legal entity that uses the System.

[0008] A "mugshot" refers to an image that shows certain features of a user's face.

[0009] A "facial recognition algorithm" refers to a computational method for extracting facial features from image data and identifying a specific individual.

[0010] An "avatar" is a digital virtual persona created based on a user's facial photograph.

[0011] "Destination" refers to a geographical destination that a user specifies within the system.

[0012] "Background image" refers to an image of scenery, buildings, etc. related to the travel destination specified by the user.

[0013] "Compositing" refers to the process of combining multiple different image elements into a single image.

[0014] "Travel photography" refers to realistic images created by combining background images with avatars.

[0015] "Sending method" refers to the means (email, SNS, LINE, etc.) used to deliver the generated travel photos to the user.

[0016] "Additional optional images" refer to images of specific local items or food included in travel photos.

[0017] "Setting information" refers to information such as the travel destination, transmission method, and transmission date and time specified by the user. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The system of the present invention generates a traveling avatar based on a facial photograph provided by the user, generates a composite image based on a travel destination selected by the user, and transmits the composite image in a specified manner. Specific embodiments of the present invention are described below.

[0040] 1. Uploading a user's photo

[0041] Users visit a website or app, register, and log in. They then select 10 photos of their face on a face photo upload screen and submit them to the server, which verifies that the selected images are in the correct format and resolution before submitting them.

[0042] 2. Facial Photo Processing and Avatar Generation

[0043] The server applies a facial recognition algorithm to the received facial photo to extract facial features (the position and shape of the eyes, nose, and mouth). It then uses a deep learning model to generate an avatar for the user. This avatar is a digital virtual persona that faithfully reflects the user's facial features. The generated avatar is then saved in the user's profile.

[0044] 3. Select your destination and set it up

[0045] The user accesses the travel destination selection screen and selects the desired travel destination from a map interface or list. They also set the sending method (email, SNS, LINE, etc.) and the date and time of receipt. The selection information is sent to the server by pressing the send button, and the server records it in the user profile.

[0046] 4. Travel Photo Generation

[0047] The server retrieves a background image of the travel destination selected by the user from a database at the specified date and time. It then composites the user's avatar onto the background image, using a deep learning model to calculate natural placement and pose. The composite image is then generated as a realistic travel photo.

[0048] 5. Send travel photos

[0049] The server sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.). For email, an email generation API is used, and for SNS and LINE, the message is sent through the corresponding API. The travel photos are delivered at the specified time.

[0050] 6. Providing additional options

[0051] Users can select additional features on the options page. For example, they can select rich food photos, luxurious accommodations, or photos of themselves with celebrities. The selected options are saved directly under the user's control, and after payment confirmation, the corresponding image generation process begins. This provides users with an even richer experience.

[0052] Specific examples

[0053] For example, let's say a user uploads a photo of their face, sets their destination as "Eiffel Tower in Paris," and selects to receive the photo every morning at 8:00. The server will combine the background image of the Eiffel Tower with the user's avatar at the time specified by the user, and send the resulting photo by email every morning at 8:00. This allows the user to have a realistic experience every day, as if they were visiting Paris themselves.

[0054] In this way, by utilizing the system of the present invention, users can easily obtain photos that look as if they were actually traveling, and enjoy them on social media and other platforms.

[0055] The processing flow will be explained below.

[0056] Step 1:

[0057] The user accesses the website or app and logs in. After logging in, the user selects 10 photos of their face on the face photo upload screen.

[0058] Step 2:

[0059] The terminal collects the facial photo files selected by the user and uploads them to a server via a network.

[0060] Step 3:

[0061] The server receives the uploaded facial photo, checks the file format and resolution of the facial photo, and converts it into a format suitable for the facial recognition algorithm.

[0062] Step 4:

[0063] The server uses a facial recognition algorithm to extract facial features (position and shape of eyes, nose, and mouth).

[0064] Step 5:

[0065] The server uses a deep learning model (e.g., StyleGAN) to generate an avatar for the user based on the extracted features.

[0066] Step 6:

[0067] The server stores the generated avatar in the user profile.

[0068] Step 7:

[0069] The user accesses a destination selection screen and selects a desired destination from a map interface or list.

[0070] Step 8:

[0071] The user sets the sending method (email, SNS, LINE, etc.) and the reception time. Once the settings are complete, press the "Settings Complete" button.

[0072] Step 9:

[0073] The terminal collects information on the travel destination and transmission method selected by the user and transmits it to the server via the network.

[0074] Step 10:

[0075] The server processes the received configuration information and stores it in a user profile.

[0076] Step 11:

[0077] When the set date and time arrives, the server retrieves a background image of the travel destination specified by the user from the database.

[0078] Step 12:

[0079] The server then composites the user's avatar with the captured background image, using a deep learning model (e.g., OpenPose) to calculate natural placement and pose.

[0080] Step 13:

[0081] The server generates the composite image as the final travel photo.

[0082] Step 14:

[0083] The server formats and sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.).

[0084] Step 15:

[0085] The device receives notifications via email, SNS, or LINE and displays the image to the user.

[0086] Step 16:

[0087] Users can access an options page and select additional options such as photos of rich cuisine, photos of luxurious accommodations, or photos of themselves with celebrities.

[0088] Step 17:

[0089] The terminal collects the selected option information and billing information and transmits them to the server via the network.

[0090] Step 18:

[0091] The server processes the received option and billing information and confirms the user's purchase.

[0092] Step 19:

[0093] The server initiates an image generation process according to the additional options, and generates and transmits an image corresponding to the specified date and time.

[0094] By going through these steps, users can easily receive realistic travel photos that make them feel as if they are traveling around the world.

[0095] Example 1

[0096] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0097] Conventional image editing systems make it difficult for users to easily and effectively create realistic travel photos using their own facial images. They also lack features such as generating realistic images based on specific travel destinations or automatically sending photos at specified dates and times. As a result, users had to manually edit and send photos one by one, which was time-consuming and laborious, resulting in a poor user experience.

[0098] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0099] In this invention, the server includes means for receiving a user's image, means for generating a virtual character for the user based on the received image using a face recognition algorithm, means for acquiring background data based on a travel destination specified by the user, means for synthesizing the acquired background data with the generated virtual character to create a realistic image, and means for transmitting the created image in a manner specified by the user. This allows users to easily generate realistic travel photos using their own face photos, and the photos are automatically transmitted at a specified time, allowing them to be enjoyed on social media and other platforms without any hassle.

[0100] "User" refers to any individual or entity that uses the System.

[0101] "Image" refers to visual information stored or displayed in digital form.

[0102] A "facial recognition algorithm" is a computational method for detecting and identifying facial features in an image.

[0103] "Virtual character" refers to a digitally generated person based on the user's facial features.

[0104] "Travel destination" refers to a place that the user would like to virtually visit.

[0105] "Background data" refers to image and video information corresponding to a travel destination.

[0106] "Synthesis" refers to the process of integrating multiple images or data to generate a single image or data.

[0107] "Realistic images" refer to images that are synthesized to appear as if they exist in reality.

[0108] "Transmission" refers to the act of sending the generated image or data to an external party in a manner designated by the user.

[0109] "Additional selection menu" refers to optional functions and content provided in addition to basic functions.

[0110] "Setting" refers to the act of saving and making available information such as the user's selected travel destination and transmission method.

[0111] "Predetermined information" refers to specific information such as the travel destination, transmission method, and transmission date and time that the user has set in advance.

[0112] The system of the present invention generates a traveling avatar based on a facial photograph provided by the user, generates a composite image based on the travel destination selected by the user, and transmits the composite image in a specified manner. Specific embodiments of the invention are described below.

[0113] Uploading a user's face photo

[0114] A user accesses a website or app. They register an account and log in. They then select 10 photos of their face on the face photo upload screen and send them to the server. The image files must be in JPEG or PNG format with a resolution of at least 400x400 pixels. The server checks the format and resolution of the received images and returns an error message if they are inappropriate.

[0115] Facial photo processing and avatar generation

[0116] After receiving the uploaded face photo, the server uses a facial recognition algorithm (e.g., OpenCV) to extract facial features. Specifically, it identifies facial landmarks such as the position and shape of the eyes, nose, and mouth. It then uses a deep learning model (e.g., a person generation model using PyTorch) to generate a digital avatar based on the user's facial features. The generated avatar is then saved in the user's profile information.

[0117] Destination selection and settings

[0118] The user is taken to a travel destination selection screen. Here, they can select their desired travel destination from a map interface or a list. They can also set how they want to send their travel photos (e.g., email, SNS, LINE) and the date and time of receipt. The information selected by the user is sent to the server by pressing the send button. The server records this information in the user's profile.

[0119] Travel photo generation

[0120] At the date and time set by the user, the server retrieves a background image of the selected travel destination from the database. The server then places the user's avatar naturally in the retrieved background image and uses a deep learning model to calculate the pose. Shadows and lighting are also adjusted to ensure consistency with the background. The resulting composite image is a realistic travel photo.

[0121] Send travel photos

[0122] The server sends the generated travel photos using the method specified by the user. For example, for email, it uses an email generation API such as SendGrid, and for SNS or LINE, it uses the corresponding API (e.g. Facebook Graph API, LINE Messaging API). The travel photos are scheduled to be sent at the specified time.

[0123] Providing additional options

[0124] Users can select additional features on the options page. For example, they can choose photos of luxurious cuisine, images of accommodations, or photos of people with celebrities. Once a user selects an option, the server saves it in the user's profile and, once payment is confirmed, starts the corresponding image generation process. For example, if a photo of a person with a celebrity is selected, an image is retrieved from a celebrity image database and naturally combined with the user's avatar.

[0125] Specific examples and prompts

[0126] Specific operation example

[0127] For example, let's say a user uploads 10 photos of their face, selects "Eiffel Tower in Paris" as their destination, and wants to receive the photos every morning at 8:00 AM.

[0128] The server uses facial recognition algorithms to extract facial features and uses deep learning models to generate an avatar for the user.

[0129] The user selects "Eiffel Tower in Paris" and sets up the app to receive photos every morning at 8am.

[0130] At the set time, the server retrieves a background image of the Eiffel Tower from the database and synthesizes it onto the user's avatar.

[0131] The server will send the generated travel photos to the specified email address every morning at 8:00.

[0132] Prompt Sentence Examples

[0133] Example prompt 1:

[0134] "If a user uploads a photo of their face, sets their destination as 'Eiffel Tower in Paris', and selects to receive the photo every morning at 8am, what does the server do?"

[0135] Example output 1:

[0136] If a user uploads 10 photos of their face, sets their destination to "The Eiffel Tower in Paris," and selects to receive travel photos every morning at 8:00, the server will perform the following process. First, it extracts facial features from the user's photo and generates an avatar for the user using a deep learning model. Next, it retrieves a background image of the Eiffel Tower from the database at the specified date and time, and composites the user's avatar onto the background image. Finally, it sends the composite photo to the user by email every morning at 8:00.

[0137] This system can provide users with a realistic travel experience, which can be shared and enjoyed on social media. Specific hardware requirements include a high-performance image processing engine on the server side and GPU resources for the deep learning model. Software requirements include face recognition algorithms (e.g., OpenCV) and deep learning libraries (e.g., TENSORFLOW (registered trademark), PyTorch).

[0138] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0139] Step 1:

[0140] Users access the website or app, register an account, and log in. Next, they select 10 photos of their face on the face photo upload screen and send them to the server.

[0141] Input: 10 photos of your face (JPEG or PNG format, minimum 400x400 pixels)

[0142] Output: Image data uploaded to the server

[0143] Specific operation: Select 10 face photos from the user's device through file selection and send them to the server as an HTTP POST request.

[0144] Step 2:

[0145] The server applies a facial recognition algorithm (e.g., OpenCV) to the 10 facial photos received to extract facial features, specifically identifying the positions and shapes of the eyes, nose, mouth, etc.

[0146] Input: 10 user-uploaded face photos

[0147] Output: Extracted facial feature data (position and shape information)

[0148] Specific operation: A facial recognition algorithm is applied on the server side to extract feature points from each facial photograph, such as the facial contours, eye positions, nose positions, and mouth positions, and obtain these as coordinate data.

[0149] Step 3:

[0150] The server generates a virtual character (avatar) for the user using a deep learning model (e.g., a person generation model using PyTorch) based on the facial feature data. The generated avatar is saved in the user's profile information.

[0151] Input: Extracted facial feature data

[0152] Output: Digital virtual character (avatar)

[0153] How it works: Facial feature data is fed into the deep learning model as input, and a virtual character resembling the user is generated based on that data. The generated virtual character data is then saved in the user's profile information.

[0154] Step 4:

[0155] The user moves to the travel destination selection screen, selects the desired travel destination from the map interface or list, and sets the method of sending the travel photos (e.g., email, SNS, LINE) and the date and time of receipt.

[0156] Input: User-selected travel destination information, sending method, and date and time of receipt

[0157] Output: Configuration information sent to the server

[0158] Specific operation: Click on a travel destination on the map interface or select it from the list, then set the sending method and receiving date and time using the drop-down menu or calendar UI. Pressing the setting button sends this information to the server.

[0159] Step 5:

[0160] The server records the user's selection information (travel destination, transmission method, date and time of receipt) in the profile.

[0161] Input: Selections received from the user

[0162] Output: Recorded user settings information

[0163] What happens: The server stores the received selection data in the user profile database for future reference.

[0164] Step 6:

[0165] At the set date and time, the server retrieves the background image of the travel destination selected by the user from the database.

[0166] Input: User selection information, background image data in the database

[0167] Output: The background image obtained.

[0168] Specific operation: For example, search and retrieve background images of a specified travel destination from cloud storage such as Amazon S3.

[0169] Step 7:

[0170] The server uses deep learning models to calculate the pose of the virtual character, placing it naturally on the background image, and also adjusts the shadows and lighting to ensure consistency with the background.

[0171] Input: background image, virtual character

[0172] Output: Synthesized, realistic travel photos

[0173] How it works: It uses deep learning models to place virtual characters naturally against backgrounds and adjust shadows and light intensity to generate composite photos.

[0174] Step 8:

[0175] The server sends the generated travel photos using the method specified by the user (email, SNS, LINE, etc.).

[0176] Input: Composite travel photos, user's sending method settings

[0177] Output: Travel photos sent to the user

[0178] Specific operation: For email, the SendGrid API is used to send emails with travel photos attached, and for SNS and LINE, the corresponding API is used to send photos along with the message.

[0179] Step 9:

[0180] Users can select additional features on the options page, such as photos of luxurious food or photos of themselves with celebrities.

[0181] Input: Additional feature selection information

[0182] Output: Configuration information for additional features sent to the server

[0183] Specific operation: Select the items displayed on the options page using checkboxes or drop-down menus, press the Settings button, and send the selected options to the server. The selected options are recorded and managed on the server side.

[0184] (Application example 1)

[0185] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0186] Today, users enjoy virtual travel experiences through various means, but these technologies lack realism and do not fully satisfy users. Furthermore, it is difficult for users to easily and effectively create digital content that reflects their own faces and share it via social media, email, etc. This creates a demand for personalized, interactive experiences. Therefore, a system is needed that can generate a realistic avatar based on a user's facial photograph, naturally combine it with the background of a specified travel destination, and deliver it in a specified manner and at a specified time.

[0187] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0188] In this invention, the server includes means for receiving a user's facial photograph, means for generating a user avatar based on the received facial photograph using a facial recognition algorithm, means for acquiring a background image based on a travel destination specified by the user, means for synthesizing the acquired background image with the generated avatar to create a realistic travel photograph, means for calculating a natural position and pose of the synthesized image using a generative AI model to create a travel photograph, and means for transmitting the created travel photograph in a manner and at a time specified by the user. This allows users to have a personalized and interactive virtual travel experience in a highly realistic way and easily share it on various digital platforms.

[0189] "User" means an individual who uses the system to upload a photo of themselves and enjoy a virtual travel experience.

[0190] A "face photo" is a digital image of the user's face, and serves as the basic data for generating an avatar.

[0191] A "facial recognition algorithm" is a computational method for extracting facial features such as eyes, nose, and mouth from a received photograph of a face and generating a digital avatar.

[0192] An "avatar" is a digital virtual persona created based on a user's facial features using facial recognition algorithms.

[0193] "Travel destination" refers to information about a place that a user specifies as the destination of a virtual trip in the system.

[0194] A "background image" is a digital image of scenery or tourist spots at a travel destination designated by the user, which is combined with the avatar.

[0195] "Generative AI model" refers to the deep learning model used to calculate natural-looking placement and pose for synthesized images.

[0196] "Transmission means" refers to tools or APIs that allow users to send the generated travel photos in a manner specified by the user (e.g., email, SNS, LINE, etc.).

[0197] "Image compositing" is the process of naturally combining background images and avatars to create realistic travel photos.

[0198] "Transmission method" refers to the communication means selected by the user for transmitting travel photos.

[0199] "Realism" refers to the quality of the synthesized images, making them appear as if they were taken during an actual trip.

[0200] This invention is a system that generates an avatar based on a user's facial photograph, combines it with a background image of a travel destination selected by the user to generate a realistic travel photograph, and transmits it in a specified manner. This system is implemented in the following steps.

[0201] First, a user accesses the website or application from a smartphone or PC, registers, and logs in. Next, the user selects 10 facial photos on the face photo upload screen and submits them to the server. These facial photos are verified to be in the correct format and resolution.

[0202] The server processes the received facial photo using a facial recognition algorithm (for example, using a deep learning library such as OpenCV or TensorFlow) to extract facial features (eyes, nose, mouth, etc.). Based on the extracted facial features, a deep learning model is used to generate an avatar for the user. This avatar is then stored in the user's profile.

[0203] Next, the user accesses the travel destination selection screen and selects their desired travel destination. Destination selection can be done using a map interface or a list of options. The user also selects the method of sending the generated travel photos (email, SNS, LINE, etc.) and the date and time of sending. This selection information is sent to the server and recorded in the user profile.

[0204] At the set date and time, the server retrieves a background image of the selected travel destination from the database. The generated avatar and background image are then composited using a deep learning model to create a natural position and pose. This process uses a generative AI model, and the composite image is then generated as a realistic travel photo.

[0205] Next, the realistic travel photos are sent in the way the user specified. For example, in the case of email, an email generation API is used, and in the case of SNS or LINE, messages are sent through the respective APIs. The travel photos are delivered at the specified time.

[0206] Additionally, users can select additional options, such as images of specific local items or meals, which are also saved in the user profile and the image generation process begins after payment confirmation.

[0207] As a concrete example, consider a case where a user selects the Eiffel Tower in Paris as a travel destination and requests to receive travel photos by email every morning at 8:00. In this case, the server will synthesize the background image of the Eiffel Tower with the user's avatar at the set time and send the travel photos by email every morning at 8:00. Throughout this process, a generative AI model is used to ensure natural synthesis and placement.

[0208] Prompt Sentence Examples

[0209] "Generate an avatar from a photo of the user's face and overlay it with a travel photo of the Eiffel Tower in the background. Background image: path / to / paris_eiffel.jpg Facial features: {Eye position: (x1, y1), Nose position: (x2, y2), Mouth shape: (x3, y3)}"

[0210] In this way, users can have a realistic virtual travel experience and easily share their travel photos on various digital platforms.

[0211] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0212] Step 1:

[0213] A user accesses a website or application using a smartphone or PC. First, the user registers and logs in to the site. Next, the user selects 10 facial photos on the facial photo upload screen and sends them to the server. The input is the 10 facial photos, and the output is the photo data stored on the server.

[0214] Step 2:

[0215] The server applies a facial recognition algorithm to the received facial photo. Specifically, it uses deep learning libraries such as OpenCV and TensorFlow to extract facial features (the position and shape of the eyes, nose, and mouth). The input is the facial photo data, and the output is data that quantifies the facial features.

[0216] Step 3:

[0217] The server uses a generative AI model to generate a digital avatar for the user based on the facial feature data extracted by the facial recognition algorithm. This avatar is virtual person data that faithfully reflects the user's facial features. The input is facial feature data, and the output is a digital avatar.

[0218] Step 4:

[0219] The user accesses the travel destination selection screen and selects the desired travel destination from a map interface or list. They also specify the method of sending the travel photos (email, SNS, LINE, etc.) and the date and time of receipt. The input is the user's selected travel destination and sending method information, and the output is the selected information recorded on the server.

[0220] Step 5:

[0221] At the set date and time, the server retrieves the background image of the travel destination selected by the user from the database. It then composites the background image with the user's avatar, using a generative AI model to calculate natural placement and poses to create realistic travel photos. The input is the background image of the travel destination and the user's digital avatar, and the output is the composite travel photo.

[0222] Step 6:

[0223] The server sends the generated travel photos in the way specified by the user. For example, in the case of email, it uses the email generation API, and in the case of SNS or LINE, it sends through the respective API. The input is the travel photo data and sending setting information, and the output is the travel photo sent to the user.

[0224] Step 7:

[0225] When a user selects an add-on feature, the server executes the corresponding image generation process, for example, adding an image of a specific local item or meal to a travel photo. This process is also performed using a generative AI model. The input is the selection information for the add-on option, and the output is the travel photo with the option added.

[0226] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0227] The system of this invention generates a traveling avatar based on a facial photo provided by the user, generates a composite image based on the travel destination selected by the user, and transmits it in a specified manner. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to suggest and select travel destinations and optional images based on the user's emotional state. Specific embodiments of the invention are described below.

[0228] 1. Uploading a user's photo

[0229] The user accesses the website or app and logs in. After logging in, the user selects 10 facial photos on the face photo upload screen and sends them to the server.

[0230] 2. Facial Photo Processing and Avatar Generation

[0231] The server applies a facial recognition algorithm to the received facial photo to extract facial features (the position and shape of the eyes, nose, and mouth), then uses a deep learning model to generate an avatar for the user, a digital virtual persona that faithfully reflects the user's facial features, and the generated avatar is saved in the user's profile.

[0232] 3. Emotion Recognition and Analysis

[0233] In addition to a photo of their face, the user provides the emotion engine with a photo of their facial expression and voice data. The emotion engine then analyzes the user's emotional state based on the data provided by the user. The analysis results are classified into emotion categories such as "joy," "sadness," "surprise," and "anger."

[0234] 4. Select your destination and set it up

[0235] Users access the travel destination selection screen and select their desired travel destination from a map interface or list. If the user requests automatic suggestions based on their emotional state, the system will suggest the most suitable travel destination based on the analysis results of the emotion engine. Users also set the sending method (email, SNS, LINE, etc.) and reception time, and press the "Settings Complete" button once the settings are complete.

[0236] 5. Travel Photo Generation

[0237] At the set date and time, the server retrieves a background image of the user's travel destination from the database. The server then composites the user's avatar onto the background image, using a deep learning model to calculate natural placement and pose. The composite image is then generated as a realistic travel photo.

[0238] 6. Send travel photos

[0239] The server sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.). For email, an email generation API is used, and for SNS or LINE, a message is sent through the corresponding API. The travel photos are delivered at the specified time.

[0240] 7. Providing additional options

[0241] Users can select additional features on the options page, such as photos of rich cuisine, luxurious accommodations, or photos of people with celebrities. The selected options are saved directly under the user's control, and the corresponding image generation process begins after payment is confirmed.

[0242] Specific examples

[0243] For example, let's consider a case where a user uploads a photo of their face and the emotion engine recognizes the emotion "joy." Based on the analysis results of the emotion engine, the server suggests "The Eiffel Tower in Paris" as the best travel destination. If the user accepts this suggestion and sets the sending method to email and the receiving time to 8:00 a.m. every morning, the server will combine a background image of the Eiffel Tower with the user's avatar, generating and sending a realistic travel photo every morning at 8:00 a.m. If the user selects a photo of "rich French cuisine" as an additional option, that image will also be sent.

[0244] In this way, by utilizing the system of the present invention, users can easily obtain realistic travel photos based on their emotional state, which can be enjoyed on social media and other platforms.

[0245] The processing flow will be explained below.

[0246] Step 1:

[0247] The user accesses the website or app and logs in. After logging in, the user selects 10 facial photos on the face photo upload screen, as well as facial expression photos and voice data for recognizing their emotional state, and sends these to the server.

[0248] Step 2:

[0249] The device collects the facial photo file and emotion recognition data selected by the user and uploads them to a server via the network.

[0250] Step 3:

[0251] The server receives the uploaded facial photo and emotion recognition data, checks the file format and resolution of the facial photo, and converts it into a format suitable for the facial recognition algorithm.

[0252] Step 4:

[0253] The server uses a facial recognition algorithm to extract facial features (position and shape of eyes, nose, and mouth).

[0254] Step 5:

[0255] The server uses a deep learning model (e.g., StyleGAN) to generate an avatar for the user based on the extracted features.

[0256] Step 6:

[0257] The server stores the generated avatar in the user profile.

[0258] Step 7:

[0259] The server uses an emotion engine to analyze the user's emotional state based on facial expressions and voice data, and classifies the results into emotion categories such as "happiness," "sadness," "surprise," and "anger."

[0260] Step 8:

[0261] The server will suggest the most suitable travel destination for the user based on the analysis results of the emotion engine. For example, if the emotion of "joy" is detected, the server will suggest the Eiffel Tower in Paris.

[0262] Step 9:

[0263] The user accesses the travel destination selection screen and selects the suggested travel destination or their own desired travel destination. They also set the sending method (email, SNS, LINE, etc.) and reception time. Once the settings are complete, they press the "Settings Complete" button.

[0264] Step 10:

[0265] The terminal collects information on the travel destination and transmission method selected by the user and transmits it to the server via the network.

[0266] Step 11:

[0267] The server processes the received configuration information and stores it in a user profile.

[0268] Step 12:

[0269] When the set date and time arrives, the server retrieves a background image of the travel destination specified by the user from the database.

[0270] Step 13:

[0271] The server then composites the user's avatar with the captured background image, using a deep learning model (e.g., OpenPose) to calculate natural placement and pose.

[0272] Step 14:

[0273] The server generates the composite image as the final travel photo.

[0274] Step 15:

[0275] The server sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.).

[0276] Step 16:

[0277] The device receives notifications via email, SNS, or LINE and displays the image to the user.

[0278] Step 17:

[0279] Users can access an options page and select additional options such as photos of rich cuisine, photos of luxurious accommodations, or photos of themselves with celebrities.

[0280] Step 18:

[0281] The terminal collects the selected option information and billing information and transmits them to the server via the network.

[0282] Step 19:

[0283] The server processes the received option and billing information and confirms the user's purchase.

[0284] Step 20:

[0285] The server initiates an image generation process according to the additional options, and generates and transmits an image corresponding to the specified date and time.

[0286] By going through these specific steps, users can receive realistic travel photos based on their own emotional state, which they can enjoy on social media and other platforms.

[0287] Example 2

[0288] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0289] In a system for generating realistic travel photos, there is a need for a means to reflect the user's emotional state and efficiently and effectively set and save travel destination suggestions and subsequent transmission methods.In addition, it has been pointed out that conventional systems need not only to synthesize background images and avatars based on the user's selection, but also to generate optional images to provide further enjoyment.

[0290] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving a user's facial photograph, means for generating a user's avatar based on the received facial photograph using a facial recognition algorithm, means for acquiring a background image based on a travel destination designated by the user, means for creating a realistic travel photograph by combining the acquired background image with the generated avatar, means for transmitting the created travel photograph in a manner designated by the user, means for analyzing and recognizing the user's emotional state and suggesting a travel destination based on the analysis results, and means for setting and saving a transmission method and transmission time. This makes it possible to suggest travel destinations and generate optional images that reflect the user's emotional state, and further to transmit the travel photograph in a manner designated by the user at a set date and time.

[0291] A "user" is an individual who uses the system to upload a photo of themselves and select a travel destination and optional images.

[0292] The "server" is a computer system that receives a user's facial photograph, generates an avatar using a facial recognition algorithm, obtains background images of the travel destination, and analyzes the user's emotional state.

[0293] A "face recognition algorithm" is a technology for extracting the position and shape of the eyes, nose, and mouth from a facial photo provided by the user.

[0294] An "avatar" is a digital virtual persona that reflects a user's facial features and is generated based on facial recognition algorithms.

[0295] A "background image" is an image of scenery or famous places at a travel destination designated by the user.

[0296] "Compositing" is the process of combining a background image with a generated avatar to create a realistic travel photo.

[0297] "Emotional state" refers to psychological states such as "joy," "sadness," "surprise," and "anger" that are analyzed from facial expression photos and voice data provided by the user.

[0298] "Suggestion" refers to the act of the server suggesting the most suitable travel destination to the user based on the analysis results of the emotion engine.

[0299] "Optional images" are images added by user selection, including images of travel destinations, specific items, and meals.

[0300] "Sending method" refers to the means by which the generated travel photos are delivered to the user, and includes email, SNS, LINE, etc.

[0301] "Send time" is the designated time for sending the generated travel photos and optional images to the user.

[0302] The system of this invention generates a traveling avatar based on a facial photograph provided by the user, suggests optimal travel destinations based on the user's emotional state, generates a composite image, and transmits it in a specified manner. Specific embodiments are described below.

[0303] First, the user accesses the website or app, logs in, selects 10 photos of their face on the face photo upload screen, and sends them to the server. At this point, the user can upload the photos from the app using their smartphone.

[0304] The server reviews the received facial photo and applies facial recognition algorithms such as OpenCV or Dlib to extract facial features. This facial recognition process specifically analyzes the position and shape of the eyes, nose, and mouth in the photo. It then uses a deep learning model such as GAN to generate an avatar that reflects the user's facial features. This avatar is then saved in the user's profile for further processing.

[0305] The user also provides the server with a photo of their facial expression and voice data. The server's emotion engine analyzes the user's emotional state based on this data. This emotion analysis uses emotion analysis APIs from Microsoft® Azure® or Google® Cloud, for example. Based on the analysis results, the user's emotional state is classified into emotion categories such as "joy," "sadness," "surprise," and "anger."

[0306] Next, the user can access the travel destination selection screen and select their desired travel destination. The system automatically suggests the most suitable travel destination based on the analysis results of the emotion engine. For example, if the emotional state is "joy," the system will suggest a travel destination that matches the emotion, such as Paris, France. The user can accept the suggested travel destination or choose a different one by themselves. At this time, the user sets the sending method (email, SNS, LINE, etc.) and reception time, and presses the "Settings Complete" button. These settings are then saved on the server.

[0307] At the set date and time, the server retrieves a background image of the travel destination from a database. This background image is taken from a common photo library such as Google Images. The server then uses a deep learning model to naturally blend the user's avatar into the background image, generating a realistic travel photo. Pose Estimation technology is used to properly position the avatar against the background image.

[0308] Finally, the server sends the generated travel photos via the method specified by the user: via email using the SendGrid API, or via the corresponding API for social media or LINE. The generated images are sent at the time specified by the user.

[0309] Users can also select additional features on the options page, such as photos of rich cuisine, luxurious accommodations, and photos of two people with celebrities. After the user confirms their selection and payment, the image generation process will begin and the generated optional images will be sent.

[0310] As a concrete example, if a user uploads a photo of their face and the emotion engine recognizes the emotion "joy," the system will suggest "The Eiffel Tower in Paris" as a travel destination, and the user will accept this suggestion and set the reception time to 8:00 a.m. every morning. The server will then combine the background image of the Eiffel Tower with the user's avatar, generate a realistic travel photo every morning, and send it by email. If the user selects a photo of "rich French cuisine" as an additional option, that image will also be sent.

[0311] Example prompt sentence:

[0312] The following prompts explain system behavior to the generative AI model:

[0313] Your task is to generate a program for a system that, based on 10 facial photos provided by the user, will combine the avatar with a background image of the travel destination selected by the user when the emotion engine recognizes the emotion "joy." The process will be as follows:

[0314] 1. Extract facial features using face recognition algorithm (OpenCV / Dlib).

[0315] 2. Generate a user avatar using GAN.

[0316] 3. Analyze user emotions using an emotion engine (Microsoft Azure / Google Cloud).

[0317] 4. Composite avatar and background image based on suggested travel destination (e.g., Eiffel Tower in Paris).

[0318] 5. Using the email sending API (SendGrid), the composite photo is sent every morning at 8:00.

[0319] In this way, by using the system of the present invention, realistic travel photos based on the user's emotional state can be easily provided.

[0320] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0321] Step 1:

[0322] A user accesses a website or app, logs in, and proceeds to the face photo upload screen. The user selects 10 face photos from the photo gallery and presses the upload button. This operation sends the face photos to the server. The input is the user's face photo, and the output is the face photo data sent to the server.

[0323] Step 2:

[0324] The server checks the received face photo and applies a facial recognition algorithm (OpenCV or Dlib) to extract facial features (the position and shape of the eyes, nose, and mouth). The algorithm converts the image data to grayscale and creates a list of detected feature points. The input is the face photo data, and the output is the extracted facial feature data.

[0325] Step 3:

[0326] The server generates a user avatar using a GAN (generative adversarial network). Using facial feature data as input, the GAN model generates a virtual avatar image based on the learning results. The generated avatar is saved in the user profile. The input is facial feature data, and the output is the generated avatar image.

[0327] Step 4:

[0328] The user then provides the server with photographs of facial expressions and voice data. The server's emotion engine then uses this data to analyze the user's emotional state. Emotion analysis APIs from Microsoft Azure and Google Cloud are used for the emotion analysis. The analysis results are categorized into "happiness," "sadness," "surprise," and "anger" and stored on the server. The input is photographs of facial expressions and voice data, and the output is analyzed emotional state data.

[0329] Step 5:

[0330] The user can access the travel destination selection screen and select their desired travel destination. The server will suggest the most suitable travel destination based on the analysis results of the emotion engine. After the user accepts the suggested travel destination or selects another one by themselves, they set the sending method (email, SNS, LINE, etc.) and reception time. These settings are saved on the server. The input is the user's emotional state and travel destination selection data, and the output is the saved setting data.

[0331] Step 6:

[0332] At the set date and time, the server retrieves a background image of the travel destination from a database. This background image can be from Google Images or a user's own photo library. The server then uses a deep learning model (using Pose Estimation technology) to naturally blend the user's avatar into the background image. The input is the background image and the avatar image, and the output is the blended travel photo.

[0333] Step 7:

[0334] The server sends the generated travel photos in the manner specified by the user. The SendGrid API is used for email transmission, and the corresponding APIs for SNS and LINE are used. The travel photos are sent to the user at the specified time. The input is the composited travel photos and transmission setting data, and the output is the sent travel photos.

[0335] Step 8:

[0336] Users can select additional features on the options page. These include photos of rich cuisine, luxurious accommodations, and photos of people with celebrities. After the user selects these options and confirms the payment, the image generation process begins. The input is the option selection data, and the output is the generated option image.

[0337] In this way, each processing step operates in cooperation with one another, thereby realizing a system that provides realistic travel photos based on the user's emotional state.

[0338] (Application example 2)

[0339] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0340] Conventional avatar generation and travel photo creation systems lack the ability to suggest travel destinations that take the user's emotional state into account, resulting in the inability to provide the optimal travel experience for the user. Furthermore, simply synthesized photos alone are unlikely to achieve the realism and personalized satisfaction that users expect. Therefore, a new system that reflects the user's emotions and provides a more personalized travel experience is needed.

[0341] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0342] In this invention, the server includes means for receiving a user's facial photograph, means for generating a user avatar based on the received facial photograph using a facial recognition algorithm, means for acquiring a background image based on a travel destination designated by the user, means for synthesizing the acquired background image with the generated avatar to create a realistic travel photograph, means for transmitting the created travel photograph in a manner designated by the user, means for analyzing the user's emotional state using an emotion recognition engine, and means for suggesting travel destinations based on the user's emotional state. This makes it possible to suggest optimal travel destinations taking the user's emotional state into consideration and create realistic travel photographs.

[0343] "User" refers to an individual who uses this system.

[0344] "Facial photo" refers to image data provided by a user that is taken of the user's own face using a camera or the like.

[0345] A "facial recognition algorithm" refers to a computational method that extracts features such as the eyes, nose, and mouth from a photograph of a face and identifies and recognizes a specific face.

[0346] An "avatar" refers to a virtual digital persona generated based on a photograph of a user's face.

[0347] "Travel destination" refers to a destination where a user wishes to take a virtual trip.

[0348] "Background image" refers to image data of scenery and famous places at travel destinations.

[0349] "Compositing" refers to the process of overlaying a user's avatar with a background image to generate a single image.

[0350] "Realistic travel photos" refer to images that are synthesized to make the user's avatar appear as if they are actually at the travel destination.

[0351] An "emotion recognition engine" refers to a system that analyzes and classifies a user's emotional state based on a photo of their face and voice data.

[0352] "Emotional state" refers to a user's current psychological state or mood.

[0353] "Suggestion" refers to the system selecting and presenting appropriate travel destinations and options to the user.

[0354] "Sending" refers to the process of delivering the generated travel photos via the communication means selected by the user.

[0355] This system generates a traveling avatar using a user's facial photograph, and uses an emotion recognition engine to suggest optimal travel destinations based on the user's emotional state. It also creates realistic travel photos by combining the user's avatar with background images of the travel destination, and sends them in a manner specified by the user, providing a virtual travel experience.

[0356] The system mainly operates in the following steps:

[0357] 1. Uploading a user's photo

[0358] Users access a dedicated application using a smartphone or PC and log in. After logging in, they select multiple facial photos on the facial photo upload screen and send them to the server.

[0359] 2. Facial Photo Processing and Avatar Generation

[0360] The server applies a facial recognition algorithm to the received face photo to extract facial features (the position and shape of the eyes, nose, and mouth). It then uses a generative AI model to generate an avatar for the user. Leveraging a deep learning model, a digital persona is created that faithfully reflects the user's facial features, and this avatar is stored in the user's profile.

[0361] 3. Emotion Recognition and Analysis

[0362] Users provide facial photos, facial expression photos, and voice data to the emotion recognition engine, which uses technologies such as DeepFace to analyze the user's emotional state and classify the results into categories such as "happiness," "sadness," "surprise," and "anger."

[0363] 4. Select your destination and set it up

[0364] Users can select their desired travel destination from a map or list through the system interface. They can also choose destinations suggested by the emotion recognition engine. If they like the suggested destination, they confirm the setting. They can also set the sending method (email, SNS, etc.) and the reception time.

[0365] 5. Travel Photo Generation

[0366] At the set date and time, the server retrieves a background image of the travel destination specified by the user from the database, then combines the generated avatar with the background image, calculating natural positioning and poses to generate a realistic travel photo.

[0367] 6. Send travel photos

[0368] The server then sends the generated travel photos via the method specified by the user: via email using a dedicated email generation API, or via the corresponding API for social media or text messages.

[0369] 7. Providing additional options

[0370] Users can select additional features on the options page within the system, such as "rich food photos" or "luxury accommodations." The options selected by the user are saved in their profile, and the corresponding image generation process is initiated after payment confirmation.

[0371] Specific examples

[0372] For example, if a user provides a photo of a "joyed" expression to the emotion recognition engine, the system will suggest "The Eiffel Tower in Paris" as the best travel destination. If the user accepts this suggestion and sets the sending method to email and the receiving time to 8:00 a.m., the server will combine the background image of the Eiffel Tower with the user's avatar, generate a realistic travel photo, and send it every morning at 8:00 a.m. If the user also selects a photo of "rich French cuisine" as an additional option, that will also be sent.

[0373] Prompt Sentence Examples

[0374] "Create realistic travel photos by combining user-uploaded photos of your face with a background image of the Eiffel Tower in Paris."

[0375] The system allows users to virtually enjoy an optimal travel experience based on their emotional state and receive personalized content tailored to their individual needs.

[0376] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0377] Step 1:

[0378] The user logs in to a dedicated application using a smartphone or computer.

[0379] Input: Login information (user ID, password)

[0380] Processing: User authentication is performed and the user profile is obtained.

[0381] Output: A successful authentication message, and the user's profile information.

[0382] Step 2:

[0383] The user selects multiple facial photos on the facial photo upload screen and sends them to the server.

[0384] Input: Multiple face photos (e.g. 10 photos)

[0385] Processing: Your face photo is uploaded to the server and temporarily stored.

[0386] Output: Face photo upload complete message.

[0387] Step 3:

[0388] The server applies a facial recognition algorithm to the received facial photo and extracts facial features.

[0389] Input: Multiple face photos

[0390] Processing: Use libraries such as OpenCV to extract facial features such as eyes, nose, and mouth.

[0391] Output: Extracted facial feature data.

[0392] Step 4:

[0393] The server uses a deep learning model to generate an avatar for the user.

[0394] Input: Extracted facial feature data

[0395] Processing: Generate an avatar based on the user's face using a generative AI model (e.g., GAN).

[0396] Output: Generated avatar data.

[0397] Step 5:

[0398] The user provides facial expression photos and voice data to the emotion recognition engine.

[0399] Input: facial expression photos and audio data

[0400] Processing: Analyze the user's emotional state using emotion recognition engines such as DeepFace.

[0401] Output: Classification of emotional state (e.g., happy, sad, surprised, angry, etc.).

[0402] Step 6:

[0403] The server suggests optimal travel destinations based on the user's emotional state.

[0404] Input: Emotional state classification result

[0405] Processing: Randomly or algorithmically select the best travel destination from a list of travel destinations that correspond to the emotional state.

[0406] Output: Suggested travel destinations.

[0407] Step 7:

[0408] The user selects a desired travel destination on the travel destination selection screen and sets the transmission method and reception time.

[0409] Input: Travel destination selection information, sending method, receiving time

[0410] Action: Save the user's selections and settings to the database.

[0411] Output: Configuration complete message.

[0412] Step 8:

[0413] When the set date and time arrives, the server retrieves a background image of the travel destination specified by the user from the database.

[0414] Input: Travel destination selection information

[0415] Processing: Search and retrieve background images corresponding to the specified travel destination from the background image database.

[0416] Output: The obtained background image.

[0417] Step 9:

[0418] The server combines the generated avatar with the acquired background image.

[0419] Input: Avatar data, background image

[0420] Processing: Uses deep learning models to calculate natural placement and pose, and then synthesizes the avatar with the background image.

[0421] Output: Realistic travel photos.

[0422] Step 10:

[0423] The server transmits the generated travel photos in a manner designated by the user.

[0424] Input: Travel photos, sending method information

[0425] Processing: Send travel photos in the specified way using email generation API or SNS sending API.

[0426] Output: A transmission completion message.

[0427] Step 11:

[0428] The user selects additional features on the options page, and after confirming payment, the corresponding image generation process begins.

[0429] Input: Additional option selection information, billing information

[0430] Processing: Performs additional image generation processing based on selected options and saves it to the user profile.

[0431] Output: The additional image generated, and a charge completion message.

[0432] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0433] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0434] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0435] [Second embodiment]

[0436] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0437] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0438] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0439] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0440] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0441] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0442] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0443] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0444] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0445] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0446] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0447] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0448] The system of the present invention generates a traveling avatar based on a facial photograph provided by the user, generates a composite image based on a travel destination selected by the user, and transmits the composite image in a specified manner. Specific embodiments of the present invention are described below.

[0449] 1. Uploading a user's photo

[0450] Users visit a website or app, register, and log in. They then select 10 photos of their face on a face photo upload screen and submit them to the server, which verifies that the selected images are in the correct format and resolution before submitting them.

[0451] 2. Facial Photo Processing and Avatar Generation

[0452] The server applies a facial recognition algorithm to the received facial photo to extract facial features (the position and shape of the eyes, nose, and mouth). It then uses a deep learning model to generate an avatar for the user. This avatar is a digital virtual persona that faithfully reflects the user's facial features. The generated avatar is then saved in the user's profile.

[0453] 3. Select your destination and set it up

[0454] The user accesses the travel destination selection screen and selects the desired travel destination from a map interface or list. They also set the sending method (email, SNS, LINE, etc.) and the date and time of receipt. The selection information is sent to the server by pressing the send button, and the server records it in the user profile.

[0455] 4. Travel Photo Generation

[0456] The server retrieves a background image of the travel destination selected by the user from a database at the specified date and time. It then composites the user's avatar onto the background image, using a deep learning model to calculate natural placement and pose. The composite image is then generated as a realistic travel photo.

[0457] 5. Send travel photos

[0458] The server sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.). For email, an email generation API is used, and for SNS and LINE, the message is sent through the corresponding API. The travel photos are delivered at the specified time.

[0459] 6. Providing additional options

[0460] Users can select additional features on the options page. For example, they can select rich food photos, luxurious accommodations, or photos of themselves with celebrities. The selected options are saved directly under the user's control, and after payment confirmation, the corresponding image generation process begins. This provides users with an even richer experience.

[0461] Specific examples

[0462] For example, let's say a user uploads a photo of their face, sets their destination as "Eiffel Tower in Paris," and selects to receive the photo every morning at 8:00. The server will combine the background image of the Eiffel Tower with the user's avatar at the time specified by the user, and send the resulting photo by email every morning at 8:00. This allows the user to have a realistic experience every day, as if they were visiting Paris themselves.

[0463] In this way, by utilizing the system of the present invention, users can easily obtain photos that look as if they were actually traveling, and enjoy them on social media and other platforms.

[0464] The processing flow will be explained below.

[0465] Step 1:

[0466] The user accesses the website or app and logs in. After logging in, the user selects 10 photos of their face on the face photo upload screen.

[0467] Step 2:

[0468] The terminal collects the facial photo files selected by the user and uploads them to a server via a network.

[0469] Step 3:

[0470] The server receives the uploaded facial photo, checks the file format and resolution of the facial photo, and converts it into a format suitable for the facial recognition algorithm.

[0471] Step 4:

[0472] The server uses a facial recognition algorithm to extract facial features (position and shape of eyes, nose, and mouth).

[0473] Step 5:

[0474] The server uses a deep learning model (e.g., StyleGAN) to generate an avatar for the user based on the extracted features.

[0475] Step 6:

[0476] The server stores the generated avatar in the user profile.

[0477] Step 7:

[0478] The user accesses a destination selection screen and selects a desired destination from a map interface or list.

[0479] Step 8:

[0480] The user sets the sending method (email, SNS, LINE, etc.) and the reception time. Once the settings are complete, press the "Settings Complete" button.

[0481] Step 9:

[0482] The terminal collects information on the travel destination and transmission method selected by the user and transmits it to the server via the network.

[0483] Step 10:

[0484] The server processes the received configuration information and stores it in a user profile.

[0485] Step 11:

[0486] When the set date and time arrives, the server retrieves a background image of the travel destination specified by the user from the database.

[0487] Step 12:

[0488] The server then composites the user's avatar with the captured background image, using a deep learning model (e.g., OpenPose) to calculate natural placement and pose.

[0489] Step 13:

[0490] The server generates the composite image as the final travel photo.

[0491] Step 14:

[0492] The server formats and sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.).

[0493] Step 15:

[0494] The device receives notifications via email, SNS, or LINE and displays the image to the user.

[0495] Step 16:

[0496] Users can access an options page and select additional options such as photos of rich cuisine, photos of luxurious accommodations, or photos of themselves with celebrities.

[0497] Step 17:

[0498] The terminal collects the selected option information and billing information and transmits them to the server via the network.

[0499] Step 18:

[0500] The server processes the received option and billing information and confirms the user's purchase.

[0501] Step 19:

[0502] The server initiates an image generation process according to the additional options, and generates and transmits an image corresponding to the specified date and time.

[0503] By going through these steps, users can easily receive realistic travel photos that make them feel as if they are traveling around the world.

[0504] Example 1

[0505] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0506] Conventional image editing systems make it difficult for users to easily and effectively create realistic travel photos using their own facial images. They also lack features such as generating realistic images based on specific travel destinations or automatically sending photos at specified dates and times. As a result, users had to manually edit and send photos one by one, which was time-consuming and laborious, resulting in a poor user experience.

[0507] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0508] In this invention, the server includes means for receiving a user's image, means for generating a virtual character for the user based on the received image using a face recognition algorithm, means for acquiring background data based on a travel destination specified by the user, means for synthesizing the acquired background data with the generated virtual character to create a realistic image, and means for transmitting the created image in a manner specified by the user. This allows users to easily generate realistic travel photos using their own face photos, and the photos are automatically transmitted at a specified time, allowing them to be enjoyed on social media and other platforms without any hassle.

[0509] "User" refers to any individual or entity that uses the System.

[0510] "Image" refers to visual information stored or displayed in digital form.

[0511] A "facial recognition algorithm" is a computational method for detecting and identifying facial features in an image.

[0512] "Virtual character" refers to a digitally generated person based on the user's facial features.

[0513] "Travel destination" refers to a place that the user would like to virtually visit.

[0514] "Background data" refers to image and video information corresponding to a travel destination.

[0515] "Synthesis" refers to the process of integrating multiple images or data to generate a single image or data.

[0516] "Realistic images" refer to images that are synthesized to appear as if they exist in reality.

[0517] "Transmission" refers to the act of sending the generated image or data to an external party in a manner designated by the user.

[0518] "Additional selection menu" refers to optional functions and content provided in addition to basic functions.

[0519] "Setting" refers to the act of saving and making available information such as the user's selected travel destination and transmission method.

[0520] "Predetermined information" refers to specific information such as the travel destination, transmission method, and transmission date and time that the user has set in advance.

[0521] The system of the present invention generates a traveling avatar based on a facial photograph provided by the user, generates a composite image based on the travel destination selected by the user, and transmits the composite image in a specified manner. Specific embodiments of the invention are described below.

[0522] Uploading a user's face photo

[0523] A user accesses a website or app. They register an account and log in. They then select 10 photos of their face on the face photo upload screen and send them to the server. The image files must be in JPEG or PNG format with a resolution of at least 400x400 pixels. The server checks the format and resolution of the received images and returns an error message if they are inappropriate.

[0524] Facial photo processing and avatar generation

[0525] After receiving the uploaded face photo, the server uses a facial recognition algorithm (e.g., OpenCV) to extract facial features. Specifically, it identifies facial landmarks such as the position and shape of the eyes, nose, and mouth. It then uses a deep learning model (e.g., a person generation model using PyTorch) to generate a digital avatar based on the user's facial features. The generated avatar is then saved in the user's profile information.

[0526] Destination selection and settings

[0527] The user is taken to a travel destination selection screen. Here, they can select their desired travel destination from a map interface or a list. They can also set how they want to send their travel photos (e.g., email, SNS, LINE) and the date and time of receipt. The information selected by the user is sent to the server by pressing the send button. The server records this information in the user's profile.

[0528] Travel photo generation

[0529] At the date and time set by the user, the server retrieves a background image of the selected travel destination from the database. The server then places the user's avatar naturally in the retrieved background image and uses a deep learning model to calculate the pose. Shadows and lighting are also adjusted to ensure consistency with the background. The resulting composite image is a realistic travel photo.

[0530] Send travel photos

[0531] The server sends the generated travel photos using the method specified by the user. For example, for email, it uses an email generation API such as SendGrid, and for SNS or LINE, it uses the corresponding API (e.g. Facebook Graph API, LINE Messaging API). The travel photos are scheduled to be sent at the specified time.

[0532] Providing additional options

[0533] Users can select additional features on the options page. For example, they can choose photos of luxurious cuisine, images of accommodations, or photos of people with celebrities. Once a user selects an option, the server saves it in the user's profile and, once payment is confirmed, starts the corresponding image generation process. For example, if a photo of a person with a celebrity is selected, an image is retrieved from a celebrity image database and naturally combined with the user's avatar.

[0534] Specific examples and prompts

[0535] Specific operation example

[0536] For example, let's say a user uploads 10 photos of their face, selects "Eiffel Tower in Paris" as their destination, and wants to receive the photos every morning at 8:00 AM.

[0537] The server uses facial recognition algorithms to extract facial features and uses deep learning models to generate an avatar for the user.

[0538] The user selects "Eiffel Tower in Paris" and sets up the app to receive photos every morning at 8am.

[0539] At the set time, the server retrieves a background image of the Eiffel Tower from the database and synthesizes it onto the user's avatar.

[0540] The server will send the generated travel photos to the specified email address every morning at 8:00.

[0541] Prompt Sentence Examples

[0542] Example prompt 1:

[0543] "If a user uploads a photo of their face, sets their destination as 'Eiffel Tower in Paris', and selects to receive the photo every morning at 8am, what does the server do?"

[0544] Example output 1:

[0545] If a user uploads 10 photos of their face, sets their destination to "The Eiffel Tower in Paris," and selects to receive travel photos every morning at 8:00, the server will perform the following process. First, it extracts facial features from the user's photo and generates an avatar for the user using a deep learning model. Next, it retrieves a background image of the Eiffel Tower from the database at the specified date and time, and composites the user's avatar onto the background image. Finally, it sends the composite photo to the user by email every morning at 8:00.

[0546] This system can provide users with a realistic travel experience, which can be shared and enjoyed on social media. Specific hardware requirements include a high-performance image processing engine on the server side and GPU resources for the deep learning model. Software requirements include facial recognition algorithms (e.g., OpenCV) and deep learning libraries (e.g., TensorFlow, PyTorch).

[0547] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0548] Step 1:

[0549] Users access the website or app, register an account, and log in. Next, they select 10 photos of their face on the face photo upload screen and send them to the server.

[0550] Input: 10 photos of your face (JPEG or PNG format, minimum 400x400 pixels)

[0551] Output: Image data uploaded to the server

[0552] Specific operation: Select 10 face photos from the user's device through file selection and send them to the server as an HTTP POST request.

[0553] Step 2:

[0554] The server applies a facial recognition algorithm (e.g., OpenCV) to the 10 facial photos received to extract facial features, specifically identifying the positions and shapes of the eyes, nose, mouth, etc.

[0555] Input: 10 user-uploaded face photos

[0556] Output: Extracted facial feature data (position and shape information)

[0557] Specific operation: A facial recognition algorithm is applied on the server side to extract feature points from each facial photograph, such as the facial contours, eye positions, nose positions, and mouth positions, and obtain these as coordinate data.

[0558] Step 3:

[0559] The server generates a virtual character (avatar) for the user using a deep learning model (e.g., a person generation model using PyTorch) based on the facial feature data. The generated avatar is saved in the user's profile information.

[0560] Input: Extracted facial feature data

[0561] Output: Digital virtual character (avatar)

[0562] How it works: Facial feature data is fed into the deep learning model as input, and a virtual character resembling the user is generated based on that data. The generated virtual character data is then saved in the user's profile information.

[0563] Step 4:

[0564] The user moves to the travel destination selection screen, selects the desired travel destination from the map interface or list, and sets the method of sending the travel photos (e.g., email, SNS, LINE) and the date and time of receipt.

[0565] Input: User-selected travel destination information, sending method, and date and time of receipt

[0566] Output: Configuration information sent to the server

[0567] Specific operation: Click on a travel destination on the map interface or select it from the list, then set the sending method and receiving date and time using the drop-down menu or calendar UI. Pressing the setting button sends this information to the server.

[0568] Step 5:

[0569] The server records the user's selection information (travel destination, transmission method, date and time of receipt) in the profile.

[0570] Input: Selections received from the user

[0571] Output: Recorded user settings information

[0572] What happens: The server stores the received selection data in the user profile database for future reference.

[0573] Step 6:

[0574] At the set date and time, the server retrieves the background image of the travel destination selected by the user from the database.

[0575] Input: User selection information, background image data in the database

[0576] Output: The background image obtained.

[0577] Specific operation: For example, search and retrieve background images of a specified travel destination from cloud storage such as Amazon S3.

[0578] Step 7:

[0579] The server uses deep learning models to calculate the pose of the virtual character, placing it naturally on the background image, and also adjusts the shadows and lighting to ensure consistency with the background.

[0580] Input: background image, virtual character

[0581] Output: Synthesized, realistic travel photos

[0582] How it works: It uses deep learning models to place virtual characters naturally against backgrounds and adjust shadows and light intensity to generate composite photos.

[0583] Step 8:

[0584] The server sends the generated travel photos using the method specified by the user (email, SNS, LINE, etc.).

[0585] Input: Composite travel photos, user's sending method settings

[0586] Output: Travel photos sent to the user

[0587] Specific operation: For email, the SendGrid API is used to send emails with travel photos attached, and for SNS and LINE, the corresponding API is used to send photos along with the message.

[0588] Step 9:

[0589] Users can select additional features on the options page, such as photos of luxurious food or photos of themselves with celebrities.

[0590] Input: Additional feature selection information

[0591] Output: Configuration information for additional features sent to the server

[0592] Specific operation: Select the items displayed on the options page using checkboxes or drop-down menus, press the Settings button, and send the selected options to the server. The selected options are recorded and managed on the server side.

[0593] (Application example 1)

[0594] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0595] Today, users enjoy virtual travel experiences through various means, but these technologies lack realism and do not fully satisfy users. Furthermore, it is difficult for users to easily and effectively create digital content that reflects their own faces and share it via social media, email, etc. This creates a demand for personalized, interactive experiences. Therefore, a system is needed that can generate a realistic avatar based on a user's facial photograph, naturally combine it with the background of a specified travel destination, and deliver it in a specified manner and at a specified time.

[0596] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0597] In this invention, the server includes means for receiving a user's facial photograph, means for generating a user avatar based on the received facial photograph using a facial recognition algorithm, means for acquiring a background image based on a travel destination specified by the user, means for synthesizing the acquired background image with the generated avatar to create a realistic travel photograph, means for calculating a natural position and pose of the synthesized image using a generative AI model to create a travel photograph, and means for transmitting the created travel photograph in a manner and at a time specified by the user. This allows users to have a personalized and interactive virtual travel experience in a highly realistic way and easily share it on various digital platforms.

[0598] "User" means an individual who uses the system to upload a photo of themselves and enjoy a virtual travel experience.

[0599] A "face photo" is a digital image of the user's face, and serves as the basic data for generating an avatar.

[0600] A "facial recognition algorithm" is a computational method for extracting facial features such as eyes, nose, and mouth from a received photograph of a face and generating a digital avatar.

[0601] An "avatar" is a digital virtual persona created based on a user's facial features using facial recognition algorithms.

[0602] "Travel destination" refers to information about a place that a user specifies as the destination of a virtual trip in the system.

[0603] A "background image" is a digital image of scenery or tourist spots at a travel destination designated by the user, which is combined with the avatar.

[0604] "Generative AI model" refers to the deep learning model used to calculate natural-looking placement and pose for synthesized images.

[0605] "Transmission means" refers to tools or APIs that allow users to send the generated travel photos in a manner specified by the user (e.g., email, SNS, LINE, etc.).

[0606] "Image compositing" is the process of naturally combining background images and avatars to create realistic travel photos.

[0607] "Transmission method" refers to the communication means selected by the user for transmitting travel photos.

[0608] "Realism" refers to the quality of the synthesized images, making them appear as if they were taken during an actual trip.

[0609] This invention is a system that generates an avatar based on a user's facial photograph, combines it with a background image of a travel destination selected by the user to generate a realistic travel photograph, and transmits it in a specified manner. This system is implemented in the following steps.

[0610] First, a user accesses the website or application from a smartphone or PC, registers, and logs in. Next, the user selects 10 facial photos on the face photo upload screen and submits them to the server. These facial photos are verified to be in the correct format and resolution.

[0611] The server processes the received facial photo using a facial recognition algorithm (for example, using a deep learning library such as OpenCV or TensorFlow) to extract facial features (eyes, nose, mouth, etc.). Based on the extracted facial features, a deep learning model is used to generate an avatar for the user. This avatar is then stored in the user's profile.

[0612] Next, the user accesses the travel destination selection screen and selects their desired travel destination. Destination selection can be done using a map interface or a list of options. The user also selects the method of sending the generated travel photos (email, SNS, LINE, etc.) and the date and time of sending. This selection information is sent to the server and recorded in the user profile.

[0613] At the set date and time, the server retrieves a background image of the selected travel destination from the database. The generated avatar and background image are then composited using a deep learning model to create a natural position and pose. This process uses a generative AI model, and the composite image is then generated as a realistic travel photo.

[0614] Next, the realistic travel photos are sent in the way the user specified. For example, in the case of email, an email generation API is used, and in the case of SNS or LINE, messages are sent through the respective APIs. The travel photos are delivered at the specified time.

[0615] Additionally, users can select additional options, such as images of specific local items or meals, which are also saved in the user profile and the image generation process begins after payment confirmation.

[0616] As a concrete example, consider a case where a user selects the Eiffel Tower in Paris as a travel destination and requests to receive travel photos by email every morning at 8:00. In this case, the server will synthesize the background image of the Eiffel Tower with the user's avatar at the set time and send the travel photos by email every morning at 8:00. Throughout this process, a generative AI model is used to ensure natural synthesis and placement.

[0617] Prompt Sentence Examples

[0618] "Generate an avatar from a photo of the user's face and overlay it with a travel photo of the Eiffel Tower in the background. Background image: path / to / paris_eiffel.jpg Facial features: {Eye position: (x1, y1), Nose position: (x2, y2), Mouth shape: (x3, y3)}"

[0619] In this way, users can have a realistic virtual travel experience and easily share their travel photos on various digital platforms.

[0620] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0621] Step 1:

[0622] A user accesses a website or application using a smartphone or PC. First, the user registers and logs in to the site. Next, the user selects 10 facial photos on the facial photo upload screen and sends them to the server. The input is the 10 facial photos, and the output is the photo data stored on the server.

[0623] Step 2:

[0624] The server applies a facial recognition algorithm to the received facial photo. Specifically, it uses deep learning libraries such as OpenCV and TensorFlow to extract facial features (the position and shape of the eyes, nose, and mouth). The input is the facial photo data, and the output is data that quantifies the facial features.

[0625] Step 3:

[0626] The server uses a generative AI model to generate a digital avatar for the user based on the facial feature data extracted by the facial recognition algorithm. This avatar is virtual person data that faithfully reflects the user's facial features. The input is facial feature data, and the output is a digital avatar.

[0627] Step 4:

[0628] The user accesses the travel destination selection screen and selects the desired travel destination from a map interface or list. They also specify the method of sending the travel photos (email, SNS, LINE, etc.) and the date and time of receipt. The input is the user's selected travel destination and sending method information, and the output is the selected information recorded on the server.

[0629] Step 5:

[0630] At the set date and time, the server retrieves the background image of the travel destination selected by the user from the database. It then composites the background image with the user's avatar, using a generative AI model to calculate natural placement and poses to create realistic travel photos. The input is the background image of the travel destination and the user's digital avatar, and the output is the composite travel photo.

[0631] Step 6:

[0632] The server sends the generated travel photos in the way specified by the user. For example, in the case of email, it uses the email generation API, and in the case of SNS or LINE, it sends through the respective API. The input is the travel photo data and sending setting information, and the output is the travel photo sent to the user.

[0633] Step 7:

[0634] When a user selects an add-on feature, the server executes the corresponding image generation process, for example, adding an image of a specific local item or meal to a travel photo. This process is also performed using a generative AI model. The input is the selection information for the add-on option, and the output is the travel photo with the option added.

[0635] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0636] The system of this invention generates a traveling avatar based on a facial photo provided by the user, generates a composite image based on the travel destination selected by the user, and transmits it in a specified manner. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to suggest and select travel destinations and optional images based on the user's emotional state. Specific embodiments of the invention are described below.

[0637] 1. Uploading a user's photo

[0638] The user accesses the website or app and logs in. After logging in, the user selects 10 facial photos on the face photo upload screen and sends them to the server.

[0639] 2. Facial Photo Processing and Avatar Generation

[0640] The server applies a facial recognition algorithm to the received facial photo to extract facial features (the position and shape of the eyes, nose, and mouth), then uses a deep learning model to generate an avatar for the user, a digital virtual persona that faithfully reflects the user's facial features, and the generated avatar is saved in the user's profile.

[0641] 3. Emotion Recognition and Analysis

[0642] In addition to a photo of their face, the user provides the emotion engine with a photo of their facial expression and voice data. The emotion engine then analyzes the user's emotional state based on the data provided by the user. The analysis results are classified into emotion categories such as "joy," "sadness," "surprise," and "anger."

[0643] 4. Select your destination and set it up

[0644] Users access the travel destination selection screen and select their desired travel destination from a map interface or list. If the user requests automatic suggestions based on their emotional state, the system will suggest the most suitable travel destination based on the analysis results of the emotion engine. Users also set the sending method (email, SNS, LINE, etc.) and reception time, and press the "Settings Complete" button once the settings are complete.

[0645] 5. Travel Photo Generation

[0646] At the set date and time, the server retrieves a background image of the user's travel destination from the database. The server then composites the user's avatar onto the background image, using a deep learning model to calculate natural placement and pose. The composite image is then generated as a realistic travel photo.

[0647] 6. Send travel photos

[0648] The server sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.). For email, an email generation API is used, and for SNS or LINE, a message is sent through the corresponding API. The travel photos are delivered at the specified time.

[0649] 7. Providing additional options

[0650] Users can select additional features on the options page, such as photos of rich cuisine, luxurious accommodations, or photos of people with celebrities. The selected options are saved directly under the user's control, and the corresponding image generation process begins after payment is confirmed.

[0651] Specific examples

[0652] For example, let's consider a case where a user uploads a photo of their face and the emotion engine recognizes the emotion "joy." Based on the analysis results of the emotion engine, the server suggests "The Eiffel Tower in Paris" as the best travel destination. If the user accepts this suggestion and sets the sending method to email and the receiving time to 8:00 a.m. every morning, the server will combine a background image of the Eiffel Tower with the user's avatar, generating and sending a realistic travel photo every morning at 8:00 a.m. If the user selects a photo of "rich French cuisine" as an additional option, that image will also be sent.

[0653] In this way, by utilizing the system of the present invention, users can easily obtain realistic travel photos based on their emotional state, which can be enjoyed on social media and other platforms.

[0654] The processing flow will be explained below.

[0655] Step 1:

[0656] The user accesses the website or app and logs in. After logging in, the user selects 10 facial photos on the face photo upload screen, as well as facial expression photos and voice data for recognizing their emotional state, and sends these to the server.

[0657] Step 2:

[0658] The device collects the facial photo file and emotion recognition data selected by the user and uploads them to a server via the network.

[0659] Step 3:

[0660] The server receives the uploaded facial photo and emotion recognition data, checks the file format and resolution of the facial photo, and converts it into a format suitable for the facial recognition algorithm.

[0661] Step 4:

[0662] The server uses a facial recognition algorithm to extract facial features (position and shape of eyes, nose, and mouth).

[0663] Step 5:

[0664] The server uses a deep learning model (e.g., StyleGAN) to generate an avatar for the user based on the extracted features.

[0665] Step 6:

[0666] The server stores the generated avatar in the user profile.

[0667] Step 7:

[0668] The server uses an emotion engine to analyze the user's emotional state based on facial expressions and voice data, and classifies the results into emotion categories such as "happiness," "sadness," "surprise," and "anger."

[0669] Step 8:

[0670] The server will suggest the most suitable travel destination for the user based on the analysis results of the emotion engine. For example, if the emotion of "joy" is detected, the server will suggest the Eiffel Tower in Paris.

[0671] Step 9:

[0672] The user accesses the travel destination selection screen and selects the suggested travel destination or their own desired travel destination. They also set the sending method (email, SNS, LINE, etc.) and reception time. Once the settings are complete, they press the "Settings Complete" button.

[0673] Step 10:

[0674] The terminal collects information on the travel destination and transmission method selected by the user and transmits it to the server via the network.

[0675] Step 11:

[0676] The server processes the received configuration information and stores it in a user profile.

[0677] Step 12:

[0678] When the set date and time arrives, the server retrieves a background image of the travel destination specified by the user from the database.

[0679] Step 13:

[0680] The server then composites the user's avatar with the captured background image, using a deep learning model (e.g., OpenPose) to calculate natural placement and pose.

[0681] Step 14:

[0682] The server generates the composite image as the final travel photo.

[0683] Step 15:

[0684] The server sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.).

[0685] Step 16:

[0686] The device receives notifications via email, SNS, or LINE and displays the image to the user.

[0687] Step 17:

[0688] Users can access an options page and select additional options such as photos of rich cuisine, photos of luxurious accommodations, or photos of themselves with celebrities.

[0689] Step 18:

[0690] The terminal collects the selected option information and billing information and transmits them to the server via the network.

[0691] Step 19:

[0692] The server processes the received option and billing information and confirms the user's purchase.

[0693] Step 20:

[0694] The server initiates an image generation process according to the additional options, and generates and transmits an image corresponding to the specified date and time.

[0695] By going through these specific steps, users can receive realistic travel photos based on their own emotional state, which they can enjoy on social media and other platforms.

[0696] Example 2

[0697] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0698] In a system for generating realistic travel photos, there is a need for a means to reflect the user's emotional state and efficiently and effectively set and save travel destination suggestions and subsequent transmission methods.In addition, it has been pointed out that conventional systems need not only to synthesize background images and avatars based on the user's selection, but also to generate optional images to provide further enjoyment.

[0699] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving a user's facial photograph, means for generating a user's avatar based on the received facial photograph using a facial recognition algorithm, means for acquiring a background image based on a travel destination designated by the user, means for creating a realistic travel photograph by combining the acquired background image with the generated avatar, means for transmitting the created travel photograph in a manner designated by the user, means for analyzing and recognizing the user's emotional state and suggesting a travel destination based on the analysis results, and means for setting and saving a transmission method and transmission time. This makes it possible to suggest travel destinations and generate optional images that reflect the user's emotional state, and further to transmit the travel photograph in a manner designated by the user at a set date and time.

[0700] A "user" is an individual who uses the system to upload a photo of themselves and select a travel destination and optional images.

[0701] The "server" is a computer system that receives a user's facial photograph, generates an avatar using a facial recognition algorithm, obtains background images of the travel destination, and analyzes the user's emotional state.

[0702] A "face recognition algorithm" is a technology for extracting the position and shape of the eyes, nose, and mouth from a facial photo provided by the user.

[0703] An "avatar" is a digital virtual persona that reflects a user's facial features and is generated based on facial recognition algorithms.

[0704] A "background image" is an image of scenery or famous places at a travel destination designated by the user.

[0705] "Compositing" is the process of combining a background image with a generated avatar to create a realistic travel photo.

[0706] "Emotional state" refers to psychological states such as "joy," "sadness," "surprise," and "anger" that are analyzed from facial expression photos and voice data provided by the user.

[0707] "Suggestion" refers to the act of the server suggesting the most suitable travel destination to the user based on the analysis results of the emotion engine.

[0708] "Optional images" are images added by user selection, including images of travel destinations, specific items, and meals.

[0709] "Sending method" refers to the means by which the generated travel photos are delivered to the user, and includes email, SNS, LINE, etc.

[0710] "Send time" is the designated time for sending the generated travel photos and optional images to the user.

[0711] The system of this invention generates a traveling avatar based on a facial photograph provided by the user, suggests optimal travel destinations based on the user's emotional state, generates a composite image, and transmits it in a specified manner. Specific embodiments are described below.

[0712] First, the user accesses the website or app, logs in, selects 10 photos of their face on the face photo upload screen, and sends them to the server. At this point, the user can upload the photos from the app using their smartphone.

[0713] The server reviews the received facial photo and applies facial recognition algorithms such as OpenCV or Dlib to extract facial features. This facial recognition process specifically analyzes the position and shape of the eyes, nose, and mouth in the photo. It then uses a deep learning model such as GAN to generate an avatar that reflects the user's facial features. This avatar is then saved in the user's profile for further processing.

[0714] The user also provides the server with a photo of their facial expression and voice data. The server's emotion engine then analyzes the user's emotional state based on this data. This emotion analysis uses emotion analysis APIs from Microsoft Azure or Google Cloud, for example. Based on the analysis results, the user's emotional state is classified into emotion categories such as "happiness," "sadness," "surprise," and "anger."

[0715] Next, the user can access the travel destination selection screen and select their desired travel destination. The system automatically suggests the most suitable travel destination based on the analysis results of the emotion engine. For example, if the emotional state is "joy," the system will suggest a travel destination that matches the emotion, such as Paris, France. The user can accept the suggested travel destination or choose a different one by themselves. At this time, the user sets the sending method (email, SNS, LINE, etc.) and reception time, and presses the "Settings Complete" button. These settings are then saved on the server.

[0716] At the set date and time, the server retrieves a background image of the travel destination from a database. This background image is taken from a common photo library such as Google Images. The server then uses a deep learning model to naturally blend the user's avatar into the background image, generating a realistic travel photo. Pose Estimation technology is used to properly position the avatar against the background image.

[0717] Finally, the server sends the generated travel photos via the method specified by the user: via email using the SendGrid API, or via the corresponding API for social media or LINE. The generated images are sent at the time specified by the user.

[0718] Users can also select additional features on the options page, such as photos of rich cuisine, luxurious accommodations, and photos of two people with celebrities. After the user confirms their selection and payment, the image generation process will begin and the generated optional images will be sent.

[0719] As a concrete example, if a user uploads a photo of their face and the emotion engine recognizes the emotion "joy," the system will suggest "The Eiffel Tower in Paris" as a travel destination, and the user will accept this suggestion and set the reception time to 8:00 a.m. every morning. The server will then combine the background image of the Eiffel Tower with the user's avatar, generate a realistic travel photo every morning, and send it by email. If the user selects a photo of "rich French cuisine" as an additional option, that image will also be sent.

[0720] Example prompt sentence:

[0721] The following prompts explain system behavior to the generative AI model:

[0722] Your task is to generate a program for a system that, based on 10 facial photos provided by the user, will combine the avatar with a background image of the travel destination selected by the user when the emotion engine recognizes the emotion "joy." The process will be as follows:

[0723] 1. Extract facial features using face recognition algorithm (OpenCV / Dlib).

[0724] 2. Generate a user avatar using GAN.

[0725] 3. Analyze user emotions using an emotion engine (Microsoft Azure / Google Cloud).

[0726] 4. Composite avatar and background image based on suggested travel destination (e.g., Eiffel Tower in Paris).

[0727] 5. Using the email sending API (SendGrid), the composite photo is sent every morning at 8:00.

[0728] In this way, by using the system of the present invention, realistic travel photos based on the user's emotional state can be easily provided.

[0729] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0730] Step 1:

[0731] A user accesses a website or app, logs in, and proceeds to the face photo upload screen. The user selects 10 face photos from the photo gallery and presses the upload button. This operation sends the face photos to the server. The input is the user's face photo, and the output is the face photo data sent to the server.

[0732] Step 2:

[0733] The server checks the received face photo and applies a facial recognition algorithm (OpenCV or Dlib) to extract facial features (the position and shape of the eyes, nose, and mouth). The algorithm converts the image data to grayscale and creates a list of detected feature points. The input is the face photo data, and the output is the extracted facial feature data.

[0734] Step 3:

[0735] The server generates a user avatar using a GAN (generative adversarial network). Using facial feature data as input, the GAN model generates a virtual avatar image based on the learning results. The generated avatar is saved in the user profile. The input is facial feature data, and the output is the generated avatar image.

[0736] Step 4:

[0737] The user then provides the server with photographs of facial expressions and voice data. The server's emotion engine then uses this data to analyze the user's emotional state. Emotion analysis APIs from Microsoft Azure and Google Cloud are used for the emotion analysis. The analysis results are categorized into "happiness," "sadness," "surprise," and "anger" and stored on the server. The input is photographs of facial expressions and voice data, and the output is analyzed emotional state data.

[0738] Step 5:

[0739] The user can access the travel destination selection screen and select their desired travel destination. The server will suggest the most suitable travel destination based on the analysis results of the emotion engine. After the user accepts the suggested travel destination or selects another one by themselves, they set the sending method (email, SNS, LINE, etc.) and reception time. These settings are saved on the server. The input is the user's emotional state and travel destination selection data, and the output is the saved setting data.

[0740] Step 6:

[0741] At the set date and time, the server retrieves a background image of the travel destination from a database. This background image can be from Google Images or a user's own photo library. The server then uses a deep learning model (using Pose Estimation technology) to naturally blend the user's avatar into the background image. The input is the background image and the avatar image, and the output is the blended travel photo.

[0742] Step 7:

[0743] The server sends the generated travel photos in the manner specified by the user. The SendGrid API is used for email transmission, and the corresponding APIs for SNS and LINE are used. The travel photos are sent to the user at the specified time. The input is the composited travel photos and transmission setting data, and the output is the sent travel photos.

[0744] Step 8:

[0745] Users can select additional features on the options page. These include photos of rich cuisine, luxurious accommodations, and photos of people with celebrities. After the user selects these options and confirms the payment, the image generation process begins. The input is the option selection data, and the output is the generated option image.

[0746] In this way, each processing step operates in cooperation with one another, thereby realizing a system that provides realistic travel photos based on the user's emotional state.

[0747] (Application example 2)

[0748] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0749] Conventional avatar generation and travel photo creation systems lack the ability to suggest travel destinations that take the user's emotional state into account, resulting in the inability to provide the optimal travel experience for the user. Furthermore, simply synthesized photos alone are unlikely to achieve the realism and personalized satisfaction that users expect. Therefore, a new system that reflects the user's emotions and provides a more personalized travel experience is needed.

[0750] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0751] In this invention, the server includes means for receiving a user's facial photograph, means for generating a user avatar based on the received facial photograph using a facial recognition algorithm, means for acquiring a background image based on a travel destination designated by the user, means for synthesizing the acquired background image with the generated avatar to create a realistic travel photograph, means for transmitting the created travel photograph in a manner designated by the user, means for analyzing the user's emotional state using an emotion recognition engine, and means for suggesting travel destinations based on the user's emotional state. This makes it possible to suggest optimal travel destinations taking the user's emotional state into consideration and create realistic travel photographs.

[0752] "User" refers to an individual who uses this system.

[0753] "Facial photo" refers to image data provided by a user that is taken of the user's own face using a camera or the like.

[0754] A "facial recognition algorithm" refers to a computational method that extracts features such as the eyes, nose, and mouth from a photograph of a face and identifies and recognizes a specific face.

[0755] An "avatar" refers to a virtual digital persona generated based on a photograph of a user's face.

[0756] "Travel destination" refers to a destination where a user wishes to take a virtual trip.

[0757] "Background image" refers to image data of scenery and famous places at travel destinations.

[0758] "Compositing" refers to the process of overlaying a user's avatar with a background image to generate a single image.

[0759] "Realistic travel photos" refer to images that are synthesized to make the user's avatar appear as if they are actually at the travel destination.

[0760] An "emotion recognition engine" refers to a system that analyzes and classifies a user's emotional state based on a photo of their face and voice data.

[0761] "Emotional state" refers to a user's current psychological state or mood.

[0762] "Suggestion" refers to the system selecting and presenting appropriate travel destinations and options to the user.

[0763] "Sending" refers to the process of delivering the generated travel photos via the communication means selected by the user.

[0764] This system generates a traveling avatar using a user's facial photograph, and uses an emotion recognition engine to suggest optimal travel destinations based on the user's emotional state. It also creates realistic travel photos by combining the user's avatar with background images of the travel destination, and sends them in a manner specified by the user, providing a virtual travel experience.

[0765] The system mainly operates in the following steps:

[0766] 1. Uploading a user's photo

[0767] Users access a dedicated application using a smartphone or PC and log in. After logging in, they select multiple facial photos on the facial photo upload screen and send them to the server.

[0768] 2. Facial Photo Processing and Avatar Generation

[0769] The server applies a facial recognition algorithm to the received face photo to extract facial features (the position and shape of the eyes, nose, and mouth). It then uses a generative AI model to generate an avatar for the user. Leveraging a deep learning model, a digital persona is created that faithfully reflects the user's facial features, and this avatar is stored in the user's profile.

[0770] 3. Emotion Recognition and Analysis

[0771] Users provide facial photos, facial expression photos, and voice data to the emotion recognition engine, which uses technologies such as DeepFace to analyze the user's emotional state and classify the results into categories such as "happiness," "sadness," "surprise," and "anger."

[0772] 4. Select your destination and set it up

[0773] Users can select their desired travel destination from a map or list through the system interface. They can also choose destinations suggested by the emotion recognition engine. If they like the suggested destination, they confirm the setting. They can also set the sending method (email, SNS, etc.) and the reception time.

[0774] 5. Travel Photo Generation

[0775] At the set date and time, the server retrieves a background image of the travel destination specified by the user from the database, then combines the generated avatar with the background image, calculating natural positioning and poses to generate a realistic travel photo.

[0776] 6. Send travel photos

[0777] The server then sends the generated travel photos via the method specified by the user: via email using a dedicated email generation API, or via the corresponding API for social media or text messages.

[0778] 7. Providing additional options

[0779] Users can select additional features on the options page within the system, such as "rich food photos" or "luxury accommodations." The options selected by the user are saved in their profile, and the corresponding image generation process is initiated after payment confirmation.

[0780] Specific examples

[0781] For example, if a user provides a photo of a "joyed" expression to the emotion recognition engine, the system will suggest "The Eiffel Tower in Paris" as the best travel destination. If the user accepts this suggestion and sets the sending method to email and the receiving time to 8:00 a.m., the server will combine the background image of the Eiffel Tower with the user's avatar, generate a realistic travel photo, and send it every morning at 8:00 a.m. If the user also selects a photo of "rich French cuisine" as an additional option, that will also be sent.

[0782] Prompt Sentence Examples

[0783] "Create realistic travel photos by combining user-uploaded photos of your face with a background image of the Eiffel Tower in Paris."

[0784] The system allows users to virtually enjoy an optimal travel experience based on their emotional state and receive personalized content tailored to their individual needs.

[0785] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0786] Step 1:

[0787] The user logs in to a dedicated application using a smartphone or computer.

[0788] Input: Login information (user ID, password)

[0789] Processing: User authentication is performed and the user profile is obtained.

[0790] Output: A successful authentication message, and the user's profile information.

[0791] Step 2:

[0792] The user selects multiple facial photos on the facial photo upload screen and sends them to the server.

[0793] Input: Multiple face photos (e.g. 10 photos)

[0794] Processing: Your face photo is uploaded to the server and temporarily stored.

[0795] Output: Face photo upload complete message.

[0796] Step 3:

[0797] The server applies a facial recognition algorithm to the received facial photo and extracts facial features.

[0798] Input: Multiple face photos

[0799] Processing: Use libraries such as OpenCV to extract facial features such as eyes, nose, and mouth.

[0800] Output: Extracted facial feature data.

[0801] Step 4:

[0802] The server uses a deep learning model to generate an avatar for the user.

[0803] Input: Extracted facial feature data

[0804] Processing: Generate an avatar based on the user's face using a generative AI model (e.g., GAN).

[0805] Output: Generated avatar data.

[0806] Step 5:

[0807] The user provides facial expression photos and voice data to the emotion recognition engine.

[0808] Input: facial expression photos and audio data

[0809] Processing: Analyze the user's emotional state using emotion recognition engines such as DeepFace.

[0810] Output: Classification of emotional state (e.g., happy, sad, surprised, angry, etc.).

[0811] Step 6:

[0812] The server suggests optimal travel destinations based on the user's emotional state.

[0813] Input: Emotional state classification result

[0814] Processing: Randomly or algorithmically select the best travel destination from a list of travel destinations that correspond to the emotional state.

[0815] Output: Suggested travel destinations.

[0816] Step 7:

[0817] The user selects a desired travel destination on the travel destination selection screen and sets the transmission method and reception time.

[0818] Input: Travel destination selection information, sending method, receiving time

[0819] Action: Save the user's selections and settings to the database.

[0820] Output: Configuration complete message.

[0821] Step 8:

[0822] When the set date and time arrives, the server retrieves a background image of the travel destination specified by the user from the database.

[0823] Input: Travel destination selection information

[0824] Processing: Search and retrieve background images corresponding to the specified travel destination from the background image database.

[0825] Output: The obtained background image.

[0826] Step 9:

[0827] The server combines the generated avatar with the acquired background image.

[0828] Input: Avatar data, background image

[0829] Processing: Uses deep learning models to calculate natural placement and pose, and then synthesizes the avatar with the background image.

[0830] Output: Realistic travel photos.

[0831] Step 10:

[0832] The server transmits the generated travel photos in a manner designated by the user.

[0833] Input: Travel photos, sending method information

[0834] Processing: Send travel photos in the specified way using email generation API or SNS sending API.

[0835] Output: A transmission completion message.

[0836] Step 11:

[0837] The user selects additional features on the options page, and after confirming payment, the corresponding image generation process begins.

[0838] Input: Additional option selection information, billing information

[0839] Processing: Performs additional image generation processing based on selected options and saves it to the user profile.

[0840] Output: The additional image generated, and a charge completion message.

[0841] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0842] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0843] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0844] [Third embodiment]

[0845] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0846] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0847] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0848] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0849] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0850] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0851] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0852] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0853] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0854] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0855] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0856] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0857] The system of the present invention generates a traveling avatar based on a facial photograph provided by the user, generates a composite image based on a travel destination selected by the user, and transmits the composite image in a specified manner. Specific embodiments of the present invention are described below.

[0858] 1. Uploading a user's photo

[0859] Users visit a website or app, register, and log in. They then select 10 photos of their face on a face photo upload screen and submit them to the server, which verifies that the selected images are in the correct format and resolution before submitting them.

[0860] 2. Facial Photo Processing and Avatar Generation

[0861] The server applies a facial recognition algorithm to the received facial photo to extract facial features (the position and shape of the eyes, nose, and mouth). It then uses a deep learning model to generate an avatar for the user. This avatar is a digital virtual persona that faithfully reflects the user's facial features. The generated avatar is then saved in the user's profile.

[0862] 3. Select your destination and set it up

[0863] The user accesses the travel destination selection screen and selects the desired travel destination from a map interface or list. They also set the sending method (email, SNS, LINE, etc.) and the date and time of receipt. The selection information is sent to the server by pressing the send button, and the server records it in the user profile.

[0864] 4. Travel Photo Generation

[0865] The server retrieves a background image of the travel destination selected by the user from a database at the specified date and time. It then composites the user's avatar onto the background image, using a deep learning model to calculate natural placement and pose. The composite image is then generated as a realistic travel photo.

[0866] 5. Send travel photos

[0867] The server sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.). For email, an email generation API is used, and for SNS and LINE, the message is sent through the corresponding API. The travel photos are delivered at the specified time.

[0868] 6. Providing additional options

[0869] Users can select additional features on the options page. For example, they can select rich food photos, luxurious accommodations, or photos of themselves with celebrities. The selected options are saved directly under the user's control, and after payment confirmation, the corresponding image generation process begins. This provides users with an even richer experience.

[0870] Specific examples

[0871] For example, let's say a user uploads a photo of their face, sets their destination as "Eiffel Tower in Paris," and selects to receive the photo every morning at 8:00. The server will combine the background image of the Eiffel Tower with the user's avatar at the time specified by the user, and send the resulting photo by email every morning at 8:00. This allows the user to have a realistic experience every day, as if they were visiting Paris themselves.

[0872] In this way, by utilizing the system of the present invention, users can easily obtain photos that look as if they were actually traveling, and enjoy them on social media and other platforms.

[0873] The processing flow will be explained below.

[0874] Step 1:

[0875] The user accesses the website or app and logs in. After logging in, the user selects 10 photos of their face on the face photo upload screen.

[0876] Step 2:

[0877] The terminal collects the facial photo files selected by the user and uploads them to a server via a network.

[0878] Step 3:

[0879] The server receives the uploaded facial photo, checks the file format and resolution of the facial photo, and converts it into a format suitable for the facial recognition algorithm.

[0880] Step 4:

[0881] The server uses a facial recognition algorithm to extract facial features (position and shape of eyes, nose, and mouth).

[0882] Step 5:

[0883] The server uses a deep learning model (e.g., StyleGAN) to generate an avatar for the user based on the extracted features.

[0884] Step 6:

[0885] The server stores the generated avatar in the user profile.

[0886] Step 7:

[0887] The user accesses a destination selection screen and selects a desired destination from a map interface or list.

[0888] Step 8:

[0889] The user sets the sending method (email, SNS, LINE, etc.) and the reception time. Once the settings are complete, press the "Settings Complete" button.

[0890] Step 9:

[0891] The terminal collects information on the travel destination and transmission method selected by the user and transmits it to the server via the network.

[0892] Step 10:

[0893] The server processes the received configuration information and stores it in a user profile.

[0894] Step 11:

[0895] When the set date and time arrives, the server retrieves a background image of the travel destination specified by the user from the database.

[0896] Step 12:

[0897] The server then composites the user's avatar with the captured background image, using a deep learning model (e.g., OpenPose) to calculate natural placement and pose.

[0898] Step 13:

[0899] The server generates the composite image as the final travel photo.

[0900] Step 14:

[0901] The server formats and sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.).

[0902] Step 15:

[0903] The device receives notifications via email, SNS, or LINE and displays the image to the user.

[0904] Step 16:

[0905] Users can access an options page and select additional options such as photos of rich cuisine, photos of luxurious accommodations, or photos of themselves with celebrities.

[0906] Step 17:

[0907] The terminal collects the selected option information and billing information and transmits them to the server via the network.

[0908] Step 18:

[0909] The server processes the received option and billing information and confirms the user's purchase.

[0910] Step 19:

[0911] The server initiates an image generation process according to the additional options, and generates and transmits an image corresponding to the specified date and time.

[0912] By going through these steps, users can easily receive realistic travel photos that make them feel as if they are traveling around the world.

[0913] Example 1

[0914] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0915] Conventional image editing systems make it difficult for users to easily and effectively create realistic travel photos using their own facial images. They also lack features such as generating realistic images based on specific travel destinations or automatically sending photos at specified dates and times. As a result, users had to manually edit and send photos one by one, which was time-consuming and laborious, resulting in a poor user experience.

[0916] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0917] In this invention, the server includes means for receiving a user's image, means for generating a virtual character for the user based on the received image using a face recognition algorithm, means for acquiring background data based on a travel destination specified by the user, means for synthesizing the acquired background data with the generated virtual character to create a realistic image, and means for transmitting the created image in a manner specified by the user. This allows users to easily generate realistic travel photos using their own face photos, and the photos are automatically transmitted at a specified time, allowing them to be enjoyed on social media and other platforms without any hassle.

[0918] "User" refers to any individual or entity that uses the System.

[0919] "Image" refers to visual information stored or displayed in digital form.

[0920] A "facial recognition algorithm" is a computational method for detecting and identifying facial features in an image.

[0921] "Virtual character" refers to a digitally generated person based on the user's facial features.

[0922] "Travel destination" refers to a place that the user would like to virtually visit.

[0923] "Background data" refers to image and video information corresponding to a travel destination.

[0924] "Synthesis" refers to the process of integrating multiple images or data to generate a single image or data.

[0925] "Realistic images" refer to images that are synthesized to appear as if they exist in reality.

[0926] "Transmission" refers to the act of sending the generated image or data to an external party in a manner designated by the user.

[0927] "Additional selection menu" refers to optional functions and content provided in addition to basic functions.

[0928] "Setting" refers to the act of saving and making available information such as the user's selected travel destination and transmission method.

[0929] "Predetermined information" refers to specific information such as the travel destination, transmission method, and transmission date and time that the user has set in advance.

[0930] The system of the present invention generates a traveling avatar based on a facial photograph provided by the user, generates a composite image based on the travel destination selected by the user, and transmits the composite image in a specified manner. Specific embodiments of the invention are described below.

[0931] Uploading a user's face photo

[0932] A user accesses a website or app. They register an account and log in. They then select 10 photos of their face on the face photo upload screen and send them to the server. The image files must be in JPEG or PNG format with a resolution of at least 400x400 pixels. The server checks the format and resolution of the received images and returns an error message if they are inappropriate.

[0933] Facial photo processing and avatar generation

[0934] After receiving the uploaded face photo, the server uses a facial recognition algorithm (e.g., OpenCV) to extract facial features. Specifically, it identifies facial landmarks such as the position and shape of the eyes, nose, and mouth. It then uses a deep learning model (e.g., a person generation model using PyTorch) to generate a digital avatar based on the user's facial features. The generated avatar is then saved in the user's profile information.

[0935] Destination selection and settings

[0936] The user is taken to a travel destination selection screen. Here, they can select their desired travel destination from a map interface or a list. They can also set how they want to send their travel photos (e.g., email, SNS, LINE) and the date and time of receipt. The information selected by the user is sent to the server by pressing the send button. The server records this information in the user's profile.

[0937] Travel photo generation

[0938] At the date and time set by the user, the server retrieves a background image of the selected travel destination from the database. The server then places the user's avatar naturally in the retrieved background image and uses a deep learning model to calculate the pose. Shadows and lighting are also adjusted to ensure consistency with the background. The resulting composite image is a realistic travel photo.

[0939] Send travel photos

[0940] The server sends the generated travel photos using the method specified by the user. For example, for email, it uses an email generation API such as SendGrid, and for SNS or LINE, it uses the corresponding API (e.g. Facebook Graph API, LINE Messaging API). The travel photos are scheduled to be sent at the specified time.

[0941] Providing additional options

[0942] Users can select additional features on the options page. For example, they can choose photos of luxurious cuisine, images of accommodations, or photos of people with celebrities. Once a user selects an option, the server saves it in the user's profile and, once payment is confirmed, starts the corresponding image generation process. For example, if a photo of a person with a celebrity is selected, an image is retrieved from a celebrity image database and naturally combined with the user's avatar.

[0943] Specific examples and prompts

[0944] Specific operation example

[0945] For example, let's say a user uploads 10 photos of their face, selects "Eiffel Tower in Paris" as their destination, and wants to receive the photos every morning at 8:00 AM.

[0946] The server uses facial recognition algorithms to extract facial features and uses deep learning models to generate an avatar for the user.

[0947] The user selects "Eiffel Tower in Paris" and sets up the app to receive photos every morning at 8am.

[0948] At the set time, the server retrieves a background image of the Eiffel Tower from the database and synthesizes it onto the user's avatar.

[0949] The server will send the generated travel photos to the specified email address every morning at 8:00.

[0950] Prompt Sentence Examples

[0951] Example prompt 1:

[0952] "If a user uploads a photo of their face, sets their destination as 'Eiffel Tower in Paris', and selects to receive the photo every morning at 8am, what does the server do?"

[0953] Example output 1:

[0954] If a user uploads 10 photos of their face, sets their destination to "The Eiffel Tower in Paris," and selects to receive travel photos every morning at 8:00, the server will perform the following process. First, it extracts facial features from the user's photo and generates an avatar for the user using a deep learning model. Next, it retrieves a background image of the Eiffel Tower from the database at the specified date and time, and composites the user's avatar onto the background image. Finally, it sends the composite photo to the user by email every morning at 8:00.

[0955] This system can provide users with a realistic travel experience, which can be shared and enjoyed on social media. Specific hardware requirements include a high-performance image processing engine on the server side and GPU resources for the deep learning model. Software requirements include facial recognition algorithms (e.g., OpenCV) and deep learning libraries (e.g., TensorFlow, PyTorch).

[0956] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0957] Step 1:

[0958] Users access the website or app, register an account, and log in. Next, they select 10 photos of their face on the face photo upload screen and send them to the server.

[0959] Input: 10 photos of your face (JPEG or PNG format, minimum 400x400 pixels)

[0960] Output: Image data uploaded to the server

[0961] Specific operation: Select 10 face photos from the user's device through file selection and send them to the server as an HTTP POST request.

[0962] Step 2:

[0963] The server applies a facial recognition algorithm (e.g., OpenCV) to the 10 facial photos received to extract facial features, specifically identifying the positions and shapes of the eyes, nose, mouth, etc.

[0964] Input: 10 user-uploaded face photos

[0965] Output: Extracted facial feature data (position and shape information)

[0966] Specific operation: A facial recognition algorithm is applied on the server side to extract feature points from each facial photograph, such as the facial contours, eye positions, nose positions, and mouth positions, and obtain these as coordinate data.

[0967] Step 3:

[0968] The server generates a virtual character (avatar) for the user using a deep learning model (e.g., a person generation model using PyTorch) based on the facial feature data. The generated avatar is saved in the user's profile information.

[0969] Input: Extracted facial feature data

[0970] Output: Digital virtual character (avatar)

[0971] How it works: Facial feature data is fed into the deep learning model as input, and a virtual character resembling the user is generated based on that data. The generated virtual character data is then saved in the user's profile information.

[0972] Step 4:

[0973] The user moves to the travel destination selection screen, selects the desired travel destination from the map interface or list, and sets the method of sending the travel photos (e.g., email, SNS, LINE) and the date and time of receipt.

[0974] Input: User-selected travel destination information, sending method, and date and time of receipt

[0975] Output: Configuration information sent to the server

[0976] Specific operation: Click on a travel destination on the map interface or select it from the list, then set the sending method and receiving date and time using the drop-down menu or calendar UI. Pressing the setting button sends this information to the server.

[0977] Step 5:

[0978] The server records the user's selection information (travel destination, transmission method, date and time of receipt) in the profile.

[0979] Input: Selections received from the user

[0980] Output: Recorded user settings information

[0981] What happens: The server stores the received selection data in the user profile database for future reference.

[0982] Step 6:

[0983] At the set date and time, the server retrieves the background image of the travel destination selected by the user from the database.

[0984] Input: User selection information, background image data in the database

[0985] Output: The background image obtained.

[0986] Specific operation: For example, search and retrieve background images of a specified travel destination from cloud storage such as Amazon S3.

[0987] Step 7:

[0988] The server uses deep learning models to calculate the pose of the virtual character, placing it naturally on the background image, and also adjusts the shadows and lighting to ensure consistency with the background.

[0989] Input: background image, virtual character

[0990] Output: Synthesized, realistic travel photos

[0991] How it works: It uses deep learning models to place virtual characters naturally against backgrounds and adjust shadows and light intensity to generate composite photos.

[0992] Step 8:

[0993] The server sends the generated travel photos using the method specified by the user (email, SNS, LINE, etc.).

[0994] Input: Composite travel photos, user's sending method settings

[0995] Output: Travel photos sent to the user

[0996] Specific operation: For email, the SendGrid API is used to send emails with travel photos attached, and for SNS and LINE, the corresponding API is used to send photos along with the message.

[0997] Step 9:

[0998] Users can select additional features on the options page, such as photos of luxurious food or photos of themselves with celebrities.

[0999] Input: Additional feature selection information

[1000] Output: Configuration information for additional features sent to the server

[1001] Specific operation: Select the items displayed on the options page using checkboxes or drop-down menus, press the Settings button, and send the selected options to the server. The selected options are recorded and managed on the server side.

[1002] (Application example 1)

[1003] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1004] Today, users enjoy virtual travel experiences through various means, but these technologies lack realism and do not fully satisfy users. Furthermore, it is difficult for users to easily and effectively create digital content that reflects their own faces and share it via social media, email, etc. This creates a demand for personalized, interactive experiences. Therefore, a system is needed that can generate a realistic avatar based on a user's facial photograph, naturally combine it with the background of a specified travel destination, and deliver it in a specified manner and at a specified time.

[1005] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1006] In this invention, the server includes means for receiving a user's facial photograph, means for generating a user avatar based on the received facial photograph using a facial recognition algorithm, means for acquiring a background image based on a travel destination specified by the user, means for synthesizing the acquired background image with the generated avatar to create a realistic travel photograph, means for calculating a natural position and pose of the synthesized image using a generative AI model to create a travel photograph, and means for transmitting the created travel photograph in a manner and at a time specified by the user. This allows users to have a personalized and interactive virtual travel experience in a highly realistic way and easily share it on various digital platforms.

[1007] "User" means an individual who uses the system to upload a photo of themselves and enjoy a virtual travel experience.

[1008] A "face photo" is a digital image of the user's face, and serves as the basic data for generating an avatar.

[1009] A "facial recognition algorithm" is a computational method for extracting facial features such as eyes, nose, and mouth from a received photograph of a face and generating a digital avatar.

[1010] An "avatar" is a digital virtual persona created based on a user's facial features using facial recognition algorithms.

[1011] "Travel destination" refers to information about a place that a user specifies as the destination of a virtual trip in the system.

[1012] A "background image" is a digital image of scenery or tourist spots at a travel destination designated by the user, which is combined with the avatar.

[1013] "Generative AI model" refers to the deep learning model used to calculate natural-looking placement and pose for synthesized images.

[1014] "Transmission means" refers to tools or APIs that allow users to send the generated travel photos in a manner specified by the user (e.g., email, SNS, LINE, etc.).

[1015] "Image compositing" is the process of naturally combining background images and avatars to create realistic travel photos.

[1016] "Transmission method" refers to the communication means selected by the user for transmitting travel photos.

[1017] "Realism" refers to the quality of the synthesized images, making them appear as if they were taken during an actual trip.

[1018] This invention is a system that generates an avatar based on a user's facial photograph, combines it with a background image of a travel destination selected by the user to generate a realistic travel photograph, and transmits it in a specified manner. This system is implemented in the following steps.

[1019] First, a user accesses the website or application from a smartphone or PC, registers, and logs in. Next, the user selects 10 facial photos on the face photo upload screen and submits them to the server. These facial photos are verified to be in the correct format and resolution.

[1020] The server processes the received facial photo using a facial recognition algorithm (for example, using a deep learning library such as OpenCV or TensorFlow) to extract facial features (eyes, nose, mouth, etc.). Based on the extracted facial features, a deep learning model is used to generate an avatar for the user. This avatar is then stored in the user's profile.

[1021] Next, the user accesses the travel destination selection screen and selects their desired travel destination. Destination selection can be done using a map interface or a list of options. The user also selects the method of sending the generated travel photos (email, SNS, LINE, etc.) and the date and time of sending. This selection information is sent to the server and recorded in the user profile.

[1022] At the set date and time, the server retrieves a background image of the selected travel destination from the database. The generated avatar and background image are then composited using a deep learning model to create a natural position and pose. This process uses a generative AI model, and the composite image is then generated as a realistic travel photo.

[1023] Next, the realistic travel photos are sent in the way the user specified. For example, in the case of email, an email generation API is used, and in the case of SNS or LINE, messages are sent through the respective APIs. The travel photos are delivered at the specified time.

[1024] Additionally, users can select additional options, such as images of specific local items or meals, which are also saved in the user profile and the image generation process begins after payment confirmation.

[1025] As a concrete example, consider a case where a user selects the Eiffel Tower in Paris as a travel destination and requests to receive travel photos by email every morning at 8:00. In this case, the server will synthesize the background image of the Eiffel Tower with the user's avatar at the set time and send the travel photos by email every morning at 8:00. Throughout this process, a generative AI model is used to ensure natural synthesis and placement.

[1026] Prompt Sentence Examples

[1027] "Generate an avatar from a photo of the user's face and overlay it with a travel photo of the Eiffel Tower in the background. Background image: path / to / paris_eiffel.jpg Facial features: {Eye position: (x1, y1), Nose position: (x2, y2), Mouth shape: (x3, y3)}"

[1028] In this way, users can have a realistic virtual travel experience and easily share their travel photos on various digital platforms.

[1029] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1030] Step 1:

[1031] A user accesses a website or application using a smartphone or PC. First, the user registers and logs in to the site. Next, the user selects 10 facial photos on the facial photo upload screen and sends them to the server. The input is the 10 facial photos, and the output is the photo data stored on the server.

[1032] Step 2:

[1033] The server applies a facial recognition algorithm to the received facial photo. Specifically, it uses deep learning libraries such as OpenCV and TensorFlow to extract facial features (the position and shape of the eyes, nose, and mouth). The input is the facial photo data, and the output is data that quantifies the facial features.

[1034] Step 3:

[1035] The server uses a generative AI model to generate a digital avatar for the user based on the facial feature data extracted by the facial recognition algorithm. This avatar is virtual person data that faithfully reflects the user's facial features. The input is facial feature data, and the output is a digital avatar.

[1036] Step 4:

[1037] The user accesses the travel destination selection screen and selects the desired travel destination from a map interface or list. They also specify the method of sending the travel photos (email, SNS, LINE, etc.) and the date and time of receipt. The input is the user's selected travel destination and sending method information, and the output is the selected information recorded on the server.

[1038] Step 5:

[1039] At the set date and time, the server retrieves the background image of the travel destination selected by the user from the database. It then composites the background image with the user's avatar, using a generative AI model to calculate natural placement and poses to create realistic travel photos. The input is the background image of the travel destination and the user's digital avatar, and the output is the composite travel photo.

[1040] Step 6:

[1041] The server sends the generated travel photos in the way specified by the user. For example, in the case of email, it uses the email generation API, and in the case of SNS or LINE, it sends through the respective API. The input is the travel photo data and sending setting information, and the output is the travel photo sent to the user.

[1042] Step 7:

[1043] When a user selects an add-on feature, the server executes the corresponding image generation process, for example, adding an image of a specific local item or meal to a travel photo. This process is also performed using a generative AI model. The input is the selection information for the add-on option, and the output is the travel photo with the option added.

[1044] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1045] The system of this invention generates a traveling avatar based on a facial photo provided by the user, generates a composite image based on the travel destination selected by the user, and transmits it in a specified manner. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to suggest and select travel destinations and optional images based on the user's emotional state. Specific embodiments of the invention are described below.

[1046] 1. Uploading a user's photo

[1047] The user accesses the website or app and logs in. After logging in, the user selects 10 facial photos on the face photo upload screen and sends them to the server.

[1048] 2. Facial Photo Processing and Avatar Generation

[1049] The server applies a facial recognition algorithm to the received facial photo to extract facial features (the position and shape of the eyes, nose, and mouth), then uses a deep learning model to generate an avatar for the user, a digital virtual persona that faithfully reflects the user's facial features, and the generated avatar is saved in the user's profile.

[1050] 3. Emotion Recognition and Analysis

[1051] In addition to a photo of their face, the user provides the emotion engine with a photo of their facial expression and voice data. The emotion engine then analyzes the user's emotional state based on the data provided by the user. The analysis results are classified into emotion categories such as "joy," "sadness," "surprise," and "anger."

[1052] 4. Select your destination and set it up

[1053] Users access the travel destination selection screen and select their desired travel destination from a map interface or list. If the user requests automatic suggestions based on their emotional state, the system will suggest the most suitable travel destination based on the analysis results of the emotion engine. Users also set the sending method (email, SNS, LINE, etc.) and reception time, and press the "Settings Complete" button once the settings are complete.

[1054] 5. Travel Photo Generation

[1055] At the set date and time, the server retrieves a background image of the user's travel destination from the database. The server then composites the user's avatar onto the background image, using a deep learning model to calculate natural placement and pose. The composite image is then generated as a realistic travel photo.

[1056] 6. Send travel photos

[1057] The server sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.). For email, an email generation API is used, and for SNS or LINE, a message is sent through the corresponding API. The travel photos are delivered at the specified time.

[1058] 7. Providing additional options

[1059] Users can select additional features on the options page, such as photos of rich cuisine, luxurious accommodations, or photos of people with celebrities. The selected options are saved directly under the user's control, and the corresponding image generation process begins after payment is confirmed.

[1060] Specific examples

[1061] For example, let's consider a case where a user uploads a photo of their face and the emotion engine recognizes the emotion "joy." Based on the analysis results of the emotion engine, the server suggests "The Eiffel Tower in Paris" as the best travel destination. If the user accepts this suggestion and sets the sending method to email and the receiving time to 8:00 a.m. every morning, the server will combine a background image of the Eiffel Tower with the user's avatar, generating and sending a realistic travel photo every morning at 8:00 a.m. If the user selects a photo of "rich French cuisine" as an additional option, that image will also be sent.

[1062] In this way, by utilizing the system of the present invention, users can easily obtain realistic travel photos based on their emotional state, which can be enjoyed on social media and other platforms.

[1063] The processing flow will be explained below.

[1064] Step 1:

[1065] The user accesses the website or app and logs in. After logging in, the user selects 10 facial photos on the face photo upload screen, as well as facial expression photos and voice data for recognizing their emotional state, and sends these to the server.

[1066] Step 2:

[1067] The device collects the facial photo file and emotion recognition data selected by the user and uploads them to a server via the network.

[1068] Step 3:

[1069] The server receives the uploaded facial photo and emotion recognition data, checks the file format and resolution of the facial photo, and converts it into a format suitable for the facial recognition algorithm.

[1070] Step 4:

[1071] The server uses a facial recognition algorithm to extract facial features (position and shape of eyes, nose, and mouth).

[1072] Step 5:

[1073] The server uses a deep learning model (e.g., StyleGAN) to generate an avatar for the user based on the extracted features.

[1074] Step 6:

[1075] The server stores the generated avatar in the user profile.

[1076] Step 7:

[1077] The server uses an emotion engine to analyze the user's emotional state based on facial expressions and voice data, and classifies the results into emotion categories such as "happiness," "sadness," "surprise," and "anger."

[1078] Step 8:

[1079] The server will suggest the most suitable travel destination for the user based on the analysis results of the emotion engine. For example, if the emotion of "joy" is detected, the server will suggest the Eiffel Tower in Paris.

[1080] Step 9:

[1081] The user accesses the travel destination selection screen and selects the suggested travel destination or their own desired travel destination. They also set the sending method (email, SNS, LINE, etc.) and reception time. Once the settings are complete, they press the "Settings Complete" button.

[1082] Step 10:

[1083] The terminal collects information on the travel destination and transmission method selected by the user and transmits it to the server via the network.

[1084] Step 11:

[1085] The server processes the received configuration information and stores it in a user profile.

[1086] Step 12:

[1087] When the set date and time arrives, the server retrieves a background image of the travel destination specified by the user from the database.

[1088] Step 13:

[1089] The server then composites the user's avatar with the captured background image, using a deep learning model (e.g., OpenPose) to calculate natural placement and pose.

[1090] Step 14:

[1091] The server generates the composite image as the final travel photo.

[1092] Step 15:

[1093] The server sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.).

[1094] Step 16:

[1095] The device receives notifications via email, SNS, or LINE and displays the image to the user.

[1096] Step 17:

[1097] Users can access an options page and select additional options such as photos of rich cuisine, photos of luxurious accommodations, or photos of themselves with celebrities.

[1098] Step 18:

[1099] The terminal collects the selected option information and billing information and transmits them to the server via the network.

[1100] Step 19:

[1101] The server processes the received option and billing information and confirms the user's purchase.

[1102] Step 20:

[1103] The server initiates an image generation process according to the additional options, and generates and transmits an image corresponding to the specified date and time.

[1104] By going through these specific steps, users can receive realistic travel photos based on their own emotional state, which they can enjoy on social media and other platforms.

[1105] Example 2

[1106] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1107] In a system for generating realistic travel photos, there is a need for a means to reflect the user's emotional state and efficiently and effectively set and save travel destination suggestions and subsequent transmission methods.In addition, it has been pointed out that conventional systems need not only to synthesize background images and avatars based on the user's selection, but also to generate optional images to provide further enjoyment.

[1108] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving a user's facial photograph, means for generating a user's avatar based on the received facial photograph using a facial recognition algorithm, means for acquiring a background image based on a travel destination designated by the user, means for creating a realistic travel photograph by combining the acquired background image with the generated avatar, means for transmitting the created travel photograph in a manner designated by the user, means for analyzing and recognizing the user's emotional state and suggesting a travel destination based on the analysis results, and means for setting and saving a transmission method and transmission time. This makes it possible to suggest travel destinations and generate optional images that reflect the user's emotional state, and further to transmit the travel photograph in a manner designated by the user at a set date and time.

[1109] A "user" is an individual who uses the system to upload a photo of themselves and select a travel destination and optional images.

[1110] The "server" is a computer system that receives a user's facial photograph, generates an avatar using a facial recognition algorithm, obtains background images of the travel destination, and analyzes the user's emotional state.

[1111] A "face recognition algorithm" is a technology for extracting the position and shape of the eyes, nose, and mouth from a facial photo provided by the user.

[1112] An "avatar" is a digital virtual persona that reflects a user's facial features and is generated based on facial recognition algorithms.

[1113] A "background image" is an image of scenery or famous places at a travel destination designated by the user.

[1114] "Compositing" is the process of combining a background image with a generated avatar to create a realistic travel photo.

[1115] "Emotional state" refers to psychological states such as "joy," "sadness," "surprise," and "anger" that are analyzed from facial expression photos and voice data provided by the user.

[1116] "Suggestion" refers to the act of the server suggesting the most suitable travel destination to the user based on the analysis results of the emotion engine.

[1117] "Optional images" are images added by user selection, including images of travel destinations, specific items, and meals.

[1118] "Sending method" refers to the means by which the generated travel photos are delivered to the user, and includes email, SNS, LINE, etc.

[1119] "Send time" is the designated time for sending the generated travel photos and optional images to the user.

[1120] The system of this invention generates a traveling avatar based on a facial photograph provided by the user, suggests optimal travel destinations based on the user's emotional state, generates a composite image, and transmits it in a specified manner. Specific embodiments are described below.

[1121] First, the user accesses the website or app, logs in, selects 10 photos of their face on the face photo upload screen, and sends them to the server. At this point, the user can upload the photos from the app using their smartphone.

[1122] The server reviews the received facial photo and applies facial recognition algorithms such as OpenCV or Dlib to extract facial features. This facial recognition process specifically analyzes the position and shape of the eyes, nose, and mouth in the photo. It then uses a deep learning model such as GAN to generate an avatar that reflects the user's facial features. This avatar is then saved in the user's profile for further processing.

[1123] The user also provides the server with a photo of their facial expression and voice data. The server's emotion engine then analyzes the user's emotional state based on this data. This emotion analysis uses emotion analysis APIs from Microsoft Azure or Google Cloud, for example. Based on the analysis results, the user's emotional state is classified into emotion categories such as "happiness," "sadness," "surprise," and "anger."

[1124] Next, the user can access the travel destination selection screen and select their desired travel destination. The system automatically suggests the most suitable travel destination based on the analysis results of the emotion engine. For example, if the emotional state is "joy," the system will suggest a travel destination that matches the emotion, such as Paris, France. The user can accept the suggested travel destination or choose a different one by themselves. At this time, the user sets the sending method (email, SNS, LINE, etc.) and reception time, and presses the "Settings Complete" button. These settings are then saved on the server.

[1125] At the set date and time, the server retrieves a background image of the travel destination from a database. This background image is taken from a common photo library such as Google Images. The server then uses a deep learning model to naturally blend the user's avatar into the background image, generating a realistic travel photo. Pose Estimation technology is used to properly position the avatar against the background image.

[1126] Finally, the server sends the generated travel photos via the method specified by the user: via email using the SendGrid API, or via the corresponding API for social media or LINE. The generated images are sent at the time specified by the user.

[1127] Users can also select additional features on the options page, such as photos of rich cuisine, luxurious accommodations, and photos of two people with celebrities. After the user confirms their selection and payment, the image generation process will begin and the generated optional images will be sent.

[1128] As a concrete example, if a user uploads a photo of their face and the emotion engine recognizes the emotion "joy," the system will suggest "The Eiffel Tower in Paris" as a travel destination, and the user will accept this suggestion and set the reception time to 8:00 a.m. every morning. The server will then combine the background image of the Eiffel Tower with the user's avatar, generate a realistic travel photo every morning, and send it by email. If the user selects a photo of "rich French cuisine" as an additional option, that image will also be sent.

[1129] Example prompt sentence:

[1130] The following prompts explain system behavior to the generative AI model:

[1131] Your task is to generate a program for a system that, based on 10 facial photos provided by the user, will combine the avatar with a background image of the travel destination selected by the user when the emotion engine recognizes the emotion "joy." The process will be as follows:

[1132] 1. Extract facial features using face recognition algorithm (OpenCV / Dlib).

[1133] 2. Generate a user avatar using GAN.

[1134] 3. Analyze user emotions using an emotion engine (Microsoft Azure / Google Cloud).

[1135] 4. Composite avatar and background image based on suggested travel destination (e.g., Eiffel Tower in Paris).

[1136] 5. Using the email sending API (SendGrid), the composite photo is sent every morning at 8:00.

[1137] In this way, by using the system of the present invention, realistic travel photos based on the user's emotional state can be easily provided.

[1138] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1139] Step 1:

[1140] A user accesses a website or app, logs in, and proceeds to the face photo upload screen. The user selects 10 face photos from the photo gallery and presses the upload button. This operation sends the face photos to the server. The input is the user's face photo, and the output is the face photo data sent to the server.

[1141] Step 2:

[1142] The server checks the received face photo and applies a facial recognition algorithm (OpenCV or Dlib) to extract facial features (the position and shape of the eyes, nose, and mouth). The algorithm converts the image data to grayscale and creates a list of detected feature points. The input is the face photo data, and the output is the extracted facial feature data.

[1143] Step 3:

[1144] The server generates a user avatar using a GAN (generative adversarial network). Using facial feature data as input, the GAN model generates a virtual avatar image based on the learning results. The generated avatar is saved in the user profile. The input is facial feature data, and the output is the generated avatar image.

[1145] Step 4:

[1146] The user then provides the server with photographs of facial expressions and voice data. The server's emotion engine then uses this data to analyze the user's emotional state. Emotion analysis APIs from Microsoft Azure and Google Cloud are used for the emotion analysis. The analysis results are categorized into "happiness," "sadness," "surprise," and "anger" and stored on the server. The input is photographs of facial expressions and voice data, and the output is analyzed emotional state data.

[1147] Step 5:

[1148] The user can access the travel destination selection screen and select their desired travel destination. The server will suggest the most suitable travel destination based on the analysis results of the emotion engine. After the user accepts the suggested travel destination or selects another one by themselves, they set the sending method (email, SNS, LINE, etc.) and reception time. These settings are saved on the server. The input is the user's emotional state and travel destination selection data, and the output is the saved setting data.

[1149] Step 6:

[1150] At the set date and time, the server retrieves a background image of the travel destination from a database. This background image can be from Google Images or a user's own photo library. The server then uses a deep learning model (using Pose Estimation technology) to naturally blend the user's avatar into the background image. The input is the background image and the avatar image, and the output is the blended travel photo.

[1151] Step 7:

[1152] The server sends the generated travel photos in the manner specified by the user. The SendGrid API is used for email transmission, and the corresponding APIs for SNS and LINE are used. The travel photos are sent to the user at the specified time. The input is the composited travel photos and transmission setting data, and the output is the sent travel photos.

[1153] Step 8:

[1154] Users can select additional features on the options page. These include photos of rich cuisine, luxurious accommodations, and photos of people with celebrities. After the user selects these options and confirms the payment, the image generation process begins. The input is the option selection data, and the output is the generated option image.

[1155] In this way, each processing step operates in cooperation with one another, thereby realizing a system that provides realistic travel photos based on the user's emotional state.

[1156] (Application example 2)

[1157] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1158] Conventional avatar generation and travel photo creation systems lack the ability to suggest travel destinations that take the user's emotional state into account, resulting in the inability to provide the optimal travel experience for the user. Furthermore, simply synthesized photos alone are unlikely to achieve the realism and personalized satisfaction that users expect. Therefore, a new system that reflects the user's emotions and provides a more personalized travel experience is needed.

[1159] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1160] In this invention, the server includes means for receiving a user's facial photograph, means for generating a user avatar based on the received facial photograph using a facial recognition algorithm, means for acquiring a background image based on a travel destination designated by the user, means for synthesizing the acquired background image with the generated avatar to create a realistic travel photograph, means for transmitting the created travel photograph in a manner designated by the user, means for analyzing the user's emotional state using an emotion recognition engine, and means for suggesting travel destinations based on the user's emotional state. This makes it possible to suggest optimal travel destinations taking the user's emotional state into consideration and create realistic travel photographs.

[1161] "User" refers to an individual who uses this system.

[1162] "Facial photo" refers to image data provided by a user that is taken of the user's own face using a camera or the like.

[1163] A "facial recognition algorithm" refers to a computational method that extracts features such as the eyes, nose, and mouth from a photograph of a face and identifies and recognizes a specific face.

[1164] An "avatar" refers to a virtual digital persona generated based on a photograph of a user's face.

[1165] "Travel destination" refers to a destination where a user wishes to take a virtual trip.

[1166] "Background image" refers to image data of scenery and famous places at travel destinations.

[1167] "Compositing" refers to the process of overlaying a user's avatar with a background image to generate a single image.

[1168] "Realistic travel photos" refer to images that are synthesized to make the user's avatar appear as if they are actually at the travel destination.

[1169] An "emotion recognition engine" refers to a system that analyzes and classifies a user's emotional state based on a photo of their face and voice data.

[1170] "Emotional state" refers to a user's current psychological state or mood.

[1171] "Suggestion" refers to the system selecting and presenting appropriate travel destinations and options to the user.

[1172] "Sending" refers to the process of delivering the generated travel photos via the communication means selected by the user.

[1173] This system generates a traveling avatar using a user's facial photograph, and uses an emotion recognition engine to suggest optimal travel destinations based on the user's emotional state. It also creates realistic travel photos by combining the user's avatar with background images of the travel destination, and sends them in a manner specified by the user, providing a virtual travel experience.

[1174] The system mainly operates in the following steps:

[1175] 1. Uploading a user's photo

[1176] Users access a dedicated application using a smartphone or PC and log in. After logging in, they select multiple facial photos on the facial photo upload screen and send them to the server.

[1177] 2. Facial Photo Processing and Avatar Generation

[1178] The server applies a facial recognition algorithm to the received face photo to extract facial features (the position and shape of the eyes, nose, and mouth). It then uses a generative AI model to generate an avatar for the user. Leveraging a deep learning model, a digital persona is created that faithfully reflects the user's facial features, and this avatar is stored in the user's profile.

[1179] 3. Emotion Recognition and Analysis

[1180] Users provide facial photos, facial expression photos, and voice data to the emotion recognition engine, which uses technologies such as DeepFace to analyze the user's emotional state and classify the results into categories such as "happiness," "sadness," "surprise," and "anger."

[1181] 4. Select your destination and set it up

[1182] Users can select their desired travel destination from a map or list through the system interface. They can also choose destinations suggested by the emotion recognition engine. If they like the suggested destination, they confirm the setting. They can also set the sending method (email, SNS, etc.) and the reception time.

[1183] 5. Travel Photo Generation

[1184] At the set date and time, the server retrieves a background image of the travel destination specified by the user from the database, then combines the generated avatar with the background image, calculating natural positioning and poses to generate a realistic travel photo.

[1185] 6. Send travel photos

[1186] The server then sends the generated travel photos via the method specified by the user: via email using a dedicated email generation API, or via the corresponding API for social media or text messages.

[1187] 7. Providing additional options

[1188] Users can select additional features on the options page within the system, such as "rich food photos" or "luxury accommodations." The options selected by the user are saved in their profile, and the corresponding image generation process is initiated after payment confirmation.

[1189] Specific examples

[1190] For example, if a user provides a photo of a "joyed" expression to the emotion recognition engine, the system will suggest "The Eiffel Tower in Paris" as the best travel destination. If the user accepts this suggestion and sets the sending method to email and the receiving time to 8:00 a.m., the server will combine the background image of the Eiffel Tower with the user's avatar, generate a realistic travel photo, and send it every morning at 8:00 a.m. If the user also selects a photo of "rich French cuisine" as an additional option, that will also be sent.

[1191] Prompt Sentence Examples

[1192] "Create realistic travel photos by combining user-uploaded photos of your face with a background image of the Eiffel Tower in Paris."

[1193] The system allows users to virtually enjoy an optimal travel experience based on their emotional state and receive personalized content tailored to their individual needs.

[1194] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1195] Step 1:

[1196] The user logs in to a dedicated application using a smartphone or computer.

[1197] Input: Login information (user ID, password)

[1198] Processing: User authentication is performed and the user profile is obtained.

[1199] Output: A successful authentication message, and the user's profile information.

[1200] Step 2:

[1201] The user selects multiple facial photos on the facial photo upload screen and sends them to the server.

[1202] Input: Multiple face photos (e.g. 10 photos)

[1203] Processing: Your face photo is uploaded to the server and temporarily stored.

[1204] Output: Face photo upload complete message.

[1205] Step 3:

[1206] The server applies a facial recognition algorithm to the received facial photo and extracts facial features.

[1207] Input: Multiple face photos

[1208] Processing: Use libraries such as OpenCV to extract facial features such as eyes, nose, and mouth.

[1209] Output: Extracted facial feature data.

[1210] Step 4:

[1211] The server uses a deep learning model to generate an avatar for the user.

[1212] Input: Extracted facial feature data

[1213] Processing: Generate an avatar based on the user's face using a generative AI model (e.g., GAN).

[1214] Output: Generated avatar data.

[1215] Step 5:

[1216] The user provides facial expression photos and voice data to the emotion recognition engine.

[1217] Input: facial expression photos and audio data

[1218] Processing: Analyze the user's emotional state using emotion recognition engines such as DeepFace.

[1219] Output: Classification of emotional state (e.g., happy, sad, surprised, angry, etc.).

[1220] Step 6:

[1221] The server suggests optimal travel destinations based on the user's emotional state.

[1222] Input: Emotional state classification result

[1223] Processing: Randomly or algorithmically select the best travel destination from a list of travel destinations that correspond to the emotional state.

[1224] Output: Suggested travel destinations.

[1225] Step 7:

[1226] The user selects a desired travel destination on the travel destination selection screen and sets the transmission method and reception time.

[1227] Input: Travel destination selection information, sending method, receiving time

[1228] Action: Save the user's selections and settings to the database.

[1229] Output: Configuration complete message.

[1230] Step 8:

[1231] When the set date and time arrives, the server retrieves a background image of the travel destination specified by the user from the database.

[1232] Input: Travel destination selection information

[1233] Processing: Search and retrieve background images corresponding to the specified travel destination from the background image database.

[1234] Output: The obtained background image.

[1235] Step 9:

[1236] The server combines the generated avatar with the acquired background image.

[1237] Input: Avatar data, background image

[1238] Processing: Uses deep learning models to calculate natural placement and pose, and then synthesizes the avatar with the background image.

[1239] Output: Realistic travel photos.

[1240] Step 10:

[1241] The server transmits the generated travel photos in a manner designated by the user.

[1242] Input: Travel photos, sending method information

[1243] Processing: Send travel photos in the specified way using email generation API or SNS sending API.

[1244] Output: A transmission completion message.

[1245] Step 11:

[1246] The user selects additional features on the options page, and after confirming payment, the corresponding image generation process begins.

[1247] Input: Additional option selection information, billing information

[1248] Processing: Performs additional image generation processing based on selected options and saves it to the user profile.

[1249] Output: The additional image generated, and a charge completion message.

[1250] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1251] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1252] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1253] [Fourth embodiment]

[1254] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1255] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1256] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1257] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1258] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1259] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1260] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1261] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1262] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1263] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1264] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1265] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1266] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1267] The system of the present invention generates a traveling avatar based on a facial photograph provided by the user, generates a composite image based on a travel destination selected by the user, and transmits the composite image in a specified manner. Specific embodiments of the present invention are described below.

[1268] 1. Uploading a user's photo

[1269] Users visit a website or app, register, and log in. They then select 10 photos of their face on a face photo upload screen and submit them to the server, which verifies that the selected images are in the correct format and resolution before submitting them.

[1270] 2. Facial Photo Processing and Avatar Generation

[1271] The server applies a facial recognition algorithm to the received facial photo to extract facial features (the position and shape of the eyes, nose, and mouth). It then uses a deep learning model to generate an avatar for the user. This avatar is a digital virtual persona that faithfully reflects the user's facial features. The generated avatar is then saved in the user's profile.

[1272] 3. Select your destination and set it up

[1273] The user accesses the travel destination selection screen and selects the desired travel destination from a map interface or list. They also set the sending method (email, SNS, LINE, etc.) and the date and time of receipt. The selection information is sent to the server by pressing the send button, and the server records it in the user profile.

[1274] 4. Travel Photo Generation

[1275] The server retrieves a background image of the travel destination selected by the user from a database at the specified date and time. It then composites the user's avatar onto the background image, using a deep learning model to calculate natural placement and pose. The composite image is then generated as a realistic travel photo.

[1276] 5. Send travel photos

[1277] The server sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.). For email, an email generation API is used, and for SNS and LINE, the message is sent through the corresponding API. The travel photos are delivered at the specified time.

[1278] 6. Providing additional options

[1279] Users can select additional features on the options page. For example, they can select rich food photos, luxurious accommodations, or photos of themselves with celebrities. The selected options are saved directly under the user's control, and after payment confirmation, the corresponding image generation process begins. This provides users with an even richer experience.

[1280] Specific examples

[1281] For example, let's say a user uploads a photo of their face, sets their destination as "Eiffel Tower in Paris," and selects to receive the photo every morning at 8:00. The server will combine the background image of the Eiffel Tower with the user's avatar at the time specified by the user, and send the resulting photo by email every morning at 8:00. This allows the user to have a realistic experience every day, as if they were visiting Paris themselves.

[1282] In this way, by utilizing the system of the present invention, users can easily obtain photos that look as if they were actually traveling, and enjoy them on social media and other platforms.

[1283] The processing flow will be explained below.

[1284] Step 1:

[1285] The user accesses the website or app and logs in. After logging in, the user selects 10 photos of their face on the face photo upload screen.

[1286] Step 2:

[1287] The terminal collects the facial photo files selected by the user and uploads them to a server via a network.

[1288] Step 3:

[1289] The server receives the uploaded facial photo, checks the file format and resolution of the facial photo, and converts it into a format suitable for the facial recognition algorithm.

[1290] Step 4:

[1291] The server uses a facial recognition algorithm to extract facial features (position and shape of eyes, nose, and mouth).

[1292] Step 5:

[1293] The server uses a deep learning model (e.g., StyleGAN) to generate an avatar for the user based on the extracted features.

[1294] Step 6:

[1295] The server stores the generated avatar in the user profile.

[1296] Step 7:

[1297] The user accesses a destination selection screen and selects a desired destination from a map interface or list.

[1298] Step 8:

[1299] The user sets the sending method (email, SNS, LINE, etc.) and the reception time. Once the settings are complete, press the "Settings Complete" button.

[1300] Step 9:

[1301] The terminal collects information on the travel destination and transmission method selected by the user and transmits it to the server via the network.

[1302] Step 10:

[1303] The server processes the received configuration information and stores it in a user profile.

[1304] Step 11:

[1305] When the set date and time arrives, the server retrieves a background image of the travel destination specified by the user from the database.

[1306] Step 12:

[1307] The server then composites the user's avatar with the captured background image, using a deep learning model (e.g., OpenPose) to calculate natural placement and pose.

[1308] Step 13:

[1309] The server generates the composite image as the final travel photo.

[1310] Step 14:

[1311] The server formats and sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.).

[1312] Step 15:

[1313] The device receives notifications via email, SNS, or LINE and displays the image to the user.

[1314] Step 16:

[1315] Users can access an options page and select additional options such as photos of rich cuisine, photos of luxurious accommodations, or photos of themselves with celebrities.

[1316] Step 17:

[1317] The terminal collects the selected option information and billing information and transmits them to the server via the network.

[1318] Step 18:

[1319] The server processes the received option and billing information and confirms the user's purchase.

[1320] Step 19:

[1321] The server initiates an image generation process according to the additional options, and generates and transmits an image corresponding to the specified date and time.

[1322] By going through these steps, users can easily receive realistic travel photos that make them feel as if they are traveling around the world.

[1323] Example 1

[1324] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1325] Conventional image editing systems make it difficult for users to easily and effectively create realistic travel photos using their own facial images. They also lack features such as generating realistic images based on specific travel destinations or automatically sending photos at specified dates and times. As a result, users had to manually edit and send photos one by one, which was time-consuming and laborious, resulting in a poor user experience.

[1326] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1327] In this invention, the server includes means for receiving a user's image, means for generating a virtual character for the user based on the received image using a face recognition algorithm, means for acquiring background data based on a travel destination specified by the user, means for synthesizing the acquired background data with the generated virtual character to create a realistic image, and means for transmitting the created image in a manner specified by the user. This allows users to easily generate realistic travel photos using their own face photos, and the photos are automatically transmitted at a specified time, allowing them to be enjoyed on social media and other platforms without any hassle.

[1328] "User" refers to any individual or entity that uses the System.

[1329] "Image" refers to visual information stored or displayed in digital form.

[1330] A "facial recognition algorithm" is a computational method for detecting and identifying facial features in an image.

[1331] "Virtual character" refers to a digitally generated person based on the user's facial features.

[1332] "Travel destination" refers to a place that the user would like to virtually visit.

[1333] "Background data" refers to image and video information corresponding to a travel destination.

[1334] "Synthesis" refers to the process of integrating multiple images or data to generate a single image or data.

[1335] "Realistic images" refer to images that are synthesized to appear as if they exist in reality.

[1336] "Transmission" refers to the act of sending the generated image or data to an external party in a manner designated by the user.

[1337] "Additional selection menu" refers to optional functions and content provided in addition to basic functions.

[1338] "Setting" refers to the act of saving and making available information such as the user's selected travel destination and transmission method.

[1339] "Predetermined information" refers to specific information such as the travel destination, transmission method, and transmission date and time that the user has set in advance.

[1340] The system of the present invention generates a traveling avatar based on a facial photograph provided by the user, generates a composite image based on the travel destination selected by the user, and transmits the composite image in a specified manner. Specific embodiments of the invention are described below.

[1341] Uploading a user's face photo

[1342] A user accesses a website or app. They register an account and log in. They then select 10 photos of their face on the face photo upload screen and send them to the server. The image files must be in JPEG or PNG format with a resolution of at least 400x400 pixels. The server checks the format and resolution of the received images and returns an error message if they are inappropriate.

[1343] Facial photo processing and avatar generation

[1344] After receiving the uploaded face photo, the server uses a facial recognition algorithm (e.g., OpenCV) to extract facial features. Specifically, it identifies facial landmarks such as the position and shape of the eyes, nose, and mouth. It then uses a deep learning model (e.g., a person generation model using PyTorch) to generate a digital avatar based on the user's facial features. The generated avatar is then saved in the user's profile information.

[1345] Destination selection and settings

[1346] The user is taken to a travel destination selection screen. Here, they can select their desired travel destination from a map interface or a list. They can also set how they want to send their travel photos (e.g., email, SNS, LINE) and the date and time of receipt. The information selected by the user is sent to the server by pressing the send button. The server records this information in the user's profile.

[1347] Travel photo generation

[1348] At the date and time set by the user, the server retrieves a background image of the selected travel destination from the database. The server then places the user's avatar naturally in the retrieved background image and uses a deep learning model to calculate the pose. Shadows and lighting are also adjusted to ensure consistency with the background. The resulting composite image is a realistic travel photo.

[1349] Send travel photos

[1350] The server sends the generated travel photos using the method specified by the user. For example, for email, it uses an email generation API such as SendGrid, and for SNS or LINE, it uses the corresponding API (e.g. Facebook Graph API, LINE Messaging API). The travel photos are scheduled to be sent at the specified time.

[1351] Providing additional options

[1352] Users can select additional features on the options page. For example, they can choose photos of luxurious cuisine, images of accommodations, or photos of people with celebrities. Once a user selects an option, the server saves it in the user's profile and, once payment is confirmed, starts the corresponding image generation process. For example, if a photo of a person with a celebrity is selected, an image is retrieved from a celebrity image database and naturally combined with the user's avatar.

[1353] Specific examples and prompts

[1354] Specific operation example

[1355] For example, let's say a user uploads 10 photos of their face, selects "Eiffel Tower in Paris" as their destination, and wants to receive the photos every morning at 8:00 AM.

[1356] The server uses facial recognition algorithms to extract facial features and uses deep learning models to generate an avatar for the user.

[1357] The user selects "Eiffel Tower in Paris" and sets up the app to receive photos every morning at 8am.

[1358] At the set time, the server retrieves a background image of the Eiffel Tower from the database and synthesizes it onto the user's avatar.

[1359] The server will send the generated travel photos to the specified email address every morning at 8:00.

[1360] Prompt Sentence Examples

[1361] Example prompt 1:

[1362] "If a user uploads a photo of their face, sets their destination as 'Eiffel Tower in Paris', and selects to receive the photo every morning at 8am, what does the server do?"

[1363] Example output 1:

[1364] If a user uploads 10 photos of their face, sets their destination to "The Eiffel Tower in Paris," and selects to receive travel photos every morning at 8:00, the server will perform the following process. First, it extracts facial features from the user's photo and generates an avatar for the user using a deep learning model. Next, it retrieves a background image of the Eiffel Tower from the database at the specified date and time, and composites the user's avatar onto the background image. Finally, it sends the composite photo to the user by email every morning at 8:00.

[1365] This system can provide users with a realistic travel experience, which can be shared and enjoyed on social media. Specific hardware requirements include a high-performance image processing engine on the server side and GPU resources for the deep learning model. Software requirements include facial recognition algorithms (e.g., OpenCV) and deep learning libraries (e.g., TensorFlow, PyTorch).

[1366] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1367] Step 1:

[1368] Users access the website or app, register an account, and log in. Next, they select 10 photos of their face on the face photo upload screen and send them to the server.

[1369] Input: 10 photos of your face (JPEG or PNG format, minimum 400x400 pixels)

[1370] Output: Image data uploaded to the server

[1371] Specific operation: Select 10 face photos from the user's device through file selection and send them to the server as an HTTP POST request.

[1372] Step 2:

[1373] The server applies a facial recognition algorithm (e.g., OpenCV) to the 10 facial photos received to extract facial features, specifically identifying the positions and shapes of the eyes, nose, mouth, etc.

[1374] Input: 10 user-uploaded face photos

[1375] Output: Extracted facial feature data (position and shape information)

[1376] Specific operation: A facial recognition algorithm is applied on the server side to extract feature points from each facial photograph, such as the facial contours, eye positions, nose positions, and mouth positions, and obtain these as coordinate data.

[1377] Step 3:

[1378] The server generates a virtual character (avatar) for the user using a deep learning model (e.g., a person generation model using PyTorch) based on the facial feature data. The generated avatar is saved in the user's profile information.

[1379] Input: Extracted facial feature data

[1380] Output: Digital virtual character (avatar)

[1381] How it works: Facial feature data is fed into the deep learning model as input, and a virtual character resembling the user is generated based on that data. The generated virtual character data is then saved in the user's profile information.

[1382] Step 4:

[1383] The user moves to the travel destination selection screen, selects the desired travel destination from the map interface or list, and sets the method of sending the travel photos (e.g., email, SNS, LINE) and the date and time of receipt.

[1384] Input: User-selected travel destination information, sending method, and date and time of receipt

[1385] Output: Configuration information sent to the server

[1386] Specific operation: Click on a travel destination on the map interface or select it from the list, then set the sending method and receiving date and time using the drop-down menu or calendar UI. Pressing the setting button sends this information to the server.

[1387] Step 5:

[1388] The server records the user's selection information (travel destination, transmission method, date and time of receipt) in the profile.

[1389] Input: Selections received from the user

[1390] Output: Recorded user settings information

[1391] What happens: The server stores the received selection data in the user profile database for future reference.

[1392] Step 6:

[1393] At the set date and time, the server retrieves the background image of the travel destination selected by the user from the database.

[1394] Input: User selection information, background image data in the database

[1395] Output: The background image obtained.

[1396] Specific operation: For example, search and retrieve background images of a specified travel destination from cloud storage such as Amazon S3.

[1397] Step 7:

[1398] The server uses deep learning models to calculate the pose of the virtual character, placing it naturally on the background image, and also adjusts the shadows and lighting to ensure consistency with the background.

[1399] Input: background image, virtual character

[1400] Output: Synthesized, realistic travel photos

[1401] How it works: It uses deep learning models to place virtual characters naturally against backgrounds and adjust shadows and light intensity to generate composite photos.

[1402] Step 8:

[1403] The server sends the generated travel photos using the method specified by the user (email, SNS, LINE, etc.).

[1404] Input: Composite travel photos, user's sending method settings

[1405] Output: Travel photos sent to the user

[1406] Specific operation: For email, the SendGrid API is used to send emails with travel photos attached, and for SNS and LINE, the corresponding API is used to send photos along with the message.

[1407] Step 9:

[1408] Users can select additional features on the options page, such as photos of luxurious food or photos of themselves with celebrities.

[1409] Input: Additional feature selection information

[1410] Output: Configuration information for additional features sent to the server

[1411] Specific operation: Select the items displayed on the options page using checkboxes or drop-down menus, press the Settings button, and send the selected options to the server. The selected options are recorded and managed on the server side.

[1412] (Application example 1)

[1413] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1414] Today, users enjoy virtual travel experiences through various means, but these technologies lack realism and do not fully satisfy users. Furthermore, it is difficult for users to easily and effectively create digital content that reflects their own faces and share it via social media, email, etc. This creates a demand for personalized, interactive experiences. Therefore, a system is needed that can generate a realistic avatar based on a user's facial photograph, naturally combine it with the background of a specified travel destination, and deliver it in a specified manner and at a specified time.

[1415] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1416] In this invention, the server includes means for receiving a user's facial photograph, means for generating a user avatar based on the received facial photograph using a facial recognition algorithm, means for acquiring a background image based on a travel destination specified by the user, means for synthesizing the acquired background image with the generated avatar to create a realistic travel photograph, means for calculating a natural position and pose of the synthesized image using a generative AI model to create a travel photograph, and means for transmitting the created travel photograph in a manner and at a time specified by the user. This allows users to have a personalized and interactive virtual travel experience in a highly realistic way and easily share it on various digital platforms.

[1417] "User" means an individual who uses the system to upload a photo of themselves and enjoy a virtual travel experience.

[1418] A "face photo" is a digital image of the user's face, and serves as the basic data for generating an avatar.

[1419] A "facial recognition algorithm" is a computational method for extracting facial features such as eyes, nose, and mouth from a received photograph of a face and generating a digital avatar.

[1420] An "avatar" is a digital virtual persona created based on a user's facial features using facial recognition algorithms.

[1421] "Travel destination" refers to information about a place that a user specifies as the destination of a virtual trip in the system.

[1422] A "background image" is a digital image of scenery or tourist spots at a travel destination designated by the user, which is combined with the avatar.

[1423] "Generative AI model" refers to the deep learning model used to calculate natural-looking placement and pose for synthesized images.

[1424] "Transmission means" refers to tools or APIs that allow users to send the generated travel photos in a manner specified by the user (e.g., email, SNS, LINE, etc.).

[1425] "Image compositing" is the process of naturally combining background images and avatars to create realistic travel photos.

[1426] "Transmission method" refers to the communication means selected by the user for transmitting travel photos.

[1427] "Realism" refers to the quality of the synthesized images, making them appear as if they were taken during an actual trip.

[1428] This invention is a system that generates an avatar based on a user's facial photograph, combines it with a background image of a travel destination selected by the user to generate a realistic travel photograph, and transmits it in a specified manner. This system is implemented in the following steps.

[1429] First, a user accesses the website or application from a smartphone or PC, registers, and logs in. Next, the user selects 10 facial photos on the face photo upload screen and submits them to the server. These facial photos are verified to be in the correct format and resolution.

[1430] The server processes the received facial photo using a facial recognition algorithm (for example, using a deep learning library such as OpenCV or TensorFlow) to extract facial features (eyes, nose, mouth, etc.). Based on the extracted facial features, a deep learning model is used to generate an avatar for the user. This avatar is then stored in the user's profile.

[1431] Next, the user accesses the travel destination selection screen and selects their desired travel destination. Destination selection can be done using a map interface or a list of options. The user also selects the method of sending the generated travel photos (email, SNS, LINE, etc.) and the date and time of sending. This selection information is sent to the server and recorded in the user profile.

[1432] At the set date and time, the server retrieves a background image of the selected travel destination from the database. The generated avatar and background image are then composited using a deep learning model to create a natural position and pose. This process uses a generative AI model, and the composite image is then generated as a realistic travel photo.

[1433] Next, the realistic travel photos are sent in the way the user specified. For example, in the case of email, an email generation API is used, and in the case of SNS or LINE, messages are sent through the respective APIs. The travel photos are delivered at the specified time.

[1434] Additionally, users can select additional options, such as images of specific local items or meals, which are also saved in the user profile and the image generation process begins after payment confirmation.

[1435] As a concrete example, consider a case where a user selects the Eiffel Tower in Paris as a travel destination and requests to receive travel photos by email every morning at 8:00. In this case, the server will synthesize the background image of the Eiffel Tower with the user's avatar at the set time and send the travel photos by email every morning at 8:00. Throughout this process, a generative AI model is used to ensure natural synthesis and placement.

[1436] Prompt Sentence Examples

[1437] "Generate an avatar from a photo of the user's face and overlay it with a travel photo of the Eiffel Tower in the background. Background image: path / to / paris_eiffel.jpg Facial features: {Eye position: (x1, y1), Nose position: (x2, y2), Mouth shape: (x3, y3)}"

[1438] In this way, users can have a realistic virtual travel experience and easily share their travel photos on various digital platforms.

[1439] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1440] Step 1:

[1441] A user accesses a website or application using a smartphone or PC. First, the user registers and logs in to the site. Next, the user selects 10 facial photos on the facial photo upload screen and sends them to the server. The input is the 10 facial photos, and the output is the photo data stored on the server.

[1442] Step 2:

[1443] The server applies a facial recognition algorithm to the received facial photo. Specifically, it uses deep learning libraries such as OpenCV and TensorFlow to extract facial features (the position and shape of the eyes, nose, and mouth). The input is the facial photo data, and the output is data that quantifies the facial features.

[1444] Step 3:

[1445] The server uses a generative AI model to generate a digital avatar for the user based on the facial feature data extracted by the facial recognition algorithm. This avatar is virtual person data that faithfully reflects the user's facial features. The input is facial feature data, and the output is a digital avatar.

[1446] Step 4:

[1447] The user accesses the travel destination selection screen and selects the desired travel destination from a map interface or list. They also specify the method of sending the travel photos (email, SNS, LINE, etc.) and the date and time of receipt. The input is the user's selected travel destination and sending method information, and the output is the selected information recorded on the server.

[1448] Step 5:

[1449] At the set date and time, the server retrieves the background image of the travel destination selected by the user from the database. It then composites the background image with the user's avatar, using a generative AI model to calculate natural placement and poses to create realistic travel photos. The input is the background image of the travel destination and the user's digital avatar, and the output is the composite travel photo.

[1450] Step 6:

[1451] The server sends the generated travel photos in the way specified by the user. For example, in the case of email, it uses the email generation API, and in the case of SNS or LINE, it sends through the respective API. The input is the travel photo data and sending setting information, and the output is the travel photo sent to the user.

[1452] Step 7:

[1453] When a user selects an add-on feature, the server executes the corresponding image generation process, for example, adding an image of a specific local item or meal to a travel photo. This process is also performed using a generative AI model. The input is the selection information for the add-on option, and the output is the travel photo with the option added.

[1454] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1455] The system of this invention generates a traveling avatar based on a facial photo provided by the user, generates a composite image based on the travel destination selected by the user, and transmits it in a specified manner. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to suggest and select travel destinations and optional images based on the user's emotional state. Specific embodiments of the invention are described below.

[1456] 1. Uploading a user's photo

[1457] The user accesses the website or app and logs in. After logging in, the user selects 10 facial photos on the face photo upload screen and sends them to the server.

[1458] 2. Facial Photo Processing and Avatar Generation

[1459] The server applies a facial recognition algorithm to the received facial photo to extract facial features (the position and shape of the eyes, nose, and mouth), then uses a deep learning model to generate an avatar for the user, a digital virtual persona that faithfully reflects the user's facial features, and the generated avatar is saved in the user's profile.

[1460] 3. Emotion Recognition and Analysis

[1461] In addition to a photo of their face, the user provides the emotion engine with a photo of their facial expression and voice data. The emotion engine then analyzes the user's emotional state based on the data provided by the user. The analysis results are classified into emotion categories such as "joy," "sadness," "surprise," and "anger."

[1462] 4. Select your destination and set it up

[1463] Users access the travel destination selection screen and select their desired travel destination from a map interface or list. If the user requests automatic suggestions based on their emotional state, the system will suggest the most suitable travel destination based on the analysis results of the emotion engine. Users also set the sending method (email, SNS, LINE, etc.) and reception time, and press the "Settings Complete" button once the settings are complete.

[1464] 5. Travel Photo Generation

[1465] At the set date and time, the server retrieves a background image of the user's travel destination from the database. The server then composites the user's avatar onto the background image, using a deep learning model to calculate natural placement and pose. The composite image is then generated as a realistic travel photo.

[1466] 6. Send travel photos

[1467] The server sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.). For email, an email generation API is used, and for SNS or LINE, a message is sent through the corresponding API. The travel photos are delivered at the specified time.

[1468] 7. Providing additional options

[1469] Users can select additional features on the options page, such as photos of rich cuisine, luxurious accommodations, or photos of people with celebrities. The selected options are saved directly under the user's control, and the corresponding image generation process begins after payment is confirmed.

[1470] Specific examples

[1471] For example, let's consider a case where a user uploads a photo of their face and the emotion engine recognizes the emotion "joy." Based on the analysis results of the emotion engine, the server suggests "The Eiffel Tower in Paris" as the best travel destination. If the user accepts this suggestion and sets the sending method to email and the receiving time to 8:00 a.m. every morning, the server will combine a background image of the Eiffel Tower with the user's avatar, generating and sending a realistic travel photo every morning at 8:00 a.m. If the user selects a photo of "rich French cuisine" as an additional option, that image will also be sent.

[1472] In this way, by utilizing the system of the present invention, users can easily obtain realistic travel photos based on their emotional state, which can be enjoyed on social media and other platforms.

[1473] The processing flow will be explained below.

[1474] Step 1:

[1475] The user accesses the website or app and logs in. After logging in, the user selects 10 facial photos on the face photo upload screen, as well as facial expression photos and voice data for recognizing their emotional state, and sends these to the server.

[1476] Step 2:

[1477] The device collects the facial photo file and emotion recognition data selected by the user and uploads them to a server via the network.

[1478] Step 3:

[1479] The server receives the uploaded facial photo and emotion recognition data, checks the file format and resolution of the facial photo, and converts it into a format suitable for the facial recognition algorithm.

[1480] Step 4:

[1481] The server uses a facial recognition algorithm to extract facial features (position and shape of eyes, nose, and mouth).

[1482] Step 5:

[1483] The server uses a deep learning model (e.g., StyleGAN) to generate an avatar for the user based on the extracted features.

[1484] Step 6:

[1485] The server stores the generated avatar in the user profile.

[1486] Step 7:

[1487] The server uses an emotion engine to analyze the user's emotional state based on facial expressions and voice data, and classifies the results into emotion categories such as "happiness," "sadness," "surprise," and "anger."

[1488] Step 8:

[1489] The server will suggest the most suitable travel destination for the user based on the analysis results of the emotion engine. For example, if the emotion of "joy" is detected, the server will suggest the Eiffel Tower in Paris.

[1490] Step 9:

[1491] The user accesses the travel destination selection screen and selects the suggested travel destination or their own desired travel destination. They also set the sending method (email, SNS, LINE, etc.) and reception time. Once the settings are complete, they press the "Settings Complete" button.

[1492] Step 10:

[1493] The terminal collects information on the travel destination and transmission method selected by the user and transmits it to the server via the network.

[1494] Step 11:

[1495] The server processes the received configuration information and stores it in a user profile.

[1496] Step 12:

[1497] When the set date and time arrives, the server retrieves a background image of the travel destination specified by the user from the database.

[1498] Step 13:

[1499] The server then composites the user's avatar with the captured background image, using a deep learning model (e.g., OpenPose) to calculate natural placement and pose.

[1500] Step 14:

[1501] The server generates the composite image as the final travel photo.

[1502] Step 15:

[1503] The server sends the generated travel photos according to the sending method specified by the user (email, SNS, LINE, etc.).

[1504] Step 16:

[1505] The device receives notifications via email, SNS, or LINE and displays the image to the user.

[1506] Step 17:

[1507] Users can access an options page and select additional options such as photos of rich cuisine, photos of luxurious accommodations, or photos of themselves with celebrities.

[1508] Step 18:

[1509] The terminal collects the selected option information and billing information and transmits them to the server via the network.

[1510] Step 19:

[1511] The server processes the received option and billing information and confirms the user's purchase.

[1512] Step 20:

[1513] The server initiates an image generation process according to the additional options, and generates and transmits an image corresponding to the specified date and time.

[1514] By going through these specific steps, users can receive realistic travel photos based on their own emotional state, which they can enjoy on social media and other platforms.

[1515] Example 2

[1516] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1517] In a system for generating realistic travel photos, there is a need for a means to reflect the user's emotional state and efficiently and effectively set and save travel destination suggestions and subsequent transmission methods.In addition, it has been pointed out that conventional systems need not only to synthesize background images and avatars based on the user's selection, but also to generate optional images to provide further enjoyment.

[1518] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving a user's facial photograph, means for generating a user's avatar based on the received facial photograph using a facial recognition algorithm, means for acquiring a background image based on a travel destination designated by the user, means for creating a realistic travel photograph by combining the acquired background image with the generated avatar, means for transmitting the created travel photograph in a manner designated by the user, means for analyzing and recognizing the user's emotional state and suggesting a travel destination based on the analysis results, and means for setting and saving a transmission method and transmission time. This makes it possible to suggest travel destinations and generate optional images that reflect the user's emotional state, and further to transmit the travel photograph in a manner designated by the user at a set date and time.

[1519] A "user" is an individual who uses the system to upload a photo of themselves and select a travel destination and optional images.

[1520] The "server" is a computer system that receives a user's facial photograph, generates an avatar using a facial recognition algorithm, obtains background images of the travel destination, and analyzes the user's emotional state.

[1521] A "face recognition algorithm" is a technology for extracting the position and shape of the eyes, nose, and mouth from a facial photo provided by the user.

[1522] An "avatar" is a digital virtual persona that reflects a user's facial features and is generated based on facial recognition algorithms.

[1523] A "background image" is an image of scenery or famous places at a travel destination designated by the user.

[1524] "Compositing" is the process of combining a background image with a generated avatar to create a realistic travel photo.

[1525] "Emotional state" refers to psychological states such as "joy," "sadness," "surprise," and "anger" that are analyzed from facial expression photos and voice data provided by the user.

[1526] "Suggestion" refers to the act of the server suggesting the most suitable travel destination to the user based on the analysis results of the emotion engine.

[1527] "Optional images" are images added by user selection, including images of travel destinations, specific items, and meals.

[1528] "Sending method" refers to the means by which the generated travel photos are delivered to the user, and includes email, SNS, LINE, etc.

[1529] "Send time" is the designated time for sending the generated travel photos and optional images to the user.

[1530] The system of this invention generates a traveling avatar based on a facial photograph provided by the user, suggests optimal travel destinations based on the user's emotional state, generates a composite image, and transmits it in a specified manner. Specific embodiments are described below.

[1531] First, the user accesses the website or app, logs in, selects 10 photos of their face on the face photo upload screen, and sends them to the server. At this point, the user can upload the photos from the app using their smartphone.

[1532] The server reviews the received facial photo and applies facial recognition algorithms such as OpenCV or Dlib to extract facial features. This facial recognition process specifically analyzes the position and shape of the eyes, nose, and mouth in the photo. It then uses a deep learning model such as GAN to generate an avatar that reflects the user's facial features. This avatar is then saved in the user's profile for further processing.

[1533] The user also provides the server with a photo of their facial expression and voice data. The server's emotion engine then analyzes the user's emotional state based on this data. This emotion analysis uses emotion analysis APIs from Microsoft Azure or Google Cloud, for example. Based on the analysis results, the user's emotional state is classified into emotion categories such as "happiness," "sadness," "surprise," and "anger."

[1534] Next, the user can access the travel destination selection screen and select their desired travel destination. The system automatically suggests the most suitable travel destination based on the analysis results of the emotion engine. For example, if the emotional state is "joy," the system will suggest a travel destination that matches the emotion, such as Paris, France. The user can accept the suggested travel destination or choose a different one by themselves. At this time, the user sets the sending method (email, SNS, LINE, etc.) and reception time, and presses the "Settings Complete" button. These settings are then saved on the server.

[1535] At the set date and time, the server retrieves a background image of the travel destination from a database. This background image is taken from a common photo library such as Google Images. The server then uses a deep learning model to naturally blend the user's avatar into the background image, generating a realistic travel photo. Pose Estimation technology is used to properly position the avatar against the background image.

[1536] Finally, the server sends the generated travel photos via the method specified by the user: via email using the SendGrid API, or via the corresponding API for social media or LINE. The generated images are sent at the time specified by the user.

[1537] Users can also select additional features on the options page, such as photos of rich cuisine, luxurious accommodations, and photos of two people with celebrities. After the user confirms their selection and payment, the image generation process will begin and the generated optional images will be sent.

[1538] As a concrete example, if a user uploads a photo of their face and the emotion engine recognizes the emotion "joy," the system will suggest "The Eiffel Tower in Paris" as a travel destination, and the user will accept this suggestion and set the reception time to 8:00 a.m. every morning. The server will then combine the background image of the Eiffel Tower with the user's avatar, generate a realistic travel photo every morning, and send it by email. If the user selects a photo of "rich French cuisine" as an additional option, that image will also be sent.

[1539] Example prompt sentence:

[1540] The following prompts explain system behavior to the generative AI model:

[1541] Your task is to generate a program for a system that, based on 10 facial photos provided by the user, will combine the avatar with a background image of the travel destination selected by the user when the emotion engine recognizes the emotion "joy." The process will be as follows:

[1542] 1. Extract facial features using face recognition algorithm (OpenCV / Dlib).

[1543] 2. Generate a user avatar using GAN.

[1544] 3. Analyze user emotions using an emotion engine (Microsoft Azure / Google Cloud).

[1545] 4. Composite avatar and background image based on suggested travel destination (e.g., Eiffel Tower in Paris).

[1546] 5. Using the email sending API (SendGrid), the composite photo is sent every morning at 8:00.

[1547] In this way, by using the system of the present invention, realistic travel photos based on the user's emotional state can be easily provided.

[1548] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1549] Step 1:

[1550] A user accesses a website or app, logs in, and proceeds to the face photo upload screen. The user selects 10 face photos from the photo gallery and presses the upload button. This operation sends the face photos to the server. The input is the user's face photo, and the output is the face photo data sent to the server.

[1551] Step 2:

[1552] The server checks the received face photo and applies a facial recognition algorithm (OpenCV or Dlib) to extract facial features (the position and shape of the eyes, nose, and mouth). The algorithm converts the image data to grayscale and creates a list of detected feature points. The input is the face photo data, and the output is the extracted facial feature data.

[1553] Step 3:

[1554] The server generates a user avatar using a GAN (generative adversarial network). Using facial feature data as input, the GAN model generates a virtual avatar image based on the learning results. The generated avatar is saved in the user profile. The input is facial feature data, and the output is the generated avatar image.

[1555] Step 4:

[1556] The user then provides the server with photographs of facial expressions and voice data. The server's emotion engine then uses this data to analyze the user's emotional state. Emotion analysis APIs from Microsoft Azure and Google Cloud are used for the emotion analysis. The analysis results are categorized into "happiness," "sadness," "surprise," and "anger" and stored on the server. The input is photographs of facial expressions and voice data, and the output is analyzed emotional state data.

[1557] Step 5:

[1558] The user can access the travel destination selection screen and select their desired travel destination. The server will suggest the most suitable travel destination based on the analysis results of the emotion engine. After the user accepts the suggested travel destination or selects another one by themselves, they set the sending method (email, SNS, LINE, etc.) and reception time. These settings are saved on the server. The input is the user's emotional state and travel destination selection data, and the output is the saved setting data.

[1559] Step 6:

[1560] At the set date and time, the server retrieves a background image of the travel destination from a database. This background image can be from Google Images or a user's own photo library. The server then uses a deep learning model (using Pose Estimation technology) to naturally blend the user's avatar into the background image. The input is the background image and the avatar image, and the output is the blended travel photo.

[1561] Step 7:

[1562] The server sends the generated travel photos in the manner specified by the user. The SendGrid API is used for email transmission, and the corresponding APIs for SNS and LINE are used. The travel photos are sent to the user at the specified time. The input is the composited travel photos and transmission setting data, and the output is the sent travel photos.

[1563] Step 8:

[1564] Users can select additional features on the options page. These include photos of rich cuisine, luxurious accommodations, and photos of people with celebrities. After the user selects these options and confirms the payment, the image generation process begins. The input is the option selection data, and the output is the generated option image.

[1565] In this way, each processing step operates in cooperation with one another, thereby realizing a system that provides realistic travel photos based on the user's emotional state.

[1566] (Application example 2)

[1567] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1568] Conventional avatar generation and travel photo creation systems lack the ability to suggest travel destinations that take the user's emotional state into account, resulting in the inability to provide the optimal travel experience for the user. Furthermore, simply synthesized photos alone are unlikely to achieve the realism and personalized satisfaction that users expect. Therefore, a new system that reflects the user's emotions and provides a more personalized travel experience is needed.

[1569] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1570] In this invention, the server includes means for receiving a user's facial photograph, means for generating a user avatar based on the received facial photograph using a facial recognition algorithm, means for acquiring a background image based on a travel destination designated by the user, means for synthesizing the acquired background image with the generated avatar to create a realistic travel photograph, means for transmitting the created travel photograph in a manner designated by the user, means for analyzing the user's emotional state using an emotion recognition engine, and means for suggesting travel destinations based on the user's emotional state. This makes it possible to suggest optimal travel destinations taking the user's emotional state into consideration and create realistic travel photographs.

[1571] "User" refers to an individual who uses this system.

[1572] "Facial photo" refers to image data provided by a user that is taken of the user's own face using a camera or the like.

[1573] A "facial recognition algorithm" refers to a computational method that extracts features such as the eyes, nose, and mouth from a photograph of a face and identifies and recognizes a specific face.

[1574] An "avatar" refers to a virtual digital persona generated based on a photograph of a user's face.

[1575] "Travel destination" refers to a destination where a user wishes to take a virtual trip.

[1576] "Background image" refers to image data of scenery and famous places at travel destinations.

[1577] "Compositing" refers to the process of overlaying a user's avatar with a background image to generate a single image.

[1578] "Realistic travel photos" refer to images that are synthesized to make the user's avatar appear as if they are actually at the travel destination.

[1579] An "emotion recognition engine" refers to a system that analyzes and classifies a user's emotional state based on a photo of their face and voice data.

[1580] "Emotional state" refers to a user's current psychological state or mood.

[1581] "Suggestion" refers to the system selecting and presenting appropriate travel destinations and options to the user.

[1582] "Sending" refers to the process of delivering the generated travel photos via the communication means selected by the user.

[1583] This system generates a traveling avatar using a user's facial photograph, and uses an emotion recognition engine to suggest optimal travel destinations based on the user's emotional state. It also creates realistic travel photos by combining the user's avatar with background images of the travel destination, and sends them in a manner specified by the user, providing a virtual travel experience.

[1584] The system mainly operates in the following steps:

[1585] 1. Uploading a user's photo

[1586] Users access a dedicated application using a smartphone or PC and log in. After logging in, they select multiple facial photos on the facial photo upload screen and send them to the server.

[1587] 2. Facial Photo Processing and Avatar Generation

[1588] The server applies a facial recognition algorithm to the received face photo to extract facial features (the position and shape of the eyes, nose, and mouth). It then uses a generative AI model to generate an avatar for the user. Leveraging a deep learning model, a digital persona is created that faithfully reflects the user's facial features, and this avatar is stored in the user's profile.

[1589] 3. Emotion Recognition and Analysis

[1590] Users provide facial photos, facial expression photos, and voice data to the emotion recognition engine, which uses technologies such as DeepFace to analyze the user's emotional state and classify the results into categories such as "happiness," "sadness," "surprise," and "anger."

[1591] 4. Select your destination and set it up

[1592] Users can select their desired travel destination from a map or list through the system interface. They can also choose destinations suggested by the emotion recognition engine. If they like the suggested destination, they confirm the setting. They can also set the sending method (email, SNS, etc.) and the reception time.

[1593] 5. Travel Photo Generation

[1594] At the set date and time, the server retrieves a background image of the travel destination specified by the user from the database, then combines the generated avatar with the background image, calculating natural positioning and poses to generate a realistic travel photo.

[1595] 6. Send travel photos

[1596] The server then sends the generated travel photos via the method specified by the user: via email using a dedicated email generation API, or via the corresponding API for social media or text messages.

[1597] 7. Providing additional options

[1598] Users can select additional features on the options page within the system, such as "rich food photos" or "luxury accommodations." The options selected by the user are saved in their profile, and the corresponding image generation process is initiated after payment confirmation.

[1599] Specific examples

[1600] For example, if a user provides a photo of a "joyed" expression to the emotion recognition engine, the system will suggest "The Eiffel Tower in Paris" as the best travel destination. If the user accepts this suggestion and sets the sending method to email and the receiving time to 8:00 a.m., the server will combine the background image of the Eiffel Tower with the user's avatar, generate a realistic travel photo, and send it every morning at 8:00 a.m. If the user also selects a photo of "rich French cuisine" as an additional option, that will also be sent.

[1601] Prompt Sentence Examples

[1602] "Create realistic travel photos by combining user-uploaded photos of your face with a background image of the Eiffel Tower in Paris."

[1603] The system allows users to virtually enjoy an optimal travel experience based on their emotional state and receive personalized content tailored to their individual needs.

[1604] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1605] Step 1:

[1606] The user logs in to a dedicated application using a smartphone or computer.

[1607] Input: Login information (user ID, password)

[1608] Processing: User authentication is performed and the user profile is obtained.

[1609] Output: A successful authentication message, and the user's profile information.

[1610] Step 2:

[1611] The user selects multiple facial photos on the facial photo upload screen and sends them to the server.

[1612] Input: Multiple face photos (e.g. 10 photos)

[1613] Processing: Your face photo is uploaded to the server and temporarily stored.

[1614] Output: Face photo upload complete message.

[1615] Step 3:

[1616] The server applies a facial recognition algorithm to the received facial photo and extracts facial features.

[1617] Input: Multiple face photos

[1618] Processing: Use libraries such as OpenCV to extract facial features such as eyes, nose, and mouth.

[1619] Output: Extracted facial feature data.

[1620] Step 4:

[1621] The server uses a deep learning model to generate an avatar for the user.

[1622] Input: Extracted facial feature data

[1623] Processing: Generate an avatar based on the user's face using a generative AI model (e.g., GAN).

[1624] Output: Generated avatar data.

[1625] Step 5:

[1626] The user provides facial expression photos and voice data to the emotion recognition engine.

[1627] Input: facial expression photos and audio data

[1628] Processing: Analyze the user's emotional state using emotion recognition engines such as DeepFace.

[1629] Output: Classification of emotional state (e.g., happy, sad, surprised, angry, etc.).

[1630] Step 6:

[1631] The server suggests optimal travel destinations based on the user's emotional state.

[1632] Input: Emotional state classification result

[1633] Processing: Randomly or algorithmically select the best travel destination from a list of travel destinations that correspond to the emotional state.

[1634] Output: Suggested travel destinations.

[1635] Step 7:

[1636] The user selects a desired travel destination on the travel destination selection screen and sets the transmission method and reception time.

[1637] Input: Travel destination selection information, sending method, receiving time

[1638] Action: Save the user's selections and settings to the database.

[1639] Output: Configuration complete message.

[1640] Step 8:

[1641] When the set date and time arrives, the server retrieves a background image of the travel destination specified by the user from the database.

[1642] Input: Travel destination selection information

[1643] Processing: Search and retrieve background images corresponding to the specified travel destination from the background image database.

[1644] Output: The obtained background image.

[1645] Step 9:

[1646] The server combines the generated avatar with the acquired background image.

[1647] Input: Avatar data, background image

[1648] Processing: Uses deep learning models to calculate natural placement and pose, and then synthesizes the avatar with the background image.

[1649] Output: Realistic travel photos.

[1650] Step 10:

[1651] The server transmits the generated travel photos in a manner designated by the user.

[1652] Input: Travel photos, sending method information

[1653] Processing: Send travel photos in the specified way using email generation API or SNS sending API.

[1654] Output: A transmission completion message.

[1655] Step 11:

[1656] The user selects additional features on the options page, and after confirming payment, the corresponding image generation process begins.

[1657] Input: Additional option selection information, billing information

[1658] Processing: Performs additional image generation processing based on selected options and saves it to the user profile.

[1659] Output: The additional image generated, and a charge completion message.

[1660] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1661] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1662] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1663] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1664] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1665] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1666] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1667] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1668] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1669] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1670] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1671] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1672] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1673] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1674] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1675] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1676] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1677] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1678] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1679] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1680] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1681] The following is further disclosed regarding the above embodiment.

[1682] (Claim 1)

[1683] means for receiving a facial photograph of the user;

[1684] means for generating an avatar of the user based on the received facial photograph using a facial recognition algorithm;

[1685] means for acquiring a background image based on a travel destination designated by a user;

[1686] A means for synthesizing the acquired background image and the generated avatar to create a realistic travel photo;

[1687] means for transmitting the created travel photos in a manner designated by the user;

[1688] A system including:

[1689] (Claim 2)

[1690] 10. The system of claim 1, further comprising means for generating additional optional images based on the travel destination, including images of specific local items or meals.

[1691] (Claim 3)

[1692] 2. The system according to claim 1, further comprising means for storing travel destination selection information and transmission method information designated by a user, and automatically generating travel photos based on the predetermined information at a set date and time.

[1693] "Example 1"

[1694] (Claim 1)

[1695] means for receiving an image of a user;

[1696] means for generating a virtual character of the user based on the received image using a facial recognition algorithm;

[1697] means for acquiring background data based on a travel destination designated by a user;

[1698] A means for synthesizing the acquired background data with the generated virtual character to create a realistic image;

[1699] means for transmitting the created image in a manner designated by the user;

[1700] A system including:

[1701] (Claim 2)

[1702] 10. The system of claim 1, further comprising means for generating additional menu selections based on the travel destination, the menu including images of specific local items and meals.

[1703] (Claim 3)

[1704] 2. The system according to claim 1, further comprising means for storing destination selection information and transmission method information designated by a user, and automatically generating an image based on the predetermined information at a set date and time.

[1705] "Application Example 1"

[1706] (Claim 1)

[1707] means for receiving a facial photograph of the user;

[1708] means for generating an avatar of the user based on the received facial photograph using a facial recognition algorithm;

[1709] means for acquiring a background image based on a travel destination designated by a user;

[1710] A means for synthesizing the acquired background image and the generated avatar to create a realistic travel photo;

[1711] A method to create travel photos by calculating natural positions and poses using an AI model to generate synthesized images, and

[1712] means for transmitting the travel photographs created in a manner and at a time designated by the user;

[1713] A system including:

[1714] (Claim 2)

[1715] 10. The system of claim 1, further comprising means for generating additional optional images based on the travel destination, including images of specific local items or meals.

[1716] (Claim 3)

[1717] 2. The system according to claim 1, further comprising means for storing travel destination selection information and transmission method information designated by a user, and automatically generating travel photos based on the predetermined information at a set date and time.

[1718] "Example 2: Combining Emotion Engines"

[1719] (Claim 1)

[1720] means for receiving a facial photograph of the user;

[1721] means for generating an avatar of the user based on the received facial photograph using a facial recognition algorithm;

[1722] means for acquiring a background image based on a travel destination designated by a user;

[1723] A means for synthesizing the acquired background image and the generated avatar to create a realistic travel photo;

[1724] means for transmitting the created travel photos in a manner designated by the user;

[1725] A means for analyzing and recognizing the user's emotional state and suggesting travel destinations based on the analysis results;

[1726] A means for setting and saving a transmission method and a transmission time;

[1727] A system including:

[1728] (Claim 2)

[1729] 10. The system of claim 1, further comprising means for generating additional optional images based on the travel destination, including images of specific local items or meals.

[1730] (Claim 3)

[1731] The system of claim 1, further comprising means for storing travel destination suggestions and transmission method information that reflect the user's emotional state, and for automatically generating travel photos based on predetermined information at a set date and time.

[1732] "Application example 2 when combining emotion engines"

[1733] (Claim 1)

[1734] means for receiving a facial photograph of the user;

[1735] means for generating an avatar of the user based on the received facial photograph using a facial recognition algorithm;

[1736] means for acquiring a background image based on a travel destination designated by a user;

[1737] A means for synthesizing the acquired background image and the generated avatar to create a realistic travel photo;

[1738] means for transmitting the created travel photos in a manner designated by the user;

[1739] means for analyzing the emotional state of a user using an emotion recognition engine;

[1740] means for suggesting travel destinations based on the emotional state of the user;

[1741] A system including:

[1742] (Claim 2)

[1743] 10. The system of claim 1, further comprising means for generating additional optional images based on the travel destination, including images of specific local items or meals.

[1744] (Claim 3)

[1745] The system according to claim 1, further comprising means for storing travel destination selection information and transmission method information designated by a user, and automatically generating and transmitting travel photos based on the predetermined information at a set date and time. [Explanation of symbols]

[1746] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving a facial photograph of the user; means for generating an avatar of the user based on the received facial photograph using a facial recognition algorithm; means for acquiring a background image based on a travel destination designated by a user; A means for synthesizing the acquired background image and the generated avatar to create a realistic travel photo; means for transmitting the created travel photos in a manner designated by the user; A system including:

2. The system of claim 1 , further comprising means for generating additional optional images based on the travel destination, including images of specific local items or meals.

3. 2. The system according to claim 1, further comprising means for storing information on selected travel destinations and transmission methods designated by a user, and for automatically generating travel photos based on predetermined information at a set date and time.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A