system

A system that generates profile icons from facial photographs addresses user resistance by creating images that preserve user features, enhancing privacy and security, thus improving online communication.

JP2026063865APending Publication Date: 2026-04-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-01
Publication Date
2026-04-13

AI Technical Summary

Technical Problem

Users have psychological resistance to using their actual face photos as profile icons in online communication tools due to privacy and security concerns, leading to unset profile icons that impair communication smoothness.

Method used

A system that allows users to upload a facial photograph, preprocess it, input it into an image generation model to create a generated image, and use that image as a profile icon, preserving user features while reducing privacy and security concerns.

Benefits of technology

Enables users to use generated images as profile icons, reducing psychological resistance and facilitating smoother online communication by addressing privacy and security issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063865000001_ABST
    Figure 2026063865000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] The means by which users can upload facial photos, A method for pre-processing uploaded facial photos, A means for inputting a pre-processed facial photograph into an image generation model and generating a generated image, A means of providing the generated image to the user, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005] ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When using an online video conferencing tool, there is psychological resistance for users to use their actual face photos as profile icons. This may cause concerns about privacy and security. Also, many users leave their profile icons unset, which is an issue that impairs the smoothness of communication. To improve such a situation, it is required to provide means that allow users to use an image generated based on their own face photos as a profile icon without resistance.

Means for Solving the Problems

[0005] This invention provides a system that includes means for a user to upload a facial photograph, means for pre-processing the uploaded facial photograph, means for inputting the pre-processed facial photograph into an image generation model to generate a generated image, and means for providing the generated image to the user. This system allows users to use a generated image instead of a real photograph as a profile icon, thereby reducing their psychological resistance and concerns about privacy and security, and facilitating smooth online communication. The generated image is designed to preserve the user's features, enabling the user to appropriately represent themselves to other participants.

[0006] A "user" is the entity that uses the system to upload a facial photograph and receives the generated image.

[0007] A "face photo" is image data that includes the user's face and is uploaded to the system.

[0008] "Uploading" refers to the act of a user sending a photo of their face from their device to the system.

[0009] "Preprocessing" is the process of converting uploaded facial photos into an appropriate format, including resolution and formatting.

[0010] An "image generation model" refers to an artificial intelligence model that generates images from input facial photographs.

[0011] A "generated image" is an image that is not a real photograph but is generated by an image generation model, while retaining the user's characteristics.

[0012] "Providing" refers to the act of notifying users of the generated images and making them available for download and use.

[0013] A "profile icon" is an image that a user uses to represent themselves in online communication tools.

[0014] A "system" is a collection of technical devices and software that perform a series of processes in which a user uploads a facial photograph and receives the generated image. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the language used in the following description will be explained.

[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention relates to a system that enables users to use generated images instead of actual photographs as profile icons in online communication tools. This system is realized through a series of processes: "uploading a user's face photo," "pre-processing the face photo," "generating an icon using an image generation AI," and "providing the generated image to the user."

[0037] The server receives and stores facial photos uploaded by users. Users access the system from their devices, select facial photos, and upload them. Next, the server preprocesses the received facial photos. Preprocessing includes resizing and noise reduction, and converting the images to the appropriate format. This allows the image generation AI to process facial photos efficiently.

[0038] The pre-processed facial photographs are sent from the server to the image generation AI. The image generation AI creates generated images based on these pre-processed facial photographs. The image generation AI used here is designed to generate icon images that are not actual photographs of the person, while preserving the characteristics of the input image.

[0039] The generated image is returned to the server, which then provides it to the user. The user receives a notification, confirms the generated icon, and sets it as their profile icon in their online communication tool. In this process, the user can mitigate privacy and security concerns by using a generated image instead of a real photograph.

[0040] As a concrete example, consider a scenario where user A uploads their own photo to the system. User A accesses the system via a browser from their device, selects a photo of their face, and clicks the upload button. The server receives this operation and saves the photo. Next, the server performs pre-processing on the saved photo, such as resizing and noise reduction. After that, the pre-processed photo is input into an image generation AI to generate a generated image that retains user A's features. The generated image is returned to the server, which sends a notification to user A saying, "Your icon is ready." User A receives this notification, downloads the generated image, and sets it as their profile icon in applications such as Zoom or Google Meet (registered trademark). In this way, user A can participate in online meetings without worrying about privacy or security.

[0041] This system allows users to use generated images instead of real photos as their profile icons, making online communication smoother and reducing psychological resistance. As a result, it provides an environment where people can communicate online with greater peace of mind.

[0042] The following describes the processing flow.

[0043] Step 1:

[0044] The user accesses the system from their device, selects a photo of their face, and clicks the upload button.

[0045] Step 2:

[0046] The server receives the facial image sent by the user and saves it to a directory corresponding to the user ID.

[0047] Step 3:

[0048] The server reads the saved facial images and performs preprocessing. Specifically, it resizes and removes noise from the facial images.

[0049] Step 4:

[0050] The server sends the pre-processed facial photograph to an AI image generation model. This AI model creates generated images while preserving the user's features.

[0051] Step 5:

[0052] The image generation AI model receives a pre-processed facial photograph as input, generates a generated image, and returns it to the server.

[0053] Step 6:

[0054] The server receives the generated image and saves it to a directory corresponding to the user ID.

[0055] Step 7:

[0056] The server notifies the user when the generated image is ready. This notification can be sent via email or system notification.

[0057] Step 8:

[0058] The user receives a notification, accesses the system through their browser, and checks the generated icon image.

[0059] Step 9:

[0060] Users download the generated image and set it as their profile icon in online communication tools such as Zoom and Google® Meet.

[0061] In this way, the process proceeds, and the user can use the generated image as their profile icon without using a real photograph.

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] In online communication tools, allowing users to use their actual photos as profile icons can raise privacy and security concerns. Furthermore, there is a psychological resistance to publicly displaying one's face. To address these issues, it's necessary to use generated images that are based on the user's face but are not actual photographs; however, there is a lack of systems to smoothly implement this.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes means for a user to upload a facial photograph, means for pre-processing the uploaded facial photograph, means for inputting the pre-processed facial photograph into an artificial intelligence model to generate a generated image, means for providing the generated image to the user, means for the user to access the system, select and upload a facial photograph, and means for receiving a notification from the system and confirming the generated image. This allows the user to use the generated image as a profile icon without worrying about privacy or security.

[0067] A "user" is an individual who uses the system to upload a facial photograph and obtain a generated image.

[0068] A "face photo" is image data of a user's face.

[0069] "Uploading" refers to the process by which a user sends a photo of their face from their device to a server.

[0070] "Preprocessing" refers to processes such as resizing and noise reduction that the server performs on uploaded facial images.

[0071] An "artificial intelligence model" is a machine learning algorithm used to create generated images based on pre-processed facial photographs.

[0072] A "generated image" is a new icon image generated by an artificial intelligence model based on a pre-processed facial photograph.

[0073] "Provision" refers to the process of sending generated images to users, allowing them to review and download them.

[0074] "Notification" refers to a means of communication used by the server to inform the user that the generated image is ready.

[0075] A "profile icon" is an image used to identify a user in online communication tools.

[0076] An "online communication tool" is an application or service that allows users to communicate with others via the internet.

[0077] This invention relates to a system that enables users to use generated images instead of actual photographs as profile icons in online communication tools. This system is realized through a series of processes: "uploading a user's face photo," "pre-processing the face photo," "generating an icon using an image generation AI," and "providing the generated image to the user."

[0078] The server receives and stores the facial photos uploaded by the user. Next, the server preprocesses the received facial photos. Preprocessing involves resizing and noise reduction of the images, and converting them to an appropriate format. This allows the image generation AI to process the facial photos efficiently. Image processing libraries such as OpenCV are used for preprocessing.

[0079] The pre-processed facial photographs are sent from the server to the image generation AI. This AI uses generative AI models such as Stable Diffusion and DALL-E. Based on the pre-processed facial photographs, the AI ​​generates icon images that retain the individual's features but are not actual photographs. The following prompts are frequently used:

[0080] "This is my profile picture. Please generate an icon image that retains my features but is not an actual photograph."

[0081] The generated image is returned to the server, which then provides it to the user. The user receives a notification, can confirm the generated icon, and download it. This generated image can then be used as a profile icon for online communication tools such as Zoom and Google Meet.

[0082] As a concrete example, consider a scenario where user A uploads their own photo to the system. User A accesses the system via a browser from their device, selects a photo of their face, and clicks the upload button. The server receives this operation and saves the photo. Next, the server performs pre-processing on the saved photo, such as resizing and noise reduction. After that, the pre-processed photo is input into an image generation AI to generate a generated image that retains user A's features. The generated image is returned to the server, which sends a notification to user A saying, "Your icon is ready." User A receives this notification, downloads the generated image, and sets it as their profile icon. In this way, user A can participate in online meetings without worrying about privacy or security.

[0083] This system reduces privacy and security concerns, allowing users to use the generated image as their profile icon. This facilitates smoother online communication and reduces psychological resistance.

[0084] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0085] Step 1: Upload a user's profile picture.

[0086] The user opens a browser and accesses the system's facial photo upload page. The user clicks the "Select Facial Photo" button on the page and selects a facial photo file saved on their device. Then, they click the "Upload" button to send the facial photo to the system. The input here is the facial photo file selected by the user, and the output is the facial photo data sent to the server.

[0087] Step 2: Receive and save the facial photo

[0088] The server receives a POST request sent from the user's terminal and saves the facial image file to the specified directory. The server verifies that the file was saved correctly and returns a success message to the user. The input is the facial image file sent by the user, and the output is the saved image file and the success message.

[0089] Step 3: Pre-processing of facial photographs

[0090] The server reads the saved facial image files and performs preprocessing. The preprocessing includes the following items:

[0091] Resizing: Resizes the face image to the specified resolution. This allows for more efficient subsequent processing. The input is the saved face image file, and the output is the resized image data.

[0092] Noise Reduction: Removes noise from facial photographs to create clearer images. The input is a resized facial photograph, and the output is the image data after noise reduction.

[0093] Specifically, this preprocessing uses the OpenCV library's resize function and the cv2.fastNlMeansDenoisingColored function.

[0094] Step 4: Input into the Generative AI Model

[0095] The pre-processed facial photographs are sent from the server to the image generation AI. Generative AI models such as Stable Diffusion and DALL-E are used as the image generation AI. The server passes the pre-processed facial photographs along with prompt text to the image generation AI. The input is the pre-processed facial photographs and prompt text, and the output is an icon image generated by the generative AI model.

[0096] Step 5: Icon generation using AI

[0097] The generative AI model generates a new icon image based on a pre-processed facial image and a prompt message. The prompt message is as follows:

[0098] "This is my profile picture. Please generate an icon image that retains my features but is not an actual photograph."

[0099] The input is a pre-processed image and a prompt message, and the output is a generated icon image.

[0100] Step 6: Provide the generated image

[0101] The generated icon image is sent to the server, which then provides this image to the user. Specifically, the server sends a notification to the user stating "Your icon is ready" and provides a link to access the generated icon image. The input is the generated icon image, and the output is the notification message and access link to the user.

[0102] Step 7: Review and download the generated image.

[0103] The user receives a notification from the system, clicks the link, and views the generated icon image. They can then download this image if needed and use it as their profile icon in online communication tools (such as Zoom or Google Meet). The input is the notification message and access link, and the output is the download of the generated image and setting it as the profile icon.

[0104] In this way, the entire process is completed, and the user can use the generated image as a profile icon without worrying about privacy or security.

[0105] (Application Example 1)

[0106] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0107] Traditionally, using real photos as profile icons in online communication tools has raised privacy and security concerns. Furthermore, it has been difficult to express individuality through icons in a virtual environment, necessitating improvements to enhance the user experience. This invention aims to solve these problems and provide a system that allows users to easily and safely utilize unique icons.

[0108] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0109] In this invention, the server includes means for a user to upload a facial photograph, means for pre-processing the uploaded facial photograph, means for inputting the pre-processed facial photograph into an image generation model and generating a generated image, means for providing the generated image to the user, and means for using the generated image as an avatar icon in a virtual environment. This enables the user to use a unique and secure icon in the virtual environment.

[0110] A "user" is an individual or legal entity that uses the system.

[0111] A "face photo" is image data of a user's face.

[0112] "Uploading" refers to the act of a user sending data from their device to a server.

[0113] "Preprocessing" refers to the act of applying processes such as resizing, noise reduction, and format conversion to image data.

[0114] An "image generation model" is an algorithm or software that generates a new image based on an input image.

[0115] A "generated image" is a new image created by an image generation model that possesses the user's characteristics.

[0116] "Providing" refers to the act of notifying the user of the generated image or sending it to the user in a downloadable format.

[0117] A "virtual environment" is a virtual space or system built on a computer system.

[0118] An "avatar icon" is an image that a user uses to represent themselves within a virtual environment.

[0119] This invention provides a system that enables users to use generated images, rather than actual photographs, as avatar icons in a virtual environment. This system is realized through a series of processes including uploading a user's face photograph, pre-processing, icon generation by an image generation model, provision of the generated image, and use as an avatar icon.

[0120] The server receives and stores facial photos uploaded by users. Users access the system from their own devices, select facial photos, and upload them. Facial photos are often taken using smart glasses or computer terminals.

[0121] Next, the server preprocesses the received facial images. This preprocessing uses OpenCV and PIL to resize and denoise the images, converting them to an appropriate format. This allows the image generation model to process the facial images efficiently.

[0122] The pre-processed facial photographs are sent from the server to the image generation model. The image generation model uses machine learning libraries such as Keras and TENSORFLOW (registered trademark) to create generated images based on these pre-processed facial photographs. The image generation model is designed to generate icon images that are not actual photographs of the person, while retaining the user's characteristics.

[0123] The generated image is returned to the server, which then provides it to the user. The user receives a notification, can confirm the generated icon, and download it. This generated image is used as an avatar icon in the virtual environment.

[0124] As a concrete example, consider a scenario where a user visits a virtual shopping mall. The user wears smart glasses and uploads a photo of their face through its interface. The server receives this photo, preprocesses it, and then passes it to an image generation model. The generated icon is provided to the user via the server, and the user sets this icon as their avatar icon in the virtual mall. Through this series of operations, the user can experience the virtual environment without using a real-life photo of their face, thus eliminating privacy and security concerns.

[0125] Examples of prompts for a generative AI model include:

[0126] "Using a user's facial photo as input, generate a virtual icon image that retains the user's features but is not a real photograph."

[0127] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0128] Step 1:

[0129] The user accesses the system from their device, takes and selects a photo of their face, and uploads it. The user's device sends the photo to the server as input. The server receives and stores this data. The output is the saved photo.

[0130] Step 2:

[0131] The server preprocesses the received facial images. This includes resizing, denoising, and format conversion using OpenCV and PIL. The input is a saved facial image, and the output is a clear image after preprocessing. This allows the image generation model to process facial images efficiently.

[0132] Step 3:

[0133] The server inputs pre-processed facial images into an image generation model and generates generated images. The image generation model used here utilizes machine learning libraries such as Keras and TensorFlow. Pre-processed facial images are provided to the model as input, and generated images that retain the user's features are produced as output.

[0134] Step 4:

[0135] The generated image is returned to the server, which then provides this image to the user. Specifically, the user is notified of a link to the generated image and download options. The input is the generated image data, and the output is the notification provided to the user's device.

[0136] Step 5:

[0137] The user receives a notification through their device, confirms the generated icon, and downloads it. The input is the generated image data notified, and the output is the generated image downloaded to the user's device.

[0138] Step 6:

[0139] The user sets the generated icon as their avatar icon in the virtual environment. Specifically, they set their avatar in virtual shopping malls, online games, etc. The input is the downloaded generated image, and the output is the avatar icon set within the virtual environment.

[0140] This process allows users to use safe and unique icons in a virtual environment without using real-life photos of their faces.

[0141] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0142] This invention relates to a system that enables users to use generated images instead of actual photographs as profile icons in online communication tools. This system is realized through a series of processes including "uploading a user's face photo," "pre-processing the face photo," "generating an icon using an image generation AI," "providing the generated image to the user," and "recognizing the user's emotions."

[0143] First, the user accesses the system from their device, selects a photo of their face, and uploads it. The server receives the photo sent by the user and saves it to a directory corresponding to the user ID. The server then reads the saved photo and performs preprocessing. This preprocessing includes resizing and noise reduction of the photo. The preprocessed photo is then sent from the server to an image generation AI model. The image generation AI model creates a generated image based on this preprocessed photo. In this process, the image generation AI generates an icon image that is not a real photograph, while retaining the user's features.

[0144] Next, an emotion engine is integrated into the image generation process. The emotion engine analyzes the user's uploaded facial photos and recognizes the user's emotions. This recognition is reflected in the expression and atmosphere of the generated image. Specifically, if a user uploads a smiling photo, the generated icon image will also be adjusted to include elements of a smile. As a result, the generated image accurately reflects the user's emotions, resulting in a more natural and appealing profile icon.

[0145] The server receives the generated icon image and saves it to a directory corresponding to the user ID. The server notifies the user when the generated image is ready. This notification can be done via email or system notification. Upon receiving the notification, the user accesses the system through their browser to view the generated icon image. The user downloads this generated image and sets it as their profile icon in online communication tools such as Zoom or Google Meet.

[0146] As a concrete example, consider a scenario where user B uploads their own photo to the system. User B accesses the system via a browser from their device, selects a photo of their face, and clicks the upload button. The server receives this operation and saves the photo. The server then performs pre-processing on the saved photo, such as resizing and noise reduction. The pre-processed photo is then fed into an image generation AI model and simultaneously analyzed by an emotion engine to recognize the user's emotional state. The generated image is returned to the server, which sends a notification to user B stating, "Your icon is ready." User B receives this notification, downloads the generated image, and sets it as their profile icon in applications such as Zoom or Google Meet. In this way, user B can participate in online meetings without worrying about privacy or security.

[0147] This system allows users to use generated images as profile icons without using actual photos, facilitating smoother online communication and reducing psychological resistance. Furthermore, the generated images, which reflect the user's emotions, naturally express the user's own feelings and atmosphere, thus improving the quality of communication.

[0148] The following describes the processing flow.

[0149] Step 1:

[0150] The user accesses the system from their device, selects a photo of their face, and clicks the upload button.

[0151] Step 2:

[0152] The server receives the facial image submitted by the user and saves it to a directory corresponding to the user ID.

[0153] Step 3:

[0154] The server reads the saved facial photos and performs pre-processing such as resizing and noise reduction.

[0155] Step 4:

[0156] The server sends the pre-processed facial image to the emotion engine, which then analyzes the user's emotions.

[0157] Step 5:

[0158] The emotion engine analyzes the user's facial image and recognizes their emotions. This recognition result is stored as facial expression data.

[0159] Step 6:

[0160] The server passes the pre-processed facial images and the emotion engine's recognition results to the image generation AI model.

[0161] Step 7:

[0162] The image generation AI model creates generated images based on pre-processed facial photographs and the recognition results of the emotion engine. The generated images reflect the user's characteristics and emotions.

[0163] Step 8:

[0164] The server receives the generated image and saves it to the directory corresponding to the user ID.

[0165] Step 9:

[0166] The server notifies the user when the generated image is ready. This notification can be sent via email or system notification.

[0167] Step 10:

[0168] The user receives a notification, accesses the system via their browser, and checks the generated icon image.

[0169] Step 11:

[0170] Users download the generated image and set it as their profile icon in online communication tools such as Zoom or Google Meet.

[0171] In this way, the process proceeds, and the user can use a generated image instead of a real photo as their profile icon. Furthermore, the generated image reflects the user's emotions, resulting in a more natural and appealing profile icon.

[0172] (Example 2)

[0173] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0174] Traditional online communication tools often raise privacy and psychological concerns when users directly use their own photos as profile icons. Furthermore, generated images often fail to reflect the user's emotions, making it difficult to create natural and appealing profile pictures. Solving these problems is essential.

[0175] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0176] In this invention, the server includes means for a user to transmit a facial image, means for processing the transmitted facial image, means for inputting the processed facial image into an image generation module to create a generated image, means for presenting the generated image to the user, and means for analyzing the user's emotions and reflecting that emotional information in the image generation process. This makes it possible for users to create natural and attractive profile pictures that reflect their own characteristics and emotions while maintaining their privacy.

[0177] A "user" refers to an individual who uploads a facial image using this system and receives a generated profile picture.

[0178] "Face image" refers to image data, including the user's face, that is sent from the device to the system.

[0179] "Means" refers to a device or method for achieving a specific function.

[0180] "Processing" refers to performing pre-processing such as resizing and noise reduction on the transmitted facial image.

[0181] An "image generation module" refers to a part or all of software that receives a processed facial image as input and creates a new generated image.

[0182] "Generated image" refers to a new profile image created by the image generation module.

[0183] "To present" refers to displaying or providing the generated image on the system so that the user can review it.

[0184] "Emotional analysis" refers to the process of identifying emotions based on the user's facial image, including their expressions.

[0185] The "image generation process" refers to a series of procedures that create a generated image based on processed facial images and emotion analysis information.

[0186] A "profile picture" refers to the icon image that a user uses in online communication tools.

[0187] "Privacy" refers to a state in which a user's personal information is protected and not leaked to others.

[0188] This invention relates to a system that enables users to use generated images as profile icons in online communication tools, instead of directly using their own facial photographs. The system includes a server, a terminal, and an image generation module.

[0189] Users access the system via a browser from their own device (e.g., a PC or smartphone) and upload a facial image. The submitted facial image is sent to the server. After receiving the uploaded facial image, the server saves it to a directory corresponding to the user's identification number.

[0190] Next, the server preprocesses the saved face images. This preprocessing includes resizing the images (e.g., changing them to 256x256 pixels) and denoising. The preprocessed face images are then treated as temporary saved files.

[0191] The pre-processed facial images are sent by the server to an image generation module. This image generation module uses a common generative AI model (such as GAN or VQ-VAE-2) to create a new generated image based on the pre-processed facial images. In this process, the image generation module generates an icon image that is not an actual photograph of the user's face, while still retaining the user's features.

[0192] Furthermore, an emotion engine is used, and the server analyzes pre-processed facial images to recognize the user's emotions. The emotion engine analyzes the user's uploaded facial images and incorporates the results into the image generation process. For example, if a user uploads a smiling photo, the generated icon image will also reflect the smiling element.

[0193] The generated icon image is returned to the server, which then saves it again in the directory corresponding to the user's identification number. The server then notifies the user that the generated image is ready via email or system notification.

[0194] Upon receiving a notification, users access the system through their browser and download the generated icon image. The downloaded image can then be set by the user as their profile icon in online communication tools such as Zoom or Google Meet.

[0195] To give a concrete example, if a user uploads a "smiling selfie" to the system, the system preprocesses it and uses an image generation module and emotion engine to create a generated image that reflects the user's smile. This generated image is returned to the server and notified to the user. The user downloads the generated image and uses it as a profile icon in online communication tools. In this way, users can use a more natural and attractive profile picture while protecting their privacy.

[0196] Examples of prompt statements include the following:

[0197] "Please upload a photo of your face. We will generate a new icon image based on the uploaded photo. The generated image will reflect your emotions."

[0198] This allows users to use generated images as profile icons without using actual photos, which is expected to facilitate smoother online communication and reduce psychological resistance.

[0199] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0200] Step 1:

[0201] User sends facial image

[0202] Input: The user selects their own face image using their device and submits it via the file upload interface.

[0203] Description: The user accesses the system through their device's browser, clicks the upload button to select a facial image, and then presses the upload button again to send the image to the server.

[0204] Data processing: The selected face image file (e.g., "selfie.jpg") is sent from the user's device to the system.

[0205] Output: Face image data is sent to the server.

[0206] Step 2:

[0207] The server receives and stores the facial image.

[0208] Input: Face image data submitted in Step 1.

[0209] Description: The server saves the received facial image to a directory corresponding to the user ID (e.g., " / user_data / user_ID / ").

[0210] Data processing: File saving is performed (e.g., " / user_data / userA / selfie.jpg").

[0211] Output: Face image file stored on the server.

[0212] Step 3:

[0213] The server performs preprocessing on the facial images.

[0214] Input: Face image files stored on the server.

[0215] Description: The server reads the facial image and performs preprocessing such as resizing (e.g., 256x256 pixels) and noise reduction. The preprocessed image is saved as a temporary file.

[0216] Data processing: Images are resized, and pixel values ​​are adjusted to remove noise.

[0217] Output: Preprocessed image file (e.g., " / tmp / preprocessed_selfie.jpg").

[0218] Step 4:

[0219] The server sends the pre-processed facial image to the image generation module.

[0220] Input: Pre-processed facial image file.

[0221] Description: The server makes an API request to send the pre-processed facial image to the image generation module (e.g., using the " / generate_icon" endpoint in an HTTP POST request).

[0222] Data processing: An API request is made, the facial image data is encoded, and then sent.

[0223] Output: The image generation module starts processing.

[0224] Step 5:

[0225] The server uses an emotion engine to recognize the user's emotions.

[0226] Input: Pre-processed facial image file.

[0227] Description: The server sends a pre-processed facial image to the emotion engine and requests emotion analysis. The emotion engine analyzes the user's facial expression and returns the emotion information to the server.

[0228] Data processing: An emotion analysis algorithm extracts facial expression features and identifies the emotional state.

[0229] Output: Analyzed emotion information (e.g., "smile").

[0230] Step 6:

[0231] The image generation module creates the generated image.

[0232] Input: Pre-processed facial image files and emotion information.

[0233] Description: The image generation module generates a new profile icon image based on pre-processed facial images and emotion information. This generation process is adjusted to reflect the user's characteristics and emotions.

[0234] Data processing: The generative AI model receives facial images and emotion information as input, calculates the data, and generates a new image.

[0235] Output: The generated image file (e.g., "generated_icon.jpg").

[0236] Step 7:

[0237] The server presents the generated image to the user.

[0238] Input: The generated image file returned from the image generation module.

[0239] Description: The server saves the generated image files to a directory corresponding to the user ID and notifies the user when generation is complete. Notification methods include email and system notifications.

[0240] Data processing: Saving image files and sending notifications.

[0241] Output: A notification is sent to the user.

[0242] Step 8:

[0243] Users download and use the generated images.

[0244] Input: Notification from the server.

[0245] Description: The user receives a notification, accesses the system, and downloads the generated image. They then set the downloaded image as their profile icon in an online communication tool.

[0246] Data processing: Download the generated image.

[0247] Output: Generated image downloaded by the user.

[0248] (Application Example 2)

[0249] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0250] Traditional online communication tools commonly use users' actual photos, raising privacy concerns. Furthermore, using users' photos directly makes it difficult to accurately convey emotions and atmosphere, sometimes hindering smooth communication. Especially in physical stores, where staff's actual emotions and atmosphere influence customer service, there was a need for a visual way to represent staff emotions.

[0251] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to upload a facial photograph, means for pre-processing the uploaded facial photograph, means for inputting the pre-processed facial photograph into an image generation model to generate a generated image, means for recognizing the user's emotions from the facial photograph using an emotion engine and reflecting those emotions in the generated image, and means for providing the generated image to the user. This makes it possible to generate a profile icon that reflects an individual's characteristics and emotions without using the user's facial photograph, and to use it for things like staff name tags in physical stores.

[0252] "User" refers to an individual person who uses the system.

[0253] A "profile picture" is a still image that shows the user's face.

[0254] "Uploading" refers to the operation of a user sending data from their device to a server.

[0255] "Pre-processing" refers to the process of modifying uploaded facial photos, such as resizing and noise reduction.

[0256] An "image generation model" is a computational model that uses AI to generate new images based on input data.

[0257] A "generated image" refers to an image newly created by an image generation model.

[0258] An "emotion engine" is a program or device that analyzes a user's emotions from a facial photograph and retrieves the analysis results.

[0259] A "profile icon" is a small image used to visually represent a user in an online or digital environment.

[0260] A "physical store" refers to a commercial facility or service provider located in a physical place.

[0261] A "staff name tag" is a name tag worn by staff at a physical store, and its purpose is to display the staff member's name and position.

[0262] A "server" is a computer system that stores, processes, and provides data over a network.

[0263] Modes for carrying out the invention

[0264] This invention relates to a system that allows users to upload their own facial photographs and use generated images that reflect their emotions as profile icons or staff name tags. This system is expected to protect user privacy and enable emotional expression, and is particularly effective in facilitating smoother communication with customers in physical stores.

[0265] The system consists of the following main components:

[0266] Hardware and software usage

[0267] Server: Performs data storage, preprocessing, and AI model execution. Example: AWS (Amazon Web Services)

[0268] Smartphone: A user device on which an application runs.

[0269] Emotion engine: Software that recognizes emotions from facial images, e.g., Affectiva SDK

[0270] Image generation AI model: Creates generated images based on pre-processed facial photographs, e.g., GAN (Generative Adversarial Network).

[0271] Notification system: Notifies the user when the generated image is complete, e.g., Firebase Cloud Messaging

[0272] Smart glasses: Display on staff name tags

[0273] Detailed processing of the system

[0274] 1. User upload of a profile picture:

[0275] The user launches the "Smile Greeting" application from their smartphone, takes or selects a face photo and uploads it. The server receives this face photo and saves it in the directory corresponding to the user ID.

[0276] 2. Preprocessing of the face photo:

[0277] The server performs preprocessing such as resizing and noise removal on the received face photo. This preprocessing enables the image generation AI model to generate high-quality generated images.

[0278] 3. Analysis by the emotion engine:

[0279] The server inputs the preprocessed face photo into the emotion engine to analyze the user's emotion. The analysis results are obtained as emotion recognition results such as "smile", "surprise", "sadness", etc.

[0280] 4. Generation of the generated image:

[0281] The server inputs the preprocessed face photo and the emotion recognition results into the image generation AI model to create a generated image. This generated image is adjusted based on the emotion recognition results while maintaining the user's characteristics.

[0282] 5. Provision of the generated image:

[0283] To provide the generated image to the user, the server saves the image in the directory corresponding to the user ID and uses the notification system to send a notification to the user saying "The icon is ready".

[0284] 6. Use of the generated image:

[0285] The user receives the notification, checks and downloads the generated image through the smartphone app or browser. In the case of the staff of a physical store, this image is used as a smart glasses or other digital name tag and utilized for customer service.

[0286] Specific Example

[0287] For example, consider the scenario where user B uploads a face photo to the system. User B launches the "Smile Greeting" app, takes a face photo, and clicks the upload button. The server receives this operation and saves the face photo. Subsequently, the server performs preprocessing such as resizing and noise removal on the saved photo. The preprocessed photo is supplied to the emotion engine, and the emotional state is recognized as "smiling". The generated image returns to the server, and a notification saying "The icon is ready" is sent to user B. User B receives the notification, downloads the generated image, and sets it as a profile icon in apps like Zoom or Google Meet or as a staff name tag in a physical store. In this way, user B can utilize the generated image that reflects emotions without worrying about privacy and security.

[0288] Examples of Prompt Sentences

[0289] User ID: staffA

[0290] Face Photo: staffA_photo.jpg

[0291] Emotion: Smiling

[0292] The flow of the specific process in Application Example 2 will be described using FIG. 14.

[0293] Step 1:

[0294] The user uploads a face photo from a smartphone. The input can be a face photo taken with the smartphone's camera or a face photo selected from an existing image library. The output includes the face photo sent to the server. The user clicks the upload button in the application to send it.

[0295] Step 2:

[0296] The server stores the received facial photographs. The input is facial photograph data received from the user. The output is the facial photographs saved in a directory corresponding to the user ID. The server saves these facial photographs to its storage area and proceeds to the next preprocessing step.

[0297] Step 3:

[0298] The server performs preprocessing on facial images. The input is stored facial image data. The output is resized and denoised preprocessed facial image data. Specifically, the server resizes the facial image and removes noise using an image processing algorithm.

[0299] Step 4:

[0300] Pre-processed facial images are input into an emotion engine to recognize emotions. The input is pre-processed facial image data. The output is the user's emotion recognition result (e.g., smile, surprise, sadness). The server uses an emotion engine (e.g., Affectiva SDK) to perform emotion analysis.

[0301] Step 5:

[0302] The emotion recognition results are input into an image generation AI model to create a generated image. The input consists of pre-processed facial image data and the emotion recognition results. The output is a generated image that reflects the emotion. Specifically, a GAN (Generative Adversarial Network) model is used to generate images that reflect emotion while preserving the features.

[0303] Step 6:

[0304] The generated image is saved to a directory corresponding to the user ID. The input is the generated image data. The output is the generated image saved in a directory accessible to the user. The server saves the generated image and proceeds to the next notification step.

[0305] Step 7:

[0306] The server uses a notification system to notify the user of the completion of the generated image. As input, there is information regarding the saving of the generated image. As output, a notification is sent to the user's smartphone. The server uses Firebase Cloud Messaging to send the notification and delivers a message "The icon is ready" to the user.

[0307] Step 8:

[0308] The user checks the notification received on the smartphone and downloads the generated image. As input, there is a notification from the server and the URL of the generated image. As output, the generated image saved on the user's smartphone is obtained. The user clicks on the link and downloads the image through a browser or an app.

[0309] Step 9:

[0310] The staff of the physical store sets the generated image on smart glasses or digital name tags. As input, there is the generated image data. As output, the generated image displayed on the digital name tag is obtained. The staff uses a dedicated application to set the generated image on the name tag to facilitate communication with customers.

[0311] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires voice indicating user input with respect to the result of the specific processing. The control unit 46A transmits voice data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0312] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0313] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0314] [Second Embodiment]

[0315] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0316] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0317] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0318] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0319] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0320] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0321] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0322] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0323] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0324] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0325] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0326] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0327] This invention relates to a system that enables users to use generated images instead of actual photographs as profile icons in online communication tools. This system is realized through a series of processes: "uploading a user's face photo," "pre-processing the face photo," "generating an icon using an image generation AI," and "providing the generated image to the user."

[0328] The server receives and stores facial photos uploaded by users. Users access the system from their devices, select facial photos, and upload them. Next, the server preprocesses the received facial photos. Preprocessing includes resizing and noise reduction, and converting the images to the appropriate format. This allows the image generation AI to process facial photos efficiently.

[0329] The pre-processed facial photographs are sent from the server to the image generation AI. The image generation AI creates generated images based on these pre-processed facial photographs. The image generation AI used here is designed to generate icon images that are not actual photographs of the person, while preserving the characteristics of the input image.

[0330] The generated image is returned to the server, which then provides it to the user. The user receives a notification, confirms the generated icon, and sets it as their profile icon in their online communication tool. In this process, the user can mitigate privacy and security concerns by using a generated image instead of a real photograph.

[0331] As a concrete example, consider a scenario where user A uploads their own photo to the system. User A accesses the system via a browser from their device, selects a photo of their face, and clicks the upload button. The server receives this operation and saves the photo. Next, the server performs pre-processing on the saved photo, such as resizing and noise reduction. After that, the pre-processed photo is input into an image generation AI to generate a generated image that retains user A's features. The generated image is returned to the server, which sends a notification to user A saying, "Your icon is ready." User A receives this notification, downloads the generated image, and sets it as their profile icon in applications such as Zoom or Google Meet. In this way, user A can participate in online meetings without worrying about privacy or security.

[0332] This system allows users to use generated images instead of real photos as their profile icons, making online communication smoother and reducing psychological resistance. As a result, it provides an environment where people can communicate online with greater peace of mind.

[0333] The following describes the processing flow.

[0334] Step 1:

[0335] The user accesses the system from their device, selects a photo of their face, and clicks the upload button.

[0336] Step 2:

[0337] The server receives the facial image sent by the user and saves it to a directory corresponding to the user ID.

[0338] Step 3:

[0339] The server reads the saved facial images and performs preprocessing. Specifically, it resizes and removes noise from the facial images.

[0340] Step 4:

[0341] The server sends the pre-processed facial photograph to an AI image generation model. This AI model creates generated images while preserving the user's features.

[0342] Step 5:

[0343] The image generation AI model receives a pre-processed facial photograph as input, generates a generated image, and returns it to the server.

[0344] Step 6:

[0345] The server receives the generated image and saves it to a directory corresponding to the user ID.

[0346] Step 7:

[0347] The server notifies the user when the generated image is ready. This notification can be sent via email or system notification.

[0348] Step 8:

[0349] The user receives a notification, accesses the system through their browser, and checks the generated icon image.

[0350] Step 9:

[0351] Users download the generated image and set it as their profile icon in online communication tools such as Zoom or Google Meet.

[0352] In this way, the process proceeds, and the user can use the generated image as their profile icon without using a real photograph.

[0353] (Example 1)

[0354] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0355] In online communication tools, allowing users to use their actual photos as profile icons can raise privacy and security concerns. Furthermore, there is a psychological resistance to publicly displaying one's face. To address these issues, it's necessary to use generated images that are based on the user's face but are not actual photographs; however, there is a lack of systems to smoothly implement this.

[0356] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0357] In this invention, the server includes means for a user to upload a facial photograph, means for pre-processing the uploaded facial photograph, means for inputting the pre-processed facial photograph into an artificial intelligence model to generate a generated image, means for providing the generated image to the user, means for the user to access the system, select and upload a facial photograph, and means for receiving a notification from the system and confirming the generated image. This allows the user to use the generated image as a profile icon without worrying about privacy or security.

[0358] A "user" is an individual who uses the system to upload a facial photograph and obtain a generated image.

[0359] A "face photo" is image data of a user's face.

[0360] "Uploading" refers to the process by which a user sends a photo of their face from their device to a server.

[0361] "Preprocessing" refers to processes such as resizing and noise reduction that the server performs on uploaded facial images.

[0362] An "artificial intelligence model" is a machine learning algorithm used to create generated images based on pre-processed facial photographs.

[0363] A "generated image" is a new icon image generated by an artificial intelligence model based on a pre-processed facial photograph.

[0364] "Provision" refers to the process of sending generated images to users, allowing them to review and download them.

[0365] "Notification" refers to a means of communication used by the server to inform the user that the generated image is ready.

[0366] A "profile icon" is an image used to identify a user in online communication tools.

[0367] An "online communication tool" is an application or service that allows users to communicate with others via the internet.

[0368] This invention relates to a system that enables users to use generated images instead of actual photographs as profile icons in online communication tools. This system is realized through a series of processes: "uploading a user's face photo," "pre-processing the face photo," "generating an icon using an image generation AI," and "providing the generated image to the user."

[0369] The server receives and stores the facial photos uploaded by the user. Next, the server preprocesses the received facial photos. Preprocessing involves resizing and noise reduction of the images, and converting them to an appropriate format. This allows the image generation AI to process the facial photos efficiently. Image processing libraries such as OpenCV are used for preprocessing.

[0370] The pre-processed facial photographs are sent from the server to the image generation AI. This AI uses generative AI models such as Stable Diffusion and DALL-E. Based on the pre-processed facial photographs, the AI ​​generates icon images that retain the individual's features but are not actual photographs. The following prompts are frequently used:

[0371] "This is my profile picture. Please generate an icon image that retains my features but is not an actual photograph."

[0372] The generated image is returned to the server, which then provides it to the user. The user receives a notification, can confirm the generated icon, and download it. This generated image can then be used as a profile icon for online communication tools such as Zoom and Google Meet.

[0373] As a concrete example, consider a scenario where user A uploads their own photo to the system. User A accesses the system via a browser from their device, selects a photo of their face, and clicks the upload button. The server receives this operation and saves the photo. Next, the server performs pre-processing on the saved photo, such as resizing and noise reduction. After that, the pre-processed photo is input into an image generation AI to generate a generated image that retains user A's features. The generated image is returned to the server, which sends a notification to user A saying, "Your icon is ready." User A receives this notification, downloads the generated image, and sets it as their profile icon. In this way, user A can participate in online meetings without worrying about privacy or security.

[0374] This system reduces privacy and security concerns, allowing users to use the generated image as their profile icon. This facilitates smoother online communication and reduces psychological resistance.

[0375] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0376] Step 1: Upload a user's profile picture.

[0377] The user opens a browser and accesses the system's facial photo upload page. The user clicks the "Select Facial Photo" button on the page and selects a facial photo file saved on their device. Then, they click the "Upload" button to send the facial photo to the system. The input here is the facial photo file selected by the user, and the output is the facial photo data sent to the server.

[0378] Step 2: Receive and save the facial photo

[0379] The server receives a POST request sent from the user's terminal and saves the facial image file to the specified directory. The server verifies that the file was saved correctly and returns a success message to the user. The input is the facial image file sent by the user, and the output is the saved image file and the success message.

[0380] Step 3: Pre-processing of facial photographs

[0381] The server reads the saved facial image files and performs preprocessing. The preprocessing includes the following items:

[0382] Resizing: Resizes the face image to the specified resolution. This allows for more efficient subsequent processing. The input is the saved face image file, and the output is the resized image data.

[0383] Noise Reduction: Removes noise from facial photographs to create clearer images. The input is a resized facial photograph, and the output is the image data after noise reduction.

[0384] Specifically, this preprocessing uses the OpenCV library's resize function and the cv2.fastNlMeansDenoisingColored function.

[0385] Step 4: Input into the Generative AI Model

[0386] The pre-processed facial photographs are sent from the server to the image generation AI. Generative AI models such as Stable Diffusion and DALL-E are used as the image generation AI. The server passes the pre-processed facial photographs along with prompt text to the image generation AI. The input is the pre-processed facial photographs and prompt text, and the output is an icon image generated by the generative AI model.

[0387] Step 5: Icon generation using AI

[0388] The generative AI model generates a new icon image based on a pre-processed facial image and a prompt message. The prompt message is as follows:

[0389] "This is my profile picture. Please generate an icon image that retains my features but is not an actual photograph."

[0390] The input is a pre-processed image and a prompt message, and the output is a generated icon image.

[0391] Step 6: Provide the generated image

[0392] The generated icon image is sent to the server, which then provides this image to the user. Specifically, the server sends a notification to the user stating "Your icon is ready" and provides a link to access the generated icon image. The input is the generated icon image, and the output is the notification message and access link to the user.

[0393] Step 7: Review and download the generated image.

[0394] The user receives a notification from the system, clicks the link, and views the generated icon image. They can then download this image if needed and use it as their profile icon in online communication tools (such as Zoom or Google Meet). The input is the notification message and access link, and the output is the download of the generated image and setting it as the profile icon.

[0395] In this way, the entire process is completed, and the user can use the generated image as a profile icon without worrying about privacy or security.

[0396] (Application Example 1)

[0397] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0398] Traditionally, using real photos as profile icons in online communication tools has raised privacy and security concerns. Furthermore, it has been difficult to express individuality through icons in a virtual environment, necessitating improvements to enhance the user experience. This invention aims to solve these problems and provide a system that allows users to easily and safely utilize unique icons.

[0399] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0400] In this invention, the server includes means for a user to upload a facial photograph, means for pre-processing the uploaded facial photograph, means for inputting the pre-processed facial photograph into an image generation model and generating a generated image, means for providing the generated image to the user, and means for using the generated image as an avatar icon in a virtual environment. This enables the user to use a unique and secure icon in the virtual environment.

[0401] A "user" is an individual or legal entity that uses the system.

[0402] A "face photo" is image data of a user's face.

[0403] "Uploading" refers to the act of a user sending data from their device to a server.

[0404] "Preprocessing" refers to the act of applying processes such as resizing, noise reduction, and format conversion to image data.

[0405] An "image generation model" is an algorithm or software that generates a new image based on an input image.

[0406] A "generated image" is a new image created by an image generation model that possesses the user's characteristics.

[0407] "Providing" refers to the act of notifying the user of the generated image or sending it to the user in a downloadable format.

[0408] A "virtual environment" is a virtual space or system built on a computer system.

[0409] An "avatar icon" is an image that a user uses to represent themselves within a virtual environment.

[0410] This invention provides a system that enables users to use generated images, rather than actual photographs, as avatar icons in a virtual environment. This system is realized through a series of processes including uploading a user's face photograph, pre-processing, icon generation by an image generation model, provision of the generated image, and use as an avatar icon.

[0411] The server receives and stores facial photos uploaded by users. Users access the system from their own devices, select facial photos, and upload them. Facial photos are often taken using smart glasses or computer terminals.

[0412] Next, the server preprocesses the received facial images. This preprocessing uses OpenCV and PIL to resize and denoise the images, converting them to an appropriate format. This allows the image generation model to process the facial images efficiently.

[0413] The pre-processed facial images are sent from the server to the image generation model. The image generation model uses machine learning libraries such as Keras and TensorFlow to create generated images based on these pre-processed facial images. The image generation model is designed to generate icon images that are not actual photographs of the person, while still retaining the user's characteristics.

[0414] The generated image is returned to the server, which then provides it to the user. The user receives a notification, can confirm the generated icon, and download it. This generated image is used as an avatar icon in the virtual environment.

[0415] As a concrete example, consider a scenario where a user visits a virtual shopping mall. The user wears smart glasses and uploads a photo of their face through its interface. The server receives this photo, preprocesses it, and then passes it to an image generation model. The generated icon is provided to the user via the server, and the user sets this icon as their avatar icon in the virtual mall. Through this series of operations, the user can experience the virtual environment without using a real-life photo of their face, thus eliminating privacy and security concerns.

[0416] Examples of prompts for a generative AI model include:

[0417] "Using a user's facial photo as input, generate a virtual icon image that retains the user's features but is not a real photograph."

[0418] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0419] Step 1:

[0420] The user accesses the system from their device, takes and selects a photo of their face, and uploads it. The user's device sends the photo to the server as input. The server receives and stores this data. The output is the saved photo.

[0421] Step 2:

[0422] The server preprocesses the received facial images. This includes resizing, denoising, and format conversion using OpenCV and PIL. The input is a saved facial image, and the output is a clear image after preprocessing. This allows the image generation model to process facial images efficiently.

[0423] Step 3:

[0424] The server inputs pre-processed facial images into an image generation model and generates generated images. The image generation model used here utilizes machine learning libraries such as Keras and TensorFlow. Pre-processed facial images are provided to the model as input, and generated images that retain the user's features are produced as output.

[0425] Step 4:

[0426] The generated image is returned to the server, which then provides this image to the user. Specifically, the user is notified of a link to the generated image and download options. The input is the generated image data, and the output is the notification provided to the user's device.

[0427] Step 5:

[0428] The user receives a notification through their device, confirms the generated icon, and downloads it. The input is the generated image data notified, and the output is the generated image downloaded to the user's device.

[0429] Step 6:

[0430] The user sets the generated icon as their avatar icon in the virtual environment. Specifically, they set their avatar in virtual shopping malls, online games, etc. The input is the downloaded generated image, and the output is the avatar icon set within the virtual environment.

[0431] This process allows users to use safe and unique icons in a virtual environment without using real-life photos of their faces.

[0432] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0433] This invention relates to a system that enables users to use generated images instead of actual photographs as profile icons in online communication tools. This system is realized through a series of processes including "uploading a user's face photo," "pre-processing the face photo," "generating an icon using an image generation AI," "providing the generated image to the user," and "recognizing the user's emotions."

[0434] First, the user accesses the system from their device, selects a photo of their face, and uploads it. The server receives the photo sent by the user and saves it to a directory corresponding to the user ID. The server then reads the saved photo and performs preprocessing. This preprocessing includes resizing and noise reduction of the photo. The preprocessed photo is then sent from the server to an image generation AI model. The image generation AI model creates a generated image based on this preprocessed photo. In this process, the image generation AI generates an icon image that is not a real photograph, while retaining the user's features.

[0435] Next, an emotion engine is integrated into the image generation process. The emotion engine analyzes the user's uploaded facial photos and recognizes the user's emotions. This recognition is reflected in the expression and atmosphere of the generated image. Specifically, if a user uploads a smiling photo, the generated icon image will also be adjusted to include elements of a smile. As a result, the generated image accurately reflects the user's emotions, resulting in a more natural and appealing profile icon.

[0436] The server receives the generated icon image and saves it to a directory corresponding to the user ID. The server notifies the user when the generated image is ready. This notification can be done via email or system notification. Upon receiving the notification, the user accesses the system through their browser to view the generated icon image. The user downloads this generated image and sets it as their profile icon in online communication tools such as Zoom or Google Meet.

[0437] As a concrete example, consider a scenario where user B uploads their own photo to the system. User B accesses the system via a browser from their device, selects a photo of their face, and clicks the upload button. The server receives this operation and saves the photo. The server then performs pre-processing on the saved photo, such as resizing and noise reduction. The pre-processed photo is then fed into an image generation AI model and simultaneously analyzed by an emotion engine to recognize the user's emotional state. The generated image is returned to the server, which sends a notification to user B stating, "Your icon is ready." User B receives this notification, downloads the generated image, and sets it as their profile icon in applications such as Zoom or Google Meet. In this way, user B can participate in online meetings without worrying about privacy or security.

[0438] This system allows users to use generated images as profile icons without using actual photos, facilitating smoother online communication and reducing psychological resistance. Furthermore, the generated images, which reflect the user's emotions, naturally express the user's own feelings and atmosphere, thus improving the quality of communication.

[0439] The following describes the processing flow.

[0440] Step 1:

[0441] The user accesses the system from their device, selects a photo of their face, and clicks the upload button.

[0442] Step 2:

[0443] The server receives the facial image submitted by the user and saves it to a directory corresponding to the user ID.

[0444] Step 3:

[0445] The server reads the saved facial photos and performs pre-processing such as resizing and noise reduction.

[0446] Step 4:

[0447] The server sends the pre-processed facial image to the emotion engine, which then analyzes the user's emotions.

[0448] Step 5:

[0449] The emotion engine analyzes the user's facial image and recognizes their emotions. This recognition result is stored as facial expression data.

[0450] Step 6:

[0451] The server passes the pre-processed facial images and the emotion engine's recognition results to the image generation AI model.

[0452] Step 7:

[0453] The image generation AI model creates generated images based on pre-processed facial photographs and the recognition results of the emotion engine. The generated images reflect the user's characteristics and emotions.

[0454] Step 8:

[0455] The server receives the generated image and saves it to the directory corresponding to the user ID.

[0456] Step 9:

[0457] The server notifies the user when the generated image is ready. This notification can be sent via email or system notification.

[0458] Step 10:

[0459] The user receives a notification, accesses the system via their browser, and checks the generated icon image.

[0460] Step 11:

[0461] Users download the generated image and set it as their profile icon in online communication tools such as Zoom or Google Meet.

[0462] In this way, the process proceeds, and the user can use a generated image instead of a real photo as their profile icon. Furthermore, the generated image reflects the user's emotions, resulting in a more natural and appealing profile icon.

[0463] (Example 2)

[0464] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0465] Traditional online communication tools often raise privacy and psychological concerns when users directly use their own photos as profile icons. Furthermore, generated images often fail to reflect the user's emotions, making it difficult to create natural and appealing profile pictures. Solving these problems is essential.

[0466] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0467] In this invention, the server includes means for a user to transmit a facial image, means for processing the transmitted facial image, means for inputting the processed facial image into an image generation module to create a generated image, means for presenting the generated image to the user, and means for analyzing the user's emotions and reflecting that emotional information in the image generation process. This makes it possible for users to create natural and attractive profile pictures that reflect their own characteristics and emotions while maintaining their privacy.

[0468] A "user" refers to an individual who uploads a facial image using this system and receives a generated profile picture.

[0469] "Face image" refers to image data, including the user's face, that is sent from the device to the system.

[0470] "Means" refers to a device or method for achieving a specific function.

[0471] "Processing" refers to performing pre-processing such as resizing and noise reduction on the transmitted facial image.

[0472] An "image generation module" refers to a part or all of software that receives a processed facial image as input and creates a new generated image.

[0473] "Generated image" refers to a new profile image created by the image generation module.

[0474] "To present" refers to displaying or providing the generated image on the system so that the user can review it.

[0475] "Emotional analysis" refers to the process of identifying emotions based on the user's facial image, including their expressions.

[0476] The "image generation process" refers to a series of procedures that create a generated image based on processed facial images and emotion analysis information.

[0477] A "profile picture" refers to the icon image that a user uses in online communication tools.

[0478] "Privacy" refers to a state in which a user's personal information is protected and not leaked to others.

[0479] This invention relates to a system that enables users to use generated images as profile icons in online communication tools, instead of directly using their own facial photographs. The system includes a server, a terminal, and an image generation module.

[0480] Users access the system via a browser from their own device (e.g., a PC or smartphone) and upload a facial image. The submitted facial image is sent to the server. After receiving the uploaded facial image, the server saves it to a directory corresponding to the user's identification number.

[0481] Next, the server preprocesses the saved face images. This preprocessing includes resizing the images (e.g., changing them to 256x256 pixels) and denoising. The preprocessed face images are then treated as temporary saved files.

[0482] The pre-processed facial images are sent by the server to an image generation module. This image generation module uses a common generative AI model (such as GAN or VQ-VAE-2) to create a new generated image based on the pre-processed facial images. In this process, the image generation module generates an icon image that is not an actual photograph of the user's face, while still retaining the user's features.

[0483] Furthermore, an emotion engine is used, and the server analyzes pre-processed facial images to recognize the user's emotions. The emotion engine analyzes the user's uploaded facial images and incorporates the results into the image generation process. For example, if a user uploads a smiling photo, the generated icon image will also reflect the smiling element.

[0484] The generated icon image is returned to the server, which then saves it again in the directory corresponding to the user's identification number. The server then notifies the user that the generated image is ready via email or system notification.

[0485] Upon receiving a notification, users access the system through their browser and download the generated icon image. The downloaded image can then be set by the user as their profile icon in online communication tools such as Zoom or Google Meet.

[0486] To give a concrete example, if a user uploads a "smiling selfie" to the system, the system preprocesses it and uses an image generation module and emotion engine to create a generated image that reflects the user's smile. This generated image is returned to the server and notified to the user. The user downloads the generated image and uses it as a profile icon in online communication tools. In this way, users can use a more natural and attractive profile picture while protecting their privacy.

[0487] Examples of prompt statements include the following:

[0488] "Please upload a photo of your face. We will generate a new icon image based on the uploaded photo. The generated image will reflect your emotions."

[0489] This allows users to use generated images as profile icons without using actual photos, which is expected to facilitate smoother online communication and reduce psychological resistance.

[0490] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0491] Step 1:

[0492] User sends facial image

[0493] Input: The user selects their own face image using their device and submits it via the file upload interface.

[0494] Description: The user accesses the system through their device's browser, clicks the upload button to select a facial image, and then presses the upload button again to send the image to the server.

[0495] Data processing: The selected face image file (e.g., "selfie.jpg") is sent from the user's device to the system.

[0496] Output: Face image data is sent to the server.

[0497] Step 2:

[0498] The server receives and stores the facial image.

[0499] Input: Face image data submitted in Step 1.

[0500] Description: The server saves the received facial image to a directory corresponding to the user ID (e.g., " / user_data / user_ID / ").

[0501] Data processing: File saving is performed (e.g., " / user_data / userA / selfie.jpg").

[0502] Output: Face image file stored on the server.

[0503] Step 3:

[0504] The server performs preprocessing on the facial images.

[0505] Input: Face image files stored on the server.

[0506] Description: The server reads the facial image and performs preprocessing such as resizing (e.g., 256x256 pixels) and noise reduction. The preprocessed image is saved as a temporary file.

[0507] Data processing: Images are resized, and pixel values ​​are adjusted to remove noise.

[0508] Output: Preprocessed image file (e.g., " / tmp / preprocessed_selfie.jpg").

[0509] Step 4:

[0510] The server sends the pre-processed facial image to the image generation module.

[0511] Input: Pre-processed facial image file.

[0512] Description: The server makes an API request to send the pre-processed facial image to the image generation module (e.g., using the " / generate_icon" endpoint in an HTTP POST request).

[0513] Data processing: An API request is made, the facial image data is encoded, and then sent.

[0514] Output: The image generation module starts processing.

[0515] Step 5:

[0516] The server uses an emotion engine to recognize the user's emotions.

[0517] Input: Pre-processed facial image file.

[0518] Description: The server sends a pre-processed facial image to the emotion engine and requests emotion analysis. The emotion engine analyzes the user's facial expression and returns the emotion information to the server.

[0519] Data processing: An emotion analysis algorithm extracts facial expression features and identifies the emotional state.

[0520] Output: Analyzed emotion information (e.g., "smile").

[0521] Step 6:

[0522] The image generation module creates the generated image.

[0523] Input: Pre-processed facial image files and emotion information.

[0524] Description: The image generation module generates a new profile icon image based on pre-processed facial images and emotion information. This generation process is adjusted to reflect the user's characteristics and emotions.

[0525] Data processing: The generative AI model receives facial images and emotion information as input, calculates the data, and generates a new image.

[0526] Output: The generated image file (e.g., "generated_icon.jpg").

[0527] Step 7:

[0528] The server presents the generated image to the user.

[0529] Input: The generated image file returned from the image generation module.

[0530] Description: The server saves the generated image files to a directory corresponding to the user ID and notifies the user when generation is complete. Notification methods include email and system notifications.

[0531] Data processing: Saving image files and sending notifications.

[0532] Output: A notification is sent to the user.

[0533] Step 8:

[0534] Users download and use the generated images.

[0535] Input: Notification from the server.

[0536] Description: The user receives a notification, accesses the system, and downloads the generated image. They then set the downloaded image as their profile icon in an online communication tool.

[0537] Data processing: Download the generated image.

[0538] Output: Generated image downloaded by the user.

[0539] (Application Example 2)

[0540] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0541] Traditional online communication tools commonly use users' actual photos, raising privacy concerns. Furthermore, using users' photos directly makes it difficult to accurately convey emotions and atmosphere, sometimes hindering smooth communication. Especially in physical stores, where staff's actual emotions and atmosphere influence customer service, there was a need for a visual way to represent staff emotions.

[0542] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to upload a facial photograph, means for pre-processing the uploaded facial photograph, means for inputting the pre-processed facial photograph into an image generation model to generate a generated image, means for recognizing the user's emotions from the facial photograph using an emotion engine and reflecting those emotions in the generated image, and means for providing the generated image to the user. This makes it possible to generate a profile icon that reflects an individual's characteristics and emotions without using the user's facial photograph, and to use it for things like staff name tags in physical stores.

[0543] "User" refers to an individual person who uses the system.

[0544] A "profile picture" is a still image that shows the user's face.

[0545] "Uploading" refers to the operation of a user sending data from their device to a server.

[0546] "Pre-processing" refers to the process of modifying uploaded facial photos, such as resizing and noise reduction.

[0547] An "image generation model" is a computational model that uses AI to generate new images based on input data.

[0548] A "generated image" refers to an image newly created by an image generation model.

[0549] An "emotion engine" is a program or device that analyzes a user's emotions from a facial photograph and retrieves the analysis results.

[0550] A "profile icon" is a small image used to visually represent a user in an online or digital environment.

[0551] A "physical store" refers to a commercial facility or service provider located in a physical place.

[0552] A "staff name tag" is a name tag worn by staff at a physical store, and its purpose is to display the staff member's name and position.

[0553] A "server" is a computer system that stores, processes, and provides data over a network.

[0554] Modes for carrying out the invention

[0555] This invention relates to a system that allows users to upload their own facial photographs and use generated images that reflect their emotions as profile icons or staff name tags. This system is expected to protect user privacy and enable emotional expression, and is particularly effective in facilitating smoother communication with customers in physical stores.

[0556] The system consists of the following main components:

[0557] Hardware and software usage

[0558] Server: Performs data storage, preprocessing, and AI model execution. Example: AWS (Amazon Web Services)

[0559] Smartphone: A user device on which an application runs.

[0560] Emotion engine: Software that recognizes emotions from facial images, e.g., Affectiva SDK

[0561] Image generation AI model: Creates generated images based on pre-processed facial photographs, e.g., GAN (Generative Adversarial Network).

[0562] Notification system: Notifies the user when the generated image is complete, e.g., Firebase Cloud Messaging

[0563] Smart glasses: Display on staff name tags

[0564] Detailed processing of the system

[0565] 1. User upload of a profile picture:

[0566] The user launches the "Smile Greeting" application on their smartphone, takes or selects a photo of their face, and uploads it. The server receives this photo and saves it to a directory corresponding to the user ID.

[0567] 2. Pre-processing of facial photographs:

[0568] The server performs preprocessing on the received facial images, such as resizing and noise reduction. This preprocessing enables the image generation AI model to produce high-quality generated images.

[0569] 3. Analysis using an emotion engine:

[0570] The server inputs pre-processed facial images into an emotion engine to analyze the user's emotions. The analysis results are obtained as emotion recognition results such as "smile," "surprise," and "sadness."

[0571] 4. Generating the generated image:

[0572] The server inputs pre-processed facial images and emotion recognition results into an image generation AI model to create generated images. These generated images are adjusted based on the emotion recognition results while retaining the user's features.

[0573] 5. Provision of generated images:

[0574] To provide the generated image to the user, the server saves the image to a directory corresponding to the user ID and uses a notification system to send a notification to the user stating, "Your icon is ready."

[0575] 6. Use of generated images:

[0576] Users receive a notification and can view and download the generated image via a smartphone app or browser. For in-store staff, this image can be used as smart glasses or other digital name tags to enhance customer service.

[0577] Specific example

[0578] For example, consider a scenario where user B uploads a photo of their face to the system. User B launches the "Smile Greeting" app, takes a photo of their face, and clicks the upload button. The server receives this action and saves the photo. The server then performs pre-processing on the saved photo, such as resizing and noise reduction. The pre-processed photo is fed to the emotion engine, which recognizes the emotional state as "smiling." The generated image is returned to the server, and user B receives a notification that "your icon is ready." User B receives the notification, downloads the generated image, and sets it as a profile icon for Zoom or Google Meet, or as a staff name tag in a physical store. In this way, user B can use an emotionally-reflecting generated image without worrying about privacy or security.

[0579] Example of a prompt

[0580] User ID: staffA

[0581] Profile picture: staffA_photo.jpg

[0582] Emotion: Smile

[0583] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0584] Step 1:

[0585] The user uploads a photo of their face from their smartphone. The input can be a photo taken with the user's smartphone camera or a photo selected from an existing image library. The output includes the photo sent to the server. The user submits the photo by clicking the upload button in the application.

[0586] Step 2:

[0587] The server stores the received facial photographs. The input is facial photograph data received from the user. The output is the facial photographs saved in a directory corresponding to the user ID. The server saves these facial photographs to its storage area and proceeds to the next preprocessing step.

[0588] Step 3:

[0589] The server performs preprocessing on facial images. The input is stored facial image data. The output is resized and denoised preprocessed facial image data. Specifically, the server resizes the facial image and removes noise using an image processing algorithm.

[0590] Step 4:

[0591] Pre-processed facial images are input into an emotion engine to recognize emotions. The input is pre-processed facial image data. The output is the user's emotion recognition result (e.g., smile, surprise, sadness). The server performs emotion analysis using an emotion engine (e.g., Affectiva SDK).

[0592] Step 5:

[0593] The emotion recognition results are input into an image generation AI model to create a generated image. The input consists of pre-processed facial image data and the emotion recognition results. The output is a generated image that reflects the emotion. Specifically, a GAN (Generative Adversarial Network) model is used to generate images that reflect emotion while preserving the features.

[0594] Step 6:

[0595] The generated image is saved to a directory corresponding to the user ID. The input is the generated image data. The output is the generated image saved to a directory accessible to the user. The server saves the generated image and proceeds to the next notification step.

[0596] Step 7:

[0597] The server uses a notification system to inform the user that the generated image is complete. The input is information indicating that the generated image has been saved. The output is a notification sent to the user's smartphone. The server uses Firebase Cloud Messaging to send the notification, delivering the message "Your icon is ready."

[0598] Step 8:

[0599] The user checks the notification received on their smartphone and downloads the generated image. The input consists of the notification from the server and the URL of the generated image. The output is the generated image saved on the user's smartphone. The user clicks the link and downloads the image via a browser or app.

[0600] Step 9:

[0601] In-store staff set the generated images on smart glasses or digital name tags. The input is generated image data. The output is the generated image displayed on the digital name tag. Staff use a dedicated application to set the generated images on their name tags, facilitating smooth communication with customers.

[0602] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0603] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0604] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0605] [Third Embodiment]

[0606] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0607] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0608] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0609] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0610] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0611] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0612] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0613] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0614] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0615] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0616] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0617] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0618] This invention relates to a system that enables users to use generated images instead of actual photographs as profile icons in online communication tools. This system is realized through a series of processes: "uploading a user's face photo," "pre-processing the face photo," "generating an icon using an image generation AI," and "providing the generated image to the user."

[0619] The server receives and stores facial photos uploaded by users. Users access the system from their devices, select facial photos, and upload them. Next, the server preprocesses the received facial photos. Preprocessing includes resizing and noise reduction, and converting the images to the appropriate format. This allows the image generation AI to process facial photos efficiently.

[0620] The pre-processed facial photographs are sent from the server to the image generation AI. The image generation AI creates generated images based on these pre-processed facial photographs. The image generation AI used here is designed to generate icon images that are not actual photographs of the person, while preserving the characteristics of the input image.

[0621] The generated image is returned to the server, which then provides it to the user. The user receives a notification, confirms the generated icon, and sets it as their profile icon in their online communication tool. In this process, the user can mitigate privacy and security concerns by using a generated image instead of a real photograph.

[0622] As a concrete example, consider a scenario where user A uploads their own photo to the system. User A accesses the system via a browser from their device, selects a photo of their face, and clicks the upload button. The server receives this operation and saves the photo. Next, the server performs pre-processing on the saved photo, such as resizing and noise reduction. After that, the pre-processed photo is input into an image generation AI to generate a generated image that retains user A's features. The generated image is returned to the server, which sends a notification to user A saying, "Your icon is ready." User A receives this notification, downloads the generated image, and sets it as their profile icon in applications such as Zoom or Google Meet. In this way, user A can participate in online meetings without worrying about privacy or security.

[0623] This system allows users to use generated images instead of real photos as their profile icons, making online communication smoother and reducing psychological resistance. As a result, it provides an environment where people can communicate online with greater peace of mind.

[0624] The following describes the processing flow.

[0625] Step 1:

[0626] The user accesses the system from their device, selects a photo of their face, and clicks the upload button.

[0627] Step 2:

[0628] The server receives the facial image sent by the user and saves it to a directory corresponding to the user ID.

[0629] Step 3:

[0630] The server reads the saved facial images and performs preprocessing. Specifically, it resizes and removes noise from the facial images.

[0631] Step 4:

[0632] The server sends the pre-processed facial photograph to an AI image generation model. This AI model creates generated images while preserving the user's features.

[0633] Step 5:

[0634] The image generation AI model receives a pre-processed facial photograph as input, generates a generated image, and returns it to the server.

[0635] Step 6:

[0636] The server receives the generated image and saves it to a directory corresponding to the user ID.

[0637] Step 7:

[0638] The server notifies the user when the generated image is ready. This notification can be sent via email or system notification.

[0639] Step 8:

[0640] The user receives a notification, accesses the system through their browser, and checks the generated icon image.

[0641] Step 9:

[0642] Users download the generated image and set it as their profile icon in online communication tools such as Zoom or Google Meet.

[0643] In this way, the process proceeds, and the user can use the generated image as their profile icon without using a real photograph.

[0644] (Example 1)

[0645] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0646] In online communication tools, allowing users to use their actual photos as profile icons can raise privacy and security concerns. Furthermore, there is a psychological resistance to publicly displaying one's face. To address these issues, it's necessary to use generated images that are based on the user's face but are not actual photographs; however, there is a lack of systems to smoothly implement this.

[0647] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0648] In this invention, the server includes means for a user to upload a facial photograph, means for pre-processing the uploaded facial photograph, means for inputting the pre-processed facial photograph into an artificial intelligence model to generate a generated image, means for providing the generated image to the user, means for the user to access the system, select and upload a facial photograph, and means for receiving a notification from the system and confirming the generated image. This allows the user to use the generated image as a profile icon without worrying about privacy or security.

[0649] A "user" is an individual who uses the system to upload a facial photograph and obtain a generated image.

[0650] A "face photo" is image data of a user's face.

[0651] "Uploading" refers to the process by which a user sends a photo of their face from their device to a server.

[0652] "Preprocessing" refers to processes such as resizing and noise reduction that the server performs on uploaded facial images.

[0653] An "artificial intelligence model" is a machine learning algorithm used to create generated images based on pre-processed facial photographs.

[0654] A "generated image" is a new icon image generated by an artificial intelligence model based on a pre-processed facial photograph.

[0655] "Provision" refers to the process of sending generated images to users, allowing them to review and download them.

[0656] "Notification" refers to a means of communication used by the server to inform the user that the generated image is ready.

[0657] A "profile icon" is an image used to identify a user in online communication tools.

[0658] An "online communication tool" is an application or service that allows users to communicate with others via the internet.

[0659] This invention relates to a system that enables users to use generated images instead of actual photographs as profile icons in online communication tools. This system is realized through a series of processes: "uploading a user's face photo," "pre-processing the face photo," "generating an icon using an image generation AI," and "providing the generated image to the user."

[0660] The server receives and stores the facial photos uploaded by the user. Next, the server preprocesses the received facial photos. Preprocessing involves resizing and noise reduction of the images, and converting them to an appropriate format. This allows the image generation AI to process the facial photos efficiently. Image processing libraries such as OpenCV are used for preprocessing.

[0661] The pre-processed facial photographs are sent from the server to the image generation AI. This AI uses generative AI models such as Stable Diffusion and DALL-E. Based on the pre-processed facial photographs, the AI ​​generates icon images that retain the individual's features but are not actual photographs. The following prompts are frequently used:

[0662] "This is my profile picture. Please generate an icon image that retains my features but is not an actual photograph."

[0663] The generated image is returned to the server, which then provides it to the user. The user receives a notification, can confirm the generated icon, and download it. This generated image can then be used as a profile icon for online communication tools such as Zoom and Google Meet.

[0664] As a concrete example, consider a scenario where user A uploads their own photo to the system. User A accesses the system via a browser from their device, selects a photo of their face, and clicks the upload button. The server receives this operation and saves the photo. Next, the server performs pre-processing on the saved photo, such as resizing and noise reduction. After that, the pre-processed photo is input into an image generation AI to generate a generated image that retains user A's features. The generated image is returned to the server, which sends a notification to user A saying, "Your icon is ready." User A receives this notification, downloads the generated image, and sets it as their profile icon. In this way, user A can participate in online meetings without worrying about privacy or security.

[0665] This system reduces privacy and security concerns, allowing users to use the generated image as their profile icon. This facilitates smoother online communication and reduces psychological resistance.

[0666] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0667] Step 1: Upload a user's profile picture.

[0668] The user opens a browser and accesses the system's facial photo upload page. The user clicks the "Select Facial Photo" button on the page and selects a facial photo file saved on their device. Then, they click the "Upload" button to send the facial photo to the system. The input here is the facial photo file selected by the user, and the output is the facial photo data sent to the server.

[0669] Step 2: Receive and save the facial photo

[0670] The server receives a POST request sent from the user's terminal and saves the facial image file to the specified directory. The server verifies that the file was saved correctly and returns a success message to the user. The input is the facial image file sent by the user, and the output is the saved image file and the success message.

[0671] Step 3: Pre-processing of facial photographs

[0672] The server reads the saved facial image files and performs preprocessing. The preprocessing includes the following items:

[0673] Resizing: Resizes the face image to the specified resolution. This allows for more efficient subsequent processing. The input is the saved face image file, and the output is the resized image data.

[0674] Noise Reduction: Removes noise from facial photographs to create clearer images. The input is a resized facial photograph, and the output is the image data after noise reduction.

[0675] Specifically, this preprocessing uses the OpenCV library's resize function and the cv2.fastNlMeansDenoisingColored function.

[0676] Step 4: Input into the Generative AI Model

[0677] The pre-processed facial photographs are sent from the server to the image generation AI. Generative AI models such as Stable Diffusion and DALL-E are used as the image generation AI. The server passes the pre-processed facial photographs along with prompt text to the image generation AI. The input is the pre-processed facial photographs and prompt text, and the output is an icon image generated by the generative AI model.

[0678] Step 5: Icon generation using AI

[0679] The generative AI model generates a new icon image based on a pre-processed facial image and a prompt message. The prompt message is as follows:

[0680] "This is my profile picture. Please generate an icon image that retains my features but is not an actual photograph."

[0681] The input is a pre-processed image and a prompt message, and the output is a generated icon image.

[0682] Step 6: Provide the generated image

[0683] The generated icon image is sent to the server, which then provides this image to the user. Specifically, the server sends a notification to the user stating "Your icon is ready" and provides a link to access the generated icon image. The input is the generated icon image, and the output is the notification message and access link to the user.

[0684] Step 7: Review and download the generated image.

[0685] The user receives a notification from the system, clicks the link, and views the generated icon image. They can then download this image if needed and use it as their profile icon in online communication tools (such as Zoom or Google Meet). The input is the notification message and access link, and the output is the download of the generated image and setting it as the profile icon.

[0686] In this way, the entire process is completed, and the user can use the generated image as a profile icon without worrying about privacy or security.

[0687] (Application Example 1)

[0688] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0689] Traditionally, using real photos as profile icons in online communication tools has raised privacy and security concerns. Furthermore, it has been difficult to express individuality through icons in a virtual environment, necessitating improvements to enhance the user experience. This invention aims to solve these problems and provide a system that allows users to easily and safely utilize unique icons.

[0690] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0691] In this invention, the server includes means for a user to upload a facial photograph, means for pre-processing the uploaded facial photograph, means for inputting the pre-processed facial photograph into an image generation model and generating a generated image, means for providing the generated image to the user, and means for using the generated image as an avatar icon in a virtual environment. This enables the user to use a unique and secure icon in the virtual environment.

[0692] A "user" is an individual or legal entity that uses the system.

[0693] A "face photo" is image data of a user's face.

[0694] "Uploading" refers to the act of a user sending data from their device to a server.

[0695] "Preprocessing" refers to the act of applying processes such as resizing, noise reduction, and format conversion to image data.

[0696] An "image generation model" is an algorithm or software that generates a new image based on an input image.

[0697] A "generated image" is a new image created by an image generation model that possesses the user's characteristics.

[0698] "Providing" refers to the act of notifying the user of the generated image or sending it to the user in a downloadable format.

[0699] A "virtual environment" is a virtual space or system built on a computer system.

[0700] An "avatar icon" is an image that a user uses to represent themselves within a virtual environment.

[0701] This invention provides a system that enables users to use generated images, rather than actual photographs, as avatar icons in a virtual environment. This system is realized through a series of processes including uploading a user's face photograph, pre-processing, icon generation by an image generation model, provision of the generated image, and use as an avatar icon.

[0702] The server receives and stores facial photos uploaded by users. Users access the system from their own devices, select facial photos, and upload them. Facial photos are often taken using smart glasses or computer terminals.

[0703] Next, the server preprocesses the received facial images. This preprocessing uses OpenCV and PIL to resize and denoise the images, converting them to an appropriate format. This allows the image generation model to process the facial images efficiently.

[0704] The pre-processed facial images are sent from the server to the image generation model. The image generation model uses machine learning libraries such as Keras and TensorFlow to create generated images based on these pre-processed facial images. The image generation model is designed to generate icon images that are not actual photographs of the person, while still retaining the user's characteristics.

[0705] The generated image is returned to the server, which then provides it to the user. The user receives a notification, can confirm the generated icon, and download it. This generated image is used as an avatar icon in the virtual environment.

[0706] As a concrete example, consider a scenario where a user visits a virtual shopping mall. The user wears smart glasses and uploads a photo of their face through its interface. The server receives this photo, preprocesses it, and then passes it to an image generation model. The generated icon is provided to the user via the server, and the user sets this icon as their avatar icon in the virtual mall. Through this series of operations, the user can experience the virtual environment without using a real-life photo of their face, thus eliminating privacy and security concerns.

[0707] Examples of prompts for a generative AI model include:

[0708] "Using a user's facial photo as input, generate a virtual icon image that retains the user's features but is not a real photograph."

[0709] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0710] Step 1:

[0711] The user accesses the system from their device, takes and selects a photo of their face, and uploads it. The user's device sends the photo to the server as input. The server receives and stores this data. The output is the saved photo.

[0712] Step 2:

[0713] The server preprocesses the received facial images. This includes resizing, denoising, and format conversion using OpenCV and PIL. The input is a saved facial image, and the output is a clear image after preprocessing. This allows the image generation model to process facial images efficiently.

[0714] Step 3:

[0715] The server inputs pre-processed facial images into an image generation model and generates generated images. The image generation model used here utilizes machine learning libraries such as Keras and TensorFlow. Pre-processed facial images are provided to the model as input, and generated images that retain the user's features are produced as output.

[0716] Step 4:

[0717] The generated image is returned to the server, which then provides this image to the user. Specifically, the user is notified of a link to the generated image and download options. The input is the generated image data, and the output is the notification provided to the user's device.

[0718] Step 5:

[0719] The user receives a notification through their device, confirms the generated icon, and downloads it. The input is the generated image data notified, and the output is the generated image downloaded to the user's device.

[0720] Step 6:

[0721] The user sets the generated icon as their avatar icon in the virtual environment. Specifically, they set their avatar in virtual shopping malls, online games, etc. The input is the downloaded generated image, and the output is the avatar icon set within the virtual environment.

[0722] This process allows users to use safe and unique icons in a virtual environment without using real-life photos of their faces.

[0723] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0724] This invention relates to a system that enables users to use generated images instead of actual photographs as profile icons in online communication tools. This system is realized through a series of processes including "uploading a user's face photo," "pre-processing the face photo," "generating an icon using an image generation AI," "providing the generated image to the user," and "recognizing the user's emotions."

[0725] First, the user accesses the system from their device, selects a photo of their face, and uploads it. The server receives the photo sent by the user and saves it to a directory corresponding to the user ID. The server then reads the saved photo and performs preprocessing. This preprocessing includes resizing and noise reduction of the photo. The preprocessed photo is then sent from the server to an image generation AI model. The image generation AI model creates a generated image based on this preprocessed photo. In this process, the image generation AI generates an icon image that is not a real photograph, while retaining the user's features.

[0726] Next, an emotion engine is integrated into the image generation process. The emotion engine analyzes the user's uploaded facial photos and recognizes the user's emotions. This recognition is reflected in the expression and atmosphere of the generated image. Specifically, if a user uploads a smiling photo, the generated icon image will also be adjusted to include elements of a smile. As a result, the generated image accurately reflects the user's emotions, resulting in a more natural and appealing profile icon.

[0727] The server receives the generated icon image and saves it to a directory corresponding to the user ID. The server notifies the user when the generated image is ready. This notification can be done via email or system notification. Upon receiving the notification, the user accesses the system through their browser to view the generated icon image. The user downloads this generated image and sets it as their profile icon in online communication tools such as Zoom or Google Meet.

[0728] As a concrete example, consider a scenario where user B uploads their own photo to the system. User B accesses the system via a browser from their device, selects a photo of their face, and clicks the upload button. The server receives this operation and saves the photo. The server then performs pre-processing on the saved photo, such as resizing and noise reduction. The pre-processed photo is then fed into an image generation AI model and simultaneously analyzed by an emotion engine to recognize the user's emotional state. The generated image is returned to the server, which sends a notification to user B stating, "Your icon is ready." User B receives this notification, downloads the generated image, and sets it as their profile icon in applications such as Zoom or Google Meet. In this way, user B can participate in online meetings without worrying about privacy or security.

[0729] This system allows users to use generated images as profile icons without using actual photos, facilitating smoother online communication and reducing psychological resistance. Furthermore, the generated images, which reflect the user's emotions, naturally express the user's own feelings and atmosphere, thus improving the quality of communication.

[0730] The following describes the processing flow.

[0731] Step 1:

[0732] The user accesses the system from their device, selects a photo of their face, and clicks the upload button.

[0733] Step 2:

[0734] The server receives the facial image submitted by the user and saves it to a directory corresponding to the user ID.

[0735] Step 3:

[0736] The server reads the saved facial photos and performs pre-processing such as resizing and noise reduction.

[0737] Step 4:

[0738] The server sends the pre-processed facial image to the emotion engine, which then analyzes the user's emotions.

[0739] Step 5:

[0740] The emotion engine analyzes the user's facial image and recognizes their emotions. This recognition result is stored as facial expression data.

[0741] Step 6:

[0742] The server passes the pre-processed facial images and the emotion engine's recognition results to the image generation AI model.

[0743] Step 7:

[0744] The image generation AI model creates generated images based on pre-processed facial photographs and the recognition results of the emotion engine. The generated images reflect the user's characteristics and emotions.

[0745] Step 8:

[0746] The server receives the generated image and saves it to the directory corresponding to the user ID.

[0747] Step 9:

[0748] The server notifies the user when the generated image is ready. This notification can be sent via email or system notification.

[0749] Step 10:

[0750] The user receives a notification, accesses the system via their browser, and checks the generated icon image.

[0751] Step 11:

[0752] Users download the generated image and set it as their profile icon in online communication tools such as Zoom or Google Meet.

[0753] In this way, the process proceeds, and the user can use a generated image instead of a real photo as their profile icon. Furthermore, the generated image reflects the user's emotions, resulting in a more natural and appealing profile icon.

[0754] (Example 2)

[0755] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0756] Traditional online communication tools often raise privacy and psychological concerns when users directly use their own photos as profile icons. Furthermore, generated images often fail to reflect the user's emotions, making it difficult to create natural and appealing profile pictures. Solving these problems is essential.

[0757] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0758] In this invention, the server includes means for a user to transmit a facial image, means for processing the transmitted facial image, means for inputting the processed facial image into an image generation module to create a generated image, means for presenting the generated image to the user, and means for analyzing the user's emotions and reflecting that emotional information in the image generation process. This makes it possible for users to create natural and attractive profile pictures that reflect their own characteristics and emotions while maintaining their privacy.

[0759] A "user" refers to an individual who uploads a facial image using this system and receives a generated profile picture.

[0760] "Face image" refers to image data, including the user's face, that is sent from the device to the system.

[0761] "Means" refers to a device or method for achieving a specific function.

[0762] "Processing" refers to performing pre-processing such as resizing and noise reduction on the transmitted facial image.

[0763] An "image generation module" refers to a part or all of software that receives a processed facial image as input and creates a new generated image.

[0764] "Generated image" refers to a new profile image created by the image generation module.

[0765] "To present" refers to displaying or providing the generated image on the system so that the user can review it.

[0766] "Emotional analysis" refers to the process of identifying emotions based on the user's facial image, including their expressions.

[0767] The "image generation process" refers to a series of procedures that create a generated image based on processed facial images and emotion analysis information.

[0768] A "profile picture" refers to the icon image that a user uses in online communication tools.

[0769] "Privacy" refers to a state in which a user's personal information is protected and not leaked to others.

[0770] This invention relates to a system that enables users to use generated images as profile icons in online communication tools, instead of directly using their own facial photographs. The system includes a server, a terminal, and an image generation module.

[0771] Users access the system via a browser from their own device (e.g., a PC or smartphone) and upload a facial image. The submitted facial image is sent to the server. After receiving the uploaded facial image, the server saves it to a directory corresponding to the user's identification number.

[0772] Next, the server preprocesses the saved face images. This preprocessing includes resizing the images (e.g., changing them to 256x256 pixels) and denoising. The preprocessed face images are then treated as temporary saved files.

[0773] The pre-processed facial images are sent by the server to an image generation module. This image generation module uses a common generative AI model (such as GAN or VQ-VAE-2) to create a new generated image based on the pre-processed facial images. In this process, the image generation module generates an icon image that is not an actual photograph of the user's face, while still retaining the user's features.

[0774] Furthermore, an emotion engine is used, and the server analyzes pre-processed facial images to recognize the user's emotions. The emotion engine analyzes the user's uploaded facial images and incorporates the results into the image generation process. For example, if a user uploads a smiling photo, the generated icon image will also reflect the smiling element.

[0775] The generated icon image is returned to the server, which then saves it again in the directory corresponding to the user's identification number. The server then notifies the user that the generated image is ready via email or system notification.

[0776] Upon receiving a notification, users access the system through their browser and download the generated icon image. The downloaded image can then be set by the user as their profile icon in online communication tools such as Zoom or Google Meet.

[0777] To give a concrete example, if a user uploads a "smiling selfie" to the system, the system preprocesses it and uses an image generation module and emotion engine to create a generated image that reflects the user's smile. This generated image is returned to the server and notified to the user. The user downloads the generated image and uses it as a profile icon in online communication tools. In this way, users can use a more natural and attractive profile picture while protecting their privacy.

[0778] Examples of prompt statements include the following:

[0779] "Please upload a photo of your face. We will generate a new icon image based on the uploaded photo. The generated image will reflect your emotions."

[0780] This allows users to use generated images as profile icons without using actual photos, which is expected to facilitate smoother online communication and reduce psychological resistance.

[0781] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0782] Step 1:

[0783] User sends facial image

[0784] Input: The user selects their own face image using their device and submits it via the file upload interface.

[0785] Description: The user accesses the system through their device's browser, clicks the upload button to select a facial image, and then presses the upload button again to send the image to the server.

[0786] Data processing: The selected face image file (e.g., "selfie.jpg") is sent from the user's device to the system.

[0787] Output: Face image data is sent to the server.

[0788] Step 2:

[0789] The server receives and stores the facial image.

[0790] Input: Face image data submitted in Step 1.

[0791] Description: The server saves the received facial image to a directory corresponding to the user ID (e.g., " / user_data / user_ID / ").

[0792] Data processing: File saving is performed (e.g., " / user_data / userA / selfie.jpg").

[0793] Output: Face image file stored on the server.

[0794] Step 3:

[0795] The server performs preprocessing on the facial images.

[0796] Input: Face image files stored on the server.

[0797] Description: The server reads the facial image and performs preprocessing such as resizing (e.g., 256x256 pixels) and noise reduction. The preprocessed image is saved as a temporary file.

[0798] Data processing: Images are resized, and pixel values ​​are adjusted to remove noise.

[0799] Output: Preprocessed image file (e.g., " / tmp / preprocessed_selfie.jpg").

[0800] Step 4:

[0801] The server sends the pre-processed facial image to the image generation module.

[0802] Input: Pre-processed facial image file.

[0803] Description: The server makes an API request to send the pre-processed facial image to the image generation module (e.g., using the " / generate_icon" endpoint in an HTTP POST request).

[0804] Data processing: An API request is made, the facial image data is encoded, and then sent.

[0805] Output: The image generation module starts processing.

[0806] Step 5:

[0807] The server uses an emotion engine to recognize the user's emotions.

[0808] Input: Pre-processed facial image file.

[0809] Description: The server sends a pre-processed facial image to the emotion engine and requests emotion analysis. The emotion engine analyzes the user's facial expression and returns the emotion information to the server.

[0810] Data processing: An emotion analysis algorithm extracts facial expression features and identifies the emotional state.

[0811] Output: Analyzed emotion information (e.g., "smile").

[0812] Step 6:

[0813] The image generation module creates the generated image.

[0814] Input: Pre-processed facial image files and emotion information.

[0815] Description: The image generation module generates a new profile icon image based on pre-processed facial images and emotion information. This generation process is adjusted to reflect the user's characteristics and emotions.

[0816] Data processing: The generative AI model receives facial images and emotion information as input, calculates the data, and generates a new image.

[0817] Output: The generated image file (e.g., "generated_icon.jpg").

[0818] Step 7:

[0819] The server presents the generated image to the user.

[0820] Input: The generated image file returned from the image generation module.

[0821] Description: The server saves the generated image files to a directory corresponding to the user ID and notifies the user when generation is complete. Notification methods include email and system notifications.

[0822] Data processing: Saving image files and sending notifications.

[0823] Output: A notification is sent to the user.

[0824] Step 8:

[0825] Users download and use the generated images.

[0826] Input: Notification from the server.

[0827] Description: The user receives a notification, accesses the system, and downloads the generated image. They then set the downloaded image as their profile icon in an online communication tool.

[0828] Data processing: Download the generated image.

[0829] Output: Generated image downloaded by the user.

[0830] (Application Example 2)

[0831] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0832] Traditional online communication tools commonly use users' actual photos, raising privacy concerns. Furthermore, using users' photos directly makes it difficult to accurately convey emotions and atmosphere, sometimes hindering smooth communication. Especially in physical stores, where staff's actual emotions and atmosphere influence customer service, there was a need for a visual way to represent staff emotions.

[0833] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to upload a facial photograph, means for pre-processing the uploaded facial photograph, means for inputting the pre-processed facial photograph into an image generation model to generate a generated image, means for recognizing the user's emotions from the facial photograph using an emotion engine and reflecting those emotions in the generated image, and means for providing the generated image to the user. This makes it possible to generate a profile icon that reflects an individual's characteristics and emotions without using the user's facial photograph, and to use it for things like staff name tags in physical stores.

[0834] "User" refers to an individual person who uses the system.

[0835] A "profile picture" is a still image that shows the user's face.

[0836] "Uploading" refers to the operation of a user sending data from their device to a server.

[0837] "Pre-processing" refers to the process of modifying uploaded facial photos, such as resizing and noise reduction.

[0838] An "image generation model" is a computational model that uses AI to generate new images based on input data.

[0839] A "generated image" refers to an image newly created by an image generation model.

[0840] An "emotion engine" is a program or device that analyzes a user's emotions from a facial photograph and retrieves the analysis results.

[0841] A "profile icon" is a small image used to visually represent a user in an online or digital environment.

[0842] A "physical store" refers to a commercial facility or service provider located in a physical place.

[0843] A "staff name tag" is a name tag worn by staff at a physical store, and its purpose is to display the staff member's name and position.

[0844] A "server" is a computer system that stores, processes, and provides data over a network.

[0845] Modes for carrying out the invention

[0846] This invention relates to a system that allows users to upload their own facial photographs and use generated images that reflect their emotions as profile icons or staff name tags. This system is expected to protect user privacy and enable emotional expression, and is particularly effective in facilitating smoother communication with customers in physical stores.

[0847] The system consists of the following main components:

[0848] Hardware and software usage

[0849] Server: Performs data storage, preprocessing, and AI model execution. Example: AWS (Amazon Web Services)

[0850] Smartphone: A user device on which an application runs.

[0851] Emotion engine: Software that recognizes emotions from facial images, e.g., Affectiva SDK

[0852] Image generation AI model: Creates generated images based on pre-processed facial photographs, e.g., GAN (Generative Adversarial Network).

[0853] Notification system: Notifies the user when the generated image is complete, e.g., Firebase Cloud Messaging

[0854] Smart glasses: Display on staff name tags

[0855] Detailed processing of the system

[0856] 1. User upload of a profile picture:

[0857] The user launches the "Smile Greeting" application on their smartphone, takes or selects a photo of their face, and uploads it. The server receives this photo and saves it to a directory corresponding to the user ID.

[0858] 2. Pre-processing of facial photographs:

[0859] The server performs preprocessing on the received facial images, such as resizing and noise reduction. This preprocessing enables the image generation AI model to produce high-quality generated images.

[0860] 3. Analysis using an emotion engine:

[0861] The server inputs pre-processed facial images into an emotion engine to analyze the user's emotions. The analysis results are obtained as emotion recognition results such as "smile," "surprise," and "sadness."

[0862] 4. Generating the generated image:

[0863] The server inputs pre-processed facial images and emotion recognition results into an image generation AI model to create generated images. These generated images are adjusted based on the emotion recognition results while retaining the user's features.

[0864] 5. Provision of generated images:

[0865] To provide the generated image to the user, the server saves the image to a directory corresponding to the user ID and uses a notification system to send a notification to the user stating, "Your icon is ready."

[0866] 6. Use of generated images:

[0867] Users receive a notification and can view and download the generated image via a smartphone app or browser. For in-store staff, this image can be used as smart glasses or other digital name tags to enhance customer service.

[0868] Specific example

[0869] For example, consider a scenario where user B uploads a photo of their face to the system. User B launches the "Smile Greeting" app, takes a photo of their face, and clicks the upload button. The server receives this action and saves the photo. The server then performs pre-processing on the saved photo, such as resizing and noise reduction. The pre-processed photo is fed to the emotion engine, which recognizes the emotional state as "smiling." The generated image is returned to the server, and user B receives a notification that "your icon is ready." User B receives the notification, downloads the generated image, and sets it as a profile icon for Zoom or Google Meet, or as a staff name tag in a physical store. In this way, user B can use an emotionally-reflecting generated image without worrying about privacy or security.

[0870] Example of a prompt

[0871] User ID: staffA

[0872] Profile picture: staffA_photo.jpg

[0873] Emotion: Smile

[0874] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0875] Step 1:

[0876] The user uploads a photo of their face from their smartphone. The input can be a photo taken with the user's smartphone camera or a photo selected from an existing image library. The output includes the photo sent to the server. The user submits the photo by clicking the upload button in the application.

[0877] Step 2:

[0878] The server stores the received facial photographs. The input is facial photograph data received from the user. The output is the facial photographs saved in a directory corresponding to the user ID. The server saves these facial photographs to its storage area and proceeds to the next preprocessing step.

[0879] Step 3:

[0880] The server performs preprocessing on facial images. The input is stored facial image data. The output is resized and denoised preprocessed facial image data. Specifically, the server resizes the facial image and removes noise using an image processing algorithm.

[0881] Step 4:

[0882] Pre-processed facial images are input into an emotion engine to recognize emotions. The input is pre-processed facial image data. The output is the user's emotion recognition result (e.g., smile, surprise, sadness). The server performs emotion analysis using an emotion engine (e.g., Affectiva SDK).

[0883] Step 5:

[0884] The emotion recognition results are input into an image generation AI model to create a generated image. The input consists of pre-processed facial image data and the emotion recognition results. The output is a generated image that reflects the emotion. Specifically, a GAN (Generative Adversarial Network) model is used to generate images that reflect emotion while preserving the features.

[0885] Step 6:

[0886] The generated image is saved to a directory corresponding to the user ID. The input is the generated image data. The output is the generated image saved to a directory accessible to the user. The server saves the generated image and proceeds to the next notification step.

[0887] Step 7:

[0888] The server uses a notification system to inform the user that the generated image is complete. The input is information indicating that the generated image has been saved. The output is a notification sent to the user's smartphone. The server uses Firebase Cloud Messaging to send the notification, delivering the message "Your icon is ready."

[0889] Step 8:

[0890] The user checks the notification received on their smartphone and downloads the generated image. The input consists of the notification from the server and the URL of the generated image. The output is the generated image saved on the user's smartphone. The user clicks the link and downloads the image via a browser or app.

[0891] Step 9:

[0892] In-store staff set the generated images on smart glasses or digital name tags. The input is generated image data. The output is the generated image displayed on the digital name tag. Staff use a dedicated application to set the generated images on their name tags, facilitating smooth communication with customers.

[0893] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0894] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0895] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0896] [Fourth Embodiment]

[0897] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0898] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0899] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0900] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0901] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0902] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0903] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0904] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0905] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0906] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0907] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0908] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0909] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0910] This invention relates to a system that enables users to use generated images instead of actual photographs as profile icons in online communication tools. This system is realized through a series of processes: "uploading a user's face photo," "pre-processing the face photo," "generating an icon using an image generation AI," and "providing the generated image to the user."

[0911] The server receives and stores facial photos uploaded by users. Users access the system from their devices, select facial photos, and upload them. Next, the server preprocesses the received facial photos. Preprocessing includes resizing and noise reduction, and converting the images to the appropriate format. This allows the image generation AI to process facial photos efficiently.

[0912] The pre-processed facial photographs are sent from the server to the image generation AI. The image generation AI creates generated images based on these pre-processed facial photographs. The image generation AI used here is designed to generate icon images that are not actual photographs of the person, while preserving the characteristics of the input image.

[0913] The generated image is returned to the server, which then provides it to the user. The user receives a notification, confirms the generated icon, and sets it as their profile icon in their online communication tool. In this process, the user can mitigate privacy and security concerns by using a generated image instead of a real photograph.

[0914] As a concrete example, consider a scenario where user A uploads their own photo to the system. User A accesses the system via a browser from their device, selects a photo of their face, and clicks the upload button. The server receives this operation and saves the photo. Next, the server performs pre-processing on the saved photo, such as resizing and noise reduction. After that, the pre-processed photo is input into an image generation AI to generate a generated image that retains user A's features. The generated image is returned to the server, which sends a notification to user A saying, "Your icon is ready." User A receives this notification, downloads the generated image, and sets it as their profile icon in applications such as Zoom or Google Meet. In this way, user A can participate in online meetings without worrying about privacy or security.

[0915] This system allows users to use generated images instead of real photos as their profile icons, making online communication smoother and reducing psychological resistance. As a result, it provides an environment where people can communicate online with greater peace of mind.

[0916] The following describes the processing flow.

[0917] Step 1:

[0918] The user accesses the system from their device, selects a photo of their face, and clicks the upload button.

[0919] Step 2:

[0920] The server receives the facial image sent by the user and saves it to a directory corresponding to the user ID.

[0921] Step 3:

[0922] The server reads the saved facial images and performs preprocessing. Specifically, it resizes and removes noise from the facial images.

[0923] Step 4:

[0924] The server sends the pre-processed facial photograph to an AI image generation model. This AI model creates generated images while preserving the user's features.

[0925] Step 5:

[0926] The image generation AI model receives a pre-processed facial photograph as input, generates a generated image, and returns it to the server.

[0927] Step 6:

[0928] The server receives the generated image and saves it to a directory corresponding to the user ID.

[0929] Step 7:

[0930] The server notifies the user when the generated image is ready. This notification can be sent via email or system notification.

[0931] Step 8:

[0932] The user receives a notification, accesses the system through their browser, and checks the generated icon image.

[0933] Step 9:

[0934] Users download the generated image and set it as their profile icon in online communication tools such as Zoom or Google Meet.

[0935] In this way, the process proceeds, and the user can use the generated image as their profile icon without using a real photograph.

[0936] (Example 1)

[0937] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0938] In online communication tools, allowing users to use their actual photos as profile icons can raise privacy and security concerns. Furthermore, there is a psychological resistance to publicly displaying one's face. To address these issues, it's necessary to use generated images that are based on the user's face but are not actual photographs; however, there is a lack of systems to smoothly implement this.

[0939] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0940] In this invention, the server includes means for a user to upload a facial photograph, means for pre-processing the uploaded facial photograph, means for inputting the pre-processed facial photograph into an artificial intelligence model to generate a generated image, means for providing the generated image to the user, means for the user to access the system, select and upload a facial photograph, and means for receiving a notification from the system and confirming the generated image. This allows the user to use the generated image as a profile icon without worrying about privacy or security.

[0941] A "user" is an individual who uses the system to upload a facial photograph and obtain a generated image.

[0942] A "face photo" is image data of a user's face.

[0943] "Uploading" refers to the process by which a user sends a photo of their face from their device to a server.

[0944] "Preprocessing" refers to processes such as resizing and noise reduction that the server performs on uploaded facial images.

[0945] An "artificial intelligence model" is a machine learning algorithm used to create generated images based on pre-processed facial photographs.

[0946] A "generated image" is a new icon image generated by an artificial intelligence model based on a pre-processed facial photograph.

[0947] "Provision" refers to the process of sending generated images to users, allowing them to review and download them.

[0948] "Notification" refers to a means of communication used by the server to inform the user that the generated image is ready.

[0949] A "profile icon" is an image used to identify a user in online communication tools.

[0950] An "online communication tool" is an application or service that allows users to communicate with others via the internet.

[0951] This invention relates to a system that enables users to use generated images instead of actual photographs as profile icons in online communication tools. This system is realized through a series of processes: "uploading a user's face photo," "pre-processing the face photo," "generating an icon using an image generation AI," and "providing the generated image to the user."

[0952] The server receives and stores the facial photos uploaded by the user. Next, the server preprocesses the received facial photos. Preprocessing involves resizing and noise reduction of the images, and converting them to an appropriate format. This allows the image generation AI to process the facial photos efficiently. Image processing libraries such as OpenCV are used for preprocessing.

[0953] The pre-processed facial photographs are sent from the server to the image generation AI. This AI uses generative AI models such as Stable Diffusion and DALL-E. Based on the pre-processed facial photographs, the AI ​​generates icon images that retain the individual's features but are not actual photographs. The following prompts are frequently used:

[0954] "This is my profile picture. Please generate an icon image that retains my features but is not an actual photograph."

[0955] The generated image is returned to the server, which then provides it to the user. The user receives a notification, can confirm the generated icon, and download it. This generated image can then be used as a profile icon for online communication tools such as Zoom and Google Meet.

[0956] As a concrete example, consider a scenario where user A uploads their own photo to the system. User A accesses the system via a browser from their device, selects a photo of their face, and clicks the upload button. The server receives this operation and saves the photo. Next, the server performs pre-processing on the saved photo, such as resizing and noise reduction. After that, the pre-processed photo is input into an image generation AI to generate a generated image that retains user A's features. The generated image is returned to the server, which sends a notification to user A saying, "Your icon is ready." User A receives this notification, downloads the generated image, and sets it as their profile icon. In this way, user A can participate in online meetings without worrying about privacy or security.

[0957] This system reduces privacy and security concerns, allowing users to use the generated image as their profile icon. This facilitates smoother online communication and reduces psychological resistance.

[0958] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0959] Step 1: Upload a user's profile picture.

[0960] The user opens a browser and accesses the system's facial photo upload page. The user clicks the "Select Facial Photo" button on the page and selects a facial photo file saved on their device. Then, they click the "Upload" button to send the facial photo to the system. The input here is the facial photo file selected by the user, and the output is the facial photo data sent to the server.

[0961] Step 2: Receive and save the facial photo

[0962] The server receives a POST request sent from the user's terminal and saves the facial image file to the specified directory. The server verifies that the file was saved correctly and returns a success message to the user. The input is the facial image file sent by the user, and the output is the saved image file and the success message.

[0963] Step 3: Pre-processing of facial photographs

[0964] The server reads the saved facial image files and performs preprocessing. The preprocessing includes the following items:

[0965] Resizing: Resizes the face image to the specified resolution. This allows for more efficient subsequent processing. The input is the saved face image file, and the output is the resized image data.

[0966] Noise Reduction: Removes noise from facial photographs to create clearer images. The input is a resized facial photograph, and the output is the image data after noise reduction.

[0967] Specifically, this preprocessing uses the OpenCV library's resize function and the cv2.fastNlMeansDenoisingColored function.

[0968] Step 4: Input into the Generative AI Model

[0969] The pre-processed facial photographs are sent from the server to the image generation AI. Generative AI models such as Stable Diffusion and DALL-E are used as the image generation AI. The server passes the pre-processed facial photographs along with prompt text to the image generation AI. The input is the pre-processed facial photographs and prompt text, and the output is an icon image generated by the generative AI model.

[0970] Step 5: Icon generation using AI

[0971] The generative AI model generates a new icon image based on a pre-processed facial image and a prompt message. The prompt message is as follows:

[0972] "This is my profile picture. Please generate an icon image that retains my features but is not an actual photograph."

[0973] The input is a pre-processed image and a prompt message, and the output is a generated icon image.

[0974] Step 6: Provide the generated image

[0975] The generated icon image is sent to the server, which then provides this image to the user. Specifically, the server sends a notification to the user stating "Your icon is ready" and provides a link to access the generated icon image. The input is the generated icon image, and the output is the notification message and access link to the user.

[0976] Step 7: Review and download the generated image.

[0977] The user receives a notification from the system, clicks the link, and views the generated icon image. They can then download this image if needed and use it as their profile icon in online communication tools (such as Zoom or Google Meet). The input is the notification message and access link, and the output is the download of the generated image and setting it as the profile icon.

[0978] In this way, the entire process is completed, and the user can use the generated image as a profile icon without worrying about privacy or security.

[0979] (Application Example 1)

[0980] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0981] Traditionally, using real photos as profile icons in online communication tools has raised privacy and security concerns. Furthermore, it has been difficult to express individuality through icons in a virtual environment, necessitating improvements to enhance the user experience. This invention aims to solve these problems and provide a system that allows users to easily and safely utilize unique icons.

[0982] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0983] In this invention, the server includes means for a user to upload a facial photograph, means for pre-processing the uploaded facial photograph, means for inputting the pre-processed facial photograph into an image generation model and generating a generated image, means for providing the generated image to the user, and means for using the generated image as an avatar icon in a virtual environment. This enables the user to use a unique and secure icon in the virtual environment.

[0984] A "user" is an individual or legal entity that uses the system.

[0985] A "face photo" is image data of a user's face.

[0986] "Uploading" refers to the act of a user sending data from their device to a server.

[0987] "Preprocessing" refers to the act of applying processes such as resizing, noise reduction, and format conversion to image data.

[0988] An "image generation model" is an algorithm or software that generates a new image based on an input image.

[0989] A "generated image" is a new image created by an image generation model that possesses the user's characteristics.

[0990] "Providing" refers to the act of notifying the user of the generated image or sending it to the user in a downloadable format.

[0991] A "virtual environment" is a virtual space or system built on a computer system.

[0992] An "avatar icon" is an image that a user uses to represent themselves within a virtual environment.

[0993] This invention provides a system that enables users to use generated images, rather than actual photographs, as avatar icons in a virtual environment. This system is realized through a series of processes including uploading a user's face photograph, pre-processing, icon generation by an image generation model, provision of the generated image, and use as an avatar icon.

[0994] The server receives and stores facial photos uploaded by users. Users access the system from their own devices, select facial photos, and upload them. Facial photos are often taken using smart glasses or computer terminals.

[0995] Next, the server preprocesses the received facial images. This preprocessing uses OpenCV and PIL to resize and denoise the images, converting them to an appropriate format. This allows the image generation model to process the facial images efficiently.

[0996] The pre-processed facial images are sent from the server to the image generation model. The image generation model uses machine learning libraries such as Keras and TensorFlow to create generated images based on these pre-processed facial images. The image generation model is designed to generate icon images that are not actual photographs of the person, while still retaining the user's characteristics.

[0997] The generated image is returned to the server, which then provides it to the user. The user receives a notification, can confirm the generated icon, and download it. This generated image is used as an avatar icon in the virtual environment.

[0998] As a concrete example, consider a scenario where a user visits a virtual shopping mall. The user wears smart glasses and uploads a photo of their face through its interface. The server receives this photo, preprocesses it, and then passes it to an image generation model. The generated icon is provided to the user via the server, and the user sets this icon as their avatar icon in the virtual mall. Through this series of operations, the user can experience the virtual environment without using a real-life photo of their face, thus eliminating privacy and security concerns.

[0999] Examples of prompts for a generative AI model include:

[1000] "Using a user's facial photo as input, generate a virtual icon image that retains the user's features but is not a real photograph."

[1001] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1002] Step 1:

[1003] The user accesses the system from their device, takes and selects a photo of their face, and uploads it. The user's device sends the photo to the server as input. The server receives and stores this data. The output is the saved photo.

[1004] Step 2:

[1005] The server preprocesses the received facial images. This includes resizing, denoising, and format conversion using OpenCV and PIL. The input is a saved facial image, and the output is a clear image after preprocessing. This allows the image generation model to process facial images efficiently.

[1006] Step 3:

[1007] The server inputs pre-processed facial images into an image generation model and generates generated images. The image generation model used here utilizes machine learning libraries such as Keras and TensorFlow. Pre-processed facial images are provided to the model as input, and generated images that retain the user's features are produced as output.

[1008] Step 4:

[1009] The generated image is returned to the server, which then provides this image to the user. Specifically, the user is notified of a link to the generated image and download options. The input is the generated image data, and the output is the notification provided to the user's device.

[1010] Step 5:

[1011] The user receives a notification through their device, confirms the generated icon, and downloads it. The input is the generated image data notified, and the output is the generated image downloaded to the user's device.

[1012] Step 6:

[1013] The user sets the generated icon as their avatar icon in the virtual environment. Specifically, they set their avatar in virtual shopping malls, online games, etc. The input is the downloaded generated image, and the output is the avatar icon set within the virtual environment.

[1014] This process allows users to use safe and unique icons in a virtual environment without using real-life photos of their faces.

[1015] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1016] This invention relates to a system that enables users to use generated images instead of actual photographs as profile icons in online communication tools. This system is realized through a series of processes including "uploading a user's face photo," "pre-processing the face photo," "generating an icon using an image generation AI," "providing the generated image to the user," and "recognizing the user's emotions."

[1017] First, the user accesses the system from their device, selects a photo of their face, and uploads it. The server receives the photo sent by the user and saves it to a directory corresponding to the user ID. The server then reads the saved photo and performs preprocessing. This preprocessing includes resizing and noise reduction of the photo. The preprocessed photo is then sent from the server to an image generation AI model. The image generation AI model creates a generated image based on this preprocessed photo. In this process, the image generation AI generates an icon image that is not a real photograph, while retaining the user's features.

[1018] Next, an emotion engine is integrated into the image generation process. The emotion engine analyzes the user's uploaded facial photos and recognizes the user's emotions. This recognition is reflected in the expression and atmosphere of the generated image. Specifically, if a user uploads a smiling photo, the generated icon image will also be adjusted to include elements of a smile. As a result, the generated image accurately reflects the user's emotions, resulting in a more natural and appealing profile icon.

[1019] The server receives the generated icon image and saves it to a directory corresponding to the user ID. The server notifies the user when the generated image is ready. This notification can be done via email or system notification. Upon receiving the notification, the user accesses the system through their browser to view the generated icon image. The user downloads this generated image and sets it as their profile icon in online communication tools such as Zoom or Google Meet.

[1020] As a concrete example, consider a scenario where user B uploads their own photo to the system. User B accesses the system via a browser from their device, selects a photo of their face, and clicks the upload button. The server receives this operation and saves the photo. The server then performs pre-processing on the saved photo, such as resizing and noise reduction. The pre-processed photo is then fed into an image generation AI model and simultaneously analyzed by an emotion engine to recognize the user's emotional state. The generated image is returned to the server, which sends a notification to user B stating, "Your icon is ready." User B receives this notification, downloads the generated image, and sets it as their profile icon in applications such as Zoom or Google Meet. In this way, user B can participate in online meetings without worrying about privacy or security.

[1021] This system allows users to use generated images as profile icons without using actual photos, facilitating smoother online communication and reducing psychological resistance. Furthermore, the generated images, which reflect the user's emotions, naturally express the user's own feelings and atmosphere, thus improving the quality of communication.

[1022] The following describes the processing flow.

[1023] Step 1:

[1024] The user accesses the system from their device, selects a photo of their face, and clicks the upload button.

[1025] Step 2:

[1026] The server receives the facial image submitted by the user and saves it to a directory corresponding to the user ID.

[1027] Step 3:

[1028] The server reads the saved facial photos and performs pre-processing such as resizing and noise reduction.

[1029] Step 4:

[1030] The server sends the pre-processed facial image to the emotion engine, which then analyzes the user's emotions.

[1031] Step 5:

[1032] The emotion engine analyzes the user's facial image and recognizes their emotions. This recognition result is stored as facial expression data.

[1033] Step 6:

[1034] The server passes the pre-processed facial images and the emotion engine's recognition results to the image generation AI model.

[1035] Step 7:

[1036] The image generation AI model creates generated images based on pre-processed facial photographs and the recognition results of the emotion engine. The generated images reflect the user's characteristics and emotions.

[1037] Step 8:

[1038] The server receives the generated image and saves it to the directory corresponding to the user ID.

[1039] Step 9:

[1040] The server notifies the user when the generated image is ready. This notification can be sent via email or system notification.

[1041] Step 10:

[1042] The user receives a notification, accesses the system via their browser, and checks the generated icon image.

[1043] Step 11:

[1044] Users download the generated image and set it as their profile icon in online communication tools such as Zoom or Google Meet.

[1045] In this way, the process proceeds, and the user can use a generated image instead of a real photo as their profile icon. Furthermore, the generated image reflects the user's emotions, resulting in a more natural and appealing profile icon.

[1046] (Example 2)

[1047] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1048] Traditional online communication tools often raise privacy and psychological concerns when users directly use their own photos as profile icons. Furthermore, generated images often fail to reflect the user's emotions, making it difficult to create natural and appealing profile pictures. Solving these problems is essential.

[1049] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1050] In this invention, the server includes means for a user to transmit a facial image, means for processing the transmitted facial image, means for inputting the processed facial image into an image generation module to create a generated image, means for presenting the generated image to the user, and means for analyzing the user's emotions and reflecting that emotional information in the image generation process. This makes it possible for users to create natural and attractive profile pictures that reflect their own characteristics and emotions while maintaining their privacy.

[1051] A "user" refers to an individual who uploads a facial image using this system and receives a generated profile picture.

[1052] "Face image" refers to image data, including the user's face, that is sent from the device to the system.

[1053] "Means" refers to a device or method for achieving a specific function.

[1054] "Processing" refers to performing pre-processing such as resizing and noise reduction on the transmitted facial image.

[1055] An "image generation module" refers to a part or all of software that receives a processed facial image as input and creates a new generated image.

[1056] "Generated image" refers to a new profile image created by the image generation module.

[1057] "To present" refers to displaying or providing the generated image on the system so that the user can review it.

[1058] "Emotional analysis" refers to the process of identifying emotions based on the user's facial image, including their expressions.

[1059] The "image generation process" refers to a series of procedures that create a generated image based on processed facial images and emotion analysis information.

[1060] A "profile picture" refers to the icon image that a user uses in online communication tools.

[1061] "Privacy" refers to a state in which a user's personal information is protected and not leaked to others.

[1062] This invention relates to a system that enables users to use generated images as profile icons in online communication tools, instead of directly using their own facial photographs. The system includes a server, a terminal, and an image generation module.

[1063] Users access the system via a browser from their own device (e.g., a PC or smartphone) and upload a facial image. The submitted facial image is sent to the server. After receiving the uploaded facial image, the server saves it to a directory corresponding to the user's identification number.

[1064] Next, the server preprocesses the saved face images. This preprocessing includes resizing the images (e.g., changing them to 256x256 pixels) and denoising. The preprocessed face images are then treated as temporary saved files.

[1065] The pre-processed facial images are sent by the server to an image generation module. This image generation module uses a common generative AI model (such as GAN or VQ-VAE-2) to create a new generated image based on the pre-processed facial images. In this process, the image generation module generates an icon image that is not an actual photograph of the user's face, while still retaining the user's features.

[1066] Furthermore, an emotion engine is used, and the server analyzes pre-processed facial images to recognize the user's emotions. The emotion engine analyzes the user's uploaded facial images and incorporates the results into the image generation process. For example, if a user uploads a smiling photo, the generated icon image will also reflect the smiling element.

[1067] The generated icon image is returned to the server, which then saves it again in the directory corresponding to the user's identification number. The server then notifies the user that the generated image is ready via email or system notification.

[1068] Upon receiving a notification, users access the system through their browser and download the generated icon image. The downloaded image can then be set by the user as their profile icon in online communication tools such as Zoom or Google Meet.

[1069] To give a concrete example, if a user uploads a "smiling selfie" to the system, the system preprocesses it and uses an image generation module and emotion engine to create a generated image that reflects the user's smile. This generated image is returned to the server and notified to the user. The user downloads the generated image and uses it as a profile icon in online communication tools. In this way, users can use a more natural and attractive profile picture while protecting their privacy.

[1070] Examples of prompt statements include the following:

[1071] "Please upload a photo of your face. We will generate a new icon image based on the uploaded photo. The generated image will reflect your emotions."

[1072] This allows users to use generated images as profile icons without using actual photos, which is expected to facilitate smoother online communication and reduce psychological resistance.

[1073] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1074] Step 1:

[1075] User sends facial image

[1076] Input: The user selects their own face image using their device and submits it via the file upload interface.

[1077] Description: The user accesses the system through their device's browser, clicks the upload button to select a facial image, and then presses the upload button again to send the image to the server.

[1078] Data processing: The selected face image file (e.g., "selfie.jpg") is sent from the user's device to the system.

[1079] Output: Face image data is sent to the server.

[1080] Step 2:

[1081] The server receives and stores the facial image.

[1082] Input: Face image data submitted in Step 1.

[1083] Description: The server saves the received facial image to a directory corresponding to the user ID (e.g., " / user_data / user_ID / ").

[1084] Data processing: File saving is performed (e.g., " / user_data / userA / selfie.jpg").

[1085] Output: Face image file stored on the server.

[1086] Step 3:

[1087] The server performs preprocessing on the facial images.

[1088] Input: Face image files stored on the server.

[1089] Description: The server reads the facial image and performs preprocessing such as resizing (e.g., 256x256 pixels) and noise reduction. The preprocessed image is saved as a temporary file.

[1090] Data processing: Images are resized, and pixel values ​​are adjusted to remove noise.

[1091] Output: Preprocessed image file (e.g., " / tmp / preprocessed_selfie.jpg").

[1092] Step 4:

[1093] The server sends the pre-processed facial image to the image generation module.

[1094] Input: Pre-processed facial image file.

[1095] Description: The server makes an API request to send the pre-processed facial image to the image generation module (e.g., using the " / generate_icon" endpoint in an HTTP POST request).

[1096] Data processing: An API request is made, the facial image data is encoded, and then sent.

[1097] Output: The image generation module starts processing.

[1098] Step 5:

[1099] The server uses an emotion engine to recognize the user's emotions.

[1100] Input: Pre-processed facial image file.

[1101] Description: The server sends a pre-processed facial image to the emotion engine and requests emotion analysis. The emotion engine analyzes the user's facial expression and returns the emotion information to the server.

[1102] Data processing: An emotion analysis algorithm extracts facial expression features and identifies the emotional state.

[1103] Output: Analyzed emotion information (e.g., "smile").

[1104] Step 6:

[1105] The image generation module creates the generated image.

[1106] Input: Pre-processed facial image files and emotion information.

[1107] Description: The image generation module generates a new profile icon image based on pre-processed facial images and emotion information. This generation process is adjusted to reflect the user's characteristics and emotions.

[1108] Data processing: The generative AI model receives facial images and emotion information as input, calculates the data, and generates a new image.

[1109] Output: The generated image file (e.g., "generated_icon.jpg").

[1110] Step 7:

[1111] The server presents the generated image to the user.

[1112] Input: The generated image file returned from the image generation module.

[1113] Description: The server saves the generated image files to a directory corresponding to the user ID and notifies the user when generation is complete. Notification methods include email and system notifications.

[1114] Data processing: Saving image files and sending notifications.

[1115] Output: A notification is sent to the user.

[1116] Step 8:

[1117] Users download and use the generated images.

[1118] Input: Notification from the server.

[1119] Description: The user receives a notification, accesses the system, and downloads the generated image. They then set the downloaded image as their profile icon in an online communication tool.

[1120] Data processing: Download the generated image.

[1121] Output: Generated image downloaded by the user.

[1122] (Application Example 2)

[1123] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1124] Traditional online communication tools commonly use users' actual photos, raising privacy concerns. Furthermore, using users' photos directly makes it difficult to accurately convey emotions and atmosphere, sometimes hindering smooth communication. Especially in physical stores, where staff's actual emotions and atmosphere influence customer service, there was a need for a visual way to represent staff emotions.

[1125] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to upload a facial photograph, means for pre-processing the uploaded facial photograph, means for inputting the pre-processed facial photograph into an image generation model to generate a generated image, means for recognizing the user's emotions from the facial photograph using an emotion engine and reflecting those emotions in the generated image, and means for providing the generated image to the user. This makes it possible to generate a profile icon that reflects an individual's characteristics and emotions without using the user's facial photograph, and to use it for things like staff name tags in physical stores.

[1126] "User" refers to an individual person who uses the system.

[1127] A "profile picture" is a still image that shows the user's face.

[1128] "Uploading" refers to the operation of a user sending data from their device to a server.

[1129] "Pre-processing" refers to the process of modifying uploaded facial photos, such as resizing and noise reduction.

[1130] An "image generation model" is a computational model that uses AI to generate new images based on input data.

[1131] A "generated image" refers to an image newly created by an image generation model.

[1132] An "emotion engine" is a program or device that analyzes a user's emotions from a facial photograph and retrieves the analysis results.

[1133] A "profile icon" is a small image used to visually represent a user in an online or digital environment.

[1134] A "physical store" refers to a commercial facility or service provider located in a physical place.

[1135] A "staff name tag" is a name tag worn by staff at a physical store, and its purpose is to display the staff member's name and position.

[1136] A "server" is a computer system that stores, processes, and provides data over a network.

[1137] Modes for carrying out the invention

[1138] This invention relates to a system that allows users to upload their own facial photographs and use generated images that reflect their emotions as profile icons or staff name tags. This system is expected to protect user privacy and enable emotional expression, and is particularly effective in facilitating smoother communication with customers in physical stores.

[1139] The system consists of the following main components:

[1140] Hardware and software usage

[1141] Server: Performs data storage, preprocessing, and AI model execution. Example: AWS (Amazon Web Services)

[1142] Smartphone: A user device on which an application runs.

[1143] Emotion engine: Software that recognizes emotions from facial images, e.g., Affectiva SDK

[1144] Image generation AI model: Creates generated images based on pre-processed facial photographs, e.g., GAN (Generative Adversarial Network).

[1145] Notification system: Notifies the user when the generated image is complete, e.g., Firebase Cloud Messaging

[1146] Smart glasses: Display on staff name tags

[1147] Detailed processing of the system

[1148] 1. User upload of a profile picture:

[1149] The user launches the "Smile Greeting" application on their smartphone, takes or selects a photo of their face, and uploads it. The server receives this photo and saves it to a directory corresponding to the user ID.

[1150] 2. Pre-processing of facial photographs:

[1151] The server performs preprocessing on the received facial images, such as resizing and noise reduction. This preprocessing enables the image generation AI model to produce high-quality generated images.

[1152] 3. Analysis using an emotion engine:

[1153] The server inputs pre-processed facial images into an emotion engine to analyze the user's emotions. The analysis results are obtained as emotion recognition results such as "smile," "surprise," and "sadness."

[1154] 4. Generating the generated image:

[1155] The server inputs pre-processed facial images and emotion recognition results into an image generation AI model to create generated images. These generated images are adjusted based on the emotion recognition results while retaining the user's features.

[1156] 5. Provision of generated images:

[1157] To provide the generated image to the user, the server saves the image to a directory corresponding to the user ID and uses a notification system to send a notification to the user stating, "Your icon is ready."

[1158] 6. Use of generated images:

[1159] Users receive a notification and can view and download the generated image via a smartphone app or browser. For in-store staff, this image can be used as smart glasses or other digital name tags to enhance customer service.

[1160] Specific example

[1161] For example, consider a scenario where user B uploads a photo of their face to the system. User B launches the "Smile Greeting" app, takes a photo of their face, and clicks the upload button. The server receives this action and saves the photo. The server then performs pre-processing on the saved photo, such as resizing and noise reduction. The pre-processed photo is fed to the emotion engine, which recognizes the emotional state as "smiling." The generated image is returned to the server, and user B receives a notification that "your icon is ready." User B receives the notification, downloads the generated image, and sets it as a profile icon for Zoom or Google Meet, or as a staff name tag in a physical store. In this way, user B can use an emotionally-reflecting generated image without worrying about privacy or security.

[1162] Example of a prompt

[1163] User ID: staffA

[1164] Profile picture: staffA_photo.jpg

[1165] Emotion: Smile

[1166] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1167] Step 1:

[1168] The user uploads a photo of their face from their smartphone. The input can be a photo taken with the user's smartphone camera or a photo selected from an existing image library. The output includes the photo sent to the server. The user submits the photo by clicking the upload button in the application.

[1169] Step 2:

[1170] The server stores the received facial photographs. The input is facial photograph data received from the user. The output is the facial photographs saved in a directory corresponding to the user ID. The server saves these facial photographs to its storage area and proceeds to the next preprocessing step.

[1171] Step 3:

[1172] The server performs preprocessing on facial images. The input is stored facial image data. The output is resized and denoised preprocessed facial image data. Specifically, the server resizes the facial image and removes noise using an image processing algorithm.

[1173] Step 4:

[1174] Pre-processed facial images are input into an emotion engine to recognize emotions. The input is pre-processed facial image data. The output is the user's emotion recognition result (e.g., smile, surprise, sadness). The server performs emotion analysis using an emotion engine (e.g., Affectiva SDK).

[1175] Step 5:

[1176] The emotion recognition results are input into an image generation AI model to create a generated image. The input consists of pre-processed facial image data and the emotion recognition results. The output is a generated image that reflects the emotion. Specifically, a GAN (Generative Adversarial Network) model is used to generate images that reflect emotion while preserving the features.

[1177] Step 6:

[1178] The generated image is saved to a directory corresponding to the user ID. The input is the generated image data. The output is the generated image saved to a directory accessible to the user. The server saves the generated image and proceeds to the next notification step.

[1179] Step 7:

[1180] The server uses a notification system to inform the user that the generated image is complete. The input is information indicating that the generated image has been saved. The output is a notification sent to the user's smartphone. The server uses Firebase Cloud Messaging to send the notification, delivering the message "Your icon is ready."

[1181] Step 8:

[1182] The user checks the notification received on their smartphone and downloads the generated image. The input consists of the notification from the server and the URL of the generated image. The output is the generated image saved on the user's smartphone. The user clicks the link and downloads the image via a browser or app.

[1183] Step 9:

[1184] In-store staff set the generated images on smart glasses or digital name tags. The input is generated image data. The output is the generated image displayed on the digital name tag. Staff use a dedicated application to set the generated images on their name tags, facilitating smooth communication with customers.

[1185] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1186] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1187] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1188] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1189] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1190] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1191] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1192] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1193] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1194] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1195] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1196] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1197] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1198] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1199] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1200] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1201] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1202] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1203] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1204] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1205] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1206] The following is further disclosed regarding the embodiments described above.

[1207] (Claim 1)

[1208] The means by which users can upload facial photos,

[1209] A method for pre-processing uploaded facial photos,

[1210] A means for inputting a pre-processed facial photograph into an image generation model and generating a generated image,

[1211] A means of providing the generated image to the user,

[1212] A system that includes this.

[1213] (Claim 2)

[1214] The system according to claim 1, wherein the image generation model is designed to preserve user characteristics when generating generated images.

[1215] (Claim 3)

[1216] The system according to claim 1, wherein the generated image is intended to be used as a user's profile icon.

[1217] "Example 1"

[1218] (Claim 1)

[1219] The means by which users can upload facial photos,

[1220] A method for pre-processing uploaded facial photos,

[1221] A means for inputting a pre-processed facial photograph into an artificial intelligence model and generating a generated image,

[1222] A means of providing the generated image to the user,

[1223] A means for the user to access the system, select a facial photograph, and upload it,

[1224] A means of receiving a notification from the system and checking the generated image,

[1225] A system that includes this.

[1226] (Claim 2)

[1227] The system according to claim 1, wherein the artificial intelligence model is designed to preserve user features when generating generated images.

[1228] (Claim 3)

[1229] The system according to claim 1, which is intended for use as a profile icon for an online communication tool.

[1230] "Application Example 1"

[1231] (Claim 1)

[1232] The means by which users can upload facial photos,

[1233] A method for pre-processing uploaded facial photos,

[1234] A means for inputting a pre-processed facial photograph into an image generation model and generating a generated image,

[1235] A means of providing the generated image to the user,

[1236] Means for using the generated image as an avatar icon in a virtual environment,

[1237] A system that includes this.

[1238] (Claim 2)

[1239] The system according to claim 1, wherein the image generation model is designed to preserve user characteristics when generating generated images.

[1240] (Claim 3)

[1241] The system according to claim 1, wherein the generated image is intended to be used as a user avatar icon in a virtual environment.

[1242] "Example 2 of combining an emotion engine"

[1243] (Claim 1)

[1244] A means by which a user sends a facial image,

[1245] Means for processing the transmitted facial image,

[1246] A means for inputting a processed facial image into an image generation module and creating a generated image,

[1247] A means of presenting the generated image to the user,

[1248] A means of analyzing user emotions and reflecting that emotional information in the image generation process,

[1249] A system that includes this.

[1250] (Claim 2)

[1251] The system according to claim 1, wherein the image generation module is designed to preserve the user's characteristics and emotions when creating the generated image.

[1252] (Claim 3)

[1253] The system according to claim 1, wherein the generated image is intended to be used as a profile picture for the user's online communication tool.

[1254] "Application example 2 when combining with an emotional engine"

[1255] (Claim 1)

[1256] The means by which users can upload facial photos,

[1257] A method for pre-processing uploaded facial photos,

[1258] A means for inputting a pre-processed facial photograph into an image generation model and generating a generated image,

[1259] A means of providing the generated image to the user,

[1260] A means of recognizing a user's emotions from a facial photograph using an emotion engine and reflecting those emotions in a generated image,

[1261] A system that includes this.

[1262] (Claim 2)

[1263] The system according to claim 1, wherein the image generation model is designed to adjust the image based on the emotion recognition result while preserving the user's characteristics when generating the generated image.

[1264] (Claim 3)

[1265] The system according to claim 1, wherein the generated image is intended to be used as a user's profile icon, and further intended to be used as a name tag displayed on a physical device. [Explanation of symbols]

[1266] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. The means by which users can upload facial photos, A method for pre-processing uploaded facial photos, A means for inputting a pre-processed facial photograph into an image generation model and generating a generated image, A means of providing the generated image to the user, A system that includes this.

2. The system according to claim 1, wherein the image generation model is designed to preserve the user's characteristics when generating the generated image.

3. The system according to claim 1, wherein the generated image is intended to be used as a user's profile icon.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A