system

A system that analyzes user images and suggests clothing and posing on a server to generate modified profile pictures addresses the mismatch between online and offline appearances, allowing users to easily create attractive images that reflect their desired impression.

JP2026070227APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Conventional methods for creating profile photos, especially in dating apps and marriage consultation agencies, often result in a mismatch between the online image and the user's actual appearance, and professional photo studios are costly and time-consuming.

Method used

A system that uses user devices to capture images, performs image analysis on a server, and generates modified images suggesting optimal clothing and posing based on user preferences, allowing easy creation of profile pictures that align with the user's desired impression.

Benefits of technology

Enables users to create profile pictures that accurately reflect their desired impression without requiring specialized skills or equipment, bridging the gap between online and offline appearances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070227000001_ABST
    Figure 2026070227000001_ABST
Patent Text Reader

Abstract

Provide a system. , , 【Solution means】 Data acquisition means for a person image in a user terminal, Selection means for a desired impression for the person image, Communication means for transmitting the acquired person image to a server, In the server, analysis means for performing image analysis, Proposal means for proposing clothing and posing based on the analysis result, Generation means for generating a modified image generated based on the proposal, Means for transmitting the modified image to the user terminal, Storage means for storing the modified image selected by the user, A system including.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional method of taking profile photos, especially in matching apps and marriage consultation agencies, there is a problem that overly processed photos are different from the actual person's image, so the other party is often disappointed when meeting in person. Also, taking photos in a professional photo studio is costly and time-consuming, and it is difficult to easily obtain photos that maximize one's own charm.

Means for Solving the Problems

[0005] This invention provides a system that uses images of people taken by users with devices such as smartphones, performs image analysis on a server, and proposes optimal clothing and posing based on the analysis results. A generation means on the server generates a modified image based on the proposal and sends it to the user's terminal, allowing the user to easily create a profile picture that matches their desired impression. This makes it easy for users to express their charm to the fullest and reduces the gap between their online image and their actual appearance in person.

[0006] A "user terminal" is an electronic device that has the means to capture images of people and communicate, and includes smartphones and tablets.

[0007] "Personal images" refer to photographs and image data that include users, and are digital data that is subject to analysis and modification.

[0008] "Data acquisition means" refers to a function on the user's terminal for taking or acquiring images of people.

[0009] "Means of selecting desired impression" refers to interfaces and methods that allow users to select the impression or style they desire.

[0010] "Communication methods" refer to network technologies and protocols used to send data from a user terminal to a server.

[0011] A "server" is a remote computer system that receives data sent from a user terminal and performs analysis and generation processing.

[0012] "Image analysis" is the process of extracting features from received images of people using machine learning algorithms and image processing techniques.

[0013] "Analysis means" refers to programs or tools used to perform image analysis on a server.

[0014] The "proposal means" is a method or mechanism for proposing appropriate clothing and posing to the user based on the analysis results.

[0015] The "generation means" is an AI model or program for generating a modified image based on the proposal.

[0016] The "modified image" is a new profile image that reflects the user's desired impression and is created by the generation means.

[0017] The "storage means" is a storage function or process for holding the modified image selected within the user terminal.

Brief Description of Drawings

[0018] [Figure 1] It is a conceptual diagram showing an example of the configuration of the data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of the data processing device and the smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of the data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of the data processing device and the smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of the data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of the data processing device and the headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of the data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of the data processing device and the robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Mode for Carrying Out the Invention

[0019] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a tagged processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.

[0022] In the following embodiments, a tagged RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0023] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0024] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0026] [First Embodiment]

[0027] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0028] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0031] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0034] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0038] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0039] In an embodiment of this invention, the user initiates the process using a user terminal such as a smartphone or tablet. The user takes a picture of themselves, and the image data is stored within the application on the user terminal.

[0040] Users select their desired impression from a list displayed within the application. The options are divided into several categories, such as "Casual," "Formal," and "Natural," allowing users to choose the one that best matches their desired impression. They can also enter more detailed requests in a text box.

[0041] Next, the user terminal compresses the captured image of the person and sends it to the server. The server temporarily stores the received image data and starts analyzing it by applying a machine learning algorithm. Through this analysis, the server extracts facial features and key components from the image.

[0042] Based on this analysis and the user's desired impression, the server uses an AI model to suggest the most suitable clothing and posing. For example, if the user selects "casual and energetic impression," the AI ​​will generate suggestions that reflect this, such as casual clothing and cheerful poses.

[0043] The server then uses a generation mechanism to create a new modified image based on the proposed clothing and posing. This creates an image that matches the desired impression without compromising the user's original features. Multiple versions of this modified image are prepared and sent from the server to the user's terminal.

[0044] The user's device displays the received modified images to the user. The user can select their favorite image from the multiple images presented and save it to their device. The saved image is configured to be used directly as a profile picture on dating apps or marriage agencies.

[0045] In this embodiment of the present invention, users can easily create and utilize attractive profile pictures while maintaining their individuality, without requiring special photography skills or a professional studio.

[0046] The following describes the processing flow.

[0047] Step 1:

[0048] The user takes pictures of themselves from multiple angles using a user device such as a smartphone or tablet. The device temporarily stores the captured images of the person within the application.

[0049] Step 2:

[0050] Users select their desired impression from several impression categories provided within the application. These categories include "Casual," "Formal," and "Natural." Users can also enter detailed requests as text if needed.

[0051] Step 3:

[0052] The device compresses the captured image and sends it to the server. Optimization is performed during transmission to avoid compromising image quality. The server saves the received image data to its storage.

[0053] Step 4:

[0054] The server applies machine learning algorithms to the received images of people to perform image analysis. The analysis extracts facial feature points to identify the user's individuality and characteristics.

[0055] Step 5:

[0056] Based on the analysis results and the user's selected desired impression, the server uses an AI model to suggest appropriate clothing and posing. The AI ​​refers to a database based on impressions to identify the best suggestions.

[0057] Step 6:

[0058] The server uses a generation mechanism to generate a modified image that reflects the proposed clothing and posing. In doing so, it emphasizes the desired impression while preserving the user's characteristics.

[0059] Step 7:

[0060] The server sends multiple modified images to the user's terminal. The terminal presents these images to the user and prompts them to make a selection.

[0061] Step 8:

[0062] The user selects the most preferred modified image from the presented options. The device saves the selected modified image to local storage.

[0063] Step 9:

[0064] The device is configured to allow the selected modified image to be used as a profile picture on dating apps and matchmaking services. It also provides options to share the image on other platforms as needed.

[0065] (Example 1)

[0066] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0067] In conventional technology, creating attractive profile pictures required specialized skills and equipment. This presented a challenge: it was difficult for ordinary users to easily create impressive images that reflected their own personality.

[0068] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0069] In this invention, the server includes an information analysis means for performing image analysis, a suggestion means for proposing appearance and posture, and an image generation means for generating modified images. This makes it possible for users to easily generate and utilize images with diverse impressions that reflect their own characteristics.

[0070] A "user device" is an electronic device used by an individual, and is a device that processes text data and image data.

[0071] "Image data acquisition means" refers to means that have the function of allowing a user device to capture or receive an image and save it as digital data.

[0072] An "impression selection method" is a function that provides an interface for users to select a specific style or theme.

[0073] "Communication means" refers to the function for sending and receiving data between a user device and a computing device (server).

[0074] A "computing device" refers to a computing device used to process input data, and specifically to a server.

[0075] "Information analysis means" refers to means that have the function of analyzing received image data and extracting facial features and major body parts.

[0076] "Proposed means" refers to a function within the system that generates an appearance and posture according to the user's wishes based on the analysis results.

[0077] The "image generation means" is a function for creating modified image data based on the proposed content.

[0078] "Transmission means" refers to a function for transmitting the generated image data to the user's device.

[0079] "Memory function" refers to a function for saving the modified image selected by the user as digital data.

[0080] In an embodiment of this invention, the user first initiates the process using a device such as a smartphone or tablet. A dedicated application is installed on the device, and the user uses this application to take an image of themselves. This image data is stored within the application and functions as an image data acquisition means.

[0081] Next, the user selects their desired impression within the application. Impression categories include "Casual," "Formal," and "Natural," allowing the user to choose the one that best matches their desired look. This function is implemented as an impression selection method. Users can also communicate specific preferences by entering prompt messages. For example, they can enter a prompt message such as, "I want a natural look that emphasizes my smile."

[0082] The captured images and selected impression information are sent to the server using the terminal's communication method. The server is equipped with a machine learning algorithm built on Python and analyzes facial features based on the received image data. Libraries such as OpenCV and TENSORFLOW® are used for this analysis and function as information analysis tools.

[0083] After obtaining the analysis results, the server uses a generative AI model to suggest the optimal appearance and posture for the user. This model utilizes technologies such as GANs and acts as a suggestion tool. Based on the analysis results and the user's preferences, the server generates appropriate modified images. This image generation is designed to reflect the user's individuality while achieving the desired impression.

[0084] The generated modified image is returned to the user terminal via the server's transmission mechanism. The user terminal presents the user with multiple images, from which the user can select one. The selected image is saved using the user terminal's storage mechanism and can be used as a profile picture.

[0085] Through this mechanism, this invention enables users to easily create and utilize attractive profile pictures that highlight their own unique characteristics, without requiring specialized skills or special equipment.

[0086] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0087] Step 1:

[0088] The user uses a smartphone or other device to launch a dedicated application. Using the application, they take a picture of themselves and save this image data to internal storage. The input is the raw image data captured by the camera, and the output is the saved image data. This saved data will be used for processing in the next section.

[0089] Step 2:

[0090] On the application screen, the user selects their desired impression from a presented list of impressions. This selection serves as the impression selection method. Additionally, the user can enter prompt text as needed to specify the desired image characteristics. This results in the selected impression and prompt text as input, and the desired data is generated as output.

[0091] Step 3:

[0092] The terminal transmits the image data acquired in the previous step and the user's desired impression to the server via network communication. In this process, the stored image data and desired data obtained in the previous step are taken as input, and a dataset sent to the server is generated as output. After transmission, analysis can be performed on the server side.

[0093] Step 4:

[0094] The server stores the received image data in analysis storage. Then, using Python libraries such as OpenCV and TensorFlow, it analyzes and identifies facial feature points from the received images. The input is the image data sent to the server, and the output is facial feature data. This process prepares the data for the next proposed step.

[0095] Step 5:

[0096] The server uses a generative AI model to suggest appropriate appearances and postures based on acquired facial feature data and the user's desired impression. Generative models such as GANs are used in this process. The inputs are facial feature data and the user's desired impression, and the output is data related to the suggested clothing and posing.

[0097] Step 6:

[0098] The server generates a modified image based on the proposal. This image generation is performed using facial feature data and proposed data, reflecting the desired impression while preserving the original individuality. The inputs used are the proposed data and facial feature data, and the output is the generated modified image.

[0099] Step 7:

[0100] The server sends the generated modified images to the user's terminal. The user's terminal displays the received modified images in the application and prompts the user to make a selection. The input is the modified images sent from the server, and the output is the image selected by the user from among the displayed images.

[0101] Step 8:

[0102] The user selects their favorite image from several presented images and saves it to their device. At this stage, the user's selection is the input, and the image data saved on the device is the output. This saved image can be immediately used as a profile picture, for example.

[0103] (Application Example 1)

[0104] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0105] Online shopping presents a problem where consumers have difficulty visually assessing what suits them when choosing clothing. In particular, the inability to try on clothes in person often leads to disappointment after purchase, resulting in increased returns and exchanges, and ultimately lowering consumer satisfaction and sales efficiency.

[0106] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0107] In this invention, the server includes means for acquiring data of a person's image from a user terminal, means for suggesting clothing and posing based on the analysis results, and means for suggesting clothing and virtually trying it on based on the impression selected by the user. This allows the user to easily virtually try on clothing using their own image and visually confirm the results.

[0108] A "user terminal" refers to a communication device used individually by a user, such as a smartphone or tablet, that is equipped with image acquisition and communication functions.

[0109] "Data acquisition means" refers to a function that uses a user terminal to capture images of people and acquire that data.

[0110] "Selection methods" refer to interfaces and functions that allow users to choose impressions based on their own preferences and desires.

[0111] "Communication means" refers to the function for sending person image data from the user terminal to the server.

[0112] "Analysis means" refers to a function that processes human images received by the server and extracts feature quantities.

[0113] The "suggestion method" refers to a function that suggests appropriate clothing and posing to the user based on the analysis results.

[0114] The "generation means" refers to a function that, based on the proposal, modifies a user's image to generate a new image.

[0115] A "clothing try-on method" is a function that allows users to virtually try on suggested clothing.

[0116] "Storage method" refers to a function that allows users to save modified images they prefer on their device.

[0117] This invention begins with a user taking an image using a user terminal such as a smartphone or tablet. The user terminal is equipped with communication means to compress the acquired image of a person and send it to a server. The server uses analysis means to analyze the received image data and extract facial features from the image.

[0118] Next, the server uses the analysis results to suggest the most suitable clothing and posing to match the impression selected by the user. By using a suggestion method that utilizes a generation AI model, for example, if the user selects "casual and energetic impression," clothing and poses suitable for that impression will be suggested. Based on this suggestion, a new modified image is generated by the generation method and sent to the user's terminal.

[0119] The generated modified images are presented to the user. The user can select the most suitable image from multiple suggestions and save it to their device. The server also provides an environment where the user can virtually try on clothes using a clothing try-on system. This allows users to have an experience on e-commerce sites that is as if they were actually trying on clothes.

[0120] The hardware used includes smartphones and tablets as user terminals, and cloud computing environments as servers. The software includes image processing libraries (e.g., OpenCV) and AI model execution environments (e.g., TensorFlow).

[0121] For example, if a salaryman in his 40s requests a "formal and trustworthy impression," the AI ​​will generate a modified image of him in a suit. This image can be used as a profile picture in business settings.

[0122] An example of a prompt would be, "Suggest formal attire for a man in his 40s and synthesize it to create an impression of trustworthiness."

[0123] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0124] Step 1:

[0125] The user takes a picture of themselves using a user device such as a smartphone or tablet. The input is the desired impression selected by the user and the captured image of the person, and the output is that image data. This image data is saved on the smartphone in the optimal format.

[0126] Step 2:

[0127] The user terminal compresses the captured image of a person and sends it to the server. The input is the uncompressed image data, and the output is the compressed image data. The user terminal uploads this compressed image to the cloud server using a communication method.

[0128] Step 3:

[0129] The server analyzes the received image data. The input is compressed image data sent from the user's terminal, and the output is extracted facial feature data. The server uses an image analysis algorithm (e.g., OpenCV's face detection function) and temporarily stores the features in a database.

[0130] Step 4:

[0131] Based on the analysis results, the server uses an AI model to generate optimal clothing and posing, taking into account the user's selected desired impression. The input is facial feature data and the user's desired impression, and the output is proposed clothing and posing data. TensorFlow is used for the AI ​​model, and the generated data is proposed based on the prompts of the generating AI model.

[0132] Step 5:

[0133] The server uses a generation mechanism to generate modified images based on the proposed data. The input is the proposed clothing and posing data, and the output is the modified image data. Image processing is performed within the server to generate the modified image.

[0134] Step 6:

[0135] The server sends the generated modified image to the user's terminal. The input is the modified image data, and the output is the image data displayed on the user's terminal. This data is then transmitted to the user's smartphone using a communication method.

[0136] Step 7:

[0137] The user selects the most suitable image from several modified images and saves it on their device. The input consists of multiple modified image data sent from the server, and the output is the image data selected and saved by the user. This allows users to use images that convey their desired impression, such as profile pictures.

[0138] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0139] In an embodiment of this invention, the user uses a user terminal such as a smartphone or tablet to take a picture of themselves and then uses an emotion engine to analyze the user's emotions. During the shooting process, the terminal senses the user's facial expressions and voice, and uses this data to determine the user's emotions.

[0140] The device sends the captured image of the person, along with emotional data analyzed by the emotion engine, to the server. The server stores the received image data and emotional data in storage and performs a detailed analysis based on them. Using the analysis method, the server extracts facial feature points and, together with the emotional data, comprehensively evaluates the user's impression.

[0141] Next, the server uses a suggestion mechanism to propose the most suitable clothing and posing based on the user's desired impression and perceived emotions. For example, if the emotion engine recognizes the user's emotion as "joy," the server will suggest bright and cheerful clothing and posing.

[0142] The server uses a generation mechanism to create modified images incorporating the suggested clothing and posing. During this process, subtle adjustments to the color scheme and atmosphere are also made to reflect the emotions. The generated modified images are prepared as multiple variations and sent to the user's terminal.

[0143] The user's device displays the received modified images to the user. The user selects the image that best matches their intended impression or emotion and saves it on the device. The saved images can then be easily used as profile pictures on dating apps, marriage agencies, etc.

[0144] This embodiment allows users to not only select their desired impression but also generate a profile picture that reflects their current emotions, enabling a more natural and appealing form of self-expression.

[0145] The following describes the processing flow.

[0146] Step 1:

[0147] The user takes images of themselves from multiple angles using a user device such as a smartphone or tablet. The emotion engine is activated to detect the user's facial expressions and voice in real time and acquire emotion data.

[0148] Step 2:

[0149] The device saves the captured image of the person and emotion data to temporary memory. The user selects their desired impression from the impression categories provided within the application and enters detailed requests as needed.

[0150] Step 3:

[0151] The device compresses the saved person images and emotion data and sends them to the server. The image data compression is optimized to improve communication efficiency while maintaining image quality.

[0152] Step 4:

[0153] The server stores received person image data and emotion data in storage. Image analysis tools extract facial feature points from the received images and analyze the user's overall impression by combining them with emotion data.

[0154] Step 5:

[0155] Based on the analysis results, the user's selected desired impression, and the perceived emotions, the server uses an AI model to suggest the most suitable clothing and posing. For example, if the server analyzes the emotion as "joy," it will suggest brightly colored clothing and dynamic posing.

[0156] Step 6:

[0157] The server uses a generation mechanism to create modified images that reflect the proposed clothing and posing. Furthermore, it makes real-time adjustments to the color tone and expression according to the emotion. The generated modified images are prepared as multiple variations.

[0158] Step 7:

[0159] The server sends multiple modified images to the user's terminal. The terminal presents these images to the user and provides an interface for the user to make a selection.

[0160] Step 8:

[0161] The user selects the modified image that best reflects their intended impression or emotion from the presented images. The device then offers the option to save the selected modified image to local storage and make it immediately available as a profile picture.

[0162] (Example 2)

[0163] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0164] In modern society, it is difficult for individuals to easily create profile pictures that effectively reflect their real emotions and desired impression when expressing themselves on online platforms. As a result, they may not be able to express themselves more naturally, potentially leading to misunderstandings. Furthermore, traditional methods do not adequately provide emotion-based image adjustments or multiple suggestion options, making it difficult for users to pursue a more ideal self-expression.

[0165] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0166] In this invention, the server includes an analysis means for analyzing features from an image, an evaluation means for evaluating impressions based on the analysis results and emotion data, and a suggestion means for generating suggestions based on the evaluation. This enables users to generate more unique and attractive profile images that reflect their real emotions. Furthermore, by providing multiple modified images and making adjustments based on the desired impression, the system can meet the diverse needs of users for self-expression.

[0167] A "user device" is a digital device used by a user that enables image capture and data manipulation.

[0168] "Means of acquiring image data" refers to functions and technologies for capturing image data such as a user's face.

[0169] "Means for generating emotional data" refers to technologies that analyze captured image data to quantify or categorize the user's emotional state.

[0170] "Communication means" refers to the technologies and protocols used to send and receive data between user devices and processing devices.

[0171] A "processing device" is a computing device used to analyze and process received data, and primarily refers to a server.

[0172] "Analysis methods for analyzing features from images" refers to technologies that identify and extract facial features and shapes from image data.

[0173] "An evaluation method for assessing impressions" refers to a technology that comprehensively estimates a user's impression based on analyzed characteristics and emotional data.

[0174] "A suggestion generation method" refers to a technology that provides users with ideas for the most suitable clothing and posing based on evaluation results.

[0175] "Generation means" refers to the technology for modifying and generating images based on the proposed content.

[0176] "Transmission means" refers to the technology and protocols used to send the generated modified image to the user's device.

[0177] "Storage method" refers to the technology or method of storing images selected by the user in a data storage device.

[0178] To implement this invention, the user first takes an image of themselves using an electronic device such as a smartphone or tablet. At this time, the camera and microphone built into the device acquire the user's facial expressions and voice as data. This data is processed by emotion analysis software installed internally, and the user's emotions are quantified or categorized.

[0179] Next, the device sends the acquired image and emotion data to the server. HTTPS, a standard protocol over the internet, is used for communication. The server utilizes high-performance computing resources to analyze the image's feature points using image processing software. This is done, for example, by facial recognition algorithms. The emotion data is also used to comprehensively evaluate the user's impression.

[0180] Subsequently, the server generates suggestions based on the user's evaluation. These suggestions utilize a generative AI model and include clothing and posing that reflect the user's emotions and impressions. The generated suggestions are then used to create modified images, which are provided to the user as multiple image variations.

[0181] The generated images are sent to the device, allowing the user to select and save the one that best suits their desired impression and emotions. Saved images can also be used in dating apps and lifestyle facilities. This system allows users to easily create unique profile pictures that not only reflect their desired impression but also their current emotions.

[0182] As a concrete example, consider a scenario where a user wants to create a relaxed profile picture during a fun music festival. The user takes a photo with their device, and if the system determines the emotion is "joyful," the server suggests casual and cheerful clothing and posing. An example of a prompt used in this process might be, "Generate a profile picture with a fun atmosphere, casual clothing, and pose."

[0183] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0184] Step 1:

[0185] The user takes a picture of their face using a smartphone or tablet. At this time, the camera and microphone built into the device capture the user's facial expressions and voice, inputting this as image and audio data into the device. The device receives this data as initial input and prepares for emotion analysis.

[0186] Step 2:

[0187] The device processes the acquired image and audio data using internal emotion analysis software. This analysis extracts facial features from images and analyzes tone patterns from audio. As output of the analysis, emotion data is generated, classifying the user's emotions into categories such as "joy" and "sadness."

[0188] Step 3:

[0189] The terminal sends image data and emotion data obtained through analysis to the server. Secure communication via the internet is used for transmission, and the data is input to the server. The server prepares the received data for the next analysis.

[0190] Step 4:

[0191] The server uses image processing software to analyze facial feature points from image data. This analysis phase applies advanced facial recognition algorithms, quantifying image details and outputting them to an internal database. Based on these analysis results and sentiment data, the server evaluates the user's impression.

[0192] Step 5:

[0193] The server uses the results of analysis and evaluation to suggest appropriate clothing and posing for the user. The generative AI model then creates prompt statements based on this process, forming the core of the suggestions. For example, it might generate specific instructions such as "casual clothing and poses in a fun atmosphere." These prompt statements are then output and input into the generation mechanism.

[0194] Step 6:

[0195] The server receives prompt text into the generation mechanism and generates a modified image according to the suggested content. At this stage, multiple image variations with different color schemes and posing are generated and prepared within the server.

[0196] Step 7:

[0197] The server sends the generated modified image to the terminal. The terminal presents the received image to the user and displays multiple options on the screen, allowing the user to select the best image.

[0198] Step 8:

[0199] The user selects the most suitable image from the presented modified images and saves it to their device. This saved image can then be immediately used as a profile picture across various online services.

[0200] (Application Example 2)

[0201] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0202] In modern brick-and-mortar stores, it is difficult to efficiently suggest products based on the subjective feelings and impressions of individual customers, and there is a need for technology that enables product selection that reflects such feelings. Furthermore, there is a need for a way for customers to easily check styles optimized for their feelings without having to try products in-store.

[0203] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0204] In this invention, the server includes an analysis means for analyzing a person's image, a suggestion means for analyzing emotional information and proposing clothing and posing, and a generation means for providing the generated modified image as a virtual try-on. This enables personalized product suggestions that take into account the user's emotional state and a virtual try-on experience.

[0205] A "user terminal" is a portable electronic device used to capture and display images of people.

[0206] A "personal image" is a digital image data that captures the user's face or entire body.

[0207] An "information processing device" is a computing system that performs analysis and makes suggestions based on received image and emotion data.

[0208] "Communication means" refers to network communication functions for sending and receiving data between a user terminal and an information processing device.

[0209] The "analysis method" is a function that performs analysis based on input human images and emotional information to extract features.

[0210] The "suggestion method" refers to a function that selects and suggests the optimal clothing and posing based on the analysis results.

[0211] The "generation means" is a function that creates modified images based on the clothing and posing selected by the proposed means.

[0212] "Storage method" refers to a digital storage medium for long-term storage of modified images selected by the user.

[0213] "Emotion analysis means" refers to technology that detects emotions from the user's facial expressions and voice, and reflects the results in other proposed means.

[0214] "Virtual try-on" is a technology that allows users to visually try on clothing digitally without actually trying on the product.

[0215] This invention is a system that acquires images of a person using a user's terminal and provides optimal clothing and style suggestions through emotion analysis based on those images. Photographs taken by the user with a smartphone or other terminal are transmitted to an information processing device via communication means. The information processing device extracts facial features from the person's image and detects the user's emotions using emotion analysis means. This process utilizes image processing libraries such as OpenCV and Microsoft's Azure Emotion API.

[0216] The information processing device, based on the analysis results, uses a generative AI model to suggest clothing and styling that best suit the user's emotions and impressions. The suggestion method utilizes the generative AI model and generates a variety of fashion styles using prompt sentences. An example of such a prompt sentence is, "Generate styling suggestions for when the user is expressing cheerful emotions. Please provide ideas that emphasize bright colors and casual clothing."

[0217] The user's device receives the modified image again via communication and displays multiple styling options on the screen. The user can then select their preferred style and save the result on their device using the saving function. This feature allows for an instant virtual try-on experience while in a physical store, supporting purchase decisions. For example, when a user is enjoying a pleasant shopping experience in a store, they may be offered suggestions for bright, casual wear, naturally facilitating a purchase.

[0218] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0219] Step 1:

[0220] The user terminal captures an image of the user with its camera and acquires the image data. The input is the image data of the person captured by the camera, and the output is the generated image data. The terminal immediately prepares to transmit this image to the information processing device via a communication means.

[0221] Step 2:

[0222] The terminal transmits captured image data to the information processing device using a communication method. The input is the acquired image data, and the output is the image data transferred to the information processing device via the network.

[0223] Step 3:

[0224] The information processing device extracts feature points from the user's face using an analysis mechanism based on the received image data. In this process, the input is image data transmitted from the terminal, and the output is facial feature point data. The processing is performed using an image processing library such as OpenCV.

[0225] Step 4:

[0226] The information processing device detects a user's emotions from facial feature point data using emotion analysis means. The input here is facial feature point data, and the output is recognized emotion information. Emotion detection utilizes Microsoft's Azure Emotion API, among others.

[0227] Step 5:

[0228] Based on emotional information, the information processing device uses a proposed means and a generative AI model to consider the optimal clothing and posing. The input is the user's emotional information, and the output is the proposed clothing and posing data. The generative AI model is given instructions via prompts to generate the optimal style.

[0229] Step 6:

[0230] Using the information determined by the proposed method, the generation method of the information processing device generates a modified image. The input is clothing and posing data, as well as the original person image, and the output is a modified virtual try-on image. This is done using a generation AI model.

[0231] Step 7:

[0232] The information processing device transmits the generated modified image to the user terminal. The input is the generated modified image, and the output is the modified image transferred to the terminal.

[0233] Step 8:

[0234] The user terminal presents the received modified image to the user and displays multiple variations. The input is a pre-generated modified image, and the output is multiple virtual try-on images that the user can visually confirm. The image selected by the user is saved and stored via a storage means for future reference or purchase decisions.

[0235] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0236] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0237] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0238] [Second Embodiment]

[0239] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0240] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0241] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0242] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0243] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0244] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0245] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0246] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0247] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0248] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0249] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0250] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0251] In an embodiment of this invention, the user initiates the process using a user terminal such as a smartphone or tablet. The user takes a picture of themselves, and the image data is stored within the application on the user terminal.

[0252] Users select their desired impression from a list displayed within the application. The options are divided into several categories, such as "Casual," "Formal," and "Natural," allowing users to choose the one that best matches their desired impression. They can also enter more detailed requests in a text box.

[0253] Next, the user terminal compresses the captured image of the person and sends it to the server. The server temporarily stores the received image data and starts analyzing it by applying a machine learning algorithm. Through this analysis, the server extracts facial features and key components from the image.

[0254] Based on this analysis and the user's desired impression, the server uses an AI model to suggest the most suitable clothing and posing. For example, if the user selects "casual and energetic impression," the AI ​​will generate suggestions that reflect this, such as casual clothing and cheerful poses.

[0255] The server then uses a generation mechanism to create a new modified image based on the proposed clothing and posing. This creates an image that matches the desired impression without compromising the user's original features. Multiple versions of this modified image are prepared and sent from the server to the user's terminal.

[0256] The user's device displays the received modified images to the user. The user can select their favorite image from the multiple images presented and save it to their device. The saved image is configured to be used directly as a profile picture on dating apps or marriage agencies.

[0257] In this embodiment of the present invention, users can easily create and utilize attractive profile pictures while maintaining their individuality, without requiring special photography skills or a professional studio.

[0258] The following describes the processing flow.

[0259] Step 1:

[0260] The user takes pictures of themselves from multiple angles using a user device such as a smartphone or tablet. The device temporarily stores the captured images of the person within the application.

[0261] Step 2:

[0262] Users select their desired impression from several impression categories provided within the application. These categories include "Casual," "Formal," and "Natural." Users can also enter detailed requests as text if needed.

[0263] Step 3:

[0264] The device compresses the captured image and sends it to the server. Optimization is performed during transmission to avoid compromising image quality. The server saves the received image data to its storage.

[0265] Step 4:

[0266] The server applies machine learning algorithms to the received images of people to perform image analysis. The analysis extracts facial feature points to identify the user's individuality and characteristics.

[0267] Step 5:

[0268] Based on the analysis results and the user's selected desired impression, the server uses an AI model to suggest appropriate clothing and posing. The AI ​​refers to a database based on impressions to identify the best suggestions.

[0269] Step 6:

[0270] The server uses a generation mechanism to generate a modified image that reflects the proposed clothing and posing. In doing so, it emphasizes the desired impression while preserving the user's characteristics.

[0271] Step 7:

[0272] The server sends multiple modified images to the user's terminal. The terminal presents these images to the user and prompts them to make a selection.

[0273] Step 8:

[0274] The user selects the most preferred modified image from the presented options. The device saves the selected modified image to local storage.

[0275] Step 9:

[0276] The device is configured to allow the selected modified image to be used as a profile picture on dating apps and matchmaking services. It also provides options to share the image on other platforms as needed.

[0277] (Example 1)

[0278] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0279] In conventional technology, creating attractive profile pictures required specialized skills and equipment. This presented a challenge: it was difficult for ordinary users to easily create impressive images that reflected their own personality.

[0280] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0281] In this invention, the server includes an information analysis means for performing image analysis, a suggestion means for proposing appearance and posture, and an image generation means for generating modified images. This makes it possible for users to easily generate and utilize images with diverse impressions that reflect their own characteristics.

[0282] A "user device" is an electronic device used by an individual, and is a device that processes text data and image data.

[0283] "Image data acquisition means" refers to means that have the function of allowing a user device to capture or receive an image and save it as digital data.

[0284] The "impression selection means" is a function that provides an interface for a user to select a specific style or theme.

[0285] The "communication means" is a function for transmitting and receiving data between a user device and a computing device (server).

[0286] The "computing device" is a computing device for processing input data, and particularly refers to a server.

[0287] The "information analysis means" is a means having a function for analyzing received image data and extracting facial features and main parts.

[0288] The "proposal means" is a function within the system for generating an appearance and pose according to the user's wishes based on the analysis result.

[0289] The "image generation means" is a function for creating modified image data based on the proposed content.

[0290] The "transmission means" is a function for transmitting the generated image data to the user device.

[0291] The "storage means" is a function for storing the modified image selected by the user as digital data.

[0292] In the mode for implementing this invention, the user first starts the process using a terminal such as a smartphone or a tablet. A dedicated application is installed on the terminal, and the user uses this application to take a picture of himself / herself. This image data is stored within the application and functions as image data acquisition means.

[0293] Next, the user selects their desired impression within the application. Impression categories include "Casual," "Formal," and "Natural," allowing the user to choose the one that best matches their desired look. This function is implemented as an impression selection method. Users can also communicate specific preferences by entering prompt messages. For example, they can enter a prompt message such as, "I want a natural look that emphasizes my smile."

[0294] The captured images and selected impression information are sent to the server using the terminal's communication method. The server is equipped with a machine learning algorithm built on Python and analyzes facial features based on the received image data. Libraries such as OpenCV and TensorFlow are used for this analysis and function as an information analysis tool.

[0295] After obtaining the analysis results, the server uses a generative AI model to suggest the optimal appearance and posture for the user. This model utilizes technologies such as GANs and acts as a suggestion tool. Based on the analysis results and the user's preferences, the server generates appropriate modified images. This image generation is designed to reflect the user's individuality while achieving the desired impression.

[0296] The generated modified image is returned to the user terminal via the server's transmission mechanism. The user terminal presents the user with multiple images, from which the user can select one. The selected image is saved using the user terminal's storage mechanism and can be used as a profile picture.

[0297] Through this mechanism, this invention enables users to easily create and utilize attractive profile pictures that highlight their own unique characteristics, without requiring specialized skills or special equipment.

[0298] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0299] Step 1:

[0300] The user uses a terminal such as a smartphone to launch a dedicated application. The application is used to take a picture of the user's own image, and this image data is saved to the internal storage. The input is the raw image data taken by the camera, and the output is the saved image data. This saved data is used for the processing in the next section.

[0301] Step 2:

[0302] The user selects the impression they desire from the presented impression list on the application screen. This selection serves as the function of the impression selection means. Also, if necessary, the user can input a prompt sentence to specify the specific features of the desired image. As a result, the selected impression and the prompt sentence are obtained as the input, and the desired data is generated as the output.

[0303] Step 3:

[0304] The terminal transmits the image data and the user's impression desire obtained in the previous step to the server using network communication. In this process, the saved image data and the desired data obtained in the previous step are the input, and the dataset transmitted to the server is generated as the output. After transmission, analysis on the server side becomes possible.

[0305] Step 4:

[0306] The server accumulates the received image data in the analysis storage. Then, using OpenCV, TensorFlow, etc. of the Python library, the feature points of the face are analyzed and identified from the received image. The input is the image data transmitted to the server, and the output is the face feature data. Through this processing, the data for the next proposed step is organized.

[0307] Step 5:

[0308] The server uses a generative AI model to suggest appropriate appearances and postures based on acquired facial feature data and the user's desired impression. Generative models such as GANs are used in this process. The inputs are facial feature data and the user's desired impression, and the output is data related to the suggested clothing and posing.

[0309] Step 6:

[0310] The server generates a modified image based on the proposal. This image generation is performed using facial feature data and proposed data, reflecting the desired impression while preserving the original individuality. The inputs used are the proposed data and facial feature data, and the output is the generated modified image.

[0311] Step 7:

[0312] The server sends the generated modified images to the user's terminal. The user's terminal displays the received modified images in the application and prompts the user to make a selection. The input is the modified images sent from the server, and the output is the image selected by the user from among the displayed images.

[0313] Step 8:

[0314] The user selects their favorite image from several presented images and saves it to their device. At this stage, the user's selection is the input, and the image data saved on the device is the output. This saved image can be immediately used as a profile picture, for example.

[0315] (Application Example 1)

[0316] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0317] Online shopping presents a problem where consumers have difficulty visually assessing what suits them when choosing clothing. In particular, the inability to try on clothes in person often leads to disappointment after purchase, resulting in increased returns and exchanges, and ultimately lowering consumer satisfaction and sales efficiency.

[0318] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0319] In this invention, the server includes means for acquiring data of a person's image from a user terminal, means for suggesting clothing and posing based on the analysis results, and means for suggesting clothing and virtually trying it on based on the impression selected by the user. This allows the user to easily virtually try on clothing using their own image and visually confirm the results.

[0320] A "user terminal" refers to a communication device used individually by a user, such as a smartphone or tablet, that is equipped with image acquisition and communication functions.

[0321] "Data acquisition means" refers to a function that uses a user terminal to capture images of people and acquire that data.

[0322] "Selection methods" refer to interfaces and functions that allow users to choose impressions based on their own preferences and desires.

[0323] "Communication means" refers to the function for sending person image data from the user terminal to the server.

[0324] "Analysis means" refers to a function that processes human images received by the server and extracts feature quantities.

[0325] The "suggestion method" refers to a function that suggests appropriate clothing and posing to the user based on the analysis results.

[0326] The "generation means" refers to a function that, based on the proposal, modifies a user's image to generate a new image.

[0327] A "clothing try-on method" is a function that allows users to virtually try on suggested clothing.

[0328] "Storage method" refers to a function that allows users to save modified images they prefer on their device.

[0329] This invention begins with a user taking an image using a user terminal such as a smartphone or tablet. The user terminal is equipped with communication means to compress the acquired image of a person and send it to a server. The server uses analysis means to analyze the received image data and extract facial features from the image.

[0330] Next, the server uses the analysis results to suggest the most suitable clothing and posing to match the impression selected by the user. By using a suggestion method that utilizes a generation AI model, for example, if the user selects "casual and energetic impression," clothing and poses suitable for that impression will be suggested. Based on this suggestion, a new modified image is generated by the generation method and sent to the user's terminal.

[0331] The generated modified images are presented to the user. The user can select the most suitable image from multiple suggestions and save it to their device. The server also provides an environment where the user can virtually try on clothes using a clothing try-on system. This allows users to have an experience on e-commerce sites that is as if they were actually trying on clothes.

[0332] The hardware used includes smartphones and tablets as user terminals, and cloud computing environments as servers. The software includes image processing libraries (e.g., OpenCV) and AI model execution environments (e.g., TensorFlow).

[0333] For example, if a salaryman in his 40s requests a "formal and trustworthy impression," the AI ​​will generate a modified image of him in a suit. This image can be used as a profile picture in business settings.

[0334] An example of a prompt would be, "Suggest formal attire for a man in his 40s and synthesize it to create an impression of trustworthiness."

[0335] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0336] Step 1:

[0337] The user takes a picture of themselves using a user device such as a smartphone or tablet. The input is the desired impression selected by the user and the captured image of the person, and the output is that image data. This image data is saved on the smartphone in the optimal format.

[0338] Step 2:

[0339] The user terminal compresses the captured image of a person and sends it to the server. The input is the uncompressed image data, and the output is the compressed image data. The user terminal uploads this compressed image to the cloud server using a communication method.

[0340] Step 3:

[0341] The server analyzes the received image data. The input is compressed image data sent from the user's terminal, and the output is extracted facial feature data. The server uses an image analysis algorithm (e.g., OpenCV's face detection function) and temporarily stores the features in a database.

[0342] Step 4:

[0343] Based on the analysis results, the server uses an AI model to generate optimal clothing and posing, taking into account the user's selected desired impression. The input is facial feature data and the user's desired impression, and the output is proposed clothing and posing data. TensorFlow is used for the AI ​​model, and the generated data is proposed based on the prompts of the generating AI model.

[0344] Step 5:

[0345] The server uses a generation mechanism to generate modified images based on the proposed data. The input is the proposed clothing and posing data, and the output is the modified image data. Image processing is performed within the server to generate the modified image.

[0346] Step 6:

[0347] The server sends the generated modified image to the user's terminal. The input is the modified image data, and the output is the image data displayed on the user's terminal. This data is then transmitted to the user's smartphone using a communication method.

[0348] Step 7:

[0349] The user selects the most suitable image from several modified images and saves it on their device. The input consists of multiple modified image data sent from the server, and the output is the image data selected and saved by the user. This allows users to use images that convey their desired impression, such as profile pictures.

[0350] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0351] In an embodiment of this invention, the user uses a user terminal such as a smartphone or tablet to take a picture of themselves and then uses an emotion engine to analyze the user's emotions. During the shooting process, the terminal senses the user's facial expressions and voice, and uses this data to determine the user's emotions.

[0352] The device sends the captured image of the person, along with emotional data analyzed by the emotion engine, to the server. The server stores the received image data and emotional data in storage and performs a detailed analysis based on them. Using the analysis method, the server extracts facial feature points and, together with the emotional data, comprehensively evaluates the user's impression.

[0353] Next, the server uses a suggestion mechanism to propose the most suitable clothing and posing based on the user's desired impression and perceived emotions. For example, if the emotion engine recognizes the user's emotion as "joy," the server will suggest bright and cheerful clothing and posing.

[0354] The server uses a generation mechanism to create modified images incorporating the suggested clothing and posing. During this process, subtle adjustments to the color scheme and atmosphere are also made to reflect the emotions. The generated modified images are prepared as multiple variations and sent to the user's terminal.

[0355] The user's device displays the received modified images to the user. The user selects the image that best matches their intended impression or emotion and saves it on the device. The saved images can then be easily used as profile pictures on dating apps, marriage agencies, etc.

[0356] This embodiment allows users to not only select their desired impression but also generate a profile picture that reflects their current emotions, enabling a more natural and appealing form of self-expression.

[0357] The following describes the processing flow.

[0358] Step 1:

[0359] The user takes images of themselves from multiple angles using a user device such as a smartphone or tablet. The emotion engine is activated to detect the user's facial expressions and voice in real time and acquire emotion data.

[0360] Step 2:

[0361] The device saves the captured image of the person and emotion data to temporary memory. The user selects their desired impression from the impression categories provided within the application and enters detailed requests as needed.

[0362] Step 3:

[0363] The device compresses the saved person images and emotion data and sends them to the server. The image data compression is optimized to improve communication efficiency while maintaining image quality.

[0364] Step 4:

[0365] The server stores received person image data and emotion data in storage. Image analysis tools extract facial feature points from the received images and analyze the user's overall impression by combining them with emotion data.

[0366] Step 5:

[0367] Based on the analysis results, the user's selected desired impression, and the perceived emotions, the server uses an AI model to suggest the most suitable clothing and posing. For example, if the server analyzes the emotion as "joy," it will suggest brightly colored clothing and dynamic posing.

[0368] Step 6:

[0369] The server uses a generation mechanism to create modified images that reflect the proposed clothing and posing. Furthermore, it makes real-time adjustments to the color tone and expression according to the emotion. The generated modified images are prepared as multiple variations.

[0370] Step 7:

[0371] The server sends multiple modified images to the user's terminal. The terminal presents these images to the user and provides an interface for the user to make a selection.

[0372] Step 8:

[0373] The user selects the modified image that best reflects their intended impression or emotion from the presented images. The device then offers the option to save the selected modified image to local storage and make it immediately available as a profile picture.

[0374] (Example 2)

[0375] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0376] In modern society, it is difficult for individuals to easily create profile pictures that effectively reflect their real emotions and desired impression when expressing themselves on online platforms. As a result, they may not be able to express themselves more naturally, potentially leading to misunderstandings. Furthermore, traditional methods do not adequately provide emotion-based image adjustments or multiple suggestion options, making it difficult for users to pursue a more ideal self-expression.

[0377] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0378] In this invention, the server includes an analysis means for analyzing features from an image, an evaluation means for evaluating impressions based on the analysis results and emotion data, and a suggestion means for generating suggestions based on the evaluation. This enables users to generate more unique and attractive profile images that reflect their real emotions. Furthermore, by providing multiple modified images and making adjustments based on the desired impression, the system can meet the diverse needs of users for self-expression.

[0379] A "user device" is a digital device used by a user that enables image capture and data manipulation.

[0380] "Means of acquiring image data" refers to functions and technologies for capturing image data such as a user's face.

[0381] "Means for generating emotional data" refers to technologies that analyze captured image data to quantify or categorize the user's emotional state.

[0382] "Communication means" refers to the technologies and protocols used to send and receive data between user devices and processing devices.

[0383] A "processing device" is a computing device used to analyze and process received data, and primarily refers to a server.

[0384] "Analysis methods for analyzing features from images" refers to technologies that identify and extract facial features and shapes from image data.

[0385] "An evaluation method for assessing impressions" refers to a technology that comprehensively estimates a user's impression based on analyzed characteristics and emotional data.

[0386] "A suggestion generation method" refers to a technology that provides users with ideas for the most suitable clothing and posing based on evaluation results.

[0387] "Generation means" refers to the technology for modifying and generating images based on the proposed content.

[0388] "Transmission means" refers to the technology and protocols used to send the generated modified image to the user's device.

[0389] "Storage method" refers to the technology or method of storing images selected by the user in a data storage device.

[0390] To implement this invention, the user first takes an image of themselves using an electronic device such as a smartphone or tablet. At this time, the camera and microphone built into the device acquire the user's facial expressions and voice as data. This data is processed by emotion analysis software installed internally, and the user's emotions are quantified or categorized.

[0391] Next, the device sends the acquired image and emotion data to the server. HTTPS, a standard protocol over the internet, is used for communication. The server utilizes high-performance computing resources to analyze the image's feature points using image processing software. This is done, for example, by facial recognition algorithms. The emotion data is also used to comprehensively evaluate the user's impression.

[0392] Subsequently, the server generates suggestions based on the user's evaluation. These suggestions utilize a generative AI model and include clothing and posing that reflect the user's emotions and impressions. The generated suggestions are then used to create modified images, which are provided to the user as multiple image variations.

[0393] The generated images are sent to the device, allowing the user to select and save the one that best suits their desired impression and emotions. Saved images can also be used in dating apps and lifestyle facilities. This system allows users to easily create unique profile pictures that not only reflect their desired impression but also their current emotions.

[0394] As a concrete example, consider a scenario where a user wants to create a relaxed profile picture during a fun music festival. The user takes a photo with their device, and if the system determines the emotion is "joyful," the server suggests casual and cheerful clothing and posing. An example of a prompt used in this process might be, "Generate a profile picture with a fun atmosphere, casual clothing, and pose."

[0395] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0396] Step 1:

[0397] The user takes a picture of their face using a smartphone or tablet. At this time, the camera and microphone built into the device capture the user's facial expressions and voice, inputting this as image and audio data into the device. The device receives this data as initial input and prepares for emotion analysis.

[0398] Step 2:

[0399] The device processes the acquired image and audio data using internal emotion analysis software. This analysis extracts facial features from images and analyzes tone patterns from audio. As output of the analysis, emotion data is generated, classifying the user's emotions into categories such as "joy" and "sadness."

[0400] Step 3:

[0401] The terminal sends image data and emotion data obtained through analysis to the server. Secure communication via the internet is used for transmission, and the data is input to the server. The server prepares the received data for the next analysis.

[0402] Step 4:

[0403] The server uses image processing software to analyze facial feature points from image data. This analysis phase applies advanced facial recognition algorithms, quantifying image details and outputting them to an internal database. Based on these analysis results and sentiment data, the server evaluates the user's impression.

[0404] Step 5:

[0405] The server uses the results of analysis and evaluation to suggest appropriate clothing and posing for the user. The generative AI model then creates prompt statements based on this process, forming the core of the suggestions. For example, it might generate specific instructions such as "casual clothing and poses in a fun atmosphere." These prompt statements are then output and input into the generation mechanism.

[0406] Step 6:

[0407] The server receives prompt text into the generation mechanism and generates a modified image according to the suggested content. At this stage, multiple image variations with different color schemes and posing are generated and prepared within the server.

[0408] Step 7:

[0409] The server sends the generated modified image to the terminal. The terminal presents the received image to the user and displays multiple options on the screen, allowing the user to select the best image.

[0410] Step 8:

[0411] The user selects the most suitable image from the presented modified images and saves it to their device. This saved image can then be immediately used as a profile picture across various online services.

[0412] (Application Example 2)

[0413] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0414] In modern brick-and-mortar stores, it is difficult to efficiently suggest products based on the subjective feelings and impressions of individual customers, and there is a need for technology that enables product selection that reflects such feelings. Furthermore, there is a need for a way for customers to easily check styles optimized for their feelings without having to try products in-store.

[0415] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0416] In this invention, the server includes an analysis means for analyzing a person's image, a suggestion means for analyzing emotional information and proposing clothing and posing, and a generation means for providing the generated modified image as a virtual try-on. This enables personalized product suggestions that take into account the user's emotional state and a virtual try-on experience.

[0417] A "user terminal" is a portable electronic device used to capture and display images of people.

[0418] A "personal image" is a digital image data that captures the user's face or entire body.

[0419] An "information processing device" is a computing system that performs analysis and makes suggestions based on received image and emotion data.

[0420] "Communication means" refers to network communication functions for sending and receiving data between a user terminal and an information processing device.

[0421] The "analysis method" is a function that performs analysis based on input human images and emotional information to extract features.

[0422] The "suggestion method" refers to a function that selects and suggests the optimal clothing and posing based on the analysis results.

[0423] The "generation means" is a function that creates modified images based on the clothing and posing selected by the proposed means.

[0424] "Storage method" refers to a digital storage medium for long-term storage of modified images selected by the user.

[0425] "Emotion analysis means" refers to technology that detects emotions from the user's facial expressions and voice, and reflects the results in other proposed means.

[0426] "Virtual try-on" is a technology that allows users to visually try on clothing digitally without actually trying on the product.

[0427] This invention is a system that acquires images of a person using a user's terminal and then provides optimal clothing and style suggestions through emotion analysis based on those images. Photographs taken by the user with a smartphone or other terminal are transmitted to an information processing device via communication means. The information processing device extracts facial features from the person's image and detects the user's emotions using emotion analysis means. This process utilizes image processing libraries such as OpenCV and Microsoft's Azure Emotion API.

[0428] The information processing device, based on the analysis results, uses a generative AI model to suggest clothing and styling that best suit the user's emotions and impressions. The suggestion method utilizes the generative AI model and generates a variety of fashion styles using prompt sentences. An example of such a prompt sentence is, "Generate styling suggestions for when the user is expressing cheerful emotions. Please provide ideas that emphasize bright colors and casual clothing."

[0429] The user's device receives the modified image again via communication and displays multiple styling options on the screen. The user can then select their preferred style and save the result on their device using the saving function. This feature allows for an instant virtual try-on experience while in a physical store, supporting purchase decisions. For example, when a user is enjoying a pleasant shopping experience in a store, they may be offered suggestions for bright, casual wear, naturally facilitating a purchase.

[0430] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0431] Step 1:

[0432] The user terminal captures an image of the user with its camera and acquires the image data. The input is the image data of the person captured by the camera, and the output is the generated image data. The terminal immediately prepares to transmit this image to the information processing device via a communication means.

[0433] Step 2:

[0434] The terminal transmits captured image data to the information processing device using a communication method. The input is the acquired image data, and the output is the image data transferred to the information processing device via the network.

[0435] Step 3:

[0436] The information processing device extracts feature points from the user's face using an analysis mechanism based on the received image data. In this process, the input is image data transmitted from the terminal, and the output is facial feature point data. The processing is performed using an image processing library such as OpenCV.

[0437] Step 4:

[0438] The information processing device detects a user's emotions from facial feature point data using emotion analysis means. The input here is facial feature point data, and the output is recognized emotion information. Emotion detection utilizes Microsoft's Azure Emotion API, among others.

[0439] Step 5:

[0440] Based on emotional information, the information processing device uses a proposed means and a generative AI model to consider the optimal clothing and posing. The input is the user's emotional information, and the output is the proposed clothing and posing data. The generative AI model is given instructions via prompts to generate the optimal style.

[0441] Step 6:

[0442] Using the information determined by the proposed method, the generation method of the information processing device generates a modified image. The input is clothing and posing data, as well as the original person image, and the output is a modified virtual try-on image. This is done using a generation AI model.

[0443] Step 7:

[0444] The information processing device transmits the generated modified image to the user terminal. The input is the generated modified image, and the output is the modified image transferred to the terminal.

[0445] Step 8:

[0446] The user terminal presents the received modified image to the user and displays multiple variations. The input is a pre-generated modified image, and the output is multiple virtual try-on images that the user can visually confirm. The image selected by the user is saved and stored via a storage means for future reference or purchase decisions.

[0447] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0448] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0449] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0450] [Third Embodiment]

[0451] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0452] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0453] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0454] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0455] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0456] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0457] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0458] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0459] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0460] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0461] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0462] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0463] In an embodiment of this invention, the user initiates the process using a user terminal such as a smartphone or tablet. The user takes a picture of themselves, and the image data is stored within the application on the user terminal.

[0464] Users select their desired impression from a list displayed within the application. The options are divided into several categories, such as "Casual," "Formal," and "Natural," allowing users to choose the one that best matches their desired impression. They can also enter more detailed requests in a text box.

[0465] Next, the user terminal compresses the captured image of the person and sends it to the server. The server temporarily stores the received image data and starts analyzing it by applying a machine learning algorithm. Through this analysis, the server extracts facial features and key components from the image.

[0466] Based on this analysis and the user's desired impression, the server uses an AI model to suggest the most suitable clothing and posing. For example, if the user selects "casual and energetic impression," the AI ​​will generate suggestions that reflect this, such as casual clothing and cheerful poses.

[0467] The server then uses a generation mechanism to create a new modified image based on the proposed clothing and posing. This creates an image that matches the desired impression without compromising the user's original features. Multiple versions of this modified image are prepared and sent from the server to the user's terminal.

[0468] The user's device displays the received modified images to the user. The user can select their favorite image from the multiple images presented and save it to their device. The saved image is configured to be used directly as a profile picture on dating apps or marriage agencies.

[0469] In this embodiment of the present invention, users can easily create and utilize attractive profile pictures while maintaining their individuality, without requiring special photography skills or a professional studio.

[0470] The following describes the processing flow.

[0471] Step 1:

[0472] The user takes pictures of themselves from multiple angles using a user device such as a smartphone or tablet. The device temporarily stores the captured images of the person within the application.

[0473] Step 2:

[0474] Users select their desired impression from several impression categories provided within the application. These categories include "Casual," "Formal," and "Natural." Users can also enter detailed requests as text if needed.

[0475] Step 3:

[0476] The device compresses the captured image and sends it to the server. Optimization is performed during transmission to avoid compromising image quality. The server saves the received image data to its storage.

[0477] Step 4:

[0478] The server applies machine learning algorithms to the received images of people to perform image analysis. The analysis extracts facial feature points to identify the user's individuality and characteristics.

[0479] Step 5:

[0480] Based on the analysis results and the user's selected desired impression, the server uses an AI model to suggest appropriate clothing and posing. The AI ​​refers to a database based on impressions to identify the best suggestions.

[0481] Step 6:

[0482] The server uses a generation mechanism to generate a modified image that reflects the proposed clothing and posing. In doing so, it emphasizes the desired impression while preserving the user's characteristics.

[0483] Step 7:

[0484] The server sends multiple modified images to the user's terminal. The terminal presents these images to the user and prompts them to make a selection.

[0485] Step 8:

[0486] The user selects the most preferred modified image from the presented options. The device saves the selected modified image to local storage.

[0487] Step 9:

[0488] The device is configured to allow the selected modified image to be used as a profile picture on dating apps and matchmaking services. It also provides options to share the image on other platforms as needed.

[0489] (Example 1)

[0490] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0491] In conventional technology, creating attractive profile pictures required specialized skills and equipment. This presented a challenge: it was difficult for ordinary users to easily create impressive images that reflected their own personality.

[0492] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0493] In this invention, the server includes an information analysis means for performing image analysis, a suggestion means for proposing appearance and posture, and an image generation means for generating modified images. This makes it possible for users to easily generate and utilize images with diverse impressions that reflect their own characteristics.

[0494] A "user device" is an electronic device used by an individual, and is a device that processes text data and image data.

[0495] "Image data acquisition means" refers to means that have the function of allowing a user device to capture or receive an image and save it as digital data.

[0496] An "impression selection method" is a function that provides an interface for users to select a specific style or theme.

[0497] "Communication means" refers to the function for sending and receiving data between a user device and a computing device (server).

[0498] A "computing device" refers to a computing device used to process input data, and specifically to a server.

[0499] "Information analysis means" refers to means that have the function of analyzing received image data and extracting facial features and major body parts.

[0500] "Proposed means" refers to a function within the system that generates an appearance and posture according to the user's wishes based on the analysis results.

[0501] The "image generation means" is a function for creating modified image data based on the proposed content.

[0502] "Transmission means" refers to a function for transmitting the generated image data to the user's device.

[0503] "Memory function" refers to a function for saving the modified image selected by the user as digital data.

[0504] In an embodiment of this invention, the user first initiates the process using a device such as a smartphone or tablet. A dedicated application is installed on the device, and the user uses this application to take an image of themselves. This image data is stored within the application and functions as an image data acquisition means.

[0505] Next, the user selects their desired impression within the application. Impression categories include "Casual," "Formal," and "Natural," allowing the user to choose the one that best matches their desired look. This function is implemented as an impression selection method. Users can also communicate specific preferences by entering prompt messages. For example, they can enter a prompt message such as, "I want a natural look that emphasizes my smile."

[0506] The captured images and selected impression information are sent to the server using the terminal's communication method. The server is equipped with a machine learning algorithm built on Python and analyzes facial features based on the received image data. Libraries such as OpenCV and TensorFlow are used for this analysis and function as an information analysis tool.

[0507] After obtaining the analysis results, the server uses a generative AI model to suggest the optimal appearance and posture for the user. This model utilizes technologies such as GANs and acts as a suggestion tool. Based on the analysis results and the user's preferences, the server generates appropriate modified images. This image generation is designed to reflect the user's individuality while achieving the desired impression.

[0508] The generated modified image is returned to the user terminal via the server's transmission mechanism. The user terminal presents the user with multiple images, from which the user can select one. The selected image is saved using the user terminal's storage mechanism and can be used as a profile picture.

[0509] Through this mechanism, this invention enables users to easily create and utilize attractive profile pictures that highlight their own unique characteristics, without requiring specialized skills or special equipment.

[0510] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0511] Step 1:

[0512] The user uses a smartphone or other device to launch a dedicated application. Using the application, they take a picture of themselves and save this image data to internal storage. The input is the raw image data captured by the camera, and the output is the saved image data. This saved data will be used for processing in the next section.

[0513] Step 2:

[0514] On the application screen, the user selects their desired impression from a presented list of impressions. This selection serves as the impression selection method. Additionally, the user can enter prompt text as needed to specify the desired image characteristics. This results in the selected impression and prompt text as input, and the desired data is generated as output.

[0515] Step 3:

[0516] The terminal transmits the image data acquired in the previous step and the user's desired impression to the server via network communication. In this process, the stored image data and desired data obtained in the previous step are taken as input, and a dataset sent to the server is generated as output. After transmission, analysis can be performed on the server side.

[0517] Step 4:

[0518] The server stores the received image data in analysis storage. Then, using Python libraries such as OpenCV and TensorFlow, it analyzes and identifies facial feature points from the received images. The input is the image data sent to the server, and the output is facial feature data. This process prepares the data for the next proposed step.

[0519] Step 5:

[0520] The server uses a generative AI model to suggest appropriate appearances and postures based on acquired facial feature data and the user's desired impression. Generative models such as GANs are used in this process. The inputs are facial feature data and the user's desired impression, and the output is data related to the suggested clothing and posing.

[0521] Step 6:

[0522] The server generates a modified image based on the proposal. This image generation is performed using facial feature data and proposed data, reflecting the desired impression while preserving the original individuality. The inputs used are the proposed data and facial feature data, and the output is the generated modified image.

[0523] Step 7:

[0524] The server sends the generated modified images to the user's terminal. The user's terminal displays the received modified images in the application and prompts the user to make a selection. The input is the modified images sent from the server, and the output is the image selected by the user from among the displayed images.

[0525] Step 8:

[0526] The user selects their favorite image from several presented images and saves it to their device. At this stage, the user's selection is the input, and the image data saved on the device is the output. This saved image can be immediately used as a profile picture, for example.

[0527] (Application Example 1)

[0528] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0529] Online shopping presents a problem where consumers have difficulty visually assessing what suits them when choosing clothing. In particular, the inability to try on clothes in person often leads to disappointment after purchase, resulting in increased returns and exchanges, and ultimately lowering consumer satisfaction and sales efficiency.

[0530] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0531] In this invention, the server includes means for acquiring data of a person's image from a user terminal, means for suggesting clothing and posing based on the analysis results, and means for suggesting clothing and virtually trying it on based on the impression selected by the user. This allows the user to easily virtually try on clothing using their own image and visually confirm the results.

[0532] A "user terminal" refers to a communication device used individually by a user, such as a smartphone or tablet, that is equipped with image acquisition and communication functions.

[0533] "Data acquisition means" refers to a function that uses a user terminal to capture images of people and acquire that data.

[0534] "Selection methods" refer to interfaces and functions that allow users to choose impressions based on their own preferences and desires.

[0535] "Communication means" refers to the function for sending person image data from the user terminal to the server.

[0536] "Analysis means" refers to a function that processes human images received by the server and extracts feature quantities.

[0537] The "suggestion method" refers to a function that suggests appropriate clothing and posing to the user based on the analysis results.

[0538] The "generation means" refers to a function that, based on the proposal, modifies a user's image to generate a new image.

[0539] A "clothing try-on method" is a function that allows users to virtually try on suggested clothing.

[0540] "Storage method" refers to a function that allows users to save modified images they prefer on their device.

[0541] This invention begins with a user taking an image using a user terminal such as a smartphone or tablet. The user terminal is equipped with communication means to compress the acquired image of a person and send it to a server. The server uses analysis means to analyze the received image data and extract facial features from the image.

[0542] Next, the server uses the analysis results to suggest the most suitable clothing and posing to match the impression selected by the user. By using a suggestion method that utilizes a generation AI model, for example, if the user selects "casual and energetic impression," clothing and poses suitable for that impression will be suggested. Based on this suggestion, a new modified image is generated by the generation method and sent to the user's terminal.

[0543] The generated modified images are presented to the user. The user can select the most suitable image from multiple suggestions and save it to their device. The server also provides an environment where the user can virtually try on clothes using a clothing try-on system. This allows users to have an experience on e-commerce sites that is as if they were actually trying on clothes.

[0544] The hardware used includes smartphones and tablets as user terminals, and cloud computing environments as servers. The software includes image processing libraries (e.g., OpenCV) and AI model execution environments (e.g., TensorFlow).

[0545] For example, if a salaryman in his 40s requests a "formal and trustworthy impression," the AI ​​will generate a modified image of him in a suit. This image can be used as a profile picture in business settings.

[0546] An example of a prompt would be, "Suggest formal attire for a man in his 40s and synthesize it to create an impression of trustworthiness."

[0547] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0548] Step 1:

[0549] The user takes a picture of themselves using a user device such as a smartphone or tablet. The input is the desired impression selected by the user and the captured image of the person, and the output is that image data. This image data is saved on the smartphone in the optimal format.

[0550] Step 2:

[0551] The user terminal compresses the captured image of a person and sends it to the server. The input is the uncompressed image data, and the output is the compressed image data. The user terminal uploads this compressed image to the cloud server using a communication method.

[0552] Step 3:

[0553] The server analyzes the received image data. The input is compressed image data sent from the user's terminal, and the output is extracted facial feature data. The server uses an image analysis algorithm (e.g., OpenCV's face detection function) and temporarily stores the features in a database.

[0554] Step 4:

[0555] Based on the analysis results, the server uses an AI model to generate optimal clothing and posing, taking into account the user's selected desired impression. The input is facial feature data and the user's desired impression, and the output is proposed clothing and posing data. TensorFlow is used for the AI ​​model, and the generated data is proposed based on the prompts of the generating AI model.

[0556] Step 5:

[0557] The server uses a generation mechanism to generate modified images based on the proposed data. The input is the proposed clothing and posing data, and the output is the modified image data. Image processing is performed within the server to generate the modified image.

[0558] Step 6:

[0559] The server sends the generated modified image to the user's terminal. The input is the modified image data, and the output is the image data displayed on the user's terminal. This data is then transmitted to the user's smartphone using a communication method.

[0560] Step 7:

[0561] The user selects the most suitable image from several modified images and saves it on their device. The input consists of multiple modified image data sent from the server, and the output is the image data selected and saved by the user. This allows users to use images that convey their desired impression, such as profile pictures.

[0562] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0563] In an embodiment of this invention, the user uses a user terminal such as a smartphone or tablet to take a picture of themselves and then uses an emotion engine to analyze the user's emotions. During the shooting process, the terminal senses the user's facial expressions and voice, and uses this data to determine the user's emotions.

[0564] The device sends the captured image of the person, along with emotional data analyzed by the emotion engine, to the server. The server stores the received image data and emotional data in storage and performs a detailed analysis based on them. Using the analysis method, the server extracts facial feature points and, together with the emotional data, comprehensively evaluates the user's impression.

[0565] Next, the server uses a suggestion mechanism to propose the most suitable clothing and posing based on the user's desired impression and perceived emotions. For example, if the emotion engine recognizes the user's emotion as "joy," the server will suggest bright and cheerful clothing and posing.

[0566] The server uses a generation mechanism to create modified images incorporating the suggested clothing and posing. During this process, subtle adjustments to the color scheme and atmosphere are also made to reflect the emotions. The generated modified images are prepared as multiple variations and sent to the user's terminal.

[0567] The user's device displays the received modified images to the user. The user selects the image that best matches their intended impression or emotion and saves it on the device. The saved images can then be easily used as profile pictures on dating apps, marriage agencies, etc.

[0568] This embodiment allows users to not only select their desired impression but also generate a profile picture that reflects their current emotions, enabling a more natural and appealing form of self-expression.

[0569] The following describes the processing flow.

[0570] Step 1:

[0571] The user takes images of themselves from multiple angles using a user device such as a smartphone or tablet. The emotion engine is activated to detect the user's facial expressions and voice in real time and acquire emotion data.

[0572] Step 2:

[0573] The device saves the captured image of the person and emotion data to temporary memory. The user selects their desired impression from the impression categories provided within the application and enters detailed requests as needed.

[0574] Step 3:

[0575] The device compresses the saved person images and emotion data and sends them to the server. The image data compression is optimized to improve communication efficiency while maintaining image quality.

[0576] Step 4:

[0577] The server stores received person image data and emotion data in storage. Image analysis tools extract facial feature points from the received images and analyze the user's overall impression by combining them with emotion data.

[0578] Step 5:

[0579] Based on the analysis results, the user's selected desired impression, and the perceived emotions, the server uses an AI model to suggest the most suitable clothing and posing. For example, if the server analyzes the emotion as "joy," it will suggest brightly colored clothing and dynamic posing.

[0580] Step 6:

[0581] The server uses a generation mechanism to create modified images that reflect the proposed clothing and posing. Furthermore, it makes real-time adjustments to the color tone and expression according to the emotion. The generated modified images are prepared as multiple variations.

[0582] Step 7:

[0583] The server sends multiple modified images to the user's terminal. The terminal presents these images to the user and provides an interface for the user to make a selection.

[0584] Step 8:

[0585] The user selects the modified image that best reflects their intended impression or emotion from the presented images. The device then offers the option to save the selected modified image to local storage and make it immediately available as a profile picture.

[0586] (Example 2)

[0587] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0588] In modern society, it is difficult for individuals to easily create profile pictures that effectively reflect their real emotions and desired impression when expressing themselves on online platforms. As a result, they may not be able to express themselves more naturally, potentially leading to misunderstandings. Furthermore, traditional methods do not adequately provide emotion-based image adjustments or multiple suggestion options, making it difficult for users to pursue a more ideal self-expression.

[0589] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0590] In this invention, the server includes an analysis means for analyzing features from an image, an evaluation means for evaluating impressions based on the analysis results and emotion data, and a suggestion means for generating suggestions based on the evaluation. This enables users to generate more unique and attractive profile images that reflect their real emotions. Furthermore, by providing multiple modified images and making adjustments based on the desired impression, the system can meet the diverse needs of users for self-expression.

[0591] A "user device" is a digital device used by a user that enables image capture and data manipulation.

[0592] "Means of acquiring image data" refers to functions and technologies for capturing image data such as a user's face.

[0593] "Means for generating emotional data" refers to technologies that analyze captured image data to quantify or categorize the user's emotional state.

[0594] "Communication means" refers to the technologies and protocols used to send and receive data between user devices and processing devices.

[0595] A "processing device" is a computing device used to analyze and process received data, and primarily refers to a server.

[0596] "Analysis methods for analyzing features from images" refers to technologies that identify and extract facial features and shapes from image data.

[0597] "An evaluation method for assessing impressions" refers to a technology that comprehensively estimates a user's impression based on analyzed characteristics and emotional data.

[0598] "A suggestion generation method" refers to a technology that provides users with ideas for the most suitable clothing and posing based on evaluation results.

[0599] "Generation means" refers to the technology for modifying and generating images based on the proposed content.

[0600] "Transmission means" refers to the technology and protocols used to send the generated modified image to the user's device.

[0601] "Storage method" refers to the technology or method of storing images selected by the user in a data storage device.

[0602] To implement this invention, the user first takes an image of themselves using an electronic device such as a smartphone or tablet. At this time, the camera and microphone built into the device acquire the user's facial expressions and voice as data. This data is processed by emotion analysis software installed internally, and the user's emotions are quantified or categorized.

[0603] Next, the device sends the acquired image and emotion data to the server. HTTPS, a standard protocol over the internet, is used for communication. The server utilizes high-performance computing resources to analyze the image's feature points using image processing software. This is done, for example, by facial recognition algorithms. The emotion data is also used to comprehensively evaluate the user's impression.

[0604] Subsequently, the server generates suggestions based on the user's evaluation. These suggestions utilize a generative AI model and include clothing and posing that reflect the user's emotions and impressions. The generated suggestions are then used to create modified images, which are provided to the user as multiple image variations.

[0605] The generated images are sent to the device, allowing the user to select and save the one that best suits their desired impression and emotions. Saved images can also be used in dating apps and lifestyle facilities. This system allows users to easily create unique profile pictures that not only reflect their desired impression but also their current emotions.

[0606] As a concrete example, consider a scenario where a user wants to create a relaxed profile picture during a fun music festival. The user takes a photo with their device, and if the system determines the emotion is "joyful," the server suggests casual and cheerful clothing and posing. An example of a prompt used in this process might be, "Generate a profile picture with a fun atmosphere, casual clothing, and pose."

[0607] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0608] Step 1:

[0609] The user takes a picture of their face using a smartphone or tablet. At this time, the camera and microphone built into the device capture the user's facial expressions and voice, inputting this as image and audio data into the device. The device receives this data as initial input and prepares for emotion analysis.

[0610] Step 2:

[0611] The device processes the acquired image and audio data using internal emotion analysis software. This analysis extracts facial features from images and analyzes tone patterns from audio. As output of the analysis, emotion data is generated, classifying the user's emotions into categories such as "joy" and "sadness."

[0612] Step 3:

[0613] The terminal sends image data and emotion data obtained through analysis to the server. Secure communication via the internet is used for transmission, and the data is input to the server. The server prepares the received data for the next analysis.

[0614] Step 4:

[0615] The server uses image processing software to analyze facial feature points from image data. This analysis phase applies advanced facial recognition algorithms, quantifying image details and outputting them to an internal database. Based on these analysis results and sentiment data, the server evaluates the user's impression.

[0616] Step 5:

[0617] The server uses the results of analysis and evaluation to suggest appropriate clothing and posing for the user. The generative AI model then creates prompt statements based on this process, forming the core of the suggestions. For example, it might generate specific instructions such as "casual clothing and poses in a fun atmosphere." These prompt statements are then output and input into the generation mechanism.

[0618] Step 6:

[0619] The server receives prompt text into the generation mechanism and generates a modified image according to the suggested content. At this stage, multiple image variations with different color schemes and posing are generated and prepared within the server.

[0620] Step 7:

[0621] The server sends the generated modified image to the terminal. The terminal presents the received image to the user and displays multiple options on the screen, allowing the user to select the best image.

[0622] Step 8:

[0623] The user selects the most suitable image from the presented modified images and saves it to their device. This saved image can then be immediately used as a profile picture across various online services.

[0624] (Application Example 2)

[0625] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0626] In modern brick-and-mortar stores, it is difficult to efficiently suggest products based on the subjective feelings and impressions of individual customers, and there is a need for technology that enables product selection that reflects such feelings. Furthermore, there is a need for a way for customers to easily check styles optimized for their feelings without having to try products in-store.

[0627] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0628] In this invention, the server includes an analysis means for analyzing a person's image, a suggestion means for analyzing emotional information and proposing clothing and posing, and a generation means for providing the generated modified image as a virtual try-on. This enables personalized product suggestions that take into account the user's emotional state and a virtual try-on experience.

[0629] A "user terminal" is a portable electronic device used to capture and display images of people.

[0630] A "personal image" is a digital image data that captures the user's face or entire body.

[0631] An "information processing device" is a computing system that performs analysis and makes suggestions based on received image and emotion data.

[0632] "Communication means" refers to network communication functions for sending and receiving data between a user terminal and an information processing device.

[0633] The "analysis method" is a function that performs analysis based on input human images and emotional information to extract features.

[0634] The "suggestion method" refers to a function that selects and suggests the optimal clothing and posing based on the analysis results.

[0635] The "generation means" is a function that creates modified images based on the clothing and posing selected by the proposed means.

[0636] "Storage method" refers to a digital storage medium for long-term storage of modified images selected by the user.

[0637] "Emotion analysis means" refers to technology that detects emotions from the user's facial expressions and voice, and reflects the results in other proposed means.

[0638] "Virtual try-on" is a technology that allows users to visually try on clothing digitally without actually trying on the product.

[0639] This invention is a system that acquires images of a person using a user's terminal and then provides optimal clothing and style suggestions through emotion analysis based on those images. Photographs taken by the user with a smartphone or other terminal are transmitted to an information processing device via communication means. The information processing device extracts facial features from the person's image and detects the user's emotions using emotion analysis means. This process utilizes image processing libraries such as OpenCV and Microsoft's Azure Emotion API.

[0640] The information processing device, based on the analysis results, uses a generative AI model to suggest clothing and styling that best suit the user's emotions and impressions. The suggestion method utilizes the generative AI model and generates a variety of fashion styles using prompt sentences. An example of such a prompt sentence is, "Generate styling suggestions for when the user is expressing cheerful emotions. Please provide ideas that emphasize bright colors and casual clothing."

[0641] The user's device receives the modified image again via communication and displays multiple styling options on the screen. The user can then select their preferred style and save the result on their device using the saving function. This feature allows for an instant virtual try-on experience while in a physical store, supporting purchase decisions. For example, when a user is enjoying a pleasant shopping experience in a store, they may be offered suggestions for bright, casual wear, naturally facilitating a purchase.

[0642] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0643] Step 1:

[0644] The user terminal captures an image of the user with its camera and acquires the image data. The input is the image data of the person captured by the camera, and the output is the generated image data. The terminal immediately prepares to transmit this image to the information processing device via a communication means.

[0645] Step 2:

[0646] The terminal transmits captured image data to the information processing device using a communication method. The input is the acquired image data, and the output is the image data transferred to the information processing device via the network.

[0647] Step 3:

[0648] The information processing device extracts feature points from the user's face using an analysis mechanism based on the received image data. In this process, the input is image data transmitted from the terminal, and the output is facial feature point data. The processing is performed using an image processing library such as OpenCV.

[0649] Step 4:

[0650] The information processing device detects a user's emotions from facial feature point data using emotion analysis means. The input here is facial feature point data, and the output is recognized emotion information. Emotion detection utilizes Microsoft's Azure Emotion API, among others.

[0651] Step 5:

[0652] Based on emotional information, the information processing device uses a proposed means and a generative AI model to consider the optimal clothing and posing. The input is the user's emotional information, and the output is the proposed clothing and posing data. The generative AI model is given instructions via prompts to generate the optimal style.

[0653] Step 6:

[0654] Using the information determined by the proposed method, the generation method of the information processing device generates a modified image. The input is clothing and posing data, as well as the original person image, and the output is a modified virtual try-on image. This is done using a generation AI model.

[0655] Step 7:

[0656] The information processing device transmits the generated modified image to the user terminal. The input is the generated modified image, and the output is the modified image transferred to the terminal.

[0657] Step 8:

[0658] The user terminal presents the received modified image to the user and displays multiple variations. The input is a pre-generated modified image, and the output is multiple virtual try-on images that the user can visually confirm. The image selected by the user is saved and stored via a storage means for future reference or purchase decisions.

[0659] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0660] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0661] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0662] [Fourth Embodiment]

[0663] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0664] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0665] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0666] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0667] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0668] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0669] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0670] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0671] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0672] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0673] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0674] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0675] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0676] In an embodiment of this invention, the user initiates the process using a user terminal such as a smartphone or tablet. The user takes a picture of themselves, and the image data is stored within the application on the user terminal.

[0677] Users select their desired impression from a list displayed within the application. The options are divided into several categories, such as "Casual," "Formal," and "Natural," allowing users to choose the one that best matches their desired impression. They can also enter more detailed requests in a text box.

[0678] Next, the user terminal compresses the captured image of the person and sends it to the server. The server temporarily stores the received image data and starts analyzing it by applying a machine learning algorithm. Through this analysis, the server extracts facial features and key components from the image.

[0679] Based on this analysis and the user's desired impression, the server uses an AI model to suggest the most suitable clothing and posing. For example, if the user selects "casual and energetic impression," the AI ​​will generate suggestions that reflect this, such as casual clothing and cheerful poses.

[0680] The server then uses a generation mechanism to create a new modified image based on the proposed clothing and posing. This creates an image that matches the desired impression without compromising the user's original features. Multiple versions of this modified image are prepared and sent from the server to the user's terminal.

[0681] The user's device displays the received modified images to the user. The user can select their favorite image from the multiple images presented and save it to their device. The saved image is configured to be used directly as a profile picture on dating apps or marriage agencies.

[0682] In this embodiment of the present invention, users can easily create and utilize attractive profile pictures while maintaining their individuality, without requiring special photography skills or a professional studio.

[0683] The following describes the processing flow.

[0684] Step 1:

[0685] The user takes pictures of themselves from multiple angles using a user device such as a smartphone or tablet. The device temporarily stores the captured images of the person within the application.

[0686] Step 2:

[0687] Users select their desired impression from several impression categories provided within the application. These categories include "Casual," "Formal," and "Natural." Users can also enter detailed requests as text if needed.

[0688] Step 3:

[0689] The device compresses the captured image and sends it to the server. Optimization is performed during transmission to avoid compromising image quality. The server saves the received image data to its storage.

[0690] Step 4:

[0691] The server applies machine learning algorithms to the received images of people to perform image analysis. The analysis extracts facial feature points to identify the user's individuality and characteristics.

[0692] Step 5:

[0693] Based on the analysis results and the user's selected desired impression, the server uses an AI model to suggest appropriate clothing and posing. The AI ​​refers to a database based on impressions to identify the best suggestions.

[0694] Step 6:

[0695] The server uses a generation mechanism to generate a modified image that reflects the proposed clothing and posing. In doing so, it emphasizes the desired impression while preserving the user's characteristics.

[0696] Step 7:

[0697] The server sends multiple modified images to the user's terminal. The terminal presents these images to the user and prompts them to make a selection.

[0698] Step 8:

[0699] The user selects the most preferred modified image from the presented options. The device saves the selected modified image to local storage.

[0700] Step 9:

[0701] The device is configured to allow the selected modified image to be used as a profile picture on dating apps and matchmaking services. It also provides options to share the image on other platforms as needed.

[0702] (Example 1)

[0703] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0704] In conventional technology, creating attractive profile pictures required specialized skills and equipment. This presented a challenge: it was difficult for ordinary users to easily create impressive images that reflected their own personality.

[0705] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0706] In this invention, the server includes an information analysis means for performing image analysis, a suggestion means for proposing appearance and posture, and an image generation means for generating modified images. This makes it possible for users to easily generate and utilize images with diverse impressions that reflect their own characteristics.

[0707] A "user device" is an electronic device used by an individual, and is a device that processes text data and image data.

[0708] "Image data acquisition means" refers to means that have the function of allowing a user device to capture or receive an image and save it as digital data.

[0709] An "impression selection method" is a function that provides an interface for users to select a specific style or theme.

[0710] "Communication means" refers to the function for sending and receiving data between a user device and a computing device (server).

[0711] A "computing device" refers to a computing device used to process input data, and specifically to a server.

[0712] "Information analysis means" refers to means that have the function of analyzing received image data and extracting facial features and major body parts.

[0713] "Proposed means" refers to a function within the system that generates an appearance and posture according to the user's wishes based on the analysis results.

[0714] The "image generation means" is a function for creating modified image data based on the proposed content.

[0715] "Transmission means" refers to a function for transmitting the generated image data to the user's device.

[0716] "Memory function" refers to a function for saving the modified image selected by the user as digital data.

[0717] In an embodiment of this invention, the user first initiates the process using a device such as a smartphone or tablet. A dedicated application is installed on the device, and the user uses this application to take an image of themselves. This image data is stored within the application and functions as an image data acquisition means.

[0718] Next, the user selects their desired impression within the application. Impression categories include "Casual," "Formal," and "Natural," allowing the user to choose the one that best matches their desired look. This function is implemented as an impression selection method. Users can also communicate specific preferences by entering prompt messages. For example, they can enter a prompt message such as, "I want a natural look that emphasizes my smile."

[0719] The captured images and selected impression information are sent to the server using the terminal's communication method. The server is equipped with a machine learning algorithm built on Python and analyzes facial features based on the received image data. Libraries such as OpenCV and TensorFlow are used for this analysis and function as an information analysis tool.

[0720] After obtaining the analysis results, the server uses a generative AI model to suggest the optimal appearance and posture for the user. This model utilizes technologies such as GANs and acts as a suggestion tool. Based on the analysis results and the user's preferences, the server generates appropriate modified images. This image generation is designed to reflect the user's individuality while achieving the desired impression.

[0721] The generated modified image is returned to the user terminal via the server's transmission mechanism. The user terminal presents the user with multiple images, from which the user can select one. The selected image is saved using the user terminal's storage mechanism and can be used as a profile picture.

[0722] Through this mechanism, this invention enables users to easily create and utilize attractive profile pictures that highlight their own unique characteristics, without requiring specialized skills or special equipment.

[0723] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0724] Step 1:

[0725] The user uses a smartphone or other device to launch a dedicated application. Using the application, they take a picture of themselves and save this image data to internal storage. The input is the raw image data captured by the camera, and the output is the saved image data. This saved data will be used for processing in the next section.

[0726] Step 2:

[0727] On the application screen, the user selects their desired impression from a presented list of impressions. This selection serves as the impression selection method. Additionally, the user can enter prompt text as needed to specify the desired image characteristics. This results in the selected impression and prompt text as input, and the desired data is generated as output.

[0728] Step 3:

[0729] The terminal transmits the image data acquired in the previous step and the user's desired impression to the server via network communication. In this process, the stored image data and desired data obtained in the previous step are taken as input, and a dataset sent to the server is generated as output. After transmission, analysis can be performed on the server side.

[0730] Step 4:

[0731] The server stores the received image data in analysis storage. Then, using Python libraries such as OpenCV and TensorFlow, it analyzes and identifies facial feature points from the received images. The input is the image data sent to the server, and the output is facial feature data. This process prepares the data for the next proposed step.

[0732] Step 5:

[0733] The server uses a generative AI model to suggest appropriate appearances and postures based on acquired facial feature data and the user's desired impression. Generative models such as GANs are used in this process. The inputs are facial feature data and the user's desired impression, and the output is data related to the suggested clothing and posing.

[0734] Step 6:

[0735] The server generates a modified image based on the proposal. This image generation is performed using facial feature data and proposed data, reflecting the desired impression while preserving the original individuality. The inputs used are the proposed data and facial feature data, and the output is the generated modified image.

[0736] Step 7:

[0737] The server sends the generated modified images to the user's terminal. The user's terminal displays the received modified images in the application and prompts the user to make a selection. The input is the modified images sent from the server, and the output is the image selected by the user from among the displayed images.

[0738] Step 8:

[0739] The user selects their favorite image from several presented images and saves it to their device. At this stage, the user's selection is the input, and the image data saved on the device is the output. This saved image can be immediately used as a profile picture, for example.

[0740] (Application Example 1)

[0741] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0742] Online shopping presents a problem where consumers have difficulty visually assessing what suits them when choosing clothing. In particular, the inability to try on clothes in person often leads to disappointment after purchase, resulting in increased returns and exchanges, and ultimately lowering consumer satisfaction and sales efficiency.

[0743] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0744] In this invention, the server includes means for acquiring data of a person's image from a user terminal, means for suggesting clothing and posing based on the analysis results, and means for suggesting clothing and virtually trying it on based on the impression selected by the user. This allows the user to easily virtually try on clothing using their own image and visually confirm the results.

[0745] A "user terminal" refers to a communication device used individually by a user, such as a smartphone or tablet, that is equipped with image acquisition and communication functions.

[0746] "Data acquisition means" refers to a function that uses a user terminal to capture images of people and acquire that data.

[0747] "Selection methods" refer to interfaces and functions that allow users to choose impressions based on their own preferences and desires.

[0748] "Communication means" refers to the function for sending person image data from the user terminal to the server.

[0749] "Analysis means" refers to a function that processes human images received by the server and extracts feature quantities.

[0750] The "suggestion method" refers to a function that suggests appropriate clothing and posing to the user based on the analysis results.

[0751] The "generation means" refers to a function that, based on the proposal, modifies a user's image to generate a new image.

[0752] A "clothing try-on method" is a function that allows users to virtually try on suggested clothing.

[0753] "Storage method" refers to a function that allows users to save modified images they prefer on their device.

[0754] This invention begins with a user taking an image using a user terminal such as a smartphone or tablet. The user terminal is equipped with communication means to compress the acquired image of a person and send it to a server. The server uses analysis means to analyze the received image data and extract facial features from the image.

[0755] Next, the server uses the analysis results to suggest the most suitable clothing and posing to match the impression selected by the user. By using a suggestion method that utilizes a generation AI model, for example, if the user selects "casual and energetic impression," clothing and poses suitable for that impression will be suggested. Based on this suggestion, a new modified image is generated by the generation method and sent to the user's terminal.

[0756] The generated modified images are presented to the user. The user can select the most suitable image from multiple suggestions and save it to their device. The server also provides an environment where the user can virtually try on clothes using a clothing try-on system. This allows users to have an experience on e-commerce sites that is as if they were actually trying on clothes.

[0757] The hardware used includes smartphones and tablets as user terminals, and cloud computing environments as servers. The software includes image processing libraries (e.g., OpenCV) and AI model execution environments (e.g., TensorFlow).

[0758] For example, if a salaryman in his 40s requests a "formal and trustworthy impression," the AI ​​will generate a modified image of him in a suit. This image can be used as a profile picture in business settings.

[0759] An example of a prompt would be, "Suggest formal attire for a man in his 40s and synthesize it to create an impression of trustworthiness."

[0760] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0761] Step 1:

[0762] The user takes a picture of themselves using a user device such as a smartphone or tablet. The input is the desired impression selected by the user and the captured image of the person, and the output is that image data. This image data is saved on the smartphone in the optimal format.

[0763] Step 2:

[0764] The user terminal compresses the captured image of a person and sends it to the server. The input is the uncompressed image data, and the output is the compressed image data. The user terminal uploads this compressed image to the cloud server using a communication method.

[0765] Step 3:

[0766] The server analyzes the received image data. The input is compressed image data sent from the user's terminal, and the output is extracted facial feature data. The server uses an image analysis algorithm (e.g., OpenCV's face detection function) and temporarily stores the features in a database.

[0767] Step 4:

[0768] Based on the analysis results, the server uses an AI model to generate optimal clothing and posing, taking into account the user's selected desired impression. The input is facial feature data and the user's desired impression, and the output is proposed clothing and posing data. TensorFlow is used for the AI ​​model, and the generated data is proposed based on the prompts of the generating AI model.

[0769] Step 5:

[0770] The server uses a generation mechanism to generate modified images based on the proposed data. The input is the proposed clothing and posing data, and the output is the modified image data. Image processing is performed within the server to generate the modified image.

[0771] Step 6:

[0772] The server sends the generated modified image to the user's terminal. The input is the modified image data, and the output is the image data displayed on the user's terminal. This data is then transmitted to the user's smartphone using a communication method.

[0773] Step 7:

[0774] The user selects the most suitable image from several modified images and saves it on their device. The input consists of multiple modified image data sent from the server, and the output is the image data selected and saved by the user. This allows users to use images that convey their desired impression, such as profile pictures.

[0775] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0776] In an embodiment of this invention, the user uses a user terminal such as a smartphone or tablet to take a picture of themselves and then uses an emotion engine to analyze the user's emotions. During the shooting process, the terminal senses the user's facial expressions and voice, and uses this data to determine the user's emotions.

[0777] The device sends the captured image of the person, along with emotional data analyzed by the emotion engine, to the server. The server stores the received image data and emotional data in storage and performs a detailed analysis based on them. Using the analysis method, the server extracts facial feature points and, together with the emotional data, comprehensively evaluates the user's impression.

[0778] Next, the server uses a suggestion mechanism to propose the most suitable clothing and posing based on the user's desired impression and perceived emotions. For example, if the emotion engine recognizes the user's emotion as "joy," the server will suggest bright and cheerful clothing and posing.

[0779] The server uses a generation mechanism to create modified images incorporating the suggested clothing and posing. During this process, subtle adjustments to the color scheme and atmosphere are also made to reflect the emotions. The generated modified images are prepared as multiple variations and sent to the user's terminal.

[0780] The user's device displays the received modified images to the user. The user selects the image that best matches their intended impression or emotion and saves it on the device. The saved images can then be easily used as profile pictures on dating apps, marriage agencies, etc.

[0781] This embodiment allows users to not only select their desired impression but also generate a profile picture that reflects their current emotions, enabling a more natural and appealing form of self-expression.

[0782] The following describes the processing flow.

[0783] Step 1:

[0784] The user takes images of themselves from multiple angles using a user device such as a smartphone or tablet. The emotion engine is activated to detect the user's facial expressions and voice in real time and acquire emotion data.

[0785] Step 2:

[0786] The device saves the captured image of the person and emotion data to temporary memory. The user selects their desired impression from the impression categories provided within the application and enters detailed requests as needed.

[0787] Step 3:

[0788] The device compresses the saved person images and emotion data and sends them to the server. The image data compression is optimized to improve communication efficiency while maintaining image quality.

[0789] Step 4:

[0790] The server stores received person image data and emotion data in storage. Image analysis tools extract facial feature points from the received images and analyze the user's overall impression by combining them with emotion data.

[0791] Step 5:

[0792] Based on the analysis results, the user's selected desired impression, and the perceived emotions, the server uses an AI model to suggest the most suitable clothing and posing. For example, if the server analyzes the emotion as "joy," it will suggest brightly colored clothing and dynamic posing.

[0793] Step 6:

[0794] The server uses a generation mechanism to create modified images that reflect the proposed clothing and posing. Furthermore, it makes real-time adjustments to the color tone and expression according to the emotion. The generated modified images are prepared as multiple variations.

[0795] Step 7:

[0796] The server sends multiple modified images to the user's terminal. The terminal presents these images to the user and provides an interface for the user to make a selection.

[0797] Step 8:

[0798] The user selects the modified image that best reflects their intended impression or emotion from the presented images. The device then offers the option to save the selected modified image to local storage and make it immediately available as a profile picture.

[0799] (Example 2)

[0800] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0801] In modern society, it is difficult for individuals to easily create profile pictures that effectively reflect their real emotions and desired impression when expressing themselves on online platforms. As a result, they may not be able to express themselves more naturally, potentially leading to misunderstandings. Furthermore, traditional methods do not adequately provide emotion-based image adjustments or multiple suggestion options, making it difficult for users to pursue a more ideal self-expression.

[0802] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0803] In this invention, the server includes an analysis means for analyzing features from an image, an evaluation means for evaluating impressions based on the analysis results and emotion data, and a suggestion means for generating suggestions based on the evaluation. This enables users to generate more unique and attractive profile images that reflect their real emotions. Furthermore, by providing multiple modified images and making adjustments based on the desired impression, the system can meet the diverse needs of users for self-expression.

[0804] A "user device" is a digital device used by a user that enables image capture and data manipulation.

[0805] "Means of acquiring image data" refers to functions and technologies for capturing image data such as a user's face.

[0806] "Means for generating emotional data" refers to technologies that analyze captured image data to quantify or categorize the user's emotional state.

[0807] "Communication means" refers to the technologies and protocols used to send and receive data between user devices and processing devices.

[0808] A "processing device" is a computing device used to analyze and process received data, and primarily refers to a server.

[0809] "Analysis methods for analyzing features from images" refers to technologies that identify and extract facial features and shapes from image data.

[0810] "An evaluation method for assessing impressions" refers to a technology that comprehensively estimates a user's impression based on analyzed characteristics and emotional data.

[0811] "A suggestion generation method" refers to a technology that provides users with ideas for the most suitable clothing and posing based on evaluation results.

[0812] "Generation means" refers to the technology for modifying and generating images based on the proposed content.

[0813] "Transmission means" refers to the technology and protocols used to send the generated modified image to the user's device.

[0814] "Storage method" refers to the technology or method of storing images selected by the user in a data storage device.

[0815] To implement this invention, the user first takes an image of themselves using an electronic device such as a smartphone or tablet. At this time, the camera and microphone built into the device acquire the user's facial expressions and voice as data. This data is processed by emotion analysis software installed internally, and the user's emotions are quantified or categorized.

[0816] Next, the device sends the acquired image and emotion data to the server. HTTPS, a standard protocol over the internet, is used for communication. The server utilizes high-performance computing resources to analyze the image's feature points using image processing software. This is done, for example, by facial recognition algorithms. The emotion data is also used to comprehensively evaluate the user's impression.

[0817] Subsequently, the server generates suggestions based on the user's evaluation. These suggestions utilize a generative AI model and include clothing and posing that reflect the user's emotions and impressions. The generated suggestions are then used to create modified images, which are provided to the user as multiple image variations.

[0818] The generated images are sent to the device, allowing the user to select and save the one that best suits their desired impression and emotions. Saved images can also be used in dating apps and lifestyle facilities. This system allows users to easily create unique profile pictures that not only reflect their desired impression but also their current emotions.

[0819] As a concrete example, consider a scenario where a user wants to create a relaxed profile picture during a fun music festival. The user takes a photo with their device, and if the system determines the emotion is "joyful," the server suggests casual and cheerful clothing and posing. An example of a prompt used in this process might be, "Generate a profile picture with a fun atmosphere, casual clothing, and pose."

[0820] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0821] Step 1:

[0822] The user takes a picture of their face using a smartphone or tablet. At this time, the camera and microphone built into the device capture the user's facial expressions and voice, inputting this as image and audio data into the device. The device receives this data as initial input and prepares for emotion analysis.

[0823] Step 2:

[0824] The device processes the acquired image and audio data using internal emotion analysis software. This analysis extracts facial features from images and analyzes tone patterns from audio. As output of the analysis, emotion data is generated, classifying the user's emotions into categories such as "joy" and "sadness."

[0825] Step 3:

[0826] The terminal sends image data and emotion data obtained through analysis to the server. Secure communication via the internet is used for transmission, and the data is input to the server. The server prepares the received data for the next analysis.

[0827] Step 4:

[0828] The server uses image processing software to analyze facial feature points from image data. This analysis phase applies advanced facial recognition algorithms, quantifying image details and outputting them to an internal database. Based on these analysis results and sentiment data, the server evaluates the user's impression.

[0829] Step 5:

[0830] The server uses the results of analysis and evaluation to suggest appropriate clothing and posing for the user. The generative AI model then creates prompt statements based on this process, forming the core of the suggestions. For example, it might generate specific instructions such as "casual clothing and poses in a fun atmosphere." These prompt statements are then output and input into the generation mechanism.

[0831] Step 6:

[0832] The server receives prompt text into the generation mechanism and generates a modified image according to the suggested content. At this stage, multiple image variations with different color schemes and posing are generated and prepared within the server.

[0833] Step 7:

[0834] The server sends the generated modified image to the terminal. The terminal presents the received image to the user and displays multiple options on the screen, allowing the user to select the best image.

[0835] Step 8:

[0836] The user selects the most suitable image from the presented modified images and saves it to their device. This saved image can then be immediately used as a profile picture across various online services.

[0837] (Application Example 2)

[0838] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0839] In modern brick-and-mortar stores, it is difficult to efficiently suggest products based on the subjective feelings and impressions of individual customers, and there is a need for technology that enables product selection that reflects such feelings. Furthermore, there is a need for a way for customers to easily check styles optimized for their feelings without having to try products in-store.

[0840] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0841] In this invention, the server includes an analysis means for analyzing a person's image, a suggestion means for analyzing emotional information and proposing clothing and posing, and a generation means for providing the generated modified image as a virtual try-on. This enables personalized product suggestions that take into account the user's emotional state and a virtual try-on experience.

[0842] A "user terminal" is a portable electronic device used to capture and display images of people.

[0843] A "personal image" is a digital image data that captures the user's face or entire body.

[0844] An "information processing device" is a computing system that performs analysis and makes suggestions based on received image and emotion data.

[0845] "Communication means" refers to network communication functions for sending and receiving data between a user terminal and an information processing device.

[0846] The "analysis method" is a function that performs analysis based on input human images and emotional information to extract features.

[0847] The "suggestion method" refers to a function that selects and suggests the optimal clothing and posing based on the analysis results.

[0848] The "generation means" is a function that creates modified images based on the clothing and posing selected by the proposed means.

[0849] "Storage method" refers to a digital storage medium for long-term storage of modified images selected by the user.

[0850] "Emotion analysis means" refers to technology that detects emotions from the user's facial expressions and voice, and reflects the results in other proposed means.

[0851] "Virtual try-on" is a technology that allows users to visually try on clothing digitally without actually trying on the product.

[0852] This invention is a system that acquires images of a person using a user's terminal and then provides optimal clothing and style suggestions through emotion analysis based on those images. Photographs taken by the user with a smartphone or other terminal are transmitted to an information processing device via communication means. The information processing device extracts facial features from the person's image and detects the user's emotions using emotion analysis means. This process utilizes image processing libraries such as OpenCV and Microsoft's Azure Emotion API.

[0853] The information processing device, based on the analysis results, uses a generative AI model to suggest clothing and styling that best suit the user's emotions and impressions. The suggestion method utilizes the generative AI model and generates a variety of fashion styles using prompt sentences. An example of such a prompt sentence is, "Generate styling suggestions for when the user is expressing cheerful emotions. Please provide ideas that emphasize bright colors and casual clothing."

[0854] The user's device receives the modified image again via communication and displays multiple styling options on the screen. The user can then select their preferred style and save the result on their device using the saving function. This feature allows for an instant virtual try-on experience while in a physical store, supporting purchase decisions. For example, when a user is enjoying a pleasant shopping experience in a store, they may be offered suggestions for bright, casual wear, naturally facilitating a purchase.

[0855] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0856] Step 1:

[0857] The user terminal captures an image of the user with its camera and acquires the image data. The input is the image data of the person captured by the camera, and the output is the generated image data. The terminal immediately prepares to transmit this image to the information processing device via a communication means.

[0858] Step 2:

[0859] The terminal transmits captured image data to the information processing device using a communication method. The input is the acquired image data, and the output is the image data transferred to the information processing device via the network.

[0860] Step 3:

[0861] The information processing device extracts feature points from the user's face using an analysis mechanism based on the received image data. In this process, the input is image data transmitted from the terminal, and the output is facial feature point data. The processing is performed using an image processing library such as OpenCV.

[0862] Step 4:

[0863] The information processing device detects a user's emotions from facial feature point data using emotion analysis means. The input here is facial feature point data, and the output is recognized emotion information. Emotion detection utilizes Microsoft's Azure Emotion API, among others.

[0864] Step 5:

[0865] Based on emotional information, the information processing device uses a proposed means and a generative AI model to consider the optimal clothing and posing. The input is the user's emotional information, and the output is the proposed clothing and posing data. The generative AI model is given instructions via prompts to generate the optimal style.

[0866] Step 6:

[0867] Using the information determined by the proposed method, the generation method of the information processing device generates a modified image. The input is clothing and posing data, as well as the original person image, and the output is a modified virtual try-on image. This is done using a generation AI model.

[0868] Step 7:

[0869] The information processing device transmits the generated modified image to the user terminal. The input is the generated modified image, and the output is the modified image transferred to the terminal.

[0870] Step 8:

[0871] The user terminal presents the received modified image to the user and displays multiple variations. The input is a pre-generated modified image, and the output is multiple virtual try-on images that the user can visually confirm. The image selected by the user is saved and stored via a storage means for future reference or purchase decisions.

[0872] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0873] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0874] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0875] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0876] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0877] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0878] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0879] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0880] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0881] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0882] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0883] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0884] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0885] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0886] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0887] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0888] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0889] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0890] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0891] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0892] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0893] The following is further disclosed regarding the embodiments described above.

[0894] (Claim 1)

[0895] A means for acquiring data of a person's image on a user terminal,

[0896] Methods for selecting desired impressions of a person's image,

[0897] A communication means for sending acquired images of people to a server,

[0898] On the server, there is an analysis means for performing image analysis,

[0899] A proposal means that suggests clothing and posing based on the analysis results,

[0900] A generation means for generating a modified image based on the proposal,

[0901] A means of sending a modified image to the user's terminal,

[0902] A means of saving the modified image selected by the user,

[0903] A system that includes this.

[0904] (Claim 2)

[0905] The system according to claim 1, further comprising means for presenting multiple generated modified images and allowing the user to select one.

[0906] (Claim 3)

[0907] The system according to claim 1, comprising means for making fine adjustments to the modified image according to the impression desired by the user.

[0908] "Example 1"

[0909] (Claim 1)

[0910] Means for acquiring image data in a user device,

[0911] A means of selecting an impression of an image,

[0912] A communication means for transmitting acquired image data to a computing device,

[0913] A computing device includes an information analysis means for performing image analysis,

[0914] A proposal means for suggesting appearance and posture based on analysis results,

[0915] Image generation means for generating modified images based on the proposal,

[0916] A transmission means for sending the modified image to the user device,

[0917] A storage means for saving the modified image selected by the user,

[0918] A system that includes this.

[0919] (Claim 2)

[0920] The system according to claim 1, further comprising means for presenting multiple generated modified images and allowing the user to select one.

[0921] (Claim 3)

[0922] The system according to claim 1, comprising means for making fine adjustments to the modified image according to the impression desired by the user.

[0923] "Application Example 1"

[0924] (Claim 1)

[0925] A means for acquiring data of a person's image on a user terminal,

[0926] Methods for selecting desired impressions of a person's image,

[0927] A communication means for sending acquired images of people to a server,

[0928] On the server, there is an analysis means for performing image analysis,

[0929] A proposal means that suggests clothing and posing based on the analysis results,

[0930] A generation means for generating a modified image based on the proposal,

[0931] A means of sending a modified image to the user's terminal,

[0932] A means of saving the modified image selected by the user,

[0933] A clothing try-on method that suggests clothing based on the impression selected by the user and allows them to virtually try it on,

[0934] A system that includes this.

[0935] (Claim 2)

[0936] The system according to claim 1, comprising means for presenting multiple generated modified images and allowing the user to select one, and means for purchasing a product through the selected clothing.

[0937] (Claim 3)

[0938] The system according to claim 1, comprising means for making fine adjustments to a modified image according to the impression desired by the user, and means for supporting the product selection process based on the clothing data from the user's try-on.

[0939] "Example 2 of combining an emotion engine"

[0940] (Claim 1)

[0941] In the user device, means for acquiring image data,

[0942] A means of generating sentiment data by analyzing acquired data,

[0943] A communication means for transmitting image data and emotion data to a processing unit,

[0944] In the processing device, an analysis means for analyzing features from an image,

[0945] An evaluation method for evaluating impressions based on analysis results and sentiment data,

[0946] A proposal means for generating proposals based on evaluation,

[0947] A generation means for generating a modified image based on the proposed information,

[0948] A transmission means for sending the generated modified image to the user device,

[0949] A means of saving the modified image selected by the user,

[0950] A system that includes this.

[0951] (Claim 2)

[0952] The system according to claim 1, comprising means for presenting multiple generated modified images and allowing the user to select one.

[0953] (Claim 3)

[0954] The system according to claim 1, comprising means for adjusting a modified image based on user sentiment data.

[0955] "Application example 2 when combining with an emotional engine"

[0956] (Claim 1)

[0957] A means for acquiring data of a person's image on a user terminal,

[0958] Methods for selecting desired impressions of a person's image,

[0959] A communication means for transmitting acquired person images to an information processing device,

[0960] An information processing device includes an analysis means for performing image analysis,

[0961] A proposal means for suggesting clothing and posing based on analysis results,

[0962] A generation means for generating a modified image based on the proposal,

[0963] A means of sending a modified image to the user's terminal,

[0964] A means of saving the modified image selected by the user,

[0965] A means of sentiment analysis that detects user sentiment information and reflects it in clothing suggestions,

[0966] A system that includes this.

[0967] (Claim 2)

[0968] The system according to claim 1, comprising means for presenting multiple generated modified images and allowing the user to select one, and implementing a virtual try-on function.

[0969] (Claim 3)

[0970] The system according to claim 1, comprising means for making fine adjustments to a modified image according to the impression desired by the user, and providing clothing suggestions that match the user's emotional state. [Explanation of Symbols]

[0971] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for acquiring data of a person's image on a user terminal, Methods for selecting desired impressions of a person's image, A communication means for sending acquired images of people to a server, On the server, there is an analysis means for performing image analysis, A proposal means that suggests clothing and posing based on the analysis results, A generation means for generating a modified image based on the proposal, A means of sending a modified image to the user's terminal, A means of saving the modified image selected by the user, A system that includes this.

2. The system according to claim 1, further comprising means for presenting multiple generated modified images and allowing the user to select one.

3. The system according to claim 1, comprising means for making fine adjustments to the modified image according to the impression desired by the user.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A