system
The system addresses the lack of personalized skin and mental health care by analyzing user data to provide tailored advice and resources, ensuring security and anonymous interaction, enhancing user well-being.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Existing systems fail to provide individually optimized advice for skin and mental health care, lacking platforms for anonymous user interaction and inadequate data security, leading to suboptimal care tailored to individual user conditions.
A system that analyzes image and text data to evaluate skin and mental states, providing personalized advice through a generation module while ensuring data security and facilitating anonymous community interaction.
Enables highly personalized care by generating user-specific advice and resources, improving mental and skin health through secure, anonymous user interaction and feedback-driven improvement.
Smart Images

Figure 2026074956000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern times, skin health and mental health are closely related, but the technology for comprehensively caring for them is underdeveloped. In particular, in normal skin care and mental care, it is impossible to provide individually optimized advice, and furthermore, there is a lack of a place where users can consult anonymously and with peace of mind. For this reason, there is a problem that it is difficult for users to receive care optimized for their own conditions.
Means for Solving the Problems
[0005] This invention is a system that analyzes image and text data provided by users and provides individually optimized advice and resources. This system uses an image processing module and a natural language processing module to evaluate the user's skin condition and mental state, and a generation module to generate advice suitable for the user. Furthermore, through a community matching module, users have the opportunity to interact anonymously with others in similar situations. In addition, data security is ensured by using a secure protocol, and the quality of advice provided can be continuously improved using a feedback tracking function.
[0006] The "image processing module" is a component that analyzes image data provided by the user and evaluates the condition of the skin.
[0007] A "natural language processing module" is a component that analyzes text data entered by a user and estimates their mental state.
[0008] A "generation module" is a component that has the function of generating and providing user-optimized advice and related resources based on analysis results.
[0009] The "Community Matching Module" is a component that identifies other users with similar statuses based on the user's status and encourages them to join the community.
[0010] A "secure protocol" is a digital communication method that maintains user anonymity and securely sends and receives data.
[0011] The "feedback tracking function" is a feature that records the adoption status of user advice, thereby improving the quality of subsequent suggestions. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0013] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0014] First, the terms used in the following description will be explained.
[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] This invention is a system that individually evaluates the user's skin condition and emotional state, and provides appropriate advice and support. In particular, it centers around an image processing module and a natural language processing module, enabling highly personalized care for the user.
[0034] The server first receives image data from the user and uses an image processing module to quantify and evaluate the skin condition. Skin condition includes factors such as humidity, oil balance, and the presence or absence of inflammation. In parallel, the server uses text data and a natural language processing module to analyze the user's emotional state. This analysis provides an indicator of the user's emotions.
[0035] Subsequently, the server uses a generation module based on the analyzed data to generate advice and resources best suited to the user's condition. This advice includes specific care methods and concrete action suggestions for improving their lifestyle. Related resources such as helpful articles and links to online support groups are also provided.
[0036] The terminal displays these analysis results and advice to the user visually, presenting them in an intuitively understandable format. Users can then use this information to improve their own lives. Furthermore, if a user wishes to participate in a community, the community matching module on the server facilitates connection with other users in similar situations. This allows users to actively interact with others and support each other while maintaining anonymity.
[0037] As a concrete example, suppose User A uses the system because they are experiencing skin problems and emotional anxiety. They upload a photo of their skin via their device and answer questions about their emotional state. Based on this information, the server analyzes and indicates that their skin is dry and their emotional state is unstable. The generated advice includes instructions on how to use moisturizing cream and breathing exercises to calm their mind. A support group is also suggested where User A can talk with others experiencing similar emotional states, allowing them to participate with confidence. As a result, User A receives care for both their skin and their emotional state, aiming for a healthier condition.
[0038] The following describes the processing flow.
[0039] Step 1:
[0040] Users take images of their skin condition with their device and upload them through the app. In addition, they answer questions within the app and input text data about their emotional state.
[0041] Step 2:
[0042] The device uses a secure protocol to send captured images and entered text data to the server. The system is designed to maintain data anonymity during transmission.
[0043] Step 3:
[0044] The server passes the received image data to an image processing module, which analyzes the skin's condition. Here, the moisture level, oil content, and presence or absence of inflammation are quantified and evaluated.
[0045] Step 4:
[0046] The server passes the text data to a natural language processing module, which analyzes the user's emotional state. It analyzes tone and keywords to estimate the user's emotional state.
[0047] Step 5:
[0048] The server integrates the results of image and text analysis and uses a generation module to create user-optimized advice. This advice includes specific care methods and resource links.
[0049] Step 6:
[0050] The server uses a community matching module to identify other users with similar statuses and suggests they join relevant communities.
[0051] Step 7:
[0052] The device displays advice and community information sent from the server to the user. Based on this, the user can select actions that are helpful in their daily life.
[0053] Step 8:
[0054] Users follow the provided advice and provide feedback via their device. The server records this feedback and uses it to further improve subsequent personalization.
[0055] (Example 1)
[0056] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0057] In modern society, the physical and mental health problems faced by individual users are diverse, but the support and advice available are general and not optimized for individual needs. Furthermore, there is a lack of platforms that allow for community interaction and mutual support while maintaining anonymity. Against this backdrop, there is a need for a system that can accurately and intuitively grasp the specific condition of users and provide appropriate advice and resources.
[0058] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0059] In this invention, the server includes image processing means for analyzing visual data collected from the user and evaluating the properties of materials; natural language processing means for analyzing text data input by the user and evaluating the user's mental state; and generation means for generating personalized advice and providing related materials to the user based on the analyzed data. This makes it possible to build a platform that provides advice tailored to the user's condition and facilitates connections with other users who have similar conditions.
[0060] "Visual data" refers to image information collected to show the user's skin condition and other physical characteristics.
[0061] "Text data" refers to text information entered or provided by the user, and the user's mental state is analyzed based on this data.
[0062] "Image processing means" refers to a module for analyzing visual data and quantifying and evaluating the properties of a substance.
[0063] A "natural language processing tool" is a module that processes text data and executes an algorithm to quantify and evaluate the user's mental state.
[0064] The "generation means" is a module that generates user-optimized advice and related materials based on the analysis results of image and text data.
[0065] A "protected communication specification" is a communication protocol used to maintain user anonymity and securely send and receive data.
[0066] The "feedback collection function" is a feedback mechanism that tracks how users' advice is being adopted and helps improve future proposals.
[0067] A "matching tool" is a module that identifies other users with similar statuses based on the user's status, thereby promoting group participation.
[0068] A description of embodiments for carrying out this invention will be given.
[0069] The server plays a central role in evaluating the user's skin and emotional state and providing appropriate advice. When a user provides images of their face and skin through their device, the server analyzes the visual data using advanced image processing software, such as open-source image processing libraries. This process quantifies and evaluates skin characteristics such as humidity, oil balance, and the presence or absence of inflammation.
[0070] In parallel, the user provides text data that expresses their emotional state. This text is analyzed by the server using natural language processing tools. For example, a natural language processing library is used to analyze the text and obtain an index of the user's emotions. These analysis results are then input into a generative AI model, which serves as a means to generate user-optimized advice and relevant resources.
[0071] The terminal visually displays analysis results and advice from the server, presenting them in a way that users can intuitively understand. Based on this information, users can reflect specific actions in their daily lives. For example, they might practice the suggested use of moisturizing cream or breathing exercises to calm their minds. The server also encourages community interaction with other users who have similar concerns, providing a space where users can support each other while maintaining anonymity.
[0072] As a concrete example, consider a scenario where a user provides image and text data regarding dry skin or mental stress. Based on this information, the server generates appropriate care methods and support group resources, which are then presented to the user on their device. This allows the user to receive the care best suited to their condition.
[0073] An example of a prompt for the generative AI model might be: "Analyze the skin image provided by the user and quantify its condition. Next, analyze the text representing the user's emotional state to generate personalized advice." This allows the system to work efficiently and provide accurate, individualized care.
[0074] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0075] Step 1:
[0076] The user takes a photo of their skin via their device and enters text about their emotional state. This visual and text data is sent as input to the server. Specifically, the device's camera function is used to acquire a high-resolution image, and the user enters free-form text about their emotions and state.
[0077] Step 2:
[0078] The server analyzes the received visual data using image processing software. After color correction and noise reduction, the input image data is quantified using a feature extraction algorithm to determine characteristics such as humidity, oil balance, and inflammation. Specific numerical data indicating the skin condition is generated as output and sent to the next processing step.
[0079] Step 3:
[0080] The server analyzes the text data provided by the user using a natural language processing library. The input text is tokenized through morphological analysis and sentiment analysis, and emotional indicators such as positive or negative are quantified. This output is numerical data that represents the user's emotional state and forms the basis for subsequent generation steps.
[0081] Step 4:
[0082] The server integrates numerical data obtained from image and text analysis and inputs it as a prompt into the generating AI model. Specifically, it sends the prompt example "Please suggest the best care method for a user with low skin moisture levels and a negative emotional index" to the model. The output of this process is personalized advice and relevant resource information.
[0083] Step 5:
[0084] The server outputs advice and resource information, which is then sent to the terminal, where it is displayed visually to the user. Specifically, graphs, charts, and text information are displayed on the screen to enable visual and intuitive understanding. This output serves as a suggestion for the user's next action and is provided as information to help improve their lifestyle.
[0085] (Application Example 1)
[0086] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0087] There is a need to accurately assess users' skin and emotional states and, based on that, suggest the most suitable care and products. However, conventional systems make it difficult to provide immediate, personalized advice in physical stores. Furthermore, there is a lack of environment that encourages interaction among users with similar concerns, fostering a sense of security and encouraging continued care.
[0088] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0089] In this invention, the server includes data processing means for analyzing image data collected from the user and evaluating the condition of the skin, natural language processing means for analyzing text data input by the user and evaluating the user's mental state, and store support means for immediately providing evaluation results and product suggestions to the user in a physical store. As a result, the user can immediately understand their own condition and receive suggestions for the most suitable products and care methods on the spot. It can also facilitate interaction with other users in similar conditions and provide support for continuous care.
[0090] A "user" refers to a person who uses the system to evaluate their own skin and mental state and receive recommendations for optimal care and products.
[0091] "Image data" refers to digital image information collected to evaluate the condition of a user's skin.
[0092] "Data processing means" refers to a device that analyzes collected image data and has the function of quantifying and evaluating the condition of the skin.
[0093] "Text data" refers to information such as sentences and words that users input to express their own mental state.
[0094] A "natural language processing system" is a device that processes text data input by a user and has the function of evaluating the user's mental state and emotions.
[0095] A "generation means" is a device that has the function of creating optimal advice and related information for the user based on the analysis results.
[0096] A "group matching method" is a system that identifies other users with similar statuses based on the user's status and has the function of promoting participation in collaboration.
[0097] "Store support means" refers to a function that allows users to receive evaluation results immediately at a physical store and receive recommendations for the most suitable products.
[0098] A "secure protocol" is a means of communication that allows for the secure transmission and reception of data while maintaining user anonymity.
[0099] The "feedback tracking function" is a feature that records the adoption status of advice received by the user and uses that information to improve subsequent suggestions.
[0100] The system that realizes this invention operates with the following configuration.
[0101] First, the user takes a picture of their skin using a dedicated terminal or smart device and uploads the photo to the system. The terminal sends the image data to a server, which analyzes the image using data processing tools. The analyzed data is converted into numerical data to evaluate specific skin conditions such as skin moisture, oil content, and the presence or absence of inflammation. Image processing libraries such as OpenCV are used for this process.
[0102] In parallel, the user answers a simple questionnaire displayed on the device, inputting their emotional state as text data. The device sends this text data to a server, which uses natural language processing to evaluate the emotions. Natural language processing libraries such as TextBlob are used for this part of the analysis, quantifying the degree of positivity or negativity of the emotions.
[0103] Based on the numerical data obtained regarding skin condition and mental state, the server generates optimal advice and suggestions for the user through a generation mechanism. For example, if moisturizing is needed, it will suggest using a moisturizing cream; if the user is feeling depressed, it will recommend relaxation techniques. By utilizing a generation AI model, it is possible to provide suggestions tailored to individual needs. Related information such as links to articles and introductions to appropriate support groups are also provided.
[0104] Furthermore, if a user is in a physical store, they can immediately view and purchase suggested products within the store using in-store support tools. This process makes it easier for users to obtain products that suit their needs in real time.
[0105] As a concrete example, suppose user C uses this system to take a picture of their skin condition in a store and answers questions about their emotional state. The system analyzes skin dryness and mild mental stress and suggests moisturizing cream and relaxation products. These products are suggested in the store and can be purchased on the spot. An example of a prompt to input into the generating AI model is, "Please analyze what emotions are most frequently expressed regarding the user's emotional state."
[0106] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0107] Step 1:
[0108] The user takes a picture of their skin using a device, and image data is acquired. The input is the captured image, which is then sent for subsequent processing. The user presses the capture button on the device, and the image is saved to the system.
[0109] Step 2:
[0110] The terminal sends image data to the server. The input is image data from the terminal, and the output is an image file stored on the server. The terminal transfers the image to the server using wireless communication.
[0111] Step 3:
[0112] The server receives image data and uses data processing tools to analyze the skin condition. The input is image data stored on the server, and the output is numerical data indicating the skin's moisture and oil balance, as well as the presence or absence of inflammation. The server uses OpenCV to analyze the image features and calculate evaluation values.
[0113] Step 4:
[0114] The user answers questions about their mental state displayed on the device and inputs text data. The input is text data entered by the user and sent to the server. The user answers the questions on the device and describes their mental state in sentences.
[0115] Step 5:
[0116] The terminal sends text data to the server. The input is the user's text data, and the output is the text data stored on the server. The terminal then transfers the data back to the server using wireless communication.
[0117] Step 6:
[0118] The server receives text data and evaluates the emotions using natural language processing. The input is text data stored on the server, and the output is numerical data indicating the degree of positivity or negativity of the emotions. The server uses TextBlob to analyze the emotions in the text data and calculate an evaluation value.
[0119] Step 7:
[0120] The server uses a generation mechanism to generate optimal advice based on the analysis results. The input is numerical data indicating the condition of the skin and mind, and the output is advice and product suggestions for the user. The server utilizes a generation AI model to generate individual suggestions.
[0121] Step 8:
[0122] When a user is in a physical store, the terminal uses store support tools to display advice and product suggestions. The input is advice from the server, and the output is information displayed to the user on the terminal. The terminal presents the results to the user through a simple interface.
[0123] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0124] This invention combines an emotional engine with a system that analyzes a user's skin and emotional state to provide optimal advice. The entire system aims to enable highly personalized care for each user and improve their quality of life.
[0125] The server first receives image data from the user and uses an image processing module to evaluate the skin condition. This evaluation includes indicators of humidity, oil content, and skin problems. Next, the server passes the text and voice data entered by the user to a natural language processing module and an emotion engine to recognize the user's mental state and emotions.
[0126] The emotion engine identifies the user's emotions from text and audio, and the server generates emotion data based on this. This emotion data is combined with other analysis results, and a generation module generates personalized emotional support advice. This advice suggests specific care methods and activities to stabilize emotions. Related online resources and community groups are also provided.
[0127] The device displays advice and emotional assessment results sent from the server to the user. Users can use this information to improve their lives and emotional state. Furthermore, they can join communities recommended based on their state, allowing them to interact with others anonymously and share their emotions.
[0128] As a concrete example, suppose user B perceives skin problems associated with daily stress. User B uploads an image on their device and answers questions about their emotional state. The server uses an image processing module to assess the dryness of the skin, an emotion engine to evaluate the stress level, and a generation module to provide suggestions for moisturizing care and meditation methods to reduce stress. It also encourages user B to join a community that helps people cope with stress. As a result, user B receives support for both their physical and mental well-being, enabling them to strive for a healthier state.
[0129] The following describes the processing flow.
[0130] Step 1:
[0131] Users use their devices to take pictures showing their skin condition and upload them through the application. They also answer questions about their emotional state and record voice messages.
[0132] Step 2:
[0133] The device transmits uploaded image data, text data, and audio data to the server via a secure protocol. User anonymity is ensured during this process.
[0134] Step 3:
[0135] The server inputs the received image data into an image processing module to analyze the skin condition. The analysis results include humidity level, oil balance, and the presence or absence of skin problems.
[0136] Step 4:
[0137] The server inputs text and audio data into a natural language processing module and an emotion engine to analyze the user's mental state and emotions. This process generates an emotion index for the user.
[0138] Step 5:
[0139] The server combines the image processing and emotion engine analysis results and passes them to the generation module. The generation module generates emotional support advice that is best suited to the user. This advice includes specific self-care methods and suggestions for emotional stabilization.
[0140] Step 6:
[0141] The server utilizes a community matching module to identify other users with similar statuses. This allows it to generate recommendations for the most suitable online community groups for each user.
[0142] Step 7:
[0143] The terminal displays analysis results and advice sent from the server, as well as suggestions for community participation. Users refer to the provided information and choose their actions based on it.
[0144] Step 8:
[0145] Users incorporate the advice provided into their lives and experience its effects. Feedback is transmitted to the server via the device, and the server uses this information to improve the quality of future suggestions.
[0146] (Example 2)
[0147] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0148] Modern cosmetic and mental health support systems offer generic recommendations to users, lacking optimized support that fully considers the specific circumstances and emotional factors of individual users. Furthermore, while there is a need to provide personalized communication and advice while ensuring user anonymity and safety, efficient and effective methods are lacking. Additionally, the inadequate mechanisms for tracking how users adopt suggested advice and the results limit opportunities for improving service quality.
[0149] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0150] In this invention, the server includes image processing means for analyzing image information collected from the user and evaluating its surface state; natural information processing means for analyzing text and voice information input by the user and evaluating its mental state; and generation means for generating support optimized for the user and providing related resources based on the analyzed information. This enables personalized and optimal support based on the user's state, secure and anonymous communication, and improvement of service quality based on feedback.
[0151] "Image information" refers to visual data such as photographs and illustrations that show the surface condition of the user.
[0152] "Image processing means" refers to technical methods and processes for analyzing the above image information and evaluating the condition of skin and other surface areas.
[0153] "Textual information" refers to text-based data entered by the user, which is used to evaluate the user's mental state.
[0154] "Audio information" refers to sound data that users input via voice, which allows for the evaluation of the user's emotions and mental state.
[0155] "Natural information processing means" refers to technical means and systems for analyzing textual and auditory information and evaluating the user's mental state.
[0156] "Generating means" refers to methods and means for providing optimized support and related resources based on the analysis results of the user's surface state and mental state.
[0157] "Methods for promoting participation" refer to technologies and methods that identify other users with similar characteristics based on the user's status and promote communication with others and participation in a community.
[0158] "Protective communication means" refers to communication technologies and protocols that maintain user anonymity and ensure security when sending and receiving data.
[0159] "Methods for gathering feedback" refer to functions and methods for tracking user feedback, understanding the adoption of support, and improving the service.
[0160] This invention is a system that analyzes the user's surface and mental state and provides optimized support. The main components of the system are a server, a terminal, and input data from the user.
[0161] The server performs various data processing and calculations using the following hardware and software. First, image information provided by the user is processed using the open-source library OpenCV. This identifies indicators of the user's skin's moisture, oil content, and problems, and evaluates them as numerical data.
[0162] Furthermore, the text and audio information obtained from users is analyzed using a natural language processing engine. A method is employed that utilizes Google® Cloud NLP to evaluate mental state from the input free-form text and audio data. Audio information is converted into text using speech recognition technology and further analyzed as text.
[0163] The generation method utilizes a generative AI model. Specifically, it uses Azure® generative AI via a cloud-based AI service to generate optimal support and advice based on the user's condition. The generated content includes not only skincare but also suggestions for maintaining mental health and related online resources.
[0164] The terminal displays information retrieved from the server in a format that is easy for the user to understand. The terminal's interface is designed to be easy for the user to operate and to quickly obtain the necessary advice.
[0165] For example, if a user is troubled by daily stress and dry skin, this system analyzes the user's image and evaluates their skin condition. It also measures the user's mental stress level through voice input. Based on this data, the server provides the user with specific skincare methods and meditation techniques for stress relief.
[0166] The following are examples of prompts to input into a generative AI model:
[0167] "The user's skin condition is dry, and their mental state is rated as high stress. Please generate optimal skincare advice and stress management activity suggestions for this user."
[0168] These processes enable users to receive effective support tailored to their individual needs.
[0169] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0170] Step 1:
[0171] User data entry
[0172] Users take photos of their skin using their device and send the image data to a server through the application. Users also answer questions about their emotions and mental state, inputting this information as voice or text data.
[0173] Input: Skin image data, text data, or audio data
[0174] The specific actions involve taking an image using the device's camera function and then pressing the image upload button to send the data. Regarding the state of mind, the user either answers questions verbally using the device's voice recognition function or enters text using the keyboard.
[0175] Step 2:
[0176] Server-based analysis of skin condition
[0177] The server receives image data sent by the user and analyzes it using an image processing module. Specifically, it uses the OpenCV library to calculate skin moisture, oil content, and trouble indicators. As a result, numerical data is generated.
[0178] Input: Skin image data submitted by the user
[0179] Output: Skin evaluation results (numerical data for humidity, oil content, and trouble indicators)
[0180] Specifically, a Python script runs on the server to calculate these values and form the evaluation results.
[0181] Step 3:
[0182] Analysis of mental state by server
[0183] The server passes text and audio data provided by the user to a natural language processing module. Text analysis is performed using Google Cloud NLP services, and emotions are identified using TEDAS. Audio data is converted to text using speech recognition technology, and this text data is analyzed in the same manner.
[0184] Input: Text or audio data sent by the user.
[0185] Output: Mental state evaluation results (emotion tags)
[0186] Specifically, the server converts the audio data into text, performs text analysis, classifies the sentiment, and tags the data.
[0187] Step 4:
[0188] Advice generation using generative AI models
[0189] The server uses collected data on skin condition and mental state to create prompts for a generative AI model and generate optimal advice. Using Azure's generative AI, it provides specific skincare methods and suggestions for improving mental well-being.
[0190] Input: Skin evaluation results, mental state evaluation results
[0191] Output: Optimized advice for the user
[0192] In terms of operation, the AI model on the server takes the prompt text as input, generates optimized advice, and returns it.
[0193] Step 5:
[0194] Advice and results displayed on the device
[0195] The device receives optimized advice and analysis results from the server and displays them to the user. This includes care instructions, links to resources, and community invitations.
[0196] Input: Generation advice and analysis results from the server.
[0197] Output: Display and notification to the user
[0198] Specifically, the device uses push notifications to inform the user of the arrival of new advice and displays detailed information within the app.
[0199] (Application Example 2)
[0200] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0201] Modern consumers face various skin and mental health issues due to daily stress and environmental changes. However, conventional skincare products and mental health support often only offer generalized advice. Therefore, it is difficult to provide highly personalized care for each individual. Furthermore, customers who use physical stores are not adequately provided with guidance on relaxation appropriate to the store environment or mental support through interaction with the community. Addressing these challenges is necessary to improve the customer experience and quality of life.
[0202] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0203] In this invention, the server includes image processing means for analyzing image data collected from the user and evaluating the condition of the skin; natural language processing means for analyzing voice data and text data input from the user and evaluating emotions and mental state; and generation means for generating advice, including care methods and emotional stabilization activities optimized for the user, based on the analyzed data, and for guiding relaxation activities applicable in the store environment. This enables personalized care and emotional support for individual users, improving the customer experience when using the store and enhancing the quality of life.
[0204] "Image processing means" refers to a processing device that analyzes image data provided by the user and evaluates the condition of the skin.
[0205] "Natural language processing means" refers to a processing device that analyzes voice data and text data input by a user and evaluates emotions and mental states.
[0206] The "generation means" is a processing device that generates care methods optimized for the user and advice including emotional stabilization activities based on analyzed data, and guides users through relaxation activities applicable to the store environment.
[0207] A "community matching means" is a processing device that identifies other users in similar states based on the user's state and facilitates anonymous interaction.
[0208] This invention is a system that analyzes the user's skin and emotional state and provides optimal care and emotional support. The system mainly consists of three components: a server, a terminal, and the user.
[0209] The server analyzes image data received from the user using image processing tools and evaluates the skin condition using software such as TENSORFLOW® and OpenCV. Furthermore, it performs natural language processing on voice and text data input from the user using the Google Cloud Natural Language API and evaluates emotions and mental state using the Emotion API. Based on the data obtained in this way, a generation tool generates advice, including care methods and emotional stabilization activities optimized for the user, and sends the results to the terminal.
[0210] The device displays received advice to the user, presenting it visually and clearly using a smartphone or smart glasses. This allows users to implement specific care methods tailored to their own condition. For example, customers can receive detailed information about dedicated relaxation spaces and activities offered within the physical store.
[0211] Through this system, users can gain a comprehensive understanding of their skin and mental state, and incorporate personalized advice into their daily lives. Furthermore, the system also has a function that facilitates anonymous communication between users with similar conditions, allowing them to receive mental health support.
[0212] A concrete example is a case where, based on a user's facial image taken in a store and input about their mental state, the server evaluates their skin's moisture level and stress level, and then provides a list of recommended skincare products and a meditation program for stress reduction. An example of a prompt used in this case would be, "Please analyze my skin condition and mental stress level today and tell me the best skincare and relaxation methods for me."
[0213] In this way, through the invention, users can receive integrated care tailored to their individual needs, and as a result, their quality of life can be improved.
[0214] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0215] Step 1:
[0216] The user takes an image of their face using a device and inputs audio or text data. This input data is sent to the server. The inputs include image data, audio data, and text data, which provide the information that forms the basis for analysis in the next step.
[0217] Step 2:
[0218] The server analyzes the received image data using image processing tools. Specifically, it uses OpenCV to detect the location and condition of skin problems at the pixel level, and then uses a model trained with TensorFlow to calculate and output indicators of skin moisture, oil content, and problems. This output is then passed on to the next step as information to evaluate the user's skin condition.
[0219] Step 3:
[0220] The server passes the voice or text data input from the user to a natural language processing system. Here, the voice data is converted to text using the Google Cloud Natural Language API, and emotion analysis is performed using the Emotion API. The user's tone of voice and text content are analyzed as input, and emotion data and indicators showing the user's emotional state are generated as output. This provides quantitative data on the user's emotional state.
[0221] Step 4:
[0222] Based on analyzed skin condition and emotional data, the server uses a generation mechanism to generate care methods and advice tailored to the user. Specifically, an AI model evaluates this data and lists appropriate skincare products and relaxation programs. It also includes guidance on relaxation activities applicable to the store environment. This output provides specific care advice that the user can implement.
[0223] Step 5:
[0224] The generated advice is sent to the device and displayed to the user. The device uses Flutter® or another native app development environment as its frontend to present the advice in an easy-to-understand visual format. Users can use this as a reference to engage in their daily care. Specifically, features such as recommended product information and the ability to book in-store activities are provided.
[0225] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0226] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0227] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0228] [Second Embodiment]
[0229] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0230] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0231] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0232] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0233] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0234] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0235] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0236] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0237] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0238] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0239] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0240] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0241] This invention is a system that individually evaluates the user's skin condition and emotional state, and provides appropriate advice and support. In particular, it centers around an image processing module and a natural language processing module, enabling highly personalized care for the user.
[0242] The server first receives image data from the user and uses an image processing module to quantify and evaluate the skin condition. Skin condition includes factors such as humidity, oil balance, and the presence or absence of inflammation. In parallel, the server uses text data and a natural language processing module to analyze the user's emotional state. This analysis provides an indicator of the user's emotions.
[0243] Subsequently, the server uses a generation module based on the analyzed data to generate advice and resources best suited to the user's condition. This advice includes specific care methods and concrete action suggestions for improving their lifestyle. Related resources such as helpful articles and links to online support groups are also provided.
[0244] The terminal displays these analysis results and advice to the user visually, presenting them in an intuitively understandable format. Users can then use this information to improve their own lives. Furthermore, if a user wishes to participate in a community, the community matching module on the server facilitates connection with other users in similar situations. This allows users to actively interact with others and support each other while maintaining anonymity.
[0245] As a concrete example, suppose User A uses the system because they are experiencing skin problems and emotional anxiety. They upload a photo of their skin via their device and answer questions about their emotional state. Based on this information, the server analyzes and indicates that their skin is dry and their emotional state is unstable. The generated advice includes instructions on how to use moisturizing cream and breathing exercises to calm their mind. A support group is also suggested where User A can talk with others experiencing similar emotional states, allowing them to participate with confidence. As a result, User A receives care for both their skin and their emotional state, aiming for a healthier condition.
[0246] The following describes the processing flow.
[0247] Step 1:
[0248] Users take images of their skin condition with their device and upload them through the app. In addition, they answer questions within the app and input text data about their emotional state.
[0249] Step 2:
[0250] The device uses a secure protocol to send captured images and entered text data to the server. The system is designed to maintain data anonymity during transmission.
[0251] Step 3:
[0252] The server passes the received image data to an image processing module, which analyzes the skin's condition. Here, the moisture level, oil content, and presence or absence of inflammation are quantified and evaluated.
[0253] Step 4:
[0254] The server passes the text data to a natural language processing module, which analyzes the user's emotional state. It analyzes tone and keywords to estimate the user's emotional state.
[0255] Step 5:
[0256] The server integrates the results of image and text analysis and uses a generation module to create user-optimized advice. This advice includes specific care methods and resource links.
[0257] Step 6:
[0258] The server uses a community matching module to identify other users with similar statuses and suggests they join relevant communities.
[0259] Step 7:
[0260] The device displays advice and community information sent from the server to the user. Based on this, the user can select actions that are helpful in their daily life.
[0261] Step 8:
[0262] Users follow the provided advice and provide feedback via their device. The server records this feedback and uses it to further improve subsequent personalization.
[0263] (Example 1)
[0264] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0265] In modern society, the physical and mental health problems faced by individual users are diverse, but the support and advice available are general and not optimized for individual needs. Furthermore, there is a lack of platforms that allow for community interaction and mutual support while maintaining anonymity. Against this backdrop, there is a need for a system that can accurately and intuitively grasp the specific condition of users and provide appropriate advice and resources.
[0266] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0267] In this invention, the server includes image processing means for analyzing visual data collected from the user and evaluating the properties of materials; natural language processing means for analyzing text data input by the user and evaluating the user's mental state; and generation means for generating personalized advice and providing related materials to the user based on the analyzed data. This makes it possible to build a platform that provides advice tailored to the user's condition and facilitates connections with other users who have similar conditions.
[0268] "Visual data" refers to image information collected to show the user's skin condition and other physical characteristics.
[0269] "Text data" refers to text information entered or provided by the user, and the user's mental state is analyzed based on this data.
[0270] "Image processing means" refers to a module for analyzing visual data and quantifying and evaluating the properties of a substance.
[0271] A "natural language processing tool" is a module that processes text data and executes an algorithm to quantify and evaluate the user's mental state.
[0272] The "generation means" is a module that generates user-optimized advice and related materials based on the analysis results of image and text data.
[0273] A "protected communication specification" is a communication protocol used to maintain user anonymity and securely send and receive data.
[0274] The "feedback collection function" is a feedback mechanism that tracks how users' advice is being adopted and helps improve future proposals.
[0275] A "matching tool" is a module that identifies other users with similar statuses based on the user's status, thereby promoting group participation.
[0276] A description of embodiments for carrying out this invention will be given.
[0277] The server plays a central role in evaluating the user's skin and emotional state and providing appropriate advice. When a user provides images of their face and skin through their device, the server analyzes the visual data using advanced image processing software, such as open-source image processing libraries. This process quantifies and evaluates skin characteristics such as humidity, oil balance, and the presence or absence of inflammation.
[0278] In parallel, the user provides text data that expresses their emotional state. This text is analyzed by the server using natural language processing tools. For example, a natural language processing library is used to analyze the text and obtain an index of the user's emotions. These analysis results are then input into a generative AI model, which serves as a means to generate user-optimized advice and relevant resources.
[0279] The terminal visually displays analysis results and advice from the server, presenting them in a way that users can intuitively understand. Based on this information, users can reflect specific actions in their daily lives. For example, they might practice the suggested use of moisturizing cream or breathing exercises to calm their minds. The server also encourages community interaction with other users who have similar concerns, providing a space where users can support each other while maintaining anonymity.
[0280] As a concrete example, consider a scenario where a user provides image and text data regarding dry skin or mental stress. Based on this information, the server generates appropriate care methods and support group resources, which are then presented to the user on their device. This allows the user to receive the care best suited to their condition.
[0281] An example of a prompt for the generative AI model might be: "Analyze the skin image provided by the user and quantify its condition. Next, analyze the text representing the user's emotional state to generate personalized advice." This allows the system to work efficiently and provide accurate, individualized care.
[0282] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0283] Step 1:
[0284] The user takes a photo of their own skin through the terminal and enters text related to their mental state. These visual data and character data are sent as inputs to the server. As specific operations, the camera function of the terminal is used to obtain a high-resolution image, and a free description regarding the user's emotions and state is entered.
[0285] Step 2:
[0286] The server analyzes the received visual data using image processing software. After performing color tone correction and noise removal on the input image data, characteristics such as humidity, oil balance, and inflammation are quantified by a feature extraction algorithm. As an output, specific numerical data indicating the skin condition is generated and sent to the next processing step.
[0287] Step 3:
[0288] The server analyzes the character data provided by the user using a natural language processing library. The input text is tokenized through morphological analysis and sentiment analysis, and sentiment indicators such as positive or negative are quantified. This output is numerical data indicating the user's mental state and serves as the basis for subsequent generation steps.
[0289] Step 4:
[0290] The server integrates the numerical data obtained from image analysis and text analysis and inputs it as a prompt into the generative AI model. As a specific operation, a prompt example sentence "Please propose the optimal care method for users with low skin moisture value and negative sentiment indicator" is sent to the model. As this output, advice optimized for the user and related resource information are generated.
[0291] Step 5:
[0292] The server outputs advice and resource information, which is then sent to the terminal, where it is displayed visually to the user. Specifically, graphs, charts, and text information are displayed on the screen to enable visual and intuitive understanding. This output serves as a suggestion for the user's next action and is provided as information to help improve their lifestyle.
[0293] (Application Example 1)
[0294] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0295] There is a need to accurately assess users' skin and emotional states and, based on that, suggest the most suitable care and products. However, conventional systems make it difficult to provide immediate, personalized advice in physical stores. Furthermore, there is a lack of environment that encourages interaction among users with similar concerns, fostering a sense of security and encouraging continued care.
[0296] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0297] In this invention, the server includes data processing means for analyzing image data collected from the user and evaluating the condition of the skin, natural language processing means for analyzing text data input by the user and evaluating the user's mental state, and store support means for immediately providing evaluation results and product suggestions to the user in a physical store. As a result, the user can immediately understand their own condition and receive suggestions for the most suitable products and care methods on the spot. It can also facilitate interaction with other users in similar conditions and provide support for continuous care.
[0298] A "user" refers to a person who uses the system to evaluate their own skin and mental state and receive recommendations for optimal care and products.
[0299] "Image data" refers to digital image information collected for evaluating the user's skin condition.
[0300] "Data processing means" has the function of analyzing the collected image data, quantifying the skin condition, and evaluating it.
[0301] "Text data" refers to information of sentences and words input by the user to represent their mental state.
[0302] "Natural language analysis means" has the function of processing the text data input by the user and evaluating the mental state and emotions.
[0303] "Generation means" has the function of creating optimal advice and relevant information for the user based on the analysis results.
[0304] "Group matching means" has the function of identifying other users with similar states based on the user's state and promoting participation in collaborators.
[0305] "Store support means" has the function of enabling the user to receive the evaluation results immediately and receive proposals for optimal products in the physical store.
[0306] "Security protocol" is a communication means for safely transmitting and receiving data while maintaining the anonymity of the user.
[0307] "Feedback tracking function" has the function of recording the adoption status of the advice received by the user and improving subsequent proposals based on that information.
[0308] The system that realizes this invention operates with the following configuration.
[0309] First, the user takes a picture of their skin using a dedicated terminal or smart device and uploads the photo to the system. The terminal sends the image data to a server, which analyzes the image using data processing tools. The analyzed data is converted into numerical data to evaluate specific skin conditions such as skin moisture, oil content, and the presence or absence of inflammation. Image processing libraries such as OpenCV are used for this process.
[0310] In parallel, the user answers a simple questionnaire displayed on the device, inputting their emotional state as text data. The device sends this text data to a server, which uses natural language processing to evaluate the emotions. Natural language processing libraries such as TextBlob are used for this part of the analysis, quantifying the degree of positivity or negativity of the emotions.
[0311] Based on the numerical data obtained regarding skin condition and mental state, the server generates optimal advice and suggestions for the user through a generation mechanism. For example, if moisturizing is needed, it will suggest using a moisturizing cream; if the user is feeling depressed, it will recommend relaxation techniques. By utilizing a generation AI model, it is possible to provide suggestions tailored to individual needs. Related information such as links to articles and introductions to appropriate support groups are also provided.
[0312] Furthermore, if a user is in a physical store, they can immediately view and purchase suggested products within the store using in-store support tools. This process makes it easier for users to obtain products that suit their needs in real time.
[0313] As a concrete example, suppose user C uses this system to take a picture of their skin condition in a store and answers questions about their emotional state. The system analyzes skin dryness and mild mental stress and suggests moisturizing cream and relaxation products. These products are suggested in the store and can be purchased on the spot. An example of a prompt to input into the generating AI model is, "Please analyze what emotions are most frequently expressed regarding the user's emotional state."
[0314] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0315] Step 1:
[0316] The user takes a picture of their skin using a device, and image data is acquired. The input is the captured image, which is then sent for subsequent processing. The user presses the capture button on the device, and the image is saved to the system.
[0317] Step 2:
[0318] The terminal sends image data to the server. The input is image data from the terminal, and the output is an image file stored on the server. The terminal transfers the image to the server using wireless communication.
[0319] Step 3:
[0320] The server receives image data and uses data processing tools to analyze the skin condition. The input is image data stored on the server, and the output is numerical data indicating the skin's moisture and oil balance, as well as the presence or absence of inflammation. The server uses OpenCV to analyze the image features and calculate evaluation values.
[0321] Step 4:
[0322] The user answers questions about their mental state displayed on the device and inputs text data. The input is text data entered by the user and sent to the server. The user answers the questions on the device and describes their mental state in sentences.
[0323] Step 5:
[0324] The terminal sends text data to the server. The input is the user's text data, and the output is the text data stored on the server. The terminal then transfers the data back to the server using wireless communication.
[0325] Step 6:
[0326] The server receives text data and evaluates the emotions using natural language processing. The input is text data stored on the server, and the output is numerical data indicating the degree of positivity or negativity of the emotions. The server uses TextBlob to analyze the emotions in the text data and calculate an evaluation value.
[0327] Step 7:
[0328] The server uses a generation mechanism to generate optimal advice based on the analysis results. The input is numerical data indicating the condition of the skin and mind, and the output is advice and product suggestions for the user. The server utilizes a generation AI model to generate individual suggestions.
[0329] Step 8:
[0330] When a user is in a physical store, the terminal uses store support tools to display advice and product suggestions. The input is advice from the server, and the output is information displayed to the user on the terminal. The terminal presents the results to the user through a simple interface.
[0331] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0332] This invention combines an emotional engine with a system that analyzes a user's skin and emotional state to provide optimal advice. The entire system aims to enable highly personalized care for each user and improve their quality of life.
[0333] The server first receives image data from the user and uses an image processing module to evaluate the skin condition. This evaluation includes indicators of humidity, oil content, and skin problems. Next, the server passes the text and voice data entered by the user to a natural language processing module and an emotion engine to recognize the user's mental state and emotions.
[0334] The emotion engine identifies the user's emotions from text and audio, and the server generates emotion data based on this. This emotion data is combined with other analysis results, and a generation module generates personalized emotional support advice. This advice suggests specific care methods and activities to stabilize emotions. Related online resources and community groups are also provided.
[0335] The device displays advice and emotional assessment results sent from the server to the user. Users can use this information to improve their lives and emotional state. Furthermore, they can join communities recommended based on their state, allowing them to interact with others anonymously and share their emotions.
[0336] As a concrete example, suppose user B perceives skin problems associated with daily stress. User B uploads an image on their device and answers questions about their emotional state. The server uses an image processing module to assess the dryness of the skin, an emotion engine to evaluate the stress level, and a generation module to provide suggestions for moisturizing care and meditation methods to reduce stress. It also encourages user B to join a community that helps people cope with stress. As a result, user B receives support for both their physical and mental well-being, enabling them to strive for a healthier state.
[0337] The following describes the processing flow.
[0338] Step 1:
[0339] Users use their devices to take pictures showing their skin condition and upload them through the application. They also answer questions about their emotional state and record voice messages.
[0340] Step 2:
[0341] The device transmits uploaded image data, text data, and audio data to the server via a secure protocol. User anonymity is ensured during this process.
[0342] Step 3:
[0343] The server inputs the received image data into an image processing module to analyze the skin condition. The analysis results include humidity level, oil balance, and the presence or absence of skin problems.
[0344] Step 4:
[0345] The server inputs text and audio data into a natural language processing module and an emotion engine to analyze the user's mental state and emotions. This process generates an emotion index for the user.
[0346] Step 5:
[0347] The server combines the image processing and emotion engine analysis results and passes them to the generation module. The generation module generates emotional support advice that is best suited to the user. This advice includes specific self-care methods and suggestions for emotional stabilization.
[0348] Step 6:
[0349] The server utilizes a community matching module to identify other users with similar statuses. This allows it to generate recommendations for the most suitable online community groups for each user.
[0350] Step 7:
[0351] The terminal displays analysis results and advice sent from the server, as well as suggestions for community participation. Users refer to the provided information and choose their actions based on it.
[0352] Step 8:
[0353] Users incorporate the advice provided into their lives and experience its effects. Feedback is transmitted to the server via the device, and the server uses this information to improve the quality of future suggestions.
[0354] (Example 2)
[0355] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0356] Modern cosmetic and mental health support systems offer generic recommendations to users, lacking optimized support that fully considers the specific circumstances and emotional factors of individual users. Furthermore, while there is a need to provide personalized communication and advice while ensuring user anonymity and safety, efficient and effective methods are lacking. Additionally, the inadequate mechanisms for tracking how users adopt suggested advice and the results limit opportunities for improving service quality.
[0357] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0358] In this invention, the server includes image processing means for analyzing image information collected from the user and evaluating its surface state; natural information processing means for analyzing text and voice information input by the user and evaluating its mental state; and generation means for generating support optimized for the user and providing related resources based on the analyzed information. This enables personalized and optimal support based on the user's state, secure and anonymous communication, and improvement of service quality based on feedback.
[0359] "Image information" refers to visual data such as photographs and illustrations that show the surface condition of the user.
[0360] "Image processing means" refers to technical methods and processes for analyzing the above image information and evaluating the condition of skin and other surface areas.
[0361] "Textual information" refers to text-based data entered by the user, which is used to evaluate the user's mental state.
[0362] "Audio information" refers to sound data that users input via voice, which allows for the evaluation of the user's emotions and mental state.
[0363] "Natural information processing means" refers to technical means and systems for analyzing textual and auditory information and evaluating the user's mental state.
[0364] "Generating means" refers to methods and means for providing optimized support and related resources based on the analysis results of the user's surface state and mental state.
[0365] "Methods for promoting participation" refer to technologies and methods that identify other users with similar characteristics based on the user's status and promote communication with others and participation in a community.
[0366] "Protective communication means" refers to communication technologies and protocols that maintain user anonymity and ensure security when sending and receiving data.
[0367] "Methods for gathering feedback" refer to functions and methods for tracking user feedback, understanding the adoption of support, and improving the service.
[0368] This invention is a system that analyzes the user's surface and mental state and provides optimized support. The main components of the system are a server, a terminal, and input data from the user.
[0369] The server performs various data processing and calculations using the following hardware and software. First, image information provided by the user is processed using the open-source library OpenCV. This identifies indicators of the user's skin's moisture, oil content, and problems, and evaluates them as numerical data.
[0370] Furthermore, the text and audio information obtained from users is analyzed using a natural language processing engine. A method is employed that utilizes Google Cloud NLP to evaluate mental state from the input free-form text and audio data. Audio information is converted into text using speech recognition technology and further analyzed as text.
[0371] The generation method utilizes a generative AI model. Specifically, it uses Azure's generative AI via a cloud-based AI service to generate optimal support and advice based on the user's condition. The generated content includes not only skincare but also suggestions for maintaining mental health and related online resources.
[0372] The terminal displays information retrieved from the server in a format that is easy for the user to understand. The terminal's interface is designed to be easy for the user to operate and to quickly obtain the necessary advice.
[0373] For example, if a user is troubled by daily stress and dry skin, this system analyzes the user's image and evaluates their skin condition. It also measures the user's mental stress level through voice input. Based on this data, the server provides the user with specific skincare methods and meditation techniques for stress relief.
[0374] The following are examples of prompts to input into a generative AI model:
[0375] "The user's skin condition is dry, and their mental state is rated as high stress. Please generate optimal skincare advice and stress management activity suggestions for this user."
[0376] These processes enable users to receive effective support tailored to their individual needs.
[0377] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0378] Step 1:
[0379] User data entry
[0380] Users take photos of their skin using their device and send the image data to a server through the application. Users also answer questions about their emotions and mental state, inputting this information as voice or text data.
[0381] Input: Skin image data, text data, or audio data
[0382] The specific actions involve taking an image using the device's camera function and then pressing the image upload button to send the data. Regarding the state of mind, the user either answers questions verbally using the device's voice recognition function or enters text using the keyboard.
[0383] Step 2:
[0384] Server-based analysis of skin condition
[0385] The server receives image data sent by the user and analyzes it using an image processing module. Specifically, it uses the OpenCV library to calculate skin moisture, oil content, and trouble indicators. As a result, numerical data is generated.
[0386] Input: Skin image data submitted by the user
[0387] Output: Skin evaluation results (numerical data for humidity, oil content, and trouble indicators)
[0388] Specifically, a Python script runs on the server to calculate these values and form the evaluation results.
[0389] Step 3:
[0390] Analysis of mental state by server
[0391] The server passes text and audio data provided by the user to a natural language processing module. Text analysis is performed using Google Cloud NLP services, and emotions are identified using TEDAS. Audio data is converted to text using speech recognition technology, and this text data is analyzed in the same manner.
[0392] Input: Text or audio data sent by the user.
[0393] Output: Mental state evaluation results (emotion tags)
[0394] Specifically, the server converts the audio data into text, performs text analysis, classifies the sentiment, and tags the data.
[0395] Step 4:
[0396] Advice generation using generative AI models
[0397] The server uses collected data on skin condition and mental state to create prompts for a generative AI model and generate optimal advice. Using Azure's generative AI, it provides specific skincare methods and suggestions for improving mental well-being.
[0398] Input: Skin evaluation results, mental state evaluation results
[0399] Output: Optimized advice for the user
[0400] In terms of operation, the AI model on the server takes the prompt text as input, generates optimized advice, and returns it.
[0401] Step 5:
[0402] Advice and results displayed on the device
[0403] The device receives optimized advice and analysis results from the server and displays them to the user. This includes care instructions, links to resources, and community invitations.
[0404] Input: Generation advice and analysis results from the server.
[0405] Output: Display and notification to the user
[0406] Specifically, the device uses push notifications to inform the user of the arrival of new advice and displays detailed information within the app.
[0407] (Application Example 2)
[0408] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0409] Modern consumers face various skin and mental health issues due to daily stress and environmental changes. However, conventional skincare products and mental health support often only offer generalized advice. Therefore, it is difficult to provide highly personalized care for each individual. Furthermore, customers who use physical stores are not adequately provided with guidance on relaxation appropriate to the store environment or mental support through interaction with the community. Addressing these challenges is necessary to improve the customer experience and quality of life.
[0410] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0411] In this invention, the server includes image processing means for analyzing image data collected from the user and evaluating the condition of the skin; natural language processing means for analyzing voice data and text data input from the user and evaluating emotions and mental state; and generation means for generating advice, including care methods and emotional stabilization activities optimized for the user, based on the analyzed data, and for guiding relaxation activities applicable in the store environment. This enables personalized care and emotional support for individual users, improving the customer experience when using the store and enhancing the quality of life.
[0412] "Image processing means" refers to a processing device that analyzes image data provided by the user and evaluates the condition of the skin.
[0413] "Natural language processing means" refers to a processing device that analyzes voice data and text data input by a user and evaluates emotions and mental states.
[0414] The "generation means" is a processing device that generates care methods optimized for the user and advice including emotional stabilization activities based on analyzed data, and guides users through relaxation activities applicable to the store environment.
[0415] A "community matching means" is a processing device that identifies other users in similar states based on the user's state and facilitates anonymous interaction.
[0416] This invention is a system that analyzes the user's skin and emotional state and provides optimal care and emotional support. The system mainly consists of three components: a server, a terminal, and the user.
[0417] The server analyzes image data received from the user using image processing tools and evaluates the skin condition using software such as TensorFlow and OpenCV. Furthermore, it performs natural language processing on voice and text data input from the user using the Google Cloud Natural Language API and evaluates emotions and mental state using the Emotion API. Based on the data obtained in this way, a generation tool generates advice, including care methods and emotional stabilization activities optimized for the user, and sends the results to the terminal.
[0418] The device displays received advice to the user, presenting it visually and clearly using a smartphone or smart glasses. This allows users to implement specific care methods tailored to their own condition. For example, customers can receive detailed information about dedicated relaxation spaces and activities offered within the physical store.
[0419] Through this system, users can gain a comprehensive understanding of their skin and mental state, and incorporate personalized advice into their daily lives. Furthermore, the system also has a function that facilitates anonymous communication between users with similar conditions, allowing them to receive mental health support.
[0420] A concrete example is a case where, based on a user's facial image taken in a store and input about their mental state, the server evaluates their skin's moisture level and stress level, and then provides a list of recommended skincare products and a meditation program for stress reduction. An example of a prompt used in this case would be, "Please analyze my skin condition and mental stress level today and tell me the best skincare and relaxation methods for me."
[0421] In this way, through the invention, users can receive integrated care tailored to their individual needs, and as a result, their quality of life can be improved.
[0422] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0423] Step 1:
[0424] The user takes an image of their face using a device and inputs audio or text data. This input data is sent to the server. The inputs include image data, audio data, and text data, which provide the information that forms the basis for analysis in the next step.
[0425] Step 2:
[0426] The server analyzes the received image data using image processing tools. Specifically, it uses OpenCV to detect the location and condition of skin problems at the pixel level, and then uses a model trained with TensorFlow to calculate and output indicators of skin moisture, oil content, and problems. This output is then passed on to the next step as information to evaluate the user's skin condition.
[0427] Step 3:
[0428] The server passes the voice or text data input from the user to a natural language processing system. Here, the voice data is converted to text using the Google Cloud Natural Language API, and emotion analysis is performed using the Emotion API. The user's tone of voice and text content are analyzed as input, and emotion data and indicators showing the user's emotional state are generated as output. This provides quantitative data on the user's emotional state.
[0429] Step 4:
[0430] Based on analyzed skin condition and emotional data, the server uses a generation mechanism to generate care methods and advice tailored to the user. Specifically, an AI model evaluates this data and lists appropriate skincare products and relaxation programs. It also includes guidance on relaxation activities applicable to the store environment. This output provides specific care advice that the user can implement.
[0431] Step 5:
[0432] The generated advice is sent to the device and displayed to the user. The device uses Flutter or other native app development environments as its frontend to present the advice in a visually easy-to-understand manner. Users can use this as a reference to engage in their daily care. Specifically, features such as recommended product information and the ability to book in-store activities are provided.
[0433] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0434] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0435] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0436] [Third Embodiment]
[0437] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0438] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0439] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0440] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0441] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0442] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0443] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0444] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0445] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0446] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0447] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0448] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0449] This invention is a system that individually evaluates the user's skin condition and emotional state, and provides appropriate advice and support. In particular, it centers around an image processing module and a natural language processing module, enabling highly personalized care for the user.
[0450] The server first receives image data from the user and uses an image processing module to quantify and evaluate the skin condition. Skin condition includes factors such as humidity, oil balance, and the presence or absence of inflammation. In parallel, the server uses text data and a natural language processing module to analyze the user's emotional state. This analysis provides an indicator of the user's emotions.
[0451] Subsequently, the server uses a generation module based on the analyzed data to generate advice and resources best suited to the user's condition. This advice includes specific care methods and concrete action suggestions for improving their lifestyle. Related resources such as helpful articles and links to online support groups are also provided.
[0452] The terminal displays these analysis results and advice to the user visually, presenting them in an intuitively understandable format. Users can then use this information to improve their own lives. Furthermore, if a user wishes to participate in a community, the community matching module on the server facilitates connection with other users in similar situations. This allows users to actively interact with others and support each other while maintaining anonymity.
[0453] As a concrete example, suppose User A uses the system because they are experiencing skin problems and emotional anxiety. They upload a photo of their skin via their device and answer questions about their emotional state. Based on this information, the server analyzes and indicates that their skin is dry and their emotional state is unstable. The generated advice includes instructions on how to use moisturizing cream and breathing exercises to calm their mind. A support group is also suggested where User A can talk with others experiencing similar emotional states, allowing them to participate with confidence. As a result, User A receives care for both their skin and their emotional state, aiming for a healthier condition.
[0454] The following describes the processing flow.
[0455] Step 1:
[0456] Users take images of their skin condition with their device and upload them through the app. In addition, they answer questions within the app and input text data about their emotional state.
[0457] Step 2:
[0458] The device uses a secure protocol to send captured images and entered text data to the server. The system is designed to maintain data anonymity during transmission.
[0459] Step 3:
[0460] The server passes the received image data to an image processing module, which analyzes the skin's condition. Here, the moisture level, oil content, and presence or absence of inflammation are quantified and evaluated.
[0461] Step 4:
[0462] The server passes the text data to a natural language processing module, which analyzes the user's emotional state. It analyzes tone and keywords to estimate the user's emotional state.
[0463] Step 5:
[0464] The server integrates the results of image and text analysis and uses a generation module to create user-optimized advice. This advice includes specific care methods and resource links.
[0465] Step 6:
[0466] The server uses a community matching module to identify other users with similar statuses and suggests they join relevant communities.
[0467] Step 7:
[0468] The device displays advice and community information sent from the server to the user. Based on this, the user can select actions that are helpful in their daily life.
[0469] Step 8:
[0470] Users follow the provided advice and provide feedback via their device. The server records this feedback and uses it to further improve subsequent personalization.
[0471] (Example 1)
[0472] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0473] In modern society, the physical and mental health problems faced by individual users are diverse, but the support and advice available are general and not optimized for individual needs. Furthermore, there is a lack of platforms that allow for community interaction and mutual support while maintaining anonymity. Against this backdrop, there is a need for a system that can accurately and intuitively grasp the specific condition of users and provide appropriate advice and resources.
[0474] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0475] In this invention, the server includes image processing means for analyzing visual data collected from the user and evaluating the properties of materials; natural language processing means for analyzing text data input by the user and evaluating the user's mental state; and generation means for generating personalized advice and providing related materials to the user based on the analyzed data. This makes it possible to build a platform that provides advice tailored to the user's condition and facilitates connections with other users who have similar conditions.
[0476] "Visual data" refers to image information collected to show the user's skin condition and other physical characteristics.
[0477] "Text data" refers to text information entered or provided by the user, and the user's mental state is analyzed based on this data.
[0478] "Image processing means" refers to a module for analyzing visual data and quantifying and evaluating the properties of a substance.
[0479] A "natural language processing tool" is a module that processes text data and executes an algorithm to quantify and evaluate the user's mental state.
[0480] The "generation means" is a module that generates user-optimized advice and related materials based on the analysis results of image and text data.
[0481] A "protected communication specification" is a communication protocol used to maintain user anonymity and securely send and receive data.
[0482] The "feedback collection function" is a feedback mechanism that tracks how users' advice is being adopted and helps improve future proposals.
[0483] A "matching tool" is a module that identifies other users with similar statuses based on the user's status, thereby promoting group participation.
[0484] A description of embodiments for carrying out this invention will be given.
[0485] The server plays a central role in evaluating the user's skin and emotional state and providing appropriate advice. When a user provides images of their face and skin through their device, the server analyzes the visual data using advanced image processing software, such as open-source image processing libraries. This process quantifies and evaluates skin characteristics such as humidity, oil balance, and the presence or absence of inflammation.
[0486] In parallel, the user provides text data that expresses their emotional state. This text is analyzed by the server using natural language processing tools. For example, a natural language processing library is used to analyze the text and obtain an index of the user's emotions. These analysis results are then input into a generative AI model, which serves as a means to generate user-optimized advice and relevant resources.
[0487] The terminal visually displays analysis results and advice from the server, presenting them in a way that users can intuitively understand. Based on this information, users can reflect specific actions in their daily lives. For example, they might practice the suggested use of moisturizing cream or breathing exercises to calm their minds. The server also encourages community interaction with other users who have similar concerns, providing a space where users can support each other while maintaining anonymity.
[0488] As a concrete example, consider a scenario where a user provides image and text data regarding dry skin or mental stress. Based on this information, the server generates appropriate care methods and support group resources, which are then presented to the user on their device. This allows the user to receive the care best suited to their condition.
[0489] An example of a prompt for the generative AI model might be: "Analyze the skin image provided by the user and quantify its condition. Next, analyze the text representing the user's emotional state to generate personalized advice." This allows the system to work efficiently and provide accurate, individualized care.
[0490] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0491] Step 1:
[0492] The user takes a photo of their skin via their device and enters text about their emotional state. This visual and text data is sent as input to the server. Specifically, the device's camera function is used to acquire a high-resolution image, and the user enters free-form text about their emotions and state.
[0493] Step 2:
[0494] The server analyzes the received visual data using image processing software. After color correction and noise reduction, the input image data is quantified using a feature extraction algorithm to determine characteristics such as humidity, oil balance, and inflammation. Specific numerical data indicating the skin condition is generated as output and sent to the next processing step.
[0495] Step 3:
[0496] The server analyzes the text data provided by the user using a natural language processing library. The input text is tokenized through morphological analysis and sentiment analysis, and emotional indicators such as positive or negative are quantified. This output is numerical data that represents the user's emotional state and forms the basis for subsequent generation steps.
[0497] Step 4:
[0498] The server integrates numerical data obtained from image and text analysis and inputs it as a prompt into the generating AI model. Specifically, it sends the prompt example "Please suggest the best care method for a user with low skin moisture levels and a negative emotional index" to the model. The output of this process is personalized advice and relevant resource information.
[0499] Step 5:
[0500] The server outputs advice and resource information, which is then sent to the terminal, where it is displayed visually to the user. Specifically, graphs, charts, and text information are displayed on the screen to enable visual and intuitive understanding. This output serves as a suggestion for the user's next action and is provided as information to help improve their lifestyle.
[0501] (Application Example 1)
[0502] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0503] There is a need to accurately assess users' skin and emotional states and, based on that, suggest the most suitable care and products. However, conventional systems make it difficult to provide immediate, personalized advice in physical stores. Furthermore, there is a lack of environment that encourages interaction among users with similar concerns, fostering a sense of security and encouraging continued care.
[0504] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0505] In this invention, the server includes data processing means for analyzing image data collected from the user and evaluating the condition of the skin, natural language processing means for analyzing text data input by the user and evaluating the user's mental state, and store support means for immediately providing evaluation results and product suggestions to the user in a physical store. As a result, the user can immediately understand their own condition and receive suggestions for the most suitable products and care methods on the spot. It can also facilitate interaction with other users in similar conditions and provide support for continuous care.
[0506] A "user" refers to a person who uses the system to evaluate their own skin and mental state and receive recommendations for optimal care and products.
[0507] "Image data" refers to digital image information collected to evaluate the condition of a user's skin.
[0508] "Data processing means" refers to a device that analyzes collected image data and has the function of quantifying and evaluating the condition of the skin.
[0509] "Text data" refers to information such as sentences and words that users input to express their own mental state.
[0510] A "natural language processing system" is a device that processes text data input by a user and has the function of evaluating the user's mental state and emotions.
[0511] A "generation means" is a device that has the function of creating optimal advice and related information for the user based on the analysis results.
[0512] A "group matching method" is a system that identifies other users with similar statuses based on the user's status and has the function of promoting participation in collaboration.
[0513] "Store support means" refers to a function that allows users to receive evaluation results immediately at a physical store and receive recommendations for the most suitable products.
[0514] A "secure protocol" is a means of communication that allows for the secure transmission and reception of data while maintaining user anonymity.
[0515] The "feedback tracking function" is a feature that records the adoption status of advice received by the user and uses that information to improve subsequent suggestions.
[0516] The system that realizes this invention operates with the following configuration.
[0517] First, the user takes a picture of their skin using a dedicated terminal or smart device and uploads the photo to the system. The terminal sends the image data to a server, which analyzes the image using data processing tools. The analyzed data is converted into numerical data to evaluate specific skin conditions such as skin moisture, oil content, and the presence or absence of inflammation. Image processing libraries such as OpenCV are used for this process.
[0518] In parallel, the user answers a simple questionnaire displayed on the device, inputting their emotional state as text data. The device sends this text data to a server, which uses natural language processing to evaluate the emotions. Natural language processing libraries such as TextBlob are used for this part of the analysis, quantifying the degree of positivity or negativity of the emotions.
[0519] Based on the numerical data obtained regarding skin condition and mental state, the server generates optimal advice and suggestions for the user through a generation mechanism. For example, if moisturizing is needed, it will suggest using a moisturizing cream; if the user is feeling depressed, it will recommend relaxation techniques. By utilizing a generation AI model, it is possible to provide suggestions tailored to individual needs. Related information such as links to articles and introductions to appropriate support groups are also provided.
[0520] Furthermore, if a user is in a physical store, they can immediately view and purchase suggested products within the store using in-store support tools. This process makes it easier for users to obtain products that suit their needs in real time.
[0521] As a concrete example, suppose user C uses this system to take a picture of their skin condition in a store and answers questions about their emotional state. The system analyzes skin dryness and mild mental stress and suggests moisturizing cream and relaxation products. These products are suggested in the store and can be purchased on the spot. An example of a prompt to input into the generating AI model is, "Please analyze what emotions are most frequently expressed regarding the user's emotional state."
[0522] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0523] Step 1:
[0524] The user takes a picture of their skin using a device, and image data is acquired. The input is the captured image, which is then sent for subsequent processing. The user presses the capture button on the device, and the image is saved to the system.
[0525] Step 2:
[0526] The terminal sends image data to the server. The input is image data from the terminal, and the output is an image file stored on the server. The terminal transfers the image to the server using wireless communication.
[0527] Step 3:
[0528] The server receives image data and uses data processing tools to analyze the skin condition. The input is image data stored on the server, and the output is numerical data indicating the skin's moisture and oil balance, as well as the presence or absence of inflammation. The server uses OpenCV to analyze the image features and calculate evaluation values.
[0529] Step 4:
[0530] The user answers questions about their mental state displayed on the device and inputs text data. The input is text data entered by the user and sent to the server. The user answers the questions on the device and describes their mental state in sentences.
[0531] Step 5:
[0532] The terminal sends text data to the server. The input is the user's text data, and the output is the text data stored on the server. The terminal then transfers the data back to the server using wireless communication.
[0533] Step 6:
[0534] The server receives text data and evaluates the emotions using natural language processing. The input is text data stored on the server, and the output is numerical data indicating the degree of positivity or negativity of the emotions. The server uses TextBlob to analyze the emotions in the text data and calculate an evaluation value.
[0535] Step 7:
[0536] The server uses a generation mechanism to generate optimal advice based on the analysis results. The input is numerical data indicating the condition of the skin and mind, and the output is advice and product suggestions for the user. The server utilizes a generation AI model to generate individual suggestions.
[0537] Step 8:
[0538] When a user is in a physical store, the terminal uses store support tools to display advice and product suggestions. The input is advice from the server, and the output is information displayed to the user on the terminal. The terminal presents the results to the user through a simple interface.
[0539] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0540] This invention combines an emotional engine with a system that analyzes a user's skin and emotional state to provide optimal advice. The entire system aims to enable highly personalized care for each user and improve their quality of life.
[0541] The server first receives image data from the user and uses an image processing module to evaluate the skin condition. This evaluation includes indicators of humidity, oil content, and skin problems. Next, the server passes the text and voice data entered by the user to a natural language processing module and an emotion engine to recognize the user's mental state and emotions.
[0542] The emotion engine identifies the user's emotions from text and audio, and the server generates emotion data based on this. This emotion data is combined with other analysis results, and a generation module generates personalized emotional support advice. This advice suggests specific care methods and activities to stabilize emotions. Related online resources and community groups are also provided.
[0543] The device displays advice and emotional assessment results sent from the server to the user. Users can use this information to improve their lives and emotional state. Furthermore, they can join communities recommended based on their state, allowing them to interact with others anonymously and share their emotions.
[0544] As a concrete example, suppose user B perceives skin problems associated with daily stress. User B uploads an image on their device and answers questions about their emotional state. The server uses an image processing module to assess the dryness of the skin, an emotion engine to evaluate the stress level, and a generation module to provide suggestions for moisturizing care and meditation methods to reduce stress. It also encourages user B to join a community that helps people cope with stress. As a result, user B receives support for both their physical and mental well-being, enabling them to strive for a healthier state.
[0545] The following describes the processing flow.
[0546] Step 1:
[0547] Users use their devices to take pictures showing their skin condition and upload them through the application. They also answer questions about their emotional state and record voice messages.
[0548] Step 2:
[0549] The device transmits uploaded image data, text data, and audio data to the server via a secure protocol. User anonymity is ensured during this process.
[0550] Step 3:
[0551] The server inputs the received image data into an image processing module to analyze the skin condition. The analysis results include humidity level, oil balance, and the presence or absence of skin problems.
[0552] Step 4:
[0553] The server inputs text and audio data into a natural language processing module and an emotion engine to analyze the user's mental state and emotions. This process generates an emotion index for the user.
[0554] Step 5:
[0555] The server combines the image processing and emotion engine analysis results and passes them to the generation module. The generation module generates emotional support advice that is best suited to the user. This advice includes specific self-care methods and suggestions for emotional stabilization.
[0556] Step 6:
[0557] The server utilizes a community matching module to identify other users with similar statuses. This allows it to generate recommendations for the most suitable online community groups for each user.
[0558] Step 7:
[0559] The terminal displays analysis results and advice sent from the server, as well as suggestions for community participation. Users refer to the provided information and choose their actions based on it.
[0560] Step 8:
[0561] Users incorporate the advice provided into their lives and experience its effects. Feedback is transmitted to the server via the device, and the server uses this information to improve the quality of future suggestions.
[0562] (Example 2)
[0563] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0564] Modern cosmetic and mental health support systems offer generic recommendations to users, lacking optimized support that fully considers the specific circumstances and emotional factors of individual users. Furthermore, while there is a need to provide personalized communication and advice while ensuring user anonymity and safety, efficient and effective methods are lacking. Additionally, the inadequate mechanisms for tracking how users adopt suggested advice and the results limit opportunities for improving service quality.
[0565] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0566] In this invention, the server includes image processing means for analyzing image information collected from the user and evaluating its surface state; natural information processing means for analyzing text and voice information input by the user and evaluating its mental state; and generation means for generating support optimized for the user and providing related resources based on the analyzed information. This enables personalized and optimal support based on the user's state, secure and anonymous communication, and improvement of service quality based on feedback.
[0567] "Image information" refers to visual data such as photographs and illustrations that show the surface condition of the user.
[0568] "Image processing means" refers to technical methods and processes for analyzing the above image information and evaluating the condition of skin and other surface areas.
[0569] "Textual information" refers to text-based data entered by the user, which is used to evaluate the user's mental state.
[0570] "Audio information" refers to sound data that users input via voice, which allows for the evaluation of the user's emotions and mental state.
[0571] "Natural information processing means" refers to technical means and systems for analyzing textual and auditory information and evaluating the user's mental state.
[0572] "Generating means" refers to methods and means for providing optimized support and related resources based on the analysis results of the user's surface state and mental state.
[0573] "Methods for promoting participation" refer to technologies and methods that identify other users with similar characteristics based on the user's status and promote communication with others and participation in a community.
[0574] "Protective communication means" refers to communication technologies and protocols that maintain user anonymity and ensure security when sending and receiving data.
[0575] "Methods for gathering feedback" refer to functions and methods for tracking user feedback, understanding the adoption of support, and improving the service.
[0576] This invention is a system that analyzes the user's surface and mental state and provides optimized support. The main components of the system are a server, a terminal, and input data from the user.
[0577] The server performs various data processing and calculations using the following hardware and software. First, image information provided by the user is processed using the open-source library OpenCV. This identifies indicators of the user's skin's moisture, oil content, and problems, and evaluates them as numerical data.
[0578] Furthermore, the text and audio information obtained from users is analyzed using a natural language processing engine. A method is employed that utilizes Google Cloud NLP to evaluate mental state from the input free-form text and audio data. Audio information is converted into text using speech recognition technology and further analyzed as text.
[0579] The generation method utilizes a generative AI model. Specifically, it uses Azure's generative AI via a cloud-based AI service to generate optimal support and advice based on the user's condition. The generated content includes not only skincare but also suggestions for maintaining mental health and related online resources.
[0580] The terminal displays information retrieved from the server in a format that is easy for the user to understand. The terminal's interface is designed to be easy for the user to operate and to quickly obtain the necessary advice.
[0581] For example, if a user is troubled by daily stress and dry skin, this system analyzes the user's image and evaluates their skin condition. It also measures the user's mental stress level through voice input. Based on this data, the server provides the user with specific skincare methods and meditation techniques for stress relief.
[0582] The following are examples of prompts to input into a generative AI model:
[0583] "The user's skin condition is dry, and their mental state is rated as high stress. Please generate optimal skincare advice and stress management activity suggestions for this user."
[0584] These processes enable users to receive effective support tailored to their individual needs.
[0585] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0586] Step 1:
[0587] User data entry
[0588] Users take photos of their skin using their device and send the image data to a server through the application. Users also answer questions about their emotions and mental state, inputting this information as voice or text data.
[0589] Input: Skin image data, text data, or audio data
[0590] The specific actions involve taking an image using the device's camera function and then pressing the image upload button to send the data. Regarding the state of mind, the user either answers questions verbally using the device's voice recognition function or enters text using the keyboard.
[0591] Step 2:
[0592] Server-based analysis of skin condition
[0593] The server receives image data sent by the user and analyzes it using an image processing module. Specifically, it uses the OpenCV library to calculate skin moisture, oil content, and trouble indicators. As a result, numerical data is generated.
[0594] Input: Skin image data submitted by the user
[0595] Output: Skin evaluation results (numerical data for humidity, oil content, and trouble indicators)
[0596] Specifically, a Python script runs on the server to calculate these values and form the evaluation results.
[0597] Step 3:
[0598] Analysis of mental state by server
[0599] The server passes text and audio data provided by the user to a natural language processing module. Text analysis is performed using Google Cloud NLP services, and emotions are identified using TEDAS. Audio data is converted to text using speech recognition technology, and this text data is analyzed in the same manner.
[0600] Input: Text or audio data sent by the user.
[0601] Output: Mental state evaluation results (emotion tags)
[0602] Specifically, the server converts the audio data into text, performs text analysis, classifies the sentiment, and tags the data.
[0603] Step 4:
[0604] Advice generation using generative AI models
[0605] The server uses collected data on skin condition and mental state to create prompts for a generative AI model and generate optimal advice. Using Azure's generative AI, it provides specific skincare methods and suggestions for improving mental well-being.
[0606] Input: Skin evaluation results, mental state evaluation results
[0607] Output: Optimized advice for the user
[0608] In terms of operation, the AI model on the server takes the prompt text as input, generates optimized advice, and returns it.
[0609] Step 5:
[0610] Advice and results displayed on the device
[0611] The device receives optimized advice and analysis results from the server and displays them to the user. This includes care instructions, links to resources, and community invitations.
[0612] Input: Generation advice and analysis results from the server.
[0613] Output: Display and notification to the user
[0614] Specifically, the device uses push notifications to inform the user of the arrival of new advice and displays detailed information within the app.
[0615] (Application Example 2)
[0616] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0617] Modern consumers face various skin and mental health issues due to daily stress and environmental changes. However, conventional skincare products and mental health support often only offer generalized advice. Therefore, it is difficult to provide highly personalized care for each individual. Furthermore, customers who use physical stores are not adequately provided with guidance on relaxation appropriate to the store environment or mental support through interaction with the community. Addressing these challenges is necessary to improve the customer experience and quality of life.
[0618] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0619] In this invention, the server includes image processing means for analyzing image data collected from the user and evaluating the condition of the skin; natural language processing means for analyzing voice data and text data input from the user and evaluating emotions and mental state; and generation means for generating advice, including care methods and emotional stabilization activities optimized for the user, based on the analyzed data, and for guiding relaxation activities applicable in the store environment. This enables personalized care and emotional support for individual users, improving the customer experience when using the store and enhancing the quality of life.
[0620] "Image processing means" refers to a processing device that analyzes image data provided by the user and evaluates the condition of the skin.
[0621] "Natural language processing means" refers to a processing device that analyzes voice data and text data input by a user and evaluates emotions and mental states.
[0622] The "generation means" is a processing device that generates care methods optimized for the user and advice including emotional stabilization activities based on analyzed data, and guides users through relaxation activities applicable to the store environment.
[0623] A "community matching means" is a processing device that identifies other users in similar states based on the user's state and facilitates anonymous interaction.
[0624] This invention is a system that analyzes the user's skin and emotional state and provides optimal care and emotional support. The system mainly consists of three components: a server, a terminal, and the user.
[0625] The server analyzes image data received from the user using image processing tools and evaluates the skin condition using software such as TensorFlow and OpenCV. Furthermore, it performs natural language processing on voice and text data input from the user using the Google Cloud Natural Language API and evaluates emotions and mental state using the Emotion API. Based on the data obtained in this way, a generation tool generates advice, including care methods and emotional stabilization activities optimized for the user, and sends the results to the terminal.
[0626] The device displays received advice to the user, presenting it visually and clearly using a smartphone or smart glasses. This allows users to implement specific care methods tailored to their own condition. For example, customers can receive detailed information about dedicated relaxation spaces and activities offered within the physical store.
[0627] Through this system, users can gain a comprehensive understanding of their skin and mental state, and incorporate personalized advice into their daily lives. Furthermore, the system also has a function that facilitates anonymous communication between users with similar conditions, allowing them to receive mental health support.
[0628] A concrete example is a case where, based on a user's facial image taken in a store and input about their mental state, the server evaluates their skin's moisture level and stress level, and then provides a list of recommended skincare products and a meditation program for stress reduction. An example of a prompt used in this case would be, "Please analyze my skin condition and mental stress level today and tell me the best skincare and relaxation methods for me."
[0629] In this way, through the invention, users can receive integrated care tailored to their individual needs, and as a result, their quality of life can be improved.
[0630] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0631] Step 1:
[0632] The user takes an image of their face using a device and inputs audio or text data. This input data is sent to the server. The inputs include image data, audio data, and text data, which provide the information that forms the basis for analysis in the next step.
[0633] Step 2:
[0634] The server analyzes the received image data using image processing tools. Specifically, it uses OpenCV to detect the location and condition of skin problems at the pixel level, and then uses a model trained with TensorFlow to calculate and output indicators of skin moisture, oil content, and problems. This output is then passed on to the next step as information to evaluate the user's skin condition.
[0635] Step 3:
[0636] The server passes the voice or text data input from the user to a natural language processing system. Here, the voice data is converted to text using the Google Cloud Natural Language API, and emotion analysis is performed using the Emotion API. The user's tone of voice and text content are analyzed as input, and emotion data and indicators showing the user's emotional state are generated as output. This provides quantitative data on the user's emotional state.
[0637] Step 4:
[0638] Based on analyzed skin condition and emotional data, the server uses a generation mechanism to generate care methods and advice tailored to the user. Specifically, an AI model evaluates this data and lists appropriate skincare products and relaxation programs. It also includes guidance on relaxation activities applicable to the store environment. This output provides specific care advice that the user can implement.
[0639] Step 5:
[0640] The generated advice is sent to the device and displayed to the user. The device uses Flutter or other native app development environments as its frontend to present the advice in a visually easy-to-understand manner. Users can use this as a reference to engage in their daily care. Specifically, features such as recommended product information and the ability to book in-store activities are provided.
[0641] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0642] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0643] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0644] [Fourth Embodiment]
[0645] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0646] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0647] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0648] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0649] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0650] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0651] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0652] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0653] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0654] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0655] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0656] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0657] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0658] This invention is a system that individually evaluates the user's skin condition and emotional state, and provides appropriate advice and support. In particular, it centers around an image processing module and a natural language processing module, enabling highly personalized care for the user.
[0659] The server first receives image data from the user and uses an image processing module to quantify and evaluate the skin condition. Skin condition includes factors such as humidity, oil balance, and the presence or absence of inflammation. In parallel, the server uses text data and a natural language processing module to analyze the user's emotional state. This analysis provides an indicator of the user's emotions.
[0660] Subsequently, the server uses a generation module based on the analyzed data to generate advice and resources best suited to the user's condition. This advice includes specific care methods and concrete action suggestions for improving their lifestyle. Related resources such as helpful articles and links to online support groups are also provided.
[0661] The terminal displays these analysis results and advice to the user visually, presenting them in an intuitively understandable format. Users can then use this information to improve their own lives. Furthermore, if a user wishes to participate in a community, the community matching module on the server facilitates connection with other users in similar situations. This allows users to actively interact with others and support each other while maintaining anonymity.
[0662] As a concrete example, suppose User A uses the system because they are experiencing skin problems and emotional anxiety. They upload a photo of their skin via their device and answer questions about their emotional state. Based on this information, the server analyzes and indicates that their skin is dry and their emotional state is unstable. The generated advice includes instructions on how to use moisturizing cream and breathing exercises to calm their mind. A support group is also suggested where User A can talk with others experiencing similar emotional states, allowing them to participate with confidence. As a result, User A receives care for both their skin and their emotional state, aiming for a healthier condition.
[0663] The following describes the processing flow.
[0664] Step 1:
[0665] Users take images of their skin condition with their device and upload them through the app. In addition, they answer questions within the app and input text data about their emotional state.
[0666] Step 2:
[0667] The device uses a secure protocol to send captured images and entered text data to the server. The system is designed to maintain data anonymity during transmission.
[0668] Step 3:
[0669] The server passes the received image data to an image processing module, which analyzes the skin's condition. Here, the moisture level, oil content, and presence or absence of inflammation are quantified and evaluated.
[0670] Step 4:
[0671] The server passes the text data to a natural language processing module, which analyzes the user's emotional state. It analyzes tone and keywords to estimate the user's emotional state.
[0672] Step 5:
[0673] The server integrates the results of image and text analysis and uses a generation module to create user-optimized advice. This advice includes specific care methods and resource links.
[0674] Step 6:
[0675] The server uses a community matching module to identify other users with similar statuses and suggests they join relevant communities.
[0676] Step 7:
[0677] The device displays advice and community information sent from the server to the user. Based on this, the user can select actions that are helpful in their daily life.
[0678] Step 8:
[0679] Users follow the provided advice and provide feedback via their device. The server records this feedback and uses it to further improve subsequent personalization.
[0680] (Example 1)
[0681] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0682] In modern society, the physical and mental health problems faced by individual users are diverse, but the support and advice available are general and not optimized for individual needs. Furthermore, there is a lack of platforms that allow for community interaction and mutual support while maintaining anonymity. Against this backdrop, there is a need for a system that can accurately and intuitively grasp the specific condition of users and provide appropriate advice and resources.
[0683] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0684] In this invention, the server includes image processing means for analyzing visual data collected from the user and evaluating the properties of materials; natural language processing means for analyzing text data input by the user and evaluating the user's mental state; and generation means for generating personalized advice and providing related materials to the user based on the analyzed data. This makes it possible to build a platform that provides advice tailored to the user's condition and facilitates connections with other users who have similar conditions.
[0685] "Visual data" refers to image information collected to show the user's skin condition and other physical characteristics.
[0686] "Text data" refers to text information entered or provided by the user, and the user's mental state is analyzed based on this data.
[0687] "Image processing means" refers to a module for analyzing visual data and quantifying and evaluating the properties of a substance.
[0688] A "natural language processing tool" is a module that processes text data and executes an algorithm to quantify and evaluate the user's mental state.
[0689] The "generation means" is a module that generates user-optimized advice and related materials based on the analysis results of image and text data.
[0690] A "protected communication specification" is a communication protocol used to maintain user anonymity and securely send and receive data.
[0691] The "feedback collection function" is a feedback mechanism that tracks how users' advice is being adopted and helps improve future proposals.
[0692] A "matching tool" is a module that identifies other users with similar statuses based on the user's status, thereby promoting group participation.
[0693] A description of embodiments for carrying out this invention will be given.
[0694] The server plays a central role in evaluating the user's skin and emotional state and providing appropriate advice. When a user provides images of their face and skin through their device, the server analyzes the visual data using advanced image processing software, such as open-source image processing libraries. This process quantifies and evaluates skin characteristics such as humidity, oil balance, and the presence or absence of inflammation.
[0695] In parallel, the user provides text data that expresses their emotional state. This text is analyzed by the server using natural language processing tools. For example, a natural language processing library is used to analyze the text and obtain an index of the user's emotions. These analysis results are then input into a generative AI model, which serves as a means to generate user-optimized advice and relevant resources.
[0696] The terminal visually displays analysis results and advice from the server, presenting them in a way that users can intuitively understand. Based on this information, users can reflect specific actions in their daily lives. For example, they might practice the suggested use of moisturizing cream or breathing exercises to calm their minds. The server also encourages community interaction with other users who have similar concerns, providing a space where users can support each other while maintaining anonymity.
[0697] As a concrete example, consider a scenario where a user provides image and text data regarding dry skin or mental stress. Based on this information, the server generates appropriate care methods and support group resources, which are then presented to the user on their device. This allows the user to receive the care best suited to their condition.
[0698] An example of a prompt for the generative AI model might be: "Analyze the skin image provided by the user and quantify its condition. Next, analyze the text representing the user's emotional state to generate personalized advice." This allows the system to work efficiently and provide accurate, individualized care.
[0699] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0700] Step 1:
[0701] The user takes a photo of their skin via their device and enters text about their emotional state. This visual and text data is sent as input to the server. Specifically, the device's camera function is used to acquire a high-resolution image, and the user enters free-form text about their emotions and state.
[0702] Step 2:
[0703] The server analyzes the received visual data using image processing software. After color correction and noise reduction, the input image data is quantified using a feature extraction algorithm to determine characteristics such as humidity, oil balance, and inflammation. Specific numerical data indicating the skin condition is generated as output and sent to the next processing step.
[0704] Step 3:
[0705] The server analyzes the text data provided by the user using a natural language processing library. The input text is tokenized through morphological analysis and sentiment analysis, and emotional indicators such as positive or negative are quantified. This output is numerical data that represents the user's emotional state and forms the basis for subsequent generation steps.
[0706] Step 4:
[0707] The server integrates numerical data obtained from image and text analysis and inputs it as a prompt into the generating AI model. Specifically, it sends the prompt example "Please suggest the best care method for a user with low skin moisture levels and a negative emotional index" to the model. The output of this process is personalized advice and relevant resource information.
[0708] Step 5:
[0709] The server outputs advice and resource information, which is then sent to the terminal, where it is displayed visually to the user. Specifically, graphs, charts, and text information are displayed on the screen to enable visual and intuitive understanding. This output serves as a suggestion for the user's next action and is provided as information to help improve their lifestyle.
[0710] (Application Example 1)
[0711] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0712] There is a need to accurately assess users' skin and emotional states and, based on that, suggest the most suitable care and products. However, conventional systems make it difficult to provide immediate, personalized advice in physical stores. Furthermore, there is a lack of environment that encourages interaction among users with similar concerns, fostering a sense of security and encouraging continued care.
[0713] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0714] In this invention, the server includes data processing means for analyzing image data collected from the user and evaluating the condition of the skin, natural language processing means for analyzing text data input by the user and evaluating the user's mental state, and store support means for immediately providing evaluation results and product suggestions to the user in a physical store. As a result, the user can immediately understand their own condition and receive suggestions for the most suitable products and care methods on the spot. It can also facilitate interaction with other users in similar conditions and provide support for continuous care.
[0715] A "user" refers to a person who uses the system to evaluate their own skin and mental state and receive recommendations for optimal care and products.
[0716] "Image data" refers to digital image information collected to evaluate the condition of a user's skin.
[0717] "Data processing means" refers to a device that analyzes collected image data and has the function of quantifying and evaluating the condition of the skin.
[0718] "Text data" refers to information such as sentences and words that users input to express their own mental state.
[0719] A "natural language processing system" is a device that processes text data input by a user and has the function of evaluating the user's mental state and emotions.
[0720] A "generation means" is a device that has the function of creating optimal advice and related information for the user based on the analysis results.
[0721] A "group matching method" is a system that identifies other users with similar statuses based on the user's status and has the function of promoting participation in collaboration.
[0722] "Store support means" refers to a function that allows users to receive evaluation results immediately at a physical store and receive recommendations for the most suitable products.
[0723] A "secure protocol" is a means of communication that allows for the secure transmission and reception of data while maintaining user anonymity.
[0724] The "feedback tracking function" is a feature that records the adoption status of advice received by the user and uses that information to improve subsequent suggestions.
[0725] The system that realizes this invention operates with the following configuration.
[0726] First, the user takes a picture of their skin using a dedicated terminal or smart device and uploads the photo to the system. The terminal sends the image data to a server, which analyzes the image using data processing tools. The analyzed data is converted into numerical data to evaluate specific skin conditions such as skin moisture, oil content, and the presence or absence of inflammation. Image processing libraries such as OpenCV are used for this process.
[0727] In parallel, the user answers a simple questionnaire displayed on the device, inputting their emotional state as text data. The device sends this text data to a server, which uses natural language processing to evaluate the emotions. Natural language processing libraries such as TextBlob are used for this part of the analysis, quantifying the degree of positivity or negativity of the emotions.
[0728] Based on the numerical data obtained regarding skin condition and mental state, the server generates optimal advice and suggestions for the user through a generation mechanism. For example, if moisturizing is needed, it will suggest using a moisturizing cream; if the user is feeling depressed, it will recommend relaxation techniques. By utilizing a generation AI model, it is possible to provide suggestions tailored to individual needs. Related information such as links to articles and introductions to appropriate support groups are also provided.
[0729] Furthermore, if a user is in a physical store, they can immediately view and purchase suggested products within the store using in-store support tools. This process makes it easier for users to obtain products that suit their needs in real time.
[0730] As a concrete example, suppose user C uses this system to take a picture of their skin condition in a store and answers questions about their emotional state. The system analyzes skin dryness and mild mental stress and suggests moisturizing cream and relaxation products. These products are suggested in the store and can be purchased on the spot. An example of a prompt to input into the generating AI model is, "Please analyze what emotions are most frequently expressed regarding the user's emotional state."
[0731] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0732] Step 1:
[0733] The user takes a picture of their skin using a device, and image data is acquired. The input is the captured image, which is then sent for subsequent processing. The user presses the capture button on the device, and the image is saved to the system.
[0734] Step 2:
[0735] The terminal sends image data to the server. The input is image data from the terminal, and the output is an image file stored on the server. The terminal transfers the image to the server using wireless communication.
[0736] Step 3:
[0737] The server receives image data and uses data processing tools to analyze the skin condition. The input is image data stored on the server, and the output is numerical data indicating the skin's moisture and oil balance, as well as the presence or absence of inflammation. The server uses OpenCV to analyze the image features and calculate evaluation values.
[0738] Step 4:
[0739] The user answers questions about their mental state displayed on the device and inputs text data. The input is text data entered by the user and sent to the server. The user answers the questions on the device and describes their mental state in sentences.
[0740] Step 5:
[0741] The terminal sends text data to the server. The input is the user's text data, and the output is the text data stored on the server. The terminal then transfers the data back to the server using wireless communication.
[0742] Step 6:
[0743] The server receives text data and evaluates the emotions using natural language processing. The input is text data stored on the server, and the output is numerical data indicating the degree of positivity or negativity of the emotions. The server uses TextBlob to analyze the emotions in the text data and calculate an evaluation value.
[0744] Step 7:
[0745] The server uses a generation mechanism to generate optimal advice based on the analysis results. The input is numerical data indicating the condition of the skin and mind, and the output is advice and product suggestions for the user. The server utilizes a generation AI model to generate individual suggestions.
[0746] Step 8:
[0747] When a user is in a physical store, the terminal uses store support tools to display advice and product suggestions. The input is advice from the server, and the output is information displayed to the user on the terminal. The terminal presents the results to the user through a simple interface.
[0748] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0749] This invention combines an emotional engine with a system that analyzes a user's skin and emotional state to provide optimal advice. The entire system aims to enable highly personalized care for each user and improve their quality of life.
[0750] The server first receives image data from the user and uses an image processing module to evaluate the skin condition. This evaluation includes indicators of humidity, oil content, and skin problems. Next, the server passes the text and voice data entered by the user to a natural language processing module and an emotion engine to recognize the user's mental state and emotions.
[0751] The emotion engine identifies the user's emotions from text and audio, and the server generates emotion data based on this. This emotion data is combined with other analysis results, and a generation module generates personalized emotional support advice. This advice suggests specific care methods and activities to stabilize emotions. Related online resources and community groups are also provided.
[0752] The device displays advice and emotional assessment results sent from the server to the user. Users can use this information to improve their lives and emotional state. Furthermore, they can join communities recommended based on their state, allowing them to interact with others anonymously and share their emotions.
[0753] As a concrete example, suppose user B perceives skin problems associated with daily stress. User B uploads an image on their device and answers questions about their emotional state. The server uses an image processing module to assess the dryness of the skin, an emotion engine to evaluate the stress level, and a generation module to provide suggestions for moisturizing care and meditation methods to reduce stress. It also encourages user B to join a community that helps people cope with stress. As a result, user B receives support for both their physical and mental well-being, enabling them to strive for a healthier state.
[0754] The following describes the processing flow.
[0755] Step 1:
[0756] Users use their devices to take pictures showing their skin condition and upload them through the application. They also answer questions about their emotional state and record voice messages.
[0757] Step 2:
[0758] The device transmits uploaded image data, text data, and audio data to the server via a secure protocol. User anonymity is ensured during this process.
[0759] Step 3:
[0760] The server inputs the received image data into an image processing module to analyze the skin condition. The analysis results include humidity level, oil balance, and the presence or absence of skin problems.
[0761] Step 4:
[0762] The server inputs text and audio data into a natural language processing module and an emotion engine to analyze the user's mental state and emotions. This process generates an emotion index for the user.
[0763] Step 5:
[0764] The server combines the image processing and emotion engine analysis results and passes them to the generation module. The generation module generates emotional support advice that is best suited to the user. This advice includes specific self-care methods and suggestions for emotional stabilization.
[0765] Step 6:
[0766] The server utilizes a community matching module to identify other users with similar statuses. This allows it to generate recommendations for the most suitable online community groups for each user.
[0767] Step 7:
[0768] The terminal displays analysis results and advice sent from the server, as well as suggestions for community participation. Users refer to the provided information and choose their actions based on it.
[0769] Step 8:
[0770] Users incorporate the advice provided into their lives and experience its effects. Feedback is transmitted to the server via the device, and the server uses this information to improve the quality of future suggestions.
[0771] (Example 2)
[0772] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0773] Modern cosmetic and mental health support systems offer generic recommendations to users, lacking optimized support that fully considers the specific circumstances and emotional factors of individual users. Furthermore, while there is a need to provide personalized communication and advice while ensuring user anonymity and safety, efficient and effective methods are lacking. Additionally, the inadequate mechanisms for tracking how users adopt suggested advice and the results limit opportunities for improving service quality.
[0774] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0775] In this invention, the server includes image processing means for analyzing image information collected from the user and evaluating its surface state; natural information processing means for analyzing text and voice information input by the user and evaluating its mental state; and generation means for generating support optimized for the user and providing related resources based on the analyzed information. This enables personalized and optimal support based on the user's state, secure and anonymous communication, and improvement of service quality based on feedback.
[0776] "Image information" refers to visual data such as photographs and illustrations that show the surface condition of the user.
[0777] "Image processing means" refers to technical methods and processes for analyzing the above image information and evaluating the condition of skin and other surface areas.
[0778] "Textual information" refers to text-based data entered by the user, which is used to evaluate the user's mental state.
[0779] "Audio information" refers to sound data that users input via voice, which allows for the evaluation of the user's emotions and mental state.
[0780] "Natural information processing means" refers to technical means and systems for analyzing textual and auditory information and evaluating the user's mental state.
[0781] "Generating means" refers to methods and means for providing optimized support and related resources based on the analysis results of the user's surface state and mental state.
[0782] "Methods for promoting participation" refer to technologies and methods that identify other users with similar characteristics based on the user's status and promote communication with others and participation in a community.
[0783] "Protective communication means" refers to communication technologies and protocols that maintain user anonymity and ensure security when sending and receiving data.
[0784] "Methods for gathering feedback" refer to functions and methods for tracking user feedback, understanding the adoption of support, and improving the service.
[0785] This invention is a system that analyzes the user's surface and mental state and provides optimized support. The main components of the system are a server, a terminal, and input data from the user.
[0786] The server performs various data processing and calculations using the following hardware and software. First, image information provided by the user is processed using the open-source library OpenCV. This identifies indicators of the user's skin's moisture, oil content, and problems, and evaluates them as numerical data.
[0787] Furthermore, the text and audio information obtained from users is analyzed using a natural language processing engine. A method is employed that utilizes Google Cloud NLP to evaluate mental state from the input free-form text and audio data. Audio information is converted into text using speech recognition technology and further analyzed as text.
[0788] The generation method utilizes a generative AI model. Specifically, it uses Azure's generative AI via a cloud-based AI service to generate optimal support and advice based on the user's condition. The generated content includes not only skincare but also suggestions for maintaining mental health and related online resources.
[0789] The terminal displays information retrieved from the server in a format that is easy for the user to understand. The terminal's interface is designed to be easy for the user to operate and to quickly obtain the necessary advice.
[0790] For example, if a user is troubled by daily stress and dry skin, this system analyzes the user's image and evaluates their skin condition. It also measures the user's mental stress level through voice input. Based on this data, the server provides the user with specific skincare methods and meditation techniques for stress relief.
[0791] The following are examples of prompts to input into a generative AI model:
[0792] "The user's skin condition is dry, and their mental state is rated as high stress. Please generate optimal skincare advice and stress management activity suggestions for this user."
[0793] These processes enable users to receive effective support tailored to their individual needs.
[0794] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0795] Step 1:
[0796] User data entry
[0797] Users take photos of their skin using their device and send the image data to a server through the application. Users also answer questions about their emotions and mental state, inputting this information as voice or text data.
[0798] Input: Skin image data, text data, or audio data
[0799] The specific actions involve taking an image using the device's camera function and then pressing the image upload button to send the data. Regarding the state of mind, the user either answers questions verbally using the device's voice recognition function or enters text using the keyboard.
[0800] Step 2:
[0801] Server-based analysis of skin condition
[0802] The server receives image data sent by the user and analyzes it using an image processing module. Specifically, it uses the OpenCV library to calculate skin moisture, oil content, and trouble indicators. As a result, numerical data is generated.
[0803] Input: Skin image data submitted by the user
[0804] Output: Skin evaluation results (numerical data for humidity, oil content, and trouble indicators)
[0805] Specifically, a Python script runs on the server to calculate these values and form the evaluation results.
[0806] Step 3:
[0807] Analysis of mental state by server
[0808] The server passes text and audio data provided by the user to a natural language processing module. Text analysis is performed using Google Cloud NLP services, and emotions are identified using TEDAS. Audio data is converted to text using speech recognition technology, and this text data is analyzed in the same manner.
[0809] Input: Text or audio data sent by the user.
[0810] Output: Mental state evaluation results (emotion tags)
[0811] Specifically, the server converts the audio data into text, performs text analysis, classifies the sentiment, and tags the data.
[0812] Step 4:
[0813] Advice generation using generative AI models
[0814] The server uses collected data on skin condition and mental state to create prompts for a generative AI model and generate optimal advice. Using Azure's generative AI, it provides specific skincare methods and suggestions for improving mental well-being.
[0815] Input: Skin evaluation results, mental state evaluation results
[0816] Output: Optimized advice for the user
[0817] In terms of operation, the AI model on the server takes the prompt text as input, generates optimized advice, and returns it.
[0818] Step 5:
[0819] Advice and results displayed on the device
[0820] The device receives optimized advice and analysis results from the server and displays them to the user. This includes care instructions, links to resources, and community invitations.
[0821] Input: Generation advice and analysis results from the server.
[0822] Output: Display and notification to the user
[0823] Specifically, the device uses push notifications to inform the user of the arrival of new advice and displays detailed information within the app.
[0824] (Application Example 2)
[0825] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0826] Modern consumers face various skin and mental health issues due to daily stress and environmental changes. However, conventional skincare products and mental health support often only offer generalized advice. Therefore, it is difficult to provide highly personalized care for each individual. Furthermore, customers who use physical stores are not adequately provided with guidance on relaxation appropriate to the store environment or mental support through interaction with the community. Addressing these challenges is necessary to improve the customer experience and quality of life.
[0827] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0828] In this invention, the server includes image processing means for analyzing image data collected from the user and evaluating the condition of the skin; natural language processing means for analyzing voice data and text data input from the user and evaluating emotions and mental state; and generation means for generating advice, including care methods and emotional stabilization activities optimized for the user, based on the analyzed data, and for guiding relaxation activities applicable in the store environment. This enables personalized care and emotional support for individual users, improving the customer experience when using the store and enhancing the quality of life.
[0829] "Image processing means" refers to a processing device that analyzes image data provided by the user and evaluates the condition of the skin.
[0830] "Natural language processing means" refers to a processing device that analyzes voice data and text data input by a user and evaluates emotions and mental states.
[0831] The "generation means" is a processing device that generates care methods optimized for the user and advice including emotional stabilization activities based on analyzed data, and guides users through relaxation activities applicable to the store environment.
[0832] A "community matching means" is a processing device that identifies other users in similar states based on the user's state and facilitates anonymous interaction.
[0833] This invention is a system that analyzes the user's skin and emotional state and provides optimal care and emotional support. The system mainly consists of three components: a server, a terminal, and the user.
[0834] The server analyzes image data received from the user using image processing tools and evaluates the skin condition using software such as TensorFlow and OpenCV. Furthermore, it performs natural language processing on voice and text data input from the user using the Google Cloud Natural Language API and evaluates emotions and mental state using the Emotion API. Based on the data obtained in this way, a generation tool generates advice, including care methods and emotional stabilization activities optimized for the user, and sends the results to the terminal.
[0835] The device displays received advice to the user, presenting it visually and clearly using a smartphone or smart glasses. This allows users to implement specific care methods tailored to their own condition. For example, customers can receive detailed information about dedicated relaxation spaces and activities offered within the physical store.
[0836] Through this system, users can gain a comprehensive understanding of their skin and mental state, and incorporate personalized advice into their daily lives. Furthermore, the system also has a function that facilitates anonymous communication between users with similar conditions, allowing them to receive mental health support.
[0837] A concrete example is a case where, based on a user's facial image taken in a store and input about their mental state, the server evaluates their skin's moisture level and stress level, and then provides a list of recommended skincare products and a meditation program for stress reduction. An example of a prompt used in this case would be, "Please analyze my skin condition and mental stress level today and tell me the best skincare and relaxation methods for me."
[0838] In this way, through the invention, users can receive integrated care tailored to their individual needs, and as a result, their quality of life can be improved.
[0839] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0840] Step 1:
[0841] The user takes an image of their face using a device and inputs audio or text data. This input data is sent to the server. The inputs include image data, audio data, and text data, which provide the information that forms the basis for analysis in the next step.
[0842] Step 2:
[0843] The server analyzes the received image data using image processing tools. Specifically, it uses OpenCV to detect the location and condition of skin problems at the pixel level, and then uses a model trained with TensorFlow to calculate and output indicators of skin moisture, oil content, and problems. This output is then passed on to the next step as information to evaluate the user's skin condition.
[0844] Step 3:
[0845] The server passes the voice or text data input from the user to a natural language processing system. Here, the voice data is converted to text using the Google Cloud Natural Language API, and emotion analysis is performed using the Emotion API. The user's tone of voice and text content are analyzed as input, and emotion data and indicators showing the user's emotional state are generated as output. This provides quantitative data on the user's emotional state.
[0846] Step 4:
[0847] Based on analyzed skin condition and emotional data, the server uses a generation mechanism to generate care methods and advice tailored to the user. Specifically, an AI model evaluates this data and lists appropriate skincare products and relaxation programs. It also includes guidance on relaxation activities applicable to the store environment. This output provides specific care advice that the user can implement.
[0848] Step 5:
[0849] The generated advice is sent to the device and displayed to the user. The device uses Flutter or other native app development environments as its frontend to present the advice in a visually easy-to-understand manner. Users can use this as a reference to engage in their daily care. Specifically, features such as recommended product information and the ability to book in-store activities are provided.
[0850] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0851] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0852] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0853] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0854] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0855] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0856] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0857] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0858] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0859] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0860] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0861] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0862] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0863] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0864] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0865] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0866] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0867] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0868] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0869] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0870] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0871] The following is further disclosed regarding the embodiments described above.
[0872] (Claim 1)
[0873] An image processing module that analyzes image data collected from users and evaluates skin condition,
[0874] A natural language processing module that analyzes text data entered by the user and evaluates their mental state,
[0875] A generation module that generates user-optimized advice and provides relevant resources based on the analyzed data,
[0876] A community matching module that identifies other users in similar states based on the user's state and facilitates their participation in the community,
[0877] A system that includes this.
[0878] (Claim 2)
[0879] The system according to claim 1, which uses a secure protocol for safely sending and receiving data while maintaining user anonymity.
[0880] (Claim 3)
[0881] The system according to claim 1, comprising a feedback tracking function for tracking the adoption status of user advice and for improving subsequent suggestions.
[0882] "Example 1"
[0883] (Claim 1)
[0884] An image processing means that analyzes visual data collected from users and evaluates the properties of a material,
[0885] A natural language processing method that analyzes text data entered by a user and evaluates their mental state,
[0886] A generation means that generates personalized advice for the user and provides related materials based on the analyzed data,
[0887] A matching mechanism that identifies other users with similar statuses based on the user's status and facilitates their participation in a group,
[0888] A terminal means for displaying visualized information,
[0889] A system that includes this.
[0890] (Claim 2)
[0891] The system according to claim 1, which uses a protected communication specification for securely sending and receiving data while maintaining the anonymity of users.
[0892] (Claim 3)
[0893] The system according to claim 1, which tracks the adoption status of user advice and includes a function for collecting feedback to further improve the proposal.
[0894] "Application Example 1"
[0895] (Claim 1)
[0896] A data processing method that analyzes image data collected from users and evaluates the condition of the skin,
[0897] A natural language processing method that analyzes text data entered by a user and evaluates their mental state,
[0898] A generation means that generates optimized advice for the user and provides related information based on the analyzed data,
[0899] A group matching mechanism that identifies other users in similar states based on the user's state and facilitates their participation as collaborators,
[0900] A store support system that provides immediate evaluation results and product suggestions to users in physical stores,
[0901] A system that includes this.
[0902] (Claim 2)
[0903] The system according to claim 1, which uses a secure protocol for securely sending and receiving data while maintaining user anonymity.
[0904] (Claim 3)
[0905] The system according to claim 1, comprising a feedback tracking function for tracking the adoption status of user suggestions and for improving subsequent suggestions.
[0906] "Example 2 of combining an emotion engine"
[0907] (Claim 1)
[0908] Image processing means for analyzing image information collected from users and evaluating surface conditions,
[0909] A natural information processing means that analyzes textual and audio information input by the user and evaluates the user's mental state,
[0910] A generation means that generates user-optimized support and provides related resources based on the analyzed information,
[0911] A means of promoting participation in the community that identifies other users in similar states based on the user's state and encourages their participation,
[0912] A system that includes this.
[0913] (Claim 2)
[0914] The system according to claim 1, which uses a secure communication means for safely sending and receiving information while maintaining the anonymity of the user.
[0915] (Claim 3)
[0916] The system according to claim 1, comprising means for tracking the user's support adoption status and collecting feedback to improve subsequent proposals.
[0917] "Application example 2 when combining with an emotional engine"
[0918] (Claim 1)
[0919] Image processing means for analyzing image data collected from users and evaluating skin condition,
[0920] A natural language processing means that analyzes voice and text data input from the user and evaluates emotions and mental state,
[0921] Based on the analyzed data, a means for generating advice including care methods and emotional stabilization activities optimized for the user, and for guiding relaxation activities applicable in the store environment,
[0922] A community matching system that identifies other users in similar situations based on the user's status and promotes participation in the community, including anonymous interaction.
[0923] A system that includes this.
[0924] (Claim 2)
[0925] The system according to claim 1, which uses a protocol for securely sending and receiving data while maintaining user anonymity.
[0926] (Claim 3)
[0927] The system according to claim 1, which tracks the adoption of user advice, has tracking capabilities to improve subsequent suggestions, and further enables feedback on the customer experience in stores. [Explanation of Symbols]
[0928] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. An image processing module that analyzes image data collected from users and evaluates skin condition, A natural language processing module that analyzes text data entered by the user and evaluates their mental state, A generation module that generates user-optimized advice and provides relevant resources based on the analyzed data, A community matching module that identifies other users in similar states based on the user's state and facilitates their participation in the community, A system that includes this.
2. The system according to claim 1, which uses a secure protocol for safely sending and receiving data while maintaining user anonymity.
3. The system according to claim 1, comprising a feedback tracking function for tracking the adoption status of user advice and for improving subsequent suggestions.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A