System

A system analyzing facial and voice data provides personalized health and fashion advice to enhance communication skills and health management.

JP2026036041APending Publication Date: 2026-03-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024138556
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Modern society faces challenges in objectively understanding one's mental and physical state, with limited time and resources for professional advice and training, necessitating a support system for effective communication skills and self-care.

Method used

A system that acquires and analyzes facial expression and voice data to evaluate mental and physical health, providing appropriate health advice and self-care recommendations, including fashion advice when needed.

Benefits of technology

Enhances users' ability to manage their health and improve communication skills through personalized advice based on comprehensive health evaluations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026036041000001_ABST
    Figure 2026036041000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes a means for acquiring expression data, a means for acquiring voice data, a means for analyzing the acquired expression data and voice data to evaluate the mental and physical health condition of the user, and a means for providing the user with appropriate health advice based on the evaluation result.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, effective communication skills are essential for living a vibrant and fulfilling life. However, it is difficult to objectively understand one's own mental and physical state and to implement appropriate self-care and communication training. In particular, busy daily lives mean that time and resources for professional advice and training are limited. Given this situation, there is a demand for a support system that allows users to properly understand their own state and improve their communication skills in their daily lives. [Means for solving the problem]

[0005] The present invention provides a system that acquires a user's facial expression data and voice data, analyzes the data, evaluates the user's mental and physical health status, and provides appropriate health advice and self-care recommendations.

[0006] 1. How to obtain facial expression data

[0007] 2. How to obtain audio data

[0008] 3. A means of analyzing acquired facial expression and voice data to evaluate the user's mental and physical health status.

[0009] 4. Means of providing appropriate health advice to users based on the evaluation results

[0010] The system also provides a means for providing fashion advice and self-care recommendations to users based on the evaluation results, thereby supporting the improvement of users' communication skills and self-management abilities. This system allows users to effectively acquire communication skills and manage their health based on specialized knowledge and methods.

[0011] "Facial expression data" refers to digital data obtained by capturing a user's facial expressions and used to assess the user's emotional and health states.

[0012] "Voice data" refers to digital data obtained by capturing the characteristics and content of a user's voice and is used to assess the user's emotional and health states.

[0013] "Analysis" refers to the use of algorithms to assess the user's mental and physical health based on acquired facial expression and voice data.

[0014] "Health advice" refers to specific guidelines and recommendations for action provided to users based on the user's mental and physical health status obtained through analysis.

[0015] "Self-care" refers to specific actions and care that users take to manage and improve their own health.

[0016] "Fashion advice" refers to recommendations about clothing and fashion provided to a user based on the user's health condition, current weather information, the user's schedule, and the like.

[0017] "User" means a person who uses the system to assess their own mental and physical health and receive advice and training. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9]1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The present invention is a system that acquires and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status and provide appropriate health advice and self-care recommendations. Here, we will explain in detail the operation of the entire system.

[0040] System Configuration

[0041] This system is broadly composed of three elements: the server, the terminal (e.g., the user's smartphone), and the user. Each element and its role are explained below.

[0042] server

[0043] The server is the core of the system and performs the following functions:

[0044] 1. Analyze the user's facial expression and voice data.

[0045] 2. Evaluate the user's mental and physical health based on the analysis results.

[0046] 3. Generate health advice and self-care recommendations to provide to users.

[0047] 4. Offer fashion advice when needed.

[0048] Terminal

[0049] The terminal provides the user interface and has the following roles:

[0050] 1. Use a camera to capture facial expression data.

[0051] 2. Use a microphone to capture audio data.

[0052] 3. The server notifies the user of the analysis results and advice.

[0053] 4. Receives input from the user and sends it to the server.

[0054] User

[0055] Users use the system to receive assessments and advice on their health status.

[0056] 1. Enter basic information (age, gender, preferences, daily activity patterns, etc.) as the initial setting.

[0057] 2. Use a camera and microphone to provide facial and voice data.

[0058] 3. Follow the health advice and self-care recommendations provided by the server.

[0059] Specific examples of processing

[0060] Here we will explain a specific scenario in which the system actually works.

[0061] 1. Initial Setup

[0062] The user downloads the smartphone app and enters basic information.

[0063] The terminal sends this information to the server.

[0064] The server stores the basic information in a database for future analysis.

[0065] 2. Acquisition and analysis of facial and speech data

[0066] A user takes a selfie in front of the smartphone camera and also captures audio data.

[0067] The terminal transmits facial expression data and voice data to the server.

[0068] The server analyzes the data and assesses the user's mental and physical health.

[0069] The server generates appropriate health advice based on the evaluation results.

[0070] 3. Providing advice

[0071] The server generates health advice and self-care recommendations and sends them to the device.

[0072] The terminal notifies the user of the advice.

[0073] The user checks the advice provided and takes the necessary action.

[0074] 4. Providing fashion advice

[0075] The server generates fashion advice based on weather information and the user's schedule.

[0076] The terminal notifies the user of this.

[0077] The user follows the advice and chooses appropriate clothing.

[0078] In this way, this system effectively acquires and analyzes facial expression and voice data and provides advice based on the results to help users improve their communication skills and health management in their daily lives.

[0079] The processing flow will be explained below.

[0080] Step 1:

[0081] A user downloads and installs a smartphone app.

[0082] Step 2:

[0083] The user launches the app and enters basic information such as age, gender, preferences, and daily activity patterns on the basic settings screen.

[0084] Step 3:

[0085] The terminal sends the entered basic information to the server.

[0086] Step 4:

[0087] The server stores the received basic information in a database for future analysis.

[0088] Step 5:

[0089] Users take selfies in front of their smartphone cameras on a daily basis.

[0090] Step 6:

[0091] The device sends the selfie (facial expression data) to the server.

[0092] Step 7:

[0093] A user speaks into a smartphone to capture audio data.

[0094] Step 8:

[0095] The terminal transmits the acquired voice data to the server.

[0096] Step 9:

[0097] The server analyzes the received facial expression data to assess the user's emotional state, for example, detecting the degree of smile or sadness.

[0098] Step 10:

[0099] The server analyzes the received audio data and evaluates the tone and patterns of the voice, such as the pitch, rate, and emotional expression.

[0100] Step 11:

[0101] The server integrates the analysis results and assesses the user's mental and physical health.

[0102] Step 12:

[0103] The server generates specific health advice based on the assessment results, for example, "You seem a little tired today. I recommend you take a break."

[0104] Step 13:

[0105] The server sends the generated health advice to the terminal.

[0106] Step 14:

[0107] The terminal notifies the user of the health advice received from the server.

[0108] Step 15:

[0109] The user checks the advice provided and takes action as necessary.

[0110] Step 16:

[0111] Users take selfies periodically to provide data for checking the condition of their hair and beard.

[0112] Step 17:

[0113] The device sends the captured selfie data to the server.

[0114] Step 18:

[0115] The server analyzes the selfie data and determines whether hair or beard care is needed.

[0116] Step 19:

[0117] The server generates recommendations for beauty salons as needed and provides them to the user.

[0118] Step 20:

[0119] The device notifies the user of beauty salon recommendations.

[0120] Step 21:

[0121] The user makes a reservation at the recommended hair salon via their smartphone.

[0122] Step 22:

[0123] The server generates fashion advice based on season, weather, and schedule information.

[0124] Step 23:

[0125] The server transmits the generated fashion advice to the terminal.

[0126] Step 24:

[0127] The terminal notifies the user of fashion advice.

[0128] Step 25:

[0129] The user checks the provided fashion advice and selects an outfit.

[0130] Step 26:

[0131] The server evaluates the user's progress and training effectiveness in real time and updates the next training menu.

[0132] Step 27:

[0133] The device will notify and guide the user about the next training menu.

[0134] Step 28:

[0135] Users follow a training menu to effectively improve their communication skills.

[0136] In this way, each step of the system works in conjunction with one another to provide comprehensive support for the user's health management and improvement of communication skills.

[0137] Example 1

[0138] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0139] Conventional health management systems lack sufficient means for comprehensively evaluating a user's mental and physical health status, and it is difficult to efficiently provide personalized advice to each user. Furthermore, there are no systems that analyze a user's facial expressions and voice data to appropriately provide self-care recommendations or fashion advice. Therefore, there is a need to improve users' health management and comfort in their daily lives.

[0140] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0141] In this invention, the server includes means for acquiring facial expression data, means for acquiring voice data, means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health status, means for generating appropriate health advice for the user based on the evaluation results, and means for generating prompt sentences using a generative AI model to provide personalized advice to the user. This enables a comprehensive evaluation of the user's mental and physical health status and the provision of appropriate and personalized advice to each user. Furthermore, by recommending self-care and providing fashion advice as needed, the server achieves health management and improves the user's quality of life.

[0142] "Facial expression data" refers to data that digitally records a user's facial expressions.

[0143] "Voice data" refers to data that is a digital recording of a user's speech or voice.

[0144] "Analysis" refers to the process of using acquired digital data to identify features and patterns and evaluate the user's mental and physical state.

[0145] "Evaluating health status" means quantitatively and qualitatively determining the user's psychological and physical health status based on the analysis results of facial expression data and voice data.

[0146] "Generating health advice" means creating appropriate courses of action or recommendations based on the assessed health status of the user.

[0147] A "generative AI model" is an algorithm or system that uses artificial intelligence to generate text or instructions.

[0148] A "prompt sentence" is text data input into a generative AI model that describes the instructions and information needed to perform a specific task.

[0149] "Personalized advice" refers to specific and appropriate advice that is customized according to the characteristics and conditions of each individual user.

[0150] The present invention is a system that acquires and analyzes a user's facial expression and voice data to evaluate the user's mental and physical health and provide appropriate health advice and self-care recommendations. The invention has three main components: a server, a terminal, and a user.

[0151] System Configuration

[0152] server

[0153] The server is the core of the system and is responsible for:

[0154] 1. Receiving and storing data: Receives facial expression data and voice data sent from the user's device and stores them in a database.

[0155] 2. Data analysis: The received data is analyzed. Facial expression data is extracted using OpenCV, and voice data is converted to text using the Google® Cloud Speech-to-Text API, followed by sentiment analysis using the spaCy library.

[0156] 3. Health assessment: Based on the analysis results, the user's mental and physical health status is assessed.

[0157] 4. Advice Generation: A generative AI model (e.g., GPT-3®) is used to generate prompts based on the assessment results, creating appropriate health advice and self-care recommendations, and providing fashion advice as needed.

[0158] Specific operation example

[0159] Initial setup and data acquisition

[0160] Users download the smartphone app and enter basic information (age, gender, preferences, daily activity patterns, etc.).

[0161] The terminal sends this information to the server, which stores it in a database.

[0162] Acquisition and analysis of facial expression and voice data

[0163] At designated times, the user takes a photo of their facial expression using the smartphone camera and records their voice using the microphone.

[0164] The device collects facial expression and voice data and sends them to the server in an appropriate format (e.g., JPEG images and WAV audio files).

[0165] The server analyzes the received data, extracting facial features using OpenCV, converting the audio data to text using the Google Cloud Speech-to-Text API, and performing emotion analysis using the spaCy library.

[0166] Health assessment and advice

[0167] Based on the analysis results, the server evaluates the user's mental and physical health.

[0168] The server sends the generated prompts to a generative AI model to create personalized health advice and self-care recommendations.

[0169] The server sends these advices to the terminal, which then notifies the user.

[0170] Specific prompt examples

[0171] “Analyze the user’s voice data and assess their emotional state. For example, if the user is tired, suggest a ‘relaxation exercise’.”

[0172] In this way, this system supports users in improving their communication skills and health management in their daily lives. By effectively analyzing facial expression and voice data and providing personalized advice based on the results, the system aims to comprehensively maintain and improve the user's health.

[0173] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0174] Step 1: Initial settings and user information registration

[0175] A user downloads and installs a smartphone app. As input, the user enters basic information (age, gender, preferences, daily activity patterns, etc.) into the app. The device collects this information and sends it to the server in JSON format. The server stores the received information in a database and uses it for future analysis. This allows user information to be accumulated in the database, making it possible to provide personalized services to each individual user.

[0176] Specific behavior:

[0177] A user installs the app on their smartphone, launches it, and enters some basic information. The device then makes an HTTP request to send this information to the server, which then stores it in a database.

[0178] Step 2: Acquire facial expression and voice data

[0179] At designated times, the user captures facial expression data using the smartphone camera and records audio data using the microphone. Camera image data and audio data are acquired as input. The device collects this data and sends it to the server as JPEG images and WAV audio files. The output is facial expression data and audio data.

[0180] Specific behavior:

[0181] When a user presses the "Get Data" button, the smartphone camera is activated and a picture of their face is taken. Then, a voice memo is recorded using the microphone. The device collects facial expression data in JPEG format and voice data in WAV format, and creates an HTTP request to send this to the server.

[0182] Step 3: Data analysis by the server

[0183] The server analyzes the received facial expression data using OpenCV and extracts feature points. It converts the voice data into text using the Google Cloud Speech-to-Text API and performs emotion analysis using an NLP library (e.g., spaCy). It receives facial expression data and voice data as input and obtains feature point data and emotion analysis results as output.

[0184] Specific behavior:

[0185] The server processes the received image data using the OpenCV library to extract facial features, converts the audio data into text using the Google Cloud Speech-to-Text API, and analyzes the resulting text data using the spaCy library to evaluate the emotional state. The results are then stored in a database.

[0186] Step 4: Health Assessment

[0187] The server evaluates the user's mental and physical health based on the analyzed facial expression and voice data. It uses feature point data and emotion analysis results as input and generates a health status evaluation result as output.

[0188] Specific behavior:

[0189] The server evaluates the user's health condition using a dedicated evaluation algorithm based on the analysis results. The evaluation results are stored in a database and used in the next step of generating advice.

[0190] Step 5: Generating and serving advice

[0191] The server sends prompts to a generative AI model (e.g., GPT-3) to generate appropriate health advice and self-care recommendations based on the assessment results. Fashion advice is also provided if necessary. The health assessment results are used as input, and the generated advice is obtained as output. The device notifies the user of this.

[0192] Specific behavior:

[0193] The server sends a prompt message to the AI ​​model saying, "The user's stress level is high. Please provide relaxation advice." The AI ​​model then collects the advice returned, converts it into JSON format, and sends it to the device. The device then generates a notification and displays it for the user to see.

[0194] The above is the specific processing flow of this system.

[0195] (Application example 1)

[0196] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0197] In traditional brick-and-mortar stores, it was difficult to grasp the mental and physical health status of customers and provide appropriate services and advice based on that. As a result, the quality of customer experience tended to be uniform, making it difficult to provide services that meet individual needs. This led to problems such as lower customer satisfaction and difficulty in securing repeat customers.

[0198] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0199] In this invention, the server includes means for acquiring facial expression data, means for acquiring voice data, means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health state, and means for providing the user with appropriate health advice based on the evaluation results. This makes it possible to grasp the mental and physical health state of customers in a physical store and provide appropriate health advice and self-care recommendations in real time.

[0200] "Facial expression data" is digital data about a person's facial expressions collected using a camera or other image capture device.

[0201] "Voice data" is digital data relating to a human voice collected using a microphone or other voice capture device.

[0202] "Analysis" refers to the act of processing acquired facial expression and voice data and applying computational processes and algorithms to understand, classify, and evaluate its content.

[0203] "Mental and physical health status" refers to the totality of information that indicates the user's psychological state and physical condition, including emotions, stress level, fatigue level, and the like.

[0204] "Health Advice" means specific recommendations or advice based on mental and physical health that are provided to users to help them maintain or improve their health in their daily lives.

[0205] "Self-care" refers to actions and efforts for health management that users themselves undertake, and includes relaxation techniques, reviewing diet and drink habits, exercise instruction, and the like.

[0206] "Fashion advice" refers to specific recommendations and suggestions to help users select appropriate clothing, accessories, etc. based on their evaluation results and circumstances.

[0207] "Real-time" refers to data acquisition, analysis, evaluation, and feedback occurring nearly simultaneously and without delay.

[0208] The present invention is a system that acquires and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status and provide appropriate health advice and self-care recommendations. Here, we will show a specific application example in a brick-and-mortar store and explain in detail an embodiment of the present invention.

[0209] System Configuration

[0210] This system is broadly composed of three elements: the server, the terminal (e.g., the user's smartphone), and the user. Each element and its role are explained below.

[0211] server

[0212] The server is the core of the system and performs the following functions:

[0213] 1. Analyze the user's facial expression and voice data.

[0214] The hardware and software used are TENSORFLOW (registered trademark) / Keras.

[0215] 2. Evaluate the user's mental and physical health based on the analysis results.

[0216] 3. Generate health advice and self-care recommendations to provide to users.

[0217] A specific example would be to generate the advice "Listen to some relaxing music."

[0218] 4. Offer fashion advice when needed.

[0219] As an example of use, it generates advice based on weather information, such as "It's going to rain today, so please bring a waterproof jacket."

[0220] Terminal

[0221] The terminal provides the user interface and has the following roles:

[0222] 1. Use a camera to capture facial expression data.

[0223] The software used is OpenCV.

[0224] 2. Use a microphone to capture audio data.

[0225] The software used is librosa.

[0226] 3. The server notifies the user of the analysis results and advice.

[0227] To provide audio advice, gTTS (Google Text-to-Speech) is used and played on playsound.

[0228] 4. Receives input from the user and sends it to the server.

[0229] User

[0230] Users use the system to receive assessments and advice on their health status.

[0231] 1. Enter basic information (age, gender, preferences, daily activity patterns, etc.) as the initial setting.

[0232] 2. Use a camera and microphone to provide facial and voice data.

[0233] 3. Follow the health advice and self-care recommendations provided by the server.

[0234] Specific examples

[0235] Here we will explain a specific scenario in which the system actually works.

[0236] 1. Initial Setup

[0237] The user downloads the smartphone app and enters basic information.

[0238] The terminal sends this information to the server.

[0239] The server stores the basic information in a database for future analysis.

[0240] 2. Acquisition and analysis of facial and speech data

[0241] A user takes a selfie in front of the smartphone camera and also captures audio data.

[0242] The terminal transmits facial expression data and voice data to the server.

[0243] The server analyzes the data and assesses the user's mental and physical health.

[0244] The server generates appropriate health advice based on the evaluation results.

[0245] 3. Providing advice

[0246] The server generates health advice and self-care recommendations and sends them to the device.

[0247] The terminal notifies the user of the advice.

[0248] The user checks the advice provided and takes the necessary action.

[0249] 4. Providing fashion advice

[0250] The server generates fashion advice based on weather information and the user's schedule.

[0251] The terminal notifies the user of this.

[0252] The user follows the advice and chooses appropriate clothing.

[0253] Prompt Sentence Examples

[0254] Prompt: "Analyze the customer's facial and voice data to assess their current mental and physical state. Based on the assessment, generate effective relaxation advice."

[0255] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0256] Step 1:

[0257] The user launches the smartphone app and enters basic information (age, gender, preferences, daily activity patterns, etc.).

[0258] Input: User's basic information (age, gender, preferences, daily activity patterns, etc.)

[0259] Output: The basic information entered is saved on the device.

[0260] Step 2:

[0261] The terminal sends the user's basic information to the server.

[0262] Input: User basic information

[0263] Output: Basic information is sent to the server.

[0264] Step 3:

[0265] The server stores the received basic information in a database.

[0266] Input: Basic information data

[0267] Output: Basic information stored in the database

[0268] Step 4:

[0269] A user takes a selfie in front of the smartphone camera and captures audio data.

[0270] Input: Facial image data and audio data

[0271] Output: The captured facial image data and recorded audio data are saved on the device.

[0272] Step 5:

[0273] The terminal transmits facial expression data and voice data to the server.

[0274] Input: Facial image data and audio data

[0275] Output: Facial expression data and voice data are sent to the server.

[0276] Step 6:

[0277] The server parses the received data.

[0278] Facial expression data: Preprocessed using OpenCV, analyzed using a Keras model, and output emotion categories (e.g., "happy," "sad," "surprised," etc.).

[0279] Input: Face image data

[0280] Data processing: Preprocessing with OpenCV (grayscale conversion, resizing, etc.)

[0281] Data Computation: Sentiment Analysis with Keras Models

[0282] Output: Emotion category

[0283] Audio data: Preprocessing is performed using librosa, features are extracted, and the data is analyzed using a Keras model. The emotional category is output.

[0284] Input: Audio data

[0285] Data processing: Feature extraction (MFCC, etc.) using librosa

[0286] Data Computation: Sentiment Analysis with Keras Models

[0287] Output: Emotion category

[0288] Step 7:

[0289] The server evaluates the user's mental and physical health based on the analysis results and generates appropriate health advice.

[0290] Input: Emotion category (facial and vocal results)

[0291] Output: Health advice

[0292] Step 8:

[0293] The server sends the generated health advice to the terminal.

[0294] Enter: Health Advice

[0295] Output: Advice data is sent to the terminal.

[0296] Step 9:

[0297] The terminal notifies the user of the generated health advice.

[0298] Input: Advice data

[0299] Output: Advice notification display, audio playback (using gTTS and playsound)

[0300] Step 10:

[0301] The user checks the health advice provided and takes necessary action.

[0302] Input: Advice content

[0303] Output: User actions (taking a break, taking a deep breath, practicing relaxation techniques, etc.)

[0304] Step 11:

[0305] The server generates fashion advice based on weather information and the user's schedule.

[0306] Input: Weather information, user schedule

[0307] Output: Fashion advice

[0308] Step 12:

[0309] The terminal notifies the user of the generated fashion advice.

[0310] Enter: fashion advice

[0311] Output: Fashion advice notification display

[0312] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0313] The present invention is a system that acquires and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status, and provides appropriate health advice and self-care recommendations, in addition to providing an emotion engine that recognizes the user's emotions. The operation of the entire system is described in detail below.

[0314] System Configuration

[0315] This system is broadly composed of three elements: the server, the terminal (user's smartphone), and the user. Each element and its role are explained below.

[0316] server

[0317] The server is the core of the system and performs the following functions:

[0318] 1. Analyze the user's facial expression and voice data.

[0319] 2. Evaluate the user's mental and physical health based on the analysis results.

[0320] 3. Generate health advice and self-care recommendations to provide to users.

[0321] 4. Offer fashion advice when needed.

[0322] 5. Analyze user sentiment in real time using an emotion engine.

[0323] 6. Provide appropriate emotional feedback based on the results of emotion analysis.

[0324] Terminal

[0325] The terminal provides the user interface and has the following roles:

[0326] 1. Use a camera to capture facial expression data.

[0327] 2. Use a microphone to capture audio data.

[0328] 3. The server notifies the user of the analysis results and advice.

[0329] 4. Receives input from the user and sends it to the server.

[0330] User

[0331] Users use the system to receive assessments and advice on their health status.

[0332] 1. Enter basic information (age, gender, preferences, daily activity patterns, etc.) as the initial setting.

[0333] 2. Use a camera and microphone to provide facial and voice data.

[0334] 3. Follow the health advice and self-care recommendations provided by the server.

[0335] Specific examples of processing

[0336] Here we will explain a specific scenario in which the system actually works.

[0337] 1. Initial Setup

[0338] The user downloads the smartphone app and enters basic information.

[0339] The terminal sends this information to the server.

[0340] The server stores the basic information in a database for future analysis.

[0341] 2. Acquisition and analysis of facial and speech data

[0342] Users take selfies in front of their smartphone cameras on a daily basis.

[0343] The device sends the selfie (facial expression data) to the server.

[0344] A user speaks into a smartphone to capture audio data.

[0345] The terminal transmits the acquired voice data to the server.

[0346] 3. Emotion analysis using an emotion engine

[0347] The server uses the received facial expression data and voice data to analyze the user's emotions in real time using an emotion engine.

[0348] The server evaluates the user's emotional state based on the emotion analysis results, detecting emotions such as joy, sadness, and anger.

[0349] The server generates appropriate emotional feedback based on the emotion analysis results, e.g., "You seem a little tired today. Please stretch to relax."

[0350] 4. Health advice and self-care recommendations

[0351] The server integrates the analysis results of the facial expression data and voice data with the emotion evaluation results to comprehensively evaluate the user's mental and physical health.

[0352] The server generates specific health advice and self-care recommendations based on the assessment results.

[0353] The server generates advice and recommendations and sends them to the device.

[0354] The terminal notifies the user of the advice.

[0355] The user checks the advice provided and takes the necessary action.

[0356] 5. Providing fashion advice

[0357] The server generates fashion advice based on season, weather, and schedule information.

[0358] The terminal notifies the user of the generated fashion advice.

[0359] The user follows the advice and chooses appropriate clothing.

[0360] 6. Training progress and feedback

[0361] The server evaluates the user's progress and training effectiveness in real time and updates the next training menu.

[0362] The device will notify and guide the user about the next training menu.

[0363] Users follow a training menu to effectively improve their communication skills.

[0364] In this way, each step of the system works in conjunction to provide comprehensive support for users' health management and improvement of communication skills. By combining the emotion engine, it is possible to evaluate mental and physical aspects with higher accuracy than before, and to provide more personalized advice and recommendations.

[0365] The processing flow will be explained below.

[0366] Step 1:

[0367] A user downloads and installs a smartphone app.

[0368] Step 2:

[0369] The user launches the app and enters basic information such as age, gender, preferences, and daily activity patterns on the basic settings screen.

[0370] Step 3:

[0371] The terminal sends the entered basic information to the server.

[0372] Step 4:

[0373] The server stores the received basic information in a database for future analysis.

[0374] Step 5:

[0375] Users take selfies in front of their smartphone cameras on a daily basis.

[0376] Step 6:

[0377] The device sends the selfie (facial expression data) to the server.

[0378] Step 7:

[0379] A user speaks into a smartphone to capture audio data.

[0380] Step 8:

[0381] The terminal transmits the acquired voice data to the server.

[0382] Step 9:

[0383] The server analyzes the received facial expression data using a facial recognition algorithm and emotion engine to assess the user's emotional state, specifically identifying emotions such as smile, sadness, anger, and surprise.

[0384] Step 10:

[0385] The voice data received by the server is also analyzed by the emotion engine to evaluate the tone, speed, and emotional expression of the voice, e.g., excited, calm, angry, etc.

[0386] Step 11:

[0387] The server combines the analysis results of facial expression data and voice data to perform an emotional evaluation, and also comprehensively evaluates the user's mental and physical health.

[0388] Step 12:

[0389] The server generates specific health advice based on the evaluation results, for example, "You seem a little tired today. Please stretch to relax."

[0390] Step 13:

[0391] The server sends the generated health advice and self-care recommendations to the terminal.

[0392] Step 14:

[0393] The terminal notifies the user of the health advice received from the server.

[0394] Step 15:

[0395] The user checks the advice provided and takes the necessary action.

[0396] Step 16:

[0397] Users take selfies periodically to provide data for checking the condition of their hair and beard.

[0398] Step 17:

[0399] The device sends the captured selfie data to the server.

[0400] Step 18:

[0401] The server analyzes the selfie data and determines if you need hair or beard care. Example: "I need a beard trim."

[0402] Step 19:

[0403] The server generates salon recommendations as needed and lists nearby salons.

[0404] Step 20:

[0405] The device notifies the user of recommended beauty salons. For example, "Here are some recommended beauty salons near you."

[0406] Step 21:

[0407] The user makes a reservation at the recommended hair salon via their smartphone.

[0408] Step 22:

[0409] The server generates fashion advice based on season, weather, and schedule information. Example: "It's cold today, so I recommend a thick coat."

[0410] Step 23:

[0411] The terminal notifies the user of the generated fashion advice.

[0412] Step 24:

[0413] The user follows the advice and chooses appropriate clothing.

[0414] Step 25:

[0415] The server evaluates the user's progress and training effectiveness in real time and updates the next training menu.

[0416] Step 26:

[0417] The device will notify and guide the user about the next training menu.

[0418] Step 27:

[0419] Users follow a training menu to effectively improve their communication skills.

[0420] In this way, each step of the system works in conjunction to provide comprehensive support for users' health management and improvement of communication skills. By combining the emotion engine, it is possible to evaluate mental and physical aspects with higher accuracy than before, and to provide more personalized advice and recommendations.

[0421] Example 2

[0422] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0423] Conventional health assessment systems have difficulty in comprehensively assessing a user's mental and physical health status in real time, and have been unable to provide appropriate feedback quickly. Furthermore, they do not adequately provide advice that takes into account the user's emotional state, making it difficult to provide personalized care.

[0424] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0425] In this invention, the server includes means for acquiring facial expression data of a user, means for acquiring voice data of the user, means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health state, means for providing the user with appropriate health advice based on the evaluation results, means for analyzing the user's emotions in real time using an emotion analysis engine, and means for providing the user with appropriate emotional feedback based on the analysis results, thereby making it possible to comprehensively evaluate the user's health state in real time and quickly provide personalized advice.

[0426] "User" refers to an individual who uses the system and receives health assessments and advice.

[0427] "Facial expression data" is digital data obtained using a camera that contains information about the user's facial expressions.

[0428] "Voice data" is digital data obtained by capturing voice information when a user speaks using a microphone.

[0429] "Analysis" refers to the act of processing information based on acquired data to quantitatively or qualitatively evaluate the user's health and emotional state.

[0430] "Health condition" is the result of a comprehensive evaluation of the user's mental and physical conditions.

[0431] "Health advice" refers to specific guidelines or recommended actions provided to the user based on the evaluation results.

[0432] An "emotion analysis engine" is software or a program for analyzing a user's emotional state in real time based on facial expression data and voice data.

[0433] "Emotional feedback" refers to specific countermeasures and advice provided to users based on the results of the emotion analysis engine.

[0434] "Self-care recommendations" are specific advice or recommended actions for self-care that a user can take by themselves.

[0435] "Fashion advice" is advice on clothing provided to users based on information such as the season, weather, and schedule.

[0436] A "terminal" is a device used by a user, and is a device that provides a user interface, including a smartphone, tablet, etc.

[0437] The "server" is the core part of this system, and is a computer system that analyzes and stores data, generates advice, and so on.

[0438] The present invention provides a system that collects and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status, and provides appropriate health advice, self-care recommendations, and emotional feedback based on real-time emotional analysis. Specific embodiments are described below.

[0439] System Configuration

[0440] This system is broadly composed of three elements: a server, a terminal (user's smartphone), and a user.

[0441] server

[0442] The server is the core of the system and performs the following functions:

[0443] 1. Analysis of the user's facial expression data and voice data: The server processes the facial expression data and voice data sent from the device using an analysis engine (e.g., Amazon Rekognition, Google Cloud Speech-to-Text).

[0444] 2. Evaluating the user's health status: Based on the results of the analysis engine, the user's mental and physical health status is evaluated.

[0445] 3. Generating health advice and self-care recommendations: Based on the health status assessment results, appropriate health advice and self-care recommendations are generated.

[0446] 4. Real-time emotion analysis using an emotion analysis engine: Facial expression data and voice data are used to activate the emotion analysis engine, which analyzes the user's emotions in real time.

[0447] 5. Providing emotional feedback: Based on the results of emotion analysis, appropriate emotional feedback is generated and sent to the device.

[0448] Device (smartphone)

[0449] The terminal is responsible for the following:

[0450] 1. Acquiring facial expression data: The user's facial expressions are captured by a camera and saved as facial expression data.

[0451] 2. Acquiring voice data: The user's speech is recorded with a microphone and saved as voice data.

[0452] 3. Data transmission: The acquired facial expression data and voice data are transmitted to the server.

[0453] 4. Notification of advice and feedback: Notify the user of the analysis results, advice, and emotional feedback received from the server.

[0454] User

[0455] The user does the following:

[0456] 1. Download the app and enter basic information: Download the smartphone app and enter basic information such as age, gender, preferences, and daily activity patterns.

[0457] 2. Providing facial expression data and voice data: Taking selfies in front of a smartphone camera in everyday life and speaking into the smartphone provides voice data.

[0458] 3. Review and act on advice: Take action based on the health advice, self-care recommendations, and emotional feedback provided by the server.

[0459] Examples of concrete examples and prompts

[0460] Specific examples

[0461] For example, when a user wakes up in the morning, they take a selfie with their smartphone and say, "Good morning, what are your plans for today?" The device immediately sends the selfie and voice data to the server. The server receives the data and analyzes it using its emotion engine. From the analysis results, it detects that the user looks a little tired, and generates advice such as, "You seem a little tired today. Please stretch to relax." The device notifies the user of this advice, and the user confirms the notification and takes time to relax early.

[0462] Prompt Sentence Examples

[0463] "Design a system where a user provides selfies and voice data, and a server analyzes this and generates appropriate health advice. For example, if the user is tired, provide advice on relaxation. Also, use an emotion analysis engine to evaluate emotions in real time."

[0464] In this way, each step of the system works in conjunction to provide comprehensive support for the user's health management and understanding of their emotional state. By combining the emotion engine, it is possible to assess mental and physical aspects more accurately than conventional systems, and to provide more personalized advice and recommendations.

[0465] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0466] Step 1:

[0467] Download the app and enter your basic information

[0468] A user downloads and installs a smartphone app.

[0469] A user opens the app and enters basic information such as age, gender, preferences, and daily activity patterns.

[0470] Input: User basic information

[0471] The terminal sends the basic information entered by the user to the server.

[0472] Output: Basic information sent to the server

[0473] Step 2:

[0474] Save basic information

[0475] The server checks the received basic information and stores it in the database.

[0476] Input: Basic information submitted

[0477] The server reviews the stored information and prepares it for future analysis.

[0478] Output: Basic information saved

[0479] Step 3:

[0480] Acquiring facial expression data

[0481] A user takes a selfie in front of the smartphone camera in daily life.

[0482] Input: Selfie image

[0483] Your device will temporarily store the selfie image in its internal memory.

[0484] Output: Facial expression data stored in temporary memory

[0485] Step 4:

[0486] Sending facial expression data

[0487] The device sends the temporarily stored selfie data to the server.

[0488] Input: Facial expression data stored in temporary memory

[0489] The server passes the received selfie data to the analysis engine.

[0490] Output: Facial expression data passed to the analysis engine

[0491] Step 5:

[0492] Acquiring audio data

[0493] The user speaks into the smartphone and voice data is captured.

[0494] Input: Audio data

[0495] The device temporarily stores the captured audio data in its internal memory.

[0496] Output: Audio data stored in temporary memory

[0497] Step 6:

[0498] Sending audio data

[0499] The terminal transmits the temporarily stored voice data to the server.

[0500] Input: Audio data stored in temporary memory

[0501] The server passes the received audio data to the analysis engine.

[0502] Output: Audio data passed to the analysis engine

[0503] Step 7:

[0504] Analysis of facial expression and voice data

[0505] The server uses the received facial expression data and voice data to instruct the emotion analysis engine to perform analysis.

[0506] Input: Facial expression and voice data passed to the analysis engine

[0507] An emotion analysis engine analyzes the data to detect emotional states such as joy, sadness, and anger.

[0508] Output: Detected emotional state

[0509] Step 8:

[0510] Generating Emotional Feedback

[0511] The server evaluates the user's emotional state based on the analysis results of the emotion engine and generates appropriate emotional feedback.

[0512] Input: Detected emotional state

[0513] The server generates emotional feedback and sends it to the device.

[0514] Output: Emotional feedback sent to the device

[0515] Step 9:

[0516] Feedback Notification

[0517] The terminal notifies the user of the received feedback.

[0518] Input: Emotional feedback sent to the device

[0519] The user checks the notification and takes the necessary action, such as stretching to relax.

[0520] Output: User takes action

[0521] Step 10:

[0522] Overall health assessment

[0523] The server integrates the analysis results of the facial expression data and voice data and the emotion evaluation results to comprehensively evaluate the user's mental and physical health.

[0524] Input: Analysis results of facial expression data and voice data, emotion evaluation results

[0525] Output: Overall health rating of the user

[0526] Step 11:

[0527] Generating health advice

[0528] The server generates specific health advice and self-care recommendations based on the assessment results.

[0529] Input: Overall health rating

[0530] The server generates advice and recommendations and sends them to the device.

[0531] Output: Health advice sent to device

[0532] Step 12:

[0533] Advice Notification

[0534] The terminal notifies the user of the generated advice.

[0535] The user checks the advice and takes necessary action, such as taking early rest.

[0536] Input: Health advice sent to device

[0537] Output: User takes action

[0538] (Application example 2)

[0539] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0540] Currently, there are no systems in factories that can monitor the mental and physical health of workers in real time and provide appropriate health advice and feedback. This increases the risk of workers suffering from overwork and stress. Furthermore, the difficulty of quickly understanding a worker's emotional state and taking appropriate measures poses a challenge to maintaining a safe and efficient work environment.

[0541] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0542] In this invention, the server includes means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health state, means for providing the user with appropriate health advice based on the evaluation results, and means for monitoring the health state of workers in the factory in real time and providing feedback, thereby enabling real-time monitoring of the health state of workers and appropriate feedback.

[0543] "Facial expression data" refers to digital information that can be analyzed by capturing a user's facial expressions in the form of an image or video.

[0544] "Voice data" refers to digital information that can be analyzed by capturing a user's voice or speech as an audio file.

[0545] "Analysis" refers to information processing and evaluation to determine the mental and physical state of the user based on the acquired facial expression data and voice data.

[0546] "User" refers to an individual who uses this system, including workers who perform work in a factory.

[0547] "Health" refers to the user's state of mental and physical well-being, including stress levels, fatigue levels, emotional state, and the like.

[0548] "Health advice" refers to specific guidelines or recommendations for action provided to users based on the analyzed assessment results.

[0549] "Self-care" refers to the self-management and self-treatment actions that users take to maintain and improve their health.

[0550] "Feedback" refers to real-time alerts and instructions provided to the user based on the analysis results.

[0551] "Factory" refers to an industrial facility and its environment where operations such as the manufacture of products take place.

[0552] "Real-time" refers to the fact that feedback is provided to the user almost immediately after data acquisition and analysis occurs.

[0553] MODE FOR CARRYING OUT THE INVENTION

[0554] The present invention is a system for analyzing the health status of workers in a factory in real time and providing appropriate health advice and self-care recommendations. A specific configuration and processing procedure for implementing the present invention are described below.

[0555] System Configuration

[0556] This system is broadly composed of three elements: a server, factory robots, and workers.

[0557] server

[0558] The server is the core of the system and performs the following functions:

[0559] 1. Analyze facial expression and voice data.

[0560] 2. Evaluate the mental and physical health of workers based on the analysis results.

[0561] 3. Generate health advice and self-care recommendations to provide to workers.

[0562] 4. Provide emotional feedback when appropriate.

[0563] 5. Use sentiment and speech analysis engines (e.g., Microsoft® Azure® Emotion API, Google Cloud Speech-to-Text API).

[0564] Factory robots

[0565] Factory robots perform the following roles while working on-site:

[0566] 1. Obtain facial expression data of the worker using a camera.

[0567] 2. Acquire the worker's voice data using a microphone.

[0568] 3. Send the acquired data to the server.

[0569] 4. The analysis results and advice from the server are notified to the worker.

[0570] Worker

[0571] Workers use the system to receive health management and feedback as follows:

[0572] 1. Data is collected automatically during daily operations.

[0573] 2. Check the health advice and feedback provided by the server.

[0574] 3. Take the necessary action.

[0575] Specific processing steps

[0576] Acquiring and Sending Data

[0577] The factory robot uses a camera and microphone to capture facial expression and voice data of the worker in real time, which is temporarily stored locally and periodically sent to a server.

[0578] Data analysis

[0579] The server analyzes the received facial expression and voice data using an emotion analysis engine and a voice analysis engine. As a result of the analysis, the worker's emotional state and health condition are evaluated.

[0580] Assessing health status and providing feedback

[0581] Based on the analysis results, the server comprehensively evaluates the mental and physical health of the worker and generates appropriate health advice and feedback. The generated feedback is provided to the worker in real time via the factory robot. For example, specific advice such as "You seem a little tired lately. I recommend you take a five-minute break. Take a deep breath and reduce stress" may be provided.

[0582] Example prompts to input to the generative AI model

[0583] For example, here is a sample prompt to input to the sentiment analysis engine:

[0584] You have received facial and audio data to analyze. Please extract the following information from it:

[0585] 1. Emotions detected from facial expressions (happiness, sadness, anger, anxiety, etc.).

[0586] 2. Emotions and stress levels extracted from audio data.

[0587] 3. Overall assessment of mental and physical health.

[0588] Use the results to generate appropriate self-care advice.

[0589] By using such prompt sentences, the emotion analysis engine and speech analysis engine on the server side can extract the desired analysis results and provide appropriate feedback.

[0590] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0591] Step 1:

[0592] A factory robot uses a camera to capture facial expression data of workers. The image data captured by the camera is input, and facial feature points are extracted through image processing. The extracted feature point data is output.

[0593] Step 2:

[0594] A factory robot uses a microphone to capture the voice data of a worker. At this time, voice recording is started and the recorded voice file becomes the input. The voice data is temporarily saved in local storage, and the saved voice file becomes the output.

[0595] Step 3:

[0596] The facial expression data and voice data stored in the local storage are sent to the server. The image data and voice data sent are input, and data transfer is performed by the client-side communication module. Data that has been received by the server is output.

[0597] Step 4:

[0598] The server analyzes the received facial expression data. At this time, image data is input and emotional information is extracted using an emotion analysis engine (e.g., Microsoft Azure Emotion API). The extracted emotional information is output.

[0599] Step 5:

[0600] The server analyzes the received voice data. The audio file is input, and a voice analysis engine (e.g., Google Cloud Speech-to-Text API) is used to analyze the voice tone and emotional state. The analyzed emotional and stress level information is output.

[0601] Step 6:

[0602] The server integrates the analysis results of the facial expression data and voice data to evaluate the overall health condition. At this time, emotional information and stress level are input, and the overall health condition is calculated by the health assessment algorithm. This evaluation result is the output.

[0603] Step 7:

[0604] The server generates appropriate health advice and self-care recommendations for the worker based on the evaluation results. The evaluation results are input, and specific action guidelines and recommendations are generated by the advice generation algorithm. This generated advice is the output.

[0605] Step 8:

[0606] The server sends the generated feedback to the factory robot. The advice content becomes the input and is sent to the factory robot via the communication module. The advice content that has been received by the factory robot becomes the output.

[0607] Step 9:

[0608] Factory robots provide advice and feedback to workers. The received feedback is input and notified to the worker via voice or display. The advice communicated to the worker is output.

[0609] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0610] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0611] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0612] [Second embodiment]

[0613] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0614] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0615] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0616] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0617] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0618] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0619] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0620] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0621] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0622] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0623] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0624] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0625] The present invention is a system that acquires and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status and provide appropriate health advice and self-care recommendations. Here, we will explain in detail the operation of the entire system.

[0626] System Configuration

[0627] This system is broadly composed of three elements: the server, the terminal (e.g., the user's smartphone), and the user. Each element and its role are explained below.

[0628] server

[0629] The server is the core of the system and performs the following functions:

[0630] 1. Analyze the user's facial expression and voice data.

[0631] 2. Evaluate the user's mental and physical health based on the analysis results.

[0632] 3. Generate health advice and self-care recommendations to provide to users.

[0633] 4. Offer fashion advice when needed.

[0634] Terminal

[0635] The terminal provides the user interface and has the following roles:

[0636] 1. Use a camera to capture facial expression data.

[0637] 2. Use a microphone to capture audio data.

[0638] 3. The server notifies the user of the analysis results and advice.

[0639] 4. Receives input from the user and sends it to the server.

[0640] User

[0641] Users use the system to receive assessments and advice on their health status.

[0642] 1. Enter basic information (age, gender, preferences, daily activity patterns, etc.) as the initial setting.

[0643] 2. Use a camera and microphone to provide facial and voice data.

[0644] 3. Follow the health advice and self-care recommendations provided by the server.

[0645] Specific examples of processing

[0646] Here we will explain a specific scenario in which the system actually works.

[0647] 1. Initial Setup

[0648] The user downloads the smartphone app and enters basic information.

[0649] The terminal sends this information to the server.

[0650] The server stores the basic information in a database for future analysis.

[0651] 2. Acquisition and analysis of facial and speech data

[0652] A user takes a selfie in front of the smartphone camera and also captures audio data.

[0653] The terminal transmits facial expression data and voice data to the server.

[0654] The server analyzes the data and assesses the user's mental and physical health.

[0655] The server generates appropriate health advice based on the evaluation results.

[0656] 3. Providing advice

[0657] The server generates health advice and self-care recommendations and sends them to the device.

[0658] The terminal notifies the user of the advice.

[0659] The user checks the advice provided and takes the necessary action.

[0660] 4. Providing fashion advice

[0661] The server generates fashion advice based on weather information and the user's schedule.

[0662] The terminal notifies the user of this.

[0663] The user follows the advice and chooses appropriate clothing.

[0664] In this way, this system effectively acquires and analyzes facial expression and voice data and provides advice based on the results to help users improve their communication skills and health management in their daily lives.

[0665] The processing flow will be explained below.

[0666] Step 1:

[0667] A user downloads and installs a smartphone app.

[0668] Step 2:

[0669] The user launches the app and enters basic information such as age, gender, preferences, and daily activity patterns on the basic settings screen.

[0670] Step 3:

[0671] The terminal sends the entered basic information to the server.

[0672] Step 4:

[0673] The server stores the received basic information in a database for future analysis.

[0674] Step 5:

[0675] Users take selfies in front of their smartphone cameras on a daily basis.

[0676] Step 6:

[0677] The device sends the selfie (facial expression data) to the server.

[0678] Step 7:

[0679] A user speaks into a smartphone to capture audio data.

[0680] Step 8:

[0681] The terminal transmits the acquired voice data to the server.

[0682] Step 9:

[0683] The server analyzes the received facial expression data to assess the user's emotional state, for example, detecting the degree of smile or sadness.

[0684] Step 10:

[0685] The server analyzes the received audio data and evaluates the tone and patterns of the voice, such as the pitch, rate, and emotional expression.

[0686] Step 11:

[0687] The server integrates the analysis results and assesses the user's mental and physical health.

[0688] Step 12:

[0689] The server generates specific health advice based on the assessment results, for example, "You seem a little tired today. I recommend you take a break."

[0690] Step 13:

[0691] The server sends the generated health advice to the terminal.

[0692] Step 14:

[0693] The terminal notifies the user of the health advice received from the server.

[0694] Step 15:

[0695] The user checks the advice provided and takes action as necessary.

[0696] Step 16:

[0697] Users take selfies periodically to provide data for checking the condition of their hair and beard.

[0698] Step 17:

[0699] The device sends the captured selfie data to the server.

[0700] Step 18:

[0701] The server analyzes the selfie data and determines whether hair or beard care is needed.

[0702] Step 19:

[0703] The server generates recommendations for beauty salons as needed and provides them to the user.

[0704] Step 20:

[0705] The device notifies the user of beauty salon recommendations.

[0706] Step 21:

[0707] The user makes a reservation at the recommended hair salon via their smartphone.

[0708] Step 22:

[0709] The server generates fashion advice based on season, weather, and schedule information.

[0710] Step 23:

[0711] The server transmits the generated fashion advice to the terminal.

[0712] Step 24:

[0713] The terminal notifies the user of fashion advice.

[0714] Step 25:

[0715] The user checks the provided fashion advice and selects an outfit.

[0716] Step 26:

[0717] The server evaluates the user's progress and training effectiveness in real time and updates the next training menu.

[0718] Step 27:

[0719] The device will notify and guide the user about the next training menu.

[0720] Step 28:

[0721] Users follow a training menu to effectively improve their communication skills.

[0722] In this way, each step of the system works in conjunction with one another to provide comprehensive support for the user's health management and improvement of communication skills.

[0723] Example 1

[0724] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0725] Conventional health management systems lack sufficient means for comprehensively evaluating a user's mental and physical health status, and it is difficult to efficiently provide personalized advice to each user. Furthermore, there are no systems that analyze a user's facial expressions and voice data to appropriately provide self-care recommendations or fashion advice. Therefore, there is a need to improve users' health management and comfort in their daily lives.

[0726] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0727] In this invention, the server includes means for acquiring facial expression data, means for acquiring voice data, means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health status, means for generating appropriate health advice for the user based on the evaluation results, and means for generating prompt sentences using a generative AI model to provide personalized advice to the user. This enables a comprehensive evaluation of the user's mental and physical health status and the provision of appropriate and personalized advice to each user. Furthermore, by recommending self-care and providing fashion advice as needed, the server achieves health management and improves the user's quality of life.

[0728] "Facial expression data" refers to data that digitally records a user's facial expressions.

[0729] "Voice data" refers to data that is a digital recording of a user's speech or voice.

[0730] "Analysis" refers to the process of using acquired digital data to identify features and patterns and evaluate the user's mental and physical state.

[0731] "Evaluating health status" means quantitatively and qualitatively determining the user's psychological and physical health status based on the analysis results of facial expression data and voice data.

[0732] "Generating health advice" means creating appropriate courses of action or recommendations based on the assessed health status of the user.

[0733] A "generative AI model" is an algorithm or system that uses artificial intelligence to generate text or instructions.

[0734] A "prompt sentence" is text data input into a generative AI model that describes the instructions and information needed to perform a specific task.

[0735] "Personalized advice" refers to specific and appropriate advice that is customized according to the characteristics and conditions of each individual user.

[0736] The present invention is a system that acquires and analyzes a user's facial expression and voice data to evaluate the user's mental and physical health and provide appropriate health advice and self-care recommendations. The invention has three main components: a server, a terminal, and a user.

[0737] System Configuration

[0738] server

[0739] The server is the core of the system and is responsible for:

[0740] 1. Receiving and storing data: Receives facial expression data and voice data sent from the user's device and stores them in a database.

[0741] 2. Data analysis: The received data is analyzed. Facial expression data is extracted using OpenCV, and voice data is converted to text using the Google Cloud Speech-to-Text API, followed by sentiment analysis using the spaCy library.

[0742] 3. Health assessment: Based on the analysis results, the user's mental and physical health status is assessed.

[0743] 4. Advice Generation: Using a generative AI model (e.g., GPT-3), prompts are generated based on the assessment results to create appropriate health advice and self-care recommendations, and provide fashion advice as needed.

[0744] Specific operation example

[0745] Initial setup and data acquisition

[0746] Users download the smartphone app and enter basic information (age, gender, preferences, daily activity patterns, etc.).

[0747] The terminal sends this information to the server, which stores it in a database.

[0748] Acquisition and analysis of facial expression and voice data

[0749] At designated times, the user takes a photo of their facial expression using the smartphone camera and records their voice using the microphone.

[0750] The device collects facial expression and voice data and sends them to the server in an appropriate format (e.g., JPEG images and WAV audio files).

[0751] The server analyzes the received data, extracting facial features using OpenCV, converting the audio data to text using the Google Cloud Speech-to-Text API, and performing emotion analysis using the spaCy library.

[0752] Health assessment and advice

[0753] Based on the analysis results, the server evaluates the user's mental and physical health.

[0754] The server sends the generated prompts to a generative AI model to create personalized health advice and self-care recommendations.

[0755] The server sends these advices to the terminal, which then notifies the user.

[0756] Specific prompt examples

[0757] “Analyze the user’s voice data and assess their emotional state. For example, if the user is tired, suggest a ‘relaxation exercise’.”

[0758] In this way, this system supports users in improving their communication skills and health management in their daily lives. By effectively analyzing facial expression and voice data and providing personalized advice based on the results, the system aims to comprehensively maintain and improve the user's health.

[0759] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0760] Step 1: Initial settings and user information registration

[0761] A user downloads and installs a smartphone app. As input, the user enters basic information (age, gender, preferences, daily activity patterns, etc.) into the app. The device collects this information and sends it to the server in JSON format. The server stores the received information in a database and uses it for future analysis. This allows user information to be accumulated in the database, making it possible to provide personalized services to each individual user.

[0762] Specific behavior:

[0763] A user installs the app on their smartphone, launches it, and enters some basic information. The device then makes an HTTP request to send this information to the server, which then stores it in a database.

[0764] Step 2: Acquire facial expression and voice data

[0765] At designated times, the user captures facial expression data using the smartphone camera and records audio data using the microphone. Camera image data and audio data are acquired as input. The device collects this data and sends it to the server as JPEG images and WAV audio files. The output is facial expression data and audio data.

[0766] Specific behavior:

[0767] When a user presses the "Get Data" button, the smartphone camera is activated and a picture of their face is taken. Then, a voice memo is recorded using the microphone. The device collects facial expression data in JPEG format and voice data in WAV format, and creates an HTTP request to send this to the server.

[0768] Step 3: Data analysis by the server

[0769] The server analyzes the received facial expression data using OpenCV and extracts feature points. It converts the voice data into text using the Google Cloud Speech-to-Text API and performs emotion analysis using an NLP library (e.g., spaCy). It receives facial expression data and voice data as input and obtains feature point data and emotion analysis results as output.

[0770] Specific behavior:

[0771] The server processes the received image data using the OpenCV library to extract facial features, converts the audio data into text using the Google Cloud Speech-to-Text API, and analyzes the resulting text data using the spaCy library to evaluate the emotional state. The results are then stored in a database.

[0772] Step 4: Health Assessment

[0773] The server evaluates the user's mental and physical health based on the analyzed facial expression and voice data. It uses feature point data and emotion analysis results as input and generates a health status evaluation result as output.

[0774] Specific behavior:

[0775] The server evaluates the user's health condition using a dedicated evaluation algorithm based on the analysis results. The evaluation results are stored in a database and used in the next step of generating advice.

[0776] Step 5: Generating and serving advice

[0777] The server sends prompts to a generative AI model (e.g., GPT-3) to generate appropriate health advice and self-care recommendations based on the assessment results. Fashion advice is also provided if necessary. The health assessment results are used as input, and the generated advice is obtained as output. The device notifies the user of this.

[0778] Specific behavior:

[0779] The server sends a prompt message to the AI ​​model saying, "The user's stress level is high. Please provide relaxation advice." The AI ​​model then collects the advice returned, converts it into JSON format, and sends it to the device. The device then generates a notification and displays it for the user to see.

[0780] The above is the specific processing flow of this system.

[0781] (Application example 1)

[0782] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0783] In traditional brick-and-mortar stores, it was difficult to grasp the mental and physical health status of customers and provide appropriate services and advice based on that. As a result, the quality of customer experience tended to be uniform, making it difficult to provide services that meet individual needs. This led to problems such as lower customer satisfaction and difficulty in securing repeat customers.

[0784] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0785] In this invention, the server includes means for acquiring facial expression data, means for acquiring voice data, means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health state, and means for providing the user with appropriate health advice based on the evaluation results. This makes it possible to grasp the mental and physical health state of customers in a physical store and provide appropriate health advice and self-care recommendations in real time.

[0786] "Facial expression data" is digital data about a person's facial expressions collected using a camera or other image capture device.

[0787] "Voice data" is digital data relating to a human voice collected using a microphone or other voice capture device.

[0788] "Analysis" refers to the act of processing acquired facial expression and voice data and applying computational processes and algorithms to understand, classify, and evaluate its content.

[0789] "Mental and physical health status" refers to the totality of information that indicates the user's psychological state and physical condition, including emotions, stress level, fatigue level, and the like.

[0790] "Health Advice" means specific recommendations or advice based on mental and physical health that are provided to users to help them maintain or improve their health in their daily lives.

[0791] "Self-care" refers to actions and efforts for health management that users themselves undertake, and includes relaxation techniques, reviewing diet and drink habits, exercise instruction, and the like.

[0792] "Fashion advice" refers to specific recommendations and suggestions to help users select appropriate clothing, accessories, etc. based on their evaluation results and circumstances.

[0793] "Real-time" refers to data acquisition, analysis, evaluation, and feedback occurring nearly simultaneously and without delay.

[0794] The present invention is a system that acquires and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status and provide appropriate health advice and self-care recommendations. Here, we will show a specific application example in a brick-and-mortar store and explain in detail an embodiment of the present invention.

[0795] System Configuration

[0796] This system is broadly composed of three elements: the server, the terminal (e.g., the user's smartphone), and the user. Each element and its role are explained below.

[0797] server

[0798] The server is the core of the system and performs the following functions:

[0799] 1. Analyze the user's facial expression and voice data.

[0800] The hardware and software used are TensorFlow / Keras.

[0801] 2. Evaluate the user's mental and physical health based on the analysis results.

[0802] 3. Generate health advice and self-care recommendations to provide to users.

[0803] A specific example would be to generate the advice "Listen to some relaxing music."

[0804] 4. Offer fashion advice when needed.

[0805] As an example of use, it generates advice based on weather information, such as "It's going to rain today, so please bring a waterproof jacket."

[0806] Terminal

[0807] The terminal provides the user interface and has the following roles:

[0808] 1. Use a camera to capture facial expression data.

[0809] The software used is OpenCV.

[0810] 2. Use a microphone to capture audio data.

[0811] The software used is librosa.

[0812] 3. The server notifies the user of the analysis results and advice.

[0813] To provide audio advice, gTTS (Google Text-to-Speech) is used and played on playsound.

[0814] 4. Receives input from the user and sends it to the server.

[0815] User

[0816] Users use the system to receive assessments and advice on their health status.

[0817] 1. Enter basic information (age, gender, preferences, daily activity patterns, etc.) as the initial setting.

[0818] 2. Use a camera and microphone to provide facial and voice data.

[0819] 3. Follow the health advice and self-care recommendations provided by the server.

[0820] Specific examples

[0821] Here we will explain a specific scenario in which the system actually works.

[0822] 1. Initial Setup

[0823] The user downloads the smartphone app and enters basic information.

[0824] The terminal sends this information to the server.

[0825] The server stores the basic information in a database for future analysis.

[0826] 2. Acquisition and analysis of facial and speech data

[0827] A user takes a selfie in front of the smartphone camera and also captures audio data.

[0828] The terminal transmits facial expression data and voice data to the server.

[0829] The server analyzes the data and assesses the user's mental and physical health.

[0830] The server generates appropriate health advice based on the evaluation results.

[0831] 3. Providing advice

[0832] The server generates health advice and self-care recommendations and sends them to the device.

[0833] The terminal notifies the user of the advice.

[0834] The user checks the advice provided and takes the necessary action.

[0835] 4. Providing fashion advice

[0836] The server generates fashion advice based on weather information and the user's schedule.

[0837] The terminal notifies the user of this.

[0838] The user follows the advice and chooses appropriate clothing.

[0839] Prompt Sentence Examples

[0840] Prompt: "Analyze the customer's facial and voice data to assess their current mental and physical state. Based on the assessment, generate effective relaxation advice."

[0841] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0842] Step 1:

[0843] The user launches the smartphone app and enters basic information (age, gender, preferences, daily activity patterns, etc.).

[0844] Input: User's basic information (age, gender, preferences, daily activity patterns, etc.)

[0845] Output: The basic information entered is saved on the device.

[0846] Step 2:

[0847] The terminal sends the user's basic information to the server.

[0848] Input: User basic information

[0849] Output: Basic information is sent to the server.

[0850] Step 3:

[0851] The server stores the received basic information in a database.

[0852] Input: Basic information data

[0853] Output: Basic information stored in the database

[0854] Step 4:

[0855] A user takes a selfie in front of the smartphone camera and captures audio data.

[0856] Input: Facial image data and audio data

[0857] Output: The captured facial image data and recorded audio data are saved on the device.

[0858] Step 5:

[0859] The terminal transmits facial expression data and voice data to the server.

[0860] Input: Facial image data and audio data

[0861] Output: Facial expression data and voice data are sent to the server.

[0862] Step 6:

[0863] The server parses the received data.

[0864] Facial expression data: Preprocessed using OpenCV, analyzed using a Keras model, and output emotion categories (e.g., "happy," "sad," "surprised," etc.).

[0865] Input: Face image data

[0866] Data processing: Preprocessing with OpenCV (grayscale conversion, resizing, etc.)

[0867] Data Computation: Sentiment Analysis with Keras Models

[0868] Output: Emotion category

[0869] Audio data: Preprocessing is performed using librosa, features are extracted, and the data is analyzed using a Keras model. The emotional category is output.

[0870] Input: Audio data

[0871] Data processing: Feature extraction (MFCC, etc.) using librosa

[0872] Data Computation: Sentiment Analysis with Keras Models

[0873] Output: Emotion category

[0874] Step 7:

[0875] The server evaluates the user's mental and physical health based on the analysis results and generates appropriate health advice.

[0876] Input: Emotion category (facial and vocal results)

[0877] Output: Health advice

[0878] Step 8:

[0879] The server sends the generated health advice to the terminal.

[0880] Enter: Health Advice

[0881] Output: Advice data is sent to the terminal.

[0882] Step 9:

[0883] The terminal notifies the user of the generated health advice.

[0884] Input: Advice data

[0885] Output: Advice notification display, audio playback (using gTTS and playsound)

[0886] Step 10:

[0887] The user checks the health advice provided and takes necessary action.

[0888] Input: Advice content

[0889] Output: User actions (taking a break, taking a deep breath, practicing relaxation techniques, etc.)

[0890] Step 11:

[0891] The server generates fashion advice based on weather information and the user's schedule.

[0892] Input: Weather information, user schedule

[0893] Output: Fashion advice

[0894] Step 12:

[0895] The terminal notifies the user of the generated fashion advice.

[0896] Enter: fashion advice

[0897] Output: Fashion advice notification display

[0898] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0899] The present invention is a system that acquires and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status, and provides appropriate health advice and self-care recommendations, in addition to providing an emotion engine that recognizes the user's emotions. The operation of the entire system is described in detail below.

[0900] System Configuration

[0901] This system is broadly composed of three elements: the server, the terminal (user's smartphone), and the user. Each element and its role are explained below.

[0902] server

[0903] The server is the core of the system and performs the following functions:

[0904] 1. Analyze the user's facial expression and voice data.

[0905] 2. Evaluate the user's mental and physical health based on the analysis results.

[0906] 3. Generate health advice and self-care recommendations to provide to users.

[0907] 4. Offer fashion advice when needed.

[0908] 5. Analyze user sentiment in real time using an emotion engine.

[0909] 6. Provide appropriate emotional feedback based on the results of emotion analysis.

[0910] Terminal

[0911] The terminal provides the user interface and has the following roles:

[0912] 1. Use a camera to capture facial expression data.

[0913] 2. Use a microphone to capture audio data.

[0914] 3. The server notifies the user of the analysis results and advice.

[0915] 4. Receives input from the user and sends it to the server.

[0916] User

[0917] Users use the system to receive assessments and advice on their health status.

[0918] 1. Enter basic information (age, gender, preferences, daily activity patterns, etc.) as the initial setting.

[0919] 2. Use a camera and microphone to provide facial and voice data.

[0920] 3. Follow the health advice and self-care recommendations provided by the server.

[0921] Specific examples of processing

[0922] Here we will explain a specific scenario in which the system actually works.

[0923] 1. Initial Setup

[0924] The user downloads the smartphone app and enters basic information.

[0925] The terminal sends this information to the server.

[0926] The server stores the basic information in a database for future analysis.

[0927] 2. Acquisition and analysis of facial and speech data

[0928] Users take selfies in front of their smartphone cameras on a daily basis.

[0929] The device sends the selfie (facial expression data) to the server.

[0930] A user speaks into a smartphone to capture audio data.

[0931] The terminal transmits the acquired voice data to the server.

[0932] 3. Emotion analysis using an emotion engine

[0933] The server uses the received facial expression data and voice data to analyze the user's emotions in real time using an emotion engine.

[0934] The server evaluates the user's emotional state based on the emotion analysis results, detecting emotions such as joy, sadness, and anger.

[0935] The server generates appropriate emotional feedback based on the emotion analysis results, e.g., "You seem a little tired today. Please stretch to relax."

[0936] 4. Health advice and self-care recommendations

[0937] The server integrates the analysis results of the facial expression data and voice data with the emotion evaluation results to comprehensively evaluate the user's mental and physical health.

[0938] The server generates specific health advice and self-care recommendations based on the assessment results.

[0939] The server generates advice and recommendations and sends them to the device.

[0940] The terminal notifies the user of the advice.

[0941] The user checks the advice provided and takes the necessary action.

[0942] 5. Providing fashion advice

[0943] The server generates fashion advice based on season, weather, and schedule information.

[0944] The terminal notifies the user of the generated fashion advice.

[0945] The user follows the advice and chooses appropriate clothing.

[0946] 6. Training progress and feedback

[0947] The server evaluates the user's progress and training effectiveness in real time and updates the next training menu.

[0948] The device will notify and guide the user about the next training menu.

[0949] Users follow a training menu to effectively improve their communication skills.

[0950] In this way, each step of the system works in conjunction to provide comprehensive support for users' health management and improvement of communication skills. By combining the emotion engine, it is possible to evaluate mental and physical aspects with higher accuracy than before, and to provide more personalized advice and recommendations.

[0951] The processing flow will be explained below.

[0952] Step 1:

[0953] A user downloads and installs a smartphone app.

[0954] Step 2:

[0955] The user launches the app and enters basic information such as age, gender, preferences, and daily activity patterns on the basic settings screen.

[0956] Step 3:

[0957] The terminal sends the entered basic information to the server.

[0958] Step 4:

[0959] The server stores the received basic information in a database for future analysis.

[0960] Step 5:

[0961] Users take selfies in front of their smartphone cameras on a daily basis.

[0962] Step 6:

[0963] The device sends the selfie (facial expression data) to the server.

[0964] Step 7:

[0965] A user speaks into a smartphone to capture audio data.

[0966] Step 8:

[0967] The terminal transmits the acquired voice data to the server.

[0968] Step 9:

[0969] The server analyzes the received facial expression data using a facial recognition algorithm and emotion engine to assess the user's emotional state, specifically identifying emotions such as smile, sadness, anger, and surprise.

[0970] Step 10:

[0971] The voice data received by the server is also analyzed by the emotion engine to evaluate the tone, speed, and emotional expression of the voice, e.g., excited, calm, angry, etc.

[0972] Step 11:

[0973] The server combines the analysis results of facial expression data and voice data to perform an emotional evaluation, and also comprehensively evaluates the user's mental and physical health.

[0974] Step 12:

[0975] The server generates specific health advice based on the evaluation results, for example, "You seem a little tired today. Please stretch to relax."

[0976] Step 13:

[0977] The server sends the generated health advice and self-care recommendations to the terminal.

[0978] Step 14:

[0979] The terminal notifies the user of the health advice received from the server.

[0980] Step 15:

[0981] The user checks the advice provided and takes the necessary action.

[0982] Step 16:

[0983] Users take selfies periodically to provide data for checking the condition of their hair and beard.

[0984] Step 17:

[0985] The device sends the captured selfie data to the server.

[0986] Step 18:

[0987] The server analyzes the selfie data and determines if you need hair or beard care. Example: "I need a beard trim."

[0988] Step 19:

[0989] The server generates salon recommendations as needed and lists nearby salons.

[0990] Step 20:

[0991] The device notifies the user of recommended beauty salons. For example, "Here are some recommended beauty salons near you."

[0992] Step 21:

[0993] The user makes a reservation at the recommended hair salon via their smartphone.

[0994] Step 22:

[0995] The server generates fashion advice based on season, weather, and schedule information. Example: "It's cold today, so I recommend a thick coat."

[0996] Step 23:

[0997] The terminal notifies the user of the generated fashion advice.

[0998] Step 24:

[0999] The user follows the advice and chooses appropriate clothing.

[1000] Step 25:

[1001] The server evaluates the user's progress and training effectiveness in real time and updates the next training menu.

[1002] Step 26:

[1003] The device will notify and guide the user about the next training menu.

[1004] Step 27:

[1005] Users follow a training menu to effectively improve their communication skills.

[1006] In this way, each step of the system works in conjunction to provide comprehensive support for users' health management and improvement of communication skills. By combining the emotion engine, it is possible to evaluate mental and physical aspects with higher accuracy than before, and to provide more personalized advice and recommendations.

[1007] Example 2

[1008] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1009] Conventional health assessment systems have difficulty in comprehensively assessing a user's mental and physical health status in real time, and have been unable to provide appropriate feedback quickly. Furthermore, they do not adequately provide advice that takes into account the user's emotional state, making it difficult to provide personalized care.

[1010] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1011] In this invention, the server includes means for acquiring facial expression data of a user, means for acquiring voice data of the user, means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health state, means for providing the user with appropriate health advice based on the evaluation results, means for analyzing the user's emotions in real time using an emotion analysis engine, and means for providing the user with appropriate emotional feedback based on the analysis results, thereby making it possible to comprehensively evaluate the user's health state in real time and quickly provide personalized advice.

[1012] "User" refers to an individual who uses the system and receives health assessments and advice.

[1013] "Facial expression data" is digital data obtained using a camera that contains information about the user's facial expressions.

[1014] "Voice data" is digital data obtained by capturing voice information when a user speaks using a microphone.

[1015] "Analysis" refers to the act of processing information based on acquired data to quantitatively or qualitatively evaluate the user's health and emotional state.

[1016] "Health condition" is the result of a comprehensive evaluation of the user's mental and physical conditions.

[1017] "Health advice" refers to specific guidelines or recommended actions provided to the user based on the evaluation results.

[1018] An "emotion analysis engine" is software or a program for analyzing a user's emotional state in real time based on facial expression data and voice data.

[1019] "Emotional feedback" refers to specific countermeasures and advice provided to users based on the results of the emotion analysis engine.

[1020] "Self-care recommendations" are specific advice or recommended actions for self-care that a user can take by themselves.

[1021] "Fashion advice" is advice on clothing provided to users based on information such as the season, weather, and schedule.

[1022] A "terminal" is a device used by a user, and is a device that provides a user interface, including a smartphone, tablet, etc.

[1023] The "server" is the core part of this system, and is a computer system that analyzes and stores data, generates advice, and so on.

[1024] The present invention provides a system that collects and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status, and provides appropriate health advice, self-care recommendations, and emotional feedback based on real-time emotional analysis. Specific embodiments are described below.

[1025] System Configuration

[1026] This system is broadly composed of three elements: a server, a terminal (user's smartphone), and a user.

[1027] server

[1028] The server is the core of the system and performs the following functions:

[1029] 1. Analysis of the user's facial expression data and voice data: The server processes the facial expression data and voice data sent from the device using an analysis engine (e.g., Amazon Rekognition, Google Cloud Speech-to-Text).

[1030] 2. Evaluating the user's health status: Based on the results of the analysis engine, the user's mental and physical health status is evaluated.

[1031] 3. Generating health advice and self-care recommendations: Based on the health status assessment results, appropriate health advice and self-care recommendations are generated.

[1032] 4. Real-time emotion analysis using an emotion analysis engine: Facial expression data and voice data are used to activate the emotion analysis engine, which analyzes the user's emotions in real time.

[1033] 5. Providing emotional feedback: Based on the results of emotion analysis, appropriate emotional feedback is generated and sent to the device.

[1034] Device (smartphone)

[1035] The terminal is responsible for the following:

[1036] 1. Acquiring facial expression data: The user's facial expressions are captured by a camera and saved as facial expression data.

[1037] 2. Acquiring voice data: The user's speech is recorded with a microphone and saved as voice data.

[1038] 3. Data transmission: The acquired facial expression data and voice data are transmitted to the server.

[1039] 4. Notification of advice and feedback: Notify the user of the analysis results, advice, and emotional feedback received from the server.

[1040] User

[1041] The user does the following:

[1042] 1. Download the app and enter basic information: Download the smartphone app and enter basic information such as age, gender, preferences, and daily activity patterns.

[1043] 2. Providing facial expression data and voice data: Taking selfies in front of a smartphone camera in everyday life and speaking into the smartphone provides voice data.

[1044] 3. Review and act on advice: Take action based on the health advice, self-care recommendations, and emotional feedback provided by the server.

[1045] Examples of concrete examples and prompts

[1046] Specific examples

[1047] For example, when a user wakes up in the morning, they take a selfie with their smartphone and say, "Good morning, what are your plans for today?" The device immediately sends the selfie and voice data to the server. The server receives the data and analyzes it using its emotion engine. From the analysis results, it detects that the user looks a little tired, and generates advice such as, "You seem a little tired today. Please stretch to relax." The device notifies the user of this advice, and the user confirms the notification and takes time to relax early.

[1048] Prompt Sentence Examples

[1049] "Design a system where a user provides selfies and voice data, and a server analyzes this and generates appropriate health advice. For example, if the user is tired, provide advice on relaxation. Also, use an emotion analysis engine to evaluate emotions in real time."

[1050] In this way, each step of the system works in conjunction to provide comprehensive support for the user's health management and understanding of their emotional state. By combining the emotion engine, it is possible to assess mental and physical aspects more accurately than conventional systems, and to provide more personalized advice and recommendations.

[1051] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1052] Step 1:

[1053] Download the app and enter your basic information

[1054] A user downloads and installs a smartphone app.

[1055] A user opens the app and enters basic information such as age, gender, preferences, and daily activity patterns.

[1056] Input: User basic information

[1057] The terminal sends the basic information entered by the user to the server.

[1058] Output: Basic information sent to the server

[1059] Step 2:

[1060] Save basic information

[1061] The server checks the received basic information and stores it in the database.

[1062] Input: Basic information submitted

[1063] The server reviews the stored information and prepares it for future analysis.

[1064] Output: Basic information saved

[1065] Step 3:

[1066] Acquiring facial expression data

[1067] A user takes a selfie in front of the smartphone camera in daily life.

[1068] Input: Selfie image

[1069] Your device will temporarily store the selfie image in its internal memory.

[1070] Output: Facial expression data stored in temporary memory

[1071] Step 4:

[1072] Sending facial expression data

[1073] The device sends the temporarily stored selfie data to the server.

[1074] Input: Facial expression data stored in temporary memory

[1075] The server passes the received selfie data to the analysis engine.

[1076] Output: Facial expression data passed to the analysis engine

[1077] Step 5:

[1078] Acquiring audio data

[1079] The user speaks into the smartphone and voice data is captured.

[1080] Input: Audio data

[1081] The device temporarily stores the captured audio data in its internal memory.

[1082] Output: Audio data stored in temporary memory

[1083] Step 6:

[1084] Sending audio data

[1085] The terminal transmits the temporarily stored voice data to the server.

[1086] Input: Audio data stored in temporary memory

[1087] The server passes the received audio data to the analysis engine.

[1088] Output: Audio data passed to the analysis engine

[1089] Step 7:

[1090] Analysis of facial expression and voice data

[1091] The server uses the received facial expression data and voice data to instruct the emotion analysis engine to perform analysis.

[1092] Input: Facial expression and voice data passed to the analysis engine

[1093] An emotion analysis engine analyzes the data to detect emotional states such as joy, sadness, and anger.

[1094] Output: Detected emotional state

[1095] Step 8:

[1096] Generating Emotional Feedback

[1097] The server evaluates the user's emotional state based on the analysis results of the emotion engine and generates appropriate emotional feedback.

[1098] Input: Detected emotional state

[1099] The server generates emotional feedback and sends it to the device.

[1100] Output: Emotional feedback sent to the device

[1101] Step 9:

[1102] Feedback Notification

[1103] The terminal notifies the user of the received feedback.

[1104] Input: Emotional feedback sent to the device

[1105] The user checks the notification and takes the necessary action, such as stretching to relax.

[1106] Output: User takes action

[1107] Step 10:

[1108] Overall health assessment

[1109] The server integrates the analysis results of the facial expression data and voice data and the emotion evaluation results to comprehensively evaluate the user's mental and physical health.

[1110] Input: Analysis results of facial expression data and voice data, emotion evaluation results

[1111] Output: Overall health rating of the user

[1112] Step 11:

[1113] Generating health advice

[1114] The server generates specific health advice and self-care recommendations based on the assessment results.

[1115] Input: Overall health rating

[1116] The server generates advice and recommendations and sends them to the device.

[1117] Output: Health advice sent to device

[1118] Step 12:

[1119] Advice Notification

[1120] The terminal notifies the user of the generated advice.

[1121] The user checks the advice and takes necessary action, such as taking early rest.

[1122] Input: Health advice sent to device

[1123] Output: User takes action

[1124] (Application example 2)

[1125] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1126] Currently, there are no systems in factories that can monitor the mental and physical health of workers in real time and provide appropriate health advice and feedback. This increases the risk of workers suffering from overwork and stress. Furthermore, the difficulty of quickly understanding a worker's emotional state and taking appropriate measures poses a challenge to maintaining a safe and efficient work environment.

[1127] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1128] In this invention, the server includes means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health state, means for providing the user with appropriate health advice based on the evaluation results, and means for monitoring the health state of workers in the factory in real time and providing feedback, thereby enabling real-time monitoring of the health state of workers and appropriate feedback.

[1129] "Facial expression data" refers to digital information that can be analyzed by capturing a user's facial expressions in the form of an image or video.

[1130] "Voice data" refers to digital information that can be analyzed by capturing a user's voice or speech as an audio file.

[1131] "Analysis" refers to information processing and evaluation to determine the mental and physical state of the user based on the acquired facial expression data and voice data.

[1132] "User" refers to an individual who uses this system, including workers who perform work in a factory.

[1133] "Health" refers to the user's state of mental and physical well-being, including stress levels, fatigue levels, emotional state, and the like.

[1134] "Health advice" refers to specific guidelines or recommendations for action provided to users based on the analyzed assessment results.

[1135] "Self-care" refers to the self-management and self-treatment actions that users take to maintain and improve their health.

[1136] "Feedback" refers to real-time alerts and instructions provided to the user based on the analysis results.

[1137] "Factory" refers to an industrial facility and its environment where operations such as the manufacture of products take place.

[1138] "Real-time" refers to the fact that feedback is provided to the user almost immediately after data acquisition and analysis occurs.

[1139] MODE FOR CARRYING OUT THE INVENTION

[1140] The present invention is a system for analyzing the health status of workers in a factory in real time and providing appropriate health advice and self-care recommendations. A specific configuration and processing procedure for implementing the present invention are described below.

[1141] System Configuration

[1142] This system is broadly composed of three elements: a server, factory robots, and workers.

[1143] server

[1144] The server is the core of the system and performs the following functions:

[1145] 1. Analyze facial expression and voice data.

[1146] 2. Evaluate the mental and physical health of workers based on the analysis results.

[1147] 3. Generate health advice and self-care recommendations to provide to workers.

[1148] 4. Provide emotional feedback when appropriate.

[1149] 5. Use sentiment and speech analysis engines (e.g., Microsoft Azure Emotion API, Google Cloud Speech-to-Text API).

[1150] Factory robots

[1151] Factory robots perform the following roles while working on-site:

[1152] 1. Obtain facial expression data of the worker using a camera.

[1153] 2. Acquire the worker's voice data using a microphone.

[1154] 3. Send the acquired data to the server.

[1155] 4. The analysis results and advice from the server are notified to the worker.

[1156] Worker

[1157] Workers use the system to receive health management and feedback as follows:

[1158] 1. Data is collected automatically during daily operations.

[1159] 2. Check the health advice and feedback provided by the server.

[1160] 3. Take the necessary action.

[1161] Specific processing steps

[1162] Acquiring and Sending Data

[1163] The factory robot uses a camera and microphone to capture facial expression and voice data of the worker in real time, which is temporarily stored locally and periodically sent to a server.

[1164] Data analysis

[1165] The server analyzes the received facial expression and voice data using an emotion analysis engine and a voice analysis engine. As a result of the analysis, the worker's emotional state and health condition are evaluated.

[1166] Assessing health status and providing feedback

[1167] Based on the analysis results, the server comprehensively evaluates the mental and physical health of the worker and generates appropriate health advice and feedback. The generated feedback is provided to the worker in real time via the factory robot. For example, specific advice such as "You seem a little tired lately. I recommend you take a five-minute break. Take a deep breath and reduce stress" may be provided.

[1168] Example prompts to input to the generative AI model

[1169] For example, here is a sample prompt to input to the sentiment analysis engine:

[1170] You have received facial and audio data to analyze. Please extract the following information from it:

[1171] 1. Emotions detected from facial expressions (happiness, sadness, anger, anxiety, etc.).

[1172] 2. Emotions and stress levels extracted from audio data.

[1173] 3. Overall assessment of mental and physical health.

[1174] Use the results to generate appropriate self-care advice.

[1175] By using such prompt sentences, the emotion analysis engine and speech analysis engine on the server side can extract the desired analysis results and provide appropriate feedback.

[1176] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1177] Step 1:

[1178] A factory robot uses a camera to capture facial expression data of workers. The image data captured by the camera is input, and facial feature points are extracted through image processing. The extracted feature point data is output.

[1179] Step 2:

[1180] A factory robot uses a microphone to capture the voice data of a worker. At this time, voice recording is started and the recorded voice file becomes the input. The voice data is temporarily saved in local storage, and the saved voice file becomes the output.

[1181] Step 3:

[1182] The facial expression data and voice data stored in the local storage are sent to the server. The image data and voice data sent are input, and data transfer is performed by the client-side communication module. Data that has been received by the server is output.

[1183] Step 4:

[1184] The server analyzes the received facial expression data. At this time, image data is input and emotional information is extracted using an emotion analysis engine (e.g., Microsoft Azure Emotion API). The extracted emotional information is output.

[1185] Step 5:

[1186] The server analyzes the received voice data. The audio file is input, and a voice analysis engine (e.g., Google Cloud Speech-to-Text API) is used to analyze the voice tone and emotional state. The analyzed emotional and stress level information is output.

[1187] Step 6:

[1188] The server integrates the analysis results of the facial expression data and voice data to evaluate the overall health condition. At this time, emotional information and stress level are input, and the overall health condition is calculated by the health assessment algorithm. This evaluation result is the output.

[1189] Step 7:

[1190] The server generates appropriate health advice and self-care recommendations for the worker based on the evaluation results. The evaluation results are input, and specific action guidelines and recommendations are generated by the advice generation algorithm. This generated advice is the output.

[1191] Step 8:

[1192] The server sends the generated feedback to the factory robot. The advice content becomes the input and is sent to the factory robot via the communication module. The advice content that has been received by the factory robot becomes the output.

[1193] Step 9:

[1194] Factory robots provide advice and feedback to workers. The received feedback is input and notified to the worker via voice or display. The advice communicated to the worker is output.

[1195] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1196] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1197] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1198] [Third embodiment]

[1199] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1200] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1201] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1202] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1203] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1204] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1205] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1206] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1207] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1208] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1209] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1210] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1211] The present invention is a system that acquires and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status and provide appropriate health advice and self-care recommendations. Here, we will explain in detail the operation of the entire system.

[1212] System Configuration

[1213] This system is broadly composed of three elements: the server, the terminal (e.g., the user's smartphone), and the user. Each element and its role are explained below.

[1214] server

[1215] The server is the core of the system and performs the following functions:

[1216] 1. Analyze the user's facial expression and voice data.

[1217] 2. Evaluate the user's mental and physical health based on the analysis results.

[1218] 3. Generate health advice and self-care recommendations to provide to users.

[1219] 4. Offer fashion advice when needed.

[1220] Terminal

[1221] The terminal provides the user interface and has the following roles:

[1222] 1. Use a camera to capture facial expression data.

[1223] 2. Use a microphone to capture audio data.

[1224] 3. The server notifies the user of the analysis results and advice.

[1225] 4. Receives input from the user and sends it to the server.

[1226] User

[1227] Users use the system to receive assessments and advice on their health status.

[1228] 1. Enter basic information (age, gender, preferences, daily activity patterns, etc.) as the initial setting.

[1229] 2. Use a camera and microphone to provide facial and voice data.

[1230] 3. Follow the health advice and self-care recommendations provided by the server.

[1231] Specific examples of processing

[1232] Here we will explain a specific scenario in which the system actually works.

[1233] 1. Initial Setup

[1234] The user downloads the smartphone app and enters basic information.

[1235] The terminal sends this information to the server.

[1236] The server stores the basic information in a database for future analysis.

[1237] 2. Acquisition and analysis of facial and speech data

[1238] A user takes a selfie in front of the smartphone camera and also captures audio data.

[1239] The terminal transmits facial expression data and voice data to the server.

[1240] The server analyzes the data and assesses the user's mental and physical health.

[1241] The server generates appropriate health advice based on the evaluation results.

[1242] 3. Providing advice

[1243] The server generates health advice and self-care recommendations and sends them to the device.

[1244] The terminal notifies the user of the advice.

[1245] The user checks the advice provided and takes the necessary action.

[1246] 4. Providing fashion advice

[1247] The server generates fashion advice based on weather information and the user's schedule.

[1248] The terminal notifies the user of this.

[1249] The user follows the advice and chooses appropriate clothing.

[1250] In this way, this system effectively acquires and analyzes facial expression and voice data and provides advice based on the results to help users improve their communication skills and health management in their daily lives.

[1251] The processing flow will be explained below.

[1252] Step 1:

[1253] A user downloads and installs a smartphone app.

[1254] Step 2:

[1255] The user launches the app and enters basic information such as age, gender, preferences, and daily activity patterns on the basic settings screen.

[1256] Step 3:

[1257] The terminal sends the entered basic information to the server.

[1258] Step 4:

[1259] The server stores the received basic information in a database for future analysis.

[1260] Step 5:

[1261] Users take selfies in front of their smartphone cameras on a daily basis.

[1262] Step 6:

[1263] The device sends the selfie (facial expression data) to the server.

[1264] Step 7:

[1265] A user speaks into a smartphone to capture audio data.

[1266] Step 8:

[1267] The terminal transmits the acquired voice data to the server.

[1268] Step 9:

[1269] The server analyzes the received facial expression data to assess the user's emotional state, for example, detecting the degree of smile or sadness.

[1270] Step 10:

[1271] The server analyzes the received audio data and evaluates the tone and patterns of the voice, such as the pitch, rate, and emotional expression.

[1272] Step 11:

[1273] The server integrates the analysis results and assesses the user's mental and physical health.

[1274] Step 12:

[1275] The server generates specific health advice based on the assessment results, for example, "You seem a little tired today. I recommend you take a break."

[1276] Step 13:

[1277] The server sends the generated health advice to the terminal.

[1278] Step 14:

[1279] The terminal notifies the user of the health advice received from the server.

[1280] Step 15:

[1281] The user checks the advice provided and takes action as necessary.

[1282] Step 16:

[1283] Users take selfies periodically to provide data for checking the condition of their hair and beard.

[1284] Step 17:

[1285] The device sends the captured selfie data to the server.

[1286] Step 18:

[1287] The server analyzes the selfie data and determines whether hair or beard care is needed.

[1288] Step 19:

[1289] The server generates recommendations for beauty salons as needed and provides them to the user.

[1290] Step 20:

[1291] The device notifies the user of beauty salon recommendations.

[1292] Step 21:

[1293] The user makes a reservation at the recommended hair salon via their smartphone.

[1294] Step 22:

[1295] The server generates fashion advice based on season, weather, and schedule information.

[1296] Step 23:

[1297] The server transmits the generated fashion advice to the terminal.

[1298] Step 24:

[1299] The terminal notifies the user of fashion advice.

[1300] Step 25:

[1301] The user checks the provided fashion advice and selects an outfit.

[1302] Step 26:

[1303] The server evaluates the user's progress and training effectiveness in real time and updates the next training menu.

[1304] Step 27:

[1305] The device will notify and guide the user about the next training menu.

[1306] Step 28:

[1307] Users follow a training menu to effectively improve their communication skills.

[1308] In this way, each step of the system works in conjunction with one another to provide comprehensive support for the user's health management and improvement of communication skills.

[1309] Example 1

[1310] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1311] Conventional health management systems lack sufficient means for comprehensively evaluating a user's mental and physical health status, and it is difficult to efficiently provide personalized advice to each user. Furthermore, there are no systems that analyze a user's facial expressions and voice data to appropriately provide self-care recommendations or fashion advice. Therefore, there is a need to improve users' health management and comfort in their daily lives.

[1312] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1313] In this invention, the server includes means for acquiring facial expression data, means for acquiring voice data, means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health status, means for generating appropriate health advice for the user based on the evaluation results, and means for generating prompt sentences using a generative AI model to provide personalized advice to the user. This enables a comprehensive evaluation of the user's mental and physical health status and the provision of appropriate and personalized advice to each user. Furthermore, by recommending self-care and providing fashion advice as needed, the server achieves health management and improves the user's quality of life.

[1314] "Facial expression data" refers to data that digitally records a user's facial expressions.

[1315] "Voice data" refers to data that is a digital recording of a user's speech or voice.

[1316] "Analysis" refers to the process of using acquired digital data to identify features and patterns and evaluate the user's mental and physical state.

[1317] "Evaluating health status" means quantitatively and qualitatively determining the user's psychological and physical health status based on the analysis results of facial expression data and voice data.

[1318] "Generating health advice" means creating appropriate courses of action or recommendations based on the assessed health status of the user.

[1319] A "generative AI model" is an algorithm or system that uses artificial intelligence to generate text or instructions.

[1320] A "prompt sentence" is text data input into a generative AI model that describes the instructions and information needed to perform a specific task.

[1321] "Personalized advice" refers to specific and appropriate advice that is customized according to the characteristics and conditions of each individual user.

[1322] The present invention is a system that acquires and analyzes a user's facial expression and voice data to evaluate the user's mental and physical health and provide appropriate health advice and self-care recommendations. The invention has three main components: a server, a terminal, and a user.

[1323] System Configuration

[1324] server

[1325] The server is the core of the system and is responsible for:

[1326] 1. Receiving and storing data: Receives facial expression data and voice data sent from the user's device and stores them in a database.

[1327] 2. Data analysis: The received data is analyzed. Facial expression data is extracted using OpenCV, and voice data is converted to text using the Google Cloud Speech-to-Text API, followed by sentiment analysis using the spaCy library.

[1328] 3. Health assessment: Based on the analysis results, the user's mental and physical health status is assessed.

[1329] 4. Advice Generation: Using a generative AI model (e.g., GPT-3), prompts are generated based on the assessment results to create appropriate health advice and self-care recommendations, and provide fashion advice as needed.

[1330] Specific operation example

[1331] Initial setup and data acquisition

[1332] Users download the smartphone app and enter basic information (age, gender, preferences, daily activity patterns, etc.).

[1333] The terminal sends this information to the server, which stores it in a database.

[1334] Acquisition and analysis of facial expression and voice data

[1335] At designated times, the user takes a photo of their facial expression using the smartphone camera and records their voice using the microphone.

[1336] The device collects facial expression and voice data and sends them to the server in an appropriate format (e.g., JPEG images and WAV audio files).

[1337] The server analyzes the received data, extracting facial features using OpenCV, converting the audio data to text using the Google Cloud Speech-to-Text API, and performing emotion analysis using the spaCy library.

[1338] Health assessment and advice

[1339] Based on the analysis results, the server evaluates the user's mental and physical health.

[1340] The server sends the generated prompts to a generative AI model to create personalized health advice and self-care recommendations.

[1341] The server sends these advices to the terminal, which then notifies the user.

[1342] Specific prompt examples

[1343] “Analyze the user’s voice data and assess their emotional state. For example, if the user is tired, suggest a ‘relaxation exercise’.”

[1344] In this way, this system supports users in improving their communication skills and health management in their daily lives. By effectively analyzing facial expression and voice data and providing personalized advice based on the results, the system aims to comprehensively maintain and improve the user's health.

[1345] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1346] Step 1: Initial settings and user information registration

[1347] A user downloads and installs a smartphone app. As input, the user enters basic information (age, gender, preferences, daily activity patterns, etc.) into the app. The device collects this information and sends it to the server in JSON format. The server stores the received information in a database and uses it for future analysis. This allows user information to be accumulated in the database, making it possible to provide personalized services to each individual user.

[1348] Specific behavior:

[1349] A user installs the app on their smartphone, launches it, and enters some basic information. The device then makes an HTTP request to send this information to the server, which then stores it in a database.

[1350] Step 2: Acquire facial expression and voice data

[1351] At designated times, the user captures facial expression data using the smartphone camera and records audio data using the microphone. Camera image data and audio data are acquired as input. The device collects this data and sends it to the server as JPEG images and WAV audio files. The output is facial expression data and audio data.

[1352] Specific behavior:

[1353] When a user presses the "Get Data" button, the smartphone camera is activated and a picture of their face is taken. Then, a voice memo is recorded using the microphone. The device collects facial expression data in JPEG format and voice data in WAV format, and creates an HTTP request to send this to the server.

[1354] Step 3: Data analysis by the server

[1355] The server analyzes the received facial expression data using OpenCV and extracts feature points. It converts the voice data into text using the Google Cloud Speech-to-Text API and performs emotion analysis using an NLP library (e.g., spaCy). It receives facial expression data and voice data as input and obtains feature point data and emotion analysis results as output.

[1356] Specific behavior:

[1357] The server processes the received image data using the OpenCV library to extract facial features, converts the audio data into text using the Google Cloud Speech-to-Text API, and analyzes the resulting text data using the spaCy library to evaluate the emotional state. The results are then stored in a database.

[1358] Step 4: Health Assessment

[1359] The server evaluates the user's mental and physical health based on the analyzed facial expression and voice data. It uses feature point data and emotion analysis results as input and generates a health status evaluation result as output.

[1360] Specific behavior:

[1361] The server evaluates the user's health condition using a dedicated evaluation algorithm based on the analysis results. The evaluation results are stored in a database and used in the next step of generating advice.

[1362] Step 5: Generating and serving advice

[1363] The server sends prompts to a generative AI model (e.g., GPT-3) to generate appropriate health advice and self-care recommendations based on the assessment results. Fashion advice is also provided if necessary. The health assessment results are used as input, and the generated advice is obtained as output. The device notifies the user of this.

[1364] Specific behavior:

[1365] The server sends a prompt message to the AI ​​model saying, "The user's stress level is high. Please provide relaxation advice." The AI ​​model then collects the advice returned, converts it into JSON format, and sends it to the device. The device then generates a notification and displays it for the user to see.

[1366] The above is the specific processing flow of this system.

[1367] (Application example 1)

[1368] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1369] In traditional brick-and-mortar stores, it was difficult to grasp the mental and physical health status of customers and provide appropriate services and advice based on that. As a result, the quality of customer experience tended to be uniform, making it difficult to provide services that meet individual needs. This led to problems such as lower customer satisfaction and difficulty in securing repeat customers.

[1370] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1371] In this invention, the server includes means for acquiring facial expression data, means for acquiring voice data, means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health state, and means for providing the user with appropriate health advice based on the evaluation results. This makes it possible to grasp the mental and physical health state of customers in a physical store and provide appropriate health advice and self-care recommendations in real time.

[1372] "Facial expression data" is digital data about a person's facial expressions collected using a camera or other image capture device.

[1373] "Voice data" is digital data relating to a human voice collected using a microphone or other voice capture device.

[1374] "Analysis" refers to the act of processing acquired facial expression and voice data and applying computational processes and algorithms to understand, classify, and evaluate its content.

[1375] "Mental and physical health status" refers to the totality of information that indicates the user's psychological state and physical condition, including emotions, stress level, fatigue level, and the like.

[1376] "Health Advice" means specific recommendations or advice based on mental and physical health that are provided to users to help them maintain or improve their health in their daily lives.

[1377] "Self-care" refers to actions and efforts for health management that users themselves undertake, and includes relaxation techniques, reviewing diet and drink habits, exercise instruction, and the like.

[1378] "Fashion advice" refers to specific recommendations and suggestions to help users select appropriate clothing, accessories, etc. based on their evaluation results and circumstances.

[1379] "Real-time" refers to data acquisition, analysis, evaluation, and feedback occurring nearly simultaneously and without delay.

[1380] The present invention is a system that acquires and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status and provide appropriate health advice and self-care recommendations. Here, we will show a specific application example in a brick-and-mortar store and explain in detail an embodiment of the present invention.

[1381] System Configuration

[1382] This system is broadly composed of three elements: the server, the terminal (e.g., the user's smartphone), and the user. Each element and its role are explained below.

[1383] server

[1384] The server is the core of the system and performs the following functions:

[1385] 1. Analyze the user's facial expression and voice data.

[1386] The hardware and software used are TensorFlow / Keras.

[1387] 2. Evaluate the user's mental and physical health based on the analysis results.

[1388] 3. Generate health advice and self-care recommendations to provide to users.

[1389] A specific example would be to generate the advice "Listen to some relaxing music."

[1390] 4. Offer fashion advice when needed.

[1391] As an example of use, it generates advice based on weather information, such as "It's going to rain today, so please bring a waterproof jacket."

[1392] Terminal

[1393] The terminal provides the user interface and has the following roles:

[1394] 1. Use a camera to capture facial expression data.

[1395] The software used is OpenCV.

[1396] 2. Use a microphone to capture audio data.

[1397] The software used is librosa.

[1398] 3. The server notifies the user of the analysis results and advice.

[1399] To provide audio advice, gTTS (Google Text-to-Speech) is used and played on playsound.

[1400] 4. Receives input from the user and sends it to the server.

[1401] User

[1402] Users use the system to receive assessments and advice on their health status.

[1403] 1. Enter basic information (age, gender, preferences, daily activity patterns, etc.) as the initial setting.

[1404] 2. Use a camera and microphone to provide facial and voice data.

[1405] 3. Follow the health advice and self-care recommendations provided by the server.

[1406] Specific examples

[1407] Here we will explain a specific scenario in which the system actually works.

[1408] 1. Initial Setup

[1409] The user downloads the smartphone app and enters basic information.

[1410] The terminal sends this information to the server.

[1411] The server stores the basic information in a database for future analysis.

[1412] 2. Acquisition and analysis of facial and speech data

[1413] A user takes a selfie in front of the smartphone camera and also captures audio data.

[1414] The terminal transmits facial expression data and voice data to the server.

[1415] The server analyzes the data and assesses the user's mental and physical health.

[1416] The server generates appropriate health advice based on the evaluation results.

[1417] 3. Providing advice

[1418] The server generates health advice and self-care recommendations and sends them to the device.

[1419] The terminal notifies the user of the advice.

[1420] The user checks the advice provided and takes the necessary action.

[1421] 4. Providing fashion advice

[1422] The server generates fashion advice based on weather information and the user's schedule.

[1423] The terminal notifies the user of this.

[1424] The user follows the advice and chooses appropriate clothing.

[1425] Prompt Sentence Examples

[1426] Prompt: "Analyze the customer's facial and voice data to assess their current mental and physical state. Based on the assessment, generate effective relaxation advice."

[1427] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1428] Step 1:

[1429] The user launches the smartphone app and enters basic information (age, gender, preferences, daily activity patterns, etc.).

[1430] Input: User's basic information (age, gender, preferences, daily activity patterns, etc.)

[1431] Output: The basic information entered is saved on the device.

[1432] Step 2:

[1433] The terminal sends the user's basic information to the server.

[1434] Input: User basic information

[1435] Output: Basic information is sent to the server.

[1436] Step 3:

[1437] The server stores the received basic information in a database.

[1438] Input: Basic information data

[1439] Output: Basic information stored in the database

[1440] Step 4:

[1441] A user takes a selfie in front of the smartphone camera and captures audio data.

[1442] Input: Facial image data and audio data

[1443] Output: The captured facial image data and recorded audio data are saved on the device.

[1444] Step 5:

[1445] The terminal transmits facial expression data and voice data to the server.

[1446] Input: Facial image data and audio data

[1447] Output: Facial expression data and voice data are sent to the server.

[1448] Step 6:

[1449] The server parses the received data.

[1450] Facial expression data: Preprocessed using OpenCV, analyzed using a Keras model, and output emotion categories (e.g., "happy," "sad," "surprised," etc.).

[1451] Input: Face image data

[1452] Data processing: Preprocessing with OpenCV (grayscale conversion, resizing, etc.)

[1453] Data Computation: Sentiment Analysis with Keras Models

[1454] Output: Emotion category

[1455] Audio data: Preprocessing is performed using librosa, features are extracted, and the data is analyzed using a Keras model. The emotional category is output.

[1456] Input: Audio data

[1457] Data processing: Feature extraction (MFCC, etc.) using librosa

[1458] Data Computation: Sentiment Analysis with Keras Models

[1459] Output: Emotion category

[1460] Step 7:

[1461] The server evaluates the user's mental and physical health based on the analysis results and generates appropriate health advice.

[1462] Input: Emotion category (facial and vocal results)

[1463] Output: Health advice

[1464] Step 8:

[1465] The server sends the generated health advice to the terminal.

[1466] Enter: Health Advice

[1467] Output: Advice data is sent to the terminal.

[1468] Step 9:

[1469] The terminal notifies the user of the generated health advice.

[1470] Input: Advice data

[1471] Output: Advice notification display, audio playback (using gTTS and playsound)

[1472] Step 10:

[1473] The user checks the health advice provided and takes necessary action.

[1474] Input: Advice content

[1475] Output: User actions (taking a break, taking a deep breath, practicing relaxation techniques, etc.)

[1476] Step 11:

[1477] The server generates fashion advice based on weather information and the user's schedule.

[1478] Input: Weather information, user schedule

[1479] Output: Fashion advice

[1480] Step 12:

[1481] The terminal notifies the user of the generated fashion advice.

[1482] Enter: fashion advice

[1483] Output: Fashion advice notification display

[1484] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1485] The present invention is a system that acquires and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status, and provides appropriate health advice and self-care recommendations, in addition to providing an emotion engine that recognizes the user's emotions. The operation of the entire system is described in detail below.

[1486] System Configuration

[1487] This system is broadly composed of three elements: the server, the terminal (user's smartphone), and the user. Each element and its role are explained below.

[1488] server

[1489] The server is the core of the system and performs the following functions:

[1490] 1. Analyze the user's facial expression and voice data.

[1491] 2. Evaluate the user's mental and physical health based on the analysis results.

[1492] 3. Generate health advice and self-care recommendations to provide to users.

[1493] 4. Offer fashion advice when needed.

[1494] 5. Analyze user sentiment in real time using an emotion engine.

[1495] 6. Provide appropriate emotional feedback based on the results of emotion analysis.

[1496] Terminal

[1497] The terminal provides the user interface and has the following roles:

[1498] 1. Use a camera to capture facial expression data.

[1499] 2. Use a microphone to capture audio data.

[1500] 3. The server notifies the user of the analysis results and advice.

[1501] 4. Receives input from the user and sends it to the server.

[1502] User

[1503] Users use the system to receive assessments and advice on their health status.

[1504] 1. Enter basic information (age, gender, preferences, daily activity patterns, etc.) as the initial setting.

[1505] 2. Use a camera and microphone to provide facial and voice data.

[1506] 3. Follow the health advice and self-care recommendations provided by the server.

[1507] Specific examples of processing

[1508] Here we will explain a specific scenario in which the system actually works.

[1509] 1. Initial Setup

[1510] The user downloads the smartphone app and enters basic information.

[1511] The terminal sends this information to the server.

[1512] The server stores the basic information in a database for future analysis.

[1513] 2. Acquisition and analysis of facial and speech data

[1514] Users take selfies in front of their smartphone cameras on a daily basis.

[1515] The device sends the selfie (facial expression data) to the server.

[1516] A user speaks into a smartphone to capture audio data.

[1517] The terminal transmits the acquired voice data to the server.

[1518] 3. Emotion analysis using an emotion engine

[1519] The server uses the received facial expression data and voice data to analyze the user's emotions in real time using an emotion engine.

[1520] The server evaluates the user's emotional state based on the emotion analysis results, detecting emotions such as joy, sadness, and anger.

[1521] The server generates appropriate emotional feedback based on the emotion analysis results, e.g., "You seem a little tired today. Please stretch to relax."

[1522] 4. Health advice and self-care recommendations

[1523] The server integrates the analysis results of the facial expression data and voice data with the emotion evaluation results to comprehensively evaluate the user's mental and physical health.

[1524] The server generates specific health advice and self-care recommendations based on the assessment results.

[1525] The server generates advice and recommendations and sends them to the device.

[1526] The terminal notifies the user of the advice.

[1527] The user checks the advice provided and takes the necessary action.

[1528] 5. Providing fashion advice

[1529] The server generates fashion advice based on season, weather, and schedule information.

[1530] The terminal notifies the user of the generated fashion advice.

[1531] The user follows the advice and chooses appropriate clothing.

[1532] 6. Training progress and feedback

[1533] The server evaluates the user's progress and training effectiveness in real time and updates the next training menu.

[1534] The device will notify and guide the user about the next training menu.

[1535] Users follow a training menu to effectively improve their communication skills.

[1536] In this way, each step of the system works in conjunction to provide comprehensive support for users' health management and improvement of communication skills. By combining the emotion engine, it is possible to evaluate mental and physical aspects with higher accuracy than before, and to provide more personalized advice and recommendations.

[1537] The processing flow will be explained below.

[1538] Step 1:

[1539] A user downloads and installs a smartphone app.

[1540] Step 2:

[1541] The user launches the app and enters basic information such as age, gender, preferences, and daily activity patterns on the basic settings screen.

[1542] Step 3:

[1543] The terminal sends the entered basic information to the server.

[1544] Step 4:

[1545] The server stores the received basic information in a database for future analysis.

[1546] Step 5:

[1547] Users take selfies in front of their smartphone cameras on a daily basis.

[1548] Step 6:

[1549] The device sends the selfie (facial expression data) to the server.

[1550] Step 7:

[1551] A user speaks into a smartphone to capture audio data.

[1552] Step 8:

[1553] The terminal transmits the acquired voice data to the server.

[1554] Step 9:

[1555] The server analyzes the received facial expression data using a facial recognition algorithm and emotion engine to assess the user's emotional state, specifically identifying emotions such as smile, sadness, anger, and surprise.

[1556] Step 10:

[1557] The voice data received by the server is also analyzed by the emotion engine to evaluate the tone, speed, and emotional expression of the voice, e.g., excited, calm, angry, etc.

[1558] Step 11:

[1559] The server combines the analysis results of facial expression data and voice data to perform an emotional evaluation, and also comprehensively evaluates the user's mental and physical health.

[1560] Step 12:

[1561] The server generates specific health advice based on the evaluation results, for example, "You seem a little tired today. Please stretch to relax."

[1562] Step 13:

[1563] The server sends the generated health advice and self-care recommendations to the terminal.

[1564] Step 14:

[1565] The terminal notifies the user of the health advice received from the server.

[1566] Step 15:

[1567] The user checks the advice provided and takes the necessary action.

[1568] Step 16:

[1569] Users take selfies periodically to provide data for checking the condition of their hair and beard.

[1570] Step 17:

[1571] The device sends the captured selfie data to the server.

[1572] Step 18:

[1573] The server analyzes the selfie data and determines if you need hair or beard care. Example: "I need a beard trim."

[1574] Step 19:

[1575] The server generates salon recommendations as needed and lists nearby salons.

[1576] Step 20:

[1577] The device notifies the user of recommended beauty salons. For example, "Here are some recommended beauty salons near you."

[1578] Step 21:

[1579] The user makes a reservation at the recommended hair salon via their smartphone.

[1580] Step 22:

[1581] The server generates fashion advice based on season, weather, and schedule information. Example: "It's cold today, so I recommend a thick coat."

[1582] Step 23:

[1583] The terminal notifies the user of the generated fashion advice.

[1584] Step 24:

[1585] The user follows the advice and chooses appropriate clothing.

[1586] Step 25:

[1587] The server evaluates the user's progress and training effectiveness in real time and updates the next training menu.

[1588] Step 26:

[1589] The device will notify and guide the user about the next training menu.

[1590] Step 27:

[1591] Users follow a training menu to effectively improve their communication skills.

[1592] In this way, each step of the system works in conjunction to provide comprehensive support for users' health management and improvement of communication skills. By combining the emotion engine, it is possible to evaluate mental and physical aspects with higher accuracy than before, and to provide more personalized advice and recommendations.

[1593] Example 2

[1594] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1595] Conventional health assessment systems have difficulty in comprehensively assessing a user's mental and physical health status in real time, and have been unable to provide appropriate feedback quickly. Furthermore, they do not adequately provide advice that takes into account the user's emotional state, making it difficult to provide personalized care.

[1596] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1597] In this invention, the server includes means for acquiring facial expression data of a user, means for acquiring voice data of the user, means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health state, means for providing the user with appropriate health advice based on the evaluation results, means for analyzing the user's emotions in real time using an emotion analysis engine, and means for providing the user with appropriate emotional feedback based on the analysis results, thereby making it possible to comprehensively evaluate the user's health state in real time and quickly provide personalized advice.

[1598] "User" refers to an individual who uses the system and receives health assessments and advice.

[1599] "Facial expression data" is digital data obtained using a camera that contains information about the user's facial expressions.

[1600] "Voice data" is digital data obtained by capturing voice information when a user speaks using a microphone.

[1601] "Analysis" refers to the act of processing information based on acquired data to quantitatively or qualitatively evaluate the user's health and emotional state.

[1602] "Health condition" is the result of a comprehensive evaluation of the user's mental and physical conditions.

[1603] "Health advice" refers to specific guidelines or recommended actions provided to the user based on the evaluation results.

[1604] An "emotion analysis engine" is software or a program for analyzing a user's emotional state in real time based on facial expression data and voice data.

[1605] "Emotional feedback" refers to specific countermeasures and advice provided to users based on the results of the emotion analysis engine.

[1606] "Self-care recommendations" are specific advice or recommended actions for self-care that a user can take by themselves.

[1607] "Fashion advice" is advice on clothing provided to users based on information such as the season, weather, and schedule.

[1608] A "terminal" is a device used by a user, and is a device that provides a user interface, including a smartphone, tablet, etc.

[1609] The "server" is the core part of this system, and is a computer system that analyzes and stores data, generates advice, and so on.

[1610] The present invention provides a system that collects and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status, and provides appropriate health advice, self-care recommendations, and emotional feedback based on real-time emotional analysis. Specific embodiments are described below.

[1611] System Configuration

[1612] This system is broadly composed of three elements: a server, a terminal (user's smartphone), and a user.

[1613] server

[1614] The server is the core of the system and performs the following functions:

[1615] 1. Analysis of the user's facial expression data and voice data: The server processes the facial expression data and voice data sent from the device using an analysis engine (e.g., Amazon Rekognition, Google Cloud Speech-to-Text).

[1616] 2. Evaluating the user's health status: Based on the results of the analysis engine, the user's mental and physical health status is evaluated.

[1617] 3. Generating health advice and self-care recommendations: Based on the health status assessment results, appropriate health advice and self-care recommendations are generated.

[1618] 4. Real-time emotion analysis using an emotion analysis engine: Facial expression data and voice data are used to activate the emotion analysis engine, which analyzes the user's emotions in real time.

[1619] 5. Providing emotional feedback: Based on the results of emotion analysis, appropriate emotional feedback is generated and sent to the device.

[1620] Device (smartphone)

[1621] The terminal is responsible for the following:

[1622] 1. Acquiring facial expression data: The user's facial expressions are captured by a camera and saved as facial expression data.

[1623] 2. Acquiring voice data: The user's speech is recorded with a microphone and saved as voice data.

[1624] 3. Data transmission: The acquired facial expression data and voice data are transmitted to the server.

[1625] 4. Notification of advice and feedback: Notify the user of the analysis results, advice, and emotional feedback received from the server.

[1626] User

[1627] The user does the following:

[1628] 1. Download the app and enter basic information: Download the smartphone app and enter basic information such as age, gender, preferences, and daily activity patterns.

[1629] 2. Providing facial expression data and voice data: Taking selfies in front of a smartphone camera in everyday life and speaking into the smartphone provides voice data.

[1630] 3. Review and act on advice: Take action based on the health advice, self-care recommendations, and emotional feedback provided by the server.

[1631] Examples of concrete examples and prompts

[1632] Specific examples

[1633] For example, when a user wakes up in the morning, they take a selfie with their smartphone and say, "Good morning, what are your plans for today?" The device immediately sends the selfie and voice data to the server. The server receives the data and analyzes it using its emotion engine. From the analysis results, it detects that the user looks a little tired, and generates advice such as, "You seem a little tired today. Please stretch to relax." The device notifies the user of this advice, and the user confirms the notification and takes time to relax early.

[1634] Prompt Sentence Examples

[1635] "Design a system where a user provides selfies and voice data, and a server analyzes this and generates appropriate health advice. For example, if the user is tired, provide advice on relaxation. Also, use an emotion analysis engine to evaluate emotions in real time."

[1636] In this way, each step of the system works in conjunction to provide comprehensive support for the user's health management and understanding of their emotional state. By combining the emotion engine, it is possible to assess mental and physical aspects more accurately than conventional systems, and to provide more personalized advice and recommendations.

[1637] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1638] Step 1:

[1639] Download the app and enter your basic information

[1640] A user downloads and installs a smartphone app.

[1641] A user opens the app and enters basic information such as age, gender, preferences, and daily activity patterns.

[1642] Input: User basic information

[1643] The terminal sends the basic information entered by the user to the server.

[1644] Output: Basic information sent to the server

[1645] Step 2:

[1646] Save basic information

[1647] The server checks the received basic information and stores it in the database.

[1648] Input: Basic information submitted

[1649] The server reviews the stored information and prepares it for future analysis.

[1650] Output: Basic information saved

[1651] Step 3:

[1652] Acquiring facial expression data

[1653] A user takes a selfie in front of the smartphone camera in daily life.

[1654] Input: Selfie image

[1655] Your device will temporarily store the selfie image in its internal memory.

[1656] Output: Facial expression data stored in temporary memory

[1657] Step 4:

[1658] Sending facial expression data

[1659] The device sends the temporarily stored selfie data to the server.

[1660] Input: Facial expression data stored in temporary memory

[1661] The server passes the received selfie data to the analysis engine.

[1662] Output: Facial expression data passed to the analysis engine

[1663] Step 5:

[1664] Acquiring audio data

[1665] The user speaks into the smartphone and voice data is captured.

[1666] Input: Audio data

[1667] The device temporarily stores the captured audio data in its internal memory.

[1668] Output: Audio data stored in temporary memory

[1669] Step 6:

[1670] Sending audio data

[1671] The terminal transmits the temporarily stored voice data to the server.

[1672] Input: Audio data stored in temporary memory

[1673] The server passes the received audio data to the analysis engine.

[1674] Output: Audio data passed to the analysis engine

[1675] Step 7:

[1676] Analysis of facial expression and voice data

[1677] The server uses the received facial expression data and voice data to instruct the emotion analysis engine to perform analysis.

[1678] Input: Facial expression and voice data passed to the analysis engine

[1679] An emotion analysis engine analyzes the data to detect emotional states such as joy, sadness, and anger.

[1680] Output: Detected emotional state

[1681] Step 8:

[1682] Generating Emotional Feedback

[1683] The server evaluates the user's emotional state based on the analysis results of the emotion engine and generates appropriate emotional feedback.

[1684] Input: Detected emotional state

[1685] The server generates emotional feedback and sends it to the device.

[1686] Output: Emotional feedback sent to the device

[1687] Step 9:

[1688] Feedback Notification

[1689] The terminal notifies the user of the received feedback.

[1690] Input: Emotional feedback sent to the device

[1691] The user checks the notification and takes the necessary action, such as stretching to relax.

[1692] Output: User takes action

[1693] Step 10:

[1694] Overall health assessment

[1695] The server integrates the analysis results of the facial expression data and voice data and the emotion evaluation results to comprehensively evaluate the user's mental and physical health.

[1696] Input: Analysis results of facial expression data and voice data, emotion evaluation results

[1697] Output: Overall health rating of the user

[1698] Step 11:

[1699] Generating health advice

[1700] The server generates specific health advice and self-care recommendations based on the assessment results.

[1701] Input: Overall health rating

[1702] The server generates advice and recommendations and sends them to the device.

[1703] Output: Health advice sent to device

[1704] Step 12:

[1705] Advice Notification

[1706] The terminal notifies the user of the generated advice.

[1707] The user checks the advice and takes necessary action, such as taking early rest.

[1708] Input: Health advice sent to device

[1709] Output: User takes action

[1710] (Application example 2)

[1711] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1712] Currently, there are no systems in factories that can monitor the mental and physical health of workers in real time and provide appropriate health advice and feedback. This increases the risk of workers suffering from overwork and stress. Furthermore, the difficulty of quickly understanding a worker's emotional state and taking appropriate measures poses a challenge to maintaining a safe and efficient work environment.

[1713] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1714] In this invention, the server includes means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health state, means for providing the user with appropriate health advice based on the evaluation results, and means for monitoring the health state of workers in the factory in real time and providing feedback, thereby enabling real-time monitoring of the health state of workers and appropriate feedback.

[1715] "Facial expression data" refers to digital information that can be analyzed by capturing a user's facial expressions in the form of an image or video.

[1716] "Voice data" refers to digital information that can be analyzed by capturing a user's voice or speech as an audio file.

[1717] "Analysis" refers to information processing and evaluation to determine the mental and physical state of the user based on the acquired facial expression data and voice data.

[1718] "User" refers to an individual who uses this system, including workers who perform work in a factory.

[1719] "Health" refers to the user's state of mental and physical well-being, including stress levels, fatigue levels, emotional state, and the like.

[1720] "Health advice" refers to specific guidelines or recommendations for action provided to users based on the analyzed assessment results.

[1721] "Self-care" refers to the self-management and self-treatment actions that users take to maintain and improve their health.

[1722] "Feedback" refers to real-time alerts and instructions provided to the user based on the analysis results.

[1723] "Factory" refers to an industrial facility and its environment where operations such as the manufacture of products take place.

[1724] "Real-time" refers to the fact that feedback is provided to the user almost immediately after data acquisition and analysis occurs.

[1725] MODE FOR CARRYING OUT THE INVENTION

[1726] The present invention is a system for analyzing the health status of workers in a factory in real time and providing appropriate health advice and self-care recommendations. A specific configuration and processing procedure for implementing the present invention are described below.

[1727] System Configuration

[1728] This system is broadly composed of three elements: a server, factory robots, and workers.

[1729] server

[1730] The server is the core of the system and performs the following functions:

[1731] 1. Analyze facial expression and voice data.

[1732] 2. Evaluate the mental and physical health of workers based on the analysis results.

[1733] 3. Generate health advice and self-care recommendations to provide to workers.

[1734] 4. Provide emotional feedback when appropriate.

[1735] 5. Use sentiment and speech analysis engines (e.g., Microsoft Azure Emotion API, Google Cloud Speech-to-Text API).

[1736] Factory robots

[1737] Factory robots perform the following roles while working on-site:

[1738] 1. Obtain facial expression data of the worker using a camera.

[1739] 2. Acquire the worker's voice data using a microphone.

[1740] 3. Send the acquired data to the server.

[1741] 4. The analysis results and advice from the server are notified to the worker.

[1742] Worker

[1743] Workers use the system to receive health management and feedback as follows:

[1744] 1. Data is collected automatically during daily operations.

[1745] 2. Check the health advice and feedback provided by the server.

[1746] 3. Take the necessary action.

[1747] Specific processing steps

[1748] Acquiring and Sending Data

[1749] The factory robot uses a camera and microphone to capture facial expression and voice data of the worker in real time, which is temporarily stored locally and periodically sent to a server.

[1750] Data analysis

[1751] The server analyzes the received facial expression and voice data using an emotion analysis engine and a voice analysis engine. As a result of the analysis, the worker's emotional state and health condition are evaluated.

[1752] Assessing health status and providing feedback

[1753] Based on the analysis results, the server comprehensively evaluates the mental and physical health of the worker and generates appropriate health advice and feedback. The generated feedback is provided to the worker in real time via the factory robot. For example, specific advice such as "You seem a little tired lately. I recommend you take a five-minute break. Take a deep breath and reduce stress" may be provided.

[1754] Example prompts to input to the generative AI model

[1755] For example, here is a sample prompt to input to the sentiment analysis engine:

[1756] You have received facial and audio data to analyze. Please extract the following information from it:

[1757] 1. Emotions detected from facial expressions (happiness, sadness, anger, anxiety, etc.).

[1758] 2. Emotions and stress levels extracted from audio data.

[1759] 3. Overall assessment of mental and physical health.

[1760] Use the results to generate appropriate self-care advice.

[1761] By using such prompt sentences, the emotion analysis engine and speech analysis engine on the server side can extract the desired analysis results and provide appropriate feedback.

[1762] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1763] Step 1:

[1764] A factory robot uses a camera to capture facial expression data of workers. The image data captured by the camera is input, and facial feature points are extracted through image processing. The extracted feature point data is output.

[1765] Step 2:

[1766] A factory robot uses a microphone to capture the voice data of a worker. At this time, voice recording is started and the recorded voice file becomes the input. The voice data is temporarily saved in local storage, and the saved voice file becomes the output.

[1767] Step 3:

[1768] The facial expression data and voice data stored in the local storage are sent to the server. The image data and voice data sent are input, and data transfer is performed by the client-side communication module. Data that has been received by the server is output.

[1769] Step 4:

[1770] The server analyzes the received facial expression data. At this time, image data is input and emotional information is extracted using an emotion analysis engine (e.g., Microsoft Azure Emotion API). The extracted emotional information is output.

[1771] Step 5:

[1772] The server analyzes the received voice data. The audio file is input, and a voice analysis engine (e.g., Google Cloud Speech-to-Text API) is used to analyze the voice tone and emotional state. The analyzed emotional and stress level information is output.

[1773] Step 6:

[1774] The server integrates the analysis results of the facial expression data and voice data to evaluate the overall health condition. At this time, emotional information and stress level are input, and the overall health condition is calculated by the health assessment algorithm. This evaluation result is the output.

[1775] Step 7:

[1776] The server generates appropriate health advice and self-care recommendations for the worker based on the evaluation results. The evaluation results are input, and specific action guidelines and recommendations are generated by the advice generation algorithm. This generated advice is the output.

[1777] Step 8:

[1778] The server sends the generated feedback to the factory robot. The advice content becomes the input and is sent to the factory robot via the communication module. The advice content that has been received by the factory robot becomes the output.

[1779] Step 9:

[1780] Factory robots provide advice and feedback to workers. The received feedback is input and notified to the worker via voice or display. The advice communicated to the worker is output.

[1781] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1782] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1783] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1784] [Fourth embodiment]

[1785] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1786] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1787] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1788] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1789] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1790] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1791] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1792] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1793] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1794] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1795] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1796] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1797] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1798] The present invention is a system that acquires and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status and provide appropriate health advice and self-care recommendations. Here, we will explain in detail the operation of the entire system.

[1799] System Configuration

[1800] This system is broadly composed of three elements: the server, the terminal (e.g., the user's smartphone), and the user. Each element and its role are explained below.

[1801] server

[1802] The server is the core of the system and performs the following functions:

[1803] 1. Analyze the user's facial expression and voice data.

[1804] 2. Evaluate the user's mental and physical health based on the analysis results.

[1805] 3. Generate health advice and self-care recommendations to provide to users.

[1806] 4. Offer fashion advice when needed.

[1807] Terminal

[1808] The terminal provides the user interface and has the following roles:

[1809] 1. Use a camera to capture facial expression data.

[1810] 2. Use a microphone to capture audio data.

[1811] 3. The server notifies the user of the analysis results and advice.

[1812] 4. Receives input from the user and sends it to the server.

[1813] User

[1814] Users use the system to receive assessments and advice on their health status.

[1815] 1. Enter basic information (age, gender, preferences, daily activity patterns, etc.) as the initial setting.

[1816] 2. Use a camera and microphone to provide facial and voice data.

[1817] 3. Follow the health advice and self-care recommendations provided by the server.

[1818] Specific examples of processing

[1819] Here we will explain a specific scenario in which the system actually works.

[1820] 1. Initial Setup

[1821] The user downloads the smartphone app and enters basic information.

[1822] The terminal sends this information to the server.

[1823] The server stores the basic information in a database for future analysis.

[1824] 2. Acquisition and analysis of facial and speech data

[1825] A user takes a selfie in front of the smartphone camera and also captures audio data.

[1826] The terminal transmits facial expression data and voice data to the server.

[1827] The server analyzes the data and assesses the user's mental and physical health.

[1828] The server generates appropriate health advice based on the evaluation results.

[1829] 3. Providing advice

[1830] The server generates health advice and self-care recommendations and sends them to the device.

[1831] The terminal notifies the user of the advice.

[1832] The user checks the advice provided and takes the necessary action.

[1833] 4. Providing fashion advice

[1834] The server generates fashion advice based on weather information and the user's schedule.

[1835] The terminal notifies the user of this.

[1836] The user follows the advice and chooses appropriate clothing.

[1837] In this way, this system effectively acquires and analyzes facial expression and voice data and provides advice based on the results to help users improve their communication skills and health management in their daily lives.

[1838] The processing flow will be explained below.

[1839] Step 1:

[1840] A user downloads and installs a smartphone app.

[1841] Step 2:

[1842] The user launches the app and enters basic information such as age, gender, preferences, and daily activity patterns on the basic settings screen.

[1843] Step 3:

[1844] The terminal sends the entered basic information to the server.

[1845] Step 4:

[1846] The server stores the received basic information in a database for future analysis.

[1847] Step 5:

[1848] Users take selfies in front of their smartphone cameras on a daily basis.

[1849] Step 6:

[1850] The device sends the selfie (facial expression data) to the server.

[1851] Step 7:

[1852] A user speaks into a smartphone to capture audio data.

[1853] Step 8:

[1854] The terminal transmits the acquired voice data to the server.

[1855] Step 9:

[1856] The server analyzes the received facial expression data to assess the user's emotional state, for example, detecting the degree of smile or sadness.

[1857] Step 10:

[1858] The server analyzes the received audio data and evaluates the tone and patterns of the voice, such as the pitch, rate, and emotional expression.

[1859] Step 11:

[1860] The server integrates the analysis results and assesses the user's mental and physical health.

[1861] Step 12:

[1862] The server generates specific health advice based on the assessment results, for example, "You seem a little tired today. I recommend you take a break."

[1863] Step 13:

[1864] The server sends the generated health advice to the terminal.

[1865] Step 14:

[1866] The terminal notifies the user of the health advice received from the server.

[1867] Step 15:

[1868] The user checks the advice provided and takes action as necessary.

[1869] Step 16:

[1870] Users take selfies periodically to provide data for checking the condition of their hair and beard.

[1871] Step 17:

[1872] The device sends the captured selfie data to the server.

[1873] Step 18:

[1874] The server analyzes the selfie data and determines whether hair or beard care is needed.

[1875] Step 19:

[1876] The server generates recommendations for beauty salons as needed and provides them to the user.

[1877] Step 20:

[1878] The device notifies the user of beauty salon recommendations.

[1879] Step 21:

[1880] The user makes a reservation at the recommended hair salon via their smartphone.

[1881] Step 22:

[1882] The server generates fashion advice based on season, weather, and schedule information.

[1883] Step 23:

[1884] The server transmits the generated fashion advice to the terminal.

[1885] Step 24:

[1886] The terminal notifies the user of fashion advice.

[1887] Step 25:

[1888] The user checks the provided fashion advice and selects an outfit.

[1889] Step 26:

[1890] The server evaluates the user's progress and training effectiveness in real time and updates the next training menu.

[1891] Step 27:

[1892] The device will notify and guide the user about the next training menu.

[1893] Step 28:

[1894] Users follow a training menu to effectively improve their communication skills.

[1895] In this way, each step of the system works in conjunction with one another to provide comprehensive support for the user's health management and improvement of communication skills.

[1896] Example 1

[1897] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1898] Conventional health management systems lack sufficient means for comprehensively evaluating a user's mental and physical health status, and it is difficult to efficiently provide personalized advice to each user. Furthermore, there are no systems that analyze a user's facial expressions and voice data to appropriately provide self-care recommendations or fashion advice. Therefore, there is a need to improve users' health management and comfort in their daily lives.

[1899] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1900] In this invention, the server includes means for acquiring facial expression data, means for acquiring voice data, means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health status, means for generating appropriate health advice for the user based on the evaluation results, and means for generating prompt sentences using a generative AI model to provide personalized advice to the user. This enables a comprehensive evaluation of the user's mental and physical health status and the provision of appropriate and personalized advice to each user. Furthermore, by recommending self-care and providing fashion advice as needed, the server achieves health management and improves the user's quality of life.

[1901] "Facial expression data" refers to data that digitally records a user's facial expressions.

[1902] "Voice data" refers to data that is a digital recording of a user's speech or voice.

[1903] "Analysis" refers to the process of using acquired digital data to identify features and patterns and evaluate the user's mental and physical state.

[1904] "Evaluating health status" means quantitatively and qualitatively determining the user's psychological and physical health status based on the analysis results of facial expression data and voice data.

[1905] "Generating health advice" means creating appropriate courses of action or recommendations based on the assessed health status of the user.

[1906] A "generative AI model" is an algorithm or system that uses artificial intelligence to generate text or instructions.

[1907] A "prompt sentence" is text data input into a generative AI model that describes the instructions and information needed to perform a specific task.

[1908] "Personalized advice" refers to specific and appropriate advice that is customized according to the characteristics and conditions of each individual user.

[1909] The present invention is a system that acquires and analyzes a user's facial expression and voice data to evaluate the user's mental and physical health and provide appropriate health advice and self-care recommendations. The invention has three main components: a server, a terminal, and a user.

[1910] System Configuration

[1911] server

[1912] The server is the core of the system and is responsible for:

[1913] 1. Receiving and storing data: Receives facial expression data and voice data sent from the user's device and stores them in a database.

[1914] 2. Data analysis: The received data is analyzed. Facial expression data is extracted using OpenCV, and voice data is converted to text using the Google Cloud Speech-to-Text API, followed by sentiment analysis using the spaCy library.

[1915] 3. Health assessment: Based on the analysis results, the user's mental and physical health status is assessed.

[1916] 4. Advice Generation: Using a generative AI model (e.g., GPT-3), prompts are generated based on the assessment results to create appropriate health advice and self-care recommendations, and provide fashion advice as needed.

[1917] Specific operation example

[1918] Initial setup and data acquisition

[1919] Users download the smartphone app and enter basic information (age, gender, preferences, daily activity patterns, etc.).

[1920] The terminal sends this information to the server, which stores it in a database.

[1921] Acquisition and analysis of facial expression and voice data

[1922] At designated times, the user takes a photo of their facial expression using the smartphone camera and records their voice using the microphone.

[1923] The device collects facial expression and voice data and sends them to the server in an appropriate format (e.g., JPEG images and WAV audio files).

[1924] The server analyzes the received data, extracting facial features using OpenCV, converting the audio data to text using the Google Cloud Speech-to-Text API, and performing emotion analysis using the spaCy library.

[1925] Health assessment and advice

[1926] Based on the analysis results, the server evaluates the user's mental and physical health.

[1927] The server sends the generated prompts to a generative AI model to create personalized health advice and self-care recommendations.

[1928] The server sends these advices to the terminal, which then notifies the user.

[1929] Specific prompt examples

[1930] “Analyze the user’s voice data and assess their emotional state. For example, if the user is tired, suggest a ‘relaxation exercise’.”

[1931] In this way, this system supports users in improving their communication skills and health management in their daily lives. By effectively analyzing facial expression and voice data and providing personalized advice based on the results, the system aims to comprehensively maintain and improve the user's health.

[1932] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1933] Step 1: Initial settings and user information registration

[1934] A user downloads and installs a smartphone app. As input, the user enters basic information (age, gender, preferences, daily activity patterns, etc.) into the app. The device collects this information and sends it to the server in JSON format. The server stores the received information in a database and uses it for future analysis. This allows user information to be accumulated in the database, making it possible to provide personalized services to each individual user.

[1935] Specific behavior:

[1936] A user installs the app on their smartphone, launches it, and enters some basic information. The device then makes an HTTP request to send this information to the server, which then stores it in a database.

[1937] Step 2: Acquire facial expression and voice data

[1938] At designated times, the user captures facial expression data using the smartphone camera and records audio data using the microphone. Camera image data and audio data are acquired as input. The device collects this data and sends it to the server as JPEG images and WAV audio files. The output is facial expression data and audio data.

[1939] Specific behavior:

[1940] When a user presses the "Get Data" button, the smartphone camera is activated and a picture of their face is taken. Then, a voice memo is recorded using the microphone. The device collects facial expression data in JPEG format and voice data in WAV format, and creates an HTTP request to send this to the server.

[1941] Step 3: Data analysis by the server

[1942] The server analyzes the received facial expression data using OpenCV and extracts feature points. It converts the voice data into text using the Google Cloud Speech-to-Text API and performs emotion analysis using an NLP library (e.g., spaCy). It receives facial expression data and voice data as input and obtains feature point data and emotion analysis results as output.

[1943] Specific behavior:

[1944] The server processes the received image data using the OpenCV library to extract facial features, converts the audio data into text using the Google Cloud Speech-to-Text API, and analyzes the resulting text data using the spaCy library to evaluate the emotional state. The results are then stored in a database.

[1945] Step 4: Health Assessment

[1946] The server evaluates the user's mental and physical health based on the analyzed facial expression and voice data. It uses feature point data and emotion analysis results as input and generates a health status evaluation result as output.

[1947] Specific behavior:

[1948] The server evaluates the user's health condition using a dedicated evaluation algorithm based on the analysis results. The evaluation results are stored in a database and used in the next step of generating advice.

[1949] Step 5: Generating and serving advice

[1950] The server sends prompts to a generative AI model (e.g., GPT-3) to generate appropriate health advice and self-care recommendations based on the assessment results. Fashion advice is also provided if necessary. The health assessment results are used as input, and the generated advice is obtained as output. The device notifies the user of this.

[1951] Specific behavior:

[1952] The server sends a prompt message to the AI ​​model saying, "The user's stress level is high. Please provide relaxation advice." The AI ​​model then collects the advice returned, converts it into JSON format, and sends it to the device. The device then generates a notification and displays it for the user to see.

[1953] The above is the specific processing flow of this system.

[1954] (Application example 1)

[1955] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1956] In traditional brick-and-mortar stores, it was difficult to grasp the mental and physical health status of customers and provide appropriate services and advice based on that. As a result, the quality of customer experience tended to be uniform, making it difficult to provide services that meet individual needs. This led to problems such as lower customer satisfaction and difficulty in securing repeat customers.

[1957] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1958] In this invention, the server includes means for acquiring facial expression data, means for acquiring voice data, means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health state, and means for providing the user with appropriate health advice based on the evaluation results. This makes it possible to grasp the mental and physical health state of customers in a physical store and provide appropriate health advice and self-care recommendations in real time.

[1959] "Facial expression data" is digital data about a person's facial expressions collected using a camera or other image capture device.

[1960] "Voice data" is digital data relating to a human voice collected using a microphone or other voice capture device.

[1961] "Analysis" refers to the act of processing acquired facial expression and voice data and applying computational processes and algorithms to understand, classify, and evaluate its content.

[1962] "Mental and physical health status" refers to the totality of information that indicates the user's psychological state and physical condition, including emotions, stress level, fatigue level, and the like.

[1963] "Health Advice" means specific recommendations or advice based on mental and physical health that are provided to users to help them maintain or improve their health in their daily lives.

[1964] "Self-care" refers to actions and efforts for health management that users themselves undertake, and includes relaxation techniques, reviewing diet and drink habits, exercise instruction, and the like.

[1965] "Fashion advice" refers to specific recommendations and suggestions to help users select appropriate clothing, accessories, etc. based on their evaluation results and circumstances.

[1966] "Real-time" refers to data acquisition, analysis, evaluation, and feedback occurring nearly simultaneously and without delay.

[1967] The present invention is a system that acquires and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status and provide appropriate health advice and self-care recommendations. Here, we will show a specific application example in a brick-and-mortar store and explain in detail an embodiment of the present invention.

[1968] System Configuration

[1969] This system is broadly composed of three elements: the server, the terminal (e.g., the user's smartphone), and the user. Each element and its role are explained below.

[1970] server

[1971] The server is the core of the system and performs the following functions:

[1972] 1. Analyze the user's facial expression and voice data.

[1973] The hardware and software used are TensorFlow / Keras.

[1974] 2. Evaluate the user's mental and physical health based on the analysis results.

[1975] 3. Generate health advice and self-care recommendations to provide to users.

[1976] A specific example would be to generate the advice "Listen to some relaxing music."

[1977] 4. Offer fashion advice when needed.

[1978] As an example of use, it generates advice based on weather information, such as "It's going to rain today, so please bring a waterproof jacket."

[1979] Terminal

[1980] The terminal provides the user interface and has the following roles:

[1981] 1. Use a camera to capture facial expression data.

[1982] The software used is OpenCV.

[1983] 2. Use a microphone to capture audio data.

[1984] The software used is librosa.

[1985] 3. The server notifies the user of the analysis results and advice.

[1986] To provide audio advice, gTTS (Google Text-to-Speech) is used and played on playsound.

[1987] 4. Receives input from the user and sends it to the server.

[1988] User

[1989] Users use the system to receive assessments and advice on their health status.

[1990] 1. Enter basic information (age, gender, preferences, daily activity patterns, etc.) as the initial setting.

[1991] 2. Use a camera and microphone to provide facial and voice data.

[1992] 3. Follow the health advice and self-care recommendations provided by the server.

[1993] Specific examples

[1994] Here we will explain a specific scenario in which the system actually works.

[1995] 1. Initial Setup

[1996] The user downloads the smartphone app and enters basic information.

[1997] The terminal sends this information to the server.

[1998] The server stores the basic information in a database for future analysis.

[1999] 2. Acquisition and analysis of facial and speech data

[2000] A user takes a selfie in front of the smartphone camera and also captures audio data.

[2001] The terminal transmits facial expression data and voice data to the server.

[2002] The server analyzes the data and assesses the user's mental and physical health.

[2003] The server generates appropriate health advice based on the evaluation results.

[2004] 3. Providing advice

[2005] The server generates health advice and self-care recommendations and sends them to the device.

[2006] The terminal notifies the user of the advice.

[2007] The user checks the advice provided and takes the necessary action.

[2008] 4. Providing fashion advice

[2009] The server generates fashion advice based on weather information and the user's schedule.

[2010] The terminal notifies the user of this.

[2011] The user follows the advice and chooses appropriate clothing.

[2012] Prompt Sentence Examples

[2013] Prompt: "Analyze the customer's facial and voice data to assess their current mental and physical state. Based on the assessment, generate effective relaxation advice."

[2014] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[2015] Step 1:

[2016] The user launches the smartphone app and enters basic information (age, gender, preferences, daily activity patterns, etc.).

[2017] Input: User's basic information (age, gender, preferences, daily activity patterns, etc.)

[2018] Output: The basic information entered is saved on the device.

[2019] Step 2:

[2020] The terminal sends the user's basic information to the server.

[2021] Input: User basic information

[2022] Output: Basic information is sent to the server.

[2023] Step 3:

[2024] The server stores the received basic information in a database.

[2025] Input: Basic information data

[2026] Output: Basic information stored in the database

[2027] Step 4:

[2028] A user takes a selfie in front of the smartphone camera and captures audio data.

[2029] Input: Facial image data and audio data

[2030] Output: The captured facial image data and recorded audio data are saved on the device.

[2031] Step 5:

[2032] The terminal transmits facial expression data and voice data to the server.

[2033] Input: Facial image data and audio data

[2034] Output: Facial expression data and voice data are sent to the server.

[2035] Step 6:

[2036] The server parses the received data.

[2037] Facial expression data: Preprocessed using OpenCV, analyzed using a Keras model, and output emotion categories (e.g., "happy," "sad," "surprised," etc.).

[2038] Input: Face image data

[2039] Data processing: Preprocessing with OpenCV (grayscale conversion, resizing, etc.)

[2040] Data Computation: Sentiment Analysis with Keras Models

[2041] Output: Emotion category

[2042] Audio data: Preprocessing is performed using librosa, features are extracted, and the data is analyzed using a Keras model. The emotional category is output.

[2043] Input: Audio data

[2044] Data processing: Feature extraction (MFCC, etc.) using librosa

[2045] Data Computation: Sentiment Analysis with Keras Models

[2046] Output: Emotion category

[2047] Step 7:

[2048] The server evaluates the user's mental and physical health based on the analysis results and generates appropriate health advice.

[2049] Input: Emotion category (facial and vocal results)

[2050] Output: Health advice

[2051] Step 8:

[2052] The server sends the generated health advice to the terminal.

[2053] Enter: Health Advice

[2054] Output: Advice data is sent to the terminal.

[2055] Step 9:

[2056] The terminal notifies the user of the generated health advice.

[2057] Input: Advice data

[2058] Output: Advice notification display, audio playback (using gTTS and playsound)

[2059] Step 10:

[2060] The user checks the health advice provided and takes necessary action.

[2061] Input: Advice content

[2062] Output: User actions (taking a break, taking a deep breath, practicing relaxation techniques, etc.)

[2063] Step 11:

[2064] The server generates fashion advice based on weather information and the user's schedule.

[2065] Input: Weather information, user schedule

[2066] Output: Fashion advice

[2067] Step 12:

[2068] The terminal notifies the user of the generated fashion advice.

[2069] Enter: fashion advice

[2070] Output: Fashion advice notification display

[2071] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2072] The present invention is a system that acquires and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status, and provides appropriate health advice and self-care recommendations, in addition to providing an emotion engine that recognizes the user's emotions. The operation of the entire system is described in detail below.

[2073] System Configuration

[2074] This system is broadly composed of three elements: the server, the terminal (user's smartphone), and the user. Each element and its role are explained below.

[2075] server

[2076] The server is the core of the system and performs the following functions:

[2077] 1. Analyze the user's facial expression and voice data.

[2078] 2. Evaluate the user's mental and physical health based on the analysis results.

[2079] 3. Generate health advice and self-care recommendations to provide to users.

[2080] 4. Offer fashion advice when needed.

[2081] 5. Analyze user sentiment in real time using an emotion engine.

[2082] 6. Provide appropriate emotional feedback based on the results of emotion analysis.

[2083] Terminal

[2084] The terminal provides the user interface and has the following roles:

[2085] 1. Use a camera to capture facial expression data.

[2086] 2. Use a microphone to capture audio data.

[2087] 3. The server notifies the user of the analysis results and advice.

[2088] 4. Receives input from the user and sends it to the server.

[2089] User

[2090] Users use the system to receive assessments and advice on their health status.

[2091] 1. Enter basic information (age, gender, preferences, daily activity patterns, etc.) as the initial setting.

[2092] 2. Use a camera and microphone to provide facial and voice data.

[2093] 3. Follow the health advice and self-care recommendations provided by the server.

[2094] Specific examples of processing

[2095] Here we will explain a specific scenario in which the system actually works.

[2096] 1. Initial Setup

[2097] The user downloads the smartphone app and enters basic information.

[2098] The terminal sends this information to the server.

[2099] The server stores the basic information in a database for future analysis.

[2100] 2. Acquisition and analysis of facial and speech data

[2101] Users take selfies in front of their smartphone cameras on a daily basis.

[2102] The device sends the selfie (facial expression data) to the server.

[2103] A user speaks into a smartphone to capture audio data.

[2104] The terminal transmits the acquired voice data to the server.

[2105] 3. Emotion analysis using an emotion engine

[2106] The server uses the received facial expression data and voice data to analyze the user's emotions in real time using an emotion engine.

[2107] The server evaluates the user's emotional state based on the emotion analysis results, detecting emotions such as joy, sadness, and anger.

[2108] The server generates appropriate emotional feedback based on the emotion analysis results, e.g., "You seem a little tired today. Please stretch to relax."

[2109] 4. Health advice and self-care recommendations

[2110] The server integrates the analysis results of the facial expression data and voice data with the emotion evaluation results to comprehensively evaluate the user's mental and physical health.

[2111] The server generates specific health advice and self-care recommendations based on the assessment results.

[2112] The server generates advice and recommendations and sends them to the device.

[2113] The terminal notifies the user of the advice.

[2114] The user checks the advice provided and takes the necessary action.

[2115] 5. Providing fashion advice

[2116] The server generates fashion advice based on season, weather, and schedule information.

[2117] The terminal notifies the user of the generated fashion advice.

[2118] The user follows the advice and chooses appropriate clothing.

[2119] 6. Training progress and feedback

[2120] The server evaluates the user's progress and training effectiveness in real time and updates the next training menu.

[2121] The device will notify and guide the user about the next training menu.

[2122] Users follow a training menu to effectively improve their communication skills.

[2123] In this way, each step of the system works in conjunction to provide comprehensive support for users' health management and improvement of communication skills. By combining the emotion engine, it is possible to evaluate mental and physical aspects with higher accuracy than before, and to provide more personalized advice and recommendations.

[2124] The processing flow will be explained below.

[2125] Step 1:

[2126] A user downloads and installs a smartphone app.

[2127] Step 2:

[2128] The user launches the app and enters basic information such as age, gender, preferences, and daily activity patterns on the basic settings screen.

[2129] Step 3:

[2130] The terminal sends the entered basic information to the server.

[2131] Step 4:

[2132] The server stores the received basic information in a database for future analysis.

[2133] Step 5:

[2134] Users take selfies in front of their smartphone cameras on a daily basis.

[2135] Step 6:

[2136] The device sends the selfie (facial expression data) to the server.

[2137] Step 7:

[2138] A user speaks into a smartphone to capture audio data.

[2139] Step 8:

[2140] The terminal transmits the acquired voice data to the server.

[2141] Step 9:

[2142] The server analyzes the received facial expression data using a facial recognition algorithm and emotion engine to assess the user's emotional state, specifically identifying emotions such as smile, sadness, anger, and surprise.

[2143] Step 10:

[2144] The voice data received by the server is also analyzed by the emotion engine to evaluate the tone, speed, and emotional expression of the voice, e.g., excited, calm, angry, etc.

[2145] Step 11:

[2146] The server combines the analysis results of facial expression data and voice data to perform an emotional evaluation, and also comprehensively evaluates the user's mental and physical health.

[2147] Step 12:

[2148] The server generates specific health advice based on the evaluation results, for example, "You seem a little tired today. Please stretch to relax."

[2149] Step 13:

[2150] The server sends the generated health advice and self-care recommendations to the terminal.

[2151] Step 14:

[2152] The terminal notifies the user of the health advice received from the server.

[2153] Step 15:

[2154] The user checks the advice provided and takes the necessary action.

[2155] Step 16:

[2156] Users take selfies periodically to provide data for checking the condition of their hair and beard.

[2157] Step 17:

[2158] The device sends the captured selfie data to the server.

[2159] Step 18:

[2160] The server analyzes the selfie data and determines if you need hair or beard care. Example: "I need a beard trim."

[2161] Step 19:

[2162] The server generates salon recommendations as needed and lists nearby salons.

[2163] Step 20:

[2164] The device notifies the user of recommended beauty salons. For example, "Here are some recommended beauty salons near you."

[2165] Step 21:

[2166] The user makes a reservation at the recommended hair salon via their smartphone.

[2167] Step 22:

[2168] The server generates fashion advice based on season, weather, and schedule information. Example: "It's cold today, so I recommend a thick coat."

[2169] Step 23:

[2170] The terminal notifies the user of the generated fashion advice.

[2171] Step 24:

[2172] The user follows the advice and chooses appropriate clothing.

[2173] Step 25:

[2174] The server evaluates the user's progress and training effectiveness in real time and updates the next training menu.

[2175] Step 26:

[2176] The device will notify and guide the user about the next training menu.

[2177] Step 27:

[2178] Users follow a training menu to effectively improve their communication skills.

[2179] In this way, each step of the system works in conjunction to provide comprehensive support for users' health management and improvement of communication skills. By combining the emotion engine, it is possible to evaluate mental and physical aspects with higher accuracy than before, and to provide more personalized advice and recommendations.

[2180] Example 2

[2181] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2182] Conventional health assessment systems have difficulty in comprehensively assessing a user's mental and physical health status in real time, and have been unable to provide appropriate feedback quickly. Furthermore, they do not adequately provide advice that takes into account the user's emotional state, making it difficult to provide personalized care.

[2183] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2184] In this invention, the server includes means for acquiring facial expression data of a user, means for acquiring voice data of the user, means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health state, means for providing the user with appropriate health advice based on the evaluation results, means for analyzing the user's emotions in real time using an emotion analysis engine, and means for providing the user with appropriate emotional feedback based on the analysis results, thereby making it possible to comprehensively evaluate the user's health state in real time and quickly provide personalized advice.

[2185] "User" refers to an individual who uses the system and receives health assessments and advice.

[2186] "Facial expression data" is digital data obtained using a camera that contains information about the user's facial expressions.

[2187] "Voice data" is digital data obtained by capturing voice information when a user speaks using a microphone.

[2188] "Analysis" refers to the act of processing information based on acquired data to quantitatively or qualitatively evaluate the user's health and emotional state.

[2189] "Health condition" is the result of a comprehensive evaluation of the user's mental and physical conditions.

[2190] "Health advice" refers to specific guidelines or recommended actions provided to the user based on the evaluation results.

[2191] An "emotion analysis engine" is software or a program for analyzing a user's emotional state in real time based on facial expression data and voice data.

[2192] "Emotional feedback" refers to specific countermeasures and advice provided to users based on the results of the emotion analysis engine.

[2193] "Self-care recommendations" are specific advice or recommended actions for self-care that a user can take by themselves.

[2194] "Fashion advice" is advice on clothing provided to users based on information such as the season, weather, and schedule.

[2195] A "terminal" is a device used by a user, and is a device that provides a user interface, including a smartphone, tablet, etc.

[2196] The "server" is the core part of this system, and is a computer system that analyzes and stores data, generates advice, and so on.

[2197] The present invention provides a system that collects and analyzes a user's facial expression data and voice data to evaluate the user's mental and physical health status, and provides appropriate health advice, self-care recommendations, and emotional feedback based on real-time emotional analysis. Specific embodiments are described below.

[2198] System Configuration

[2199] This system is broadly composed of three elements: a server, a terminal (user's smartphone), and a user.

[2200] server

[2201] The server is the core of the system and performs the following functions:

[2202] 1. Analysis of the user's facial expression data and voice data: The server processes the facial expression data and voice data sent from the device using an analysis engine (e.g., Amazon Rekognition, Google Cloud Speech-to-Text).

[2203] 2. Evaluating the user's health status: Based on the results of the analysis engine, the user's mental and physical health status is evaluated.

[2204] 3. Generating health advice and self-care recommendations: Based on the health status assessment results, appropriate health advice and self-care recommendations are generated.

[2205] 4. Real-time emotion analysis using an emotion analysis engine: Facial expression data and voice data are used to activate the emotion analysis engine, which analyzes the user's emotions in real time.

[2206] 5. Providing emotional feedback: Based on the results of emotion analysis, appropriate emotional feedback is generated and sent to the device.

[2207] Device (smartphone)

[2208] The terminal is responsible for the following:

[2209] 1. Acquiring facial expression data: The user's facial expressions are captured by a camera and saved as facial expression data.

[2210] 2. Acquiring voice data: The user's speech is recorded with a microphone and saved as voice data.

[2211] 3. Data transmission: The acquired facial expression data and voice data are transmitted to the server.

[2212] 4. Notification of advice and feedback: Notify the user of the analysis results, advice, and emotional feedback received from the server.

[2213] User

[2214] The user does the following:

[2215] 1. Download the app and enter basic information: Download the smartphone app and enter basic information such as age, gender, preferences, and daily activity patterns.

[2216] 2. Providing facial expression data and voice data: Taking selfies in front of a smartphone camera in everyday life and speaking into the smartphone provides voice data.

[2217] 3. Review and act on advice: Take action based on the health advice, self-care recommendations, and emotional feedback provided by the server.

[2218] Examples of concrete examples and prompts

[2219] Specific examples

[2220] For example, when a user wakes up in the morning, they take a selfie with their smartphone and say, "Good morning, what are your plans for today?" The device immediately sends the selfie and voice data to the server. The server receives the data and analyzes it using its emotion engine. From the analysis results, it detects that the user looks a little tired, and generates advice such as, "You seem a little tired today. Please stretch to relax." The device notifies the user of this advice, and the user confirms the notification and takes time to relax early.

[2221] Prompt Sentence Examples

[2222] "Design a system where a user provides selfies and voice data, and a server analyzes this and generates appropriate health advice. For example, if the user is tired, provide advice on relaxation. Also, use an emotion analysis engine to evaluate emotions in real time."

[2223] In this way, each step of the system works in conjunction to provide comprehensive support for the user's health management and understanding of their emotional state. By combining the emotion engine, it is possible to assess mental and physical aspects more accurately than conventional systems, and to provide more personalized advice and recommendations.

[2224] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2225] Step 1:

[2226] Download the app and enter your basic information

[2227] A user downloads and installs a smartphone app.

[2228] A user opens the app and enters basic information such as age, gender, preferences, and daily activity patterns.

[2229] Input: User basic information

[2230] The terminal sends the basic information entered by the user to the server.

[2231] Output: Basic information sent to the server

[2232] Step 2:

[2233] Save basic information

[2234] The server checks the received basic information and stores it in the database.

[2235] Input: Basic information submitted

[2236] The server reviews the stored information and prepares it for future analysis.

[2237] Output: Basic information saved

[2238] Step 3:

[2239] Acquiring facial expression data

[2240] A user takes a selfie in front of the smartphone camera in daily life.

[2241] Input: Selfie image

[2242] Your device will temporarily store the selfie image in its internal memory.

[2243] Output: Facial expression data stored in temporary memory

[2244] Step 4:

[2245] Sending facial expression data

[2246] The device sends the temporarily stored selfie data to the server.

[2247] Input: Facial expression data stored in temporary memory

[2248] The server passes the received selfie data to the analysis engine.

[2249] Output: Facial expression data passed to the analysis engine

[2250] Step 5:

[2251] Acquiring audio data

[2252] The user speaks into the smartphone and voice data is captured.

[2253] Input: Audio data

[2254] The device temporarily stores the captured audio data in its internal memory.

[2255] Output: Audio data stored in temporary memory

[2256] Step 6:

[2257] Sending audio data

[2258] The terminal transmits the temporarily stored voice data to the server.

[2259] Input: Audio data stored in temporary memory

[2260] The server passes the received audio data to the analysis engine.

[2261] Output: Audio data passed to the analysis engine

[2262] Step 7:

[2263] Analysis of facial expression and voice data

[2264] The server uses the received facial expression data and voice data to instruct the emotion analysis engine to perform analysis.

[2265] Input: Facial expression and voice data passed to the analysis engine

[2266] An emotion analysis engine analyzes the data to detect emotional states such as joy, sadness, and anger.

[2267] Output: Detected emotional state

[2268] Step 8:

[2269] Generating Emotional Feedback

[2270] The server evaluates the user's emotional state based on the analysis results of the emotion engine and generates appropriate emotional feedback.

[2271] Input: Detected emotional state

[2272] The server generates emotional feedback and sends it to the device.

[2273] Output: Emotional feedback sent to the device

[2274] Step 9:

[2275] Feedback Notification

[2276] The terminal notifies the user of the received feedback.

[2277] Input: Emotional feedback sent to the device

[2278] The user checks the notification and takes the necessary action, such as stretching to relax.

[2279] Output: User takes action

[2280] Step 10:

[2281] Overall health assessment

[2282] The server integrates the analysis results of the facial expression data and voice data and the emotion evaluation results to comprehensively evaluate the user's mental and physical health.

[2283] Input: Analysis results of facial expression data and voice data, emotion evaluation results

[2284] Output: Overall health rating of the user

[2285] Step 11:

[2286] Generating health advice

[2287] The server generates specific health advice and self-care recommendations based on the assessment results.

[2288] Input: Overall health rating

[2289] The server generates advice and recommendations and sends them to the device.

[2290] Output: Health advice sent to device

[2291] Step 12:

[2292] Advice Notification

[2293] The terminal notifies the user of the generated advice.

[2294] The user checks the advice and takes necessary action, such as taking early rest.

[2295] Input: Health advice sent to device

[2296] Output: User takes action

[2297] (Application example 2)

[2298] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2299] Currently, there are no systems in factories that can monitor the mental and physical health of workers in real time and provide appropriate health advice and feedback. This increases the risk of workers suffering from overwork and stress. Furthermore, the difficulty of quickly understanding a worker's emotional state and taking appropriate measures poses a challenge to maintaining a safe and efficient work environment.

[2300] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[2301] In this invention, the server includes means for analyzing the acquired facial expression data and voice data to evaluate the user's mental and physical health state, means for providing the user with appropriate health advice based on the evaluation results, and means for monitoring the health state of workers in the factory in real time and providing feedback, thereby enabling real-time monitoring of the health state of workers and appropriate feedback.

[2302] "Facial expression data" refers to digital information that can be analyzed by capturing a user's facial expressions in the form of an image or video.

[2303] "Voice data" refers to digital information that can be analyzed by capturing a user's voice or speech as an audio file.

[2304] "Analysis" refers to information processing and evaluation to determine the mental and physical state of the user based on the acquired facial expression data and voice data.

[2305] "User" refers to an individual who uses this system, including workers who perform work in a factory.

[2306] "Health" refers to the user's state of mental and physical well-being, including stress levels, fatigue levels, emotional state, and the like.

[2307] "Health advice" refers to specific guidelines or recommendations for action provided to users based on the analyzed assessment results.

[2308] "Self-care" refers to the self-management and self-treatment actions that users take to maintain and improve their health.

[2309] "Feedback" refers to real-time alerts and instructions provided to the user based on the analysis results.

[2310] "Factory" refers to an industrial facility and its environment where operations such as the manufacture of products take place.

[2311] "Real-time" refers to the fact that feedback is provided to the user almost immediately after data acquisition and analysis occurs.

[2312] MODE FOR CARRYING OUT THE INVENTION

[2313] The present invention is a system for analyzing the health status of workers in a factory in real time and providing appropriate health advice and self-care recommendations. A specific configuration and processing procedure for implementing the present invention are described below.

[2314] System Configuration

[2315] This system is broadly composed of three elements: a server, factory robots, and workers.

[2316] server

[2317] The server is the core of the system and performs the following functions:

[2318] 1. Analyze facial expression and voice data.

[2319] 2. Evaluate the mental and physical health of workers based on the analysis results.

[2320] 3. Generate health advice and self-care recommendations to provide to workers.

[2321] 4. Provide emotional feedback when appropriate.

[2322] 5. Use sentiment and speech analysis engines (e.g., Microsoft Azure Emotion API, Google Cloud Speech-to-Text API).

[2323] Factory robots

[2324] Factory robots perform the following roles while working on-site:

[2325] 1. Obtain facial expression data of the worker using a camera.

[2326] 2. Acquire the worker's voice data using a microphone.

[2327] 3. Send the acquired data to the server.

[2328] 4. The analysis results and advice from the server are notified to the worker.

[2329] Worker

[2330] Workers use the system to receive health management and feedback as follows:

[2331] 1. Data is collected automatically during daily operations.

[2332] 2. Check the health advice and feedback provided by the server.

[2333] 3. Take the necessary action.

[2334] Specific processing steps

[2335] Acquiring and Sending Data

[2336] The factory robot uses a camera and microphone to capture facial expression and voice data of the worker in real time, which is temporarily stored locally and periodically sent to a server.

[2337] Data analysis

[2338] The server analyzes the received facial expression and voice data using an emotion analysis engine and a voice analysis engine. As a result of the analysis, the worker's emotional state and health condition are evaluated.

[2339] Assessing health status and providing feedback

[2340] Based on the analysis results, the server comprehensively evaluates the mental and physical health of the worker and generates appropriate health advice and feedback. The generated feedback is provided to the worker in real time via the factory robot. For example, specific advice such as "You seem a little tired lately. I recommend you take a five-minute break. Take a deep breath and reduce stress" may be provided.

[2341] Example prompts to input to the generative AI model

[2342] For example, here is a sample prompt to input to the sentiment analysis engine:

[2343] You have received facial and audio data to analyze. Please extract the following information from it:

[2344] 1. Emotions detected from facial expressions (happiness, sadness, anger, anxiety, etc.).

[2345] 2. Emotions and stress levels extracted from audio data.

[2346] 3. Overall assessment of mental and physical health.

[2347] Use the results to generate appropriate self-care advice.

[2348] By using such prompt sentences, the emotion analysis engine and speech analysis engine on the server side can extract the desired analysis results and provide appropriate feedback.

[2349] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2350] Step 1:

[2351] A factory robot uses a camera to capture facial expression data of workers. The image data captured by the camera is input, and facial feature points are extracted through image processing. The extracted feature point data is output.

[2352] Step 2:

[2353] A factory robot uses a microphone to capture the voice data of a worker. At this time, voice recording is started and the recorded voice file becomes the input. The voice data is temporarily saved in local storage, and the saved voice file becomes the output.

[2354] Step 3:

[2355] The facial expression data and voice data stored in the local storage are sent to the server. The image data and voice data sent are input, and data transfer is performed by the client-side communication module. Data that has been received by the server is output.

[2356] Step 4:

[2357] The server analyzes the received facial expression data. At this time, image data is input and emotional information is extracted using an emotion analysis engine (e.g., Microsoft Azure Emotion API). The extracted emotional information is output.

[2358] Step 5:

[2359] The server analyzes the received voice data. The audio file is input, and a voice analysis engine (e.g., Google Cloud Speech-to-Text API) is used to analyze the voice tone and emotional state. The analyzed emotional and stress level information is output.

[2360] Step 6:

[2361] The server integrates the analysis results of the facial expression data and voice data to evaluate the overall health condition. At this time, emotional information and stress level are input, and the overall health condition is calculated by the health assessment algorithm. This evaluation result is the output.

[2362] Step 7:

[2363] The server generates appropriate health advice and self-care recommendations for the worker based on the evaluation results. The evaluation results are input, and specific action guidelines and recommendations are generated by the advice generation algorithm. This generated advice is the output.

[2364] Step 8:

[2365] The server sends the generated feedback to the factory robot. The advice content becomes the input and is sent to the factory robot via the communication module. The advice content that has been received by the factory robot becomes the output.

[2366] Step 9:

[2367] Factory robots provide advice and feedback to workers. The received feedback is input and notified to the worker via voice or display. The advice communicated to the worker is output.

[2368] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2369] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2370] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2371] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2372] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2373] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2374] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2375] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2376] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2377] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2378] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2379] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2380] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2381] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2382] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2383] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2384] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2385] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2386] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2387] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2388] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2389] The following is further disclosed regarding the above embodiment.

[2390] (Claim 1)

[2391] A means for acquiring facial expression data;

[2392] means for acquiring audio data;

[2393] A means for analyzing the acquired facial expression data and voice data to evaluate the mental and physical health state of the user;

[2394] means for providing appropriate health advice to the user based on the evaluation results;

[2395] A system including:

[2396] (Claim 2)

[2397] 10. The system of claim 1, further comprising means for providing self-care recommendations to the user based on the facial expression data and voice data.

[2398] (Claim 3)

[2399] 10. The system according to claim 1, further comprising means for providing fashion advice to the user based on the evaluation result.

[2400] "Example 1"

[2401] (Claim 1)

[2402] A means for acquiring facial expression data;

[2403] means for acquiring audio data;

[2404] A means for analyzing the acquired facial expression data and voice data to evaluate the mental and physical health state of the user;

[2405] means for generating appropriate health advice for the user based on the evaluation results;

[2406] a means for generating prompts using a generative AI model to provide personalized advice to a user;

[2407] A system including:

[2408] (Claim 2)

[2409] 10. The system of claim 1, wherein the system provides self-care recommendations to the user based on facial expression data and voice data.

[2410] (Claim 3)

[2411] 2. The system according to claim 1, wherein fashion advice is provided to the user based on the evaluation results.

[2412] "Application Example 1"

[2413] (Claim 1)

[2414] A means for acquiring facial expression data;

[2415] means for acquiring audio data;

[2416] A means for analyzing the acquired facial expression data and voice data to evaluate the mental and physical health state of the user;

[2417] means for providing appropriate health advice to the user based on the evaluation results;

[2418] A means for providing appropriate self-care advice to users in real time based on the analysis results;

[2419] A system including:

[2420] (Claim 2)

[2421] The system according to claim 1, wherein self-care recommendations are made to the user based on the facial expression data and voice data.

[2422] (Claim 3)

[2423] The system according to claim 1, wherein fashion advice is provided to the user based on the evaluation results.

[2424] "Example 2: Combining Emotion Engines"

[2425] (Claim 1)

[2426] means for acquiring facial expression data of a user;

[2427] means for acquiring user voice data;

[2428] A means for analyzing the acquired facial expression data and voice data to evaluate the mental and physical health state of the user;

[2429] means for providing appropriate health advice to the user based on the evaluation results;

[2430] a means for analyzing user sentiment in real time using a sentiment analysis engine;

[2431] means for providing appropriate emotional feedback to the user based on the analysis results;

[2432] A system including:

[2433] (Claim 2)

[2434] 10. The system of claim 1, further comprising means for providing self-care recommendations to the user based on the user's facial expression data and voice data.

[2435] (Claim 3)

[2436] 10. The system according to claim 1, further comprising means for providing fashion advice to the user based on the evaluation result.

[2437] "Application example 2 when combining emotion engines"

[2438] (Claim 1)

[2439] A means for acquiring facial expression data;

[2440] means for acquiring audio data;

[2441] A means for analyzing the acquired facial expression data and voice data to evaluate the mental and physical health state of the user;

[2442] means for providing appropriate health advice to the user based on the evaluation results;

[2443] a means for monitoring and providing feedback on the health status of workers in a factory in real time;

[2444] A system including:

[2445] (Claim 2)

[2446] 10. The system of claim 1, further comprising means for assessing the health status of workers in a factory and providing appropriate self-care advice.

[2447] (Claim 3)

[2448] 10. The system of claim 1, further comprising means for providing emotional feedback based on voice data and facial expression data of the worker. [Explanation of symbols]

[2449] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for acquiring facial expression data; means for acquiring audio data; A means for analyzing the acquired facial expression data and voice data to evaluate the mental and physical health state of the user; means for providing appropriate health advice to the user based on the evaluation results; A system including:

2. The system of claim 1 , further comprising means for providing self-care recommendations to the user based on the facial expression data and voice data.

3. The system according to claim 1 , further comprising means for providing fashion advice to the user based on the evaluation result.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A