System

The AI-equipped mirror system addresses the challenge of real-time health monitoring by integrating facial image and voice analysis to provide comprehensive health assessments, improving daily health management through continuous and early detection of abnormalities.

JP2026026906APending Publication Date: 2026-02-18SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024129327
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-05
Publication Date
2026-02-18

AI Technical Summary

Technical Problem

Conventional health monitoring methods rely on periodic examinations and self-reporting, making it difficult to monitor health status in real time and continuously, and there is a lack of technology that can centrally manage multiple health indicators in an easy-to-use manner.

Method used

An AI-equipped mirror system that captures facial images, analyzes biometric information, engages in voice dialogue, integrates data for comprehensive health assessment, and generates health reports, allowing users to receive a multifaceted health assessment in real time by standing in front of the mirror.

Benefits of technology

Enables continuous, real-time health monitoring and early detection of abnormalities by integrating facial image analysis, voice interaction, and data integration to provide personalized health reports, enhancing daily health management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026026906000001_ABST
    Figure 2026026906000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes a means for acquiring a video image of a face, a means for analyzing the video image of the face to measure biological information, a means for performing voice interaction with a user, a means for analyzing voice data to evaluate a health state, a means for integrating acquired data to perform comprehensive health evaluation, a means for generating a health report, and a means for displaying the generated health report.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, it is becoming increasingly important for individuals and companies to understand their health status in real time and detect abnormalities early. However, conventional health monitoring methods rely on periodic examinations and self-reporting, making it difficult to monitor health status in real time and continuously. Furthermore, there is a lack of technology that can centrally manage many health indicators and provide them to users in an easy-to-use manner. Therefore, the purpose of this invention is to provide a system that can be naturally integrated into daily life and enables continuous, multifaceted health assessment. [Means for solving the problem]

[0005] The present invention provides an AI-equipped mirror system. This system includes a means for acquiring a facial image, a means for analyzing the facial image to measure biometric information, a means for engaging in voice dialogue with a user, a means for analyzing the voice data to evaluate the user's health, a means for integrating the acquired data to perform a comprehensive health assessment, a means for generating a health report, and a means for displaying the generated health report. The system also includes a means for acquiring the user's lifestyle information through voice dialogue, and a means for analyzing the facial image to measure the user's heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, hemoglobin concentration, stress, and fatigue level. In this way, a user can receive a multifaceted health assessment in real time simply by standing in front of the mirror, enabling health maintenance and early detection of abnormalities in daily life.

[0006] "Means for acquiring facial images" refers to technical means for capturing the user's face using a camera or other photographic device.

[0007] "Means for measuring biometric information by analyzing facial images" refers to a technical means for analyzing captured facial images and deriving biometric information such as heart rate, blood pressure, and body temperature.

[0008] "Means for conducting voice dialogue with the user" refers to technical means including a microphone, speaker, voice recognition technology, voice synthesis technology, etc. for communicating with the user by voice.

[0009] The "means for analyzing voice data to assess health status" refers to a technical means for analyzing acquired voice data and assessing health status such as stress and fatigue level.

[0010] "Means for integrating acquired data to conduct a comprehensive health assessment" refers to a technical means for integrating facial image data, audio data, lifestyle habit data, etc. to conduct a comprehensive health assessment.

[0011] The "means for generating a health report" refers to a technical means for generating a report for reporting the user's health condition based on the integrated data.

[0012] The "means for displaying the generated health report" refers to technical means including a display, an audio output device, etc. for visually or audibly presenting the generated health report to the user.

[0013] The "means for acquiring lifestyle information of a user through voice dialogue" refers to a technical means for collecting information on lifestyle habits such as sleep, diet, and exercise through voice dialogue with a user.

[0014] "Means for measuring heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, hemoglobin concentration, stress, and fatigue level" refers to a technological means for analyzing facial images to measure these biometric indicators. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] This invention relates to an AI-equipped mirror system that analyzes facial images and audio to assess the user's health condition. This system allows the user to receive a multifaceted health assessment in real time by standing in front of the mirror.

[0037] System Overview

[0038] The system includes the following main elements:

[0039] 1. Camera that captures facial images

[0040] 2. Processor for video analysis and biometric measurement

[0041] 3. A microphone and speaker for voice interaction with the user

[0042] 4. Software that analyzes voice data to assess health status

[0043] 5. Server that integrates data and provides comprehensive health assessment

[0044] 6. Display to generate and display health reports

[0045] Program processing flow

[0046] The program of this system is designed as follows:

[0047] 1. The user stands in front of the AI-enabled mirror

[0048] Every morning, the user stands in front of the mirror and looks at their face, which automatically activates the system.

[0049] 2. The device recognizes the user's face

[0050] The device uses a built-in camera to capture an image of your face and uses a facial recognition algorithm to identify you, which includes matching it against a historical database.

[0051] 3. The device analyzes your face and muscle movements

[0052] The device uses a high-resolution camera to capture detailed facial video data, allowing it to analyze and measure biometric information such as heart rate, blood pressure, and body temperature.

[0053] 4. The device starts a dialogue with the user

[0054] The device initiates a voice dialogue, asking a question such as, "How much sleep did you get last night?" When the user responds, the voice data is acquired.

[0055] 5. The device analyzes the voice data

[0056] The device analyzes the user's voice and assesses stress and fatigue levels based on the tone and speed of the voice.

[0057] 6. The server integrates and analyzes the data

[0058] The server integrates the facial video data, audio data, and question-and-answer data sent from the device to perform a comprehensive health assessment.

[0059] 7. The server generates a health report

[0060] The server generates a health report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[0061] 8. The device will display a health report

[0062] The device will display a health report on the mirror and issue warnings as needed, such as telling the user, "You're feeling stressed today. We recommend taking some time to relax."

[0063] Specific examples

[0064] 1. The user stands in front of the mirror

[0065] Every morning, when the user stands in front of the mirror, facial images are automatically captured.

[0066] 2. Facial Recognition

[0067] The device captures the user's face through the camera and identifies them as "Yamada-san" by comparing it with a previously registered database.

[0068] 3. Video Analysis

[0069] The camera captures detailed images of the user's face, and the processor measures heart rate, blood pressure, body temperature, etc. At the same time, subtle changes in facial color and facial expressions are analyzed.

[0070] 4. Starting a voice conversation

[0071] The device uses a microphone and speaker to ask, "How much sleep did you get last night?" The user responds, "I slept for 6 hours."

[0072] 5. Analysis of audio data

[0073] The device analyzes the user's voice data and detects stress levels from the tone and speed of the voice.

[0074] 6. Data integration and analysis

[0075] The collected facial video data, audio data, and lifestyle data are sent to a server for a comprehensive health assessment.

[0076] 7. Generate Health Reports

[0077] Based on the evaluation results, the server generates a health report for the user, including a warning such as, "Your heart rate is high today. You may need to rest."

[0078] 8. View Health Report

[0079] The device displays a health report on the mirror and notifies the user with an automated voice.

[0080] In this way, users can easily monitor their daily health status and take necessary measures. The present invention significantly improves daily health management without relying solely on regular medical checkups.

[0081] The processing flow will be explained below.

[0082] Step 1:

[0083] The user stands in front of the mirror, which automatically activates the system and prepares the camera to capture the user's face.

[0084] Step 2:

[0085] The device uses its built-in camera to capture the user's face, then uses a facial recognition algorithm to identify the user and match them with a historical database.

[0086] Step 3:

[0087] The device uses a high-resolution camera to capture detailed facial video data, including technology to capture facial biometric information such as heart rate, blood pressure, and temperature.

[0088] Step 4:

[0089] The device starts a voice dialogue with the user. For example, it asks, "Good morning. How long did you sleep today?" The user replies, "I slept for 6 hours."

[0090] Step 5:

[0091] The device analyzes the voice data it acquires, and uses voice analysis technology to assess stress and fatigue levels based on the tone and speed of the user's voice.

[0092] Step 6:

[0093] The device analyzes facial expressions and changes in color tone to measure heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, hemoglobin concentration, stress, and fatigue level.

[0094] Step 7:

[0095] The facial image data, audio data, and question and answer data acquired by the terminal are transmitted to the server.

[0096] Step 8:

[0097] The server then combines the received data to provide a comprehensive health assessment, including heart rate and blood pressure variability, voice analysis results, and lifestyle information.

[0098] Step 9:

[0099] The server generates a health report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[0100] Step 10:

[0101] The device receives the health report sent from the server and displays it on the mirror, for example, displaying a message such as "Your heart rate is high today. You may need to rest."

[0102] Step 11:

[0103] The device will notify the user with an automated voice, providing feedback such as, "Good morning. You're feeling stressed today. We recommend you relax."

[0104] Example 1

[0105] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0106] Conventional health management systems require users to manually input information, which can lead to problems with accuracy and effort. In addition, the long intervals between regular health checkups make it difficult to quickly detect changes in health status, making it impossible to grasp health status in real time.

[0107] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0108] In this invention, the server includes means for acquiring a facial image, means for analyzing the acquired facial image to measure biometric information, means for engaging in voice dialogue with the user, means for analyzing the voice data to evaluate the health condition, means for integrating the acquired video data and voice data to perform a comprehensive health evaluation, means for generating a health report, and means for displaying the generated health report, thereby enabling the user to receive a comprehensive health evaluation in real time simply by standing in front of the mirror.

[0109] "Means for capturing facial images" refers to a camera or other imaging device that captures an image or video of a user's face.

[0110] "Means for analyzing captured facial images and measuring biometric information" refers to processors and algorithms that analyze and measure biometric information such as heart rate, blood pressure, and body temperature from captured facial image data.

[0111] "Means for conducting voice dialogue with the user" refers to a microphone and speaker for communicating with the user by voice, and software for controlling the dialogue.

[0112] "Means for analyzing voice data to assess health status" refers to algorithms or analytical software for analyzing voice data and assessing a user's stress level, fatigue level, etc.

[0113] "Means for integrating acquired video and audio data to perform a comprehensive health assessment" refers to a server and analysis program for integrating facial video and audio data to assess overall health status.

[0114] "Means for generating a health report" refers to software or algorithms that automatically generate a report showing the user's health status based on the analysis results.

[0115] "Means for displaying the generated health report" refers to a display or screen for visually presenting the generated health report to the user.

[0116] This invention relates to an AI-equipped mirror system that analyzes the user's facial image and voice to evaluate their health condition in real time. To implement this system, the following hardware and software are used:

[0117] Hardware and software used

[0118] 1. Camera: Used to capture video of the user's face. For example, a 1080p HD camera or a 4K camera is suitable.

[0119] 2. Processor: Used for video analysis and biometric measurement. Examples include an Intel Core i7 processor or an NVIDIA GPU.

[0120] 3. Microphone and speaker: Used for voice interaction with the user, a high-sensitivity microphone and stereo speakers are desirable to provide clear sound quality.

[0121] 4. Analysis software: Software used to analyze video and audio data often uses OpenCV, TensorFlow, or Google Cloud Speech-to-Text API.

[0122] 5. Server: A backend server that integrates data and performs comprehensive health assessments. This is where AI models using TensorFlow and PyTorch run.

[0123] 6. Display: Used to display the generated health report. A display that also functions as a mirror is desirable.

[0124] How the system works

[0125] This system automatically starts when the user stands in front of the mirror and performs the following processes.

[0126] Image acquisition by camera

[0127] When a user stands in front of the mirror, the camera automatically captures an image of the user's face.

[0128] Facial image analysis

[0129] The processor analyzes the captured facial video data. For example, it detects subtle color changes and facial expressions, and uses this data to estimate heart rate, blood pressure, and body temperature. This analysis uses the OpenCV library.

[0130] Voice interaction with the user

[0131] The device uses a microphone and speaker to ask the user questions, such as prompting them with questions like, "How much sleep did you get last night?" If the user responds, "I slept for six hours," the device collects the audio data.

[0132] Analysis of audio data

[0133] The device converts voice data into text using the Google Cloud Speech-to-Text API, and then analyzes the content, tone, and speed of the voice, allowing it to assess stress and fatigue levels.

[0134] Data integration and analysis

[0135] The server receives the facial video data, audio data, and question-and-answer data sent from the device and performs an integrated analysis. A comprehensive health assessment is performed using an AI model built with TensorFlow and PyTorch.

[0136] Generate health reports

[0137] The server generates a detailed health report based on the analysis results, including a warning such as "Your heart rate is high today. You may need to rest" if the analysis results show a higher than normal heart rate.

[0138] View Health Report

[0139] Finally, the device displays the generated health report on the screen and, if necessary, notifies the user by voice.

[0140] Prompt Sentence Examples

[0141] Examples of prompts the system might give to the user include:

[0142] "Good morning. How much sleep did you get last night?"

[0143] "How are you feeling today?"

[0144] "Are you stressed?"

[0145] These specific actions allow users to easily monitor their daily health status and quickly take necessary measures.

[0146] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0147] Step 1:

[0148] When a user stands in front of the mirror, the device detects movement and automatically activates the camera and microphone. The input is the user's movement, and the output is the activation of the camera and microphone. Specifically, the sensor detects the user's presence, and the camera begins capturing video.

[0149] Step 2:

[0150] The device recognizes the user's face. The camera captures an image of the user's face, and this video data is input into a facial recognition algorithm. The algorithm compares it with a previously saved database to identify the user, and the resulting username is output. A specific example of how this works is to identify the user as "Yamada-san" using facial recognition technology.

[0151] Step 3:

[0152] The device analyzes images of the user's face and measures their biometric information. The input is high-resolution video data acquired from a camera, which the processor analyzes to estimate heart rate, blood pressure, and body temperature. The output is this biometric information. Specifically, it uses computer vision technology to detect subtle color changes and muscle movements on the face to measure heart rate and blood pressure.

[0153] Step 4:

[0154] The device starts a voice dialogue with the user. It uses a microphone and speaker to ask a question and obtains the answer. The input is a prompt such as "How much sleep did you get last night?" and the user's voice response, and the output is voice data. Specifically, the system makes the user respond with "I slept for 6 hours."

[0155] Step 5:

[0156] The device analyzes the voice data. The input is the captured voice data, which the software converts into text and then analyzes the tone and speed. The output is an assessment of the user's stress level and fatigue. Specifically, it uses voice analysis technology to determine that a high-pitched voice and fast speed indicate high stress.

[0157] Step 6:

[0158] The server integrates and analyzes the data. Facial video data, audio data, and question-and-answer data sent from the device are input, and the server integrates these to perform a comprehensive health assessment. The output is a detailed health assessment result. Specifically, the system uses an AI model to analyze multidimensional data and evaluate the overall health condition.

[0159] Step 7:

[0160] The server generates a health report. The input is the integrated and analyzed data, and the software uses this to create a health report. The output is the user's health report. Specific operations include generating a report with a warning, such as "Your heart rate is high today. You may need to rest."

[0161] Step 8:

[0162] The device displays the generated health report. The input is the health report, which is visually displayed on the display. The output is visual feedback to the user. Specifically, the device displays "Your heart rate is high today. We recommend you rest" and notifies you by voice if necessary.

[0163] Through the above processing steps, users can easily understand their daily health condition in real time and take necessary measures promptly.

[0164] (Application example 1)

[0165] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0166] Conventional health management systems require users to voluntarily collect and input data, limiting their practicality and continuity. Furthermore, it is difficult to monitor detailed health conditions outside of regular health checkups. The present invention aims to provide a system that more efficiently and instantly assesses a user's health condition. In particular, the system can be installed in a brick-and-mortar store, solving the problem of enabling users to naturally manage their health as part of their daily activities.

[0167] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0168] In this invention, the server includes means for acquiring a facial image, means for analyzing the facial image to measure biometric information, means for engaging in voice dialogue with the user, means for analyzing the voice data to evaluate the health condition, means for integrating the acquired data to perform a comprehensive health evaluation, means for generating a health report, means for displaying the generated health report, and means for instantly evaluating the user's health using a mirror installed in the physical store. This allows the user's health condition to be automatically evaluated every time they visit the store, providing instant feedback and facilitating daily health management.

[0169] The "means for acquiring a facial image" refers to hardware such as a camera or image sensor for capturing an image of the user's face, and software for controlling the hardware.

[0170] The "means for analyzing facial images to measure biometric information" refers to software and algorithms for analyzing captured facial images and calculating biometric information such as heart rate, blood pressure, and body temperature.

[0171] The "means for conducting a voice dialogue with the user" refers to hardware and software for controlling the hardware and the speaker for conducting a voice dialogue with the user using a microphone and speaker.

[0172] The "means for analyzing voice data to assess health status" refers to software and algorithms for analyzing acquired voice data and assessing the user's stress level and fatigue level.

[0173] The "means for integrating acquired data to perform a comprehensive health assessment" refers to software and a server system for integrating facial video data and audio data to assess the user's overall health condition.

[0174] The "means for generating a health report" is software for creating a report that reports the user's health status based on the integrated health data.

[0175] The "means for displaying the generated health report" refers to a display and its control software for visually displaying the generated health report to the user.

[0176] The "means for instantly assessing a user's health using a mirror installed in a physical store" is a system for instantly assessing the health of users who visit the store and displaying the results through a mirror-like device installed in the store.

[0177] This invention relates to an "AI Health Check Mirror System" installed in a brick-and-mortar store, where users can stand in front of the mirror and have their facial images and voices analyzed, and receive instant feedback on their health condition for the day. The specific implementation of this system is described below.

[0178] System configuration

[0179] The system includes the following main elements:

[0180] 1. Camera: A means of capturing facial images, such as a high-resolution camera like the Logitech C920.

[0181] 2. Processor: A hardware device that analyzes facial images and measures biometric information. Various algorithms are embedded in the processor.

[0182] 3. Microphone and speaker: A means of audio interaction with the user. For example, a high-performance microphone such as the Blue Yeti.

[0183] 4. Voice analysis software: A means of analyzing voice data to assess health status. Custom algorithms for voice analysis.

[0184] 5. Server: A means of integrating acquired data to provide a comprehensive health assessment. Often, a cloud-based server is used.

[0185] 6. Display: Means for generating and displaying health reports. Display device that acts as a mirror.

[0186] 7. Physical store mirror: Installed in a physical store, it provides an instant health assessment of the user.

[0187] Data collection and analysis

[0188] 1. Acquiring facial images

[0189] When a user stands in front of the AI ​​health check mirror installed in a physical store, the camera automatically captures the user's facial image, which automatically activates the system.

[0190] 2. Measurement of biological information

[0191] The captured facial image is analyzed by a processor to measure biometric information such as heart rate, blood pressure, and body temperature. Subtle changes in facial color and facial expressions are also analyzed.

[0192] 3. Starting a voice conversation

[0193] When a user speaks to the system, a microphone picks up the voice and a voice interaction takes place through the speaker. The system asks questions, such as, "How much sleep did you get last night?"

[0194] 4. Analysis of audio data

[0195] The captured voice data is analyzed using voice analysis software, and stress and fatigue levels are assessed based on the tone and speed of the voice.

[0196] 5. Data integration and comprehensive health assessment

[0197] The server integrates facial video data, audio data, and voice dialogue data to provide a comprehensive health assessment, which includes various biometric data and voice analysis results.

[0198] 6. Generate and view health reports

[0199] The server generates a health report for the user based on the analysis results. If any abnormalities are detected, a report including a warning message is generated. The generated health report is displayed on the mirror and a warning is issued as needed. For example, feedback is given in the form of "Your heart rate is high today. You may need to rest."

[0200] Specific examples

[0201] When you stand in front of an AI health check mirror installed in a physical store, the system will automatically activate and capture an image of your face.

[0202] The images captured by the camera are analyzed by a processor to measure heart rate and blood pressure.

[0203] When the user looks into the mirror and says, "I slept six hours last night," voice analysis is performed to assess stress and fatigue levels.

[0204] The server integrates this data, performs a comprehensive health assessment, and then displays advice such as, "Your heart rate is high today. You may need to rest."

[0205] Example prompts for generative AI models

[0206] "Please explain in detail about the system that analyzes facial images and audio to assess the user's health condition. The system estimates biometric information such as heart rate, blood pressure, and body temperature by having the user stand in front of a mirror, and evaluates stress levels from the tone and rate of voice. Based on the assessment results, the system generates a health report and provides advice."

[0207] This allows users to easily monitor their daily health status and improve their ability to take necessary measures. It is expected that the present invention will significantly improve health management not only through conventional periodic health checkups but also through everyday activities.

[0208] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0209] Step 1:

[0210] When a user stands in front of the mirror, the device activates the built-in camera to capture an image of the user's face.

[0211] Input: Face image taken by camera

[0212] Output: Acquired facial image data

[0213] What it does: The system automatically activates the camera and captures high-resolution facial footage, which is then stored for further processing.

[0214] Step 2:

[0215] The facial image captured by the device is sent to a processor, which analyzes the facial image and measures biometric information.

[0216] Input: Acquired facial video data

[0217] Output: vital signs such as heart rate, blood pressure, and temperature

[0218] How it works: Video analysis algorithms analyze subtle changes in facial color and facial expressions to measure biometric information. A processor extracts and stores this data.

[0219] Step 3:

[0220] The device initiates a voice dialogue with the user. For example, the device asks, "How much sleep did you get last night?"

[0221] Input: User voice input

[0222] Output: Acquired audio data

[0223] What it does: A microphone picks up the user's voice and records the audio in an interactive format. A speaker is used to ask questions and record the user's responses.

[0224] Step 4:

[0225] The device analyzes the voice data it acquires and assesses the user's health condition.

[0226] Input: User's voice data

[0227] Output: User's voice analysis results (stress level, fatigue level, etc.)

[0228] What it does: Voice analysis software analyzes the tone and rate of the user's voice to assess stress and fatigue levels. These assessments are used in the next step.

[0229] Step 5:

[0230] The device sends facial video and audio data to a server, which then integrates the data to provide a comprehensive health assessment.

[0231] Input: Facial video data, audio data, audio analysis results

[0232] Output: Comprehensive health assessment data

[0233] How it works: The server combines facial video and audio data, and evaluates the user's overall health based on biometric and audio analysis results. An algorithm performs this evaluation and generates a result.

[0234] Step 6:

[0235] The server generates a health report. The server creates a report based on the evaluation results and adds warning messages if necessary.

[0236] Input: Comprehensive health assessment data

[0237] Output: Health report

[0238] What happens: The server generates a personalized health report based on the evaluation results, including a warning message if any abnormalities are detected. The report is displayed in the next step.

[0239] Step 7:

[0240] The health report generated by the device is displayed on the mirror and notified to the user.

[0241] Input: Health Report

[0242] Output: Health report displayed on mirror, audio notification

[0243] What it does: The device displays the generated health report on the mirror and notifies the user through the speaker, providing feedback such as, "Your heart rate is high today. You may need to rest."

[0244] This series of processes allows the user to quickly know their health condition and take measures as needed.

[0245] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0246] This invention relates to an AI-equipped mirror system that analyzes facial images and voices to evaluate the user's health and emotional state. This system allows the user to receive a multifaceted health and emotional state evaluation in real time by standing in front of the mirror.

[0247] System Overview

[0248] The system includes the following main elements:

[0249] 1. Camera that captures facial images

[0250] 2. Processor for video analysis and biometric measurement

[0251] 3. A microphone and speaker for voice interaction with the user

[0252] 4. Software that analyzes voice data to assess health status and emotions

[0253] 5. Server that integrates data and provides comprehensive health assessment

[0254] 6. Display to generate and display health reports

[0255] 7. Emotion engine that recognizes user emotions

[0256] Program processing flow

[0257] The program of this system is designed as follows:

[0258] 1. The user stands in front of the AI-enabled mirror

[0259] Every morning, the user stands in front of the mirror and looks at their face, which automatically activates the system.

[0260] 2. The device recognizes the user's face

[0261] The device uses a built-in camera to capture an image of your face and uses a facial recognition algorithm to identify you, which includes matching it against a historical database.

[0262] 3. The device analyzes your face and muscle movements

[0263] The device uses a high-resolution camera to capture detailed facial video data, allowing it to analyze and measure biometric information such as heart rate, blood pressure, and body temperature.

[0264] 4. The device starts a dialogue with the user

[0265] The terminal starts a voice dialogue, asking, for example, "Good morning. How long did you sleep today?" When the user responds, the voice data is acquired.

[0266] 5. The device analyzes the voice data

[0267] The device analyzes the user's voice and assesses stress and fatigue levels based on the tone and speed of the voice.

[0268] 6. The device analyzes facial images and recognizes emotions

[0269] The emotion engine analyzes the user's facial expressions and subtle muscle movements to recognize emotions such as joy, anger, and surprise.

[0270] 7. The server integrates and analyzes the data

[0271] The server integrates the facial video data, audio data, question and answer data, and emotional data sent from the device to perform a comprehensive health and emotional assessment.

[0272] 8. The server generates health and emotion reports

[0273] The server generates a health and emotional report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[0274] 9. The device displays the report

[0275] The device will display health and emotional reports on the mirror and issue warnings as needed, such as telling the user, "You're feeling stressed today. We recommend taking some time to relax."

[0276] Specific examples

[0277] 1. The user stands in front of the mirror

[0278] Every morning, when the user stands in front of the mirror, facial images are automatically captured.

[0279] 2. Facial Recognition

[0280] The device captures the user's face through the camera and identifies them as "Yamada-san" by comparing it with a previously registered database.

[0281] 3. Video Analysis

[0282] The camera captures detailed images of the user's face, and the processor measures heart rate, blood pressure, body temperature, etc. At the same time, subtle changes in facial color and facial expressions are analyzed.

[0283] 4. Starting a voice conversation

[0284] The device uses a microphone and speaker to ask, "How much sleep did you get last night?" The user responds, "I slept for 6 hours."

[0285] 5. Analysis of audio data

[0286] The device analyzes the user's voice data and detects stress levels from the tone and speed of the voice.

[0287] 6. Emotional Recognition

[0288] The emotion engine analyzes the user's facial expressions and recognizes their current emotions, such as whether they are happy, angry, or surprised.

[0289] 7. Data integration and analysis

[0290] The collected facial video data, audio data, lifestyle data, and emotional data are sent to a server, and a comprehensive health and emotional assessment is performed.

[0291] 8. Generate health and emotional reports

[0292] Based on the evaluation results, the server generates a health and mood report for the user, including a warning such as, "Your heart rate is high today. You may need to rest."

[0293] 9. View the report

[0294] The device displays health and emotion reports on the mirror and provides automated voice feedback to the user, such as "Good morning. You're feeling stressed today. We recommend you relax."

[0295] In this way, users can easily monitor their daily health and emotional state and take necessary measures. The present invention significantly improves daily health management and emotional management without relying solely on regular medical checkups.

[0296] The processing flow will be explained below.

[0297] Step 1:

[0298] The user stands in front of the mirror, which automatically activates the system and prepares the camera to capture the user's face.

[0299] Step 2:

[0300] The device uses its built-in camera to capture the user's face, then uses a facial recognition algorithm to identify the user and match them against a database to identify the user.

[0301] Step 3:

[0302] The device uses a high-resolution camera to capture detailed video data of the user's face, including technology to measure biometric information such as heart rate, blood pressure, and temperature from the face.

[0303] Step 4:

[0304] The terminal initiates a voice dialogue with the user, for example asking, "Good morning. How are you feeling today?" The user responds, "I'm feeling great."

[0305] Step 5:

[0306] The device analyzes the voice data it acquires, and uses voice analysis technology to assess stress and fatigue levels based on the tone and speed of the user's voice.

[0307] Step 6:

[0308] The device analyzes facial expressions and changes in color tone to measure heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, hemoglobin concentration, stress, and fatigue level.

[0309] Step 7:

[0310] The emotion engine analyzes the user's facial expressions and subtle muscle movements to recognize the user's emotional state. For example, if the user is smiling, it will recognize that they are feeling "happy."

[0311] Step 8:

[0312] The facial image data, audio data, question and answer data, and emotion data acquired by the terminal are transmitted to the server.

[0313] Step 9:

[0314] The server then integrates the received data to provide a comprehensive health and emotional assessment, including heart rate and blood pressure variability, voice analysis results, lifestyle information, and emotional state.

[0315] Step 10:

[0316] The server generates a health and emotion report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[0317] Step 11:

[0318] The device receives the health and emotion reports sent from the server and displays them on the mirror, for example, displaying a message such as "Your heart rate is high today. You may need to rest."

[0319] Step 12:

[0320] The device will notify the user with an automated voice, providing feedback such as, "Good morning. You seem to be feeling a little stressed today. I recommend you relax."

[0321] Example 2

[0322] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0323] Conventional health management systems often require specialized equipment at specific medical facilities to accurately assess a user's health status. This makes them inconvenient for everyday use and places a burden on users. Furthermore, they are unable to simultaneously assess the user's emotional state, making comprehensive health management difficult.

[0324] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring a facial image, means for analyzing the facial image to measure biometric information, means for conducting a voice dialogue with the user, means for analyzing the voice data to evaluate the health state and emotions, means for integrating the acquired data to perform a comprehensive health evaluation and emotion evaluation, means for generating a health and emotion report, and means for displaying the generated health and emotion report. This enables the user to easily and accurately monitor their own health and emotional state every day.

[0325] The "means for acquiring a facial image" is a device such as a camera for capturing an image of the user's face.

[0326] The "means for analyzing facial images and measuring biometric information" refers to software and hardware for analyzing biometric indicators such as heart rate, blood pressure, and body temperature based on the captured facial images.

[0327] The "means for performing voice interaction with a user" is a device for performing voice communication with a user using a microphone and a speaker.

[0328] The "means for analyzing voice data to assess health status and emotions" is software for analyzing the user's voice data and inferring health status and emotional state from the tone and speed of the voice.

[0329] "Means for integrating acquired data to provide a comprehensive health and emotional assessment" refers to a processor or algorithm that integrates collected biometric information, audio data, and other relevant data to provide a comprehensive assessment of the user's health and emotional state.

[0330] The "means for generating health and emotional reports" is software for automatically generating health and emotional reports based on the integrated data.

[0331] The "means for displaying the generated health and emotion report" is a display or monitor for visually presenting the generated report to the user.

[0332] The present invention relates to an AI-equipped mirror system that evaluates a user's health and emotional state using facial video and audio data. This system allows a user to receive multifaceted health and emotional evaluations in real time by standing in front of the mirror. Specific means and methods for implementing the present invention are described below.

[0333] Hardware used

[0334] Camera: Used to capture high-resolution footage, specifically capturing detailed footage of the face.

[0335] Processor: A processing device for analyzing video and audio data. For example, analyzing biometric information (heart rate, blood pressure, body temperature, etc.) from facial images.

[0336] Microphone and speaker: Used for voice interaction with the user, collecting information about the user's lifestyle habits and current emotional state through the interaction.

[0337] Display: A display device for presenting the generated health and emotion reports to the user.

[0338] Software used

[0339] Facial recognition algorithm: Analyzes facial images captured by the camera to recognize and identify the user.

[0340] Biometric analysis software: Biometric information is measured based on captured video data. This analysis uses changes in facial color and subtle muscle movements.

[0341] Voice analysis software: Analyzes voice data collected through a microphone to assess stress and fatigue levels.

[0342] Emotion engine: Analyzes facial expressions and subtle muscle movements to recognize the user's emotional state (happiness, anger, surprise, etc.).

[0343] Data integration algorithm: Integrates facial video data, audio data, and emotion data to perform comprehensive health and emotion assessments.

[0344] Report generation software: Generates health and emotional reports based on the integrated data.

[0345] System Operation

[0346] When a user stands in front of the mirror, the system captures facial images through a camera and identifies the user using a facial recognition algorithm. The processor analyzes the image data and measures biometric information (heart rate, blood pressure, body temperature, etc.). Next, a voice dialogue with the user is conducted through a microphone and speaker to collect information about the user's lifestyle habits and emotional state. The collected voice data is analyzed to evaluate the user's stress and fatigue levels. An emotion engine analyzes facial expressions and subtle muscle movements to recognize the user's emotional state. The server integrates this data to provide a comprehensive health and emotion assessment. Finally, report generation software creates health and emotion reports and displays them on the display. If any abnormalities are detected, a report is created with appropriate warning messages.

[0347] Specific examples

[0348] 1. The user stands in front of the mirror

[0349] Every morning, when the user stands in front of the mirror, facial images are automatically captured.

[0350] 2. Facial Recognition

[0351] The device captures the user's face through the camera and identifies them as "Yamada-san" by comparing it with a previously registered database.

[0352] 3. Video Analysis

[0353] The camera captures detailed images of the user's face, and the processor measures heart rate, blood pressure, body temperature, etc. At the same time, subtle changes in facial color and facial expressions are analyzed.

[0354] 4. Starting a voice conversation

[0355] The device uses a microphone and speaker to ask, "How much sleep did you get last night?" The user responds, "I slept for 6 hours."

[0356] 5. Analysis of audio data

[0357] The device analyzes the user's voice data and detects stress levels from the tone and speed of the voice.

[0358] 6. Emotional Recognition

[0359] The emotion engine analyzes the user's facial expressions and recognizes their current emotions, such as whether they are happy, angry, or surprised.

[0360] 7. Data integration and analysis

[0361] The collected facial video data, audio data, lifestyle data, and emotional data are sent to a server, and a comprehensive health and emotional assessment is performed.

[0362] 8. Generate health and emotional reports

[0363] Based on the evaluation results, the server generates a health and mood report for the user, including a warning such as, "Your heart rate is high today. You may need to rest."

[0364] 9. View the report

[0365] The device displays health and emotion reports on the mirror and provides automated voice feedback to the user, such as "Good morning. You're feeling stressed today. We recommend you relax."

[0366] Prompt Sentence Examples

[0367] Enter the following prompt into the generative AI model:

[0368] "This system uses an AI-equipped mirror to assess the user's health and emotional state. When the user stands in front of the mirror, a camera captures an image of their face and a voice dialogue begins. The acquired data is sent to a server for integrated analysis. Can you give us a concrete example?"

[0369] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0370] Step 1:

[0371] The system automatically activates when a user stands in front of the mirror. The input is the user's presence, which activates the camera and the device captures an image of the user's face. This action allows the system to capture an image of the user's face when the user says "Good morning."

[0372] Step 2:

[0373] The device uses a built-in camera to recognize the user's face. The input is a captured image of the face, and the output is the user's identification. Specifically, the facial recognition algorithm analyzes the image data and compares it with a database to identify the user as "Yamada-san."

[0374] Step 3:

[0375] The device uses a high-resolution camera to capture detailed facial images and analyze biometric information. The input is the recognized user's video data, and the output is biometric information such as heart rate, blood pressure, and body temperature. The processor analyzes subtle changes in facial color and muscle movements to make a diagnosis, such as whether the user has a pale complexion.

[0376] Step 4:

[0377] The device initiates a voice dialogue with the user using a microphone and speaker. The input is the user's voice, and the output is voice data. For example, the system asks, "Good morning. How much sleep did you get last night?" and the user replies, "I slept for six hours."

[0378] Step 5:

[0379] The device analyzes the acquired voice data and evaluates stress and fatigue levels based on the tone and speed of the voice. The input is the acquired voice data, and the output is the evaluation result of the user's stress and fatigue. For example, if the voice is quiet and low-pitched, the device will diagnose, "You seem a little tired."

[0380] Step 6:

[0381] The device uses an emotion engine to analyze facial expressions and subtle muscle movements to recognize emotions. The input is video data of the user's face, and the output is the recognition result of the user's emotional state (happiness, anger, surprise, etc.). For example, if the user's brow is furrowed, the device will evaluate the user as "slightly anxious."

[0382] Step 7:

[0383] The server integrates and analyzes the facial video data, audio data, question-and-answer data, and emotional data sent from the device. The input is multiple data sets, and the output is a comprehensive health and emotional assessment result. Based on this, the server determines that "the user is experiencing accumulated fatigue."

[0384] Step 8:

[0385] The server generates a health and emotion report based on the assessment results. The input is the integrated and analyzed assessment results, and the output is a health and emotion report. For example, it may include a warning message such as "Your heart rate is high today. Please take a rest."

[0386] Step 9:

[0387] The terminal displays the generated report on the mirror and issues a warning if necessary. The input is the generated report, and the output is the displayed report and audio feedback. Specifically, it tells the user, "Good morning. You are feeling stressed today. We recommend that you relax."

[0388] (Application example 2)

[0389] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0390] In modern society, especially in brick-and-mortar stores, there is a need to understand customers' health and emotional states in real time and provide services accordingly. However, conventional methods require customers to self-report their health and emotional states or have specialized staff manually monitor them, which has issues with inefficiency and inaccuracy. The present invention aims to provide a system that solves these issues and significantly improves customer experience.

[0391] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a facial image, means for analyzing the facial image to measure biometric information, means for engaging in voice dialogue with the user, means for analyzing voice data to evaluate a health condition, means for integrating the acquired data to perform a comprehensive health evaluation, means for generating a health report, means for displaying the generated health report, and means for displaying the health evaluation and emotional evaluation on a display device in the store or on a staff member's device to provide personalized service. This allows customers to understand their own health and emotional state in real time, and enables store staff to provide appropriate and personalized service to customers.

[0392] "Means for acquiring facial images" refers to a device or equipment for capturing a user's facial image in real time.

[0393] "Means for analyzing facial images to measure biometric information" refers to algorithms and software for analyzing captured facial images and measuring biometric information such as heart rate, blood pressure, and body temperature.

[0394] "Means for audio interaction with the user" refers to a microphone and speaker, and associated interaction software, for asking questions of the user and receiving responses.

[0395] "Means for analyzing voice data to assess health status" refers to software or algorithms for analyzing a user's voice data and assessing health status from the tone and rate of the voice data.

[0396] "Means for integrating acquired data to conduct a comprehensive health assessment" refers to a server or analysis system for integrating multiple data, such as facial video data and audio data, to conduct a comprehensive health assessment.

[0397] "Means for generating a health report" refers to software or algorithms that automatically generate a report summarizing health status based on the results of a comprehensive health assessment.

[0398] "Means for displaying the generated health report" refers to a display or display device, and associated software, for displaying the generated health report so that the user can review it.

[0399] "Means for displaying health and emotional assessments on in-store display devices or staff devices to provide personalized services" refers to systems and software that grasp customers' health and emotional status in real time, display this information on in-store displays or staff smart devices, and provide personalized services.

[0400] The present invention relates to a system that assesses a customer's health and emotional state in real time and provides personalized services in physical stores. When a customer enters a store, the system captures and analyzes facial images and evaluates their heart rate, blood pressure, body temperature, emotional state, etc. to provide an optimal customer experience.

[0401] Key elements of the system

[0402] The system includes the following main elements:

[0403] 1. Camera that captures facial images:

[0404] This camera uses a high-resolution camera (e.g., a high-resolution USB camera) that is placed on a mirror or other suitable location within the store.

[0405] 2. Video analysis and biometrics processor:

[0406] The video is analyzed using a facial recognition algorithm using the OpenCV library, and biometric information is measured using algorithms such as photoplethysmography.

[0407] 3. Microphone and speaker for voice interaction with the user:

[0408] For voice interaction, a high-precision microphone (e.g., a high-performance USB microphone) and a speaker are used, and voice analysis is performed via the Google Speech-to-Text API.

[0409] 4. Software that analyzes voice data to assess health and emotions:

[0410] To analyze voice data, the Google Speech-to-Text API and AI models such as TensorFlow are used.

[0411] 5. Server that integrates data to provide a comprehensive health assessment:

[0412] The server uses a cloud-based server (e.g., a public cloud service provider) to consolidate the collected data and perform real-time evaluation.

[0413] 6. Display to generate and show health reports:

[0414] Health reports are generated using a Python Flask-based web application and are displayed on displays in-store and on staff's smart devices.

[0415] 7. System for displaying health and emotional assessments and providing personalized services:

[0416] Use a system (e.g., personalized assistance application) that allows store staff to provide personalized service based on customer ratings.

[0417] Specific Examples

[0418] When a customer enters a store, a high-resolution camera inside the store captures an image of the customer's face, and at the same time, a voice conversation begins using a microphone and speaker.

[0419] The resulting facial and audio data is sent to video and audio analysis software, where the facial data is used to measure vital signs such as heart rate, blood pressure, and body temperature, while the audio data is analyzed to assess emotional state.

[0420] The server integrates this data to perform a comprehensive evaluation and generate health and emotion reports, which are displayed in real time on displays in the store and on staff's smart devices.

[0421] Staff can use this information to provide customers with personalized service, for example, suggesting products that will help them relax if they are feeling tired.

[0422] Prompt Sentence Examples

[0423] text

[0424] Develop a system that will capture the customer's facial image and voice in real time using a mirror in the store when they enter the store, evaluate their health and emotional state based on their heart rate, blood pressure, body temperature, tone and speed of voice, and integrate the collected data on a server to provide personalized services. The specifications will use OpenCV for facial recognition, Google Speech-to-Text API for voice analysis, Affectiva SDK for emotion recognition, and Python Flask for data integration.

[0425] In this way, the present invention allows for real-time assessment of a customer's health and emotional state to provide optimal service, thereby improving the customer experience and the service quality of the store.

[0426] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0427] Step 1:

[0428] When a user enters a store, the device's camera captures the user's facial image. This camera has high resolution and captures images in real time. The input is the user's facial image, and the output is the captured facial image data.

[0429] Step 2:

[0430] The facial image data acquired by the device is sent to a facial recognition algorithm. This algorithm uses the OpenCV library to identify the user's face. The input is the facial image data, and the output is facial recognition data in which the user's face is identified.

[0431] Step 3:

[0432] Based on the facial recognition data, the device measures the user's biometric information. Specifically, it uses video analysis software to measure heart rate, blood pressure, body temperature, etc. Photoplethysmography technology is used for this purpose. The input is facial recognition data, and the output is biometric data.

[0433] Step 4:

[0434] The terminal starts a voice dialogue with the user, asks a question through a microphone, and obtains the user's response. For example, the terminal asks, "What kind of product are you looking for today?" The input is the user's voice response, and the output is the obtained voice data.

[0435] Step 5:

[0436] The device sends the captured voice data to the Google Speech-to-Text API, which converts the voice into text data. Furthermore, speech analysis software is used to evaluate the user's emotional state based on the tone and speed of the voice. The input is voice data, and the output is textual voice data and emotional state data.

[0437] Step 6:

[0438] The device sends facial video data and audio data to a server, which then integrates the data to perform comprehensive health and emotional assessments. The inputs are biometric data, audio text data, and emotional state data, and the output is an integrated health and emotional assessment report.

[0439] Step 7:

[0440] The server returns the generated health and emotion evaluation reports to the terminal, which then displays these reports on displays in the store or on staff members' smart devices. The input is the health and emotion evaluation reports, and the output is the displayed reports.

[0441] Step 8:

[0442] Store staff provide personalized services based on the user's state. For example, if the user is stressed, they will suggest products that will help them relax. The staff responds to the customer while referring to the display on the terminal. The input is the displayed report, and the output is the service provided.

[0443] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0444] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0445] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0446] [Second embodiment]

[0447] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0448] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0449] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0450] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0451] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0452] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0453] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0454] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0455] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0456] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0457] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0458] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0459] This invention relates to an AI-equipped mirror system that analyzes facial images and audio to assess the user's health condition. This system allows the user to receive a multifaceted health assessment in real time by standing in front of the mirror.

[0460] System Overview

[0461] The system includes the following main elements:

[0462] 1. Camera that captures facial images

[0463] 2. Processor for video analysis and biometric measurement

[0464] 3. A microphone and speaker for voice interaction with the user

[0465] 4. Software that analyzes voice data to assess health status

[0466] 5. Server that integrates data and provides comprehensive health assessment

[0467] 6. Display to generate and display health reports

[0468] Program processing flow

[0469] The program of this system is designed as follows:

[0470] 1. The user stands in front of the AI-enabled mirror

[0471] Every morning, the user stands in front of the mirror and looks at their face, which automatically activates the system.

[0472] 2. The device recognizes the user's face

[0473] The device uses a built-in camera to capture an image of your face and uses a facial recognition algorithm to identify you, which includes matching it against a historical database.

[0474] 3. The device analyzes your face and muscle movements

[0475] The device uses a high-resolution camera to capture detailed facial video data, allowing it to analyze and measure biometric information such as heart rate, blood pressure, and body temperature.

[0476] 4. The device starts a dialogue with the user

[0477] The device initiates a voice dialogue, asking a question such as, "How much sleep did you get last night?" When the user responds, the voice data is acquired.

[0478] 5. The device analyzes the voice data

[0479] The device analyzes the user's voice and assesses stress and fatigue levels based on the tone and speed of the voice.

[0480] 6. The server integrates and analyzes the data

[0481] The server integrates the facial video data, audio data, and question-and-answer data sent from the device to perform a comprehensive health assessment.

[0482] 7. The server generates a health report

[0483] The server generates a health report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[0484] 8. The device will display a health report

[0485] The device will display a health report on the mirror and issue warnings as needed, such as telling the user, "You're feeling stressed today. We recommend taking some time to relax."

[0486] Specific examples

[0487] 1. The user stands in front of the mirror

[0488] Every morning, when the user stands in front of the mirror, facial images are automatically captured.

[0489] 2. Facial Recognition

[0490] The device captures the user's face through the camera and identifies them as "Yamada-san" by comparing it with a previously registered database.

[0491] 3. Video Analysis

[0492] The camera captures detailed images of the user's face, and the processor measures heart rate, blood pressure, body temperature, etc. At the same time, subtle changes in facial color and facial expressions are analyzed.

[0493] 4. Starting a voice conversation

[0494] The device uses a microphone and speaker to ask, "How much sleep did you get last night?" The user responds, "I slept for 6 hours."

[0495] 5. Analysis of audio data

[0496] The device analyzes the user's voice data and detects stress levels from the tone and speed of the voice.

[0497] 6. Data integration and analysis

[0498] The collected facial video data, audio data, and lifestyle data are sent to a server for a comprehensive health assessment.

[0499] 7. Generate Health Reports

[0500] Based on the evaluation results, the server generates a health report for the user, including a warning such as, "Your heart rate is high today. You may need to rest."

[0501] 8. View Health Report

[0502] The device displays a health report on the mirror and notifies the user with an automated voice.

[0503] In this way, users can easily monitor their daily health status and take necessary measures. The present invention significantly improves daily health management without relying solely on regular medical checkups.

[0504] The processing flow will be explained below.

[0505] Step 1:

[0506] The user stands in front of the mirror, which automatically activates the system and prepares the camera to capture the user's face.

[0507] Step 2:

[0508] The device uses its built-in camera to capture the user's face, then uses a facial recognition algorithm to identify the user and match them with a historical database.

[0509] Step 3:

[0510] The device uses a high-resolution camera to capture detailed facial video data, including technology to capture facial biometric information such as heart rate, blood pressure, and temperature.

[0511] Step 4:

[0512] The device starts a voice dialogue with the user. For example, it asks, "Good morning. How long did you sleep today?" The user replies, "I slept for 6 hours."

[0513] Step 5:

[0514] The device analyzes the voice data it acquires, and uses voice analysis technology to assess stress and fatigue levels based on the tone and speed of the user's voice.

[0515] Step 6:

[0516] The device analyzes facial expressions and changes in color tone to measure heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, hemoglobin concentration, stress, and fatigue level.

[0517] Step 7:

[0518] The facial image data, audio data, and question and answer data acquired by the terminal are transmitted to the server.

[0519] Step 8:

[0520] The server then combines the received data to provide a comprehensive health assessment, including heart rate and blood pressure variability, voice analysis results, and lifestyle information.

[0521] Step 9:

[0522] The server generates a health report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[0523] Step 10:

[0524] The device receives the health report sent from the server and displays it on the mirror, for example, displaying a message such as "Your heart rate is high today. You may need to rest."

[0525] Step 11:

[0526] The device will notify the user with an automated voice, providing feedback such as, "Good morning. You're feeling stressed today. We recommend you relax."

[0527] Example 1

[0528] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0529] Conventional health management systems require users to manually input information, which can lead to problems with accuracy and effort. In addition, the long intervals between regular health checkups make it difficult to quickly detect changes in health status, making it impossible to grasp health status in real time.

[0530] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0531] In this invention, the server includes means for acquiring a facial image, means for analyzing the acquired facial image to measure biometric information, means for engaging in voice dialogue with the user, means for analyzing the voice data to evaluate the health condition, means for integrating the acquired video data and voice data to perform a comprehensive health evaluation, means for generating a health report, and means for displaying the generated health report, thereby enabling the user to receive a comprehensive health evaluation in real time simply by standing in front of the mirror.

[0532] "Means for capturing facial images" refers to a camera or other imaging device that captures an image or video of a user's face.

[0533] "Means for analyzing captured facial images and measuring biometric information" refers to processors and algorithms that analyze and measure biometric information such as heart rate, blood pressure, and body temperature from captured facial image data.

[0534] "Means for conducting voice dialogue with the user" refers to a microphone and speaker for communicating with the user by voice, and software for controlling the dialogue.

[0535] "Means for analyzing voice data to assess health status" refers to algorithms or analytical software for analyzing voice data and assessing a user's stress level, fatigue level, etc.

[0536] "Means for integrating acquired video and audio data to perform a comprehensive health assessment" refers to a server and analysis program for integrating facial video and audio data to assess overall health status.

[0537] "Means for generating a health report" refers to software or algorithms that automatically generate a report showing the user's health status based on the analysis results.

[0538] "Means for displaying the generated health report" refers to a display or screen for visually presenting the generated health report to the user.

[0539] This invention relates to an AI-equipped mirror system that analyzes the user's facial image and voice to evaluate their health condition in real time. To implement this system, the following hardware and software are used:

[0540] Hardware and software used

[0541] 1. Camera: Used to capture video of the user's face. For example, a 1080p HD camera or a 4K camera is suitable.

[0542] 2. Processor: Used for video analysis and biometric measurement. Examples include an Intel Core i7 processor or an NVIDIA GPU.

[0543] 3. Microphone and speaker: Used for voice interaction with the user, a high-sensitivity microphone and stereo speakers are desirable to provide clear sound quality.

[0544] 4. Analysis software: Software used to analyze video and audio data often uses OpenCV, TensorFlow, or Google Cloud Speech-to-Text API.

[0545] 5. Server: A backend server that integrates data and performs comprehensive health assessments. This is where AI models using TensorFlow and PyTorch run.

[0546] 6. Display: Used to display the generated health report. A display that also functions as a mirror is desirable.

[0547] How the system works

[0548] This system automatically starts when the user stands in front of the mirror and performs the following processes.

[0549] Image acquisition by camera

[0550] When a user stands in front of the mirror, the camera automatically captures an image of the user's face.

[0551] Facial image analysis

[0552] The processor analyzes the captured facial video data. For example, it detects subtle color changes and facial expressions, and uses this data to estimate heart rate, blood pressure, and body temperature. This analysis uses the OpenCV library.

[0553] Voice interaction with the user

[0554] The device uses a microphone and speaker to ask the user questions, such as prompting them with questions like, "How much sleep did you get last night?" If the user responds, "I slept for six hours," the device collects the audio data.

[0555] Analysis of audio data

[0556] The device converts voice data into text using the Google Cloud Speech-to-Text API, and then analyzes the content, tone, and speed of the voice, allowing it to assess stress and fatigue levels.

[0557] Data integration and analysis

[0558] The server receives the facial video data, audio data, and question-and-answer data sent from the device and performs an integrated analysis. A comprehensive health assessment is performed using an AI model built with TensorFlow and PyTorch.

[0559] Generate health reports

[0560] The server generates a detailed health report based on the analysis results, including a warning such as "Your heart rate is high today. You may need to rest" if the analysis results show a higher than normal heart rate.

[0561] View Health Report

[0562] Finally, the device displays the generated health report on the screen and, if necessary, notifies the user by voice.

[0563] Prompt Sentence Examples

[0564] Examples of prompts the system might give to the user include:

[0565] "Good morning. How much sleep did you get last night?"

[0566] "How are you feeling today?"

[0567] "Are you stressed?"

[0568] These specific actions allow users to easily monitor their daily health status and quickly take necessary measures.

[0569] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0570] Step 1:

[0571] When a user stands in front of the mirror, the device detects movement and automatically activates the camera and microphone. The input is the user's movement, and the output is the activation of the camera and microphone. Specifically, the sensor detects the user's presence, and the camera begins capturing video.

[0572] Step 2:

[0573] The device recognizes the user's face. The camera captures an image of the user's face, and this video data is input into a facial recognition algorithm. The algorithm compares it with a previously saved database to identify the user, and the resulting username is output. A specific example of how this works is to identify the user as "Yamada-san" using facial recognition technology.

[0574] Step 3:

[0575] The device analyzes images of the user's face and measures their biometric information. The input is high-resolution video data acquired from a camera, which the processor analyzes to estimate heart rate, blood pressure, and body temperature. The output is this biometric information. Specifically, it uses computer vision technology to detect subtle color changes and muscle movements on the face to measure heart rate and blood pressure.

[0576] Step 4:

[0577] The device starts a voice dialogue with the user. It uses a microphone and speaker to ask a question and obtains the answer. The input is a prompt such as "How much sleep did you get last night?" and the user's voice response, and the output is voice data. Specifically, the system makes the user respond with "I slept for 6 hours."

[0578] Step 5:

[0579] The device analyzes the voice data. The input is the captured voice data, which the software converts into text and then analyzes the tone and speed. The output is an assessment of the user's stress level and fatigue. Specifically, it uses voice analysis technology to determine that a high-pitched voice and fast speed indicate high stress.

[0580] Step 6:

[0581] The server integrates and analyzes the data. Facial video data, audio data, and question-and-answer data sent from the device are input, and the server integrates these to perform a comprehensive health assessment. The output is a detailed health assessment result. Specifically, the system uses an AI model to analyze multidimensional data and evaluate the overall health condition.

[0582] Step 7:

[0583] The server generates a health report. The input is the integrated and analyzed data, and the software uses this to create a health report. The output is the user's health report. Specific operations include generating a report with a warning, such as "Your heart rate is high today. You may need to rest."

[0584] Step 8:

[0585] The device displays the generated health report. The input is the health report, which is visually displayed on the display. The output is visual feedback to the user. Specifically, the device displays "Your heart rate is high today. We recommend you rest" and notifies you by voice if necessary.

[0586] Through the above processing steps, users can easily understand their daily health condition in real time and take necessary measures promptly.

[0587] (Application example 1)

[0588] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0589] Conventional health management systems require users to voluntarily collect and input data, limiting their practicality and continuity. Furthermore, it is difficult to monitor detailed health conditions outside of regular health checkups. The present invention aims to provide a system that more efficiently and instantly assesses a user's health condition. In particular, the system can be installed in a brick-and-mortar store, solving the problem of enabling users to naturally manage their health as part of their daily activities.

[0590] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0591] In this invention, the server includes means for acquiring a facial image, means for analyzing the facial image to measure biometric information, means for engaging in voice dialogue with the user, means for analyzing the voice data to evaluate the health condition, means for integrating the acquired data to perform a comprehensive health evaluation, means for generating a health report, means for displaying the generated health report, and means for instantly evaluating the user's health using a mirror installed in the physical store. This allows the user's health condition to be automatically evaluated every time they visit the store, providing instant feedback and facilitating daily health management.

[0592] The "means for acquiring a facial image" refers to hardware such as a camera or image sensor for capturing an image of the user's face, and software for controlling the hardware.

[0593] The "means for analyzing facial images to measure biometric information" refers to software and algorithms for analyzing captured facial images and calculating biometric information such as heart rate, blood pressure, and body temperature.

[0594] The "means for conducting a voice dialogue with the user" refers to hardware and software for controlling the hardware and the speaker for conducting a voice dialogue with the user using a microphone and speaker.

[0595] The "means for analyzing voice data to assess health status" refers to software and algorithms for analyzing acquired voice data and assessing the user's stress level and fatigue level.

[0596] The "means for integrating acquired data to perform a comprehensive health assessment" refers to software and a server system for integrating facial video data and audio data to assess the user's overall health condition.

[0597] The "means for generating a health report" is software for creating a report that reports the user's health status based on the integrated health data.

[0598] The "means for displaying the generated health report" refers to a display and its control software for visually displaying the generated health report to the user.

[0599] The "means for instantly assessing a user's health using a mirror installed in a physical store" is a system for instantly assessing the health of users who visit the store and displaying the results through a mirror-like device installed in the store.

[0600] This invention relates to an "AI Health Check Mirror System" installed in a brick-and-mortar store, where users can stand in front of the mirror and have their facial images and voices analyzed, and receive instant feedback on their health condition for the day. The specific implementation of this system is described below.

[0601] System configuration

[0602] The system includes the following main elements:

[0603] 1. Camera: A means of capturing facial images, such as a high-resolution camera like the Logitech C920.

[0604] 2. Processor: A hardware device that analyzes facial images and measures biometric information. Various algorithms are embedded in the processor.

[0605] 3. Microphone and speaker: A means of audio interaction with the user. For example, a high-performance microphone such as the Blue Yeti.

[0606] 4. Voice analysis software: A means of analyzing voice data to assess health status. Custom algorithms for voice analysis.

[0607] 5. Server: A means of integrating acquired data to provide a comprehensive health assessment. Often, a cloud-based server is used.

[0608] 6. Display: Means for generating and displaying health reports. Display device that acts as a mirror.

[0609] 7. Physical store mirror: Installed in a physical store, it provides an instant health assessment of the user.

[0610] Data collection and analysis

[0611] 1. Acquiring facial images

[0612] When a user stands in front of the AI ​​health check mirror installed in a physical store, the camera automatically captures the user's facial image, which automatically activates the system.

[0613] 2. Measurement of biological information

[0614] The captured facial image is analyzed by a processor to measure biometric information such as heart rate, blood pressure, and body temperature. Subtle changes in facial color and facial expressions are also analyzed.

[0615] 3. Starting a voice conversation

[0616] When a user speaks to the system, a microphone picks up the voice and a voice interaction takes place through the speaker. The system asks questions, such as, "How much sleep did you get last night?"

[0617] 4. Analysis of audio data

[0618] The captured voice data is analyzed using voice analysis software, and stress and fatigue levels are assessed based on the tone and speed of the voice.

[0619] 5. Data integration and comprehensive health assessment

[0620] The server integrates facial video data, audio data, and voice dialogue data to provide a comprehensive health assessment, which includes various biometric data and voice analysis results.

[0621] 6. Generate and view health reports

[0622] The server generates a health report for the user based on the analysis results. If any abnormalities are detected, a report including a warning message is generated. The generated health report is displayed on the mirror and a warning is issued as needed. For example, feedback is given in the form of "Your heart rate is high today. You may need to rest."

[0623] Specific examples

[0624] When you stand in front of an AI health check mirror installed in a physical store, the system will automatically activate and capture an image of your face.

[0625] The images captured by the camera are analyzed by a processor to measure heart rate and blood pressure.

[0626] When the user looks into the mirror and says, "I slept six hours last night," voice analysis is performed to assess stress and fatigue levels.

[0627] The server integrates this data, performs a comprehensive health assessment, and then displays advice such as, "Your heart rate is high today. You may need to rest."

[0628] Example prompts for generative AI models

[0629] "Please explain in detail about the system that analyzes facial images and audio to assess the user's health condition. The system estimates biometric information such as heart rate, blood pressure, and body temperature by having the user stand in front of a mirror, and evaluates stress levels from the tone and rate of voice. Based on the assessment results, the system generates a health report and provides advice."

[0630] This allows users to easily monitor their daily health status and improve their ability to take necessary measures. It is expected that the present invention will significantly improve health management not only through conventional periodic health checkups but also through everyday activities.

[0631] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0632] Step 1:

[0633] When a user stands in front of the mirror, the device activates the built-in camera to capture an image of the user's face.

[0634] Input: Face image taken by camera

[0635] Output: Acquired facial image data

[0636] What it does: The system automatically activates the camera and captures high-resolution facial footage, which is then stored for further processing.

[0637] Step 2:

[0638] The facial image captured by the device is sent to a processor, which analyzes the facial image and measures biometric information.

[0639] Input: Acquired facial video data

[0640] Output: vital signs such as heart rate, blood pressure, and temperature

[0641] How it works: Video analysis algorithms analyze subtle changes in facial color and facial expressions to measure biometric information. A processor extracts and stores this data.

[0642] Step 3:

[0643] The device initiates a voice dialogue with the user. For example, the device asks, "How much sleep did you get last night?"

[0644] Input: User voice input

[0645] Output: Acquired audio data

[0646] What it does: A microphone picks up the user's voice and records the audio in an interactive format. A speaker is used to ask questions and record the user's responses.

[0647] Step 4:

[0648] The device analyzes the voice data it acquires and assesses the user's health condition.

[0649] Input: User's voice data

[0650] Output: User's voice analysis results (stress level, fatigue level, etc.)

[0651] What it does: Voice analysis software analyzes the tone and rate of the user's voice to assess stress and fatigue levels. These assessments are used in the next step.

[0652] Step 5:

[0653] The device sends facial video and audio data to a server, which then integrates the data to provide a comprehensive health assessment.

[0654] Input: Facial video data, audio data, audio analysis results

[0655] Output: Comprehensive health assessment data

[0656] How it works: The server combines facial video and audio data, and evaluates the user's overall health based on biometric and audio analysis results. An algorithm performs this evaluation and generates a result.

[0657] Step 6:

[0658] The server generates a health report. The server creates a report based on the evaluation results and adds warning messages if necessary.

[0659] Input: Comprehensive health assessment data

[0660] Output: Health report

[0661] What happens: The server generates a personalized health report based on the evaluation results, including a warning message if any abnormalities are detected. The report is displayed in the next step.

[0662] Step 7:

[0663] The health report generated by the device is displayed on the mirror and notified to the user.

[0664] Input: Health Report

[0665] Output: Health report displayed on mirror, audio notification

[0666] What it does: The device displays the generated health report on the mirror and notifies the user through the speaker, providing feedback such as, "Your heart rate is high today. You may need to rest."

[0667] This series of processes allows the user to quickly know their health condition and take measures as needed.

[0668] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0669] This invention relates to an AI-equipped mirror system that analyzes facial images and voices to evaluate the user's health and emotional state. This system allows the user to receive a multifaceted health and emotional state evaluation in real time by standing in front of the mirror.

[0670] System Overview

[0671] The system includes the following main elements:

[0672] 1. Camera that captures facial images

[0673] 2. Processor for video analysis and biometric measurement

[0674] 3. A microphone and speaker for voice interaction with the user

[0675] 4. Software that analyzes voice data to assess health status and emotions

[0676] 5. Server that integrates data and provides comprehensive health assessment

[0677] 6. Display to generate and display health reports

[0678] 7. Emotion engine that recognizes user emotions

[0679] Program processing flow

[0680] The program of this system is designed as follows:

[0681] 1. The user stands in front of the AI-enabled mirror

[0682] Every morning, the user stands in front of the mirror and looks at their face, which automatically activates the system.

[0683] 2. The device recognizes the user's face

[0684] The device uses a built-in camera to capture an image of your face and uses a facial recognition algorithm to identify you, which includes matching it against a historical database.

[0685] 3. The device analyzes your face and muscle movements

[0686] The device uses a high-resolution camera to capture detailed facial video data, allowing it to analyze and measure biometric information such as heart rate, blood pressure, and body temperature.

[0687] 4. The device starts a dialogue with the user

[0688] The terminal starts a voice dialogue, asking, for example, "Good morning. How long did you sleep today?" When the user responds, the voice data is acquired.

[0689] 5. The device analyzes the voice data

[0690] The device analyzes the user's voice and assesses stress and fatigue levels based on the tone and speed of the voice.

[0691] 6. The device analyzes facial images and recognizes emotions

[0692] The emotion engine analyzes the user's facial expressions and subtle muscle movements to recognize emotions such as joy, anger, and surprise.

[0693] 7. The server integrates and analyzes the data

[0694] The server integrates the facial video data, audio data, question and answer data, and emotional data sent from the device to perform a comprehensive health and emotional assessment.

[0695] 8. The server generates health and emotion reports

[0696] The server generates a health and emotional report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[0697] 9. The device displays the report

[0698] The device will display health and emotional reports on the mirror and issue warnings as needed, such as telling the user, "You're feeling stressed today. We recommend taking some time to relax."

[0699] Specific examples

[0700] 1. The user stands in front of the mirror

[0701] Every morning, when the user stands in front of the mirror, facial images are automatically captured.

[0702] 2. Facial Recognition

[0703] The device captures the user's face through the camera and identifies them as "Yamada-san" by comparing it with a previously registered database.

[0704] 3. Video Analysis

[0705] The camera captures detailed images of the user's face, and the processor measures heart rate, blood pressure, body temperature, etc. At the same time, subtle changes in facial color and facial expressions are analyzed.

[0706] 4. Starting a voice conversation

[0707] The device uses a microphone and speaker to ask, "How much sleep did you get last night?" The user responds, "I slept for 6 hours."

[0708] 5. Analysis of audio data

[0709] The device analyzes the user's voice data and detects stress levels from the tone and speed of the voice.

[0710] 6. Emotional Recognition

[0711] The emotion engine analyzes the user's facial expressions and recognizes their current emotions, such as whether they are happy, angry, or surprised.

[0712] 7. Data integration and analysis

[0713] The collected facial video data, audio data, lifestyle data, and emotional data are sent to a server, and a comprehensive health and emotional assessment is performed.

[0714] 8. Generate health and emotional reports

[0715] Based on the evaluation results, the server generates a health and mood report for the user, including a warning such as, "Your heart rate is high today. You may need to rest."

[0716] 9. View the report

[0717] The device displays health and emotion reports on the mirror and provides automated voice feedback to the user, such as "Good morning. You're feeling stressed today. We recommend you relax."

[0718] In this way, users can easily monitor their daily health and emotional state and take necessary measures. The present invention significantly improves daily health management and emotional management without relying solely on regular medical checkups.

[0719] The processing flow will be explained below.

[0720] Step 1:

[0721] The user stands in front of the mirror, which automatically activates the system and prepares the camera to capture the user's face.

[0722] Step 2:

[0723] The device uses its built-in camera to capture the user's face, then uses a facial recognition algorithm to identify the user and match them against a database to identify the user.

[0724] Step 3:

[0725] The device uses a high-resolution camera to capture detailed video data of the user's face, including technology to measure biometric information such as heart rate, blood pressure, and temperature from the face.

[0726] Step 4:

[0727] The terminal initiates a voice dialogue with the user, for example asking, "Good morning. How are you feeling today?" The user responds, "I'm feeling great."

[0728] Step 5:

[0729] The device analyzes the voice data it acquires, and uses voice analysis technology to assess stress and fatigue levels based on the tone and speed of the user's voice.

[0730] Step 6:

[0731] The device analyzes facial expressions and changes in color tone to measure heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, hemoglobin concentration, stress, and fatigue level.

[0732] Step 7:

[0733] The emotion engine analyzes the user's facial expressions and subtle muscle movements to recognize the user's emotional state. For example, if the user is smiling, it will recognize that they are feeling "happy."

[0734] Step 8:

[0735] The facial image data, audio data, question and answer data, and emotion data acquired by the terminal are transmitted to the server.

[0736] Step 9:

[0737] The server then integrates the received data to provide a comprehensive health and emotional assessment, including heart rate and blood pressure variability, voice analysis results, lifestyle information, and emotional state.

[0738] Step 10:

[0739] The server generates a health and emotion report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[0740] Step 11:

[0741] The device receives the health and emotion reports sent from the server and displays them on the mirror, for example, displaying a message such as "Your heart rate is high today. You may need to rest."

[0742] Step 12:

[0743] The device will notify the user with an automated voice, providing feedback such as, "Good morning. You seem to be feeling a little stressed today. I recommend you relax."

[0744] Example 2

[0745] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0746] Conventional health management systems often require specialized equipment at specific medical facilities to accurately assess a user's health status. This makes them inconvenient for everyday use and places a burden on users. Furthermore, they are unable to simultaneously assess the user's emotional state, making comprehensive health management difficult.

[0747] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring a facial image, means for analyzing the facial image to measure biometric information, means for conducting a voice dialogue with the user, means for analyzing the voice data to evaluate the health state and emotions, means for integrating the acquired data to perform a comprehensive health evaluation and emotion evaluation, means for generating a health and emotion report, and means for displaying the generated health and emotion report. This enables the user to easily and accurately monitor their own health and emotional state every day.

[0748] The "means for acquiring a facial image" is a device such as a camera for capturing an image of the user's face.

[0749] The "means for analyzing facial images and measuring biometric information" refers to software and hardware for analyzing biometric indicators such as heart rate, blood pressure, and body temperature based on the captured facial images.

[0750] The "means for performing voice interaction with a user" is a device for performing voice communication with a user using a microphone and a speaker.

[0751] The "means for analyzing voice data to assess health status and emotions" is software for analyzing the user's voice data and inferring health status and emotional state from the tone and speed of the voice.

[0752] "Means for integrating acquired data to provide a comprehensive health and emotional assessment" refers to a processor or algorithm that integrates collected biometric information, audio data, and other relevant data to provide a comprehensive assessment of the user's health and emotional state.

[0753] The "means for generating health and emotional reports" is software for automatically generating health and emotional reports based on the integrated data.

[0754] The "means for displaying the generated health and emotion report" is a display or monitor for visually presenting the generated report to the user.

[0755] The present invention relates to an AI-equipped mirror system that evaluates a user's health and emotional state using facial video and audio data. This system allows a user to receive multifaceted health and emotional evaluations in real time by standing in front of the mirror. Specific means and methods for implementing the present invention are described below.

[0756] Hardware used

[0757] Camera: Used to capture high-resolution footage, specifically capturing detailed footage of the face.

[0758] Processor: A processing device for analyzing video and audio data. For example, analyzing biometric information (heart rate, blood pressure, body temperature, etc.) from facial images.

[0759] Microphone and speaker: Used for voice interaction with the user, collecting information about the user's lifestyle habits and current emotional state through the interaction.

[0760] Display: A display device for presenting the generated health and emotion reports to the user.

[0761] Software used

[0762] Facial recognition algorithm: Analyzes facial images captured by the camera to recognize and identify the user.

[0763] Biometric analysis software: Biometric information is measured based on captured video data. This analysis uses changes in facial color and subtle muscle movements.

[0764] Voice analysis software: Analyzes voice data collected through a microphone to assess stress and fatigue levels.

[0765] Emotion engine: Analyzes facial expressions and subtle muscle movements to recognize the user's emotional state (happiness, anger, surprise, etc.).

[0766] Data integration algorithm: Integrates facial video data, audio data, and emotion data to perform comprehensive health and emotion assessments.

[0767] Report generation software: Generates health and emotional reports based on the integrated data.

[0768] System Operation

[0769] When a user stands in front of the mirror, the system captures facial images through a camera and identifies the user using a facial recognition algorithm. The processor analyzes the image data and measures biometric information (heart rate, blood pressure, body temperature, etc.). Next, a voice dialogue with the user is conducted through a microphone and speaker to collect information about the user's lifestyle habits and emotional state. The collected voice data is analyzed to evaluate the user's stress and fatigue levels. An emotion engine analyzes facial expressions and subtle muscle movements to recognize the user's emotional state. The server integrates this data to provide a comprehensive health and emotion assessment. Finally, report generation software creates health and emotion reports and displays them on the display. If any abnormalities are detected, a report is created with appropriate warning messages.

[0770] Specific examples

[0771] 1. The user stands in front of the mirror

[0772] Every morning, when the user stands in front of the mirror, facial images are automatically captured.

[0773] 2. Facial Recognition

[0774] The device captures the user's face through the camera and identifies them as "Yamada-san" by comparing it with a previously registered database.

[0775] 3. Video Analysis

[0776] The camera captures detailed images of the user's face, and the processor measures heart rate, blood pressure, body temperature, etc. At the same time, subtle changes in facial color and facial expressions are analyzed.

[0777] 4. Starting a voice conversation

[0778] The device uses a microphone and speaker to ask, "How much sleep did you get last night?" The user responds, "I slept for 6 hours."

[0779] 5. Analysis of audio data

[0780] The device analyzes the user's voice data and detects stress levels from the tone and speed of the voice.

[0781] 6. Emotional Recognition

[0782] The emotion engine analyzes the user's facial expressions and recognizes their current emotions, such as whether they are happy, angry, or surprised.

[0783] 7. Data integration and analysis

[0784] The collected facial video data, audio data, lifestyle data, and emotional data are sent to a server, and a comprehensive health and emotional assessment is performed.

[0785] 8. Generate health and emotional reports

[0786] Based on the evaluation results, the server generates a health and mood report for the user, including a warning such as, "Your heart rate is high today. You may need to rest."

[0787] 9. View the report

[0788] The device displays health and emotion reports on the mirror and provides automated voice feedback to the user, such as "Good morning. You're feeling stressed today. We recommend you relax."

[0789] Prompt Sentence Examples

[0790] Enter the following prompt into the generative AI model:

[0791] "This system uses an AI-equipped mirror to assess the user's health and emotional state. When the user stands in front of the mirror, a camera captures an image of their face and a voice dialogue begins. The acquired data is sent to a server for integrated analysis. Can you give us a concrete example?"

[0792] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0793] Step 1:

[0794] The system automatically activates when a user stands in front of the mirror. The input is the user's presence, which activates the camera and the device captures an image of the user's face. This action allows the system to capture an image of the user's face when the user says "Good morning."

[0795] Step 2:

[0796] The device uses a built-in camera to recognize the user's face. The input is a captured image of the face, and the output is the user's identification. Specifically, the facial recognition algorithm analyzes the image data and compares it with a database to identify the user as "Yamada-san."

[0797] Step 3:

[0798] The device uses a high-resolution camera to capture detailed facial images and analyze biometric information. The input is the recognized user's video data, and the output is biometric information such as heart rate, blood pressure, and body temperature. The processor analyzes subtle changes in facial color and muscle movements to make a diagnosis, such as whether the user has a pale complexion.

[0799] Step 4:

[0800] The device initiates a voice dialogue with the user using a microphone and speaker. The input is the user's voice, and the output is voice data. For example, the system asks, "Good morning. How much sleep did you get last night?" and the user replies, "I slept for six hours."

[0801] Step 5:

[0802] The device analyzes the acquired voice data and evaluates stress and fatigue levels based on the tone and speed of the voice. The input is the acquired voice data, and the output is the evaluation result of the user's stress and fatigue. For example, if the voice is quiet and low-pitched, the device will diagnose, "You seem a little tired."

[0803] Step 6:

[0804] The device uses an emotion engine to analyze facial expressions and subtle muscle movements to recognize emotions. The input is video data of the user's face, and the output is the recognition result of the user's emotional state (happiness, anger, surprise, etc.). For example, if the user's brow is furrowed, the device will evaluate the user as "slightly anxious."

[0805] Step 7:

[0806] The server integrates and analyzes the facial video data, audio data, question-and-answer data, and emotional data sent from the device. The input is multiple data sets, and the output is a comprehensive health and emotional assessment result. Based on this, the server determines that "the user is experiencing accumulated fatigue."

[0807] Step 8:

[0808] The server generates a health and emotion report based on the assessment results. The input is the integrated and analyzed assessment results, and the output is a health and emotion report. For example, it may include a warning message such as "Your heart rate is high today. Please take a rest."

[0809] Step 9:

[0810] The terminal displays the generated report on the mirror and issues a warning if necessary. The input is the generated report, and the output is the displayed report and audio feedback. Specifically, it tells the user, "Good morning. You are feeling stressed today. We recommend that you relax."

[0811] (Application example 2)

[0812] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0813] In modern society, especially in brick-and-mortar stores, there is a need to understand customers' health and emotional states in real time and provide services accordingly. However, conventional methods require customers to self-report their health and emotional states or have specialized staff manually monitor them, which has issues with inefficiency and inaccuracy. The present invention aims to provide a system that solves these issues and significantly improves customer experience.

[0814] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a facial image, means for analyzing the facial image to measure biometric information, means for engaging in voice dialogue with the user, means for analyzing voice data to evaluate a health condition, means for integrating the acquired data to perform a comprehensive health evaluation, means for generating a health report, means for displaying the generated health report, and means for displaying the health evaluation and emotional evaluation on a display device in the store or on a staff member's device to provide personalized service. This allows customers to understand their own health and emotional state in real time, and enables store staff to provide appropriate and personalized service to customers.

[0815] "Means for acquiring facial images" refers to a device or equipment for capturing a user's facial image in real time.

[0816] "Means for analyzing facial images to measure biometric information" refers to algorithms and software for analyzing captured facial images and measuring biometric information such as heart rate, blood pressure, and body temperature.

[0817] "Means for audio interaction with the user" refers to a microphone and speaker, and associated interaction software, for asking questions of the user and receiving responses.

[0818] "Means for analyzing voice data to assess health status" refers to software or algorithms for analyzing a user's voice data and assessing health status from the tone and rate of the voice data.

[0819] "Means for integrating acquired data to conduct a comprehensive health assessment" refers to a server or analysis system for integrating multiple data, such as facial video data and audio data, to conduct a comprehensive health assessment.

[0820] "Means for generating a health report" refers to software or algorithms that automatically generate a report summarizing health status based on the results of a comprehensive health assessment.

[0821] "Means for displaying the generated health report" refers to a display or display device, and associated software, for displaying the generated health report so that the user can review it.

[0822] "Means for displaying health and emotional assessments on in-store display devices or staff devices to provide personalized services" refers to systems and software that grasp customers' health and emotional status in real time, display this information on in-store displays or staff smart devices, and provide personalized services.

[0823] The present invention relates to a system that assesses a customer's health and emotional state in real time and provides personalized services in physical stores. When a customer enters a store, the system captures and analyzes facial images and evaluates their heart rate, blood pressure, body temperature, emotional state, etc. to provide an optimal customer experience.

[0824] Key elements of the system

[0825] The system includes the following main elements:

[0826] 1. Camera that captures facial images:

[0827] This camera uses a high-resolution camera (e.g., a high-resolution USB camera) that is placed on a mirror or other suitable location within the store.

[0828] 2. Video analysis and biometrics processor:

[0829] The video is analyzed using a facial recognition algorithm using the OpenCV library, and biometric information is measured using algorithms such as photoplethysmography.

[0830] 3. Microphone and speaker for voice interaction with the user:

[0831] For voice interaction, a high-precision microphone (e.g., a high-performance USB microphone) and a speaker are used, and voice analysis is performed via the Google Speech-to-Text API.

[0832] 4. Software that analyzes voice data to assess health and emotions:

[0833] To analyze voice data, the Google Speech-to-Text API and AI models such as TensorFlow are used.

[0834] 5. Server that integrates data to provide a comprehensive health assessment:

[0835] The server uses a cloud-based server (e.g., a public cloud service provider) to consolidate the collected data and perform real-time evaluation.

[0836] 6. Display to generate and show health reports:

[0837] Health reports are generated using a Python Flask-based web application and are displayed on displays in-store and on staff's smart devices.

[0838] 7. System for displaying health and emotional assessments and providing personalized services:

[0839] Use a system (e.g., personalized assistance application) that allows store staff to provide personalized service based on customer ratings.

[0840] Specific Examples

[0841] When a customer enters a store, a high-resolution camera inside the store captures an image of the customer's face, and at the same time, a voice conversation begins using a microphone and speaker.

[0842] The resulting facial and audio data is sent to video and audio analysis software, where the facial data is used to measure vital signs such as heart rate, blood pressure, and body temperature, while the audio data is analyzed to assess emotional state.

[0843] The server integrates this data to perform a comprehensive evaluation and generate health and emotion reports, which are displayed in real time on displays in the store and on staff's smart devices.

[0844] Staff can use this information to provide customers with personalized service, for example, suggesting products that will help them relax if they are feeling tired.

[0845] Prompt Sentence Examples

[0846] text

[0847] Develop a system that will capture the customer's facial image and voice in real time using a mirror in the store when they enter the store, evaluate their health and emotional state based on their heart rate, blood pressure, body temperature, tone and speed of voice, and integrate the collected data on a server to provide personalized services. The specifications will use OpenCV for facial recognition, Google Speech-to-Text API for voice analysis, Affectiva SDK for emotion recognition, and Python Flask for data integration.

[0848] In this way, the present invention allows for real-time assessment of a customer's health and emotional state to provide optimal service, thereby improving the customer experience and the service quality of the store.

[0849] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0850] Step 1:

[0851] When a user enters a store, the device's camera captures the user's facial image. This camera has high resolution and captures images in real time. The input is the user's facial image, and the output is the captured facial image data.

[0852] Step 2:

[0853] The facial image data acquired by the device is sent to a facial recognition algorithm. This algorithm uses the OpenCV library to identify the user's face. The input is the facial image data, and the output is facial recognition data in which the user's face is identified.

[0854] Step 3:

[0855] Based on the facial recognition data, the device measures the user's biometric information. Specifically, it uses video analysis software to measure heart rate, blood pressure, body temperature, etc. Photoplethysmography technology is used for this purpose. The input is facial recognition data, and the output is biometric data.

[0856] Step 4:

[0857] The terminal starts a voice dialogue with the user, asks a question through a microphone, and obtains the user's response. For example, the terminal asks, "What kind of product are you looking for today?" The input is the user's voice response, and the output is the obtained voice data.

[0858] Step 5:

[0859] The device sends the captured voice data to the Google Speech-to-Text API, which converts the voice into text data. Furthermore, speech analysis software is used to evaluate the user's emotional state based on the tone and speed of the voice. The input is voice data, and the output is textual voice data and emotional state data.

[0860] Step 6:

[0861] The device sends facial video data and audio data to a server, which then integrates the data to perform comprehensive health and emotional assessments. The inputs are biometric data, audio text data, and emotional state data, and the output is an integrated health and emotional assessment report.

[0862] Step 7:

[0863] The server returns the generated health and emotion evaluation reports to the terminal, which then displays these reports on displays in the store or on staff members' smart devices. The input is the health and emotion evaluation reports, and the output is the displayed reports.

[0864] Step 8:

[0865] Store staff provide personalized services based on the user's state. For example, if the user is stressed, they will suggest products that will help them relax. The staff responds to the customer while referring to the display on the terminal. The input is the displayed report, and the output is the service provided.

[0866] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0867] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0868] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0869] [Third embodiment]

[0870] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0871] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0872] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0873] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0874] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0875] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0876] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0877] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0878] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0879] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0880] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0881] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0882] This invention relates to an AI-equipped mirror system that analyzes facial images and audio to assess the user's health condition. This system allows the user to receive a multifaceted health assessment in real time by standing in front of the mirror.

[0883] System Overview

[0884] The system includes the following main elements:

[0885] 1. Camera that captures facial images

[0886] 2. Processor for video analysis and biometric measurement

[0887] 3. A microphone and speaker for voice interaction with the user

[0888] 4. Software that analyzes voice data to assess health status

[0889] 5. Server that integrates data and provides comprehensive health assessment

[0890] 6. Display to generate and display health reports

[0891] Program processing flow

[0892] The program of this system is designed as follows:

[0893] 1. The user stands in front of the AI-enabled mirror

[0894] Every morning, the user stands in front of the mirror and looks at their face, which automatically activates the system.

[0895] 2. The device recognizes the user's face

[0896] The device uses a built-in camera to capture an image of your face and uses a facial recognition algorithm to identify you, which includes matching it against a historical database.

[0897] 3. The device analyzes your face and muscle movements

[0898] The device uses a high-resolution camera to capture detailed facial video data, allowing it to analyze and measure biometric information such as heart rate, blood pressure, and body temperature.

[0899] 4. The device starts a dialogue with the user

[0900] The device initiates a voice dialogue, asking a question such as, "How much sleep did you get last night?" When the user responds, the voice data is acquired.

[0901] 5. The device analyzes the voice data

[0902] The device analyzes the user's voice and assesses stress and fatigue levels based on the tone and speed of the voice.

[0903] 6. The server integrates and analyzes the data

[0904] The server integrates the facial video data, audio data, and question-and-answer data sent from the device to perform a comprehensive health assessment.

[0905] 7. The server generates a health report

[0906] The server generates a health report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[0907] 8. The device will display a health report

[0908] The device will display a health report on the mirror and issue warnings as needed, such as telling the user, "You're feeling stressed today. We recommend taking some time to relax."

[0909] Specific examples

[0910] 1. The user stands in front of the mirror

[0911] Every morning, when the user stands in front of the mirror, facial images are automatically captured.

[0912] 2. Facial Recognition

[0913] The device captures the user's face through the camera and identifies them as "Yamada-san" by comparing it with a previously registered database.

[0914] 3. Video Analysis

[0915] The camera captures detailed images of the user's face, and the processor measures heart rate, blood pressure, body temperature, etc. At the same time, subtle changes in facial color and facial expressions are analyzed.

[0916] 4. Starting a voice conversation

[0917] The device uses a microphone and speaker to ask, "How much sleep did you get last night?" The user responds, "I slept for 6 hours."

[0918] 5. Analysis of audio data

[0919] The device analyzes the user's voice data and detects stress levels from the tone and speed of the voice.

[0920] 6. Data integration and analysis

[0921] The collected facial video data, audio data, and lifestyle data are sent to a server for a comprehensive health assessment.

[0922] 7. Generate Health Reports

[0923] Based on the evaluation results, the server generates a health report for the user, including a warning such as, "Your heart rate is high today. You may need to rest."

[0924] 8. View Health Report

[0925] The device displays a health report on the mirror and notifies the user with an automated voice.

[0926] In this way, users can easily monitor their daily health status and take necessary measures. The present invention significantly improves daily health management without relying solely on regular medical checkups.

[0927] The processing flow will be explained below.

[0928] Step 1:

[0929] The user stands in front of the mirror, which automatically activates the system and prepares the camera to capture the user's face.

[0930] Step 2:

[0931] The device uses its built-in camera to capture the user's face, then uses a facial recognition algorithm to identify the user and match them with a historical database.

[0932] Step 3:

[0933] The device uses a high-resolution camera to capture detailed facial video data, including technology to capture facial biometric information such as heart rate, blood pressure, and temperature.

[0934] Step 4:

[0935] The device starts a voice dialogue with the user. For example, it asks, "Good morning. How long did you sleep today?" The user replies, "I slept for 6 hours."

[0936] Step 5:

[0937] The device analyzes the voice data it acquires, and uses voice analysis technology to assess stress and fatigue levels based on the tone and speed of the user's voice.

[0938] Step 6:

[0939] The device analyzes facial expressions and changes in color tone to measure heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, hemoglobin concentration, stress, and fatigue level.

[0940] Step 7:

[0941] The facial image data, audio data, and question and answer data acquired by the terminal are transmitted to the server.

[0942] Step 8:

[0943] The server then combines the received data to provide a comprehensive health assessment, including heart rate and blood pressure variability, voice analysis results, and lifestyle information.

[0944] Step 9:

[0945] The server generates a health report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[0946] Step 10:

[0947] The device receives the health report sent from the server and displays it on the mirror, for example, displaying a message such as "Your heart rate is high today. You may need to rest."

[0948] Step 11:

[0949] The device will notify the user with an automated voice, providing feedback such as, "Good morning. You're feeling stressed today. We recommend you relax."

[0950] Example 1

[0951] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0952] Conventional health management systems require users to manually input information, which can lead to problems with accuracy and effort. In addition, the long intervals between regular health checkups make it difficult to quickly detect changes in health status, making it impossible to grasp health status in real time.

[0953] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0954] In this invention, the server includes means for acquiring a facial image, means for analyzing the acquired facial image to measure biometric information, means for engaging in voice dialogue with the user, means for analyzing the voice data to evaluate the health condition, means for integrating the acquired video data and voice data to perform a comprehensive health evaluation, means for generating a health report, and means for displaying the generated health report, thereby enabling the user to receive a comprehensive health evaluation in real time simply by standing in front of the mirror.

[0955] "Means for capturing facial images" refers to a camera or other imaging device that captures an image or video of a user's face.

[0956] "Means for analyzing captured facial images and measuring biometric information" refers to processors and algorithms that analyze and measure biometric information such as heart rate, blood pressure, and body temperature from captured facial image data.

[0957] "Means for conducting voice dialogue with the user" refers to a microphone and speaker for communicating with the user by voice, and software for controlling the dialogue.

[0958] "Means for analyzing voice data to assess health status" refers to algorithms or analytical software for analyzing voice data and assessing a user's stress level, fatigue level, etc.

[0959] "Means for integrating acquired video and audio data to perform a comprehensive health assessment" refers to a server and analysis program for integrating facial video and audio data to assess overall health status.

[0960] "Means for generating a health report" refers to software or algorithms that automatically generate a report showing the user's health status based on the analysis results.

[0961] "Means for displaying the generated health report" refers to a display or screen for visually presenting the generated health report to the user.

[0962] This invention relates to an AI-equipped mirror system that analyzes the user's facial image and voice to evaluate their health condition in real time. To implement this system, the following hardware and software are used:

[0963] Hardware and software used

[0964] 1. Camera: Used to capture video of the user's face. For example, a 1080p HD camera or a 4K camera is suitable.

[0965] 2. Processor: Used for video analysis and biometric measurement. Examples include an Intel Core i7 processor or an NVIDIA GPU.

[0966] 3. Microphone and speaker: Used for voice interaction with the user, a high-sensitivity microphone and stereo speakers are desirable to provide clear sound quality.

[0967] 4. Analysis software: Software used to analyze video and audio data often uses OpenCV, TensorFlow, or Google Cloud Speech-to-Text API.

[0968] 5. Server: A backend server that integrates data and performs comprehensive health assessments. This is where AI models using TensorFlow and PyTorch run.

[0969] 6. Display: Used to display the generated health report. A display that also functions as a mirror is desirable.

[0970] How the system works

[0971] This system automatically starts when the user stands in front of the mirror and performs the following processes.

[0972] Image acquisition by camera

[0973] When a user stands in front of the mirror, the camera automatically captures an image of the user's face.

[0974] Facial image analysis

[0975] The processor analyzes the captured facial video data. For example, it detects subtle color changes and facial expressions, and uses this data to estimate heart rate, blood pressure, and body temperature. This analysis uses the OpenCV library.

[0976] Voice interaction with the user

[0977] The device uses a microphone and speaker to ask the user questions, such as prompting them with questions like, "How much sleep did you get last night?" If the user responds, "I slept for six hours," the device collects the audio data.

[0978] Analysis of audio data

[0979] The device converts voice data into text using the Google Cloud Speech-to-Text API, and then analyzes the content, tone, and speed of the voice, allowing it to assess stress and fatigue levels.

[0980] Data integration and analysis

[0981] The server receives the facial video data, audio data, and question-and-answer data sent from the device and performs an integrated analysis. A comprehensive health assessment is performed using an AI model built with TensorFlow and PyTorch.

[0982] Generate health reports

[0983] The server generates a detailed health report based on the analysis results, including a warning such as "Your heart rate is high today. You may need to rest" if the analysis results show a higher than normal heart rate.

[0984] View Health Report

[0985] Finally, the device displays the generated health report on the screen and, if necessary, notifies the user by voice.

[0986] Prompt Sentence Examples

[0987] Examples of prompts the system might give to the user include:

[0988] "Good morning. How much sleep did you get last night?"

[0989] "How are you feeling today?"

[0990] "Are you stressed?"

[0991] These specific actions allow users to easily monitor their daily health status and quickly take necessary measures.

[0992] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0993] Step 1:

[0994] When a user stands in front of the mirror, the device detects movement and automatically activates the camera and microphone. The input is the user's movement, and the output is the activation of the camera and microphone. Specifically, the sensor detects the user's presence, and the camera begins capturing video.

[0995] Step 2:

[0996] The device recognizes the user's face. The camera captures an image of the user's face, and this video data is input into a facial recognition algorithm. The algorithm compares it with a previously saved database to identify the user, and the resulting username is output. A specific example of how this works is to identify the user as "Yamada-san" using facial recognition technology.

[0997] Step 3:

[0998] The device analyzes images of the user's face and measures their biometric information. The input is high-resolution video data acquired from a camera, which the processor analyzes to estimate heart rate, blood pressure, and body temperature. The output is this biometric information. Specifically, it uses computer vision technology to detect subtle color changes and muscle movements on the face to measure heart rate and blood pressure.

[0999] Step 4:

[1000] The device starts a voice dialogue with the user. It uses a microphone and speaker to ask a question and obtains the answer. The input is a prompt such as "How much sleep did you get last night?" and the user's voice response, and the output is voice data. Specifically, the system makes the user respond with "I slept for 6 hours."

[1001] Step 5:

[1002] The device analyzes the voice data. The input is the captured voice data, which the software converts into text and then analyzes the tone and speed. The output is an assessment of the user's stress level and fatigue. Specifically, it uses voice analysis technology to determine that a high-pitched voice and fast speed indicate high stress.

[1003] Step 6:

[1004] The server integrates and analyzes the data. Facial video data, audio data, and question-and-answer data sent from the device are input, and the server integrates these to perform a comprehensive health assessment. The output is a detailed health assessment result. Specifically, the system uses an AI model to analyze multidimensional data and evaluate the overall health condition.

[1005] Step 7:

[1006] The server generates a health report. The input is the integrated and analyzed data, and the software uses this to create a health report. The output is the user's health report. Specific operations include generating a report with a warning, such as "Your heart rate is high today. You may need to rest."

[1007] Step 8:

[1008] The device displays the generated health report. The input is the health report, which is visually displayed on the display. The output is visual feedback to the user. Specifically, the device displays "Your heart rate is high today. We recommend you rest" and notifies you by voice if necessary.

[1009] Through the above processing steps, users can easily understand their daily health condition in real time and take necessary measures promptly.

[1010] (Application example 1)

[1011] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1012] Conventional health management systems require users to voluntarily collect and input data, limiting their practicality and continuity. Furthermore, it is difficult to monitor detailed health conditions outside of regular health checkups. The present invention aims to provide a system that more efficiently and instantly assesses a user's health condition. In particular, the system can be installed in a brick-and-mortar store, solving the problem of enabling users to naturally manage their health as part of their daily activities.

[1013] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1014] In this invention, the server includes means for acquiring a facial image, means for analyzing the facial image to measure biometric information, means for engaging in voice dialogue with the user, means for analyzing the voice data to evaluate the health condition, means for integrating the acquired data to perform a comprehensive health evaluation, means for generating a health report, means for displaying the generated health report, and means for instantly evaluating the user's health using a mirror installed in the physical store. This allows the user's health condition to be automatically evaluated every time they visit the store, providing instant feedback and facilitating daily health management.

[1015] The "means for acquiring a facial image" refers to hardware such as a camera or image sensor for capturing an image of the user's face, and software for controlling the hardware.

[1016] The "means for analyzing facial images to measure biometric information" refers to software and algorithms for analyzing captured facial images and calculating biometric information such as heart rate, blood pressure, and body temperature.

[1017] The "means for conducting a voice dialogue with the user" refers to hardware and software for controlling the hardware and the speaker for conducting a voice dialogue with the user using a microphone and speaker.

[1018] The "means for analyzing voice data to assess health status" refers to software and algorithms for analyzing acquired voice data and assessing the user's stress level and fatigue level.

[1019] The "means for integrating acquired data to perform a comprehensive health assessment" refers to software and a server system for integrating facial video data and audio data to assess the user's overall health condition.

[1020] The "means for generating a health report" is software for creating a report that reports the user's health status based on the integrated health data.

[1021] The "means for displaying the generated health report" refers to a display and its control software for visually displaying the generated health report to the user.

[1022] The "means for instantly assessing a user's health using a mirror installed in a physical store" is a system for instantly assessing the health of users who visit the store and displaying the results through a mirror-like device installed in the store.

[1023] This invention relates to an "AI Health Check Mirror System" installed in a brick-and-mortar store, where users can stand in front of the mirror and have their facial images and voices analyzed, and receive instant feedback on their health condition for the day. The specific implementation of this system is described below.

[1024] System configuration

[1025] The system includes the following main elements:

[1026] 1. Camera: A means of capturing facial images, such as a high-resolution camera like the Logitech C920.

[1027] 2. Processor: A hardware device that analyzes facial images and measures biometric information. Various algorithms are embedded in the processor.

[1028] 3. Microphone and speaker: A means of audio interaction with the user. For example, a high-performance microphone such as the Blue Yeti.

[1029] 4. Voice analysis software: A means of analyzing voice data to assess health status. Custom algorithms for voice analysis.

[1030] 5. Server: A means of integrating acquired data to provide a comprehensive health assessment. Often, a cloud-based server is used.

[1031] 6. Display: Means for generating and displaying health reports. Display device that acts as a mirror.

[1032] 7. Physical store mirror: Installed in a physical store, it provides an instant health assessment of the user.

[1033] Data collection and analysis

[1034] 1. Acquiring facial images

[1035] When a user stands in front of the AI ​​health check mirror installed in a physical store, the camera automatically captures the user's facial image, which automatically activates the system.

[1036] 2. Measurement of biological information

[1037] The captured facial image is analyzed by a processor to measure biometric information such as heart rate, blood pressure, and body temperature. Subtle changes in facial color and facial expressions are also analyzed.

[1038] 3. Starting a voice conversation

[1039] When a user speaks to the system, a microphone picks up the voice and a voice interaction takes place through the speaker. The system asks questions, such as, "How much sleep did you get last night?"

[1040] 4. Analysis of audio data

[1041] The captured voice data is analyzed using voice analysis software, and stress and fatigue levels are assessed based on the tone and speed of the voice.

[1042] 5. Data integration and comprehensive health assessment

[1043] The server integrates facial video data, audio data, and voice dialogue data to provide a comprehensive health assessment, which includes various biometric data and voice analysis results.

[1044] 6. Generate and view health reports

[1045] The server generates a health report for the user based on the analysis results. If any abnormalities are detected, a report including a warning message is generated. The generated health report is displayed on the mirror and a warning is issued as needed. For example, feedback is given in the form of "Your heart rate is high today. You may need to rest."

[1046] Specific examples

[1047] When you stand in front of an AI health check mirror installed in a physical store, the system will automatically activate and capture an image of your face.

[1048] The images captured by the camera are analyzed by a processor to measure heart rate and blood pressure.

[1049] When the user looks into the mirror and says, "I slept six hours last night," voice analysis is performed to assess stress and fatigue levels.

[1050] The server integrates this data, performs a comprehensive health assessment, and then displays advice such as, "Your heart rate is high today. You may need to rest."

[1051] Example prompts for generative AI models

[1052] "Please explain in detail about the system that analyzes facial images and audio to assess the user's health condition. The system estimates biometric information such as heart rate, blood pressure, and body temperature by having the user stand in front of a mirror, and evaluates stress levels from the tone and rate of voice. Based on the assessment results, the system generates a health report and provides advice."

[1053] This allows users to easily monitor their daily health status and improve their ability to take necessary measures. It is expected that the present invention will significantly improve health management not only through conventional periodic health checkups but also through everyday activities.

[1054] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1055] Step 1:

[1056] When a user stands in front of the mirror, the device activates the built-in camera to capture an image of the user's face.

[1057] Input: Face image taken by camera

[1058] Output: Acquired facial image data

[1059] What it does: The system automatically activates the camera and captures high-resolution facial footage, which is then stored for further processing.

[1060] Step 2:

[1061] The facial image captured by the device is sent to a processor, which analyzes the facial image and measures biometric information.

[1062] Input: Acquired facial video data

[1063] Output: vital signs such as heart rate, blood pressure, and temperature

[1064] How it works: Video analysis algorithms analyze subtle changes in facial color and facial expressions to measure biometric information. A processor extracts and stores this data.

[1065] Step 3:

[1066] The device initiates a voice dialogue with the user. For example, the device asks, "How much sleep did you get last night?"

[1067] Input: User voice input

[1068] Output: Acquired audio data

[1069] What it does: A microphone picks up the user's voice and records the audio in an interactive format. A speaker is used to ask questions and record the user's responses.

[1070] Step 4:

[1071] The device analyzes the voice data it acquires and assesses the user's health condition.

[1072] Input: User's voice data

[1073] Output: User's voice analysis results (stress level, fatigue level, etc.)

[1074] What it does: Voice analysis software analyzes the tone and rate of the user's voice to assess stress and fatigue levels. These assessments are used in the next step.

[1075] Step 5:

[1076] The device sends facial video and audio data to a server, which then integrates the data to provide a comprehensive health assessment.

[1077] Input: Facial video data, audio data, audio analysis results

[1078] Output: Comprehensive health assessment data

[1079] How it works: The server combines facial video and audio data, and evaluates the user's overall health based on biometric and audio analysis results. An algorithm performs this evaluation and generates a result.

[1080] Step 6:

[1081] The server generates a health report. The server creates a report based on the evaluation results and adds warning messages if necessary.

[1082] Input: Comprehensive health assessment data

[1083] Output: Health report

[1084] What happens: The server generates a personalized health report based on the evaluation results, including a warning message if any abnormalities are detected. The report is displayed in the next step.

[1085] Step 7:

[1086] The health report generated by the device is displayed on the mirror and notified to the user.

[1087] Input: Health Report

[1088] Output: Health report displayed on mirror, audio notification

[1089] What it does: The device displays the generated health report on the mirror and notifies the user through the speaker, providing feedback such as, "Your heart rate is high today. You may need to rest."

[1090] This series of processes allows the user to quickly know their health condition and take measures as needed.

[1091] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1092] This invention relates to an AI-equipped mirror system that analyzes facial images and voices to evaluate the user's health and emotional state. This system allows the user to receive a multifaceted health and emotional state evaluation in real time by standing in front of the mirror.

[1093] System Overview

[1094] The system includes the following main elements:

[1095] 1. Camera that captures facial images

[1096] 2. Processor for video analysis and biometric measurement

[1097] 3. A microphone and speaker for voice interaction with the user

[1098] 4. Software that analyzes voice data to assess health status and emotions

[1099] 5. Server that integrates data and provides comprehensive health assessment

[1100] 6. Display to generate and display health reports

[1101] 7. Emotion engine that recognizes user emotions

[1102] Program processing flow

[1103] The program of this system is designed as follows:

[1104] 1. The user stands in front of the AI-enabled mirror

[1105] Every morning, the user stands in front of the mirror and looks at their face, which automatically activates the system.

[1106] 2. The device recognizes the user's face

[1107] The device uses a built-in camera to capture an image of your face and uses a facial recognition algorithm to identify you, which includes matching it against a historical database.

[1108] 3. The device analyzes your face and muscle movements

[1109] The device uses a high-resolution camera to capture detailed facial video data, allowing it to analyze and measure biometric information such as heart rate, blood pressure, and body temperature.

[1110] 4. The device starts a dialogue with the user

[1111] The terminal starts a voice dialogue, asking, for example, "Good morning. How long did you sleep today?" When the user responds, the voice data is acquired.

[1112] 5. The device analyzes the voice data

[1113] The device analyzes the user's voice and assesses stress and fatigue levels based on the tone and speed of the voice.

[1114] 6. The device analyzes facial images and recognizes emotions

[1115] The emotion engine analyzes the user's facial expressions and subtle muscle movements to recognize emotions such as joy, anger, and surprise.

[1116] 7. The server integrates and analyzes the data

[1117] The server integrates the facial video data, audio data, question and answer data, and emotional data sent from the device to perform a comprehensive health and emotional assessment.

[1118] 8. The server generates health and emotion reports

[1119] The server generates a health and emotional report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[1120] 9. The device displays the report

[1121] The device will display health and emotional reports on the mirror and issue warnings as needed, such as telling the user, "You're feeling stressed today. We recommend taking some time to relax."

[1122] Specific examples

[1123] 1. The user stands in front of the mirror

[1124] Every morning, when the user stands in front of the mirror, facial images are automatically captured.

[1125] 2. Facial Recognition

[1126] The device captures the user's face through the camera and identifies them as "Yamada-san" by comparing it with a previously registered database.

[1127] 3. Video Analysis

[1128] The camera captures detailed images of the user's face, and the processor measures heart rate, blood pressure, body temperature, etc. At the same time, subtle changes in facial color and facial expressions are analyzed.

[1129] 4. Starting a voice conversation

[1130] The device uses a microphone and speaker to ask, "How much sleep did you get last night?" The user responds, "I slept for 6 hours."

[1131] 5. Analysis of audio data

[1132] The device analyzes the user's voice data and detects stress levels from the tone and speed of the voice.

[1133] 6. Emotional Recognition

[1134] The emotion engine analyzes the user's facial expressions and recognizes their current emotions, such as whether they are happy, angry, or surprised.

[1135] 7. Data integration and analysis

[1136] The collected facial video data, audio data, lifestyle data, and emotional data are sent to a server, and a comprehensive health and emotional assessment is performed.

[1137] 8. Generate health and emotional reports

[1138] Based on the evaluation results, the server generates a health and mood report for the user, including a warning such as, "Your heart rate is high today. You may need to rest."

[1139] 9. View the report

[1140] The device displays health and emotion reports on the mirror and provides automated voice feedback to the user, such as "Good morning. You're feeling stressed today. We recommend you relax."

[1141] In this way, users can easily monitor their daily health and emotional state and take necessary measures. The present invention significantly improves daily health management and emotional management without relying solely on regular medical checkups.

[1142] The processing flow will be explained below.

[1143] Step 1:

[1144] The user stands in front of the mirror, which automatically activates the system and prepares the camera to capture the user's face.

[1145] Step 2:

[1146] The device uses its built-in camera to capture the user's face, then uses a facial recognition algorithm to identify the user and match them against a database to identify the user.

[1147] Step 3:

[1148] The device uses a high-resolution camera to capture detailed video data of the user's face, including technology to measure biometric information such as heart rate, blood pressure, and temperature from the face.

[1149] Step 4:

[1150] The terminal initiates a voice dialogue with the user, for example asking, "Good morning. How are you feeling today?" The user responds, "I'm feeling great."

[1151] Step 5:

[1152] The device analyzes the voice data it acquires, and uses voice analysis technology to assess stress and fatigue levels based on the tone and speed of the user's voice.

[1153] Step 6:

[1154] The device analyzes facial expressions and changes in color tone to measure heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, hemoglobin concentration, stress, and fatigue level.

[1155] Step 7:

[1156] The emotion engine analyzes the user's facial expressions and subtle muscle movements to recognize the user's emotional state. For example, if the user is smiling, it will recognize that they are feeling "happy."

[1157] Step 8:

[1158] The facial image data, audio data, question and answer data, and emotion data acquired by the terminal are transmitted to the server.

[1159] Step 9:

[1160] The server then integrates the received data to provide a comprehensive health and emotional assessment, including heart rate and blood pressure variability, voice analysis results, lifestyle information, and emotional state.

[1161] Step 10:

[1162] The server generates a health and emotion report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[1163] Step 11:

[1164] The device receives the health and emotion reports sent from the server and displays them on the mirror, for example, displaying a message such as "Your heart rate is high today. You may need to rest."

[1165] Step 12:

[1166] The device will notify the user with an automated voice, providing feedback such as, "Good morning. You seem to be feeling a little stressed today. I recommend you relax."

[1167] Example 2

[1168] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1169] Conventional health management systems often require specialized equipment at specific medical facilities to accurately assess a user's health status. This makes them inconvenient for everyday use and places a burden on users. Furthermore, they are unable to simultaneously assess the user's emotional state, making comprehensive health management difficult.

[1170] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring a facial image, means for analyzing the facial image to measure biometric information, means for conducting a voice dialogue with the user, means for analyzing the voice data to evaluate the health state and emotions, means for integrating the acquired data to perform a comprehensive health evaluation and emotion evaluation, means for generating a health and emotion report, and means for displaying the generated health and emotion report. This enables the user to easily and accurately monitor their own health and emotional state every day.

[1171] The "means for acquiring a facial image" is a device such as a camera for capturing an image of the user's face.

[1172] The "means for analyzing facial images and measuring biometric information" refers to software and hardware for analyzing biometric indicators such as heart rate, blood pressure, and body temperature based on the captured facial images.

[1173] The "means for performing voice interaction with a user" is a device for performing voice communication with a user using a microphone and a speaker.

[1174] The "means for analyzing voice data to assess health status and emotions" is software for analyzing the user's voice data and inferring health status and emotional state from the tone and speed of the voice.

[1175] "Means for integrating acquired data to provide a comprehensive health and emotional assessment" refers to a processor or algorithm that integrates collected biometric information, audio data, and other relevant data to provide a comprehensive assessment of the user's health and emotional state.

[1176] The "means for generating health and emotional reports" is software for automatically generating health and emotional reports based on the integrated data.

[1177] The "means for displaying the generated health and emotion report" is a display or monitor for visually presenting the generated report to the user.

[1178] The present invention relates to an AI-equipped mirror system that evaluates a user's health and emotional state using facial video and audio data. This system allows a user to receive multifaceted health and emotional evaluations in real time by standing in front of the mirror. Specific means and methods for implementing the present invention are described below.

[1179] Hardware used

[1180] Camera: Used to capture high-resolution footage, specifically capturing detailed footage of the face.

[1181] Processor: A processing device for analyzing video and audio data. For example, analyzing biometric information (heart rate, blood pressure, body temperature, etc.) from facial images.

[1182] Microphone and speaker: Used for voice interaction with the user, collecting information about the user's lifestyle habits and current emotional state through the interaction.

[1183] Display: A display device for presenting the generated health and emotion reports to the user.

[1184] Software used

[1185] Facial recognition algorithm: Analyzes facial images captured by the camera to recognize and identify the user.

[1186] Biometric analysis software: Biometric information is measured based on captured video data. This analysis uses changes in facial color and subtle muscle movements.

[1187] Voice analysis software: Analyzes voice data collected through a microphone to assess stress and fatigue levels.

[1188] Emotion engine: Analyzes facial expressions and subtle muscle movements to recognize the user's emotional state (happiness, anger, surprise, etc.).

[1189] Data integration algorithm: Integrates facial video data, audio data, and emotion data to perform comprehensive health and emotion assessments.

[1190] Report generation software: Generates health and emotional reports based on the integrated data.

[1191] System Operation

[1192] When a user stands in front of the mirror, the system captures facial images through a camera and identifies the user using a facial recognition algorithm. The processor analyzes the image data and measures biometric information (heart rate, blood pressure, body temperature, etc.). Next, a voice dialogue with the user is conducted through a microphone and speaker to collect information about the user's lifestyle habits and emotional state. The collected voice data is analyzed to evaluate the user's stress and fatigue levels. An emotion engine analyzes facial expressions and subtle muscle movements to recognize the user's emotional state. The server integrates this data to provide a comprehensive health and emotion assessment. Finally, report generation software creates health and emotion reports and displays them on the display. If any abnormalities are detected, a report is created with appropriate warning messages.

[1193] Specific examples

[1194] 1. The user stands in front of the mirror

[1195] Every morning, when the user stands in front of the mirror, facial images are automatically captured.

[1196] 2. Facial Recognition

[1197] The device captures the user's face through the camera and identifies them as "Yamada-san" by comparing it with a previously registered database.

[1198] 3. Video Analysis

[1199] The camera captures detailed images of the user's face, and the processor measures heart rate, blood pressure, body temperature, etc. At the same time, subtle changes in facial color and facial expressions are analyzed.

[1200] 4. Starting a voice conversation

[1201] The device uses a microphone and speaker to ask, "How much sleep did you get last night?" The user responds, "I slept for 6 hours."

[1202] 5. Analysis of audio data

[1203] The device analyzes the user's voice data and detects stress levels from the tone and speed of the voice.

[1204] 6. Emotional Recognition

[1205] The emotion engine analyzes the user's facial expressions and recognizes their current emotions, such as whether they are happy, angry, or surprised.

[1206] 7. Data integration and analysis

[1207] The collected facial video data, audio data, lifestyle data, and emotional data are sent to a server, and a comprehensive health and emotional assessment is performed.

[1208] 8. Generate health and emotional reports

[1209] Based on the evaluation results, the server generates a health and mood report for the user, including a warning such as, "Your heart rate is high today. You may need to rest."

[1210] 9. View the report

[1211] The device displays health and emotion reports on the mirror and provides automated voice feedback to the user, such as "Good morning. You're feeling stressed today. We recommend you relax."

[1212] Prompt Sentence Examples

[1213] Enter the following prompt into the generative AI model:

[1214] "This system uses an AI-equipped mirror to assess the user's health and emotional state. When the user stands in front of the mirror, a camera captures an image of their face and a voice dialogue begins. The acquired data is sent to a server for integrated analysis. Can you give us a concrete example?"

[1215] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1216] Step 1:

[1217] The system automatically activates when a user stands in front of the mirror. The input is the user's presence, which activates the camera and the device captures an image of the user's face. This action allows the system to capture an image of the user's face when the user says "Good morning."

[1218] Step 2:

[1219] The device uses a built-in camera to recognize the user's face. The input is a captured image of the face, and the output is the user's identification. Specifically, the facial recognition algorithm analyzes the image data and compares it with a database to identify the user as "Yamada-san."

[1220] Step 3:

[1221] The device uses a high-resolution camera to capture detailed facial images and analyze biometric information. The input is the recognized user's video data, and the output is biometric information such as heart rate, blood pressure, and body temperature. The processor analyzes subtle changes in facial color and muscle movements to make a diagnosis, such as whether the user has a pale complexion.

[1222] Step 4:

[1223] The device initiates a voice dialogue with the user using a microphone and speaker. The input is the user's voice, and the output is voice data. For example, the system asks, "Good morning. How much sleep did you get last night?" and the user replies, "I slept for six hours."

[1224] Step 5:

[1225] The device analyzes the acquired voice data and evaluates stress and fatigue levels based on the tone and speed of the voice. The input is the acquired voice data, and the output is the evaluation result of the user's stress and fatigue. For example, if the voice is quiet and low-pitched, the device will diagnose, "You seem a little tired."

[1226] Step 6:

[1227] The device uses an emotion engine to analyze facial expressions and subtle muscle movements to recognize emotions. The input is video data of the user's face, and the output is the recognition result of the user's emotional state (happiness, anger, surprise, etc.). For example, if the user's brow is furrowed, the device will evaluate the user as "slightly anxious."

[1228] Step 7:

[1229] The server integrates and analyzes the facial video data, audio data, question-and-answer data, and emotional data sent from the device. The input is multiple data sets, and the output is a comprehensive health and emotional assessment result. Based on this, the server determines that "the user is experiencing accumulated fatigue."

[1230] Step 8:

[1231] The server generates a health and emotion report based on the assessment results. The input is the integrated and analyzed assessment results, and the output is a health and emotion report. For example, it may include a warning message such as "Your heart rate is high today. Please take a rest."

[1232] Step 9:

[1233] The terminal displays the generated report on the mirror and issues a warning if necessary. The input is the generated report, and the output is the displayed report and audio feedback. Specifically, it tells the user, "Good morning. You are feeling stressed today. We recommend that you relax."

[1234] (Application example 2)

[1235] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1236] In modern society, especially in brick-and-mortar stores, there is a need to understand customers' health and emotional states in real time and provide services accordingly. However, conventional methods require customers to self-report their health and emotional states or have specialized staff manually monitor them, which has issues with inefficiency and inaccuracy. The present invention aims to provide a system that solves these issues and significantly improves customer experience.

[1237] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a facial image, means for analyzing the facial image to measure biometric information, means for engaging in voice dialogue with the user, means for analyzing voice data to evaluate a health condition, means for integrating the acquired data to perform a comprehensive health evaluation, means for generating a health report, means for displaying the generated health report, and means for displaying the health evaluation and emotional evaluation on a display device in the store or on a staff member's device to provide personalized service. This allows customers to understand their own health and emotional state in real time, and enables store staff to provide appropriate and personalized service to customers.

[1238] "Means for acquiring facial images" refers to a device or equipment for capturing a user's facial image in real time.

[1239] "Means for analyzing facial images to measure biometric information" refers to algorithms and software for analyzing captured facial images and measuring biometric information such as heart rate, blood pressure, and body temperature.

[1240] "Means for audio interaction with the user" refers to a microphone and speaker, and associated interaction software, for asking questions of the user and receiving responses.

[1241] "Means for analyzing voice data to assess health status" refers to software or algorithms for analyzing a user's voice data and assessing health status from the tone and rate of the voice data.

[1242] "Means for integrating acquired data to conduct a comprehensive health assessment" refers to a server or analysis system for integrating multiple data, such as facial video data and audio data, to conduct a comprehensive health assessment.

[1243] "Means for generating a health report" refers to software or algorithms that automatically generate a report summarizing health status based on the results of a comprehensive health assessment.

[1244] "Means for displaying the generated health report" refers to a display or display device, and associated software, for displaying the generated health report so that the user can review it.

[1245] "Means for displaying health and emotional assessments on in-store display devices or staff devices to provide personalized services" refers to systems and software that grasp customers' health and emotional status in real time, display this information on in-store displays or staff smart devices, and provide personalized services.

[1246] The present invention relates to a system that assesses a customer's health and emotional state in real time and provides personalized services in physical stores. When a customer enters a store, the system captures and analyzes facial images and evaluates their heart rate, blood pressure, body temperature, emotional state, etc. to provide an optimal customer experience.

[1247] Key elements of the system

[1248] The system includes the following main elements:

[1249] 1. Camera that captures facial images:

[1250] This camera uses a high-resolution camera (e.g., a high-resolution USB camera) that is placed on a mirror or other suitable location within the store.

[1251] 2. Video analysis and biometrics processor:

[1252] The video is analyzed using a facial recognition algorithm using the OpenCV library, and biometric information is measured using algorithms such as photoplethysmography.

[1253] 3. Microphone and speaker for voice interaction with the user:

[1254] For voice interaction, a high-precision microphone (e.g., a high-performance USB microphone) and a speaker are used, and voice analysis is performed via the Google Speech-to-Text API.

[1255] 4. Software that analyzes voice data to assess health and emotions:

[1256] To analyze voice data, the Google Speech-to-Text API and AI models such as TensorFlow are used.

[1257] 5. Server that integrates data to provide a comprehensive health assessment:

[1258] The server uses a cloud-based server (e.g., a public cloud service provider) to consolidate the collected data and perform real-time evaluation.

[1259] 6. Display to generate and show health reports:

[1260] Health reports are generated using a Python Flask-based web application and are displayed on displays in-store and on staff's smart devices.

[1261] 7. System for displaying health and emotional assessments and providing personalized services:

[1262] Use a system (e.g., personalized assistance application) that allows store staff to provide personalized service based on customer ratings.

[1263] Specific Examples

[1264] When a customer enters a store, a high-resolution camera inside the store captures an image of the customer's face, and at the same time, a voice conversation begins using a microphone and speaker.

[1265] The resulting facial and audio data is sent to video and audio analysis software, where the facial data is used to measure vital signs such as heart rate, blood pressure, and body temperature, while the audio data is analyzed to assess emotional state.

[1266] The server integrates this data to perform a comprehensive evaluation and generate health and emotion reports, which are displayed in real time on displays in the store and on staff's smart devices.

[1267] Staff can use this information to provide customers with personalized service, for example, suggesting products that will help them relax if they are feeling tired.

[1268] Prompt Sentence Examples

[1269] text

[1270] Develop a system that will capture the customer's facial image and voice in real time using a mirror in the store when they enter the store, evaluate their health and emotional state based on their heart rate, blood pressure, body temperature, tone and speed of voice, and integrate the collected data on a server to provide personalized services. The specifications will use OpenCV for facial recognition, Google Speech-to-Text API for voice analysis, Affectiva SDK for emotion recognition, and Python Flask for data integration.

[1271] In this way, the present invention allows for real-time assessment of a customer's health and emotional state to provide optimal service, thereby improving the customer experience and the service quality of the store.

[1272] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1273] Step 1:

[1274] When a user enters a store, the device's camera captures the user's facial image. This camera has high resolution and captures images in real time. The input is the user's facial image, and the output is the captured facial image data.

[1275] Step 2:

[1276] The facial image data acquired by the device is sent to a facial recognition algorithm. This algorithm uses the OpenCV library to identify the user's face. The input is the facial image data, and the output is facial recognition data in which the user's face is identified.

[1277] Step 3:

[1278] Based on the facial recognition data, the device measures the user's biometric information. Specifically, it uses video analysis software to measure heart rate, blood pressure, body temperature, etc. Photoplethysmography technology is used for this purpose. The input is facial recognition data, and the output is biometric data.

[1279] Step 4:

[1280] The terminal starts a voice dialogue with the user, asks a question through a microphone, and obtains the user's response. For example, the terminal asks, "What kind of product are you looking for today?" The input is the user's voice response, and the output is the obtained voice data.

[1281] Step 5:

[1282] The device sends the captured voice data to the Google Speech-to-Text API, which converts the voice into text data. Furthermore, speech analysis software is used to evaluate the user's emotional state based on the tone and speed of the voice. The input is voice data, and the output is textual voice data and emotional state data.

[1283] Step 6:

[1284] The device sends facial video data and audio data to a server, which then integrates the data to perform comprehensive health and emotional assessments. The inputs are biometric data, audio text data, and emotional state data, and the output is an integrated health and emotional assessment report.

[1285] Step 7:

[1286] The server returns the generated health and emotion evaluation reports to the terminal, which then displays these reports on displays in the store or on staff members' smart devices. The input is the health and emotion evaluation reports, and the output is the displayed reports.

[1287] Step 8:

[1288] Store staff provide personalized services based on the user's state. For example, if the user is stressed, they will suggest products that will help them relax. The staff responds to the customer while referring to the display on the terminal. The input is the displayed report, and the output is the service provided.

[1289] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1290] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1291] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1292] [Fourth embodiment]

[1293] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1294] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1295] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1296] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1297] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1298] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1299] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1300] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1301] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1302] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1303] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1304] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1305] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1306] This invention relates to an AI-equipped mirror system that analyzes facial images and audio to assess the user's health condition. This system allows the user to receive a multifaceted health assessment in real time by standing in front of the mirror.

[1307] System Overview

[1308] The system includes the following main elements:

[1309] 1. Camera that captures facial images

[1310] 2. Processor for video analysis and biometric measurement

[1311] 3. A microphone and speaker for voice interaction with the user

[1312] 4. Software that analyzes voice data to assess health status

[1313] 5. Server that integrates data and provides comprehensive health assessment

[1314] 6. Display to generate and display health reports

[1315] Program processing flow

[1316] The program of this system is designed as follows:

[1317] 1. The user stands in front of the AI-enabled mirror

[1318] Every morning, the user stands in front of the mirror and looks at their face, which automatically activates the system.

[1319] 2. The device recognizes the user's face

[1320] The device uses a built-in camera to capture an image of your face and uses a facial recognition algorithm to identify you, which includes matching it against a historical database.

[1321] 3. The device analyzes your face and muscle movements

[1322] The device uses a high-resolution camera to capture detailed facial video data, allowing it to analyze and measure biometric information such as heart rate, blood pressure, and body temperature.

[1323] 4. The device starts a dialogue with the user

[1324] The device initiates a voice dialogue, asking a question such as, "How much sleep did you get last night?" When the user responds, the voice data is acquired.

[1325] 5. The device analyzes the voice data

[1326] The device analyzes the user's voice and assesses stress and fatigue levels based on the tone and speed of the voice.

[1327] 6. The server integrates and analyzes the data

[1328] The server integrates the facial video data, audio data, and question-and-answer data sent from the device to perform a comprehensive health assessment.

[1329] 7. The server generates a health report

[1330] The server generates a health report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[1331] 8. The device will display a health report

[1332] The device will display a health report on the mirror and issue warnings as needed, such as telling the user, "You're feeling stressed today. We recommend taking some time to relax."

[1333] Specific examples

[1334] 1. The user stands in front of the mirror

[1335] Every morning, when the user stands in front of the mirror, facial images are automatically captured.

[1336] 2. Facial Recognition

[1337] The device captures the user's face through the camera and identifies them as "Yamada-san" by comparing it with a previously registered database.

[1338] 3. Video Analysis

[1339] The camera captures detailed images of the user's face, and the processor measures heart rate, blood pressure, body temperature, etc. At the same time, subtle changes in facial color and facial expressions are analyzed.

[1340] 4. Starting a voice conversation

[1341] The device uses a microphone and speaker to ask, "How much sleep did you get last night?" The user responds, "I slept for 6 hours."

[1342] 5. Analysis of audio data

[1343] The device analyzes the user's voice data and detects stress levels from the tone and speed of the voice.

[1344] 6. Data integration and analysis

[1345] The collected facial video data, audio data, and lifestyle data are sent to a server for a comprehensive health assessment.

[1346] 7. Generate Health Reports

[1347] Based on the evaluation results, the server generates a health report for the user, including a warning such as, "Your heart rate is high today. You may need to rest."

[1348] 8. View Health Report

[1349] The device displays a health report on the mirror and notifies the user with an automated voice.

[1350] In this way, users can easily monitor their daily health status and take necessary measures. The present invention significantly improves daily health management without relying solely on regular medical checkups.

[1351] The processing flow will be explained below.

[1352] Step 1:

[1353] The user stands in front of the mirror, which automatically activates the system and prepares the camera to capture the user's face.

[1354] Step 2:

[1355] The device uses its built-in camera to capture the user's face, then uses a facial recognition algorithm to identify the user and match them with a historical database.

[1356] Step 3:

[1357] The device uses a high-resolution camera to capture detailed facial video data, including technology to capture facial biometric information such as heart rate, blood pressure, and temperature.

[1358] Step 4:

[1359] The device starts a voice dialogue with the user. For example, it asks, "Good morning. How long did you sleep today?" The user replies, "I slept for 6 hours."

[1360] Step 5:

[1361] The device analyzes the voice data it acquires, and uses voice analysis technology to assess stress and fatigue levels based on the tone and speed of the user's voice.

[1362] Step 6:

[1363] The device analyzes facial expressions and changes in color tone to measure heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, hemoglobin concentration, stress, and fatigue level.

[1364] Step 7:

[1365] The facial image data, audio data, and question and answer data acquired by the terminal are transmitted to the server.

[1366] Step 8:

[1367] The server then combines the received data to provide a comprehensive health assessment, including heart rate and blood pressure variability, voice analysis results, and lifestyle information.

[1368] Step 9:

[1369] The server generates a health report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[1370] Step 10:

[1371] The device receives the health report sent from the server and displays it on the mirror, for example, displaying a message such as "Your heart rate is high today. You may need to rest."

[1372] Step 11:

[1373] The device will notify the user with an automated voice, providing feedback such as, "Good morning. You're feeling stressed today. We recommend you relax."

[1374] Example 1

[1375] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1376] Conventional health management systems require users to manually input information, which can lead to problems with accuracy and effort. In addition, the long intervals between regular health checkups make it difficult to quickly detect changes in health status, making it impossible to grasp health status in real time.

[1377] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1378] In this invention, the server includes means for acquiring a facial image, means for analyzing the acquired facial image to measure biometric information, means for engaging in voice dialogue with the user, means for analyzing the voice data to evaluate the health condition, means for integrating the acquired video data and voice data to perform a comprehensive health evaluation, means for generating a health report, and means for displaying the generated health report, thereby enabling the user to receive a comprehensive health evaluation in real time simply by standing in front of the mirror.

[1379] "Means for capturing facial images" refers to a camera or other imaging device that captures an image or video of a user's face.

[1380] "Means for analyzing captured facial images and measuring biometric information" refers to processors and algorithms that analyze and measure biometric information such as heart rate, blood pressure, and body temperature from captured facial image data.

[1381] "Means for conducting voice dialogue with the user" refers to a microphone and speaker for communicating with the user by voice, and software for controlling the dialogue.

[1382] "Means for analyzing voice data to assess health status" refers to algorithms or analytical software for analyzing voice data and assessing a user's stress level, fatigue level, etc.

[1383] "Means for integrating acquired video and audio data to perform a comprehensive health assessment" refers to a server and analysis program for integrating facial video and audio data to assess overall health status.

[1384] "Means for generating a health report" refers to software or algorithms that automatically generate a report showing the user's health status based on the analysis results.

[1385] "Means for displaying the generated health report" refers to a display or screen for visually presenting the generated health report to the user.

[1386] This invention relates to an AI-equipped mirror system that analyzes the user's facial image and voice to evaluate their health condition in real time. To implement this system, the following hardware and software are used:

[1387] Hardware and software used

[1388] 1. Camera: Used to capture video of the user's face. For example, a 1080p HD camera or a 4K camera is suitable.

[1389] 2. Processor: Used for video analysis and biometric measurement. Examples include an Intel Core i7 processor or an NVIDIA GPU.

[1390] 3. Microphone and speaker: Used for voice interaction with the user, a high-sensitivity microphone and stereo speakers are desirable to provide clear sound quality.

[1391] 4. Analysis software: Software used to analyze video and audio data often uses OpenCV, TensorFlow, or Google Cloud Speech-to-Text API.

[1392] 5. Server: A backend server that integrates data and performs comprehensive health assessments. This is where AI models using TensorFlow and PyTorch run.

[1393] 6. Display: Used to display the generated health report. A display that also functions as a mirror is desirable.

[1394] How the system works

[1395] This system automatically starts when the user stands in front of the mirror and performs the following processes.

[1396] Image acquisition by camera

[1397] When a user stands in front of the mirror, the camera automatically captures an image of the user's face.

[1398] Facial image analysis

[1399] The processor analyzes the captured facial video data. For example, it detects subtle color changes and facial expressions, and uses this data to estimate heart rate, blood pressure, and body temperature. This analysis uses the OpenCV library.

[1400] Voice interaction with the user

[1401] The device uses a microphone and speaker to ask the user questions, such as prompting them with questions like, "How much sleep did you get last night?" If the user responds, "I slept for six hours," the device collects the audio data.

[1402] Analysis of audio data

[1403] The device converts voice data into text using the Google Cloud Speech-to-Text API, and then analyzes the content, tone, and speed of the voice, allowing it to assess stress and fatigue levels.

[1404] Data integration and analysis

[1405] The server receives the facial video data, audio data, and question-and-answer data sent from the device and performs an integrated analysis. A comprehensive health assessment is performed using an AI model built with TensorFlow and PyTorch.

[1406] Generate health reports

[1407] The server generates a detailed health report based on the analysis results, including a warning such as "Your heart rate is high today. You may need to rest" if the analysis results show a higher than normal heart rate.

[1408] View Health Report

[1409] Finally, the device displays the generated health report on the screen and, if necessary, notifies the user by voice.

[1410] Prompt Sentence Examples

[1411] Examples of prompts the system might give to the user include:

[1412] "Good morning. How much sleep did you get last night?"

[1413] "How are you feeling today?"

[1414] "Are you stressed?"

[1415] These specific actions allow users to easily monitor their daily health status and quickly take necessary measures.

[1416] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1417] Step 1:

[1418] When a user stands in front of the mirror, the device detects movement and automatically activates the camera and microphone. The input is the user's movement, and the output is the activation of the camera and microphone. Specifically, the sensor detects the user's presence, and the camera begins capturing video.

[1419] Step 2:

[1420] The device recognizes the user's face. The camera captures an image of the user's face, and this video data is input into a facial recognition algorithm. The algorithm compares it with a previously saved database to identify the user, and the resulting username is output. A specific example of how this works is to identify the user as "Yamada-san" using facial recognition technology.

[1421] Step 3:

[1422] The device analyzes images of the user's face and measures their biometric information. The input is high-resolution video data acquired from a camera, which the processor analyzes to estimate heart rate, blood pressure, and body temperature. The output is this biometric information. Specifically, it uses computer vision technology to detect subtle color changes and muscle movements on the face to measure heart rate and blood pressure.

[1423] Step 4:

[1424] The device starts a voice dialogue with the user. It uses a microphone and speaker to ask a question and obtains the answer. The input is a prompt such as "How much sleep did you get last night?" and the user's voice response, and the output is voice data. Specifically, the system makes the user respond with "I slept for 6 hours."

[1425] Step 5:

[1426] The device analyzes the voice data. The input is the captured voice data, which the software converts into text and then analyzes the tone and speed. The output is an assessment of the user's stress level and fatigue. Specifically, it uses voice analysis technology to determine that a high-pitched voice and fast speed indicate high stress.

[1427] Step 6:

[1428] The server integrates and analyzes the data. Facial video data, audio data, and question-and-answer data sent from the device are input, and the server integrates these to perform a comprehensive health assessment. The output is a detailed health assessment result. Specifically, the system uses an AI model to analyze multidimensional data and evaluate the overall health condition.

[1429] Step 7:

[1430] The server generates a health report. The input is the integrated and analyzed data, and the software uses this to create a health report. The output is the user's health report. Specific operations include generating a report with a warning, such as "Your heart rate is high today. You may need to rest."

[1431] Step 8:

[1432] The device displays the generated health report. The input is the health report, which is visually displayed on the display. The output is visual feedback to the user. Specifically, the device displays "Your heart rate is high today. We recommend you rest" and notifies you by voice if necessary.

[1433] Through the above processing steps, users can easily understand their daily health condition in real time and take necessary measures promptly.

[1434] (Application example 1)

[1435] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1436] Conventional health management systems require users to voluntarily collect and input data, limiting their practicality and continuity. Furthermore, it is difficult to monitor detailed health conditions outside of regular health checkups. The present invention aims to provide a system that more efficiently and instantly assesses a user's health condition. In particular, the system can be installed in a brick-and-mortar store, solving the problem of enabling users to naturally manage their health as part of their daily activities.

[1437] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1438] In this invention, the server includes means for acquiring a facial image, means for analyzing the facial image to measure biometric information, means for engaging in voice dialogue with the user, means for analyzing the voice data to evaluate the health condition, means for integrating the acquired data to perform a comprehensive health evaluation, means for generating a health report, means for displaying the generated health report, and means for instantly evaluating the user's health using a mirror installed in the physical store. This allows the user's health condition to be automatically evaluated every time they visit the store, providing instant feedback and facilitating daily health management.

[1439] The "means for acquiring a facial image" refers to hardware such as a camera or image sensor for capturing an image of the user's face, and software for controlling the hardware.

[1440] The "means for analyzing facial images to measure biometric information" refers to software and algorithms for analyzing captured facial images and calculating biometric information such as heart rate, blood pressure, and body temperature.

[1441] The "means for conducting a voice dialogue with the user" refers to hardware and software for controlling the hardware and the speaker for conducting a voice dialogue with the user using a microphone and speaker.

[1442] The "means for analyzing voice data to assess health status" refers to software and algorithms for analyzing acquired voice data and assessing the user's stress level and fatigue level.

[1443] The "means for integrating acquired data to perform a comprehensive health assessment" refers to software and a server system for integrating facial video data and audio data to assess the user's overall health condition.

[1444] The "means for generating a health report" is software for creating a report that reports the user's health status based on the integrated health data.

[1445] The "means for displaying the generated health report" refers to a display and its control software for visually displaying the generated health report to the user.

[1446] The "means for instantly assessing a user's health using a mirror installed in a physical store" is a system for instantly assessing the health of users who visit the store and displaying the results through a mirror-like device installed in the store.

[1447] This invention relates to an "AI Health Check Mirror System" installed in a brick-and-mortar store, where users can stand in front of the mirror and have their facial images and voices analyzed, and receive instant feedback on their health condition for the day. The specific implementation of this system is described below.

[1448] System configuration

[1449] The system includes the following main elements:

[1450] 1. Camera: A means of capturing facial images, such as a high-resolution camera like the Logitech C920.

[1451] 2. Processor: A hardware device that analyzes facial images and measures biometric information. Various algorithms are embedded in the processor.

[1452] 3. Microphone and speaker: A means of audio interaction with the user. For example, a high-performance microphone such as the Blue Yeti.

[1453] 4. Voice analysis software: A means of analyzing voice data to assess health status. Custom algorithms for voice analysis.

[1454] 5. Server: A means of integrating acquired data to provide a comprehensive health assessment. Often, a cloud-based server is used.

[1455] 6. Display: Means for generating and displaying health reports. Display device that acts as a mirror.

[1456] 7. Physical store mirror: Installed in a physical store, it provides an instant health assessment of the user.

[1457] Data collection and analysis

[1458] 1. Acquiring facial images

[1459] When a user stands in front of the AI ​​health check mirror installed in a physical store, the camera automatically captures the user's facial image, which automatically activates the system.

[1460] 2. Measurement of biological information

[1461] The captured facial image is analyzed by a processor to measure biometric information such as heart rate, blood pressure, and body temperature. Subtle changes in facial color and facial expressions are also analyzed.

[1462] 3. Starting a voice conversation

[1463] When a user speaks to the system, a microphone picks up the voice and a voice interaction takes place through the speaker. The system asks questions, such as, "How much sleep did you get last night?"

[1464] 4. Analysis of audio data

[1465] The captured voice data is analyzed using voice analysis software, and stress and fatigue levels are assessed based on the tone and speed of the voice.

[1466] 5. Data integration and comprehensive health assessment

[1467] The server integrates facial video data, audio data, and voice dialogue data to provide a comprehensive health assessment, which includes various biometric data and voice analysis results.

[1468] 6. Generate and view health reports

[1469] The server generates a health report for the user based on the analysis results. If any abnormalities are detected, a report including a warning message is generated. The generated health report is displayed on the mirror and a warning is issued as needed. For example, feedback is given in the form of "Your heart rate is high today. You may need to rest."

[1470] Specific examples

[1471] When you stand in front of an AI health check mirror installed in a physical store, the system will automatically activate and capture an image of your face.

[1472] The images captured by the camera are analyzed by a processor to measure heart rate and blood pressure.

[1473] When the user looks into the mirror and says, "I slept six hours last night," voice analysis is performed to assess stress and fatigue levels.

[1474] The server integrates this data, performs a comprehensive health assessment, and then displays advice such as, "Your heart rate is high today. You may need to rest."

[1475] Example prompts for generative AI models

[1476] "Please explain in detail about the system that analyzes facial images and audio to assess the user's health condition. The system estimates biometric information such as heart rate, blood pressure, and body temperature by having the user stand in front of a mirror, and evaluates stress levels from the tone and rate of voice. Based on the assessment results, the system generates a health report and provides advice."

[1477] This allows users to easily monitor their daily health status and improve their ability to take necessary measures. It is expected that the present invention will significantly improve health management not only through conventional periodic health checkups but also through everyday activities.

[1478] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1479] Step 1:

[1480] When a user stands in front of the mirror, the device activates the built-in camera to capture an image of the user's face.

[1481] Input: Face image taken by camera

[1482] Output: Acquired facial image data

[1483] What it does: The system automatically activates the camera and captures high-resolution facial footage, which is then stored for further processing.

[1484] Step 2:

[1485] The facial image captured by the device is sent to a processor, which analyzes the facial image and measures biometric information.

[1486] Input: Acquired facial video data

[1487] Output: vital signs such as heart rate, blood pressure, and temperature

[1488] How it works: Video analysis algorithms analyze subtle changes in facial color and facial expressions to measure biometric information. A processor extracts and stores this data.

[1489] Step 3:

[1490] The device initiates a voice dialogue with the user. For example, the device asks, "How much sleep did you get last night?"

[1491] Input: User voice input

[1492] Output: Acquired audio data

[1493] What it does: A microphone picks up the user's voice and records the audio in an interactive format. A speaker is used to ask questions and record the user's responses.

[1494] Step 4:

[1495] The device analyzes the voice data it acquires and assesses the user's health condition.

[1496] Input: User's voice data

[1497] Output: User's voice analysis results (stress level, fatigue level, etc.)

[1498] What it does: Voice analysis software analyzes the tone and rate of the user's voice to assess stress and fatigue levels. These assessments are used in the next step.

[1499] Step 5:

[1500] The device sends facial video and audio data to a server, which then integrates the data to provide a comprehensive health assessment.

[1501] Input: Facial video data, audio data, audio analysis results

[1502] Output: Comprehensive health assessment data

[1503] How it works: The server combines facial video and audio data, and evaluates the user's overall health based on biometric and audio analysis results. An algorithm performs this evaluation and generates a result.

[1504] Step 6:

[1505] The server generates a health report. The server creates a report based on the evaluation results and adds warning messages if necessary.

[1506] Input: Comprehensive health assessment data

[1507] Output: Health report

[1508] What happens: The server generates a personalized health report based on the evaluation results, including a warning message if any abnormalities are detected. The report is displayed in the next step.

[1509] Step 7:

[1510] The health report generated by the device is displayed on the mirror and notified to the user.

[1511] Input: Health Report

[1512] Output: Health report displayed on mirror, audio notification

[1513] What it does: The device displays the generated health report on the mirror and notifies the user through the speaker, providing feedback such as, "Your heart rate is high today. You may need to rest."

[1514] This series of processes allows the user to quickly know their health condition and take measures as needed.

[1515] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1516] This invention relates to an AI-equipped mirror system that analyzes facial images and voices to evaluate the user's health and emotional state. This system allows the user to receive a multifaceted health and emotional state evaluation in real time by standing in front of the mirror.

[1517] System Overview

[1518] The system includes the following main elements:

[1519] 1. Camera that captures facial images

[1520] 2. Processor for video analysis and biometric measurement

[1521] 3. A microphone and speaker for voice interaction with the user

[1522] 4. Software that analyzes voice data to assess health status and emotions

[1523] 5. Server that integrates data and provides comprehensive health assessment

[1524] 6. Display to generate and display health reports

[1525] 7. Emotion engine that recognizes user emotions

[1526] Program processing flow

[1527] The program of this system is designed as follows:

[1528] 1. The user stands in front of the AI-enabled mirror

[1529] Every morning, the user stands in front of the mirror and looks at their face, which automatically activates the system.

[1530] 2. The device recognizes the user's face

[1531] The device uses a built-in camera to capture an image of your face and uses a facial recognition algorithm to identify you, which includes matching it against a historical database.

[1532] 3. The device analyzes your face and muscle movements

[1533] The device uses a high-resolution camera to capture detailed facial video data, allowing it to analyze and measure biometric information such as heart rate, blood pressure, and body temperature.

[1534] 4. The device starts a dialogue with the user

[1535] The terminal starts a voice dialogue, asking, for example, "Good morning. How long did you sleep today?" When the user responds, the voice data is acquired.

[1536] 5. The device analyzes the voice data

[1537] The device analyzes the user's voice and assesses stress and fatigue levels based on the tone and speed of the voice.

[1538] 6. The device analyzes facial images and recognizes emotions

[1539] The emotion engine analyzes the user's facial expressions and subtle muscle movements to recognize emotions such as joy, anger, and surprise.

[1540] 7. The server integrates and analyzes the data

[1541] The server integrates the facial video data, audio data, question and answer data, and emotional data sent from the device to perform a comprehensive health and emotional assessment.

[1542] 8. The server generates health and emotion reports

[1543] The server generates a health and emotional report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[1544] 9. The device displays the report

[1545] The device will display health and emotional reports on the mirror and issue warnings as needed, such as telling the user, "You're feeling stressed today. We recommend taking some time to relax."

[1546] Specific examples

[1547] 1. The user stands in front of the mirror

[1548] Every morning, when the user stands in front of the mirror, facial images are automatically captured.

[1549] 2. Facial Recognition

[1550] The device captures the user's face through the camera and identifies them as "Yamada-san" by comparing it with a previously registered database.

[1551] 3. Video Analysis

[1552] The camera captures detailed images of the user's face, and the processor measures heart rate, blood pressure, body temperature, etc. At the same time, subtle changes in facial color and facial expressions are analyzed.

[1553] 4. Starting a voice conversation

[1554] The device uses a microphone and speaker to ask, "How much sleep did you get last night?" The user responds, "I slept for 6 hours."

[1555] 5. Analysis of audio data

[1556] The device analyzes the user's voice data and detects stress levels from the tone and speed of the voice.

[1557] 6. Emotional Recognition

[1558] The emotion engine analyzes the user's facial expressions and recognizes their current emotions, such as whether they are happy, angry, or surprised.

[1559] 7. Data integration and analysis

[1560] The collected facial video data, audio data, lifestyle data, and emotional data are sent to a server, and a comprehensive health and emotional assessment is performed.

[1561] 8. Generate health and emotional reports

[1562] Based on the evaluation results, the server generates a health and mood report for the user, including a warning such as, "Your heart rate is high today. You may need to rest."

[1563] 9. View the report

[1564] The device displays health and emotion reports on the mirror and provides automated voice feedback to the user, such as "Good morning. You're feeling stressed today. We recommend you relax."

[1565] In this way, users can easily monitor their daily health and emotional state and take necessary measures. The present invention significantly improves daily health management and emotional management without relying solely on regular medical checkups.

[1566] The processing flow will be explained below.

[1567] Step 1:

[1568] The user stands in front of the mirror, which automatically activates the system and prepares the camera to capture the user's face.

[1569] Step 2:

[1570] The device uses its built-in camera to capture the user's face, then uses a facial recognition algorithm to identify the user and match them against a database to identify the user.

[1571] Step 3:

[1572] The device uses a high-resolution camera to capture detailed video data of the user's face, including technology to measure biometric information such as heart rate, blood pressure, and temperature from the face.

[1573] Step 4:

[1574] The terminal initiates a voice dialogue with the user, for example asking, "Good morning. How are you feeling today?" The user responds, "I'm feeling great."

[1575] Step 5:

[1576] The device analyzes the voice data it acquires, and uses voice analysis technology to assess stress and fatigue levels based on the tone and speed of the user's voice.

[1577] Step 6:

[1578] The device analyzes facial expressions and changes in color tone to measure heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, hemoglobin concentration, stress, and fatigue level.

[1579] Step 7:

[1580] The emotion engine analyzes the user's facial expressions and subtle muscle movements to recognize the user's emotional state. For example, if the user is smiling, it will recognize that they are feeling "happy."

[1581] Step 8:

[1582] The facial image data, audio data, question and answer data, and emotion data acquired by the terminal are transmitted to the server.

[1583] Step 9:

[1584] The server then integrates the received data to provide a comprehensive health and emotional assessment, including heart rate and blood pressure variability, voice analysis results, lifestyle information, and emotional state.

[1585] Step 10:

[1586] The server generates a health and emotion report for the user based on the analysis results. If an abnormality is detected, a report including a warning message is created.

[1587] Step 11:

[1588] The device receives the health and emotion reports sent from the server and displays them on the mirror, for example, displaying a message such as "Your heart rate is high today. You may need to rest."

[1589] Step 12:

[1590] The device will notify the user with an automated voice, providing feedback such as, "Good morning. You seem to be feeling a little stressed today. I recommend you relax."

[1591] Example 2

[1592] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1593] Conventional health management systems often require specialized equipment at specific medical facilities to accurately assess a user's health status. This makes them inconvenient for everyday use and places a burden on users. Furthermore, they are unable to simultaneously assess the user's emotional state, making comprehensive health management difficult.

[1594] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring a facial image, means for analyzing the facial image to measure biometric information, means for conducting a voice dialogue with the user, means for analyzing the voice data to evaluate the health state and emotions, means for integrating the acquired data to perform a comprehensive health evaluation and emotion evaluation, means for generating a health and emotion report, and means for displaying the generated health and emotion report. This enables the user to easily and accurately monitor their own health and emotional state every day.

[1595] The "means for acquiring a facial image" is a device such as a camera for capturing an image of the user's face.

[1596] The "means for analyzing facial images and measuring biometric information" refers to software and hardware for analyzing biometric indicators such as heart rate, blood pressure, and body temperature based on the captured facial images.

[1597] The "means for performing voice interaction with a user" is a device for performing voice communication with a user using a microphone and a speaker.

[1598] The "means for analyzing voice data to assess health status and emotions" is software for analyzing the user's voice data and inferring health status and emotional state from the tone and speed of the voice.

[1599] "Means for integrating acquired data to provide a comprehensive health and emotional assessment" refers to a processor or algorithm that integrates collected biometric information, audio data, and other relevant data to provide a comprehensive assessment of the user's health and emotional state.

[1600] The "means for generating health and emotional reports" is software for automatically generating health and emotional reports based on the integrated data.

[1601] The "means for displaying the generated health and emotion report" is a display or monitor for visually presenting the generated report to the user.

[1602] The present invention relates to an AI-equipped mirror system that evaluates a user's health and emotional state using facial video and audio data. This system allows a user to receive multifaceted health and emotional evaluations in real time by standing in front of the mirror. Specific means and methods for implementing the present invention are described below.

[1603] Hardware used

[1604] Camera: Used to capture high-resolution footage, specifically capturing detailed footage of the face.

[1605] Processor: A processing device for analyzing video and audio data. For example, analyzing biometric information (heart rate, blood pressure, body temperature, etc.) from facial images.

[1606] Microphone and speaker: Used for voice interaction with the user, collecting information about the user's lifestyle habits and current emotional state through the interaction.

[1607] Display: A display device for presenting the generated health and emotion reports to the user.

[1608] Software used

[1609] Facial recognition algorithm: Analyzes facial images captured by the camera to recognize and identify the user.

[1610] Biometric analysis software: Biometric information is measured based on captured video data. This analysis uses changes in facial color and subtle muscle movements.

[1611] Voice analysis software: Analyzes voice data collected through a microphone to assess stress and fatigue levels.

[1612] Emotion engine: Analyzes facial expressions and subtle muscle movements to recognize the user's emotional state (happiness, anger, surprise, etc.).

[1613] Data integration algorithm: Integrates facial video data, audio data, and emotion data to perform comprehensive health and emotion assessments.

[1614] Report generation software: Generates health and emotional reports based on the integrated data.

[1615] System Operation

[1616] When a user stands in front of the mirror, the system captures facial images through a camera and identifies the user using a facial recognition algorithm. The processor analyzes the image data and measures biometric information (heart rate, blood pressure, body temperature, etc.). Next, a voice dialogue with the user is conducted through a microphone and speaker to collect information about the user's lifestyle habits and emotional state. The collected voice data is analyzed to evaluate the user's stress and fatigue levels. An emotion engine analyzes facial expressions and subtle muscle movements to recognize the user's emotional state. The server integrates this data to provide a comprehensive health and emotion assessment. Finally, report generation software creates health and emotion reports and displays them on the display. If any abnormalities are detected, a report is created with appropriate warning messages.

[1617] Specific examples

[1618] 1. The user stands in front of the mirror

[1619] Every morning, when the user stands in front of the mirror, facial images are automatically captured.

[1620] 2. Facial Recognition

[1621] The device captures the user's face through the camera and identifies them as "Yamada-san" by comparing it with a previously registered database.

[1622] 3. Video Analysis

[1623] The camera captures detailed images of the user's face, and the processor measures heart rate, blood pressure, body temperature, etc. At the same time, subtle changes in facial color and facial expressions are analyzed.

[1624] 4. Starting a voice conversation

[1625] The device uses a microphone and speaker to ask, "How much sleep did you get last night?" The user responds, "I slept for 6 hours."

[1626] 5. Analysis of audio data

[1627] The device analyzes the user's voice data and detects stress levels from the tone and speed of the voice.

[1628] 6. Emotional Recognition

[1629] The emotion engine analyzes the user's facial expressions and recognizes their current emotions, such as whether they are happy, angry, or surprised.

[1630] 7. Data integration and analysis

[1631] The collected facial video data, audio data, lifestyle data, and emotional data are sent to a server, and a comprehensive health and emotional assessment is performed.

[1632] 8. Generate health and emotional reports

[1633] Based on the evaluation results, the server generates a health and mood report for the user, including a warning such as, "Your heart rate is high today. You may need to rest."

[1634] 9. View the report

[1635] The device displays health and emotion reports on the mirror and provides automated voice feedback to the user, such as "Good morning. You're feeling stressed today. We recommend you relax."

[1636] Prompt Sentence Examples

[1637] Enter the following prompt into the generative AI model:

[1638] "This system uses an AI-equipped mirror to assess the user's health and emotional state. When the user stands in front of the mirror, a camera captures an image of their face and a voice dialogue begins. The acquired data is sent to a server for integrated analysis. Can you give us a concrete example?"

[1639] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1640] Step 1:

[1641] The system automatically activates when a user stands in front of the mirror. The input is the user's presence, which activates the camera and the device captures an image of the user's face. This action allows the system to capture an image of the user's face when the user says "Good morning."

[1642] Step 2:

[1643] The device uses a built-in camera to recognize the user's face. The input is a captured image of the face, and the output is the user's identification. Specifically, the facial recognition algorithm analyzes the image data and compares it with a database to identify the user as "Yamada-san."

[1644] Step 3:

[1645] The device uses a high-resolution camera to capture detailed facial images and analyze biometric information. The input is the recognized user's video data, and the output is biometric information such as heart rate, blood pressure, and body temperature. The processor analyzes subtle changes in facial color and muscle movements to make a diagnosis, such as whether the user has a pale complexion.

[1646] Step 4:

[1647] The device initiates a voice dialogue with the user using a microphone and speaker. The input is the user's voice, and the output is voice data. For example, the system asks, "Good morning. How much sleep did you get last night?" and the user replies, "I slept for six hours."

[1648] Step 5:

[1649] The device analyzes the acquired voice data and evaluates stress and fatigue levels based on the tone and speed of the voice. The input is the acquired voice data, and the output is the evaluation result of the user's stress and fatigue. For example, if the voice is quiet and low-pitched, the device will diagnose, "You seem a little tired."

[1650] Step 6:

[1651] The device uses an emotion engine to analyze facial expressions and subtle muscle movements to recognize emotions. The input is video data of the user's face, and the output is the recognition result of the user's emotional state (happiness, anger, surprise, etc.). For example, if the user's brow is furrowed, the device will evaluate the user as "slightly anxious."

[1652] Step 7:

[1653] The server integrates and analyzes the facial video data, audio data, question-and-answer data, and emotional data sent from the device. The input is multiple data sets, and the output is a comprehensive health and emotional assessment result. Based on this, the server determines that "the user is experiencing accumulated fatigue."

[1654] Step 8:

[1655] The server generates a health and emotion report based on the assessment results. The input is the integrated and analyzed assessment results, and the output is a health and emotion report. For example, it may include a warning message such as "Your heart rate is high today. Please take a rest."

[1656] Step 9:

[1657] The terminal displays the generated report on the mirror and issues a warning if necessary. The input is the generated report, and the output is the displayed report and audio feedback. Specifically, it tells the user, "Good morning. You are feeling stressed today. We recommend that you relax."

[1658] (Application example 2)

[1659] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1660] In modern society, especially in brick-and-mortar stores, there is a need to understand customers' health and emotional states in real time and provide services accordingly. However, conventional methods require customers to self-report their health and emotional states or have specialized staff manually monitor them, which has issues with inefficiency and inaccuracy. The present invention aims to provide a system that solves these issues and significantly improves customer experience.

[1661] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a facial image, means for analyzing the facial image to measure biometric information, means for engaging in voice dialogue with the user, means for analyzing voice data to evaluate a health condition, means for integrating the acquired data to perform a comprehensive health evaluation, means for generating a health report, means for displaying the generated health report, and means for displaying the health evaluation and emotional evaluation on a display device in the store or on a staff member's device to provide personalized service. This allows customers to understand their own health and emotional state in real time, and enables store staff to provide appropriate and personalized service to customers.

[1662] "Means for acquiring facial images" refers to a device or equipment for capturing a user's facial image in real time.

[1663] "Means for analyzing facial images to measure biometric information" refers to algorithms and software for analyzing captured facial images and measuring biometric information such as heart rate, blood pressure, and body temperature.

[1664] "Means for audio interaction with the user" refers to a microphone and speaker, and associated interaction software, for asking questions of the user and receiving responses.

[1665] "Means for analyzing voice data to assess health status" refers to software or algorithms for analyzing a user's voice data and assessing health status from the tone and rate of the voice data.

[1666] "Means for integrating acquired data to conduct a comprehensive health assessment" refers to a server or analysis system for integrating multiple data, such as facial video data and audio data, to conduct a comprehensive health assessment.

[1667] "Means for generating a health report" refers to software or algorithms that automatically generate a report summarizing health status based on the results of a comprehensive health assessment.

[1668] "Means for displaying the generated health report" refers to a display or display device, and associated software, for displaying the generated health report so that the user can review it.

[1669] "Means for displaying health and emotional assessments on in-store display devices or staff devices to provide personalized services" refers to systems and software that grasp customers' health and emotional status in real time, display this information on in-store displays or staff smart devices, and provide personalized services.

[1670] The present invention relates to a system that assesses a customer's health and emotional state in real time and provides personalized services in physical stores. When a customer enters a store, the system captures and analyzes facial images and evaluates their heart rate, blood pressure, body temperature, emotional state, etc. to provide an optimal customer experience.

[1671] Key elements of the system

[1672] The system includes the following main elements:

[1673] 1. Camera that captures facial images:

[1674] This camera uses a high-resolution camera (e.g., a high-resolution USB camera) that is placed on a mirror or other suitable location within the store.

[1675] 2. Video analysis and biometrics processor:

[1676] The video is analyzed using a facial recognition algorithm using the OpenCV library, and biometric information is measured using algorithms such as photoplethysmography.

[1677] 3. Microphone and speaker for voice interaction with the user:

[1678] For voice interaction, a high-precision microphone (e.g., a high-performance USB microphone) and a speaker are used, and voice analysis is performed via the Google Speech-to-Text API.

[1679] 4. Software that analyzes voice data to assess health and emotions:

[1680] To analyze voice data, the Google Speech-to-Text API and AI models such as TensorFlow are used.

[1681] 5. Server that integrates data to provide a comprehensive health assessment:

[1682] The server uses a cloud-based server (e.g., a public cloud service provider) to consolidate the collected data and perform real-time evaluation.

[1683] 6. Display to generate and show health reports:

[1684] Health reports are generated using a Python Flask-based web application and are displayed on displays in-store and on staff's smart devices.

[1685] 7. System for displaying health and emotional assessments and providing personalized services:

[1686] Use a system (e.g., personalized assistance application) that allows store staff to provide personalized service based on customer ratings.

[1687] Specific Examples

[1688] When a customer enters a store, a high-resolution camera inside the store captures an image of the customer's face, and at the same time, a voice conversation begins using a microphone and speaker.

[1689] The resulting facial and audio data is sent to video and audio analysis software, where the facial data is used to measure vital signs such as heart rate, blood pressure, and body temperature, while the audio data is analyzed to assess emotional state.

[1690] The server integrates this data to perform a comprehensive evaluation and generate health and emotion reports, which are displayed in real time on displays in the store and on staff's smart devices.

[1691] Staff can use this information to provide customers with personalized service, for example, suggesting products that will help them relax if they are feeling tired.

[1692] Prompt Sentence Examples

[1693] text

[1694] Develop a system that will capture the customer's facial image and voice in real time using a mirror in the store when they enter the store, evaluate their health and emotional state based on their heart rate, blood pressure, body temperature, tone and speed of voice, and integrate the collected data on a server to provide personalized services. The specifications will use OpenCV for facial recognition, Google Speech-to-Text API for voice analysis, Affectiva SDK for emotion recognition, and Python Flask for data integration.

[1695] In this way, the present invention allows for real-time assessment of a customer's health and emotional state to provide optimal service, thereby improving the customer experience and the service quality of the store.

[1696] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1697] Step 1:

[1698] When a user enters a store, the device's camera captures the user's facial image. This camera has high resolution and captures images in real time. The input is the user's facial image, and the output is the captured facial image data.

[1699] Step 2:

[1700] The facial image data acquired by the device is sent to a facial recognition algorithm. This algorithm uses the OpenCV library to identify the user's face. The input is the facial image data, and the output is facial recognition data in which the user's face is identified.

[1701] Step 3:

[1702] Based on the facial recognition data, the device measures the user's biometric information. Specifically, it uses video analysis software to measure heart rate, blood pressure, body temperature, etc. Photoplethysmography technology is used for this purpose. The input is facial recognition data, and the output is biometric data.

[1703] Step 4:

[1704] The terminal starts a voice dialogue with the user, asks a question through a microphone, and obtains the user's response. For example, the terminal asks, "What kind of product are you looking for today?" The input is the user's voice response, and the output is the obtained voice data.

[1705] Step 5:

[1706] The device sends the captured voice data to the Google Speech-to-Text API, which converts the voice into text data. Furthermore, speech analysis software is used to evaluate the user's emotional state based on the tone and speed of the voice. The input is voice data, and the output is textual voice data and emotional state data.

[1707] Step 6:

[1708] The device sends facial video data and audio data to a server, which then integrates the data to perform comprehensive health and emotional assessments. The inputs are biometric data, audio text data, and emotional state data, and the output is an integrated health and emotional assessment report.

[1709] Step 7:

[1710] The server returns the generated health and emotion evaluation reports to the terminal, which then displays these reports on displays in the store or on staff members' smart devices. The input is the health and emotion evaluation reports, and the output is the displayed reports.

[1711] Step 8:

[1712] Store staff provide personalized services based on the user's state. For example, if the user is stressed, they will suggest products that will help them relax. The staff responds to the customer while referring to the display on the terminal. The input is the displayed report, and the output is the service provided.

[1713] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1714] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1715] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1716] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1717] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1718] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1719] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1720] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1721] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1722] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1723] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1724] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1725] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1726] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1727] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1728] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1729] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1730] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1731] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1732] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1733] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1734] The following is further disclosed regarding the above embodiment.

[1735] (Claim 1)

[1736] a means for acquiring facial images;

[1737] A means of analyzing facial images and measuring biometric information,

[1738] means for conducting a voice interaction with a user;

[1739] A means for analyzing voice data to assess health status;

[1740] a means of integrating the acquired data to provide a comprehensive health assessment;

[1741] means for generating a health report;

[1742] A means to view the generated health report

[1743] A system including:

[1744] (Claim 2)

[1745] 10. The system according to claim 1, further comprising means for acquiring lifestyle information of the user through voice dialogue.

[1746] (Claim 3)

[1747] 10. The system according to claim 1, further comprising means for analyzing the facial image to measure heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, hemoglobin concentration, stress, and fatigue level.

[1748] "Example 1"

[1749] (Claim 1)

[1750] a means for acquiring facial images;

[1751] A means for analyzing the acquired facial image and measuring biometric information;

[1752] means for conducting a voice interaction with a user;

[1753] A means for analyzing voice data to assess health status;

[1754] a means for integrating the acquired video data and audio data to perform a comprehensive health assessment;

[1755] means for generating a health report;

[1756] A means to view the generated health report

[1757] A system including:

[1758] (Claim 2)

[1759] 10. The system according to claim 1, further comprising means for acquiring lifestyle habit information of the user through voice dialogue and analyzing the information.

[1760] (Claim 3)

[1761] 10. The system according to claim 1, further comprising means for analyzing facial images to measure biometric information including heart rate, blood pressure, and body temperature.

[1762] "Application Example 1"

[1763] (Claim 1)

[1764] a means for acquiring facial images;

[1765] A means of analyzing facial images and measuring biometric information,

[1766] means for conducting a voice interaction with a user;

[1767] A means for analyzing voice data to assess health status;

[1768] a means of integrating the acquired data to provide a comprehensive health assessment;

[1769] means for generating a health report;

[1770] a means for displaying the generated health report;

[1771] A means of instantly assessing users' health using mirrors installed in physical stores

[1772] A system including:

[1773] (Claim 2)

[1774] 10. The system according to claim 1, further comprising means for acquiring lifestyle information of the user through voice dialogue.

[1775] (Claim 3)

[1776] 10. The system according to claim 1, further comprising means for analyzing the facial image to measure heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, hemoglobin concentration, stress, and fatigue level.

[1777] "Example 2: Combining Emotion Engines"

[1778] (Claim 1)

[1779] a means for acquiring facial images;

[1780] A means of analyzing facial images and measuring biometric information,

[1781] means for conducting a voice interaction with a user;

[1782] means for analyzing the voice data to assess health status and emotions;

[1783] a means of integrating the acquired data to provide a comprehensive health and emotional assessment;

[1784] means for generating health and emotional reports;

[1785] A means to display the generated health and emotion reports

[1786] A system including:

[1787] (Claim 2)

[1788] 10. The system according to claim 1, further comprising means for acquiring lifestyle information of the user through voice dialogue.

[1789] (Claim 3)

[1790] 10. The system according to claim 1, further comprising means for analyzing the facial image to measure heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, hemoglobin concentration, emotion, stress, and fatigue level.

[1791] "Application example 2 when combining emotion engines"

[1792] (Claim 1)

[1793] a means for acquiring facial images;

[1794] A means of analyzing facial images and measuring biometric information,

[1795] means for conducting a voice interaction with a user;

[1796] A means for analyzing voice data to assess health status;

[1797] a means of integrating the acquired data to provide a comprehensive health assessment;

[1798] means for generating a health report;

[1799] a means for displaying the generated health report;

[1800] Displaying health and emotional assessments on in-store displays and staff devices to provide personalized service

[1801] A system including:

[1802] (Claim 2)

[1803] 10. The system according to claim 1, further comprising means for acquiring lifestyle information of the user through voice dialogue.

[1804] (Claim 3)

[1805] 10. The system according to claim 1, further comprising means for analyzing the facial image to measure heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, hemoglobin concentration, stress, and fatigue level. [Explanation of symbols]

[1806] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for acquiring facial images; A means of analyzing facial images and measuring biometric information, means for conducting a voice interaction with a user; A means for analyzing voice data to assess health status; a means of integrating the acquired data to provide a comprehensive health assessment; means for generating a health report; A means to view the generated health report A system including:

2. 10. The system according to claim 1, further comprising means for acquiring lifestyle information of the user through voice dialogue.

3. 2. The system according to claim 1, further comprising means for analyzing the facial image to measure heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, hemoglobin concentration, stress, and fatigue level.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A