System
The system addresses the limitations of conventional health monitoring by integrating user information input, voice-to-text conversion, health analysis, and medical institution collaboration to provide comprehensive and timely healthcare support.
Patent Information
- Application Number
- JP2024124053
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Conventional health monitoring systems fail to provide comprehensive and timely health status monitoring, early detection of abnormalities, and linkage with medical institutions, leading to inadequate healthcare support for individuals.
A system that integrates user information input, voice-to-text conversion, health analysis, anomaly detection, and notification, with the ability to collaborate with medical institutions and adjust based on user feedback.
Enables comprehensive health monitoring, early detection of abnormalities, and prompt professional support by integrating user information input, voice-to-text conversion, health analysis, and collaboration with medical institutions.
Smart Images

Figure 2026022536000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, the importance of mental and physical healthcare is increasingly recognized, yet many individuals find it difficult to properly monitor their own health status and detect abnormalities early. Conventional methods, such as individual devices and applications, only provide fragmented information, which is insufficient for overall health monitoring, early detection of abnormalities, and linkage with appropriate medical institutions and consultation centers. Therefore, there is a need for a system that allows individuals to comprehensively manage their health status and receive professional support when needed. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for inputting a user's basic information via a terminal, transmitting it to a server for storage, a means for converting voice input into text data and transmitting it to the server, and a means for analyzing the text data received by the server and evaluating the user's condition. The system also includes a means for detecting abnormalities based on the user's condition and notifying the user, as well as a means for generating notifications and advice for the user and transmitting them to the terminal. If an abnormality is detected, the system also includes a means for invoking an API for linking with a medical institution or consultation center and transmitting the necessary data with the user's permission. The system also includes a means for collecting user feedback, analyzing it on the server, and evaluating and adjusting system performance. In this way, the system can comprehensively monitor the user's health condition and provide appropriate advice and professional support as needed.
[0006] "User" refers to an individual who utilizes this system to manage their own mental and physical health care.
[0007] "Basic information" refers to personal identification information such as the user's name, age, and gender, as well as initial setup data required for health monitoring.
[0008] "Terminal" refers to a device used by a user to input voice or data, such as an earphone, smartphone, or speaker.
[0009] "Server" refers to a central system for receiving and analyzing data sent by users, and storing and transmitting necessary information.
[0010] "Voice input" refers to voice data that a user speaks to the system through a terminal.
[0011] "Text data" refers to data that has been analyzed using voice recognition technology and converted into text information from voice input.
[0012] "Database" refers to a structured data storage system within a server for storing basic user information and analyzed text data.
[0013] "Anomaly detection" refers to the process of determining if there is something abnormal in a user's condition, and is done using machine learning algorithms.
[0014] "Notification" refers to messages or alerts sent from the server to the terminal, informing the user of abnormalities or advice.
[0015] "API" stands for Application Programming Interface and refers to a standardized interface for communication and data exchange between different software systems.
[0016] "Feedback" refers to the impressions and opinions that users receive in response to advice and notifications from the system, and by sending these back to the system, provides information that will be useful in improving the system. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] The system of the present invention supports the mental and physical health care of users and is composed of the following main means:
[0039] User information registration and initial settings
[0040] server:
[0041] Provides an interface for identifying the device (earphones, smartphone, speaker, etc.) used by the user.
[0042] Manages a database for entering and storing basic user information (name, age, gender, etc.).
[0043] Device:
[0044] Displays an initial setup screen and prompts the user to enter basic information.
[0045] The entered basic information is sent to the server.
[0046] Displays to the user the confirmation message received from the server.
[0047] Daily conversations and data collection
[0048] Device:
[0049] At set times, a generative AI voice will ask the user questions (e.g., "What did you eat today?", "How are you feeling?").
[0050] The user's voice input is recorded and converted into text data using voice recognition technology.
[0051] The converted text data is sent to the server.
[0052] server:
[0053] Receives text data sent from the terminal and stores it in a database.
[0054] The received data is analyzed periodically to assess the user's health condition.
[0055] Trend analysis and anomaly detection
[0056] server:
[0057] The stored data is analyzed using machine learning algorithms to identify trends in the user's mental and physical health.
[0058] Based on trend data, an alert is generated if an abnormal change is detected.
[0059] Send the alert content to the device.
[0060] Device:
[0061] Notify the user of alerts received from the server by voice or text.
[0062] Presents the user with options to contact a healthcare provider.
[0063] Collaboration and advice
[0064] server:
[0065] Generate appropriate advice based on the user's health status and send it to the device (e.g., "Drink more water today.").
[0066] If an abnormality is detected, a collaboration API with a pre-configured medical institution or consultation center will be called.
[0067] If the user gives permission, the necessary data will be sent to a medical institution or consultation center.
[0068] Device:
[0069] The user is notified of the advice received from the server and prompted to take action.
[0070] With the user's permission, the necessary data is sent to medical institutions and consultation centers.
[0071] Feedback collection and system adjustments
[0072] Device:
[0073] An interface is displayed to the user requesting feedback on the usefulness of the system's advice and notifications.
[0074] Sends the user-entered feedback to the server.
[0075] server:
[0076] Analyze the feedback received and evaluate the performance of the system.
[0077] If necessary, adjust the questions the AI asks and how it generates advice.
[0078] For example, if a user answers a morning question with "I didn't sleep well last night," the system analyzes this information and determines that the user has reported a continuous lack of sleep. The server detects the abnormality and generates advice to the user, such as "We recommend that you review your nightly routine to improve the quality of your sleep," and notifies the user via their device. In this way, the system can comprehensively monitor the user's health and provide prompt and accurate advice.
[0079] The processing flow will be explained below.
[0080] Specific processing steps of the program
[0081] User information registration and initial settings
[0082] Step 1:
[0083] User: Start the device and open the initial setup screen.
[0084] Step 2:
[0085] Terminal: Presents an interface for the user to enter basic information (e.g., name, age, gender).
[0086] Step 3:
[0087] User: Enter the required basic information.
[0088] Step 4:
[0089] Terminal: Sends the entered basic information to the server.
[0090] Step 5:
[0091] Server: Receives basic information sent from the device and stores it in a database.
[0092] Server: Sends a confirmation message to the device that the device was successfully registered.
[0093] Step 6:
[0094] Terminal: Display a confirmation message to the user.
[0095] Daily conversations and data collection
[0096] Step 7:
[0097] Device: Automatically wakes up at a set time and asks the user questions using a generative AI voice (e.g., "What did you eat today?").
[0098] Step 8:
[0099] User: Answers questions by voice.
[0100] Step 9:
[0101] Terminal: Records the user's voice and converts it into text data using voice recognition technology.
[0102] Step 10:
[0103] Terminal: Sends the converted text data to the server.
[0104] Step 11:
[0105] Server: Receives text data sent from the device and stores it in a database.
[0106] Trend analysis and anomaly detection
[0107] Step 12:
[0108] Server: Analyzes the stored data using machine learning algorithms to generate trends in the user's mental and physical health.
[0109] Step 13:
[0110] Server: Generates alerts based on trend data if abnormal changes are detected.
[0111] Step 14:
[0112] Server: Sends alert content to the device.
[0113] Step 15:
[0114] Terminal: Notifies the user of alerts received from the server via voice or text.
[0115] Step 16:
[0116] Device: Presents the user with options to contact a healthcare provider.
[0117] Collaboration and advice
[0118] Step 17:
[0119] Server: Generates appropriate advice based on the user's health status and sends it to the device (e.g., "Drink more water today.").
[0120] Step 18:
[0121] Server: If an abnormality is detected, it calls a link API with a pre-configured medical institution or consultation center.
[0122] Step 19:
[0123] Server: If the user gives permission, it sends the necessary data to a medical institution or consultation center.
[0124] Step 20:
[0125] Terminal: Notifies the user of the advice received from the server and encourages them to take action.
[0126] Step 21:
[0127] Terminal: With the user's permission, transmits the necessary data to medical institutions and consultation centers.
[0128] Feedback collection and system adjustments
[0129] Step 22:
[0130] Terminal: Display an interface that asks the user for feedback on the usefulness of system advice and notifications.
[0131] Step 23:
[0132] Users: Enter feedback.
[0133] Step 24:
[0134] Terminal: Sends the entered feedback to the server.
[0135] Step 25:
[0136] Server: Analyzes the feedback received and evaluates the performance of the system.
[0137] Server: Adjusts the questions the AI asks and how it generates advice as needed.
[0138] Through the above processing steps, it is possible to comprehensively monitor the user's health condition and provide appropriate advice and support.
[0139] Example 1
[0140] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0141] In modern society, people face the challenge of properly monitoring their own health status and responding promptly in their busy daily lives. In particular, in the areas of mental and physical health, people often delay self-awareness and are unable to identify appropriate countermeasures, creating a need for more efficient and effective healthcare support systems. Furthermore, when an abnormality is detected, it is necessary to quickly collaborate with medical institutions and take appropriate measures, but achieving this requires advanced technology and rapid data analysis.
[0142] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0143] In this invention, the server includes a means for inputting basic information of the user via a terminal and transmitting and storing the information to the server, a means for converting voice input into text data and transmitting the text data to the server, and a means for analyzing the received text data on the server side and evaluating the user's condition. This enables comprehensive monitoring of the user's mental and physical health and prompt cooperation with a medical institution when an abnormality is detected.
[0144] "Basic user information" refers to basic information for identifying an individual, such as the user's name, age, and gender.
[0145] "Terminal" refers to a device used by a user, such as a smartphone, earphones, or speaker.
[0146] "Server" refers to a computer system that manages and processes user information, and stores, analyzes, and notifies various data.
[0147] "Speech input" is the sound data that the user provides to the system through speech synthesis.
[0148] "Text data" refers to data that has been converted from voice input into text information.
[0149] "Analysis" refers to the process of analyzing received data and assessing the user's health status.
[0150] "Evaluation" is the act of determining the user's current health condition based on the analyzed data.
[0151] A "machine learning algorithm" is a method for using large amounts of data to train models to identify trends and make predictions.
[0152] "Trend" refers to the pattern or tendency of fluctuations in a user's health condition over a certain period of time.
[0153] "Abnormal" refers to data or conditions that deviate from the normal range and may have adverse health effects.
[0154] "Notification" is the act of notifying the user of information detected on the server side.
[0155] "Advice" refers to specific advice or suggestions provided based on the user's health status.
[0156] "API" refers to an interface for exchanging data between different systems.
[0157] "Feedback" refers to the evaluations and opinions that users give regarding the services and functions provided by the system.
[0158] A "generative AI model" refers to a model that uses artificial intelligence technology to generate questions and advice.
[0159] The system of the present invention supports users' mental and physical healthcare and includes the following processing steps: registering user information, daily conversations and data collection, trend analysis and anomaly detection, collaboration and advice provision, feedback collection and system adjustment.
[0160] Hardware and software used
[0161] 1. Server: Used for reliable data management and analysis. Specifically, it provides a REST API and uses a MySQL database to store basic user information and collected data. It also uses Python and Pandas for data analysis, and scikit-learn and TensorFlow for implementing machine learning algorithms.
[0162] 2. Terminal: A device through which a user directly inputs information. This includes smartphones, earphones, speakers, etc. It uses the Google Speech-to-Text API to convert voice input into text data and the Google Text-to-Speech API to play the audio.
[0163] System Operation
[0164] The server receives basic information entered by the user via the device and stores it in a database. The device periodically asks the user questions and accepts voice input. The voice is converted into text data and sent to the server. The server analyzes the received text data and evaluates the user's health condition. A machine learning algorithm is used to analyze health trends, and if an abnormality is detected, an alert is quickly generated and sent to the device.
[0165] The server then generates advice based on the user's condition and any abnormalities and sends it to the device. If an abnormality is detected, it calls an API to connect with medical institutions or consultation centers and, with the user's permission, transmits the necessary data. The device then displays the received advice and notifications to the user, encouraging them to take action. It also collects user feedback and uses it to evaluate and adjust the system's performance.
[0166] As a concrete example, consider the case where a user answers a morning question with, "I didn't sleep well last night." The server analyzes this information and confirms that the user has reported a continuous lack of sleep. If the server detects an abnormality, it generates advice such as, "We recommend that you review your nightly routine to improve the quality of your sleep," and notifies the device. The user receives the advice through their device and can contact a medical institution or consultation center if necessary.
[0167] Example prompt sentence:
[0168] "Write a prompt to generate appropriate advice based on the user's health condition, such as when they say they didn't sleep well last night."
[0169] In this way, the system of the present invention can comprehensively monitor the user's health condition and provide prompt and accurate advice.
[0170] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0171] Step 1:
[0172] User information registration and initial settings
[0173] server:
[0174] The server provides an interface (REST API) that identifies the device used by the user. The input is the user's basic information (name, age, gender, etc.) sent from the device, which the server receives and saves in a database. For example, the information is stored in a MySQL database. The output is a confirmation message indicating that the save was successful.
[0175] Device:
[0176] The device displays an initial setup screen and prompts the user to enter basic information. Specifically, the smartphone app displays a form, the user enters the required information, and taps the "Submit" button. The input is information provided by the user, such as name, age, and gender, and the output is the transmission of the information to the server.
[0177] The terminal displays to the user the confirmation message received from the server.
[0178] Step 2:
[0179] Daily conversations and data collection
[0180] Device:
[0181] At a set time, a generative AI voice is used to ask the user questions. For example, questions such as "What did you eat today?" or "How are you feeling?" are played back aloud using the Google Text-to-Speech API. The input is the question sent from the server, and the output is the user's voice input.
[0182] The user responds by voice, and the device records this voice. The recorded voice data is converted into text data using the Google Speech-to-Text API. The input is the user's voice data, and the output is the converted text data.
[0183] The converted text data is sent to the server via HTTPS. The input is the text data, and the output is a transmission confirmation message.
[0184] server:
[0185] The server receives text data sent from the terminal and stores it in a database. The input is the text data, and the output is a confirmation message indicating that the data was saved successfully.
[0186] Step 3:
[0187] Trend analysis and anomaly detection
[0188] server:
[0189] The server analyzes the stored data using machine learning algorithms (e.g., scikit-learn or TensorFlow) to find trends in the user's mental and physical health. The input is a collection of past text data, and the output is the analyzed trend information.
[0190] Generates an alert when an abnormal change is detected. The input is the analyzed trend information, and the output is an alert message indicating the abnormality.
[0191] Sends alert content to the terminal. The input is the alert message, and the output is a transmission confirmation message.
[0192] Device:
[0193] The device notifies the user of the alert received from the server by voice or text, for example by displaying a pop-up on the screen or by reading it aloud. The input is the alert message, and the output is the notified alert.
[0194] Presents the user with options to contact the healthcare provider, allowing them to select "Contact now" or "Later."
[0195] Step 4:
[0196] Collaboration and advice
[0197] server:
[0198] The server generates appropriate advice based on the user's health condition and sends it to the device. For example, the advice may be "Drink more water today." The input is the health condition assessment result, and the output is the generated advice message.
[0199] If an abnormality is detected, a pre-configured API for linking with a medical institution or consultation center is called. For example, a notification is sent using the Twilio API. The input is an alert message indicating the abnormality, and the output is a message confirming that the notification has been sent.
[0200] If the user gives permission, the necessary data is sent to the medical institution or consultation center. The input is the user's consent and the necessary data, and the output is a transmission confirmation message.
[0201] Device:
[0202] The device notifies the user of the advice received from the server and prompts them to take action, for example by displaying it on the screen or reading it out loud. The input is the advice message, and the output is the notified advice.
[0203] With the user's permission, the necessary data will be sent to the medical institution or consultation center. The user taps the "Send data" button.
[0204] Step 5:
[0205] Feedback collection and system adjustments
[0206] Device:
[0207] The terminal displays an interface to ask the user for feedback on the usefulness of the system's advice or notifications. For example, it displays a rating slider or a comment input form. The input is the feedback request, and the output is the feedback data entered by the user.
[0208] Sends the feedback entered by the user to the server. The input is the feedback data, and the output is a transmission confirmation message.
[0209] server:
[0210] The server analyzes the received feedback and evaluates the system's performance. Here, NLP techniques are used to analyze text. The input is the feedback data, and the output is the performance evaluation results.
[0211] If necessary, adjust the questions the AI asks and how it generates advice. For example, retrain the model based on the latest user data. The input is the performance evaluation results and the latest user data, and the output is the adjusted model.
[0212] Example prompt sentence:
[0213] "Write a prompt to generate appropriate advice based on the user's health condition, such as when they say they didn't sleep well last night."
[0214] (Application example 1)
[0215] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0216] Conventional health management systems have had difficulty accurately monitoring a user's daily health status and providing appropriate advice tailored to each individual's situation. Furthermore, they lacked mechanisms for early detection of abnormalities and prompt communication with medical institutions, resulting in delayed responses to maintain optimal user health. To address these issues, there is a need for technology that uses the user's voice input to check daily health status and generate appropriate prompts.
[0217] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0218] In this invention, the server includes means for inputting basic information of a user via a terminal and transmitting it to the server for storage, means for converting voice input into text data and transmitting it to the server, means for analyzing the text data received by the server and evaluating the user's condition, means for detecting abnormalities based on the user's condition and notifying the user, means for generating notifications and advice for the user and transmitting them to the terminal, and means for converting voice input into text data via the terminal, analyzing the text data, and providing appropriate advice using prompt sentences generated by a generative AI model according to the user's condition. This makes it possible to accurately monitor the user's daily health condition and quickly provide appropriate advice according to each individual situation.
[0219] "Basic information" refers to basic personal data such as a user's name, age, and gender.
[0220] A "terminal" is an electronic device used by a user, and refers to devices such as smartphones, earphones, and speakers.
[0221] "Server" refers to a central processing unit for receiving and processing information sent by users.
[0222] "Voice input" refers to the act of a user providing information by voice.
[0223] "Text data" refers to data that has been converted from voice input into text information.
[0224] "Analysis" refers to the act of processing received data and finding meaning and trends.
[0225] "Abnormality" refers to any unusual change or sign in the user's health condition.
[0226] "Notification" refers to the act of informing a user of information.
[0227] "Advice" refers to specific instructions or advice provided based on the user's condition.
[0228] A "collaboration API" refers to an interface for sharing data and functions between different systems and services.
[0229] "Generative AI model" refers to an artificial intelligence model that generates text or advice based on user input.
[0230] A "prompt sentence" is an input sentence to a generative AI model, and refers to trigger text that elicits appropriate answers or advice.
[0231] This invention is a system that supports users' health management and includes a server, a terminal, and a generative AI model. The system collects basic information about the user, analyzes health data based on that information, and provides appropriate advice as needed. Furthermore, if an abnormality is detected, the system also cooperates with medical institutions.
[0232] User information registration and initial settings
[0233] The terminal provides an interface for entering basic information about the user (such as name, age, and gender). Once the user has entered the information, it is sent to the server and stored. At this point, the terminal receives a confirmation message from the server and displays it to the user.
[0234] Daily conversations and data collection
[0235] The device uses the generative AI model to ask the user questions at set times (e.g., "What did you eat today?", "How are you feeling?"). The user responds verbally, and the speech is converted into text data using the Google Cloud Speech-to-Text API. The converted text data is sent to a server and stored in a database.
[0236] Trend analysis and anomaly detection
[0237] The server uses machine learning algorithms to analyze the stored text data. These algorithms can identify trends in the user's mental and physical health. If an abnormality is detected based on the trend data, the server generates an alert and sends it to the device. The device notifies the user of the alert and provides options for contacting a medical institution, if necessary.
[0238] Collaboration and advice
[0239] The server uses a generative AI model to generate appropriate advice based on the user's health status. This generation utilizes the prompt format. For example, if the user answers, "I didn't sleep well last night," the server generates advice such as, "I recommend trying some light stretching before bed to relax."
[0240] Advice is sent to the user via the device. If an abnormality is detected, the server calls the API to connect with the medical institution and, with the user's permission, sends the necessary data.
[0241] Feedback collection and system adjustments
[0242] The device provides an interface that asks the user for feedback on the usefulness of the system's advice and notifications. The user's input feedback is sent to a server for analysis. The server uses this information to evaluate the system's performance and adjusts the questions the generative AI model asks and how it generates advice, if necessary.
[0243] As a specific example of use, if a user answers a morning question with "I didn't sleep well last night," the system analyzes this information and determines that the user has reported poor sleep for several consecutive nights. The server detects the anomaly and generates advice to the user, such as "We recommend that you review your nightly routine to improve the quality of your sleep," and notifies the user via their device.
[0244] Example prompt sentence:
[0245] User Input: I didn't sleep very well last night.
[0246] Health Tip: To help you relax, try some gentle stretching before bed.
[0247] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0248] Step 1:
[0249] The terminal inputs the user's basic information (name, age, gender, etc.) and sends it to the server. The input data is sent from the terminal to the server, and the server saves the data in a database. This registers the user's basic information in the system.
[0250] Step 2:
[0251] At the set time, the device uses the generative AI model to ask the user questions (e.g., "What did you eat today?", "How are you feeling?"). This voice input is input as voice data from the user.
[0252] Step 3:
[0253] The device uses the Google Cloud Speech-to-Text API to convert voice input into text data, which is then sent to a server for analysis.
[0254] Step 4:
[0255] The server analyzes the received text data using machine learning algorithms to assess the user's health status. The output of this analysis step is the assessment of the user's health status.
[0256] Step 5:
[0257] The server detects abnormalities based on the user's health status. To do this, it compares saved data with new input data. If an abnormality is detected, it generates an alert and sends it to the device. The content of the alert indicates that an abnormality exists.
[0258] Step 6:
[0259] The terminal notifies the user of the alert received from the server and presents an option to contact a medical institution if necessary. The output of the notification is that the user receives the alert information.
[0260] Step 7:
[0261] The server uses a generative AI model to generate appropriate advice based on the user's condition. This advice generation uses prompt sentences. For example, if a user inputs "I didn't sleep well last night," the server generates advice such as "I recommend you try some light stretching before bed to relax."
[0262] Step 8:
[0263] The advice is sent from the server to the terminal, and the terminal notifies the user, allowing the user to receive specific health advice.
[0264] Step 9:
[0265] If necessary, the server will call the linkage API and, with the user's permission, send the necessary data to a medical institution or consultation center. This allows the user's health information to be provided to the appropriate institution, enabling a prompt response.
[0266] Step 10:
[0267] The device asks the user for feedback on the usefulness of the system's advice and notifications. The user enters the feedback, which is then sent to the server. The server analyzes the feedback data to evaluate and adjust the system's performance. This adjustment improves the accuracy of the generative AI model and the overall system performance.
[0268] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0269] The system of the present invention supports the user's mental and physical healthcare and is composed of the following main means. Furthermore, by combining it with an emotion engine for recognizing the user's emotions, it is possible to provide more advanced care.
[0270] User information registration and initial settings
[0271] server:
[0272] Provides an interface for identifying the device (earphones, smartphone, speaker, etc.) used by the user.
[0273] Manages a database for entering and storing basic user information (name, age, gender, etc.).
[0274] Device:
[0275] Displays an initial setup screen and prompts the user to enter basic information.
[0276] The entered basic information is sent to the server.
[0277] Displays to the user the confirmation message received from the server.
[0278] Daily conversations and data collection
[0279] Device:
[0280] It automatically starts at the set time and asks the user questions using a generative AI voice (e.g., "What did you eat today?", "How are you feeling?").
[0281] The user's voice input is recorded and converted into text data using voice recognition technology.
[0282] The converted text data is sent to the server.
[0283] server:
[0284] Receives text data sent from the terminal and stores it in a database.
[0285] Emotion Recognition and State Assessment
[0286] Emotion engine (can run on device or server):
[0287] The user's voice tone, pitch, and rate are analyzed to identify emotions.
[0288] The identified emotion data is transmitted to a server.
[0289] server:
[0290] The emotion data received from the emotion engine is analyzed together with the text data to comprehensively evaluate the user's state.
[0291] The stored data is analyzed using machine learning algorithms to generate trends in the user's mental and physical health.
[0292] Based on trend data, an alert is generated if an abnormal change is detected.
[0293] Send the alert content to the device.
[0294] Device:
[0295] Notify the user of alerts received from the server by voice or text.
[0296] Presents the user with options to contact a healthcare provider.
[0297] Collaboration and advice
[0298] server:
[0299] Generate appropriate advice based on the user's health status and emotional data and send it to the device (e.g., "We recommend listening to music to relax today.").
[0300] If an abnormality is detected, a collaboration API with a pre-configured medical institution or consultation center will be called.
[0301] If the user gives permission, the necessary data will be sent to a medical institution or consultation center.
[0302] Device:
[0303] The user is notified of the advice received from the server and prompted to take action.
[0304] With the user's permission, the necessary data is sent to medical institutions and consultation centers.
[0305] Feedback collection and system adjustments
[0306] Device:
[0307] An interface is displayed to the user requesting feedback on the usefulness of the system's advice and notifications.
[0308] Sends the user-entered feedback to the server.
[0309] server:
[0310] Analyze the feedback received and evaluate the performance of the system.
[0311] If necessary, adjust the questions the AI asks and how it generates advice.
[0312] For example, if a user responds to a morning question with "I didn't sleep well last night," and the emotion engine detects a depressed tone, the system analyzes this information and determines that the user has continued to have poor sleep. The server detects the abnormality and generates advice for the user, such as "To improve the quality of your sleep, we recommend listening to relaxing music tonight," and notifies the user via their device. In this way, the system comprehensively monitors the user's health status, including emotion recognition, and can provide prompt and accurate advice.
[0313] The processing flow will be explained below.
[0314] Specific processing steps of the program
[0315] User information registration and initial settings
[0316] Step 1:
[0317] User: Start the device and open the initial setup screen.
[0318] Step 2:
[0319] Terminal: Presents an interface for the user to enter basic information (e.g., name, age, gender).
[0320] Step 3:
[0321] User: Enter the required basic information.
[0322] Step 4:
[0323] Terminal: Sends the entered basic information to the server.
[0324] Step 5:
[0325] Server: Receives basic information sent from the device and stores it in a database.
[0326] Server: Sends a confirmation message to the device that the device was successfully registered.
[0327] Step 6:
[0328] Terminal: Display a confirmation message to the user.
[0329] Daily conversations and data collection
[0330] Step 7:
[0331] Device: Automatically wakes up at a set time and asks questions to the user using a generative AI voice (e.g., "What did you eat today?", "How are you feeling?").
[0332] Step 8:
[0333] User: Answers questions by voice.
[0334] Step 9:
[0335] Terminal: Records the user's voice and converts it into text data using voice recognition technology.
[0336] Step 10:
[0337] Terminal: Sends the converted text data to the server.
[0338] Step 11:
[0339] Server: Receives text data sent from the device and stores it in a database.
[0340] Emotion Recognition and State Assessment
[0341] Step 12:
[0342] Terminal or Server (Emotion Engine): Analyzes the user's voice tone, pitch, and rate to identify emotions.
[0343] Step 13:
[0344] Emotion engine (terminal or server): Sends identified emotion data to the server.
[0345] Step 14:
[0346] Server: Analyzes the emotion data received from the emotion engine along with the text data, and makes a comprehensive assessment of the user's state.
[0347] Step 15:
[0348] Server: Analyzes the stored data using machine learning algorithms to generate trends in the user's mental and physical health.
[0349] Step 16:
[0350] Server: Generates alerts based on trend data if abnormal changes are detected.
[0351] Step 17:
[0352] Server: Sends alert content to the device.
[0353] Step 18:
[0354] Terminal: Notifies the user of alerts received from the server via voice or text.
[0355] Step 19:
[0356] Device: Presents the user with options to contact a healthcare provider.
[0357] Collaboration and advice
[0358] Step 20:
[0359] Server: Generates appropriate advice based on the user's health status and emotional data and sends it to the device (e.g., "We recommend listening to music to relax today.").
[0360] Step 21:
[0361] Server: If an abnormality is detected, it calls a link API with a pre-configured medical institution or consultation center.
[0362] Step 22:
[0363] Server: If the user gives permission, it sends the necessary data to a medical institution or consultation center.
[0364] Step 23:
[0365] Terminal: Notifies the user of the advice received from the server and encourages them to take action.
[0366] Step 24:
[0367] Terminal: With the user's permission, transmits the necessary data to medical institutions and consultation centers.
[0368] Feedback collection and system adjustments
[0369] Step 25:
[0370] Terminal: Display an interface that asks the user for feedback on the usefulness of system advice and notifications.
[0371] Step 26:
[0372] Users: Enter feedback.
[0373] Step 27:
[0374] Terminal: Sends the entered feedback to the server.
[0375] Step 28:
[0376] Server: Analyzes the feedback received and evaluates the performance of the system.
[0377] Server: Adjusts the questions the AI asks and how it generates advice as needed.
[0378] The above processing steps enable comprehensive monitoring of a user's health status and provide appropriate advice and support, including emotion recognition. For example, if a user responds to a morning question with "I didn't sleep well last night," and the emotion engine detects a depressed tone, the system analyzes this information, considers the user's mental and physical health trends, and generates advice such as "To improve the quality of your sleep, we recommend listening to relaxing music tonight," and notifies the user via the device. In this way, this system realizes comprehensive health management that incorporates emotion recognition.
[0379] Example 2
[0380] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0381] In recent years, stress and anxiety have increased in modern society, making mental and physical healthcare increasingly important. However, conventional healthcare systems lack the ability to accurately identify a user's emotional state and provide appropriate advice. As a result, it is difficult to detect abnormalities in a user's state early and take appropriate measures.
[0382] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0383] In this invention, the server includes means for inputting basic information of a user via a terminal and transmitting the information to the server for storage, means for converting voice input into text data and transmitting the text data to the server, means for analyzing the text data received on the server side and evaluating the user's condition, means for analyzing the user's voice tone, pitch, and speed using an emotion engine and identifying emotions, means for transmitting the identified emotion data to the server and analyzing it together with the text data, means for detecting abnormalities based on the user's condition and notifying the user, and means for generating notifications and advice for the user and transmitting the generated notifications and advice to the terminal. This makes it possible to comprehensively monitor the user's emotional state and health condition, detect abnormalities early, and quickly provide appropriate advice and responses.
[0384] "Means for inputting basic user information via a terminal and sending it to a server for storage" refers to the process and infrastructure that allows a user to input basic personal information such as name, age, and gender using the terminal interface, and then transmit the input data to a server and store it in a database.
[0385] "Means for converting voice input into text data and transmitting it to a server" refers to technology and devices for recording a user's voice input, converting it into text data using voice recognition technology, and transmitting the text data to a server.
[0386] "Means for analyzing text data received on the server side and evaluating the user's condition" refers to a system and analysis algorithm that analyzes and processes text data received by the server from the terminal and evaluates the user's mental and physical health condition.
[0387] "Means for identifying emotions by analyzing a user's voice tone, pitch, and rate using an emotion engine" refers to a technology and method for using an emotion engine to analyze characteristics of a user's speech, such as voice tone, pitch, and rate, to identify the user's emotional state.
[0388] The "means for transmitting identified emotion data to a server for analysis together with text data" refers to a process and system for transmitting emotion data identified by the emotion engine to a server for analysis of the data together with text data.
[0389] "Means for detecting abnormalities based on the user's condition and notifying the user" refers to a system and method that determines whether there are any abnormalities in the user's health or emotional state based on data analyzed by the server and notifies the user of the results.
[0390] "Means for generating notifications and advice for the user and sending them to the terminal" refers to the technology and platform for generating appropriate notifications and advice based on the results of analyzing the user's condition and sending them to the terminal.
[0391] "Means for calling APIs to link with medical institutions and consultation centers" refers to systems and technologies for calling application programming interfaces (APIs) to link with pre-set medical institutions and consultation centers when an abnormality is detected.
[0392] "Means for transmitting necessary data to a medical institution or consultation center with the user's permission" refers to the process and system for transmitting necessary health data and emotion data to a designated medical institution or consultation center with the user's consent.
[0393] "Means for collecting user feedback, analyzing it on the server side, and evaluating and adjusting system performance" refers to a mechanism for collecting feedback from users regarding system usage, analyzing it on the server to evaluate system performance, and adjusting the system's functions and service content as necessary.
[0394] This invention is a system for supporting users' mental and physical healthcare, which provides advice and notifications according to the situation by collecting and analyzing the user's basic information, daily activity data, and emotional data. To achieve this, a server, a terminal, and an emotion engine work in cooperation with each other.
[0395] User information registration and initial settings
[0396] Server: The server provides an interface to identify the device used by the user (earphones, smartphone, speaker, etc.). When the user connects to the device, the server collects the information and prepares it to be stored in a database. The server also manages a database for entering and storing basic information about the user (name, age, gender, etc.).
[0397] Terminal: The terminal displays an initial setup screen and prompts the user to enter basic information. After the user enters the information, the terminal sends the information to the server. The server stores the received information in a database and sends a confirmation message to the terminal. The terminal displays a confirmation message to the user, informing them that the initial setup is complete.
[0398] Daily conversations and data collection
[0399] Device: The device automatically wakes up at the set time and asks the user questions using a generative AI voice (e.g., "What did you eat today?", "How are you feeling?"). When the user responds, the device records the voice and converts it into text data using voice recognition technology (e.g., Google Speech-to-Text API).
[0400] Server: The device sends the converted text data to the server, which stores it in a database for further analysis.
[0401] Emotion Recognition and State Assessment
[0402] Emotion engine (can run on the device or on the server): The emotion engine analyzes the user's voice tone, pitch, and rate to identify emotions. The identified emotion data is sent to the server.
[0403] Server: The server analyzes the emotion data received from the emotion engine along with the text data to comprehensively evaluate the user's condition. It uses machine learning algorithms (e.g., TensorFlow, PyTorch) to generate trends in the user's mental and physical health. If any abnormal changes are detected based on the trends, it generates an alert and sends it to the device.
[0404] Collaboration and advice
[0405] Server: The server generates appropriate advice based on the user's health condition and emotional data. For example, if the user's recent sleep data is poor, the server generates advice such as "We recommend listening to relaxing music" and sends it to the device. If an abnormality is detected, the server calls a link API with a pre-set medical institution or consultation center. If the user gives permission, the server sends the necessary data to the medical institution or consultation center.
[0406] Device: The device notifies the user of the advice received from the server. It displays the advice in audio or text and prompts the user to take specific action (e.g., displaying a button to play music). With the user's permission, it also sends the necessary data to medical institutions or consultation centers.
[0407] Feedback collection and system adjustments
[0408] Terminal: The terminal provides an interface to collect user feedback on advice and notifications from the system. When the user enters feedback, it is sent to the server.
[0409] Server: The server analyzes the feedback received from users, evaluates the system's performance, and adjusts the questions the AI asks and how it generates advice to improve the system, if necessary.
[0410] Specific examples
[0411] Example 1: Morning Questions
[0412] The device starts up every morning at 7:00 and asks the user, "Good morning. How was the quality of your sleep last night?" If the user replies, "I didn't sleep well," the device records the voice, converts it into text data, and sends it to the server.
[0413] Example 2: Emotion Identification
[0414] The emotion engine analyzes the user's voice and identifies the emotion "depressed" from the low tone and slow rate. This data is sent to a server and analyzed together with the text data.
[0415] Example 3: Alert and Advice Notifications
[0416] The server continuously checks whether the quality of sleep is poor and detects any abnormalities. It generates advice such as "To improve the quality of your sleep, we recommend listening to relaxing music tonight" and sends it to the device. The device notifies the user of the advice and offers the option to play music.
[0417] Prompt Sentence Examples
[0418] "How was your sleep quality last night?"
[0419] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0420] Step 1:
[0421] The user enters basic information into the terminal and sends it to the server.
[0422] Specific operation: The device displays an initial setup screen and asks the user to enter basic information such as name, age, gender, etc. After the user enters the information, the device sends it to the server.
[0423] Input: Basic information entered by the user, such as name, age, and gender.
[0424] Output: Basic user information data sent from the device to the server.
[0425] Data processing: The server receives these data and stores them in a database.
[0426] Step 2:
[0427] The server stores the basic information in a database and sends a confirmation message to the device.
[0428] Specific operation: The server saves the received basic information in the database, generates a confirmation message for the save completion and sends it to the terminal, which then displays the confirmation message to the user.
[0429] Input: User basic information data.
[0430] Output: Basic information stored in the database and a confirmation message displayed to the user.
[0431] Data processing: The server stores the basic information in a database as structured data.
[0432] Step 3:
[0433] The device will start up at the set time and ask the user questions using a generated AI voice.
[0434] Specific operation: The device automatically starts up at a pre-set time and asks the user questions using a generated AI voice, such as, "How was the quality of your sleep last night?"
[0435] Input: None (automatic start by scheduled trigger).
[0436] Output: A spoken question to the user.
[0437] Data processing: Generate appropriate question voice data using a generative AI model.
[0438] Step 4:
[0439] The user responds by voice, and the terminal records the voice and converts it into text data.
[0440] What it does: When a user responds to a question, the device records the audio and then uses voice recognition technology to convert the audio into text data (e.g., using the Google Speech-to-Text API).
[0441] Input: The user's spoken response.
[0442] Output: User answers converted to text data.
[0443] Data processing: Converting audio data into text format.
[0444] Step 5:
[0445] The terminal transmits the text data to the server.
[0446] Specific operation: The converted text data is sent from the terminal to the server.
[0447] Input: Text data.
[0448] Output: The text data sent to the server.
[0449] Data processing: Encodes text data into transmission packets.
[0450] Step 6:
[0451] The server receives the text data and stores it in a database.
[0452] Specific operation: The server receives text data from the terminal and stores it in a database.
[0453] Input: Text data.
[0454] Output: Text data stored in a database.
[0455] Data processing: Structure the text data and record it in a database.
[0456] Step 7:
[0457] The emotion engine (running on the device or server) analyzes the voice data and identifies emotions.
[0458] How it works: The emotion engine analyzes the user's voice tone, pitch, and rate to identify emotions. The emotion data is sent to the server and analyzed along with the text data.
[0459] Input: User's voice data.
[0460] Output: Emotion data.
[0461] Data processing: Analyze voice characteristics such as tone, pitch, and speed to identify emotions.
[0462] Step 8:
[0463] The server analyzes the emotional data and text data to evaluate the user's state.
[0464] How it works: The server uses emotion data and text data to analyze the user's mental and physical state. The analysis uses machine learning algorithms (e.g., TensorFlow, PyTorch).
[0465] Input: Emotion data, text data.
[0466] Output: User status assessment results.
[0467] Data processing: Integrate both sets of data and perform statistical analysis using machine learning models.
[0468] Step 9:
[0469] The server detects anomalies based on the user's status and generates an alert.
[0470] Specific operation: The server detects anomalies based on the analysis results, generates an alert if an anomaly is detected, and sends the alert content to the user's device.
[0471] Input: User condition assessment results.
[0472] Output: Alert data.
[0473] Data processing: Conditions are judged based on the evaluation results, and alert data is generated and sent.
[0474] Step 10:
[0475] The device will notify the user of the alert and provide advice if necessary.
[0476] Specific behavior: The device notifies the user of the alert received from the server via voice or text, and generates and provides additional advice to the user if necessary.
[0477] Input: Alert data.
[0478] Output: Notification and advice to the user.
[0479] Data processing: Received alerts are notified to the user in an appropriate format and advice is generated.
[0480] Step 11:
[0481] The server and terminal work together to contact medical institutions and consultation centers.
[0482] Specific operation: If an abnormality is detected, the server calls the API for linking with the configured medical institution or consultation center. If the user gives permission, the device sends the necessary data to the medical institution or consultation center.
[0483] Input: User consent, anomaly detection results.
[0484] Output: Sending data to medical institutions and consultation centers.
[0485] Data processing: Formatting and sending the necessary data to medical institutions and consultation centers.
[0486] Step 12:
[0487] Gathering feedback and tuning system performance.
[0488] What it does: The device collects feedback from users and sends that data to the server, which analyzes the feedback, evaluates system performance, and makes adjustments as needed.
[0489] Input: User feedback data.
[0490] Output: Evaluation results and adjustments.
[0491] Data processing: Analyze feedback data to evaluate and adjust system performance.
[0492] (Application example 2)
[0493] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0494] Currently, there is no system in virtual stores that adequately supports users' mental and physical healthcare. As a result, users may feel stressed and their health may deteriorate. In particular, the shopping experience in a virtual store has a significant impact on users' psychological state, so there is a need to introduce emotion analysis and provide appropriate health advice.
[0495] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0496] In this invention, the server includes means for inputting basic information of a user via a device and transmitting it to the server for storage, means for converting voice input into text data and transmitting it to the server, means for analyzing the text data received by the server and evaluating the user's condition, means for detecting abnormalities based on the user's condition and notifying the user, means for generating notifications and advice for the user and transmitting them to the device, means for analyzing the user's emotions and generating healthy advice based on the emotions and notifying the user, and means for collecting feedback, analyzing it on the server, and evaluating and adjusting system performance. This reduces the stress felt by users during their shopping experience in a virtual store, enabling a healthy and comfortable experience.
[0497] "Basic user information" is information that can be used to identify an individual, such as the user's name, age, and gender.
[0498] A "device" is a device that a user operates to input and receive information, such as a smartphone, smart glasses, or headset.
[0499] "Voice input" is voice information that a user utters into a device.
[0500] "Text data" is data obtained by converting voice input into characters.
[0501] The "server" is a central control device that stores and analyzes basic user information and text data.
[0502] "Analysis" is the process of assessing the user's state and emotions based on the received data.
[0503] An "abnormality" is a state that deviates from the user's normal state and may have a negative impact on the user's health or psychology.
[0504] "Notification" refers to the act of the server informing the user of an abnormality or advice.
[0505] The "advice content" is a specific suggestion to encourage healthy behavior based on the user's condition.
[0506] "Emotion analysis" is the process of identifying a user's emotional state based on their voice and text data.
[0507] "Healthy advice" is suggestions and guidance for improving a user's mental and physical health.
[0508] "Feedback" is the act of a user providing opinions or evaluations regarding notifications or advice from the system.
[0509] "System performance" is an indicator of how effectively the system is supporting the user's healthcare.
[0510] A "virtual store" is a virtual environment that provides products and services over the Internet and allows users to have a shopping experience.
[0511] These definitions clarify each element in the claims and make it easier to understand the system configuration and functionality.
[0512] The embodiment of the present invention is to construct a mental and physical healthcare system for a user. The main components of the system include a device, a server, an emotion analysis engine, an advice generation function, and a feedback collection function.
[0513] System configuration
[0514] 1. Device
[0515] Examples of devices include smartphones, smart glasses, headsets, etc. These devices are used to input user information, collect voice data, receive notifications, etc.
[0516] The device provides an interface for inputting basic information about the user and sends that information to the server.
[0517] 2. Server
[0518] The server stores and manages the user's basic information and text data sent from the device.
[0519] The server uses speech recognition technology to convert voice data into text data, and an emotion analysis engine to analyze the user's emotional state.
[0520] The server evaluates the user's state based on the emotion data and generates a notification if an abnormality is detected.
[0521] The server also generates notifications and advice for the user and sends them to the device.
[0522] 3. Sentiment Analysis Engine
[0523] The emotion analysis engine has the ability to analyze user emotions from voice and text data, and this analysis can measure the stress and happiness felt by the user.
[0524] 4. Advice generation function
[0525] The server generates healthy advice based on the emotion data, which is then sent to the device and notified to the user.
[0526] As a specific example, if the emotion analysis engine determines that the user is "nervous," the server generates advice such as, "We recommend that you go to a rest area so that you can relax a bit."
[0527] 5. Feedback collection feature
[0528] The device displays an interface where the user can provide feedback on the advice and notifications. The collected feedback is sent to a server and used to evaluate and tune the system's performance.
[0529] Specific hardware and software used:
[0530] Hardware:
[0531] Smart glasses, smartphone, microphone
[0532] software:
[0533] speech_recognition library: Converts speech to text.
[0534] emotion_recognition module: A proprietary library for identifying emotions from text.
[0535] virtual_store_interface module: A module for interfacing with virtual stores.
[0536] health_advisor module: Generates health advice based on emotion data.
[0537] feedback_module module: Provides feedback collection functionality.
[0538] Examples of specific examples and prompts
[0539] Examples:
[0540] If a user answers "I feel nervous" in a virtual store and the sentiment analysis engine detects stress in the user's tone of voice, the health advisor will offer advice such as "I recommend you go to the rest area so you can relax a bit."
[0541] Example prompt sentence:
[0542] "Please tell me a concrete method to convert Japanese speech data into text, analyze the user's emotions from the text, and generate healthy advice based on those emotions."
[0543] This reduces the stress users feel during their shopping experience in a virtual store, allowing for a healthier and more comfortable experience.
[0544] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0545] Program processing steps
[0546] Step 1:
[0547] User information registration and initial settings
[0548] The user uses the device to input basic information (such as name, age, and gender). The terminal sends this input information to the server, which stores the information in a database. Based on the input, the server stores the user information, generates a confirmation message and sends it to the terminal, which displays the confirmation message to the user. Input: User's basic information. Output: Saving basic information to the database, confirmation message.
[0549] Step 2:
[0550] Daily conversations and data collection
[0551] The device will automatically start up at the set time and ask the user questions using a generated AI voice. The user will respond verbally, and the device will collect the voice data. The collected voice data will be converted into text data using voice recognition technology (speech_recognition library). The device will send the text data to the server, which will store it in a database. Input: User's voice response. Output: Text data will be generated and stored in the database.
[0552] Step 3:
[0553] Emotion Recognition and State Assessment
[0554] The server analyzes the received text data using an emotion analysis engine (emotion_recognition module) to assess the user's emotional state. The server then analyzes this emotion data again along with the text data to generate trends of the user's mental and physical health. Input: Collected text data. Output: Emotion data and health trends.
[0555] Step 4:
[0556] Anomaly detection and notification
[0557] The server runs an algorithm to detect abnormal changes based on the generated trend data. If an abnormality is detected, the server generates an alert message and sends it to the terminal. The terminal notifies the user of the alert. Input: Trend data. Output: Notification of alert message.
[0558] Step 5:
[0559] Generating and delivering health advice
[0560] The server generates healthy advice based on the user's emotional state and health data (health_advisor module). The generated advice is sent to the device, which notifies the user. For example, if the emotion analysis engine determines that the user is feeling stressed, the server will provide advice such as "We recommend that you go to a rest area to relax." Input: Emotional data. Output: Generation and notification of health advice.
[0561] Step 6:
[0562] Feedback collection and system adjustments
[0563] The terminal displays an interface to the user to ask for feedback on the effectiveness of the advice and notifications. The user enters the feedback, and the terminal sends it to the server. The server analyzes the feedback and evaluates and adjusts the system performance. Input: User feedback. Output: System performance adjustment.
[0564] Prompt Sentence Examples
[0565] Example prompt sentence:
[0566] "Please tell me a concrete method to convert Japanese speech data into text, analyze the user's emotions from the text, and generate healthy advice based on those emotions."
[0567] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0568] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0569] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0570] [Second embodiment]
[0571] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0572] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0573] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0574] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0575] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0576] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0577] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0578] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0579] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0580] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0581] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0582] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0583] The system of the present invention supports the mental and physical health care of users and is composed of the following main means:
[0584] User information registration and initial settings
[0585] server:
[0586] Provides an interface for identifying the device (earphones, smartphone, speaker, etc.) used by the user.
[0587] Manages a database for entering and storing basic user information (name, age, gender, etc.).
[0588] Device:
[0589] Displays an initial setup screen and prompts the user to enter basic information.
[0590] The entered basic information is sent to the server.
[0591] Displays to the user the confirmation message received from the server.
[0592] Daily conversations and data collection
[0593] Device:
[0594] At set times, a generative AI voice will ask the user questions (e.g., "What did you eat today?", "How are you feeling?").
[0595] The user's voice input is recorded and converted into text data using voice recognition technology.
[0596] The converted text data is sent to the server.
[0597] server:
[0598] Receives text data sent from the terminal and stores it in a database.
[0599] The received data is analyzed periodically to assess the user's health condition.
[0600] Trend analysis and anomaly detection
[0601] server:
[0602] The stored data is analyzed using machine learning algorithms to identify trends in the user's mental and physical health.
[0603] Based on trend data, an alert is generated if an abnormal change is detected.
[0604] Send the alert content to the device.
[0605] Device:
[0606] Notify the user of alerts received from the server by voice or text.
[0607] Presents the user with options to contact a healthcare provider.
[0608] Collaboration and advice
[0609] server:
[0610] Generate appropriate advice based on the user's health status and send it to the device (e.g., "Drink more water today.").
[0611] If an abnormality is detected, a collaboration API with a pre-configured medical institution or consultation center will be called.
[0612] If the user gives permission, the necessary data will be sent to a medical institution or consultation center.
[0613] Device:
[0614] The user is notified of the advice received from the server and prompted to take action.
[0615] With the user's permission, the necessary data is sent to medical institutions and consultation centers.
[0616] Feedback collection and system adjustments
[0617] Device:
[0618] An interface is displayed to the user requesting feedback on the usefulness of the system's advice and notifications.
[0619] Sends the user-entered feedback to the server.
[0620] server:
[0621] Analyze the feedback received and evaluate the performance of the system.
[0622] If necessary, adjust the questions the AI asks and how it generates advice.
[0623] For example, if a user answers a morning question with "I didn't sleep well last night," the system analyzes this information and determines that the user has reported a continuous lack of sleep. The server detects the abnormality and generates advice to the user, such as "We recommend that you review your nightly routine to improve the quality of your sleep," and notifies the user via their device. In this way, the system can comprehensively monitor the user's health and provide prompt and accurate advice.
[0624] The processing flow will be explained below.
[0625] Specific processing steps of the program
[0626] User information registration and initial settings
[0627] Step 1:
[0628] User: Start the device and open the initial setup screen.
[0629] Step 2:
[0630] Terminal: Presents an interface for the user to enter basic information (e.g., name, age, gender).
[0631] Step 3:
[0632] User: Enter the required basic information.
[0633] Step 4:
[0634] Terminal: Sends the entered basic information to the server.
[0635] Step 5:
[0636] Server: Receives basic information sent from the device and stores it in a database.
[0637] Server: Sends a confirmation message to the device that the device was successfully registered.
[0638] Step 6:
[0639] Terminal: Display a confirmation message to the user.
[0640] Daily conversations and data collection
[0641] Step 7:
[0642] Device: Automatically wakes up at a set time and asks the user questions using a generative AI voice (e.g., "What did you eat today?").
[0643] Step 8:
[0644] User: Answers questions by voice.
[0645] Step 9:
[0646] Terminal: Records the user's voice and converts it into text data using voice recognition technology.
[0647] Step 10:
[0648] Terminal: Sends the converted text data to the server.
[0649] Step 11:
[0650] Server: Receives text data sent from the device and stores it in a database.
[0651] Trend analysis and anomaly detection
[0652] Step 12:
[0653] Server: Analyzes the stored data using machine learning algorithms to generate trends in the user's mental and physical health.
[0654] Step 13:
[0655] Server: Generates alerts based on trend data if abnormal changes are detected.
[0656] Step 14:
[0657] Server: Sends alert content to the device.
[0658] Step 15:
[0659] Terminal: Notifies the user of alerts received from the server via voice or text.
[0660] Step 16:
[0661] Device: Presents the user with options to contact a healthcare provider.
[0662] Collaboration and advice
[0663] Step 17:
[0664] Server: Generates appropriate advice based on the user's health status and sends it to the device (e.g., "Drink more water today.").
[0665] Step 18:
[0666] Server: If an abnormality is detected, it calls a link API with a pre-configured medical institution or consultation center.
[0667] Step 19:
[0668] Server: If the user gives permission, it sends the necessary data to a medical institution or consultation center.
[0669] Step 20:
[0670] Terminal: Notifies the user of the advice received from the server and encourages them to take action.
[0671] Step 21:
[0672] Terminal: With the user's permission, transmits the necessary data to medical institutions and consultation centers.
[0673] Feedback collection and system adjustments
[0674] Step 22:
[0675] Terminal: Display an interface that asks the user for feedback on the usefulness of system advice and notifications.
[0676] Step 23:
[0677] Users: Enter feedback.
[0678] Step 24:
[0679] Terminal: Sends the entered feedback to the server.
[0680] Step 25:
[0681] Server: Analyzes the feedback received and evaluates the performance of the system.
[0682] Server: Adjusts the questions the AI asks and how it generates advice as needed.
[0683] Through the above processing steps, it is possible to comprehensively monitor the user's health condition and provide appropriate advice and support.
[0684] Example 1
[0685] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0686] In modern society, people face the challenge of properly monitoring their own health status and responding promptly in their busy daily lives. In particular, in the areas of mental and physical health, people often delay self-awareness and are unable to identify appropriate countermeasures, creating a need for more efficient and effective healthcare support systems. Furthermore, when an abnormality is detected, it is necessary to quickly collaborate with medical institutions and take appropriate measures, but achieving this requires advanced technology and rapid data analysis.
[0687] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0688] In this invention, the server includes a means for inputting basic information of the user via a terminal and transmitting and storing the information to the server, a means for converting voice input into text data and transmitting the text data to the server, and a means for analyzing the received text data on the server side and evaluating the user's condition. This enables comprehensive monitoring of the user's mental and physical health and prompt cooperation with a medical institution when an abnormality is detected.
[0689] "Basic user information" refers to basic information for identifying an individual, such as the user's name, age, and gender.
[0690] "Terminal" refers to a device used by a user, such as a smartphone, earphones, or speaker.
[0691] "Server" refers to a computer system that manages and processes user information, and stores, analyzes, and notifies various data.
[0692] "Speech input" is the sound data that the user provides to the system through speech synthesis.
[0693] "Text data" refers to data that has been converted from voice input into text information.
[0694] "Analysis" refers to the process of analyzing received data and assessing the user's health status.
[0695] "Evaluation" is the act of determining the user's current health condition based on the analyzed data.
[0696] A "machine learning algorithm" is a method for using large amounts of data to train models to identify trends and make predictions.
[0697] "Trend" refers to the pattern or tendency of fluctuations in a user's health condition over a certain period of time.
[0698] "Abnormal" refers to data or conditions that deviate from the normal range and may have adverse health effects.
[0699] "Notification" is the act of notifying the user of information detected on the server side.
[0700] "Advice" refers to specific advice or suggestions provided based on the user's health status.
[0701] "API" refers to an interface for exchanging data between different systems.
[0702] "Feedback" refers to the evaluations and opinions that users give regarding the services and functions provided by the system.
[0703] A "generative AI model" refers to a model that uses artificial intelligence technology to generate questions and advice.
[0704] The system of the present invention supports users' mental and physical healthcare and includes the following processing steps: registering user information, daily conversations and data collection, trend analysis and anomaly detection, collaboration and advice provision, feedback collection and system adjustment.
[0705] Hardware and software used
[0706] 1. Server: Used for reliable data management and analysis. Specifically, it provides a REST API and uses a MySQL database to store basic user information and collected data. It also uses Python and Pandas for data analysis, and scikit-learn and TensorFlow for implementing machine learning algorithms.
[0707] 2. Terminal: A device through which a user directly inputs information. This includes smartphones, earphones, speakers, etc. It uses the Google Speech-to-Text API to convert voice input into text data and the Google Text-to-Speech API to play the audio.
[0708] System Operation
[0709] The server receives basic information entered by the user via the device and stores it in a database. The device periodically asks the user questions and accepts voice input. The voice is converted into text data and sent to the server. The server analyzes the received text data and evaluates the user's health condition. A machine learning algorithm is used to analyze health trends, and if an abnormality is detected, an alert is quickly generated and sent to the device.
[0710] The server then generates advice based on the user's condition and any abnormalities and sends it to the device. If an abnormality is detected, it calls an API to connect with medical institutions or consultation centers and, with the user's permission, transmits the necessary data. The device then displays the received advice and notifications to the user, encouraging them to take action. It also collects user feedback and uses it to evaluate and adjust the system's performance.
[0711] As a concrete example, consider the case where a user answers a morning question with, "I didn't sleep well last night." The server analyzes this information and confirms that the user has reported a continuous lack of sleep. If the server detects an abnormality, it generates advice such as, "We recommend that you review your nightly routine to improve the quality of your sleep," and notifies the device. The user receives the advice through their device and can contact a medical institution or consultation center if necessary.
[0712] Example prompt sentence:
[0713] "Write a prompt to generate appropriate advice based on the user's health condition, such as when they say they didn't sleep well last night."
[0714] In this way, the system of the present invention can comprehensively monitor the user's health condition and provide prompt and accurate advice.
[0715] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0716] Step 1:
[0717] User information registration and initial settings
[0718] server:
[0719] The server provides an interface (REST API) that identifies the device used by the user. The input is the user's basic information (name, age, gender, etc.) sent from the device, which the server receives and saves in a database. For example, the information is stored in a MySQL database. The output is a confirmation message indicating that the save was successful.
[0720] Device:
[0721] The device displays an initial setup screen and prompts the user to enter basic information. Specifically, the smartphone app displays a form, the user enters the required information, and taps the "Submit" button. The input is information provided by the user, such as name, age, and gender, and the output is the transmission of the information to the server.
[0722] The terminal displays to the user the confirmation message received from the server.
[0723] Step 2:
[0724] Daily conversations and data collection
[0725] Device:
[0726] At a set time, a generative AI voice is used to ask the user questions. For example, questions such as "What did you eat today?" or "How are you feeling?" are played back aloud using the Google Text-to-Speech API. The input is the question sent from the server, and the output is the user's voice input.
[0727] The user responds by voice, and the device records this voice. The recorded voice data is converted into text data using the Google Speech-to-Text API. The input is the user's voice data, and the output is the converted text data.
[0728] The converted text data is sent to the server via HTTPS. The input is the text data, and the output is a transmission confirmation message.
[0729] server:
[0730] The server receives text data sent from the terminal and stores it in a database. The input is the text data, and the output is a confirmation message indicating that the data was saved successfully.
[0731] Step 3:
[0732] Trend analysis and anomaly detection
[0733] server:
[0734] The server analyzes the stored data using machine learning algorithms (e.g., scikit-learn or TensorFlow) to find trends in the user's mental and physical health. The input is a collection of past text data, and the output is the analyzed trend information.
[0735] Generates an alert when an abnormal change is detected. The input is the analyzed trend information, and the output is an alert message indicating the abnormality.
[0736] Sends alert content to the terminal. The input is the alert message, and the output is a transmission confirmation message.
[0737] Device:
[0738] The device notifies the user of the alert received from the server by voice or text, for example by displaying a pop-up on the screen or by reading it aloud. The input is the alert message, and the output is the notified alert.
[0739] Presents the user with options to contact the healthcare provider, allowing them to select "Contact now" or "Later."
[0740] Step 4:
[0741] Collaboration and advice
[0742] server:
[0743] The server generates appropriate advice based on the user's health condition and sends it to the device. For example, the advice may be "Drink more water today." The input is the health condition assessment result, and the output is the generated advice message.
[0744] If an abnormality is detected, a pre-configured API for linking with a medical institution or consultation center is called. For example, a notification is sent using the Twilio API. The input is an alert message indicating the abnormality, and the output is a message confirming that the notification has been sent.
[0745] If the user gives permission, the necessary data is sent to a medical institution or consultation center. The input is the user's consent and the necessary data, and the output is a transmission confirmation message.
[0746] Device:
[0747] The device notifies the user of the advice received from the server and prompts them to take action, for example by displaying it on the screen or reading it aloud. The input is the advice message, and the output is the notified advice.
[0748] With the user's permission, the necessary data will be sent to the medical institution or consultation center. The user taps the "Send data" button.
[0749] Step 5:
[0750] Feedback collection and system adjustments
[0751] Device:
[0752] The terminal displays an interface to ask the user for feedback on the usefulness of the system's advice or notifications. For example, it displays a rating slider or a comment input form. The input is the feedback request, and the output is the feedback data entered by the user.
[0753] Sends the feedback entered by the user to the server. The input is the feedback data, and the output is a transmission confirmation message.
[0754] server:
[0755] The server analyzes the received feedback and evaluates the system's performance. Here, NLP techniques are used for text analysis. The input is the feedback data, and the output is the performance evaluation results.
[0756] If necessary, adjust the questions the AI asks and how it generates advice. For example, retrain the model based on the latest user data. The input is the performance evaluation results and the latest user data, and the output is the adjusted model.
[0757] Example prompt sentence:
[0758] "Write a prompt to generate appropriate advice based on the user's health condition, such as when they say they didn't sleep well last night."
[0759] (Application example 1)
[0760] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0761] Conventional health management systems have had difficulty accurately monitoring a user's daily health status and providing appropriate advice tailored to each individual's situation. Furthermore, they lacked mechanisms for early detection of abnormalities and prompt communication with medical institutions, resulting in delayed responses to maintain optimal user health. To address these issues, there is a need for technology that uses the user's voice input to check daily health status and generate appropriate prompts.
[0762] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0763] In this invention, the server includes means for inputting basic information of a user via a terminal and transmitting it to the server for storage, means for converting voice input into text data and transmitting it to the server, means for analyzing the text data received by the server and evaluating the user's condition, means for detecting abnormalities based on the user's condition and notifying the user, means for generating notifications and advice for the user and transmitting them to the terminal, and means for converting voice input into text data via the terminal, analyzing the text data, and providing appropriate advice using prompt sentences generated by a generative AI model according to the user's condition. This makes it possible to accurately monitor the user's daily health condition and quickly provide appropriate advice according to each individual situation.
[0764] "Basic information" refers to basic personal data such as a user's name, age, and gender.
[0765] A "terminal" is an electronic device used by a user, and refers to devices such as smartphones, earphones, and speakers.
[0766] "Server" refers to a central processing unit for receiving and processing information sent by users.
[0767] "Voice input" refers to the act of a user providing information by voice.
[0768] "Text data" refers to data that has been converted from voice input into text information.
[0769] "Analysis" refers to the act of processing received data and finding meaning and trends.
[0770] "Abnormality" refers to any unusual change or sign in the user's health condition.
[0771] "Notification" refers to the act of informing a user of information.
[0772] "Advice" refers to specific instructions or advice provided based on the user's condition.
[0773] A "collaboration API" refers to an interface for sharing data and functions between different systems and services.
[0774] "Generative AI model" refers to an artificial intelligence model that generates text or advice based on user input.
[0775] A "prompt sentence" is an input sentence to a generative AI model, and refers to trigger text that elicits appropriate answers or advice.
[0776] This invention is a system that supports users' health management and includes a server, a terminal, and a generative AI model. The system collects basic information about the user, analyzes health data based on that information, and provides appropriate advice as needed. Furthermore, if an abnormality is detected, the system also cooperates with medical institutions.
[0777] User information registration and initial settings
[0778] The terminal provides an interface for entering basic information about the user (such as name, age, and gender). Once the user has entered the information, it is sent to the server and stored. At this point, the terminal receives a confirmation message from the server and displays it to the user.
[0779] Daily conversations and data collection
[0780] The device uses the generative AI model to ask the user questions at set times (e.g., "What did you eat today?", "How are you feeling?"). The user responds verbally, and the speech is converted into text data using the Google Cloud Speech-to-Text API. The converted text data is sent to a server and stored in a database.
[0781] Trend analysis and anomaly detection
[0782] The server uses machine learning algorithms to analyze the stored text data. These algorithms can identify trends in the user's mental and physical health. If an abnormality is detected based on the trend data, the server generates an alert and sends it to the device. The device notifies the user of the alert and provides options for contacting a medical institution, if necessary.
[0783] Collaboration and advice
[0784] The server uses a generative AI model to generate appropriate advice based on the user's health status. This generation utilizes the prompt format. For example, if the user answers, "I didn't sleep well last night," the server generates advice such as, "I recommend trying some light stretching before bed to relax."
[0785] Advice is sent to the user via the device. If an abnormality is detected, the server calls the API to connect with the medical institution and, with the user's permission, sends the necessary data.
[0786] Feedback collection and system adjustments
[0787] The device provides an interface that asks the user for feedback on the usefulness of the system's advice and notifications. The user's input feedback is sent to a server for analysis. The server uses this information to evaluate the system's performance and adjusts the questions the generative AI model asks and how it generates advice, if necessary.
[0788] As a specific example of use, if a user answers a morning question with "I didn't sleep well last night," the system analyzes this information and determines that the user has reported poor sleep for several consecutive nights. The server detects the anomaly and generates advice to the user, such as "We recommend that you review your nightly routine to improve the quality of your sleep," and notifies the user via their device.
[0789] Example prompt sentence:
[0790] User Input: I didn't sleep very well last night.
[0791] Health Tip: To help you relax, try some gentle stretching before bed.
[0792] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0793] Step 1:
[0794] The terminal inputs the user's basic information (name, age, gender, etc.) and sends it to the server. The input data is sent from the terminal to the server, and the server saves the data in a database. This registers the user's basic information in the system.
[0795] Step 2:
[0796] At the set time, the device uses the generative AI model to ask the user questions (e.g., "What did you eat today?", "How are you feeling?"). This voice input is input as voice data from the user.
[0797] Step 3:
[0798] The device uses the Google Cloud Speech-to-Text API to convert voice input into text data, which is then sent to a server for analysis.
[0799] Step 4:
[0800] The server analyzes the received text data using machine learning algorithms to assess the user's health status. The output of this analysis step is the assessment of the user's health status.
[0801] Step 5:
[0802] The server detects abnormalities based on the user's health status. To do this, it compares saved data with new input data. If an abnormality is detected, it generates an alert and sends it to the device. The content of the alert indicates that an abnormality exists.
[0803] Step 6:
[0804] The terminal notifies the user of the alert received from the server and presents an option to contact a medical institution if necessary. The output of the notification is that the user receives the alert information.
[0805] Step 7:
[0806] The server uses a generative AI model to generate appropriate advice based on the user's condition. This advice generation uses prompt sentences. For example, if a user inputs "I didn't sleep well last night," the server generates advice such as "I recommend you try some light stretching before bed to relax."
[0807] Step 8:
[0808] The advice is sent from the server to the terminal, and the terminal notifies the user, allowing the user to receive specific health advice.
[0809] Step 9:
[0810] If necessary, the server will call the linkage API and, with the user's permission, send the necessary data to a medical institution or consultation center. This allows the user's health information to be provided to the appropriate institution, enabling a prompt response.
[0811] Step 10:
[0812] The device asks the user for feedback on the usefulness of the system's advice and notifications. The user enters the feedback, which is then sent to the server. The server analyzes the feedback data to evaluate and adjust the system's performance. This adjustment improves the accuracy of the generative AI model and the overall system performance.
[0813] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0814] The system of the present invention supports the user's mental and physical healthcare and is composed of the following main means. In addition, by combining it with an emotion engine that recognizes the user's emotions, it provides more advanced care.
[0815] User information registration and initial settings
[0816] server:
[0817] Provides an interface for identifying the device (earphones, smartphone, speaker, etc.) used by the user.
[0818] Manages a database for entering and storing basic user information (name, age, gender, etc.).
[0819] Device:
[0820] Displays an initial setup screen and prompts the user to enter basic information.
[0821] The entered basic information is sent to the server.
[0822] Displays to the user the confirmation message received from the server.
[0823] Daily conversations and data collection
[0824] Device:
[0825] It automatically starts at the set time and asks the user questions using a generative AI voice (e.g., "What did you eat today?", "How are you feeling?").
[0826] The user's voice input is recorded and converted into text data using voice recognition technology.
[0827] The converted text data is sent to the server.
[0828] server:
[0829] Receives text data sent from the terminal and stores it in a database.
[0830] Emotion Recognition and State Assessment
[0831] Emotion engine (can run on device or server):
[0832] The user's voice tone, pitch, and rate are analyzed to identify emotions.
[0833] The identified emotion data is transmitted to a server.
[0834] server:
[0835] The emotion data received from the emotion engine is analyzed together with the text data to comprehensively evaluate the user's state.
[0836] The stored data is analyzed using machine learning algorithms to generate trends in the user's mental and physical health.
[0837] Based on trend data, an alert is generated if an abnormal change is detected.
[0838] Send the alert content to the device.
[0839] Device:
[0840] Notify the user of alerts received from the server by voice or text.
[0841] Presents the user with options to contact a healthcare provider.
[0842] Collaboration and advice
[0843] server:
[0844] Generate appropriate advice based on the user's health status and emotional data and send it to the device (e.g., "We recommend listening to music to relax today.").
[0845] If an abnormality is detected, a collaboration API with a pre-configured medical institution or consultation center will be called.
[0846] If the user gives permission, the necessary data will be sent to a medical institution or consultation center.
[0847] Device:
[0848] The user is notified of the advice received from the server and prompted to take action.
[0849] With the user's permission, the necessary data is sent to medical institutions and consultation centers.
[0850] Feedback collection and system adjustments
[0851] Device:
[0852] An interface is displayed to the user requesting feedback on the usefulness of the system's advice and notifications.
[0853] Sends the user-entered feedback to the server.
[0854] server:
[0855] Analyze the feedback received and evaluate the performance of the system.
[0856] If necessary, adjust the questions the AI asks and how it generates advice.
[0857] For example, if a user responds to a morning question with "I didn't sleep well last night," and the emotion engine detects a depressed tone, the system analyzes this information and determines that the user has continued to have poor sleep. The server detects the abnormality and generates advice for the user, such as "To improve the quality of your sleep, we recommend listening to relaxing music tonight," and notifies the user via their device. In this way, the system comprehensively monitors the user's health status, including emotion recognition, and can provide prompt and accurate advice.
[0858] The processing flow will be explained below.
[0859] Specific processing steps of the program
[0860] User information registration and initial settings
[0861] Step 1:
[0862] User: Start the device and open the initial setup screen.
[0863] Step 2:
[0864] Terminal: Presents an interface for the user to enter basic information (e.g., name, age, gender).
[0865] Step 3:
[0866] User: Enter the required basic information.
[0867] Step 4:
[0868] Terminal: Sends the entered basic information to the server.
[0869] Step 5:
[0870] Server: Receives basic information sent from the device and stores it in a database.
[0871] Server: Sends a confirmation message to the device that the device was successfully registered.
[0872] Step 6:
[0873] Terminal: Display a confirmation message to the user.
[0874] Daily conversations and data collection
[0875] Step 7:
[0876] Device: Automatically wakes up at a set time and asks questions to the user using a generative AI voice (e.g., "What did you eat today?", "How are you feeling?").
[0877] Step 8:
[0878] User: Answers questions by voice.
[0879] Step 9:
[0880] Terminal: Records the user's voice and converts it into text data using voice recognition technology.
[0881] Step 10:
[0882] Terminal: Sends the converted text data to the server.
[0883] Step 11:
[0884] Server: Receives text data sent from the device and stores it in a database.
[0885] Emotion Recognition and State Assessment
[0886] Step 12:
[0887] Terminal or Server (Emotion Engine): Analyzes the user's voice tone, pitch, and rate to identify emotions.
[0888] Step 13:
[0889] Emotion engine (terminal or server): Sends identified emotion data to the server.
[0890] Step 14:
[0891] Server: Analyzes the emotion data received from the emotion engine along with the text data, and makes a comprehensive assessment of the user's state.
[0892] Step 15:
[0893] Server: Analyzes the stored data using machine learning algorithms to generate trends in the user's mental and physical health.
[0894] Step 16:
[0895] Server: Generates alerts based on trend data if abnormal changes are detected.
[0896] Step 17:
[0897] Server: Sends alert content to the device.
[0898] Step 18:
[0899] Terminal: Notifies the user of alerts received from the server via voice or text.
[0900] Step 19:
[0901] Device: Presents the user with options to contact a healthcare provider.
[0902] Collaboration and advice
[0903] Step 20:
[0904] Server: Generates appropriate advice based on the user's health status and emotional data and sends it to the device (e.g., "We recommend listening to music to relax today.").
[0905] Step 21:
[0906] Server: If an abnormality is detected, it calls a link API with a pre-configured medical institution or consultation center.
[0907] Step 22:
[0908] Server: If the user gives permission, it sends the necessary data to a medical institution or consultation center.
[0909] Step 23:
[0910] Terminal: Notifies the user of the advice received from the server and encourages them to take action.
[0911] Step 24:
[0912] Terminal: With the user's permission, transmits the necessary data to medical institutions and consultation centers.
[0913] Feedback collection and system adjustments
[0914] Step 25:
[0915] Terminal: Display an interface that asks the user for feedback on the usefulness of system advice and notifications.
[0916] Step 26:
[0917] Users: Enter feedback.
[0918] Step 27:
[0919] Terminal: Sends the entered feedback to the server.
[0920] Step 28:
[0921] Server: Analyzes the feedback received and evaluates the performance of the system.
[0922] Server: Adjusts the questions the AI asks and how it generates advice as needed.
[0923] The above processing steps enable comprehensive monitoring of a user's health status and provide appropriate advice and support, including emotion recognition. For example, if a user responds to a morning question with "I didn't sleep well last night," and the emotion engine detects a depressed tone, the system analyzes this information, considers the user's mental and physical health trends, and generates advice such as "To improve the quality of your sleep, we recommend listening to relaxing music tonight," and notifies the user via the device. In this way, this system realizes comprehensive health management that incorporates emotion recognition.
[0924] Example 2
[0925] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0926] In recent years, stress and anxiety have increased in modern society, making mental and physical healthcare increasingly important. However, conventional healthcare systems lack the ability to accurately identify a user's emotional state and provide appropriate advice. As a result, it is difficult to detect abnormalities in a user's state early and take appropriate measures.
[0927] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0928] In this invention, the server includes means for inputting basic information of a user via a terminal and transmitting the information to the server for storage, means for converting voice input into text data and transmitting the text data to the server, means for analyzing the text data received on the server side and evaluating the user's condition, means for analyzing the user's voice tone, pitch, and speed using an emotion engine and identifying emotions, means for transmitting the identified emotion data to the server and analyzing it together with the text data, means for detecting abnormalities based on the user's condition and notifying the user, and means for generating notifications and advice for the user and transmitting the generated notifications and advice to the terminal. This makes it possible to comprehensively monitor the user's emotional state and health condition, detect abnormalities early, and quickly provide appropriate advice and responses.
[0929] "Means for inputting basic user information via a terminal and sending it to a server for storage" refers to the process and infrastructure that allows a user to input basic personal information such as name, age, and gender using the terminal interface, and then transmit the input data to a server and store it in a database.
[0930] "Means for converting voice input into text data and transmitting it to a server" refers to technology and devices for recording a user's voice input, converting it into text data using voice recognition technology, and transmitting the text data to a server.
[0931] The "means for analyzing text data received on the server side and evaluating the user's condition" refers to a system and analysis algorithm that analyzes and processes text data received by the server from the terminal and evaluates the user's mental and physical health condition.
[0932] "Means for identifying emotions by analyzing a user's voice tone, pitch, and rate using an emotion engine" refers to a technology and method for using an emotion engine to analyze characteristics of a user's speech, such as voice tone, pitch, and rate, to identify the user's emotional state.
[0933] The "means for transmitting identified emotion data to a server for analysis together with text data" refers to a process and system for transmitting emotion data identified by the emotion engine to a server for analysis of the data together with text data.
[0934] "Means for detecting abnormalities based on the user's condition and notifying the user" refers to a system and method that determines whether there are any abnormalities in the user's health or emotional state based on data analyzed by the server and notifies the user of the results.
[0935] "Means for generating notifications and advice for the user and sending them to the terminal" refers to the technology and platform for generating appropriate notifications and advice based on the results of analyzing the user's condition and sending them to the terminal.
[0936] "Means for calling APIs to link with medical institutions and consultation centers" refers to systems and technologies for calling application programming interfaces (APIs) to link with pre-set medical institutions and consultation centers when an abnormality is detected.
[0937] "Means for transmitting necessary data to a medical institution or consultation center with the user's permission" refers to the process and system for transmitting necessary health data and emotion data to a designated medical institution or consultation center with the user's consent.
[0938] "Means for collecting user feedback, analyzing it on the server side, and evaluating and adjusting system performance" refers to a mechanism for collecting feedback from users regarding system usage, analyzing it on the server to evaluate system performance, and adjusting the system's functions and service content as necessary.
[0939] This invention is a system for supporting users' mental and physical healthcare, which provides advice and notifications according to the situation by collecting and analyzing the user's basic information, daily activity data, and emotional data. To achieve this, a server, a terminal, and an emotion engine work in cooperation with each other.
[0940] User information registration and initial settings
[0941] Server: The server provides an interface to identify the device used by the user (earphones, smartphone, speaker, etc.). When the user connects to the device, the server collects the information and prepares it to be stored in a database. The server also manages a database for entering and storing basic information about the user (name, age, gender, etc.).
[0942] Terminal: The terminal displays an initial setup screen and prompts the user to enter basic information. After the user enters the information, the terminal sends the information to the server. The server stores the received information in a database and sends a confirmation message to the terminal. The terminal displays a confirmation message to the user, informing them that the initial setup is complete.
[0943] Daily conversations and data collection
[0944] Device: The device automatically wakes up at the set time and asks the user questions using a generative AI voice (e.g., "What did you eat today?", "How are you feeling?"). When the user responds, the device records the voice and converts it into text data using voice recognition technology (e.g., Google Speech-to-Text API).
[0945] Server: The device sends the converted text data to the server, which stores it in a database for further analysis.
[0946] Emotion Recognition and State Assessment
[0947] Emotion engine (can run on the device or on the server): The emotion engine analyzes the user's voice tone, pitch, and rate to identify emotions. The identified emotion data is sent to the server.
[0948] Server: The server analyzes the emotion data received from the emotion engine along with the text data to comprehensively evaluate the user's condition. It uses machine learning algorithms (e.g., TensorFlow, PyTorch) to generate trends in the user's mental and physical health. If any abnormal changes are detected based on the trends, it generates an alert and sends it to the device.
[0949] Collaboration and advice
[0950] Server: The server generates appropriate advice based on the user's health condition and emotional data. For example, if the user's recent sleep data is poor, the server generates advice such as "We recommend listening to relaxing music" and sends it to the device. If an abnormality is detected, the server calls a link API with a pre-set medical institution or consultation center. If the user gives permission, the server sends the necessary data to the medical institution or consultation center.
[0951] Device: The device notifies the user of the advice received from the server. It displays the advice in audio or text and prompts the user to take specific action (e.g., displaying a button to play music). With the user's permission, it also sends the necessary data to medical institutions or consultation centers.
[0952] Feedback collection and system adjustments
[0953] Terminal: The terminal provides an interface to collect user feedback on advice and notifications from the system. When the user enters feedback, it is sent to the server.
[0954] Server: The server analyzes the feedback received from users, evaluates the system's performance, and adjusts the questions the AI asks and how it generates advice to improve the system, if necessary.
[0955] Specific examples
[0956] Example 1: Morning Questions
[0957] The device starts up every morning at 7:00 and asks the user, "Good morning. How was the quality of your sleep last night?" If the user replies, "I didn't sleep well," the device records the voice, converts it into text data, and sends it to the server.
[0958] Example 2: Emotion Identification
[0959] The emotion engine analyzes the user's voice and identifies the emotion "depressed" from the low tone and slow rate. This data is sent to a server and analyzed together with the text data.
[0960] Example 3: Alert and Advice Notifications
[0961] The server continuously checks whether the quality of sleep is poor and detects any abnormalities. It generates advice such as "To improve the quality of your sleep, we recommend listening to relaxing music tonight" and sends it to the device. The device notifies the user of the advice and offers the option to play music.
[0962] Prompt Sentence Examples
[0963] "How was your sleep quality last night?"
[0964] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0965] Step 1:
[0966] The user enters basic information into the terminal and sends it to the server.
[0967] Specific operation: The device displays an initial setup screen and asks the user to enter basic information such as name, age, gender, etc. After the user enters the information, the device sends it to the server.
[0968] Input: Basic information entered by the user, such as name, age, and gender.
[0969] Output: Basic user information data sent from the device to the server.
[0970] Data processing: The server receives these data and stores them in a database.
[0971] Step 2:
[0972] The server stores the basic information in a database and sends a confirmation message to the device.
[0973] Specific operation: The server saves the received basic information in the database, generates a confirmation message for the save completion and sends it to the terminal, which then displays the confirmation message to the user.
[0974] Input: User basic information data.
[0975] Output: Basic information stored in the database and a confirmation message displayed to the user.
[0976] Data processing: The server stores the basic information in a database as structured data.
[0977] Step 3:
[0978] The device will start up at the set time and ask the user questions using a generated AI voice.
[0979] Specific operation: The device automatically starts up at a pre-set time and asks the user questions using a generated AI voice, such as, "How was the quality of your sleep last night?"
[0980] Input: None (automatic start by scheduled trigger).
[0981] Output: A spoken question to the user.
[0982] Data processing: Generate appropriate question voice data using a generative AI model.
[0983] Step 4:
[0984] The user responds by voice, and the terminal records the voice and converts it into text data.
[0985] What it does: When a user responds to a question, the device records the audio and then uses voice recognition technology to convert the audio into text data (e.g., using the Google Speech-to-Text API).
[0986] Input: The user's spoken response.
[0987] Output: User answers converted to text data.
[0988] Data processing: Converting audio data into text format.
[0989] Step 5:
[0990] The terminal transmits the text data to the server.
[0991] Specific operation: The converted text data is sent from the terminal to the server.
[0992] Input: Text data.
[0993] Output: The text data sent to the server.
[0994] Data processing: Encodes text data into transmission packets.
[0995] Step 6:
[0996] The server receives the text data and stores it in a database.
[0997] Specific operation: The server receives text data from the terminal and stores it in a database.
[0998] Input: Text data.
[0999] Output: Text data stored in a database.
[1000] Data processing: Structure the text data and record it in a database.
[1001] Step 7:
[1002] The emotion engine (running on the device or server) analyzes the voice data and identifies emotions.
[1003] How it works: The emotion engine analyzes the user's voice tone, pitch, and rate to identify emotions. The emotion data is sent to the server and analyzed along with the text data.
[1004] Input: User's voice data.
[1005] Output: Emotion data.
[1006] Data processing: Analyze voice characteristics such as tone, pitch, and speed to identify emotions.
[1007] Step 8:
[1008] The server analyzes the emotional data and text data to evaluate the user's state.
[1009] How it works: The server uses emotion data and text data to analyze the user's mental and physical state. The analysis uses machine learning algorithms (e.g., TensorFlow, PyTorch).
[1010] Input: Emotion data, text data.
[1011] Output: User status assessment results.
[1012] Data processing: Integrate both sets of data and perform statistical analysis using machine learning models.
[1013] Step 9:
[1014] The server detects anomalies based on the user's status and generates an alert.
[1015] Specific operation: The server detects anomalies based on the analysis results, generates an alert if an anomaly is detected, and sends the alert content to the user's device.
[1016] Input: User condition assessment result.
[1017] Output: Alert data.
[1018] Data processing: Conditions are judged based on the evaluation results, and alert data is generated and sent.
[1019] Step 10:
[1020] The device will notify the user of the alert and provide advice if necessary.
[1021] Specific behavior: The device notifies the user of the alert received from the server via voice or text, and generates and provides additional advice to the user if necessary.
[1022] Input: Alert data.
[1023] Output: Notification and advice to the user.
[1024] Data processing: Received alerts are notified to the user in an appropriate format and advice is generated.
[1025] Step 11:
[1026] The server and terminal work together to contact medical institutions and consultation centers.
[1027] Specific operation: If an abnormality is detected, the server calls the API for linking with the configured medical institution or consultation center. If the user gives permission, the device sends the necessary data to the medical institution or consultation center.
[1028] Input: User consent, anomaly detection results.
[1029] Output: Sending data to medical institutions and consultation centers.
[1030] Data processing: Formatting and sending the necessary data to medical institutions and consultation centers.
[1031] Step 12:
[1032] Gathering feedback and tuning system performance.
[1033] What it does: The device collects feedback from users and sends that data to the server, which analyzes the feedback, evaluates system performance, and makes adjustments as needed.
[1034] Input: User feedback data.
[1035] Output: Evaluation results and adjustments.
[1036] Data processing: Analyze feedback data to evaluate and adjust system performance.
[1037] (Application example 2)
[1038] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1039] Currently, there is no system in virtual stores that adequately supports users' mental and physical healthcare. As a result, users may feel stressed and their health may deteriorate. In particular, the shopping experience in a virtual store has a significant impact on users' psychological state, so there is a need to introduce emotion analysis and provide appropriate health advice.
[1040] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1041] In this invention, the server includes means for inputting basic information of a user via a device and transmitting it to the server for storage, means for converting voice input into text data and transmitting it to the server, means for analyzing the text data received by the server and evaluating the user's condition, means for detecting abnormalities based on the user's condition and notifying the user, means for generating notifications and advice for the user and transmitting them to the device, means for analyzing the user's emotions and generating healthy advice based on the emotions and notifying the user, and means for collecting feedback, analyzing it on the server, and evaluating and adjusting system performance. This reduces the stress felt by users during their shopping experience in a virtual store, enabling a healthy and comfortable experience.
[1042] "Basic user information" is information that can be used to identify an individual, such as the user's name, age, and gender.
[1043] A "device" is a device that a user operates to input and receive information, such as a smartphone, smart glasses, or headset.
[1044] "Voice input" is voice information that a user utters into a device.
[1045] "Text data" is data obtained by converting voice input into characters.
[1046] The "server" is a central control device that stores and analyzes basic user information and text data.
[1047] "Analysis" is the process of assessing the user's state and emotions based on the received data.
[1048] An "abnormality" is a state that deviates from the user's normal state and may have a negative impact on the user's health or psychology.
[1049] "Notification" refers to the act of the server informing the user of an abnormality or advice.
[1050] The "advice content" is a specific suggestion to encourage healthy behavior based on the user's condition.
[1051] "Emotion analysis" is the process of identifying a user's emotional state based on their voice and text data.
[1052] "Healthy advice" is suggestions and guidance for improving a user's mental and physical health.
[1053] "Feedback" is the act of a user providing opinions or evaluations regarding notifications or advice from the system.
[1054] "System performance" is an indicator of how effectively the system is supporting the user's healthcare.
[1055] A "virtual store" is a virtual environment that provides products and services over the Internet and allows users to have a shopping experience.
[1056] These definitions clarify each element in the claims and make it easier to understand the system configuration and functionality.
[1057] The embodiment of the present invention is to construct a mental and physical healthcare system for a user. The main components of the system include a device, a server, an emotion analysis engine, an advice generation function, and a feedback collection function.
[1058] System configuration
[1059] 1. Device
[1060] Examples of devices include smartphones, smart glasses, headsets, etc. These devices are used to input user information, collect voice data, receive notifications, etc.
[1061] The device provides an interface for inputting basic information about the user and sends that information to the server.
[1062] 2. Server
[1063] The server stores and manages the user's basic information and text data sent from the device.
[1064] The server uses speech recognition technology to convert voice data into text data, and an emotion analysis engine to analyze the user's emotional state.
[1065] The server evaluates the user's state based on the emotion data and generates a notification if an abnormality is detected.
[1066] The server also generates notifications and advice for the user and sends them to the device.
[1067] 3. Sentiment Analysis Engine
[1068] The emotion analysis engine has the ability to analyze user emotions from voice and text data, and this analysis can measure the stress and happiness felt by the user.
[1069] 4. Advice Generation Function
[1070] The server generates healthy advice based on the emotion data, which is then sent to the device and notified to the user.
[1071] As a specific example, if the emotion analysis engine determines that the user is "nervous," the server generates advice such as, "We recommend that you go to a rest area so that you can relax a little."
[1072] 5. Feedback collection feature
[1073] The device displays an interface where the user can provide feedback on the advice and notifications. The collected feedback is sent to a server and used to evaluate and tune the system's performance.
[1074] Specific hardware and software used:
[1075] Hardware:
[1076] Smart glasses, smartphone, microphone
[1077] software:
[1078] speech_recognition library: Converts speech to text.
[1079] emotion_recognition module: A proprietary library for identifying emotions from text.
[1080] virtual_store_interface module: A module for interfacing with virtual stores.
[1081] health_advisor module: Generates health advice based on emotion data.
[1082] feedback_module module: Provides feedback collection functionality.
[1083] Examples of concrete examples and prompts
[1084] Examples:
[1085] If a user answers "I feel nervous" in a virtual store and the sentiment analysis engine detects stress in the user's tone of voice, the health advisor will offer advice such as "I recommend you go to the rest area so you can relax a bit."
[1086] Example prompt sentence:
[1087] "Please tell me a concrete method to convert Japanese speech data into text, analyze the user's emotions from the text, and generate healthy advice based on those emotions."
[1088] This reduces the stress users feel during their shopping experience in a virtual store, allowing for a healthier and more comfortable experience.
[1089] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1090] Program processing steps
[1091] Step 1:
[1092] User information registration and initial settings
[1093] The user uses the device to input basic information (such as name, age, and gender). The terminal sends this input information to the server, which stores the information in a database. Based on the input, the server stores the user information, generates a confirmation message and sends it to the terminal, which displays the confirmation message to the user. Input: User's basic information. Output: Saving basic information to the database, confirmation message.
[1094] Step 2:
[1095] Daily conversations and data collection
[1096] The device will automatically start up at the set time and ask the user questions using a generated AI voice. The user will respond verbally, and the device will collect the voice data. The collected voice data will be converted into text data using voice recognition technology (speech_recognition library). The device will send the text data to the server, which will store it in a database. Input: User's voice response. Output: Text data will be generated and stored in the database.
[1097] Step 3:
[1098] Emotion Recognition and State Assessment
[1099] The server analyzes the received text data using an emotion analysis engine (emotion_recognition module) to assess the user's emotional state. The server then analyzes this emotion data again along with the text data to generate trends of the user's mental and physical health. Input: Collected text data. Output: Emotion data and health trends.
[1100] Step 4:
[1101] Anomaly detection and notification
[1102] The server runs an algorithm to detect abnormal changes based on the generated trend data. If an abnormality is detected, the server generates an alert message and sends it to the terminal. The terminal notifies the user of the alert. Input: Trend data. Output: Notification of alert message.
[1103] Step 5:
[1104] Generating and delivering health advice
[1105] The server generates healthy advice based on the user's emotional state and health data (health_advisor module). The generated advice is sent to the device, which notifies the user. For example, if the emotion analysis engine determines that the user is feeling stressed, the server will provide advice such as "We recommend that you go to a rest area to relax." Input: Emotional data. Output: Generation and notification of health advice.
[1106] Step 6:
[1107] Feedback collection and system adjustments
[1108] The terminal displays an interface to the user to ask for feedback on the effectiveness of the advice and notifications. The user enters the feedback, and the terminal sends it to the server. The server analyzes the feedback and evaluates and adjusts the system performance. Input: User feedback. Output: System performance adjustment.
[1109] Prompt Sentence Examples
[1110] Example prompt sentence:
[1111] "Please tell me a concrete method to convert Japanese speech data into text, analyze the user's emotions from the text, and generate healthy advice based on those emotions."
[1112] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1113] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1114] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1115] [Third embodiment]
[1116] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1117] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1118] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1119] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1120] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1121] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1122] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1123] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1124] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1125] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1126] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1127] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1128] The system of the present invention supports the mental and physical health care of users and is composed of the following main means:
[1129] User information registration and initial settings
[1130] server:
[1131] Provides an interface for identifying the device (earphones, smartphone, speaker, etc.) used by the user.
[1132] Manages a database for entering and storing basic user information (name, age, gender, etc.).
[1133] Device:
[1134] Displays an initial setup screen and prompts the user to enter basic information.
[1135] The entered basic information is sent to the server.
[1136] Displays to the user a confirmation message received from the server.
[1137] Daily conversations and data collection
[1138] Device:
[1139] At set times, a generative AI voice will ask the user questions (e.g., "What did you eat today?", "How are you feeling?").
[1140] The user's voice input is recorded and converted into text data using voice recognition technology.
[1141] The converted text data is sent to the server.
[1142] server:
[1143] Receives text data sent from the terminal and stores it in a database.
[1144] The received data is analyzed periodically to assess the user's health condition.
[1145] Trend analysis and anomaly detection
[1146] server:
[1147] The stored data is analyzed using machine learning algorithms to identify trends in the user's mental and physical health.
[1148] Based on trend data, an alert is generated if an abnormal change is detected.
[1149] Send the alert content to the device.
[1150] Device:
[1151] Notify the user of alerts received from the server by voice or text.
[1152] Presents the user with options to contact a healthcare provider.
[1153] Collaboration and advice
[1154] server:
[1155] Generate appropriate advice based on the user's health status and send it to the device (e.g., "Drink more water today.").
[1156] If an abnormality is detected, a collaboration API with a pre-configured medical institution or consultation center will be called.
[1157] If the user gives permission, the necessary data will be sent to a medical institution or consultation center.
[1158] Device:
[1159] The user is notified of the advice received from the server and prompted to take action.
[1160] With the user's permission, the necessary data is sent to medical institutions and consultation centers.
[1161] Feedback collection and system adjustments
[1162] Device:
[1163] An interface is displayed to the user requesting feedback on the usefulness of the system's advice and notifications.
[1164] Sends the user-entered feedback to the server.
[1165] server:
[1166] Analyze the feedback received and evaluate the performance of the system.
[1167] If necessary, adjust the questions the AI asks and how it generates advice.
[1168] For example, if a user answers a morning question with "I didn't sleep well last night," the system analyzes this information and determines that the user has reported a continuous lack of sleep. The server detects the abnormality and generates advice for the user, such as "We recommend that you review your nightly routine to improve the quality of your sleep," and notifies the user via their device. In this way, the system can comprehensively monitor the user's health status and provide prompt and accurate advice.
[1169] The processing flow will be explained below.
[1170] Specific processing steps of the program
[1171] User information registration and initial settings
[1172] Step 1:
[1173] User: Start the device and open the initial setup screen.
[1174] Step 2:
[1175] Terminal: Presents an interface for the user to enter basic information (e.g., name, age, gender).
[1176] Step 3:
[1177] User: Enter the required basic information.
[1178] Step 4:
[1179] Terminal: Sends the entered basic information to the server.
[1180] Step 5:
[1181] Server: Receives basic information sent from the device and stores it in a database.
[1182] Server: Sends a confirmation message to the device that the device was successfully registered.
[1183] Step 6:
[1184] Terminal: Display a confirmation message to the user.
[1185] Daily conversations and data collection
[1186] Step 7:
[1187] Device: Automatically wakes up at a set time and asks the user questions using a generative AI voice (e.g., "What did you eat today?").
[1188] Step 8:
[1189] User: Answers questions by voice.
[1190] Step 9:
[1191] Terminal: Records the user's voice and converts it into text data using voice recognition technology.
[1192] Step 10:
[1193] Terminal: Sends the converted text data to the server.
[1194] Step 11:
[1195] Server: Receives text data sent from the device and stores it in a database.
[1196] Trend analysis and anomaly detection
[1197] Step 12:
[1198] Server: Analyzes the stored data using machine learning algorithms to generate trends in the user's mental and physical health.
[1199] Step 13:
[1200] Server: Generates alerts based on trend data if abnormal changes are detected.
[1201] Step 14:
[1202] Server: Sends alert content to the device.
[1203] Step 15:
[1204] Terminal: Notifies the user of alerts received from the server via voice or text.
[1205] Step 16:
[1206] Device: Presents the user with options to contact a healthcare provider.
[1207] Collaboration and advice
[1208] Step 17:
[1209] Server: Generates appropriate advice based on the user's health status and sends it to the device (e.g., "Drink more water today.").
[1210] Step 18:
[1211] Server: If an abnormality is detected, it calls a link API with a pre-configured medical institution or consultation center.
[1212] Step 19:
[1213] Server: If the user gives permission, it sends the necessary data to a medical institution or consultation center.
[1214] Step 20:
[1215] Terminal: Notifies the user of the advice received from the server and encourages them to take action.
[1216] Step 21:
[1217] Terminal: With the user's permission, transmits the necessary data to medical institutions and consultation centers.
[1218] Feedback collection and system adjustments
[1219] Step 22:
[1220] Terminal: Display an interface that asks the user for feedback on the usefulness of system advice and notifications.
[1221] Step 23:
[1222] Users: Enter feedback.
[1223] Step 24:
[1224] Terminal: Sends the entered feedback to the server.
[1225] Step 25:
[1226] Server: Analyzes the feedback received and evaluates the performance of the system.
[1227] Server: Adjusts the questions the AI asks and how it generates advice as needed.
[1228] Through the above processing steps, it is possible to comprehensively monitor the user's health condition and provide appropriate advice and support.
[1229] Example 1
[1230] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1231] In modern society, people face the challenge of properly monitoring their own health status and responding promptly in their busy daily lives. In particular, in the areas of mental and physical health, people often delay self-awareness and are unable to identify appropriate countermeasures, creating a need for more efficient and effective healthcare support systems. Furthermore, when an abnormality is detected, it is necessary to quickly collaborate with medical institutions and take appropriate measures, but achieving this requires advanced technology and rapid data analysis.
[1232] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1233] In this invention, the server includes a means for inputting basic information of the user via a terminal and transmitting and storing the information to the server, a means for converting voice input into text data and transmitting the text data to the server, and a means for analyzing the received text data on the server side and evaluating the user's condition. This enables comprehensive monitoring of the user's mental and physical health and prompt cooperation with a medical institution when an abnormality is detected.
[1234] "Basic user information" refers to basic information for identifying an individual, such as the user's name, age, and gender.
[1235] "Terminal" refers to a device used by a user, such as a smartphone, earphones, or speaker.
[1236] "Server" refers to a computer system that manages and processes user information, and stores, analyzes, and notifies various data.
[1237] "Speech input" is the sound data that the user provides to the system through speech synthesis.
[1238] "Text data" refers to data that has been converted from voice input into text information.
[1239] "Analysis" refers to the process of analyzing received data and assessing the user's health status.
[1240] "Evaluation" is the act of determining the user's current health condition based on the analyzed data.
[1241] A "machine learning algorithm" is a method for using large amounts of data to train models to identify trends and make predictions.
[1242] "Trend" refers to the pattern or tendency of fluctuations in a user's health condition over a certain period of time.
[1243] "Abnormal" refers to data or conditions that deviate from the normal range and may have adverse health effects.
[1244] "Notification" is the act of notifying the user of information detected on the server side.
[1245] "Advice" refers to specific advice or suggestions provided based on the user's health status.
[1246] "API" refers to an interface for exchanging data between different systems.
[1247] "Feedback" refers to the evaluations and opinions that users give regarding the services and functions provided by the system.
[1248] A "generative AI model" refers to a model that uses artificial intelligence technology to generate questions and advice.
[1249] The system of the present invention supports users' mental and physical healthcare and includes the following processing steps: registering user information, daily conversations and data collection, trend analysis and anomaly detection, collaboration and advice provision, feedback collection and system adjustment.
[1250] Hardware and software used
[1251] 1. Server: Used for reliable data management and analysis. Specifically, it provides a REST API and uses a MySQL database to store basic user information and collected data. It also uses Python and Pandas for data analysis, and scikit-learn and TensorFlow for implementing machine learning algorithms.
[1252] 2. Terminal: A device through which a user directly inputs information. This includes smartphones, earphones, speakers, etc. It uses the Google Speech-to-Text API to convert voice input into text data and the Google Text-to-Speech API to play the audio.
[1253] System Operation
[1254] The server receives basic information entered by the user via the device and stores it in a database. The device periodically asks the user questions and accepts voice input. The voice is converted into text data and sent to the server. The server analyzes the received text data and evaluates the user's health condition. A machine learning algorithm is used to analyze health trends, and if an abnormality is detected, an alert is quickly generated and sent to the device.
[1255] The server then generates advice based on the user's condition and any abnormalities and sends it to the device. If an abnormality is detected, it calls an API to connect with medical institutions or consultation centers and, with the user's permission, transmits the necessary data. The device then displays the received advice and notifications to the user, encouraging them to take action. It also collects user feedback and uses it to evaluate and adjust the system's performance.
[1256] As a concrete example, consider the case where a user answers a morning question with, "I didn't sleep well last night." The server analyzes this information and confirms that the user has reported a continuous lack of sleep. If the server detects an abnormality, it generates advice such as, "We recommend that you review your nightly routine to improve the quality of your sleep," and notifies the device. The user receives the advice through their device and can contact a medical institution or consultation center if necessary.
[1257] Example prompt sentence:
[1258] "Write a prompt to generate appropriate advice based on the user's health condition, such as when they say they didn't sleep well last night."
[1259] In this way, the system of the present invention can comprehensively monitor the user's health condition and provide prompt and accurate advice.
[1260] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1261] Step 1:
[1262] User information registration and initial settings
[1263] server:
[1264] The server provides an interface (REST API) that identifies the device used by the user. The input is the user's basic information (name, age, gender, etc.) sent from the device, which the server receives and saves in a database. For example, the information is stored in a MySQL database. The output is a confirmation message indicating that the save was successful.
[1265] Device:
[1266] The device displays an initial setup screen and prompts the user to enter basic information. Specifically, the smartphone app displays a form, the user enters the required information, and taps the "Submit" button. The input is information provided by the user, such as name, age, and gender, and the output is the transmission of the information to the server.
[1267] The terminal displays to the user the confirmation message received from the server.
[1268] Step 2:
[1269] Daily conversations and data collection
[1270] Device:
[1271] At a set time, a generative AI voice is used to ask the user questions. For example, questions such as "What did you eat today?" or "How are you feeling?" are played back aloud using the Google Text-to-Speech API. The input is the question sent from the server, and the output is the user's voice input.
[1272] The user responds by voice, and the device records this voice. The recorded voice data is converted into text data using the Google Speech-to-Text API. The input is the user's voice data, and the output is the converted text data.
[1273] The converted text data is sent to the server via HTTPS. The input is the text data, and the output is a transmission confirmation message.
[1274] server:
[1275] The server receives text data sent from the terminal and stores it in a database. The input is the text data, and the output is a confirmation message indicating that the data was saved successfully.
[1276] Step 3:
[1277] Trend analysis and anomaly detection
[1278] server:
[1279] The server analyzes the stored data using machine learning algorithms (e.g., scikit-learn or TensorFlow) to find trends in the user's mental and physical health. The input is a collection of past text data, and the output is the analyzed trend information.
[1280] Generates an alert when an abnormal change is detected. The input is the analyzed trend information, and the output is an alert message indicating the abnormality.
[1281] Sends alert content to the terminal. The input is the alert message, and the output is a transmission confirmation message.
[1282] Device:
[1283] The device notifies the user of the alert received from the server by voice or text, for example by displaying a pop-up on the screen or by reading it aloud. The input is the alert message, and the output is the notified alert.
[1284] Presents the user with options to contact the healthcare provider, allowing them to select "Contact now" or "Later."
[1285] Step 4:
[1286] Collaboration and advice
[1287] server:
[1288] The server generates appropriate advice based on the user's health condition and sends it to the device. For example, the advice may be "Drink more water today." The input is the health condition assessment result, and the output is the generated advice message.
[1289] If an abnormality is detected, a pre-configured API for linking with a medical institution or consultation center is called. For example, a notification is sent using the Twilio API. The input is an alert message indicating the abnormality, and the output is a message confirming that the notification has been sent.
[1290] If the user gives permission, the necessary data is sent to the medical institution or consultation center. The input is the user's consent and the necessary data, and the output is a transmission confirmation message.
[1291] Device:
[1292] The device notifies the user of the advice received from the server and prompts them to take action, for example by displaying it on the screen or reading it aloud. The input is the advice message, and the output is the notified advice.
[1293] With the user's permission, the necessary data will be sent to the medical institution or consultation center. The user taps the "Send data" button.
[1294] Step 5:
[1295] Feedback collection and system adjustments
[1296] Device:
[1297] The terminal displays an interface to ask the user for feedback on the usefulness of the system's advice or notifications. For example, it displays a rating slider or a comment input form. The input is the feedback request, and the output is the feedback data entered by the user.
[1298] Sends the feedback entered by the user to the server. The input is the feedback data, and the output is a transmission confirmation message.
[1299] server:
[1300] The server analyzes the received feedback and evaluates the system's performance. Here, NLP techniques are used to analyze text. The input is the feedback data, and the output is the performance evaluation results.
[1301] If necessary, adjust the questions the AI asks and how it generates advice. For example, retrain the model based on the latest user data. The input is the performance evaluation results and the latest user data, and the output is the adjusted model.
[1302] Example prompt sentence:
[1303] "Write a prompt to generate appropriate advice based on the user's health condition, such as when they say they didn't sleep well last night."
[1304] (Application example 1)
[1305] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1306] Conventional health management systems have had difficulty accurately monitoring a user's daily health status and providing appropriate advice tailored to each individual's situation. Furthermore, they lacked mechanisms for early detection of abnormalities and prompt communication with medical institutions, resulting in delayed responses to maintain optimal user health. To address these issues, there is a need for technology that uses the user's voice input to check daily health status and generate appropriate prompts.
[1307] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1308] In this invention, the server includes means for inputting basic information of a user via a terminal and transmitting it to the server for storage, means for converting voice input into text data and transmitting it to the server, means for analyzing the text data received by the server and evaluating the user's condition, means for detecting abnormalities based on the user's condition and notifying the user, means for generating notifications and advice for the user and transmitting them to the terminal, and means for converting voice input into text data via the terminal, analyzing the text data, and providing appropriate advice using prompt sentences generated by a generative AI model according to the user's condition. This makes it possible to accurately monitor the user's daily health condition and quickly provide appropriate advice according to each individual situation.
[1309] "Basic information" refers to basic personal data such as a user's name, age, and gender.
[1310] A "terminal" is an electronic device used by a user, and refers to devices such as smartphones, earphones, and speakers.
[1311] "Server" refers to a central processing unit for receiving and processing information sent by users.
[1312] "Voice input" refers to the act of a user providing information by voice.
[1313] "Text data" refers to data that has been converted from voice input into text information.
[1314] "Analysis" refers to the act of processing received data and finding meaning and trends.
[1315] "Abnormality" refers to any unusual change or sign in the user's health condition.
[1316] "Notification" refers to the act of informing a user of information.
[1317] "Advice" refers to specific instructions or advice provided based on the user's condition.
[1318] A "collaboration API" refers to an interface for sharing data and functions between different systems and services.
[1319] "Generative AI model" refers to an artificial intelligence model that generates text or advice based on user input.
[1320] A "prompt sentence" is an input sentence to a generative AI model, and refers to trigger text that elicits appropriate answers or advice.
[1321] This invention is a system that supports users' health management and includes a server, a terminal, and a generative AI model. The system collects basic information about the user, analyzes health data based on that information, and provides appropriate advice as needed. Furthermore, if an abnormality is detected, the system also cooperates with medical institutions.
[1322] User information registration and initial settings
[1323] The terminal provides an interface for entering basic information about the user (such as name, age, and gender). Once the user has entered the information, it is sent to the server and stored. At this point, the terminal receives a confirmation message from the server and displays it to the user.
[1324] Daily conversations and data collection
[1325] The device uses the generative AI model to ask the user questions at set times (e.g., "What did you eat today?", "How are you feeling?"). The user responds verbally, and the speech is converted into text data using the Google Cloud Speech-to-Text API. The converted text data is sent to a server and stored in a database.
[1326] Trend analysis and anomaly detection
[1327] The server uses machine learning algorithms to analyze the stored text data. These algorithms can identify trends in the user's mental and physical health. If an abnormality is detected based on the trend data, the server generates an alert and sends it to the device. The device notifies the user of the alert and provides options for contacting a medical institution, if necessary.
[1328] Collaboration and advice
[1329] The server uses a generative AI model to generate appropriate advice based on the user's health status. This generation utilizes the prompt format. For example, if the user answers, "I didn't sleep well last night," the server generates advice such as, "I recommend trying some light stretching before bed to relax."
[1330] Advice is sent to the user via the device. If an abnormality is detected, the server calls the API to connect with the medical institution and, with the user's permission, sends the necessary data.
[1331] Feedback collection and system adjustments
[1332] The device provides an interface that asks the user for feedback on the usefulness of the system's advice and notifications. The user's input feedback is sent to a server for analysis. The server uses this information to evaluate the system's performance and adjusts the questions the generative AI model asks and how it generates advice, if necessary.
[1333] As a specific example of use, if a user answers a morning question with "I didn't sleep well last night," the system analyzes this information and determines that the user has reported poor sleep for several consecutive nights. The server detects the anomaly and generates advice to the user, such as "We recommend that you review your nightly routine to improve the quality of your sleep," and notifies the user via their device.
[1334] Example prompt sentence:
[1335] User Input: I didn't sleep very well last night.
[1336] Health Tip: To help you relax, try some gentle stretching before bed.
[1337] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1338] Step 1:
[1339] The terminal inputs the user's basic information (name, age, gender, etc.) and sends it to the server. The input data is sent from the terminal to the server, and the server saves the data in a database. This registers the user's basic information in the system.
[1340] Step 2:
[1341] At the set time, the device uses the generative AI model to ask the user questions (e.g., "What did you eat today?", "How are you feeling?"). This voice input is input as voice data from the user.
[1342] Step 3:
[1343] The device uses the Google Cloud Speech-to-Text API to convert voice input into text data, which is then sent to a server for analysis.
[1344] Step 4:
[1345] The server analyzes the received text data using machine learning algorithms to assess the user's health status. The output of this analysis step is the assessment of the user's health status.
[1346] Step 5:
[1347] The server detects abnormalities based on the user's health status. To do this, it compares saved data with new input data. If an abnormality is detected, it generates an alert and sends it to the device. The content of the alert indicates that an abnormality exists.
[1348] Step 6:
[1349] The terminal notifies the user of the alert received from the server and presents an option to contact a medical institution if necessary. The output of the notification is that the user receives the alert information.
[1350] Step 7:
[1351] The server uses a generative AI model to generate appropriate advice based on the user's condition. This advice is generated using prompt sentences. For example, if a user inputs "I didn't sleep well last night," the server generates advice such as "I recommend you try some light stretching before bed to relax."
[1352] Step 8:
[1353] The advice is sent from the server to the terminal, and the terminal notifies the user, allowing the user to receive specific health advice.
[1354] Step 9:
[1355] If necessary, the server will call the linkage API and, with the user's permission, send the necessary data to a medical institution or consultation center. This allows the user's health information to be provided to the appropriate institution, enabling a prompt response.
[1356] Step 10:
[1357] The device asks the user for feedback on the usefulness of the system's advice and notifications. The user enters the feedback, which is then sent to the server. The server analyzes the feedback data to evaluate and adjust the system's performance. This adjustment improves the accuracy of the generative AI model and the overall system performance.
[1358] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1359] The system of the present invention supports the user's mental and physical healthcare and is composed of the following main means. Furthermore, by combining it with an emotion engine for recognizing the user's emotions, it is possible to provide more advanced care.
[1360] User information registration and initial settings
[1361] server:
[1362] Provides an interface for identifying the device (earphones, smartphone, speaker, etc.) used by the user.
[1363] Manages a database for entering and storing basic user information (name, age, gender, etc.).
[1364] Device:
[1365] Displays an initial setup screen and prompts the user to enter basic information.
[1366] The entered basic information is sent to the server.
[1367] Displays to the user a confirmation message received from the server.
[1368] Daily conversations and data collection
[1369] Device:
[1370] It automatically starts at the set time and asks the user questions using a generative AI voice (e.g., "What did you eat today?", "How are you feeling?").
[1371] The user's voice input is recorded and converted into text data using voice recognition technology.
[1372] The converted text data is sent to the server.
[1373] server:
[1374] Receives text data sent from the terminal and stores it in a database.
[1375] Emotion Recognition and State Assessment
[1376] Emotion engine (can run on device or server):
[1377] The user's voice tone, pitch, and rate are analyzed to identify emotions.
[1378] The identified emotion data is transmitted to a server.
[1379] server:
[1380] The emotion data received from the emotion engine is analyzed together with the text data to comprehensively evaluate the user's state.
[1381] The stored data is analyzed using machine learning algorithms to generate trends in the user's mental and physical health.
[1382] Based on trend data, an alert is generated if an abnormal change is detected.
[1383] Send the alert content to the device.
[1384] Device:
[1385] Notify the user of alerts received from the server by voice or text.
[1386] Presents the user with options to contact a healthcare provider.
[1387] Collaboration and advice
[1388] server:
[1389] Generate appropriate advice based on the user's health status and emotional data and send it to the device (e.g., "We recommend listening to music to relax today.").
[1390] If an abnormality is detected, a collaboration API with a pre-configured medical institution or consultation center will be called.
[1391] If the user gives permission, the necessary data will be sent to a medical institution or consultation center.
[1392] Device:
[1393] The user is notified of the advice received from the server and prompted to take action.
[1394] With the user's permission, the necessary data is sent to medical institutions and consultation centers.
[1395] Feedback collection and system adjustments
[1396] Device:
[1397] An interface is displayed to the user requesting feedback on the usefulness of the system's advice and notifications.
[1398] Sends the user-entered feedback to the server.
[1399] server:
[1400] Analyze the feedback received and evaluate the performance of the system.
[1401] If necessary, adjust the questions the AI asks and how it generates advice.
[1402] For example, if a user responds to a morning question with "I didn't sleep well last night," and the emotion engine detects a depressed tone, the system analyzes this information and determines that the user has continued to have poor sleep. The server detects the abnormality and generates advice for the user, such as "To improve the quality of your sleep, we recommend listening to relaxing music tonight," and notifies the user via their device. In this way, the system can comprehensively monitor the user's health, including emotion recognition, and provide prompt and accurate advice.
[1403] The processing flow will be explained below.
[1404] Specific processing steps of the program
[1405] User information registration and initial settings
[1406] Step 1:
[1407] User: Start the device and open the initial setup screen.
[1408] Step 2:
[1409] Terminal: Presents an interface for the user to enter basic information (e.g., name, age, gender).
[1410] Step 3:
[1411] User: Enter the required basic information.
[1412] Step 4:
[1413] Terminal: Sends the entered basic information to the server.
[1414] Step 5:
[1415] Server: Receives basic information sent from the device and stores it in a database.
[1416] Server: Sends a confirmation message to the device that the device was successfully registered.
[1417] Step 6:
[1418] Terminal: Display a confirmation message to the user.
[1419] Daily conversations and data collection
[1420] Step 7:
[1421] Device: Automatically wakes up at a set time and asks questions to the user using a generative AI voice (e.g., "What did you eat today?", "How are you feeling?").
[1422] Step 8:
[1423] User: Answers questions by voice.
[1424] Step 9:
[1425] Terminal: Records the user's voice and converts it into text data using voice recognition technology.
[1426] Step 10:
[1427] Terminal: Sends the converted text data to the server.
[1428] Step 11:
[1429] Server: Receives text data sent from the device and stores it in a database.
[1430] Emotion Recognition and State Assessment
[1431] Step 12:
[1432] Terminal or Server (Emotion Engine): Analyzes the user's voice tone, pitch, and rate to identify emotions.
[1433] Step 13:
[1434] Emotion engine (terminal or server): Sends identified emotion data to the server.
[1435] Step 14:
[1436] Server: Analyzes the emotion data received from the emotion engine along with the text data, and makes a comprehensive assessment of the user's state.
[1437] Step 15:
[1438] Server: Analyzes the stored data using machine learning algorithms to generate trends in the user's mental and physical health.
[1439] Step 16:
[1440] Server: Generates alerts based on trend data if abnormal changes are detected.
[1441] Step 17:
[1442] Server: Sends alert content to the device.
[1443] Step 18:
[1444] Terminal: Notifies the user of alerts received from the server via voice or text.
[1445] Step 19:
[1446] Device: Presents the user with options to contact a healthcare provider.
[1447] Collaboration and advice
[1448] Step 20:
[1449] Server: Generates appropriate advice based on the user's health status and emotional data and sends it to the device (e.g., "We recommend listening to music to relax today.").
[1450] Step 21:
[1451] Server: If an abnormality is detected, it calls a link API with a pre-configured medical institution or consultation center.
[1452] Step 22:
[1453] Server: If the user gives permission, it sends the necessary data to a medical institution or consultation center.
[1454] Step 23:
[1455] Terminal: Notifies the user of the advice received from the server and encourages them to take action.
[1456] Step 24:
[1457] Terminal: With the user's permission, transmits the necessary data to medical institutions and consultation centers.
[1458] Feedback collection and system adjustments
[1459] Step 25:
[1460] Terminal: Display an interface that asks the user for feedback on the usefulness of system advice and notifications.
[1461] Step 26:
[1462] Users: Enter feedback.
[1463] Step 27:
[1464] Terminal: Sends the entered feedback to the server.
[1465] Step 28:
[1466] Server: Analyzes the feedback received and evaluates the performance of the system.
[1467] Server: Adjusts the questions the AI asks and how it generates advice as needed.
[1468] The above processing steps enable comprehensive monitoring of a user's health status and provide appropriate advice and support, including emotion recognition. For example, if a user responds to a morning question with "I didn't sleep well last night," and the emotion engine detects a depressed tone, the system analyzes this information, considers the user's mental and physical health trends, and generates advice such as "To improve the quality of your sleep, we recommend listening to relaxing music tonight," and notifies the user via the device. In this way, this system realizes comprehensive health management that incorporates emotion recognition.
[1469] Example 2
[1470] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1471] In recent years, stress and anxiety have increased in modern society, making mental and physical healthcare increasingly important. However, conventional healthcare systems lack the ability to accurately identify a user's emotional state and provide appropriate advice. As a result, it is difficult to detect abnormalities in a user's state early and take appropriate measures.
[1472] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1473] In this invention, the server includes means for inputting basic information of a user via a terminal and transmitting the information to the server for storage, means for converting voice input into text data and transmitting the text data to the server, means for analyzing the text data received on the server side and evaluating the user's condition, means for analyzing the user's voice tone, pitch, and speed using an emotion engine and identifying emotions, means for transmitting the identified emotion data to the server and analyzing it together with the text data, means for detecting abnormalities based on the user's condition and notifying the user, and means for generating notifications and advice for the user and transmitting the generated notifications and advice to the terminal. This makes it possible to comprehensively monitor the user's emotional state and health condition, detect abnormalities early, and quickly provide appropriate advice and responses.
[1474] "Means for inputting basic user information via a terminal and sending it to a server for storage" refers to the process and infrastructure that allows a user to input basic personal information such as name, age, and gender using the terminal interface, and then transmit the input data to a server and store it in a database.
[1475] "Means for converting voice input into text data and transmitting it to a server" refers to technology and devices for recording a user's voice input, converting it into text data using voice recognition technology, and transmitting the text data to a server.
[1476] The "means for analyzing text data received on the server side and evaluating the user's condition" refers to a system and analysis algorithm that analyzes and processes text data received by the server from the terminal and evaluates the user's mental and physical health condition.
[1477] "Means for identifying emotions by analyzing a user's voice tone, pitch, and rate using an emotion engine" refers to a technology and method for using an emotion engine to analyze characteristics of a user's speech, such as voice tone, pitch, and rate, to identify the user's emotional state.
[1478] The "means for transmitting identified emotion data to a server for analysis together with text data" refers to a process and system for transmitting emotion data identified by the emotion engine to a server for analysis of the data together with text data.
[1479] "Means for detecting abnormalities based on the user's condition and notifying the user" refers to a system and method that determines whether there are any abnormalities in the user's health or emotional state based on data analyzed by the server and notifies the user of the results.
[1480] "Means for generating notifications and advice for the user and sending them to the terminal" refers to the technology and platform for generating appropriate notifications and advice based on the results of analyzing the user's condition and sending them to the terminal.
[1481] "Means for calling APIs to link with medical institutions and consultation centers" refers to systems and technologies for calling application programming interfaces (APIs) to link with pre-set medical institutions and consultation centers when an abnormality is detected.
[1482] "Means for transmitting necessary data to a medical institution or consultation center with the user's permission" refers to the process and system for transmitting necessary health data and emotion data to a designated medical institution or consultation center with the user's consent.
[1483] "Means for collecting user feedback, analyzing it on the server side, and evaluating and adjusting system performance" refers to a mechanism for collecting feedback from users regarding system usage, analyzing it on the server to evaluate system performance, and adjusting the system's functions and service content as necessary.
[1484] This invention is a system for supporting users' mental and physical healthcare, which provides advice and notifications according to the situation by collecting and analyzing the user's basic information, daily activity data, and emotional data. To achieve this, a server, a terminal, and an emotion engine work in cooperation with each other.
[1485] User information registration and initial settings
[1486] Server: The server provides an interface to identify the device used by the user (earphones, smartphone, speaker, etc.). When the user connects to the device, the server collects the information and prepares it to be stored in a database. The server also manages a database for entering and storing basic information about the user (name, age, gender, etc.).
[1487] Terminal: The terminal displays an initial setup screen and prompts the user to enter basic information. After the user enters the information, the terminal sends the information to the server. The server stores the received information in a database and sends a confirmation message to the terminal. The terminal displays a confirmation message to the user, informing them that the initial setup is complete.
[1488] Daily conversations and data collection
[1489] Device: The device automatically wakes up at the set time and asks the user questions using a generative AI voice (e.g., "What did you eat today?", "How are you feeling?"). When the user responds, the device records the voice and converts it into text data using voice recognition technology (e.g., Google Speech-to-Text API).
[1490] Server: The device sends the converted text data to the server, which stores it in a database for further analysis.
[1491] Emotion Recognition and State Assessment
[1492] Emotion engine (can run on the device or on the server): The emotion engine analyzes the user's voice tone, pitch, and rate to identify emotions. The identified emotion data is sent to the server.
[1493] Server: The server analyzes the emotion data received from the emotion engine along with the text data to comprehensively evaluate the user's condition. It uses machine learning algorithms (e.g., TensorFlow, PyTorch) to generate trends in the user's mental and physical health. If any abnormal changes are detected based on the trends, it generates an alert and sends it to the device.
[1494] Collaboration and advice
[1495] Server: The server generates appropriate advice based on the user's health condition and emotional data. For example, if the user's recent sleep data is poor, the server generates advice such as "We recommend listening to relaxing music" and sends it to the device. If an abnormality is detected, the server calls a link API with a pre-set medical institution or consultation center. If the user gives permission, the server sends the necessary data to the medical institution or consultation center.
[1496] Device: The device notifies the user of the advice received from the server. It displays the advice in audio or text and prompts the user to take specific action (e.g., displaying a button to play music). With the user's permission, it also sends the necessary data to medical institutions or consultation centers.
[1497] Feedback collection and system adjustments
[1498] Terminal: The terminal provides an interface to collect user feedback on advice and notifications from the system. When the user enters feedback, it is sent to the server.
[1499] Server: The server analyzes the feedback received from users, evaluates the system's performance, and adjusts the questions the AI asks and how it generates advice to improve the system, if necessary.
[1500] Specific examples
[1501] Example 1: Morning Questions
[1502] The device starts up every morning at 7:00 and asks the user, "Good morning. How was the quality of your sleep last night?" If the user replies, "I didn't sleep well," the device records the voice, converts it into text data, and sends it to the server.
[1503] Example 2: Emotion Identification
[1504] The emotion engine analyzes the user's voice and identifies the emotion "depressed" from the low tone and slow rate. This data is sent to a server and analyzed together with the text data.
[1505] Example 3: Alert and Advice Notifications
[1506] The server continuously checks whether the quality of sleep is poor and detects any abnormalities. It generates advice such as "To improve the quality of your sleep, we recommend listening to relaxing music tonight" and sends it to the device. The device notifies the user of the advice and offers the option to play music.
[1507] Prompt Sentence Examples
[1508] "How was your sleep quality last night?"
[1509] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1510] Step 1:
[1511] The user enters basic information into the terminal and sends it to the server.
[1512] Specific operation: The device displays an initial setup screen and asks the user to enter basic information such as name, age, gender, etc. After the user enters the information, the device sends it to the server.
[1513] Input: Basic information entered by the user, such as name, age, and gender.
[1514] Output: Basic user information data sent from the device to the server.
[1515] Data processing: The server receives these data and stores them in a database.
[1516] Step 2:
[1517] The server stores the basic information in a database and sends a confirmation message to the device.
[1518] Specific operation: The server saves the received basic information in the database, generates a confirmation message for the save completion and sends it to the terminal, which then displays the confirmation message to the user.
[1519] Input: User basic information data.
[1520] Output: Basic information stored in the database and a confirmation message displayed to the user.
[1521] Data processing: The server stores the basic information in a database as structured data.
[1522] Step 3:
[1523] The device will start up at the set time and ask the user questions using a generated AI voice.
[1524] Specific operation: The device automatically starts up at a pre-set time and asks the user questions using a generated AI voice, such as, "How was the quality of your sleep last night?"
[1525] Input: None (automatic start by scheduled trigger).
[1526] Output: A spoken question to the user.
[1527] Data processing: Generate appropriate question voice data using a generative AI model.
[1528] Step 4:
[1529] The user responds by voice, and the terminal records the voice and converts it into text data.
[1530] What it does: When a user responds to a question, the device records the audio and then uses voice recognition technology to convert the audio into text data (e.g., using the Google Speech-to-Text API).
[1531] Input: The user's spoken response.
[1532] Output: User answers converted to text data.
[1533] Data processing: Converting audio data into text format.
[1534] Step 5:
[1535] The terminal transmits the text data to the server.
[1536] Specific operation: The converted text data is sent from the terminal to the server.
[1537] Input: Text data.
[1538] Output: The text data sent to the server.
[1539] Data processing: Encodes text data into transmission packets.
[1540] Step 6:
[1541] The server receives the text data and stores it in a database.
[1542] Specific operation: The server receives text data from the terminal and stores it in a database.
[1543] Input: Text data.
[1544] Output: Text data stored in a database.
[1545] Data processing: Structure the text data and record it in a database.
[1546] Step 7:
[1547] The emotion engine (running on the device or server) analyzes the voice data and identifies emotions.
[1548] How it works: The emotion engine analyzes the user's voice tone, pitch, and rate to identify emotions. The emotion data is sent to the server and analyzed along with the text data.
[1549] Input: User's voice data.
[1550] Output: Emotion data.
[1551] Data processing: Analyze voice characteristics such as tone, pitch, and speed to identify emotions.
[1552] Step 8:
[1553] The server analyzes the emotional data and text data to evaluate the user's state.
[1554] How it works: The server uses emotion data and text data to analyze the user's mental and physical state. The analysis uses machine learning algorithms (e.g., TensorFlow, PyTorch).
[1555] Input: Emotion data, text data.
[1556] Output: User status assessment results.
[1557] Data processing: Integrate both sets of data and perform statistical analysis using machine learning models.
[1558] Step 9:
[1559] The server detects anomalies based on the user's status and generates an alert.
[1560] Specific operation: The server detects anomalies based on the analysis results, generates an alert if an anomaly is detected, and sends the alert content to the user's device.
[1561] Input: User condition assessment result.
[1562] Output: Alert data.
[1563] Data processing: Conditions are judged based on the evaluation results, and alert data is generated and sent.
[1564] Step 10:
[1565] The device will notify the user of the alert and provide advice if necessary.
[1566] Specific behavior: The device notifies the user of the alert received from the server via voice or text, and generates and provides additional advice to the user if necessary.
[1567] Input: Alert data.
[1568] Output: Notification and advice to the user.
[1569] Data processing: Received alerts are notified to the user in an appropriate format and advice is generated.
[1570] Step 11:
[1571] The server and terminal work together to contact medical institutions and consultation centers.
[1572] Specific operation: If an abnormality is detected, the server calls the API for linking with the configured medical institution or consultation center. If the user gives permission, the device sends the necessary data to the medical institution or consultation center.
[1573] Input: User consent, anomaly detection results.
[1574] Output: Sending data to medical institutions and consultation centers.
[1575] Data processing: Formatting and sending the necessary data to medical institutions and consultation centers.
[1576] Step 12:
[1577] Gathering feedback and tuning system performance.
[1578] What it does: The device collects feedback from users and sends that data to the server, which analyzes the feedback, evaluates system performance, and makes adjustments as needed.
[1579] Input: User feedback data.
[1580] Output: Evaluation results and adjustments.
[1581] Data processing: Analyze feedback data to evaluate and adjust system performance.
[1582] (Application example 2)
[1583] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1584] Currently, there is no system in virtual stores that adequately supports users' mental and physical healthcare. As a result, users may feel stressed and their health may deteriorate. In particular, the shopping experience in a virtual store has a significant impact on users' psychological state, so there is a need to introduce emotion analysis and provide appropriate health advice.
[1585] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1586] In this invention, the server includes means for inputting basic information of a user via a device and transmitting it to the server for storage, means for converting voice input into text data and transmitting it to the server, means for analyzing the text data received by the server and evaluating the user's condition, means for detecting abnormalities based on the user's condition and notifying the user, means for generating notifications and advice for the user and transmitting them to the device, means for analyzing the user's emotions and generating healthy advice based on the emotions and notifying the user, and means for collecting feedback, analyzing it on the server, and evaluating and adjusting system performance. This reduces the stress felt by users during their shopping experience in a virtual store, enabling a healthy and comfortable experience.
[1587] "Basic user information" is information that can be used to identify an individual, such as the user's name, age, and gender.
[1588] A "device" is a device that a user operates to input and receive information, such as a smartphone, smart glasses, or headset.
[1589] "Voice input" is voice information that a user utters into a device.
[1590] "Text data" is data obtained by converting voice input into characters.
[1591] The "server" is a central control device that stores and analyzes basic user information and text data.
[1592] "Analysis" is the process of assessing the user's state and emotions based on the received data.
[1593] An "abnormality" is a state that deviates from the user's normal state and may have a negative impact on the user's health or psychology.
[1594] "Notification" refers to the act of the server informing the user of an abnormality or advice.
[1595] The "advice content" is a specific suggestion to encourage healthy behavior based on the user's condition.
[1596] "Emotion analysis" is the process of identifying a user's emotional state based on their voice and text data.
[1597] "Healthy advice" is suggestions and guidance for improving a user's mental and physical health.
[1598] "Feedback" is the act of a user providing opinions or evaluations regarding notifications or advice from the system.
[1599] "System performance" is an indicator of how effectively the system is supporting the user's healthcare.
[1600] A "virtual store" is a virtual environment that provides products and services over the Internet and allows users to have a shopping experience.
[1601] These definitions clarify each element in the claims and make it easier to understand the system configuration and functionality.
[1602] The embodiment of the present invention is to construct a mental and physical healthcare system for a user. The main components of the system include a device, a server, an emotion analysis engine, an advice generation function, and a feedback collection function.
[1603] System configuration
[1604] 1. Device
[1605] Examples of devices include smartphones, smart glasses, headsets, etc. These devices are used to input user information, collect voice data, receive notifications, etc.
[1606] The device provides an interface for inputting basic information about the user and sends that information to the server.
[1607] 2. Server
[1608] The server stores and manages the user's basic information and text data sent from the device.
[1609] The server uses speech recognition technology to convert voice data into text data, and an emotion analysis engine to analyze the user's emotional state.
[1610] The server evaluates the user's state based on the emotion data and generates a notification if an abnormality is detected.
[1611] The server also generates notifications and advice for the user and sends them to the device.
[1612] 3. Sentiment Analysis Engine
[1613] The emotion analysis engine has the ability to analyze user emotions from voice and text data, and this analysis can measure the stress and happiness felt by the user.
[1614] 4. Advice Generation Function
[1615] The server generates healthy advice based on the emotion data, which is then sent to the device and notified to the user.
[1616] As a specific example, if the emotion analysis engine determines that the user is "nervous," the server generates advice such as, "We recommend that you go to a rest area so that you can relax a little."
[1617] 5. Feedback collection feature
[1618] The device displays an interface where the user can provide feedback on the advice and notifications. The collected feedback is sent to a server and used to evaluate and tune the system's performance.
[1619] Specific hardware and software used:
[1620] Hardware:
[1621] Smart glasses, smartphone, microphone
[1622] software:
[1623] speech_recognition library: Converts speech to text.
[1624] emotion_recognition module: A proprietary library for identifying emotions from text.
[1625] virtual_store_interface module: A module for interfacing with virtual stores.
[1626] health_advisor module: Generates health advice based on emotion data.
[1627] feedback_module module: Provides feedback collection functionality.
[1628] Examples of concrete examples and prompts
[1629] Examples:
[1630] If a user answers "I feel nervous" in a virtual store and the sentiment analysis engine detects stress in the user's tone of voice, the health advisor will offer advice such as "I recommend you go to the rest area so you can relax a bit."
[1631] Example prompt sentence:
[1632] "Please tell me a concrete method to convert Japanese speech data into text, analyze the user's emotions from the text, and generate healthy advice based on those emotions."
[1633] This reduces the stress users feel during their shopping experience in a virtual store, allowing for a healthier and more comfortable experience.
[1634] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1635] Program processing steps
[1636] Step 1:
[1637] User information registration and initial settings
[1638] The user uses the device to input basic information (such as name, age, and gender). The terminal sends this input information to the server, which stores the information in a database. Based on the input, the server stores the user information, generates a confirmation message and sends it to the terminal, which displays the confirmation message to the user. Input: User's basic information. Output: Saving basic information to the database, confirmation message.
[1639] Step 2:
[1640] Daily conversations and data collection
[1641] The device will automatically start up at the set time and ask the user questions using a generated AI voice. The user will respond verbally, and the device will collect the voice data. The collected voice data will be converted into text data using voice recognition technology (speech_recognition library). The device will send the text data to the server, which will store it in a database. Input: User's voice response. Output: Text data will be generated and stored in the database.
[1642] Step 3:
[1643] Emotion Recognition and State Assessment
[1644] The server analyzes the received text data using an emotion analysis engine (emotion_recognition module) to assess the user's emotional state. The server then analyzes this emotion data again along with the text data to generate trends of the user's mental and physical health. Input: Collected text data. Output: Emotion data and health trends.
[1645] Step 4:
[1646] Anomaly detection and notification
[1647] The server runs an algorithm to detect abnormal changes based on the generated trend data. If an abnormality is detected, the server generates an alert message and sends it to the terminal. The terminal notifies the user of the alert. Input: Trend data. Output: Notification of alert message.
[1648] Step 5:
[1649] Generating and delivering health advice
[1650] The server generates healthy advice based on the user's emotional state and health data (health_advisor module). The generated advice is sent to the device, which notifies the user. For example, if the emotion analysis engine determines that the user is feeling stressed, the server will provide advice such as "We recommend that you go to a rest area to relax." Input: Emotional data. Output: Generation and notification of health advice.
[1651] Step 6:
[1652] Feedback collection and system adjustments
[1653] The terminal displays an interface to the user to ask for feedback on the effectiveness of the advice and notifications. The user enters the feedback, and the terminal sends it to the server. The server analyzes the feedback and evaluates and adjusts the system performance. Input: User feedback. Output: System performance adjustment.
[1654] Prompt Sentence Examples
[1655] Example prompt sentence:
[1656] "Please tell me a concrete method to convert Japanese speech data into text, analyze the user's emotions from the text, and generate healthy advice based on those emotions."
[1657] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1658] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1659] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1660] [Fourth embodiment]
[1661] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1662] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1663] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1664] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1665] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1666] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1667] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1668] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1669] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1670] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1671] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1672] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1673] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1674] The system of the present invention supports the mental and physical health care of users and is composed of the following main means:
[1675] User information registration and initial settings
[1676] server:
[1677] Provides an interface for identifying the device (earphones, smartphone, speaker, etc.) used by the user.
[1678] Manages a database for entering and storing basic user information (name, age, gender, etc.).
[1679] Device:
[1680] Displays an initial setup screen and prompts the user to enter basic information.
[1681] The entered basic information is sent to the server.
[1682] Displays to the user a confirmation message received from the server.
[1683] Daily conversations and data collection
[1684] Device:
[1685] At set times, a generative AI voice will ask the user questions (e.g., "What did you eat today?", "How are you feeling?").
[1686] The user's voice input is recorded and converted into text data using voice recognition technology.
[1687] The converted text data is sent to the server.
[1688] server:
[1689] Receives text data sent from the terminal and stores it in a database.
[1690] The received data is analyzed periodically to assess the user's health condition.
[1691] Trend analysis and anomaly detection
[1692] server:
[1693] The stored data is analyzed using machine learning algorithms to identify trends in the user's mental and physical health.
[1694] Based on trend data, an alert is generated if an abnormal change is detected.
[1695] Send the alert content to the device.
[1696] Device:
[1697] Notify the user of alerts received from the server by voice or text.
[1698] Presents the user with options to contact a healthcare provider.
[1699] Collaboration and advice
[1700] server:
[1701] Generate appropriate advice based on the user's health status and send it to the device (e.g., "Drink more water today.").
[1702] If an abnormality is detected, a collaboration API with a pre-configured medical institution or consultation center will be called.
[1703] If the user gives permission, the necessary data will be sent to a medical institution or consultation center.
[1704] Device:
[1705] The user is notified of the advice received from the server and prompted to take action.
[1706] With the user's permission, the necessary data is sent to medical institutions and consultation centers.
[1707] Feedback collection and system adjustments
[1708] Device:
[1709] An interface is displayed to the user requesting feedback on the usefulness of the system's advice and notifications.
[1710] Sends the user-entered feedback to the server.
[1711] server:
[1712] Analyze the feedback received and evaluate the performance of the system.
[1713] If necessary, adjust the questions the AI asks and how it generates advice.
[1714] For example, if a user answers a morning question with "I didn't sleep well last night," the system analyzes this information and determines that the user has reported a continuous lack of sleep. The server detects the abnormality and generates advice for the user, such as "We recommend that you review your nightly routine to improve the quality of your sleep," and notifies the user via their device. In this way, the system can comprehensively monitor the user's health status and provide prompt and accurate advice.
[1715] The processing flow will be explained below.
[1716] Specific processing steps of the program
[1717] User information registration and initial settings
[1718] Step 1:
[1719] User: Start the device and open the initial setup screen.
[1720] Step 2:
[1721] Terminal: Presents an interface for the user to enter basic information (e.g., name, age, gender).
[1722] Step 3:
[1723] User: Enter the required basic information.
[1724] Step 4:
[1725] Terminal: Sends the entered basic information to the server.
[1726] Step 5:
[1727] Server: Receives basic information sent from the device and stores it in a database.
[1728] Server: Sends a confirmation message to the device that the device was successfully registered.
[1729] Step 6:
[1730] Terminal: Display a confirmation message to the user.
[1731] Daily conversations and data collection
[1732] Step 7:
[1733] Device: Automatically wakes up at a set time and asks the user questions using a generative AI voice (e.g., "What did you eat today?").
[1734] Step 8:
[1735] User: Answers questions by voice.
[1736] Step 9:
[1737] Terminal: Records the user's voice and converts it into text data using voice recognition technology.
[1738] Step 10:
[1739] Terminal: Sends the converted text data to the server.
[1740] Step 11:
[1741] Server: Receives text data sent from the device and stores it in a database.
[1742] Trend analysis and anomaly detection
[1743] Step 12:
[1744] Server: Analyzes the stored data using machine learning algorithms to generate trends in the user's mental and physical health.
[1745] Step 13:
[1746] Server: Generates alerts based on trend data if abnormal changes are detected.
[1747] Step 14:
[1748] Server: Sends alert content to the device.
[1749] Step 15:
[1750] Terminal: Notifies the user of alerts received from the server via voice or text.
[1751] Step 16:
[1752] Device: Presents the user with options to contact a healthcare provider.
[1753] Collaboration and advice
[1754] Step 17:
[1755] Server: Generates appropriate advice based on the user's health status and sends it to the device (e.g., "Drink more water today.").
[1756] Step 18:
[1757] Server: If an abnormality is detected, it calls a link API with a pre-configured medical institution or consultation center.
[1758] Step 19:
[1759] Server: If the user gives permission, it sends the necessary data to a medical institution or consultation center.
[1760] Step 20:
[1761] Terminal: Notifies the user of the advice received from the server and encourages them to take action.
[1762] Step 21:
[1763] Terminal: With the user's permission, transmits the necessary data to medical institutions and consultation centers.
[1764] Feedback collection and system adjustments
[1765] Step 22:
[1766] Terminal: Display an interface that asks the user for feedback on the usefulness of system advice and notifications.
[1767] Step 23:
[1768] Users: Enter feedback.
[1769] Step 24:
[1770] Terminal: Sends the entered feedback to the server.
[1771] Step 25:
[1772] Server: Analyzes the feedback received and evaluates the performance of the system.
[1773] Server: Adjusts the questions the AI asks and how it generates advice as needed.
[1774] Through the above processing steps, it is possible to comprehensively monitor the user's health condition and provide appropriate advice and support.
[1775] Example 1
[1776] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1777] In modern society, people face the challenge of properly monitoring their own health status and responding promptly in their busy daily lives. In particular, in the areas of mental and physical health, people often delay self-awareness and are unable to identify appropriate countermeasures, creating a need for more efficient and effective healthcare support systems. Furthermore, when an abnormality is detected, it is necessary to quickly collaborate with medical institutions and take appropriate measures, but achieving this requires advanced technology and rapid data analysis.
[1778] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1779] In this invention, the server includes a means for inputting basic information of the user via a terminal and transmitting and storing the information to the server, a means for converting voice input into text data and transmitting the text data to the server, and a means for analyzing the received text data on the server side and evaluating the user's condition. This enables comprehensive monitoring of the user's mental and physical health and prompt cooperation with a medical institution when an abnormality is detected.
[1780] "Basic user information" refers to basic information for identifying an individual, such as the user's name, age, and gender.
[1781] "Terminal" refers to a device used by a user, such as a smartphone, earphones, or speaker.
[1782] "Server" refers to a computer system that manages and processes user information, and stores, analyzes, and notifies various data.
[1783] "Speech input" is the sound data that the user provides to the system through speech synthesis.
[1784] "Text data" refers to data that has been converted from voice input into text information.
[1785] "Analysis" refers to the process of analyzing received data and assessing the user's health status.
[1786] "Evaluation" is the act of determining the user's current health condition based on the analyzed data.
[1787] A "machine learning algorithm" is a method for using large amounts of data to train models to identify trends and make predictions.
[1788] "Trend" refers to the pattern or tendency of fluctuations in a user's health condition over a certain period of time.
[1789] "Abnormal" refers to data or conditions that deviate from the normal range and may have adverse health effects.
[1790] "Notification" is the act of notifying the user of information detected on the server side.
[1791] "Advice" refers to specific advice or suggestions provided based on the user's health status.
[1792] "API" refers to an interface for exchanging data between different systems.
[1793] "Feedback" refers to the evaluations and opinions that users give regarding the services and functions provided by the system.
[1794] A "generative AI model" refers to a model that uses artificial intelligence technology to generate questions and advice.
[1795] The system of the present invention supports users' mental and physical healthcare and includes the following processing steps: registering user information, daily conversations and data collection, trend analysis and anomaly detection, collaboration and advice provision, feedback collection and system adjustment.
[1796] Hardware and software used
[1797] 1. Server: Used for reliable data management and analysis. Specifically, it provides a REST API and uses a MySQL database to store basic user information and collected data. It also uses Python and Pandas for data analysis, and scikit-learn and TensorFlow for implementing machine learning algorithms.
[1798] 2. Terminal: A device through which a user directly inputs information. This includes smartphones, earphones, speakers, etc. It uses the Google Speech-to-Text API to convert voice input into text data and the Google Text-to-Speech API to play the audio.
[1799] System Operation
[1800] The server receives basic information entered by the user via the device and stores it in a database. The device periodically asks the user questions and accepts voice input. The voice is converted into text data and sent to the server. The server analyzes the received text data and evaluates the user's health condition. A machine learning algorithm is used to analyze health trends, and if an abnormality is detected, an alert is quickly generated and sent to the device.
[1801] The server then generates advice based on the user's condition and any abnormalities and sends it to the device. If an abnormality is detected, it calls an API to connect with medical institutions or consultation centers and, with the user's permission, transmits the necessary data. The device then displays the received advice and notifications to the user, encouraging them to take action. It also collects user feedback and uses it to evaluate and adjust the system's performance.
[1802] As a concrete example, consider the case where a user answers a morning question with, "I didn't sleep well last night." The server analyzes this information and confirms that the user has reported a continuous lack of sleep. If the server detects an abnormality, it generates advice such as, "We recommend that you review your nightly routine to improve the quality of your sleep," and notifies the device. The user receives the advice through their device and can contact a medical institution or consultation center if necessary.
[1803] Example prompt sentence:
[1804] "Write a prompt to generate appropriate advice based on the user's health condition, such as when they say they didn't sleep well last night."
[1805] In this way, the system of the present invention can comprehensively monitor the user's health condition and provide prompt and accurate advice.
[1806] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1807] Step 1:
[1808] User information registration and initial settings
[1809] server:
[1810] The server provides an interface (REST API) that identifies the device used by the user. The input is the user's basic information (name, age, gender, etc.) sent from the device, which the server receives and saves in a database. For example, the information is stored in a MySQL database. The output is a confirmation message indicating that the save was successful.
[1811] Device:
[1812] The device displays an initial setup screen and prompts the user to enter basic information. Specifically, the smartphone app displays a form, the user enters the required information, and taps the "Submit" button. The input is information provided by the user, such as name, age, and gender, and the output is the transmission of the information to the server.
[1813] The terminal displays to the user the confirmation message received from the server.
[1814] Step 2:
[1815] Daily conversations and data collection
[1816] Device:
[1817] At a set time, a generative AI voice is used to ask the user questions. For example, questions such as "What did you eat today?" or "How are you feeling?" are played back aloud using the Google Text-to-Speech API. The input is the question sent from the server, and the output is the user's voice input.
[1818] The user responds by voice, and the device records this voice. The recorded voice data is converted into text data using the Google Speech-to-Text API. The input is the user's voice data, and the output is the converted text data.
[1819] The converted text data is sent to the server via HTTPS. The input is the text data, and the output is a transmission confirmation message.
[1820] server:
[1821] The server receives text data sent from the terminal and stores it in a database. The input is the text data, and the output is a confirmation message indicating that the data was saved successfully.
[1822] Step 3:
[1823] Trend analysis and anomaly detection
[1824] server:
[1825] The server analyzes the stored data using machine learning algorithms (e.g., scikit-learn or TensorFlow) to find trends in the user's mental and physical health. The input is a collection of past text data, and the output is the analyzed trend information.
[1826] Generates an alert when an abnormal change is detected. The input is the analyzed trend information, and the output is an alert message indicating the abnormality.
[1827] Sends alert content to the terminal. The input is the alert message, and the output is a transmission confirmation message.
[1828] Device:
[1829] The device notifies the user of the alert received from the server by voice or text, for example by displaying a pop-up on the screen or by reading it aloud. The input is the alert message, and the output is the notified alert.
[1830] Presents the user with options to contact the healthcare provider, allowing them to select "Contact now" or "Later."
[1831] Step 4:
[1832] Collaboration and advice
[1833] server:
[1834] The server generates appropriate advice based on the user's health condition and sends it to the device. For example, the advice may be "Drink more water today." The input is the health condition assessment result, and the output is the generated advice message.
[1835] If an abnormality is detected, a pre-configured API for linking with a medical institution or consultation center is called. For example, a notification is sent using the Twilio API. The input is an alert message indicating the abnormality, and the output is a message confirming that the notification has been sent.
[1836] If the user gives permission, the necessary data is sent to the medical institution or consultation center. The input is the user's consent and the necessary data, and the output is a transmission confirmation message.
[1837] Device:
[1838] The device notifies the user of the advice received from the server and prompts them to take action, for example by displaying it on the screen or reading it aloud. The input is the advice message, and the output is the notified advice.
[1839] With the user's permission, the necessary data will be sent to the medical institution or consultation center. The user taps the "Send data" button.
[1840] Step 5:
[1841] Feedback collection and system adjustments
[1842] Device:
[1843] The terminal displays an interface to ask the user for feedback on the usefulness of the system's advice or notifications. For example, it displays a rating slider or a comment input form. The input is the feedback request, and the output is the feedback data entered by the user.
[1844] Sends the feedback entered by the user to the server. The input is the feedback data, and the output is a transmission confirmation message.
[1845] server:
[1846] The server analyzes the received feedback and evaluates the system's performance. Here, NLP techniques are used to analyze text. The input is the feedback data, and the output is the performance evaluation results.
[1847] If necessary, adjust the questions the AI asks and how it generates advice. For example, retrain the model based on the latest user data. The input is the performance evaluation results and the latest user data, and the output is the adjusted model.
[1848] Example prompt sentence:
[1849] "Write a prompt to generate appropriate advice based on the user's health condition, such as when they say they didn't sleep well last night."
[1850] (Application example 1)
[1851] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1852] Conventional health management systems have had difficulty accurately monitoring a user's daily health status and providing appropriate advice tailored to each individual's situation. Furthermore, they lacked mechanisms for early detection of abnormalities and prompt communication with medical institutions, resulting in delayed responses to maintain optimal user health. To address these issues, there is a need for technology that uses the user's voice input to check daily health status and generate appropriate prompts.
[1853] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1854] In this invention, the server includes means for inputting basic information of a user via a terminal and transmitting it to the server for storage, means for converting voice input into text data and transmitting it to the server, means for analyzing the text data received by the server and evaluating the user's condition, means for detecting abnormalities based on the user's condition and notifying the user, means for generating notifications and advice for the user and transmitting them to the terminal, and means for converting voice input into text data via the terminal, analyzing the text data, and providing appropriate advice using prompt sentences generated by a generative AI model according to the user's condition. This makes it possible to accurately monitor the user's daily health condition and quickly provide appropriate advice according to each individual situation.
[1855] "Basic information" refers to basic personal data such as a user's name, age, and gender.
[1856] A "terminal" is an electronic device used by a user, and refers to devices such as smartphones, earphones, and speakers.
[1857] "Server" refers to a central processing unit for receiving and processing information sent by users.
[1858] "Voice input" refers to the act of a user providing information by voice.
[1859] "Text data" refers to data that has been converted from voice input into text information.
[1860] "Analysis" refers to the act of processing received data and finding meaning and trends.
[1861] "Abnormality" refers to any unusual change or sign in the user's health condition.
[1862] "Notification" refers to the act of informing a user of information.
[1863] "Advice" refers to specific instructions or advice provided based on the user's condition.
[1864] A "collaboration API" refers to an interface for sharing data and functions between different systems and services.
[1865] "Generative AI model" refers to an artificial intelligence model that generates text or advice based on user input.
[1866] A "prompt sentence" is an input sentence to a generative AI model, and refers to trigger text that elicits appropriate answers or advice.
[1867] This invention is a system that supports users' health management and includes a server, a terminal, and a generative AI model. The system collects basic information about the user, analyzes health data based on that information, and provides appropriate advice as needed. Furthermore, if an abnormality is detected, the system also cooperates with medical institutions.
[1868] User information registration and initial settings
[1869] The terminal provides an interface for entering basic information about the user (such as name, age, and gender). Once the user has entered the information, it is sent to the server and stored. At this point, the terminal receives a confirmation message from the server and displays it to the user.
[1870] Daily conversations and data collection
[1871] The device uses the generative AI model to ask the user questions at set times (e.g., "What did you eat today?", "How are you feeling?"). The user responds verbally, and the speech is converted into text data using the Google Cloud Speech-to-Text API. The converted text data is sent to a server and stored in a database.
[1872] Trend analysis and anomaly detection
[1873] The server uses machine learning algorithms to analyze the stored text data. These algorithms can identify trends in the user's mental and physical health. If an abnormality is detected based on the trend data, the server generates an alert and sends it to the device. The device notifies the user of the alert and provides options for contacting a medical institution, if necessary.
[1874] Collaboration and advice
[1875] The server uses a generative AI model to generate appropriate advice based on the user's health status. This generation utilizes the prompt format. For example, if the user answers, "I didn't sleep well last night," the server generates advice such as, "I recommend trying some light stretching before bed to relax."
[1876] Advice is sent to the user via the device. If an abnormality is detected, the server calls the API to connect with the medical institution and, with the user's permission, sends the necessary data.
[1877] Feedback collection and system adjustments
[1878] The device provides an interface that asks the user for feedback on the usefulness of the system's advice and notifications. The user's input feedback is sent to a server for analysis. The server uses this information to evaluate the system's performance and adjusts the questions the generative AI model asks and how it generates advice, if necessary.
[1879] As a specific example of use, if a user answers a morning question with "I didn't sleep well last night," the system analyzes this information and determines that the user has reported poor sleep for several consecutive nights. The server detects the anomaly and generates advice to the user, such as "We recommend that you review your nightly routine to improve the quality of your sleep," and notifies the user via their device.
[1880] Example prompt sentence:
[1881] User Input: I didn't sleep very well last night.
[1882] Health Tip: To help you relax, try some gentle stretching before bed.
[1883] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1884] Step 1:
[1885] The terminal inputs the user's basic information (name, age, gender, etc.) and sends it to the server. The input data is sent from the terminal to the server, and the server saves the data in a database. This registers the user's basic information in the system.
[1886] Step 2:
[1887] At the set time, the device uses the generative AI model to ask the user questions (e.g., "What did you eat today?", "How are you feeling?"). This voice input is input as voice data from the user.
[1888] Step 3:
[1889] The device uses the Google Cloud Speech-to-Text API to convert voice input into text data, which is then sent to a server for analysis.
[1890] Step 4:
[1891] The server analyzes the received text data using machine learning algorithms to assess the user's health status. The output of this analysis step is the assessment of the user's health status.
[1892] Step 5:
[1893] The server detects abnormalities based on the user's health status. To do this, it compares saved data with new input data. If an abnormality is detected, it generates an alert and sends it to the device. The content of the alert indicates that an abnormality exists.
[1894] Step 6:
[1895] The terminal notifies the user of the alert received from the server and presents an option to contact a medical institution if necessary. The output of the notification is that the user receives the alert information.
[1896] Step 7:
[1897] The server uses a generative AI model to generate appropriate advice based on the user's condition. This advice is generated using prompt sentences. For example, if a user inputs "I didn't sleep well last night," the server generates advice such as "I recommend you try some light stretching before bed to relax."
[1898] Step 8:
[1899] The advice is sent from the server to the terminal, and the terminal notifies the user, allowing the user to receive specific health advice.
[1900] Step 9:
[1901] If necessary, the server will call the linkage API and, with the user's permission, send the necessary data to a medical institution or consultation center. This allows the user's health information to be provided to the appropriate institution, enabling a prompt response.
[1902] Step 10:
[1903] The device asks the user for feedback on the usefulness of the system's advice and notifications. The user enters the feedback, which is then sent to the server. The server analyzes the feedback data to evaluate and adjust the system's performance. This adjustment improves the accuracy of the generative AI model and the overall system performance.
[1904] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1905] The system of the present invention supports the user's mental and physical healthcare and is composed of the following main means. Furthermore, by combining it with an emotion engine for recognizing the user's emotions, it is possible to provide more advanced care.
[1906] User information registration and initial settings
[1907] server:
[1908] Provides an interface for identifying the device (earphones, smartphone, speaker, etc.) used by the user.
[1909] Manages a database for entering and storing basic user information (name, age, gender, etc.).
[1910] Device:
[1911] Displays an initial setup screen and prompts the user to enter basic information.
[1912] The entered basic information is sent to the server.
[1913] Displays to the user a confirmation message received from the server.
[1914] Daily conversations and data collection
[1915] Device:
[1916] It automatically starts at the set time and asks the user questions using a generative AI voice (e.g., "What did you eat today?", "How are you feeling?").
[1917] The user's voice input is recorded and converted into text data using voice recognition technology.
[1918] The converted text data is sent to the server.
[1919] server:
[1920] Receives text data sent from the terminal and stores it in a database.
[1921] Emotion Recognition and State Assessment
[1922] Emotion engine (can run on device or server):
[1923] The user's voice tone, pitch, and rate are analyzed to identify emotions.
[1924] The identified emotion data is transmitted to a server.
[1925] server:
[1926] The emotion data received from the emotion engine is analyzed together with the text data to comprehensively evaluate the user's state.
[1927] The stored data is analyzed using machine learning algorithms to generate trends in the user's mental and physical health.
[1928] Based on trend data, an alert is generated if an abnormal change is detected.
[1929] Send the alert content to the device.
[1930] Device:
[1931] Notify the user of alerts received from the server by voice or text.
[1932] Presents the user with options to contact a healthcare provider.
[1933] Collaboration and advice
[1934] server:
[1935] Generate appropriate advice based on the user's health status and emotional data and send it to the device (e.g., "We recommend listening to music to relax today.").
[1936] If an abnormality is detected, a collaboration API with a pre-configured medical institution or consultation center will be called.
[1937] If the user gives permission, the necessary data will be sent to a medical institution or consultation center.
[1938] Device:
[1939] The user is notified of the advice received from the server and prompted to take action.
[1940] With the user's permission, the necessary data is sent to medical institutions and consultation centers.
[1941] Feedback collection and system adjustments
[1942] Device:
[1943] An interface is displayed to the user requesting feedback on the usefulness of the system's advice and notifications.
[1944] Sends the user-entered feedback to the server.
[1945] server:
[1946] Analyze the feedback received and evaluate the performance of the system.
[1947] If necessary, adjust the questions the AI asks and how it generates advice.
[1948] For example, if a user responds to a morning question with "I didn't sleep well last night," and the emotion engine detects a depressed tone, the system analyzes this information and determines that the user has continued to have poor sleep. The server detects the abnormality and generates advice for the user, such as "To improve the quality of your sleep, we recommend listening to relaxing music tonight," and notifies the user via their device. In this way, the system can comprehensively monitor the user's health, including emotion recognition, and provide prompt and accurate advice.
[1949] The processing flow will be explained below.
[1950] Specific processing steps of the program
[1951] User information registration and initial settings
[1952] Step 1:
[1953] User: Start the device and open the initial setup screen.
[1954] Step 2:
[1955] Terminal: Presents an interface for the user to enter basic information (e.g., name, age, gender).
[1956] Step 3:
[1957] User: Enter the required basic information.
[1958] Step 4:
[1959] Terminal: Sends the entered basic information to the server.
[1960] Step 5:
[1961] Server: Receives basic information sent from the device and stores it in a database.
[1962] Server: Sends a confirmation message to the device that the device was successfully registered.
[1963] Step 6:
[1964] Terminal: Display a confirmation message to the user.
[1965] Daily conversations and data collection
[1966] Step 7:
[1967] Device: Automatically wakes up at a set time and asks questions to the user using a generative AI voice (e.g., "What did you eat today?", "How are you feeling?").
[1968] Step 8:
[1969] User: Answers questions by voice.
[1970] Step 9:
[1971] Terminal: Records the user's voice and converts it into text data using voice recognition technology.
[1972] Step 10:
[1973] Terminal: Sends the converted text data to the server.
[1974] Step 11:
[1975] Server: Receives text data sent from the device and stores it in a database.
[1976] Emotion Recognition and State Assessment
[1977] Step 12:
[1978] Terminal or Server (Emotion Engine): Analyzes the user's voice tone, pitch, and rate to identify emotions.
[1979] Step 13:
[1980] Emotion engine (terminal or server): Sends identified emotion data to the server.
[1981] Step 14:
[1982] Server: Analyzes the emotion data received from the emotion engine along with the text data, and makes a comprehensive assessment of the user's state.
[1983] Step 15:
[1984] Server: Analyzes the stored data using machine learning algorithms to generate trends in the user's mental and physical health.
[1985] Step 16:
[1986] Server: Generates alerts based on trend data if abnormal changes are detected.
[1987] Step 17:
[1988] Server: Sends alert content to the device.
[1989] Step 18:
[1990] Terminal: Notifies the user of alerts received from the server via voice or text.
[1991] Step 19:
[1992] Device: Presents the user with options to contact a healthcare provider.
[1993] Collaboration and advice
[1994] Step 20:
[1995] Server: Generates appropriate advice based on the user's health status and emotional data and sends it to the device (e.g., "We recommend listening to music to relax today.").
[1996] Step 21:
[1997] Server: If an abnormality is detected, it calls a link API with a pre-configured medical institution or consultation center.
[1998] Step 22:
[1999] Server: If the user gives permission, it sends the necessary data to a medical institution or consultation center.
[2000] Step 23:
[2001] Terminal: Notifies the user of the advice received from the server and encourages them to take action.
[2002] Step 24:
[2003] Terminal: With the user's permission, transmits the necessary data to medical institutions and consultation centers.
[2004] Feedback collection and system adjustments
[2005] Step 25:
[2006] Terminal: Display an interface that asks the user for feedback on the usefulness of system advice and notifications.
[2007] Step 26:
[2008] Users: Enter feedback.
[2009] Step 27:
[2010] Terminal: Sends the entered feedback to the server.
[2011] Step 28:
[2012] Server: Analyzes the feedback received and evaluates the performance of the system.
[2013] Server: Adjusts the questions the AI asks and how it generates advice as needed.
[2014] The above processing steps enable comprehensive monitoring of a user's health status and provide appropriate advice and support, including emotion recognition. For example, if a user responds to a morning question with "I didn't sleep well last night," and the emotion engine detects a depressed tone, the system analyzes this information, considers the user's mental and physical health trends, and generates advice such as "To improve the quality of your sleep, we recommend listening to relaxing music tonight," and notifies the user via the device. In this way, this system realizes comprehensive health management that incorporates emotion recognition.
[2015] Example 2
[2016] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2017] In recent years, stress and anxiety have increased in modern society, making mental and physical healthcare increasingly important. However, conventional healthcare systems lack the ability to accurately identify a user's emotional state and provide appropriate advice. As a result, it is difficult to detect abnormalities in a user's state early and take appropriate measures.
[2018] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[2019] In this invention, the server includes means for inputting basic information of a user via a terminal and transmitting the information to the server for storage, means for converting voice input into text data and transmitting the text data to the server, means for analyzing the text data received on the server side and evaluating the user's condition, means for analyzing the user's voice tone, pitch, and speed using an emotion engine and identifying emotions, means for transmitting the identified emotion data to the server and analyzing it together with the text data, means for detecting abnormalities based on the user's condition and notifying the user, and means for generating notifications and advice for the user and transmitting the generated notifications and advice to the terminal. This makes it possible to comprehensively monitor the user's emotional state and health condition, detect abnormalities early, and quickly provide appropriate advice and responses.
[2020] "Means for inputting basic user information via a terminal and sending it to a server for storage" refers to the process and infrastructure that allows a user to input basic personal information such as name, age, and gender using the terminal interface, and then transmit the input data to a server and store it in a database.
[2021] "Means for converting voice input into text data and transmitting it to a server" refers to technology and devices for recording a user's voice input, converting it into text data using voice recognition technology, and transmitting the text data to a server.
[2022] The "means for analyzing text data received on the server side and evaluating the user's condition" refers to a system and analysis algorithm that analyzes and processes text data received by the server from the terminal and evaluates the user's mental and physical health condition.
[2023] "Means for identifying emotions by analyzing a user's voice tone, pitch, and rate using an emotion engine" refers to a technology and method for using an emotion engine to analyze characteristics of a user's speech, such as voice tone, pitch, and rate, to identify the user's emotional state.
[2024] The "means for transmitting identified emotion data to a server for analysis together with text data" refers to a process and system for transmitting emotion data identified by the emotion engine to a server for analysis of the data together with text data.
[2025] "Means for detecting abnormalities based on the user's condition and notifying the user" refers to a system and method that determines whether there are any abnormalities in the user's health or emotional state based on data analyzed by the server and notifies the user of the results.
[2026] "Means for generating notifications and advice for the user and sending them to the terminal" refers to the technology and platform for generating appropriate notifications and advice based on the results of analyzing the user's condition and sending them to the terminal.
[2027] "Means for calling APIs to link with medical institutions and consultation centers" refers to systems and technologies for calling application programming interfaces (APIs) to link with pre-set medical institutions and consultation centers when an abnormality is detected.
[2028] "Means for transmitting necessary data to a medical institution or consultation center with the user's permission" refers to the process and system for transmitting necessary health data and emotion data to a designated medical institution or consultation center with the user's consent.
[2029] "Means for collecting user feedback, analyzing it on the server side, and evaluating and adjusting system performance" refers to a mechanism for collecting feedback from users regarding system usage, analyzing it on the server to evaluate system performance, and adjusting the system's functions and service content as necessary.
[2030] This invention is a system for supporting the mental and physical health care of users, and provides advice and notifications according to the situation by collecting and analyzing the user's basic information, daily activity data, and emotional data. To achieve this, a server, a terminal, and an emotion engine work in cooperation with each other.
[2031] User information registration and initial settings
[2032] Server: The server provides an interface to identify the device used by the user (earphones, smartphone, speaker, etc.). When the user connects to the device, the server collects the information and prepares it to be stored in a database. The server also manages a database for entering and storing basic information about the user (name, age, gender, etc.).
[2033] Terminal: The terminal displays an initial setup screen and prompts the user to enter basic information. After the user enters the information, the terminal sends the information to the server. The server stores the received information in a database and sends a confirmation message to the terminal. The terminal displays a confirmation message to the user, informing them that the initial setup is complete.
[2034] Daily conversations and data collection
[2035] Device: The device automatically wakes up at the set time and asks the user questions using a generative AI voice (e.g., "What did you eat today?", "How are you feeling?"). When the user responds, the device records the voice and converts it into text data using voice recognition technology (e.g., Google Speech-to-Text API).
[2036] Server: The device sends the converted text data to the server, which stores it in a database for further analysis.
[2037] Emotion Recognition and State Assessment
[2038] Emotion engine (can run on the device or on the server): The emotion engine analyzes the user's voice tone, pitch, and rate to identify emotions. The identified emotion data is sent to the server.
[2039] Server: The server analyzes the emotion data received from the emotion engine along with the text data to comprehensively evaluate the user's condition. It uses machine learning algorithms (e.g., TensorFlow, PyTorch) to generate trends in the user's mental and physical health. If any abnormal changes are detected based on the trends, it generates an alert and sends it to the device.
[2040] Collaboration and advice
[2041] Server: The server generates appropriate advice based on the user's health condition and emotional data. For example, if the user's recent sleep data is poor, the server generates advice such as "We recommend listening to relaxing music" and sends it to the device. If an abnormality is detected, the server calls a link API with a pre-set medical institution or consultation center. If the user gives permission, the server sends the necessary data to the medical institution or consultation center.
[2042] Device: The device notifies the user of the advice received from the server. It displays the advice in audio or text and prompts the user to take specific action (e.g., displaying a button to play music). With the user's permission, it also sends the necessary data to medical institutions or consultation centers.
[2043] Feedback collection and system adjustments
[2044] Terminal: The terminal provides an interface to collect user feedback on advice and notifications from the system. When the user enters feedback, it is sent to the server.
[2045] Server: The server analyzes the feedback received from users, evaluates the system's performance, and adjusts the questions the AI asks and how it generates advice to improve the system, if necessary.
[2046] Specific examples
[2047] Example 1: Morning Questions
[2048] The device starts up every morning at 7:00 and asks the user, "Good morning. How was the quality of your sleep last night?" If the user replies, "I didn't sleep well," the device records the voice, converts it into text data, and sends it to the server.
[2049] Example 2: Emotion Identification
[2050] The emotion engine analyzes the user's voice and identifies the emotion "depressed" from the low tone and slow rate. This data is sent to a server and analyzed together with the text data.
[2051] Example 3: Alert and Advice Notifications
[2052] The server continuously checks whether the quality of sleep is poor and detects any abnormalities. It generates advice such as "To improve the quality of your sleep, we recommend listening to relaxing music tonight" and sends it to the device. The device notifies the user of the advice and offers the option to play music.
[2053] Prompt Sentence Examples
[2054] "How was your sleep quality last night?"
[2055] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2056] Step 1:
[2057] The user enters basic information into the terminal and sends it to the server.
[2058] Specific operation: The device displays an initial setup screen and asks the user to enter basic information such as name, age, gender, etc. After the user enters the information, the device sends it to the server.
[2059] Input: Basic information entered by the user, such as name, age, and gender.
[2060] Output: Basic user information data sent from the device to the server.
[2061] Data processing: The server receives these data and stores them in a database.
[2062] Step 2:
[2063] The server stores the basic information in a database and sends a confirmation message to the device.
[2064] Specific operation: The server saves the received basic information in the database, generates a confirmation message to confirm the save completion and sends it to the terminal, which then displays the confirmation message to the user.
[2065] Input: User basic information data.
[2066] Output: Basic information stored in the database and a confirmation message displayed to the user.
[2067] Data processing: The server stores the basic information in a database as structured data.
[2068] Step 3:
[2069] The device will start up at the set time and ask the user questions using a generated AI voice.
[2070] Specific operation: The device automatically starts up at a pre-set time and asks the user questions using a generated AI voice, such as, "How was the quality of your sleep last night?"
[2071] Input: None (automatic start by scheduled trigger).
[2072] Output: A spoken question to the user.
[2073] Data processing: Generate appropriate question voice data using a generative AI model.
[2074] Step 4:
[2075] The user responds by voice, and the terminal records the voice and converts it into text data.
[2076] What it does: When a user responds to a question, the device records the audio and then uses voice recognition technology to convert the audio into text data (e.g., using the Google Speech-to-Text API).
[2077] Input: The user's spoken response.
[2078] Output: User answers converted to text data.
[2079] Data processing: Converting audio data into text format.
[2080] Step 5:
[2081] The terminal transmits the text data to the server.
[2082] Specific operation: The converted text data is sent from the terminal to the server.
[2083] Input: Text data.
[2084] Output: The text data sent to the server.
[2085] Data processing: Encodes text data into transmission packets.
[2086] Step 6:
[2087] The server receives the text data and stores it in a database.
[2088] Specific operation: The server receives text data from the terminal and stores it in a database.
[2089] Input: Text data.
[2090] Output: Text data stored in a database.
[2091] Data processing: Structure the text data and record it in a database.
[2092] Step 7:
[2093] The emotion engine (running on the device or server) analyzes the voice data and identifies emotions.
[2094] How it works: The emotion engine analyzes the user's voice tone, pitch, and rate to identify emotions. The emotion data is sent to the server and analyzed along with the text data.
[2095] Input: User's voice data.
[2096] Output: Emotion data.
[2097] Data processing: Analyze voice characteristics such as tone, pitch, and speed to identify emotions.
[2098] Step 8:
[2099] The server analyzes the emotional data and text data to evaluate the user's state.
[2100] How it works: The server uses emotion data and text data to analyze the user's mental and physical state. The analysis uses machine learning algorithms (e.g., TensorFlow, PyTorch).
[2101] Input: Emotion data, text data.
[2102] Output: User status assessment results.
[2103] Data processing: Integrate both sets of data and perform statistical analysis using machine learning models.
[2104] Step 9:
[2105] The server detects anomalies based on the user's status and generates an alert.
[2106] Specific operation: The server detects anomalies based on the analysis results, generates an alert if an anomaly is detected, and sends the alert content to the user's device.
[2107] Input: User condition assessment result.
[2108] Output: Alert data.
[2109] Data processing: Conditions are judged based on the evaluation results, and alert data is generated and sent.
[2110] Step 10:
[2111] The device will notify the user of the alert and provide advice if necessary.
[2112] Specific behavior: The device notifies the user of the alert received from the server via voice or text, and generates and provides additional advice to the user if necessary.
[2113] Input: Alert data.
[2114] Output: Notification and advice to the user.
[2115] Data processing: Received alerts are notified to the user in an appropriate format and advice is generated.
[2116] Step 11:
[2117] The server and terminal work together to contact medical institutions and consultation centers.
[2118] Specific operation: If an abnormality is detected, the server calls the API for linking with the configured medical institution or consultation center. If the user gives permission, the device sends the necessary data to the medical institution or consultation center.
[2119] Input: User consent, anomaly detection results.
[2120] Output: Sending data to medical institutions and consultation centers.
[2121] Data processing: Formatting and sending the necessary data to medical institutions and consultation centers.
[2122] Step 12:
[2123] Gathering feedback and tuning system performance.
[2124] What it does: The device collects feedback from users and sends that data to the server, which analyzes the feedback, evaluates system performance, and makes adjustments as needed.
[2125] Input: User feedback data.
[2126] Output: Evaluation results and adjustments.
[2127] Data processing: Analyze feedback data to evaluate and adjust system performance.
[2128] (Application example 2)
[2129] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2130] Currently, there is no system in virtual stores that adequately supports users' mental and physical healthcare. As a result, users may feel stressed and their health may deteriorate. In particular, the shopping experience in a virtual store has a significant impact on users' psychological state, so there is a need to introduce emotion analysis and provide appropriate health advice.
[2131] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[2132] In this invention, the server includes means for inputting basic information of a user via a device and transmitting it to the server for storage, means for converting voice input into text data and transmitting it to the server, means for analyzing the text data received by the server and evaluating the user's condition, means for detecting abnormalities based on the user's condition and notifying the user, means for generating notifications and advice for the user and transmitting them to the device, means for analyzing the user's emotions and generating healthy advice based on the emotions and notifying the user, and means for collecting feedback, analyzing it on the server, and evaluating and adjusting system performance. This reduces the stress felt by users during their shopping experience in a virtual store, enabling a healthy and comfortable experience.
[2133] "Basic user information" is information that can be used to identify an individual, such as the user's name, age, and gender.
[2134] A "device" is a device that a user operates to input and receive information, such as a smartphone, smart glasses, or headset.
[2135] "Voice input" is voice information that a user utters into a device.
[2136] "Text data" is data obtained by converting voice input into characters.
[2137] The "server" is a central control device that stores and analyzes basic user information and text data.
[2138] "Analysis" is the process of assessing the user's state and emotions based on the received data.
[2139] An "abnormality" is a state that deviates from the user's normal state and may have a negative impact on the user's health or psychology.
[2140] "Notification" refers to the act of the server informing the user of an abnormality or advice.
[2141] The "advice content" is a specific suggestion to encourage healthy behavior based on the user's condition.
[2142] "Emotion analysis" is the process of identifying a user's emotional state based on their voice and text data.
[2143] "Healthy advice" is suggestions and guidance for improving a user's mental and physical health.
[2144] "Feedback" is the act of a user providing opinions or evaluations regarding notifications or advice from the system.
[2145] "System performance" is an indicator of how effectively the system is supporting the user's healthcare.
[2146] A "virtual store" is a virtual environment that provides products and services over the Internet and allows users to have a shopping experience.
[2147] These definitions clarify each element in the claims and make it easier to understand the system configuration and functionality.
[2148] The embodiment of the present invention is to construct a mental and physical healthcare system for a user. The main components of the system include a device, a server, an emotion analysis engine, an advice generation function, and a feedback collection function.
[2149] System configuration
[2150] 1. Device
[2151] Examples of devices include smartphones, smart glasses, headsets, etc. These devices are used to input user information, collect voice data, receive notifications, etc.
[2152] The device provides an interface for inputting basic information about the user and sends that information to the server.
[2153] 2. Server
[2154] The server stores and manages the user's basic information and text data sent from the device.
[2155] The server uses speech recognition technology to convert voice data into text data, and an emotion analysis engine to analyze the user's emotional state.
[2156] The server evaluates the user's state based on the emotion data and generates a notification if an abnormality is detected.
[2157] The server also generates notifications and advice for the user and sends them to the device.
[2158] 3. Sentiment Analysis Engine
[2159] The emotion analysis engine has the ability to analyze user emotions from voice and text data, and this analysis can measure the stress and happiness felt by the user.
[2160] 4. Advice Generation Function
[2161] The server generates healthy advice based on the emotion data, which is then sent to the device and notified to the user.
[2162] As a specific example, if the emotion analysis engine determines that the user is "nervous," the server generates advice such as, "We recommend that you go to a rest area so that you can relax a little."
[2163] 5. Feedback collection feature
[2164] The device displays an interface where the user can provide feedback on the advice and notifications. The collected feedback is sent to a server and used to evaluate and tune the system's performance.
[2165] Specific hardware and software used:
[2166] Hardware:
[2167] Smart glasses, smartphone, microphone
[2168] software:
[2169] speech_recognition library: Converts speech to text.
[2170] emotion_recognition module: A proprietary library for identifying emotions from text.
[2171] virtual_store_interface module: A module for interfacing with virtual stores.
[2172] health_advisor module: Generates health advice based on emotion data.
[2173] feedback_module module: Provides feedback collection functionality.
[2174] Examples of concrete examples and prompts
[2175] Examples:
[2176] If a user answers "I feel nervous" in a virtual store and the sentiment analysis engine detects stress in the user's tone of voice, the health advisor will offer advice such as "I recommend you go to the rest area so you can relax a bit."
[2177] Example prompt sentence:
[2178] "Please tell me a concrete method to convert Japanese speech data into text, analyze the user's emotions from the text, and generate healthy advice based on those emotions."
[2179] This reduces the stress users feel during their shopping experience in a virtual store, allowing for a healthier and more comfortable experience.
[2180] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2181] Program processing steps
[2182] Step 1:
[2183] User information registration and initial settings
[2184] The user uses the device to input basic information (such as name, age, and gender). The terminal sends this input information to the server, which stores the information in a database. Based on the input, the server stores the user information, generates a confirmation message and sends it to the terminal, which displays the confirmation message to the user. Input: User's basic information. Output: Saving basic information to the database, confirmation message.
[2185] Step 2:
[2186] Daily conversations and data collection
[2187] The device will automatically start up at the set time and ask the user questions using a generated AI voice. The user will respond verbally, and the device will collect the voice data. The collected voice data will be converted into text data using voice recognition technology (speech_recognition library). The device will send the text data to the server, which will store it in a database. Input: User's voice response. Output: Text data will be generated and stored in the database.
[2188] Step 3:
[2189] Emotion Recognition and State Assessment
[2190] The server analyzes the received text data using an emotion analysis engine (emotion_recognition module) to assess the user's emotional state. The server then analyzes this emotion data again along with the text data to generate trends of the user's mental and physical health. Input: Collected text data. Output: Emotion data and health trends.
[2191] Step 4:
[2192] Anomaly detection and notification
[2193] The server runs an algorithm to detect abnormal changes based on the generated trend data. If an abnormality is detected, the server generates an alert message and sends it to the terminal. The terminal notifies the user of the alert. Input: Trend data. Output: Notification of alert message.
[2194] Step 5:
[2195] Generating and delivering health advice
[2196] The server generates healthy advice based on the user's emotional state and health data (health_advisor module). The generated advice is sent to the device, which notifies the user. For example, if the emotion analysis engine determines that the user is feeling stressed, the server will provide advice such as "We recommend that you go to a rest area to relax." Input: Emotional data. Output: Generation and notification of health advice.
[2197] Step 6:
[2198] Feedback collection and system adjustments
[2199] The terminal displays an interface to the user to ask for feedback on the effectiveness of the advice and notifications. The user enters the feedback, and the terminal sends it to the server. The server analyzes the feedback and evaluates and adjusts the system performance. Input: User feedback. Output: System performance adjustment.
[2200] Prompt Sentence Examples
[2201] Example prompt sentence:
[2202] "Please tell me a concrete method to convert Japanese speech data into text, analyze the user's emotions from the text, and generate healthy advice based on those emotions."
[2203] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2204] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2205] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2206] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2207] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2208] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2209] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2210] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2211] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2212] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2213] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2214] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2215] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2216] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2217] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2218] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2219] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2220] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2221] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such ...
Claims
1. A means for inputting basic user information via a terminal and sending it to a server for storage; means for converting voice input into text data and transmitting the text data to a server; A means for analyzing the received text data on the server side and evaluating the user's status; means for detecting an abnormality based on the user's state and notifying the user; A system including a means for generating notifications and advice for users and sending them to terminals.
2. When an abnormality is detected, a means to call an API to connect with medical institutions and consultation centers, 2. The system according to claim 1, further comprising means for transmitting the necessary data to a medical institution or a consultation center with the user's permission.
3. 10. The system of claim 1, further comprising means for collecting user feedback, analyzing it on the server side, and evaluating and adjusting the performance of the system.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A