System

A system using a generative AI model to analyze voice data for patient health assessment addresses the challenge of remote health evaluation, providing quick and accurate medical responses.

JP2026035464APending Publication Date: 2026-03-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Patients in remote locations or during emergencies face challenges in quickly and accurately assessing their health conditions without immediate medical assistance, leading to potential health risks due to delayed responses.

Method used

A system utilizing a generative AI model to analyze voice data for breathing, voice intonation, and speaking rate to evaluate physical conditions and symptoms, automatically providing appropriate medical responses such as prescriptions, dispatching doctors, or contacting emergency services.

Benefits of technology

Enables rapid and accurate assessment of patient health, ensuring timely medical interventions and reducing health risks by automating medical decision-making based on voice data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035464000001_ABST
    Figure 2026035464000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving and storing audio data; means for executing a generative AI model for analyzing the received audio data; means for assessing the physical condition and symptoms of a patient from the analyzed data; and means for suggesting an appropriate medical response based on the assessment results.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] When patients are in places where there is no doctor on-site, or when a doctor cannot arrive immediately in an emergency, it is difficult for them to analyze their own health condition. Furthermore, a shortage of doctors or delays in emergency response can have a negative impact on patients' health. For this reason, there is a need for a system that uses voice data to quickly and accurately analyze an individual's health condition and symptoms, and provide appropriate medical care. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for receiving and storing voice data, a means for executing a generative AI model to analyze the received voice data, a means for evaluating a patient's physical condition and symptoms from the analyzed data, and a means for proposing appropriate medical treatment based on the evaluation results. The generative AI model analyzes breathing, voice intonation, and speaking rate from the voice data and evaluates the urgency of the patient's physical condition and symptoms based on the results. Based on the evaluation results, the system automatically sends a prescription, dispatches a doctor, or contacts emergency services, thereby achieving prompt and accurate medical treatment.

[0006] "Audio data" refers to digital or analog data that records sound.

[0007] "Reception" means that the server receives and processes data sent from the terminal.

[0008] "Storage" means recording the received data in a database or storage.

[0009] A "generative AI model" is an algorithm or model that uses artificial intelligence technology to analyze voice data and extract specific information.

[0010] "Analysis" refers to analyzing characteristics contained in audio data, such as breathing, voice intonation, and speaking rate, and extracting information.

[0011] "Evaluation" refers to determining the patient's physical condition, the state of symptoms, and the urgency of the condition based on the analyzed data.

[0012] "Recommendation" refers to providing appropriate medical responses (e.g., sending a prescription, dispatching a doctor, contacting emergency services, etc.) based on the assessment results. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2]1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0021] [First embodiment]

[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0034] System Overview

[0035] The present invention relates to a system that automatically evaluates a patient's physical condition and symptoms using voice data and proposes appropriate medical treatment. This system is mainly composed of means for receiving, analyzing, and evaluating the voice data and proposing medical treatment.

[0036] System Configuration

[0037] The system consists of the following main components:

[0038] 1. Terminal: This includes the device used by the user (e.g., a smartphone). The user is responsible for recording audio and sending the data to the server.

[0039] 2. Server: This is the central system that stores and analyzes the received voice data, runs the generative AI model, and makes evaluations and suggestions.

[0040] 3. Database: Serves as data storage for saving voice data and analysis results.

[0041] Explanation of program processing

[0042] Recording and sending audio

[0043] 1. User: If the user feels unwell, they speak to the terminal about their symptoms.

[0044] 2. Device: The device (e.g., a smartphone) starts a recording application and records the user's voice. When the recording is finished, the voice data is sent to the server.

[0045] Receiving and storing audio data

[0046] 1. Server: Receives the voice data sent from the device and stores it in a database.

[0047] Analysis of audio data

[0048] 1. Server: The stored voice data is input into the generative AI model for analysis. The generative AI model extracts information such as breathing, voice intonation, and speaking rate.

[0049] 2. Server: Extracts features from the analyzed data and records them.

[0050] Assessment of the patient's physical condition and symptoms

[0051] 1. Server: Evaluates the patient's physical condition and symptoms based on the extracted features. The evaluation also involves comparison with past voice data.

[0052] 2. Server: Based on the assessment results, classify the patient's physical condition and the urgency of their symptoms (mild, severe, urgent).

[0053] Suggested actions needed

[0054] 1. Server: Based on the evaluation results, propose appropriate medical responses. Specific responses are as follows:

[0055] Mild cases: Send a prescription to the pharmacy and notify the user to pick up the medication.

[0056] In severe cases: The nearest doctor will be contacted and dispatched. The user will be notified of the date and time of the doctor's visit.

[0057] In case of emergency: Call emergency services and arrange for an ambulance. Notify the user that immediate emergency response is required.

[0058] Specific examples

[0059] Example 1: Mild illness

[0060] 1. User: Feeling unwell, speaks about symptoms into the terminal.

[0061] 2. Device: Records and sends the audio data to the server.

[0062] 3. Server: Analyzes the voice data and evaluates the symptoms as mild.

[0063] 4. Server: Sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[0064] Example 2: Urgent illness

[0065] 1. User: Feeling a sudden deterioration in their health, they talk about their symptoms into the device.

[0066] 2. Device: Records and sends the audio data to the server.

[0067] 3. Server: Analyzes the voice data and detects high levels of urgency.

[0068] 4. Server: Immediately contacts emergency services and dispatches an ambulance. The user is notified that emergency response is required.

[0069] In this way, the present invention is a system that can efficiently and quickly evaluate a patient's physical condition and symptoms using voice data and provide appropriate medical treatment.

[0070] The processing flow will be explained below.

[0071] Step 1:

[0072] User: If the user feels unwell, they talk to the device about their symptoms.

[0073] Step 2:

[0074] Terminal: The terminal (e.g., a smartphone) launches a recording application and records what the user says.

[0075] Step 3:

[0076] Device: Once recording is complete, the audio data is sent to the server.

[0077] Step 4:

[0078] Server: Receives the voice data sent from the terminal.

[0079] Step 5:

[0080] Server: Stores the received voice data in a database.

[0081] Step 6:

[0082] Server: Inputs the stored voice data into the generative AI model.

[0083] Step 7:

[0084] Server (generative AI model): The generative AI model analyzes the audio data and extracts features such as breathing, voice intonation, and speaking rate.

[0085] Step 8:

[0086] Server (analysis system): Records features based on the analyzed data.

[0087] Step 9:

[0088] Server: Evaluates the patient's physical condition and symptoms based on the features. This also compares the results with past voice data.

[0089] Step 10:

[0090] Server: As a result of the evaluation, classify the patient's physical condition and the urgency of their symptoms (mild, severe, urgent).

[0091] Step 11:

[0092] Server: Selects appropriate medical response based on the evaluation results.

[0093] Step 12:

[0094] Server: If the condition is mild, it sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[0095] Step 13:

[0096] Server: If the condition is severe, contact the nearest doctor and arrange for a doctor to be dispatched. The user is notified of the date and time of the doctor's visit.

[0097] Step 14:

[0098] Server: In case of an emergency, contacts emergency services and arranges for an ambulance based on the user's location information. The user is notified that emergency response is required.

[0099] Through the above steps, the system of the present invention can utilize voice data to quickly evaluate the patient's physical condition and symptoms and provide appropriate medical care.

[0100] Example 1

[0101] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0102] Conventional medical systems have struggled to remotely and quickly and accurately assess a patient's physical condition and symptoms, and to propose appropriate medical treatment. They also struggled to respond immediately to particularly urgent symptoms, posing a significant risk to the patient's health and safety. Furthermore, they lacked advanced technology for analyzing voice data and extracting features, making it impossible to accurately identify changes in physical condition.

[0103] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0104] In this invention, the server includes means for receiving and storing voice data, means for executing a generative AI model to analyze the received voice data, means for extracting features from the analyzed data and recording them in a database, means for evaluating the patient's physical condition and symptoms from the extracted features and comparing them with past data, means for classifying the urgency of the symptoms based on the evaluation results, and means for proposing appropriate medical responses based on the evaluation results. This makes it possible to quickly and accurately evaluate the patient's physical condition and symptoms and propose appropriate medical measures. Specific responses can include automatically sending prescriptions, dispatching medical personnel, and contacting emergency services.

[0105] "Audio data" refers to data in which sound is recorded and stored in digital form.

[0106] "Reception" is the process of receiving data or signals sent from the outside.

[0107] "Storage" means storing the received data in a database or storage so that it can be accessed later if necessary.

[0108] A "generative AI model" is an artificial intelligence model that has been trained using machine learning or deep learning to perform specific tasks automatically.

[0109] "Analysis" is the processing of data and the extraction of meaningful information.

[0110] A "feature" is a numerical value or index extracted from data that represents a specific pattern or characteristic.

[0111] A "database" is an electronic data management system that stores data systematically and enables efficient searching and manipulation.

[0112] "Evaluation" refers to determining the content or state of data based on specific criteria.

[0113] "Urgency" is a scale that indicates the severity of the evaluated symptoms and the degree of need for response.

[0114] "Proposals" refer to the presentation of optimal actions or measures based on the evaluation results.

[0115] "Medical response" refers to medical services and treatments provided according to the patient's physical condition and symptoms.

[0116] MODE FOR CARRYING OUT THE INVENTION

[0117] The present invention relates to a system that automatically evaluates a patient's physical condition and symptoms using voice data and proposes appropriate medical treatment. This system mainly includes means for receiving, analyzing, and evaluating the voice data and proposing medical treatment.

[0118] System Overview

[0119] The system consists of the following main components:

[0120] 1. Terminal (e.g., smartphone): A device that allows a user to record audio and send the data to a server.

[0121] 2. Server: A central system for storing and analyzing received voice data, running generative AI models, and providing a means for evaluation and recommendations.

[0122] 3. Database: A storage system that stores voice data and analysis results.

[0123] Recording and sending audio

[0124] (User): When a user feels unwell, they can talk about their symptoms using a device such as a smartphone. For example, they can say, "I have chest pain."

[0125] (Device): The device (smartphone) activates the voice recording function and records the user's voice. When the recording is finished, the voice data is automatically sent to the server using the HTTPS protocol.

[0126] Receiving and storing audio data

[0127] (Server): The server receives the voice data sent from the device. The received voice data is stored in a temporary storage area, and once it is confirmed that it has been received completely, it is permanently stored in the database.

[0128] Analysis of audio data

[0129] (Server): The server inputs the saved voice data into a generative AI model (e.g., an artificial intelligence model trained using machine learning or deep learning) and begins the analysis process. Specifically, the AI ​​model analyzes features such as breathing, voice intonation, and speaking rate from the voice data.

[0130] (Server): Extracts features from the analyzed data and records them in a database. These features are basic data for evaluating the patient's physical condition.

[0131] Assessment of the patient's physical condition and symptoms

[0132] (Server): Evaluates the patient's physical condition and symptoms based on the extracted features. Executes evaluation procedures to compare with past voice data and analyze changes. For example, compares past voice data with current data and uses a machine learning algorithm to determine whether there are any abnormal changes.

[0133] (Server): Based on the evaluation results, the urgency of the symptoms is classified into three categories: mild, severe, and urgent.

[0134] Suggested actions needed

[0135] (Server): Based on the evaluation results, the following appropriate medical responses are proposed:

[0136] For mild cases: Email the prescription to the nearest pharmacy and notify the user to collect the medication. For example, send the prescription information to the pharmacy's email address.

[0137] In severe cases: Contact the nearest healthcare professional and arrange for a visit. Notify the user about the date and time of the healthcare professional's visit, for example by contacting the healthcare professional using their contact information.

[0138] In case of emergency: Contact emergency services and dispatch an ambulance. The user is notified immediately and informed that emergency response is required, for example by auto-dialing an emergency number and providing location and symptom information.

[0139] Specific examples

[0140] Example 1: Mild illness

[0141] (User): The user experiences a slight headache and says to their smartphone, "My head hurts a little."

[0142] (Device): The recording application records the audio and sends the audio data to the server.

[0143] (Server): Receives the voice data, analyzes it using a generative AI model, and determines that the symptoms are mild.

[0144] (Server): Sends the prescription to the pharmacy and notifies the user that the medicine has been received. For example, the server sends the prescription information to the pharmacy's email address.

[0145] Example 2: Urgent illness

[0146] (User): Feeling sudden chest pain, he says into his smartphone, "My chest hurts, I can't breathe."

[0147] (Device): The recording application records the audio and sends the audio data to the server.

[0148] (Server): Receives voice data, analyzes it using a generative AI model, and determines the urgency of the data.

[0149] (Server): Contacts emergency services and dispatches an ambulance. The user is notified immediately and informed that emergency response is required. For example, by automatically dialing an emergency number and providing location and symptom information.

[0150] Prompt Sentence Examples

[0151] "Analyze breathing, voice intonation, and speech rate to determine the patient's health condition. For example, analyze the audio data of 'I have a severe stomachache' and assess the urgency of the situation."

[0152] In this way, the present invention is a system that utilizes voice data to quickly and accurately evaluate a patient's physical condition and symptoms, and provide appropriate medical care.

[0153] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0154] Step 1:

[0155] (Input): If the user feels unwell, they can speak to the terminal about their symptoms. For example, they can say, "My chest hurts."

[0156] (Action): The device activates the voice recording function and records the user's voice.

[0157] (Output): Recorded audio data.

[0158] Step 2:

[0159] (Input): Recorded audio data.

[0160] (Operation): After the device finishes recording, it sends the audio data to the server using the HTTPS protocol.

[0161] (Output): The audio data sent to the server.

[0162] Step 3:

[0163] (Input): Audio data sent from the device.

[0164] (Operation): The server receives the voice data and stores it in a temporary storage area. Once reception is complete, the voice data is permanently stored in the database.

[0165] (Output): The audio data stored in the database.

[0166] Step 4:

[0167] (Input): Audio data stored in the database.

[0168] (Operation): The server inputs the voice data into the generative AI model and performs an analysis process. The generative AI model analyzes features such as breathing, voice intonation, and speaking rate from the voice data.

[0169] (Output): Parsed features.

[0170] Step 5:

[0171] (Input): Parsed features.

[0172] (Operation): The server extracts features and records them in a database.

[0173] (Output): Features recorded in the database.

[0174] Step 6:

[0175] (Input): Features recorded in the database.

[0176] (Operation): The server evaluates the patient's physical condition and symptoms based on the extracted features, compares them with past data, and analyzes changes. Specifically, it uses machine learning algorithms to identify abnormal changes.

[0177] (Output): Evaluation results (physical condition evaluation and urgency determination).

[0178] Step 7:

[0179] (Input): Evaluation result.

[0180] (Operation): Based on the evaluation results, the server classifies the urgency of the symptoms into three categories: mild, severe, and urgent.

[0181] (Output): Urgency classification result.

[0182] Step 8:

[0183] (Input): Urgency classification result.

[0184] (Operation): The server suggests appropriate medical responses, specifically sending a prescription to a pharmacy if the condition is mild, dispatching medical personnel if the condition is severe, and contacting emergency services to dispatch an ambulance if the condition is urgent.

[0185] (Output): Notification to the user and medical response.

[0186] This enables the system to use voice data to quickly and accurately assess a patient's physical condition and symptoms, enabling appropriate medical treatment.

[0187] (Application example 1)

[0188] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0189] Conventional security systems have had difficulty quickly and accurately detecting physical security threats and anomalies. Furthermore, delayed response in emergencies can lead to serious damage to human life and property. To solve these problems, there is a need for a system that can analyze voice data, automatically assess security situations, and quickly take appropriate action.

[0190] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0191] In this invention, the server includes means for receiving and storing voice data, means for executing a generative AI model to analyze the received voice data, means for evaluating the target situation or abnormality from the analyzed data, and means for proposing or implementing appropriate countermeasures based on the evaluation results, thereby enabling quick and accurate detection of physical security threats and abnormalities and prompt implementation of appropriate countermeasures.

[0192] "Audio data" means digital or analog information that is an electronic recording of sound.

[0193] "Reception" is the act of a specific device or system taking in data or signals from the outside.

[0194] A "generative AI model" is an artificial intelligence algorithm or network designed to analyze data and make predictions.

[0195] "Analysis" is the process of examining data or information in detail to understand its structure and meaning.

[0196] "Subject" refers to an object or situation that is the subject of observation or analysis under specific conditions or circumstances.

[0197] A "situation" refers to the state or condition of an environment or event at a given moment.

[0198] An "abnormality" is a problem or malfunction that deviates from normal conditions or norms.

[0199] "Evaluation" is the act of making judgments or analyses based on data or information in accordance with specific criteria.

[0200] "Countermeasures" refer to specific actions or measures taken to address a problem or abnormality.

[0201] A "suggestion" is a recommendation for a particular action or solution.

[0202] "Implementation" means actually carrying out the proposed measures or actions.

[0203] "Urgency" is a measure of the seriousness of a situation or problem and the degree to which a response is necessary.

[0204] "Notification" is the act of conveying specific information or instructions to interested parties.

[0205] "Guard" refers to a security guard or security service deployed to protect the security of a particular place or object.

[0206] A "police agency" is a government agency established to maintain public order and safety.

[0207] System Overview

[0208] The present invention relates to a system for enhancing security in offices and homes using voice data, which mainly comprises means for receiving, analyzing, and evaluating the voice data, and proposing or implementing emergency response measures.

[0209] System Configuration

[0210] The system consists of the following main components:

[0211] 1. Terminal (smartphone, smart glasses, security robot): The user is responsible for recording voice and sending the data to the server.

[0212] 2. Server: This is the central system that stores and analyzes the received voice data, runs the generative AI model, and evaluates it, proposing or implementing countermeasures.

[0213] 3. Database (MySQL (registered trademark), PostgreSQL): Serves as data storage for saving voice data and analysis results.

[0214] Explanation of program processing

[0215] Recording and sending audio

[0216] If a user senses a physical security threat or an abnormality, they report it by voice into the device (smartphone, smart glasses, robot), which then records the voice and sends the recorded data to the server.

[0217] Receiving and storing audio data

[0218] The server receives the voice data sent from the terminal and stores it in a database for subsequent processing.

[0219] Analysis of audio data

[0220] The server inputs the saved voice data into a generative AI model for analysis. The AI ​​model extracts voice tension, abnormal sounds, alarm sounds, etc. Features are extracted from the analyzed data and recorded in a database.

[0221] Security Status Assessment

[0222] The server evaluates the security situation based on the extracted features. The evaluation also involves comparison with past voice data. Based on the evaluation results, the urgency of the security situation is classified (normal, caution, emergency).

[0223] Proposing or implementing measures

[0224] The server will propose or implement appropriate measures based on the evaluation results. Specific actions include:

[0225] Caution: Notify the user and call for caution.

[0226] In case of emergency: Contact the police or security company and arrange for appropriate response. Inform users to evacuate immediately.

[0227] Specific examples

[0228] Example 1: When caution is required

[0229] A user hears an unusual sound around the house and reports it to the device. The device records the sound and sends it to the server. The server analyzes the sound data and determines that the situation requires attention. The server then sends a notification to the user to alert them.

[0230] Example 2: When emergency response is required

[0231] The user suspects an intruder in their home and reports the incident to the device. The device records the audio and sends it to the server. The server analyzes the audio data and determines that the situation requires emergency response. The server immediately contacts the police and security companies and arranges for countermeasures. The user is notified to evacuate immediately.

[0232] Example prompts for generative AI models

[0233] Enter the following prompts into the generative AI model:

[0234] "The server has recorded some unusual sounds around the house. Please analyze whether this sound should be a cause for alarm or ignored."

[0235] "We have recorded a possible intruder on our server. Please analyze whether this audio requires immediate attention or should be ignored."

[0236] In this way, the present invention is a system that can use voice data to efficiently and quickly detect physical security threats and anomalies and provide appropriate responses.

[0237] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0238] Step 1:

[0239] When a user senses a physical security threat or anomaly, they report the details by voice into a device such as a smartphone, smart glasses, or security robot. The input is voice data, and the output is a recorded voice file. The device then performs the specific action of recording this voice.

[0240] Step 2:

[0241] When the recording is finished, the device sends the audio data to the server. The input is the recorded audio file, and the output is the audio data transferred to the server. The server receives and stores this data.

[0242] Step 3:

[0243] The server inputs the received voice data into the generative AI model for analysis. The input is the voice data stored on the server, and the output is the analyzed features (voice tension, abnormal sounds, alarm sounds, etc.). The server analyzes the data and performs specific operations to extract important features.

[0244] Step 4:

[0245] After features are extracted from the voice data analyzed by the generative AI model, the server evaluates the security situation based on these features. The input is the extracted features, and the output is the evaluation result (normal, caution, emergency). The server compares it with past data and performs specific actions to determine the security situation.

[0246] Step 5:

[0247] Based on the evaluation results, the server proposes or automatically executes appropriate countermeasures. The input is the evaluation results, and the output is notifications and the execution of countermeasures. Specifically, if caution is required, a notification is sent to the user to warn them, and in the case of an emergency, the police or security company is contacted and a response is arranged. The server performs the specific operations of generating notifications and arranging contact.

[0248] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0249] System Overview

[0250] The present invention relates to a system that uses voice data to automatically evaluate a patient's emotional state in addition to their physical condition and symptoms, and proposes appropriate medical treatment. This system is composed of various means, including receiving, analyzing, and evaluating voice data, proposing medical treatment, and an emotion engine.

[0251] System Configuration

[0252] The system consists of the following main components:

[0253] 1. Terminal: This includes the device used by the user (e.g., a smartphone). The user is responsible for recording audio and sending the data to the server.

[0254] 2. Server: This is the central system that stores and analyzes the received voice data, runs the generative AI model, and makes evaluations and suggestions.

[0255] 3. Emotion engine: Analyzes the user's emotional state from voice data and reflects the results in the evaluation of physical condition and symptoms.

[0256] 4. Database: Serves as data storage for saving voice data and analysis results.

[0257] Explanation of program processing

[0258] Recording and sending audio

[0259] 1. User: When a user feels unwell, they talk to the device about their symptoms and feelings.

[0260] 2. Device: The device (e.g., a smartphone) starts a recording application and records what the user says. When the recording is finished, the audio data is sent to the server.

[0261] Receiving and storing audio data

[0262] 1. Server: Receives the voice data sent from the terminal.

[0263] 2. Server: Stores the received voice data in a database.

[0264] Analysis of audio data

[0265] 1. Server: Inputs the stored voice data into the generative AI model.

[0266] 2. Server (generative AI model): The generative AI model analyzes the audio data and extracts features such as breathing, voice intonation, and speaking rate.

[0267] 3. Server (Emotion Engine): The emotion engine analyzes the voice data and identifies the user's emotional state.

[0268] 4. Server (analysis system): Records features and emotional states based on the analyzed data.

[0269] Assessment of the patient's physical condition and symptoms

[0270] 1. Server: Evaluates the patient's physical condition and symptoms based on the extracted features and emotional state. This also compares with past voice data.

[0271] 2. Server: As a result of the assessment, classify the patient's physical condition and the urgency of their symptoms (mild, severe, urgent).

[0272] Suggested actions needed

[0273] 1. Server: Selects appropriate medical response based on the evaluation results and the user's emotional state.

[0274] Mild cases:

[0275] 1. Server: Sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[0276] Severe cases:

[0277] 1. Server: Contacts the nearest doctor and arranges for a doctor to be dispatched. The user is notified of the doctor's visit date and time.

[0278] In case of emergency:

[0279] 1. Server: Contacts emergency services and dispatches an ambulance based on the user's location. The server notifies the user that an emergency response is required.

[0280] Specific examples

[0281] Example 1: Mild illness

[0282] 1. User: Feeling unwell, they talk to the device about their symptoms, such as a slight headache or feeling tired.

[0283] 2. Device: Records and sends the audio data to the server.

[0284] 3. Server: Analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the stress level is low.

[0285] 4. Server: Sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[0286] Example 2: Urgent illness

[0287] 1. User: Feeling a sudden deterioration in their physical condition, they talk about their symptoms into the device, such as chest pain or shortness of breath.

[0288] 2. Device: Records and sends the audio data to the server.

[0289] 3. Server: Analyzes the voice data and detects high levels of urgency. At the same time, the emotion engine recognizes strong anxiety and fear.

[0290] 4. Server: Immediately contacts emergency services and dispatches an ambulance. The user is notified that emergency response is required.

[0291] In this way, the system of the present invention can utilize voice data and take emotional state into account to provide a more accurate and faster medical response.

[0292] The processing flow will be explained below.

[0293] Step 1:

[0294] User: If the user feels unwell, they talk to the device about their symptoms and feelings.

[0295] Step 2:

[0296] Terminal: The terminal (e.g., a smartphone) launches a recording application and records what the user says.

[0297] Step 3:

[0298] Device: Once recording is complete, the audio data is sent to the server.

[0299] Step 4:

[0300] Server: Receives the voice data sent from the terminal.

[0301] Step 5:

[0302] Server: Stores the received voice data in a database.

[0303] Step 6:

[0304] Server: Inputs the stored voice data into the generative AI model.

[0305] Step 7:

[0306] Server (generative AI model): The generative AI model analyzes the audio data and extracts features such as breathing, voice intonation, and speaking rate.

[0307] Step 8:

[0308] Server (Emotion Engine): The emotion engine analyzes the voice data and identifies the user's emotional state, which can include joy, sadness, anger, anxiety, etc.

[0309] Step 9:

[0310] Server (analysis system): Records features and emotional states based on the analyzed data.

[0311] Step 10:

[0312] Server: Evaluates the user's physical condition and symptoms based on features and emotional state, and compares them with past voice data.

[0313] Step 11:

[0314] Server: As a result of the evaluation, the server classifies the urgency of the user's physical condition and symptoms (mild, severe, urgent).

[0315] Step 12:

[0316] Server: Selects appropriate medical response based on the evaluation results and the user's emotional state.

[0317] Mild cases:

[0318] Step 13:

[0319] Server: Sends the prescription to the pharmacy and notifies the user to pick up the medication.

[0320] Severe cases:

[0321] Step 14:

[0322] Server: Contacts the nearest doctor and arranges for the doctor to be dispatched. The user is notified of the date and time of the doctor's visit.

[0323] In case of emergency:

[0324] Step 15:

[0325] Server: Contacts emergency services and arranges for an ambulance based on the user's location. The user is notified that an emergency response is required.

[0326] Specific examples

[0327] Example 1: Mild illness

[0328] Step 1:

[0329] User: Feeling unwell, he / she talks about his / her symptoms into the device, such as a slight headache or feeling tired.

[0330] Step 2:

[0331] Device: Records and sends the audio data to the server.

[0332] Step 3:

[0333] Server: Analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the stress level is low.

[0334] Step 4:

[0335] Server: Sends the prescription to the pharmacy and notifies the user to pick up the medication.

[0336] Example 2: Urgent illness

[0337] Step 1:

[0338] User: Feeling a sudden deterioration in their physical condition, they talk about their symptoms into the device, such as chest pain or shortness of breath.

[0339] Step 2:

[0340] Device: Records and sends the audio data to the server.

[0341] Step 3:

[0342] Server: Analyzes voice data to detect high levels of urgency. At the same time, the emotion engine recognizes strong anxiety and fear.

[0343] Step 4:

[0344] Server: Immediately contacts emergency services and dispatches an ambulance. The user is notified that emergency response is required.

[0345] In this way, the system of the present invention can utilize voice data and take emotional state into account to provide a more accurate and faster medical response.

[0346] Example 2

[0347] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0348] In today's world, there is a demand for providing prompt and accurate medical treatment appropriate to a user's physical condition and symptoms. However, conventional systems have difficulty taking into account the user's emotional state in their assessment, which can result in inappropriate medical treatment. Furthermore, prompt treatment for urgent illnesses can be delayed, which can pose serious health risks. To solve these problems, it is necessary to use voice data to assess a user's physical condition and symptoms and propose appropriate medical treatment that also takes into account their emotional state.

[0349] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0350] In this invention, the server includes a means for the user to record voice data and send it to the server, a means for receiving and saving the voice data, a means for inputting the saved voice data into a generative AI model and extracting and analyzing features such as breathing, voice intonation, and speaking rate, a means for evaluating the user's physical condition, symptoms, and emotional state from the analyzed data, and a means for proposing appropriate medical treatment based on the evaluation results. This enables prompt and accurate medical treatment using voice data and taking emotional state into consideration.

[0351] "User" refers to an individual who uses this system.

[0352] "Server" refers to the central processing unit that receives, stores, analyzes, evaluates, and recommends medical responses to voice data.

[0353] "Audio data" refers to audio files in which a user records their physical condition, symptoms, and emotional state.

[0354] A "generative AI model" refers to an artificial intelligence model that analyzes voice data and extracts features such as breathing, voice intonation, and speaking rate.

[0355] An "emotion engine" is a system that analyzes a user's emotional state from voice data and uses the results to help evaluate their physical condition and symptoms.

[0356] "Means for receiving and storing voice data" refers to the function of receiving voice data sent from a terminal and storing it in a database or storage.

[0357] "Means for extracting and analyzing features" refers to the function of passing the received voice data through a generative AI model to extract features such as breathing, voice intonation, and speaking rate.

[0358] "Means for evaluation" refers to a function that evaluates the user's physical condition and symptoms based on the analyzed features and emotional state, and determines the urgency of the condition.

[0359] "Means for proposing medical treatment" refers to a function for proposing appropriate medical treatment to the user based on the evaluation results.

[0360] "Prescription transmission" refers to the act of transmitting prescription information to a pharmacy based on the evaluation results.

[0361] "Dispatch of a doctor" refers to the act of dispatching a doctor to the user based on the evaluation results and the degree of urgency.

[0362] "Contacting emergency services" refers to the act of automatically contacting emergency services in the event of an emergency based on the evaluation results to provide a prompt response.

[0363] MODE FOR CARRYING OUT THE INVENTION

[0364] The present invention provides a system that uses voice data from a user to analyze the user's physical condition, symptoms, and even emotional state, and suggests appropriate medical treatment. Specific embodiments for realizing the present invention will be described below.

[0365] System configuration

[0366] The system consists of the following main components:

[0367] 1. Terminal: This includes the device used by the user (e.g., a smartphone). The user records audio and sends the data to the server. Specific recording applications include the smartphone's built-in recording app and a custom app.

[0368] 2. Server: A central processing unit that stores and analyzes received voice data. It runs generative AI models and makes evaluations and recommendations. The server can be a high-performance cloud server or an on-premise server.

[0369] 3. Emotion engine: This engine analyzes the user's emotional state from voice data. The emotion engine uses machine learning algorithms to extract emotional features from voice data.

[0370] 4. Database: This serves as data storage for saving voice data and analysis results. The database can be an SQL-based database (e.g., MySQL, PostgreSQL) or a NoSQL database (e.g., MongoDB).

[0371] System Features

[0372] 1. Recording and sending audio

[0373] User: When a user feels unwell, they talk to the device about their symptoms and feelings. For example, they might say, "I have a headache and my body feels tired."

[0374] On the device, a recording application is used to record the user's voice. After recording, the voice data is sent to the server. The data is sent via an internet connection.

[0375] 2. Receiving and storing audio data

[0376] Server: Receives the voice data sent from the device and stores it in a database. The saved voice data is appended with a timestamp and user ID.

[0377] 3. Analysis of audio data

[0378] Server: The stored voice data is input into the generative AI model, which analyzes the voice data and extracts features such as breathing, voice intonation, and speaking rate.

[0379] Server (emotion engine): The emotion engine identifies the user's emotional state based on the analysis results. This emotional data is reflected in the evaluation of physical condition and symptoms.

[0380] 4. Assessment of physical condition and symptoms

[0381] Server: Based on the extracted features and emotional state, the server comprehensively evaluates the user's physical condition and symptoms. The evaluation also compares with past data and uses an algorithm to determine the urgency of the condition.

[0382] 5. Recommendation of necessary actions

[0383] Server: Based on the evaluation results, it proposes appropriate medical responses. For example, if the condition is mild, it sends a prescription to a pharmacy and notifies the user to pick up the medicine. If the condition is severe, it dispatches a doctor, and if it is an emergency, it contacts emergency services.

[0384] Specific examples

[0385] Example 1: Mild illness

[0386] 1. The user feels unwell and says to the device, "I have a slight headache."

[0387] 2. The device records the audio and sends it to the server.

[0388] 3. The server analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the patient is experiencing low stress.

[0389] 4. The server sends the prescription to the pharmacy and notifies the user.

[0390] Example 2: In the case of an emergency illness

[0391] 1. The user suddenly feels unwell and says to the device, "My chest hurts and I'm having trouble breathing."

[0392] 2. The device records the audio and sends it to the server.

[0393] 3. The server analyzes the voice data to detect high levels of urgency, and the emotion engine recognizes high levels of anxiety.

[0394] 4. The server immediately contacts emergency services and dispatches an ambulance, notifying the user that an emergency response is required and an ambulance is on the way.

[0395] Prompt Sentence Examples

[0396] Prompt for mild illness:

[0397] "The user will be asked to describe in voice that they are experiencing a slight headache or fatigue. The system will analyze the voice data and suggest appropriate medical treatment."

[0398] Emergency Illness Prompt:

[0399] "Please explain to the user in a voice that you suddenly felt chest pain or shortness of breath. Please analyze the voice data and quickly suggest the necessary emergency response."

[0400] In this way, the system of the present invention can utilize voice data and take emotional state into account to provide a more accurate and faster medical response.

[0401] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0402] Step 1:

[0403] User: Feeling unwell, the user talks about their symptoms and feelings into the device. Specifically, they input something like "I have a headache and my body feels tired." The input voice data is captured by the device's recording application. The output is an audio data file.

[0404] Step 2:

[0405] Device: Use a recording application to record the user's voice. Specifically, the user presses the "Start Recording" button, and then presses the "Stop Recording" button after finishing describing the symptoms. After recording is complete, the device saves the voice data in a digital file format (e.g., WAV format) and sends it to the server. The input is the user's voice, and the output is the voice data file sent to the server.

[0406] Step 3:

[0407] Server: Receives audio data sent from the device. Specifically, it receives audio files via HTTP requests. The audio data is stored in a database, and a timestamp and user ID are added to the stored data. The input is the sent audio data, and the output is the saved audio data file and its metadata.

[0408] Step 4:

[0409] Server: The saved voice data is input into the generative AI model and analyzed. Specifically, the voice data is spectrally analyzed and features such as breathing, voice intonation, and speaking rate are extracted. The input is the saved voice data, and the output is the extracted voice features.

[0410] Step 5:

[0411] Server (Emotion Engine): The emotion engine identifies the user's emotional state based on the analyzed voice data. Specifically, it uses a machine learning algorithm to identify emotional states such as stress and anxiety. The input is voice features, and the output is the identified emotional state.

[0412] Step 6:

[0413] Server: Evaluates the user's physical condition and symptoms based on the extracted features and emotional state. Specifically, it compares with past data and executes algorithms to evaluate the progression of symptoms. The input is voice features and emotion identification results, and the output is an evaluation of the user's physical condition and symptoms.

[0414] Step 7:

[0415] Server: Based on the evaluation results, selects and proposes appropriate medical responses. Specifically, if symptoms are mild, it sends a prescription to a pharmacy, if symptoms are severe, it dispatches a doctor, and if there is an emergency, it contacts emergency services. The input is the evaluation results of physical condition and symptoms, and the output is a medical response proposal.

[0416] Step 8:

[0417] Server: Notifies the user of the selected medical response. Specifically, it contacts the user via email or SMS and instructs them on the necessary response. The input is the medical response proposal, and the output is a notification message to the user.

[0418] (Application example 2)

[0419] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0420] Until now, there has been no system in brick-and-mortar stores that can quickly and accurately evaluate the physical and emotional state of customers and provide optimal medical care. As a result, customers have had to wait a long time to receive appropriate medical care, which could worsen their symptoms. To solve this problem, the present invention aims to provide a system that can analyze the voice of customers in brick-and-mortar stores, evaluate their physical and emotional state, and quickly propose and implement appropriate medical care.

[0421] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving and storing voice data, means for executing a generative AI model for analyzing the received voice data, means for evaluating the user's physical condition and emotional state from the analyzed data, and means for proposing and implementing appropriate medical treatment based on the evaluation results. This makes it possible to quickly evaluate the physical condition and emotional state of customers in a physical store and provide optimal medical treatment.

[0422] "Voice data" refers to data in which the contents of a user's speech are recorded in digital format.

[0423] A "generative AI model" is an algorithm that uses machine learning technology to analyze voice data and extract specific features and emotional states.

[0424] The "means for analyzing" refers to the means for inputting received voice data into a generative AI model and carrying out the process of analyzing the characteristics of the data.

[0425] "Physical and emotional state" refers to the user's physical health and psychological feelings and moods.

[0426] The "means for evaluation" is a means for carrying out a process of determining the user's physical condition and emotional state based on the analyzed data.

[0427] "Medical response" means taking action based on the user's physical or emotional state, including suggesting appropriate medication, recommending a doctor's appointment, or contacting emergency services.

[0428] The "means for proposing and implementing" is a means for carrying out the process of proposing specific medical measures to the user based on the evaluation results and implementing those measures.

[0429] A "physical store" refers to a physical store where a user actually visits and receives face-to-face service.

[0430] System Overview

[0431] This invention relates to a system that automatically evaluates the physical and emotional state of customers in brick-and-mortar stores and proposes and implements appropriate medical treatment. This system collects and analyzes voice data, evaluates their emotional state, and then proposes and implements the most appropriate medical treatment for the user.

[0432] System Configuration

[0433] The system consists of the following main components:

[0434] 1. Terminal: This includes a recording device (e.g., a tablet or dedicated microphone at the reception desk) installed in the physical store and used by customers. Customers talk to this terminal about their symptoms and condition.

[0435] 2. Server: This is the central system that receives and analyzes the voice data, runs the generative AI model, and makes evaluations and recommendations.

[0436] 3. Emotion engine: Analyzes the emotional state of customers from voice data and reflects the results in assessing their physical condition and symptoms.

[0437] 4. Database: Serves as data storage for saving voice data and analysis results.

[0438] How it works

[0439] Recording and sending audio

[0440] The user speaks about their symptoms and feelings into a device installed in a physical store. The device launches a recording application and records what is said. After recording, the audio data is sent to the server.

[0441] Receiving and storing audio data

[0442] The server receives the voice data sent from the terminal and stores it in a database.

[0443] Analysis of audio data

[0444] The server inputs the stored voice data into a generative AI model for analysis. The generative AI model extracts features such as breathing, voice intonation, speaking rate, and emotional state. An emotion engine analyzes the data and identifies the user's emotional state.

[0445] Assessment of physical condition and symptoms

[0446] The server evaluates the customer's physical condition and symptoms based on the extracted features and emotional state, and also compares this with past voice data.

[0447] Propose and take necessary actions

[0448] Based on the evaluation results and the user's emotional state, the server selects and implements appropriate medical responses, such as suggesting medication, recommending a doctor's visit, or contacting emergency services.

[0449] Specific examples

[0450] Example 1: Mild illness

[0451] 1. User: Complains of a slight headache and fatigue and describes his symptoms into the terminal.

[0452] 2. Server: Analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the stress level is low.

[0453] 3. Server: Display a notification on the terminal to customers saying, "Please use over-the-counter headache medicine."

[0454] Example 2: Urgent illness

[0455] 1. User: Complains of chest pain and shortness of breath and describes his symptoms into the terminal.

[0456] 2. Server: Analyzes the voice data and detects high levels of urgency. At the same time, the emotion engine recognizes strong anxiety and fear.

[0457] 3. Server: Immediately dispatch emergency services and display a notification to the customer on their device saying, "Emergency response required. An ambulance is on the way."

[0458] Prompt Sentence Examples

[0459] User Voice:

[0460] "I've been having terrible headaches lately and I'm feeling stressed."

[0461] System response:

[0462] Emotional state: Stressed, high

[0463] Symptom assessment: mild

[0464] Suggestion: "Take advantage of over-the-counter headache medication. I recommend relaxation techniques to help relieve stress."

[0465] In this way, the system of the present invention can quickly assess the physical and emotional state of customers in a physical store and provide optimal medical care.

[0466] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0467] Step 1:

[0468] The user talks about their symptoms and feelings into a device installed in a physical store. The device launches a recording application and records the user's voice. When the recording is finished, the device sends the voice data (input) to the server (output).

[0469] Step 2:

[0470] The server receives the voice data sent from the terminal and stores it in a database. Specifically, after receiving the voice data (input), it records (outputs) it in the database.

[0471] Step 3:

[0472] The server inputs the saved voice data into the generative AI model for analysis. The server analyzes the voice data (input) and extracts features such as breathing, voice intonation, speaking rate, and emotional state (output). Specifically, it runs the generative AI model and obtains the results of analyzing each feature.

[0473] Step 4:

[0474] The server uses an emotion engine to identify the user's emotional state from the analyzed data. The emotion engine performs data calculations based on the input voice features and evaluates (outputs) the user's emotional state.

[0475] Step 5:

[0476] The server evaluates the user's physical condition and symptoms based on the features and emotional state extracted from the voice data. It compares the current condition with past data and evaluates it. In this process, the server derives the evaluation results (output) of the physical condition and symptoms based on the features and emotional state (input).

[0477] Step 6:

[0478] The server selects and implements appropriate medical responses based on the evaluation results and the user's emotional state. Specifically, it suggests medication, recommends a doctor's consultation, or contacts emergency services depending on the user's physical condition and the urgency of the symptoms. The input in this step is the evaluation results and emotional state, and the output is the proposal or implementation of specific medical responses.

[0479] Step 7:

[0480] The server notifies the user either through the server itself or the terminal. For example, if a medication is suggested, a notification saying "Please use over-the-counter headache medicine" is displayed on the terminal. In this step, a notification (output) to the user is created based on the medical treatment suggestion (input).

[0481] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0482] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search<url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0483] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0484] [Second embodiment]

[0485] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0486] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0487] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0488] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0489] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0490] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0491] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0492] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0493] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0494] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0495] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0496] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0497] System Overview

[0498] The present invention relates to a system that automatically evaluates a patient's physical condition and symptoms using voice data and proposes appropriate medical treatment. This system is mainly composed of means for receiving, analyzing, and evaluating the voice data and proposing medical treatment.

[0499] System Configuration

[0500] The system consists of the following main components:

[0501] 1. Terminal: This includes the device used by the user (e.g., a smartphone). The user is responsible for recording audio and sending the data to the server.

[0502] 2. Server: This is the central system that stores and analyzes the received voice data, runs the generative AI model, and makes evaluations and suggestions.

[0503] 3. Database: Serves as data storage for saving voice data and analysis results.

[0504] Explanation of program processing

[0505] Recording and sending audio

[0506] 1. User: If the user feels unwell, they speak to the terminal about their symptoms.

[0507] 2. Device: The device (e.g., a smartphone) starts a recording application and records the user's voice. When the recording is finished, the voice data is sent to the server.

[0508] Receiving and storing audio data

[0509] 1. Server: Receives the voice data sent from the device and stores it in a database.

[0510] Analysis of audio data

[0511] 1. Server: The stored voice data is input into the generative AI model for analysis. The generative AI model extracts information such as breathing, voice intonation, and speaking rate.

[0512] 2. Server: Extracts features from the analyzed data and records them.

[0513] Assessment of the patient's physical condition and symptoms

[0514] 1. Server: Evaluates the patient's physical condition and symptoms based on the extracted features. The evaluation also involves comparison with past voice data.

[0515] 2. Server: Based on the assessment results, classify the patient's physical condition and the urgency of their symptoms (mild, severe, urgent).

[0516] Suggested actions needed

[0517] 1. Server: Based on the evaluation results, propose appropriate medical responses. Specific responses are as follows:

[0518] Mild cases: Send a prescription to the pharmacy and notify the user to pick up the medication.

[0519] In severe cases: The nearest doctor will be contacted and dispatched. The user will be notified of the date and time of the doctor's visit.

[0520] In case of emergency: Call emergency services and arrange for an ambulance. Notify the user that immediate emergency response is required.

[0521] Specific examples

[0522] Example 1: Mild illness

[0523] 1. User: Feeling unwell, speaks about symptoms into the terminal.

[0524] 2. Device: Records and sends the audio data to the server.

[0525] 3. Server: Analyzes the voice data and evaluates the symptoms as mild.

[0526] 4. Server: Sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[0527] Example 2: Urgent illness

[0528] 1. User: Feeling a sudden deterioration in their health, they talk about their symptoms into the device.

[0529] 2. Device: Records and sends the audio data to the server.

[0530] 3. Server: Analyzes the voice data and detects high levels of urgency.

[0531] 4. Server: Immediately contacts emergency services and dispatches an ambulance. The user is notified that emergency response is required.

[0532] In this way, the present invention is a system that can efficiently and quickly evaluate a patient's physical condition and symptoms using voice data and provide appropriate medical treatment.

[0533] The processing flow will be explained below.

[0534] Step 1:

[0535] User: If the user feels unwell, they talk to the device about their symptoms.

[0536] Step 2:

[0537] Terminal: The terminal (e.g., a smartphone) launches a recording application and records what the user says.

[0538] Step 3:

[0539] Device: Once recording is complete, the audio data is sent to the server.

[0540] Step 4:

[0541] Server: Receives the voice data sent from the terminal.

[0542] Step 5:

[0543] Server: Stores the received voice data in a database.

[0544] Step 6:

[0545] Server: Inputs the stored voice data into the generative AI model.

[0546] Step 7:

[0547] Server (generative AI model): The generative AI model analyzes the audio data and extracts features such as breathing, voice intonation, and speaking rate.

[0548] Step 8:

[0549] Server (analysis system): Records features based on the analyzed data.

[0550] Step 9:

[0551] Server: Evaluates the patient's physical condition and symptoms based on the features. This also compares the results with past voice data.

[0552] Step 10:

[0553] Server: As a result of the evaluation, classify the patient's physical condition and the urgency of their symptoms (mild, severe, urgent).

[0554] Step 11:

[0555] Server: Selects appropriate medical response based on the evaluation results.

[0556] Step 12:

[0557] Server: If the condition is mild, it sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[0558] Step 13:

[0559] Server: If the condition is severe, contact the nearest doctor and arrange for a doctor to be dispatched. The user is notified of the date and time of the doctor's visit.

[0560] Step 14:

[0561] Server: In case of an emergency, contacts emergency services and arranges for an ambulance based on the user's location information. The user is notified that emergency response is required.

[0562] Through the above steps, the system of the present invention can utilize voice data to quickly evaluate the patient's physical condition and symptoms and provide appropriate medical care.

[0563] Example 1

[0564] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0565] Conventional medical systems have struggled to remotely and quickly and accurately assess a patient's physical condition and symptoms, and to propose appropriate medical treatment. They also struggled to respond immediately to particularly urgent symptoms, posing a significant risk to the patient's health and safety. Furthermore, they lacked advanced technology for analyzing voice data and extracting features, making it impossible to accurately identify changes in physical condition.

[0566] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0567] In this invention, the server includes means for receiving and storing voice data, means for executing a generative AI model to analyze the received voice data, means for extracting features from the analyzed data and recording them in a database, means for evaluating the patient's physical condition and symptoms from the extracted features and comparing them with past data, means for classifying the urgency of the symptoms based on the evaluation results, and means for proposing appropriate medical responses based on the evaluation results. This makes it possible to quickly and accurately evaluate the patient's physical condition and symptoms and propose appropriate medical measures. Specific responses can include automatically sending prescriptions, dispatching medical personnel, and contacting emergency services.

[0568] "Audio data" refers to data in which sound is recorded and stored in digital form.

[0569] "Reception" is the process of receiving data or signals sent from the outside.

[0570] "Storage" means storing the received data in a database or storage so that it can be accessed later if necessary.

[0571] A "generative AI model" is an artificial intelligence model that has been trained using machine learning or deep learning to perform specific tasks automatically.

[0572] "Analysis" is the processing of data and the extraction of meaningful information.

[0573] A "feature" is a numerical value or index extracted from data that represents a specific pattern or characteristic.

[0574] A "database" is an electronic data management system that stores data systematically and enables efficient searching and manipulation.

[0575] "Evaluation" refers to determining the content or state of data based on specific criteria.

[0576] "Urgency" is a scale that indicates the severity of the evaluated symptoms and the degree of need for response.

[0577] "Proposals" refer to the presentation of optimal actions or measures based on the evaluation results.

[0578] "Medical response" refers to medical services and treatments provided according to the patient's physical condition and symptoms.

[0579] MODE FOR CARRYING OUT THE INVENTION

[0580] The present invention relates to a system that automatically evaluates a patient's physical condition and symptoms using voice data and proposes appropriate medical treatment. This system mainly includes means for receiving, analyzing, and evaluating the voice data and proposing medical treatment.

[0581] System Overview

[0582] The system consists of the following main components:

[0583] 1. Terminal (e.g., smartphone): A device that allows a user to record audio and send the data to a server.

[0584] 2. Server: A central system for storing and analyzing received voice data, running generative AI models, and providing a means for evaluation and recommendations.

[0585] 3. Database: A storage system that stores voice data and analysis results.

[0586] Recording and sending audio

[0587] (User): When a user feels unwell, they can talk about their symptoms using a device such as a smartphone. For example, they can say, "I have chest pain."

[0588] (Device): The device (smartphone) activates the voice recording function and records the user's voice. When the recording is finished, the voice data is automatically sent to the server using the HTTPS protocol.

[0589] Receiving and storing audio data

[0590] (Server): The server receives the voice data sent from the device. The received voice data is stored in a temporary storage area, and once it is confirmed that it has been received completely, it is permanently stored in the database.

[0591] Analysis of audio data

[0592] (Server): The server inputs the saved voice data into a generative AI model (e.g., an artificial intelligence model trained using machine learning or deep learning) and begins the analysis process. Specifically, the AI ​​model analyzes features such as breathing, voice intonation, and speaking rate from the voice data.

[0593] (Server): Extracts features from the analyzed data and records them in a database. These features are basic data for evaluating the patient's physical condition.

[0594] Assessment of the patient's physical condition and symptoms

[0595] (Server): Evaluates the patient's physical condition and symptoms based on the extracted features. Executes evaluation procedures to compare with past voice data and analyze changes. For example, compares past voice data with current data and uses a machine learning algorithm to determine whether there are any abnormal changes.

[0596] (Server): Based on the evaluation results, the urgency of the symptoms is classified into three categories: mild, severe, and urgent.

[0597] Suggested actions needed

[0598] (Server): Based on the evaluation results, the following appropriate medical responses are proposed:

[0599] For mild cases: Email the prescription to the nearest pharmacy and notify the user to collect the medication. For example, send the prescription information to the pharmacy's email address.

[0600] In severe cases: Contact the nearest healthcare professional and arrange for a visit. Notify the user about the date and time of the healthcare professional's visit, for example by contacting the healthcare professional using their contact information.

[0601] In case of emergency: Contact emergency services and dispatch an ambulance. The user is notified immediately and informed that emergency response is required, for example by auto-dialing an emergency number and providing location and symptom information.

[0602] Specific examples

[0603] Example 1: Mild illness

[0604] (User): The user experiences a slight headache and says to their smartphone, "My head hurts a little."

[0605] (Device): The recording application records the audio and sends the audio data to the server.

[0606] (Server): Receives the voice data, analyzes it using a generative AI model, and determines that the symptoms are mild.

[0607] (Server): Sends the prescription to the pharmacy and notifies the user that the medicine has been received. For example, the server sends the prescription information to the pharmacy's email address.

[0608] Example 2: Urgent illness

[0609] (User): Feeling sudden chest pain, he says into his smartphone, "My chest hurts, I can't breathe."

[0610] (Device): The recording application records the audio and sends the audio data to the server.

[0611] (Server): Receives voice data, analyzes it using a generative AI model, and determines the urgency of the data.

[0612] (Server): Contacts emergency services and dispatches an ambulance. The user is notified immediately and informed that emergency response is required. For example, by automatically dialing an emergency number and providing location and symptom information.

[0613] Prompt Sentence Examples

[0614] "Analyze breathing, voice intonation, and speech rate to determine the patient's health condition. For example, analyze the audio data of 'I have a severe stomachache' and assess the urgency of the situation."

[0615] In this way, the present invention is a system that utilizes voice data to quickly and accurately evaluate a patient's physical condition and symptoms, and provide appropriate medical care.

[0616] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0617] Step 1:

[0618] (Input): If the user feels unwell, they can speak to the terminal about their symptoms. For example, they can say, "My chest hurts."

[0619] (Action): The device activates the voice recording function and records the user's voice.

[0620] (Output): Recorded audio data.

[0621] Step 2:

[0622] (Input): Recorded audio data.

[0623] (Operation): After the device finishes recording, it sends the audio data to the server using the HTTPS protocol.

[0624] (Output): The audio data sent to the server.

[0625] Step 3:

[0626] (Input): Audio data sent from the device.

[0627] (Operation): The server receives the voice data and stores it in a temporary storage area. Once reception is complete, the voice data is permanently stored in the database.

[0628] (Output): The audio data stored in the database.

[0629] Step 4:

[0630] (Input): Audio data stored in the database.

[0631] (Operation): The server inputs the voice data into the generative AI model and performs an analysis process. The generative AI model analyzes features such as breathing, voice intonation, and speaking rate from the voice data.

[0632] (Output): Parsed features.

[0633] Step 5:

[0634] (Input): Parsed features.

[0635] (Operation): The server extracts features and records them in a database.

[0636] (Output): Features recorded in the database.

[0637] Step 6:

[0638] (Input): Features recorded in the database.

[0639] (Operation): The server evaluates the patient's physical condition and symptoms based on the extracted features, compares them with past data, and analyzes changes. Specifically, it uses machine learning algorithms to identify abnormal changes.

[0640] (Output): Evaluation results (physical condition evaluation and urgency determination).

[0641] Step 7:

[0642] (Input): Evaluation result.

[0643] (Operation): Based on the evaluation results, the server classifies the urgency of the symptoms into three categories: mild, severe, and urgent.

[0644] (Output): Urgency classification result.

[0645] Step 8:

[0646] (Input): Urgency classification result.

[0647] (Operation): The server suggests appropriate medical responses, specifically sending a prescription to a pharmacy if the condition is mild, dispatching medical personnel if the condition is severe, and contacting emergency services to dispatch an ambulance if the condition is urgent.

[0648] (Output): Notification to the user and medical response.

[0649] This enables the system to use voice data to quickly and accurately assess a patient's physical condition and symptoms, enabling appropriate medical treatment.

[0650] (Application example 1)

[0651] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0652] Conventional security systems have had difficulty quickly and accurately detecting physical security threats and anomalies. Furthermore, delayed response in emergencies can lead to serious damage to human life and property. To solve these problems, there is a need for a system that can analyze voice data, automatically assess security situations, and quickly take appropriate action.

[0653] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0654] In this invention, the server includes means for receiving and storing voice data, means for executing a generative AI model to analyze the received voice data, means for evaluating the target situation or abnormality from the analyzed data, and means for proposing or implementing appropriate countermeasures based on the evaluation results, thereby enabling quick and accurate detection of physical security threats and abnormalities and prompt implementation of appropriate countermeasures.

[0655] "Audio data" means digital or analog information that is an electronic recording of sound.

[0656] "Reception" is the act of a specific device or system taking in data or signals from the outside.

[0657] A "generative AI model" is an artificial intelligence algorithm or network designed to analyze data and make predictions.

[0658] "Analysis" is the process of examining data or information in detail to understand its structure and meaning.

[0659] "Subject" refers to an object or situation that is the subject of observation or analysis under specific conditions or circumstances.

[0660] A "situation" refers to the state or condition of an environment or event at a given moment.

[0661] An "abnormality" is a problem or malfunction that deviates from normal conditions or norms.

[0662] "Evaluation" is the act of making judgments or analyses based on data or information in accordance with specific criteria.

[0663] "Countermeasures" refer to specific actions or measures taken to address a problem or abnormality.

[0664] A "suggestion" is a recommendation for a particular action or solution.

[0665] "Implementation" means actually carrying out the proposed measures or actions.

[0666] "Urgency" is a measure of the seriousness of a situation or problem and the degree to which a response is necessary.

[0667] "Notification" is the act of conveying specific information or instructions to interested parties.

[0668] "Guard" refers to a security guard or security service deployed to protect the security of a particular place or object.

[0669] A "police agency" is a government agency established to maintain public order and safety.

[0670] System Overview

[0671] The present invention relates to a system for enhancing security in offices and homes using voice data, which mainly comprises means for receiving, analyzing, and evaluating the voice data, and proposing or implementing emergency response measures.

[0672] System Configuration

[0673] The system consists of the following main components:

[0674] 1. Terminal (smartphone, smart glasses, security robot): The user is responsible for recording voice and sending the data to the server.

[0675] 2. Server: This is the central system that stores and analyzes the received voice data, runs the generative AI model, and evaluates it, proposing or implementing countermeasures.

[0676] 3. Database (MySQL, PostgreSQL): Serves as data storage for saving voice data and analysis results.

[0677] Explanation of program processing

[0678] Recording and sending audio

[0679] If a user senses a physical security threat or an abnormality, they report it by voice into the device (smartphone, smart glasses, robot), which then records the voice and sends the recorded data to the server.

[0680] Receiving and storing audio data

[0681] The server receives the voice data sent from the terminal and stores it in a database for subsequent processing.

[0682] Analysis of audio data

[0683] The server inputs the saved voice data into a generative AI model for analysis. The AI ​​model extracts voice tension, abnormal sounds, alarm sounds, etc. Features are extracted from the analyzed data and recorded in a database.

[0684] Security Status Assessment

[0685] The server evaluates the security situation based on the extracted features. The evaluation also involves comparison with past voice data. Based on the evaluation results, the urgency of the security situation is classified (normal, caution, emergency).

[0686] Proposing or implementing measures

[0687] The server will propose or implement appropriate measures based on the evaluation results. Specific actions include:

[0688] Caution: Notify the user and call for caution.

[0689] In case of emergency: Contact the police or security company and arrange for appropriate response. Inform users to evacuate immediately.

[0690] Specific examples

[0691] Example 1: When caution is required

[0692] A user hears an unusual sound around the house and reports it to the device. The device records the sound and sends it to the server. The server analyzes the sound data and determines that the situation requires attention. The server then sends a notification to the user to alert them.

[0693] Example 2: When emergency response is required

[0694] The user suspects an intruder in their home and reports the incident to the device. The device records the audio and sends it to the server. The server analyzes the audio data and determines that the situation requires emergency response. The server immediately contacts the police and security companies and arranges for countermeasures. The user is notified to evacuate immediately.

[0695] Example prompts for generative AI models

[0696] Enter the following prompts into the generative AI model:

[0697] "The server has recorded some unusual sounds around the house. Please analyze whether this sound should be a cause for alarm or ignored."

[0698] "We have recorded a possible intruder on our server. Please analyze whether this audio requires immediate attention or should be ignored."

[0699] In this way, the present invention is a system that can use voice data to efficiently and quickly detect physical security threats and anomalies and provide appropriate responses.

[0700] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0701] Step 1:

[0702] When a user senses a physical security threat or anomaly, they report the details by voice into a device such as a smartphone, smart glasses, or security robot. The input is voice data, and the output is a recorded voice file. The device then performs the specific action of recording this voice.

[0703] Step 2:

[0704] When the recording is finished, the device sends the audio data to the server. The input is the recorded audio file, and the output is the audio data transferred to the server. The server receives and stores this data.

[0705] Step 3:

[0706] The server inputs the received voice data into the generative AI model for analysis. The input is the voice data stored on the server, and the output is the analyzed features (voice tension, abnormal sounds, alarm sounds, etc.). The server analyzes the data and performs specific operations to extract important features.

[0707] Step 4:

[0708] After features are extracted from the voice data analyzed by the generative AI model, the server evaluates the security situation based on these features. The input is the extracted features, and the output is the evaluation result (normal, caution, emergency). The server compares it with past data and performs specific actions to determine the security situation.

[0709] Step 5:

[0710] Based on the evaluation results, the server proposes or automatically executes appropriate countermeasures. The input is the evaluation results, and the output is notifications and the execution of countermeasures. Specifically, if caution is required, a notification is sent to the user to warn them, and in the case of an emergency, the police or security company is contacted and a response is arranged. The server performs the specific operations of generating notifications and arranging contact.

[0711] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0712] System Overview

[0713] The present invention relates to a system that uses voice data to automatically evaluate a patient's emotional state in addition to their physical condition and symptoms, and proposes appropriate medical treatment. This system is composed of various means, including receiving, analyzing, and evaluating voice data, proposing medical treatment, and an emotion engine.

[0714] System Configuration

[0715] The system consists of the following main components:

[0716] 1. Terminal: This includes the device used by the user (e.g., a smartphone). The user is responsible for recording audio and sending the data to the server.

[0717] 2. Server: This is the central system that stores and analyzes the received voice data, runs the generative AI model, and makes evaluations and suggestions.

[0718] 3. Emotion engine: Analyzes the user's emotional state from voice data and reflects the results in the evaluation of physical condition and symptoms.

[0719] 4. Database: Serves as data storage for saving voice data and analysis results.

[0720] Explanation of program processing

[0721] Recording and sending audio

[0722] 1. User: When a user feels unwell, they talk to the device about their symptoms and feelings.

[0723] 2. Device: The device (e.g., a smartphone) starts a recording application and records what the user says. When the recording is finished, the audio data is sent to the server.

[0724] Receiving and storing audio data

[0725] 1. Server: Receives the voice data sent from the terminal.

[0726] 2. Server: Stores the received voice data in a database.

[0727] Analysis of audio data

[0728] 1. Server: Inputs the stored voice data into the generative AI model.

[0729] 2. Server (generative AI model): The generative AI model analyzes the audio data and extracts features such as breathing, voice intonation, and speaking rate.

[0730] 3. Server (Emotion Engine): The emotion engine analyzes the voice data and identifies the user's emotional state.

[0731] 4. Server (analysis system): Records features and emotional states based on the analyzed data.

[0732] Assessment of the patient's physical condition and symptoms

[0733] 1. Server: Evaluates the patient's physical condition and symptoms based on the extracted features and emotional state. This also compares with past voice data.

[0734] 2. Server: As a result of the assessment, classify the patient's physical condition and the urgency of their symptoms (mild, severe, urgent).

[0735] Suggested actions needed

[0736] 1. Server: Selects appropriate medical response based on the evaluation results and the user's emotional state.

[0737] Mild cases:

[0738] 1. Server: Sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[0739] Severe cases:

[0740] 1. Server: Contacts the nearest doctor and arranges for a doctor to be dispatched. The user is notified of the doctor's visit date and time.

[0741] In case of emergency:

[0742] 1. Server: Contacts emergency services and dispatches an ambulance based on the user's location. The server notifies the user that an emergency response is required.

[0743] Specific examples

[0744] Example 1: Mild illness

[0745] 1. User: Feeling unwell, they talk to the device about their symptoms, such as a slight headache or feeling tired.

[0746] 2. Device: Records and sends the audio data to the server.

[0747] 3. Server: Analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the stress level is low.

[0748] 4. Server: Sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[0749] Example 2: Urgent illness

[0750] 1. User: Feeling a sudden deterioration in their physical condition, they talk about their symptoms into the device, such as chest pain or shortness of breath.

[0751] 2. Device: Records and sends the audio data to the server.

[0752] 3. Server: Analyzes the voice data and detects high levels of urgency. At the same time, the emotion engine recognizes strong anxiety and fear.

[0753] 4. Server: Immediately contacts emergency services and dispatches an ambulance. The user is notified that emergency response is required.

[0754] In this way, the system of the present invention can utilize voice data and take emotional state into account to provide a more accurate and faster medical response.

[0755] The processing flow will be explained below.

[0756] Step 1:

[0757] User: If the user feels unwell, they talk to the device about their symptoms and feelings.

[0758] Step 2:

[0759] Terminal: The terminal (e.g., a smartphone) launches a recording application and records what the user says.

[0760] Step 3:

[0761] Device: Once recording is complete, the audio data is sent to the server.

[0762] Step 4:

[0763] Server: Receives the voice data sent from the terminal.

[0764] Step 5:

[0765] Server: Stores the received voice data in a database.

[0766] Step 6:

[0767] Server: Inputs the stored voice data into the generative AI model.

[0768] Step 7:

[0769] Server (generative AI model): The generative AI model analyzes the audio data and extracts features such as breathing, voice intonation, and speaking rate.

[0770] Step 8:

[0771] Server (Emotion Engine): The emotion engine analyzes the voice data and identifies the user's emotional state, which can include joy, sadness, anger, anxiety, etc.

[0772] Step 9:

[0773] Server (analysis system): Records features and emotional states based on the analyzed data.

[0774] Step 10:

[0775] Server: Evaluates the user's physical condition and symptoms based on features and emotional state, and compares them with past voice data.

[0776] Step 11:

[0777] Server: As a result of the evaluation, the server classifies the urgency of the user's physical condition and symptoms (mild, severe, urgent).

[0778] Step 12:

[0779] Server: Selects appropriate medical response based on the evaluation results and the user's emotional state.

[0780] Mild cases:

[0781] Step 13:

[0782] Server: Sends the prescription to the pharmacy and notifies the user to pick up the medication.

[0783] Severe cases:

[0784] Step 14:

[0785] Server: Contacts the nearest doctor and arranges for the doctor to be dispatched. The user is notified of the date and time of the doctor's visit.

[0786] In case of emergency:

[0787] Step 15:

[0788] Server: Contacts emergency services and arranges for an ambulance based on the user's location. The user is notified that an emergency response is required.

[0789] Specific examples

[0790] Example 1: Mild illness

[0791] Step 1:

[0792] User: Feeling unwell, he / she talks about his / her symptoms into the device, such as a slight headache or feeling tired.

[0793] Step 2:

[0794] Device: Records and sends the audio data to the server.

[0795] Step 3:

[0796] Server: Analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the stress level is low.

[0797] Step 4:

[0798] Server: Sends the prescription to the pharmacy and notifies the user to pick up the medication.

[0799] Example 2: Urgent illness

[0800] Step 1:

[0801] User: Feeling a sudden deterioration in their physical condition, they talk about their symptoms into the device, such as chest pain or shortness of breath.

[0802] Step 2:

[0803] Device: Records and sends the audio data to the server.

[0804] Step 3:

[0805] Server: Analyzes voice data to detect high levels of urgency. At the same time, the emotion engine recognizes strong anxiety and fear.

[0806] Step 4:

[0807] Server: Immediately contacts emergency services and dispatches an ambulance. The user is notified that emergency response is required.

[0808] In this way, the system of the present invention can utilize voice data and take emotional state into account to provide a more accurate and faster medical response.

[0809] Example 2

[0810] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0811] In today's world, there is a demand for providing prompt and accurate medical treatment appropriate to a user's physical condition and symptoms. However, conventional systems have difficulty taking into account the user's emotional state in their assessment, which can result in inappropriate medical treatment. Furthermore, prompt treatment for urgent illnesses can be delayed, which can pose serious health risks. To solve these problems, it is necessary to use voice data to assess a user's physical condition and symptoms and propose appropriate medical treatment that also takes into account their emotional state.

[0812] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0813] In this invention, the server includes a means for the user to record voice data and send it to the server, a means for receiving and saving the voice data, a means for inputting the saved voice data into a generative AI model and extracting and analyzing features such as breathing, voice intonation, and speaking rate, a means for evaluating the user's physical condition, symptoms, and emotional state from the analyzed data, and a means for proposing appropriate medical treatment based on the evaluation results. This enables prompt and accurate medical treatment using voice data and taking emotional state into consideration.

[0814] "User" refers to an individual who uses this system.

[0815] "Server" refers to the central processing unit that receives, stores, analyzes, evaluates, and recommends medical responses to voice data.

[0816] "Audio data" refers to audio files in which a user records their physical condition, symptoms, and emotional state.

[0817] A "generative AI model" refers to an artificial intelligence model that analyzes voice data and extracts features such as breathing, voice intonation, and speaking rate.

[0818] An "emotion engine" is a system that analyzes a user's emotional state from voice data and uses the results to help evaluate their physical condition and symptoms.

[0819] "Means for receiving and storing voice data" refers to the function of receiving voice data sent from a terminal and storing it in a database or storage.

[0820] "Means for extracting and analyzing features" refers to the function of passing the received voice data through a generative AI model to extract features such as breathing, voice intonation, and speaking rate.

[0821] "Means for evaluation" refers to a function that evaluates the user's physical condition and symptoms based on the analyzed features and emotional state, and determines the urgency of the condition.

[0822] "Means for proposing medical treatment" refers to a function for proposing appropriate medical treatment to the user based on the evaluation results.

[0823] "Prescription transmission" refers to the act of transmitting prescription information to a pharmacy based on the evaluation results.

[0824] "Dispatch of a doctor" refers to the act of dispatching a doctor to the user based on the evaluation results and the degree of urgency.

[0825] "Contacting emergency services" refers to the act of automatically contacting emergency services in the event of an emergency based on the evaluation results to provide a prompt response.

[0826] MODE FOR CARRYING OUT THE INVENTION

[0827] The present invention provides a system that uses voice data from a user to analyze the user's physical condition, symptoms, and even emotional state, and suggests appropriate medical treatment. Specific embodiments for realizing the present invention will be described below.

[0828] System configuration

[0829] The system consists of the following main components:

[0830] 1. Terminal: This includes the device used by the user (e.g., a smartphone). The user records audio and sends the data to the server. Specific recording applications include the smartphone's built-in recording app and a custom app.

[0831] 2. Server: A central processing unit that stores and analyzes received voice data. It runs generative AI models and makes evaluations and recommendations. The server can be a high-performance cloud server or an on-premise server.

[0832] 3. Emotion engine: This engine analyzes the user's emotional state from voice data. The emotion engine uses machine learning algorithms to extract emotional features from voice data.

[0833] 4. Database: This serves as data storage for saving voice data and analysis results. The database can be an SQL-based database (e.g., MySQL, PostgreSQL) or a NoSQL database (e.g., MongoDB).

[0834] System Features

[0835] 1. Recording and sending audio

[0836] User: When a user feels unwell, they talk to the device about their symptoms and feelings. For example, they might say, "I have a headache and my body feels tired."

[0837] On the device, a recording application is used to record the user's voice. After recording, the voice data is sent to the server. The data is sent via an internet connection.

[0838] 2. Receiving and storing audio data

[0839] Server: Receives the voice data sent from the device and stores it in a database. The saved voice data is appended with a timestamp and user ID.

[0840] 3. Analysis of audio data

[0841] Server: The stored voice data is input into the generative AI model, which analyzes the voice data and extracts features such as breathing, voice intonation, and speaking rate.

[0842] Server (emotion engine): The emotion engine identifies the user's emotional state based on the analysis results. This emotional data is reflected in the evaluation of physical condition and symptoms.

[0843] 4. Assessment of physical condition and symptoms

[0844] Server: Based on the extracted features and emotional state, the server comprehensively evaluates the user's physical condition and symptoms. The evaluation also compares with past data and uses an algorithm to determine the urgency of the condition.

[0845] 5. Recommendation of necessary actions

[0846] Server: Based on the evaluation results, it proposes appropriate medical responses. For example, if the condition is mild, it sends a prescription to a pharmacy and notifies the user to pick up the medicine. If the condition is severe, it dispatches a doctor, and if it is an emergency, it contacts emergency services.

[0847] Specific examples

[0848] Example 1: Mild illness

[0849] 1. The user feels unwell and says to the device, "I have a slight headache."

[0850] 2. The device records the audio and sends it to the server.

[0851] 3. The server analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the patient is experiencing low stress.

[0852] 4. The server sends the prescription to the pharmacy and notifies the user.

[0853] Example 2: In the case of an emergency illness

[0854] 1. The user suddenly feels unwell and says to the device, "My chest hurts and I'm having trouble breathing."

[0855] 2. The device records the audio and sends it to the server.

[0856] 3. The server analyzes the voice data to detect high levels of urgency, and the emotion engine recognizes high levels of anxiety.

[0857] 4. The server immediately contacts emergency services and dispatches an ambulance, notifying the user that an emergency response is required and an ambulance is on the way.

[0858] Prompt Sentence Examples

[0859] Prompt for mild illness:

[0860] "The user will be asked to describe in voice that they are experiencing a slight headache or fatigue. The system will analyze the voice data and suggest appropriate medical treatment."

[0861] Emergency Illness Prompt:

[0862] "Please explain to the user in a voice that you suddenly felt chest pain or shortness of breath. Please analyze the voice data and quickly suggest the necessary emergency response."

[0863] In this way, the system of the present invention can utilize voice data and take emotional state into account to provide a more accurate and faster medical response.

[0864] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0865] Step 1:

[0866] User: Feeling unwell, the user talks about their symptoms and feelings into the device. Specifically, they input something like "I have a headache and my body feels tired." The input voice data is captured by the device's recording application. The output is an audio data file.

[0867] Step 2:

[0868] Device: Use a recording application to record the user's voice. Specifically, the user presses the "Start Recording" button, and then presses the "Stop Recording" button after finishing describing the symptoms. After recording is complete, the device saves the voice data in a digital file format (e.g., WAV format) and sends it to the server. The input is the user's voice, and the output is the voice data file sent to the server.

[0869] Step 3:

[0870] Server: Receives audio data sent from the device. Specifically, it receives audio files via HTTP requests. The audio data is stored in a database, and a timestamp and user ID are added to the stored data. The input is the sent audio data, and the output is the saved audio data file and its metadata.

[0871] Step 4:

[0872] Server: The saved voice data is input into the generative AI model and analyzed. Specifically, the voice data is spectrally analyzed and features such as breathing, voice intonation, and speaking rate are extracted. The input is the saved voice data, and the output is the extracted voice features.

[0873] Step 5:

[0874] Server (Emotion Engine): The emotion engine identifies the user's emotional state based on the analyzed voice data. Specifically, it uses a machine learning algorithm to identify emotional states such as stress and anxiety. The input is voice features, and the output is the identified emotional state.

[0875] Step 6:

[0876] Server: Evaluates the user's physical condition and symptoms based on the extracted features and emotional state. Specifically, it compares with past data and executes algorithms to evaluate the progression of symptoms. The input is voice features and emotion identification results, and the output is an evaluation of the user's physical condition and symptoms.

[0877] Step 7:

[0878] Server: Based on the evaluation results, selects and proposes appropriate medical responses. Specifically, if symptoms are mild, it sends a prescription to a pharmacy, if symptoms are severe, it dispatches a doctor, and if there is an emergency, it contacts emergency services. The input is the evaluation results of physical condition and symptoms, and the output is a medical response proposal.

[0879] Step 8:

[0880] Server: Notifies the user of the selected medical response. Specifically, it contacts the user via email or SMS and instructs them on the necessary response. The input is the medical response proposal, and the output is a notification message to the user.

[0881] (Application example 2)

[0882] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0883] Until now, there has been no system in brick-and-mortar stores that can quickly and accurately evaluate the physical and emotional state of customers and provide optimal medical care. As a result, customers have had to wait a long time to receive appropriate medical care, which could worsen their symptoms. To solve this problem, the present invention aims to provide a system that can analyze the voice of customers in brick-and-mortar stores, evaluate their physical and emotional state, and quickly propose and implement appropriate medical care.

[0884] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving and storing voice data, means for executing a generative AI model for analyzing the received voice data, means for evaluating the user's physical condition and emotional state from the analyzed data, and means for proposing and implementing appropriate medical treatment based on the evaluation results. This makes it possible to quickly evaluate the physical condition and emotional state of customers in a physical store and provide optimal medical treatment.

[0885] "Voice data" refers to data in which the contents of a user's speech are recorded in digital format.

[0886] A "generative AI model" is an algorithm that uses machine learning technology to analyze voice data and extract specific features and emotional states.

[0887] The "means for analyzing" refers to the means for inputting received voice data into a generative AI model and carrying out the process of analyzing the characteristics of the data.

[0888] "Physical and emotional state" refers to the user's physical health and psychological feelings and moods.

[0889] The "means for evaluation" is a means for carrying out a process of determining the user's physical condition and emotional state based on the analyzed data.

[0890] "Medical response" means taking action based on the user's physical or emotional state, including suggesting appropriate medication, recommending a doctor's appointment, or contacting emergency services.

[0891] The "means for proposing and implementing" is a means for carrying out the process of proposing specific medical measures to the user based on the evaluation results and implementing those measures.

[0892] A "physical store" refers to a physical store where a user actually visits and receives face-to-face service.

[0893] System Overview

[0894] This invention relates to a system that automatically evaluates the physical and emotional state of customers in brick-and-mortar stores and proposes and implements appropriate medical treatment. This system collects and analyzes voice data, evaluates their emotional state, and then proposes and implements the most appropriate medical treatment for the user.

[0895] System Configuration

[0896] The system consists of the following main components:

[0897] 1. Terminal: This includes a recording device (e.g., a tablet or dedicated microphone at the reception desk) installed in the physical store and used by customers. Customers talk to this terminal about their symptoms and condition.

[0898] 2. Server: This is the central system that receives and analyzes the voice data, runs the generative AI model, and makes evaluations and recommendations.

[0899] 3. Emotion engine: Analyzes the emotional state of customers from voice data and reflects the results in assessing their physical condition and symptoms.

[0900] 4. Database: Serves as data storage for saving voice data and analysis results.

[0901] How it works

[0902] Recording and sending audio

[0903] The user speaks about their symptoms and feelings into a device installed in a physical store. The device launches a recording application and records what is said. After recording, the audio data is sent to the server.

[0904] Receiving and storing audio data

[0905] The server receives the voice data sent from the terminal and stores it in a database.

[0906] Analysis of audio data

[0907] The server inputs the stored voice data into a generative AI model for analysis. The generative AI model extracts features such as breathing, voice intonation, speaking rate, and emotional state. An emotion engine analyzes the data and identifies the user's emotional state.

[0908] Assessment of physical condition and symptoms

[0909] The server evaluates the customer's physical condition and symptoms based on the extracted features and emotional state, and also compares this with past voice data.

[0910] Propose and take necessary actions

[0911] Based on the evaluation results and the user's emotional state, the server selects and implements appropriate medical responses, such as suggesting medication, recommending a doctor's visit, or contacting emergency services.

[0912] Specific examples

[0913] Example 1: Mild illness

[0914] 1. User: Complains of a slight headache and fatigue and describes his symptoms into the terminal.

[0915] 2. Server: Analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the stress level is low.

[0916] 3. Server: Display a notification on the terminal to customers saying, "Please use over-the-counter headache medicine."

[0917] Example 2: Urgent illness

[0918] 1. User: Complains of chest pain and shortness of breath and describes his symptoms into the terminal.

[0919] 2. Server: Analyzes the voice data and detects high levels of urgency. At the same time, the emotion engine recognizes strong anxiety and fear.

[0920] 3. Server: Immediately dispatch emergency services and display a notification to the customer on their device saying, "Emergency response required. An ambulance is on the way."

[0921] Prompt Sentence Examples

[0922] User Voice:

[0923] "I've been having terrible headaches lately and I'm feeling stressed."

[0924] System response:

[0925] Emotional state: Stressed, high

[0926] Symptom assessment: mild

[0927] Suggestion: "Take advantage of over-the-counter headache medication. I recommend relaxation techniques to help relieve stress."

[0928] In this way, the system of the present invention can quickly assess the physical and emotional state of customers in a physical store and provide optimal medical care.

[0929] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0930] Step 1:

[0931] The user talks about their symptoms and feelings into a device installed in a physical store. The device launches a recording application and records the user's voice. When the recording is finished, the device sends the voice data (input) to the server (output).

[0932] Step 2:

[0933] The server receives the voice data sent from the terminal and stores it in a database. Specifically, after receiving the voice data (input), it records (outputs) it in the database.

[0934] Step 3:

[0935] The server inputs the saved voice data into the generative AI model for analysis. The server analyzes the voice data (input) and extracts features such as breathing, voice intonation, speaking rate, and emotional state (output). Specifically, it runs the generative AI model and obtains the results of analyzing each feature.

[0936] Step 4:

[0937] The server uses an emotion engine to identify the user's emotional state from the analyzed data. The emotion engine performs data calculations based on the input voice features and evaluates (outputs) the user's emotional state.

[0938] Step 5:

[0939] The server evaluates the user's physical condition and symptoms based on the features and emotional state extracted from the voice data. It compares the current condition with past data and evaluates it. In this process, the server derives the evaluation results (output) of the physical condition and symptoms based on the features and emotional state (input).

[0940] Step 6:

[0941] The server selects and implements appropriate medical responses based on the evaluation results and the user's emotional state. Specifically, it suggests medication, recommends a doctor's consultation, or contacts emergency services depending on the user's physical condition and the urgency of the symptoms. The input in this step is the evaluation results and emotional state, and the output is the proposal or implementation of specific medical responses.

[0942] Step 7:

[0943] The server notifies the user either through the server itself or the terminal. For example, if a medication is suggested, a notification saying "Please use over-the-counter headache medicine" is displayed on the terminal. In this step, a notification (output) to the user is created based on the medical treatment suggestion (input).

[0944] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0945] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0946] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0947] [Third embodiment]

[0948] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0949] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0950] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0951] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0952] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0953] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0954] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0955] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0956] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0957] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0958] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0959] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0960] System Overview

[0961] The present invention relates to a system that automatically evaluates a patient's physical condition and symptoms using voice data and proposes appropriate medical treatment. This system is mainly composed of means for receiving, analyzing, and evaluating the voice data and proposing medical treatment.

[0962] System Configuration

[0963] The system consists of the following main components:

[0964] 1. Terminal: This includes the device used by the user (e.g., a smartphone). The user is responsible for recording audio and sending the data to the server.

[0965] 2. Server: This is the central system that stores and analyzes the received voice data, runs the generative AI model, and makes evaluations and suggestions.

[0966] 3. Database: Serves as data storage for saving voice data and analysis results.

[0967] Explanation of program processing

[0968] Recording and sending audio

[0969] 1. User: If the user feels unwell, they speak to the terminal about their symptoms.

[0970] 2. Device: The device (e.g., a smartphone) starts a recording application and records the user's voice. When the recording is finished, the voice data is sent to the server.

[0971] Receiving and storing audio data

[0972] 1. Server: Receives the voice data sent from the device and stores it in a database.

[0973] Analysis of audio data

[0974] 1. Server: The stored voice data is input into the generative AI model for analysis. The generative AI model extracts information such as breathing, voice intonation, and speaking rate.

[0975] 2. Server: Extracts features from the analyzed data and records them.

[0976] Assessment of the patient's physical condition and symptoms

[0977] 1. Server: Evaluates the patient's physical condition and symptoms based on the extracted features. The evaluation also involves comparison with past voice data.

[0978] 2. Server: Based on the assessment results, classify the patient's physical condition and the urgency of their symptoms (mild, severe, urgent).

[0979] Suggested actions needed

[0980] 1. Server: Based on the evaluation results, propose appropriate medical responses. Specific responses are as follows:

[0981] Mild cases: Send a prescription to the pharmacy and notify the user to pick up the medication.

[0982] In severe cases: The nearest doctor will be contacted and dispatched. The user will be notified of the date and time of the doctor's visit.

[0983] In case of emergency: Call emergency services and arrange for an ambulance. Notify the user that immediate emergency response is required.

[0984] Specific examples

[0985] Example 1: Mild illness

[0986] 1. User: Feeling unwell, speaks about symptoms into the terminal.

[0987] 2. Device: Records and sends the audio data to the server.

[0988] 3. Server: Analyzes the voice data and evaluates the symptoms as mild.

[0989] 4. Server: Sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[0990] Example 2: Urgent illness

[0991] 1. User: Feeling a sudden deterioration in their health, they talk about their symptoms into the device.

[0992] 2. Device: Records and sends the audio data to the server.

[0993] 3. Server: Analyzes the voice data and detects high levels of urgency.

[0994] 4. Server: Immediately contacts emergency services and dispatches an ambulance. The user is notified that emergency response is required.

[0995] In this way, the present invention is a system that can efficiently and quickly evaluate a patient's physical condition and symptoms using voice data and provide appropriate medical treatment.

[0996] The processing flow will be explained below.

[0997] Step 1:

[0998] User: If the user feels unwell, they talk to the device about their symptoms.

[0999] Step 2:

[1000] Terminal: The terminal (e.g., a smartphone) launches a recording application and records what the user says.

[1001] Step 3:

[1002] Device: Once recording is complete, the audio data is sent to the server.

[1003] Step 4:

[1004] Server: Receives the voice data sent from the terminal.

[1005] Step 5:

[1006] Server: Stores the received voice data in a database.

[1007] Step 6:

[1008] Server: Inputs the stored voice data into the generative AI model.

[1009] Step 7:

[1010] Server (generative AI model): The generative AI model analyzes the audio data and extracts features such as breathing, voice intonation, and speaking rate.

[1011] Step 8:

[1012] Server (analysis system): Records features based on the analyzed data.

[1013] Step 9:

[1014] Server: Evaluates the patient's physical condition and symptoms based on the features. This also compares the results with past voice data.

[1015] Step 10:

[1016] Server: As a result of the evaluation, classify the patient's physical condition and the urgency of their symptoms (mild, severe, urgent).

[1017] Step 11:

[1018] Server: Selects appropriate medical response based on the evaluation results.

[1019] Step 12:

[1020] Server: If the condition is mild, it sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[1021] Step 13:

[1022] Server: If the condition is severe, contact the nearest doctor and arrange for a doctor to be dispatched. The user is notified of the date and time of the doctor's visit.

[1023] Step 14:

[1024] Server: In case of an emergency, contacts emergency services and arranges for an ambulance based on the user's location information. The user is notified that emergency response is required.

[1025] Through the above steps, the system of the present invention can utilize voice data to quickly evaluate the patient's physical condition and symptoms and provide appropriate medical care.

[1026] Example 1

[1027] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1028] Conventional medical systems have struggled to remotely and quickly and accurately assess a patient's physical condition and symptoms, and to propose appropriate medical treatment. They also struggled to respond immediately to particularly urgent symptoms, posing a significant risk to the patient's health and safety. Furthermore, they lacked advanced technology for analyzing voice data and extracting features, making it impossible to accurately identify changes in physical condition.

[1029] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1030] In this invention, the server includes means for receiving and storing voice data, means for executing a generative AI model to analyze the received voice data, means for extracting features from the analyzed data and recording them in a database, means for evaluating the patient's physical condition and symptoms from the extracted features and comparing them with past data, means for classifying the urgency of the symptoms based on the evaluation results, and means for proposing appropriate medical responses based on the evaluation results. This makes it possible to quickly and accurately evaluate the patient's physical condition and symptoms and propose appropriate medical measures. Specific responses can include automatically sending prescriptions, dispatching medical personnel, and contacting emergency services.

[1031] "Audio data" refers to data in which sound is recorded and stored in digital form.

[1032] "Reception" is the process of receiving data or signals sent from the outside.

[1033] "Storage" means storing the received data in a database or storage so that it can be accessed later if necessary.

[1034] A "generative AI model" is an artificial intelligence model that has been trained using machine learning or deep learning to perform specific tasks automatically.

[1035] "Analysis" is the processing of data and the extraction of meaningful information.

[1036] A "feature" is a numerical value or index extracted from data that represents a specific pattern or characteristic.

[1037] A "database" is an electronic data management system that stores data systematically and enables efficient searching and manipulation.

[1038] "Evaluation" refers to determining the content or state of data based on specific criteria.

[1039] "Urgency" is a scale that indicates the severity of the evaluated symptoms and the degree of need for response.

[1040] "Proposals" refer to the presentation of optimal actions or measures based on the evaluation results.

[1041] "Medical response" refers to medical services and treatments provided according to the patient's physical condition and symptoms.

[1042] MODE FOR CARRYING OUT THE INVENTION

[1043] The present invention relates to a system that automatically evaluates a patient's physical condition and symptoms using voice data and proposes appropriate medical treatment. This system mainly includes means for receiving, analyzing, and evaluating the voice data and proposing medical treatment.

[1044] System Overview

[1045] The system consists of the following main components:

[1046] 1. Terminal (e.g., smartphone): A device that allows a user to record audio and send the data to a server.

[1047] 2. Server: A central system for storing and analyzing received voice data, running generative AI models, and providing a means for evaluation and recommendations.

[1048] 3. Database: A storage system that stores voice data and analysis results.

[1049] Recording and sending audio

[1050] (User): When a user feels unwell, they can talk about their symptoms using a device such as a smartphone. For example, they can say, "I have chest pain."

[1051] (Device): The device (smartphone) activates the voice recording function and records the user's voice. When the recording is finished, the voice data is automatically sent to the server using the HTTPS protocol.

[1052] Receiving and storing audio data

[1053] (Server): The server receives the voice data sent from the device. The received voice data is stored in a temporary storage area, and once it is confirmed that it has been received completely, it is permanently stored in the database.

[1054] Analysis of audio data

[1055] (Server): The server inputs the saved voice data into a generative AI model (e.g., an artificial intelligence model trained using machine learning or deep learning) and begins the analysis process. Specifically, the AI ​​model analyzes features such as breathing, voice intonation, and speaking rate from the voice data.

[1056] (Server): Extracts features from the analyzed data and records them in a database. These features are basic data for evaluating the patient's physical condition.

[1057] Assessment of the patient's physical condition and symptoms

[1058] (Server): Evaluates the patient's physical condition and symptoms based on the extracted features. Executes evaluation procedures to compare with past voice data and analyze changes. For example, compares past voice data with current data and uses a machine learning algorithm to determine whether there are any abnormal changes.

[1059] (Server): Based on the evaluation results, the urgency of the symptoms is classified into three categories: mild, severe, and urgent.

[1060] Suggested actions needed

[1061] (Server): Based on the evaluation results, the following appropriate medical responses are proposed:

[1062] For mild cases: Email the prescription to the nearest pharmacy and notify the user to collect the medication. For example, send the prescription information to the pharmacy's email address.

[1063] In severe cases: Contact the nearest healthcare professional and arrange for a visit. Notify the user about the date and time of the healthcare professional's visit, for example by contacting the healthcare professional using their contact information.

[1064] In case of emergency: Contact emergency services and dispatch an ambulance. The user is notified immediately and informed that emergency response is required, for example by auto-dialing an emergency number and providing location and symptom information.

[1065] Specific examples

[1066] Example 1: Mild illness

[1067] (User): The user experiences a slight headache and says to their smartphone, "My head hurts a little."

[1068] (Device): The recording application records the audio and sends the audio data to the server.

[1069] (Server): Receives the voice data, analyzes it using a generative AI model, and determines that the symptoms are mild.

[1070] (Server): Sends the prescription to the pharmacy and notifies the user that the medicine has been received. For example, the server sends the prescription information to the pharmacy's email address.

[1071] Example 2: Urgent illness

[1072] (User): Feeling sudden chest pain, he says into his smartphone, "My chest hurts, I can't breathe."

[1073] (Device): The recording application records the audio and sends the audio data to the server.

[1074] (Server): Receives voice data, analyzes it using a generative AI model, and determines the urgency of the data.

[1075] (Server): Contacts emergency services and dispatches an ambulance. The user is notified immediately and informed that emergency response is required. For example, by automatically dialing an emergency number and providing location and symptom information.

[1076] Prompt Sentence Examples

[1077] "Analyze breathing, voice intonation, and speech rate to determine the patient's health condition. For example, analyze the audio data of 'I have a severe stomachache' and assess the urgency of the situation."

[1078] In this way, the present invention is a system that utilizes voice data to quickly and accurately evaluate a patient's physical condition and symptoms, and provide appropriate medical care.

[1079] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1080] Step 1:

[1081] (Input): If the user feels unwell, they can speak to the terminal about their symptoms. For example, they can say, "My chest hurts."

[1082] (Action): The device activates the voice recording function and records the user's voice.

[1083] (Output): Recorded audio data.

[1084] Step 2:

[1085] (Input): Recorded audio data.

[1086] (Operation): After the device finishes recording, it sends the audio data to the server using the HTTPS protocol.

[1087] (Output): The audio data sent to the server.

[1088] Step 3:

[1089] (Input): Audio data sent from the device.

[1090] (Operation): The server receives the voice data and stores it in a temporary storage area. Once reception is complete, the voice data is permanently stored in the database.

[1091] (Output): The audio data stored in the database.

[1092] Step 4:

[1093] (Input): Audio data stored in the database.

[1094] (Operation): The server inputs the voice data into the generative AI model and performs an analysis process. The generative AI model analyzes features such as breathing, voice intonation, and speaking rate from the voice data.

[1095] (Output): Parsed features.

[1096] Step 5:

[1097] (Input): Parsed features.

[1098] (Operation): The server extracts features and records them in a database.

[1099] (Output): Features recorded in the database.

[1100] Step 6:

[1101] (Input): Features recorded in the database.

[1102] (Operation): The server evaluates the patient's physical condition and symptoms based on the extracted features, compares them with past data, and analyzes changes. Specifically, it uses machine learning algorithms to identify abnormal changes.

[1103] (Output): Evaluation results (physical condition evaluation and urgency determination).

[1104] Step 7:

[1105] (Input): Evaluation result.

[1106] (Operation): Based on the evaluation results, the server classifies the urgency of the symptoms into three categories: mild, severe, and urgent.

[1107] (Output): Urgency classification result.

[1108] Step 8:

[1109] (Input): Urgency classification result.

[1110] (Operation): The server suggests appropriate medical responses, specifically sending a prescription to a pharmacy if the condition is mild, dispatching medical personnel if the condition is severe, and contacting emergency services to dispatch an ambulance if the condition is urgent.

[1111] (Output): Notification to the user and medical response.

[1112] This enables the system to use voice data to quickly and accurately assess a patient's physical condition and symptoms, enabling appropriate medical treatment.

[1113] (Application example 1)

[1114] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1115] Conventional security systems have had difficulty quickly and accurately detecting physical security threats and anomalies. Furthermore, delayed response in emergencies can lead to serious damage to human life and property. To solve these problems, there is a need for a system that can analyze voice data, automatically assess security situations, and quickly take appropriate action.

[1116] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1117] In this invention, the server includes means for receiving and storing voice data, means for executing a generative AI model to analyze the received voice data, means for evaluating the target situation or abnormality from the analyzed data, and means for proposing or implementing appropriate countermeasures based on the evaluation results, thereby enabling quick and accurate detection of physical security threats and abnormalities and prompt implementation of appropriate countermeasures.

[1118] "Audio data" means digital or analog information that is an electronic recording of sound.

[1119] "Reception" is the act of a specific device or system taking in data or signals from the outside.

[1120] A "generative AI model" is an artificial intelligence algorithm or network designed to analyze data and make predictions.

[1121] "Analysis" is the process of examining data or information in detail to understand its structure and meaning.

[1122] "Subject" refers to an object or situation that is the subject of observation or analysis under specific conditions or circumstances.

[1123] A "situation" refers to the state or condition of an environment or event at a given moment.

[1124] An "abnormality" is a problem or malfunction that deviates from normal conditions or norms.

[1125] "Evaluation" is the act of making judgments or analyses based on data or information in accordance with specific criteria.

[1126] "Countermeasures" refer to specific actions or measures taken to address a problem or abnormality.

[1127] A "suggestion" is a recommendation for a particular action or solution.

[1128] "Implementation" means actually carrying out the proposed measures or actions.

[1129] "Urgency" is a measure of the seriousness of a situation or problem and the degree to which a response is necessary.

[1130] "Notification" is the act of conveying specific information or instructions to interested parties.

[1131] "Guard" refers to a security guard or security service deployed to protect the security of a particular place or object.

[1132] A "police agency" is a government agency established to maintain public order and safety.

[1133] System Overview

[1134] The present invention relates to a system for enhancing security in offices and homes using voice data, which mainly comprises means for receiving, analyzing, and evaluating the voice data, and proposing or implementing emergency response measures.

[1135] System Configuration

[1136] The system consists of the following main components:

[1137] 1. Terminal (smartphone, smart glasses, security robot): The user is responsible for recording voice and sending the data to the server.

[1138] 2. Server: This is the central system that stores and analyzes the received voice data, runs the generative AI model, and evaluates it, proposing or implementing countermeasures.

[1139] 3. Database (MySQL, PostgreSQL): Serves as data storage for saving voice data and analysis results.

[1140] Explanation of program processing

[1141] Recording and sending audio

[1142] If a user senses a physical security threat or an abnormality, they report it by voice into the device (smartphone, smart glasses, robot), which then records the voice and sends the recorded data to the server.

[1143] Receiving and storing audio data

[1144] The server receives the voice data sent from the terminal and stores it in a database for subsequent processing.

[1145] Analysis of audio data

[1146] The server inputs the saved voice data into a generative AI model for analysis. The AI ​​model extracts voice tension, abnormal sounds, alarm sounds, etc. Features are extracted from the analyzed data and recorded in a database.

[1147] Security Status Assessment

[1148] The server evaluates the security situation based on the extracted features. The evaluation also involves comparison with past voice data. Based on the evaluation results, the urgency of the security situation is classified (normal, caution, emergency).

[1149] Proposing or implementing measures

[1150] The server will propose or implement appropriate measures based on the evaluation results. Specific actions include:

[1151] Caution: Notify the user and call for caution.

[1152] In case of emergency: Contact the police or security company and arrange for appropriate response. Inform users to evacuate immediately.

[1153] Specific examples

[1154] Example 1: When caution is required

[1155] A user hears an unusual sound around the house and reports it to the device. The device records the sound and sends it to the server. The server analyzes the sound data and determines that the situation requires attention. The server then sends a notification to the user to alert them.

[1156] Example 2: When emergency response is required

[1157] The user suspects an intruder in their home and reports the incident to the device. The device records the audio and sends it to the server. The server analyzes the audio data and determines that the situation requires emergency response. The server immediately contacts the police and security companies and arranges for countermeasures. The user is notified to evacuate immediately.

[1158] Example prompts for generative AI models

[1159] Enter the following prompts into the generative AI model:

[1160] "The server has recorded some unusual sounds around the house. Please analyze whether this sound should be a cause for alarm or ignored."

[1161] "We have recorded a possible intruder on our server. Please analyze whether this audio requires immediate attention or should be ignored."

[1162] In this way, the present invention is a system that can use voice data to efficiently and quickly detect physical security threats and anomalies and provide appropriate responses.

[1163] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1164] Step 1:

[1165] When a user senses a physical security threat or anomaly, they report the details by voice into a device such as a smartphone, smart glasses, or security robot. The input is voice data, and the output is a recorded voice file. The device then performs the specific action of recording this voice.

[1166] Step 2:

[1167] When the recording is finished, the device sends the audio data to the server. The input is the recorded audio file, and the output is the audio data transferred to the server. The server receives and stores this data.

[1168] Step 3:

[1169] The server inputs the received voice data into the generative AI model for analysis. The input is the voice data stored on the server, and the output is the analyzed features (voice tension, abnormal sounds, alarm sounds, etc.). The server analyzes the data and performs specific operations to extract important features.

[1170] Step 4:

[1171] After features are extracted from the voice data analyzed by the generative AI model, the server evaluates the security situation based on these features. The input is the extracted features, and the output is the evaluation result (normal, caution, emergency). The server compares it with past data and performs specific actions to determine the security situation.

[1172] Step 5:

[1173] Based on the evaluation results, the server proposes or automatically executes appropriate countermeasures. The input is the evaluation results, and the output is notifications and the execution of countermeasures. Specifically, if caution is required, a notification is sent to the user to warn them, and in the case of an emergency, the police or security company is contacted and a response is arranged. The server performs the specific operations of generating notifications and arranging contact.

[1174] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1175] System Overview

[1176] The present invention relates to a system that uses voice data to automatically evaluate a patient's emotional state in addition to their physical condition and symptoms, and proposes appropriate medical treatment. This system is composed of various means, including receiving, analyzing, and evaluating voice data, proposing medical treatment, and an emotion engine.

[1177] System Configuration

[1178] The system consists of the following main components:

[1179] 1. Terminal: This includes the device used by the user (e.g., a smartphone). The user is responsible for recording audio and sending the data to the server.

[1180] 2. Server: This is the central system that stores and analyzes the received voice data, runs the generative AI model, and makes evaluations and suggestions.

[1181] 3. Emotion engine: Analyzes the user's emotional state from voice data and reflects the results in the evaluation of physical condition and symptoms.

[1182] 4. Database: Serves as data storage for saving voice data and analysis results.

[1183] Explanation of program processing

[1184] Recording and sending audio

[1185] 1. User: When a user feels unwell, they talk to the device about their symptoms and feelings.

[1186] 2. Device: The device (e.g., a smartphone) starts a recording application and records what the user says. When the recording is finished, the audio data is sent to the server.

[1187] Receiving and storing audio data

[1188] 1. Server: Receives the voice data sent from the terminal.

[1189] 2. Server: Stores the received voice data in a database.

[1190] Analysis of audio data

[1191] 1. Server: Inputs the stored voice data into the generative AI model.

[1192] 2. Server (generative AI model): The generative AI model analyzes the audio data and extracts features such as breathing, voice intonation, and speaking rate.

[1193] 3. Server (Emotion Engine): The emotion engine analyzes the voice data and identifies the user's emotional state.

[1194] 4. Server (analysis system): Records features and emotional states based on the analyzed data.

[1195] Assessment of the patient's physical condition and symptoms

[1196] 1. Server: Evaluates the patient's physical condition and symptoms based on the extracted features and emotional state. This also compares with past voice data.

[1197] 2. Server: As a result of the assessment, classify the patient's physical condition and the urgency of their symptoms (mild, severe, urgent).

[1198] Suggested actions needed

[1199] 1. Server: Selects appropriate medical response based on the evaluation results and the user's emotional state.

[1200] Mild cases:

[1201] 1. Server: Sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[1202] Severe cases:

[1203] 1. Server: Contacts the nearest doctor and arranges for a doctor to be dispatched. The user is notified of the doctor's visit date and time.

[1204] In case of emergency:

[1205] 1. Server: Contacts emergency services and dispatches an ambulance based on the user's location. The server notifies the user that an emergency response is required.

[1206] Specific examples

[1207] Example 1: Mild illness

[1208] 1. User: Feeling unwell, they talk to the device about their symptoms, such as a slight headache or feeling tired.

[1209] 2. Device: Records and sends the audio data to the server.

[1210] 3. Server: Analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the stress level is low.

[1211] 4. Server: Sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[1212] Example 2: Urgent illness

[1213] 1. User: Feeling a sudden deterioration in their physical condition, they talk about their symptoms into the device, such as chest pain or shortness of breath.

[1214] 2. Device: Records and sends the audio data to the server.

[1215] 3. Server: Analyzes the voice data and detects high levels of urgency. At the same time, the emotion engine recognizes strong anxiety and fear.

[1216] 4. Server: Immediately contacts emergency services and dispatches an ambulance. The user is notified that emergency response is required.

[1217] In this way, the system of the present invention can utilize voice data and take emotional state into account to provide a more accurate and faster medical response.

[1218] The processing flow will be explained below.

[1219] Step 1:

[1220] User: If the user feels unwell, they talk to the device about their symptoms and feelings.

[1221] Step 2:

[1222] Terminal: The terminal (e.g., a smartphone) launches a recording application and records what the user says.

[1223] Step 3:

[1224] Device: Once recording is complete, the audio data is sent to the server.

[1225] Step 4:

[1226] Server: Receives the voice data sent from the terminal.

[1227] Step 5:

[1228] Server: Stores the received voice data in a database.

[1229] Step 6:

[1230] Server: Inputs the stored voice data into the generative AI model.

[1231] Step 7:

[1232] Server (generative AI model): The generative AI model analyzes the audio data and extracts features such as breathing, voice intonation, and speaking rate.

[1233] Step 8:

[1234] Server (Emotion Engine): The emotion engine analyzes the voice data and identifies the user's emotional state, which can include joy, sadness, anger, anxiety, etc.

[1235] Step 9:

[1236] Server (analysis system): Records features and emotional states based on the analyzed data.

[1237] Step 10:

[1238] Server: Evaluates the user's physical condition and symptoms based on features and emotional state, and compares them with past voice data.

[1239] Step 11:

[1240] Server: As a result of the evaluation, the server classifies the urgency of the user's physical condition and symptoms (mild, severe, urgent).

[1241] Step 12:

[1242] Server: Selects appropriate medical response based on the evaluation results and the user's emotional state.

[1243] Mild cases:

[1244] Step 13:

[1245] Server: Sends the prescription to the pharmacy and notifies the user to pick up the medication.

[1246] Severe cases:

[1247] Step 14:

[1248] Server: Contacts the nearest doctor and arranges for the doctor to be dispatched. The user is notified of the date and time of the doctor's visit.

[1249] In case of emergency:

[1250] Step 15:

[1251] Server: Contacts emergency services and arranges for an ambulance based on the user's location. The user is notified that an emergency response is required.

[1252] Specific examples

[1253] Example 1: Mild illness

[1254] Step 1:

[1255] User: Feeling unwell, he / she talks about his / her symptoms into the device, such as a slight headache or feeling tired.

[1256] Step 2:

[1257] Device: Records and sends the audio data to the server.

[1258] Step 3:

[1259] Server: Analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the stress level is low.

[1260] Step 4:

[1261] Server: Sends the prescription to the pharmacy and notifies the user to pick up the medication.

[1262] Example 2: Urgent illness

[1263] Step 1:

[1264] User: Feeling a sudden deterioration in their physical condition, they talk about their symptoms into the device, such as chest pain or shortness of breath.

[1265] Step 2:

[1266] Device: Records and sends the audio data to the server.

[1267] Step 3:

[1268] Server: Analyzes voice data to detect high levels of urgency. At the same time, the emotion engine recognizes strong anxiety and fear.

[1269] Step 4:

[1270] Server: Immediately contacts emergency services and dispatches an ambulance. The user is notified that emergency response is required.

[1271] In this way, the system of the present invention can utilize voice data and take emotional state into account to provide a more accurate and faster medical response.

[1272] Example 2

[1273] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1274] In today's world, there is a demand for providing prompt and accurate medical treatment appropriate to a user's physical condition and symptoms. However, conventional systems have difficulty taking into account the user's emotional state in their assessment, which can result in inappropriate medical treatment. Furthermore, prompt treatment for urgent illnesses can be delayed, which can pose serious health risks. To solve these problems, it is necessary to use voice data to assess a user's physical condition and symptoms and propose appropriate medical treatment that also takes into account their emotional state.

[1275] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1276] In this invention, the server includes a means for the user to record voice data and send it to the server, a means for receiving and saving the voice data, a means for inputting the saved voice data into a generative AI model and extracting and analyzing features such as breathing, voice intonation, and speaking rate, a means for evaluating the user's physical condition, symptoms, and emotional state from the analyzed data, and a means for proposing appropriate medical treatment based on the evaluation results. This enables prompt and accurate medical treatment using voice data and taking emotional state into consideration.

[1277] "User" refers to an individual who uses this system.

[1278] "Server" refers to the central processing unit that receives, stores, analyzes, evaluates, and recommends medical responses to voice data.

[1279] "Audio data" refers to audio files in which a user records their physical condition, symptoms, and emotional state.

[1280] A "generative AI model" refers to an artificial intelligence model that analyzes voice data and extracts features such as breathing, voice intonation, and speaking rate.

[1281] An "emotion engine" is a system that analyzes a user's emotional state from voice data and uses the results to help evaluate their physical condition and symptoms.

[1282] "Means for receiving and storing voice data" refers to the function of receiving voice data sent from a terminal and storing it in a database or storage.

[1283] "Means for extracting and analyzing features" refers to the function of passing the received voice data through a generative AI model to extract features such as breathing, voice intonation, and speaking rate.

[1284] "Means for evaluation" refers to a function that evaluates the user's physical condition and symptoms based on the analyzed features and emotional state, and determines the urgency of the condition.

[1285] "Means for proposing medical treatment" refers to a function for proposing appropriate medical treatment to the user based on the evaluation results.

[1286] "Prescription transmission" refers to the act of transmitting prescription information to a pharmacy based on the evaluation results.

[1287] "Dispatch of a doctor" refers to the act of dispatching a doctor to the user based on the evaluation results and the degree of urgency.

[1288] "Contacting emergency services" refers to the act of automatically contacting emergency services in the event of an emergency based on the evaluation results to provide a prompt response.

[1289] MODE FOR CARRYING OUT THE INVENTION

[1290] The present invention provides a system that uses voice data from a user to analyze the user's physical condition, symptoms, and even emotional state, and suggests appropriate medical treatment. Specific embodiments for realizing the present invention will be described below.

[1291] System configuration

[1292] The system consists of the following main components:

[1293] 1. Terminal: This includes the device used by the user (e.g., a smartphone). The user records audio and sends the data to the server. Specific recording applications include the smartphone's built-in recording app and a custom app.

[1294] 2. Server: A central processing unit that stores and analyzes received voice data. It runs generative AI models and makes evaluations and recommendations. The server can be a high-performance cloud server or an on-premise server.

[1295] 3. Emotion engine: This engine analyzes the user's emotional state from voice data. The emotion engine uses machine learning algorithms to extract emotional features from voice data.

[1296] 4. Database: This serves as data storage for saving voice data and analysis results. The database can be an SQL-based database (e.g., MySQL, PostgreSQL) or a NoSQL database (e.g., MongoDB).

[1297] System Features

[1298] 1. Recording and sending audio

[1299] User: When a user feels unwell, they talk to the device about their symptoms and feelings. For example, they might say, "I have a headache and my body feels tired."

[1300] On the device, a recording application is used to record the user's voice. After recording, the voice data is sent to the server. The data is sent via an internet connection.

[1301] 2. Receiving and storing audio data

[1302] Server: Receives the voice data sent from the device and stores it in a database. The saved voice data is appended with a timestamp and user ID.

[1303] 3. Analysis of audio data

[1304] Server: The stored voice data is input into the generative AI model, which analyzes the voice data and extracts features such as breathing, voice intonation, and speaking rate.

[1305] Server (emotion engine): The emotion engine identifies the user's emotional state based on the analysis results. This emotional data is reflected in the evaluation of physical condition and symptoms.

[1306] 4. Assessment of physical condition and symptoms

[1307] Server: Based on the extracted features and emotional state, the server comprehensively evaluates the user's physical condition and symptoms. The evaluation also compares with past data and uses an algorithm to determine the urgency of the condition.

[1308] 5. Recommendation of necessary actions

[1309] Server: Based on the evaluation results, it proposes appropriate medical responses. For example, if the condition is mild, it sends a prescription to a pharmacy and notifies the user to pick up the medicine. If the condition is severe, it dispatches a doctor, and if it is an emergency, it contacts emergency services.

[1310] Specific examples

[1311] Example 1: Mild illness

[1312] 1. The user feels unwell and says to the device, "I have a slight headache."

[1313] 2. The device records the audio and sends it to the server.

[1314] 3. The server analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the patient is experiencing low stress.

[1315] 4. The server sends the prescription to the pharmacy and notifies the user.

[1316] Example 2: In the case of an emergency illness

[1317] 1. The user suddenly feels unwell and says to the device, "My chest hurts and I'm having trouble breathing."

[1318] 2. The device records the audio and sends it to the server.

[1319] 3. The server analyzes the voice data to detect high levels of urgency, and the emotion engine recognizes high levels of anxiety.

[1320] 4. The server immediately contacts emergency services and dispatches an ambulance, notifying the user that an emergency response is required and an ambulance is on the way.

[1321] Prompt Sentence Examples

[1322] Prompt for mild illness:

[1323] "The user will be asked to describe in voice that they are experiencing a slight headache or fatigue. The system will analyze the voice data and suggest appropriate medical treatment."

[1324] Emergency Illness Prompt:

[1325] "Please explain to the user in a voice that you suddenly felt chest pain or shortness of breath. Please analyze the voice data and quickly suggest the necessary emergency response."

[1326] In this way, the system of the present invention can utilize voice data and take emotional state into account to provide a more accurate and faster medical response.

[1327] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1328] Step 1:

[1329] User: Feeling unwell, the user talks about their symptoms and feelings into the device. Specifically, they input something like "I have a headache and my body feels tired." The input voice data is captured by the device's recording application. The output is an audio data file.

[1330] Step 2:

[1331] Device: Use a recording application to record the user's voice. Specifically, the user presses the "Start Recording" button, and then presses the "Stop Recording" button after finishing describing the symptoms. After recording is complete, the device saves the voice data in a digital file format (e.g., WAV format) and sends it to the server. The input is the user's voice, and the output is the voice data file sent to the server.

[1332] Step 3:

[1333] Server: Receives audio data sent from the device. Specifically, it receives audio files via HTTP requests. The audio data is stored in a database, and a timestamp and user ID are added to the stored data. The input is the sent audio data, and the output is the saved audio data file and its metadata.

[1334] Step 4:

[1335] Server: The saved voice data is input into the generative AI model and analyzed. Specifically, the voice data is spectrally analyzed and features such as breathing, voice intonation, and speaking rate are extracted. The input is the saved voice data, and the output is the extracted voice features.

[1336] Step 5:

[1337] Server (Emotion Engine): The emotion engine identifies the user's emotional state based on the analyzed voice data. Specifically, it uses a machine learning algorithm to identify emotional states such as stress and anxiety. The input is voice features, and the output is the identified emotional state.

[1338] Step 6:

[1339] Server: Evaluates the user's physical condition and symptoms based on the extracted features and emotional state. Specifically, it compares with past data and executes algorithms to evaluate the progression of symptoms. The input is voice features and emotion identification results, and the output is an evaluation of the user's physical condition and symptoms.

[1340] Step 7:

[1341] Server: Based on the evaluation results, selects and proposes appropriate medical responses. Specifically, if symptoms are mild, it sends a prescription to a pharmacy, if symptoms are severe, it dispatches a doctor, and if there is an emergency, it contacts emergency services. The input is the evaluation results of physical condition and symptoms, and the output is a medical response proposal.

[1342] Step 8:

[1343] Server: Notifies the user of the selected medical response. Specifically, it contacts the user via email or SMS and instructs them on the necessary response. The input is the medical response proposal, and the output is a notification message to the user.

[1344] (Application example 2)

[1345] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1346] Until now, there has been no system in brick-and-mortar stores that can quickly and accurately evaluate the physical and emotional state of customers and provide optimal medical care. As a result, customers have had to wait a long time to receive appropriate medical care, which could worsen their symptoms. To solve this problem, the present invention aims to provide a system that can analyze the voice of customers in brick-and-mortar stores, evaluate their physical and emotional state, and quickly propose and implement appropriate medical care.

[1347] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving and storing voice data, means for executing a generative AI model for analyzing the received voice data, means for evaluating the user's physical condition and emotional state from the analyzed data, and means for proposing and implementing appropriate medical treatment based on the evaluation results. This makes it possible to quickly evaluate the physical condition and emotional state of customers in a physical store and provide optimal medical treatment.

[1348] "Voice data" refers to data in which the contents of a user's speech are recorded in digital format.

[1349] A "generative AI model" is an algorithm that uses machine learning technology to analyze voice data and extract specific features and emotional states.

[1350] The "means for analyzing" refers to the means for inputting received voice data into a generative AI model and carrying out the process of analyzing the characteristics of the data.

[1351] "Physical and emotional state" refers to the user's physical health and psychological feelings and moods.

[1352] The "means for evaluation" is a means for carrying out a process of determining the user's physical condition and emotional state based on the analyzed data.

[1353] "Medical response" means taking action based on the user's physical or emotional state, including suggesting appropriate medication, recommending a doctor's appointment, or contacting emergency services.

[1354] The "means for proposing and implementing" is a means for carrying out the process of proposing specific medical measures to the user based on the evaluation results and implementing those measures.

[1355] A "physical store" refers to a physical store where a user actually visits and receives face-to-face service.

[1356] System Overview

[1357] This invention relates to a system that automatically evaluates the physical and emotional state of customers in brick-and-mortar stores and proposes and implements appropriate medical treatment. This system collects and analyzes voice data, evaluates their emotional state, and then proposes and implements the most appropriate medical treatment for the user.

[1358] System Configuration

[1359] The system consists of the following main components:

[1360] 1. Terminal: This includes a recording device (e.g., a tablet or dedicated microphone at the reception desk) installed in the physical store and used by customers. Customers talk to this terminal about their symptoms and condition.

[1361] 2. Server: This is the central system that receives and analyzes the voice data, runs the generative AI model, and makes evaluations and recommendations.

[1362] 3. Emotion engine: Analyzes the emotional state of customers from voice data and reflects the results in assessing their physical condition and symptoms.

[1363] 4. Database: Serves as data storage for saving voice data and analysis results.

[1364] How it works

[1365] Recording and sending audio

[1366] The user speaks about their symptoms and feelings into a device installed in a physical store. The device launches a recording application and records what is said. After recording, the audio data is sent to the server.

[1367] Receiving and storing audio data

[1368] The server receives the voice data sent from the terminal and stores it in a database.

[1369] Analysis of audio data

[1370] The server inputs the stored voice data into a generative AI model for analysis. The generative AI model extracts features such as breathing, voice intonation, speaking rate, and emotional state. An emotion engine analyzes the data and identifies the user's emotional state.

[1371] Assessment of physical condition and symptoms

[1372] The server evaluates the customer's physical condition and symptoms based on the extracted features and emotional state, and also compares this with past voice data.

[1373] Propose and take necessary actions

[1374] Based on the evaluation results and the user's emotional state, the server selects and implements appropriate medical responses, such as suggesting medication, recommending a doctor's visit, or contacting emergency services.

[1375] Specific examples

[1376] Example 1: Mild illness

[1377] 1. User: Complains of a slight headache and fatigue and describes his symptoms into the terminal.

[1378] 2. Server: Analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the stress level is low.

[1379] 3. Server: Display a notification on the terminal to customers saying, "Please use over-the-counter headache medicine."

[1380] Example 2: Urgent illness

[1381] 1. User: Complains of chest pain and shortness of breath and describes his symptoms into the terminal.

[1382] 2. Server: Analyzes the voice data and detects high levels of urgency. At the same time, the emotion engine recognizes strong anxiety and fear.

[1383] 3. Server: Immediately dispatch emergency services and display a notification to the customer on their device saying, "Emergency response required. An ambulance is on the way."

[1384] Prompt Sentence Examples

[1385] User Voice:

[1386] "I've been having terrible headaches lately and I'm feeling stressed."

[1387] System response:

[1388] Emotional state: Stressed, high

[1389] Symptom assessment: mild

[1390] Suggestion: "Take advantage of over-the-counter headache medication. I recommend relaxation techniques to help relieve stress."

[1391] In this way, the system of the present invention can quickly assess the physical and emotional state of customers in a physical store and provide optimal medical care.

[1392] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1393] Step 1:

[1394] The user talks about their symptoms and feelings into a device installed in a physical store. The device launches a recording application and records the user's voice. When the recording is finished, the device sends the voice data (input) to the server (output).

[1395] Step 2:

[1396] The server receives the voice data sent from the terminal and stores it in a database. Specifically, after receiving the voice data (input), it records (outputs) it in the database.

[1397] Step 3:

[1398] The server inputs the saved voice data into the generative AI model for analysis. The server analyzes the voice data (input) and extracts features such as breathing, voice intonation, speaking rate, and emotional state (output). Specifically, it runs the generative AI model and obtains the results of analyzing each feature.

[1399] Step 4:

[1400] The server uses an emotion engine to identify the user's emotional state from the analyzed data. The emotion engine performs data calculations based on the input voice features and evaluates (outputs) the user's emotional state.

[1401] Step 5:

[1402] The server evaluates the user's physical condition and symptoms based on the features and emotional state extracted from the voice data. It compares the current condition with past data and evaluates it. In this process, the server derives the evaluation results (output) of the physical condition and symptoms based on the features and emotional state (input).

[1403] Step 6:

[1404] The server selects and implements appropriate medical responses based on the evaluation results and the user's emotional state. Specifically, it suggests medication, recommends a doctor's consultation, or contacts emergency services depending on the user's physical condition and the urgency of the symptoms. The input in this step is the evaluation results and emotional state, and the output is the proposal or implementation of specific medical responses.

[1405] Step 7:

[1406] The server notifies the user either through the server itself or the terminal. For example, if a medication is suggested, a notification saying "Please use over-the-counter headache medicine" is displayed on the terminal. In this step, a notification (output) to the user is created based on the medical treatment suggestion (input).

[1407] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1408] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1409] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1410] [Fourth embodiment]

[1411] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1412] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1413] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1414] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1415] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1416] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1417] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1418] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1419] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1420] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1421] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1422] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1423] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1424] System Overview

[1425] The present invention relates to a system that automatically evaluates a patient's physical condition and symptoms using voice data and proposes appropriate medical treatment. This system is mainly composed of means for receiving, analyzing, and evaluating the voice data and proposing medical treatment.

[1426] System Configuration

[1427] The system consists of the following main components:

[1428] 1. Terminal: This includes the device used by the user (e.g., a smartphone). The user is responsible for recording audio and sending the data to the server.

[1429] 2. Server: This is the central system that stores and analyzes the received voice data, runs the generative AI model, and makes evaluations and suggestions.

[1430] 3. Database: Serves as data storage for saving voice data and analysis results.

[1431] Explanation of program processing

[1432] Recording and sending audio

[1433] 1. User: If the user feels unwell, they speak to the terminal about their symptoms.

[1434] 2. Device: The device (e.g., a smartphone) starts a recording application and records the user's voice. When the recording is finished, the voice data is sent to the server.

[1435] Receiving and storing audio data

[1436] 1. Server: Receives the voice data sent from the device and stores it in a database.

[1437] Analysis of audio data

[1438] 1. Server: The stored voice data is input into the generative AI model for analysis. The generative AI model extracts information such as breathing, voice intonation, and speaking rate.

[1439] 2. Server: Extracts features from the analyzed data and records them.

[1440] Assessment of the patient's physical condition and symptoms

[1441] 1. Server: Evaluates the patient's physical condition and symptoms based on the extracted features. The evaluation also involves comparison with past voice data.

[1442] 2. Server: Based on the assessment results, classify the patient's physical condition and the urgency of their symptoms (mild, severe, urgent).

[1443] Suggested actions needed

[1444] 1. Server: Based on the evaluation results, propose appropriate medical responses. Specific responses are as follows:

[1445] Mild cases: Send a prescription to the pharmacy and notify the user to pick up the medication.

[1446] In severe cases: The nearest doctor will be contacted and dispatched. The user will be notified of the date and time of the doctor's visit.

[1447] In case of emergency: Call emergency services and arrange for an ambulance. Notify the user that immediate emergency response is required.

[1448] Specific examples

[1449] Example 1: Mild illness

[1450] 1. User: Feeling unwell, speaks about symptoms into the terminal.

[1451] 2. Device: Records and sends the audio data to the server.

[1452] 3. Server: Analyzes the voice data and evaluates the symptoms as mild.

[1453] 4. Server: Sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[1454] Example 2: Urgent illness

[1455] 1. User: Feeling a sudden deterioration in their health, they talk about their symptoms into the device.

[1456] 2. Device: Records and sends the audio data to the server.

[1457] 3. Server: Analyzes the voice data and detects high levels of urgency.

[1458] 4. Server: Immediately contacts emergency services and dispatches an ambulance. The user is notified that emergency response is required.

[1459] In this way, the present invention is a system that can efficiently and quickly evaluate a patient's physical condition and symptoms using voice data and provide appropriate medical treatment.

[1460] The processing flow will be explained below.

[1461] Step 1:

[1462] User: If the user feels unwell, they talk to the device about their symptoms.

[1463] Step 2:

[1464] Terminal: The terminal (e.g., a smartphone) launches a recording application and records what the user says.

[1465] Step 3:

[1466] Device: Once recording is complete, the audio data is sent to the server.

[1467] Step 4:

[1468] Server: Receives the voice data sent from the terminal.

[1469] Step 5:

[1470] Server: Stores the received voice data in a database.

[1471] Step 6:

[1472] Server: Inputs the stored voice data into the generative AI model.

[1473] Step 7:

[1474] Server (generative AI model): The generative AI model analyzes the audio data and extracts features such as breathing, voice intonation, and speaking rate.

[1475] Step 8:

[1476] Server (analysis system): Records features based on the analyzed data.

[1477] Step 9:

[1478] Server: Evaluates the patient's physical condition and symptoms based on the features. This also compares the results with past voice data.

[1479] Step 10:

[1480] Server: As a result of the evaluation, classify the patient's physical condition and the urgency of their symptoms (mild, severe, urgent).

[1481] Step 11:

[1482] Server: Selects appropriate medical response based on the evaluation results.

[1483] Step 12:

[1484] Server: If the condition is mild, it sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[1485] Step 13:

[1486] Server: If the condition is severe, contact the nearest doctor and arrange for a doctor to be dispatched. The user is notified of the date and time of the doctor's visit.

[1487] Step 14:

[1488] Server: In case of an emergency, contacts emergency services and arranges for an ambulance based on the user's location information. The user is notified that emergency response is required.

[1489] Through the above steps, the system of the present invention can utilize voice data to quickly evaluate the patient's physical condition and symptoms and provide appropriate medical care.

[1490] Example 1

[1491] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1492] Conventional medical systems have struggled to remotely and quickly and accurately assess a patient's physical condition and symptoms, and to propose appropriate medical treatment. They also struggled to respond immediately to particularly urgent symptoms, posing a significant risk to the patient's health and safety. Furthermore, they lacked advanced technology for analyzing voice data and extracting features, making it impossible to accurately identify changes in physical condition.

[1493] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1494] In this invention, the server includes means for receiving and storing voice data, means for executing a generative AI model to analyze the received voice data, means for extracting features from the analyzed data and recording them in a database, means for evaluating the patient's physical condition and symptoms from the extracted features and comparing them with past data, means for classifying the urgency of the symptoms based on the evaluation results, and means for proposing appropriate medical responses based on the evaluation results. This makes it possible to quickly and accurately evaluate the patient's physical condition and symptoms and propose appropriate medical measures. Specific responses can include automatically sending prescriptions, dispatching medical personnel, and contacting emergency services.

[1495] "Audio data" refers to data in which sound is recorded and stored in digital form.

[1496] "Reception" is the process of receiving data or signals sent from the outside.

[1497] "Storage" means storing the received data in a database or storage so that it can be accessed later if necessary.

[1498] A "generative AI model" is an artificial intelligence model that has been trained using machine learning or deep learning to perform specific tasks automatically.

[1499] "Analysis" is the processing of data and the extraction of meaningful information.

[1500] A "feature" is a numerical value or index extracted from data that represents a specific pattern or characteristic.

[1501] A "database" is an electronic data management system that stores data systematically and enables efficient searching and manipulation.

[1502] "Evaluation" refers to determining the content or state of data based on specific criteria.

[1503] "Urgency" is a scale that indicates the severity of the evaluated symptoms and the degree of need for response.

[1504] "Proposals" refer to the presentation of optimal actions or measures based on the evaluation results.

[1505] "Medical response" refers to medical services and treatments provided according to the patient's physical condition and symptoms.

[1506] MODE FOR CARRYING OUT THE INVENTION

[1507] The present invention relates to a system that automatically evaluates a patient's physical condition and symptoms using voice data and proposes appropriate medical treatment. This system mainly includes means for receiving, analyzing, and evaluating the voice data and proposing medical treatment.

[1508] System Overview

[1509] The system consists of the following main components:

[1510] 1. Terminal (e.g., smartphone): A device that allows a user to record audio and send the data to a server.

[1511] 2. Server: A central system for storing and analyzing received voice data, running generative AI models, and providing a means for evaluation and recommendations.

[1512] 3. Database: A storage system that stores voice data and analysis results.

[1513] Recording and sending audio

[1514] (User): When a user feels unwell, they can talk about their symptoms using a device such as a smartphone. For example, they can say, "I have chest pain."

[1515] (Device): The device (smartphone) activates the voice recording function and records the user's voice. When the recording is finished, the voice data is automatically sent to the server using the HTTPS protocol.

[1516] Receiving and storing audio data

[1517] (Server): The server receives the voice data sent from the device. The received voice data is stored in a temporary storage area, and once it is confirmed that it has been received completely, it is permanently stored in the database.

[1518] Analysis of audio data

[1519] (Server): The server inputs the saved voice data into a generative AI model (e.g., an artificial intelligence model trained using machine learning or deep learning) and begins the analysis process. Specifically, the AI ​​model analyzes features such as breathing, voice intonation, and speaking rate from the voice data.

[1520] (Server): Extracts features from the analyzed data and records them in a database. These features are basic data for evaluating the patient's physical condition.

[1521] Assessment of the patient's physical condition and symptoms

[1522] (Server): Evaluates the patient's physical condition and symptoms based on the extracted features. Executes evaluation procedures to compare with past voice data and analyze changes. For example, compares past voice data with current data and uses a machine learning algorithm to determine whether there are any abnormal changes.

[1523] (Server): Based on the evaluation results, the urgency of the symptoms is classified into three categories: mild, severe, and urgent.

[1524] Suggested actions needed

[1525] (Server): Based on the evaluation results, the following appropriate medical responses are proposed:

[1526] For mild cases: Email the prescription to the nearest pharmacy and notify the user to collect the medication. For example, send the prescription information to the pharmacy's email address.

[1527] In severe cases: Contact the nearest healthcare professional and arrange for a visit. Notify the user about the date and time of the healthcare professional's visit, for example by contacting the healthcare professional using their contact information.

[1528] In case of emergency: Contact emergency services and dispatch an ambulance. The user is notified immediately and informed that emergency response is required, for example by auto-dialing an emergency number and providing location and symptom information.

[1529] Specific examples

[1530] Example 1: Mild illness

[1531] (User): The user experiences a slight headache and says to their smartphone, "My head hurts a little."

[1532] (Device): The recording application records the audio and sends the audio data to the server.

[1533] (Server): Receives the voice data, analyzes it using a generative AI model, and determines that the symptoms are mild.

[1534] (Server): Sends the prescription to the pharmacy and notifies the user that the medicine has been received. For example, the server sends the prescription information to the pharmacy's email address.

[1535] Example 2: Urgent illness

[1536] (User): Feeling sudden chest pain, he says into his smartphone, "My chest hurts, I can't breathe."

[1537] (Device): The recording application records the audio and sends the audio data to the server.

[1538] (Server): Receives voice data, analyzes it using a generative AI model, and determines the urgency of the data.

[1539] (Server): Contacts emergency services and dispatches an ambulance. The user is notified immediately and informed that emergency response is required. For example, by automatically dialing an emergency number and providing location and symptom information.

[1540] Prompt Sentence Examples

[1541] "Analyze breathing, voice intonation, and speech rate to determine the patient's health condition. For example, analyze the audio data of 'I have a severe stomachache' and assess the urgency of the situation."

[1542] In this way, the present invention is a system that utilizes voice data to quickly and accurately evaluate a patient's physical condition and symptoms, and provide appropriate medical care.

[1543] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1544] Step 1:

[1545] (Input): If the user feels unwell, they can speak to the terminal about their symptoms. For example, they can say, "My chest hurts."

[1546] (Action): The device activates the voice recording function and records the user's voice.

[1547] (Output): Recorded audio data.

[1548] Step 2:

[1549] (Input): Recorded audio data.

[1550] (Operation): After the device finishes recording, it sends the audio data to the server using the HTTPS protocol.

[1551] (Output): The audio data sent to the server.

[1552] Step 3:

[1553] (Input): Audio data sent from the device.

[1554] (Operation): The server receives the voice data and stores it in a temporary storage area. Once reception is complete, the voice data is permanently stored in the database.

[1555] (Output): The audio data stored in the database.

[1556] Step 4:

[1557] (Input): Audio data stored in the database.

[1558] (Operation): The server inputs the voice data into the generative AI model and performs an analysis process. The generative AI model analyzes features such as breathing, voice intonation, and speaking rate from the voice data.

[1559] (Output): Parsed features.

[1560] Step 5:

[1561] (Input): Parsed features.

[1562] (Operation): The server extracts features and records them in a database.

[1563] (Output): Features recorded in the database.

[1564] Step 6:

[1565] (Input): Features recorded in the database.

[1566] (Operation): The server evaluates the patient's physical condition and symptoms based on the extracted features, compares them with past data, and analyzes changes. Specifically, it uses machine learning algorithms to identify abnormal changes.

[1567] (Output): Evaluation results (physical condition evaluation and urgency determination).

[1568] Step 7:

[1569] (Input): Evaluation result.

[1570] (Operation): Based on the evaluation results, the server classifies the urgency of the symptoms into three categories: mild, severe, and urgent.

[1571] (Output): Urgency classification result.

[1572] Step 8:

[1573] (Input): Urgency classification result.

[1574] (Operation): The server suggests appropriate medical responses, specifically sending a prescription to a pharmacy if the condition is mild, dispatching medical personnel if the condition is severe, and contacting emergency services to dispatch an ambulance if the condition is urgent.

[1575] (Output): Notification to the user and medical response.

[1576] This enables the system to use voice data to quickly and accurately assess a patient's physical condition and symptoms, enabling appropriate medical treatment.

[1577] (Application example 1)

[1578] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1579] Conventional security systems have had difficulty quickly and accurately detecting physical security threats and anomalies. Furthermore, delayed response in emergencies can lead to serious damage to human life and property. To solve these problems, there is a need for a system that can analyze voice data, automatically assess security situations, and quickly take appropriate action.

[1580] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1581] In this invention, the server includes means for receiving and storing voice data, means for executing a generative AI model to analyze the received voice data, means for evaluating the target situation or abnormality from the analyzed data, and means for proposing or implementing appropriate countermeasures based on the evaluation results, thereby enabling quick and accurate detection of physical security threats and abnormalities and prompt implementation of appropriate countermeasures.

[1582] "Audio data" means digital or analog information that is an electronic recording of sound.

[1583] "Reception" is the act of a specific device or system taking in data or signals from the outside.

[1584] A "generative AI model" is an artificial intelligence algorithm or network designed to analyze data and make predictions.

[1585] "Analysis" is the process of examining data or information in detail to understand its structure and meaning.

[1586] "Subject" refers to an object or situation that is the subject of observation or analysis under specific conditions or circumstances.

[1587] A "situation" refers to the state or condition of an environment or event at a given moment.

[1588] An "abnormality" is a problem or malfunction that deviates from normal conditions or norms.

[1589] "Evaluation" is the act of making judgments or analyses based on data or information in accordance with specific criteria.

[1590] "Countermeasures" refer to specific actions or measures taken to address a problem or abnormality.

[1591] A "suggestion" is a recommendation for a particular action or solution.

[1592] "Implementation" means actually carrying out the proposed measures or actions.

[1593] "Urgency" is a measure of the seriousness of a situation or problem and the degree to which a response is necessary.

[1594] "Notification" is the act of conveying specific information or instructions to interested parties.

[1595] "Guard" refers to a security guard or security service deployed to protect the security of a particular place or object.

[1596] A "police agency" is a government agency established to maintain public order and safety.

[1597] System Overview

[1598] The present invention relates to a system for enhancing security in offices and homes using voice data, which mainly comprises means for receiving, analyzing, and evaluating the voice data, and proposing or implementing emergency response measures.

[1599] System Configuration

[1600] The system consists of the following main components:

[1601] 1. Terminal (smartphone, smart glasses, security robot): The user is responsible for recording voice and sending the data to the server.

[1602] 2. Server: This is the central system that stores and analyzes the received voice data, runs the generative AI model, and evaluates it, proposing or implementing countermeasures.

[1603] 3. Database (MySQL, PostgreSQL): Serves as data storage for saving voice data and analysis results.

[1604] Explanation of program processing

[1605] Recording and sending audio

[1606] If a user senses a physical security threat or an abnormality, they report it by voice into the device (smartphone, smart glasses, robot), which then records the voice and sends the recorded data to the server.

[1607] Receiving and storing audio data

[1608] The server receives the voice data sent from the terminal and stores it in a database for subsequent processing.

[1609] Analysis of audio data

[1610] The server inputs the saved voice data into a generative AI model for analysis. The AI ​​model extracts voice tension, abnormal sounds, alarm sounds, etc. Features are extracted from the analyzed data and recorded in a database.

[1611] Security Status Assessment

[1612] The server evaluates the security situation based on the extracted features. The evaluation also involves comparison with past voice data. Based on the evaluation results, the urgency of the security situation is classified (normal, caution, emergency).

[1613] Proposing or implementing measures

[1614] The server will propose or implement appropriate measures based on the evaluation results. Specific actions include:

[1615] Caution: Notify the user and call for caution.

[1616] In case of emergency: Contact the police or security company and arrange for appropriate response. Inform users to evacuate immediately.

[1617] Specific examples

[1618] Example 1: When caution is required

[1619] A user hears an unusual sound around the house and reports it to the device. The device records the sound and sends it to the server. The server analyzes the sound data and determines that the situation requires attention. The server then sends a notification to the user to alert them.

[1620] Example 2: When emergency response is required

[1621] The user suspects an intruder in their home and reports the incident to the device. The device records the audio and sends it to the server. The server analyzes the audio data and determines that the situation requires emergency response. The server immediately contacts the police and security companies and arranges for countermeasures. The user is notified to evacuate immediately.

[1622] Example prompts for generative AI models

[1623] Enter the following prompts into the generative AI model:

[1624] "The server has recorded some unusual sounds around the house. Please analyze whether this sound should be a cause for alarm or ignored."

[1625] "We have recorded a possible intruder on our server. Please analyze whether this audio requires immediate attention or should be ignored."

[1626] In this way, the present invention is a system that can use voice data to efficiently and quickly detect physical security threats and anomalies and provide appropriate responses.

[1627] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1628] Step 1:

[1629] When a user senses a physical security threat or anomaly, they report the details by voice into a device such as a smartphone, smart glasses, or security robot. The input is voice data, and the output is a recorded voice file. The device then performs the specific action of recording this voice.

[1630] Step 2:

[1631] When the recording is finished, the device sends the audio data to the server. The input is the recorded audio file, and the output is the audio data transferred to the server. The server receives and stores this data.

[1632] Step 3:

[1633] The server inputs the received voice data into the generative AI model for analysis. The input is the voice data stored on the server, and the output is the analyzed features (voice tension, abnormal sounds, alarm sounds, etc.). The server analyzes the data and performs specific operations to extract important features.

[1634] Step 4:

[1635] After features are extracted from the voice data analyzed by the generative AI model, the server evaluates the security situation based on these features. The input is the extracted features, and the output is the evaluation result (normal, caution, emergency). The server compares it with past data and performs specific actions to determine the security situation.

[1636] Step 5:

[1637] Based on the evaluation results, the server proposes or automatically executes appropriate countermeasures. The input is the evaluation results, and the output is notifications and the execution of countermeasures. Specifically, if caution is required, a notification is sent to the user to warn them, and in the case of an emergency, the police or security company is contacted and a response is arranged. The server performs the specific operations of generating notifications and arranging contact.

[1638] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1639] System Overview

[1640] The present invention relates to a system that uses voice data to automatically evaluate a patient's emotional state in addition to their physical condition and symptoms, and proposes appropriate medical treatment. This system is composed of various means, including receiving, analyzing, and evaluating voice data, proposing medical treatment, and an emotion engine.

[1641] System Configuration

[1642] The system consists of the following main components:

[1643] 1. Terminal: This includes the device used by the user (e.g., a smartphone). The user is responsible for recording audio and sending the data to the server.

[1644] 2. Server: This is the central system that stores and analyzes the received voice data, runs the generative AI model, and makes evaluations and suggestions.

[1645] 3. Emotion engine: Analyzes the user's emotional state from voice data and reflects the results in the evaluation of physical condition and symptoms.

[1646] 4. Database: Serves as data storage for saving voice data and analysis results.

[1647] Explanation of program processing

[1648] Recording and sending audio

[1649] 1. User: When a user feels unwell, they talk to the device about their symptoms and feelings.

[1650] 2. Device: The device (e.g., a smartphone) starts a recording application and records what the user says. When the recording is finished, the audio data is sent to the server.

[1651] Receiving and storing audio data

[1652] 1. Server: Receives the voice data sent from the terminal.

[1653] 2. Server: Stores the received voice data in a database.

[1654] Analysis of audio data

[1655] 1. Server: Inputs the stored voice data into the generative AI model.

[1656] 2. Server (generative AI model): The generative AI model analyzes the audio data and extracts features such as breathing, voice intonation, and speaking rate.

[1657] 3. Server (Emotion Engine): The emotion engine analyzes the voice data and identifies the user's emotional state.

[1658] 4. Server (analysis system): Records features and emotional states based on the analyzed data.

[1659] Assessment of the patient's physical condition and symptoms

[1660] 1. Server: Evaluates the patient's physical condition and symptoms based on the extracted features and emotional state. This also compares with past voice data.

[1661] 2. Server: As a result of the assessment, classify the patient's physical condition and the urgency of their symptoms (mild, severe, urgent).

[1662] Suggested actions needed

[1663] 1. Server: Selects appropriate medical response based on the evaluation results and the user's emotional state.

[1664] Mild cases:

[1665] 1. Server: Sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[1666] Severe cases:

[1667] 1. Server: Contacts the nearest doctor and arranges for a doctor to be dispatched. The user is notified of the doctor's visit date and time.

[1668] In case of emergency:

[1669] 1. Server: Contacts emergency services and dispatches an ambulance based on the user's location. The server notifies the user that an emergency response is required.

[1670] Specific examples

[1671] Example 1: Mild illness

[1672] 1. User: Feeling unwell, they talk to the device about their symptoms, such as a slight headache or feeling tired.

[1673] 2. Device: Records and sends the audio data to the server.

[1674] 3. Server: Analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the stress level is low.

[1675] 4. Server: Sends the prescription to the pharmacy and notifies the user to pick up the medicine.

[1676] Example 2: Urgent illness

[1677] 1. User: Feeling a sudden deterioration in their physical condition, they talk about their symptoms into the device, such as chest pain or shortness of breath.

[1678] 2. Device: Records and sends the audio data to the server.

[1679] 3. Server: Analyzes the voice data and detects high levels of urgency. At the same time, the emotion engine recognizes strong anxiety and fear.

[1680] 4. Server: Immediately contacts emergency services and dispatches an ambulance. The user is notified that emergency response is required.

[1681] In this way, the system of the present invention can utilize voice data and take emotional state into account to provide a more accurate and faster medical response.

[1682] The processing flow will be explained below.

[1683] Step 1:

[1684] User: If the user feels unwell, they talk to the device about their symptoms and feelings.

[1685] Step 2:

[1686] Terminal: The terminal (e.g., a smartphone) launches a recording application and records what the user says.

[1687] Step 3:

[1688] Device: Once recording is complete, the audio data is sent to the server.

[1689] Step 4:

[1690] Server: Receives the voice data sent from the terminal.

[1691] Step 5:

[1692] Server: Stores the received voice data in a database.

[1693] Step 6:

[1694] Server: Inputs the stored voice data into the generative AI model.

[1695] Step 7:

[1696] Server (generative AI model): The generative AI model analyzes the audio data and extracts features such as breathing, voice intonation, and speaking rate.

[1697] Step 8:

[1698] Server (Emotion Engine): The emotion engine analyzes the voice data and identifies the user's emotional state, which can include joy, sadness, anger, anxiety, etc.

[1699] Step 9:

[1700] Server (analysis system): Records features and emotional states based on the analyzed data.

[1701] Step 10:

[1702] Server: Evaluates the user's physical condition and symptoms based on features and emotional state, and compares them with past voice data.

[1703] Step 11:

[1704] Server: As a result of the evaluation, the server classifies the urgency of the user's physical condition and symptoms (mild, severe, urgent).

[1705] Step 12:

[1706] Server: Selects appropriate medical response based on the evaluation results and the user's emotional state.

[1707] Mild cases:

[1708] Step 13:

[1709] Server: Sends the prescription to the pharmacy and notifies the user to pick up the medication.

[1710] Severe cases:

[1711] Step 14:

[1712] Server: Contacts the nearest doctor and arranges for the doctor to be dispatched. The user is notified of the date and time of the doctor's visit.

[1713] In case of emergency:

[1714] Step 15:

[1715] Server: Contacts emergency services and arranges for an ambulance based on the user's location. The user is notified that an emergency response is required.

[1716] Specific examples

[1717] Example 1: Mild illness

[1718] Step 1:

[1719] User: Feeling unwell, he / she talks about his / her symptoms into the device, such as a slight headache or feeling tired.

[1720] Step 2:

[1721] Device: Records and sends the audio data to the server.

[1722] Step 3:

[1723] Server: Analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the stress level is low.

[1724] Step 4:

[1725] Server: Sends the prescription to the pharmacy and notifies the user to pick up the medication.

[1726] Example 2: Urgent illness

[1727] Step 1:

[1728] User: Feeling a sudden deterioration in their physical condition, they talk about their symptoms into the device, such as chest pain or shortness of breath.

[1729] Step 2:

[1730] Device: Records and sends the audio data to the server.

[1731] Step 3:

[1732] Server: Analyzes voice data to detect high levels of urgency. At the same time, the emotion engine recognizes strong anxiety and fear.

[1733] Step 4:

[1734] Server: Immediately contacts emergency services and dispatches an ambulance. The user is notified that emergency response is required.

[1735] In this way, the system of the present invention can utilize voice data and take emotional state into account to provide a more accurate and faster medical response.

[1736] Example 2

[1737] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1738] In today's world, there is a demand for providing prompt and accurate medical treatment appropriate to a user's physical condition and symptoms. However, conventional systems have difficulty taking into account the user's emotional state in their assessment, which can result in inappropriate medical treatment. Furthermore, prompt treatment for urgent illnesses can be delayed, which can pose serious health risks. To solve these problems, it is necessary to use voice data to assess a user's physical condition and symptoms and propose appropriate medical treatment that also takes into account their emotional state.

[1739] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1740] In this invention, the server includes a means for the user to record voice data and send it to the server, a means for receiving and saving the voice data, a means for inputting the saved voice data into a generative AI model and extracting and analyzing features such as breathing, voice intonation, and speaking rate, a means for evaluating the user's physical condition, symptoms, and emotional state from the analyzed data, and a means for proposing appropriate medical treatment based on the evaluation results. This enables prompt and accurate medical treatment using voice data and taking emotional state into consideration.

[1741] "User" refers to an individual who uses this system.

[1742] "Server" refers to the central processing unit that receives, stores, analyzes, evaluates, and recommends medical responses to voice data.

[1743] "Audio data" refers to audio files in which a user records their physical condition, symptoms, and emotional state.

[1744] A "generative AI model" refers to an artificial intelligence model that analyzes voice data and extracts features such as breathing, voice intonation, and speaking rate.

[1745] An "emotion engine" is a system that analyzes a user's emotional state from voice data and uses the results to help evaluate their physical condition and symptoms.

[1746] "Means for receiving and storing voice data" refers to the function of receiving voice data sent from a terminal and storing it in a database or storage.

[1747] "Means for extracting and analyzing features" refers to the function of passing the received voice data through a generative AI model to extract features such as breathing, voice intonation, and speaking rate.

[1748] "Means for evaluation" refers to a function that evaluates the user's physical condition and symptoms based on the analyzed features and emotional state, and determines the urgency of the condition.

[1749] "Means for proposing medical treatment" refers to a function for proposing appropriate medical treatment to the user based on the evaluation results.

[1750] "Prescription transmission" refers to the act of transmitting prescription information to a pharmacy based on the evaluation results.

[1751] "Dispatch of a doctor" refers to the act of dispatching a doctor to the user based on the evaluation results and the degree of urgency.

[1752] "Contacting emergency services" refers to the act of automatically contacting emergency services in the event of an emergency based on the evaluation results to provide a prompt response.

[1753] MODE FOR CARRYING OUT THE INVENTION

[1754] The present invention provides a system that uses voice data from a user to analyze the user's physical condition, symptoms, and even emotional state, and suggests appropriate medical treatment. Specific embodiments for realizing the present invention will be described below.

[1755] System configuration

[1756] The system consists of the following main components:

[1757] 1. Terminal: This includes the device used by the user (e.g., a smartphone). The user records audio and sends the data to the server. Specific recording applications include the smartphone's built-in recording app and a custom app.

[1758] 2. Server: A central processing unit that stores and analyzes received voice data. It runs generative AI models and makes evaluations and recommendations. The server can be a high-performance cloud server or an on-premise server.

[1759] 3. Emotion engine: This engine analyzes the user's emotional state from voice data. The emotion engine uses machine learning algorithms to extract emotional features from voice data.

[1760] 4. Database: This serves as data storage for saving voice data and analysis results. The database can be an SQL-based database (e.g., MySQL, PostgreSQL) or a NoSQL database (e.g., MongoDB).

[1761] System Features

[1762] 1. Recording and sending audio

[1763] User: When a user feels unwell, they talk to the device about their symptoms and feelings. For example, they might say, "I have a headache and my body feels tired."

[1764] On the device, a recording application is used to record the user's voice. After recording, the voice data is sent to the server. The data is sent via an internet connection.

[1765] 2. Receiving and storing audio data

[1766] Server: Receives the voice data sent from the device and stores it in a database. The saved voice data is appended with a timestamp and user ID.

[1767] 3. Analysis of audio data

[1768] Server: The stored voice data is input into the generative AI model, which analyzes the voice data and extracts features such as breathing, voice intonation, and speaking rate.

[1769] Server (emotion engine): The emotion engine identifies the user's emotional state based on the analysis results. This emotional data is reflected in the evaluation of physical condition and symptoms.

[1770] 4. Assessment of physical condition and symptoms

[1771] Server: Based on the extracted features and emotional state, the server comprehensively evaluates the user's physical condition and symptoms. The evaluation also compares with past data and uses an algorithm to determine the urgency of the condition.

[1772] 5. Recommendation of necessary actions

[1773] Server: Based on the evaluation results, it proposes appropriate medical responses. For example, if the condition is mild, it sends a prescription to a pharmacy and notifies the user to pick up the medicine. If the condition is severe, it dispatches a doctor, and if it is an emergency, it contacts emergency services.

[1774] Specific examples

[1775] Example 1: Mild illness

[1776] 1. The user feels unwell and says to the device, "I have a slight headache."

[1777] 2. The device records the audio and sends it to the server.

[1778] 3. The server analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the patient is experiencing low stress.

[1779] 4. The server sends the prescription to the pharmacy and notifies the user.

[1780] Example 2: In the case of an emergency illness

[1781] 1. The user suddenly feels unwell and says to the device, "My chest hurts and I'm having trouble breathing."

[1782] 2. The device records the audio and sends it to the server.

[1783] 3. The server analyzes the voice data to detect high levels of urgency, and the emotion engine recognizes high levels of anxiety.

[1784] 4. The server immediately contacts emergency services and dispatches an ambulance, notifying the user that an emergency response is required and an ambulance is on the way.

[1785] Prompt Sentence Examples

[1786] Prompt for mild illness:

[1787] "The user will be asked to describe in voice that they are experiencing a slight headache or fatigue. The system will analyze the voice data and suggest appropriate medical treatment."

[1788] Emergency Illness Prompt:

[1789] "Please explain to the user in a voice that you suddenly felt chest pain or shortness of breath. Please analyze the voice data and quickly suggest the necessary emergency response."

[1790] In this way, the system of the present invention can utilize voice data and take emotional state into account to provide a more accurate and faster medical response.

[1791] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1792] Step 1:

[1793] User: Feeling unwell, the user talks about their symptoms and feelings into the device. Specifically, they input something like "I have a headache and my body feels tired." The input voice data is captured by the device's recording application. The output is an audio data file.

[1794] Step 2:

[1795] Device: Use a recording application to record the user's voice. Specifically, the user presses the "Start Recording" button, and then presses the "Stop Recording" button after finishing describing the symptoms. After recording is complete, the device saves the voice data in a digital file format (e.g., WAV format) and sends it to the server. The input is the user's voice, and the output is the voice data file sent to the server.

[1796] Step 3:

[1797] Server: Receives audio data sent from the device. Specifically, it receives audio files via HTTP requests. The audio data is stored in a database, and a timestamp and user ID are added to the stored data. The input is the sent audio data, and the output is the saved audio data file and its metadata.

[1798] Step 4:

[1799] Server: The saved voice data is input into the generative AI model and analyzed. Specifically, the voice data is spectrally analyzed and features such as breathing, voice intonation, and speaking rate are extracted. The input is the saved voice data, and the output is the extracted voice features.

[1800] Step 5:

[1801] Server (Emotion Engine): The emotion engine identifies the user's emotional state based on the analyzed voice data. Specifically, it uses a machine learning algorithm to identify emotional states such as stress and anxiety. The input is voice features, and the output is the identified emotional state.

[1802] Step 6:

[1803] Server: Evaluates the user's physical condition and symptoms based on the extracted features and emotional state. Specifically, it compares with past data and executes algorithms to evaluate the progression of symptoms. The input is voice features and emotion identification results, and the output is an evaluation of the user's physical condition and symptoms.

[1804] Step 7:

[1805] Server: Based on the evaluation results, selects and proposes appropriate medical responses. Specifically, if symptoms are mild, it sends a prescription to a pharmacy, if symptoms are severe, it dispatches a doctor, and if there is an emergency, it contacts emergency services. The input is the evaluation results of physical condition and symptoms, and the output is a medical response proposal.

[1806] Step 8:

[1807] Server: Notifies the user of the selected medical response. Specifically, it contacts the user via email or SMS and instructs them on the necessary response. The input is the medical response proposal, and the output is a notification message to the user.

[1808] (Application example 2)

[1809] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1810] Until now, there has been no system in brick-and-mortar stores that can quickly and accurately evaluate the physical and emotional state of customers and provide optimal medical care. As a result, customers have had to wait a long time to receive appropriate medical care, which could worsen their symptoms. To solve this problem, the present invention aims to provide a system that can analyze the voice of customers in brick-and-mortar stores, evaluate their physical and emotional state, and quickly propose and implement appropriate medical care.

[1811] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving and storing voice data, means for executing a generative AI model for analyzing the received voice data, means for evaluating the user's physical condition and emotional state from the analyzed data, and means for proposing and implementing appropriate medical treatment based on the evaluation results. This makes it possible to quickly evaluate the physical condition and emotional state of customers in a physical store and provide optimal medical treatment.

[1812] "Voice data" refers to data in which the contents of a user's speech are recorded in digital format.

[1813] A "generative AI model" is an algorithm that uses machine learning technology to analyze voice data and extract specific features and emotional states.

[1814] The "means for analyzing" refers to the means for inputting received voice data into a generative AI model and carrying out the process of analyzing the characteristics of the data.

[1815] "Physical and emotional state" refers to the user's physical health and psychological feelings and moods.

[1816] The "means for evaluation" is a means for carrying out a process of determining the user's physical condition and emotional state based on the analyzed data.

[1817] "Medical response" means taking action based on the user's physical or emotional state, including suggesting appropriate medication, recommending a doctor's appointment, or contacting emergency services.

[1818] The "means for proposing and implementing" is a means for carrying out the process of proposing specific medical measures to the user based on the evaluation results and implementing those measures.

[1819] A "physical store" refers to a physical store where a user actually visits and receives face-to-face service.

[1820] System Overview

[1821] This invention relates to a system that automatically evaluates the physical and emotional state of customers in brick-and-mortar stores and proposes and implements appropriate medical treatment. This system collects and analyzes voice data, evaluates their emotional state, and then proposes and implements the most appropriate medical treatment for the user.

[1822] System Configuration

[1823] The system consists of the following main components:

[1824] 1. Terminal: This includes a recording device (e.g., a tablet or dedicated microphone at the reception desk) installed in the physical store and used by customers. Customers talk to this terminal about their symptoms and condition.

[1825] 2. Server: This is the central system that receives and analyzes the voice data, runs the generative AI model, and makes evaluations and recommendations.

[1826] 3. Emotion engine: Analyzes the emotional state of customers from voice data and reflects the results in assessing their physical condition and symptoms.

[1827] 4. Database: Serves as data storage for saving voice data and analysis results.

[1828] How it works

[1829] Recording and sending audio

[1830] The user speaks about their symptoms and feelings into a device installed in a physical store. The device launches a recording application and records what is said. After recording, the audio data is sent to the server.

[1831] Receiving and storing audio data

[1832] The server receives the voice data sent from the terminal and stores it in a database.

[1833] Analysis of audio data

[1834] The server inputs the stored voice data into a generative AI model for analysis. The generative AI model extracts features such as breathing, voice intonation, speaking rate, and emotional state. An emotion engine analyzes the data and identifies the user's emotional state.

[1835] Assessment of physical condition and symptoms

[1836] The server evaluates the customer's physical condition and symptoms based on the extracted features and emotional state, and also compares this with past voice data.

[1837] Propose and take necessary actions

[1838] Based on the evaluation results and the user's emotional state, the server selects and implements appropriate medical responses, such as suggesting medication, recommending a doctor's visit, or contacting emergency services.

[1839] Specific examples

[1840] Example 1: Mild illness

[1841] 1. User: Complains of a slight headache and fatigue and describes his symptoms into the terminal.

[1842] 2. Server: Analyzes the voice data and evaluates the symptoms as mild. The emotion engine also determines that the stress level is low.

[1843] 3. Server: Display a notification on the terminal to customers saying, "Please use over-the-counter headache medicine."

[1844] Example 2: Urgent illness

[1845] 1. User: Complains of chest pain and shortness of breath and describes his symptoms into the terminal.

[1846] 2. Server: Analyzes the voice data and detects high levels of urgency. At the same time, the emotion engine recognizes strong anxiety and fear.

[1847] 3. Server: Immediately dispatch emergency services and display a notification to the customer on their device saying, "Emergency response required. An ambulance is on the way."

[1848] Prompt Sentence Examples

[1849] User Voice:

[1850] "I've been having terrible headaches lately and I'm feeling stressed."

[1851] System response:

[1852] Emotional state: Stressed, high

[1853] Symptom assessment: mild

[1854] Suggestion: "Take advantage of over-the-counter headache medication. I recommend relaxation techniques to help relieve stress."

[1855] In this way, the system of the present invention can quickly assess the physical and emotional state of customers in a physical store and provide optimal medical care.

[1856] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1857] Step 1:

[1858] The user talks about their symptoms and feelings into a device installed in a physical store. The device launches a recording application and records the user's voice. When the recording is finished, the device sends the voice data (input) to the server (output).

[1859] Step 2:

[1860] The server receives the voice data sent from the terminal and stores it in a database. Specifically, after receiving the voice data (input), it records (outputs) it in the database.

[1861] Step 3:

[1862] The server inputs the saved voice data into the generative AI model for analysis. The server analyzes the voice data (input) and extracts features such as breathing, voice intonation, speaking rate, and emotional state (output). Specifically, it runs the generative AI model and obtains the results of analyzing each feature.

[1863] Step 4:

[1864] The server uses an emotion engine to identify the user's emotional state from the analyzed data. The emotion engine performs data calculations based on the input voice features and evaluates (outputs) the user's emotional state.

[1865] Step 5:

[1866] The server evaluates the user's physical condition and symptoms based on the features and emotional state extracted from the voice data. It compares the current condition with past data and evaluates it. In this process, the server derives the evaluation results (output) of the physical condition and symptoms based on the features and emotional state (input).

[1867] Step 6:

[1868] The server selects and implements appropriate medical responses based on the evaluation results and the user's emotional state. Specifically, it suggests medication, recommends a doctor's consultation, or contacts emergency services depending on the user's physical condition and the urgency of the symptoms. The input in this step is the evaluation results and emotional state, and the output is the proposal or implementation of specific medical responses.

[1869] Step 7:

[1870] The server notifies the user either through the server itself or the terminal. For example, if a medication is suggested, a notification saying "Please use over-the-counter headache medicine" is displayed on the terminal. In this step, a notification (output) to the user is created based on the medical treatment suggestion (input).

[1871] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1872] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1873] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1874] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1875] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1876] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1877] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1878] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1879] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1880] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1881] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1882] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1883] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1884] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1885] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1886] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1887] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1888] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1889] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1890] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1891] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1892] The following is further disclosed regarding the above embodiment.

[1893] (Claim 1)

[1894] means for receiving and storing audio data;

[1895] means for executing a generative AI model to analyze the received audio data;

[1896] A means of evaluating the patient's physical condition and symptoms from the analyzed data;

[1897] A means of proposing appropriate medical responses based on the evaluation results;

[1898] A system including:

[1899] (Claim 2)

[1900] 2. The system of claim 1, wherein the generative AI model analyzes breathing, voice intonation, and speaking rate from audio data.

[1901] (Claim 3)

[1902] 10. The system of claim 1, wherein the system automatically sends a prescription, dispatches a doctor, or contacts emergency services depending on the assessed patient's physical condition and the urgency of the symptoms.

[1903] "Example 1"

[1904] (Claim 1)

[1905] means for receiving and storing audio data;

[1906] means for executing a generative AI model to analyze the received audio data;

[1907] A means for extracting features from the analyzed data and recording them in a database;

[1908] A method for evaluating the patient's physical condition and symptoms from the extracted features and comparing them with past data.

[1909] a means for classifying the urgency of the symptoms based on the evaluation results;

[1910] A means of proposing appropriate medical responses based on the evaluation results;

[1911] A system including:

[1912] (Claim 2)

[1913] The system of claim 1, wherein the generative AI model analyzes breathing, voice intonation, and speaking rate from audio data and extracts features.

[1914] (Claim 3)

[1915] 10. The system of claim 1, wherein the system automatically sends a prescription, dispatches a medical professional, or contacts emergency services depending on the assessed patient condition and the urgency of the symptoms.

[1916] "Application Example 1"

[1917] (Claim 1)

[1918] means for receiving and storing audio data;

[1919] means for executing a generative AI model to analyze the received audio data;

[1920] A means for evaluating the target situation or abnormality from the analyzed data;

[1921] A means of proposing or implementing appropriate measures based on the results of the assessment;

[1922] the systems that contain them.

[1923] (Claim 2)

[1924] The system of claim 1, wherein the generative AI model analyzes voice tension, abnormal sounds, and alarm sounds from voice data.

[1925] (Claim 3)

[1926] The system according to claim 1, which automatically sends a notification, dispatches a guard, or contacts the police depending on the evaluated condition of the object and the urgency of the abnormality.

[1927] "Example 2: Combining Emotion Engines"

[1928] (Claim 1)

[1929] A means for a user to record voice data and transmit the voice data to a server;

[1930] means for receiving and storing audio data;

[1931] The stored voice data is input into a generative AI model to extract and analyze characteristics such as breathing, voice intonation, and speaking speed.

[1932] A means for assessing the user's physical condition, symptoms, and emotional state from the analyzed data;

[1933] A means of suggesting appropriate medical responses (e.g., sending a prescription, dispatching a doctor, or contacting emergency services) based on the results of the assessment;

[1934] A system including:

[1935] (Claim 2)

[1936] 2. The system of claim 1, wherein the generative AI model analyzes breathing, voice intonation, and speaking rate from the voice data, and analyzes the user's emotional state using an emotion engine.

[1937] (Claim 3)

[1938] 10. The system of claim 1, further comprising means for automatically sending a prescription, dispatching a doctor, or contacting emergency services depending on the user's assessed condition and the urgency of the symptoms.

[1939] "Application example 2 when combining emotion engines"

[1940] (Claim 1)

[1941] means for receiving and storing audio data;

[1942] means for executing a generative AI model to analyze the received audio data;

[1943] A means for evaluating the user's physical and emotional state from the analyzed data;

[1944] A means to propose and implement appropriate medical responses based on the assessment results;

[1945] A system including:

[1946] (Claim 2)

[1947] 10. The system of claim 1, wherein the generative AI model analyzes breathing, voice intonation, speaking rate, and emotional state from the audio data.

[1948] (Claim 3)

[1949] 10. The system of claim 1, wherein the system automatically suggests medication, recommends a doctor's visit, or contacts emergency services depending on the urgency of the analyzed user's physical and emotional state. [Explanation of symbols]

[1950] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving and storing audio data; means for executing a generative AI model to analyze the received audio data; A means of evaluating the patient's physical condition and symptoms from the analyzed data; A means of proposing appropriate medical responses based on the evaluation results; A system including:

2. The system of claim 1 , wherein the generative AI model analyzes breathing, voice intonation, and speaking rate from audio data.

3. 10. The system of claim 1, wherein the system automatically sends a prescription, dispatches a doctor, or contacts emergency services depending on the assessed patient condition and the urgency of the symptoms.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A