System

A system integrating natural language processing and image analysis with medical databases addresses inefficiencies in medical record entry and diagnostic support by automating the collection and organization of subjective and visual patient data, enhancing diagnostic accuracy.

JP2026018059APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119120
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Medical records in patient settings are inefficiently and inaccurately entered, with subjective and visual information difficult to integrate, leading to complex information organization and insufficient diagnostic support accuracy.

Method used

A system combining natural language processing, image analysis, and medical record database technologies to extract and store subjective and visual patient information, generating diagnostic support based on past data and guidelines.

Benefits of technology

Automates medical record entry, reduces provider burden, and improves diagnostic accuracy by efficiently and accurately collecting and organizing patient information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026018059000001_ABST
    Figure 2026018059000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: natural language processing means; image analysis means; medical record database means; means for extracting subjective information of a patient using the natural language processing means; means for extracting visual information of the patient using the image analysis means; means for storing the extracted subjective information and visual information in the medical record database means; and means for generating diagnosis support information based on the stored information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In medical settings, patient medical records must be entered efficiently and accurately, but this task is time-consuming and places a heavy burden on medical professionals. It is also difficult to accurately capture patients' subjective symptoms and visual information, which can affect the accuracy of diagnoses. Current systems struggle to integrate subjective and visual information into medical records, making the organization and analysis of information complex and resulting in insufficient diagnostic support accuracy. There is a need to resolve these issues, automate medical record entry, reduce the burden on medical providers, and improve diagnostic accuracy. [Means for solving the problem]

[0005] The present invention solves the above-mentioned problems by providing a system that combines natural language processing means, image analysis means, and medical record database means. Specifically, the natural language processing means is used to extract subjective information about the patient, and the image analysis means is used to extract visual information about the patient. This information is stored in the medical record database means, and diagnostic support information is generated based on the stored information, thereby achieving efficient and accurate information organization and diagnostic support. In addition, the natural language processing means includes a function for converting voice input into text, and the image analysis means includes a function for analyzing the patient's complexion, facial expression, and movements. Furthermore, because the diagnostic support information is generated based on past data and medical guidelines, the accuracy of diagnosis can be improved.

[0006] "Natural language processing means" refers to technology that analyzes natural language text and extracts meaning and information.

[0007] "Image analysis means" refers to technology that analyzes image data and extracts specific features or information from it.

[0008] "Medical record database means" refers to a database system for storing and managing patient medical records.

[0009] "Subjective information" refers to information based on the symptoms and emotions felt by the patient themselves.

[0010] "Visual information" refers to information related to vision, such as the patient's complexion, facial expression, and movements.

[0011] "Diagnostic support information" refers to information that provides assistance in diagnosis based on collected medical information.

[0012] "Voice input" refers to the act of collecting voice data using a device such as a microphone.

[0013] "Convert to text" refers to the process of converting audio data into written information.

[0014] "Complexion" refers to the color and complexion of the patient's face.

[0015] "Facial expression" refers to emotions and states shown by the movement of facial muscles.

[0016] "Movement" refers to the actions or states indicated by the patient's body movements.

[0017] "Historical Data" refers to medical information and patient data that has been previously collected and recorded.

[0018] "Medical guidelines" refer to standards and guidelines that define standard treatment methods and responses in the medical industry. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0027] [First embodiment]

[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0040] The system of the present invention is composed of a combination of natural language processing means, image analysis means, and medical record database means, and is capable of efficiently and accurately collecting subjective and visual information from patients and reflecting it in medical records.

[0041] Implementation of natural language processing methods

[0042] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device converts this speech into text data using a speech recognition module. The converted text data is sent to a server, which analyzes it using a natural language processing engine to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[0043] Implementation of image analysis procedures

[0044] The user uses the device's camera to capture visual information about the patient. For example, the camera can be used to record the patient's complexion and facial expression. This image data is sent from the device to a server, which then uses an image analysis engine to analyze the data and extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[0045] Implementation of medical record database measures

[0046] The medical record database means organizes and stores the collected subjective and visual information. The server generates diagnostic support information based on the information stored in this database. For example, based on subjective information such as "My headaches have been getting worse recently" and visual information such as "Your face is pale" and "Your expression looks painful," the server refers to past data and medical guidelines and generates diagnostic support information such as "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[0047] Specific examples

[0048] Specific examples are shown below.

[0049] Example of subjective information extraction

[0050] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[0051] Example of information extraction using image analysis

[0052] The image of the patient's face taken with a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[0053] Example of generating diagnostic support information

[0054] Based on subjective information such as "My back has been hurting recently" and visual information such as "My face is red and swollen," the server determines that "there is a high possibility of an inflammatory disease." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[0055] In this way, the system of the present invention efficiently and accurately collects subjective and visual information and reflects it in medical records, automating the process of entering medical records, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[0056] The processing flow will be explained below.

[0057] Extracting subjective information

[0058] Step 1:

[0059] The user inputs the patient's subjective symptoms into the terminal by voice or text. For example, consider the case where the user inputs by voice, "My headache has been getting worse recently."

[0060] Step 2:

[0061] The device converts the voice data into text data using a speech recognition module, and the resulting text data is "My headache has been getting worse recently."

[0062] Step 3:

[0063] The device sends the generated text data to the server, securely using the HTTPS protocol.

[0064] Step 4:

[0065] The server uses a natural language processing engine to analyze the text data and extract key subjective information, such as "headache" and "getting worse."

[0066] Step 5:

[0067] The server organizes the extracted subjective information and stores it in a medical record database, which electronically records the patient's complaints.

[0068] Image analysis

[0069] Step 1:

[0070] The user uses the device's camera to capture visual information of the patient, for example, recording the patient's complexion and facial expressions.

[0071] Step 2:

[0072] The device sends the captured image data to the server. The image is sent in a compressed format and the data is transferred securely using the HTTPS protocol.

[0073] Step 3:

[0074] The server uses an image analysis engine to analyze the image data, specifically analyzing the patient's complexion, facial expression, and movements to extract characteristic information.

[0075] Step 4:

[0076] For example, information such as "his face is pale" and "his expression looks like he's in pain" is extracted.

[0077] Step 5:

[0078] The server stores the extracted visual information in the medical record database, so that the visual information is reflected in the medical record.

[0079] Diagnostic support

[0080] Step 1:

[0081] The server integrates and analyzes the subjective and visual information stored in the medical record database.

[0082] Step 2:

[0083] The server generates diagnostic support information by referencing past data and medical guidelines. For example, it analyzes information such as "Your headaches have gotten worse recently" and "Your face is pale" and "You look like you're in pain."

[0084] Step 3:

[0085] Based on the integrated information, the server generates diagnostic support information such as "high possibility of migraine."

[0086] Step 4:

[0087] The server transmits the generated diagnostic assistance information to the terminal.

[0088] Step 5:

[0089] The device notifies the user of the diagnostic support information, for example, by providing information to a healthcare provider via a pop-up notification or email notification.

[0090] Through this process, the AISOAP system can efficiently and accurately collect patients' subjective and visual information and reflect it in medical records and diagnostic support information, automating the medical record entry process, reducing the burden on medical providers and improving the accuracy of diagnoses.

[0091] Example 1

[0092] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0093] In the modern medical environment, entering medical records is extremely time-consuming, placing a strain on healthcare providers. Furthermore, accurate diagnosis requires reliable collection of patient subjective and visual information, which must be integrated to generate diagnostic support information. However, conventional methods have proven difficult to achieve this efficiently and accurately.

[0094] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0095] In this invention, the server includes a natural language processing unit, an image analysis unit, and a medical record database unit, which enables efficient and accurate collection of subjective and visual information from patients and reflecting it in medical records.

[0096] A "natural language processing means" is a means for extracting subjective patient information from speech or text input.

[0097] The "image analysis means" is a means for extracting visual information of the patient from the camera image.

[0098] A "medical record database means" is a means for storing and managing extracted subjective and visual information.

[0099] "Voice input" refers to voice data uttered by a user through a microphone.

[0100] "Text input" refers to character data that a user inputs using a keyboard or the like.

[0101] A "voice recognition module" is a software or hardware component for converting voice input into text data.

[0102] "Camera footage" refers to video and still image data captured by a camera.

[0103] "Diagnostic support information" is information that helps with diagnosis and is generated based on stored information.

[0104] "User" refers to a healthcare provider or any individual or organization operating the System.

[0105] The system of the present invention is composed of a combination of natural language processing means, image analysis means, and medical record database means, and is capable of efficiently and accurately collecting subjective and visual information from patients and reflecting it in medical records.

[0106] Implementation of natural language processing methods

[0107] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device converts this speech into text data using the Google Cloud Speech-to-Text API. The converted text data is sent to the server. The server then analyzes the text data using a natural language processing engine (e.g., SpaCy or GPT-4) to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[0108] Implementation of image analysis procedures

[0109] The user uses the device's camera to capture visual information about the patient. For example, the camera can be used to record the patient's complexion and facial expression. This image data is sent from the device to a server, which then analyzes the data using an image analysis engine (e.g., OpenCV or TensorFlow) to extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in the medical record database.

[0110] Implementation of medical record database measures

[0111] The medical record database means organizes and stores the collected subjective and visual information. The server generates diagnostic support information based on the information stored in this database. For example, based on subjective information such as "My headaches have been getting worse recently" and visual information such as "Your face is pale" and "Your expression looks painful," the server refers to past data and medical guidelines and generates diagnostic support information such as "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[0112] Specific examples

[0113] Specific examples are shown below.

[0114] Example of subjective information extraction

[0115] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[0116] Example of information extraction using image analysis

[0117] The image of the patient's face taken with a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[0118] Example of generating diagnostic support information

[0119] Based on subjective information such as "My back has been hurting recently" and visual information such as "My face is red and swollen," the server determines that "there is a high possibility of an inflammatory disease." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[0120] Example prompt sentence:

[0121] "Convert the following speech input into text data and extract subjective information. Example: 'My back has been hurting lately.'"

[0122] By utilizing generative AI models, the system described above can collect patient information and provide diagnostic support information more efficiently and accurately than existing medical systems, reducing the burden on medical providers and improving diagnostic accuracy.

[0123] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0124] Step 1:

[0125] The user inputs the patient's symptoms into the terminal in voice or text format. When inputting by voice, the user speaks into a microphone, for example, "My headache has been getting worse recently." The input in this case is voice data or text data.

[0126] Step 2:

[0127] The device uses a speech recognition module to convert speech into text data. Specifically, it uses the Google Cloud Speech-to-Text API. This module receives speech data and converts it into text based on a language model. For example, the device outputs text data such as "My headache has been getting worse recently." This text data is then sent from the device to the server.

[0128] Step 3:

[0129] The server analyzes the received text data using a natural language processing engine (e.g., SpaCy or GPT-4). By inputting and analyzing the text data, subjective information such as "headache" and "getting worse" is extracted. The extracted information is stored in a medical record database. Specifically, the server receives the text data and uses the SpaCy library to extract keywords.

[0130] Step 4:

[0131] The user captures visual information of the patient using the device's camera. For example, they can take a picture of the patient's face and record their facial expression. The input image and video data are then sent from the device to the server.

[0132] Step 5:

[0133] The server analyzes the received image data using an image analysis engine (e.g., OpenCV or TensorFlow). The image data is input and visual information such as "pale face" or "pained expression" is extracted. The extracted information is also stored in the medical record database. Specifically, the server receives the image data and analyzes it using OpenCV.

[0134] Step 6:

[0135] The server generates diagnostic support information based on the subjective and visual information stored in the medical record database. For example, based on subjective information such as "Your headaches have gotten worse recently" and visual information such as "Your face is pale" and "Your expression looks painful," it references past data and medical guidelines to generate diagnostic support information such as "There is a high possibility of a migraine." The generated information is then stored again in the database and simultaneously sent to the device.

[0136] Step 7:

[0137] The device displays the received diagnostic support information on the user interface. The user is notified of new information and can confirm its contents. For example, the diagnostic support information "High possibility of migraine" may be displayed. This allows the user to quickly take necessary measures.

[0138] Through the above steps, patients' subjective and visual information can be collected efficiently and accurately and reflected in medical records, thereby reducing the burden on medical providers and improving diagnostic accuracy.

[0139] (Application example 1)

[0140] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0141] Conventional health management systems have had difficulty efficiently and accurately collecting subjective and visual information from patients to provide diagnostic support. Furthermore, they lacked a means to quickly notify users of the collected information, making it difficult to respond quickly in emergencies. The present invention aims to solve these problems and provide a health management system that is easy for healthcare providers and users to use.

[0142] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0143] In this invention, the server includes means for extracting subjective information of a patient using natural language processing means, means for extracting visual information of the patient using image analysis means, means for storing the extracted subjective information and visual information in a medical record database means, means for generating diagnostic support information based on the stored information, and means for notifying the smartphone terminal of the generated diagnostic support information, thereby enabling a quick analysis of the patient's health condition and a prompt response if necessary.

[0144] "Natural language processing means" is a technology for analyzing a patient's subjective information from voice or text input and extracting specific symptoms and conditions.

[0145] "Image analysis means" is a technology that analyzes the visual information of a patient obtained using a camera and quantitatively evaluates their health condition, such as their complexion and facial expression.

[0146] The "medical record database means" is a database technology that efficiently organizes and stores extracted subjective and visual information.

[0147] "Patient subjective information" is subjective data in audio or text format in which a patient describes their symptoms or physical condition.

[0148] "Patient visual information" refers to visual data such as the patient's appearance and facial expression obtained using a visual device such as a camera.

[0149] The "means for notifying a smartphone terminal" is a technology for sending the generated diagnostic assistance information to a smartphone and notifying the user in real time.

[0150] The "means for generating diagnostic support information" refers to a technology that generates information that assists in specific diagnoses based on stored subjective and visual information, and refers to medical data and guidelines.

[0151] The present invention provides a system for efficiently and accurately collecting subjective and visual information from a patient and providing diagnostic support based on that information. Specific embodiments for carrying out the present invention will be described below.

[0152] Hardware or software used

[0153] Hardware

[0154] Smartphone: Equipped with a microphone, speaker, and camera.

[0155] Server: For data analysis and database management.

[0156] software

[0157] Voice Recognition:

[0158] API: Google Speech Recognition API

[0159] Library: speech_recognition

[0160] Image analysis:

[0161] Deep Learning Model: A pre-trained image classification model using TensorFlow

[0162] Libraries: cv2 (OpenCV), PIL (Python Imaging Library), tensorflow

[0163] Generation of diagnostic support information:

[0164] API: Endpoint for external diagnostic support systems

[0165] Data processing and calculation

[0166] Audio data processing

[0167] The user launches the smartphone application and reports their current symptoms and physical condition by voice input. The smartphone uses a voice recognition module to convert the voice data into text, and then sends the text data to the server. The server uses a natural language processing engine to extract specific symptoms and conditions (e.g., "I have a headache" or "I have a fever") from the text data and stores them in a medical record database.

[0168] Image data processing

[0169] The user records their own facial expression and complexion using the smartphone camera. This image data is sent to a server, which then analyzes it using an image analysis engine. As a result, visual information such as "pale complexion" or "pained facial expression" is extracted and stored in a medical record database.

[0170] Generation and notification of diagnostic support information

[0171] The server generates diagnostic support information based on the subjective and visual information stored in the medical record database. The accuracy of this diagnostic support information is improved by referencing medical data and guidelines. The generated diagnostic support information is then notified to the user via a smartphone application.

[0172] Specific examples

[0173] Usage example

[0174] For example, a user wakes up in the morning feeling unwell and launches the app, then speaks the following:

[0175] Recently, I've been feeling tired and feverish.

[0176] Based on this voice input, the app performs an analysis, then uses the smartphone camera to take and analyze a photograph of the complexion. If the result is a diagnosis of "visual problems," the app will finally notify you as follows:

[0177] Your health may be deteriorating. We recommend that you consult a medical institution.

[0178] In this way, the system can quickly analyze the patient's health condition and help them take prompt action if necessary.

[0179] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0180] Step 1:

[0181] The user starts the smartphone application and reports their current symptoms and physical condition by voice input. At this time, the user speaks into the microphone. The input is the user's voice data, and the output is the voice data recorded by the smartphone device.

[0182] Step 2:

[0183] The device uses a speech recognition module (Google Speech Recognition API) to convert voice data into text data. The input is the user's voice data, and the output is symptom information in text format. The data processing performed here is text conversion based on speech recognition.

[0184] Step 3:

[0185] The terminal sends the converted text data to the server. The input is text data, and the output is data transmission to the server. There is no data processing, but communication is included.

[0186] Step 4:

[0187] The server uses a natural language processing engine to analyze the text data and extract specific symptoms and conditions (e.g., "I have a headache" or "I have a fever"). The input is text data, and the output is the analyzed symptom information. Data processing is an extraction process using natural language processing.

[0188] Step 5:

[0189] The server saves the analyzed symptom information to the medical record database. The input is the analyzed symptom information, and the output is the saving operation to the database. There is no data processing, but the saving operation is included.

[0190] Step 6:

[0191] A user records their facial expression and complexion using a smartphone camera. The device then acquires the image data. The input is the image data captured by the camera, and the output is the recorded image.

[0192] Step 7:

[0193] The terminal sends image data to the server. The input is image data, and the output is data transmission to the server. There is no data processing, but communication is included.

[0194] Step 8:

[0195] The server analyzes the image data using an image analysis engine (a pre-trained model using TensorFlow) and extracts visual information such as "the face looks pale" or "the face looks painful." The input is image data, and the output is the analyzed visual information. Data processing is the extraction process using image analysis.

[0196] Step 9:

[0197] The server stores the analyzed visual information in a medical record database. The input is the analyzed visual information, and the output is the saving operation to the database. There is no data processing, but the saving operation is included.

[0198] Step 10:

[0199] The server generates diagnostic support information based on the subjective and visual information stored in the medical record database. The input is subjective and visual information, and the output is diagnostic support information. Data processing involves information integration and analysis based on diagnostic algorithms.

[0200] Step 11:

[0201] The server notifies the smartphone of the generated diagnostic support information. The input is the diagnostic support information, and the output is a notification to the smartphone terminal. There is no data processing, but communication is involved.

[0202] Step 12:

[0203] The smartphone terminal displays the received diagnostic support information to the user. The input is the diagnostic support information, and the output is the presentation of information to the user. The specific operation is to display a notification message on the screen.

[0204] In this way, by going through each step, the user's health condition can be quickly analyzed and prompt action can be taken if necessary.

[0205] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0206] The system of the present invention is composed of a natural language processing means, an image analysis means, a medical record database means, and an emotion engine that recognizes the user's emotions. This system allows the efficient and accurate collection of subjective, visual, and emotional information of patients and reflects it in medical records.

[0207] Implementation of natural language processing methods

[0208] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device converts this speech into text data using a speech recognition module. The converted text data is sent to a server, which analyzes it using a natural language processing engine to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[0209] Implementation of image analysis procedures

[0210] The user uses the device's camera to capture visual information about the patient. For example, they can record the patient's complexion and facial expression. This image data is sent from the device to a server, which then uses an image analysis engine to analyze the data and extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[0211] Emotion Engine Implementation

[0212] The user inputs the patient's symptoms in voice or text format, and the emotion engine simultaneously analyzes the input data and extracts emotions. Specifically, for voice data, the server analyzes the tone and tempo of the voice to recognize emotions such as "irritability" or "sadness." For text data, the server works with a natural language processing engine to extract emotions from the text. The resulting emotional information is also stored in the medical record database.

[0213] Implementation of medical record database measures

[0214] The medical record database means organizes and stores the collected subjective, visual, and emotional information. The server generates diagnostic support information based on the information stored in this database. For example, if the subjective information "My headaches have been getting worse recently" is recognized, along with visual information such as "Your face is pale" and "Your facial expression looks painful," and a "feeling of irritability" is also recognized, the server will refer to past data and medical guidelines to generate diagnostic support information stating "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[0215] Specific examples

[0216] Specific examples are shown below.

[0217] Example of subjective information extraction

[0218] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[0219] Example of information extraction using image analysis

[0220] The image of the patient's face taken with a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[0221] Example of information extraction using emotion engine

[0222] The user speaks, "I've been feeling sad lately." The device converts this into text data, "I've been feeling sad lately," and sends it to the server. The server uses an emotion engine to extract the emotion "sadness" and stores it in the medical record database.

[0223] Example of generating diagnostic support information

[0224] Based on the subjective information "I've had a backache recently," the visual information "my face is red and swollen," and the emotional information "sadness," the server determines that "the back pain may be related to psychological or mental stress." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[0225] In this way, the system of the present invention efficiently and accurately collects subjective, visual, and emotional information and reflects it in medical records, thereby automating the process of entering medical records, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[0226] The processing flow will be explained below.

[0227] Implementation of natural language processing methods

[0228] Step 1:

[0229] The user inputs the patient's subjective symptoms into the terminal by voice or text. For example, consider the case where the user inputs by voice, "My headache has been getting worse recently."

[0230] Step 2:

[0231] The device converts the voice data into text data using a speech recognition module, and the resulting text data is "My headache has been getting worse recently."

[0232] Step 3:

[0233] The device sends the generated text data to the server, securely using the HTTPS protocol.

[0234] Step 4:

[0235] The server uses a natural language processing engine to analyze the text data and extract key subjective information, such as the phrases "headache" and "getting worse."

[0236] Step 5:

[0237] The server organizes the extracted subjective information and stores it in a medical record database, which electronically records the patient's complaints.

[0238] Implementation of image analysis procedures

[0239] Step 1:

[0240] The user uses the device's camera to capture visual information of the patient, for example, recording the patient's complexion and facial expressions.

[0241] Step 2:

[0242] The device sends the captured image data to the server. The image is sent in a compressed format and the data is transferred securely using the HTTPS protocol.

[0243] Step 3:

[0244] The server uses an image analysis engine to analyze the image data, specifically analyzing the patient's complexion, facial expression, and movements to extract characteristic information.

[0245] Step 4:

[0246] For example, information such as "his face is pale" and "his expression looks like he's in pain" is extracted.

[0247] Step 5:

[0248] The server stores the extracted visual information in the medical record database, so that the visual information is reflected in the medical record.

[0249] Emotion Engine Implementation

[0250] Step 1:

[0251] As the user inputs the patient's symptoms via voice or text, the emotion engine simultaneously analyzes the input data and extracts emotions.

[0252] Step 2:

[0253] In the case of voice data, the terminal transmits the voice data to the server, and in the case of text data, the text data is also transmitted to the server.

[0254] Step 3:

[0255] The server analyzes the tone and tempo of the voice to recognize emotions such as "anxiety" or "sadness." It also extracts emotions from text data.

[0256] Step 4:

[0257] The extracted emotional information is stored in the medical record database, so that the emotional information is also reflected in the medical record.

[0258] Diagnostic support

[0259] Step 1:

[0260] The server integrates and analyzes the subjective, visual, and emotional information stored in the medical record database.

[0261] Step 2:

[0262] The server generates diagnostic support information by referencing past data and medical guidelines. For example, it analyzes information such as "My headache has gotten worse recently," "My face is pale," and "I feel irritable."

[0263] Step 3:

[0264] Based on the integrated information, the server generates diagnostic support information such as "high possibility of migraine."

[0265] Step 4:

[0266] The server transmits the generated diagnostic assistance information to the terminal.

[0267] Step 5:

[0268] The device notifies the user of the diagnostic support information, for example, by providing information to a healthcare provider via a pop-up notification or email notification.

[0269] Through this process, the AISOAP system can efficiently and accurately collect patients' subjective, visual, and emotional information and reflect it in medical records and diagnostic support information, automating the input work of medical records, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[0270] Example 2

[0271] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0272] Modern healthcare requires efficient and accurate collection of patients' subjective, visual, and emotional information and integration of it into medical records. However, traditional methods make it difficult to centrally collect and integrate this information into medical records, which increases the burden on healthcare providers and potentially impacts diagnostic accuracy. To address this challenge, a system that can automatically analyze and integrate data collected from multiple sources is needed.

[0273] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a natural language processing means, an image analysis means, a medical record database means, and an emotion recognition means. This makes it possible to efficiently and accurately collect subjective information, visual information, and emotion information of the patient and reflect it in the medical record.

[0274] "Natural language processing means" is a technology that analyzes input data in voice or text format and extracts meaning and keywords.

[0275] "Image analysis means" is a technology that analyzes video data acquired from a camera or other device and extracts visual information such as the patient's complexion, facial expression, and movements.

[0276] The "medical record database means" is a database system that organizes and stores collected subjective, visual, and emotional information.

[0277] "Emotion recognition means" is a technology that analyzes voice and text data and extracts emotions from them.

[0278] "Subjective information" refers to information such as symptoms and discomfort reported by the patient themselves.

[0279] "Visual information" refers to information extracted from video data acquired using a camera or the like, such as the patient's complexion, facial expression, and movements.

[0280] "Emotion information" is information about the patient's emotions extracted from voice and text data.

[0281] "Diagnostic support information" is information that assists diagnosis and is generated based on collected subjective information, visual information, and emotional information.

[0282] The system of the present invention is composed of a combination of natural language processing means, image analysis means, medical record database means, and emotion recognition means, and is capable of efficiently and accurately collecting subjective, visual, and emotional information from patients and reflecting it in medical records.

[0283] Implementation of natural language processing methods

[0284] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device uses a voice recognition module (e.g., a voice recognition API) to convert the speech into text data. The converted text data is sent to a server, which then analyzes the text data using a natural language processing engine (e.g., SpaCy) to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[0285] Implementation of image analysis procedures

[0286] The user uses the device's camera to capture visual information about the patient. For example, they can record the patient's complexion and facial expression. This image data is sent from the device to a server, which then analyzes the data using an image analysis engine (e.g., OpenCV) to extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[0287] Implementing emotion recognition measures

[0288] The user inputs the patient's symptoms in voice or text format, and at the same time, the emotion recognition means analyzes the input data and extracts emotions. Specifically, in the case of voice data, the server analyzes the tone and tempo of the voice to recognize emotions such as "irritability" or "sadness." In the case of text data, emotions are extracted from the text in cooperation with a natural language processing engine. The emotional information obtained in this way is also stored in the medical record database.

[0289] Implementation of medical record database measures

[0290] The medical record database means organizes and stores the collected subjective, visual, and emotional information. The server generates diagnostic support information based on the information stored in this database. For example, if the subjective information "My headaches have been getting worse recently" is recognized, along with visual information such as "Your face is pale" and "Your facial expression looks painful," and a "feeling of irritability" is also recognized, the server will refer to past data and medical guidelines to generate diagnostic support information stating "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[0291] Specific examples

[0292] Example of subjective information extraction

[0293] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[0294] Example of information extraction using image analysis

[0295] The camera takes a picture of the patient's face and sends the image to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[0296] Example of information extraction using emotion recognition methods

[0297] The user speaks, "I've been feeling sad lately." The device converts this into text data, "I've been feeling sad lately," and sends it to the server. The server uses emotion recognition means to extract the emotion "sadness" and stores it in the medical record database.

[0298] Example of generating diagnostic support information

[0299] Based on the subjective information "I've had a backache recently," the visual information "my face is red and swollen," and the emotional information "sadness," the server determines that "the back pain may be related to psychological or mental stress." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[0300] Through the above process, this system can efficiently and accurately collect patients' subjective, visual, and emotional information and reflect it in medical records, which is an important tool for automating medical record entry work, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[0301] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0302] Step 1: User Input

[0303] The user inputs the patient's subjective symptoms into the terminal in voice or text format. This input becomes the basic data for the system. Specifically, the user says, "My headache has been getting worse recently."

[0304] input:

[0305] Speech data (e.g., "My headaches have been getting worse lately.")

[0306] output:

[0307] Audio data (as is)

[0308] Step 2: Voice Recognition

[0309] The device uses a speech recognition module (e.g., speech recognition API) to convert the speech into text data. At this stage, the speech becomes text data, which is used in the next step.

[0310] input:

[0311] Speech data (e.g., "My headaches have been getting worse lately.")

[0312] output:

[0313] Text data (e.g., "My headaches have been getting worse lately.")

[0314] Operation:

[0315] The audio waveform is analyzed and converted into string data.

[0316] Step 3: Send to the server

[0317] The device then sends the converted text data to the server, where it is analyzed, so fast and accurate transmission is essential.

[0318] input:

[0319] Text data (e.g., "My headaches have been getting worse lately.")

[0320] output:

[0321] Text data (sent as is to the server)

[0322] Operation:

[0323] The text data is uploaded to a server via a network.

[0324] Step 4: Natural Language Processing

[0325] The server uses a natural language processing engine (e.g., SpaCy) to analyze the received text data, and extracts subjective information (e.g., "headache" and "getting worse") as a result of the analysis.

[0326] input:

[0327] Text data (e.g., "My headaches have been getting worse lately.")

[0328] output:

[0329] Subjective information (e.g., "headache" or "it's getting worse")

[0330] Operation:

[0331] Morphological analysis of text data is performed to extract important words and phrases.

[0332] Step 5: Preserving subjective information

[0333] The server stores the extracted subjective information in a medical record database, which is used for subsequent analysis and medical support.

[0334] input:

[0335] Subjective information (e.g., "headache" or "it's getting worse")

[0336] output:

[0337] Saving to a database

[0338] Operation:

[0339] The extracted information is written to a medical record database.

[0340] Step 6: Capture visual information

[0341] The user uses the device's camera to capture visual information of the patient, for example, recording the patient's complexion and facial expressions, and saves this data.

[0342] input:

[0343] Video data (e.g., patient's complexion and facial expression)

[0344] output:

[0345] Video data (as is)

[0346] Operation:

[0347] Use a camera to capture video data.

[0348] Step 7: Sending visual data to the server

[0349] The device sends the recorded image data to a server, where it is analyzed, so it must be sent quickly and accurately.

[0350] input:

[0351] Video data (e.g., patient's complexion and facial expression)

[0352] output:

[0353] Video data (sent as is to the server)

[0354] Operation:

[0355] The video data is uploaded to a server via a network.

[0356] Step 8: Image analysis

[0357] The server analyzes the received video data using an image analysis engine (e.g., OpenCV), and extracts visual information (e.g., "the face is pale" or "the facial expression looks painful").

[0358] input:

[0359] Video data (e.g., patient's complexion and facial expression)

[0360] output:

[0361] Visual information (e.g., "He looks pale" or "He looks like he's in pain")

[0362] Operation:

[0363] Frame analysis of video data is performed to extract important visual information.

[0364] Step 9: Save the visual information

[0365] The server stores the extracted visual information in a medical record database, which is used for subsequent analysis and medical support.

[0366] input:

[0367] Visual information (e.g., "He looks pale" or "He looks like he's in pain")

[0368] output:

[0369] Saving to a database

[0370] Operation:

[0371] The extracted information is written to a medical record database.

[0372] Step 10: Emotion Recognition

[0373] The server analyzes the voice or text data entered by the user using an emotion recognition engine (e.g., emotion recognition API) to extract emotional information.

[0374] input:

[0375] Voice or text data (e.g., "I've been feeling sad lately")

[0376] output:

[0377] Emotional information (e.g., "sadness")

[0378] Operation:

[0379] Analyze the tone and content of voice and text data to identify emotions.

[0380] Step 11: Storing Emotional Information

[0381] The server stores the extracted emotion information in a medical record database, which is used for subsequent analysis and medical support.

[0382] input:

[0383] Emotional information (e.g., "sadness")

[0384] output:

[0385] Saving to a database

[0386] Operation:

[0387] The extracted emotion information is written to a medical record database.

[0388] Step 12: Generating diagnostic support information

[0389] The server generates diagnostic support information based on the stored subjective, visual, and emotional information. For example, if the server recognizes subjective information such as "My headaches have been getting worse recently," along with visual information such as "Your face is pale" and "Your facial expression looks painful," as well as a sense of irritability, the server will refer to past data and medical guidelines and generate diagnostic support information such as "There is a high possibility that you have a migraine."

[0390] input:

[0391] Subjective information, visual information, emotional information

[0392] output:

[0393] Diagnostic support information (e.g., "Probable migraine").

[0394] Operation:

[0395] Integrated analysis is performed based on the stored data to generate diagnostic support information.

[0396] Step 13: Send to device

[0397] The server sends the generated diagnostic support information to the terminal and notifies the user. This information serves as a diagnostic support tool for medical providers.

[0398] input:

[0399] Diagnostic support information (e.g., "Probable migraine").

[0400] output:

[0401] Diagnostic support information (sent directly to the device)

[0402] Operation:

[0403] The diagnostic assistance information is transmitted to a terminal via a network.

[0404] (Application example 2)

[0405] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0406] Conventional health monitoring systems only analyze subjective and visual information, making it difficult to accurately grasp an employee's emotional state. This makes it difficult to detect abnormalities in health status early and take appropriate measures. Furthermore, because it takes time for employees to report their health status in detail, there is a risk of incomplete information or false reports. The present invention aims to solve these problems and provide a system that efficiently and accurately monitors the health status of employees in factories.

[0407] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes natural language processing means, image analysis means, medical record database means, emotion recognition means, means for extracting patient emotion information using the natural language processing means and emotion recognition means, and means for generating diagnostic support information based on the stored information. This makes it possible to comprehensively collect and analyze employee subjective information, visual information, and emotion information. Ultimately, it is possible to more accurately grasp the employee's health status, detect abnormalities early, and take appropriate action.

[0408] "Natural language processing means" is a technology that analyzes voice and text data and extracts meaning from its content.

[0409] "Image analysis means" is a technique for analyzing image data and extracting features contained therein.

[0410] "Medical record database means" refers to a database system that collects and manages health information of patients and employees.

[0411] "Means for extracting subjective patient information" refers to a technology that uses natural language processing to extract symptoms and sensations reported by patients from text data.

[0412] The "means for extracting visual information of a patient" is a technology that uses image analysis means to extract information related to the patient's health condition from their appearance, such as their complexion and facial expression.

[0413] "Emotion recognition means" is a technology that analyzes emotions from voice or text and extracts that emotional information.

[0414] The "means for generating diagnostic support information" is a system that automatically proposes appropriate diagnoses and countermeasures based on collected health information.

[0415] The system of the present invention combines natural language processing, image analysis, a medical record database, and emotion recognition, and is capable of efficiently and accurately collecting employee subjective health information, visual health information, and emotion information to monitor health status and detect abnormalities.

[0416] Implementation of natural language processing methods

[0417] The user inputs their physical condition and symptoms in voice or text format. If the user says, "My headache has been getting worse recently," the device converts this voice data into text using a voice recognition module. The converted text data is sent to a server, which analyzes it using a natural language processing engine to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[0418] Implementation of image analysis procedures

[0419] The device uses a built-in camera to capture visual information about the employee, such as their complexion and facial expression. This image data is sent from the device to a server, which then uses an image analysis engine to analyze the data and extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[0420] Implementing emotion recognition measures

[0421] When a user reports their symptoms in voice or text format, the emotion recognition means simultaneously analyzes the input data and extracts emotions. Specifically, for voice data, the server analyzes the tone and tempo of the voice to recognize emotions such as "irritability" or "sadness." For text data, the server works with a natural language processing engine to extract emotions from the text. The emotional information obtained in this way is also stored in the medical record database.

[0422] Implementation of medical record database measures

[0423] The medical record database means organizes and stores the collected subjective, visual, and emotional information. The server generates diagnostic support information based on the information stored in this database. For example, if the subjective information "My headaches have been getting worse recently" is recognized, along with visual information such as "Your face is pale" and "Your facial expression looks painful," and a "feeling of irritability" is also recognized, the server will refer to past data and medical guidelines to generate diagnostic support information stating "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[0424] Specific examples

[0425] Specific examples are shown below.

[0426] Example of subjective information extraction

[0427] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[0428] Example of information extraction using image analysis

[0429] The image of the employee's face captured by a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling" and stores it in a medical record database.

[0430] Example of information extraction using emotion engine

[0431] The user speaks, "I've been feeling sad lately." The device converts this into text data, "I've been feeling sad lately," and sends it to the server. The server uses emotion recognition means to extract the emotion "sadness" and stores it in the medical record database.

[0432] Example of generating diagnostic support information

[0433] Based on the subjective information "I've had a backache recently," the visual information "my face is red and swollen," and the emotional information "sadness," the server determines that "the back pain may be related to psychological or mental stress." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[0434] In this way, the system of the present invention performs employee health monitoring and early abnormality detection based on the integrated collection of subjective, visual, and emotional information. The following are examples of prompt sentences that can be input to the generative AI model:

[0435] I haven't been feeling well lately. My shoulders feel heavy and I just don't feel motivated.

[0436] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0437] Step 1:

[0438] The user inputs their physical condition and symptoms into the terminal by voice or text. The input voice data is converted into text data using a voice recognition module. At this stage, the input is voice data or text data, and the output is the converted text data. A specific example of operation is when the user speaks into the terminal, saying, "My headache has been getting worse recently."

[0439] Step 2:

[0440] The device sends the converted text data to the server. The server uses a natural language processing engine to analyze the text data and extract subjective information. At this stage, the input is text data, and the output is the analyzed subjective information. Specifically, the natural language processing engine extracts information such as "headache" and "getting worse."

[0441] Step 3:

[0442] The camera installed on the device is used to capture the employee's facial expression and facial complexion. The captured image data is sent to a server, which then uses an image analysis engine to extract visual information. At this stage, the input is image data, and the output is analyzed visual information. Specific actions such as "the employee's face is pale" and "they look like they're in pain" are extracted by the image analysis engine.

[0443] Step 4:

[0444] The data reported by the user through voice or text is simultaneously analyzed by the emotion recognition means to extract emotional information. At this stage, the input is voice data or text data, and the output is the analyzed emotional information. Specifically, the emotion recognition means analyzes the tone and tempo of the voice to recognize "anxiety" or "sadness."

[0445] Step 5:

[0446] The server stores the subjective, visual, and emotional information obtained above in a medical record database. The input at this stage is subjective, visual, and emotional information, and the output is a database in which this information is stored. Specifically, the information is integrated and organized in the medical record database.

[0447] Step 6:

[0448] The server generates diagnostic support information based on the information stored in the medical record database. The input at this stage is the information stored in the medical record database, and the output is the generated diagnostic support information. Specifically, by referencing past data and medical guidelines, diagnostic support information such as "high possibility of migraine" is generated.

[0449] Step 7:

[0450] The generated diagnostic assistance information is sent to the terminal and notified to the user. The input at this stage is the generated diagnostic assistance information, and the output is the diagnostic assistance information displayed on the terminal. As a specific operation, the diagnostic assistance information is notified to the user, and the user is prompted to take appropriate action.

[0451] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0452] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0453] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0454] [Second embodiment]

[0455] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0456] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0457] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0458] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0459] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0460] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0461] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0462] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0463] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0464] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0465] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0466] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0467] The system of the present invention is composed of a combination of natural language processing means, image analysis means, and medical record database means, and is capable of efficiently and accurately collecting subjective and visual information from patients and reflecting it in medical records.

[0468] Implementation of natural language processing methods

[0469] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device converts this speech into text data using a speech recognition module. The converted text data is sent to a server, which analyzes it using a natural language processing engine to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[0470] Implementation of image analysis procedures

[0471] The user uses the device's camera to capture visual information about the patient. For example, the camera can be used to record the patient's complexion and facial expression. This image data is sent from the device to a server, which then uses an image analysis engine to analyze the data and extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[0472] Implementation of medical record database measures

[0473] The medical record database means organizes and stores the collected subjective and visual information. The server generates diagnostic support information based on the information stored in this database. For example, based on subjective information such as "My headaches have been getting worse recently" and visual information such as "Your face is pale" and "Your expression looks painful," the server refers to past data and medical guidelines and generates diagnostic support information such as "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[0474] Specific examples

[0475] Specific examples are shown below.

[0476] Example of subjective information extraction

[0477] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[0478] Example of information extraction using image analysis

[0479] The image of the patient's face taken with a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[0480] Example of generating diagnostic support information

[0481] Based on subjective information such as "My back has been hurting recently" and visual information such as "My face is red and swollen," the server determines that "there is a high possibility of an inflammatory disease." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[0482] In this way, the system of the present invention efficiently and accurately collects subjective and visual information and reflects it in medical records, automating the process of entering medical records, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[0483] The processing flow will be explained below.

[0484] Extracting subjective information

[0485] Step 1:

[0486] The user inputs the patient's subjective symptoms into the terminal by voice or text. For example, consider the case where the user inputs by voice, "My headache has been getting worse recently."

[0487] Step 2:

[0488] The device converts the voice data into text data using a speech recognition module, and the resulting text data is "My headache has been getting worse recently."

[0489] Step 3:

[0490] The device sends the generated text data to the server, securely using the HTTPS protocol.

[0491] Step 4:

[0492] The server uses a natural language processing engine to analyze the text data and extract key subjective information, such as "headache" and "getting worse."

[0493] Step 5:

[0494] The server organizes the extracted subjective information and stores it in a medical record database, which electronically records the patient's complaints.

[0495] Image analysis

[0496] Step 1:

[0497] The user uses the device's camera to capture visual information of the patient, for example, recording the patient's complexion and facial expressions.

[0498] Step 2:

[0499] The device sends the captured image data to the server. The image is sent in a compressed format and the data is transferred securely using the HTTPS protocol.

[0500] Step 3:

[0501] The server uses an image analysis engine to analyze the image data, specifically analyzing the patient's complexion, facial expression, and movements to extract characteristic information.

[0502] Step 4:

[0503] For example, information such as "his face is pale" and "his expression looks like he's in pain" is extracted.

[0504] Step 5:

[0505] The server stores the extracted visual information in the medical record database, so that the visual information is reflected in the medical record.

[0506] Diagnostic support

[0507] Step 1:

[0508] The server integrates and analyzes the subjective and visual information stored in the medical record database.

[0509] Step 2:

[0510] The server generates diagnostic support information by referencing past data and medical guidelines. For example, it analyzes information such as "Your headaches have gotten worse recently" and "Your face is pale" and "You look like you're in pain."

[0511] Step 3:

[0512] Based on the integrated information, the server generates diagnostic support information such as "high possibility of migraine."

[0513] Step 4:

[0514] The server transmits the generated diagnostic assistance information to the terminal.

[0515] Step 5:

[0516] The device notifies the user of the diagnostic support information, for example, by providing information to a healthcare provider via a pop-up notification or email notification.

[0517] Through this process, the AISOAP system can efficiently and accurately collect patients' subjective and visual information and reflect it in medical records and diagnostic support information, automating the medical record entry process, reducing the burden on medical providers and improving the accuracy of diagnoses.

[0518] Example 1

[0519] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0520] In the modern medical environment, entering medical records is extremely time-consuming, placing a strain on healthcare providers. Furthermore, accurate diagnosis requires reliable collection of patient subjective and visual information, which must be integrated to generate diagnostic support information. However, conventional methods have proven difficult to achieve this efficiently and accurately.

[0521] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0522] In this invention, the server includes a natural language processing unit, an image analysis unit, and a medical record database unit, which enables efficient and accurate collection of subjective and visual information from patients and reflecting it in medical records.

[0523] A "natural language processing means" is a means for extracting subjective patient information from speech or text input.

[0524] The "image analysis means" is a means for extracting visual information of the patient from the camera image.

[0525] A "medical record database means" is a means for storing and managing extracted subjective and visual information.

[0526] "Voice input" refers to voice data uttered by a user through a microphone.

[0527] "Text input" refers to character data that a user inputs using a keyboard or the like.

[0528] A "voice recognition module" is a software or hardware component for converting voice input into text data.

[0529] "Camera footage" refers to video and still image data captured by a camera.

[0530] "Diagnostic support information" is information that helps with diagnosis and is generated based on stored information.

[0531] "User" refers to a healthcare provider or any individual or organization operating the System.

[0532] The system of the present invention is composed of a combination of natural language processing means, image analysis means, and medical record database means, and is capable of efficiently and accurately collecting subjective and visual information from patients and reflecting it in medical records.

[0533] Implementation of natural language processing methods

[0534] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device converts this speech into text data using the Google Cloud Speech-to-Text API. The converted text data is sent to the server. The server then analyzes the text data using a natural language processing engine (e.g., SpaCy or GPT-4) to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[0535] Implementation of image analysis procedures

[0536] The user uses the device's camera to capture visual information about the patient. For example, the camera can be used to record the patient's complexion and facial expression. This image data is sent from the device to a server, which then analyzes the data using an image analysis engine (e.g., OpenCV or TensorFlow) to extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in the medical record database.

[0537] Implementation of medical record database measures

[0538] The medical record database means organizes and stores the collected subjective and visual information. The server generates diagnostic support information based on the information stored in this database. For example, based on subjective information such as "My headaches have been getting worse recently" and visual information such as "Your face is pale" and "Your expression looks painful," the server refers to past data and medical guidelines and generates diagnostic support information such as "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[0539] Specific examples

[0540] Specific examples are shown below.

[0541] Example of subjective information extraction

[0542] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[0543] Example of information extraction using image analysis

[0544] The image of the patient's face taken with a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[0545] Example of generating diagnostic support information

[0546] Based on subjective information such as "My back has been hurting recently" and visual information such as "My face is red and swollen," the server determines that "there is a high possibility of an inflammatory disease." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[0547] Example prompt sentence:

[0548] "Convert the following speech input into text data and extract subjective information. Example: 'My back has been hurting lately.'"

[0549] By utilizing generative AI models, the system described above can collect patient information and provide diagnostic support information more efficiently and accurately than existing medical systems, reducing the burden on medical providers and improving diagnostic accuracy.

[0550] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0551] Step 1:

[0552] The user inputs the patient's symptoms into the terminal in voice or text format. When inputting by voice, the user speaks into a microphone, for example, "My headache has been getting worse recently." The input in this case is voice data or text data.

[0553] Step 2:

[0554] The device uses a speech recognition module to convert speech into text data. Specifically, it uses the Google Cloud Speech-to-Text API. This module receives speech data and converts it into text based on a language model. For example, the device outputs text data such as "My headache has been getting worse recently." This text data is then sent from the device to the server.

[0555] Step 3:

[0556] The server analyzes the received text data using a natural language processing engine (e.g., SpaCy or GPT-4). By inputting and analyzing the text data, subjective information such as "headache" and "getting worse" is extracted. The extracted information is stored in a medical record database. Specifically, the server receives the text data and uses the SpaCy library to extract keywords.

[0557] Step 4:

[0558] The user captures visual information of the patient using the device's camera. For example, they can take a picture of the patient's face and record their facial expression. The input image and video data are then sent from the device to the server.

[0559] Step 5:

[0560] The server analyzes the received image data using an image analysis engine (e.g., OpenCV or TensorFlow). The image data is input and visual information such as "pale face" or "pained expression" is extracted. The extracted information is also stored in the medical record database. Specifically, the server receives the image data and analyzes it using OpenCV.

[0561] Step 6:

[0562] The server generates diagnostic support information based on the subjective and visual information stored in the medical record database. For example, based on subjective information such as "Your headaches have gotten worse recently" and visual information such as "Your face is pale" and "Your expression looks painful," it references past data and medical guidelines to generate diagnostic support information such as "There is a high possibility of a migraine." The generated information is then stored again in the database and simultaneously sent to the device.

[0563] Step 7:

[0564] The device displays the received diagnostic support information on the user interface. The user is notified of new information and can confirm its contents. For example, the diagnostic support information "High possibility of migraine" may be displayed. This allows the user to quickly take necessary measures.

[0565] Through the above steps, patients' subjective and visual information can be collected efficiently and accurately and reflected in medical records, thereby reducing the burden on medical providers and improving diagnostic accuracy.

[0566] (Application example 1)

[0567] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0568] Conventional health management systems have had difficulty efficiently and accurately collecting subjective and visual information from patients to provide diagnostic support. Furthermore, they lacked a means to quickly notify users of the collected information, making it difficult to respond quickly in emergencies. The present invention aims to solve these problems and provide a health management system that is easy for healthcare providers and users to use.

[0569] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0570] In this invention, the server includes means for extracting subjective information of a patient using natural language processing means, means for extracting visual information of the patient using image analysis means, means for storing the extracted subjective information and visual information in a medical record database means, means for generating diagnostic support information based on the stored information, and means for notifying the smartphone terminal of the generated diagnostic support information, thereby enabling a quick analysis of the patient's health condition and a prompt response if necessary.

[0571] "Natural language processing means" is a technology for analyzing a patient's subjective information from voice or text input and extracting specific symptoms and conditions.

[0572] "Image analysis means" is a technology that analyzes the visual information of a patient obtained using a camera and quantitatively evaluates their health condition, such as their complexion and facial expression.

[0573] The "medical record database means" is a database technology that efficiently organizes and stores extracted subjective and visual information.

[0574] "Patient subjective information" is subjective data in audio or text format in which a patient describes their symptoms or physical condition.

[0575] "Patient visual information" refers to visual data such as the patient's appearance and facial expression obtained using a visual device such as a camera.

[0576] The "means for notifying a smartphone terminal" is a technology for sending the generated diagnostic assistance information to a smartphone and notifying the user in real time.

[0577] The "means for generating diagnostic support information" refers to a technology that generates information that assists in specific diagnoses based on stored subjective and visual information, and refers to medical data and guidelines.

[0578] The present invention provides a system for efficiently and accurately collecting subjective and visual information from a patient and providing diagnostic support based on that information. Specific embodiments for carrying out the present invention will be described below.

[0579] Hardware or software used

[0580] Hardware

[0581] Smartphone: Equipped with a microphone, speaker, and camera.

[0582] Server: For data analysis and database management.

[0583] software

[0584] Voice Recognition:

[0585] API: Google Speech Recognition API

[0586] Library: speech_recognition

[0587] Image analysis:

[0588] Deep Learning Model: A pre-trained image classification model using TensorFlow

[0589] Libraries: cv2 (OpenCV), PIL (Python Imaging Library), tensorflow

[0590] Generation of diagnostic support information:

[0591] API: Endpoint for external diagnostic support systems

[0592] Data processing and calculation

[0593] Audio data processing

[0594] The user launches the smartphone application and reports their current symptoms and physical condition by voice input. The smartphone uses a voice recognition module to convert the voice data into text, and then sends the text data to the server. The server uses a natural language processing engine to extract specific symptoms and conditions (e.g., "I have a headache" or "I have a fever") from the text data and stores them in a medical record database.

[0595] Image data processing

[0596] The user records their own facial expression and complexion using the smartphone camera. This image data is sent to a server, which then analyzes it using an image analysis engine. As a result, visual information such as "pale complexion" or "pained facial expression" is extracted and stored in a medical record database.

[0597] Generation and notification of diagnostic support information

[0598] The server generates diagnostic support information based on the subjective and visual information stored in the medical record database. The accuracy of this diagnostic support information is improved by referencing medical data and guidelines. The generated diagnostic support information is then notified to the user via a smartphone application.

[0599] Specific examples

[0600] Usage example

[0601] For example, a user wakes up in the morning feeling unwell and launches the app, then speaks the following:

[0602] Recently, I've been feeling tired and feverish.

[0603] Based on this voice input, the app performs an analysis, then uses the smartphone camera to take and analyze a photograph of the complexion. If the result is a diagnosis of "visual problems," the app will finally notify you as follows:

[0604] Your health may be deteriorating. We recommend that you consult a medical institution.

[0605] In this way, the system can quickly analyze the patient's health condition and help them take prompt action if necessary.

[0606] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0607] Step 1:

[0608] The user starts the smartphone application and reports their current symptoms and physical condition by voice input. At this time, the user speaks into the microphone. The input is the user's voice data, and the output is the voice data recorded by the smartphone device.

[0609] Step 2:

[0610] The device uses a speech recognition module (Google Speech Recognition API) to convert voice data into text data. The input is the user's voice data, and the output is symptom information in text format. The data processing performed here is text conversion based on speech recognition.

[0611] Step 3:

[0612] The terminal sends the converted text data to the server. The input is text data, and the output is data transmission to the server. There is no data processing, but communication is included.

[0613] Step 4:

[0614] The server uses a natural language processing engine to analyze the text data and extract specific symptoms and conditions (e.g., "I have a headache" or "I have a fever"). The input is text data, and the output is the analyzed symptom information. Data processing is an extraction process using natural language processing.

[0615] Step 5:

[0616] The server saves the analyzed symptom information to the medical record database. The input is the analyzed symptom information, and the output is the saving operation to the database. There is no data processing, but the saving operation is included.

[0617] Step 6:

[0618] A user records their facial expression and complexion using a smartphone camera. The device then acquires the image data. The input is the image data captured by the camera, and the output is the recorded image.

[0619] Step 7:

[0620] The terminal sends image data to the server. The input is image data, and the output is data transmission to the server. There is no data processing, but communication is included.

[0621] Step 8:

[0622] The server analyzes the image data using an image analysis engine (a pre-trained model using TensorFlow) and extracts visual information such as "the face looks pale" or "the face looks painful." The input is image data, and the output is the analyzed visual information. Data processing is the extraction process using image analysis.

[0623] Step 9:

[0624] The server stores the analyzed visual information in a medical record database. The input is the analyzed visual information, and the output is the saving operation to the database. There is no data processing, but the saving operation is included.

[0625] Step 10:

[0626] The server generates diagnostic support information based on the subjective and visual information stored in the medical record database. The input is subjective and visual information, and the output is diagnostic support information. Data processing involves information integration and analysis based on diagnostic algorithms.

[0627] Step 11:

[0628] The server notifies the smartphone of the generated diagnostic support information. The input is the diagnostic support information, and the output is a notification to the smartphone terminal. There is no data processing, but communication is involved.

[0629] Step 12:

[0630] The smartphone terminal displays the received diagnostic support information to the user. The input is the diagnostic support information, and the output is the presentation of information to the user. The specific operation is to display a notification message on the screen.

[0631] In this way, by going through each step, the user's health condition can be quickly analyzed and prompt action can be taken if necessary.

[0632] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0633] The system of the present invention is composed of a natural language processing means, an image analysis means, a medical record database means, and an emotion engine that recognizes the user's emotions. This system allows the efficient and accurate collection of subjective, visual, and emotional information of patients and reflects it in medical records.

[0634] Implementation of natural language processing methods

[0635] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device converts this speech into text data using a speech recognition module. The converted text data is sent to a server, which analyzes it using a natural language processing engine to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[0636] Implementation of image analysis procedures

[0637] The user uses the device's camera to capture visual information about the patient. For example, they can record the patient's complexion and facial expression. This image data is sent from the device to a server, which then uses an image analysis engine to analyze the data and extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[0638] Emotion Engine Implementation

[0639] The user inputs the patient's symptoms in voice or text format, and the emotion engine simultaneously analyzes the input data and extracts emotions. Specifically, for voice data, the server analyzes the tone and tempo of the voice to recognize emotions such as "irritability" or "sadness." For text data, the server works with a natural language processing engine to extract emotions from the text. The resulting emotional information is also stored in the medical record database.

[0640] Implementation of medical record database measures

[0641] The medical record database means organizes and stores the collected subjective, visual, and emotional information. The server generates diagnostic support information based on the information stored in this database. For example, if the subjective information "My headaches have been getting worse recently" is recognized, along with visual information such as "Your face is pale" and "Your facial expression looks painful," and a "feeling of irritability" is also recognized, the server will refer to past data and medical guidelines to generate diagnostic support information stating "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[0642] Specific examples

[0643] Specific examples are shown below.

[0644] Example of subjective information extraction

[0645] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[0646] Example of information extraction using image analysis

[0647] The image of the patient's face taken with a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[0648] Example of information extraction using emotion engine

[0649] The user speaks, "I've been feeling sad lately." The device converts this into text data, "I've been feeling sad lately," and sends it to the server. The server uses an emotion engine to extract the emotion "sadness" and stores it in the medical record database.

[0650] Example of generating diagnostic support information

[0651] Based on the subjective information "I've had a backache recently," the visual information "my face is red and swollen," and the emotional information "sadness," the server determines that "the back pain may be related to psychological or mental stress." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[0652] In this way, the system of the present invention efficiently and accurately collects subjective, visual, and emotional information and reflects it in medical records, thereby automating the process of entering medical records, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[0653] The processing flow will be explained below.

[0654] Implementation of natural language processing methods

[0655] Step 1:

[0656] The user inputs the patient's subjective symptoms into the terminal by voice or text. For example, consider the case where the user inputs by voice, "My headache has been getting worse recently."

[0657] Step 2:

[0658] The device converts the voice data into text data using a speech recognition module, and the resulting text data is "My headache has been getting worse recently."

[0659] Step 3:

[0660] The device sends the generated text data to the server, securely using the HTTPS protocol.

[0661] Step 4:

[0662] The server uses a natural language processing engine to analyze the text data and extract key subjective information, such as the phrases "headache" and "getting worse."

[0663] Step 5:

[0664] The server organizes the extracted subjective information and stores it in a medical record database, which electronically records the patient's complaints.

[0665] Implementation of image analysis procedures

[0666] Step 1:

[0667] The user uses the device's camera to capture visual information of the patient, for example, recording the patient's complexion and facial expressions.

[0668] Step 2:

[0669] The device sends the captured image data to the server. The image is sent in a compressed format and the data is transferred securely using the HTTPS protocol.

[0670] Step 3:

[0671] The server uses an image analysis engine to analyze the image data, specifically analyzing the patient's complexion, facial expression, and movements to extract characteristic information.

[0672] Step 4:

[0673] For example, information such as "his face is pale" and "his expression looks like he's in pain" is extracted.

[0674] Step 5:

[0675] The server stores the extracted visual information in the medical record database, so that the visual information is reflected in the medical record.

[0676] Emotion Engine Implementation

[0677] Step 1:

[0678] As the user inputs the patient's symptoms via voice or text, the emotion engine simultaneously analyzes the input data and extracts emotions.

[0679] Step 2:

[0680] In the case of voice data, the terminal transmits the voice data to the server, and in the case of text data, the text data is also transmitted to the server.

[0681] Step 3:

[0682] The server analyzes the tone and tempo of the voice to recognize emotions such as "anxiety" or "sadness." It also extracts emotions from text data.

[0683] Step 4:

[0684] The extracted emotional information is stored in the medical record database, so that the emotional information is also reflected in the medical record.

[0685] Diagnostic support

[0686] Step 1:

[0687] The server integrates and analyzes the subjective, visual, and emotional information stored in the medical record database.

[0688] Step 2:

[0689] The server generates diagnostic support information by referencing past data and medical guidelines. For example, it analyzes information such as "My headache has gotten worse recently," "My face is pale," and "I feel irritable."

[0690] Step 3:

[0691] Based on the integrated information, the server generates diagnostic support information such as "high possibility of migraine."

[0692] Step 4:

[0693] The server transmits the generated diagnostic assistance information to the terminal.

[0694] Step 5:

[0695] The device notifies the user of the diagnostic support information, for example, by providing information to a healthcare provider via a pop-up notification or email notification.

[0696] Through this process, the AISOAP system can efficiently and accurately collect patients' subjective, visual, and emotional information and reflect it in medical records and diagnostic support information, automating the input work of medical records, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[0697] Example 2

[0698] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0699] Modern healthcare requires efficient and accurate collection of patients' subjective, visual, and emotional information and integration of it into medical records. However, traditional methods make it difficult to centrally collect and integrate this information into medical records, which increases the burden on healthcare providers and potentially impacts diagnostic accuracy. To address this challenge, a system that can automatically analyze and integrate data collected from multiple sources is needed.

[0700] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a natural language processing means, an image analysis means, a medical record database means, and an emotion recognition means. This makes it possible to efficiently and accurately collect subjective information, visual information, and emotion information of the patient and reflect it in the medical record.

[0701] "Natural language processing means" is a technology that analyzes input data in voice or text format and extracts meaning and keywords.

[0702] "Image analysis means" is a technology that analyzes video data acquired from a camera or other device and extracts visual information such as the patient's complexion, facial expression, and movements.

[0703] The "medical record database means" is a database system that organizes and stores collected subjective, visual, and emotional information.

[0704] "Emotion recognition means" is a technology that analyzes voice and text data and extracts emotions from them.

[0705] "Subjective information" refers to information such as symptoms and discomfort reported by the patient themselves.

[0706] "Visual information" refers to information extracted from video data acquired using a camera or the like, such as the patient's complexion, facial expression, and movements.

[0707] "Emotion information" is information about the patient's emotions extracted from voice and text data.

[0708] "Diagnostic support information" is information that assists diagnosis and is generated based on collected subjective information, visual information, and emotional information.

[0709] The system of the present invention is composed of a combination of natural language processing means, image analysis means, medical record database means, and emotion recognition means, and is capable of efficiently and accurately collecting subjective, visual, and emotional information from patients and reflecting it in medical records.

[0710] Implementation of natural language processing methods

[0711] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device uses a voice recognition module (e.g., a voice recognition API) to convert the speech into text data. The converted text data is sent to a server, which then analyzes the text data using a natural language processing engine (e.g., SpaCy) to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[0712] Implementation of image analysis procedures

[0713] The user uses the device's camera to capture visual information about the patient. For example, they can record the patient's complexion and facial expression. This image data is sent from the device to a server, which then analyzes the data using an image analysis engine (e.g., OpenCV) to extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[0714] Implementing emotion recognition measures

[0715] The user inputs the patient's symptoms in voice or text format, and at the same time, the emotion recognition means analyzes the input data and extracts emotions. Specifically, in the case of voice data, the server analyzes the tone and tempo of the voice to recognize emotions such as "irritability" or "sadness." In the case of text data, emotions are extracted from the text in cooperation with a natural language processing engine. The emotional information obtained in this way is also stored in the medical record database.

[0716] Implementation of medical record database measures

[0717] The medical record database means organizes and stores the collected subjective, visual, and emotional information. The server generates diagnostic support information based on the information stored in this database. For example, if the subjective information "My headaches have been getting worse recently" is recognized, along with visual information such as "Your face is pale" and "Your facial expression looks painful," and a "feeling of irritability" is also recognized, the server will refer to past data and medical guidelines to generate diagnostic support information stating "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[0718] Specific examples

[0719] Example of subjective information extraction

[0720] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[0721] Example of information extraction using image analysis

[0722] The camera takes a picture of the patient's face and sends the image to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[0723] Example of information extraction using emotion recognition methods

[0724] The user speaks, "I've been feeling sad lately." The device converts this into text data, "I've been feeling sad lately," and sends it to the server. The server uses emotion recognition means to extract the emotion "sadness" and stores it in the medical record database.

[0725] Example of generating diagnostic support information

[0726] Based on the subjective information "I've had a backache recently," the visual information "my face is red and swollen," and the emotional information "sadness," the server determines that "the back pain may be related to psychological or mental stress." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[0727] Through the above process, this system can efficiently and accurately collect patients' subjective, visual, and emotional information and reflect it in medical records, which is an important tool for automating medical record entry work, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[0728] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0729] Step 1: User Input

[0730] The user inputs the patient's subjective symptoms into the terminal in voice or text format. This input becomes the basic data for the system. Specifically, the user says, "My headache has been getting worse recently."

[0731] input:

[0732] Speech data (e.g., "My headaches have been getting worse lately.")

[0733] output:

[0734] Audio data (as is)

[0735] Step 2: Voice Recognition

[0736] The device uses a speech recognition module (e.g., speech recognition API) to convert the speech into text data. At this stage, the speech becomes text data, which is used in the next step.

[0737] input:

[0738] Speech data (e.g., "My headaches have been getting worse lately.")

[0739] output:

[0740] Text data (e.g., "My headaches have been getting worse lately.")

[0741] Operation:

[0742] The audio waveform is analyzed and converted into string data.

[0743] Step 3: Send to the server

[0744] The device then sends the converted text data to the server, where it is analyzed, so fast and accurate transmission is essential.

[0745] input:

[0746] Text data (e.g., "My headaches have been getting worse lately.")

[0747] output:

[0748] Text data (sent as is to the server)

[0749] Operation:

[0750] The text data is uploaded to a server via a network.

[0751] Step 4: Natural Language Processing

[0752] The server uses a natural language processing engine (e.g., SpaCy) to analyze the received text data, and extracts subjective information (e.g., "headache" and "getting worse") as a result of the analysis.

[0753] input:

[0754] Text data (e.g., "My headaches have been getting worse lately.")

[0755] output:

[0756] Subjective information (e.g., "headache" or "it's getting worse")

[0757] Operation:

[0758] Morphological analysis of text data is performed to extract important words and phrases.

[0759] Step 5: Preserving subjective information

[0760] The server stores the extracted subjective information in a medical record database, which is used for subsequent analysis and medical support.

[0761] input:

[0762] Subjective information (e.g., "headache" or "it's getting worse")

[0763] output:

[0764] Saving to a database

[0765] Operation:

[0766] The extracted information is written to a medical record database.

[0767] Step 6: Capture visual information

[0768] The user uses the device's camera to capture visual information of the patient, for example, recording the patient's complexion and facial expressions, and saves this data.

[0769] input:

[0770] Video data (e.g., patient's complexion and facial expression)

[0771] output:

[0772] Video data (as is)

[0773] Operation:

[0774] Use a camera to capture video data.

[0775] Step 7: Sending visual data to the server

[0776] The device sends the recorded image data to a server, where it is analyzed, so it must be sent quickly and accurately.

[0777] input:

[0778] Video data (e.g., patient's complexion and facial expression)

[0779] output:

[0780] Video data (sent as is to the server)

[0781] Operation:

[0782] The video data is uploaded to a server via a network.

[0783] Step 8: Image analysis

[0784] The server analyzes the received video data using an image analysis engine (e.g., OpenCV), and extracts visual information (e.g., "the face is pale" or "the facial expression looks painful").

[0785] input:

[0786] Video data (e.g., patient's complexion and facial expression)

[0787] output:

[0788] Visual information (e.g., "He looks pale" or "He looks like he's in pain")

[0789] Operation:

[0790] Frame analysis of video data is performed to extract important visual information.

[0791] Step 9: Save the visual information

[0792] The server stores the extracted visual information in a medical record database, which is used for subsequent analysis and medical support.

[0793] input:

[0794] Visual information (e.g., "He looks pale" or "He looks like he's in pain")

[0795] output:

[0796] Saving to a database

[0797] Operation:

[0798] The extracted information is written to a medical record database.

[0799] Step 10: Emotion Recognition

[0800] The server analyzes the voice or text data entered by the user using an emotion recognition engine (e.g., emotion recognition API) to extract emotional information.

[0801] input:

[0802] Voice or text data (e.g., "I've been feeling sad lately")

[0803] output:

[0804] Emotional information (e.g., "sadness")

[0805] Operation:

[0806] Analyze the tone and content of voice and text data to identify emotions.

[0807] Step 11: Storing Emotional Information

[0808] The server stores the extracted emotion information in a medical record database, which is used for subsequent analysis and medical support.

[0809] input:

[0810] Emotional information (e.g., "sadness")

[0811] output:

[0812] Saving to a database

[0813] Operation:

[0814] The extracted emotion information is written to a medical record database.

[0815] Step 12: Generating diagnostic support information

[0816] The server generates diagnostic support information based on the stored subjective, visual, and emotional information. For example, if the server recognizes subjective information such as "My headaches have been getting worse recently," along with visual information such as "Your face is pale" and "Your facial expression looks painful," as well as a sense of irritability, the server will refer to past data and medical guidelines and generate diagnostic support information such as "There is a high possibility that you have a migraine."

[0817] input:

[0818] Subjective information, visual information, emotional information

[0819] output:

[0820] Diagnostic support information (e.g., "Probable migraine").

[0821] Operation:

[0822] Integrated analysis is performed based on the stored data to generate diagnostic support information.

[0823] Step 13: Send to device

[0824] The server sends the generated diagnostic support information to the terminal and notifies the user. This information serves as a diagnostic support tool for medical providers.

[0825] input:

[0826] Diagnostic support information (e.g., "Probable migraine").

[0827] output:

[0828] Diagnostic support information (sent directly to the device)

[0829] Operation:

[0830] The diagnostic assistance information is transmitted to a terminal via a network.

[0831] (Application example 2)

[0832] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0833] Conventional health monitoring systems only analyze subjective and visual information, making it difficult to accurately grasp an employee's emotional state. This makes it difficult to detect abnormalities in health status early and take appropriate measures. Furthermore, because it takes time for employees to report their health status in detail, there is a risk of incomplete information or false reports. The present invention aims to solve these problems and provide a system that efficiently and accurately monitors the health status of employees in factories.

[0834] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes natural language processing means, image analysis means, medical record database means, emotion recognition means, means for extracting patient emotion information using the natural language processing means and emotion recognition means, and means for generating diagnostic support information based on the stored information. This makes it possible to comprehensively collect and analyze employee subjective information, visual information, and emotion information. Ultimately, it is possible to more accurately grasp the employee's health status, detect abnormalities early, and take appropriate action.

[0835] "Natural language processing means" is a technology that analyzes voice and text data and extracts meaning from its content.

[0836] "Image analysis means" is a technique for analyzing image data and extracting features contained therein.

[0837] "Medical record database means" refers to a database system that collects and manages health information of patients and employees.

[0838] "Means for extracting subjective patient information" refers to a technology that uses natural language processing to extract symptoms and sensations reported by patients from text data.

[0839] The "means for extracting visual information of a patient" is a technology that uses image analysis means to extract information related to the patient's health condition from their appearance, such as their complexion and facial expression.

[0840] "Emotion recognition means" is a technology that analyzes emotions from voice or text and extracts that emotional information.

[0841] The "means for generating diagnostic support information" is a system that automatically proposes appropriate diagnoses and countermeasures based on collected health information.

[0842] The system of the present invention combines natural language processing, image analysis, a medical record database, and emotion recognition, and is capable of efficiently and accurately collecting employee subjective health information, visual health information, and emotion information to monitor health status and detect abnormalities.

[0843] Implementation of natural language processing methods

[0844] The user inputs their physical condition and symptoms in voice or text format. If the user says, "My headache has been getting worse recently," the device converts this voice data into text using a voice recognition module. The converted text data is sent to a server, which analyzes it using a natural language processing engine to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[0845] Implementation of image analysis procedures

[0846] The device uses a built-in camera to capture visual information about the employee, such as their complexion and facial expression. This image data is sent from the device to a server, which then uses an image analysis engine to analyze the data and extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[0847] Implementing emotion recognition measures

[0848] When a user reports their symptoms in voice or text format, the emotion recognition means simultaneously analyzes the input data and extracts emotions. Specifically, for voice data, the server analyzes the tone and tempo of the voice to recognize emotions such as "irritability" or "sadness." For text data, the server works with a natural language processing engine to extract emotions from the text. The emotional information obtained in this way is also stored in the medical record database.

[0849] Implementation of medical record database measures

[0850] The medical record database means organizes and stores the collected subjective, visual, and emotional information. The server generates diagnostic support information based on the information stored in this database. For example, if the subjective information "My headaches have been getting worse recently" is recognized, along with visual information such as "Your face is pale" and "Your facial expression looks painful," and a "feeling of irritability" is also recognized, the server will refer to past data and medical guidelines to generate diagnostic support information stating "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[0851] Specific examples

[0852] Specific examples are shown below.

[0853] Example of subjective information extraction

[0854] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[0855] Example of information extraction using image analysis

[0856] The image of the employee's face captured by a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling" and stores it in a medical record database.

[0857] Example of information extraction using emotion engine

[0858] The user speaks, "I've been feeling sad lately." The device converts this into text data, "I've been feeling sad lately," and sends it to the server. The server uses emotion recognition means to extract the emotion "sadness" and stores it in the medical record database.

[0859] Example of generating diagnostic support information

[0860] Based on the subjective information "I've had a backache recently," the visual information "my face is red and swollen," and the emotional information "sadness," the server determines that "the back pain may be related to psychological or mental stress." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[0861] In this way, the system of the present invention performs employee health monitoring and early abnormality detection based on the integrated collection of subjective, visual, and emotional information. The following are examples of prompt sentences that can be input to the generative AI model:

[0862] I haven't been feeling well lately. My shoulders feel heavy and I just don't feel motivated.

[0863] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0864] Step 1:

[0865] The user inputs their physical condition and symptoms into the terminal by voice or text. The input voice data is converted into text data using a voice recognition module. At this stage, the input is voice data or text data, and the output is the converted text data. A specific example of operation is when the user speaks into the terminal, saying, "My headache has been getting worse recently."

[0866] Step 2:

[0867] The device sends the converted text data to the server. The server uses a natural language processing engine to analyze the text data and extract subjective information. At this stage, the input is text data, and the output is the analyzed subjective information. Specifically, the natural language processing engine extracts information such as "headache" and "getting worse."

[0868] Step 3:

[0869] The camera installed on the device is used to capture the employee's facial expression and facial complexion. The captured image data is sent to a server, which then uses an image analysis engine to extract visual information. At this stage, the input is image data, and the output is analyzed visual information. Specific actions such as "the employee's face is pale" and "they look like they're in pain" are extracted by the image analysis engine.

[0870] Step 4:

[0871] The data reported by the user through voice or text is simultaneously analyzed by the emotion recognition means to extract emotional information. At this stage, the input is voice data or text data, and the output is the analyzed emotional information. Specifically, the emotion recognition means analyzes the tone and tempo of the voice to recognize "anxiety" or "sadness."

[0872] Step 5:

[0873] The server stores the subjective, visual, and emotional information obtained above in a medical record database. The input at this stage is subjective, visual, and emotional information, and the output is a database in which this information is stored. Specifically, the information is integrated and organized in the medical record database.

[0874] Step 6:

[0875] The server generates diagnostic support information based on the information stored in the medical record database. The input at this stage is the information stored in the medical record database, and the output is the generated diagnostic support information. Specifically, by referencing past data and medical guidelines, diagnostic support information such as "high possibility of migraine" is generated.

[0876] Step 7:

[0877] The generated diagnostic assistance information is sent to the terminal and notified to the user. The input at this stage is the generated diagnostic assistance information, and the output is the diagnostic assistance information displayed on the terminal. As a specific operation, the diagnostic assistance information is notified to the user, and the user is prompted to take appropriate action.

[0878] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0879] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0880] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0881] [Third embodiment]

[0882] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0883] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0884] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0885] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0886] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0887] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0888] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0889] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0890] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0891] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0892] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0893] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0894] The system of the present invention is composed of a combination of natural language processing means, image analysis means, and medical record database means, and is capable of efficiently and accurately collecting subjective and visual information from patients and reflecting it in medical records.

[0895] Implementation of natural language processing methods

[0896] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device converts this speech into text data using a speech recognition module. The converted text data is sent to a server, which analyzes it using a natural language processing engine to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[0897] Implementation of image analysis procedures

[0898] The user uses the device's camera to capture visual information about the patient. For example, the camera can be used to record the patient's complexion and facial expression. This image data is sent from the device to a server, which then uses an image analysis engine to analyze the data and extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[0899] Implementation of medical record database measures

[0900] The medical record database means organizes and stores the collected subjective and visual information. The server generates diagnostic support information based on the information stored in this database. For example, based on subjective information such as "My headaches have been getting worse recently" and visual information such as "Your face is pale" and "Your expression looks painful," the server refers to past data and medical guidelines and generates diagnostic support information such as "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[0901] Specific examples

[0902] Specific examples are shown below.

[0903] Example of subjective information extraction

[0904] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[0905] Example of information extraction using image analysis

[0906] The image of the patient's face taken with a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[0907] Example of generating diagnostic support information

[0908] Based on subjective information such as "My back has been hurting recently" and visual information such as "My face is red and swollen," the server determines that "there is a high possibility of an inflammatory disease." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[0909] In this way, the system of the present invention efficiently and accurately collects subjective and visual information and reflects it in medical records, automating the process of entering medical records, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[0910] The processing flow will be explained below.

[0911] Extracting subjective information

[0912] Step 1:

[0913] The user inputs the patient's subjective symptoms into the terminal by voice or text. For example, consider the case where the user inputs by voice, "My headache has been getting worse recently."

[0914] Step 2:

[0915] The device converts the voice data into text data using a speech recognition module, and the resulting text data is "My headache has been getting worse recently."

[0916] Step 3:

[0917] The device sends the generated text data to the server, securely using the HTTPS protocol.

[0918] Step 4:

[0919] The server uses a natural language processing engine to analyze the text data and extract key subjective information, such as "headache" and "getting worse."

[0920] Step 5:

[0921] The server organizes the extracted subjective information and stores it in a medical record database, which electronically records the patient's complaints.

[0922] Image analysis

[0923] Step 1:

[0924] The user uses the device's camera to capture visual information of the patient, for example, recording the patient's complexion and facial expressions.

[0925] Step 2:

[0926] The device sends the captured image data to the server. The image is sent in a compressed format and the data is transferred securely using the HTTPS protocol.

[0927] Step 3:

[0928] The server uses an image analysis engine to analyze the image data, specifically analyzing the patient's complexion, facial expression, and movements to extract characteristic information.

[0929] Step 4:

[0930] For example, information such as "his face is pale" and "his expression looks like he's in pain" is extracted.

[0931] Step 5:

[0932] The server stores the extracted visual information in the medical record database, so that the visual information is reflected in the medical record.

[0933] Diagnostic support

[0934] Step 1:

[0935] The server integrates and analyzes the subjective and visual information stored in the medical record database.

[0936] Step 2:

[0937] The server generates diagnostic support information by referencing past data and medical guidelines. For example, it analyzes information such as "Your headaches have gotten worse recently" and "Your face is pale" and "You look like you're in pain."

[0938] Step 3:

[0939] Based on the integrated information, the server generates diagnostic support information such as "high possibility of migraine."

[0940] Step 4:

[0941] The server transmits the generated diagnostic assistance information to the terminal.

[0942] Step 5:

[0943] The device notifies the user of the diagnostic support information, for example, by providing information to a healthcare provider via a pop-up notification or email notification.

[0944] Through this process, the AISOAP system can efficiently and accurately collect patients' subjective and visual information and reflect it in medical records and diagnostic support information, automating the medical record entry process, reducing the burden on medical providers and improving the accuracy of diagnoses.

[0945] Example 1

[0946] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0947] In the modern medical environment, entering medical records is extremely time-consuming, placing a strain on healthcare providers. Furthermore, accurate diagnosis requires reliable collection of patient subjective and visual information, which must be integrated to generate diagnostic support information. However, conventional methods have proven difficult to achieve this efficiently and accurately.

[0948] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0949] In this invention, the server includes a natural language processing unit, an image analysis unit, and a medical record database unit, which enables efficient and accurate collection of subjective and visual information from patients and reflecting it in medical records.

[0950] A "natural language processing means" is a means for extracting subjective patient information from speech or text input.

[0951] The "image analysis means" is a means for extracting visual information of the patient from the camera image.

[0952] A "medical record database means" is a means for storing and managing extracted subjective and visual information.

[0953] "Voice input" refers to voice data uttered by a user through a microphone.

[0954] "Text input" refers to character data that a user inputs using a keyboard or the like.

[0955] A "voice recognition module" is a software or hardware component for converting voice input into text data.

[0956] "Camera footage" refers to video and still image data captured by a camera.

[0957] "Diagnostic support information" is information that helps with diagnosis and is generated based on stored information.

[0958] "User" refers to a healthcare provider or any individual or organization operating the System.

[0959] The system of the present invention is composed of a combination of natural language processing means, image analysis means, and medical record database means, and is capable of efficiently and accurately collecting subjective and visual information from patients and reflecting it in medical records.

[0960] Implementation of natural language processing methods

[0961] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device converts this speech into text data using the Google Cloud Speech-to-Text API. The converted text data is sent to the server. The server then analyzes the text data using a natural language processing engine (e.g., SpaCy or GPT-4) to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[0962] Implementation of image analysis procedures

[0963] The user uses the device's camera to capture visual information about the patient. For example, the camera can be used to record the patient's complexion and facial expression. This image data is sent from the device to a server, which then analyzes the data using an image analysis engine (e.g., OpenCV or TensorFlow) to extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in the medical record database.

[0964] Implementation of medical record database measures

[0965] The medical record database means organizes and stores the collected subjective and visual information. The server generates diagnostic support information based on the information stored in this database. For example, based on subjective information such as "My headaches have been getting worse recently" and visual information such as "Your face is pale" and "Your expression looks painful," the server refers to past data and medical guidelines and generates diagnostic support information such as "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[0966] Specific examples

[0967] Specific examples are shown below.

[0968] Example of subjective information extraction

[0969] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[0970] Example of information extraction using image analysis

[0971] The image of the patient's face taken with a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[0972] Example of generating diagnostic support information

[0973] Based on subjective information such as "My back has been hurting recently" and visual information such as "My face is red and swollen," the server determines that "there is a high possibility of an inflammatory disease." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[0974] Example prompt sentence:

[0975] "Convert the following speech input into text data and extract subjective information. Example: 'My back has been hurting lately.'"

[0976] By utilizing generative AI models, the system described above can collect patient information and provide diagnostic support information more efficiently and accurately than existing medical systems, reducing the burden on medical providers and improving diagnostic accuracy.

[0977] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0978] Step 1:

[0979] The user inputs the patient's symptoms into the terminal in voice or text format. When inputting by voice, the user speaks into a microphone, for example, "My headache has been getting worse recently." The input in this case is voice data or text data.

[0980] Step 2:

[0981] The device uses a speech recognition module to convert speech into text data. Specifically, it uses the Google Cloud Speech-to-Text API. This module receives speech data and converts it into text based on a language model. For example, the device outputs text data such as "My headache has been getting worse recently." This text data is then sent from the device to the server.

[0982] Step 3:

[0983] The server analyzes the received text data using a natural language processing engine (e.g., SpaCy or GPT-4). By inputting and analyzing the text data, subjective information such as "headache" and "getting worse" is extracted. The extracted information is stored in a medical record database. Specifically, the server receives the text data and uses the SpaCy library to extract keywords.

[0984] Step 4:

[0985] The user captures visual information of the patient using the device's camera. For example, they can take a picture of the patient's face and record their facial expression. The input image and video data are then sent from the device to the server.

[0986] Step 5:

[0987] The server analyzes the received image data using an image analysis engine (e.g., OpenCV or TensorFlow). The image data is input and visual information such as "pale face" or "pained expression" is extracted. The extracted information is also stored in the medical record database. Specifically, the server receives the image data and analyzes it using OpenCV.

[0988] Step 6:

[0989] The server generates diagnostic support information based on the subjective and visual information stored in the medical record database. For example, based on subjective information such as "Your headaches have gotten worse recently" and visual information such as "Your face is pale" and "Your expression looks painful," it references past data and medical guidelines to generate diagnostic support information such as "There is a high possibility of a migraine." The generated information is then stored again in the database and simultaneously sent to the device.

[0990] Step 7:

[0991] The device displays the received diagnostic support information on the user interface. The user is notified of new information and can confirm its contents. For example, the diagnostic support information "High possibility of migraine" may be displayed. This allows the user to quickly take necessary measures.

[0992] Through the above steps, patients' subjective and visual information can be collected efficiently and accurately and reflected in medical records, thereby reducing the burden on medical providers and improving diagnostic accuracy.

[0993] (Application example 1)

[0994] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0995] Conventional health management systems have had difficulty efficiently and accurately collecting subjective and visual information from patients to provide diagnostic support. Furthermore, they lacked a means to quickly notify users of the collected information, making it difficult to respond quickly in emergencies. The present invention aims to solve these problems and provide a health management system that is easy for healthcare providers and users to use.

[0996] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0997] In this invention, the server includes means for extracting subjective information of a patient using natural language processing means, means for extracting visual information of the patient using image analysis means, means for storing the extracted subjective information and visual information in a medical record database means, means for generating diagnostic support information based on the stored information, and means for notifying the smartphone terminal of the generated diagnostic support information, thereby enabling a quick analysis of the patient's health condition and a prompt response if necessary.

[0998] "Natural language processing means" is a technology for analyzing a patient's subjective information from voice or text input and extracting specific symptoms and conditions.

[0999] "Image analysis means" is a technology that analyzes the visual information of a patient obtained using a camera and quantitatively evaluates their health condition, such as their complexion and facial expression.

[1000] The "medical record database means" is a database technology that efficiently organizes and stores extracted subjective and visual information.

[1001] "Patient subjective information" is subjective data in audio or text format in which a patient describes their symptoms or physical condition.

[1002] "Patient visual information" refers to visual data such as the patient's appearance and facial expression obtained using a visual device such as a camera.

[1003] The "means for notifying a smartphone terminal" is a technology for sending the generated diagnostic assistance information to a smartphone and notifying the user in real time.

[1004] The "means for generating diagnostic support information" refers to a technology that generates information that assists in specific diagnoses based on stored subjective and visual information, and refers to medical data and guidelines.

[1005] The present invention provides a system for efficiently and accurately collecting subjective and visual information from a patient and providing diagnostic support based on that information. Specific embodiments for carrying out the present invention will be described below.

[1006] Hardware or software used

[1007] Hardware

[1008] Smartphone: Equipped with a microphone, speaker, and camera.

[1009] Server: For data analysis and database management.

[1010] software

[1011] Voice Recognition:

[1012] API: Google Speech Recognition API

[1013] Library: speech_recognition

[1014] Image analysis:

[1015] Deep Learning Model: A pre-trained image classification model using TensorFlow

[1016] Libraries: cv2 (OpenCV), PIL (Python Imaging Library), tensorflow

[1017] Generation of diagnostic support information:

[1018] API: Endpoint for external diagnostic support systems

[1019] Data processing and calculation

[1020] Audio data processing

[1021] The user launches the smartphone application and reports their current symptoms and physical condition by voice input. The smartphone uses a voice recognition module to convert the voice data into text, and then sends the text data to the server. The server uses a natural language processing engine to extract specific symptoms and conditions (e.g., "I have a headache" or "I have a fever") from the text data and stores them in a medical record database.

[1022] Image data processing

[1023] The user records their own facial expression and complexion using the smartphone camera. This image data is sent to a server, which then analyzes it using an image analysis engine. As a result, visual information such as "pale complexion" or "pained facial expression" is extracted and stored in a medical record database.

[1024] Generation and notification of diagnostic support information

[1025] The server generates diagnostic support information based on the subjective and visual information stored in the medical record database. The accuracy of this diagnostic support information is improved by referencing medical data and guidelines. The generated diagnostic support information is then notified to the user via a smartphone application.

[1026] Specific examples

[1027] Usage example

[1028] For example, a user wakes up in the morning feeling unwell and launches the app, then speaks the following:

[1029] Recently, I've been feeling tired and feverish.

[1030] Based on this voice input, the app performs an analysis, then uses the smartphone camera to take and analyze a photograph of the complexion. If the result is a diagnosis of "visual problems," the app will finally notify you as follows:

[1031] Your health may be deteriorating. We recommend that you consult a medical institution.

[1032] In this way, the system can quickly analyze the patient's health condition and help them take prompt action if necessary.

[1033] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1034] Step 1:

[1035] The user starts the smartphone application and reports their current symptoms and physical condition by voice input. At this time, the user speaks into the microphone. The input is the user's voice data, and the output is the voice data recorded by the smartphone device.

[1036] Step 2:

[1037] The device uses a speech recognition module (Google Speech Recognition API) to convert voice data into text data. The input is the user's voice data, and the output is symptom information in text format. The data processing performed here is text conversion based on speech recognition.

[1038] Step 3:

[1039] The terminal sends the converted text data to the server. The input is text data, and the output is data transmission to the server. There is no data processing, but communication is included.

[1040] Step 4:

[1041] The server uses a natural language processing engine to analyze the text data and extract specific symptoms and conditions (e.g., "I have a headache" or "I have a fever"). The input is text data, and the output is the analyzed symptom information. Data processing is an extraction process using natural language processing.

[1042] Step 5:

[1043] The server saves the analyzed symptom information to the medical record database. The input is the analyzed symptom information, and the output is the saving operation to the database. There is no data processing, but the saving operation is included.

[1044] Step 6:

[1045] A user records their facial expression and complexion using a smartphone camera. The device then acquires the image data. The input is the image data captured by the camera, and the output is the recorded image.

[1046] Step 7:

[1047] The terminal sends image data to the server. The input is image data, and the output is data transmission to the server. There is no data processing, but communication is included.

[1048] Step 8:

[1049] The server analyzes the image data using an image analysis engine (a pre-trained model using TensorFlow) and extracts visual information such as "the face looks pale" or "the face looks painful." The input is image data, and the output is the analyzed visual information. Data processing is the extraction process using image analysis.

[1050] Step 9:

[1051] The server stores the analyzed visual information in a medical record database. The input is the analyzed visual information, and the output is the saving operation to the database. There is no data processing, but the saving operation is included.

[1052] Step 10:

[1053] The server generates diagnostic support information based on the subjective and visual information stored in the medical record database. The input is subjective and visual information, and the output is diagnostic support information. Data processing involves information integration and analysis based on diagnostic algorithms.

[1054] Step 11:

[1055] The server notifies the smartphone of the generated diagnostic support information. The input is the diagnostic support information, and the output is a notification to the smartphone terminal. There is no data processing, but communication is involved.

[1056] Step 12:

[1057] The smartphone terminal displays the received diagnostic support information to the user. The input is the diagnostic support information, and the output is the presentation of information to the user. The specific operation is to display a notification message on the screen.

[1058] In this way, by going through each step, the user's health condition can be quickly analyzed and prompt action can be taken if necessary.

[1059] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1060] The system of the present invention is composed of a natural language processing means, an image analysis means, a medical record database means, and an emotion engine that recognizes the user's emotions. This system allows the efficient and accurate collection of subjective, visual, and emotional information of patients and reflects it in medical records.

[1061] Implementation of natural language processing methods

[1062] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device converts this speech into text data using a speech recognition module. The converted text data is sent to a server, which analyzes it using a natural language processing engine to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[1063] Implementation of image analysis procedures

[1064] The user uses the device's camera to capture visual information about the patient. For example, they can record the patient's complexion and facial expression. This image data is sent from the device to a server, which then uses an image analysis engine to analyze the data and extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[1065] Emotion Engine Implementation

[1066] The user inputs the patient's symptoms in voice or text format, and the emotion engine simultaneously analyzes the input data and extracts emotions. Specifically, for voice data, the server analyzes the tone and tempo of the voice to recognize emotions such as "irritability" or "sadness." For text data, the server works with a natural language processing engine to extract emotions from the text. The resulting emotional information is also stored in the medical record database.

[1067] Implementation of medical record database measures

[1068] The medical record database means organizes and stores the collected subjective, visual, and emotional information. The server generates diagnostic support information based on the information stored in this database. For example, if the subjective information "My headaches have been getting worse recently" is recognized, along with visual information such as "Your face is pale" and "Your facial expression looks painful," and a "feeling of irritability" is also recognized, the server will refer to past data and medical guidelines to generate diagnostic support information stating "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[1069] Specific examples

[1070] Specific examples are shown below.

[1071] Example of subjective information extraction

[1072] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[1073] Example of information extraction using image analysis

[1074] The image of the patient's face taken with a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[1075] Example of information extraction using emotion engine

[1076] The user speaks, "I've been feeling sad lately." The device converts this into text data, "I've been feeling sad lately," and sends it to the server. The server uses an emotion engine to extract the emotion "sadness" and stores it in the medical record database.

[1077] Example of generating diagnostic support information

[1078] Based on the subjective information "I've had a backache recently," the visual information "my face is red and swollen," and the emotional information "sadness," the server determines that "the back pain may be related to psychological or mental stress." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[1079] In this way, the system of the present invention efficiently and accurately collects subjective, visual, and emotional information and reflects it in medical records, thereby automating the process of entering medical records, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[1080] The processing flow will be explained below.

[1081] Implementation of natural language processing methods

[1082] Step 1:

[1083] The user inputs the patient's subjective symptoms into the terminal by voice or text. For example, consider the case where the user inputs by voice, "My headache has been getting worse recently."

[1084] Step 2:

[1085] The device converts the voice data into text data using a speech recognition module, and the resulting text data is "My headache has been getting worse recently."

[1086] Step 3:

[1087] The device sends the generated text data to the server, securely using the HTTPS protocol.

[1088] Step 4:

[1089] The server uses a natural language processing engine to analyze the text data and extract key subjective information, such as the phrases "headache" and "getting worse."

[1090] Step 5:

[1091] The server organizes the extracted subjective information and stores it in a medical record database, which electronically records the patient's complaints.

[1092] Implementation of image analysis procedures

[1093] Step 1:

[1094] The user uses the device's camera to capture visual information of the patient, for example, recording the patient's complexion and facial expressions.

[1095] Step 2:

[1096] The device sends the captured image data to the server. The image is sent in a compressed format and the data is transferred securely using the HTTPS protocol.

[1097] Step 3:

[1098] The server uses an image analysis engine to analyze the image data, specifically analyzing the patient's complexion, facial expression, and movements to extract characteristic information.

[1099] Step 4:

[1100] For example, information such as "his face is pale" and "his expression looks like he's in pain" is extracted.

[1101] Step 5:

[1102] The server stores the extracted visual information in the medical record database, so that the visual information is reflected in the medical record.

[1103] Emotion Engine Implementation

[1104] Step 1:

[1105] As the user inputs the patient's symptoms via voice or text, the emotion engine simultaneously analyzes the input data and extracts emotions.

[1106] Step 2:

[1107] In the case of voice data, the terminal transmits the voice data to the server, and in the case of text data, the text data is also transmitted to the server.

[1108] Step 3:

[1109] The server analyzes the tone and tempo of the voice to recognize emotions such as "anxiety" or "sadness." It also extracts emotions from text data.

[1110] Step 4:

[1111] The extracted emotional information is stored in the medical record database, so that the emotional information is also reflected in the medical record.

[1112] Diagnostic support

[1113] Step 1:

[1114] The server integrates and analyzes the subjective, visual, and emotional information stored in the medical record database.

[1115] Step 2:

[1116] The server generates diagnostic support information by referencing past data and medical guidelines. For example, it analyzes information such as "My headache has gotten worse recently," "My face is pale," and "I feel irritable."

[1117] Step 3:

[1118] Based on the integrated information, the server generates diagnostic support information such as "high possibility of migraine."

[1119] Step 4:

[1120] The server transmits the generated diagnostic assistance information to the terminal.

[1121] Step 5:

[1122] The device notifies the user of the diagnostic support information, for example, by providing information to a healthcare provider via a pop-up notification or email notification.

[1123] Through this process, the AISOAP system can efficiently and accurately collect patients' subjective, visual, and emotional information and reflect it in medical records and diagnostic support information, automating the input work of medical records, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[1124] Example 2

[1125] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1126] Modern healthcare requires efficient and accurate collection of patients' subjective, visual, and emotional information and integration of it into medical records. However, traditional methods make it difficult to centrally collect and integrate this information into medical records, which increases the burden on healthcare providers and potentially impacts diagnostic accuracy. To address this challenge, a system that can automatically analyze and integrate data collected from multiple sources is needed.

[1127] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a natural language processing means, an image analysis means, a medical record database means, and an emotion recognition means. This makes it possible to efficiently and accurately collect subjective information, visual information, and emotion information of the patient and reflect it in the medical record.

[1128] "Natural language processing means" is a technology that analyzes input data in voice or text format and extracts meaning and keywords.

[1129] "Image analysis means" is a technology that analyzes video data acquired from a camera or other device and extracts visual information such as the patient's complexion, facial expression, and movements.

[1130] The "medical record database means" is a database system that organizes and stores collected subjective, visual, and emotional information.

[1131] "Emotion recognition means" is a technology that analyzes voice and text data and extracts emotions from them.

[1132] "Subjective information" refers to information such as symptoms and discomfort reported by the patient themselves.

[1133] "Visual information" refers to information extracted from video data acquired using a camera or the like, such as the patient's complexion, facial expression, and movements.

[1134] "Emotion information" is information about the patient's emotions extracted from voice and text data.

[1135] "Diagnostic support information" is information that assists diagnosis and is generated based on collected subjective information, visual information, and emotional information.

[1136] The system of the present invention is composed of a combination of natural language processing means, image analysis means, medical record database means, and emotion recognition means, and is capable of efficiently and accurately collecting subjective, visual, and emotional information from patients and reflecting it in medical records.

[1137] Implementation of natural language processing methods

[1138] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device uses a voice recognition module (e.g., a voice recognition API) to convert the speech into text data. The converted text data is sent to a server, which then analyzes the text data using a natural language processing engine (e.g., SpaCy) to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[1139] Implementation of image analysis procedures

[1140] The user uses the device's camera to capture visual information about the patient. For example, they can record the patient's complexion and facial expression. This image data is sent from the device to a server, which then analyzes the data using an image analysis engine (e.g., OpenCV) to extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[1141] Implementing emotion recognition measures

[1142] The user inputs the patient's symptoms in voice or text format, and at the same time, the emotion recognition means analyzes the input data and extracts emotions. Specifically, in the case of voice data, the server analyzes the tone and tempo of the voice to recognize emotions such as "irritability" or "sadness." In the case of text data, emotions are extracted from the text in cooperation with a natural language processing engine. The emotional information obtained in this way is also stored in the medical record database.

[1143] Implementation of medical record database measures

[1144] The medical record database means organizes and stores the collected subjective, visual, and emotional information. The server generates diagnostic support information based on the information stored in this database. For example, if the subjective information "My headaches have been getting worse recently" is recognized, along with visual information such as "Your face is pale" and "Your facial expression looks painful," and a "feeling of irritability" is also recognized, the server will refer to past data and medical guidelines to generate diagnostic support information stating "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[1145] Specific examples

[1146] Example of subjective information extraction

[1147] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[1148] Example of information extraction using image analysis

[1149] The camera takes a picture of the patient's face and sends the image to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[1150] Example of information extraction using emotion recognition methods

[1151] The user speaks, "I've been feeling sad lately." The device converts this into text data, "I've been feeling sad lately," and sends it to the server. The server uses emotion recognition means to extract the emotion "sadness" and stores it in the medical record database.

[1152] Example of generating diagnostic support information

[1153] Based on the subjective information "I've had a backache recently," the visual information "my face is red and swollen," and the emotional information "sadness," the server determines that "the back pain may be related to psychological or mental stress." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[1154] Through the above process, this system can efficiently and accurately collect patients' subjective, visual, and emotional information and reflect it in medical records, which is an important tool for automating medical record entry work, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[1155] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1156] Step 1: User Input

[1157] The user inputs the patient's subjective symptoms into the terminal in voice or text format. This input becomes the basic data for the system. Specifically, the user says, "My headache has been getting worse recently."

[1158] input:

[1159] Speech data (e.g., "My headaches have been getting worse lately.")

[1160] output:

[1161] Audio data (as is)

[1162] Step 2: Voice Recognition

[1163] The device uses a speech recognition module (e.g., speech recognition API) to convert the speech into text data. At this stage, the speech becomes text data, which is used in the next step.

[1164] input:

[1165] Speech data (e.g., "My headaches have been getting worse lately.")

[1166] output:

[1167] Text data (e.g., "My headaches have been getting worse lately.")

[1168] Operation:

[1169] The audio waveform is analyzed and converted into string data.

[1170] Step 3: Send to the server

[1171] The device then sends the converted text data to the server, where it is analyzed, so fast and accurate transmission is essential.

[1172] input:

[1173] Text data (e.g., "My headaches have been getting worse lately.")

[1174] output:

[1175] Text data (sent as is to the server)

[1176] Operation:

[1177] The text data is uploaded to a server via a network.

[1178] Step 4: Natural Language Processing

[1179] The server uses a natural language processing engine (e.g., SpaCy) to analyze the received text data, and extracts subjective information (e.g., "headache" and "getting worse") as a result of the analysis.

[1180] input:

[1181] Text data (e.g., "My headaches have been getting worse lately.")

[1182] output:

[1183] Subjective information (e.g., "headache" or "it's getting worse")

[1184] Operation:

[1185] Morphological analysis of text data is performed to extract important words and phrases.

[1186] Step 5: Preserving subjective information

[1187] The server stores the extracted subjective information in a medical record database, which is used for subsequent analysis and medical support.

[1188] input:

[1189] Subjective information (e.g., "headache" or "it's getting worse")

[1190] output:

[1191] Saving to a database

[1192] Operation:

[1193] The extracted information is written to a medical record database.

[1194] Step 6: Capture visual information

[1195] The user uses the device's camera to capture visual information of the patient, for example, recording the patient's complexion and facial expressions, and saves this data.

[1196] input:

[1197] Video data (e.g., patient's complexion and facial expression)

[1198] output:

[1199] Video data (as is)

[1200] Operation:

[1201] Use a camera to capture video data.

[1202] Step 7: Sending visual data to the server

[1203] The device sends the recorded image data to a server, where it is analyzed, so it must be sent quickly and accurately.

[1204] input:

[1205] Video data (e.g., patient's complexion and facial expression)

[1206] output:

[1207] Video data (sent as is to the server)

[1208] Operation:

[1209] The video data is uploaded to a server via a network.

[1210] Step 8: Image analysis

[1211] The server analyzes the received video data using an image analysis engine (e.g., OpenCV), and extracts visual information (e.g., "the face is pale" or "the facial expression looks painful").

[1212] input:

[1213] Video data (e.g., patient's complexion and facial expression)

[1214] output:

[1215] Visual information (e.g., "He looks pale" or "He looks like he's in pain")

[1216] Operation:

[1217] Frame analysis of video data is performed to extract important visual information.

[1218] Step 9: Save the visual information

[1219] The server stores the extracted visual information in a medical record database, which is used for subsequent analysis and medical support.

[1220] input:

[1221] Visual information (e.g., "He looks pale" or "He looks like he's in pain")

[1222] output:

[1223] Saving to a database

[1224] Operation:

[1225] The extracted information is written to a medical record database.

[1226] Step 10: Emotion Recognition

[1227] The server analyzes the voice or text data entered by the user using an emotion recognition engine (e.g., emotion recognition API) to extract emotional information.

[1228] input:

[1229] Voice or text data (e.g., "I've been feeling sad lately")

[1230] output:

[1231] Emotional information (e.g., "sadness")

[1232] Operation:

[1233] Analyze the tone and content of voice and text data to identify emotions.

[1234] Step 11: Storing Emotional Information

[1235] The server stores the extracted emotion information in a medical record database, which is used for subsequent analysis and medical support.

[1236] input:

[1237] Emotional information (e.g., "sadness")

[1238] output:

[1239] Saving to a database

[1240] Operation:

[1241] The extracted emotion information is written to a medical record database.

[1242] Step 12: Generating diagnostic support information

[1243] The server generates diagnostic support information based on the stored subjective, visual, and emotional information. For example, if the server recognizes subjective information such as "My headaches have been getting worse recently," along with visual information such as "Your face is pale" and "Your facial expression looks painful," as well as a sense of irritability, the server will refer to past data and medical guidelines and generate diagnostic support information such as "There is a high possibility that you have a migraine."

[1244] input:

[1245] Subjective information, visual information, emotional information

[1246] output:

[1247] Diagnostic support information (e.g., "Probable migraine").

[1248] Operation:

[1249] Integrated analysis is performed based on the stored data to generate diagnostic support information.

[1250] Step 13: Send to device

[1251] The server sends the generated diagnostic support information to the terminal and notifies the user. This information serves as a diagnostic support tool for medical providers.

[1252] input:

[1253] Diagnostic support information (e.g., "Probable migraine").

[1254] output:

[1255] Diagnostic support information (sent directly to the device)

[1256] Operation:

[1257] The diagnostic assistance information is transmitted to a terminal via a network.

[1258] (Application example 2)

[1259] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1260] Conventional health monitoring systems only analyze subjective and visual information, making it difficult to accurately grasp an employee's emotional state. This makes it difficult to detect abnormalities in health status early and take appropriate measures. Furthermore, because it takes time for employees to report their health status in detail, there is a risk of incomplete information or false reports. The present invention aims to solve these problems and provide a system that efficiently and accurately monitors the health status of employees in factories.

[1261] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes natural language processing means, image analysis means, medical record database means, emotion recognition means, means for extracting patient emotion information using the natural language processing means and emotion recognition means, and means for generating diagnostic support information based on the stored information. This makes it possible to comprehensively collect and analyze employee subjective information, visual information, and emotion information. Ultimately, it is possible to more accurately grasp the employee's health status, detect abnormalities early, and take appropriate action.

[1262] "Natural language processing means" is a technology that analyzes voice and text data and extracts meaning from its content.

[1263] "Image analysis means" is a technique for analyzing image data and extracting features contained therein.

[1264] "Medical record database means" refers to a database system that collects and manages health information of patients and employees.

[1265] "Means for extracting subjective patient information" refers to a technology that uses natural language processing to extract symptoms and sensations reported by patients from text data.

[1266] The "means for extracting visual information of a patient" is a technology that uses image analysis means to extract information related to the patient's health condition from their appearance, such as their complexion and facial expression.

[1267] "Emotion recognition means" is a technology that analyzes emotions from voice or text and extracts that emotional information.

[1268] The "means for generating diagnostic support information" is a system that automatically proposes appropriate diagnoses and countermeasures based on collected health information.

[1269] The system of the present invention combines natural language processing, image analysis, a medical record database, and emotion recognition, and is capable of efficiently and accurately collecting employee subjective health information, visual health information, and emotion information to monitor health status and detect abnormalities.

[1270] Implementation of natural language processing methods

[1271] The user inputs their physical condition and symptoms in voice or text format. If the user says, "My headache has been getting worse recently," the device converts this voice data into text using a voice recognition module. The converted text data is sent to a server, which analyzes it using a natural language processing engine to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[1272] Implementation of image analysis procedures

[1273] The device uses a built-in camera to capture visual information about the employee, such as their complexion and facial expression. This image data is sent from the device to a server, which then uses an image analysis engine to analyze the data and extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[1274] Implementing emotion recognition measures

[1275] When a user reports their symptoms in voice or text format, the emotion recognition means simultaneously analyzes the input data and extracts emotions. Specifically, for voice data, the server analyzes the tone and tempo of the voice to recognize emotions such as "irritability" or "sadness." For text data, the server works with a natural language processing engine to extract emotions from the text. The emotional information obtained in this way is also stored in the medical record database.

[1276] Implementation of medical record database measures

[1277] The medical record database means organizes and stores the collected subjective, visual, and emotional information. The server generates diagnostic support information based on the information stored in this database. For example, if the subjective information "My headaches have been getting worse recently" is recognized, along with visual information such as "Your face is pale" and "Your facial expression looks painful," and a "feeling of irritability" is also recognized, the server will refer to past data and medical guidelines to generate diagnostic support information stating "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[1278] Specific examples

[1279] Specific examples are shown below.

[1280] Example of subjective information extraction

[1281] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[1282] Example of information extraction using image analysis

[1283] The image of the employee's face captured by a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling" and stores it in a medical record database.

[1284] Example of information extraction using emotion engine

[1285] The user speaks, "I've been feeling sad lately." The device converts this into text data, "I've been feeling sad lately," and sends it to the server. The server uses emotion recognition means to extract the emotion "sadness" and stores it in the medical record database.

[1286] Example of generating diagnostic support information

[1287] Based on the subjective information "I've had a backache recently," the visual information "my face is red and swollen," and the emotional information "sadness," the server determines that "the back pain may be related to psychological or mental stress." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[1288] In this way, the system of the present invention performs employee health monitoring and early abnormality detection based on the integrated collection of subjective, visual, and emotional information. The following are examples of prompt sentences that can be input to the generative AI model:

[1289] I haven't been feeling well lately. My shoulders feel heavy and I just don't feel motivated.

[1290] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1291] Step 1:

[1292] The user inputs their physical condition and symptoms into the terminal by voice or text. The input voice data is converted into text data using a voice recognition module. At this stage, the input is voice data or text data, and the output is the converted text data. A specific example of operation is when the user speaks into the terminal, saying, "My headache has been getting worse recently."

[1293] Step 2:

[1294] The device sends the converted text data to the server. The server uses a natural language processing engine to analyze the text data and extract subjective information. At this stage, the input is text data, and the output is the analyzed subjective information. Specifically, the natural language processing engine extracts information such as "headache" and "getting worse."

[1295] Step 3:

[1296] The camera installed on the device is used to capture the employee's facial expression and facial complexion. The captured image data is sent to a server, which then uses an image analysis engine to extract visual information. At this stage, the input is image data, and the output is analyzed visual information. Specific actions such as "the employee's face is pale" and "they look like they're in pain" are extracted by the image analysis engine.

[1297] Step 4:

[1298] The data reported by the user through voice or text is simultaneously analyzed by the emotion recognition means to extract emotional information. At this stage, the input is voice data or text data, and the output is the analyzed emotional information. Specifically, the emotion recognition means analyzes the tone and tempo of the voice to recognize "anxiety" or "sadness."

[1299] Step 5:

[1300] The server stores the subjective, visual, and emotional information obtained above in a medical record database. The input at this stage is subjective, visual, and emotional information, and the output is a database in which this information is stored. Specifically, the information is integrated and organized in the medical record database.

[1301] Step 6:

[1302] The server generates diagnostic support information based on the information stored in the medical record database. The input at this stage is the information stored in the medical record database, and the output is the generated diagnostic support information. Specifically, by referencing past data and medical guidelines, diagnostic support information such as "high possibility of migraine" is generated.

[1303] Step 7:

[1304] The generated diagnostic assistance information is sent to the terminal and notified to the user. The input at this stage is the generated diagnostic assistance information, and the output is the diagnostic assistance information displayed on the terminal. As a specific operation, the diagnostic assistance information is notified to the user, and the user is prompted to take appropriate action.

[1305] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1306] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1307] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1308] [Fourth embodiment]

[1309] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1310] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1311] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1312] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1313] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1314] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1315] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1316] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1317] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1318] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1319] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1320] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1321] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1322] The system of the present invention is composed of a combination of natural language processing means, image analysis means, and medical record database means, and is capable of efficiently and accurately collecting subjective and visual information from patients and reflecting it in medical records.

[1323] Implementation of natural language processing methods

[1324] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device converts this speech into text data using a speech recognition module. The converted text data is sent to a server, which analyzes it using a natural language processing engine to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[1325] Implementation of image analysis procedures

[1326] The user uses the device's camera to capture visual information about the patient. For example, the camera can be used to record the patient's complexion and facial expression. This image data is sent from the device to a server, which then uses an image analysis engine to analyze the data and extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[1327] Implementation of medical record database measures

[1328] The medical record database means organizes and stores the collected subjective and visual information. The server generates diagnostic support information based on the information stored in this database. For example, based on subjective information such as "My headaches have been getting worse recently" and visual information such as "Your face is pale" and "Your expression looks painful," the server refers to past data and medical guidelines and generates diagnostic support information such as "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[1329] Specific examples

[1330] Specific examples are shown below.

[1331] Example of subjective information extraction

[1332] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[1333] Example of information extraction using image analysis

[1334] The image of the patient's face taken with a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[1335] Example of generating diagnostic support information

[1336] Based on subjective information such as "My back has been hurting recently" and visual information such as "My face is red and swollen," the server determines that "there is a high possibility of an inflammatory disease." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[1337] In this way, the system of the present invention efficiently and accurately collects subjective and visual information and reflects it in medical records, automating the process of entering medical records, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[1338] The processing flow will be explained below.

[1339] Extracting subjective information

[1340] Step 1:

[1341] The user inputs the patient's subjective symptoms into the terminal by voice or text. For example, consider the case where the user inputs by voice, "My headache has been getting worse recently."

[1342] Step 2:

[1343] The device converts the voice data into text data using a speech recognition module, and the resulting text data is "My headache has been getting worse recently."

[1344] Step 3:

[1345] The device sends the generated text data to the server, securely using the HTTPS protocol.

[1346] Step 4:

[1347] The server uses a natural language processing engine to analyze the text data and extract key subjective information, such as "headache" and "getting worse."

[1348] Step 5:

[1349] The server organizes the extracted subjective information and stores it in a medical record database, which electronically records the patient's complaints.

[1350] Image analysis

[1351] Step 1:

[1352] The user uses the device's camera to capture visual information of the patient, for example, recording the patient's complexion and facial expressions.

[1353] Step 2:

[1354] The device sends the captured image data to the server. The image is sent in a compressed format and the data is transferred securely using the HTTPS protocol.

[1355] Step 3:

[1356] The server uses an image analysis engine to analyze the image data, specifically analyzing the patient's complexion, facial expression, and movements to extract characteristic information.

[1357] Step 4:

[1358] For example, information such as "his face is pale" and "his expression looks like he's in pain" is extracted.

[1359] Step 5:

[1360] The server stores the extracted visual information in the medical record database, so that the visual information is reflected in the medical record.

[1361] Diagnostic support

[1362] Step 1:

[1363] The server integrates and analyzes the subjective and visual information stored in the medical record database.

[1364] Step 2:

[1365] The server generates diagnostic support information by referencing past data and medical guidelines. For example, it analyzes information such as "Your headaches have gotten worse recently" and "Your face is pale" and "You look like you're in pain."

[1366] Step 3:

[1367] Based on the integrated information, the server generates diagnostic support information such as "high possibility of migraine."

[1368] Step 4:

[1369] The server transmits the generated diagnostic assistance information to the terminal.

[1370] Step 5:

[1371] The device notifies the user of the diagnostic support information, for example, by providing information to a healthcare provider via a pop-up notification or email notification.

[1372] Through this process, the AISOAP system can efficiently and accurately collect patients' subjective and visual information and reflect it in medical records and diagnostic support information, automating the medical record entry process, reducing the burden on medical providers and improving the accuracy of diagnoses.

[1373] Example 1

[1374] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1375] In the modern medical environment, entering medical records is extremely time-consuming, placing a strain on healthcare providers. Furthermore, accurate diagnosis requires reliable collection of patient subjective and visual information, which must be integrated to generate diagnostic support information. However, conventional methods have proven difficult to achieve this efficiently and accurately.

[1376] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1377] In this invention, the server includes a natural language processing unit, an image analysis unit, and a medical record database unit, which enables efficient and accurate collection of subjective and visual information from patients and reflecting it in medical records.

[1378] A "natural language processing means" is a means for extracting subjective patient information from speech or text input.

[1379] The "image analysis means" is a means for extracting visual information of the patient from the camera image.

[1380] A "medical record database means" is a means for storing and managing extracted subjective and visual information.

[1381] "Voice input" refers to voice data uttered by a user through a microphone.

[1382] "Text input" refers to character data that a user inputs using a keyboard or the like.

[1383] A "voice recognition module" is a software or hardware component for converting voice input into text data.

[1384] "Camera footage" refers to video and still image data captured by a camera.

[1385] "Diagnostic support information" is information that helps with diagnosis and is generated based on stored information.

[1386] "User" refers to a healthcare provider or any individual or organization operating the System.

[1387] The system of the present invention is composed of a combination of natural language processing means, image analysis means, and medical record database means, and is capable of efficiently and accurately collecting subjective and visual information from patients and reflecting it in medical records.

[1388] Implementation of natural language processing methods

[1389] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device converts this speech into text data using the Google Cloud Speech-to-Text API. The converted text data is sent to the server. The server then analyzes the text data using a natural language processing engine (e.g., SpaCy or GPT-4) to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[1390] Implementation of image analysis procedures

[1391] The user uses the device's camera to capture visual information about the patient. For example, the camera can be used to record the patient's complexion and facial expression. This image data is sent from the device to a server, which then analyzes the data using an image analysis engine (e.g., OpenCV or TensorFlow) to extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in the medical record database.

[1392] Implementation of medical record database measures

[1393] The medical record database means organizes and stores the collected subjective and visual information. The server generates diagnostic support information based on the information stored in this database. For example, based on subjective information such as "My headaches have been getting worse recently" and visual information such as "Your face is pale" and "Your expression looks painful," the server refers to past data and medical guidelines and generates diagnostic support information such as "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[1394] Specific examples

[1395] Specific examples are shown below.

[1396] Example of subjective information extraction

[1397] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[1398] Example of information extraction using image analysis

[1399] The image of the patient's face taken with a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[1400] Example of generating diagnostic support information

[1401] Based on subjective information such as "My back has been hurting recently" and visual information such as "My face is red and swollen," the server determines that "there is a high possibility of an inflammatory disease." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[1402] Example prompt sentence:

[1403] "Convert the following speech input into text data and extract subjective information. Example: 'My back has been hurting lately.'"

[1404] By utilizing generative AI models, the system described above can collect patient information and provide diagnostic support information more efficiently and accurately than existing medical systems, reducing the burden on medical providers and improving diagnostic accuracy.

[1405] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1406] Step 1:

[1407] The user inputs the patient's symptoms into the terminal in voice or text format. When inputting by voice, the user speaks into a microphone, for example, "My headache has been getting worse recently." The input in this case is voice data or text data.

[1408] Step 2:

[1409] The device uses a speech recognition module to convert speech into text data. Specifically, it uses the Google Cloud Speech-to-Text API. This module receives speech data and converts it into text based on a language model. For example, the device outputs text data such as "My headache has been getting worse recently." This text data is then sent from the device to the server.

[1410] Step 3:

[1411] The server analyzes the received text data using a natural language processing engine (e.g., SpaCy or GPT-4). By inputting and analyzing the text data, subjective information such as "headache" and "getting worse" is extracted. The extracted information is stored in a medical record database. Specifically, the server receives the text data and uses the SpaCy library to extract keywords.

[1412] Step 4:

[1413] The user captures visual information of the patient using the device's camera. For example, they can take a picture of the patient's face and record their facial expression. The input image and video data are then sent from the device to the server.

[1414] Step 5:

[1415] The server analyzes the received image data using an image analysis engine (e.g., OpenCV or TensorFlow). The image data is input and visual information such as "pale face" or "pained expression" is extracted. The extracted information is also stored in the medical record database. Specifically, the server receives the image data and analyzes it using OpenCV.

[1416] Step 6:

[1417] The server generates diagnostic support information based on the subjective and visual information stored in the medical record database. For example, based on subjective information such as "Your headaches have gotten worse recently" and visual information such as "Your face is pale" and "Your expression looks painful," it references past data and medical guidelines to generate diagnostic support information such as "There is a high possibility of a migraine." The generated information is then stored again in the database and simultaneously sent to the device.

[1418] Step 7:

[1419] The device displays the received diagnostic support information on the user interface. The user is notified of new information and can confirm its contents. For example, the diagnostic support information "High possibility of migraine" may be displayed. This allows the user to quickly take necessary measures.

[1420] Through the above steps, patients' subjective and visual information can be collected efficiently and accurately and reflected in medical records, thereby reducing the burden on medical providers and improving diagnostic accuracy.

[1421] (Application example 1)

[1422] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1423] Conventional health management systems have had difficulty efficiently and accurately collecting subjective and visual information from patients to provide diagnostic support. Furthermore, they lacked a means to quickly notify users of the collected information, making it difficult to respond quickly in emergencies. The present invention aims to solve these problems and provide a health management system that is easy for healthcare providers and users to use.

[1424] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1425] In this invention, the server includes means for extracting subjective information of a patient using natural language processing means, means for extracting visual information of the patient using image analysis means, means for storing the extracted subjective information and visual information in a medical record database means, means for generating diagnostic support information based on the stored information, and means for notifying the smartphone terminal of the generated diagnostic support information, thereby enabling a quick analysis of the patient's health condition and a prompt response if necessary.

[1426] "Natural language processing means" is a technology for analyzing a patient's subjective information from voice or text input and extracting specific symptoms and conditions.

[1427] "Image analysis means" is a technology that analyzes the visual information of a patient obtained using a camera and quantitatively evaluates their health condition, such as their complexion and facial expression.

[1428] The "medical record database means" is a database technology that efficiently organizes and stores extracted subjective and visual information.

[1429] "Patient subjective information" is subjective data in audio or text format in which a patient describes their symptoms or physical condition.

[1430] "Patient visual information" refers to visual data such as the patient's appearance and facial expression obtained using a visual device such as a camera.

[1431] The "means for notifying a smartphone terminal" is a technology for sending the generated diagnostic assistance information to a smartphone and notifying the user in real time.

[1432] The "means for generating diagnostic support information" refers to a technology that generates information that assists in specific diagnoses based on stored subjective and visual information, and refers to medical data and guidelines.

[1433] The present invention provides a system for efficiently and accurately collecting subjective and visual information from a patient and providing diagnostic support based on that information. Specific embodiments for carrying out the present invention will be described below.

[1434] Hardware or software used

[1435] Hardware

[1436] Smartphone: Equipped with a microphone, speaker, and camera.

[1437] Server: For data analysis and database management.

[1438] software

[1439] Voice Recognition:

[1440] API: Google Speech Recognition API

[1441] Library: speech_recognition

[1442] Image analysis:

[1443] Deep Learning Model: A pre-trained image classification model using TensorFlow

[1444] Libraries: cv2 (OpenCV), PIL (Python Imaging Library), tensorflow

[1445] Generation of diagnostic support information:

[1446] API: Endpoint for external diagnostic support systems

[1447] Data processing and calculation

[1448] Audio data processing

[1449] The user launches the smartphone application and reports their current symptoms and physical condition by voice input. The smartphone uses a voice recognition module to convert the voice data into text, and then sends the text data to the server. The server uses a natural language processing engine to extract specific symptoms and conditions (e.g., "I have a headache" or "I have a fever") from the text data and stores them in a medical record database.

[1450] Image data processing

[1451] The user records their own facial expression and complexion using the smartphone camera. This image data is sent to a server, which then analyzes it using an image analysis engine. As a result, visual information such as "pale complexion" or "pained facial expression" is extracted and stored in a medical record database.

[1452] Generation and notification of diagnostic support information

[1453] The server generates diagnostic support information based on the subjective and visual information stored in the medical record database. The accuracy of this diagnostic support information is improved by referencing medical data and guidelines. The generated diagnostic support information is then notified to the user via a smartphone application.

[1454] Specific examples

[1455] Usage example

[1456] For example, a user wakes up in the morning feeling unwell and launches the app, then speaks the following:

[1457] Recently, I've been feeling tired and feverish.

[1458] Based on this voice input, the app performs an analysis, then uses the smartphone camera to take and analyze a photograph of the complexion. If the result is a diagnosis of "visual problems," the app will finally notify you as follows:

[1459] Your health may be deteriorating. We recommend that you consult a medical institution.

[1460] In this way, the system can quickly analyze the patient's health condition and help them take prompt action if necessary.

[1461] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1462] Step 1:

[1463] The user starts the smartphone application and reports their current symptoms and physical condition by voice input. At this time, the user speaks into the microphone. The input is the user's voice data, and the output is the voice data recorded by the smartphone device.

[1464] Step 2:

[1465] The device uses a speech recognition module (Google Speech Recognition API) to convert voice data into text data. The input is the user's voice data, and the output is symptom information in text format. The data processing performed here is text conversion based on speech recognition.

[1466] Step 3:

[1467] The terminal sends the converted text data to the server. The input is text data, and the output is data transmission to the server. There is no data processing, but communication is included.

[1468] Step 4:

[1469] The server uses a natural language processing engine to analyze the text data and extract specific symptoms and conditions (e.g., "I have a headache" or "I have a fever"). The input is text data, and the output is the analyzed symptom information. Data processing is an extraction process using natural language processing.

[1470] Step 5:

[1471] The server saves the analyzed symptom information to the medical record database. The input is the analyzed symptom information, and the output is the saving operation to the database. There is no data processing, but the saving operation is included.

[1472] Step 6:

[1473] A user records their facial expression and complexion using a smartphone camera. The device then acquires the image data. The input is the image data captured by the camera, and the output is the recorded image.

[1474] Step 7:

[1475] The terminal sends image data to the server. The input is image data, and the output is data transmission to the server. There is no data processing, but communication is included.

[1476] Step 8:

[1477] The server analyzes the image data using an image analysis engine (a pre-trained model using TensorFlow) and extracts visual information such as "the face looks pale" or "the face looks painful." The input is image data, and the output is the analyzed visual information. Data processing is the extraction process using image analysis.

[1478] Step 9:

[1479] The server stores the analyzed visual information in a medical record database. The input is the analyzed visual information, and the output is the saving operation to the database. There is no data processing, but the saving operation is included.

[1480] Step 10:

[1481] The server generates diagnostic support information based on the subjective and visual information stored in the medical record database. The input is subjective and visual information, and the output is diagnostic support information. Data processing involves information integration and analysis based on diagnostic algorithms.

[1482] Step 11:

[1483] The server notifies the smartphone of the generated diagnostic support information. The input is the diagnostic support information, and the output is a notification to the smartphone terminal. There is no data processing, but communication is involved.

[1484] Step 12:

[1485] The smartphone terminal displays the received diagnostic support information to the user. The input is the diagnostic support information, and the output is the presentation of information to the user. The specific operation is to display a notification message on the screen.

[1486] In this way, by going through each step, the user's health condition can be quickly analyzed and prompt action can be taken if necessary.

[1487] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1488] The system of the present invention is composed of a natural language processing means, an image analysis means, a medical record database means, and an emotion engine that recognizes the user's emotions. This system allows the efficient and accurate collection of subjective, visual, and emotional information of patients and reflects it in medical records.

[1489] Implementation of natural language processing methods

[1490] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device converts this speech into text data using a speech recognition module. The converted text data is sent to a server, which analyzes it using a natural language processing engine to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[1491] Implementation of image analysis procedures

[1492] The user uses the device's camera to capture visual information about the patient. For example, they can record the patient's complexion and facial expression. This image data is sent from the device to a server, which then uses an image analysis engine to analyze the data and extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[1493] Emotion Engine Implementation

[1494] The user inputs the patient's symptoms in voice or text format, and the emotion engine simultaneously analyzes the input data and extracts emotions. Specifically, for voice data, the server analyzes the tone and tempo of the voice to recognize emotions such as "irritability" or "sadness." For text data, the server works with a natural language processing engine to extract emotions from the text. The resulting emotional information is also stored in the medical record database.

[1495] Implementation of medical record database measures

[1496] The medical record database means organizes and stores the collected subjective, visual, and emotional information. The server generates diagnostic support information based on the information stored in this database. For example, if the subjective information "My headaches have been getting worse recently" is recognized, along with visual information such as "Your face is pale" and "Your facial expression looks painful," and a "feeling of irritability" is also recognized, the server will refer to past data and medical guidelines to generate diagnostic support information stating "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[1497] Specific examples

[1498] Specific examples are shown below.

[1499] Example of subjective information extraction

[1500] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[1501] Example of information extraction using image analysis

[1502] The image of the patient's face taken with a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[1503] Example of information extraction using emotion engine

[1504] The user speaks, "I've been feeling sad lately." The device converts this into text data, "I've been feeling sad lately," and sends it to the server. The server uses an emotion engine to extract the emotion "sadness" and stores it in the medical record database.

[1505] Example of generating diagnostic support information

[1506] Based on the subjective information "I've had a backache recently," the visual information "my face is red and swollen," and the emotional information "sadness," the server determines that "the back pain may be related to psychological or mental stress." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[1507] In this way, the system of the present invention efficiently and accurately collects subjective, visual, and emotional information and reflects it in medical records, thereby automating the process of entering medical records, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[1508] The processing flow will be explained below.

[1509] Implementation of natural language processing methods

[1510] Step 1:

[1511] The user inputs the patient's subjective symptoms into the terminal by voice or text. For example, consider the case where the user inputs by voice, "My headache has been getting worse recently."

[1512] Step 2:

[1513] The device converts the voice data into text data using a speech recognition module, and the resulting text data is "My headache has been getting worse recently."

[1514] Step 3:

[1515] The device sends the generated text data to the server, securely using the HTTPS protocol.

[1516] Step 4:

[1517] The server uses a natural language processing engine to analyze the text data and extract key subjective information, such as the phrases "headache" and "getting worse."

[1518] Step 5:

[1519] The server organizes the extracted subjective information and stores it in a medical record database, which electronically records the patient's complaints.

[1520] Implementation of image analysis procedures

[1521] Step 1:

[1522] The user uses the device's camera to capture visual information of the patient, for example, recording the patient's complexion and facial expressions.

[1523] Step 2:

[1524] The device sends the captured image data to the server. The image is sent in a compressed format and the data is transferred securely using the HTTPS protocol.

[1525] Step 3:

[1526] The server uses an image analysis engine to analyze the image data, specifically analyzing the patient's complexion, facial expression, and movements to extract characteristic information.

[1527] Step 4:

[1528] For example, information such as "his face is pale" and "his expression looks like he's in pain" is extracted.

[1529] Step 5:

[1530] The server stores the extracted visual information in the medical record database, so that the visual information is reflected in the medical record.

[1531] Emotion Engine Implementation

[1532] Step 1:

[1533] As the user inputs the patient's symptoms via voice or text, the emotion engine simultaneously analyzes the input data and extracts emotions.

[1534] Step 2:

[1535] In the case of voice data, the terminal transmits the voice data to the server, and in the case of text data, the text data is also transmitted to the server.

[1536] Step 3:

[1537] The server analyzes the tone and tempo of the voice to recognize emotions such as "anxiety" or "sadness." It also extracts emotions from text data.

[1538] Step 4:

[1539] The extracted emotional information is stored in the medical record database, so that the emotional information is also reflected in the medical record.

[1540] Diagnostic support

[1541] Step 1:

[1542] The server integrates and analyzes the subjective, visual, and emotional information stored in the medical record database.

[1543] Step 2:

[1544] The server generates diagnostic support information by referencing past data and medical guidelines. For example, it analyzes information such as "My headache has gotten worse recently," "My face is pale," and "I feel irritable."

[1545] Step 3:

[1546] Based on the integrated information, the server generates diagnostic support information such as "high possibility of migraine."

[1547] Step 4:

[1548] The server transmits the generated diagnostic assistance information to the terminal.

[1549] Step 5:

[1550] The device notifies the user of the diagnostic support information, for example, by providing information to a healthcare provider via a pop-up notification or email notification.

[1551] Through this process, the AISOAP system can efficiently and accurately collect patients' subjective, visual, and emotional information and reflect it in medical records and diagnostic support information, automating the input work of medical records, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[1552] Example 2

[1553] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1554] Modern healthcare requires efficient and accurate collection of patients' subjective, visual, and emotional information and integration of it into medical records. However, traditional methods make it difficult to centrally collect and integrate this information into medical records, which increases the burden on healthcare providers and potentially impacts diagnostic accuracy. To address this challenge, a system that can automatically analyze and integrate data collected from multiple sources is needed.

[1555] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a natural language processing means, an image analysis means, a medical record database means, and an emotion recognition means. This makes it possible to efficiently and accurately collect subjective information, visual information, and emotion information of the patient and reflect it in the medical record.

[1556] "Natural language processing means" is a technology that analyzes input data in voice or text format and extracts meaning and keywords.

[1557] "Image analysis means" is a technology that analyzes video data acquired from a camera or other device and extracts visual information such as the patient's complexion, facial expression, and movements.

[1558] The "medical record database means" is a database system that organizes and stores collected subjective, visual, and emotional information.

[1559] "Emotion recognition means" is a technology that analyzes voice and text data and extracts emotions from them.

[1560] "Subjective information" refers to information such as symptoms and discomfort reported by the patient themselves.

[1561] "Visual information" refers to information extracted from video data acquired using a camera or the like, such as the patient's complexion, facial expression, and movements.

[1562] "Emotion information" is information about the patient's emotions extracted from voice and text data.

[1563] "Diagnostic support information" is information that assists diagnosis and is generated based on collected subjective information, visual information, and emotional information.

[1564] The system of the present invention is composed of a combination of natural language processing means, image analysis means, medical record database means, and emotion recognition means, and is capable of efficiently and accurately collecting subjective, visual, and emotional information from patients and reflecting it in medical records.

[1565] Implementation of natural language processing methods

[1566] The user inputs the patient's subjective symptoms into the device in voice or text format. For example, if the user says, "My headache has been getting worse recently," the device uses a voice recognition module (e.g., a voice recognition API) to convert the speech into text data. The converted text data is sent to a server, which then analyzes the text data using a natural language processing engine (e.g., SpaCy) to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[1567] Implementation of image analysis procedures

[1568] The user uses the device's camera to capture visual information about the patient. For example, they can record the patient's complexion and facial expression. This image data is sent from the device to a server, which then analyzes the data using an image analysis engine (e.g., OpenCV) to extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[1569] Implementing emotion recognition measures

[1570] The user inputs the patient's symptoms in voice or text format, and at the same time, the emotion recognition means analyzes the input data and extracts emotions. Specifically, in the case of voice data, the server analyzes the tone and tempo of the voice to recognize emotions such as "irritability" or "sadness." In the case of text data, emotions are extracted from the text in cooperation with a natural language processing engine. The emotional information obtained in this way is also stored in the medical record database.

[1571] Implementation of medical record database measures

[1572] The medical record database means organizes and stores the collected subjective, visual, and emotional information. The server generates diagnostic support information based on the information stored in this database. For example, if the subjective information "My headaches have been getting worse recently" is recognized, along with visual information such as "Your face is pale" and "Your facial expression looks painful," and a "feeling of irritability" is also recognized, the server will refer to past data and medical guidelines to generate diagnostic support information stating "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[1573] Specific examples

[1574] Example of subjective information extraction

[1575] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[1576] Example of information extraction using image analysis

[1577] The camera takes a picture of the patient's face and sends the image to a server, which uses an image analysis engine to extract information such as "redness and swelling of the face" and store it in a medical record database.

[1578] Example of information extraction using emotion recognition methods

[1579] The user speaks, "I've been feeling sad lately." The device converts this into text data, "I've been feeling sad lately," and sends it to the server. The server uses emotion recognition means to extract the emotion "sadness" and stores it in the medical record database.

[1580] Example of generating diagnostic support information

[1581] Based on the subjective information "I've had a backache recently," the visual information "my face is red and swollen," and the emotional information "sadness," the server determines that "the back pain may be related to psychological or mental stress." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[1582] Through the above process, this system can efficiently and accurately collect patients' subjective, visual, and emotional information and reflect it in medical records, which is an important tool for automating medical record entry work, reducing the burden on medical providers, and improving the accuracy of diagnoses.

[1583] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1584] Step 1: User Input

[1585] The user inputs the patient's subjective symptoms into the terminal in voice or text format. This input becomes the basic data for the system. Specifically, the user says, "My headache has been getting worse recently."

[1586] input:

[1587] Speech data (e.g., "My headaches have been getting worse lately.")

[1588] output:

[1589] Audio data (as is)

[1590] Step 2: Voice Recognition

[1591] The device uses a speech recognition module (e.g., speech recognition API) to convert the speech into text data. At this stage, the speech becomes text data, which is used in the next step.

[1592] input:

[1593] Speech data (e.g., "My headaches have been getting worse lately.")

[1594] output:

[1595] Text data (e.g., "My headaches have been getting worse lately.")

[1596] Operation:

[1597] The audio waveform is analyzed and converted into string data.

[1598] Step 3: Send to the server

[1599] The device then sends the converted text data to the server, where it is analyzed, so fast and accurate transmission is essential.

[1600] input:

[1601] Text data (e.g., "My headaches have been getting worse lately.")

[1602] output:

[1603] Text data (sent as is to the server)

[1604] Operation:

[1605] The text data is uploaded to a server via a network.

[1606] Step 4: Natural Language Processing

[1607] The server uses a natural language processing engine (e.g., SpaCy) to analyze the received text data, and extracts subjective information (e.g., "headache" and "getting worse") as a result of the analysis.

[1608] input:

[1609] Text data (e.g., "My headaches have been getting worse lately.")

[1610] output:

[1611] Subjective information (e.g., "headache" or "it's getting worse")

[1612] Operation:

[1613] Morphological analysis of text data is performed to extract important words and phrases.

[1614] Step 5: Preserving subjective information

[1615] The server stores the extracted subjective information in a medical record database, which is used for subsequent analysis and medical support.

[1616] input:

[1617] Subjective information (e.g., "headache" or "it's getting worse")

[1618] output:

[1619] Saving to a database

[1620] Operation:

[1621] The extracted information is written to a medical record database.

[1622] Step 6: Capture visual information

[1623] The user uses the device's camera to capture visual information of the patient, for example, recording the patient's complexion and facial expressions, and saves this data.

[1624] input:

[1625] Video data (e.g., patient's complexion and facial expression)

[1626] output:

[1627] Video data (as is)

[1628] Operation:

[1629] Use a camera to capture video data.

[1630] Step 7: Sending visual data to the server

[1631] The device sends the recorded image data to a server, where it is analyzed, so it must be sent quickly and accurately.

[1632] input:

[1633] Video data (e.g., patient's complexion and facial expression)

[1634] output:

[1635] Video data (sent as is to the server)

[1636] Operation:

[1637] The video data is uploaded to a server via a network.

[1638] Step 8: Image analysis

[1639] The server analyzes the received video data using an image analysis engine (e.g., OpenCV), and extracts visual information (e.g., "the face is pale" or "the facial expression looks painful").

[1640] input:

[1641] Video data (e.g., patient's complexion and facial expression)

[1642] output:

[1643] Visual information (e.g., "He looks pale" or "He looks like he's in pain")

[1644] Operation:

[1645] Frame analysis of video data is performed to extract important visual information.

[1646] Step 9: Save the visual information

[1647] The server stores the extracted visual information in a medical record database, which is used for subsequent analysis and medical support.

[1648] input:

[1649] Visual information (e.g., "He looks pale" or "He looks like he's in pain")

[1650] output:

[1651] Saving to a database

[1652] Operation:

[1653] The extracted information is written to a medical record database.

[1654] Step 10: Emotion Recognition

[1655] The server analyzes the voice or text data entered by the user using an emotion recognition engine (e.g., emotion recognition API) to extract emotional information.

[1656] input:

[1657] Voice or text data (e.g., "I've been feeling sad lately")

[1658] output:

[1659] Emotional information (e.g., "sadness")

[1660] Operation:

[1661] Analyze the tone and content of voice and text data to identify emotions.

[1662] Step 11: Storing Emotional Information

[1663] The server stores the extracted emotion information in a medical record database, which is used for subsequent analysis and medical support.

[1664] input:

[1665] Emotional information (e.g., "sadness")

[1666] output:

[1667] Saving to a database

[1668] Operation:

[1669] The extracted emotion information is written to a medical record database.

[1670] Step 12: Generating diagnostic support information

[1671] The server generates diagnostic support information based on the stored subjective, visual, and emotional information. For example, if the server recognizes subjective information such as "My headaches have been getting worse recently," along with visual information such as "Your face is pale" and "Your facial expression looks painful," as well as a sense of irritability, the server will refer to past data and medical guidelines and generate diagnostic support information such as "There is a high possibility that you have a migraine."

[1672] input:

[1673] Subjective information, visual information, emotional information

[1674] output:

[1675] Diagnostic support information (e.g., "Probable migraine").

[1676] Operation:

[1677] Integrated analysis is performed based on the stored data to generate diagnostic support information.

[1678] Step 13: Send to device

[1679] The server sends the generated diagnostic support information to the terminal and notifies the user. This information serves as a diagnostic support tool for medical providers.

[1680] input:

[1681] Diagnostic support information (e.g., "Probable migraine").

[1682] output:

[1683] Diagnostic support information (sent directly to the device)

[1684] Operation:

[1685] The diagnostic assistance information is transmitted to a terminal via a network.

[1686] (Application example 2)

[1687] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1688] Conventional health monitoring systems only analyze subjective and visual information, making it difficult to accurately grasp an employee's emotional state. This makes it difficult to detect abnormalities in health status early and take appropriate measures. Furthermore, because it takes time for employees to report their health status in detail, there is a risk of incomplete information or false reports. The present invention aims to solve these problems and provide a system that efficiently and accurately monitors the health status of employees in factories.

[1689] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes natural language processing means, image analysis means, medical record database means, emotion recognition means, means for extracting patient emotion information using the natural language processing means and emotion recognition means, and means for generating diagnostic support information based on the stored information. This makes it possible to comprehensively collect and analyze employee subjective information, visual information, and emotion information. Ultimately, it is possible to more accurately grasp the employee's health status, detect abnormalities early, and take appropriate action.

[1690] "Natural language processing means" is a technology that analyzes voice and text data and extracts meaning from its content.

[1691] "Image analysis means" is a technique for analyzing image data and extracting features contained therein.

[1692] "Medical record database means" refers to a database system that collects and manages health information of patients and employees.

[1693] "Means for extracting subjective patient information" refers to a technology that uses natural language processing to extract symptoms and sensations reported by patients from text data.

[1694] The "means for extracting visual information of a patient" is a technology that uses image analysis means to extract information related to the patient's health condition from their appearance, such as their complexion and facial expression.

[1695] "Emotion recognition means" is a technology that analyzes emotions from voice or text and extracts that emotional information.

[1696] The "means for generating diagnostic support information" is a system that automatically proposes appropriate diagnoses and countermeasures based on collected health information.

[1697] The system of the present invention combines natural language processing, image analysis, a medical record database, and emotion recognition, and is capable of efficiently and accurately collecting employee subjective health information, visual health information, and emotion information to monitor health status and detect abnormalities.

[1698] Implementation of natural language processing methods

[1699] The user inputs their physical condition and symptoms in voice or text format. If the user says, "My headache has been getting worse recently," the device converts this voice data into text using a voice recognition module. The converted text data is sent to a server, which analyzes it using a natural language processing engine to extract subjective information such as "headache" and "getting worse." This extracted information is then stored in a medical record database.

[1700] Implementation of image analysis procedures

[1701] The device uses a built-in camera to capture visual information about the employee, such as their complexion and facial expression. This image data is sent from the device to a server, which then uses an image analysis engine to analyze the data and extract visual information such as "pale complexion" or "pained facial expression." This information is also stored in a medical record database.

[1702] Implementing emotion recognition measures

[1703] When a user reports their symptoms in voice or text format, the emotion recognition means simultaneously analyzes the input data and extracts emotions. Specifically, for voice data, the server analyzes the tone and tempo of the voice to recognize emotions such as "irritability" or "sadness." For text data, the server works with a natural language processing engine to extract emotions from the text. The emotional information obtained in this way is also stored in the medical record database.

[1704] Implementation of medical record database measures

[1705] The medical record database means organizes and stores the collected subjective, visual, and emotional information. The server generates diagnostic support information based on the information stored in this database. For example, if the subjective information "My headaches have been getting worse recently" is recognized, along with visual information such as "Your face is pale" and "Your facial expression looks painful," and a "feeling of irritability" is also recognized, the server will refer to past data and medical guidelines to generate diagnostic support information stating "There is a high possibility of a migraine." This information is sent to the terminal and notified to the user.

[1706] Specific examples

[1707] Specific examples are shown below.

[1708] Example of subjective information extraction

[1709] The user speaks, "My back hurts recently." The device converts this into text data, "My back hurts recently," and sends it to the server. The server uses a natural language processing engine to extract "My back hurts," and stores it in the medical record database.

[1710] Example of information extraction using image analysis

[1711] The image of the employee's face captured by a camera is sent to a server, which uses an image analysis engine to extract information such as "redness and swelling" and stores it in a medical record database.

[1712] Example of information extraction using emotion engine

[1713] The user speaks, "I've been feeling sad lately." The device converts this into text data, "I've been feeling sad lately," and sends it to the server. The server uses emotion recognition means to extract the emotion "sadness" and stores it in the medical record database.

[1714] Example of generating diagnostic support information

[1715] Based on the subjective information "I've had a backache recently," the visual information "my face is red and swollen," and the emotional information "sadness," the server determines that "the back pain may be related to psychological or mental stress." It generates diagnostic support information, sends it to the terminal, and notifies the user.

[1716] In this way, the system of the present invention performs employee health monitoring and early abnormality detection based on the integrated collection of subjective, visual, and emotional information. The following are examples of prompt sentences that can be input to the generative AI model:

[1717] I haven't been feeling well lately. My shoulders feel heavy and I just don't feel motivated.

[1718] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1719] Step 1:

[1720] The user inputs their physical condition and symptoms into the terminal by voice or text. The input voice data is converted into text data using a voice recognition module. At this stage, the input is voice data or text data, and the output is the converted text data. A specific example of operation is when the user speaks into the terminal, saying, "My headache has been getting worse recently."

[1721] Step 2:

[1722] The device sends the converted text data to the server. The server uses a natural language processing engine to analyze the text data and extract subjective information. At this stage, the input is text data, and the output is the analyzed subjective information. Specifically, the natural language processing engine extracts information such as "headache" and "getting worse."

[1723] Step 3:

[1724] The camera installed on the device is used to capture the employee's facial expression and facial complexion. The captured image data is sent to a server, which then uses an image analysis engine to extract visual information. At this stage, the input is image data, and the output is analyzed visual information. Specific actions such as "the employee's face is pale" and "they look like they're in pain" are extracted by the image analysis engine.

[1725] Step 4:

[1726] The data reported by the user through voice or text is simultaneously analyzed by the emotion recognition means to extract emotional information. At this stage, the input is voice data or text data, and the output is the analyzed emotional information. Specifically, the emotion recognition means analyzes the tone and tempo of the voice to recognize "anxiety" or "sadness."

[1727] Step 5:

[1728] The server stores the subjective, visual, and emotional information obtained above in a medical record database. The input at this stage is subjective, visual, and emotional information, and the output is a database in which this information is stored. Specifically, the information is integrated and organized in the medical record database.

[1729] Step 6:

[1730] The server generates diagnostic support information based on the information stored in the medical record database. The input at this stage is the information stored in the medical record database, and the output is the generated diagnostic support information. Specifically, by referencing past data and medical guidelines, diagnostic support information such as "high possibility of migraine" is generated.

[1731] Step 7:

[1732] The generated diagnostic assistance information is sent to the terminal and notified to the user. The input at this stage is the generated diagnostic assistance information, and the output is the diagnostic assistance information displayed on the terminal. As a specific operation, the diagnostic assistance information is notified to the user, and the user is prompted to take appropriate action.

[1733] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1734] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1735] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1736] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1737] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1738] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1739] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1740] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1741] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1742] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1743] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1744] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1745] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1746] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1747] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1748] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1749] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1750] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1751] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1752] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1753] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1754] The following is further disclosed regarding the above embodiment.

[1755] The draft claims are shown below.

[1756] (Claim 1)

[1757] natural language processing means;

[1758] Image analysis means;

[1759] a medical record database means;

[1760] means for extracting subjective information of a patient using the natural language processing means;

[1761] means for extracting visual information of a patient using said image analysis means;

[1762] means for storing the extracted subjective and visual information in the medical record database means;

[1763] means for generating diagnostic assistance information based on the stored information;

[1764] Including system.

[1765] (Claim 2)

[1766] 10. The system of claim 1, wherein the natural language processing means includes means for converting speech input to text.

[1767] (Claim 3)

[1768] 10. The system of claim 1, wherein the image analysis means includes means for analyzing the patient's complexion, facial expression, and movements.

[1769] (Claim 4)

[1770] The system of claim 1 , wherein the diagnostic support information is generated based on historical data and medical guidelines.

[1771] "Example 1"

[1772] (Claim 1)

[1773] natural language processing means;

[1774] Image analysis means;

[1775] a medical record database means;

[1776] means for extracting subjective patient information from speech or text input using said natural language processing means;

[1777] means for extracting visual information of the patient from the camera image using the image analysis means;

[1778] means for storing the extracted subjective and visual information in the medical record database means;

[1779] means for generating diagnostic assistance information based on the stored information and notifying the information to a terminal;

[1780] Including system.

[1781] (Claim 2)

[1782] 10. The system of claim 1, wherein the natural language processing means includes means for converting voice input to text using a speech recognition module.

[1783] (Claim 3)

[1784] 2. The system according to claim 1, wherein the image analysis means includes means for analyzing the patient's complexion and facial expression.

[1785] "Application Example 1"

[1786] (Claim 1)

[1787] natural language processing means;

[1788] Image analysis means;

[1789] a medical record database means;

[1790] means for extracting subjective information of a patient using the natural language processing means;

[1791] means for extracting visual information of a patient using said image analysis means;

[1792] means for storing the extracted subjective and visual information in the medical record database means;

[1793] a means for generating diagnostic support information based on the stored information;

[1794] a means for notifying a smartphone terminal of the generated diagnostic assistance information,

[1795] Including system.

[1796] (Claim 2)

[1797] 10. The system of claim 1, wherein the natural language processing means includes means for converting speech input to text.

[1798] (Claim 3)

[1799] 10. The system of claim 1, wherein the image analysis means includes means for analyzing the patient's complexion, facial expression, and movements.

[1800] "Example 2: Combining Emotion Engines"

[1801] (Claim 1)

[1802] natural language processing means;

[1803] Image analysis means;

[1804] a medical record database means;

[1805] An emotion recognition means;

[1806] means for extracting subjective information of a patient using the natural language processing means;

[1807] means for extracting visual information of a patient using said image analysis means;

[1808] means for extracting emotional information of a patient using the emotion recognition means;

[1809] means for storing the extracted subjective, visual and affective information in said medical record database means;

[1810] means for generating diagnostic assistance information based on the stored information;

[1811] Including system.

[1812] (Claim 2)

[1813] 10. The system of claim 1, wherein the natural language processing means includes means for converting speech input to text.

[1814] (Claim 3)

[1815] 10. The system of claim 1, wherein the image analysis means includes means for analyzing the patient's complexion, facial expression, and movements.

[1816] "Application example 2 when combining emotion engines"

[1817] (Claim 1)

[1818] natural language processing means;

[1819] Image analysis means;

[1820] a medical record database means;

[1821] means for extracting subjective information of a patient using the natural language processing means;

[1822] means for extracting visual information of a patient using said image analysis means;

[1823] means for storing the extracted subjective and visual information in the medical record database means;

[1824] An emotion recognition means;

[1825] means for extracting patient emotion information using the natural language processing means and emotion recognition means;

[1826] means for generating diagnostic assistance information based on the stored information;

[1827] Including system.

[1828] (Claim 2)

[1829] 10. The system of claim 1, wherein the natural language processing means includes means for converting speech input to text.

[1830] (Claim 3)

[1831] 10. The system of claim 1, wherein the image analysis means includes means for analyzing the patient's complexion and facial expression. [Explanation of symbols]

[1832] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. natural language processing means; Image analysis means; a medical record database means; means for extracting subjective information of a patient using the natural language processing means; means for extracting visual information of a patient using said image analysis means; means for storing the extracted subjective and visual information in the medical record database means; means for generating diagnostic assistance information based on the stored information; Including system.

2. 2. The system of claim 1, wherein the natural language processing means includes means for converting speech input to text.

3. 2. The system of claim 1, wherein said image analysis means includes means for analyzing the patient's complexion, facial expression, and movements.

4. The system of claim 1 , wherein the diagnostic support information is generated based on historical data and medical guidelines.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A