System
A voice recognition system automates the conversion of patient audio to text for medical records, reducing manual effort and errors, thus enhancing efficiency and accuracy in medical recordkeeping.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2026-03-06
AI Technical Summary
Manual entry of patient information into medical records is time-consuming and prone to errors, placing a heavy burden on medical professionals and reducing work efficiency.
A system that utilizes voice recognition technology to convert patient audio data into text, automatically complete medical records, and allows users to confirm and correct the content, thereby streamlining the process from voice input to final record storage.
The system reduces the burden on medical professionals by efficiently and accurately reflecting patient information in medical records, minimizing errors and improving work efficiency.
Smart Images

Figure 2026037344000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Previously, when medical professionals entered information obtained from patients into medical records, it was often done manually, which was time-consuming and laborious. Furthermore, to accurately reflect the information obtained in the medical records, the voice data had to be converted into text manually, which could lead to input errors or missing information. Furthermore, because the work of completing the medical records was also done manually, it placed a heavy burden on medical professionals and reduced work efficiency. To solve these issues, a system was needed that automatically converted patient information into text data using voice recognition technology, combined with an automatic medical record completion function. [Means for solving the problem]
[0005] This invention provides a system that allows medical professionals to efficiently and accurately reflect patient interview content in medical records. Specifically, the system includes a means for acquiring patient dictation as audio data and a means for transmitting the audio data to a server, which then converts the audio data into text data. The server also includes a means for analyzing the converted text data and automatically completing the medical record content. The system also includes a means for displaying the automatically completed medical record content, allowing the user to confirm and correct it, and a means for transmitting and saving the final corrected content to the server. This configuration allows the interview content to be quickly and accurately reflected in the medical record, reducing the burden on medical professionals and improving work efficiency.
[0006] "Patient" means a person who receives medical advice or treatment from a healthcare professional for a health or illness condition.
[0007] "Oral information" refers to information such as symptoms and medical history that is communicated by the patient verbally.
[0008] "Audio data" refers to data that records audio in digital format.
[0009] A "server" is a computer system that stores and processes data over a network.
[0010] A "terminal" is a device used by medical professionals that acquires and displays voice data.
[0011] A "voice recognition engine" is software or a system that analyzes voice data and converts the voice into text data.
[0012] "Text data" refers to data expressed in character string format.
[0013] A "medical record" is a medical record that contains information such as the patient's medical examination results, treatment details, and medical history.
[0014] "Auto-completion" is a function that automatically adds or supplements necessary information based on the input text data.
[0015] "Users" refer to medical professionals (doctors, nurses, etc.) who operate the system.
[0016] "Confirmation and correction" refers to the process in which the user checks the automatically completed medical record content displayed and makes corrections or amendments as necessary.
[0017] "Storing" refers to the process of recording the modified data on a server so that it can be used later. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] As an embodiment of this invention, we provide a system that acquires oral information from conversations with patients in real time and automatically reflects it in medical records. The system combines voice recognition technology with an automatic medical record completion function based on text data, reducing the burden on medical professionals and improving work efficiency.
[0040] Program processing overview
[0041] Voice input
[0042] The user interviews the patient about their symptoms and medical history. The device collects this conversation in real time as audio data. For example, a medical professional might ask the patient, "What symptoms do you have?" and the patient might reply, "I have a headache, and it's been going on since yesterday." The device records this conversation as audio data.
[0043] Voice Recognition
[0044] The recorded voice data is sent from the device to the server. The server uses a voice recognition engine to convert this voice data into text data. For example, the voice "I have a headache, and it has been going on since yesterday" is converted into the text "I have a headache, and it has been going on since yesterday."
[0045] Auto-completion
[0046] The server analyzes the converted text and extracts relevant keywords. Based on the extracted keywords, the server automatically completes the medical record contents using the patient's existing data and general medical knowledge. For example, based on the keyword "headache," information such as "Date of onset: yesterday. Pain level: moderate" is added.
[0047] Check and fix
[0048] The automatically completed medical record information is sent to the terminal and displayed to the user. The user can check the information and make corrections as necessary. For example, a doctor can change the "pain level" from "moderate" to "severe."
[0049] keep
[0050] Once the changes are confirmed, the device sends the revised medical record information to the server. The server saves the received data in the patient's electronic medical record. As a result, for example, data such as "I have a headache that has been going on since yesterday. Onset date: yesterday. Pain level: severe" is recorded in the electronic medical record.
[0051] This system automates the entire process from voice input to final medical record storage, enabling medical professionals to efficiently and accurately record patient information. This invention is particularly useful when doctors and nurses need to quickly and accurately create medical records based on their conversations with patients.
[0052] The processing flow will be explained below.
[0053] Step 1:
[0054] The user interviews the patient about their symptoms and medical history. For example, a medical professional asks, "What symptoms do you have?" The patient answers, "I have a headache, and it's been going on since yesterday."
[0055] Step 2:
[0056] The terminal collects the conversation between the user and the patient as voice data. The terminal starts recording and records the voice data in real time.
[0057] Step 3:
[0058] The device converts the recorded audio data into a digital format and sends it to the server, where it is compressed to optimize bandwidth.
[0059] Step 4:
[0060] The server analyzes the received voice data and converts it into text data using a voice recognition engine. For example, a voice saying "I have a headache, and it's been going on since yesterday" is converted into text "I have a headache, and it's been going on since yesterday."
[0061] Step 5:
[0062] The server analyzes the text data and extracts keywords. Based on the extracted keywords, the server references the patient's existing information and medical databases to automatically complete the information. For example, the keyword "headache" can be supplemented with the information "Date of onset: yesterday, Pain level: moderate."
[0063] Step 6:
[0064] The server generates automatically completed medical record information and sends it to the terminal. For example, medical record data such as "I have a headache that has continued since yesterday. Onset: yesterday. Pain level: moderate" is generated.
[0065] Step 7:
[0066] The terminal displays the received medical record information to the user, who then checks the displayed information and makes any necessary corrections.
[0067] Step 8:
[0068] The user checks the medical record information and makes corrections as necessary, for example, changing the "pain level" from "moderate" to "severe."
[0069] Step 9:
[0070] The terminal sends the medical record information corrected by the user to the server, and after final confirmation, the data is sent.
[0071] Step 10:
[0072] The server stores the received corrected medical record information as the patient's electronic medical record. For example, information such as "I have a headache that has continued since yesterday. Onset: yesterday. Pain level: severe" is recorded in the electronic medical record.
[0073] Example 1
[0074] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0075] In conventional medical settings, creating medical records based on conversations with patients requires a great deal of time and effort. Furthermore, manual data entry often carries the risk of errors and omissions, increasing the burden on medical professionals. The purpose of this invention is to provide a system that reduces the burden on medical professionals, improves work efficiency, and prevents errors and omissions.
[0076] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0077] In this invention, the server includes a means for converting voice data into text data using a voice recognition engine, a means for analyzing the text data and extracting related keywords, and a means for automatically completing the medical record content based on the keywords. This automates the process from voice input to medical record creation, reducing the burden on medical professionals and improving work efficiency and data accuracy.
[0078] "Oral information from patients" refers to audio information such as symptoms and medical history obtained directly from patients by medical professionals.
[0079] "Audio data" refers to digital audio files of recorded conversations with patients.
[0080] "Terminal" refers to an electronic device for collecting and transmitting voice data to a server.
[0081] "Server" refers to the computer system that processes voice data and automatically completes medical records.
[0082] "Speech recognition engine" refers to software or algorithms that analyze voice data and convert it into text data.
[0083] "Text data" refers to character information converted by a voice recognition engine.
[0084] "Natural language processing technology" refers to computer technology for analyzing text data and extracting keywords.
[0085] "Keywords" refer to important words and phrases extracted from text data that are necessary for automatically completing the contents of medical records.
[0086] A "medical record" refers to a medical record that describes a patient's symptoms, medical history, treatment details, etc.
[0087] "Automatic completion" refers to the process in which the system automatically adds and supplements medical record content based on keywords.
[0088] "User" refers to a medical professional who operates the system to check and correct medical records.
[0089] As an embodiment of this invention, we provide a system that acquires oral information from conversations with patients in real time and automatically reflects it in medical records. The system combines voice recognition technology with an automatic medical record completion function based on text data, reducing the burden on medical professionals and improving work efficiency.
[0090] Collecting voice input
[0091] Users, or medical professionals, use devices such as smartphones and tablets to obtain voice data of symptoms and medical history directly from patients. The devices are equipped with highly sensitive microphones that record conversations in real time. For example, if a medical professional asks, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday," this content is collected as voice data on the device.
[0092] Sending audio data
[0093] The device sends the recorded audio data to a server over the Internet using the HTTPS protocol to ensure data security. A program on the device converts the audio data into a digital format and sends it to the server.
[0094] Speech Recognition Processing
[0095] The server passes the received voice data to a voice recognition engine, which converts the voice into text data. A typical voice recognition engine used here is a third-party voice recognition service. For example, voice data such as "I have a headache, and it's been going on since yesterday" is converted directly into text data such as "I have a headache, and it's been going on since yesterday."
[0096] Text analysis and keyword extraction
[0097] The server receives the converted text data and analyzes it using natural language processing technology. Related keywords are extracted through the analysis. For example, the keywords "headache," "yesterday," and "continuing" are extracted from the sentence "I have a headache, and it has been going on since yesterday."
[0098] Auto-completion
[0099] The server automatically completes the medical record based on the extracted keywords. It references the medical database and adds information related to the keywords to the medical record. For example, based on the keyword "headache," the information "Date of onset: yesterday. Pain level: moderate" is automatically added to the medical record.
[0100] Sending and checking medical record contents
[0101] The completed medical record information is sent to the terminal and displayed to the user. The user can check the medical record information and make corrections as necessary. For example, a doctor may change the "pain level" from "moderate" to "severe."
[0102] Saving medical record contents
[0103] Once the user has confirmed the medical record contents after making the edits, the terminal sends them back to the server. The server then saves the received medical record contents in the patient's electronic medical record. For example, information such as "I have a headache that has been going on since yesterday. Onset date: yesterday. Pain level: severe" is recorded in the electronic medical record.
[0104] Examples of concrete examples and prompts
[0105] Example: A healthcare professional asks a patient, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday." The audio is recorded in real time and converted to text. Auto-complete adds, "Date of onset: yesterday. Pain level: moderate." The doctor then corrects "moderate" to "severe." The information is saved in the electronic medical record.
[0106] Example prompt: "Please run a program that will automatically record a conversation about your headache symptoms that have continued since yesterday into the electronic medical record."
[0107] In this way, the system automates the process from voice input to final medical record storage, enabling medical professionals to efficiently and accurately record patient information.
[0108] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0109] Step 1: Collecting voice input
[0110] The user listens to the patient's symptoms and medical history and engages in a conversation. The device uses a built-in microphone to collect this audio in real time and records it as audio data. The input is the conversation between the user and the patient, and the output is digital audio data. Specifically, the medical professional asks, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday." This audio is collected by the device.
[0111] Step 2: Sending audio data
[0112] The audio data collected by the device is sent to a server via the Internet. The input is digital audio data, and the output is audio data transferred to the server. Specifically, a program on the device sends the audio data to the server using the HTTPS protocol. This ensures that the data arrives securely at the server.
[0113] Step 3: Speech recognition processing
[0114] The server passes the received voice data to a voice recognition engine, which converts the voice into text data. The input is the voice data sent to the server, and the output is text data. Specifically, the server calls a voice recognition engine (for example, a natural language processing library) and converts the voice data "I have a headache, and it's been going on since yesterday" into the text "I have a headache, and it's been going on since yesterday."
[0115] Step 4: Text analysis and keyword extraction
[0116] The server analyzes the converted text data and extracts related keywords. The input is text data obtained from the speech recognition engine, and the output is the extracted keywords. Specifically, the server uses natural language processing technology to extract keywords such as "headache," "yesterday," and "continuing" from the text "I have a headache, and it has been going on since yesterday."
[0117] Step 5: Auto-completion
[0118] The server automatically completes the medical record contents based on the extracted keywords. The input is the extracted keywords, and the output is the automatically completed medical record contents. Specifically, the server references the medical database and adds information such as "Date of onset: yesterday. Pain level: moderate" to the medical record based on the keywords "headache," "yesterday," and "continuing."
[0119] Step 6: Send and view medical records
[0120] The server sends the auto-completed medical record contents to the terminal and displays them to the user. The input is the auto-completed medical record contents, and the output is the medical record information displayed on the terminal. Specifically, the server sends the medical record contents to the terminal via the HTTPS protocol, and the terminal displays the information on the user interface.
[0121] Step 7: Check and correct the medical record
[0122] The user checks the medical record displayed on the terminal and makes any necessary corrections. The input is the medical record information displayed on the terminal, and the output is the medical record content corrected by the user. Specifically, the doctor changes the "pain level" on the terminal screen from "moderate" to "severe."
[0123] Step 8: Save the medical record
[0124] The user confirms the medical record contents after making the edits, and the terminal sends them back to the server. The server saves the received medical record contents in the patient's electronic medical record. The input is the edited medical record contents, and the output is the information saved in the electronic medical record. Specifically, when the user presses the "Save" button on the terminal, the terminal sends the edited contents to the server, and the server finally records the data "Date of onset: yesterday. Pain level: severe" in the electronic medical record.
[0125] Through these steps, this invention automates the process from voice input to final medical record storage, reducing the burden on medical professionals and achieving efficient and accurate information management.
[0126] (Application example 1)
[0127] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0128] In the traditional medical record creation process, medical professionals manually record patient information, consuming a great deal of time and effort. Human errors, such as clerical errors and omissions, are also common. Meanwhile, in autonomous vehicles, systems for responding quickly and appropriately to vehicle abnormalities or emergencies may be inadequate, potentially reducing passenger safety and vehicle operational efficiency. A system that solves these problems and improves automation and efficiency is needed.
[0129] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0130] In this invention, the server includes means for acquiring oral information from the patient as voice data, means for transmitting the voice data to a data processing device, means for converting the voice data into text data in the data processing device, means for analyzing the text data and automatically completing the medical record content, means for displaying the automatically completed medical record content and allowing the user to confirm and correct it, means for transmitting the confirmed and corrected content to the data processing device and saving it, means for collecting information on the vehicle condition and external environment from the conversation content of the occupants, and means for analyzing the collected information and proposing countermeasures if an abnormality is detected. This enables improved efficiency and accuracy in creating medical records in the medical field, and enables quick and appropriate responses to abnormalities and emergencies in self-driving vehicles.
[0131] "Patient" refers to anyone who uses a medical institution or healthcare system.
[0132] "Oral information" refers to information or content that is spoken aloud.
[0133] "Audio data" is data converted from audio into digital form.
[0134] A "data processing device" is a machine or system that processes voice data, converts it into text data, and analyzes it.
[0135] "Text data" refers to written information in digital form.
[0136] "Analysis" is the process of extracting meaning and keywords from text data and organizing the information.
[0137] A "medical record" is a document that records a patient's medical treatment at a medical institution.
[0138] "Auto-completion" is the process by which the system automatically fills in missing data based on retrieved information.
[0139] A "user" is someone who operates the system and reviews and modifies the results.
[0140] "Display" means the visual presentation of information through a data processing device or other output device.
[0141] "Occupant" refers to any person riding in an autonomous vehicle.
[0142] "Conversation content" refers to verbal exchanges between multiple people.
[0143] "Vehicle status" is information indicating the operating status of the vehicle and the operating status of each function.
[0144] The "external environment" refers to the surrounding circumstances and conditions in which the vehicle is traveling.
[0145] "Collection" is the act of gathering specific information.
[0146] An "abnormality" is an event that indicates a state or malfunction that is different from the normal state.
[0147] A "solution" is a method or means for dealing with a particular situation.
[0148] As an embodiment of the present invention, there is provided a system in which a voice recognition system is installed in an autonomous driving vehicle, collects information on the vehicle's condition and the external environment from the conversation of the occupants, and proposes countermeasures when an abnormality is detected. The server includes means for acquiring oral information from the patient as voice data, means for transmitting the voice data to a data processing device, means for converting the voice data into text data in the data processing device, means for analyzing the text data and automatically completing the medical record content, means for displaying the automatically completed medical record content and allowing the user to confirm and correct it, means for transmitting the confirmed and corrected content to the data processing device and saving it, means for collecting information on the vehicle's condition and the external environment from the conversation of the occupants, and means for analyzing the collected information and proposing countermeasures when an abnormality is detected.
[0149] Hardware and software used
[0150] The server processes the audio data using the following hardware and software:
[0151] Microphone: Microphones installed inside the vehicle are used to collect passenger conversations in real time.
[0152] Data processing device: Receives and processes voice data. Here, a cloud-based server plays a key role.
[0153] Speech recognition engine: Uses Google® Cloud Speech-to-Text API to convert voice data into text data.
[0154] Text analysis engine: Uses NLTK (Natural Language Toolkit) to analyze the collected text data and extract relevant keywords.
[0155] Database: AWS (registered trademark) RDS (Relational Database Service) will be used to store analysis results and medical record information.
[0156] Navigation API: Uses Google Maps API to get real-time traffic information and external environment data.
[0157] Data processing and calculation
[0158] 1. Voice collection: Microphones inside the vehicle collect the passengers' conversations as voice data.
[0159] 2. Speech Recognition: The collected voice data is sent to a data processing device and converted into text data using the Google Cloud Speech-to-Text API.
[0160] 3. Keyword Extraction: The text data is analyzed using NLTK to extract relevant keywords.
[0161] 4. Data matching and auto-completion: Based on the extracted keywords, they are matched with historical data stored in AWS RDS or data retrieved from the navigation API, and then auto-completion is performed.
[0162] 5. Display and correction: The auto-completed content is displayed on the vehicle's display, where the user can confirm and correct it.
[0163] 6. Storage: The final confirmed and corrected content is sent to the data processing device and stored in AWS RDS.
[0164] Specific examples
[0165] For example, if a passenger says, "The air conditioner is not working properly," the conversation is collected as voice data via a microphone. This voice data is then sent to a server, where a voice recognition engine converts it into text data such as "The air conditioner is not working properly." Next, a text analysis engine analyzes this text data and extracts keywords such as "air conditioner," "working," and "not good." Finally, the system checks the condition of the air conditioner, and if there is an abnormality, it suggests a solution such as "Air conditioner abnormality: Please check."
[0166] Prompt Sentence Examples
[0167] "As a voice-activated vehicle assistant, please analyze the following conversation and suggest an appropriate response. Conversation: 'The air conditioning isn't working. Please add a new restaurant to Maps.'"
[0168] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0169] Step 1:
[0170] Audio Collection
[0171] Subject: Terminal
[0172] How it works: A microphone installed in the vehicle collects passenger conversations in real time. For example, if a passenger says, "The air conditioning isn't working," the voice is collected as voice data through the microphone.
[0173] Input: Passenger voice
[0174] Output: Collected audio data
[0175] Step 2:
[0176] Sending audio data
[0177] Subject: Terminal
[0178] Specific operation: The collected voice data is sent to a data processing device (cloud server). This transmission is done in real time so that the voice data can be processed immediately.
[0179] Input: Collected audio data
[0180] Output: Audio data sent to the data processing device
[0181] Step 3:
[0182] Voice Recognition
[0183] Subject: Server
[0184] Specific operation: The data processing device uses the Google Cloud Speech-to-Text API to convert the received voice data into text data. For example, the voice saying "The air conditioner is not working" is converted into the text "The air conditioner is not working."
[0185] Input: Transmitted audio data
[0186] Output: Converted text data
[0187] Step 4:
[0188] Keyword extraction
[0189] Subject: Server
[0190] Specific operation: The server analyzes the converted text data using NLTK and extracts related keywords. For example, from the text "The air conditioner is not working," keywords such as "air conditioner" and "not working" are extracted.
[0191] Input: Converted text data
[0192] Output: Extracted keywords
[0193] Step 5:
[0194] Data matching and auto-completion
[0195] Subject: Server
[0196] Specific operation: Based on the extracted keywords, the server compares them with past data stored in AWS RDS and information obtained from the navigation API (Google Maps API) and performs auto-completion. For example, based on the keywords "air conditioner" and "not working," the server will automatically complete the search results by checking the air conditioner's status and suggesting a solution such as "Air conditioner malfunction: check."
[0197] Input: Extracted keywords
[0198] Output: Auto-completed solutions and information
[0199] Step 6:
[0200] View and Modify
[0201] Subject: Terminal
[0202] Specific operation: The supplemented information is displayed on the vehicle's display. The user (passenger) can check this information and make corrections as necessary. For example, the passenger can correct the information by saying, "I checked the air conditioning, and there is actually no problem."
[0203] Input: Auto-completed solutions and information
[0204] Output: Displayed completions and user corrections
[0205] Step 7:
[0206] Submitting and saving your modifications
[0207] Subject: Terminal
[0208] Specific operation: The content modified by the user is sent back to the data processing device and stored in AWS RDS, so that the latest modified information is reflected in the system and can be used as reference data in the future.
[0209] Input: User-modified content
[0210] Output: Saved modifications
[0211] In this way, the system achieves its objective by performing data processing and calculations based on the input data at each step and obtaining the final output.
[0212] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0213] As an embodiment of this invention, we provide a system that acquires oral information from conversations with patients in real time and automatically reflects it in medical records, as well as a system that recognizes the user's emotions by combining an emotion engine. This system combines voice recognition technology, an automatic medical record completion function based on text data, and an emotion recognition function to reduce the burden on medical professionals and improve work efficiency.
[0214] Program processing overview
[0215] Voice input
[0216] The user interviews the patient about their symptoms and medical history. The device collects this conversation in real time as audio data. For example, a medical professional might ask, "What symptoms do you have?" The patient might reply, "I have a headache, and it's been going on since yesterday." The device then records this conversation as audio data.
[0217] Voice Recognition
[0218] The recorded voice data is sent from the device to the server. The server uses a voice recognition engine to convert this voice data into text data. For example, the voice "I have a headache, and it's been going on since yesterday" is converted into text "I have a headache, and it's been going on since yesterday."
[0219] Auto-completion
[0220] The server analyzes the converted text and extracts keywords. Based on the extracted keywords, it references the patient's existing information and medical databases to automatically complete the information. For example, based on the keyword "headache," the information "Date of onset: yesterday, Pain level: moderate" is added.
[0221] emotion recognition
[0222] The voice data is also sent to the emotion engine. The server uses the emotion engine to analyze the user's emotions from the voice data. For example, the analysis may reveal that the user is feeling "acute anxiety." The emotion engine adds this to the text data and reflects it in the patient's medical record.
[0223] Check and fix
[0224] The automatically completed medical record contents and emotion analysis results are sent to the terminal and displayed to the user. The user can check the contents and make corrections as necessary. For example, a doctor may change the "pain level" from "moderate" to "severe." Emotional information is also checked, and appropriate responses for patients with strong anxiety are considered.
[0225] keep
[0226] Once the changes are confirmed, the device sends the revised medical record information and emotional information to the server. The server saves this as the patient's electronic medical record. For example, the following data is recorded in the electronic medical record: "I have a headache that has been going on since yesterday. Onset: yesterday. Pain level: severe. Emotion: acute anxiety."
[0227] This system automates everything from voice input to emotion recognition and final medical record storage. It not only enables medical professionals to efficiently and accurately record patient information, but also allows them to understand the patient's emotional state, enabling them to provide appropriate medical care. This invention is particularly useful in interactive medical environments that emphasize the patient's emotional state.
[0228] The processing flow will be explained below.
[0229] Step 1:
[0230] The user interviews the patient about their symptoms and medical history. For example, a medical professional asks, "What symptoms do you have?" The patient answers, "I have a headache, and it's been going on since yesterday."
[0231] Step 2:
[0232] The terminal collects the conversation between the user and the patient as voice data. The terminal starts recording and records the voice data in real time.
[0233] Step 3:
[0234] The device converts the recorded audio data into a digital format and sends it to the server, where it is compressed to optimize bandwidth.
[0235] Step 4:
[0236] The server analyzes the received voice data and converts it into text data using a voice recognition engine. For example, a voice saying "I have a headache, and it's been going on since yesterday" is converted into text "I have a headache, and it's been going on since yesterday."
[0237] Step 5:
[0238] The server analyzes the text data and extracts keywords. Based on the extracted keywords, the server references the patient's existing information and medical databases to automatically complete the information. For example, the keyword "headache" can be supplemented with the information "Date of onset: yesterday, Pain level: moderate."
[0239] Step 6:
[0240] The server also sends the voice data to the emotion engine.
[0241] Step 7:
[0242] The emotion engine on the server analyzes the user's emotions from the voice data. For example, the emotion analysis may reveal that the user is experiencing "acute anxiety."
[0243] Step 8:
[0244] The server adds the results of the emotion analysis to the text data and reflects it in the medical record. For example, information such as "Emotion: acute anxiety" is added.
[0245] Step 9:
[0246] The server then sends the automatically completed medical record information and the emotion analysis results to the terminal. For example, it might generate the following: "I have a headache that has been going on since yesterday. Onset: yesterday. Pain level: moderate. Emotion: acute anxiety."
[0247] Step 10:
[0248] The terminal displays the received medical record information and emotion information to the user, who can then check the displayed information and make corrections as necessary.
[0249] Step 11:
[0250] The user reviews the patient record information and makes corrections as needed, for example, changing the "pain level" from "moderate" to "severe." Emotional information is also reviewed, and measures are planned to address patients with high anxiety.
[0251] Step 12:
[0252] The terminal transmits the medical record information and emotion information corrected by the user to the server. The confirmed corrections are then transmitted to the server.
[0253] Step 13:
[0254] The server stores the corrected medical record information and emotion information received as the patient's electronic medical record. For example, information such as "I have a headache that has continued since yesterday. Onset: yesterday. Pain level: severe. Emotion: acute anxiety" is recorded in the electronic medical record.
[0255] Example 2
[0256] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0257] In current medical settings, it is difficult to accurately and quickly record a patient's oral information and simultaneously manage the patient's emotional state. When medical professionals record information to create medical records, work efficiency decreases and human error is likely to occur. In addition, additional effort is required to understand the patient's emotional state, making it difficult to improve the overall quality of medical care.
[0258] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0259] In this invention, the server includes means for acquiring oral information from the patient as voice data, means for transmitting the voice data to a central computer, means for converting the voice data to text data in the central computer, means for analyzing the text data and automatically completing the medical record content, means for analyzing the emotional state from the voice data in the central computer, means for displaying the automatically completed medical record content and the emotion analysis results so that the user can confirm and correct them, and means for transmitting the confirmed and corrected content to the central computer and saving it. This automates the entire process from voice input to text conversion, emotion analysis, and medical record updating and saving, reducing the burden on medical professionals and enabling them to understand the patient's emotional state, enabling appropriate medical care to be provided.
[0260] "Oral information" refers to verbal information such as symptoms and medical history that a patient tells a medical professional.
[0261] "Audio data" refers to audio files that digitally record oral information obtained from a patient.
[0262] The "central computer" is a server that receives, analyzes, and stores voice and text data sent from the terminal.
[0263] "Speech recognition software" is a program for converting voice data into text data.
[0264] "Text data" is data in which voice data is expressed as a string of characters.
[0265] A "medical record" is a medical document that details a patient's symptoms, medical history, and treatment.
[0266] "Automatic completion" is a process that automatically adds missing parts based on converted text data by referencing past data and databases.
[0267] "Emotion analysis" is the process of analyzing the patient's emotional state from the audio data and adding the results to the text data.
[0268] "User" refers to a medical professional who uses the system to review and modify patient information.
[0269] To implement this invention, voice input technology, voice recognition technology, text data analysis technology, emotion recognition technology, and record management technology are required. This system captures oral information from conversations with patients in real time and automatically reflects it in medical records. It also simultaneously analyzes the patient's emotional state and adds it to the medical record. In this way, it reduces the workload of medical professionals and enables them to manage patient information efficiently and accurately.
[0270] Collecting voice input
[0271] The device has a built-in microphone that collects conversations between patients and medical professionals in real time. For example, a medical professional might ask, "What symptoms do you have?" and the patient might reply, "I have a headache, and it's been going on since yesterday." This conversation is recorded by the device's microphone and saved as audio data.
[0272] Sending audio data
[0273] The device sends the collected voice data to a central computer (server) over the Internet, for example, securely using the HTTPS protocol, where it can be processed for speech recognition and emotion engines.
[0274] Converting audio data to text
[0275] The server uses speech recognition software such as Google Cloud Speech-to-Text to convert the transmitted voice data into text data. For example, the voice data "I have a headache, and it's been going on since yesterday" is converted into text data "I have a headache, and it's been going on since yesterday."
[0276] Text data analysis and auto-completion
[0277] The server analyzes the converted text data and extracts keywords. For example, it recognizes the keyword "headache." It then references the patient's existing information and medical databases to automatically supplement the information with relevant information. For example, information such as "Date of onset: yesterday" and "Pain level: moderate" is added.
[0278] Conducting sentiment analysis
[0279] The server uses emotion recognition software such as IBM Watson® Tone Analyzer to analyze the patient's emotional state from the voice data. For example, the server obtains emotional information such as "acute anxiety" as the analysis result and adds this information to the text data.
[0280] Check and fix
[0281] The device displays the automatically completed medical record and emotion analysis results to the user. The user can review the information and make corrections as necessary. For example, a doctor can change the "pain level" from "moderate" to "severe," and then review the emotion information to consider the necessary response.
[0282] Save the final data
[0283] Once the changes are confirmed, the device sends the revised medical record information and emotion information back to the server. The server saves this as the patient's electronic medical record. For example, the final data recorded would be "I have a headache that has been going on since yesterday. Onset: yesterday. Pain level: severe. Emotion: acute anxiety."
[0284] This system automates the entire process from voice input to text conversion, emotion analysis, and medical record updating and storage, reducing the burden on medical professionals and enabling efficient and accurate information management. It also makes it possible to grasp the patient's emotional state, enabling more appropriate medical care.
[0285] Specific examples and examples of prompts for generative AI models
[0286] 1. Conversation between healthcare providers and patients:
[0287] Medical professional: "What symptoms do you have?"
[0288] Patient: "I have a headache that has been going on since yesterday."
[0289] 2. Audio to text conversion:
[0290] Voice: "I have a headache that has been going on since yesterday."
[0291] Text: "I have a headache that has been going on since yesterday."
[0292] 3. Auto-completion:
[0293] Keyword: "headache"
[0294] Supplementary data: "Date of onset: yesterday" "Pain level: moderate"
[0295] 4. Emotion analysis:
[0296] Speech data → Sentiment analysis : As in, “Acute anxiety .”
[0297] 5. Modification:
[0298] User modification: "Pain level: Moderate" → "Severe"
[0299] 6. Save:
[0300] Medical record data: "Headache, ongoing since yesterday. Onset: yesterday. Pain level: severe. Emotions: acute anxiety."
[0301] Example prompts for generative AI models
[0302] "Write a program that generates text data from audio data, extracts keywords for auto-completion, and adds sentiment analysis using an emotion engine. Also, include a process to finally save this data as an electronic medical record."
[0303] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0304] Step 1:
[0305] Collecting voice input
[0306] The device has a built-in microphone that collects conversations between patients and medical professionals in real time. Specifically, the medical professional asks, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday." While this conversation is taking place, the device records audio data. The input is the audio of the conversation, and the output is digital audio data.
[0307] Step 2:
[0308] Sending voice data to the server
[0309] The device transmits the collected voice data to a central computer (server) via the Internet. For example, the data is transmitted securely using the HTTPS protocol. This allows it to be processed for speech recognition and emotion engines. The input is digital voice data, and the output is the voice data transmitted to the server.
[0310] Step 3:
[0311] Converting audio data to text
[0312] The server uses speech recognition software such as Google Cloud Speech-to-Text to convert the transmitted voice data into text data. For example, the voice data "I have a headache, and it's been going on since yesterday" is converted into text data "I have a headache, and it's been going on since yesterday." The input is the voice data received by the server, and the output is the conversation in text format.
[0313] Step 4:
[0314] Text data analysis and auto-completion
[0315] The server analyzes the converted text data and extracts keywords. For example, it recognizes the keyword "headache." It then references the patient's existing information and medical databases to automatically complete the relevant information. For example, information such as "onset of illness: yesterday" and "pain level: moderate" is added. The input is text data, and the output is the completed text data.
[0316] Step 5:
[0317] Conducting sentiment analysis
[0318] The server uses emotion recognition software such as IBM Watson Tone Analyzer to analyze the patient's emotional state from the voice data. For example, the analysis results in emotional information such as "acute anxiety" and adds this to the text data. The input is voice data, and the output is text data with emotion analysis added.
[0319] Step 6:
[0320] Check and fix
[0321] The device displays the automatically completed medical record content and the emotion analysis results to the user. The user checks this content and makes corrections as necessary. For example, a doctor may change the "pain level" from "moderate" to "strong," check the emotion information, and consider the necessary response. The input is the completed text data and the emotion analysis results, and the output is the corrected medical record data.
[0322] Step 7:
[0323] Save the final data
[0324] Once the corrections are confirmed, the device sends the corrected medical record content and emotion information back to the server. The server saves this as the patient's electronic medical record. For example, the final data recorded is "I have a headache that has been going on since yesterday. Onset date: yesterday. Pain level: severe. Emotion: acute anxiety." The input is the corrected medical record data, and the output is the final data saved in the patient's electronic medical record.
[0325] (Application example 2)
[0326] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0327] Conventional food delivery services face the problem of difficulty in analyzing customer sentiment when taking orders, making it difficult to improve service quality. In particular, when customers are in a hurry or have special requests, they often cannot make prompt and appropriate suggestions, which can lead to a decline in customer satisfaction.
[0328] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0329] In this invention, the server includes means for acquiring dictation information from a customer as voice data, means for transmitting the voice data to the server, means for converting the voice data to text data in the server, means for analyzing the text data and automatically completing the order details, means for displaying the automatically completed order details and emotion analysis results for confirmation and correction by the user, means for transmitting the confirmed and corrected details to the server and saving them, means for analyzing the user's emotions from the voice data, and means for making appropriate suggestions based on the emotion analysis results, thereby enabling quick and accurate order confirmation and suggestions that take customer emotions into consideration.
[0330] "Patient" means an individual receiving medical services.
[0331] "Dicted information" is information provided by the user through speech.
[0332] "Audio data" refers to data in which audio is recorded in digital format.
[0333] A "server" is a computer system that processes and stores data.
[0334] "Text data" refers to data obtained by converting voice data into character information.
[0335] A "medical record" is a digital or paper-based document for recording medical information.
[0336] "Auto-completion" is the process of automatically adding missing data or content based on existing information.
[0337] A "displaying means" is a device or software that presents digital information to a user.
[0338] "Verify and correct" is the process by which a user checks the accuracy of information and corrects it if necessary.
[0339] "Emotion analysis" is a process of analyzing emotional states from voice data.
[0340] The "means for making suggestions" is a means for providing appropriate advice or recommendations to the user based on the analyzed information.
[0341] A "digital format" is a way of representing information in electronic form.
[0342] A "voice recognition engine" is software or hardware for converting voice data into text data.
[0343] This invention is a system that acquires oral information from patients or customers (hereinafter referred to as users) in real time, analyzes and supplements that information, and provides appropriate services quickly and accurately. This system combines technologies such as voice recognition technology, emotion analysis engine, text analysis, and recommendation engine.
[0344] Hardware and software used
[0345] 1. Hardware:
[0346] Devices such as smartphones, smart glasses, and head-mounted displays
[0347] Cloud Server
[0348] 2. Software:
[0349] Speech recognition engine: Google Speech-to-Text API, Amazon Transcribe, etc.
[0350] Emotion recognition engine: IBM Watson Tone Analyzer, Microsoft® Azure® Emotion API, etc.
[0351] Backend: Node.js, Python (Flask)
[0352] Database: MongoDB, Firebase
[0353] Program processing overview
[0354] Voice input
[0355] The user dictates information into the terminal. For example, in a food delivery service, the user might say, "I'd like to order a hamburger set, and a Coke to drink."
[0356] Voice Recognition
[0357] The voice data collected by the device is sent to a cloud server. The server uses a voice recognition engine to convert this voice data into text data. For example, a voice saying "I'd like to order a hamburger set, and a Coke to drink" is converted into the text "Hamburger set, Coke."
[0358] Auto-completion
[0359] The server analyzes the converted text and automatically completes the order details. For example, it generates an order list such as "Hamburger Set (Contents: Cheeseburger, Fries) + Coke." This automatically adds detailed information according to the user's request.
[0360] emotion recognition
[0361] The voice data is also sent to an emotion analysis engine, which analyzes the user's emotions. For example, the engine may detect that the user is in a hurry. Based on this, a menu of options for responding quickly is presented.
[0362] Check and fix
[0363] The auto-completed order details and the sentiment analysis results are displayed on the device, and the user can confirm and modify them. For example, the user can confirm an order displayed as "Hamburger Set (Cheeseburger, Fries) + Coke." At this time, the user can make modifications as necessary.
[0364] keep
[0365] The final order details and emotion data after confirmation and correction are sent to the server and stored in the database. For example, data such as "Hamburger set (cheeseburger, fries) + Coke. Emotion: In a hurry" is stored.
[0366] Examples of concrete examples and prompts
[0367] For example, suppose a user inputs dictation information such as "This hamburger looks very tasty. I'd like some fries with it, please."
[0368] Speech recognition prompt: transcribe audio
[0369] Emotion recognition prompt: analyze sentiment from text: "This burger looks delicious. I'd like some fries with it, please."
[0370] Recommendation prompt: get recommendations for quick service items based on "interest"
[0371] By implementing this invention, users can enjoy a fast and accurate ordering experience through voice input, and service providers can respond appropriately based on the user's emotions.
[0372] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0373] Step 1:
[0374] The user inputs dictation information into the terminal.
[0375] Specifically, the user speaks to a device such as a smartphone or smart glasses, saying, "I'd like to order a hamburger set, and a Coke to drink."
[0376] The input data is stored in the terminal as voice data.
[0377] The data that is output is the collected voice data.
[0378] Step 2:
[0379] The device sends the collected voice data to a cloud server.
[0380] Specifically, the terminal program uploads the collected voice data to a designated cloud server via the Internet.
[0381] The input data is audio data.
[0382] The output data is the audio data sent to the cloud server.
[0383] Step 3:
[0384] The server uses a speech recognition engine to convert the voice data into text data.
[0385] Specifically, the server calls a speech recognition engine such as the Google Speech-to-Text API or Amazon Transcribe to convert the voice data into text data.
[0386] The input data is audio data.
[0387] The output data is the converted text data.
[0388] Step 4:
[0389] The server analyzes the converted text data and automatically completes the order details.
[0390] Specifically, the server program analyzes the text data, extracts and completes the order details. For example, from the text data "hamburger set, cola," it generates "hamburger set (cheeseburger, fries) + cola."
[0391] The input data is converted text data.
[0392] The output data is the supplemented order details.
[0393] Step 5:
[0394] The voice data is sent to an emotion analysis engine to analyze the user's emotions.
[0395] Specifically, the server uses an emotion analysis engine such as IBM Watson Tone Analyzer or Microsoft Azure Emotion API to analyze the emotional state from the voice data.
[0396] The input data is audio data.
[0397] The output data is the emotion analysis results.
[0398] Step 6:
[0399] The automatically completed order details and sentiment analysis results are displayed on the terminal.
[0400] Specifically, the server sends the generated order details and the emotion analysis results to the terminal, which then displays them to the user. For example, the terminal displays "Hamburger set (cheeseburger, fries) + Coke. Emotion: In a hurry."
[0401] The input data is the supplemented order details and the results of sentiment analysis.
[0402] The output data is the information that is displayed on the terminal.
[0403] Step 7:
[0404] The user reviews and corrects what is displayed.
[0405] Specifically, the user checks the displayed order details and emotion information and makes corrections as necessary. For example, the user may make a correction such as "change fries to large size."
[0406] The data to be input is the information displayed on the terminal.
[0407] The data that is output is the content that has been confirmed and corrected by the user.
[0408] Step 8:
[0409] The confirmed and corrected content is sent to the server and stored in the database.
[0410] Specifically, the terminal sends the confirmed and corrected content back to the cloud server, and the server stores it in the database.
[0411] The data to be input is the content that has been confirmed and corrected by the user.
[0412] The output data is the order details and emotion information stored in the database.
[0413] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0414] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0415] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0416] [Second embodiment]
[0417] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0418] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0419] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0420] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0421] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0422] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0423] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0424] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0425] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0426] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0427] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0428] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0429] As an embodiment of this invention, we provide a system that acquires oral information from conversations with patients in real time and automatically reflects it in medical records. The system combines voice recognition technology with an automatic medical record completion function based on text data, reducing the burden on medical professionals and improving work efficiency.
[0430] Program processing overview
[0431] Voice input
[0432] The user interviews the patient about their symptoms and medical history. The device collects this conversation in real time as audio data. For example, a medical professional might ask the patient, "What symptoms do you have?" and the patient might reply, "I have a headache, and it's been going on since yesterday." The device records this conversation as audio data.
[0433] Voice Recognition
[0434] The recorded voice data is sent from the device to the server. The server uses a voice recognition engine to convert this voice data into text data. For example, the voice "I have a headache, and it has been going on since yesterday" is converted into the text "I have a headache, and it has been going on since yesterday."
[0435] Auto-completion
[0436] The server analyzes the converted text and extracts relevant keywords. Based on the extracted keywords, the server automatically completes the medical record contents using the patient's existing data and general medical knowledge. For example, based on the keyword "headache," information such as "Date of onset: yesterday. Pain level: moderate" is added.
[0437] Check and fix
[0438] The automatically completed medical record information is sent to the terminal and displayed to the user. The user can check the information and make corrections as necessary. For example, a doctor can change the "pain level" from "moderate" to "severe."
[0439] keep
[0440] Once the changes are confirmed, the device sends the revised medical record information to the server. The server saves the received data in the patient's electronic medical record. As a result, for example, data such as "I have a headache that has been going on since yesterday. Onset date: yesterday. Pain level: severe" is recorded in the electronic medical record.
[0441] This system automates the entire process from voice input to final medical record storage, enabling medical professionals to efficiently and accurately record patient information. This invention is particularly useful when doctors and nurses need to quickly and accurately create medical records based on their conversations with patients.
[0442] The processing flow will be explained below.
[0443] Step 1:
[0444] The user interviews the patient about their symptoms and medical history. For example, a medical professional asks, "What symptoms do you have?" The patient answers, "I have a headache, and it's been going on since yesterday."
[0445] Step 2:
[0446] The terminal collects the conversation between the user and the patient as voice data. The terminal starts recording and records the voice data in real time.
[0447] Step 3:
[0448] The device converts the recorded audio data into a digital format and sends it to the server, where it is compressed to optimize bandwidth.
[0449] Step 4:
[0450] The server analyzes the received voice data and converts it into text data using a voice recognition engine. For example, a voice saying "I have a headache, and it's been going on since yesterday" is converted into text "I have a headache, and it's been going on since yesterday."
[0451] Step 5:
[0452] The server analyzes the text data and extracts keywords. Based on the extracted keywords, the server references the patient's existing information and medical databases to automatically complete the information. For example, the keyword "headache" can be supplemented with the information "Date of onset: yesterday, Pain level: moderate."
[0453] Step 6:
[0454] The server generates automatically completed medical record information and sends it to the terminal. For example, medical record data such as "I have a headache that has continued since yesterday. Onset: yesterday. Pain level: moderate" is generated.
[0455] Step 7:
[0456] The terminal displays the received medical record information to the user, who then checks the displayed information and makes any necessary corrections.
[0457] Step 8:
[0458] The user checks the medical record information and makes corrections as necessary, for example, changing the "pain level" from "moderate" to "severe."
[0459] Step 9:
[0460] The terminal sends the medical record information corrected by the user to the server, and after final confirmation, the data is sent.
[0461] Step 10:
[0462] The server stores the received corrected medical record information as the patient's electronic medical record. For example, information such as "I have a headache that has continued since yesterday. Onset: yesterday. Pain level: severe" is recorded in the electronic medical record.
[0463] Example 1
[0464] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0465] In conventional medical settings, creating medical records based on conversations with patients requires a great deal of time and effort. Furthermore, manual data entry often carries the risk of errors and omissions, increasing the burden on medical professionals. The purpose of this invention is to provide a system that reduces the burden on medical professionals, improves work efficiency, and prevents errors and omissions.
[0466] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0467] In this invention, the server includes a means for converting voice data into text data using a voice recognition engine, a means for analyzing the text data and extracting related keywords, and a means for automatically completing the medical record content based on the keywords. This automates the process from voice input to medical record creation, reducing the burden on medical professionals and improving work efficiency and data accuracy.
[0468] "Oral information from patients" refers to audio information such as symptoms and medical history obtained directly from patients by medical professionals.
[0469] "Audio data" refers to digital audio files of recorded conversations with patients.
[0470] "Terminal" refers to an electronic device for collecting and transmitting voice data to a server.
[0471] "Server" refers to the computer system that processes voice data and automatically completes medical records.
[0472] "Speech recognition engine" refers to software or algorithms that analyze voice data and convert it into text data.
[0473] "Text data" refers to character information converted by a voice recognition engine.
[0474] "Natural language processing technology" refers to computer technology for analyzing text data and extracting keywords.
[0475] "Keywords" refer to important words and phrases extracted from text data that are necessary for automatically completing the contents of medical records.
[0476] A "medical record" refers to a medical record that describes a patient's symptoms, medical history, treatment details, etc.
[0477] "Automatic completion" refers to the process in which the system automatically adds and supplements medical record content based on keywords.
[0478] "User" refers to a medical professional who operates the system to check and correct medical records.
[0479] As an embodiment of this invention, we provide a system that acquires oral information from conversations with patients in real time and automatically reflects it in medical records. The system combines voice recognition technology with an automatic medical record completion function based on text data, reducing the burden on medical professionals and improving work efficiency.
[0480] Collecting voice input
[0481] Users, or medical professionals, use devices such as smartphones and tablets to obtain voice data of symptoms and medical history directly from patients. The devices are equipped with highly sensitive microphones that record conversations in real time. For example, if a medical professional asks, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday," this content is collected as voice data on the device.
[0482] Sending audio data
[0483] The device sends the recorded audio data to a server over the Internet using the HTTPS protocol to ensure data security. A program on the device converts the audio data into a digital format and sends it to the server.
[0484] Speech Recognition Processing
[0485] The server passes the received voice data to a voice recognition engine, which converts the voice into text data. A typical voice recognition engine used here is a third-party voice recognition service. For example, voice data such as "I have a headache, and it's been going on since yesterday" is converted directly into text data such as "I have a headache, and it's been going on since yesterday."
[0486] Text analysis and keyword extraction
[0487] The server receives the converted text data and analyzes it using natural language processing technology. Related keywords are extracted through the analysis. For example, the keywords "headache," "yesterday," and "continuing" are extracted from the sentence "I have a headache, and it has been going on since yesterday."
[0488] Auto-completion
[0489] The server automatically completes the medical record based on the extracted keywords. It references the medical database and adds information related to the keywords to the medical record. For example, based on the keyword "headache," the information "Date of onset: yesterday. Pain level: moderate" is automatically added to the medical record.
[0490] Sending and checking medical record contents
[0491] The completed medical record information is sent to the terminal and displayed to the user. The user can check the medical record information and make corrections as necessary. For example, a doctor may change the "pain level" from "moderate" to "severe."
[0492] Saving medical record contents
[0493] Once the user has confirmed the medical record contents after making the edits, the terminal sends them back to the server. The server then saves the received medical record contents in the patient's electronic medical record. For example, information such as "I have a headache that has been going on since yesterday. Onset date: yesterday. Pain level: severe" is recorded in the electronic medical record.
[0494] Examples of concrete examples and prompts
[0495] Example: A healthcare professional asks a patient, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday." The audio is recorded in real time and converted to text. Auto-complete adds, "Date of onset: yesterday. Pain level: moderate." The doctor then corrects "moderate" to "severe." The information is saved in the electronic medical record.
[0496] Example prompt: "Please run a program that will automatically record a conversation about your headache symptoms that have continued since yesterday into the electronic medical record."
[0497] In this way, the system automates the process from voice input to final medical record storage, enabling medical professionals to efficiently and accurately record patient information.
[0498] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0499] Step 1: Collecting voice input
[0500] The user listens to the patient's symptoms and medical history and engages in a conversation. The device uses a built-in microphone to collect this audio in real time and records it as audio data. The input is the conversation between the user and the patient, and the output is digital audio data. Specifically, the medical professional asks, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday." This audio is collected by the device.
[0501] Step 2: Sending audio data
[0502] The audio data collected by the device is sent to a server via the Internet. The input is digital audio data, and the output is audio data transferred to the server. Specifically, a program on the device sends the audio data to the server using the HTTPS protocol. This ensures that the data arrives securely at the server.
[0503] Step 3: Speech recognition processing
[0504] The server passes the received voice data to a voice recognition engine, which converts the voice into text data. The input is the voice data sent to the server, and the output is text data. Specifically, the server calls a voice recognition engine (for example, a natural language processing library) and converts the voice data "I have a headache, and it's been going on since yesterday" into the text "I have a headache, and it's been going on since yesterday."
[0505] Step 4: Text analysis and keyword extraction
[0506] The server analyzes the converted text data and extracts related keywords. The input is text data obtained from the speech recognition engine, and the output is the extracted keywords. Specifically, the server uses natural language processing technology to extract keywords such as "headache," "yesterday," and "continuing" from the text "I have a headache, and it has been going on since yesterday."
[0507] Step 5: Auto-completion
[0508] The server automatically completes the medical record contents based on the extracted keywords. The input is the extracted keywords, and the output is the automatically completed medical record contents. Specifically, the server references the medical database and adds information such as "Date of onset: yesterday. Pain level: moderate" to the medical record based on the keywords "headache," "yesterday," and "continuing."
[0509] Step 6: Send and view medical records
[0510] The server sends the auto-completed medical record contents to the terminal and displays them to the user. The input is the auto-completed medical record contents, and the output is the medical record information displayed on the terminal. Specifically, the server sends the medical record contents to the terminal via the HTTPS protocol, and the terminal displays the information on the user interface.
[0511] Step 7: Check and correct the medical record
[0512] The user checks the medical record displayed on the terminal and makes any necessary corrections. The input is the medical record information displayed on the terminal, and the output is the medical record content corrected by the user. Specifically, the doctor changes the "pain level" on the terminal screen from "moderate" to "severe."
[0513] Step 8: Save the medical record
[0514] The user confirms the medical record contents after making the edits, and the terminal sends them back to the server. The server saves the received medical record contents in the patient's electronic medical record. The input is the edited medical record contents, and the output is the information saved in the electronic medical record. Specifically, when the user presses the "Save" button on the terminal, the terminal sends the edited contents to the server, and the server finally records the data "Date of onset: yesterday. Pain level: severe" in the electronic medical record.
[0515] Through these steps, this invention automates the process from voice input to final medical record storage, reducing the burden on medical professionals and achieving efficient and accurate information management.
[0516] (Application example 1)
[0517] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0518] In the traditional medical record creation process, medical professionals manually record patient information, consuming a great deal of time and effort. Human errors, such as clerical errors and omissions, are also common. Meanwhile, in autonomous vehicles, systems for responding quickly and appropriately to vehicle abnormalities or emergencies may be inadequate, potentially reducing passenger safety and vehicle operational efficiency. A system that solves these problems and improves automation and efficiency is needed.
[0519] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0520] In this invention, the server includes means for acquiring oral information from the patient as voice data, means for transmitting the voice data to a data processing device, means for converting the voice data into text data in the data processing device, means for analyzing the text data and automatically completing the medical record content, means for displaying the automatically completed medical record content and allowing the user to confirm and correct it, means for transmitting the confirmed and corrected content to the data processing device and saving it, means for collecting information on the vehicle condition and external environment from the conversation content of the occupants, and means for analyzing the collected information and proposing countermeasures if an abnormality is detected. This enables improved efficiency and accuracy in creating medical records in the medical field, and enables quick and appropriate responses to abnormalities and emergencies in self-driving vehicles.
[0521] "Patient" refers to anyone who uses a medical institution or healthcare system.
[0522] "Oral information" refers to information or content that is spoken aloud.
[0523] "Audio data" is data converted from audio into digital form.
[0524] A "data processing device" is a machine or system that processes voice data, converts it into text data, and analyzes it.
[0525] "Text data" refers to written information in digital form.
[0526] "Analysis" is the process of extracting meaning and keywords from text data and organizing the information.
[0527] A "medical record" is a document that records a patient's medical treatment at a medical institution.
[0528] "Auto-completion" is the process by which the system automatically fills in missing data based on retrieved information.
[0529] A "user" is someone who operates the system and reviews and modifies the results.
[0530] "Display" means the visual presentation of information through a data processing device or other output device.
[0531] "Occupant" refers to any person riding in an autonomous vehicle.
[0532] "Conversation content" refers to verbal exchanges between multiple people.
[0533] "Vehicle status" is information indicating the operating status of the vehicle and the operating status of each function.
[0534] The "external environment" refers to the surrounding circumstances and conditions in which the vehicle is traveling.
[0535] "Collection" is the act of gathering specific information.
[0536] An "abnormality" is an event that indicates a state or malfunction that is different from the normal state.
[0537] A "solution" is a method or means for dealing with a particular situation.
[0538] As an embodiment of the present invention, there is provided a system in which a voice recognition system is installed in an autonomous driving vehicle, collects information on the vehicle's condition and the external environment from the conversation of the occupants, and proposes countermeasures when an abnormality is detected. The server includes means for acquiring oral information from the patient as voice data, means for transmitting the voice data to a data processing device, means for converting the voice data into text data in the data processing device, means for analyzing the text data and automatically completing the medical record content, means for displaying the automatically completed medical record content and allowing the user to confirm and correct it, means for transmitting the confirmed and corrected content to the data processing device and saving it, means for collecting information on the vehicle's condition and the external environment from the conversation of the occupants, and means for analyzing the collected information and proposing countermeasures when an abnormality is detected.
[0539] Hardware and software used
[0540] The server processes the audio data using the following hardware and software:
[0541] Microphone: Microphones installed inside the vehicle are used to collect passenger conversations in real time.
[0542] Data processing device: Receives and processes voice data. Here, a cloud-based server plays a key role.
[0543] Speech recognition engine: Uses the Google Cloud Speech-to-Text API to convert voice data into text data.
[0544] Text analysis engine: Uses NLTK (Natural Language Toolkit) to analyze the collected text data and extract relevant keywords.
[0545] Database: AWS RDS (Relational Database Service) is used to store analysis results and medical record information.
[0546] Navigation API: Uses Google Maps API to get real-time traffic information and external environment data.
[0547] Data processing and calculation
[0548] 1. Voice collection: Microphones inside the vehicle collect the passengers' conversations as voice data.
[0549] 2. Speech Recognition: The collected voice data is sent to a data processing device and converted into text data using the Google Cloud Speech-to-Text API.
[0550] 3. Keyword Extraction: The text data is analyzed using NLTK to extract relevant keywords.
[0551] 4. Data matching and auto-completion: Based on the extracted keywords, they are matched with historical data stored in AWS RDS or data retrieved from the navigation API, and then auto-completion is performed.
[0552] 5. Display and correction: The auto-completed content is displayed on the vehicle's display, where the user can confirm and correct it.
[0553] 6. Storage: The final confirmed and corrected content is sent to the data processing device and stored in AWS RDS.
[0554] Specific examples
[0555] For example, if a passenger says, "The air conditioner is not working properly," the conversation is collected as voice data via a microphone. This voice data is then sent to a server, where a voice recognition engine converts it into text data such as "The air conditioner is not working properly." Next, a text analysis engine analyzes this text data and extracts keywords such as "air conditioner," "working," and "not good." Finally, the system checks the condition of the air conditioner, and if there is an abnormality, it suggests a solution such as "Air conditioner abnormality: Please check."
[0556] Prompt Sentence Examples
[0557] "As a voice-activated vehicle assistant, please analyze the following conversation and suggest an appropriate response. Conversation: 'The air conditioning isn't working. Please add a new restaurant to Maps.'"
[0558] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0559] Step 1:
[0560] Audio Collection
[0561] Subject: Terminal
[0562] How it works: A microphone installed in the vehicle collects passenger conversations in real time. For example, if a passenger says, "The air conditioning isn't working," the voice is collected as voice data through the microphone.
[0563] Input: Passenger voice
[0564] Output: Collected audio data
[0565] Step 2:
[0566] Sending audio data
[0567] Subject: Terminal
[0568] Specific operation: The collected voice data is sent to a data processing device (cloud server). This transmission is done in real time so that the voice data can be processed immediately.
[0569] Input: Collected audio data
[0570] Output: Audio data sent to the data processing device
[0571] Step 3:
[0572] Voice Recognition
[0573] Subject: Server
[0574] Specific operation: The data processing device uses the Google Cloud Speech-to-Text API to convert the received voice data into text data. For example, the voice saying "The air conditioner is not working" is converted into the text "The air conditioner is not working."
[0575] Input: Transmitted audio data
[0576] Output: Converted text data
[0577] Step 4:
[0578] Keyword extraction
[0579] Subject: Server
[0580] Specific operation: The server analyzes the converted text data using NLTK and extracts related keywords. For example, from the text "The air conditioner is not working," keywords such as "air conditioner" and "not working" are extracted.
[0581] Input: Converted text data
[0582] Output: Extracted keywords
[0583] Step 5:
[0584] Data matching and auto-completion
[0585] Subject: Server
[0586] Specific operation: Based on the extracted keywords, the server compares them with past data stored in AWS RDS and information obtained from the navigation API (Google Maps API) and performs auto-completion. For example, based on the keywords "air conditioner" and "not working," the server will automatically complete the search results by checking the air conditioner's status and suggesting a solution such as "Air conditioner malfunction: check."
[0587] Input: Extracted keywords
[0588] Output: Auto-completed solutions and information
[0589] Step 6:
[0590] View and Modify
[0591] Subject: Terminal
[0592] Specific operation: The supplemented information is displayed on the vehicle's display. The user (passenger) can check this information and make corrections as necessary. For example, the passenger can correct the information by saying, "I checked the air conditioning, and there is actually no problem."
[0593] Input: Auto-completed solutions and information
[0594] Output: Displayed completions and user corrections
[0595] Step 7:
[0596] Submitting and saving your modifications
[0597] Subject: Terminal
[0598] Specific operation: The content modified by the user is sent back to the data processing device and stored in AWS RDS, so that the latest modified information is reflected in the system and can be used as reference data in the future.
[0599] Input: User-modified content
[0600] Output: Saved modifications
[0601] In this way, the system achieves its objective by performing data processing and calculations based on the input data at each step and obtaining the final output.
[0602] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0603] As an embodiment of this invention, we provide a system that acquires oral information from conversations with patients in real time and automatically reflects it in medical records, as well as a system that recognizes the user's emotions by combining an emotion engine. This system combines voice recognition technology, an automatic medical record completion function based on text data, and an emotion recognition function to reduce the burden on medical professionals and improve work efficiency.
[0604] Program processing overview
[0605] Voice input
[0606] The user interviews the patient about their symptoms and medical history. The device collects this conversation in real time as audio data. For example, a medical professional might ask, "What symptoms do you have?" The patient might reply, "I have a headache, and it's been going on since yesterday." The device then records this conversation as audio data.
[0607] Voice Recognition
[0608] The recorded voice data is sent from the device to the server. The server uses a voice recognition engine to convert this voice data into text data. For example, the voice "I have a headache, and it's been going on since yesterday" is converted into text "I have a headache, and it's been going on since yesterday."
[0609] Auto-completion
[0610] The server analyzes the converted text and extracts keywords. Based on the extracted keywords, it references the patient's existing information and medical databases to automatically complete the information. For example, based on the keyword "headache," the information "Date of onset: yesterday, Pain level: moderate" is added.
[0611] emotion recognition
[0612] The voice data is also sent to the emotion engine. The server uses the emotion engine to analyze the user's emotions from the voice data. For example, the analysis may reveal that the user is feeling "acute anxiety." The emotion engine adds this to the text data and reflects it in the patient's medical record.
[0613] Check and fix
[0614] The automatically completed medical record contents and emotion analysis results are sent to the terminal and displayed to the user. The user can check the contents and make corrections as necessary. For example, a doctor may change the "pain level" from "moderate" to "severe." Emotional information is also checked, and appropriate responses for patients with strong anxiety are considered.
[0615] keep
[0616] Once the changes are confirmed, the device sends the revised medical record information and emotional information to the server. The server saves this as the patient's electronic medical record. For example, the following data is recorded in the electronic medical record: "I have a headache that has been going on since yesterday. Onset: yesterday. Pain level: severe. Emotion: acute anxiety."
[0617] This system automates everything from voice input to emotion recognition and final medical record storage. It not only enables medical professionals to efficiently and accurately record patient information, but also allows them to understand the patient's emotional state, enabling them to provide appropriate medical care. This invention is particularly useful in interactive medical environments that emphasize the patient's emotional state.
[0618] The processing flow will be explained below.
[0619] Step 1:
[0620] The user interviews the patient about their symptoms and medical history. For example, a medical professional asks, "What symptoms do you have?" The patient answers, "I have a headache, and it's been going on since yesterday."
[0621] Step 2:
[0622] The terminal collects the conversation between the user and the patient as voice data. The terminal starts recording and records the voice data in real time.
[0623] Step 3:
[0624] The device converts the recorded audio data into a digital format and sends it to the server, where it is compressed to optimize bandwidth.
[0625] Step 4:
[0626] The server analyzes the received voice data and converts it into text data using a voice recognition engine. For example, a voice saying "I have a headache, and it's been going on since yesterday" is converted into text "I have a headache, and it's been going on since yesterday."
[0627] Step 5:
[0628] The server analyzes the text data and extracts keywords. Based on the extracted keywords, the server references the patient's existing information and medical databases to automatically complete the information. For example, the keyword "headache" can be supplemented with the information "Date of onset: yesterday, Pain level: moderate."
[0629] Step 6:
[0630] The server also sends the voice data to the emotion engine.
[0631] Step 7:
[0632] The emotion engine on the server analyzes the user's emotions from the voice data. For example, the emotion analysis may reveal that the user is experiencing "acute anxiety."
[0633] Step 8:
[0634] The server adds the results of the emotion analysis to the text data and reflects it in the medical record. For example, information such as "Emotion: acute anxiety" is added.
[0635] Step 9:
[0636] The server then sends the automatically completed medical record information and the emotion analysis results to the terminal. For example, it might generate the following: "I have a headache that has been going on since yesterday. Onset: yesterday. Pain level: moderate. Emotion: acute anxiety."
[0637] Step 10:
[0638] The terminal displays the received medical record information and emotion information to the user, who can then check the displayed information and make corrections as necessary.
[0639] Step 11:
[0640] The user reviews the patient record information and makes corrections as needed, for example, changing the "pain level" from "moderate" to "severe." Emotional information is also reviewed, and measures are planned to address patients with high anxiety.
[0641] Step 12:
[0642] The terminal transmits the medical record information and emotion information corrected by the user to the server. The confirmed corrections are then transmitted to the server.
[0643] Step 13:
[0644] The server stores the corrected medical record information and emotion information received as the patient's electronic medical record. For example, information such as "I have a headache that has continued since yesterday. Onset: yesterday. Pain level: severe. Emotion: acute anxiety" is recorded in the electronic medical record.
[0645] Example 2
[0646] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0647] In current medical settings, it is difficult to accurately and quickly record a patient's oral information and simultaneously manage the patient's emotional state. When medical professionals record information to create medical records, work efficiency decreases and human error is likely to occur. In addition, additional effort is required to understand the patient's emotional state, making it difficult to improve the overall quality of medical care.
[0648] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0649] In this invention, the server includes means for acquiring oral information from the patient as voice data, means for transmitting the voice data to a central computer, means for converting the voice data to text data in the central computer, means for analyzing the text data and automatically completing the medical record content, means for analyzing the emotional state from the voice data in the central computer, means for displaying the automatically completed medical record content and the emotion analysis results so that the user can confirm and correct them, and means for transmitting the confirmed and corrected content to the central computer and saving it. This automates the entire process from voice input to text conversion, emotion analysis, and medical record updating and saving, reducing the burden on medical professionals and enabling them to understand the patient's emotional state, enabling appropriate medical care to be provided.
[0650] "Oral information" refers to verbal information such as symptoms and medical history that a patient tells a medical professional.
[0651] "Audio data" refers to audio files that digitally record oral information obtained from a patient.
[0652] The "central computer" is a server that receives, analyzes, and stores voice and text data sent from the terminal.
[0653] "Speech recognition software" is a program for converting voice data into text data.
[0654] "Text data" is data in which voice data is expressed as a string of characters.
[0655] A "medical record" is a medical document that details a patient's symptoms, medical history, and treatment.
[0656] "Automatic completion" is a process that automatically adds missing parts based on converted text data by referencing past data and databases.
[0657] "Emotion analysis" is the process of analyzing the patient's emotional state from the audio data and adding the results to the text data.
[0658] "User" refers to a medical professional who uses the system to review and modify patient information.
[0659] To implement this invention, voice input technology, voice recognition technology, text data analysis technology, emotion recognition technology, and record management technology are required. This system captures oral information from conversations with patients in real time and automatically reflects it in medical records. It also simultaneously analyzes the patient's emotional state and adds it to the medical record. In this way, it reduces the workload of medical professionals and enables them to manage patient information efficiently and accurately.
[0660] Collecting voice input
[0661] The device has a built-in microphone that collects conversations between patients and medical professionals in real time. For example, a medical professional might ask, "What symptoms do you have?" and the patient might reply, "I have a headache, and it's been going on since yesterday." This conversation is recorded by the device's microphone and saved as audio data.
[0662] Sending audio data
[0663] The device sends the collected voice data to a central computer (server) over the Internet, for example, securely using the HTTPS protocol, where it can be processed for speech recognition and emotion engines.
[0664] Converting audio data to text
[0665] The server uses speech recognition software such as Google Cloud Speech-to-Text to convert the transmitted voice data into text data. For example, the voice data "I have a headache, and it's been going on since yesterday" is converted into text data "I have a headache, and it's been going on since yesterday."
[0666] Text data analysis and auto-completion
[0667] The server analyzes the converted text data and extracts keywords. For example, it recognizes the keyword "headache." It then references the patient's existing information and medical databases to automatically supplement the information with relevant information. For example, information such as "Date of onset: yesterday" and "Pain level: moderate" is added.
[0668] Conducting sentiment analysis
[0669] The server uses emotion recognition software such as IBM Watson Tone Analyzer to analyze the patient's emotional state from the voice data. For example, it obtains emotional information such as "acute anxiety" as a result of the analysis and adds this information to the text data.
[0670] Check and fix
[0671] The device displays the automatically completed medical record and emotion analysis results to the user. The user can review the information and make corrections as necessary. For example, a doctor can change the "pain level" from "moderate" to "severe," and then review the emotion information to consider the necessary response.
[0672] Save the final data
[0673] Once the changes are confirmed, the device sends the revised medical record information and emotion information back to the server. The server saves this as the patient's electronic medical record. For example, the final data recorded would be "I have a headache that has been going on since yesterday. Onset: yesterday. Pain level: severe. Emotion: acute anxiety."
[0674] This system automates the entire process from voice input to text conversion, emotion analysis, and medical record updating and storage, reducing the burden on medical professionals and enabling efficient and accurate information management. It also makes it possible to grasp the patient's emotional state, enabling more appropriate medical care.
[0675] Specific examples and examples of prompts for generative AI models
[0676] 1. Conversation between healthcare providers and patients:
[0677] Medical professional: "What symptoms do you have?"
[0678] Patient: "I have a headache that has been going on since yesterday."
[0679] 2. Audio to text conversion:
[0680] Voice: "I have a headache that has been going on since yesterday."
[0681] Text: "I have a headache that has been going on since yesterday."
[0682] 3. Auto-completion:
[0683] Keyword: "headache"
[0684] Supplementary data: "Date of onset: yesterday" "Pain level: moderate"
[0685] 4. Emotion analysis:
[0686] Speech data → Sentiment analysis : As in, “Acute anxiety .”
[0687] 5. Modification:
[0688] User modification: "Pain level: Moderate" → "Severe"
[0689] 6. Save:
[0690] Medical record data: "Headache, ongoing since yesterday. Onset: yesterday. Pain level: severe. Emotions: acute anxiety."
[0691] Example prompts for generative AI models
[0692] "Write a program that generates text data from audio data, extracts keywords for auto-completion, and adds sentiment analysis using an emotion engine. Also, include a process to finally save this data as an electronic medical record."
[0693] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0694] Step 1:
[0695] Collecting voice input
[0696] The device has a built-in microphone that collects conversations between patients and medical professionals in real time. Specifically, the medical professional asks, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday." While this conversation is taking place, the device records audio data. The input is the audio of the conversation, and the output is digital audio data.
[0697] Step 2:
[0698] Sending voice data to the server
[0699] The device transmits the collected voice data to a central computer (server) via the Internet. For example, the data is transmitted securely using the HTTPS protocol. This allows it to be processed for speech recognition and emotion engines. The input is digital voice data, and the output is the voice data transmitted to the server.
[0700] Step 3:
[0701] Converting audio data to text
[0702] The server uses speech recognition software such as Google Cloud Speech-to-Text to convert the transmitted voice data into text data. For example, the voice data "I have a headache, and it's been going on since yesterday" is converted into text data "I have a headache, and it's been going on since yesterday." The input is the voice data received by the server, and the output is the conversation in text format.
[0703] Step 4:
[0704] Text data analysis and auto-completion
[0705] The server analyzes the converted text data and extracts keywords. For example, it recognizes the keyword "headache." It then references the patient's existing information and medical databases to automatically complete the relevant information. For example, information such as "onset of illness: yesterday" and "pain level: moderate" is added. The input is text data, and the output is the completed text data.
[0706] Step 5:
[0707] Conducting sentiment analysis
[0708] The server uses emotion recognition software such as IBM Watson Tone Analyzer to analyze the patient's emotional state from the voice data. For example, the analysis results in emotional information such as "acute anxiety" and adds this to the text data. The input is voice data, and the output is text data with emotion analysis added.
[0709] Step 6:
[0710] Check and fix
[0711] The device displays the automatically completed medical record content and the emotion analysis results to the user. The user checks this content and makes corrections as necessary. For example, a doctor may change the "pain level" from "moderate" to "strong," check the emotion information, and consider the necessary response. The input is the completed text data and the emotion analysis results, and the output is the corrected medical record data.
[0712] Step 7:
[0713] Save the final data
[0714] Once the corrections are confirmed, the device sends the corrected medical record content and emotion information back to the server. The server saves this as the patient's electronic medical record. For example, the final data recorded is "I have a headache that has been going on since yesterday. Onset date: yesterday. Pain level: severe. Emotion: acute anxiety." The input is the corrected medical record data, and the output is the final data saved in the patient's electronic medical record.
[0715] (Application example 2)
[0716] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0717] Conventional food delivery services face the problem of difficulty in analyzing customer sentiment when taking orders, making it difficult to improve service quality. In particular, when customers are in a hurry or have special requests, they often cannot make prompt and appropriate suggestions, which can lead to a decline in customer satisfaction.
[0718] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0719] In this invention, the server includes means for acquiring dictation information from a customer as voice data, means for transmitting the voice data to the server, means for converting the voice data to text data in the server, means for analyzing the text data and automatically completing the order details, means for displaying the automatically completed order details and emotion analysis results for confirmation and correction by the user, means for transmitting the confirmed and corrected details to the server and saving them, means for analyzing the user's emotions from the voice data, and means for making appropriate suggestions based on the emotion analysis results, thereby enabling quick and accurate order confirmation and suggestions that take customer emotions into consideration.
[0720] "Patient" means an individual receiving medical services.
[0721] "Dicted information" is information provided by the user through speech.
[0722] "Audio data" refers to data in which audio is recorded in digital format.
[0723] A "server" is a computer system that processes and stores data.
[0724] "Text data" refers to data obtained by converting voice data into character information.
[0725] A "medical record" is a digital or paper-based document for recording medical information.
[0726] "Auto-completion" is the process of automatically adding missing data or content based on existing information.
[0727] A "displaying means" is a device or software that presents digital information to a user.
[0728] "Verify and correct" is the process by which a user checks the accuracy of information and corrects it if necessary.
[0729] "Emotion analysis" is a process of analyzing emotional states from voice data.
[0730] The "means for making suggestions" is a means for providing appropriate advice or recommendations to the user based on the analyzed information.
[0731] A "digital format" is a way of representing information in electronic form.
[0732] A "voice recognition engine" is software or hardware for converting voice data into text data.
[0733] This invention is a system that acquires oral information from patients or customers (hereinafter referred to as users) in real time, analyzes and supplements that information, and provides appropriate services quickly and accurately. This system combines technologies such as voice recognition technology, emotion analysis engine, text analysis, and recommendation engine.
[0734] Hardware and software used
[0735] 1. Hardware:
[0736] Devices such as smartphones, smart glasses, and head-mounted displays
[0737] Cloud Server
[0738] 2. Software:
[0739] Speech recognition engine: Google Speech-to-Text API, Amazon Transcribe, etc.
[0740] Emotion recognition engine: IBM Watson Tone Analyzer, Microsoft Azure Emotion API, etc.
[0741] Backend: Node.js, Python (Flask)
[0742] Database: MongoDB, Firebase
[0743] Program processing overview
[0744] Voice input
[0745] The user dictates information into the terminal. For example, in a food delivery service, the user might say, "I'd like to order a hamburger set, and a Coke to drink."
[0746] Voice Recognition
[0747] The voice data collected by the device is sent to a cloud server. The server uses a voice recognition engine to convert this voice data into text data. For example, a voice saying "I'd like to order a hamburger set, and a Coke to drink" is converted into the text "Hamburger set, Coke."
[0748] Auto-completion
[0749] The server analyzes the converted text and automatically completes the order details. For example, it generates an order list such as "Hamburger Set (Contents: Cheeseburger, Fries) + Coke." This automatically adds detailed information according to the user's request.
[0750] emotion recognition
[0751] The voice data is also sent to an emotion analysis engine, which analyzes the user's emotions. For example, the engine may detect that the user is in a hurry. Based on this, a menu of options for responding quickly is presented.
[0752] Check and fix
[0753] The auto-completed order details and the sentiment analysis results are displayed on the device, and the user can confirm and modify them. For example, the user can confirm an order displayed as "Hamburger Set (Cheeseburger, Fries) + Coke." At this time, the user can make modifications as necessary.
[0754] keep
[0755] The final order details and emotion data after confirmation and correction are sent to the server and stored in the database. For example, data such as "Hamburger set (cheeseburger, fries) + Coke. Emotion: In a hurry" is stored.
[0756] Examples of concrete examples and prompts
[0757] For example, suppose a user inputs dictation information such as "This hamburger looks very tasty. I'd like some fries with it, please."
[0758] Speech recognition prompt: transcribe audio
[0759] Emotion recognition prompt: analyze sentiment from text: "This burger looks delicious. I'd like some fries with it, please."
[0760] Recommendation prompt: get recommendations for quick service items based on "interest"
[0761] By implementing this invention, users can enjoy a fast and accurate ordering experience through voice input, and service providers can respond appropriately based on the user's emotions.
[0762] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0763] Step 1:
[0764] The user inputs dictation information into the terminal.
[0765] Specifically, the user speaks to a device such as a smartphone or smart glasses, saying, "I'd like to order a hamburger set, and a Coke to drink."
[0766] The input data is stored in the terminal as voice data.
[0767] The data that is output is the collected voice data.
[0768] Step 2:
[0769] The device sends the collected voice data to a cloud server.
[0770] Specifically, the terminal program uploads the collected voice data to a designated cloud server via the Internet.
[0771] The input data is audio data.
[0772] The output data is the audio data sent to the cloud server.
[0773] Step 3:
[0774] The server uses a speech recognition engine to convert the voice data into text data.
[0775] Specifically, the server calls a speech recognition engine such as the Google Speech-to-Text API or Amazon Transcribe to convert the voice data into text data.
[0776] The input data is audio data.
[0777] The output data is the converted text data.
[0778] Step 4:
[0779] The server analyzes the converted text data and automatically completes the order details.
[0780] Specifically, the server program analyzes the text data, extracts and completes the order details. For example, from the text data "hamburger set, cola," it generates "hamburger set (cheeseburger, fries) + cola."
[0781] The input data is converted text data.
[0782] The output data is the supplemented order details.
[0783] Step 5:
[0784] The voice data is sent to an emotion analysis engine to analyze the user's emotions.
[0785] Specifically, the server uses an emotion analysis engine such as IBM Watson Tone Analyzer or Microsoft Azure Emotion API to analyze the emotional state from the voice data.
[0786] The input data is audio data.
[0787] The output data is the emotion analysis results.
[0788] Step 6:
[0789] The automatically completed order details and sentiment analysis results are displayed on the terminal.
[0790] Specifically, the server sends the generated order details and the emotion analysis results to the terminal, which then displays them to the user. For example, the terminal displays "Hamburger set (cheeseburger, fries) + Coke. Emotion: In a hurry."
[0791] The input data is the supplemented order details and the results of sentiment analysis.
[0792] The output data is the information that is displayed on the terminal.
[0793] Step 7:
[0794] The user reviews and corrects what is displayed.
[0795] Specifically, the user checks the displayed order details and emotion information and makes corrections as necessary. For example, the user may make a correction such as "change fries to large size."
[0796] The data to be input is the information displayed on the terminal.
[0797] The data that is output is the content that has been confirmed and corrected by the user.
[0798] Step 8:
[0799] The confirmed and corrected content is sent to the server and stored in the database.
[0800] Specifically, the terminal sends the confirmed and corrected content back to the cloud server, and the server stores it in the database.
[0801] The data to be input is the content that has been confirmed and corrected by the user.
[0802] The output data is the order details and emotion information stored in the database.
[0803] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0804] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0805] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0806] [Third embodiment]
[0807] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0808] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0809] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0810] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0811] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0812] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0813] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0814] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0815] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0816] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0817] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0818] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0819] As an embodiment of this invention, we provide a system that acquires oral information from conversations with patients in real time and automatically reflects it in medical records. The system combines voice recognition technology with an automatic medical record completion function based on text data, reducing the burden on medical professionals and improving work efficiency.
[0820] Program processing overview
[0821] Voice input
[0822] The user interviews the patient about their symptoms and medical history. The device collects this conversation in real time as audio data. For example, a medical professional might ask the patient, "What symptoms do you have?" and the patient might reply, "I have a headache, and it's been going on since yesterday." The device records this conversation as audio data.
[0823] Voice Recognition
[0824] The recorded voice data is sent from the device to the server. The server uses a voice recognition engine to convert this voice data into text data. For example, the voice "I have a headache, and it has been going on since yesterday" is converted into the text "I have a headache, and it has been going on since yesterday."
[0825] Auto-completion
[0826] The server analyzes the converted text and extracts relevant keywords. Based on the extracted keywords, the server automatically completes the medical record contents using the patient's existing data and general medical knowledge. For example, based on the keyword "headache," information such as "Date of onset: yesterday. Pain level: moderate" is added.
[0827] Check and fix
[0828] The automatically completed medical record information is sent to the terminal and displayed to the user. The user can check the information and make corrections as necessary. For example, a doctor can change the "pain level" from "moderate" to "severe."
[0829] keep
[0830] Once the changes are confirmed, the device sends the revised medical record information to the server. The server saves the received data in the patient's electronic medical record. As a result, for example, data such as "I have a headache that has been going on since yesterday. Onset date: yesterday. Pain level: severe" is recorded in the electronic medical record.
[0831] This system automates the entire process from voice input to final medical record storage, enabling medical professionals to efficiently and accurately record patient information. This invention is particularly useful when doctors and nurses need to quickly and accurately create medical records based on their conversations with patients.
[0832] The processing flow will be explained below.
[0833] Step 1:
[0834] The user interviews the patient about their symptoms and medical history. For example, a medical professional asks, "What symptoms do you have?" The patient answers, "I have a headache, and it's been going on since yesterday."
[0835] Step 2:
[0836] The terminal collects the conversation between the user and the patient as voice data. The terminal starts recording and records the voice data in real time.
[0837] Step 3:
[0838] The device converts the recorded audio data into a digital format and sends it to the server, where it is compressed to optimize bandwidth.
[0839] Step 4:
[0840] The server analyzes the received voice data and converts it into text data using a voice recognition engine. For example, a voice saying "I have a headache, and it's been going on since yesterday" is converted into text "I have a headache, and it's been going on since yesterday."
[0841] Step 5:
[0842] The server analyzes the text data and extracts keywords. Based on the extracted keywords, the server references the patient's existing information and medical databases to automatically complete the information. For example, the keyword "headache" can be supplemented with the information "Date of onset: yesterday, Pain level: moderate."
[0843] Step 6:
[0844] The server generates automatically completed medical record information and sends it to the terminal. For example, medical record data such as "I have a headache that has continued since yesterday. Onset: yesterday. Pain level: moderate" is generated.
[0845] Step 7:
[0846] The terminal displays the received medical record information to the user, who then checks the displayed information and makes any necessary corrections.
[0847] Step 8:
[0848] The user checks the medical record information and makes corrections as necessary, for example, changing the "pain level" from "moderate" to "severe."
[0849] Step 9:
[0850] The terminal sends the medical record information corrected by the user to the server, and after final confirmation, the data is sent.
[0851] Step 10:
[0852] The server stores the received corrected medical record information as the patient's electronic medical record. For example, information such as "I have a headache that has continued since yesterday. Onset: yesterday. Pain level: severe" is recorded in the electronic medical record.
[0853] Example 1
[0854] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0855] In conventional medical settings, creating medical records based on conversations with patients requires a great deal of time and effort. Furthermore, manual data entry often carries the risk of errors and omissions, increasing the burden on medical professionals. The purpose of this invention is to provide a system that reduces the burden on medical professionals, improves work efficiency, and prevents errors and omissions.
[0856] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0857] In this invention, the server includes a means for converting voice data into text data using a voice recognition engine, a means for analyzing the text data and extracting related keywords, and a means for automatically completing the medical record content based on the keywords. This automates the process from voice input to medical record creation, reducing the burden on medical professionals and improving work efficiency and data accuracy.
[0858] "Oral information from patients" refers to audio information such as symptoms and medical history obtained directly from patients by medical professionals.
[0859] "Audio data" refers to digital audio files of recorded conversations with patients.
[0860] "Terminal" refers to an electronic device for collecting and transmitting voice data to a server.
[0861] "Server" refers to the computer system that processes voice data and automatically completes medical records.
[0862] "Speech recognition engine" refers to software or algorithms that analyze voice data and convert it into text data.
[0863] "Text data" refers to character information converted by a voice recognition engine.
[0864] "Natural language processing technology" refers to computer technology for analyzing text data and extracting keywords.
[0865] "Keywords" refer to important words and phrases extracted from text data that are necessary for automatically completing the contents of medical records.
[0866] A "medical record" refers to a medical record that describes a patient's symptoms, medical history, treatment details, etc.
[0867] "Automatic completion" refers to the process in which the system automatically adds and supplements medical record content based on keywords.
[0868] "User" refers to a medical professional who operates the system to check and correct medical records.
[0869] As an embodiment of this invention, we provide a system that acquires oral information from conversations with patients in real time and automatically reflects it in medical records. The system combines voice recognition technology with an automatic medical record completion function based on text data, reducing the burden on medical professionals and improving work efficiency.
[0870] Collecting voice input
[0871] Users, or medical professionals, use devices such as smartphones and tablets to obtain voice data of symptoms and medical history directly from patients. The devices are equipped with highly sensitive microphones that record conversations in real time. For example, if a medical professional asks, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday," this content is collected as voice data on the device.
[0872] Sending audio data
[0873] The device sends the recorded audio data to a server over the Internet using the HTTPS protocol to ensure data security. A program on the device converts the audio data into a digital format and sends it to the server.
[0874] Speech Recognition Processing
[0875] The server passes the received voice data to a voice recognition engine, which converts the voice into text data. A typical voice recognition engine used here is a third-party voice recognition service. For example, voice data such as "I have a headache, and it's been going on since yesterday" is converted directly into text data such as "I have a headache, and it's been going on since yesterday."
[0876] Text analysis and keyword extraction
[0877] The server receives the converted text data and analyzes it using natural language processing technology. Related keywords are extracted through the analysis. For example, the keywords "headache," "yesterday," and "continuing" are extracted from the sentence "I have a headache, and it has been going on since yesterday."
[0878] Auto-completion
[0879] The server automatically completes the medical record based on the extracted keywords. It references the medical database and adds information related to the keywords to the medical record. For example, based on the keyword "headache," the information "Date of onset: yesterday. Pain level: moderate" is automatically added to the medical record.
[0880] Sending and checking medical record contents
[0881] The completed medical record information is sent to the terminal and displayed to the user. The user can check the medical record information and make corrections as necessary. For example, a doctor may change the "pain level" from "moderate" to "severe."
[0882] Saving medical record contents
[0883] Once the user has confirmed the medical record contents after making the edits, the terminal sends them back to the server. The server then saves the received medical record contents in the patient's electronic medical record. For example, information such as "I have a headache that has been going on since yesterday. Onset date: yesterday. Pain level: severe" is recorded in the electronic medical record.
[0884] Examples of concrete examples and prompts
[0885] Example: A healthcare professional asks a patient, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday." The audio is recorded in real time and converted to text. Auto-complete adds, "Date of onset: yesterday. Pain level: moderate." The doctor then corrects "moderate" to "severe." The information is saved in the electronic medical record.
[0886] Example prompt: "Please run a program that will automatically record a conversation about your headache symptoms that have continued since yesterday into the electronic medical record."
[0887] In this way, the system automates the process from voice input to final medical record storage, enabling medical professionals to efficiently and accurately record patient information.
[0888] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0889] Step 1: Collecting voice input
[0890] The user listens to the patient's symptoms and medical history and engages in a conversation. The device uses a built-in microphone to collect this audio in real time and records it as audio data. The input is the conversation between the user and the patient, and the output is digital audio data. Specifically, the medical professional asks, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday." This audio is collected by the device.
[0891] Step 2: Sending audio data
[0892] The audio data collected by the device is sent to a server via the Internet. The input is digital audio data, and the output is audio data transferred to the server. Specifically, a program on the device sends the audio data to the server using the HTTPS protocol. This ensures that the data arrives securely at the server.
[0893] Step 3: Speech recognition processing
[0894] The server passes the received voice data to a voice recognition engine, which converts the voice into text data. The input is the voice data sent to the server, and the output is text data. Specifically, the server calls a voice recognition engine (for example, a natural language processing library) and converts the voice data "I have a headache, and it's been going on since yesterday" into the text "I have a headache, and it's been going on since yesterday."
[0895] Step 4: Text analysis and keyword extraction
[0896] The server analyzes the converted text data and extracts related keywords. The input is text data obtained from the speech recognition engine, and the output is the extracted keywords. Specifically, the server uses natural language processing technology to extract keywords such as "headache," "yesterday," and "continuing" from the text "I have a headache, and it has been going on since yesterday."
[0897] Step 5: Auto-completion
[0898] The server automatically completes the medical record contents based on the extracted keywords. The input is the extracted keywords, and the output is the automatically completed medical record contents. Specifically, the server references the medical database and adds information such as "Date of onset: yesterday. Pain level: moderate" to the medical record based on the keywords "headache," "yesterday," and "continuing."
[0899] Step 6: Send and view medical records
[0900] The server sends the auto-completed medical record contents to the terminal and displays them to the user. The input is the auto-completed medical record contents, and the output is the medical record information displayed on the terminal. Specifically, the server sends the medical record contents to the terminal via the HTTPS protocol, and the terminal displays the information on the user interface.
[0901] Step 7: Check and correct the medical record
[0902] The user checks the medical record displayed on the terminal and makes any necessary corrections. The input is the medical record information displayed on the terminal, and the output is the medical record content corrected by the user. Specifically, the doctor changes the "pain level" on the terminal screen from "moderate" to "severe."
[0903] Step 8: Save the medical record
[0904] The user confirms the medical record contents after making the edits, and the terminal sends them back to the server. The server saves the received medical record contents in the patient's electronic medical record. The input is the edited medical record contents, and the output is the information saved in the electronic medical record. Specifically, when the user presses the "Save" button on the terminal, the terminal sends the edited contents to the server, and the server finally records the data "Date of onset: yesterday. Pain level: severe" in the electronic medical record.
[0905] Through these steps, this invention automates the process from voice input to final medical record storage, reducing the burden on medical professionals and achieving efficient and accurate information management.
[0906] (Application example 1)
[0907] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0908] In the traditional medical record creation process, medical professionals manually record patient information, consuming a great deal of time and effort. Human errors, such as clerical errors and omissions, are also common. Meanwhile, in autonomous vehicles, systems for responding quickly and appropriately to vehicle abnormalities or emergencies may be inadequate, potentially reducing passenger safety and vehicle operational efficiency. A system that solves these problems and improves automation and efficiency is needed.
[0909] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0910] In this invention, the server includes means for acquiring oral information from the patient as voice data, means for transmitting the voice data to a data processing device, means for converting the voice data into text data in the data processing device, means for analyzing the text data and automatically completing the medical record content, means for displaying the automatically completed medical record content and allowing the user to confirm and correct it, means for transmitting the confirmed and corrected content to the data processing device and saving it, means for collecting information on the vehicle condition and external environment from the conversation content of the occupants, and means for analyzing the collected information and proposing countermeasures if an abnormality is detected. This enables improved efficiency and accuracy in creating medical records in the medical field, and enables quick and appropriate responses to abnormalities and emergencies in self-driving vehicles.
[0911] "Patient" refers to anyone who uses a medical institution or healthcare system.
[0912] "Oral information" refers to information or content that is spoken aloud.
[0913] "Audio data" is data converted from audio into digital form.
[0914] A "data processing device" is a machine or system that processes voice data, converts it into text data, and analyzes it.
[0915] "Text data" refers to written information in digital form.
[0916] "Analysis" is the process of extracting meaning and keywords from text data and organizing the information.
[0917] A "medical record" is a document that records a patient's medical treatment at a medical institution.
[0918] "Auto-completion" is the process by which the system automatically fills in missing data based on retrieved information.
[0919] A "user" is someone who operates the system and reviews and modifies the results.
[0920] "Display" means the visual presentation of information through a data processing device or other output device.
[0921] "Occupant" refers to any person riding in an autonomous vehicle.
[0922] "Conversation content" refers to verbal exchanges between multiple people.
[0923] "Vehicle status" is information indicating the operating status of the vehicle and the operating status of each function.
[0924] The "external environment" refers to the surrounding circumstances and conditions in which the vehicle is traveling.
[0925] "Collection" is the act of gathering specific information.
[0926] An "abnormality" is an event that indicates a state or malfunction that is different from the normal state.
[0927] A "solution" is a method or means for dealing with a particular situation.
[0928] As an embodiment of the present invention, there is provided a system in which a voice recognition system is installed in an autonomous driving vehicle, collects information on the vehicle's condition and the external environment from the conversation of the occupants, and proposes countermeasures when an abnormality is detected. The server includes means for acquiring oral information from the patient as voice data, means for transmitting the voice data to a data processing device, means for converting the voice data into text data in the data processing device, means for analyzing the text data and automatically completing the medical record content, means for displaying the automatically completed medical record content and allowing the user to confirm and correct it, means for transmitting the confirmed and corrected content to the data processing device and saving it, means for collecting information on the vehicle's condition and the external environment from the conversation of the occupants, and means for analyzing the collected information and proposing countermeasures when an abnormality is detected.
[0929] Hardware and software used
[0930] The server processes the audio data using the following hardware and software:
[0931] Microphone: Microphones installed inside the vehicle are used to collect passenger conversations in real time.
[0932] Data processing device: Receives and processes voice data. Here, a cloud-based server plays a key role.
[0933] Speech recognition engine: Uses the Google Cloud Speech-to-Text API to convert voice data into text data.
[0934] Text analysis engine: Uses NLTK (Natural Language Toolkit) to analyze the collected text data and extract relevant keywords.
[0935] Database: AWS RDS (Relational Database Service) is used to store analysis results and medical record information.
[0936] Navigation API: Uses Google Maps API to get real-time traffic information and external environment data.
[0937] Data processing and calculation
[0938] 1. Voice collection: Microphones inside the vehicle collect the passengers' conversations as voice data.
[0939] 2. Speech Recognition: The collected voice data is sent to a data processing device and converted into text data using the Google Cloud Speech-to-Text API.
[0940] 3. Keyword Extraction: The text data is analyzed using NLTK to extract relevant keywords.
[0941] 4. Data matching and auto-completion: Based on the extracted keywords, they are matched with historical data stored in AWS RDS or data retrieved from the navigation API, and then auto-completion is performed.
[0942] 5. Display and correction: The auto-completed content is displayed on the vehicle's display, where the user can confirm and correct it.
[0943] 6. Storage: The final confirmed and corrected content is sent to the data processing device and stored in AWS RDS.
[0944] Specific examples
[0945] For example, if a passenger says, "The air conditioner is not working properly," the conversation is collected as voice data via a microphone. This voice data is then sent to a server, where a voice recognition engine converts it into text data such as "The air conditioner is not working properly." Next, a text analysis engine analyzes this text data and extracts keywords such as "air conditioner," "working," and "not good." Finally, the system checks the condition of the air conditioner, and if there is an abnormality, it suggests a solution such as "Air conditioner abnormality: Please check."
[0946] Prompt Sentence Examples
[0947] "As a voice-activated vehicle assistant, please analyze the following conversation and suggest an appropriate response. Conversation: 'The air conditioning isn't working. Please add a new restaurant to Maps.'"
[0948] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0949] Step 1:
[0950] Audio Collection
[0951] Subject: Terminal
[0952] How it works: A microphone installed in the vehicle collects passenger conversations in real time. For example, if a passenger says, "The air conditioning isn't working," the voice is collected as voice data through the microphone.
[0953] Input: Passenger voice
[0954] Output: Collected audio data
[0955] Step 2:
[0956] Sending audio data
[0957] Subject: Terminal
[0958] Specific operation: The collected voice data is sent to a data processing device (cloud server). This transmission is done in real time so that the voice data can be processed immediately.
[0959] Input: Collected audio data
[0960] Output: Audio data sent to the data processing device
[0961] Step 3:
[0962] Voice Recognition
[0963] Subject: Server
[0964] Specific operation: The data processing device uses the Google Cloud Speech-to-Text API to convert the received voice data into text data. For example, the voice saying "The air conditioner is not working" is converted into the text "The air conditioner is not working."
[0965] Input: Transmitted audio data
[0966] Output: Converted text data
[0967] Step 4:
[0968] Keyword extraction
[0969] Subject: Server
[0970] Specific operation: The server analyzes the converted text data using NLTK and extracts related keywords. For example, from the text "The air conditioner is not working," keywords such as "air conditioner" and "not working" are extracted.
[0971] Input: Converted text data
[0972] Output: Extracted keywords
[0973] Step 5:
[0974] Data matching and auto-completion
[0975] Subject: Server
[0976] Specific operation: Based on the extracted keywords, the server compares them with past data stored in AWS RDS and information obtained from the navigation API (Google Maps API) and performs auto-completion. For example, based on the keywords "air conditioner" and "not working," the server will automatically complete the search results by checking the air conditioner's status and suggesting a solution such as "Air conditioner malfunction: check."
[0977] Input: Extracted keywords
[0978] Output: Auto-completed solutions and information
[0979] Step 6:
[0980] View and Modify
[0981] Subject: Terminal
[0982] Specific operation: The supplemented information is displayed on the vehicle's display. The user (passenger) can check this information and make corrections as necessary. For example, the passenger can correct the information by saying, "I checked the air conditioning, and there is actually no problem."
[0983] Input: Auto-completed solutions and information
[0984] Output: Displayed completions and user corrections
[0985] Step 7:
[0986] Submitting and saving your modifications
[0987] Subject: Terminal
[0988] Specific operation: The content modified by the user is sent back to the data processing device and stored in AWS RDS, so that the latest modified information is reflected in the system and can be used as reference data in the future.
[0989] Input: User-modified content
[0990] Output: Saved modifications
[0991] In this way, the system achieves its objective by performing data processing and calculations based on the input data at each step and obtaining the final output.
[0992] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0993] As an embodiment of this invention, we provide a system that acquires oral information from conversations with patients in real time and automatically reflects it in medical records, as well as a system that recognizes the user's emotions by combining an emotion engine. This system combines voice recognition technology, an automatic medical record completion function based on text data, and an emotion recognition function to reduce the burden on medical professionals and improve work efficiency.
[0994] Program processing overview
[0995] Voice input
[0996] The user interviews the patient about their symptoms and medical history. The device collects this conversation in real time as audio data. For example, a medical professional might ask, "What symptoms do you have?" The patient might reply, "I have a headache, and it's been going on since yesterday." The device then records this conversation as audio data.
[0997] Voice Recognition
[0998] The recorded voice data is sent from the device to the server. The server uses a voice recognition engine to convert this voice data into text data. For example, the voice "I have a headache, and it's been going on since yesterday" is converted into text "I have a headache, and it's been going on since yesterday."
[0999] Auto-completion
[1000] The server analyzes the converted text and extracts keywords. Based on the extracted keywords, it references the patient's existing information and medical databases to automatically complete the information. For example, based on the keyword "headache," the information "Date of onset: yesterday, Pain level: moderate" is added.
[1001] emotion recognition
[1002] The voice data is also sent to the emotion engine. The server uses the emotion engine to analyze the user's emotions from the voice data. For example, the analysis may reveal that the user is feeling "acute anxiety." The emotion engine adds this to the text data and reflects it in the patient's medical record.
[1003] Check and fix
[1004] The automatically completed medical record contents and emotion analysis results are sent to the terminal and displayed to the user. The user can check the contents and make corrections as necessary. For example, a doctor may change the "pain level" from "moderate" to "severe." Emotional information is also checked, and appropriate responses for patients with strong anxiety are considered.
[1005] keep
[1006] Once the changes are confirmed, the device sends the revised medical record information and emotional information to the server. The server saves this as the patient's electronic medical record. For example, the following data is recorded in the electronic medical record: "I have a headache that has been going on since yesterday. Onset: yesterday. Pain level: severe. Emotion: acute anxiety."
[1007] This system automates everything from voice input to emotion recognition and final medical record storage. It not only enables medical professionals to efficiently and accurately record patient information, but also allows them to understand the patient's emotional state, enabling them to provide appropriate medical care. This invention is particularly useful in interactive medical environments that emphasize the patient's emotional state.
[1008] The processing flow will be explained below.
[1009] Step 1:
[1010] The user interviews the patient about their symptoms and medical history. For example, a medical professional asks, "What symptoms do you have?" The patient answers, "I have a headache, and it's been going on since yesterday."
[1011] Step 2:
[1012] The terminal collects the conversation between the user and the patient as voice data. The terminal starts recording and records the voice data in real time.
[1013] Step 3:
[1014] The device converts the recorded audio data into a digital format and sends it to the server, where it is compressed to optimize bandwidth.
[1015] Step 4:
[1016] The server analyzes the received voice data and converts it into text data using a voice recognition engine. For example, a voice saying "I have a headache, and it's been going on since yesterday" is converted into text "I have a headache, and it's been going on since yesterday."
[1017] Step 5:
[1018] The server analyzes the text data and extracts keywords. Based on the extracted keywords, the server references the patient's existing information and medical databases to automatically complete the information. For example, the keyword "headache" can be supplemented with the information "Date of onset: yesterday, Pain level: moderate."
[1019] Step 6:
[1020] The server also sends the voice data to the emotion engine.
[1021] Step 7:
[1022] The emotion engine on the server analyzes the user's emotions from the voice data. For example, the emotion analysis may reveal that the user is experiencing "acute anxiety."
[1023] Step 8:
[1024] The server adds the results of the emotion analysis to the text data and reflects it in the medical record. For example, information such as "Emotion: acute anxiety" is added.
[1025] Step 9:
[1026] The server then sends the automatically completed medical record information and the emotion analysis results to the terminal. For example, it might generate the following: "I have a headache that has been going on since yesterday. Onset: yesterday. Pain level: moderate. Emotion: acute anxiety."
[1027] Step 10:
[1028] The terminal displays the received medical record information and emotion information to the user, who can then check the displayed information and make corrections as necessary.
[1029] Step 11:
[1030] The user reviews the patient record information and makes corrections as needed, for example, changing the "pain level" from "moderate" to "severe." Emotional information is also reviewed, and measures are planned to address patients with high anxiety.
[1031] Step 12:
[1032] The terminal transmits the medical record information and emotion information corrected by the user to the server. The confirmed corrections are then transmitted to the server.
[1033] Step 13:
[1034] The server stores the corrected medical record information and emotion information received as the patient's electronic medical record. For example, information such as "I have a headache that has continued since yesterday. Onset: yesterday. Pain level: severe. Emotion: acute anxiety" is recorded in the electronic medical record.
[1035] Example 2
[1036] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1037] In current medical settings, it is difficult to accurately and quickly record a patient's oral information and simultaneously manage the patient's emotional state. When medical professionals record information to create medical records, work efficiency decreases and human error is likely to occur. In addition, additional effort is required to understand the patient's emotional state, making it difficult to improve the overall quality of medical care.
[1038] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1039] In this invention, the server includes means for acquiring oral information from the patient as voice data, means for transmitting the voice data to a central computer, means for converting the voice data to text data in the central computer, means for analyzing the text data and automatically completing the medical record content, means for analyzing the emotional state from the voice data in the central computer, means for displaying the automatically completed medical record content and the emotion analysis results so that the user can confirm and correct them, and means for transmitting the confirmed and corrected content to the central computer and saving it. This automates the entire process from voice input to text conversion, emotion analysis, and medical record updating and saving, reducing the burden on medical professionals and enabling them to understand the patient's emotional state, enabling appropriate medical care to be provided.
[1040] "Oral information" refers to verbal information such as symptoms and medical history that a patient tells a medical professional.
[1041] "Audio data" refers to audio files that digitally record oral information obtained from a patient.
[1042] The "central computer" is a server that receives, analyzes, and stores voice and text data sent from the terminal.
[1043] "Speech recognition software" is a program for converting voice data into text data.
[1044] "Text data" is data in which voice data is expressed as a string of characters.
[1045] A "medical record" is a medical document that details a patient's symptoms, medical history, and treatment.
[1046] "Automatic completion" is a process that automatically adds missing parts based on converted text data by referencing past data and databases.
[1047] "Emotion analysis" is the process of analyzing the patient's emotional state from the audio data and adding the results to the text data.
[1048] "User" refers to a medical professional who uses the system to review and modify patient information.
[1049] To implement this invention, voice input technology, voice recognition technology, text data analysis technology, emotion recognition technology, and record management technology are required. This system captures oral information from conversations with patients in real time and automatically reflects it in medical records. It also simultaneously analyzes the patient's emotional state and adds it to the medical record. In this way, it reduces the workload of medical professionals and enables them to manage patient information efficiently and accurately.
[1050] Collecting voice input
[1051] The device has a built-in microphone that collects conversations between patients and medical professionals in real time. For example, a medical professional might ask, "What symptoms do you have?" and the patient might reply, "I have a headache, and it's been going on since yesterday." This conversation is recorded by the device's microphone and saved as audio data.
[1052] Sending audio data
[1053] The device sends the collected voice data to a central computer (server) over the Internet, for example, securely using the HTTPS protocol, where it can be processed for speech recognition and emotion engines.
[1054] Converting audio data to text
[1055] The server uses speech recognition software such as Google Cloud Speech-to-Text to convert the transmitted voice data into text data. For example, the voice data "I have a headache, and it's been going on since yesterday" is converted into text data "I have a headache, and it's been going on since yesterday."
[1056] Text data analysis and auto-completion
[1057] The server analyzes the converted text data and extracts keywords. For example, it recognizes the keyword "headache." It then references the patient's existing information and medical databases to automatically supplement the information with relevant information. For example, information such as "Date of onset: yesterday" and "Pain level: moderate" is added.
[1058] Conducting sentiment analysis
[1059] The server uses emotion recognition software such as IBM Watson Tone Analyzer to analyze the patient's emotional state from the voice data. For example, it obtains emotional information such as "acute anxiety" as a result of the analysis and adds this information to the text data.
[1060] Check and fix
[1061] The device displays the automatically completed medical record and emotion analysis results to the user. The user can review the information and make corrections as necessary. For example, a doctor can change the "pain level" from "moderate" to "severe," and then review the emotion information to consider the necessary response.
[1062] Save the final data
[1063] Once the changes are confirmed, the device sends the revised medical record information and emotion information back to the server. The server saves this as the patient's electronic medical record. For example, the final data recorded would be "I have a headache that has been going on since yesterday. Onset: yesterday. Pain level: severe. Emotion: acute anxiety."
[1064] This system automates the entire process from voice input to text conversion, emotion analysis, and medical record updating and storage, reducing the burden on medical professionals and enabling efficient and accurate information management. It also makes it possible to grasp the patient's emotional state, enabling more appropriate medical care.
[1065] Specific examples and examples of prompts for generative AI models
[1066] 1. Conversation between healthcare providers and patients:
[1067] Medical professional: "What symptoms do you have?"
[1068] Patient: "I have a headache that has been going on since yesterday."
[1069] 2. Audio to text conversion:
[1070] Voice: "I have a headache that has been going on since yesterday."
[1071] Text: "I have a headache that has been going on since yesterday."
[1072] 3. Auto-completion:
[1073] Keyword: "headache"
[1074] Supplementary data: "Date of onset: yesterday" "Pain level: moderate"
[1075] 4. Emotion analysis:
[1076] Speech data → Sentiment analysis : As in, “Acute anxiety .”
[1077] 5. Modification:
[1078] User modification: "Pain level: Moderate" → "Severe"
[1079] 6. Save:
[1080] Medical record data: "Headache, ongoing since yesterday. Onset: yesterday. Pain level: severe. Emotions: acute anxiety."
[1081] Example prompts for generative AI models
[1082] "Write a program that generates text data from audio data, extracts keywords for auto-completion, and adds sentiment analysis using an emotion engine. Also, include a process to finally save this data as an electronic medical record."
[1083] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1084] Step 1:
[1085] Collecting voice input
[1086] The device has a built-in microphone that collects conversations between patients and medical professionals in real time. Specifically, the medical professional asks, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday." While this conversation is taking place, the device records audio data. The input is the audio of the conversation, and the output is digital audio data.
[1087] Step 2:
[1088] Sending voice data to the server
[1089] The device transmits the collected voice data to a central computer (server) via the Internet. For example, the data is transmitted securely using the HTTPS protocol. This allows it to be processed for speech recognition and emotion engines. The input is digital voice data, and the output is the voice data transmitted to the server.
[1090] Step 3:
[1091] Converting audio data to text
[1092] The server uses speech recognition software such as Google Cloud Speech-to-Text to convert the transmitted voice data into text data. For example, the voice data "I have a headache, and it's been going on since yesterday" is converted into text data "I have a headache, and it's been going on since yesterday." The input is the voice data received by the server, and the output is the conversation in text format.
[1093] Step 4:
[1094] Text data analysis and auto-completion
[1095] The server analyzes the converted text data and extracts keywords. For example, it recognizes the keyword "headache." It then references the patient's existing information and medical databases to automatically complete the relevant information. For example, information such as "onset of illness: yesterday" and "pain level: moderate" is added. The input is text data, and the output is the completed text data.
[1096] Step 5:
[1097] Conducting sentiment analysis
[1098] The server uses emotion recognition software such as IBM Watson Tone Analyzer to analyze the patient's emotional state from the voice data. For example, the analysis results in emotional information such as "acute anxiety" and adds this to the text data. The input is voice data, and the output is text data with emotion analysis added.
[1099] Step 6:
[1100] Check and fix
[1101] The device displays the automatically completed medical record content and the emotion analysis results to the user. The user checks this content and makes corrections as necessary. For example, a doctor may change the "pain level" from "moderate" to "strong," check the emotion information, and consider the necessary response. The input is the completed text data and the emotion analysis results, and the output is the corrected medical record data.
[1102] Step 7:
[1103] Save the final data
[1104] Once the corrections are confirmed, the device sends the corrected medical record content and emotion information back to the server. The server saves this as the patient's electronic medical record. For example, the final data recorded is "I have a headache that has been going on since yesterday. Onset date: yesterday. Pain level: severe. Emotion: acute anxiety." The input is the corrected medical record data, and the output is the final data saved in the patient's electronic medical record.
[1105] (Application example 2)
[1106] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1107] Conventional food delivery services face the problem of difficulty in analyzing customer sentiment when taking orders, making it difficult to improve service quality. In particular, when customers are in a hurry or have special requests, they often cannot make prompt and appropriate suggestions, which can lead to a decline in customer satisfaction.
[1108] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1109] In this invention, the server includes means for acquiring dictation information from a customer as voice data, means for transmitting the voice data to the server, means for converting the voice data to text data in the server, means for analyzing the text data and automatically completing the order details, means for displaying the automatically completed order details and emotion analysis results for confirmation and correction by the user, means for transmitting the confirmed and corrected details to the server and saving them, means for analyzing the user's emotions from the voice data, and means for making appropriate suggestions based on the emotion analysis results, thereby enabling quick and accurate order confirmation and suggestions that take customer emotions into consideration.
[1110] "Patient" means an individual receiving medical services.
[1111] "Dicted information" is information provided by the user through speech.
[1112] "Audio data" refers to data in which audio is recorded in digital format.
[1113] A "server" is a computer system that processes and stores data.
[1114] "Text data" refers to data obtained by converting voice data into character information.
[1115] A "medical record" is a digital or paper-based document for recording medical information.
[1116] "Auto-completion" is the process of automatically adding missing data or content based on existing information.
[1117] A "displaying means" is a device or software that presents digital information to a user.
[1118] "Verify and correct" is the process by which a user checks the accuracy of information and corrects it if necessary.
[1119] "Emotion analysis" is a process of analyzing emotional states from voice data.
[1120] The "means for making suggestions" is a means for providing appropriate advice or recommendations to the user based on the analyzed information.
[1121] A "digital format" is a way of representing information in electronic form.
[1122] A "voice recognition engine" is software or hardware for converting voice data into text data.
[1123] This invention is a system that acquires oral information from patients or customers (hereinafter referred to as users) in real time, analyzes and supplements that information, and provides appropriate services quickly and accurately. This system combines technologies such as voice recognition technology, emotion analysis engine, text analysis, and recommendation engine.
[1124] Hardware and software used
[1125] 1. Hardware:
[1126] Devices such as smartphones, smart glasses, and head-mounted displays
[1127] Cloud Server
[1128] 2. Software:
[1129] Speech recognition engine: Google Speech-to-Text API, Amazon Transcribe, etc.
[1130] Emotion recognition engine: IBM Watson Tone Analyzer, Microsoft Azure Emotion API, etc.
[1131] Backend: Node.js, Python (Flask)
[1132] Database: MongoDB, Firebase
[1133] Program processing overview
[1134] Voice input
[1135] The user dictates information into the terminal. For example, in a food delivery service, the user might say, "I'd like to order a hamburger set, and a Coke to drink."
[1136] Voice Recognition
[1137] The voice data collected by the device is sent to a cloud server. The server uses a voice recognition engine to convert this voice data into text data. For example, a voice saying "I'd like to order a hamburger set, and a Coke to drink" is converted into the text "Hamburger set, Coke."
[1138] Auto-completion
[1139] The server analyzes the converted text and automatically completes the order details. For example, it generates an order list such as "Hamburger Set (Contents: Cheeseburger, Fries) + Coke." This automatically adds detailed information according to the user's request.
[1140] emotion recognition
[1141] The voice data is also sent to an emotion analysis engine, which analyzes the user's emotions. For example, the engine may detect that the user is in a hurry. Based on this, a menu of options for responding quickly is presented.
[1142] Check and fix
[1143] The auto-completed order details and the sentiment analysis results are displayed on the device, and the user can confirm and modify them. For example, the user can confirm an order displayed as "Hamburger Set (Cheeseburger, Fries) + Coke." At this time, the user can make modifications as necessary.
[1144] keep
[1145] The final order details and emotion data after confirmation and correction are sent to the server and stored in the database. For example, data such as "Hamburger set (cheeseburger, fries) + Coke. Emotion: In a hurry" is stored.
[1146] Examples of concrete examples and prompts
[1147] For example, suppose a user inputs dictation information such as "This hamburger looks very tasty. I'd like some fries with it, please."
[1148] Speech recognition prompt: transcribe audio
[1149] Emotion recognition prompt: analyze sentiment from text: "This burger looks delicious. I'd like some fries with it, please."
[1150] Recommendation prompt: get recommendations for quick service items based on "interest"
[1151] By implementing this invention, users can enjoy a fast and accurate ordering experience through voice input, and service providers can respond appropriately based on the user's emotions.
[1152] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1153] Step 1:
[1154] The user inputs dictation information into the terminal.
[1155] Specifically, the user speaks to a device such as a smartphone or smart glasses, saying, "I'd like to order a hamburger set, and a Coke to drink."
[1156] The input data is stored in the terminal as voice data.
[1157] The data that is output is the collected voice data.
[1158] Step 2:
[1159] The device sends the collected voice data to a cloud server.
[1160] Specifically, the terminal program uploads the collected voice data to a designated cloud server via the Internet.
[1161] The input data is audio data.
[1162] The output data is the audio data sent to the cloud server.
[1163] Step 3:
[1164] The server uses a speech recognition engine to convert the voice data into text data.
[1165] Specifically, the server calls a speech recognition engine such as the Google Speech-to-Text API or Amazon Transcribe to convert the voice data into text data.
[1166] The input data is audio data.
[1167] The output data is the converted text data.
[1168] Step 4:
[1169] The server analyzes the converted text data and automatically completes the order details.
[1170] Specifically, the server program analyzes the text data, extracts and completes the order details. For example, from the text data "hamburger set, cola," it generates "hamburger set (cheeseburger, fries) + cola."
[1171] The input data is converted text data.
[1172] The output data is the supplemented order details.
[1173] Step 5:
[1174] The voice data is sent to an emotion analysis engine to analyze the user's emotions.
[1175] Specifically, the server uses an emotion analysis engine such as IBM Watson Tone Analyzer or Microsoft Azure Emotion API to analyze the emotional state from the voice data.
[1176] The input data is audio data.
[1177] The output data is the emotion analysis results.
[1178] Step 6:
[1179] The automatically completed order details and sentiment analysis results are displayed on the terminal.
[1180] Specifically, the server sends the generated order details and the emotion analysis results to the terminal, which then displays them to the user. For example, the terminal displays "Hamburger set (cheeseburger, fries) + Coke. Emotion: In a hurry."
[1181] The input data is the supplemented order details and the results of sentiment analysis.
[1182] The output data is the information that is displayed on the terminal.
[1183] Step 7:
[1184] The user reviews and corrects what is displayed.
[1185] Specifically, the user checks the displayed order details and emotion information and makes corrections as necessary. For example, the user may make a correction such as "change fries to large size."
[1186] The data to be input is the information displayed on the terminal.
[1187] The data that is output is the content that has been confirmed and corrected by the user.
[1188] Step 8:
[1189] The confirmed and corrected content is sent to the server and stored in the database.
[1190] Specifically, the terminal sends the confirmed and corrected content back to the cloud server, and the server stores it in the database.
[1191] The data to be input is the content that has been confirmed and corrected by the user.
[1192] The output data is the order details and emotion information stored in the database.
[1193] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1194] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1195] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1196] [Fourth embodiment]
[1197] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1198] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1199] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1200] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1201] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1202] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1203] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1204] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1205] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1206] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1207] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1208] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1209] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1210] As an embodiment of this invention, we provide a system that acquires oral information from conversations with patients in real time and automatically reflects it in medical records. The system combines voice recognition technology with an automatic medical record completion function based on text data, reducing the burden on medical professionals and improving work efficiency.
[1211] Program processing overview
[1212] Voice input
[1213] The user interviews the patient about their symptoms and medical history. The device collects this conversation in real time as audio data. For example, a medical professional might ask the patient, "What symptoms do you have?" and the patient might reply, "I have a headache, and it's been going on since yesterday." The device records this conversation as audio data.
[1214] Voice Recognition
[1215] The recorded voice data is sent from the device to the server. The server uses a voice recognition engine to convert this voice data into text data. For example, the voice "I have a headache, and it has been going on since yesterday" is converted into the text "I have a headache, and it has been going on since yesterday."
[1216] Auto-completion
[1217] The server analyzes the converted text and extracts relevant keywords. Based on the extracted keywords, the server automatically completes the medical record contents using the patient's existing data and general medical knowledge. For example, based on the keyword "headache," information such as "Date of onset: yesterday. Pain level: moderate" is added.
[1218] Check and fix
[1219] The automatically completed medical record information is sent to the terminal and displayed to the user. The user can check the information and make corrections as necessary. For example, a doctor can change the "pain level" from "moderate" to "severe."
[1220] keep
[1221] Once the changes are confirmed, the device sends the revised medical record information to the server. The server saves the received data in the patient's electronic medical record. As a result, for example, data such as "I have a headache that has been going on since yesterday. Onset date: yesterday. Pain level: severe" is recorded in the electronic medical record.
[1222] This system automates the entire process from voice input to final medical record storage, enabling medical professionals to efficiently and accurately record patient information. This invention is particularly useful when doctors and nurses need to quickly and accurately create medical records based on their conversations with patients.
[1223] The processing flow will be explained below.
[1224] Step 1:
[1225] The user interviews the patient about their symptoms and medical history. For example, a medical professional asks, "What symptoms do you have?" The patient answers, "I have a headache, and it's been going on since yesterday."
[1226] Step 2:
[1227] The terminal collects the conversation between the user and the patient as voice data. The terminal starts recording and records the voice data in real time.
[1228] Step 3:
[1229] The device converts the recorded audio data into a digital format and sends it to the server, where it is compressed to optimize bandwidth.
[1230] Step 4:
[1231] The server analyzes the received voice data and converts it into text data using a voice recognition engine. For example, a voice saying "I have a headache, and it's been going on since yesterday" is converted into text "I have a headache, and it's been going on since yesterday."
[1232] Step 5:
[1233] The server analyzes the text data and extracts keywords. Based on the extracted keywords, the server references the patient's existing information and medical databases to automatically complete the information. For example, the keyword "headache" can be supplemented with the information "Date of onset: yesterday, Pain level: moderate."
[1234] Step 6:
[1235] The server generates automatically completed medical record information and sends it to the terminal. For example, medical record data such as "I have a headache that has continued since yesterday. Onset: yesterday. Pain level: moderate" is generated.
[1236] Step 7:
[1237] The terminal displays the received medical record information to the user, who then checks the displayed information and makes any necessary corrections.
[1238] Step 8:
[1239] The user checks the medical record information and makes corrections as necessary, for example, changing the "pain level" from "moderate" to "severe."
[1240] Step 9:
[1241] The terminal sends the medical record information corrected by the user to the server, and after final confirmation, the data is sent.
[1242] Step 10:
[1243] The server stores the received corrected medical record information as the patient's electronic medical record. For example, information such as "I have a headache that has continued since yesterday. Onset: yesterday. Pain level: severe" is recorded in the electronic medical record.
[1244] Example 1
[1245] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1246] In conventional medical settings, creating medical records based on conversations with patients requires a great deal of time and effort. Furthermore, manual data entry often carries the risk of errors and omissions, increasing the burden on medical professionals. The purpose of this invention is to provide a system that reduces the burden on medical professionals, improves work efficiency, and prevents errors and omissions.
[1247] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1248] In this invention, the server includes a means for converting voice data into text data using a voice recognition engine, a means for analyzing the text data and extracting related keywords, and a means for automatically completing the medical record content based on the keywords. This automates the process from voice input to medical record creation, reducing the burden on medical professionals and improving work efficiency and data accuracy.
[1249] "Oral information from patients" refers to audio information such as symptoms and medical history obtained directly from patients by medical professionals.
[1250] "Audio data" refers to digital audio files of recorded conversations with patients.
[1251] "Terminal" refers to an electronic device for collecting and transmitting voice data to a server.
[1252] "Server" refers to the computer system that processes voice data and automatically completes medical records.
[1253] "Speech recognition engine" refers to software or algorithms that analyze voice data and convert it into text data.
[1254] "Text data" refers to character information converted by a voice recognition engine.
[1255] "Natural language processing technology" refers to computer technology for analyzing text data and extracting keywords.
[1256] "Keywords" refer to important words and phrases extracted from text data that are necessary for automatically completing the contents of medical records.
[1257] A "medical record" refers to a medical record that describes a patient's symptoms, medical history, treatment details, etc.
[1258] "Automatic completion" refers to the process in which the system automatically adds and supplements medical record content based on keywords.
[1259] "User" refers to a medical professional who operates the system to check and correct medical records.
[1260] As an embodiment of this invention, we provide a system that acquires oral information from conversations with patients in real time and automatically reflects it in medical records. The system combines voice recognition technology with an automatic medical record completion function based on text data, reducing the burden on medical professionals and improving work efficiency.
[1261] Collecting voice input
[1262] Users, or medical professionals, use devices such as smartphones and tablets to obtain voice data of symptoms and medical history directly from patients. The devices are equipped with highly sensitive microphones that record conversations in real time. For example, if a medical professional asks, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday," this content is collected as voice data on the device.
[1263] Sending audio data
[1264] The device sends the recorded audio data to a server over the Internet using the HTTPS protocol to ensure data security. A program on the device converts the audio data into a digital format and sends it to the server.
[1265] Speech Recognition Processing
[1266] The server passes the received voice data to a voice recognition engine, which converts the voice into text data. A typical voice recognition engine used here is a third-party voice recognition service. For example, voice data such as "I have a headache, and it's been going on since yesterday" is converted directly into text data such as "I have a headache, and it's been going on since yesterday."
[1267] Text analysis and keyword extraction
[1268] The server receives the converted text data and analyzes it using natural language processing technology. Related keywords are extracted through the analysis. For example, the keywords "headache," "yesterday," and "continuing" are extracted from the sentence "I have a headache, and it has been going on since yesterday."
[1269] Auto-completion
[1270] The server automatically completes the medical record based on the extracted keywords. It references the medical database and adds information related to the keywords to the medical record. For example, based on the keyword "headache," the information "Date of onset: yesterday. Pain level: moderate" is automatically added to the medical record.
[1271] Sending and checking medical record contents
[1272] The completed medical record information is sent to the terminal and displayed to the user. The user can check the medical record information and make corrections as necessary. For example, a doctor may change the "pain level" from "moderate" to "severe."
[1273] Saving medical record contents
[1274] Once the user has confirmed the medical record contents after making the edits, the terminal sends them back to the server. The server then saves the received medical record contents in the patient's electronic medical record. For example, information such as "I have a headache that has been going on since yesterday. Onset date: yesterday. Pain level: severe" is recorded in the electronic medical record.
[1275] Examples of concrete examples and prompts
[1276] Example: A healthcare professional asks a patient, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday." The audio is recorded in real time and converted to text. Auto-complete adds, "Date of onset: yesterday. Pain level: moderate." The doctor then corrects "moderate" to "severe." The information is saved in the electronic medical record.
[1277] Example prompt: "Please run a program that will automatically record a conversation about your headache symptoms that have continued since yesterday into the electronic medical record."
[1278] In this way, the system automates the process from voice input to final medical record storage, enabling medical professionals to efficiently and accurately record patient information.
[1279] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1280] Step 1: Collecting voice input
[1281] The user listens to the patient's symptoms and medical history and engages in a conversation. The device uses a built-in microphone to collect this audio in real time and records it as audio data. The input is the conversation between the user and the patient, and the output is digital audio data. Specifically, the medical professional asks, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday." This audio is collected by the device.
[1282] Step 2: Sending audio data
[1283] The audio data collected by the device is sent to a server via the Internet. The input is digital audio data, and the output is audio data transferred to the server. Specifically, a program on the device sends the audio data to the server using the HTTPS protocol. This ensures that the data arrives securely at the server.
[1284] Step 3: Speech recognition processing
[1285] The server passes the received voice data to a voice recognition engine, which converts the voice into text data. The input is the voice data sent to the server, and the output is text data. Specifically, the server calls a voice recognition engine (for example, a natural language processing library) and converts the voice data "I have a headache, and it's been going on since yesterday" into the text "I have a headache, and it's been going on since yesterday."
[1286] Step 4: Text analysis and keyword extraction
[1287] The server analyzes the converted text data and extracts related keywords. The input is text data obtained from the speech recognition engine, and the output is the extracted keywords. Specifically, the server uses natural language processing technology to extract keywords such as "headache," "yesterday," and "continuing" from the text "I have a headache, and it has been going on since yesterday."
[1288] Step 5: Auto-completion
[1289] The server automatically completes the medical record contents based on the extracted keywords. The input is the extracted keywords, and the output is the automatically completed medical record contents. Specifically, the server references the medical database and adds information such as "Date of onset: yesterday. Pain level: moderate" to the medical record based on the keywords "headache," "yesterday," and "continuing."
[1290] Step 6: Send and view medical records
[1291] The server sends the auto-completed medical record contents to the terminal and displays them to the user. The input is the auto-completed medical record contents, and the output is the medical record information displayed on the terminal. Specifically, the server sends the medical record contents to the terminal via the HTTPS protocol, and the terminal displays the information on the user interface.
[1292] Step 7: Check and correct the medical record
[1293] The user checks the medical record displayed on the terminal and makes any necessary corrections. The input is the medical record information displayed on the terminal, and the output is the medical record content corrected by the user. Specifically, the doctor changes the "pain level" on the terminal screen from "moderate" to "severe."
[1294] Step 8: Save the medical record
[1295] The user confirms the medical record contents after making the edits, and the terminal sends them back to the server. The server saves the received medical record contents in the patient's electronic medical record. The input is the edited medical record contents, and the output is the information saved in the electronic medical record. Specifically, when the user presses the "Save" button on the terminal, the terminal sends the edited contents to the server, and the server finally records the data "Date of onset: yesterday. Pain level: severe" in the electronic medical record.
[1296] Through these steps, this invention automates the process from voice input to final medical record storage, reducing the burden on medical professionals and achieving efficient and accurate information management.
[1297] (Application example 1)
[1298] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1299] In the traditional medical record creation process, medical professionals manually record patient information, consuming a great deal of time and effort. Human errors, such as clerical errors and omissions, are also common. Meanwhile, in autonomous vehicles, systems for responding quickly and appropriately to vehicle abnormalities or emergencies may be inadequate, potentially reducing passenger safety and vehicle operational efficiency. A system that solves these problems and improves automation and efficiency is needed.
[1300] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1301] In this invention, the server includes means for acquiring oral information from the patient as voice data, means for transmitting the voice data to a data processing device, means for converting the voice data into text data in the data processing device, means for analyzing the text data and automatically completing the medical record content, means for displaying the automatically completed medical record content and allowing the user to confirm and correct it, means for transmitting the confirmed and corrected content to the data processing device and saving it, means for collecting information on the vehicle condition and external environment from the conversation content of the occupants, and means for analyzing the collected information and proposing countermeasures if an abnormality is detected. This enables improved efficiency and accuracy in creating medical records in the medical field, and enables quick and appropriate responses to abnormalities and emergencies in self-driving vehicles.
[1302] "Patient" refers to anyone who uses a medical institution or healthcare system.
[1303] "Oral information" refers to information or content that is spoken aloud.
[1304] "Audio data" is data converted from audio into digital form.
[1305] A "data processing device" is a machine or system that processes voice data, converts it into text data, and analyzes it.
[1306] "Text data" refers to written information in digital form.
[1307] "Analysis" is the process of extracting meaning and keywords from text data and organizing the information.
[1308] A "medical record" is a document that records a patient's medical treatment at a medical institution.
[1309] "Auto-completion" is the process by which the system automatically fills in missing data based on retrieved information.
[1310] A "user" is someone who operates the system and reviews and modifies the results.
[1311] "Display" means the visual presentation of information through a data processing device or other output device.
[1312] "Occupant" refers to any person riding in an autonomous vehicle.
[1313] "Conversation content" refers to verbal exchanges between multiple people.
[1314] "Vehicle status" is information indicating the operating status of the vehicle and the operating status of each function.
[1315] The "external environment" refers to the surrounding circumstances and conditions in which the vehicle is traveling.
[1316] "Collection" is the act of gathering specific information.
[1317] An "abnormality" is an event that indicates a state or malfunction that is different from the normal state.
[1318] A "solution" is a method or means for dealing with a particular situation.
[1319] As an embodiment of the present invention, there is provided a system in which a voice recognition system is installed in an autonomous driving vehicle, collects information on the vehicle's condition and the external environment from the conversation of the occupants, and proposes countermeasures when an abnormality is detected. The server includes means for acquiring oral information from the patient as voice data, means for transmitting the voice data to a data processing device, means for converting the voice data into text data in the data processing device, means for analyzing the text data and automatically completing the medical record content, means for displaying the automatically completed medical record content and allowing the user to confirm and correct it, means for transmitting the confirmed and corrected content to the data processing device and saving it, means for collecting information on the vehicle's condition and the external environment from the conversation of the occupants, and means for analyzing the collected information and proposing countermeasures when an abnormality is detected.
[1320] Hardware and software used
[1321] The server processes the audio data using the following hardware and software:
[1322] Microphone: Microphones installed inside the vehicle are used to collect passenger conversations in real time.
[1323] Data processing device: Receives and processes voice data. Here, a cloud-based server plays a key role.
[1324] Speech recognition engine: Uses the Google Cloud Speech-to-Text API to convert voice data into text data.
[1325] Text analysis engine: Uses NLTK (Natural Language Toolkit) to analyze the collected text data and extract relevant keywords.
[1326] Database: AWS RDS (Relational Database Service) is used to store analysis results and medical record information.
[1327] Navigation API: Uses Google Maps API to get real-time traffic information and external environment data.
[1328] Data processing and calculation
[1329] 1. Voice collection: Microphones inside the vehicle collect the passengers' conversations as voice data.
[1330] 2. Speech Recognition: The collected voice data is sent to a data processing device and converted into text data using the Google Cloud Speech-to-Text API.
[1331] 3. Keyword Extraction: The text data is analyzed using NLTK to extract relevant keywords.
[1332] 4. Data matching and auto-completion: Based on the extracted keywords, they are matched with historical data stored in AWS RDS or data retrieved from the navigation API, and then auto-completion is performed.
[1333] 5. Display and correction: The auto-completed content is displayed on the vehicle's display, where the user can confirm and correct it.
[1334] 6. Storage: The final confirmed and corrected content is sent to the data processing device and stored in AWS RDS.
[1335] Specific examples
[1336] For example, if a passenger says, "The air conditioner is not working properly," the conversation is collected as voice data via a microphone. This voice data is then sent to a server, where a voice recognition engine converts it into text data such as "The air conditioner is not working properly." Next, a text analysis engine analyzes this text data and extracts keywords such as "air conditioner," "working," and "not good." Finally, the system checks the condition of the air conditioner, and if there is an abnormality, it suggests a solution such as "Air conditioner abnormality: Please check."
[1337] Prompt Sentence Examples
[1338] "As a voice-activated vehicle assistant, please analyze the following conversation and suggest an appropriate response. Conversation: 'The air conditioning isn't working. Please add a new restaurant to Maps.'"
[1339] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1340] Step 1:
[1341] Audio Collection
[1342] Subject: Terminal
[1343] How it works: A microphone installed in the vehicle collects passenger conversations in real time. For example, if a passenger says, "The air conditioning isn't working," the voice is collected as voice data through the microphone.
[1344] Input: Passenger voice
[1345] Output: Collected audio data
[1346] Step 2:
[1347] Sending audio data
[1348] Subject: Terminal
[1349] Specific operation: The collected voice data is sent to a data processing device (cloud server). This transmission is done in real time so that the voice data can be processed immediately.
[1350] Input: Collected audio data
[1351] Output: Audio data sent to the data processing device
[1352] Step 3:
[1353] Voice Recognition
[1354] Subject: Server
[1355] Specific operation: The data processing device uses the Google Cloud Speech-to-Text API to convert the received voice data into text data. For example, the voice saying "The air conditioner is not working" is converted into the text "The air conditioner is not working."
[1356] Input: Transmitted audio data
[1357] Output: Converted text data
[1358] Step 4:
[1359] Keyword extraction
[1360] Subject: Server
[1361] Specific operation: The server analyzes the converted text data using NLTK and extracts related keywords. For example, from the text "The air conditioner is not working," keywords such as "air conditioner" and "not working" are extracted.
[1362] Input: Converted text data
[1363] Output: Extracted keywords
[1364] Step 5:
[1365] Data matching and auto-completion
[1366] Subject: Server
[1367] Specific operation: Based on the extracted keywords, the server compares them with past data stored in AWS RDS and information obtained from the navigation API (Google Maps API) and performs auto-completion. For example, based on the keywords "air conditioner" and "not working," the server will automatically complete the search results by checking the air conditioner's status and suggesting a solution such as "Air conditioner malfunction: check."
[1368] Input: Extracted keywords
[1369] Output: Auto-completed solutions and information
[1370] Step 6:
[1371] View and Modify
[1372] Subject: Terminal
[1373] Specific operation: The supplemented information is displayed on the vehicle's display. The user (passenger) can check this information and make corrections as necessary. For example, the passenger can correct the information by saying, "I checked the air conditioning, and there is actually no problem."
[1374] Input: Auto-completed solutions and information
[1375] Output: Displayed completions and user corrections
[1376] Step 7:
[1377] Submitting and saving your modifications
[1378] Subject: Terminal
[1379] Specific operation: The content modified by the user is sent back to the data processing device and stored in AWS RDS, so that the latest modified information is reflected in the system and can be used as reference data in the future.
[1380] Input: User-modified content
[1381] Output: Saved modifications
[1382] In this way, the system achieves its objective by performing data processing and calculations based on the input data at each step and obtaining the final output.
[1383] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1384] As an embodiment of this invention, we provide a system that acquires oral information from conversations with patients in real time and automatically reflects it in medical records, as well as a system that recognizes the user's emotions by combining an emotion engine. This system combines voice recognition technology, an automatic medical record completion function based on text data, and an emotion recognition function to reduce the burden on medical professionals and improve work efficiency.
[1385] Program processing overview
[1386] Voice input
[1387] The user interviews the patient about their symptoms and medical history. The device collects this conversation in real time as audio data. For example, a medical professional might ask, "What symptoms do you have?" The patient might reply, "I have a headache, and it's been going on since yesterday." The device then records this conversation as audio data.
[1388] Voice Recognition
[1389] The recorded voice data is sent from the device to the server. The server uses a voice recognition engine to convert this voice data into text data. For example, the voice "I have a headache, and it's been going on since yesterday" is converted into text "I have a headache, and it's been going on since yesterday."
[1390] Auto-completion
[1391] The server analyzes the converted text and extracts keywords. Based on the extracted keywords, it references the patient's existing information and medical databases to automatically complete the information. For example, based on the keyword "headache," the information "Date of onset: yesterday, Pain level: moderate" is added.
[1392] emotion recognition
[1393] The voice data is also sent to the emotion engine. The server uses the emotion engine to analyze the user's emotions from the voice data. For example, the analysis may reveal that the user is feeling "acute anxiety." The emotion engine adds this to the text data and reflects it in the patient's medical record.
[1394] Check and fix
[1395] The automatically completed medical record contents and emotion analysis results are sent to the terminal and displayed to the user. The user can check the contents and make corrections as necessary. For example, a doctor may change the "pain level" from "moderate" to "severe." Emotional information is also checked, and appropriate responses for patients with strong anxiety are considered.
[1396] keep
[1397] Once the changes are confirmed, the device sends the revised medical record information and emotional information to the server. The server saves this as the patient's electronic medical record. For example, the following data is recorded in the electronic medical record: "I have a headache that has been going on since yesterday. Onset: yesterday. Pain level: severe. Emotion: acute anxiety."
[1398] This system automates everything from voice input to emotion recognition and final medical record storage. It not only enables medical professionals to efficiently and accurately record patient information, but also allows them to understand the patient's emotional state, enabling them to provide appropriate medical care. This invention is particularly useful in interactive medical environments that emphasize the patient's emotional state.
[1399] The processing flow will be explained below.
[1400] Step 1:
[1401] The user interviews the patient about their symptoms and medical history. For example, a medical professional asks, "What symptoms do you have?" The patient answers, "I have a headache, and it's been going on since yesterday."
[1402] Step 2:
[1403] The terminal collects the conversation between the user and the patient as voice data. The terminal starts recording and records the voice data in real time.
[1404] Step 3:
[1405] The device converts the recorded audio data into a digital format and sends it to the server, where it is compressed to optimize bandwidth.
[1406] Step 4:
[1407] The server analyzes the received voice data and converts it into text data using a voice recognition engine. For example, a voice saying "I have a headache, and it's been going on since yesterday" is converted into text "I have a headache, and it's been going on since yesterday."
[1408] Step 5:
[1409] The server analyzes the text data and extracts keywords. Based on the extracted keywords, the server references the patient's existing information and medical databases to automatically complete the information. For example, the keyword "headache" can be supplemented with the information "Date of onset: yesterday, Pain level: moderate."
[1410] Step 6:
[1411] The server also sends the voice data to the emotion engine.
[1412] Step 7:
[1413] The emotion engine on the server analyzes the user's emotions from the voice data. For example, the emotion analysis may reveal that the user is experiencing "acute anxiety."
[1414] Step 8:
[1415] The server adds the results of the emotion analysis to the text data and reflects it in the medical record. For example, information such as "Emotion: acute anxiety" is added.
[1416] Step 9:
[1417] The server then sends the automatically completed medical record information and the emotion analysis results to the terminal. For example, it might generate the following: "I have a headache that has been going on since yesterday. Onset: yesterday. Pain level: moderate. Emotion: acute anxiety."
[1418] Step 10:
[1419] The terminal displays the received medical record information and emotion information to the user, who can then check the displayed information and make corrections as necessary.
[1420] Step 11:
[1421] The user reviews the patient record information and makes corrections as needed, for example, changing the "pain level" from "moderate" to "severe." Emotional information is also reviewed, and measures are planned to address patients with high anxiety.
[1422] Step 12:
[1423] The terminal transmits the medical record information and emotion information corrected by the user to the server. The confirmed corrections are then transmitted to the server.
[1424] Step 13:
[1425] The server stores the corrected medical record information and emotion information received as the patient's electronic medical record. For example, information such as "I have a headache that has continued since yesterday. Onset: yesterday. Pain level: severe. Emotion: acute anxiety" is recorded in the electronic medical record.
[1426] Example 2
[1427] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1428] In current medical settings, it is difficult to accurately and quickly record a patient's oral information and simultaneously manage the patient's emotional state. When medical professionals record information to create medical records, work efficiency decreases and human error is likely to occur. In addition, additional effort is required to understand the patient's emotional state, making it difficult to improve the overall quality of medical care.
[1429] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1430] In this invention, the server includes means for acquiring oral information from the patient as voice data, means for transmitting the voice data to a central computer, means for converting the voice data to text data in the central computer, means for analyzing the text data and automatically completing the medical record content, means for analyzing the emotional state from the voice data in the central computer, means for displaying the automatically completed medical record content and the emotion analysis results so that the user can confirm and correct them, and means for transmitting the confirmed and corrected content to the central computer and saving it. This automates the entire process from voice input to text conversion, emotion analysis, and medical record updating and saving, reducing the burden on medical professionals and enabling them to understand the patient's emotional state, enabling appropriate medical care to be provided.
[1431] "Oral information" refers to verbal information such as symptoms and medical history that a patient tells a medical professional.
[1432] "Audio data" refers to audio files that digitally record oral information obtained from a patient.
[1433] The "central computer" is a server that receives, analyzes, and stores voice and text data sent from the terminal.
[1434] "Speech recognition software" is a program for converting voice data into text data.
[1435] "Text data" is data in which voice data is expressed as a string of characters.
[1436] A "medical record" is a medical document that details a patient's symptoms, medical history, and treatment.
[1437] "Automatic completion" is a process that automatically adds missing parts based on converted text data by referencing past data and databases.
[1438] "Emotion analysis" is the process of analyzing the patient's emotional state from the audio data and adding the results to the text data.
[1439] "User" refers to a medical professional who uses the system to review and modify patient information.
[1440] To implement this invention, voice input technology, voice recognition technology, text data analysis technology, emotion recognition technology, and record management technology are required. This system captures oral information from conversations with patients in real time and automatically reflects it in medical records. It also simultaneously analyzes the patient's emotional state and adds it to the medical record. In this way, it reduces the workload of medical professionals and enables them to manage patient information efficiently and accurately.
[1441] Collecting voice input
[1442] The device has a built-in microphone that collects conversations between patients and medical professionals in real time. For example, a medical professional might ask, "What symptoms do you have?" and the patient might reply, "I have a headache, and it's been going on since yesterday." This conversation is recorded by the device's microphone and saved as audio data.
[1443] Sending audio data
[1444] The device sends the collected voice data to a central computer (server) over the Internet, for example, securely using the HTTPS protocol, where it can be processed for speech recognition and emotion engines.
[1445] Converting audio data to text
[1446] The server uses speech recognition software such as Google Cloud Speech-to-Text to convert the transmitted voice data into text data. For example, the voice data "I have a headache, and it's been going on since yesterday" is converted into text data "I have a headache, and it's been going on since yesterday."
[1447] Text data analysis and auto-completion
[1448] The server analyzes the converted text data and extracts keywords. For example, it recognizes the keyword "headache." It then references the patient's existing information and medical databases to automatically supplement the information with relevant information. For example, information such as "Date of onset: yesterday" and "Pain level: moderate" is added.
[1449] Conducting sentiment analysis
[1450] The server uses emotion recognition software such as IBM Watson Tone Analyzer to analyze the patient's emotional state from the voice data. For example, it obtains emotional information such as "acute anxiety" as a result of the analysis and adds this information to the text data.
[1451] Check and fix
[1452] The device displays the automatically completed medical record and emotion analysis results to the user. The user can review the information and make corrections as necessary. For example, a doctor can change the "pain level" from "moderate" to "severe," and then review the emotion information to consider the necessary response.
[1453] Save the final data
[1454] Once the changes are confirmed, the device sends the revised medical record information and emotion information back to the server. The server saves this as the patient's electronic medical record. For example, the final data recorded would be "I have a headache that has been going on since yesterday. Onset: yesterday. Pain level: severe. Emotion: acute anxiety."
[1455] This system automates the entire process from voice input to text conversion, emotion analysis, and medical record updating and storage, reducing the burden on medical professionals and enabling efficient and accurate information management. It also makes it possible to grasp the patient's emotional state, enabling more appropriate medical care.
[1456] Specific examples and examples of prompts for generative AI models
[1457] 1. Conversation between healthcare providers and patients:
[1458] Medical professional: "What symptoms do you have?"
[1459] Patient: "I have a headache that has been going on since yesterday."
[1460] 2. Audio to text conversion:
[1461] Voice: "I have a headache that has been going on since yesterday."
[1462] Text: "I have a headache that has been going on since yesterday."
[1463] 3. Auto-completion:
[1464] Keyword: "headache"
[1465] Supplementary data: "Date of onset: yesterday" "Pain level: moderate"
[1466] 4. Emotion analysis:
[1467] Speech data → Sentiment analysis : As in, “Acute anxiety .”
[1468] 5. Modification:
[1469] User modification: "Pain level: Moderate" → "Severe"
[1470] 6. Save:
[1471] Medical record data: "Headache, ongoing since yesterday. Onset: yesterday. Pain level: severe. Emotions: acute anxiety."
[1472] Example prompts for generative AI models
[1473] "Write a program that generates text data from audio data, extracts keywords for auto-completion, and adds sentiment analysis using an emotion engine. Also, include a process to finally save this data as an electronic medical record."
[1474] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1475] Step 1:
[1476] Collecting voice input
[1477] The device has a built-in microphone that collects conversations between patients and medical professionals in real time. Specifically, the medical professional asks, "What symptoms do you have?" and the patient replies, "I have a headache, and it's been going on since yesterday." While this conversation is taking place, the device records audio data. The input is the audio of the conversation, and the output is digital audio data.
[1478] Step 2:
[1479] Sending voice data to the server
[1480] The device transmits the collected voice data to a central computer (server) via the Internet. For example, the data is transmitted securely using the HTTPS protocol. This allows it to be processed for speech recognition and emotion engines. The input is digital voice data, and the output is the voice data transmitted to the server.
[1481] Step 3:
[1482] Converting audio data to text
[1483] The server uses speech recognition software such as Google Cloud Speech-to-Text to convert the transmitted voice data into text data. For example, the voice data "I have a headache, and it's been going on since yesterday" is converted into text data "I have a headache, and it's been going on since yesterday." The input is the voice data received by the server, and the output is the conversation in text format.
[1484] Step 4:
[1485] Text data analysis and auto-completion
[1486] The server analyzes the converted text data and extracts keywords. For example, it recognizes the keyword "headache." It then references the patient's existing information and medical databases to automatically complete the relevant information. For example, information such as "onset of illness: yesterday" and "pain level: moderate" is added. The input is text data, and the output is the completed text data.
[1487] Step 5:
[1488] Conducting sentiment analysis
[1489] The server uses emotion recognition software such as IBM Watson Tone Analyzer to analyze the patient's emotional state from the voice data. For example, the analysis results in emotional information such as "acute anxiety" and adds this to the text data. The input is voice data, and the output is text data with emotion analysis added.
[1490] Step 6:
[1491] Check and fix
[1492] The device displays the automatically completed medical record content and the emotion analysis results to the user. The user checks this content and makes corrections as necessary. For example, a doctor may change the "pain level" from "moderate" to "strong," check the emotion information, and consider the necessary response. The input is the completed text data and the emotion analysis results, and the output is the corrected medical record data.
[1493] Step 7:
[1494] Save the final data
[1495] Once the corrections are confirmed, the device sends the corrected medical record content and emotion information back to the server. The server saves this as the patient's electronic medical record. For example, the final data recorded is "I have a headache that has been going on since yesterday. Onset date: yesterday. Pain level: severe. Emotion: acute anxiety." The input is the corrected medical record data, and the output is the final data saved in the patient's electronic medical record.
[1496] (Application example 2)
[1497] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1498] Conventional food delivery services face the problem of difficulty in analyzing customer sentiment when taking orders, making it difficult to improve service quality. In particular, when customers are in a hurry or have special requests, they often cannot make prompt and appropriate suggestions, which can lead to a decline in customer satisfaction.
[1499] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1500] In this invention, the server includes means for acquiring dictation information from a customer as voice data, means for transmitting the voice data to the server, means for converting the voice data to text data in the server, means for analyzing the text data and automatically completing the order details, means for displaying the automatically completed order details and emotion analysis results for confirmation and correction by the user, means for transmitting the confirmed and corrected details to the server and saving them, means for analyzing the user's emotions from the voice data, and means for making appropriate suggestions based on the emotion analysis results, thereby enabling quick and accurate order confirmation and suggestions that take customer emotions into consideration.
[1501] "Patient" means an individual receiving medical services.
[1502] "Dicted information" is information provided by the user through speech.
[1503] "Audio data" refers to data in which audio is recorded in digital format.
[1504] A "server" is a computer system that processes and stores data.
[1505] "Text data" refers to data obtained by converting voice data into character information.
[1506] A "medical record" is a digital or paper-based document for recording medical information.
[1507] "Auto-completion" is the process of automatically adding missing data or content based on existing information.
[1508] A "displaying means" is a device or software that presents digital information to a user.
[1509] "Verify and correct" is the process by which a user checks the accuracy of information and corrects it if necessary.
[1510] "Emotion analysis" is a process of analyzing emotional states from voice data.
[1511] The "means for making suggestions" is a means for providing appropriate advice or recommendations to the user based on the analyzed information.
[1512] A "digital format" is a way of representing information in electronic form.
[1513] A "voice recognition engine" is software or hardware for converting voice data into text data.
[1514] This invention is a system that acquires oral information from patients or customers (hereinafter referred to as users) in real time, analyzes and supplements that information, and provides appropriate services quickly and accurately. This system combines technologies such as voice recognition technology, emotion analysis engine, text analysis, and recommendation engine.
[1515] Hardware and software used
[1516] 1. Hardware:
[1517] Devices such as smartphones, smart glasses, and head-mounted displays
[1518] Cloud Server
[1519] 2. Software:
[1520] Speech recognition engine: Google Speech-to-Text API, Amazon Transcribe, etc.
[1521] Emotion recognition engine: IBM Watson Tone Analyzer, Microsoft Azure Emotion API, etc.
[1522] Backend: Node.js, Python (Flask)
[1523] Database: MongoDB, Firebase
[1524] Program processing overview
[1525] Voice input
[1526] The user dictates information into the terminal. For example, in a food delivery service, the user might say, "I'd like to order a hamburger set, and a Coke to drink."
[1527] Voice Recognition
[1528] The voice data collected by the device is sent to a cloud server. The server uses a voice recognition engine to convert this voice data into text data. For example, a voice saying "I'd like to order a hamburger set, and a Coke to drink" is converted into the text "Hamburger set, Coke."
[1529] Auto-completion
[1530] The server analyzes the converted text and automatically completes the order details. For example, it generates an order list such as "Hamburger Set (Contents: Cheeseburger, Fries) + Coke." This automatically adds detailed information according to the user's request.
[1531] emotion recognition
[1532] The voice data is also sent to an emotion analysis engine, which analyzes the user's emotions. For example, the engine may detect that the user is in a hurry. Based on this, a menu of options for responding quickly is presented.
[1533] Check and fix
[1534] The auto-completed order details and the sentiment analysis results are displayed on the device, and the user can confirm and modify them. For example, the user can confirm an order displayed as "Hamburger Set (Cheeseburger, Fries) + Coke." At this time, the user can make modifications as necessary.
[1535] keep
[1536] The final order details and emotion data after confirmation and correction are sent to the server and stored in the database. For example, data such as "Hamburger set (cheeseburger, fries) + Coke. Emotion: In a hurry" is stored.
[1537] Examples of concrete examples and prompts
[1538] For example, suppose a user inputs dictation information such as "This hamburger looks very tasty. I'd like some fries with it, please."
[1539] Speech recognition prompt: transcribe audio
[1540] Emotion recognition prompt: analyze sentiment from text: "This burger looks delicious. I'd like some fries with it, please."
[1541] Recommendation prompt: get recommendations for quick service items based on "interest"
[1542] By implementing this invention, users can enjoy a fast and accurate ordering experience through voice input, and service providers can respond appropriately based on the user's emotions.
[1543] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1544] Step 1:
[1545] The user inputs dictation information into the terminal.
[1546] Specifically, the user speaks to a device such as a smartphone or smart glasses, saying, "I'd like to order a hamburger set, and a Coke to drink."
[1547] The input data is stored in the terminal as voice data.
[1548] The data that is output is the collected voice data.
[1549] Step 2:
[1550] The device sends the collected voice data to a cloud server.
[1551] Specifically, the terminal program uploads the collected voice data to a designated cloud server via the Internet.
[1552] The input data is audio data.
[1553] The output data is the audio data sent to the cloud server.
[1554] Step 3:
[1555] The server uses a speech recognition engine to convert the voice data into text data.
[1556] Specifically, the server calls a speech recognition engine such as the Google Speech-to-Text API or Amazon Transcribe to convert the voice data into text data.
[1557] The input data is audio data.
[1558] The output data is the converted text data.
[1559] Step 4:
[1560] The server analyzes the converted text data and automatically completes the order details.
[1561] Specifically, the server program analyzes the text data, extracts and completes the order details. For example, from the text data "hamburger set, cola," it generates "hamburger set (cheeseburger, fries) + cola."
[1562] The input data is converted text data.
[1563] The output data is the supplemented order details.
[1564] Step 5:
[1565] The voice data is sent to an emotion analysis engine to analyze the user's emotions.
[1566] Specifically, the server uses an emotion analysis engine such as IBM Watson Tone Analyzer or Microsoft Azure Emotion API to analyze the emotional state from the voice data.
[1567] The input data is audio data.
[1568] The output data is the emotion analysis results.
[1569] Step 6:
[1570] The automatically completed order details and sentiment analysis results are displayed on the terminal.
[1571] Specifically, the server sends the generated order details and the emotion analysis results to the terminal, which then displays them to the user. For example, the terminal displays "Hamburger set (cheeseburger, fries) + Coke. Emotion: In a hurry."
[1572] The input data is the supplemented order details and the results of sentiment analysis.
[1573] The output data is the information that is displayed on the terminal.
[1574] Step 7:
[1575] The user reviews and corrects what is displayed.
[1576] Specifically, the user checks the displayed order details and emotion information and makes corrections as necessary. For example, the user may make a correction such as "change fries to large size."
[1577] The data to be input is the information displayed on the terminal.
[1578] The data that is output is the content that has been confirmed and corrected by the user.
[1579] Step 8:
[1580] The confirmed and corrected content is sent to the server and stored in the database.
[1581] Specifically, the terminal sends the confirmed and corrected content back to the cloud server, and the server stores it in the database.
[1582] The data to be input is the content that has been confirmed and corrected by the user.
[1583] The output data is the order details and emotion information stored in the database.
[1584] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1585] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1586] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1587] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1588] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1589] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1590] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1591] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1592] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1593] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1594] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1595] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1596] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1597] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1598] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1599] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1600] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1601] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1602] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1603] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1604] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1605] The following is further disclosed regarding the above embodiment.
[1606] (Claim 1)
[1607] A means for acquiring oral information from a patient as voice data;
[1608] means for transmitting the voice data to a server;
[1609] means for converting voice data into text data in the server;
[1610] means for analyzing the text data and automatically completing the contents of the medical record;
[1611] a means for displaying the automatically completed medical record contents and allowing the user to confirm and correct them;
[1612] The system includes a means for transmitting and storing the confirmed and corrected content on a server.
[1613] (Claim 2)
[1614] 10. The system of claim 1, further comprising means for transmitting said audio data in digital form to a server.
[1615] (Claim 3)
[1616] 2. The system according to claim 1, further comprising means for converting voice data into text data using a voice recognition engine in the server.
[1617] "Example 1"
[1618] (Claim 1)
[1619] A means for acquiring oral information from a patient as voice data;
[1620] means for transmitting the voice data to a server;
[1621] means for converting voice data into text data using a voice recognition engine in the server;
[1622] means for analyzing the text data and extracting related keywords;
[1623] means for automatically completing the contents of the medical record based on the keywords;
[1624] a means for transmitting the automatically completed medical record contents to a terminal and allowing a user to confirm and correct the contents;
[1625] The system includes a means for transmitting and storing the confirmed and corrected content on a server.
[1626] (Claim 2)
[1627] 10. The system of claim 1, further comprising means for transmitting the audio data in digital form to a server.
[1628] (Claim 3)
[1629] 2. The system according to claim 1, further comprising means for analyzing text data using natural language processing technology and extracting keywords.
[1630] "Application Example 1"
[1631] (Claim 1)
[1632] A means for acquiring oral information from a patient as voice data;
[1633] means for transmitting the audio data to a data processing device;
[1634] means for converting voice data into text data in the data processing device;
[1635] means for analyzing the text data and automatically completing the contents of the medical record;
[1636] a means for displaying the automatically completed medical record contents and allowing the user to confirm and correct them;
[1637] means for transmitting and storing said confirmed and corrected content in a data processing device;
[1638] A means for collecting information on the vehicle's condition and external environment from the conversations of the occupants;
[1639] A system including a means for analyzing the collected information and proposing countermeasures when an abnormality is detected.
[1640] (Claim 2)
[1641] 10. The system of claim 1, further comprising means for transmitting said audio data in digital form to a data processing device.
[1642] (Claim 3)
[1643] 2. The system according to claim 1, further comprising means for converting voice data into text data using a voice recognition engine in the data processing device.
[1644] "Example 2: Combining Emotion Engines"
[1645] (Claim 1)
[1646] A means for acquiring oral information from a patient as voice data;
[1647] means for transmitting said audio data to a central computer;
[1648] means for converting voice data into text data at said central computer;
[1649] means for analyzing the text data and automatically completing the medical record contents;
[1650] means for analyzing an emotional state from the speech data at said central computer;
[1651] a means for displaying the automatically completed medical record content and emotion analysis results, and allowing the user to confirm and correct them;
[1652] The system includes means for transmitting and storing said verified and corrected content in a central computer.
[1653] (Claim 2)
[1654] 10. The system of claim 1, further comprising means for transmitting said audio data in digital form to a central computer.
[1655] (Claim 3)
[1656] 10. The system of claim 1, further comprising means for converting voice data into text data using voice recognition software at said central computer.
[1657] "Application example 2 when combining emotion engines"
[1658] (Claim 1)
[1659] A means for acquiring oral information from a patient as voice data;
[1660] means for transmitting the voice data to a server;
[1661] means for converting voice data into text data in the server;
[1662] means for analyzing the text data and automatically completing the contents of the medical record;
[1663] a means for displaying the automatically completed medical record contents and emotion analysis results, and allowing the user to confirm and correct them;
[1664] means for transmitting and storing the confirmed and corrected content to a server;
[1665] A means for analyzing user emotions from voice data;
[1666] The system includes a means for making appropriate suggestions based on the emotion analysis results.
[1667] (Claim 2)
[1668] 10. The system of claim 1, further comprising means for transmitting said audio data in digital form to a server.
[1669] (Claim 3)
[1670] 2. The system according to claim 1, further comprising means for converting voice data into text data using a voice recognition engine in the server. [Explanation of symbols]
[1671] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for acquiring oral information from a patient as voice data; means for transmitting the voice data to a server; means for converting voice data into text data in the server; means for analyzing the text data and automatically completing the contents of the medical record; a means for displaying the automatically completed medical record contents and allowing the user to confirm and correct them; The system includes a means for transmitting and storing the confirmed and corrected content on a server.
2. 10. The system of claim 1, further comprising means for transmitting said audio data in digital form to a server.
3. 2. The system according to claim 1, further comprising means for converting voice data into text data using a voice recognition engine in the server.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A