System
The system addresses data management and security challenges in the medical industry by digitizing records, using generative AI, and ensuring privacy, enhancing diagnostic accuracy and resource allocation.
Patent Information
- Application Number
- JP2024118129
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-02-04
AI Technical Summary
The medical industry faces challenges in efficiently managing fragmented data from paper and electronic medical records, insufficient medical staff training, and inadequate data security and privacy, which hinders accurate diagnosis and resource allocation.
A system that digitizes medical records, uses generative AI models for data management, anonymizes patient data, and ensures security through access control and encryption, supporting diagnostic predictions and resource allocation.
Enables efficient management and utilization of medical records, supports medical staff, and strengthens data security, improving diagnostic accuracy and resource allocation.
Smart Images

Figure 2026017347000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In the medical industry, there is a large amount of fragmented data, such as paper medical records and electronic records, and efficiently managing and utilizing this data is a challenge. Furthermore, medical staff training and education are insufficient, making it difficult to efficiently diagnose and allocate medical resources appropriately. Furthermore, data management is difficult from the perspectives of privacy and security, requiring appropriate security measures. In response to these challenges, the objective of this invention is to provide a system that digitizes medical records and uses generative AI models to efficiently and securely manage and utilize data. [Means for solving the problem]
[0005] The present invention provides a system that solves problems in the medical industry by the following means. First, it provides a means for digitizing past medical records and storing them in a database. Next, it provides a means for learning medical datasets using a generative AI model and supporting patient interviews through voice recognition. It also has a means for anonymizing patient data to enhance security, and includes a means for supporting diagnostic prediction and the appropriate allocation of medical resources. It also has a means for ensuring data security through access control and data encryption. This enables efficient management and utilization of medical records, supports medical staff, and strengthens data security.
[0006] "Past medical records" refers to data such as a patient's medical history, diagnosis results, prescriptions, and test results previously created by a medical institution.
[0007] "Digitization" is the process of converting information stored on paper or in other non-digital formats into electronic data format.
[0008] A "database" is a system for structuring and storing large amounts of digital information so that it can be efficiently managed, searched, and used.
[0009] A "generative AI model" is a mathematical model that uses artificial intelligence technology to learn from large amounts of data and make predictions and classifications.
[0010] "Speech recognition" is a technology that converts speech into text data, making it possible to give instructions to computers and systems through speech.
[0011] "Patient interview" is the process in which medical professionals directly ask patients about their symptoms, medical history, lifestyle habits, and other information.
[0012] "Anonymization" is the process of removing or transforming personally identifiable information to make the data subject anonymous.
[0013] "Security" refers to the technology and means for protecting information and data from unauthorized access, tampering, leakage, etc.
[0014] "Diagnostic prediction" is a technology that predicts future health conditions and the likelihood of disease based on a patient's symptoms and medical data.
[0015] "Medical resources" is a general term for the goods, equipment, medicines, and personnel required to provide medical services.
[0016] "Access control" is a management technique for restricting access to data and systems to authorized users and preventing unauthorized access.
[0017] "Data encryption" is a technology that transforms data using a specific algorithm so that only authorized users can decrypt it. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] This invention is a system that promotes digitalization and cloud computing in the medical industry. Below, we will create a program for this system and explain its processing in natural language. Specific examples will also be provided.
[0040] Data collection and digitization
[0041] User:
[0042] First, users collect the patient's paper medical records or existing electronic medical records, which are then digitized by scanning them, and then upload the digitized data to the server via a terminal.
[0043] Device:
[0044] The terminal receives the scanned paper medical records and converts them into text data using OCR (optical character recognition) technology, which is then sent to the server.
[0045] server:
[0046] The server converts the received text data into an appropriate database format and stores it in the database, allowing past medical records to be digitized and efficiently managed.
[0047] Training generative AI models
[0048] server:
[0049] The server extracts historical patient data and medical records from the database and divides them into training and validation datasets for the generative AI model. The training dataset is used to train the AI model, completing the model learning process. This process also includes adjusting hyperparameters such as the number of epochs and batch size. The model's performance is evaluated and adjustments are made as necessary.
[0050] Building a voice support system
[0051] Device:
[0052] The device uses a speech recognition API to enable voice input. When the user interviews a patient, the device converts the speech into text data in real time. It also generates optimal questions based on a pre-set question list and past response history.
[0053] server:
[0054] The server receives voice data in real time, converts it into text using speech recognition technology, and then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user.
[0055] Data partitioning and sharing
[0056] server:
[0057] The server will appropriately anonymize patient data to enhance privacy and security, and will partition and share this data with relevant medical institutions and researchers as needed, with security protocols in place and appropriate access rights managed.
[0058] Diagnostic prediction and optimal allocation of medical resources
[0059] server:
[0060] When new patient data is input, the server uses the generative AI model to predict diagnoses. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes them as necessary.
[0061] Security and Privacy Measures
[0062] server:
[0063] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology to prevent unauthorized access and data leaks. In addition, security audits are conducted regularly to ensure the safety of the system.
[0064] Specific examples
[0065] 1. The user scans Patient A's paper medical record and uploads it to the server via their device. The server converts it into text data using OCR technology and stores it in the database.
[0066] 2. When interviewing new patient B, the user uses voice input, and the device generates appropriate questions and converts them into text in real time. The server analyzes them using a generative AI model and provides feedback to the user.
[0067] 3. The server makes a diagnosis prediction and diagnoses Patient C as having a high possibility of influenza. The server checks the inventory of necessary medical equipment and medicines and allocates them appropriately.
[0068] 4. Patient data is anonymized and securely shared with other medical institutions for research purposes. The server applies access controls and data encryption to ensure privacy and security.
[0069] As described above, by specifically implementing the present invention, it is possible to realize efficient management and utilization of medical records, support for medical staff, and strengthen data security.
[0070] The processing flow will be explained below.
[0071] Step 1:
[0072] The user collects the patient's paper medical record or existing electronic medical record and scans the paper medical record, which then digitizes the data and loads it into the terminal.
[0073] Step 2:
[0074] The terminal converts scanned paper chart images into text data using OCR (Optical Character Recognition) technology, which is then sent to a server for further processing.
[0075] Step 3:
[0076] The server converts the received text data into an appropriate database format and stores it in the database. This standardizes the data and enables efficient management.
[0077] Step 4:
[0078] The server extracts historical patient data and medical records from the database and splits them into training and validation datasets.
[0079] Step 5:
[0080] The server trains the generative AI model using the training dataset, sets hyperparameters such as the number of epochs and batch size, and executes the model learning process.
[0081] Step 6:
[0082] The server evaluates the performance of the trained model using a validation dataset and adjusts the model as needed.
[0083] Step 7:
[0084] The device uses the speech recognition API to enable the voice input function, preparing the device for the user to use voice input.
[0085] Step 8:
[0086] When interviewing patients, the device automatically generates the most appropriate questions based on a pre-set list of questions and past response history.
[0087] Step 9:
[0088] The device converts the user's voice input into text in real time and immediately sends the text to the server.
[0089] Step 10:
[0090] The server analyzes the text data sent and uses a generative AI model to provide appropriate feedback to the user.
[0091] Step 11:
[0092] The server anonymizes patient data, removing or transforming personally identifiable information to protect the confidentiality of the data.
[0093] Step 12:
[0094] The server will divide the anonymized data and share it with medical institutions and researchers who need it, and will also set up security protocols and manage access rights.
[0095] Step 13:
[0096] When new patient data is entered into the server, the generative AI model is used to make a diagnostic prediction, and the results are notified to the user in real time.
[0097] Step 14:
[0098] The server checks the inventory status of medical resources (medical equipment and medicines) based on patient data and diagnosis predictions, and allocates or replenishes them as necessary.
[0099] Step 15:
[0100] The server manages the access control list (ACL) and sets the access rights for each user. Data is protected using AES-256 encryption technology.
[0101] Step 16:
[0102] We regularly conduct security audits on our servers to ensure the safety of our systems. In the event of a security incident, we respond immediately and take corrective measures.
[0103] Example 1
[0104] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0105] Problems with conventional medical systems include the lack of progress in digitizing medical records and the ineffective use of collected data. Furthermore, there is a lack of systems that support the accuracy of diagnostic predictions and the appropriate allocation of medical resources. There are also many shortcomings in terms of security and privacy protection. These issues are major obstacles to improving the efficiency and quality of medical care, and they require rapid and accurate responses.
[0106] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0107] In this invention, the server includes a means for scanning and digitizing past medical records, converting them into text data using OCR technology, and storing them in a database, a means for extracting and training data sets using a generative AI model, and a means for anonymizing patient data and using technology to enhance security, thereby enabling efficient management and utilization of medical records, improving the accuracy of diagnostic predictions, and optimizing the allocation of medical resources.
[0108] A "database format" is a structure or format for efficiently storing, retrieving, and managing digital data.
[0109] "OCR technology" is an abbreviation for optical character recognition technology, which extracts character information from scanned images and converts it into text data.
[0110] A "generative AI model" refers to an algorithm or model that uses artificial intelligence techniques to generate new information or predictions from a given dataset.
[0111] "Voice recognition technology" is a technology that converts voice input into text data in real time.
[0112] "Data anonymization" refers to the process of removing or modifying personally identifiable information to protect the privacy of the data.
[0113] A "security protocol" is a set of procedures and rules that are used to prevent unauthorized access when transmitting or receiving data.
[0114] "Diagnostic prediction" is the process of using generative AI models based on medical data to predict illness and treatment.
[0115] An "access control list (ACL)" is a list used to manage the rights and permissions of each user accessing a system.
[0116] "Data encryption" is a technology that encrypts important information using a specific algorithm to prevent unauthorized access and data leaks.
[0117] "Hyperparameter tuning" is the process of fine-tuning the parameters of a machine learning model to optimize its learning efficiency and performance.
[0118] The present invention is a system that promotes digitalization and cloud computing in the medical industry. The program processing of this system will be explained below in natural language.
[0119] Data collection and digitization
[0120] User:
[0121] First, the user collects the patient's paper medical records or existing electronic medical records. They use a general-purpose scanner to scan the paper medical records and convert them into digital data. For example, they use a general-purpose high-performance scanner. Next, the user uploads the scanned digital data to a server via a terminal. This is done using the hospital's internal network.
[0122] Device:
[0123] The terminal receives the scanned image data of the paper medical record and converts it into text data using OCR technology. This technology uses commonly used OCR software. The converted text data is sent to the server. When sending, the communication is encrypted using SSL / TLS.
[0124] server:
[0125] The server receives the text data sent from the terminal. The received data is converted into an appropriate database format and stored in a database. For example, a relational database management system (RDBMS) is used as the database.
[0126] Training generative AI models
[0127] server:
[0128] The server extracts past patient data and medical records from the database. The extracted data is divided into a training dataset and a validation dataset for the generative AI model. The server trains the generative AI model using the training dataset. Deep learning libraries such as PyTorch and TensorFlow are used for this training. Hyperparameters such as the number of epochs and batch size are also adjusted. The model's performance is evaluated on the validation dataset, and readjustments are made as necessary.
[0129] Building a voice support system
[0130] Device:
[0131] The device uses a speech recognition API to enable voice input. For example, a general speech recognition API can be used. When the user interviews a patient, the device converts the speech into text data in real time. It also generates optimal questions based on a pre-set question list and past response history.
[0132] server:
[0133] The server receives the voice data in real time, converts it into text using speech recognition technology, and then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user.
[0134] Data partitioning and sharing
[0135] server:
[0136] The server will appropriately anonymize patient data to enhance privacy and security. This will involve techniques to remove or modify personally identifiable information. Furthermore, the anonymized data will be split up as needed and shared with relevant medical institutions and researchers. This sharing will be subject to security protocols and appropriate access rights management.
[0137] Diagnostic prediction and optimal allocation of medical resources
[0138] server:
[0139] When new patient data is input, the server uses the generative AI model to predict diagnoses. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes these resources as needed.
[0140] Security and Privacy Measures
[0141] server:
[0142] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology to prevent unauthorized access and data leaks. In addition, security audits are conducted regularly to ensure the safety of the system.
[0143] Specific examples
[0144] 1. The user scans Patient A's paper medical record and uploads it to the server via their device. The server converts it into text data using OCR technology and stores it in the database.
[0145] 2. When interviewing new patient B, the user uses voice input, and the device generates appropriate questions and converts them into text in real time. The server analyzes them using a generative AI model and provides feedback to the user.
[0146] 3. The server makes a diagnosis prediction and diagnoses Patient C as having a high possibility of influenza. The server checks the inventory of necessary medical equipment and medicines and allocates them appropriately.
[0147] 4. Patient data is anonymized and securely shared with other medical institutions for research purposes. The server applies access controls and data encryption to ensure privacy and security.
[0148] Prompt Sentence Examples
[0149] "How do you digitize and store Patient A's paper medical records?"
[0150] "How do I generate the best questions for a new patient, Patient B?"
[0151] "How can we predict diagnoses and allocate necessary medical resources appropriately?"
[0152] As described above, by specifically implementing the present invention, it is possible to realize efficient management and utilization of medical records, support for medical staff, and strengthen data security.
[0153] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0154] Step 1:
[0155] User: The user obtains digital image data by scanning a patient's paper chart. This input data is in an image file format (e.g., JPEG or PNG). The specific actions of scanning using a scanner include placing the paper chart on the scanner and pressing the Scan button. The output is the scanned digital image data.
[0156] Step 2:
[0157] User: The user uploads the scanned digital image data to the terminal. The uploading process involves sending the digital image to a designated folder using the hospital network. The input is the scanned digital image, and the output is the uploaded digital image in the terminal.
[0158] Step 3:
[0159] Terminal: The terminal converts the received digital image data into text data using OCR technology. Specific operations include launching the OCR software, reading the digital image, and performing character recognition. The input is digital image data, and the output is the converted text data.
[0160] Step 4:
[0161] Terminal: The terminal sends the text data generated by OCR to the server. This transmission process uses SSL / TLS encrypted communication. The input is the converted text data, and the output is the text data sent to the server.
[0162] Step 5:
[0163] Server: The server converts the received text data into a database format and stores it in the database. Specific operations include converting the text data into an SQL query format and performing an insert operation on the database. The input is the text data sent to the server, and the output is the medical record data stored in the database.
[0164] Step 6:
[0165] Server: The server extracts historical patient data and medical records from the database and splits them into training and validation datasets for the generative AI model. Specific operations include executing SQL queries to retrieve the data and splitting the extracted data into training and validation datasets. The input is the medical record data stored in the database, and the output is the split dataset.
[0166] Step 7:
[0167] Server: The server trains the generative AI model using the training dataset. This is done using deep learning libraries such as PyTorch or TensorFlow. It also adjusts hyperparameters such as the number of epochs and batch size. The input is the training dataset, and the output is the trained generative AI model.
[0168] Step 8:
[0169] Server: The server evaluates the model's performance on the validation dataset and adjusts hyperparameters or modifies the model structure as necessary. Specific operations include calculating metrics for model evaluation and retraining. The input is the validation dataset and the existing generative AI model, and the output is an optimized generative AI model.
[0170] Step 9:
[0171] Terminal: The terminal uses a speech recognition API to enable voice input and converts the user's voice during patient interviews into text data in real time. It also generates optimal questions based on a pre-set question list and past response history. The input is voice data, and the output is text data and the generated questions.
[0172] Step 10:
[0173] Server: The server receives voice data in real time and converts it into text data using speech recognition technology. It then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user. The input is voice data, and the output is interpreted text data and feedback.
[0174] Step 11:
[0175] Server: The server anonymizes patient data, removes or modifies any personally identifiable information from the data, and applies security protocols and manages access rights to share the anonymized data with relevant medical institutions and researchers. The input is raw patient data, and the output is anonymized data.
[0176] Step 12:
[0177] Server: When new patient data is input, the server uses the generative AI model to make a diagnosis prediction. The prediction results are notified to the user in real time, and at the same time, the server checks the inventory status of medical resources and allocates or replenishes them as needed. The input is the new patient data, and the output is the diagnosis prediction results and medical resource management information.
[0178] Step 13:
[0179] Server: The server manages the access control list (ACL) and sets the access rights for each user. It also protects data with AES-256 encryption technology and performs regular security audits to prevent unauthorized access and data leakage. The input is user information and data, and the output is controlled access rights and protected data.
[0180] Through these steps, the system achieves efficient digitization of medical records, diagnostic prediction through generative AI models, and data security management.
[0181] (Application example 1)
[0182] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0183] The modern healthcare industry requires the digitization of medical records and efficient data management. However, several challenges exist, including the digitization of paper medical records, the centralized management of existing electronic medical records, the anonymization and security of medical data, and the improvement of diagnostic prediction accuracy. In particular, there is a lack of support systems for quickly and accurately recording patient interactions and making appropriate diagnoses. It is also important to safely anonymize collected data and share it with other medical institutions and researchers. It is necessary to resolve these challenges and realize digitalization and enhanced security in the healthcare industry.
[0184] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0185] In this invention, the server includes means for digitizing past medical records and storing them in a database, means for training medical datasets using a generative AI model, means for supporting patient interviews using voice recognition, means for anonymizing patient data to enhance security, means for making diagnostic predictions and supporting the appropriate allocation of medical resources, means for ensuring data security through access control and data encryption, means for generating text data from scanned images using optical character recognition technology, means for converting voice input into text data in real time using a voice recognition API, and means for checking the availability of medical resources based on diagnostic predictions using the generative AI model. This enables efficient digitization and management of medical records, rapid recording of patient interactions, appropriate diagnostic support, and secure data sharing.
[0186] "Prior medical records" are previously collected medical data, including a patient's medical history and treatment history.
[0187] "Digitalization" is the process of converting information in analog form, such as paper, into digital form.
[0188] A "database" is a system for storing and managing data in an organized structure.
[0189] A "generative AI model" is a machine learning algorithm that learns from large amounts of data and is optimized to perform a specific task.
[0190] A "medical dataset" is a collection of various medical-related data used for learning and analysis.
[0191] "Speech recognition" is the technology that takes voice input and converts it into a digital format such as text.
[0192] A "patient interview" is a process in which a medical professional directly obtains information from a patient, such as symptoms and medical history.
[0193] "Anonymization" is the process of processing data so that individuals cannot be identified.
[0194] "Security" refers to a group of technologies aimed at protecting data and preventing unauthorized access.
[0195] "Diagnostic prediction" is the process of predicting future disease states and appropriate treatments based on collected medical data.
[0196] "Medical resources" is a general term for all equipment, medicines, personnel, etc. used in medical institutions.
[0197] "Proper allocation" means properly allocating the necessary resources to the necessary locations.
[0198] "Access control" is a mechanism that allows only authorized users to access specific data.
[0199] "Data encryption" is the technology that encrypts data and converts it into a form that is unintelligible to unauthorized users.
[0200] "Optical character recognition technology" is a technology that analyzes character information in an image and converts it into text data.
[0201] "Real-time" means that data input, processing, and output are carried out instantly.
[0202] A "speech recognition API" is an application programming interface that provides the functionality to convert voice input into text.
[0203] "Inventory status" is information that indicates the current level of medical resources available for use.
[0204] This invention is a system that promotes digitalization and cloud computing in the medical industry, and provides a wide range of functions, including digitization of past medical records, diagnostic prediction using generative AI models, patient interview support using voice recognition, and enhanced data security using anonymization technology.
[0205] Data collection and digitization
[0206] User:
[0207] Users first scan the patient's paper medical records using their smartphone camera and upload them to the server via the device, allowing them to be digitized and managed efficiently.
[0208] Application of OCR technology
[0209] Device:
[0210] The terminal uses optical character recognition (OCR) technology to generate text data from scanned images of paper medical records, which is then sent to a server and stored in a database.
[0211] Building a voice support system
[0212] Device:
[0213] The device uses a speech recognition API to convert voice input into text data in real time. The user records conversations with patients and converts the data into text in real time. The device also generates optimal questions based on a pre-set question list and past response history.
[0214] Data storage and anonymization
[0215] server:
[0216] The server converts the received text data into an appropriate database format and stores it in the database. The server anonymizes the data and uses data encryption technology (e.g., AES-256) to enhance privacy and security. This data is split as needed and shared with relevant medical institutions and researchers.
[0217] Training generative AI models
[0218] server:
[0219] The server extracts historical patient data and medical records from the database and divides them into training and validation datasets for the generative AI model. The training dataset is used to train the AI model, completing the model learning process. The model's performance is evaluated and adjusted as necessary.
[0220] Diagnostic prediction and optimal allocation of medical resources
[0221] server:
[0222] When new patient data is entered, the server uses the generative AI model to predict a diagnosis. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (medical equipment, drugs, etc.) and allocates or replenishes them as needed.
[0223] Examples of concrete examples and prompts
[0224] 1. A user scans a patient's paper chart and uploads it to the server via a terminal. An example of the prompt is as follows:
[0225] "Take images of medical records, digitize them, and upload them to the cloud."
[0226] 2. The user records the conversation with the patient and uses a speech recognition API to transcribe it in real time, with prompts such as:
[0227] "Record conversations with patients, convert them into text, and save them."
[0228] 3. The server performs diagnosis prediction, checks the availability of medical resources, and allocates them appropriately. An example of the prompt is as follows:
[0229] "Using generative AI models to predict diagnoses and allocate necessary medical resources"
[0230] This enables efficient digitization and management of medical records, enables prompt recording of patient interactions, and ensures proper diagnostic support and secure data sharing, significantly improving operational efficiency and security in the healthcare industry.
[0231] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0232] Step 1:
[0233] The user scans the patient's paper medical record using the smartphone camera and uploads it to the server via the terminal. The input is an image of the paper medical record, and the output is digitized text data. The user launches the scanning app on the terminal and takes a picture of the paper medical record.
[0234] Step 2:
[0235] The device generates text data from the scanned image using optical character recognition (OCR). The input is the scanned image from step 1, and the output is the text data generated by OCR. The device sends the scanned image to the OCR engine and obtains the text data.
[0236] Step 3:
[0237] The device sends the generated text data to the server. The input is the text data, and the output is the data sent to the server. The device then uploads the data to the cloud server via the network.
[0238] Step 4:
[0239] The server converts the received text data into an appropriate database format and stores it in the database. The input is the text data received by the server, and the output is the data stored in the database. The server formats the text data into the database format and stores it.
[0240] Step 5:
[0241] The server uses a speech recognition API to convert the voice data recorded by the user into text data in real time. The input is voice data and the output is text data. The server sends the voice data to the API and saves the resulting text data.
[0242] Step 6:
[0243] The server anonymizes the received data and uses data encryption techniques (e.g., AES-256) to enhance privacy and security. The input is text data stored in a database, and the output is anonymized and encrypted data. The server anonymizes the data and applies encryption techniques.
[0244] Step 7:
[0245] The server uses a generative AI model to split the past patient data and medical records extracted from the database into a training dataset and a validation dataset. The input is the past patient data and medical records, and the output is the training dataset and the validation dataset. The server splits the data into training and validation datasets.
[0246] Step 8:
[0247] The server trains the generative AI model using the training dataset, completing the model learning process. The input is the training dataset, and the output is the trained generative AI model. The server uses the dataset to train the model.
[0248] Step 9:
[0249] When new patient data is input, the server uses the generative AI model to make a diagnosis prediction. The input is the new patient data, and the output is the diagnosis prediction result. The server inputs the new patient data into the model and obtains the diagnosis prediction result.
[0250] Step 10:
[0251] The server checks the inventory status of medical resources (medical equipment, medicines, etc.) based on the diagnosis prediction, and allocates or replenishes them as needed. The input is the diagnosis prediction result, and the output is confirmation of the inventory status and the appropriate allocation of resources. The server works with the inventory system based on the prediction results to optimally allocate resources.
[0252] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0253] This invention combines an emotion engine with a system that promotes digitalization and cloud computing in the medical industry. Below, we will generate a program for this system and explain its processing in natural language. Specific examples will also be provided.
[0254] Data collection and digitization
[0255] User:
[0256] First, the user collects the patient's paper medical records or existing electronic medical records. The paper medical records are digitized through scanning and imported into the terminal. The digitized data is then uploaded to the server via the terminal.
[0257] Device:
[0258] The terminal receives the scanned paper medical records and converts them into text data using OCR (optical character recognition) technology, which is then sent to the server.
[0259] server:
[0260] The server converts the received text data into an appropriate database format and stores it in the database, allowing past medical records to be digitized and efficiently managed.
[0261] Training generative AI models
[0262] server:
[0263] The server extracts historical patient data and medical records from the database and splits them into training and validation datasets for the generative AI model. The training dataset is used to train the AI model, completing the model learning process. The model's performance is evaluated and adjustments are made as needed.
[0264] Building a voice support system
[0265] Device:
[0266] The device uses a speech recognition API to enable voice input. When the user interviews a patient, the device converts the speech into text data in real time. It also generates optimal questions based on a pre-set question list and past response history.
[0267] server:
[0268] The server receives voice data in real time, converts it into text using speech recognition technology, and then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user.
[0269] Incorporating an emotion engine
[0270] Device:
[0271] The device has an emotion engine that recognizes emotions from the user's voice and input data in real time, for example, determining the user's emotional state from the tone of their voice and the words they choose.
[0272] server:
[0273] The server analyzes the emotion data sent from the emotion engine and generates feedback according to the user's emotions. If the user is feeling stressed or anxious, the system will suggest appropriate responses and support.
[0274] Data partitioning and sharing
[0275] server:
[0276] The server will appropriately anonymize patient data to enhance privacy and security, and will partition and share this data with relevant medical institutions and researchers as needed, with security protocols in place and appropriate access rights managed.
[0277] Diagnostic prediction and optimal allocation of medical resources
[0278] server:
[0279] When new patient data is input, the server uses the generative AI model to predict diagnoses. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes them as necessary.
[0280] Security and Privacy Measures
[0281] server:
[0282] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology to prevent unauthorized access and data leaks. In addition, security audits are conducted regularly to ensure the safety of the system.
[0283] Specific examples
[0284] 1. The user scans Patient D's paper medical record and uploads it to the server via a terminal. The server uses OCR technology to convert the medical record into text data and saves it in the database.
[0285] 2. When interviewing Patient E, the user uses voice input, and the device generates appropriate questions and converts them into text in real time. The server analyzes the results using a generative AI model and provides feedback to the user. If the emotion engine detects stress in the user, it provides appropriate support.
[0286] 3. The server makes a diagnosis prediction and checks the inventory of necessary medical equipment and medicines for Patient F, who is diagnosed with possible pneumonia, and allocates them appropriately.
[0287] 4. Patient data is anonymized and securely shared with other medical institutions for research purposes. The server applies access controls and data encryption to ensure privacy and security.
[0288] As described above, by specifically implementing the present invention, it is possible to efficiently manage and utilize medical records, support medical staff, strengthen data security, and even provide appropriate feedback according to the user's emotions.
[0289] The processing flow will be explained below.
[0290] Step 1:
[0291] The user collects the patient's paper or electronic records and scans the paper records, which then digitizes the data and loads it into the terminal.
[0292] Step 2:
[0293] The terminal converts the scanned paper chart image into text data using OCR (optical character recognition) technology, which is then sent to the server.
[0294] Step 3:
[0295] The server converts the received text data into an appropriate database format and stores it in the database, thereby digitizing the medical records.
[0296] Step 4:
[0297] The server extracts historical patient data and medical records from the database and splits them into training and validation datasets.
[0298] Step 5:
[0299] The server trains the generative AI model using the training dataset, setting parameters such as the number of epochs and batch size, and running the model's learning process.
[0300] Step 6:
[0301] The server evaluates the performance of the trained model using a validation dataset and adjusts the model as needed.
[0302] Step 7:
[0303] The device uses the speech recognition API to enable voice input, preparing to convert what the user says into text in real time.
[0304] Step 8:
[0305] When interviewing patients, the device automatically generates the most appropriate questions based on a set list of questions and past response history.
[0306] Step 9:
[0307] The device converts the user's voice input into text in real time and immediately sends the converted text to the server.
[0308] Step 10:
[0309] The server analyzes the text data sent and uses a generative AI model to provide appropriate feedback to the user.
[0310] Step 11:
[0311] The device uses an emotion engine to recognize emotions from the user's voice and text data in real time, and transmits the emotional state to the server.
[0312] Step 12:
[0313] The server analyzes the emotion data sent from the emotion engine and adjusts the content of the feedback based on the user's emotions.
[0314] Step 13:
[0315] The server anonymizes patient data to enhance privacy and security, and the data is segmented and shared with relevant medical institutions and researchers as needed.
[0316] Step 14:
[0317] The server uses the generative AI model to make a diagnostic prediction based on the new patient data entered, and notifies the user of the diagnostic prediction results in real time.
[0318] Step 15:
[0319] The server checks the inventory status of medical resources (e.g., medical equipment and medicines) based on patient data and diagnosis predictions, and allocates or replenishes them as needed.
[0320] Step 16:
[0321] The server manages the access control list (ACL) and sets the access rights for each user. Data is protected using AES-256 encryption technology.
[0322] Step 17:
[0323] We regularly conduct security audits of our servers to ensure the safety of our systems. If a security incident occurs, we will respond immediately and take corrective measures.
[0324] Example 2
[0325] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0326] In modern healthcare, vast amounts of patient data are continually being generated, and managing, analyzing, and sharing this data presents major challenges. In particular, where many past medical records, including paper charts, remain, it is important to digitize and appropriately utilize them. There is also a need to understand the emotional state of patients through their voices and provide appropriate feedback and support. Furthermore, effective methods are needed to improve diagnostic accuracy and efficiently allocate medical resources.
[0327] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0328] In this invention, the server includes means for digitizing past medical records and storing them in a database, means for training medical datasets using a generative AI model, means for supporting patient interviews using voice recognition, means for anonymizing patient data to enhance security, means for making diagnostic predictions and supporting the optimal allocation of medical resources, means for ensuring data security through access control and data encryption, means for recognizing user emotions in real time and providing corresponding feedback, and means for sharing medical data with other medical institutions and researchers using security protocols. This enables efficient management of massive amounts of patient data, rapid feedback based on patient emotions, improved diagnostic accuracy, and optimal allocation of medical resources.
[0329] "Method of digitizing past medical records and storing them in a database" refers to the process of scanning paper medical records or existing electronic medical records, converting them into text data using OCR technology, and storing them in a database.
[0330] "Method of using a generative AI model to train a medical dataset" refers to the process of collecting medical data, dividing it into a training dataset and a validation dataset to train an AI model, and evaluating and adjusting the model's performance.
[0331] "Means for supporting patient hearings using voice recognition" refers to a process that uses a voice recognition API to convert voice input from patients into text data and support the hearing process.
[0332] "Methods for de-identifying and enhancing security of patient data" refers to technologies and processes used to de-identify personal patient information and enhance privacy and security.
[0333] "Means for making diagnostic predictions and supporting the appropriate allocation of medical resources" refers to the process of using generative AI models to make diagnostic predictions from new patient data and then allocating or replenishing necessary medical resources based on the results.
[0334] "Data security through access control and data encryption" refers to the process of using access control lists to set each user's access rights and protecting the data with AES-256 encryption technology.
[0335] "Means for recognizing a user's emotions in real time and providing corresponding feedback" refers to the technology and process for using an emotion engine to recognize emotions from a user's voice or input data and providing corresponding feedback in real time.
[0336] "Means of sharing medical data with other medical institutions and researchers using security protocols" refers to the process of appropriately anonymizing medical data and applying security protocols to safely share it with other medical institutions and researchers.
[0337] This invention is a system that promotes digitalization and cloud computing in the medical field and incorporates an emotion engine to provide efficient support to patients and medical staff. Specific embodiments of the present invention are described below.
[0338] Data collection and digitization
[0339] The user collects the patient's paper medical records and digitizes them using a scanner. The scanner captures the paper medical records as image data, and the terminal then converts them into text data using OCR (optical character recognition) technology. This text data is uploaded from the terminal to the server. The server converts the text data into an appropriate database format and stores it in the database.
[0340] Example: A user scans a medical record, and the device converts the image data into text using OCR software (e.g., ABBYY FineReader). The converted data is then uploaded to a cloud server and stored in a database.
[0341] Training generative AI models
[0342] The server extracts historical patient data and medical records from the database and splits them into a training dataset and a validation dataset for the generative AI model. The training dataset is used to train the AI model, verify its performance, and make adjustments as needed.
[0343] Example: The server uses frameworks such as Python and TensorFlow to input past diagnostic data into an AI model for learning. For example, training using random forests or neural networks is conceivable.
[0344] Building a voice support system
[0345] The device uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to enable voice input. When the user interviews a patient, the device converts the speech into text data in real time and sends it to the server. The server uses a generative AI model to analyze the context of the conversation and provide optimal feedback.
[0346] Example: When a user speaks to the device, the device converts the voice data into text data and sends it to the server. The server analyzes it and presents the next question.
[0347] Incorporating an emotion engine
[0348] The device uses an emotion engine to recognize emotions from the user's voice and input data. For example, it determines the user's emotional state in real time from the user's tone of voice and selected words. The server analyzes this data and generates feedback according to the user's emotions.
[0349] Example: If a user asks a question in an anxious tone, the server generates reassuring feedback and provides it to the user through the terminal.
[0350] Data partitioning and sharing
[0351] The server will appropriately anonymize patient data and enhance security. This anonymized data will be shared with relevant medical institutions and researchers as needed. Security protocols will be applied to sharing, and access rights will be strictly controlled.
[0352] Example: A server processes data using an anonymization algorithm (e.g., K-anonymization or differential privacy) and shares it over a secure communication protocol (e.g., HTTPs or TLS).
[0353] Diagnostic prediction and optimal allocation of medical resources
[0354] When new patient data is input, the server uses the generative AI model to predict a diagnosis. The prediction results are notified to the user in real time, and at the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes them as needed.
[0355] Example: The server inputs the symptom data of a new patient into an AI model, and if it predicts pneumonia, it notifies the user of the result and allocates the necessary medical resources.
[0356] Security and Privacy Measures
[0357] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology, and regular security audits ensure the system is secure.
[0358] For example: The server manages access privileges for each user based on ACLs, and stored data is encrypted with AES-256. Vulnerability checks are performed regularly using security audit tools (e.g., Nessus or OpenVAS).
[0359] Prompt Sentence Examples
[0360] Specifically, the following prompt sentences could be input into the generative AI model:
[0361] "Enter new patient data."
[0362] "Generate the next question to ask."
[0363] "Do sentiment analysis and provide appropriate feedback."
[0364] As described above, the system of the present invention integrates a wide range of functions to realize efficient data management, patient support, diagnostic prediction, and enhanced security in medical settings.
[0365] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0366] Step 1: Collect and digitize data
[0367] Input: Patient's paper chart
[0368] Output: Digitized text data
[0369] The user collects the patient's paper chart and places it into the scanner.
[0370] How it works: When the user presses the scan button, the scanner reads the paper medical record as image data.
[0371] The device receives the image data from the scanner and converts it into text data using OCR technology. For example, ABBYY FineReader installed on the device is used.
[0372] What happens: The device launches the OCR software and converts the image file into a text file.
[0373] The terminal uploads the converted text data to the server.
[0374] Operation: The device sends text data to the server's API endpoint.
[0375] Step 2: Saving to the database
[0376] Input: Converted text data
[0377] Output: Medical records stored in a database
[0378] The server converts the received text data into the appropriate database format.
[0379] How it works: The server parses the data format and generates SQL queries to insert into the database.
[0380] The server stores the text data in a database.
[0381] Step 3: Prepare training data for the generative AI model
[0382] Input: Medical records stored in a database
[0383] Output: Training and validation datasets
[0384] The server extracts medical record data from the database and splits it into a training dataset and a validation dataset.
[0385] How it works: The server executes SQL queries to extract data and runs an algorithm to randomly split the dataset.
[0386] Step 4: Training the generative AI model
[0387] Input: Teacher dataset
[0388] Output: A trained generative AI model
[0389] The server trains the AI model using the training dataset.
[0390] How it works: The server runs a script to train an AI model using Python or TensorFlow. It inputs the training dataset into the model and trains it.
[0391] Step 5: Building a voice support system
[0392] Input: User voice input
[0393] Output: Real-time audio data converted to text
[0394] The device uses the speech recognition API to enable the voice input function.
[0395] What it does: The device creates an instance of a speech recognition API (e.g., Google Cloud Speech-to-Text) and displays a UI to initiate voice input.
[0396] The user interviews the patient and inputs the information by voice.
[0397] The device converts the voice into text data in real time and sends it to the server.
[0398] Step 6: Analysis of audio data and feedback
[0399] Input: Real-time audio data converted to text
[0400] Output: Analysis results and feedback
[0401] The server analyzes the context of the conversation using a generative AI model.
[0402] How it works: The server inputs text data into the generative AI model and obtains the analysis results.
[0403] The server generates optimal feedback and provides it to the user through the terminal.
[0404] Step 7: Recognize emotions and provide feedback
[0405] Input: User voice and input data
[0406] Output: Emotional feedback
[0407] The terminal uses an emotion engine to recognize the user's emotion.
[0408] How it works: The emotion engine analyzes the user's tone of voice and word choice to determine their emotional state.
[0409] The server analyzes the emotional data and generates appropriate feedback for the user.
[0410] Operation: The server generates a feedback message based on the emotion data and provides it to the user via the terminal.
[0411] Step 8: Anonymize and share data
[0412] Input: Patient Data
[0413] Output: Anonymized data
[0414] The server appropriately anonymizes the patient data.
[0415] How it works: The server processes the data using an anonymization algorithm (e.g., K-anonymization or differential privacy).
[0416] The server will share the anonymized data with other medical institutions and researchers as needed.
[0417] How it works: Data is transmitted securely using security protocols (e.g., HTTPs or TLS).
[0418] Step 9: Diagnostic prediction and resource allocation
[0419] Input: New patient data
[0420] Output: Diagnosis results and medical resource allocation information
[0421] The server uses the generated AI model to make diagnostic predictions for new patient data.
[0422] How it works: New patient data is fed into a generative AI model to obtain a diagnosis.
[0423] The server notifies the user of the diagnosis results and checks the inventory status of medical resources (e.g., medical equipment and medicines) and allocates them appropriately.
[0424] Operation: The server queries the inventory management system and allocates and replenishes needed medical resources.
[0425] Step 10: Security and Privacy
[0426] Input: User ID and data
[0427] Output: Access permission settings and encrypted data
[0428] The server manages the access control list (ACL) and sets the access rights for each user.
[0429] How it works: The server updates the ACL database to manage user identities and access privileges.
[0430] The server protects your data with AES-256 encryption technology.
[0431] How it works: Before the server stores the data, it encrypts it using the AES-256 algorithm.
[0432] The server performs regular security audits to ensure the system is secure.
[0433] What happens: The server runs a vulnerability check using a security audit tool (e.g., Nessus or OpenVAS).
[0434] (Application example 2)
[0435] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0436] While conventional digital and cloud-based systems can efficiently manage and utilize medical records, they are inadequate for real-time support, such as displaying work instructions in real time or assessing workers' emotional states. They also have difficulty providing feedback and appropriate work support through voice recognition. Furthermore, they have faced problems with the psychological burden on workers and reduced work efficiency due to errors.
[0437] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0438] In this invention, the server includes means for digitizing past medical records and storing them in a database, means for training medical datasets using a generative AI model, means for supporting patient interviews using voice recognition, means for anonymizing patient data to enhance security, means for making diagnostic predictions and supporting the appropriate allocation of medical resources, means for ensuring data security through access control and data encryption, means for having a visual device for displaying work instructions in real time, means for providing feedback to workers using a voice response system, and means for assessing the emotional state of workers through emotion recognition. This makes it possible to evaluate the emotional state of workers while displaying work instructions in real time and providing feedback through voice recognition.
[0439] "Digitizing past medical records and storing them in a database" means converting paper medical records and existing electronic medical records into a digital format using scanning or optical character recognition technology (OCR) and storing them in a database.
[0440] "Learning medical datasets using a generative AI model" means training a generative AI model using datasets based on past patient data and medical records, enabling it to analyze medical data and make diagnostic predictions.
[0441] "Supporting patient interviews using voice recognition" means using voice recognition technology to convert interviews with patients into text in real time, and providing appropriate support to medical staff.
[0442] "Anonymizing patient data to enhance security" means protecting the privacy and enhancing the security of patient data by removing or transforming personally identifiable information.
[0443] "Making diagnostic predictions and supporting the appropriate allocation of medical resources" means using a generative AI model to make diagnostic predictions for patients and then appropriately allocating medical resources such as medical equipment and medicines based on the results.
[0444] "Ensuring data security through access control and data encryption" means setting access rights for each user and protecting data using encryption technology such as AES-256.
[0445] "Having a visual device for displaying work instructions in real time" means displaying work instructions in real time using a device such as smart glasses or a head-up display.
[0446] "Using a voice response system to provide feedback to workers" means using a microphone and voice recognition technology to recognize voice instructions from workers and provide appropriate responses or feedback.
[0447] "Assessing the emotional state of workers through emotion recognition" means using voice analysis and facial expression recognition technology to assess the emotional state of workers in real time and provide appropriate support based on the results.
[0448] This invention relates to a work support system for factories. This system exchanges information in real time between a server, terminals, and users, and provides multifunctional support to improve work efficiency.
[0449] Data collection and digitization
[0450] server:
[0451] Past medical records and work history are digitized and stored in a database. Paper medical records and handwritten work records are scanned and converted into text data using OCR (optical character recognition) technology. The converted data is then stored in a database for efficient management.
[0452] Training generative AI models
[0453] server:
[0454] Past work data and medical records are extracted from the database and divided into a training dataset and a validation dataset for the generative AI model. The training dataset is used to train the AI model, and its performance is evaluated and adjusted. This enables accurate diagnosis prediction and work instruction generation.
[0455] Building a voice support system
[0456] Device:
[0457] It uses a speech recognition API to enable voice input functionality. When users give voice instructions during work, the device converts the speech into text data in real time and provides appropriate feedback. It also uses generative AI models to automatically generate work instructions and questions.
[0458] Incorporating an emotion engine
[0459] Device:
[0460] It has an emotion engine that recognizes emotions from the user's voice and input data in real time, determines the emotional state from the tone of voice and words used, and provides feedback according to the emotion.
[0461] View work instructions in real time
[0462] Device:
[0463] Using smart glasses or head-up displays, work instructions are displayed in real time, allowing workers to receive visual information immediately and work efficiently.
[0464] ★Example:
[0465] For example, when a worker needs to install the next part in a factory, the smart glasses will display work instructions such as "Please take the next part, B, and install it on machine C." If the worker asks by voice, "What should I do next?", the next work instruction will be instantly provided via voice and text.
[0466] Emotion recognition for worker support
[0467] server:
[0468] The system analyzes the emotional data received from the emotion engine and provides appropriate support if a worker is feeling stressed or fatigued, such as suggesting slowing down work speed or taking a break.
[0469] Predicting abnormalities and presenting countermeasures
[0470] server:
[0471] If an abnormality occurs during work, the generative AI model is used to analyze the anomaly and suggest countermeasures, allowing for swift and appropriate measures to be taken.
[0472] ★Example prompt:
[0473] The current task is assembling parts. Please generate the next task instruction.
[0474] In this way, by implementing the present invention, not only can workers receive work instructions in real time and work efficiently, but it also enables support according to emotional states and the prediction of abnormalities, thereby improving overall work efficiency and safety.
[0475] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0476] Step 1:
[0477] Data collection and digitization
[0478] Users scan paper medical records or handwritten work records and input them into the terminal. The terminal then converts the scanned data into text data using OCR (optical character recognition) technology. This text data is sent to the server and stored in a database.
[0479] Input: Scanned paper charts and handwritten work records
[0480] Data processing: Text conversion using OCR technology
[0481] Output: Text data stored in the database
[0482] Step 2:
[0483] Training generative AI models
[0484] The server extracts past medical records and work data from the database and divides them into training and validation datasets for the generative AI model. The server trains the generative AI model using the training dataset and evaluates and adjusts the model's performance.
[0485] Input: Medical records and work data in a database
[0486] Data Computing: Learning and Evaluating Generative AI Models
[0487] Output: A trained generative AI model
[0488] Step 3:
[0489] Building a voice support system
[0490] The device uses a speech recognition API to enable voice input. The user gives voice instructions while working, and the device converts the speech into text data in real time. The server then uses a generative AI model to analyze the text data and generate appropriate feedback and work instructions.
[0491] Input: User's voice commands
[0492] Data processing: Text conversion using voice recognition, analysis using generative AI models
[0493] Output: Text data of feedback and work instructions
[0494] Step 4:
[0495] Incorporating emotion recognition
[0496] The device uses an emotion engine to recognize emotions from the user's voice and input data in real time. The recognized emotion data is sent to a server, which analyzes it and provides appropriate feedback to the user. For example, if the user is feeling stressed, the server may suggest taking a break.
[0497] Input: User voice and input data
[0498] Data calculation: Emotion recognition by emotion engine, analysis by server
[0499] Output: Emotional feedback
[0500] Step 5:
[0501] View work instructions in real time
[0502] The device displays work instructions in real time using smart glasses or a head-up display. The server uses a generative AI model to predict the next task and sends it to the device. The user receives the work instructions through the visual device, allowing them to work efficiently.
[0503] Input: Work instructions from the server
[0504] Data Computation: Task Prediction with Generative AI Models
[0505] Output: Work instructions displayed on a visual device
[0506] Step 6:
[0507] Predicting abnormalities and presenting countermeasures
[0508] The server monitors data generated during work in real time, and if an abnormality occurs, it analyzes it using a generative AI model. Based on the analysis results, it sends appropriate countermeasures to the device. The user can quickly resolve the abnormality by following the countermeasures provided by the device.
[0509] Input: Real-time data as you work
[0510] Data Computation: Anomaly Analysis with Generative AI Models
[0511] Output: Actions sent to the device
[0512] Step 7:
[0513] Emotion recognition for worker support
[0514] The server analyzes the emotion data received from the emotion engine and provides appropriate support if the worker is feeling stressed or fatigued, such as suggesting slowing down work speed or taking a break.
[0515] Input: Emotion data from the emotion engine
[0516] Data Computing: Server-Based Emotion Analysis
[0517] Output: Emotionally appropriate support and feedback
[0518] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0519] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0520] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0521] [Second embodiment]
[0522] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0523] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0524] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0525] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0526] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0527] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0528] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0529] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0530] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0531] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0532] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0533] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0534] This invention is a system that promotes digitalization and cloud computing in the medical industry. Below, we will create a program for this system and explain its processing in natural language. Specific examples will also be provided.
[0535] Data collection and digitization
[0536] User:
[0537] First, users collect the patient's paper medical records or existing electronic medical records, which are then digitized by scanning them, and then upload the digitized data to the server via a terminal.
[0538] Device:
[0539] The terminal receives the scanned paper medical records and converts them into text data using OCR (optical character recognition) technology, which is then sent to the server.
[0540] server:
[0541] The server converts the received text data into an appropriate database format and stores it in the database, allowing past medical records to be digitized and efficiently managed.
[0542] Training generative AI models
[0543] server:
[0544] The server extracts historical patient data and medical records from the database and divides them into training and validation datasets for the generative AI model. The training dataset is used to train the AI model, completing the model learning process. This process also includes adjusting hyperparameters such as the number of epochs and batch size. The model's performance is evaluated and adjustments are made as necessary.
[0545] Building a voice support system
[0546] Device:
[0547] The device uses a speech recognition API to enable voice input. When the user interviews a patient, the device converts the speech into text data in real time. It also generates optimal questions based on a pre-set question list and past response history.
[0548] server:
[0549] The server receives voice data in real time, converts it into text using speech recognition technology, and then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user.
[0550] Data partitioning and sharing
[0551] server:
[0552] The server will appropriately anonymize patient data to enhance privacy and security, and will partition and share this data with relevant medical institutions and researchers as needed, with security protocols in place and appropriate access rights managed.
[0553] Diagnostic prediction and optimal allocation of medical resources
[0554] server:
[0555] When new patient data is input, the server uses the generative AI model to predict diagnoses. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes them as necessary.
[0556] Security and Privacy Measures
[0557] server:
[0558] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology to prevent unauthorized access and data leaks. In addition, security audits are conducted regularly to ensure the safety of the system.
[0559] Specific examples
[0560] 1. The user scans Patient A's paper medical record and uploads it to the server via their device. The server converts it into text data using OCR technology and stores it in the database.
[0561] 2. When interviewing new patient B, the user uses voice input, and the device generates appropriate questions and converts them into text in real time. The server analyzes them using a generative AI model and provides feedback to the user.
[0562] 3. The server makes a diagnosis prediction and diagnoses Patient C as having a high possibility of influenza. The server checks the inventory of necessary medical equipment and medicines and allocates them appropriately.
[0563] 4. Patient data is anonymized and securely shared with other medical institutions for research purposes. The server applies access controls and data encryption to ensure privacy and security.
[0564] As described above, by specifically implementing the present invention, it is possible to realize efficient management and utilization of medical records, support for medical staff, and strengthen data security.
[0565] The processing flow will be explained below.
[0566] Step 1:
[0567] The user collects the patient's paper medical record or existing electronic medical record and scans the paper medical record, which then digitizes the data and loads it into the terminal.
[0568] Step 2:
[0569] The terminal converts scanned paper chart images into text data using OCR (Optical Character Recognition) technology, which is then sent to a server for further processing.
[0570] Step 3:
[0571] The server converts the received text data into an appropriate database format and stores it in the database. This standardizes the data and enables efficient management.
[0572] Step 4:
[0573] The server extracts historical patient data and medical records from the database and splits them into training and validation datasets.
[0574] Step 5:
[0575] The server trains the generative AI model using the training dataset, sets hyperparameters such as the number of epochs and batch size, and executes the model learning process.
[0576] Step 6:
[0577] The server evaluates the performance of the trained model using a validation dataset and adjusts the model as needed.
[0578] Step 7:
[0579] The device uses the speech recognition API to enable the voice input function, preparing the device for the user to use voice input.
[0580] Step 8:
[0581] When interviewing patients, the device automatically generates the most appropriate questions based on a pre-set list of questions and past response history.
[0582] Step 9:
[0583] The device converts the user's voice input into text in real time and immediately sends the text to the server.
[0584] Step 10:
[0585] The server analyzes the text data sent and uses a generative AI model to provide appropriate feedback to the user.
[0586] Step 11:
[0587] The server anonymizes patient data, removing or transforming personally identifiable information to protect the confidentiality of the data.
[0588] Step 12:
[0589] The server will divide the anonymized data and share it with medical institutions and researchers who need it, and will also set up security protocols and manage access rights.
[0590] Step 13:
[0591] When new patient data is entered into the server, the generative AI model is used to make a diagnostic prediction, and the results are notified to the user in real time.
[0592] Step 14:
[0593] The server checks the inventory status of medical resources (medical equipment and medicines) based on patient data and diagnosis predictions, and allocates or replenishes them as necessary.
[0594] Step 15:
[0595] The server manages the access control list (ACL) and sets the access rights for each user. Data is protected using AES-256 encryption technology.
[0596] Step 16:
[0597] We regularly conduct security audits on our servers to ensure the safety of our systems. In the event of a security incident, we respond immediately and take corrective measures.
[0598] Example 1
[0599] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0600] Problems with conventional medical systems include the lack of progress in digitizing medical records and the ineffective use of collected data. Furthermore, there is a lack of systems that support the accuracy of diagnostic predictions and the appropriate allocation of medical resources. There are also many shortcomings in terms of security and privacy protection. These issues are major obstacles to improving the efficiency and quality of medical care, and they require rapid and accurate responses.
[0601] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0602] In this invention, the server includes a means for scanning and digitizing past medical records, converting them into text data using OCR technology, and storing them in a database, a means for extracting and training data sets using a generative AI model, and a means for anonymizing patient data and using technology to enhance security, thereby enabling efficient management and utilization of medical records, improving the accuracy of diagnostic predictions, and optimizing the allocation of medical resources.
[0603] A "database format" is a structure or format for efficiently storing, retrieving, and managing digital data.
[0604] "OCR technology" is an abbreviation for optical character recognition technology, which extracts character information from scanned images and converts it into text data.
[0605] A "generative AI model" refers to an algorithm or model that uses artificial intelligence techniques to generate new information or predictions from a given dataset.
[0606] "Voice recognition technology" is a technology that converts voice input into text data in real time.
[0607] "Data anonymization" refers to the process of removing or modifying personally identifiable information to protect the privacy of the data.
[0608] A "security protocol" is a set of procedures and rules that are used to prevent unauthorized access when transmitting or receiving data.
[0609] "Diagnostic prediction" is the process of using generative AI models based on medical data to predict illness and treatment.
[0610] An "access control list (ACL)" is a list used to manage the rights and permissions of each user accessing a system.
[0611] "Data encryption" is a technology that encrypts important information using a specific algorithm to prevent unauthorized access and data leaks.
[0612] "Hyperparameter tuning" is the process of fine-tuning the parameters of a machine learning model to optimize its learning efficiency and performance.
[0613] The present invention is a system that promotes digitalization and cloud computing in the medical industry. The program processing of this system will be explained below in natural language.
[0614] Data collection and digitization
[0615] User:
[0616] First, the user collects the patient's paper medical records or existing electronic medical records. They use a general-purpose scanner to scan the paper medical records and convert them into digital data. For example, they use a general-purpose high-performance scanner. Next, the user uploads the scanned digital data to a server via a terminal. This is done using the hospital's internal network.
[0617] Device:
[0618] The terminal receives the scanned image data of the paper medical record and converts it into text data using OCR technology. This technology uses commonly used OCR software. The converted text data is sent to the server. When sending, the communication is encrypted using SSL / TLS.
[0619] server:
[0620] The server receives the text data sent from the terminal. The received data is converted into an appropriate database format and stored in a database. For example, a relational database management system (RDBMS) is used as the database.
[0621] Training generative AI models
[0622] server:
[0623] The server extracts past patient data and medical records from the database. The extracted data is divided into a training dataset and a validation dataset for the generative AI model. The server trains the generative AI model using the training dataset. Deep learning libraries such as PyTorch and TensorFlow are used for this training. Hyperparameters such as the number of epochs and batch size are also adjusted. The model's performance is evaluated on the validation dataset, and readjustments are made as necessary.
[0624] Building a voice support system
[0625] Device:
[0626] The device uses a speech recognition API to enable voice input. For example, a general speech recognition API can be used. When the user interviews a patient, the device converts the speech into text data in real time. It also generates optimal questions based on a pre-set question list and past response history.
[0627] server:
[0628] The server receives the voice data in real time, converts it into text using speech recognition technology, and then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user.
[0629] Data partitioning and sharing
[0630] server:
[0631] The server will appropriately anonymize patient data to enhance privacy and security. This will involve techniques to remove or modify personally identifiable information. Furthermore, the anonymized data will be split up as needed and shared with relevant medical institutions and researchers. This sharing will be subject to security protocols and appropriate access rights management.
[0632] Diagnostic prediction and optimal allocation of medical resources
[0633] server:
[0634] When new patient data is input, the server uses the generative AI model to predict diagnoses. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes these resources as needed.
[0635] Security and Privacy Measures
[0636] server:
[0637] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology to prevent unauthorized access and data leaks. In addition, security audits are conducted regularly to ensure the safety of the system.
[0638] Specific examples
[0639] 1. The user scans Patient A's paper medical record and uploads it to the server via their device. The server converts it into text data using OCR technology and stores it in the database.
[0640] 2. When interviewing new patient B, the user uses voice input, and the device generates appropriate questions and converts them into text in real time. The server analyzes them using a generative AI model and provides feedback to the user.
[0641] 3. The server makes a diagnosis prediction and diagnoses Patient C as having a high possibility of influenza. The server checks the inventory of necessary medical equipment and medicines and allocates them appropriately.
[0642] 4. Patient data is anonymized and securely shared with other medical institutions for research purposes. The server applies access controls and data encryption to ensure privacy and security.
[0643] Prompt Sentence Examples
[0644] "How do you digitize and store Patient A's paper medical records?"
[0645] "How do I generate the best questions for a new patient, Patient B?"
[0646] "How can we predict diagnoses and allocate necessary medical resources appropriately?"
[0647] As described above, by specifically implementing the present invention, it is possible to realize efficient management and utilization of medical records, support for medical staff, and strengthen data security.
[0648] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0649] Step 1:
[0650] User: The user obtains digital image data by scanning a patient's paper chart. This input data is in an image file format (e.g., JPEG or PNG). The specific actions of scanning using a scanner include placing the paper chart on the scanner and pressing the Scan button. The output is the scanned digital image data.
[0651] Step 2:
[0652] User: The user uploads the scanned digital image data to the terminal. The uploading process involves sending the digital image to a designated folder using the hospital network. The input is the scanned digital image, and the output is the uploaded digital image in the terminal.
[0653] Step 3:
[0654] Terminal: The terminal converts the received digital image data into text data using OCR technology. Specific operations include launching the OCR software, reading the digital image, and performing character recognition. The input is digital image data, and the output is the converted text data.
[0655] Step 4:
[0656] Terminal: The terminal sends the text data generated by OCR to the server. This transmission process uses SSL / TLS encrypted communication. The input is the converted text data, and the output is the text data sent to the server.
[0657] Step 5:
[0658] Server: The server converts the received text data into a database format and stores it in the database. Specific operations include converting the text data into an SQL query format and performing an insert operation on the database. The input is the text data sent to the server, and the output is the medical record data stored in the database.
[0659] Step 6:
[0660] Server: The server extracts historical patient data and medical records from the database and splits them into training and validation datasets for the generative AI model. Specific operations include executing SQL queries to retrieve the data and splitting the extracted data into training and validation datasets. The input is the medical record data stored in the database, and the output is the split dataset.
[0661] Step 7:
[0662] Server: The server trains the generative AI model using the training dataset. This is done using deep learning libraries such as PyTorch or TensorFlow. It also adjusts hyperparameters such as the number of epochs and batch size. The input is the training dataset, and the output is the trained generative AI model.
[0663] Step 8:
[0664] Server: The server evaluates the model's performance on the validation dataset and adjusts hyperparameters or modifies the model structure as necessary. Specific operations include calculating metrics for model evaluation and retraining. The input is the validation dataset and the existing generative AI model, and the output is an optimized generative AI model.
[0665] Step 9:
[0666] Terminal: The terminal uses a speech recognition API to enable voice input and converts the user's voice during patient interviews into text data in real time. It also generates optimal questions based on a pre-set question list and past response history. The input is voice data, and the output is text data and the generated questions.
[0667] Step 10:
[0668] Server: The server receives voice data in real time and converts it into text data using speech recognition technology. It then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user. The input is voice data, and the output is interpreted text data and feedback.
[0669] Step 11:
[0670] Server: The server anonymizes patient data, removes or modifies any personally identifiable information from the data, and applies security protocols and manages access rights to share the anonymized data with relevant medical institutions and researchers. The input is raw patient data, and the output is anonymized data.
[0671] Step 12:
[0672] Server: When new patient data is input, the server uses the generative AI model to make a diagnosis prediction. The prediction results are notified to the user in real time, and at the same time, the server checks the inventory status of medical resources and allocates or replenishes them as needed. The input is the new patient data, and the output is the diagnosis prediction results and medical resource management information.
[0673] Step 13:
[0674] Server: The server manages the access control list (ACL) and sets the access rights for each user. It also protects data with AES-256 encryption technology and performs regular security audits to prevent unauthorized access and data leakage. The input is user information and data, and the output is controlled access rights and protected data.
[0675] Through these steps, the system achieves efficient digitization of medical records, diagnostic prediction through generative AI models, and data security management.
[0676] (Application example 1)
[0677] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0678] The modern healthcare industry requires the digitization of medical records and efficient data management. However, several challenges exist, including the digitization of paper medical records, the centralized management of existing electronic medical records, the anonymization and security of medical data, and the improvement of diagnostic prediction accuracy. In particular, there is a lack of support systems for quickly and accurately recording patient interactions and making appropriate diagnoses. It is also important to safely anonymize collected data and share it with other medical institutions and researchers. It is necessary to resolve these challenges and realize digitalization and enhanced security in the healthcare industry.
[0679] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0680] In this invention, the server includes means for digitizing past medical records and storing them in a database, means for training medical datasets using a generative AI model, means for supporting patient interviews using voice recognition, means for anonymizing patient data to enhance security, means for making diagnostic predictions and supporting the appropriate allocation of medical resources, means for ensuring data security through access control and data encryption, means for generating text data from scanned images using optical character recognition technology, means for converting voice input into text data in real time using a voice recognition API, and means for checking the availability of medical resources based on diagnostic predictions using the generative AI model. This enables efficient digitization and management of medical records, rapid recording of patient interactions, appropriate diagnostic support, and secure data sharing.
[0681] "Prior medical records" are previously collected medical data, including a patient's medical history and treatment history.
[0682] "Digitalization" is the process of converting information in analog form, such as paper, into digital form.
[0683] A "database" is a system for storing and managing data in an organized structure.
[0684] A "generative AI model" is a machine learning algorithm that learns from large amounts of data and is optimized to perform a specific task.
[0685] A "medical dataset" is a collection of various medical-related data used for learning and analysis.
[0686] "Speech recognition" is the technology that takes voice input and converts it into a digital format such as text.
[0687] A "patient interview" is a process in which a medical professional directly obtains information from a patient, such as symptoms and medical history.
[0688] "Anonymization" is the process of processing data so that individuals cannot be identified.
[0689] "Security" refers to a group of technologies aimed at protecting data and preventing unauthorized access.
[0690] "Diagnostic prediction" is the process of predicting future disease states and appropriate treatments based on collected medical data.
[0691] "Medical resources" is a general term for all equipment, medicines, personnel, etc. used in medical institutions.
[0692] "Proper allocation" means properly allocating the necessary resources to the necessary locations.
[0693] "Access control" is a mechanism that allows only authorized users to access specific data.
[0694] "Data encryption" is the technology that encrypts data and converts it into a form that is unintelligible to unauthorized users.
[0695] "Optical character recognition technology" is a technology that analyzes character information in an image and converts it into text data.
[0696] "Real-time" means that data input, processing, and output are carried out instantly.
[0697] A "speech recognition API" is an application programming interface that provides the functionality to convert voice input into text.
[0698] "Inventory status" is information that indicates the current level of medical resources available for use.
[0699] This invention is a system that promotes digitalization and cloud computing in the medical industry, and provides a wide range of functions, including digitization of past medical records, diagnostic prediction using generative AI models, patient interview support using voice recognition, and enhanced data security using anonymization technology.
[0700] Data collection and digitization
[0701] User:
[0702] Users first scan the patient's paper medical records using their smartphone camera and upload them to the server via the device, allowing them to be digitized and managed efficiently.
[0703] Application of OCR technology
[0704] Device:
[0705] The terminal uses optical character recognition (OCR) technology to generate text data from scanned images of paper medical records, which is then sent to a server and stored in a database.
[0706] Building a voice support system
[0707] Device:
[0708] The device uses a speech recognition API to convert voice input into text data in real time. The user records conversations with patients and converts the data into text in real time. The device also generates optimal questions based on a pre-set question list and past response history.
[0709] Data storage and anonymization
[0710] server:
[0711] The server converts the received text data into an appropriate database format and stores it in the database. The server anonymizes the data and uses data encryption technology (e.g., AES-256) to enhance privacy and security. This data is split as needed and shared with relevant medical institutions and researchers.
[0712] Training generative AI models
[0713] server:
[0714] The server extracts historical patient data and medical records from the database and divides them into training and validation datasets for the generative AI model. The training dataset is used to train the AI model, completing the model learning process. The model's performance is evaluated and adjusted as necessary.
[0715] Diagnostic prediction and optimal allocation of medical resources
[0716] server:
[0717] When new patient data is entered, the server uses the generative AI model to predict a diagnosis. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (medical equipment, drugs, etc.) and allocates or replenishes them as needed.
[0718] Examples of concrete examples and prompts
[0719] 1. A user scans a patient's paper chart and uploads it to the server via a terminal. An example of the prompt is as follows:
[0720] "Take images of medical records, digitize them, and upload them to the cloud."
[0721] 2. The user records the conversation with the patient and uses a speech recognition API to transcribe it in real time, with prompts such as:
[0722] "Record conversations with patients, convert them into text, and save them."
[0723] 3. The server performs diagnosis prediction, checks the availability of medical resources, and allocates them appropriately. An example of the prompt is as follows:
[0724] "Using generative AI models to predict diagnoses and allocate necessary medical resources"
[0725] This enables efficient digitization and management of medical records, enables prompt recording of patient interactions, and ensures proper diagnostic support and secure data sharing, significantly improving operational efficiency and security in the healthcare industry.
[0726] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0727] Step 1:
[0728] The user scans the patient's paper medical record using the smartphone camera and uploads it to the server via the terminal. The input is an image of the paper medical record, and the output is digitized text data. The user launches the scanning app on the terminal and takes a picture of the paper medical record.
[0729] Step 2:
[0730] The device generates text data from the scanned image using optical character recognition (OCR). The input is the scanned image from step 1, and the output is the text data generated by OCR. The device sends the scanned image to the OCR engine and obtains the text data.
[0731] Step 3:
[0732] The device sends the generated text data to the server. The input is the text data, and the output is the data sent to the server. The device then uploads the data to the cloud server via the network.
[0733] Step 4:
[0734] The server converts the received text data into an appropriate database format and stores it in the database. The input is the text data received by the server, and the output is the data stored in the database. The server formats the text data into the database format and stores it.
[0735] Step 5:
[0736] The server uses a speech recognition API to convert the voice data recorded by the user into text data in real time. The input is voice data and the output is text data. The server sends the voice data to the API and saves the resulting text data.
[0737] Step 6:
[0738] The server anonymizes the received data and uses data encryption techniques (e.g., AES-256) to enhance privacy and security. The input is text data stored in a database, and the output is anonymized and encrypted data. The server anonymizes the data and applies encryption techniques.
[0739] Step 7:
[0740] The server uses a generative AI model to split the past patient data and medical records extracted from the database into a training dataset and a validation dataset. The input is the past patient data and medical records, and the output is the training dataset and the validation dataset. The server splits the data into training and validation datasets.
[0741] Step 8:
[0742] The server trains the generative AI model using the training dataset, completing the model learning process. The input is the training dataset, and the output is the trained generative AI model. The server uses the dataset to train the model.
[0743] Step 9:
[0744] When new patient data is input, the server uses the generative AI model to make a diagnosis prediction. The input is the new patient data, and the output is the diagnosis prediction result. The server inputs the new patient data into the model and obtains the diagnosis prediction result.
[0745] Step 10:
[0746] The server checks the inventory status of medical resources (medical equipment, medicines, etc.) based on the diagnosis prediction, and allocates or replenishes them as needed. The input is the diagnosis prediction result, and the output is confirmation of the inventory status and the appropriate allocation of resources. The server works with the inventory system based on the prediction results to optimally allocate resources.
[0747] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0748] This invention combines an emotion engine with a system that promotes digitalization and cloud computing in the medical industry. Below, we will generate a program for this system and explain its processing in natural language. Specific examples will also be provided.
[0749] Data collection and digitization
[0750] User:
[0751] First, the user collects the patient's paper medical records or existing electronic medical records. The paper medical records are digitized through scanning and imported into the terminal. The digitized data is then uploaded to the server via the terminal.
[0752] Device:
[0753] The terminal receives the scanned paper medical records and converts them into text data using OCR (optical character recognition) technology, which is then sent to the server.
[0754] server:
[0755] The server converts the received text data into an appropriate database format and stores it in the database, allowing past medical records to be digitized and efficiently managed.
[0756] Training generative AI models
[0757] server:
[0758] The server extracts historical patient data and medical records from the database and splits them into training and validation datasets for the generative AI model. The training dataset is used to train the AI model, completing the model learning process. The model's performance is evaluated and adjustments are made as needed.
[0759] Building a voice support system
[0760] Device:
[0761] The device uses a speech recognition API to enable voice input. When the user interviews a patient, the device converts the speech into text data in real time. It also generates optimal questions based on a pre-set question list and past response history.
[0762] server:
[0763] The server receives voice data in real time, converts it into text using speech recognition technology, and then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user.
[0764] Incorporating an emotion engine
[0765] Device:
[0766] The device has an emotion engine that recognizes emotions from the user's voice and input data in real time, for example, determining the user's emotional state from the tone of their voice and the words they choose.
[0767] server:
[0768] The server analyzes the emotion data sent from the emotion engine and generates feedback according to the user's emotions. If the user is feeling stressed or anxious, the system will suggest appropriate responses and support.
[0769] Data partitioning and sharing
[0770] server:
[0771] The server will appropriately anonymize patient data to enhance privacy and security, and will partition and share this data with relevant medical institutions and researchers as needed, with security protocols in place and appropriate access rights managed.
[0772] Diagnostic prediction and optimal allocation of medical resources
[0773] server:
[0774] When new patient data is input, the server uses the generative AI model to predict diagnoses. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes them as necessary.
[0775] Security and Privacy Measures
[0776] server:
[0777] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology to prevent unauthorized access and data leaks. In addition, security audits are conducted regularly to ensure the safety of the system.
[0778] Specific examples
[0779] 1. The user scans Patient D's paper medical record and uploads it to the server via a terminal. The server uses OCR technology to convert the medical record into text data and saves it in the database.
[0780] 2. When interviewing Patient E, the user uses voice input, and the device generates appropriate questions and converts them into text in real time. The server analyzes the results using a generative AI model and provides feedback to the user. If the emotion engine detects stress in the user, it provides appropriate support.
[0781] 3. The server makes a diagnosis prediction and checks the inventory of necessary medical equipment and medicines for Patient F, who is diagnosed with possible pneumonia, and allocates them appropriately.
[0782] 4. Patient data is anonymized and securely shared with other medical institutions for research purposes. The server applies access controls and data encryption to ensure privacy and security.
[0783] As described above, by specifically implementing the present invention, it is possible to efficiently manage and utilize medical records, support medical staff, strengthen data security, and even provide appropriate feedback according to the user's emotions.
[0784] The processing flow will be explained below.
[0785] Step 1:
[0786] The user collects the patient's paper or electronic records and scans the paper records, which then digitizes the data and loads it into the terminal.
[0787] Step 2:
[0788] The terminal converts the scanned paper chart image into text data using OCR (optical character recognition) technology, which is then sent to the server.
[0789] Step 3:
[0790] The server converts the received text data into an appropriate database format and stores it in the database, thereby digitizing the medical records.
[0791] Step 4:
[0792] The server extracts historical patient data and medical records from the database and splits them into training and validation datasets.
[0793] Step 5:
[0794] The server trains the generative AI model using the training dataset, setting parameters such as the number of epochs and batch size, and running the model's learning process.
[0795] Step 6:
[0796] The server evaluates the performance of the trained model using a validation dataset and adjusts the model as needed.
[0797] Step 7:
[0798] The device uses the speech recognition API to enable voice input, preparing to convert what the user says into text in real time.
[0799] Step 8:
[0800] When interviewing patients, the device automatically generates the most appropriate questions based on a set list of questions and past response history.
[0801] Step 9:
[0802] The device converts the user's voice input into text in real time and immediately sends the converted text to the server.
[0803] Step 10:
[0804] The server analyzes the text data sent and uses a generative AI model to provide appropriate feedback to the user.
[0805] Step 11:
[0806] The device uses an emotion engine to recognize emotions from the user's voice and text data in real time, and transmits the emotional state to the server.
[0807] Step 12:
[0808] The server analyzes the emotion data sent from the emotion engine and adjusts the content of the feedback based on the user's emotions.
[0809] Step 13:
[0810] The server anonymizes patient data to enhance privacy and security, and the data is segmented and shared with relevant medical institutions and researchers as needed.
[0811] Step 14:
[0812] The server uses the generative AI model to make a diagnostic prediction based on the new patient data entered, and notifies the user of the diagnostic prediction results in real time.
[0813] Step 15:
[0814] The server checks the inventory status of medical resources (e.g., medical equipment and medicines) based on patient data and diagnosis predictions, and allocates or replenishes them as needed.
[0815] Step 16:
[0816] The server manages the access control list (ACL) and sets the access rights for each user. Data is protected using AES-256 encryption technology.
[0817] Step 17:
[0818] We regularly conduct security audits of our servers to ensure the safety of our systems. If a security incident occurs, we will respond immediately and take corrective measures.
[0819] Example 2
[0820] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0821] In modern healthcare, vast amounts of patient data are continually being generated, and managing, analyzing, and sharing this data presents major challenges. In particular, where many past medical records, including paper charts, remain, it is important to digitize and appropriately utilize them. There is also a need to understand the emotional state of patients through their voices and provide appropriate feedback and support. Furthermore, effective methods are needed to improve diagnostic accuracy and efficiently allocate medical resources.
[0822] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0823] In this invention, the server includes means for digitizing past medical records and storing them in a database, means for training medical datasets using a generative AI model, means for supporting patient interviews using voice recognition, means for anonymizing patient data to enhance security, means for making diagnostic predictions and supporting the optimal allocation of medical resources, means for ensuring data security through access control and data encryption, means for recognizing user emotions in real time and providing corresponding feedback, and means for sharing medical data with other medical institutions and researchers using security protocols. This enables efficient management of massive amounts of patient data, rapid feedback based on patient emotions, improved diagnostic accuracy, and optimal allocation of medical resources.
[0824] "Method of digitizing past medical records and storing them in a database" refers to the process of scanning paper medical records or existing electronic medical records, converting them into text data using OCR technology, and storing them in a database.
[0825] "Method of using a generative AI model to train a medical dataset" refers to the process of collecting medical data, dividing it into a training dataset and a validation dataset to train an AI model, and evaluating and adjusting the model's performance.
[0826] "Means for supporting patient hearings using voice recognition" refers to a process that uses a voice recognition API to convert voice input from patients into text data and support the hearing process.
[0827] "Methods for de-identifying and enhancing security of patient data" refers to technologies and processes used to de-identify personal patient information and enhance privacy and security.
[0828] "Means for making diagnostic predictions and supporting the appropriate allocation of medical resources" refers to the process of using generative AI models to make diagnostic predictions from new patient data and then allocating or replenishing necessary medical resources based on the results.
[0829] "Data security through access control and data encryption" refers to the process of using access control lists to set each user's access rights and protecting the data with AES-256 encryption technology.
[0830] "Means for recognizing a user's emotions in real time and providing corresponding feedback" refers to the technology and process for using an emotion engine to recognize emotions from a user's voice or input data and providing corresponding feedback in real time.
[0831] "Means of sharing medical data with other medical institutions and researchers using security protocols" refers to the process of appropriately anonymizing medical data and applying security protocols to safely share it with other medical institutions and researchers.
[0832] This invention is a system that promotes digitalization and cloud computing in the medical field and incorporates an emotion engine to provide efficient support to patients and medical staff. Specific embodiments of the present invention are described below.
[0833] Data collection and digitization
[0834] The user collects the patient's paper medical records and digitizes them using a scanner. The scanner captures the paper medical records as image data, and the terminal then converts them into text data using OCR (optical character recognition) technology. This text data is uploaded from the terminal to the server. The server converts the text data into an appropriate database format and stores it in the database.
[0835] Example: A user scans a medical record, and the device converts the image data into text using OCR software (e.g., ABBYY FineReader). The converted data is then uploaded to a cloud server and stored in a database.
[0836] Training generative AI models
[0837] The server extracts historical patient data and medical records from the database and splits them into a training dataset and a validation dataset for the generative AI model. The training dataset is used to train the AI model, verify its performance, and make adjustments as needed.
[0838] Example: The server uses frameworks such as Python and TensorFlow to input past diagnostic data into an AI model for learning. For example, training using random forests or neural networks is conceivable.
[0839] Building a voice support system
[0840] The device uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to enable voice input. When the user interviews a patient, the device converts the speech into text data in real time and sends it to the server. The server uses a generative AI model to analyze the context of the conversation and provide optimal feedback.
[0841] Example: When a user speaks to the device, the device converts the voice data into text data and sends it to the server. The server analyzes it and presents the next question.
[0842] Incorporating an emotion engine
[0843] The device uses an emotion engine to recognize emotions from the user's voice and input data. For example, it determines the user's emotional state in real time from the user's tone of voice and selected words. The server analyzes this data and generates feedback according to the user's emotions.
[0844] Example: If a user asks a question in an anxious tone, the server generates reassuring feedback and provides it to the user through the terminal.
[0845] Data partitioning and sharing
[0846] The server will appropriately anonymize patient data and enhance security. This anonymized data will be shared with relevant medical institutions and researchers as needed. Security protocols will be applied to sharing, and access rights will be strictly controlled.
[0847] Example: A server processes data using an anonymization algorithm (e.g., K-anonymization or differential privacy) and shares it over a secure communication protocol (e.g., HTTPs or TLS).
[0848] Diagnostic prediction and optimal allocation of medical resources
[0849] When new patient data is input, the server uses the generative AI model to predict a diagnosis. The prediction results are notified to the user in real time, and at the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes them as needed.
[0850] Example: The server inputs the symptom data of a new patient into an AI model, and if it predicts pneumonia, it notifies the user of the result and allocates the necessary medical resources.
[0851] Security and Privacy Measures
[0852] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology, and regular security audits ensure the system is secure.
[0853] For example: The server manages access privileges for each user based on ACLs, and stored data is encrypted with AES-256. Vulnerability checks are performed regularly using security audit tools (e.g., Nessus or OpenVAS).
[0854] Prompt Sentence Examples
[0855] Specifically, the following prompt sentences could be input into the generative AI model:
[0856] "Enter new patient data."
[0857] "Generate the next question to ask."
[0858] "Do sentiment analysis and provide appropriate feedback."
[0859] As described above, the system of the present invention integrates a wide range of functions to realize efficient data management, patient support, diagnostic prediction, and enhanced security in medical settings.
[0860] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0861] Step 1: Collect and digitize data
[0862] Input: Patient's paper chart
[0863] Output: Digitized text data
[0864] The user collects the patient's paper chart and places it into the scanner.
[0865] How it works: When the user presses the scan button, the scanner reads the paper medical record as image data.
[0866] The device receives the image data from the scanner and converts it into text data using OCR technology. For example, ABBYY FineReader installed on the device is used.
[0867] What happens: The device launches the OCR software and converts the image file into a text file.
[0868] The terminal uploads the converted text data to the server.
[0869] Operation: The device sends text data to the server's API endpoint.
[0870] Step 2: Saving to the database
[0871] Input: Converted text data
[0872] Output: Medical records stored in a database
[0873] The server converts the received text data into the appropriate database format.
[0874] How it works: The server parses the data format and generates SQL queries to insert into the database.
[0875] The server stores the text data in a database.
[0876] Step 3: Prepare training data for the generative AI model
[0877] Input: Medical records stored in a database
[0878] Output: Training and validation datasets
[0879] The server extracts medical record data from the database and splits it into a training dataset and a validation dataset.
[0880] How it works: The server executes SQL queries to extract data and runs an algorithm to randomly split the dataset.
[0881] Step 4: Training the generative AI model
[0882] Input: Teacher dataset
[0883] Output: A trained generative AI model
[0884] The server trains the AI model using the training dataset.
[0885] How it works: The server runs a script to train an AI model using Python or TensorFlow. It inputs the training dataset into the model and trains it.
[0886] Step 5: Building a voice support system
[0887] Input: User voice input
[0888] Output: Real-time audio data converted to text
[0889] The device uses the speech recognition API to enable the voice input function.
[0890] What it does: The device creates an instance of a speech recognition API (e.g., Google Cloud Speech-to-Text) and displays a UI to initiate voice input.
[0891] The user interviews the patient and inputs the information by voice.
[0892] The device converts the voice into text data in real time and sends it to the server.
[0893] Step 6: Analysis of audio data and feedback
[0894] Input: Real-time audio data converted to text
[0895] Output: Analysis results and feedback
[0896] The server analyzes the context of the conversation using a generative AI model.
[0897] How it works: The server inputs text data into the generative AI model and obtains the analysis results.
[0898] The server generates optimal feedback and provides it to the user through the terminal.
[0899] Step 7: Recognize emotions and provide feedback
[0900] Input: User voice and input data
[0901] Output: Emotional feedback
[0902] The terminal uses an emotion engine to recognize the user's emotion.
[0903] How it works: The emotion engine analyzes the user's tone of voice and word choice to determine their emotional state.
[0904] The server analyzes the emotional data and generates appropriate feedback for the user.
[0905] Operation: The server generates a feedback message based on the emotion data and provides it to the user via the terminal.
[0906] Step 8: Anonymize and share data
[0907] Input: Patient Data
[0908] Output: Anonymized data
[0909] The server appropriately anonymizes the patient data.
[0910] How it works: The server processes the data using an anonymization algorithm (e.g., K-anonymization or differential privacy).
[0911] The server will share the anonymized data with other medical institutions and researchers as needed.
[0912] How it works: Data is transmitted securely using security protocols (e.g., HTTPs or TLS).
[0913] Step 9: Diagnostic prediction and resource allocation
[0914] Input: New patient data
[0915] Output: Diagnosis results and medical resource allocation information
[0916] The server uses the generated AI model to make diagnostic predictions for new patient data.
[0917] How it works: New patient data is fed into a generative AI model to obtain a diagnosis.
[0918] The server notifies the user of the diagnosis results and checks the inventory status of medical resources (e.g., medical equipment and medicines) and allocates them appropriately.
[0919] Operation: The server queries the inventory management system and allocates and replenishes needed medical resources.
[0920] Step 10: Security and Privacy
[0921] Input: User ID and data
[0922] Output: Access permission settings and encrypted data
[0923] The server manages the access control list (ACL) and sets the access rights for each user.
[0924] How it works: The server updates the ACL database to manage user identities and access privileges.
[0925] The server protects your data with AES-256 encryption technology.
[0926] How it works: Before the server stores the data, it encrypts it using the AES-256 algorithm.
[0927] The server performs regular security audits to ensure the system is secure.
[0928] What happens: The server runs a vulnerability check using a security audit tool (e.g., Nessus or OpenVAS).
[0929] (Application example 2)
[0930] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0931] While conventional digital and cloud-based systems can efficiently manage and utilize medical records, they are inadequate for real-time support, such as displaying work instructions in real time or assessing workers' emotional states. They also have difficulty providing feedback and appropriate work support through voice recognition. Furthermore, they have faced problems with the psychological burden on workers and reduced work efficiency due to errors.
[0932] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0933] In this invention, the server includes means for digitizing past medical records and storing them in a database, means for training medical datasets using a generative AI model, means for supporting patient interviews using voice recognition, means for anonymizing patient data to enhance security, means for making diagnostic predictions and supporting the appropriate allocation of medical resources, means for ensuring data security through access control and data encryption, means for having a visual device for displaying work instructions in real time, means for providing feedback to workers using a voice response system, and means for assessing the emotional state of workers through emotion recognition. This makes it possible to evaluate the emotional state of workers while displaying work instructions in real time and providing feedback through voice recognition.
[0934] "Digitizing past medical records and storing them in a database" means converting paper medical records and existing electronic medical records into a digital format using scanning or optical character recognition technology (OCR) and storing them in a database.
[0935] "Learning medical datasets using a generative AI model" means training a generative AI model using datasets based on past patient data and medical records, enabling it to analyze medical data and make diagnostic predictions.
[0936] "Supporting patient interviews using voice recognition" means using voice recognition technology to convert interviews with patients into text in real time, and providing appropriate support to medical staff.
[0937] "Anonymizing patient data to enhance security" means protecting the privacy and enhancing the security of patient data by removing or transforming personally identifiable information.
[0938] "Making diagnostic predictions and supporting the appropriate allocation of medical resources" means using a generative AI model to make diagnostic predictions for patients and then appropriately allocating medical resources such as medical equipment and medicines based on the results.
[0939] "Ensuring data security through access control and data encryption" means setting access rights for each user and protecting data using encryption technology such as AES-256.
[0940] "Having a visual device for displaying work instructions in real time" means displaying work instructions in real time using a device such as smart glasses or a head-up display.
[0941] "Using a voice response system to provide feedback to workers" means using a microphone and voice recognition technology to recognize voice instructions from workers and provide appropriate responses or feedback.
[0942] "Assessing the emotional state of workers through emotion recognition" means using voice analysis and facial expression recognition technology to assess the emotional state of workers in real time and provide appropriate support based on the results.
[0943] This invention relates to a work support system for factories. This system exchanges information in real time between a server, terminals, and users, and provides multifunctional support to improve work efficiency.
[0944] Data collection and digitization
[0945] server:
[0946] Past medical records and work history are digitized and stored in a database. Paper medical records and handwritten work records are scanned and converted into text data using OCR (optical character recognition) technology. The converted data is then stored in a database for efficient management.
[0947] Training generative AI models
[0948] server:
[0949] Past work data and medical records are extracted from the database and divided into a training dataset and a validation dataset for the generative AI model. The training dataset is used to train the AI model, and its performance is evaluated and adjusted. This enables accurate diagnosis prediction and work instruction generation.
[0950] Building a voice support system
[0951] Device:
[0952] It uses a speech recognition API to enable voice input functionality. When users give voice instructions during work, the device converts the speech into text data in real time and provides appropriate feedback. It also uses generative AI models to automatically generate work instructions and questions.
[0953] Incorporating an emotion engine
[0954] Device:
[0955] It has an emotion engine that recognizes emotions from the user's voice and input data in real time, determines the emotional state from the tone of voice and words used, and provides feedback according to the emotion.
[0956] View work instructions in real time
[0957] Device:
[0958] Using smart glasses or head-up displays, work instructions are displayed in real time, allowing workers to receive visual information immediately and work efficiently.
[0959] ★Example:
[0960] For example, when a worker needs to install the next part in a factory, the smart glasses will display work instructions such as "Please take the next part, B, and install it on machine C." If the worker asks by voice, "What should I do next?", the next work instruction will be instantly provided via voice and text.
[0961] Emotion recognition for worker support
[0962] server:
[0963] The system analyzes the emotional data received from the emotion engine and provides appropriate support if a worker is feeling stressed or fatigued, such as suggesting slowing down work speed or taking a break.
[0964] Predicting abnormalities and presenting countermeasures
[0965] server:
[0966] If an abnormality occurs during work, the generative AI model is used to analyze the anomaly and suggest countermeasures, allowing for swift and appropriate measures to be taken.
[0967] ★Example prompt:
[0968] The current task is assembling parts. Please generate the next task instruction.
[0969] In this way, by implementing the present invention, not only can workers receive work instructions in real time and work efficiently, but it also enables support according to emotional states and the prediction of abnormalities, thereby improving overall work efficiency and safety.
[0970] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0971] Step 1:
[0972] Data collection and digitization
[0973] Users scan paper medical records or handwritten work records and input them into the terminal. The terminal then converts the scanned data into text data using OCR (optical character recognition) technology. This text data is sent to the server and stored in a database.
[0974] Input: Scanned paper charts and handwritten work records
[0975] Data processing: Text conversion using OCR technology
[0976] Output: Text data stored in the database
[0977] Step 2:
[0978] Training generative AI models
[0979] The server extracts past medical records and work data from the database and divides them into training and validation datasets for the generative AI model. The server trains the generative AI model using the training dataset and evaluates and adjusts the model's performance.
[0980] Input: Medical records and work data in a database
[0981] Data Computing: Learning and Evaluating Generative AI Models
[0982] Output: A trained generative AI model
[0983] Step 3:
[0984] Building a voice support system
[0985] The device uses a speech recognition API to enable voice input. The user gives voice instructions while working, and the device converts the speech into text data in real time. The server then uses a generative AI model to analyze the text data and generate appropriate feedback and work instructions.
[0986] Input: User's voice commands
[0987] Data processing: Text conversion using voice recognition, analysis using generative AI models
[0988] Output: Text data of feedback and work instructions
[0989] Step 4:
[0990] Incorporating emotion recognition
[0991] The device uses an emotion engine to recognize emotions from the user's voice and input data in real time. The recognized emotion data is sent to a server, which analyzes it and provides appropriate feedback to the user. For example, if the user is feeling stressed, the server may suggest taking a break.
[0992] Input: User voice and input data
[0993] Data calculation: Emotion recognition by emotion engine, analysis by server
[0994] Output: Emotional feedback
[0995] Step 5:
[0996] View work instructions in real time
[0997] The device displays work instructions in real time using smart glasses or a head-up display. The server uses a generative AI model to predict the next task and sends it to the device. The user receives the work instructions through the visual device, allowing them to work efficiently.
[0998] Input: Work instructions from the server
[0999] Data Computation: Task Prediction with Generative AI Models
[1000] Output: Work instructions displayed on a visual device
[1001] Step 6:
[1002] Predicting abnormalities and presenting countermeasures
[1003] The server monitors data generated during work in real time, and if an abnormality occurs, it analyzes it using a generative AI model. Based on the analysis results, it sends appropriate countermeasures to the device. The user can quickly resolve the abnormality by following the countermeasures provided by the device.
[1004] Input: Real-time data as you work
[1005] Data Computation: Anomaly Analysis with Generative AI Models
[1006] Output: Actions sent to the device
[1007] Step 7:
[1008] Emotion recognition for worker support
[1009] The server analyzes the emotion data received from the emotion engine and provides appropriate support if the worker is feeling stressed or fatigued, such as suggesting slowing down work speed or taking a break.
[1010] Input: Emotion data from the emotion engine
[1011] Data Computing: Server-Based Emotion Analysis
[1012] Output: Emotionally appropriate support and feedback
[1013] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1014] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1015] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1016] [Third embodiment]
[1017] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1018] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1019] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1020] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1021] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1022] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1023] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1024] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1025] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1026] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1027] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1028] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1029] This invention is a system that promotes digitalization and cloud computing in the medical industry. Below, we will create a program for this system and explain its processing in natural language. Specific examples will also be provided.
[1030] Data collection and digitization
[1031] User:
[1032] First, users collect the patient's paper medical records or existing electronic medical records, which are then digitized by scanning them, and then upload the digitized data to the server via a terminal.
[1033] Device:
[1034] The terminal receives the scanned paper medical records and converts them into text data using OCR (optical character recognition) technology, which is then sent to the server.
[1035] server:
[1036] The server converts the received text data into an appropriate database format and stores it in the database, allowing past medical records to be digitized and efficiently managed.
[1037] Training generative AI models
[1038] server:
[1039] The server extracts historical patient data and medical records from the database and divides them into training and validation datasets for the generative AI model. The training dataset is used to train the AI model, completing the model learning process. This process also includes adjusting hyperparameters such as the number of epochs and batch size. The model's performance is evaluated and adjustments are made as necessary.
[1040] Building a voice support system
[1041] Device:
[1042] The device uses a speech recognition API to enable voice input. When the user interviews a patient, the device converts the speech into text data in real time. It also generates optimal questions based on a pre-set question list and past response history.
[1043] server:
[1044] The server receives voice data in real time, converts it into text using speech recognition technology, and then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user.
[1045] Data partitioning and sharing
[1046] server:
[1047] The server will appropriately anonymize patient data to enhance privacy and security, and will partition and share this data with relevant medical institutions and researchers as needed, with security protocols in place and appropriate access rights managed.
[1048] Diagnostic prediction and optimal allocation of medical resources
[1049] server:
[1050] When new patient data is input, the server uses the generative AI model to predict diagnoses. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes them as necessary.
[1051] Security and Privacy Measures
[1052] server:
[1053] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology to prevent unauthorized access and data leaks. In addition, security audits are conducted regularly to ensure the safety of the system.
[1054] Specific examples
[1055] 1. The user scans Patient A's paper medical record and uploads it to the server via their device. The server converts it into text data using OCR technology and stores it in the database.
[1056] 2. When interviewing new patient B, the user uses voice input, and the device generates appropriate questions and converts them into text in real time. The server analyzes them using a generative AI model and provides feedback to the user.
[1057] 3. The server makes a diagnosis prediction and diagnoses Patient C as having a high possibility of influenza. The server checks the inventory of necessary medical equipment and medicines and allocates them appropriately.
[1058] 4. Patient data is anonymized and securely shared with other medical institutions for research purposes. The server applies access controls and data encryption to ensure privacy and security.
[1059] As described above, by specifically implementing the present invention, it is possible to realize efficient management and utilization of medical records, support for medical staff, and strengthen data security.
[1060] The processing flow will be explained below.
[1061] Step 1:
[1062] The user collects the patient's paper medical record or existing electronic medical record and scans the paper medical record, which then digitizes the data and loads it into the terminal.
[1063] Step 2:
[1064] The terminal converts scanned paper chart images into text data using OCR (Optical Character Recognition) technology, which is then sent to a server for further processing.
[1065] Step 3:
[1066] The server converts the received text data into an appropriate database format and stores it in the database. This standardizes the data and enables efficient management.
[1067] Step 4:
[1068] The server extracts historical patient data and medical records from the database and splits them into training and validation datasets.
[1069] Step 5:
[1070] The server trains the generative AI model using the training dataset, sets hyperparameters such as the number of epochs and batch size, and executes the model learning process.
[1071] Step 6:
[1072] The server evaluates the performance of the trained model using a validation dataset and adjusts the model as needed.
[1073] Step 7:
[1074] The device uses the speech recognition API to enable the voice input function, preparing the device for the user to use voice input.
[1075] Step 8:
[1076] When interviewing patients, the device automatically generates the most appropriate questions based on a pre-set list of questions and past response history.
[1077] Step 9:
[1078] The device converts the user's voice input into text in real time and immediately sends the text to the server.
[1079] Step 10:
[1080] The server analyzes the text data sent and uses a generative AI model to provide appropriate feedback to the user.
[1081] Step 11:
[1082] The server anonymizes patient data, removing or transforming personally identifiable information to protect the confidentiality of the data.
[1083] Step 12:
[1084] The server will divide the anonymized data and share it with medical institutions and researchers who need it, and will also set up security protocols and manage access rights.
[1085] Step 13:
[1086] When new patient data is entered into the server, the generative AI model is used to make a diagnostic prediction, and the results are notified to the user in real time.
[1087] Step 14:
[1088] The server checks the inventory status of medical resources (medical equipment and medicines) based on patient data and diagnosis predictions, and allocates or replenishes them as necessary.
[1089] Step 15:
[1090] The server manages the access control list (ACL) and sets the access rights for each user. Data is protected using AES-256 encryption technology.
[1091] Step 16:
[1092] We regularly conduct security audits on our servers to ensure the safety of our systems. In the event of a security incident, we respond immediately and take corrective measures.
[1093] Example 1
[1094] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1095] Problems with conventional medical systems include the lack of progress in digitizing medical records and the ineffective use of collected data. Furthermore, there is a lack of systems that support the accuracy of diagnostic predictions and the appropriate allocation of medical resources. There are also many shortcomings in terms of security and privacy protection. These issues are major obstacles to improving the efficiency and quality of medical care, and they require rapid and accurate responses.
[1096] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1097] In this invention, the server includes a means for scanning and digitizing past medical records, converting them into text data using OCR technology, and storing them in a database, a means for extracting and training data sets using a generative AI model, and a means for anonymizing patient data and using technology to enhance security, thereby enabling efficient management and utilization of medical records, improving the accuracy of diagnostic predictions, and optimizing the allocation of medical resources.
[1098] A "database format" is a structure or format for efficiently storing, retrieving, and managing digital data.
[1099] "OCR technology" is an abbreviation for optical character recognition technology, which extracts character information from scanned images and converts it into text data.
[1100] A "generative AI model" refers to an algorithm or model that uses artificial intelligence techniques to generate new information or predictions from a given dataset.
[1101] "Voice recognition technology" is a technology that converts voice input into text data in real time.
[1102] "Data anonymization" refers to the process of removing or modifying personally identifiable information to protect the privacy of the data.
[1103] A "security protocol" is a set of procedures and rules that are used to prevent unauthorized access when transmitting or receiving data.
[1104] "Diagnostic prediction" is the process of using generative AI models based on medical data to predict illness and treatment.
[1105] An "access control list (ACL)" is a list used to manage the rights and permissions of each user accessing a system.
[1106] "Data encryption" is a technology that encrypts important information using a specific algorithm to prevent unauthorized access and data leaks.
[1107] "Hyperparameter tuning" is the process of fine-tuning the parameters of a machine learning model to optimize its learning efficiency and performance.
[1108] The present invention is a system that promotes digitalization and cloud computing in the medical industry. The program processing of this system will be explained below in natural language.
[1109] Data collection and digitization
[1110] User:
[1111] First, the user collects the patient's paper medical records or existing electronic medical records. They use a general-purpose scanner to scan the paper medical records and convert them into digital data. For example, they use a general-purpose high-performance scanner. Next, the user uploads the scanned digital data to a server via a terminal. This is done using the hospital's internal network.
[1112] Device:
[1113] The terminal receives the scanned image data of the paper medical record and converts it into text data using OCR technology. This technology uses commonly used OCR software. The converted text data is sent to the server. When sending, the communication is encrypted using SSL / TLS.
[1114] server:
[1115] The server receives the text data sent from the terminal. The received data is converted into an appropriate database format and stored in a database. For example, a relational database management system (RDBMS) is used as the database.
[1116] Training generative AI models
[1117] server:
[1118] The server extracts past patient data and medical records from the database. The extracted data is divided into a training dataset and a validation dataset for the generative AI model. The server trains the generative AI model using the training dataset. Deep learning libraries such as PyTorch and TensorFlow are used for this training. Hyperparameters such as the number of epochs and batch size are also adjusted. The model's performance is evaluated on the validation dataset, and readjustments are made as necessary.
[1119] Building a voice support system
[1120] Device:
[1121] The device uses a speech recognition API to enable voice input. For example, a general speech recognition API can be used. When the user interviews a patient, the device converts the speech into text data in real time. It also generates optimal questions based on a pre-set question list and past response history.
[1122] server:
[1123] The server receives the voice data in real time, converts it into text using speech recognition technology, and then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user.
[1124] Data partitioning and sharing
[1125] server:
[1126] The server will appropriately anonymize patient data to enhance privacy and security. This will involve techniques to remove or modify personally identifiable information. Furthermore, the anonymized data will be split up as needed and shared with relevant medical institutions and researchers. This sharing will be subject to security protocols and appropriate access rights management.
[1127] Diagnostic prediction and optimal allocation of medical resources
[1128] server:
[1129] When new patient data is input, the server uses the generative AI model to predict diagnoses. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes these resources as needed.
[1130] Security and Privacy Measures
[1131] server:
[1132] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology to prevent unauthorized access and data leaks. In addition, security audits are conducted regularly to ensure the safety of the system.
[1133] Specific examples
[1134] 1. The user scans Patient A's paper medical record and uploads it to the server via their device. The server converts it into text data using OCR technology and stores it in the database.
[1135] 2. When interviewing new patient B, the user uses voice input, and the device generates appropriate questions and converts them into text in real time. The server analyzes them using a generative AI model and provides feedback to the user.
[1136] 3. The server makes a diagnosis prediction and diagnoses Patient C as having a high possibility of influenza. The server checks the inventory of necessary medical equipment and medicines and allocates them appropriately.
[1137] 4. Patient data is anonymized and securely shared with other medical institutions for research purposes. The server applies access controls and data encryption to ensure privacy and security.
[1138] Prompt Sentence Examples
[1139] "How do you digitize and store Patient A's paper medical records?"
[1140] "How do I generate the best questions for a new patient, Patient B?"
[1141] "How can we predict diagnoses and allocate necessary medical resources appropriately?"
[1142] As described above, by specifically implementing the present invention, it is possible to realize efficient management and utilization of medical records, support for medical staff, and strengthen data security.
[1143] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1144] Step 1:
[1145] User: The user obtains digital image data by scanning a patient's paper chart. This input data is in an image file format (e.g., JPEG or PNG). The specific actions of scanning using a scanner include placing the paper chart on the scanner and pressing the Scan button. The output is the scanned digital image data.
[1146] Step 2:
[1147] User: The user uploads the scanned digital image data to the terminal. The uploading process involves sending the digital image to a designated folder using the hospital network. The input is the scanned digital image, and the output is the uploaded digital image in the terminal.
[1148] Step 3:
[1149] Terminal: The terminal converts the received digital image data into text data using OCR technology. Specific operations include launching the OCR software, reading the digital image, and performing character recognition. The input is digital image data, and the output is the converted text data.
[1150] Step 4:
[1151] Terminal: The terminal sends the text data generated by OCR to the server. This transmission process uses SSL / TLS encrypted communication. The input is the converted text data, and the output is the text data sent to the server.
[1152] Step 5:
[1153] Server: The server converts the received text data into a database format and stores it in the database. Specific operations include converting the text data into an SQL query format and performing an insert operation on the database. The input is the text data sent to the server, and the output is the medical record data stored in the database.
[1154] Step 6:
[1155] Server: The server extracts historical patient data and medical records from the database and splits them into training and validation datasets for the generative AI model. Specific operations include executing SQL queries to retrieve the data and splitting the extracted data into training and validation datasets. The input is the medical record data stored in the database, and the output is the split dataset.
[1156] Step 7:
[1157] Server: The server trains the generative AI model using the training dataset. This is done using deep learning libraries such as PyTorch or TensorFlow. It also adjusts hyperparameters such as the number of epochs and batch size. The input is the training dataset, and the output is the trained generative AI model.
[1158] Step 8:
[1159] Server: The server evaluates the model's performance on the validation dataset and adjusts hyperparameters or modifies the model structure as necessary. Specific operations include calculating metrics for model evaluation and retraining. The input is the validation dataset and the existing generative AI model, and the output is an optimized generative AI model.
[1160] Step 9:
[1161] Terminal: The terminal uses a speech recognition API to enable voice input and converts the user's voice during patient interviews into text data in real time. It also generates optimal questions based on a pre-set question list and past response history. The input is voice data, and the output is text data and the generated questions.
[1162] Step 10:
[1163] Server: The server receives voice data in real time and converts it into text data using speech recognition technology. It then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user. The input is voice data, and the output is interpreted text data and feedback.
[1164] Step 11:
[1165] Server: The server anonymizes patient data, removes or modifies any personally identifiable information from the data, and applies security protocols and manages access rights to share the anonymized data with relevant medical institutions and researchers. The input is raw patient data, and the output is anonymized data.
[1166] Step 12:
[1167] Server: When new patient data is input, the server uses the generative AI model to make a diagnosis prediction. The prediction results are notified to the user in real time, and at the same time, the server checks the inventory status of medical resources and allocates or replenishes them as needed. The input is the new patient data, and the output is the diagnosis prediction results and medical resource management information.
[1168] Step 13:
[1169] Server: The server manages the access control list (ACL) and sets the access rights for each user. It also protects data with AES-256 encryption technology and performs regular security audits to prevent unauthorized access and data leakage. The input is user information and data, and the output is controlled access rights and protected data.
[1170] Through these steps, the system achieves efficient digitization of medical records, diagnostic prediction through generative AI models, and data security management.
[1171] (Application example 1)
[1172] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1173] The modern healthcare industry requires the digitization of medical records and efficient data management. However, several challenges exist, including the digitization of paper medical records, the centralized management of existing electronic medical records, the anonymization and security of medical data, and the improvement of diagnostic prediction accuracy. In particular, there is a lack of support systems for quickly and accurately recording patient interactions and making appropriate diagnoses. It is also important to safely anonymize collected data and share it with other medical institutions and researchers. It is necessary to resolve these challenges and realize digitalization and enhanced security in the healthcare industry.
[1174] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1175] In this invention, the server includes means for digitizing past medical records and storing them in a database, means for training medical datasets using a generative AI model, means for supporting patient interviews using voice recognition, means for anonymizing patient data to enhance security, means for making diagnostic predictions and supporting the appropriate allocation of medical resources, means for ensuring data security through access control and data encryption, means for generating text data from scanned images using optical character recognition technology, means for converting voice input into text data in real time using a voice recognition API, and means for checking the availability of medical resources based on diagnostic predictions using the generative AI model. This enables efficient digitization and management of medical records, rapid recording of patient interactions, appropriate diagnostic support, and secure data sharing.
[1176] "Prior medical records" are previously collected medical data, including a patient's medical history and treatment history.
[1177] "Digitalization" is the process of converting information in analog form, such as paper, into digital form.
[1178] A "database" is a system for storing and managing data in an organized structure.
[1179] A "generative AI model" is a machine learning algorithm that learns from large amounts of data and is optimized to perform a specific task.
[1180] A "medical dataset" is a collection of various medical-related data used for learning and analysis.
[1181] "Speech recognition" is the technology that takes voice input and converts it into a digital format such as text.
[1182] A "patient interview" is a process in which a medical professional directly obtains information from a patient, such as symptoms and medical history.
[1183] "Anonymization" is the process of processing data so that individuals cannot be identified.
[1184] "Security" refers to a group of technologies aimed at protecting data and preventing unauthorized access.
[1185] "Diagnostic prediction" is the process of predicting future disease states and appropriate treatments based on collected medical data.
[1186] "Medical resources" is a general term for all equipment, medicines, personnel, etc. used in medical institutions.
[1187] "Proper allocation" means properly allocating the necessary resources to the necessary locations.
[1188] "Access control" is a mechanism that allows only authorized users to access specific data.
[1189] "Data encryption" is the technology that encrypts data and converts it into a form that is unintelligible to unauthorized users.
[1190] "Optical character recognition technology" is a technology that analyzes character information in an image and converts it into text data.
[1191] "Real-time" means that data input, processing, and output are carried out instantly.
[1192] A "speech recognition API" is an application programming interface that provides the functionality to convert voice input into text.
[1193] "Inventory status" is information that indicates the current level of medical resources available for use.
[1194] This invention is a system that promotes digitalization and cloud computing in the medical industry, and provides a wide range of functions, including digitization of past medical records, diagnostic prediction using generative AI models, patient interview support using voice recognition, and enhanced data security using anonymization technology.
[1195] Data collection and digitization
[1196] User:
[1197] Users first scan the patient's paper medical records using their smartphone camera and upload them to the server via the device, allowing them to be digitized and managed efficiently.
[1198] Application of OCR technology
[1199] Device:
[1200] The terminal uses optical character recognition (OCR) technology to generate text data from scanned images of paper medical records, which is then sent to a server and stored in a database.
[1201] Building a voice support system
[1202] Device:
[1203] The device uses a speech recognition API to convert voice input into text data in real time. The user records conversations with patients and converts the data into text in real time. The device also generates optimal questions based on a pre-set question list and past response history.
[1204] Data storage and anonymization
[1205] server:
[1206] The server converts the received text data into an appropriate database format and stores it in the database. The server anonymizes the data and uses data encryption technology (e.g., AES-256) to enhance privacy and security. This data is split as needed and shared with relevant medical institutions and researchers.
[1207] Training generative AI models
[1208] server:
[1209] The server extracts historical patient data and medical records from the database and divides them into training and validation datasets for the generative AI model. The training dataset is used to train the AI model, completing the model learning process. The model's performance is evaluated and adjusted as necessary.
[1210] Diagnostic prediction and optimal allocation of medical resources
[1211] server:
[1212] When new patient data is entered, the server uses the generative AI model to predict a diagnosis. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (medical equipment, drugs, etc.) and allocates or replenishes them as needed.
[1213] Examples of concrete examples and prompts
[1214] 1. A user scans a patient's paper chart and uploads it to the server via a terminal. An example of the prompt is as follows:
[1215] "Take images of medical records, digitize them, and upload them to the cloud."
[1216] 2. The user records the conversation with the patient and uses a speech recognition API to transcribe it in real time, with prompts such as:
[1217] "Record conversations with patients, convert them into text, and save them."
[1218] 3. The server performs diagnosis prediction, checks the availability of medical resources, and allocates them appropriately. An example of the prompt is as follows:
[1219] "Using generative AI models to predict diagnoses and allocate necessary medical resources"
[1220] This enables efficient digitization and management of medical records, enables prompt recording of patient interactions, and ensures proper diagnostic support and secure data sharing, significantly improving operational efficiency and security in the healthcare industry.
[1221] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1222] Step 1:
[1223] The user scans the patient's paper medical record using the smartphone camera and uploads it to the server via the terminal. The input is an image of the paper medical record, and the output is digitized text data. The user launches the scanning app on the terminal and takes a picture of the paper medical record.
[1224] Step 2:
[1225] The device generates text data from the scanned image using optical character recognition (OCR). The input is the scanned image from step 1, and the output is the text data generated by OCR. The device sends the scanned image to the OCR engine and obtains the text data.
[1226] Step 3:
[1227] The device sends the generated text data to the server. The input is the text data, and the output is the data sent to the server. The device then uploads the data to the cloud server via the network.
[1228] Step 4:
[1229] The server converts the received text data into an appropriate database format and stores it in the database. The input is the text data received by the server, and the output is the data stored in the database. The server formats the text data into the database format and stores it.
[1230] Step 5:
[1231] The server uses a speech recognition API to convert the voice data recorded by the user into text data in real time. The input is voice data and the output is text data. The server sends the voice data to the API and saves the resulting text data.
[1232] Step 6:
[1233] The server anonymizes the received data and uses data encryption techniques (e.g., AES-256) to enhance privacy and security. The input is text data stored in a database, and the output is anonymized and encrypted data. The server anonymizes the data and applies encryption techniques.
[1234] Step 7:
[1235] The server uses a generative AI model to split the past patient data and medical records extracted from the database into a training dataset and a validation dataset. The input is the past patient data and medical records, and the output is the training dataset and the validation dataset. The server splits the data into training and validation datasets.
[1236] Step 8:
[1237] The server trains the generative AI model using the training dataset, completing the model learning process. The input is the training dataset, and the output is the trained generative AI model. The server uses the dataset to train the model.
[1238] Step 9:
[1239] When new patient data is input, the server uses the generative AI model to make a diagnosis prediction. The input is the new patient data, and the output is the diagnosis prediction result. The server inputs the new patient data into the model and obtains the diagnosis prediction result.
[1240] Step 10:
[1241] The server checks the inventory status of medical resources (medical equipment, medicines, etc.) based on the diagnosis prediction, and allocates or replenishes them as needed. The input is the diagnosis prediction result, and the output is confirmation of the inventory status and the appropriate allocation of resources. The server works with the inventory system based on the prediction results to optimally allocate resources.
[1242] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1243] This invention combines an emotion engine with a system that promotes digitalization and cloud computing in the medical industry. Below, we will generate a program for this system and explain its processing in natural language. Specific examples will also be provided.
[1244] Data collection and digitization
[1245] User:
[1246] First, the user collects the patient's paper medical records or existing electronic medical records. The paper medical records are digitized through scanning and imported into the terminal. The digitized data is then uploaded to the server via the terminal.
[1247] Device:
[1248] The terminal receives the scanned paper medical records and converts them into text data using OCR (optical character recognition) technology, which is then sent to the server.
[1249] server:
[1250] The server converts the received text data into an appropriate database format and stores it in the database, allowing past medical records to be digitized and efficiently managed.
[1251] Training generative AI models
[1252] server:
[1253] The server extracts historical patient data and medical records from the database and splits them into training and validation datasets for the generative AI model. The training dataset is used to train the AI model, completing the model learning process. The model's performance is evaluated and adjustments are made as needed.
[1254] Building a voice support system
[1255] Device:
[1256] The device uses a speech recognition API to enable voice input. When the user interviews a patient, the device converts the speech into text data in real time. It also generates optimal questions based on a pre-set question list and past response history.
[1257] server:
[1258] The server receives voice data in real time, converts it into text using speech recognition technology, and then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user.
[1259] Incorporating an emotion engine
[1260] Device:
[1261] The device has an emotion engine that recognizes emotions from the user's voice and input data in real time, for example, determining the user's emotional state from the tone of their voice and the words they choose.
[1262] server:
[1263] The server analyzes the emotion data sent from the emotion engine and generates feedback according to the user's emotions. If the user is feeling stressed or anxious, the system will suggest appropriate responses and support.
[1264] Data partitioning and sharing
[1265] server:
[1266] The server will appropriately anonymize patient data to enhance privacy and security, and will partition and share this data with relevant medical institutions and researchers as needed, with security protocols in place and appropriate access rights managed.
[1267] Diagnostic prediction and optimal allocation of medical resources
[1268] server:
[1269] When new patient data is input, the server uses the generative AI model to predict diagnoses. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes them as necessary.
[1270] Security and Privacy Measures
[1271] server:
[1272] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology to prevent unauthorized access and data leaks. In addition, security audits are conducted regularly to ensure the safety of the system.
[1273] Specific examples
[1274] 1. The user scans Patient D's paper medical record and uploads it to the server via a terminal. The server uses OCR technology to convert the medical record into text data and saves it in the database.
[1275] 2. When interviewing Patient E, the user uses voice input, and the device generates appropriate questions and converts them into text in real time. The server analyzes the results using a generative AI model and provides feedback to the user. If the emotion engine detects stress in the user, it provides appropriate support.
[1276] 3. The server makes a diagnosis prediction and checks the inventory of necessary medical equipment and medicines for Patient F, who is diagnosed with possible pneumonia, and allocates them appropriately.
[1277] 4. Patient data is anonymized and securely shared with other medical institutions for research purposes. The server applies access controls and data encryption to ensure privacy and security.
[1278] As described above, by specifically implementing the present invention, it is possible to efficiently manage and utilize medical records, support medical staff, strengthen data security, and even provide appropriate feedback according to the user's emotions.
[1279] The processing flow will be explained below.
[1280] Step 1:
[1281] The user collects the patient's paper or electronic records and scans the paper records, which then digitizes the data and loads it into the terminal.
[1282] Step 2:
[1283] The terminal converts the scanned paper chart image into text data using OCR (optical character recognition) technology, which is then sent to the server.
[1284] Step 3:
[1285] The server converts the received text data into an appropriate database format and stores it in the database, thereby digitizing the medical records.
[1286] Step 4:
[1287] The server extracts historical patient data and medical records from the database and splits them into training and validation datasets.
[1288] Step 5:
[1289] The server trains the generative AI model using the training dataset, setting parameters such as the number of epochs and batch size, and running the model's learning process.
[1290] Step 6:
[1291] The server evaluates the performance of the trained model using a validation dataset and adjusts the model as needed.
[1292] Step 7:
[1293] The device uses the speech recognition API to enable voice input, preparing to convert what the user says into text in real time.
[1294] Step 8:
[1295] When interviewing patients, the device automatically generates the most appropriate questions based on a set list of questions and past response history.
[1296] Step 9:
[1297] The device converts the user's voice input into text in real time and immediately sends the converted text to the server.
[1298] Step 10:
[1299] The server analyzes the text data sent and uses a generative AI model to provide appropriate feedback to the user.
[1300] Step 11:
[1301] The device uses an emotion engine to recognize emotions from the user's voice and text data in real time, and transmits the emotional state to the server.
[1302] Step 12:
[1303] The server analyzes the emotion data sent from the emotion engine and adjusts the content of the feedback based on the user's emotions.
[1304] Step 13:
[1305] The server anonymizes patient data to enhance privacy and security, and the data is segmented and shared with relevant medical institutions and researchers as needed.
[1306] Step 14:
[1307] The server uses the generative AI model to make a diagnostic prediction based on the new patient data entered, and notifies the user of the diagnostic prediction results in real time.
[1308] Step 15:
[1309] The server checks the inventory status of medical resources (e.g., medical equipment and medicines) based on patient data and diagnosis predictions, and allocates or replenishes them as needed.
[1310] Step 16:
[1311] The server manages the access control list (ACL) and sets the access rights for each user. Data is protected using AES-256 encryption technology.
[1312] Step 17:
[1313] We regularly conduct security audits of our servers to ensure the safety of our systems. If a security incident occurs, we will respond immediately and take corrective measures.
[1314] Example 2
[1315] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1316] In modern healthcare, vast amounts of patient data are continually being generated, and managing, analyzing, and sharing this data presents major challenges. In particular, where many past medical records, including paper charts, remain, it is important to digitize and appropriately utilize them. There is also a need to understand the emotional state of patients through their voices and provide appropriate feedback and support. Furthermore, effective methods are needed to improve diagnostic accuracy and efficiently allocate medical resources.
[1317] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1318] In this invention, the server includes means for digitizing past medical records and storing them in a database, means for training medical datasets using a generative AI model, means for supporting patient interviews using voice recognition, means for anonymizing patient data to enhance security, means for making diagnostic predictions and supporting the optimal allocation of medical resources, means for ensuring data security through access control and data encryption, means for recognizing user emotions in real time and providing corresponding feedback, and means for sharing medical data with other medical institutions and researchers using security protocols. This enables efficient management of massive amounts of patient data, rapid feedback based on patient emotions, improved diagnostic accuracy, and optimal allocation of medical resources.
[1319] "Method of digitizing past medical records and storing them in a database" refers to the process of scanning paper medical records or existing electronic medical records, converting them into text data using OCR technology, and storing them in a database.
[1320] "Method of using a generative AI model to train a medical dataset" refers to the process of collecting medical data, dividing it into a training dataset and a validation dataset to train an AI model, and evaluating and adjusting the model's performance.
[1321] "Means for supporting patient hearings using voice recognition" refers to a process that uses a voice recognition API to convert voice input from patients into text data and support the hearing process.
[1322] "Methods for de-identifying and enhancing security of patient data" refers to technologies and processes used to de-identify personal patient information and enhance privacy and security.
[1323] "Means for making diagnostic predictions and supporting the appropriate allocation of medical resources" refers to the process of using generative AI models to make diagnostic predictions from new patient data and then allocating or replenishing necessary medical resources based on the results.
[1324] "Data security through access control and data encryption" refers to the process of using access control lists to set each user's access rights and protecting the data with AES-256 encryption technology.
[1325] "Means for recognizing a user's emotions in real time and providing corresponding feedback" refers to the technology and process for using an emotion engine to recognize emotions from a user's voice or input data and providing corresponding feedback in real time.
[1326] "Means of sharing medical data with other medical institutions and researchers using security protocols" refers to the process of appropriately anonymizing medical data and applying security protocols to safely share it with other medical institutions and researchers.
[1327] This invention is a system that promotes digitalization and cloud computing in the medical field and incorporates an emotion engine to provide efficient support to patients and medical staff. Specific embodiments of the present invention are described below.
[1328] Data collection and digitization
[1329] The user collects the patient's paper medical records and digitizes them using a scanner. The scanner captures the paper medical records as image data, and the terminal then converts them into text data using OCR (optical character recognition) technology. This text data is uploaded from the terminal to the server. The server converts the text data into an appropriate database format and stores it in the database.
[1330] Example: A user scans a medical record, and the device converts the image data into text using OCR software (e.g., ABBYY FineReader). The converted data is then uploaded to a cloud server and stored in a database.
[1331] Training generative AI models
[1332] The server extracts historical patient data and medical records from the database and splits them into a training dataset and a validation dataset for the generative AI model. The training dataset is used to train the AI model, verify its performance, and make adjustments as needed.
[1333] Example: The server uses frameworks such as Python and TensorFlow to input past diagnostic data into an AI model for learning. For example, training using random forests or neural networks is conceivable.
[1334] Building a voice support system
[1335] The device uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to enable voice input. When the user interviews a patient, the device converts the speech into text data in real time and sends it to the server. The server uses a generative AI model to analyze the context of the conversation and provide optimal feedback.
[1336] Example: When a user speaks to the device, the device converts the voice data into text data and sends it to the server. The server analyzes it and presents the next question.
[1337] Incorporating an emotion engine
[1338] The device uses an emotion engine to recognize emotions from the user's voice and input data. For example, it determines the user's emotional state in real time from the user's tone of voice and selected words. The server analyzes this data and generates feedback according to the user's emotions.
[1339] Example: If a user asks a question in an anxious tone, the server generates reassuring feedback and provides it to the user through the terminal.
[1340] Data partitioning and sharing
[1341] The server will appropriately anonymize patient data and enhance security. This anonymized data will be shared with relevant medical institutions and researchers as needed. Security protocols will be applied to sharing, and access rights will be strictly controlled.
[1342] Example: A server processes data using an anonymization algorithm (e.g., K-anonymization or differential privacy) and shares it over a secure communication protocol (e.g., HTTPs or TLS).
[1343] Diagnostic prediction and optimal allocation of medical resources
[1344] When new patient data is input, the server uses the generative AI model to predict a diagnosis. The prediction results are notified to the user in real time, and at the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes them as needed.
[1345] Example: The server inputs the symptom data of a new patient into an AI model, and if it predicts pneumonia, it notifies the user of the result and allocates the necessary medical resources.
[1346] Security and Privacy Measures
[1347] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology, and regular security audits ensure the system is secure.
[1348] For example: The server manages access privileges for each user based on ACLs, and stored data is encrypted with AES-256. Vulnerability checks are performed regularly using security audit tools (e.g., Nessus or OpenVAS).
[1349] Prompt Sentence Examples
[1350] Specifically, the following prompt sentences could be input into the generative AI model:
[1351] "Enter new patient data."
[1352] "Generate the next question to ask."
[1353] "Do sentiment analysis and provide appropriate feedback."
[1354] As described above, the system of the present invention integrates a wide range of functions to realize efficient data management, patient support, diagnostic prediction, and enhanced security in medical settings.
[1355] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1356] Step 1: Collect and digitize data
[1357] Input: Patient's paper chart
[1358] Output: Digitized text data
[1359] The user collects the patient's paper chart and places it into the scanner.
[1360] How it works: When the user presses the scan button, the scanner reads the paper medical record as image data.
[1361] The device receives the image data from the scanner and converts it into text data using OCR technology. For example, ABBYY FineReader installed on the device is used.
[1362] What happens: The device launches the OCR software and converts the image file into a text file.
[1363] The terminal uploads the converted text data to the server.
[1364] Operation: The device sends text data to the server's API endpoint.
[1365] Step 2: Saving to the database
[1366] Input: Converted text data
[1367] Output: Medical records stored in a database
[1368] The server converts the received text data into the appropriate database format.
[1369] How it works: The server parses the data format and generates SQL queries to insert into the database.
[1370] The server stores the text data in a database.
[1371] Step 3: Prepare training data for the generative AI model
[1372] Input: Medical records stored in a database
[1373] Output: Training and validation datasets
[1374] The server extracts medical record data from the database and splits it into a training dataset and a validation dataset.
[1375] How it works: The server executes SQL queries to extract data and runs an algorithm to randomly split the dataset.
[1376] Step 4: Training the generative AI model
[1377] Input: Teacher dataset
[1378] Output: A trained generative AI model
[1379] The server trains the AI model using the training dataset.
[1380] How it works: The server runs a script to train an AI model using Python or TensorFlow. It inputs the training dataset into the model and trains it.
[1381] Step 5: Building a voice support system
[1382] Input: User voice input
[1383] Output: Real-time audio data converted to text
[1384] The device uses the speech recognition API to enable the voice input function.
[1385] What it does: The device creates an instance of a speech recognition API (e.g., Google Cloud Speech-to-Text) and displays a UI to initiate voice input.
[1386] The user interviews the patient and inputs the information by voice.
[1387] The device converts the voice into text data in real time and sends it to the server.
[1388] Step 6: Analysis of audio data and feedback
[1389] Input: Real-time audio data converted to text
[1390] Output: Analysis results and feedback
[1391] The server analyzes the context of the conversation using a generative AI model.
[1392] How it works: The server inputs text data into the generative AI model and obtains the analysis results.
[1393] The server generates optimal feedback and provides it to the user through the terminal.
[1394] Step 7: Recognize emotions and provide feedback
[1395] Input: User voice and input data
[1396] Output: Emotional feedback
[1397] The terminal uses an emotion engine to recognize the user's emotion.
[1398] How it works: The emotion engine analyzes the user's tone of voice and word choice to determine their emotional state.
[1399] The server analyzes the emotional data and generates appropriate feedback for the user.
[1400] Operation: The server generates a feedback message based on the emotion data and provides it to the user via the terminal.
[1401] Step 8: Anonymize and share data
[1402] Input: Patient Data
[1403] Output: Anonymized data
[1404] The server appropriately anonymizes the patient data.
[1405] How it works: The server processes the data using an anonymization algorithm (e.g., K-anonymization or differential privacy).
[1406] The server will share the anonymized data with other medical institutions and researchers as needed.
[1407] How it works: Data is transmitted securely using security protocols (e.g., HTTPs or TLS).
[1408] Step 9: Diagnostic prediction and resource allocation
[1409] Input: New patient data
[1410] Output: Diagnosis results and medical resource allocation information
[1411] The server uses the generated AI model to make diagnostic predictions for new patient data.
[1412] How it works: New patient data is fed into a generative AI model to obtain a diagnosis.
[1413] The server notifies the user of the diagnosis results and checks the inventory status of medical resources (e.g., medical equipment and medicines) and allocates them appropriately.
[1414] Operation: The server queries the inventory management system and allocates and replenishes needed medical resources.
[1415] Step 10: Security and Privacy
[1416] Input: User ID and data
[1417] Output: Access permission settings and encrypted data
[1418] The server manages the access control list (ACL) and sets the access rights for each user.
[1419] How it works: The server updates the ACL database to manage user identities and access privileges.
[1420] The server protects your data with AES-256 encryption technology.
[1421] How it works: Before the server stores the data, it encrypts it using the AES-256 algorithm.
[1422] The server performs regular security audits to ensure the system is secure.
[1423] What happens: The server runs a vulnerability check using a security audit tool (e.g., Nessus or OpenVAS).
[1424] (Application example 2)
[1425] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1426] While conventional digital and cloud-based systems can efficiently manage and utilize medical records, they are inadequate for real-time support, such as displaying work instructions in real time or assessing workers' emotional states. They also have difficulty providing feedback and appropriate work support through voice recognition. Furthermore, they have faced problems with the psychological burden on workers and reduced work efficiency due to errors.
[1427] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1428] In this invention, the server includes means for digitizing past medical records and storing them in a database, means for training medical datasets using a generative AI model, means for supporting patient interviews using voice recognition, means for anonymizing patient data to enhance security, means for making diagnostic predictions and supporting the appropriate allocation of medical resources, means for ensuring data security through access control and data encryption, means for having a visual device for displaying work instructions in real time, means for providing feedback to workers using a voice response system, and means for assessing the emotional state of workers through emotion recognition. This makes it possible to evaluate the emotional state of workers while displaying work instructions in real time and providing feedback through voice recognition.
[1429] "Digitizing past medical records and storing them in a database" means converting paper medical records and existing electronic medical records into a digital format using scanning or optical character recognition technology (OCR) and storing them in a database.
[1430] "Learning medical datasets using a generative AI model" means training a generative AI model using datasets based on past patient data and medical records, enabling it to analyze medical data and make diagnostic predictions.
[1431] "Supporting patient interviews using voice recognition" means using voice recognition technology to convert interviews with patients into text in real time, and providing appropriate support to medical staff.
[1432] "Anonymizing patient data to enhance security" means protecting the privacy and enhancing the security of patient data by removing or transforming personally identifiable information.
[1433] "Making diagnostic predictions and supporting the appropriate allocation of medical resources" means using a generative AI model to make diagnostic predictions for patients and then appropriately allocating medical resources such as medical equipment and medicines based on the results.
[1434] "Ensuring data security through access control and data encryption" means setting access rights for each user and protecting data using encryption technology such as AES-256.
[1435] "Having a visual device for displaying work instructions in real time" means displaying work instructions in real time using a device such as smart glasses or a head-up display.
[1436] "Using a voice response system to provide feedback to workers" means using a microphone and voice recognition technology to recognize voice instructions from workers and provide appropriate responses or feedback.
[1437] "Assessing the emotional state of workers through emotion recognition" means using voice analysis and facial expression recognition technology to assess the emotional state of workers in real time and provide appropriate support based on the results.
[1438] This invention relates to a work support system for factories. This system exchanges information in real time between a server, terminals, and users, and provides multifunctional support to improve work efficiency.
[1439] Data collection and digitization
[1440] server:
[1441] Past medical records and work history are digitized and stored in a database. Paper medical records and handwritten work records are scanned and converted into text data using OCR (optical character recognition) technology. The converted data is then stored in a database for efficient management.
[1442] Training generative AI models
[1443] server:
[1444] Past work data and medical records are extracted from the database and divided into a training dataset and a validation dataset for the generative AI model. The training dataset is used to train the AI model, and its performance is evaluated and adjusted. This enables accurate diagnosis prediction and work instruction generation.
[1445] Building a voice support system
[1446] Device:
[1447] It uses a speech recognition API to enable voice input functionality. When users give voice instructions during work, the device converts the speech into text data in real time and provides appropriate feedback. It also uses generative AI models to automatically generate work instructions and questions.
[1448] Incorporating an emotion engine
[1449] Device:
[1450] It has an emotion engine that recognizes emotions from the user's voice and input data in real time, determines the emotional state from the tone of voice and words used, and provides feedback according to the emotion.
[1451] View work instructions in real time
[1452] Device:
[1453] Using smart glasses or head-up displays, work instructions are displayed in real time, allowing workers to receive visual information immediately and work efficiently.
[1454] ★Example:
[1455] For example, when a worker needs to install the next part in a factory, the smart glasses will display work instructions such as "Please take the next part, B, and install it on machine C." If the worker asks by voice, "What should I do next?", the next work instruction will be instantly provided via voice and text.
[1456] Emotion recognition for worker support
[1457] server:
[1458] The system analyzes the emotional data received from the emotion engine and provides appropriate support if a worker is feeling stressed or fatigued, such as suggesting slowing down work speed or taking a break.
[1459] Predicting abnormalities and presenting countermeasures
[1460] server:
[1461] If an abnormality occurs during work, the generative AI model is used to analyze the anomaly and suggest countermeasures, allowing for swift and appropriate measures to be taken.
[1462] ★Example prompt:
[1463] The current task is assembling parts. Please generate the next task instruction.
[1464] In this way, by implementing the present invention, not only can workers receive work instructions in real time and work efficiently, but it also enables support according to emotional states and the prediction of abnormalities, thereby improving overall work efficiency and safety.
[1465] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1466] Step 1:
[1467] Data collection and digitization
[1468] Users scan paper medical records or handwritten work records and input them into the terminal. The terminal then converts the scanned data into text data using OCR (optical character recognition) technology. This text data is sent to the server and stored in a database.
[1469] Input: Scanned paper charts and handwritten work records
[1470] Data processing: Text conversion using OCR technology
[1471] Output: Text data stored in the database
[1472] Step 2:
[1473] Training generative AI models
[1474] The server extracts past medical records and work data from the database and divides them into training and validation datasets for the generative AI model. The server trains the generative AI model using the training dataset and evaluates and adjusts the model's performance.
[1475] Input: Medical records and work data in a database
[1476] Data Computing: Learning and Evaluating Generative AI Models
[1477] Output: A trained generative AI model
[1478] Step 3:
[1479] Building a voice support system
[1480] The device uses a speech recognition API to enable voice input. The user gives voice instructions while working, and the device converts the speech into text data in real time. The server then uses a generative AI model to analyze the text data and generate appropriate feedback and work instructions.
[1481] Input: User's voice commands
[1482] Data processing: Text conversion using voice recognition, analysis using generative AI models
[1483] Output: Text data of feedback and work instructions
[1484] Step 4:
[1485] Incorporating emotion recognition
[1486] The device uses an emotion engine to recognize emotions from the user's voice and input data in real time. The recognized emotion data is sent to a server, which analyzes it and provides appropriate feedback to the user. For example, if the user is feeling stressed, the server may suggest taking a break.
[1487] Input: User voice and input data
[1488] Data calculation: Emotion recognition by emotion engine, analysis by server
[1489] Output: Emotional feedback
[1490] Step 5:
[1491] View work instructions in real time
[1492] The device displays work instructions in real time using smart glasses or a head-up display. The server uses a generative AI model to predict the next task and sends it to the device. The user receives the work instructions through the visual device, allowing them to work efficiently.
[1493] Input: Work instructions from the server
[1494] Data Computation: Task Prediction with Generative AI Models
[1495] Output: Work instructions displayed on a visual device
[1496] Step 6:
[1497] Predicting abnormalities and presenting countermeasures
[1498] The server monitors data generated during work in real time, and if an abnormality occurs, it analyzes it using a generative AI model. Based on the analysis results, it sends appropriate countermeasures to the device. The user can quickly resolve the abnormality by following the countermeasures provided by the device.
[1499] Input: Real-time data as you work
[1500] Data Computation: Anomaly Analysis with Generative AI Models
[1501] Output: Actions sent to the device
[1502] Step 7:
[1503] Emotion recognition for worker support
[1504] The server analyzes the emotion data received from the emotion engine and provides appropriate support if the worker is feeling stressed or fatigued, such as suggesting slowing down work speed or taking a break.
[1505] Input: Emotion data from the emotion engine
[1506] Data Computing: Server-Based Emotion Analysis
[1507] Output: Emotionally appropriate support and feedback
[1508] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1509] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1510] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1511] [Fourth embodiment]
[1512] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1513] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1514] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1515] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1516] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1517] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1518] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1519] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1520] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1521] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1522] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1523] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1524] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1525] This invention is a system that promotes digitalization and cloud computing in the medical industry. Below, we will create a program for this system and explain its processing in natural language. Specific examples will also be provided.
[1526] Data collection and digitization
[1527] User:
[1528] First, users collect the patient's paper medical records or existing electronic medical records, which are then digitized by scanning them, and then upload the digitized data to the server via a terminal.
[1529] Device:
[1530] The terminal receives the scanned paper medical records and converts them into text data using OCR (optical character recognition) technology, which is then sent to the server.
[1531] server:
[1532] The server converts the received text data into an appropriate database format and stores it in the database, allowing past medical records to be digitized and efficiently managed.
[1533] Training generative AI models
[1534] server:
[1535] The server extracts historical patient data and medical records from the database and divides them into training and validation datasets for the generative AI model. The training dataset is used to train the AI model, completing the model learning process. This process also includes adjusting hyperparameters such as the number of epochs and batch size. The model's performance is evaluated and adjustments are made as necessary.
[1536] Building a voice support system
[1537] Device:
[1538] The device uses a speech recognition API to enable voice input. When the user interviews a patient, the device converts the speech into text data in real time. It also generates optimal questions based on a pre-set question list and past response history.
[1539] server:
[1540] The server receives voice data in real time, converts it into text using speech recognition technology, and then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user.
[1541] Data partitioning and sharing
[1542] server:
[1543] The server will appropriately anonymize patient data to enhance privacy and security, and will partition and share this data with relevant medical institutions and researchers as needed, with security protocols in place and appropriate access rights managed.
[1544] Diagnostic prediction and optimal allocation of medical resources
[1545] server:
[1546] When new patient data is input, the server uses the generative AI model to predict diagnoses. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes them as necessary.
[1547] Security and Privacy Measures
[1548] server:
[1549] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology to prevent unauthorized access and data leaks. In addition, security audits are conducted regularly to ensure the safety of the system.
[1550] Specific examples
[1551] 1. The user scans Patient A's paper medical record and uploads it to the server via their device. The server converts it into text data using OCR technology and stores it in the database.
[1552] 2. When interviewing new patient B, the user uses voice input, and the device generates appropriate questions and converts them into text in real time. The server analyzes them using a generative AI model and provides feedback to the user.
[1553] 3. The server makes a diagnosis prediction and diagnoses Patient C as having a high possibility of influenza. The server checks the inventory of necessary medical equipment and medicines and allocates them appropriately.
[1554] 4. Patient data is anonymized and securely shared with other medical institutions for research purposes. The server applies access controls and data encryption to ensure privacy and security.
[1555] As described above, by specifically implementing the present invention, it is possible to realize efficient management and utilization of medical records, support for medical staff, and strengthen data security.
[1556] The processing flow will be explained below.
[1557] Step 1:
[1558] The user collects the patient's paper medical record or existing electronic medical record and scans the paper medical record, which then digitizes the data and loads it into the terminal.
[1559] Step 2:
[1560] The terminal converts scanned paper chart images into text data using OCR (Optical Character Recognition) technology, which is then sent to a server for further processing.
[1561] Step 3:
[1562] The server converts the received text data into an appropriate database format and stores it in the database. This standardizes the data and enables efficient management.
[1563] Step 4:
[1564] The server extracts historical patient data and medical records from the database and splits them into training and validation datasets.
[1565] Step 5:
[1566] The server trains the generative AI model using the training dataset, sets hyperparameters such as the number of epochs and batch size, and executes the model learning process.
[1567] Step 6:
[1568] The server evaluates the performance of the trained model using a validation dataset and adjusts the model as needed.
[1569] Step 7:
[1570] The device uses the speech recognition API to enable the voice input function, preparing the device for the user to use voice input.
[1571] Step 8:
[1572] When interviewing patients, the device automatically generates the most appropriate questions based on a pre-set list of questions and past response history.
[1573] Step 9:
[1574] The device converts the user's voice input into text in real time and immediately sends the text to the server.
[1575] Step 10:
[1576] The server analyzes the text data sent and uses a generative AI model to provide appropriate feedback to the user.
[1577] Step 11:
[1578] The server anonymizes patient data, removing or transforming personally identifiable information to protect the confidentiality of the data.
[1579] Step 12:
[1580] The server will divide the anonymized data and share it with medical institutions and researchers who need it, and will also set up security protocols and manage access rights.
[1581] Step 13:
[1582] When new patient data is entered into the server, the generative AI model is used to make a diagnostic prediction, and the results are notified to the user in real time.
[1583] Step 14:
[1584] The server checks the inventory status of medical resources (medical equipment and medicines) based on patient data and diagnosis predictions, and allocates or replenishes them as necessary.
[1585] Step 15:
[1586] The server manages the access control list (ACL) and sets the access rights for each user. Data is protected using AES-256 encryption technology.
[1587] Step 16:
[1588] We regularly conduct security audits on our servers to ensure the safety of our systems. In the event of a security incident, we respond immediately and take corrective measures.
[1589] Example 1
[1590] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1591] Problems with conventional medical systems include the lack of progress in digitizing medical records and the ineffective use of collected data. Furthermore, there is a lack of systems that support the accuracy of diagnostic predictions and the appropriate allocation of medical resources. There are also many shortcomings in terms of security and privacy protection. These issues are major obstacles to improving the efficiency and quality of medical care, and they require rapid and accurate responses.
[1592] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1593] In this invention, the server includes a means for scanning and digitizing past medical records, converting them into text data using OCR technology, and storing them in a database, a means for extracting and training data sets using a generative AI model, and a means for anonymizing patient data and using technology to enhance security, thereby enabling efficient management and utilization of medical records, improving the accuracy of diagnostic predictions, and optimizing the allocation of medical resources.
[1594] A "database format" is a structure or format for efficiently storing, retrieving, and managing digital data.
[1595] "OCR technology" is an abbreviation for optical character recognition technology, which extracts character information from scanned images and converts it into text data.
[1596] A "generative AI model" refers to an algorithm or model that uses artificial intelligence techniques to generate new information or predictions from a given dataset.
[1597] "Voice recognition technology" is a technology that converts voice input into text data in real time.
[1598] "Data anonymization" refers to the process of removing or modifying personally identifiable information to protect the privacy of the data.
[1599] A "security protocol" is a set of procedures and rules that are used to prevent unauthorized access when transmitting or receiving data.
[1600] "Diagnostic prediction" is the process of using generative AI models based on medical data to predict illness and treatment.
[1601] An "access control list (ACL)" is a list used to manage the rights and permissions of each user accessing a system.
[1602] "Data encryption" is a technology that encrypts important information using a specific algorithm to prevent unauthorized access and data leaks.
[1603] "Hyperparameter tuning" is the process of fine-tuning the parameters of a machine learning model to optimize its learning efficiency and performance.
[1604] The present invention is a system that promotes digitalization and cloud computing in the medical industry. The program processing of this system will be explained below in natural language.
[1605] Data collection and digitization
[1606] User:
[1607] First, the user collects the patient's paper medical records or existing electronic medical records. They use a general-purpose scanner to scan the paper medical records and convert them into digital data. For example, they use a general-purpose high-performance scanner. Next, the user uploads the scanned digital data to a server via a terminal. This is done using the hospital's internal network.
[1608] Device:
[1609] The terminal receives the scanned image data of the paper medical record and converts it into text data using OCR technology. This technology uses commonly used OCR software. The converted text data is sent to the server. When sending, the communication is encrypted using SSL / TLS.
[1610] server:
[1611] The server receives the text data sent from the terminal. The received data is converted into an appropriate database format and stored in a database. For example, a relational database management system (RDBMS) is used as the database.
[1612] Training generative AI models
[1613] server:
[1614] The server extracts past patient data and medical records from the database. The extracted data is divided into a training dataset and a validation dataset for the generative AI model. The server trains the generative AI model using the training dataset. Deep learning libraries such as PyTorch and TensorFlow are used for this training. Hyperparameters such as the number of epochs and batch size are also adjusted. The model's performance is evaluated on the validation dataset, and readjustments are made as necessary.
[1615] Building a voice support system
[1616] Device:
[1617] The device uses a speech recognition API to enable voice input. For example, a general speech recognition API can be used. When the user interviews a patient, the device converts the speech into text data in real time. It also generates optimal questions based on a pre-set question list and past response history.
[1618] server:
[1619] The server receives the voice data in real time, converts it into text using speech recognition technology, and then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user.
[1620] Data partitioning and sharing
[1621] server:
[1622] The server will appropriately anonymize patient data to enhance privacy and security. This will involve techniques to remove or modify personally identifiable information. Furthermore, the anonymized data will be split up as needed and shared with relevant medical institutions and researchers. This sharing will be subject to security protocols and appropriate access rights management.
[1623] Diagnostic prediction and optimal allocation of medical resources
[1624] server:
[1625] When new patient data is input, the server uses the generative AI model to predict diagnoses. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes these resources as needed.
[1626] Security and Privacy Measures
[1627] server:
[1628] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology to prevent unauthorized access and data leaks. In addition, security audits are conducted regularly to ensure the safety of the system.
[1629] Specific examples
[1630] 1. The user scans Patient A's paper medical record and uploads it to the server via their device. The server converts it into text data using OCR technology and stores it in the database.
[1631] 2. When interviewing new patient B, the user uses voice input, and the device generates appropriate questions and converts them into text in real time. The server analyzes them using a generative AI model and provides feedback to the user.
[1632] 3. The server makes a diagnosis prediction and diagnoses Patient C as having a high possibility of influenza. The server checks the inventory of necessary medical equipment and medicines and allocates them appropriately.
[1633] 4. Patient data is anonymized and securely shared with other medical institutions for research purposes. The server applies access controls and data encryption to ensure privacy and security.
[1634] Prompt Sentence Examples
[1635] "How do you digitize and store Patient A's paper medical records?"
[1636] "How do I generate the best questions for a new patient, Patient B?"
[1637] "How can we predict diagnoses and allocate necessary medical resources appropriately?"
[1638] As described above, by specifically implementing the present invention, it is possible to realize efficient management and utilization of medical records, support for medical staff, and strengthen data security.
[1639] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1640] Step 1:
[1641] User: The user obtains digital image data by scanning a patient's paper chart. This input data is in an image file format (e.g., JPEG or PNG). The specific actions of scanning using a scanner include placing the paper chart on the scanner and pressing the Scan button. The output is the scanned digital image data.
[1642] Step 2:
[1643] User: The user uploads the scanned digital image data to the terminal. The uploading process involves sending the digital image to a designated folder using the hospital network. The input is the scanned digital image, and the output is the uploaded digital image in the terminal.
[1644] Step 3:
[1645] Terminal: The terminal converts the received digital image data into text data using OCR technology. Specific operations include launching the OCR software, reading the digital image, and performing character recognition. The input is digital image data, and the output is the converted text data.
[1646] Step 4:
[1647] Terminal: The terminal sends the text data generated by OCR to the server. This transmission process uses SSL / TLS encrypted communication. The input is the converted text data, and the output is the text data sent to the server.
[1648] Step 5:
[1649] Server: The server converts the received text data into a database format and stores it in the database. Specific operations include converting the text data into an SQL query format and performing an insert operation on the database. The input is the text data sent to the server, and the output is the medical record data stored in the database.
[1650] Step 6:
[1651] Server: The server extracts historical patient data and medical records from the database and splits them into training and validation datasets for the generative AI model. Specific operations include executing SQL queries to retrieve the data and splitting the extracted data into training and validation datasets. The input is the medical record data stored in the database, and the output is the split dataset.
[1652] Step 7:
[1653] Server: The server trains the generative AI model using the training dataset. This is done using deep learning libraries such as PyTorch or TensorFlow. It also adjusts hyperparameters such as the number of epochs and batch size. The input is the training dataset, and the output is the trained generative AI model.
[1654] Step 8:
[1655] Server: The server evaluates the model's performance on the validation dataset and adjusts hyperparameters or modifies the model structure as necessary. Specific operations include calculating metrics for model evaluation and retraining. The input is the validation dataset and the existing generative AI model, and the output is an optimized generative AI model.
[1656] Step 9:
[1657] Terminal: The terminal uses a speech recognition API to enable voice input and converts the user's voice during patient interviews into text data in real time. It also generates optimal questions based on a pre-set question list and past response history. The input is voice data, and the output is text data and the generated questions.
[1658] Step 10:
[1659] Server: The server receives voice data in real time and converts it into text data using speech recognition technology. It then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user. The input is voice data, and the output is interpreted text data and feedback.
[1660] Step 11:
[1661] Server: The server anonymizes patient data, removes or modifies any personally identifiable information from the data, and applies security protocols and manages access rights to share the anonymized data with relevant medical institutions and researchers. The input is raw patient data, and the output is anonymized data.
[1662] Step 12:
[1663] Server: When new patient data is input, the server uses the generative AI model to make a diagnosis prediction. The prediction results are notified to the user in real time, and at the same time, the server checks the inventory status of medical resources and allocates or replenishes them as needed. The input is the new patient data, and the output is the diagnosis prediction results and medical resource management information.
[1664] Step 13:
[1665] Server: The server manages the access control list (ACL) and sets the access rights for each user. It also protects data with AES-256 encryption technology and performs regular security audits to prevent unauthorized access and data leakage. The input is user information and data, and the output is controlled access rights and protected data.
[1666] Through these steps, the system achieves efficient digitization of medical records, diagnostic prediction through generative AI models, and data security management.
[1667] (Application example 1)
[1668] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1669] The modern healthcare industry requires the digitization of medical records and efficient data management. However, several challenges exist, including the digitization of paper medical records, the centralized management of existing electronic medical records, the anonymization and security of medical data, and the improvement of diagnostic prediction accuracy. In particular, there is a lack of support systems for quickly and accurately recording patient interactions and making appropriate diagnoses. It is also important to safely anonymize collected data and share it with other medical institutions and researchers. It is necessary to resolve these challenges and realize digitalization and enhanced security in the healthcare industry.
[1670] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1671] In this invention, the server includes means for digitizing past medical records and storing them in a database, means for training medical datasets using a generative AI model, means for supporting patient interviews using voice recognition, means for anonymizing patient data to enhance security, means for making diagnostic predictions and supporting the appropriate allocation of medical resources, means for ensuring data security through access control and data encryption, means for generating text data from scanned images using optical character recognition technology, means for converting voice input into text data in real time using a voice recognition API, and means for checking the availability of medical resources based on diagnostic predictions using the generative AI model. This enables efficient digitization and management of medical records, rapid recording of patient interactions, appropriate diagnostic support, and secure data sharing.
[1672] "Prior medical records" are previously collected medical data, including a patient's medical history and treatment history.
[1673] "Digitalization" is the process of converting information in analog form, such as paper, into digital form.
[1674] A "database" is a system for storing and managing data in an organized structure.
[1675] A "generative AI model" is a machine learning algorithm that learns from large amounts of data and is optimized to perform a specific task.
[1676] A "medical dataset" is a collection of various medical-related data used for learning and analysis.
[1677] "Speech recognition" is the technology that takes voice input and converts it into a digital format such as text.
[1678] A "patient interview" is a process in which a medical professional directly obtains information from a patient, such as symptoms and medical history.
[1679] "Anonymization" is the process of processing data so that individuals cannot be identified.
[1680] "Security" refers to a group of technologies aimed at protecting data and preventing unauthorized access.
[1681] "Diagnostic prediction" is the process of predicting future disease states and appropriate treatments based on collected medical data.
[1682] "Medical resources" is a general term for all equipment, medicines, personnel, etc. used in medical institutions.
[1683] "Proper allocation" means properly allocating the necessary resources to the necessary locations.
[1684] "Access control" is a mechanism that allows only authorized users to access specific data.
[1685] "Data encryption" is the technology that encrypts data and converts it into a form that is unintelligible to unauthorized users.
[1686] "Optical character recognition technology" is a technology that analyzes character information in an image and converts it into text data.
[1687] "Real-time" means that data input, processing, and output are carried out instantly.
[1688] A "speech recognition API" is an application programming interface that provides the functionality to convert voice input into text.
[1689] "Inventory status" is information that indicates the current level of medical resources available for use.
[1690] This invention is a system that promotes digitalization and cloud computing in the medical industry, and provides a wide range of functions, including digitization of past medical records, diagnostic prediction using generative AI models, patient interview support using voice recognition, and enhanced data security using anonymization technology.
[1691] Data collection and digitization
[1692] User:
[1693] Users first scan the patient's paper medical records using their smartphone camera and upload them to the server via the device, allowing them to be digitized and managed efficiently.
[1694] Application of OCR technology
[1695] Device:
[1696] The terminal uses optical character recognition (OCR) technology to generate text data from scanned images of paper medical records, which is then sent to a server and stored in a database.
[1697] Building a voice support system
[1698] Device:
[1699] The device uses a speech recognition API to convert voice input into text data in real time. The user records conversations with patients and converts the data into text in real time. The device also generates optimal questions based on a pre-set question list and past response history.
[1700] Data storage and anonymization
[1701] server:
[1702] The server converts the received text data into an appropriate database format and stores it in the database. The server anonymizes the data and uses data encryption technology (e.g., AES-256) to enhance privacy and security. This data is split as needed and shared with relevant medical institutions and researchers.
[1703] Training generative AI models
[1704] server:
[1705] The server extracts historical patient data and medical records from the database and divides them into training and validation datasets for the generative AI model. The training dataset is used to train the AI model, completing the model learning process. The model's performance is evaluated and adjusted as necessary.
[1706] Diagnostic prediction and optimal allocation of medical resources
[1707] server:
[1708] When new patient data is entered, the server uses the generative AI model to predict a diagnosis. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (medical equipment, drugs, etc.) and allocates or replenishes them as needed.
[1709] Examples of concrete examples and prompts
[1710] 1. A user scans a patient's paper chart and uploads it to the server via a terminal. An example of the prompt is as follows:
[1711] "Take images of medical records, digitize them, and upload them to the cloud."
[1712] 2. The user records the conversation with the patient and uses a speech recognition API to transcribe it in real time, with prompts such as:
[1713] "Record conversations with patients, convert them into text, and save them."
[1714] 3. The server performs diagnosis prediction, checks the availability of medical resources, and allocates them appropriately. An example of the prompt is as follows:
[1715] "Using generative AI models to predict diagnoses and allocate necessary medical resources"
[1716] This enables efficient digitization and management of medical records, enables prompt recording of patient interactions, and ensures proper diagnostic support and secure data sharing, significantly improving operational efficiency and security in the healthcare industry.
[1717] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1718] Step 1:
[1719] The user scans the patient's paper medical record using the smartphone camera and uploads it to the server via the terminal. The input is an image of the paper medical record, and the output is digitized text data. The user launches the scanning app on the terminal and takes a picture of the paper medical record.
[1720] Step 2:
[1721] The device generates text data from the scanned image using optical character recognition (OCR). The input is the scanned image from step 1, and the output is the text data generated by OCR. The device sends the scanned image to the OCR engine and obtains the text data.
[1722] Step 3:
[1723] The device sends the generated text data to the server. The input is the text data, and the output is the data sent to the server. The device then uploads the data to the cloud server via the network.
[1724] Step 4:
[1725] The server converts the received text data into an appropriate database format and stores it in the database. The input is the text data received by the server, and the output is the data stored in the database. The server formats the text data into the database format and stores it.
[1726] Step 5:
[1727] The server uses a speech recognition API to convert the voice data recorded by the user into text data in real time. The input is voice data and the output is text data. The server sends the voice data to the API and saves the resulting text data.
[1728] Step 6:
[1729] The server anonymizes the received data and uses data encryption techniques (e.g., AES-256) to enhance privacy and security. The input is text data stored in a database, and the output is anonymized and encrypted data. The server anonymizes the data and applies encryption techniques.
[1730] Step 7:
[1731] The server uses a generative AI model to split the past patient data and medical records extracted from the database into a training dataset and a validation dataset. The input is the past patient data and medical records, and the output is the training dataset and the validation dataset. The server splits the data into training and validation datasets.
[1732] Step 8:
[1733] The server trains the generative AI model using the training dataset, completing the model learning process. The input is the training dataset, and the output is the trained generative AI model. The server uses the dataset to train the model.
[1734] Step 9:
[1735] When new patient data is input, the server uses the generative AI model to make a diagnosis prediction. The input is the new patient data, and the output is the diagnosis prediction result. The server inputs the new patient data into the model and obtains the diagnosis prediction result.
[1736] Step 10:
[1737] The server checks the inventory status of medical resources (medical equipment, medicines, etc.) based on the diagnosis prediction, and allocates or replenishes them as needed. The input is the diagnosis prediction result, and the output is confirmation of the inventory status and the appropriate allocation of resources. The server works with the inventory system based on the prediction results to optimally allocate resources.
[1738] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1739] This invention combines an emotion engine with a system that promotes digitalization and cloud computing in the medical industry. Below, we will generate a program for this system and explain its processing in natural language. Specific examples will also be provided.
[1740] Data collection and digitization
[1741] User:
[1742] First, the user collects the patient's paper medical records or existing electronic medical records. The paper medical records are digitized through scanning and imported into the terminal. The digitized data is then uploaded to the server via the terminal.
[1743] Device:
[1744] The terminal receives the scanned paper medical records and converts them into text data using OCR (optical character recognition) technology, which is then sent to the server.
[1745] server:
[1746] The server converts the received text data into an appropriate database format and stores it in the database, allowing past medical records to be digitized and efficiently managed.
[1747] Training generative AI models
[1748] server:
[1749] The server extracts historical patient data and medical records from the database and splits them into training and validation datasets for the generative AI model. The training dataset is used to train the AI model, completing the model learning process. The model's performance is evaluated and adjustments are made as needed.
[1750] Building a voice support system
[1751] Device:
[1752] The device uses a speech recognition API to enable voice input. When the user interviews a patient, the device converts the speech into text data in real time. It also generates optimal questions based on a pre-set question list and past response history.
[1753] server:
[1754] The server receives voice data in real time, converts it into text using speech recognition technology, and then uses a generative AI model to analyze the context of the conversation and provide appropriate feedback to the user.
[1755] Incorporating an emotion engine
[1756] Device:
[1757] The device has an emotion engine that recognizes emotions from the user's voice and input data in real time, for example, determining the user's emotional state from the tone of their voice and the words they choose.
[1758] server:
[1759] The server analyzes the emotion data sent from the emotion engine and generates feedback according to the user's emotions. If the user is feeling stressed or anxious, the system will suggest appropriate responses and support.
[1760] Data partitioning and sharing
[1761] server:
[1762] The server will appropriately anonymize patient data to enhance privacy and security, and will partition and share this data with relevant medical institutions and researchers as needed, with security protocols in place and appropriate access rights managed.
[1763] Diagnostic prediction and optimal allocation of medical resources
[1764] server:
[1765] When new patient data is input, the server uses the generative AI model to predict diagnoses. The prediction results are notified to the user in real time. At the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes them as necessary.
[1766] Security and Privacy Measures
[1767] server:
[1768] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology to prevent unauthorized access and data leaks. In addition, security audits are conducted regularly to ensure the safety of the system.
[1769] Specific examples
[1770] 1. The user scans Patient D's paper medical record and uploads it to the server via a terminal. The server uses OCR technology to convert the medical record into text data and saves it in the database.
[1771] 2. When interviewing Patient E, the user uses voice input, and the device generates appropriate questions and converts them into text in real time. The server analyzes the results using a generative AI model and provides feedback to the user. If the emotion engine detects stress in the user, it provides appropriate support.
[1772] 3. The server makes a diagnosis prediction and checks the inventory of necessary medical equipment and medicines for Patient F, who is diagnosed with possible pneumonia, and allocates them appropriately.
[1773] 4. Patient data is anonymized and securely shared with other medical institutions for research purposes. The server applies access controls and data encryption to ensure privacy and security.
[1774] As described above, by specifically implementing the present invention, it is possible to efficiently manage and utilize medical records, support medical staff, strengthen data security, and even provide appropriate feedback according to the user's emotions.
[1775] The processing flow will be explained below.
[1776] Step 1:
[1777] The user collects the patient's paper or electronic records and scans the paper records, which then digitizes the data and loads it into the terminal.
[1778] Step 2:
[1779] The terminal converts the scanned paper chart image into text data using OCR (optical character recognition) technology, which is then sent to the server.
[1780] Step 3:
[1781] The server converts the received text data into an appropriate database format and stores it in the database, thereby digitizing the medical records.
[1782] Step 4:
[1783] The server extracts historical patient data and medical records from the database and splits them into training and validation datasets.
[1784] Step 5:
[1785] The server trains the generative AI model using the training dataset, setting parameters such as the number of epochs and batch size, and running the model's learning process.
[1786] Step 6:
[1787] The server evaluates the performance of the trained model using a validation dataset and adjusts the model as needed.
[1788] Step 7:
[1789] The device uses the speech recognition API to enable voice input, preparing to convert what the user says into text in real time.
[1790] Step 8:
[1791] When interviewing patients, the device automatically generates the most appropriate questions based on a set list of questions and past response history.
[1792] Step 9:
[1793] The device converts the user's voice input into text in real time and immediately sends the converted text to the server.
[1794] Step 10:
[1795] The server analyzes the text data sent and uses a generative AI model to provide appropriate feedback to the user.
[1796] Step 11:
[1797] The device uses an emotion engine to recognize emotions from the user's voice and text data in real time, and transmits the emotional state to the server.
[1798] Step 12:
[1799] The server analyzes the emotion data sent from the emotion engine and adjusts the content of the feedback based on the user's emotions.
[1800] Step 13:
[1801] The server anonymizes patient data to enhance privacy and security, and the data is segmented and shared with relevant medical institutions and researchers as needed.
[1802] Step 14:
[1803] The server uses the generative AI model to make a diagnostic prediction based on the new patient data entered, and notifies the user of the diagnostic prediction results in real time.
[1804] Step 15:
[1805] The server checks the inventory status of medical resources (e.g., medical equipment and medicines) based on patient data and diagnosis predictions, and allocates or replenishes them as needed.
[1806] Step 16:
[1807] The server manages the access control list (ACL) and sets the access rights for each user. Data is protected using AES-256 encryption technology.
[1808] Step 17:
[1809] We regularly conduct security audits of our servers to ensure the safety of our systems. If a security incident occurs, we will respond immediately and take corrective measures.
[1810] Example 2
[1811] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1812] In modern healthcare, vast amounts of patient data are continually being generated, and managing, analyzing, and sharing this data presents major challenges. In particular, where many past medical records, including paper charts, remain, it is important to digitize and appropriately utilize them. There is also a need to understand the emotional state of patients through their voices and provide appropriate feedback and support. Furthermore, effective methods are needed to improve diagnostic accuracy and efficiently allocate medical resources.
[1813] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1814] In this invention, the server includes means for digitizing past medical records and storing them in a database, means for training medical datasets using a generative AI model, means for supporting patient interviews using voice recognition, means for anonymizing patient data to enhance security, means for making diagnostic predictions and supporting the optimal allocation of medical resources, means for ensuring data security through access control and data encryption, means for recognizing user emotions in real time and providing corresponding feedback, and means for sharing medical data with other medical institutions and researchers using security protocols. This enables efficient management of massive amounts of patient data, rapid feedback based on patient emotions, improved diagnostic accuracy, and optimal allocation of medical resources.
[1815] "Method of digitizing past medical records and storing them in a database" refers to the process of scanning paper medical records or existing electronic medical records, converting them into text data using OCR technology, and storing them in a database.
[1816] "Method of using a generative AI model to train a medical dataset" refers to the process of collecting medical data, dividing it into a training dataset and a validation dataset to train an AI model, and evaluating and adjusting the model's performance.
[1817] "Means for supporting patient hearings using voice recognition" refers to a process that uses a voice recognition API to convert voice input from patients into text data and support the hearing process.
[1818] "Methods for de-identifying and enhancing security of patient data" refers to technologies and processes used to de-identify personal patient information and enhance privacy and security.
[1819] "Means for making diagnostic predictions and supporting the appropriate allocation of medical resources" refers to the process of using generative AI models to make diagnostic predictions from new patient data and then allocating or replenishing necessary medical resources based on the results.
[1820] "Data security through access control and data encryption" refers to the process of using access control lists to set each user's access rights and protecting the data with AES-256 encryption technology.
[1821] "Means for recognizing a user's emotions in real time and providing corresponding feedback" refers to the technology and process for using an emotion engine to recognize emotions from a user's voice or input data and providing corresponding feedback in real time.
[1822] "Means of sharing medical data with other medical institutions and researchers using security protocols" refers to the process of appropriately anonymizing medical data and applying security protocols to safely share it with other medical institutions and researchers.
[1823] This invention is a system that promotes digitalization and cloud computing in the medical field and incorporates an emotion engine to provide efficient support to patients and medical staff. Specific embodiments of the present invention are described below.
[1824] Data collection and digitization
[1825] The user collects the patient's paper medical records and digitizes them using a scanner. The scanner captures the paper medical records as image data, and the terminal then converts them into text data using OCR (optical character recognition) technology. This text data is uploaded from the terminal to the server. The server converts the text data into an appropriate database format and stores it in the database.
[1826] Example: A user scans a medical record, and the device converts the image data into text using OCR software (e.g., ABBYY FineReader). The converted data is then uploaded to a cloud server and stored in a database.
[1827] Training generative AI models
[1828] The server extracts historical patient data and medical records from the database and splits them into a training dataset and a validation dataset for the generative AI model. The training dataset is used to train the AI model, verify its performance, and make adjustments as needed.
[1829] Example: The server uses frameworks such as Python and TensorFlow to input past diagnostic data into an AI model for learning. For example, training using random forests or neural networks is conceivable.
[1830] Building a voice support system
[1831] The device uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to enable voice input. When the user interviews a patient, the device converts the speech into text data in real time and sends it to the server. The server uses a generative AI model to analyze the context of the conversation and provide optimal feedback.
[1832] Example: When a user speaks to the device, the device converts the voice data into text data and sends it to the server. The server analyzes it and presents the next question.
[1833] Incorporating an emotion engine
[1834] The device uses an emotion engine to recognize emotions from the user's voice and input data. For example, it determines the user's emotional state in real time from the user's tone of voice and selected words. The server analyzes this data and generates feedback according to the user's emotions.
[1835] Example: If a user asks a question in an anxious tone, the server generates reassuring feedback and provides it to the user through the terminal.
[1836] Data partitioning and sharing
[1837] The server will appropriately anonymize patient data and enhance security. This anonymized data will be shared with relevant medical institutions and researchers as needed. Security protocols will be applied to sharing, and access rights will be strictly controlled.
[1838] Example: A server processes data using an anonymization algorithm (e.g., K-anonymization or differential privacy) and shares it over a secure communication protocol (e.g., HTTPs or TLS).
[1839] Diagnostic prediction and optimal allocation of medical resources
[1840] When new patient data is input, the server uses the generative AI model to predict a diagnosis. The prediction results are notified to the user in real time, and at the same time, the server checks the inventory status of medical resources (e.g., medical equipment and drugs) and allocates or replenishes them as needed.
[1841] Example: The server inputs the symptom data of a new patient into an AI model, and if it predicts pneumonia, it notifies the user of the result and allocates the necessary medical resources.
[1842] Security and Privacy Measures
[1843] The server manages access control lists (ACLs) and sets access privileges for each user. Data is protected by AES-256 encryption technology, and regular security audits ensure the system is secure.
[1844] For example: The server manages access privileges for each user based on ACLs, and stored data is encrypted with AES-256. Vulnerability checks are performed regularly using security audit tools (e.g., Nessus or OpenVAS).
[1845] Prompt Sentence Examples
[1846] Specifically, the following prompt sentences could be input into the generative AI model:
[1847] "Enter new patient data."
[1848] "Generate the next question to ask."
[1849] "Do sentiment analysis and provide appropriate feedback."
[1850] As described above, the system of the present invention integrates a wide range of functions to realize efficient data management, patient support, diagnostic prediction, and enhanced security in medical settings.
[1851] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1852] Step 1: Collect and digitize data
[1853] Input: Patient's paper chart
[1854] Output: Digitized text data
[1855] The user collects the patient's paper chart and places it into the scanner.
[1856] How it works: When the user presses the scan button, the scanner reads the paper medical record as image data.
[1857] The device receives the image data from the scanner and converts it into text data using OCR technology. For example, ABBYY FineReader installed on the device is used.
[1858] What happens: The device launches the OCR software and converts the image file into a text file.
[1859] The terminal uploads the converted text data to the server.
[1860] Operation: The device sends text data to the server's API endpoint.
[1861] Step 2: Saving to the database
[1862] Input: Converted text data
[1863] Output: Medical records stored in a database
[1864] The server converts the received text data into the appropriate database format.
[1865] How it works: The server parses the data format and generates SQL queries to insert into the database.
[1866] The server stores the text data in a database.
[1867] Step 3: Prepare training data for the generative AI model
[1868] Input: Medical records stored in a database
[1869] Output: Training and validation datasets
[1870] The server extracts medical record data from the database and splits it into a training dataset and a validation dataset.
[1871] How it works: The server executes SQL queries to extract data and runs an algorithm to randomly split the dataset.
[1872] Step 4: Training the generative AI model
[1873] Input: Teacher dataset
[1874] Output: A trained generative AI model
[1875] The server trains the AI model using the training dataset.
[1876] How it works: The server runs a script to train an AI model using Python or TensorFlow. It inputs the training dataset into the model and trains it.
[1877] Step 5: Building a voice support system
[1878] Input: User voice input
[1879] Output: Real-time audio data converted to text
[1880] The device uses the speech recognition API to enable the voice input function.
[1881] What it does: The device creates an instance of a speech recognition API (e.g., Google Cloud Speech-to-Text) and displays a UI to initiate voice input.
[1882] The user interviews the patient and inputs the information by voice.
[1883] The device converts the voice into text data in real time and sends it to the server.
[1884] Step 6: Analysis of audio data and feedback
[1885] Input: Real-time audio data converted to text
[1886] Output: Analysis results and feedback
[1887] The server analyzes the context of the conversation using a generative AI model.
[1888] How it works: The server inputs text data into the generative AI model and obtains the analysis results.
[1889] The server generates optimal feedback and provides it to the user through the terminal.
[1890] Step 7: Recognize emotions and provide feedback
[1891] Input: User voice and input data
[1892] Output: Emotional feedback
[1893] The terminal uses an emotion engine to recognize the user's emotion.
[1894] How it works: The emotion engine analyzes the user's tone of voice and word choice to determine their emotional state.
[1895] The server analyzes the emotional data and generates appropriate feedback for the user.
[1896] Operation: The server generates a feedback message based on the emotion data and provides it to the user via the terminal.
[1897] Step 8: Anonymize and share data
[1898] Input: Patient Data
[1899] Output: Anonymized data
[1900] The server appropriately anonymizes the patient data.
[1901] How it works: The server processes the data using an anonymization algorithm (e.g., K-anonymization or differential privacy).
[1902] The server will share the anonymized data with other medical institutions and researchers as needed.
[1903] How it works: Data is transmitted securely using security protocols (e.g., HTTPs or TLS).
[1904] Step 9: Diagnostic prediction and resource allocation
[1905] Input: New patient data
[1906] Output: Diagnosis results and medical resource allocation information
[1907] The server uses the generated AI model to make diagnostic predictions for new patient data.
[1908] How it works: New patient data is fed into a generative AI model to obtain a diagnosis.
[1909] The server notifies the user of the diagnosis results and checks the inventory status of medical resources (e.g., medical equipment and medicines) and allocates them appropriately.
[1910] Operation: The server queries the inventory management system and allocates and replenishes needed medical resources.
[1911] Step 10: Security and Privacy
[1912] Input: User ID and data
[1913] Output: Access permission settings and encrypted data
[1914] The server manages the access control list (ACL) and sets the access rights for each user.
[1915] How it works: The server updates the ACL database to manage user identities and access privileges.
[1916] The server protects your data with AES-256 encryption technology.
[1917] How it works: Before the server stores the data, it encrypts it using the AES-256 algorithm.
[1918] The server performs regular security audits to ensure the system is secure.
[1919] What happens: The server runs a vulnerability check using a security audit tool (e.g., Nessus or OpenVAS).
[1920] (Application example 2)
[1921] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1922] While conventional digital and cloud-based systems can efficiently manage and utilize medical records, they are inadequate for real-time support, such as displaying work instructions in real time or assessing workers' emotional states. They also have difficulty providing feedback and appropriate work support through voice recognition. Furthermore, they have faced problems with the psychological burden on workers and reduced work efficiency due to errors.
[1923] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1924] In this invention, the server includes means for digitizing past medical records and storing them in a database, means for training medical datasets using a generative AI model, means for supporting patient interviews using voice recognition, means for anonymizing patient data to enhance security, means for making diagnostic predictions and supporting the appropriate allocation of medical resources, means for ensuring data security through access control and data encryption, means for having a visual device for displaying work instructions in real time, means for providing feedback to workers using a voice response system, and means for assessing the emotional state of workers through emotion recognition. This makes it possible to evaluate the emotional state of workers while displaying work instructions in real time and providing feedback through voice recognition.
[1925] "Digitizing past medical records and storing them in a database" means converting paper medical records and existing electronic medical records into a digital format using scanning or optical character recognition technology (OCR) and storing them in a database.
[1926] "Learning medical datasets using a generative AI model" means training a generative AI model using datasets based on past patient data and medical records, enabling it to analyze medical data and make diagnostic predictions.
[1927] "Supporting patient interviews using voice recognition" means using voice recognition technology to convert interviews with patients into text in real time, and providing appropriate support to medical staff.
[1928] "Anonymizing patient data to enhance security" means protecting the privacy and enhancing the security of patient data by removing or transforming personally identifiable information.
[1929] "Making diagnostic predictions and supporting the appropriate allocation of medical resources" means using a generative AI model to make diagnostic predictions for patients and then appropriately allocating medical resources such as medical equipment and medicines based on the results.
[1930] "Ensuring data security through access control and data encryption" means setting access rights for each user and protecting data using encryption technology such as AES-256.
[1931] "Having a visual device for displaying work instructions in real time" means displaying work instructions in real time using a device such as smart glasses or a head-up display.
[1932] "Using a voice response system to provide feedback to workers" means using a microphone and voice recognition technology to recognize voice instructions from workers and provide appropriate responses or feedback.
[1933] "Assessing the emotional state of workers through emotion recognition" means using voice analysis and facial expression recognition technology to assess the emotional state of workers in real time and provide appropriate support based on the results.
[1934] This invention relates to a work support system for factories. This system exchanges information in real time between a server, terminals, and users, and provides multifunctional support to improve work efficiency.
[1935] Data collection and digitization
[1936] server:
[1937] Past medical records and work history are digitized and stored in a database. Paper medical records and handwritten work records are scanned and converted into text data using OCR (optical character recognition) technology. The converted data is then stored in a database for efficient management.
[1938] Training generative AI models
[1939] server:
[1940] Past work data and medical records are extracted from the database and divided into a training dataset and a validation dataset for the generative AI model. The training dataset is used to train the AI model, and its performance is evaluated and adjusted. This enables accurate diagnosis prediction and work instruction generation.
[1941] Building a voice support system
[1942] Device:
[1943] It uses a speech recognition API to enable voice input functionality. When users give voice instructions during work, the device converts the speech into text data in real time and provides appropriate feedback. It also uses generative AI models to automatically generate work instructions and questions.
[1944] Incorporating an emotion engine
[1945] Device:
[1946] It has an emotion engine that recognizes emotions from the user's voice and input data in real time, determines the emotional state from the tone of voice and words used, and provides feedback according to the emotion.
[1947] View work instructions in real time
[1948] Device:
[1949] Using smart glasses or head-up displays, work instructions are displayed in real time, allowing workers to receive visual information immediately and work efficiently.
[1950] ★Example:
[1951] For example, when a worker needs to install the next part in a factory, the smart glasses will display work instructions such as "Please take the next part, B, and install it on machine C." If the worker asks by voice, "What should I do next?", the next work instruction will be instantly provided via voice and text.
[1952] Emotion recognition for worker support
[1953] server:
[1954] The system analyzes the emotional data received from the emotion engine and provides appropriate support if a worker is feeling stressed or fatigued, such as suggesting slowing down work speed or taking a break.
[1955] Predicting abnormalities and presenting countermeasures
[1956] server:
[1957] If an abnormality occurs during work, the generative AI model is used to analyze the anomaly and suggest countermeasures, allowing for swift and appropriate measures to be taken.
[1958] ★Example prompt:
[1959] The current task is assembling parts. Please generate the next task instruction.
[1960] In this way, by implementing the present invention, not only can workers receive work instructions in real time and work efficiently, but it also enables support according to emotional states and the prediction of abnormalities, thereby improving overall work efficiency and safety.
[1961] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1962] Step 1:
[1963] Data collection and digitization
[1964] Users scan paper medical records or handwritten work records and input them into the terminal. The terminal then converts the scanned data into text data using OCR (optical character recognition) technology. This text data is sent to the server and stored in a database.
[1965] Input: Scanned paper charts and handwritten work records
[1966] Data processing: Text conversion using OCR technology
[1967] Output: Text data stored in the database
[1968] Step 2:
[1969] Training generative AI models
[1970] The server extracts past medical records and work data from the database and divides them into training and validation datasets for the generative AI model. The server trains the generative AI model using the training dataset and evaluates and adjusts the model's performance.
[1971] Input: Medical records and work data in a database
[1972] Data Computing: Learning and Evaluating Generative AI Models
[1973] Output: A trained generative AI model
[1974] Step 3:
[1975] Building a voice support system
[1976] The device uses a speech recognition API to enable voice input. The user gives voice instructions while working, and the device converts the speech into text data in real time. The server then uses a generative AI model to analyze the text data and generate appropriate feedback and work instructions.
[1977] Input: User's voice commands
[1978] Data processing: Text conversion using voice recognition, analysis using generative AI models
[1979] Output: Text data of feedback and work instructions
[1980] Step 4:
[1981] Incorporating emotion recognition
[1982] The device uses an emotion engine to recognize emotions from the user's voice and input data in real time. The recognized emotion data is sent to a server, which analyzes it and provides appropriate feedback to the user. For example, if the user is feeling stressed, the server may suggest taking a break.
[1983] Input: User voice and input data
[1984] Data calculation: Emotion recognition by emotion engine, analysis by server
[1985] Output: Emotional feedback
[1986] Step 5:
[1987] View work instructions in real time
[1988] The device displays work instructions in real time using smart glasses or a head-up display. The server uses a generative AI model to predict the next task and sends it to the device. The user receives the work instructions through the visual device, allowing them to work efficiently.
[1989] Input: Work instructions from the server
[1990] Data Computation: Task Prediction with Generative AI Models
[1991] Output: Work instructions displayed on a visual device
[1992] Step 6:
[1993] Predicting abnormalities and presenting countermeasures
[1994] The server monitors data generated during work in real time, and if an abnormality occurs, it analyzes it using a generative AI model. Based on the analysis results, it sends appropriate countermeasures to the device. The user can quickly resolve the abnormality by following the countermeasures provided by the device.
[1995] Input: Real-time data as you work
[1996] Data Computation: Anomaly Analysis with Generative AI Models
[1997] Output: Actions sent to the device
[1998] Step 7:
[1999] Emotion recognition for worker support
[2000] The server analyzes the emotion data received from the emotion engine and provides appropriate support if the worker is feeling stressed or fatigued, such as suggesting slowing down work speed or taking a break.
[2001] Input: Emotion data from the emotion engine
[2002] Data Computing: Server-Based Emotion Analysis
[2003] Output: Emotionally appropriate support and feedback
[2004] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2005] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the da...
Claims
1. A means of digitizing past medical records and storing them in a database; a means for training a medical dataset using a generative AI model; A means for supporting patient hearing using speech recognition; measures to anonymize patient data to enhance security; A means of making diagnostic predictions and supporting the appropriate allocation of medical resources; a means of ensuring data security through access control and data encryption; A system including:
2. A means for splitting and managing training datasets and validation datasets for generative AI models; A means to validate and adjust the performance of generative AI models; and The system of claim 1 further comprising:
3. means for processing patient data in real time and providing feedback to the user; a means for making diagnostic predictions using machine learning algorithms; means for automatically generating questions based on voice input; The system of claim 1 further comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A