system
A system using cloud storage and AI for medical image analysis addresses the inefficiencies in determining the cause of death by providing rapid and accurate diagnostic reports, tailored to the user's emotional state, reducing the need for autopsies and improving diagnostic accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-16
- Publication Date
- 2026-04-28
AI Technical Summary
In modern medical fields, determining the cause of death is challenging due to low autopsy rates, cultural factors, and inefficiencies in conventional methods, which require significant time and labor, and often lack detailed information when autopsies cannot be performed.
A system that stores medical image data from devices like CT scanners and MRIs in cloud storage, preprocesses the data, and uses an artificial intelligence model for diagnostic analysis to generate a report without the need for autopsy, improving accuracy through data comparison and storage for future research.
Enables rapid and accurate determination of the cause of death, reducing the burden of autopsy and enhancing diagnostic accuracy through data analysis and emotional intelligence adjustments.
Smart Images

Figure 2026071004000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern medical fields, determination of the cause of death is mainly carried out by autopsy. However, due to cultural factors and a shortage of specialists, the autopsy rate is low, and there is a problem that it is difficult to identify the true cause of death. In addition, the conventional method for determining the cause of death requires a great deal of time and labor, and there are also problems with efficiency. As a further problem, when an autopsy cannot be performed, it is difficult to obtain detailed information regarding the cause of death.
Means for Solving the Problems
[0005] This invention provides a system that stores medical image data acquired from an image acquisition device in cloud storage, standardizes the stored data through preprocessing, and then performs diagnostic analysis using an artificial intelligence model. Furthermore, by combining this with a means of generating a diagnostic report based on the analysis results and providing it to the user terminal, the cause of death can be efficiently identified without autopsy. By improving the accuracy of the analysis through comparison with past medical image data, and by storing the diagnostic analysis results in the cloud, enabling future data analysis and research use, this system achieves efficient and accurate determination of the cause of death.
[0006] "Image acquisition equipment" refers to devices such as CT scanners and MRI machines used to acquire medical images.
[0007] "Medical image data" refers to digital information that visualizes the internal structure of the human body, obtained using imaging devices such as CT and MRI.
[0008] "Cloud storage" is a system for storing data on remote servers accessible via the internet.
[0009] "Preprocessing" refers to processes such as noise reduction and resolution standardization performed to make image data easier to analyze.
[0010] "Standardization" is a process or method for unifying data into a certain format to facilitate analysis and comparison.
[0011] An "artificial intelligence model" is a collection of algorithms designed to perform pattern recognition and data analysis using machine learning and deep learning.
[0012] "Diagnostic analysis" is the process of processing information to identify diseases or abnormalities based on acquired medical image data.
[0013] A "diagnostic report" is a document created based on analysis results, providing medical professionals with information about the cause of death and the patient's condition.
[0014] A "user terminal" refers to a computer or mobile device used by medical professionals to review diagnostic information.
[0015] "Comparison" is the act of comparing data obtained at different times or under different conditions and analyzing their differences and similarities.
[0016] "Analysis accuracy" is an indicator that shows the accuracy of the analysis results, representing the low rate of misdiagnosis and the accurate identification of the disease state.
[0017] "Data analysis" is the process of processing collected data using statistical methods and machine learning to derive information that meets a specific purpose.
[0018] "Research use" refers to activities that aim to gain new insights or improve technology based on stored data. [Brief explanation of the drawing]
[0019] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8]It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying Out the Invention
[0020] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be described.
[0022] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0023] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0024] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0025] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0027] [First Embodiment]
[0028] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0029] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0032] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0035] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0039] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0040] This invention provides a system for efficiently determining the cause of death using medical image data from an image acquisition device. This system includes a mechanism for storing image data in cloud storage, performing preprocessing, and analyzing the data using an artificial intelligence model.
[0041] The server has the function of receiving medical image data transmitted from hospital CT and MRI scanners and saving it to cloud storage. The saved data is preprocessed, including noise reduction and resolution standardization, and is ready for analysis.
[0042] Next, the server inputs the pre-processed data into an artificial intelligence model. This model has been trained on numerous training datasets and is designed to identify congenital abnormalities and pathological findings. For example, it can detect abnormal artifacts around the heart and suggest the possibility of a myocardial infarction.
[0043] After the analysis is performed, the server generates a diagnostic report and sends the results to a terminal for the user to use. This report contains detailed cause-of-death information obtained without an autopsy and serves as valuable material for medical professionals to review.
[0044] Furthermore, users can improve diagnostic accuracy by comparing current data with past image data of the same patient. This enables a multifaceted analysis that includes the progression of chronic diseases and the effectiveness of medications.
[0045] Furthermore, the diagnostic analysis results will be stored in the cloud and used for future data analysis and other research purposes. This accumulated database will serve as a foundation to support improvements in diagnostic technology and information sharing across regions.
[0046] The system according to the present invention enables rapid and accurate determination of the cause of death while reducing the burden of autopsy.
[0047] The following describes the processing flow.
[0048] Step 1:
[0049] The server receives medical image data from the hospital's CT and MRI machines. This data is typically provided in DICOM format and organized by patient ID and imaging date.
[0050] Step 2:
[0051] The server stores the received medical image data in cloud storage. This allows for efficient management of large amounts of data and easy access to it later.
[0052] Step 3:
[0053] The server preprocesses the stored image data. By removing noise and maintaining a consistent resolution, it prepares the data for the AI model to analyze efficiently.
[0054] Step 4:
[0055] The server inputs pre-processed image data into an artificial intelligence model. This model is trained using a vast amount of historical data and is tuned to detect abnormal findings and disease features.
[0056] Step 5:
[0057] The server generates a diagnostic report based on the analysis results from the AI model. This report includes identified lesions and possible causes of death, providing information useful for medical decision-making.
[0058] Step 6:
[0059] The server sends the generated diagnostic report to the terminal. In this case, the terminal is a PC or mobile device used by a medical professional.
[0060] Step 7:
[0061] Users can view diagnostic reports on their devices and share information with other medical professionals as needed. They can also compare past medical image data with current data to further improve the accuracy of their diagnoses.
[0062] Step 8:
[0063] The server stores the analysis results in the cloud. This makes the data available for future data analysis and research, creating a database that contributes to the advancement of medicine.
[0064] (Example 1)
[0065] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0066] In modern medical practice, there is a demand for rapid and accurate pathological diagnoses, but conventional systems lack the means to improve analytical accuracy and safely deliver diagnostic results. Furthermore, mechanisms for improving diagnostic accuracy through effective comparison with past medical data, and efficient methods for data storage and utilization, are insufficient. Therefore, challenges remain in formulating treatment plans for patients and managing their long-term health.
[0067] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0068] In this invention, the server includes means for storing medical image data acquired from an image acquisition device in an information storage device, means for processing and standardizing the information on the stored medical image data, means for performing diagnostic analysis using a learning model with the pre-processed medical image data, means for generating a diagnostic document based on the diagnostic analysis results, and means for encrypting and transmitting the generated diagnostic document to provide it to the user securely. This improves the accuracy of the diagnosis and makes it possible to provide diagnostic results quickly and with peace of mind. Furthermore, the accuracy of the diagnosis can be further enhanced by comparing it with past data, and at the same time, it becomes possible to build a database that can be used for future data analysis and research.
[0069] "Image acquisition equipment" refers to devices used to generate medical image data, and examples include CT scanners and MRI scanners.
[0070] "Medical image data" refers to image information generated by image acquisition equipment and used for diagnosis and treatment.
[0071] An "information storage device" is a digital storage system for storing medical image data and analysis results, and is a part of cloud storage, among other things.
[0072] "Processing to prepare information" refers to applying pre-processing steps such as homogenization and noise reduction to image data to make it suitable for analysis.
[0073] A "learning model" is an artificial intelligence system used to analyze medical image data, possessing pattern recognition capabilities based on past training data.
[0074] A "diagnostic document" is a report generated based on the results of analyzing medical imaging data, and includes information for identifying and diagnosing abnormalities.
[0075] A "user terminal" refers to a device such as a computer or tablet used by medical professionals to receive and view diagnostic documents.
[0076] "Encryption" is a technology that protects data using specific algorithms to safeguard information from unauthorized access.
[0077] The system of the present invention is designed to perform image analysis efficiently and accurately in medical institutions. Specific embodiments are described below.
[0078] The server first receives medical image data transmitted from imaging devices, such as CT scanners and MRI machines. At this time, the server acquires the data in real time via the network and stores it in cloud storage, which is an information storage device. Examples of cloud storage used include Amazon Web Services (AWS®) and Google Cloud Platform (GCP). The stored data is organized for reuse and backup purposes.
[0079] Next, the server processes the stored image data to prepare it for analysis. Specifically, it uses image processing libraries such as OpenCV to remove noise and equalize resolution, making the data suitable for analysis. Once this preprocessing is complete, the server inputs the data into a training model. The training model is built using AI frameworks such as TENSORFLOW® and PyTorch, and is trained on a large number of medical datasets. This model detects pathological findings and abnormalities with high accuracy.
[0080] Based on the analysis results, the server automatically generates a diagnostic document. This document includes identifying abnormalities and detailed information for diagnosis. The generated document is then encrypted and securely transmitted to the user's terminal. The user terminal, typically a computer or tablet, is used by medical professionals to open the diagnostic document and consider treatment plans for the patient based on its contents.
[0081] As a concrete example, if a cardiac CT image is input into a learning model and abnormal spots are detected, the model will reflect this as a possible myocardial infarction in the diagnostic document. This analysis result is compared with stored past data of the same patient to clarify the progression of the disease. An example of a prompt message to the generating AI model might be, "Based on the analysis results of the cardiac CT image, diagnose the possibility of myocardial infarction and create a comparison report with past data." Based on this prompt message, the generating AI model will perform an accurate analysis and report.
[0082] Thus, by using this system, medical institutions can perform rapid and accurate diagnoses non-invasively, significantly improving the quality of medical care.
[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0084] Step 1:
[0085] The server receives medical image data from image acquisition equipment. This data is transferred from devices such as CT and MRI scanners and transmitted to the server via the network. The input is unprocessed image data. The server saves the received data to cloud storage. This is to create a backup while retaining the original image data. The output is medical image data securely stored in the data storage device.
[0086] Step 2:
[0087] The server preprocesses the stored medical image data. The input is image data obtained from cloud storage. The server uses image processing libraries such as OpenCV to perform noise reduction and resolution equalization. Specifically, it applies a noise filter and processes the image to improve its detail. This results in data of optimal quality for AI models. The output is preprocessed, high-quality image data.
[0088] Step 3:
[0089] The server inputs pre-processed image data into a learning model. The input is pre-processed image data. The server executes an AI model built using TensorFlow or PyTorch, and determines the presence or absence of anomalies based on the data. For example, it recognizes unusual patterns or shapes within the image and makes a diagnosis based on them. The output is analysis data including the anomaly detection results.
[0090] Step 4:
[0091] The server generates a diagnostic document based on the analysis results. The input is the analysis results obtained by the AI model. The server processes this data and creates a diagnostic report summarizing the results. The report includes details of the detected anomalies and recommendations for diagnosis. Natural language processing technology is used in this generation process. The output is a completed report as a diagnostic document.
[0092] Step 5:
[0093] The server encrypts the generated diagnostic document and sends it to the user's terminal. The input is the generated diagnostic document. The server encrypts the document to enhance security and sends it to the medical professional's terminal using an appropriate communication protocol. Specifically, it establishes a secure communication channel using SSL / TLS. The output is a diagnostic report that can be decrypted on the recipient medical professional's terminal.
[0094] Step 6:
[0095] The user, a medical professional, reviews the received diagnostic documents using a terminal. The input is the diagnostic report received on the terminal. The user reviews the report, assesses the patient's condition, and obtains guidance for determining the necessary treatment plan. Specifically, they open the report using dedicated viewer software and examine the details. The output at this stage is the specific diagnostic result that will be incorporated into the treatment plan.
[0096] (Application Example 1)
[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0098] In the manufacturing industry, product quality control is extremely important, but traditional inspection methods often fail to detect defects in real time, leading to decreased production efficiency and quality. Furthermore, human visual inspection has limitations and is susceptible to human error. Under these circumstances, there is a need to automate the detection and reporting of defects on the production line to improve productivity.
[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0100] In this invention, the server includes means for storing visualization data acquired from an image acquisition mechanism in a remote storage device, means for pre-processing and standardizing the stored visualization data, means for performing analysis using a machine learning model with the pre-processed visualization data, means for detecting defective products on the manufacturing line using a device that photographs the components of the product, and means for detecting and notifying abnormalities in real time. This enables rapid automatic detection and notification of defective products on the manufacturing line, thereby improving production efficiency and quality.
[0101] An "image acquisition mechanism" is a device used to capture or acquire images of a product or other important visualization data.
[0102] "Visualized data" refers to images and other data that can be visually analyzed, and is used for product inspection and evaluation.
[0103] A "remote storage device" is a device used to store and manage data via a network, such as cloud storage.
[0104] "Preprocessing" refers to initial data processing, such as standardization and noise reduction, performed on acquired data to improve the accuracy of the analysis.
[0105] A "machine learning model" is an algorithm that learns specific patterns and features from large amounts of data to classify and predict data.
[0106] A "device for photographing product components" is a device that photographs the components of each product on the manufacturing line and uses those images for inspection.
[0107] "Means for detecting defective products" refers to a method or apparatus for determining and identifying whether a product on a manufacturing line is defective based on collected data.
[0108] "Means for detecting and notifying anomalies in real time" refers to a method or device that detects the results of an analysis almost immediately and immediately notifies the relevant parties if an anomaly is found.
[0109] The system realizing this invention consists of an image acquisition mechanism, a remote storage device, a preprocessing module, and a program including a machine learning model, all positioned on a manufacturing line. The server first acquires visualization data of the product using the image acquisition mechanism. This data is transferred in real time to a remote storage device, such as cloud storage, for rapid access.
[0110] Next, the server uses a preprocessing module to perform initial processing on this data, such as denoising and standardization. By using OpenCV or equivalent image processing software for this process, the data is prepared for analysis.
[0111] The preprocessed data is input into a machine learning model such as TensorFlow. This model is programmed to detect product anomalies based on a large amount of training data. Specifically, it learns the characteristics of defective products and determines whether or not a product's components are defective.
[0112] For example, if a manufactured smartphone case is cracked, a machine learning model can detect the anomaly in real time and immediately notify the product line manager. This real-time notification uses a web application framework such as Flask, and relevant information is sent to the manager's terminal each time an anomaly is detected.
[0113] An example of a prompt message used when a generative AI model is employed is as follows: "Using image data of the product captured by the smart camera, please use the AI model to detect surface anomalies and defects in real time."
[0114] In this way, the server enables the immediate identification and reporting of defective products during the manufacturing process, thereby improving the efficiency of quality control.
[0115] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0116] Step 1:
[0117] The server uses an image acquisition mechanism to obtain visualization data from products on the manufacturing line. The input is the manufactured product, and the output is high-resolution image data of the product. This image data is acquired in real time and used for rapid analysis.
[0118] Step 2:
[0119] The server transmits the acquired image data to a remote storage device for storage. The input is the high-resolution image data obtained in step 1, and the output is the image file stored on cloud storage. This storage operation makes the data accessible at any time.
[0120] Step 3:
[0121] The server uses a preprocessing module to perform standardization and denoising on the stored image data. The input is the original image on cloud storage, and the output is standardized, clean image data. This preprocessing is performed using OpenCV.
[0122] Step 4:
[0123] Using pre-processed image data, the server performs anomaly detection using a machine learning model. The input is the clean image data from step 3, and the output is the location and type of anomaly if detected. The trained model performs this using TensorFlow.
[0124] Step 5:
[0125] The server instantly notifies the administrator terminal based on the detected anomaly. The input is the anomaly information from step 4, and the output is a notification message detailing the anomaly. The notification is sent via a web application framework such as Flask.
[0126] This series of processes enables high-speed detection of defective products on the manufacturing line and appropriate reactions.
[0127] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0128] This invention combines an emotion engine with a system that analyzes the cause of death using medical image data to provide diagnostic information that takes into account the user's emotional state. In addition to acquiring, storing, preprocessing, analyzing, and generating reports on medical image data, this system includes an emotion engine that recognizes the user's emotions.
[0129] The server first receives medical image data from the image acquisition device and saves it to cloud storage. The saved images undergo preprocessing, including noise reduction and resolution adjustment, to prepare them for analysis by the AI model.
[0130] Next, the server inputs the pre-processed images into a deep learning model to extract features related to the cause of death. This process allows the model to determine the location of abnormalities and the likelihood of disease. Based on the analysis results, a detailed diagnostic report is generated.
[0131] The generated diagnostic report is provided to the user, and this is where the device utilizes an emotion engine. This emotion engine senses the user's facial expressions and tone of voice as they review the report and recognizes their emotions. Based on the recognized emotions, adjustments are made to how the report is presented and its content. For example, if the user is feeling anxious, additional explanations or supplementary information may be provided.
[0132] Users can view the adjusted reports through their devices and share information with other experts as needed. Furthermore, the emotional data recognized by the emotion engine is stored in cloud storage and used to improve future services.
[0133] This system goes beyond mere diagnosis, providing a new interface to improve the quality of medical services. It enables flexible responses tailored to the user's emotional state, improving the patient experience and delivering more personalized medical information.
[0134] The following describes the processing flow.
[0135] Step 1:
[0136] The server receives medical image data from the image acquisition device. This data is stored in DICOM format and organized and saved in cloud storage.
[0137] Step 2:
[0138] The server preprocesses the stored image data. Specifically, it removes noise and standardizes the image resolution. This process creates a uniform dataset suitable for analysis.
[0139] Step 3:
[0140] The server inputs pre-processed image data into a deep learning model. Based on knowledge learned from past training data, the model identifies abnormalities and lesions within the images. The resulting diagnostic results are then analyzed, and a report is generated.
[0141] Step 4:
[0142] The device displays the diagnostic report received from the server and simultaneously activates an emotion engine to recognize the user's emotions. The engine uses the camera and microphone to analyze the user's facial expressions and tone of voice to identify their emotions.
[0143] Step 5:
[0144] The device adjusts the report display based on the emotions it recognizes. For example, if the emotion engine detects user anxiety, the report will be adjusted to include additional explanations and reassuring information.
[0145] Step 6:
[0146] Users can review the adjusted report and, if necessary, share their diagnosis with other healthcare professionals. Additionally, emotional data recorded by the emotion engine is stored in the cloud for analysis.
[0147] Step 7:
[0148] The server will use the stored emotional data to improve future medical services. This will enable the provision of medical information more tailored to individual patients, aiming to improve patient satisfaction.
[0149] (Example 2)
[0150] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0151] Current medical imaging diagnostic systems have shortcomings, such as insufficient accuracy in diagnostic results and inadequate information provision tailored to the user's emotional state. Specifically, there is a need to improve the user experience because data comparison to improve the accuracy of medical image analysis and appropriate information provision tailored to the emotions of users receiving diagnostic results are not being carried out.
[0152] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0153] In this invention, the server includes means for recording medical image information acquired from an image acquisition means in a remote data storage means, means for performing data preprocessing on the recorded medical image information to standardize it, and means for performing diagnostic analysis using a machine learning model with the preprocessed medical image information. This makes it possible to recognize the user's emotional state and provide a diagnostic report that is adjusted according to that state, as well as to improve the accuracy of the analysis through comparison with past data.
[0154] "Image acquisition means" refers to devices and methods for collecting image data from medical imaging equipment.
[0155] "Remote data storage means" refers to servers or cloud storage that can store acquired data via the internet.
[0156] "Data preprocessing" refers to processes, including noise reduction and resolution adjustment, performed to improve the quality of acquired image data.
[0157] A "machine learning model" refers to an algorithm or system that extracts features from observed data and makes judgments or predictions.
[0158] "User equipment" refers to computers or mobile devices used by users to receive and view diagnostic reports.
[0159] "Means of recognizing emotional states" refers to technologies and devices that analyze a user's facial expressions and voice to determine their emotions.
[0160] A "diagnostic report" refers to a document or data that summarizes the results of medical image analysis and provides them to the user.
[0161] The system in this invention processes medical image information and provides a diagnostic report that takes into account the user's emotional state based on the analysis results. The embodiments of this system are described in detail below.
[0162] The server receives image data from the medical imaging device and records it in a remote data storage system. The received data is securely stored, for example, using a cloud-based storage service. Next, noise reduction is performed using the OpenCV library, and data preprocessing such as resolution adjustment is performed using the Python Imaging Library. This enables data analysis in a standardized format.
[0163] The server feeds pre-processed medical image data into machine learning models trained with TensorFlow or PyTorch to perform diagnostic analysis. The models utilize technologies such as Convolutional Neural Networks (CNNs) to identify abnormalities in the images and determine the likelihood of disease.
[0164] Based on the generated diagnostic analysis results, the server prepares to create a diagnostic report and send it to the user's device. Specifically, it uses visualization tools such as Matplotlib and Seaborn to graphically represent the diagnostic results and uses natural language generation technology to create a document that is easy for the user to understand.
[0165] When presenting a diagnostic report to the user, the device analyzes the user's facial expressions and voice using emotion recognition technology. This allows the device to determine in real time whether the user is experiencing anxiety or doubts, and adjust the content and presentation method of the report accordingly. For emotion recognition, APIs such as Microsoft® Azure® are used.
[0166] As a concrete example, consider a scenario where a user undergoes an MRI scan of their knee, and this system analyzes the results. The server performs noise reduction on the received image and begins analyzing for abnormalities in the knee area. As a result, a diagnostic report is generated stating, "No abnormalities were found in the knee." If the user is detected to be feeling anxious while viewing the report, the terminal displays a supplementary message saying, "There's no need to worry."
[0167] An example of a prompt message might be: "Analyze the MRI images of the knee, compile a diagnostic report indicating whether or not there are any abnormalities, and add information to alleviate the user's anxiety."
[0168] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0169] Step 1:
[0170] The server receives image data from a medical imaging device using an image acquisition method. The received data is received in a standardized format such as DICOM format. Since the input image data cannot be used directly for analysis, it is first recorded in a remote data storage system. During storage, the integrity of the data is checked, and if it is normal, the process proceeds to the next step.
[0171] Step 2:
[0172] The server performs data preprocessing on stored medical image information, including noise reduction and resolution adjustment. It takes the received image data as input and applies an OpenCV Gaussian filter to remove noise. Next, it adjusts the resolution using the Python Imaging Library and outputs standardized, high-quality image data. This process improves the accuracy of subsequent analyses.
[0173] Step 3:
[0174] The server feeds pre-processed image data into a machine learning model. This model is a generative AI model trained using TensorFlow, PyTorch, etc. The input is a pre-processed image, and the output is the anomaly detection result in the image. The model uses a Convolutional Neural Network (CNN) to extract features of the anomaly and make them ready for diagnosis.
[0175] Step 4:
[0176] The server generates a diagnostic report based on the results of a machine learning model's diagnostic analysis. The input is the analysis data, and the output is a diagnostic report in a user-friendly format. It uses tools like Matplotlib and Seaborn to graphically represent the data, and natural language generation technology to create the report content.
[0177] Step 5:
[0178] The device recognizes the user's emotions when providing the generated diagnostic report to the user. When the report is displayed, the built-in camera and microphone receive the user's facial expressions and voice as input, and the emotion engine analyzes their state. Using the Microsoft Azure API, the device obtains user emotion data (e.g., anxiety, reassurance) as output.
[0179] Step 6:
[0180] The device adjusts the content and presentation of the diagnostic report based on data obtained from the emotion engine. If the user is feeling anxious based on the emotional data input, a reassuring message such as "There's no need to worry" is added and displayed. This enables more personalized information delivery and improves the user experience.
[0181] Step 7:
[0182] Users review the diagnostic report provided through their device. They carefully examine the output report and, if necessary, request supplementary information or decide to share the information with other medical professionals. Throughout this process, the information is securely managed, allowing users to continue using the service with peace of mind.
[0183] (Application Example 2)
[0184] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0185] Traditional security systems simply record and play back video and audio data, but they have difficulty detecting abnormal situations or changes in people's emotional states in real time and prompting appropriate responses. Furthermore, they fail to provide reports that take user emotions into account, resulting in missed opportunities for effective countermeasures. There is a need to address these issues and provide a higher level of management and security.
[0186] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0187] In this invention, the server includes means for storing image data acquired from an image acquisition unit in network storage, means for preprocessing and standardizing the stored image data, and means for performing diagnostic analysis using a machine learning model with the preprocessed image data. This makes it possible to detect abnormal situations and changes in emotional states at the site in real time and provide information reports to user terminals. Furthermore, by analyzing the user's emotions and adjusting the report, it is possible to encourage appropriate responses.
[0188] The "image acquisition unit" refers to the devices and sensors used to collect video data within the system.
[0189] "Network storage" refers to an online storage system for storing and managing data via the internet.
[0190] "Preprocessing" refers to the preparatory steps for improving the accuracy of analysis by standardizing and denoising the acquired data.
[0191] A "machine learning model" refers to an algorithm that learns patterns from data and performs predictions and classifications.
[0192] "Diagnostic analysis" is the process of evaluating and determining specific events or conditions based on data.
[0193] An "information report" is a document or data that organizes analysis results and related information and provides it to users.
[0194] "User terminal" refers to a device used to display and operate information, and includes PCs, smartphones, and other similar devices.
[0195] "Emotional analysis" is the process of identifying and evaluating a person's emotional state based on video and audio data.
[0196] An "abnormal situation" refers to an unusual or irregular phenomenon or circumstance that is considered to require attention or action.
[0197] The system that realizes this invention consists of multiple servers, terminals, and users. First, the servers save the video data obtained from the image acquisition unit to network storage. The saved video data is preprocessed to remove noise and adjust the resolution. This preprocessing standardizes the data and makes it suitable for analysis. The preprocessed data is input into a machine learning model to detect abnormal situations and changes in emotional state in the video.
[0198] This analysis utilizes deep learning algorithms such as YOLO (You Only Look Once) and speech analysis using TensorFlow. The server generates an information report based on the analysis results and provides it to the user's device. This report also includes the analyzed emotional state. The device performs emotion analysis and flexibly adjusts the report content by sensing the user's facial expressions and voice nuances. OpenCV and Face API are used for emotion analysis, which enables the provision of additional information tailored to the user's emotions.
[0199] As a concrete example, in large commercial facilities, surveillance systems may use multiple cameras and microphones to perform real-time analysis and detect increases in anxiety or tension within the premises. When detected, an alert is sent to security staff, allowing for a swift response.
[0200] In this way, this system improves on-site safety and enables efficient management. An example of a prompt message used for this purpose is: "Provide real-time sentiment analysis data within the security area to detect situations where anxiety or tension is rising early. If an anomaly is detected, provide specific action guidelines as an alert."
[0201] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0202] Step 1:
[0203] The server receives video data from the image acquisition unit and stores it in network storage. Specifically, it collects video streams transmitted from cameras and surveillance sensors and saves each frame as digital data. The input is video data, and the output is the saved digital data.
[0204] Step 2:
[0205] The server performs preprocessing on the stored video data. This includes noise reduction and resolution adjustment, and aims to standardize the data. The input for preprocessing is the stored raw data, and the output is standardized data. Specifically, it applies filters to reduce noise and resizes frames.
[0206] Step 3:
[0207] The server inputs pre-processed data into a machine learning model, performs analysis, and detects abnormal situations and changes in emotional states. At this stage, algorithms such as YOLO are used to analyze the position and movement of people and objects. The input is pre-processed data, and the output is the detected features.
[0208] Step 4:
[0209] The server generates an information report based on the analysis results. The generated report includes identification of anomalies and an evaluation of detected emotional states. The input to the report is the analysis results, and the output is text information provided to the user. Specifically, this includes automatically generated comments and recommended actions.
[0210] Step 5:
[0211] The device receives the generated information report and analyzes the user's emotions. Using OpenCV and the Face API, it analyzes the user's facial expressions and voice tone to identify their emotional state. The input is data from the device's camera and microphone, and the output is parameters indicating the user's emotional state.
[0212] Step 6:
[0213] The device adjusts and presents an information report based on the recognized user's emotions. For example, if anxiety is detected, additional explanations or reassuring information are added. The input is the user's emotion parameters, and the output is the adjusted report. At this stage, the text is restructured and displayed on the screen.
[0214] Step 7:
[0215] Users review information reports provided through their devices and adjust their actions as needed. Based on the information presented, users understand the situation on-site and decide on corrective actions. The input is the adjusted report, and the output is the user's decision to take action. Specific examples include implementing safety measures and improvement plans tailored to the situation.
[0216] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0217] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0218] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0219] [Second Embodiment]
[0220] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0221] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0222] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0223] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0224] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0225] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0226] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0227] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0228] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0229] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0230] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0231] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0232] This invention provides a system for efficiently determining the cause of death using medical image data from an image acquisition device. This system includes a mechanism for storing image data in cloud storage, performing preprocessing, and analyzing the data using an artificial intelligence model.
[0233] The server has the function of receiving medical image data transmitted from hospital CT and MRI scanners and saving it to cloud storage. The saved data is preprocessed, including noise reduction and resolution standardization, and is ready for analysis.
[0234] Next, the server inputs the pre-processed data into an artificial intelligence model. This model has been trained on numerous training datasets and is designed to identify congenital abnormalities and pathological findings. For example, it can detect abnormal artifacts around the heart and suggest the possibility of a myocardial infarction.
[0235] After the analysis is performed, the server generates a diagnostic report and sends the results to a terminal for the user to use. This report contains detailed cause-of-death information obtained without an autopsy and serves as valuable material for medical professionals to review.
[0236] Furthermore, users can improve diagnostic accuracy by comparing current data with past image data of the same patient. This enables a multifaceted analysis that includes the progression of chronic diseases and the effectiveness of medications.
[0237] Furthermore, the diagnostic analysis results will be stored in the cloud and used for future data analysis and other research purposes. This accumulated database will serve as a foundation to support improvements in diagnostic technology and information sharing across regions.
[0238] The system according to the present invention enables rapid and accurate determination of the cause of death while reducing the burden of autopsy.
[0239] The following describes the processing flow.
[0240] Step 1:
[0241] The server receives medical image data from the hospital's CT and MRI machines. This data is typically provided in DICOM format and organized by patient ID and imaging date.
[0242] Step 2:
[0243] The server stores the received medical image data in cloud storage. This allows for efficient management of large amounts of data and easy access to it later.
[0244] Step 3:
[0245] The server preprocesses the stored image data. By removing noise and maintaining a consistent resolution, it prepares the data for the AI model to analyze efficiently.
[0246] Step 4:
[0247] The server inputs pre-processed image data into an artificial intelligence model. This model is trained using a vast amount of historical data and is tuned to detect abnormal findings and disease features.
[0248] Step 5:
[0249] The server generates a diagnostic report based on the analysis results from the AI model. This report includes identified lesions and possible causes of death, providing information useful for medical decision-making.
[0250] Step 6:
[0251] The server sends the generated diagnostic report to the terminal. In this case, the terminal is a PC or mobile device used by a medical professional.
[0252] Step 7:
[0253] Users can view diagnostic reports on their devices and share information with other medical professionals as needed. They can also compare past medical image data with current data to further improve the accuracy of their diagnoses.
[0254] Step 8:
[0255] The server stores the analysis results in the cloud. This makes the data available for future data analysis and research, creating a database that contributes to the advancement of medicine.
[0256] (Example 1)
[0257] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0258] In modern medical practice, there is a demand for rapid and accurate pathological diagnoses, but conventional systems lack the means to improve analytical accuracy and safely deliver diagnostic results. Furthermore, mechanisms for improving diagnostic accuracy through effective comparison with past medical data, and efficient methods for data storage and utilization, are insufficient. Therefore, challenges remain in formulating treatment plans for patients and managing their long-term health.
[0259] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0260] In this invention, the server includes means for storing medical image data acquired from an image acquisition device in an information storage device, means for processing and standardizing the information on the stored medical image data, means for performing diagnostic analysis using a learning model with the pre-processed medical image data, means for generating a diagnostic document based on the diagnostic analysis results, and means for encrypting and transmitting the generated diagnostic document to provide it to the user securely. This improves the accuracy of the diagnosis and makes it possible to provide diagnostic results quickly and with peace of mind. Furthermore, the accuracy of the diagnosis can be further enhanced by comparing it with past data, and at the same time, it becomes possible to build a database that can be used for future data analysis and research.
[0261] "Image acquisition equipment" refers to devices used to generate medical image data, and examples include CT scanners and MRI scanners.
[0262] "Medical image data" refers to image information generated by image acquisition equipment and used for diagnosis and treatment.
[0263] An "information storage device" is a digital storage system for storing medical image data and analysis results, and is a part of cloud storage, among other things.
[0264] "Processing to prepare information" refers to applying pre-processing steps such as homogenization and noise reduction to image data to make it suitable for analysis.
[0265] A "learning model" is an artificial intelligence system used to analyze medical image data, possessing pattern recognition capabilities based on past training data.
[0266] A "diagnostic document" is a report generated based on the results of analyzing medical imaging data, and includes information for identifying and diagnosing abnormalities.
[0267] A "user terminal" refers to a device such as a computer or tablet used by medical professionals to receive and view diagnostic documents.
[0268] "Encryption" is a technology that protects data using specific algorithms to safeguard information from unauthorized access.
[0269] The system of the present invention is designed to perform image analysis efficiently and accurately in medical institutions. Specific embodiments are described below.
[0270] The server first receives medical image data transmitted from imaging devices, such as CT scanners and MRI machines. At this time, the server acquires the data in real time via the network and stores it in cloud storage, which is an information storage device. Examples of cloud storage used include Amazon Web Services (AWS) and Google Cloud Platform (GCP). The stored data is organized for reuse and backup purposes.
[0271] Next, the server processes the stored image data to prepare it for analysis. Specifically, it uses image processing libraries such as OpenCV to remove noise and equalize resolution, making the data suitable for analysis. Once this preprocessing is complete, the server inputs the data into a training model. The training model is built using AI frameworks such as TensorFlow and PyTorch and is trained on a large number of medical datasets. This model detects pathological findings and abnormalities with high accuracy.
[0272] Based on the analysis results, the server automatically generates a diagnostic document. This document includes identifying abnormalities and detailed information for diagnosis. The generated document is then encrypted and securely transmitted to the user's terminal. The user terminal, typically a computer or tablet, is used by medical professionals to open the diagnostic document and consider treatment plans for the patient based on its contents.
[0273] As a concrete example, if a cardiac CT image is input into a learning model and abnormal spots are detected, the model will reflect this as a possible myocardial infarction in the diagnostic document. This analysis result is compared with stored past data of the same patient to clarify the progression of the disease. An example of a prompt message to the generating AI model might be, "Based on the analysis results of the cardiac CT image, diagnose the possibility of myocardial infarction and create a comparison report with past data." Based on this prompt message, the generating AI model will perform an accurate analysis and report.
[0274] Thus, by using this system, medical institutions can perform rapid and accurate diagnoses non-invasively, significantly improving the quality of medical care.
[0275] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0276] Step 1:
[0277] The server receives medical image data from image acquisition equipment. This data is transferred from devices such as CT and MRI scanners and transmitted to the server via the network. The input is unprocessed image data. The server saves the received data to cloud storage. This is to create a backup while retaining the original image data. The output is medical image data securely stored in the data storage device.
[0278] Step 2:
[0279] The server preprocesses the stored medical image data. The input is image data obtained from cloud storage. The server uses image processing libraries such as OpenCV to perform noise reduction and resolution equalization. Specifically, it applies a noise filter and processes the image to improve its detail. This results in data of optimal quality for AI models. The output is preprocessed, high-quality image data.
[0280] Step 3:
[0281] The server inputs pre-processed image data into a learning model. The input is pre-processed image data. The server executes an AI model built using TensorFlow or PyTorch, and determines the presence or absence of anomalies based on the data. For example, it recognizes unusual patterns or shapes within the image and makes a diagnosis based on them. The output is analysis data including the anomaly detection results.
[0282] Step 4:
[0283] The server generates a diagnostic document based on the analysis results. The input is the analysis results obtained by the AI model. The server processes this data and creates a diagnostic report summarizing the results. The report includes details of the detected anomalies and recommendations for diagnosis. Natural language processing technology is used in this generation process. The output is a completed report as a diagnostic document.
[0284] Step 5:
[0285] The server encrypts the generated diagnostic document and sends it to the user terminal. The input is the generated diagnostic document. The server encrypts the document to enhance security and sends it to the medical expert's terminal using an appropriate communication protocol. As a specific operation, a secure communication channel using SSL / TLS is established. The output is a diagnostic report that can be decoded on the terminal of the medical expert who is the recipient.
[0286] Step 6:
[0287] The medical expert who is the user checks the diagnostic document received using the terminal. The input is the diagnostic report received on the terminal. The user checks the report content, evaluates the patient's condition, and obtains guidelines for determining the necessary treatment plan. As a specific operation, the report is opened using dedicated viewer software and the details are examined. The output at this stage is the specific diagnostic result that is translated into a treatment plan.
[0288] (Application Example 1)
[0289] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0290] In the manufacturing industry, product quality control is very important. However, with conventional inspection methods, defective products are often not detected in real time, which easily causes a decrease in production efficiency and quality. In addition, there are limitations to visual inspection by humans, and human errors may occur. Under such circumstances, it is required to automate the detection and reporting of defective products on the production line to improve productivity.
[0291] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0292] In this invention, the server includes means for storing visualization data acquired from an image acquisition mechanism in a remote storage device, means for pre-processing and standardizing the stored visualization data, means for performing analysis using a machine learning model with the pre-processed visualization data, means for detecting defective products on the manufacturing line using a device that photographs the components of the product, and means for detecting and notifying abnormalities in real time. This enables rapid automatic detection and notification of defective products on the manufacturing line, thereby improving production efficiency and quality.
[0293] An "image acquisition mechanism" is a device used to capture or acquire images of a product or other important visualization data.
[0294] "Visualized data" refers to images and other data that can be visually analyzed, and is used for product inspection and evaluation.
[0295] A "remote storage device" is a device used to store and manage data via a network, such as cloud storage.
[0296] "Preprocessing" refers to initial data processing, such as standardization and noise reduction, performed on acquired data to improve the accuracy of the analysis.
[0297] A "machine learning model" is an algorithm that learns specific patterns and features from large amounts of data to classify and predict data.
[0298] A "device for photographing product components" is a device that photographs the components of each product on the manufacturing line and uses those images for inspection.
[0299] "Means for detecting defective products" refers to a method or apparatus for determining and identifying whether a product on a manufacturing line is defective based on collected data.
[0300] "Means for detecting and notifying anomalies in real time" refers to a method or device that detects the results of an analysis almost immediately and immediately notifies the relevant parties if an anomaly is found.
[0301] The system realizing this invention consists of an image acquisition mechanism, a remote storage device, a preprocessing module, and a program including a machine learning model, all positioned on a manufacturing line. The server first acquires visualization data of the product using the image acquisition mechanism. This data is transferred in real time to a remote storage device, such as cloud storage, for rapid access.
[0302] Next, the server uses a preprocessing module to perform initial processing on this data, such as denoising and standardization. By using OpenCV or equivalent image processing software for this process, the data is prepared for analysis.
[0303] The preprocessed data is input into a machine learning model such as TensorFlow. This model is programmed to detect product anomalies based on a large amount of training data. Specifically, it learns the characteristics of defective products and determines whether or not a product's components are defective.
[0304] For example, if a manufactured smartphone case is cracked, a machine learning model can detect the anomaly in real time and immediately notify the product line manager. This real-time notification uses a web application framework such as Flask, and relevant information is sent to the manager's terminal each time an anomaly is detected.
[0305] An example of a prompt message used when a generative AI model is employed is as follows: "Using image data of the product captured by the smart camera, please use the AI model to detect surface anomalies and defects in real time."
[0306] In this way, the server enables the immediate identification and reporting of defective products during the manufacturing process, thereby improving the efficiency of quality control.
[0307] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0308] Step 1:
[0309] The server uses an image acquisition mechanism to obtain visualization data from the products on the production line. The input is the manufactured product, and the output is the high-resolution image data of the product. This image data is acquired in real time and used for rapid analysis.
[0310] Step 2:
[0311] The server transmits the acquired image data to a remote storage device for storage. The input is the high-resolution image data obtained in Step 1, and the output is the image file stored on cloud storage. By this storage operation, the data can be accessed at any time.
[0312] Step 3:
[0313] The server performs normalization and noise removal on the stored image data using a preprocessing module. The input is the original image on cloud storage, and the output is the normalized clean image data. This preprocessing is executed using OpenCV.
[0314] Step 4:
[0315] Using the preprocessed image data, the server performs anomaly detection using a machine learning model. The input is the clean image data from Step 3, and the output is the location information and type of anomaly when an anomaly is detected. A trained model executes this using TensorFlow.
[0316] Step 5:
[0317] The server instantly notifies the administrator terminal based on the detected anomaly. The input is the anomaly information from step 4, and the output is a notification message detailing the anomaly. The notification is sent via a web application framework such as Flask.
[0318] This series of processes enables high-speed detection of defective products on the manufacturing line and appropriate reactions.
[0319] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0320] This invention combines an emotion engine with a system that analyzes the cause of death using medical image data to provide diagnostic information that takes into account the user's emotional state. In addition to acquiring, storing, preprocessing, analyzing, and generating reports on medical image data, this system includes an emotion engine that recognizes the user's emotions.
[0321] The server first receives medical image data from the image acquisition device and saves it to cloud storage. The saved images undergo preprocessing, including noise reduction and resolution adjustment, to prepare them for analysis by the AI model.
[0322] Next, the server inputs the pre-processed images into a deep learning model to extract features related to the cause of death. This process allows the model to determine the location of abnormalities and the likelihood of disease. Based on the analysis results, a detailed diagnostic report is generated.
[0323] The generated diagnostic report is provided to the user, and this is where the device utilizes an emotion engine. This emotion engine senses the user's facial expressions and tone of voice as they review the report and recognizes their emotions. Based on the recognized emotions, adjustments are made to how the report is presented and its content. For example, if the user is feeling anxious, additional explanations or supplementary information may be provided.
[0324] Users can view the adjusted reports through their devices and share information with other experts as needed. Furthermore, the emotional data recognized by the emotion engine is stored in cloud storage and used to improve future services.
[0325] This system goes beyond mere diagnosis, providing a new interface to improve the quality of medical services. It enables flexible responses tailored to the user's emotional state, improving the patient experience and delivering more personalized medical information.
[0326] The following describes the processing flow.
[0327] Step 1:
[0328] The server receives medical image data from the image acquisition device. This data is stored in DICOM format and organized and saved in cloud storage.
[0329] Step 2:
[0330] The server preprocesses the stored image data. Specifically, it removes noise and standardizes the image resolution. This process creates a uniform dataset suitable for analysis.
[0331] Step 3:
[0332] The server inputs pre-processed image data into a deep learning model. Based on knowledge learned from past training data, the model identifies abnormalities and lesions within the images. The resulting diagnostic results are then analyzed, and a report is generated.
[0333] Step 4:
[0334] The device displays the diagnostic report received from the server and simultaneously activates an emotion engine to recognize the user's emotions. The engine uses the camera and microphone to analyze the user's facial expressions and tone of voice to identify their emotions.
[0335] Step 5:
[0336] The device adjusts the report display based on the emotions it recognizes. For example, if the emotion engine detects user anxiety, the report will be adjusted to include additional explanations and reassuring information.
[0337] Step 6:
[0338] Users can review the adjusted report and, if necessary, share their diagnosis with other healthcare professionals. Additionally, emotional data recorded by the emotion engine is stored in the cloud for analysis.
[0339] Step 7:
[0340] The server will use the stored emotional data to improve future medical services. This will enable the provision of medical information more tailored to individual patients, aiming to improve patient satisfaction.
[0341] (Example 2)
[0342] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0343] Current medical imaging diagnostic systems have shortcomings, such as insufficient accuracy in diagnostic results and inadequate information provision tailored to the user's emotional state. Specifically, there is a need to improve the user experience because data comparison to improve the accuracy of medical image analysis and appropriate information provision tailored to the emotions of users receiving diagnostic results are not being carried out.
[0344] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0345] In this invention, the server includes means for recording medical image information acquired from an image acquisition means in a remote data storage means, means for performing data preprocessing on the recorded medical image information to standardize it, and means for performing diagnostic analysis using a machine learning model with the preprocessed medical image information. This makes it possible to recognize the user's emotional state and provide a diagnostic report that is adjusted according to that state, as well as to improve the accuracy of the analysis through comparison with past data.
[0346] "Image acquisition means" refers to devices and methods for collecting image data from medical imaging equipment.
[0347] "Remote data storage means" refers to servers or cloud storage that can store acquired data via the internet.
[0348] "Data preprocessing" refers to processes, including noise reduction and resolution adjustment, performed to improve the quality of acquired image data.
[0349] A "machine learning model" refers to an algorithm or system that extracts features from observed data and makes judgments or predictions.
[0350] "User equipment" refers to computers or mobile devices used by users to receive and view diagnostic reports.
[0351] "Means of recognizing emotional states" refers to technologies and devices that analyze a user's facial expressions and voice to determine their emotions.
[0352] A "diagnostic report" refers to a document or data that summarizes the results of medical image analysis and provides them to the user.
[0353] The system in this invention processes medical image information and provides a diagnostic report that takes into account the user's emotional state based on the analysis results. The embodiments of this system are described in detail below.
[0354] The server receives image data from the medical imaging device and records it in a remote data storage system. The received data is securely stored, for example, using a cloud-based storage service. Next, noise reduction is performed using the OpenCV library, and data preprocessing such as resolution adjustment is performed using the Python Imaging Library. This enables data analysis in a standardized format.
[0355] The server feeds pre-processed medical image data into machine learning models trained with TensorFlow or PyTorch to perform diagnostic analysis. The models utilize technologies such as Convolutional Neural Networks (CNNs) to identify abnormalities in the images and determine the likelihood of disease.
[0356] Based on the generated diagnostic analysis results, the server prepares to create a diagnostic report and send it to the user's device. Specifically, it uses visualization tools such as Matplotlib and Seaborn to graphically represent the diagnostic results and uses natural language generation technology to create a document that is easy for the user to understand.
[0357] When presenting a diagnostic report to the user, the device analyzes the user's facial expressions and voice using emotion recognition technology. This allows it to determine in real time whether the user is experiencing anxiety or doubts, and adjust the content and presentation method of the report accordingly. For emotion recognition, for example, Microsoft Azure APIs are used.
[0358] As a concrete example, consider a scenario where a user undergoes an MRI scan of their knee, and this system analyzes the results. The server performs noise reduction on the received image and begins analyzing for abnormalities in the knee area. As a result, a diagnostic report is generated stating, "No abnormalities were found in the knee." If the user is detected to be feeling anxious while viewing the report, the terminal displays a supplementary message saying, "There's no need to worry."
[0359] An example of a prompt message might be: "Analyze the MRI images of the knee, compile a diagnostic report indicating whether or not there are any abnormalities, and add information to alleviate the user's anxiety."
[0360] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0361] Step 1:
[0362] The server receives image data from a medical imaging device using an image acquisition method. The received data is received in a standardized format such as DICOM format. Since the input image data cannot be used directly for analysis, it is first recorded in a remote data storage system. During storage, the integrity of the data is checked, and if it is normal, the process proceeds to the next step.
[0363] Step 2:
[0364] The server performs data preprocessing on stored medical image information, including noise reduction and resolution adjustment. It takes the received image data as input and applies an OpenCV Gaussian filter to remove noise. Next, it adjusts the resolution using the Python Imaging Library and outputs standardized, high-quality image data. This process improves the accuracy of subsequent analyses.
[0365] Step 3:
[0366] The server feeds pre-processed image data into a machine learning model. This model is a generative AI model trained using TensorFlow, PyTorch, etc. The input is a pre-processed image, and the output is the anomaly detection result in the image. The model uses a Convolutional Neural Network (CNN) to extract features of the anomaly and make them ready for diagnosis.
[0367] Step 4:
[0368] The server generates a diagnostic report based on the results of a machine learning model's diagnostic analysis. The input is the analysis data, and the output is a diagnostic report in a user-friendly format. It uses tools like Matplotlib and Seaborn to graphically represent the data, and natural language generation technology to create the report content.
[0369] Step 5:
[0370] The device recognizes the user's emotions when providing the generated diagnostic report to the user. When the report is displayed, the built-in camera and microphone receive the user's facial expressions and voice as input, and the emotion engine analyzes their state. Using the Microsoft Azure API, the device obtains user emotion data (e.g., anxiety, reassurance) as output.
[0371] Step 6:
[0372] The device adjusts the content and presentation of the diagnostic report based on data obtained from the emotion engine. If the user is feeling anxious based on the emotional data input, a reassuring message such as "There's no need to worry" is added and displayed. This enables more personalized information delivery and improves the user experience.
[0373] Step 7:
[0374] Users review the diagnostic report provided through their device. They carefully examine the output report and, if necessary, request supplementary information or decide to share the information with other medical professionals. Throughout this process, the information is securely managed, allowing users to continue using the service with peace of mind.
[0375] (Application Example 2)
[0376] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0377] Traditional security systems simply record and play back video and audio data, but they have difficulty detecting abnormal situations or changes in people's emotional states in real time and prompting appropriate responses. Furthermore, they fail to provide reports that take user emotions into account, resulting in missed opportunities for effective countermeasures. There is a need to address these issues and provide a higher level of management and security.
[0378] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0379] In this invention, the server includes means for storing image data acquired from an image acquisition unit in network storage, means for preprocessing and standardizing the stored image data, and means for performing diagnostic analysis using a machine learning model with the preprocessed image data. This makes it possible to detect abnormal situations and changes in emotional states at the site in real time and provide information reports to user terminals. Furthermore, by analyzing the user's emotions and adjusting the report, it is possible to encourage appropriate responses.
[0380] The "image acquisition unit" refers to the devices and sensors used to collect video data within the system.
[0381] "Network storage" refers to an online storage system for storing and managing data via the internet.
[0382] "Preprocessing" refers to the preparatory steps for improving the accuracy of analysis by standardizing and denoising the acquired data.
[0383] A "machine learning model" refers to an algorithm that learns patterns from data and performs predictions and classifications.
[0384] "Diagnostic analysis" is the process of evaluating and determining specific events or conditions based on data.
[0385] An "information report" is a document or data that organizes analysis results and related information and provides it to users.
[0386] "User terminal" refers to a device used to display and operate information, and includes PCs, smartphones, and other similar devices.
[0387] "Emotional analysis" is the process of identifying and evaluating a person's emotional state based on video and audio data.
[0388] An "abnormal situation" refers to an unusual or irregular phenomenon or circumstance that is considered to require attention or action.
[0389] The system that realizes this invention consists of multiple servers, terminals, and users. First, the servers save the video data obtained from the image acquisition unit to network storage. The saved video data is preprocessed to remove noise and adjust the resolution. This preprocessing standardizes the data and makes it suitable for analysis. The preprocessed data is input into a machine learning model to detect abnormal situations and changes in emotional state in the video.
[0390] This analysis utilizes deep learning algorithms such as YOLO (You Only Look Once) and speech analysis using TensorFlow. The server generates an information report based on the analysis results and provides it to the user's device. This report also includes the analyzed emotional state. The device performs emotion analysis and flexibly adjusts the report content by sensing the user's facial expressions and voice nuances. OpenCV and Face API are used for emotion analysis, which enables the provision of additional information tailored to the user's emotions.
[0391] As a concrete example, in large commercial facilities, surveillance systems may use multiple cameras and microphones to perform real-time analysis and detect increases in anxiety or tension within the premises. When detected, an alert is sent to security staff, allowing for a swift response.
[0392] In this way, this system improves on-site safety and enables efficient management. An example of a prompt message used for this purpose is: "Provide real-time sentiment analysis data within the security area to detect situations where anxiety or tension is rising early. If an anomaly is detected, provide specific action guidelines as an alert."
[0393] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0394] Step 1:
[0395] The server receives video data from the image acquisition unit and stores it in network storage. Specifically, it collects video streams transmitted from cameras and surveillance sensors and saves each frame as digital data. The input is video data, and the output is the saved digital data.
[0396] Step 2:
[0397] The server performs preprocessing on the stored video data. This includes noise reduction and resolution adjustment, and aims to standardize the data. The input for preprocessing is the stored raw data, and the output is standardized data. Specifically, it applies filters to reduce noise and resizes frames.
[0398] Step 3:
[0399] The server inputs pre-processed data into a machine learning model, performs analysis, and detects abnormal situations and changes in emotional states. At this stage, algorithms such as YOLO are used to analyze the position and movement of people and objects. The input is pre-processed data, and the output is the detected features.
[0400] Step 4:
[0401] The server generates an information report based on the analysis results. The generated report includes identification of anomalies and an evaluation of detected emotional states. The input to the report is the analysis results, and the output is text information provided to the user. Specifically, this includes automatically generated comments and recommended actions.
[0402] Step 5:
[0403] The device receives the generated information report and analyzes the user's emotions. Using OpenCV and the Face API, it analyzes the user's facial expressions and voice tone to identify their emotional state. The input is data from the device's camera and microphone, and the output is parameters indicating the user's emotional state.
[0404] Step 6:
[0405] The device adjusts and presents an information report based on the recognized user's emotions. For example, if anxiety is detected, additional explanations or reassuring information are added. The input is the user's emotion parameters, and the output is the adjusted report. At this stage, the text is restructured and displayed on the screen.
[0406] Step 7:
[0407] Users review information reports provided through their devices and adjust their actions as needed. Based on the information presented, users understand the situation on-site and decide on corrective actions. The input is the adjusted report, and the output is the user's decision to take action. Specific examples include implementing safety measures and improvement plans tailored to the situation.
[0408] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0409] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0410] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0411] [Third Embodiment]
[0412] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0413] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0414] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0415] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0416] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0417] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0418] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0419] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0420] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0421] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0422] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0423] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0424] This invention provides a system for efficiently determining the cause of death using medical image data from an image acquisition device. This system includes a mechanism for storing image data in cloud storage, performing preprocessing, and analyzing the data using an artificial intelligence model.
[0425] The server has the function of receiving medical image data transmitted from hospital CT and MRI scanners and saving it to cloud storage. The saved data is preprocessed, including noise reduction and resolution standardization, and is ready for analysis.
[0426] Next, the server inputs the pre-processed data into an artificial intelligence model. This model has been trained on numerous training datasets and is designed to identify congenital abnormalities and pathological findings. For example, it can detect abnormal artifacts around the heart and suggest the possibility of a myocardial infarction.
[0427] After the analysis is performed, the server generates a diagnostic report and sends the results to a terminal for the user to use. This report contains detailed cause-of-death information obtained without an autopsy and serves as valuable material for medical professionals to review.
[0428] Furthermore, users can improve diagnostic accuracy by comparing current data with past image data of the same patient. This enables a multifaceted analysis that includes the progression of chronic diseases and the effectiveness of medications.
[0429] Furthermore, the diagnostic analysis results will be stored in the cloud and used for future data analysis and other research purposes. This accumulated database will serve as a foundation to support improvements in diagnostic technology and information sharing across regions.
[0430] The system according to the present invention enables rapid and accurate determination of the cause of death while reducing the burden of autopsy.
[0431] The following describes the processing flow.
[0432] Step 1:
[0433] The server receives medical image data from the hospital's CT and MRI machines. This data is typically provided in DICOM format and organized by patient ID and imaging date.
[0434] Step 2:
[0435] The server stores the received medical image data in cloud storage. This allows for efficient management of large amounts of data and easy access to it later.
[0436] Step 3:
[0437] The server preprocesses the stored image data. By removing noise and maintaining a consistent resolution, it prepares the data for the AI model to analyze efficiently.
[0438] Step 4:
[0439] The server inputs pre-processed image data into an artificial intelligence model. This model is trained using a vast amount of historical data and is tuned to detect abnormal findings and disease features.
[0440] Step 5:
[0441] The server generates a diagnostic report based on the analysis results from the AI model. This report includes identified lesions and possible causes of death, providing information useful for medical decision-making.
[0442] Step 6:
[0443] The server sends the generated diagnostic report to the terminal. In this case, the terminal is a PC or mobile device used by a medical professional.
[0444] Step 7:
[0445] Users can view diagnostic reports on their devices and share information with other medical professionals as needed. They can also compare past medical image data with current data to further improve the accuracy of their diagnoses.
[0446] Step 8:
[0447] The server stores the analysis results in the cloud. This makes the data available for future data analysis and research, creating a database that contributes to the advancement of medicine.
[0448] (Example 1)
[0449] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0450] In modern medical practice, there is a demand for rapid and accurate pathological diagnoses, but conventional systems lack the means to improve analytical accuracy and safely deliver diagnostic results. Furthermore, mechanisms for improving diagnostic accuracy through effective comparison with past medical data, and efficient methods for data storage and utilization, are insufficient. Therefore, challenges remain in formulating treatment plans for patients and managing their long-term health.
[0451] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0452] In this invention, the server includes means for storing medical image data acquired from an image acquisition device in an information storage device, means for processing and standardizing the information on the stored medical image data, means for performing diagnostic analysis using a learning model with the pre-processed medical image data, means for generating a diagnostic document based on the diagnostic analysis results, and means for encrypting and transmitting the generated diagnostic document to provide it to the user securely. This improves the accuracy of the diagnosis and makes it possible to provide diagnostic results quickly and with peace of mind. Furthermore, the accuracy of the diagnosis can be further enhanced by comparing it with past data, and at the same time, it becomes possible to build a database that can be used for future data analysis and research.
[0453] "Image acquisition equipment" refers to devices used to generate medical image data, and examples include CT scanners and MRI scanners.
[0454] "Medical image data" refers to image information generated by image acquisition equipment and used for diagnosis and treatment.
[0455] An "information storage device" is a digital storage system for storing medical image data and analysis results, and is a part of cloud storage, among other things.
[0456] "Processing to prepare information" refers to applying pre-processing steps such as homogenization and noise reduction to image data to make it suitable for analysis.
[0457] A "learning model" is an artificial intelligence system used to analyze medical image data, possessing pattern recognition capabilities based on past training data.
[0458] A "diagnostic document" is a report generated based on the results of analyzing medical imaging data, and includes information for identifying and diagnosing abnormalities.
[0459] A "user terminal" refers to a device such as a computer or tablet used by medical professionals to receive and view diagnostic documents.
[0460] "Encryption" is a technology that protects data using specific algorithms to safeguard information from unauthorized access.
[0461] The system of the present invention is designed to perform image analysis efficiently and accurately in medical institutions. Specific embodiments are described below.
[0462] The server first receives medical image data transmitted from imaging devices, such as CT scanners and MRI machines. At this time, the server acquires the data in real time via the network and stores it in cloud storage, which is an information storage device. Examples of cloud storage used include Amazon Web Services (AWS) and Google Cloud Platform (GCP). The stored data is organized for reuse and backup purposes.
[0463] Next, the server processes the stored image data to prepare it for analysis. Specifically, it uses image processing libraries such as OpenCV to remove noise and equalize resolution, making the data suitable for analysis. Once this preprocessing is complete, the server inputs the data into a training model. The training model is built using AI frameworks such as TensorFlow and PyTorch and is trained on a large number of medical datasets. This model detects pathological findings and abnormalities with high accuracy.
[0464] Based on the analysis results, the server automatically generates a diagnostic document. This document includes identifying abnormalities and detailed information for diagnosis. The generated document is then encrypted and securely transmitted to the user's terminal. The user terminal, typically a computer or tablet, is used by medical professionals to open the diagnostic document and consider treatment plans for the patient based on its contents.
[0465] As a concrete example, if a cardiac CT image is input into a learning model and abnormal spots are detected, the model will reflect this as a possible myocardial infarction in the diagnostic document. This analysis result is compared with stored past data of the same patient to clarify the progression of the disease. An example of a prompt message to the generating AI model might be, "Based on the analysis results of the cardiac CT image, diagnose the possibility of myocardial infarction and create a comparison report with past data." Based on this prompt message, the generating AI model will perform an accurate analysis and report.
[0466] Thus, by using this system, medical institutions can perform rapid and accurate diagnoses non-invasively, significantly improving the quality of medical care.
[0467] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0468] Step 1:
[0469] The server receives medical image data from image acquisition equipment. This data is transferred from devices such as CT and MRI scanners and transmitted to the server via the network. The input is unprocessed image data. The server saves the received data to cloud storage. This is to create a backup while retaining the original image data. The output is medical image data securely stored in the data storage device.
[0470] Step 2:
[0471] The server preprocesses the stored medical image data. The input is image data obtained from cloud storage. The server uses image processing libraries such as OpenCV to perform noise reduction and resolution equalization. Specifically, it applies a noise filter and processes the image to improve its detail. This results in data of optimal quality for AI models. The output is preprocessed, high-quality image data.
[0472] Step 3:
[0473] The server inputs pre-processed image data into a learning model. The input is pre-processed image data. The server executes an AI model built using TensorFlow or PyTorch, and determines the presence or absence of anomalies based on the data. For example, it recognizes unusual patterns or shapes within the image and makes a diagnosis based on them. The output is analysis data including the anomaly detection results.
[0474] Step 4:
[0475] The server generates a diagnostic document based on the analysis results. The input is the analysis results obtained by the AI model. The server processes this data and creates a diagnostic report summarizing the results. The report includes details of the detected anomalies and recommendations for diagnosis. Natural language processing technology is used in this generation process. The output is a completed report as a diagnostic document.
[0476] Step 5:
[0477] The server encrypts the generated diagnostic document and sends it to the user's terminal. The input is the generated diagnostic document. The server encrypts the document to enhance security and sends it to the medical professional's terminal using an appropriate communication protocol. Specifically, it establishes a secure communication channel using SSL / TLS. The output is a diagnostic report that can be decrypted on the recipient medical professional's terminal.
[0478] Step 6:
[0479] The user, a medical professional, reviews the received diagnostic documents using a terminal. The input is the diagnostic report received on the terminal. The user reviews the report, assesses the patient's condition, and obtains guidance for determining the necessary treatment plan. Specifically, they open the report using dedicated viewer software and examine the details. The output at this stage is the specific diagnostic result that will be incorporated into the treatment plan.
[0480] (Application Example 1)
[0481] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0482] In the manufacturing industry, product quality control is extremely important, but traditional inspection methods often fail to detect defects in real time, leading to decreased production efficiency and quality. Furthermore, human visual inspection has limitations and is susceptible to human error. Under these circumstances, there is a need to automate the detection and reporting of defects on the production line to improve productivity.
[0483] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0484] In this invention, the server includes means for storing visualization data acquired from an image acquisition mechanism in a remote storage device, means for pre-processing and standardizing the stored visualization data, means for performing analysis using a machine learning model with the pre-processed visualization data, means for detecting defective products on the manufacturing line using a device that photographs the components of the product, and means for detecting and notifying abnormalities in real time. This enables rapid automatic detection and notification of defective products on the manufacturing line, thereby improving production efficiency and quality.
[0485] An "image acquisition mechanism" is a device used to capture or acquire images of a product or other important visualization data.
[0486] "Visualized data" refers to images and other data that can be visually analyzed, and is used for product inspection and evaluation.
[0487] A "remote storage device" is a device used to store and manage data via a network, such as cloud storage.
[0488] "Preprocessing" refers to initial data processing, such as standardization and noise reduction, performed on acquired data to improve the accuracy of the analysis.
[0489] A "machine learning model" is an algorithm that learns specific patterns and features from large amounts of data to classify and predict data.
[0490] A "device for photographing product components" is a device that photographs the components of each product on the manufacturing line and uses those images for inspection.
[0491] "Means for detecting defective products" refers to a method or apparatus for determining and identifying whether a product on a manufacturing line is defective based on collected data.
[0492] "Means for detecting and notifying anomalies in real time" refers to a method or device that detects the results of an analysis almost immediately and immediately notifies the relevant parties if an anomaly is found.
[0493] The system realizing this invention consists of an image acquisition mechanism, a remote storage device, a preprocessing module, and a program including a machine learning model, all positioned on a manufacturing line. The server first acquires visualization data of the product using the image acquisition mechanism. This data is transferred in real time to a remote storage device, such as cloud storage, for rapid access.
[0494] Next, the server uses a preprocessing module to perform initial processing on this data, such as denoising and standardization. By using OpenCV or equivalent image processing software for this process, the data is prepared for analysis.
[0495] The preprocessed data is input into a machine learning model such as TensorFlow. This model is programmed to detect product anomalies based on a large amount of training data. Specifically, it learns the characteristics of defective products and determines whether or not a product's components are defective.
[0496] For example, if a manufactured smartphone case is cracked, a machine learning model can detect the anomaly in real time and immediately notify the product line manager. This real-time notification uses a web application framework such as Flask, and relevant information is sent to the manager's terminal each time an anomaly is detected.
[0497] An example of a prompt message used when a generative AI model is employed is as follows: "Using image data of the product captured by the smart camera, please use the AI model to detect surface anomalies and defects in real time."
[0498] In this way, the server enables the immediate identification and reporting of defective products during the manufacturing process, thereby improving the efficiency of quality control.
[0499] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0500] Step 1:
[0501] The server uses an image acquisition mechanism to obtain visualization data from products on the manufacturing line. The input is the manufactured product, and the output is high-resolution image data of the product. This image data is acquired in real time and used for rapid analysis.
[0502] Step 2:
[0503] The server transmits the acquired image data to a remote storage device for storage. The input is the high-resolution image data obtained in step 1, and the output is the image file stored on cloud storage. This storage operation makes the data accessible at any time.
[0504] Step 3:
[0505] The server uses a preprocessing module to perform standardization and denoising on the stored image data. The input is the original image on cloud storage, and the output is standardized, clean image data. This preprocessing is performed using OpenCV.
[0506] Step 4:
[0507] Using pre-processed image data, the server performs anomaly detection using a machine learning model. The input is the clean image data from step 3, and the output is the location and type of anomaly if detected. The trained model performs this using TensorFlow.
[0508] Step 5:
[0509] The server instantly notifies the administrator terminal based on the detected anomaly. The input is the anomaly information from step 4, and the output is a notification message detailing the anomaly. The notification is sent via a web application framework such as Flask.
[0510] This series of processes enables high-speed detection of defective products on the manufacturing line and appropriate reactions.
[0511] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0512] This invention combines an emotion engine with a system that analyzes the cause of death using medical image data to provide diagnostic information that takes into account the user's emotional state. In addition to acquiring, storing, preprocessing, analyzing, and generating reports on medical image data, this system includes an emotion engine that recognizes the user's emotions.
[0513] The server first receives medical image data from the image acquisition device and saves it to cloud storage. The saved images undergo preprocessing, including noise reduction and resolution adjustment, to prepare them for analysis by the AI model.
[0514] Next, the server inputs the pre-processed images into a deep learning model to extract features related to the cause of death. This process allows the model to determine the location of abnormalities and the likelihood of disease. Based on the analysis results, a detailed diagnostic report is generated.
[0515] The generated diagnostic report is provided to the user, and this is where the device utilizes an emotion engine. This emotion engine senses the user's facial expressions and tone of voice as they review the report and recognizes their emotions. Based on the recognized emotions, adjustments are made to how the report is presented and its content. For example, if the user is feeling anxious, additional explanations or supplementary information may be provided.
[0516] Users can view the adjusted reports through their devices and share information with other experts as needed. Furthermore, the emotional data recognized by the emotion engine is stored in cloud storage and used to improve future services.
[0517] This system goes beyond mere diagnosis, providing a new interface to improve the quality of medical services. It enables flexible responses tailored to the user's emotional state, improving the patient experience and delivering more personalized medical information.
[0518] The following describes the processing flow.
[0519] Step 1:
[0520] The server receives medical image data from the image acquisition device. This data is stored in DICOM format and organized and saved in cloud storage.
[0521] Step 2:
[0522] The server preprocesses the stored image data. Specifically, it removes noise and standardizes the image resolution. This process creates a uniform dataset suitable for analysis.
[0523] Step 3:
[0524] The server inputs pre-processed image data into a deep learning model. Based on knowledge learned from past training data, the model identifies abnormalities and lesions within the images. The resulting diagnostic results are then analyzed, and a report is generated.
[0525] Step 4:
[0526] The device displays the diagnostic report received from the server and simultaneously activates an emotion engine to recognize the user's emotions. The engine uses the camera and microphone to analyze the user's facial expressions and tone of voice to identify their emotions.
[0527] Step 5:
[0528] The device adjusts the report display based on the emotions it recognizes. For example, if the emotion engine detects user anxiety, the report will be adjusted to include additional explanations and reassuring information.
[0529] Step 6:
[0530] Users can review the adjusted report and, if necessary, share their diagnosis with other healthcare professionals. Additionally, emotional data recorded by the emotion engine is stored in the cloud for analysis.
[0531] Step 7:
[0532] The server will use the stored emotional data to improve future medical services. This will enable the provision of medical information more tailored to individual patients, aiming to improve patient satisfaction.
[0533] (Example 2)
[0534] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0535] Current medical imaging diagnostic systems have shortcomings, such as insufficient accuracy in diagnostic results and inadequate information provision tailored to the user's emotional state. Specifically, there is a need to improve the user experience because data comparison to improve the accuracy of medical image analysis and appropriate information provision tailored to the emotions of users receiving diagnostic results are not being carried out.
[0536] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0537] In this invention, the server includes means for recording medical image information acquired from an image acquisition means in a remote data storage means, means for performing data preprocessing on the recorded medical image information to standardize it, and means for performing diagnostic analysis using a machine learning model with the preprocessed medical image information. This makes it possible to recognize the user's emotional state and provide a diagnostic report that is adjusted according to that state, as well as to improve the accuracy of the analysis through comparison with past data.
[0538] "Image acquisition means" refers to devices and methods for collecting image data from medical imaging equipment.
[0539] "Remote data storage means" refers to servers or cloud storage that can store acquired data via the internet.
[0540] "Data preprocessing" refers to processes, including noise reduction and resolution adjustment, performed to improve the quality of acquired image data.
[0541] A "machine learning model" refers to an algorithm or system that extracts features from observed data and makes judgments or predictions.
[0542] "User equipment" refers to computers or mobile devices used by users to receive and view diagnostic reports.
[0543] "Means of recognizing emotional states" refers to technologies and devices that analyze a user's facial expressions and voice to determine their emotions.
[0544] A "diagnostic report" refers to a document or data that summarizes the results of medical image analysis and provides them to the user.
[0545] The system in this invention processes medical image information and provides a diagnostic report that takes into account the user's emotional state based on the analysis results. The embodiments of this system are described in detail below.
[0546] The server receives image data from the medical imaging device and records it in a remote data storage system. The received data is securely stored, for example, using a cloud-based storage service. Next, noise reduction is performed using the OpenCV library, and data preprocessing such as resolution adjustment is performed using the Python Imaging Library. This enables data analysis in a standardized format.
[0547] The server feeds pre-processed medical image data into machine learning models trained with TensorFlow or PyTorch to perform diagnostic analysis. The models utilize technologies such as Convolutional Neural Networks (CNNs) to identify abnormalities in the images and determine the likelihood of disease.
[0548] Based on the generated diagnostic analysis results, the server prepares to create a diagnostic report and send it to the user's device. Specifically, it uses visualization tools such as Matplotlib and Seaborn to graphically represent the diagnostic results and uses natural language generation technology to create a document that is easy for the user to understand.
[0549] When presenting a diagnostic report to the user, the device analyzes the user's facial expressions and voice using emotion recognition technology. This allows it to determine in real time whether the user is experiencing anxiety or doubts, and adjust the content and presentation method of the report accordingly. For emotion recognition, for example, Microsoft Azure APIs are used.
[0550] As a concrete example, consider a scenario where a user undergoes an MRI scan of their knee, and this system analyzes the results. The server performs noise reduction on the received image and begins analyzing for abnormalities in the knee area. As a result, a diagnostic report is generated stating, "No abnormalities were found in the knee." If the user is detected to be feeling anxious while viewing the report, the terminal displays a supplementary message saying, "There's no need to worry."
[0551] An example of a prompt message might be: "Analyze the MRI images of the knee, compile a diagnostic report indicating whether or not there are any abnormalities, and add information to alleviate the user's anxiety."
[0552] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0553] Step 1:
[0554] The server receives image data from a medical imaging device using an image acquisition method. The received data is received in a standardized format such as DICOM format. Since the input image data cannot be used directly for analysis, it is first recorded in a remote data storage system. During storage, the integrity of the data is checked, and if it is normal, the process proceeds to the next step.
[0555] Step 2:
[0556] The server performs data preprocessing on stored medical image information, including noise reduction and resolution adjustment. It takes the received image data as input and applies an OpenCV Gaussian filter to remove noise. Next, it adjusts the resolution using the Python Imaging Library and outputs standardized, high-quality image data. This process improves the accuracy of subsequent analyses.
[0557] Step 3:
[0558] The server feeds pre-processed image data into a machine learning model. This model is a generative AI model trained using TensorFlow, PyTorch, etc. The input is a pre-processed image, and the output is the anomaly detection result in the image. The model uses a Convolutional Neural Network (CNN) to extract features of the anomaly and make them ready for diagnosis.
[0559] Step 4:
[0560] The server generates a diagnostic report based on the results of a machine learning model's diagnostic analysis. The input is the analysis data, and the output is a diagnostic report in a user-friendly format. It uses tools like Matplotlib and Seaborn to graphically represent the data, and natural language generation technology to create the report content.
[0561] Step 5:
[0562] The device recognizes the user's emotions when providing the generated diagnostic report to the user. When the report is displayed, the built-in camera and microphone receive the user's facial expressions and voice as input, and the emotion engine analyzes their state. Using the Microsoft Azure API, the device obtains user emotion data (e.g., anxiety, reassurance) as output.
[0563] Step 6:
[0564] The device adjusts the content and presentation of the diagnostic report based on data obtained from the emotion engine. If the user is feeling anxious based on the emotional data input, a reassuring message such as "There's no need to worry" is added and displayed. This enables more personalized information delivery and improves the user experience.
[0565] Step 7:
[0566] Users review the diagnostic report provided through their device. They carefully examine the output report and, if necessary, request supplementary information or decide to share the information with other medical professionals. Throughout this process, the information is securely managed, allowing users to continue using the service with peace of mind.
[0567] (Application Example 2)
[0568] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0569] Traditional security systems simply record and play back video and audio data, but they have difficulty detecting abnormal situations or changes in people's emotional states in real time and prompting appropriate responses. Furthermore, they fail to provide reports that take user emotions into account, resulting in missed opportunities for effective countermeasures. There is a need to address these issues and provide a higher level of management and security.
[0570] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0571] In this invention, the server includes means for storing image data acquired from an image acquisition unit in network storage, means for preprocessing and standardizing the stored image data, and means for performing diagnostic analysis using a machine learning model with the preprocessed image data. This makes it possible to detect abnormal situations and changes in emotional states at the site in real time and provide information reports to user terminals. Furthermore, by analyzing the user's emotions and adjusting the report, it is possible to encourage appropriate responses.
[0572] The "image acquisition unit" refers to the devices and sensors used to collect video data within the system.
[0573] "Network storage" refers to an online storage system for storing and managing data via the internet.
[0574] "Preprocessing" refers to the preparatory steps for improving the accuracy of analysis by standardizing and denoising the acquired data.
[0575] A "machine learning model" refers to an algorithm that learns patterns from data and performs predictions and classifications.
[0576] "Diagnostic analysis" is the process of evaluating and determining specific events or conditions based on data.
[0577] An "information report" is a document or data that organizes analysis results and related information and provides it to users.
[0578] "User terminal" refers to a device used to display and operate information, and includes PCs, smartphones, and other similar devices.
[0579] "Emotional analysis" is the process of identifying and evaluating a person's emotional state based on video and audio data.
[0580] An "abnormal situation" refers to an unusual or irregular phenomenon or circumstance that is considered to require attention or action.
[0581] The system that realizes this invention consists of multiple servers, terminals, and users. First, the servers save the video data obtained from the image acquisition unit to network storage. The saved video data is preprocessed to remove noise and adjust the resolution. This preprocessing standardizes the data and makes it suitable for analysis. The preprocessed data is input into a machine learning model to detect abnormal situations and changes in emotional state in the video.
[0582] This analysis utilizes deep learning algorithms such as YOLO (You Only Look Once) and speech analysis using TensorFlow. The server generates an information report based on the analysis results and provides it to the user's device. This report also includes the analyzed emotional state. The device performs emotion analysis and flexibly adjusts the report content by sensing the user's facial expressions and voice nuances. OpenCV and Face API are used for emotion analysis, which enables the provision of additional information tailored to the user's emotions.
[0583] As a concrete example, in large commercial facilities, surveillance systems may use multiple cameras and microphones to perform real-time analysis and detect increases in anxiety or tension within the premises. When detected, an alert is sent to security staff, allowing for a swift response.
[0584] In this way, this system improves on-site safety and enables efficient management. An example of a prompt message used for this purpose is: "Provide real-time sentiment analysis data within the security area to detect situations where anxiety or tension is rising early. If an anomaly is detected, provide specific action guidelines as an alert."
[0585] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0586] Step 1:
[0587] The server receives video data from the image acquisition unit and stores it in network storage. Specifically, it collects video streams transmitted from cameras and surveillance sensors and saves each frame as digital data. The input is video data, and the output is the saved digital data.
[0588] Step 2:
[0589] The server performs preprocessing on the stored video data. This includes noise reduction and resolution adjustment, and aims to standardize the data. The input for preprocessing is the stored raw data, and the output is standardized data. Specifically, it applies filters to reduce noise and resizes frames.
[0590] Step 3:
[0591] The server inputs pre-processed data into a machine learning model, performs analysis, and detects abnormal situations and changes in emotional states. At this stage, algorithms such as YOLO are used to analyze the position and movement of people and objects. The input is pre-processed data, and the output is the detected features.
[0592] Step 4:
[0593] The server generates an information report based on the analysis results. The generated report includes identification of anomalies and an evaluation of detected emotional states. The input to the report is the analysis results, and the output is text information provided to the user. Specifically, this includes automatically generated comments and recommended actions.
[0594] Step 5:
[0595] The device receives the generated information report and analyzes the user's emotions. Using OpenCV and the Face API, it analyzes the user's facial expressions and voice tone to identify their emotional state. The input is data from the device's camera and microphone, and the output is parameters indicating the user's emotional state.
[0596] Step 6:
[0597] The device adjusts and presents an information report based on the recognized user's emotions. For example, if anxiety is detected, additional explanations or reassuring information are added. The input is the user's emotion parameters, and the output is the adjusted report. At this stage, the text is restructured and displayed on the screen.
[0598] Step 7:
[0599] Users review information reports provided through their devices and adjust their actions as needed. Based on the information presented, users understand the situation on-site and decide on corrective actions. The input is the adjusted report, and the output is the user's decision to take action. Specific examples include implementing safety measures and improvement plans tailored to the situation.
[0600] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0601] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0602] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0603] [Fourth Embodiment]
[0604] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0605] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0606] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0607] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0608] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0609] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0610] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0611] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0612] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0613] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0614] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0615] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0616] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0617] This invention provides a system for efficiently determining the cause of death using medical image data from an image acquisition device. This system includes a mechanism for storing image data in cloud storage, performing preprocessing, and analyzing the data using an artificial intelligence model.
[0618] The server has the function of receiving medical image data transmitted from hospital CT and MRI scanners and saving it to cloud storage. The saved data is preprocessed, including noise reduction and resolution standardization, and is ready for analysis.
[0619] Next, the server inputs the pre-processed data into an artificial intelligence model. This model has been trained on numerous training datasets and is designed to identify congenital abnormalities and pathological findings. For example, it can detect abnormal artifacts around the heart and suggest the possibility of a myocardial infarction.
[0620] After the analysis is performed, the server generates a diagnostic report and sends the results to a terminal for the user to use. This report contains detailed cause-of-death information obtained without an autopsy and serves as valuable material for medical professionals to review.
[0621] Furthermore, users can improve diagnostic accuracy by comparing current data with past image data of the same patient. This enables a multifaceted analysis that includes the progression of chronic diseases and the effectiveness of medications.
[0622] Furthermore, the diagnostic analysis results will be stored in the cloud and used for future data analysis and other research purposes. This accumulated database will serve as a foundation to support improvements in diagnostic technology and information sharing across regions.
[0623] The system according to the present invention enables rapid and accurate determination of the cause of death while reducing the burden of autopsy.
[0624] The following describes the processing flow.
[0625] Step 1:
[0626] The server receives medical image data from the hospital's CT and MRI machines. This data is typically provided in DICOM format and organized by patient ID and imaging date.
[0627] Step 2:
[0628] The server stores the received medical image data in cloud storage. This allows for efficient management of large amounts of data and easy access to it later.
[0629] Step 3:
[0630] The server preprocesses the stored image data. By removing noise and maintaining a consistent resolution, it prepares the data for the AI model to analyze efficiently.
[0631] Step 4:
[0632] The server inputs pre-processed image data into an artificial intelligence model. This model is trained using a vast amount of historical data and is tuned to detect abnormal findings and disease features.
[0633] Step 5:
[0634] The server generates a diagnostic report based on the analysis results from the AI model. This report includes identified lesions and possible causes of death, providing information useful for medical decision-making.
[0635] Step 6:
[0636] The server sends the generated diagnostic report to the terminal. In this case, the terminal is a PC or mobile device used by a medical professional.
[0637] Step 7:
[0638] Users can view diagnostic reports on their devices and share information with other medical professionals as needed. They can also compare past medical image data with current data to further improve the accuracy of their diagnoses.
[0639] Step 8:
[0640] The server stores the analysis results in the cloud. This makes the data available for future data analysis and research, creating a database that contributes to the advancement of medicine.
[0641] (Example 1)
[0642] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0643] In modern medical practice, there is a demand for rapid and accurate pathological diagnoses, but conventional systems lack the means to improve analytical accuracy and safely deliver diagnostic results. Furthermore, mechanisms for improving diagnostic accuracy through effective comparison with past medical data, and efficient methods for data storage and utilization, are insufficient. Therefore, challenges remain in formulating treatment plans for patients and managing their long-term health.
[0644] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0645] In this invention, the server includes means for storing medical image data acquired from an image acquisition device in an information storage device, means for processing and standardizing the information on the stored medical image data, means for performing diagnostic analysis using a learning model with the pre-processed medical image data, means for generating a diagnostic document based on the diagnostic analysis results, and means for encrypting and transmitting the generated diagnostic document to provide it to the user securely. This improves the accuracy of the diagnosis and makes it possible to provide diagnostic results quickly and with peace of mind. Furthermore, the accuracy of the diagnosis can be further enhanced by comparing it with past data, and at the same time, it becomes possible to build a database that can be used for future data analysis and research.
[0646] "Image acquisition equipment" refers to devices used to generate medical image data, and examples include CT scanners and MRI scanners.
[0647] "Medical image data" refers to image information generated by image acquisition equipment and used for diagnosis and treatment.
[0648] An "information storage device" is a digital storage system for storing medical image data and analysis results, and is a part of cloud storage, among other things.
[0649] "Processing to prepare information" refers to applying pre-processing steps such as homogenization and noise reduction to image data to make it suitable for analysis.
[0650] A "learning model" is an artificial intelligence system used to analyze medical image data, possessing pattern recognition capabilities based on past training data.
[0651] A "diagnostic document" is a report generated based on the results of analyzing medical imaging data, and includes information for identifying and diagnosing abnormalities.
[0652] A "user terminal" refers to a device such as a computer or tablet used by medical professionals to receive and view diagnostic documents.
[0653] "Encryption" is a technology that protects data using specific algorithms to safeguard information from unauthorized access.
[0654] The system of the present invention is designed to perform image analysis efficiently and accurately in medical institutions. Specific embodiments are described below.
[0655] The server first receives medical image data transmitted from imaging devices, such as CT scanners and MRI machines. At this time, the server acquires the data in real time via the network and stores it in cloud storage, which is an information storage device. Examples of cloud storage used include Amazon Web Services (AWS) and Google Cloud Platform (GCP). The stored data is organized for reuse and backup purposes.
[0656] Next, the server processes the stored image data to prepare it for analysis. Specifically, it uses image processing libraries such as OpenCV to remove noise and equalize resolution, making the data suitable for analysis. Once this preprocessing is complete, the server inputs the data into a training model. The training model is built using AI frameworks such as TensorFlow and PyTorch and is trained on a large number of medical datasets. This model detects pathological findings and abnormalities with high accuracy.
[0657] Based on the analysis results, the server automatically generates a diagnostic document. This document includes identifying abnormalities and detailed information for diagnosis. The generated document is then encrypted and securely transmitted to the user's terminal. The user terminal, typically a computer or tablet, is used by medical professionals to open the diagnostic document and consider treatment plans for the patient based on its contents.
[0658] As a concrete example, if a cardiac CT image is input into a learning model and abnormal spots are detected, the model will reflect this as a possible myocardial infarction in the diagnostic document. This analysis result is compared with stored past data of the same patient to clarify the progression of the disease. An example of a prompt message to the generating AI model might be, "Based on the analysis results of the cardiac CT image, diagnose the possibility of myocardial infarction and create a comparison report with past data." Based on this prompt message, the generating AI model will perform an accurate analysis and report.
[0659] Thus, by using this system, medical institutions can perform rapid and accurate diagnoses non-invasively, significantly improving the quality of medical care.
[0660] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0661] Step 1:
[0662] The server receives medical image data from image acquisition equipment. This data is transferred from devices such as CT and MRI scanners and transmitted to the server via the network. The input is unprocessed image data. The server saves the received data to cloud storage. This is to create a backup while retaining the original image data. The output is medical image data securely stored in the data storage device.
[0663] Step 2:
[0664] The server preprocesses the stored medical image data. The input is image data obtained from cloud storage. The server uses image processing libraries such as OpenCV to perform noise reduction and resolution equalization. Specifically, it applies a noise filter and processes the image to improve its detail. This results in data of optimal quality for AI models. The output is preprocessed, high-quality image data.
[0665] Step 3:
[0666] The server inputs pre-processed image data into a learning model. The input is pre-processed image data. The server executes an AI model built using TensorFlow or PyTorch, and determines the presence or absence of anomalies based on the data. For example, it recognizes unusual patterns or shapes within the image and makes a diagnosis based on them. The output is analysis data including the anomaly detection results.
[0667] Step 4:
[0668] The server generates a diagnostic document based on the analysis results. The input is the analysis results obtained by the AI model. The server processes this data and creates a diagnostic report summarizing the results. The report includes details of the detected anomalies and recommendations for diagnosis. Natural language processing technology is used in this generation process. The output is a completed report as a diagnostic document.
[0669] Step 5:
[0670] The server encrypts the generated diagnostic document and sends it to the user's terminal. The input is the generated diagnostic document. The server encrypts the document to enhance security and sends it to the medical professional's terminal using an appropriate communication protocol. Specifically, it establishes a secure communication channel using SSL / TLS. The output is a diagnostic report that can be decrypted on the recipient medical professional's terminal.
[0671] Step 6:
[0672] The user, a medical professional, reviews the received diagnostic documents using a terminal. The input is the diagnostic report received on the terminal. The user reviews the report, assesses the patient's condition, and obtains guidance for determining the necessary treatment plan. Specifically, they open the report using dedicated viewer software and examine the details. The output at this stage is the specific diagnostic result that will be incorporated into the treatment plan.
[0673] (Application Example 1)
[0674] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0675] In the manufacturing industry, product quality control is extremely important, but traditional inspection methods often fail to detect defects in real time, leading to decreased production efficiency and quality. Furthermore, human visual inspection has limitations and is susceptible to human error. Under these circumstances, there is a need to automate the detection and reporting of defects on the production line to improve productivity.
[0676] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0677] In this invention, the server includes means for storing visualization data acquired from an image acquisition mechanism in a remote storage device, means for pre-processing and standardizing the stored visualization data, means for performing analysis using a machine learning model with the pre-processed visualization data, means for detecting defective products on the manufacturing line using a device that photographs the components of the product, and means for detecting and notifying abnormalities in real time. This enables rapid automatic detection and notification of defective products on the manufacturing line, thereby improving production efficiency and quality.
[0678] An "image acquisition mechanism" is a device used to capture or acquire images of a product or other important visualization data.
[0679] "Visualized data" refers to images and other data that can be visually analyzed, and is used for product inspection and evaluation.
[0680] A "remote storage device" is a device used to store and manage data via a network, such as cloud storage.
[0681] "Preprocessing" refers to initial data processing, such as standardization and noise reduction, performed on acquired data to improve the accuracy of the analysis.
[0682] A "machine learning model" is an algorithm that learns specific patterns and features from large amounts of data to classify and predict data.
[0683] A "device for photographing product components" is a device that photographs the components of each product on the manufacturing line and uses those images for inspection.
[0684] "Means for detecting defective products" refers to a method or apparatus for determining and identifying whether a product on a manufacturing line is defective based on collected data.
[0685] "Means for detecting and notifying anomalies in real time" refers to a method or device that detects the results of an analysis almost immediately and immediately notifies the relevant parties if an anomaly is found.
[0686] The system realizing this invention consists of an image acquisition mechanism, a remote storage device, a preprocessing module, and a program including a machine learning model, all positioned on a manufacturing line. The server first acquires visualization data of the product using the image acquisition mechanism. This data is transferred in real time to a remote storage device, such as cloud storage, for rapid access.
[0687] Next, the server uses a preprocessing module to perform initial processing on this data, such as denoising and standardization. By using OpenCV or equivalent image processing software for this process, the data is prepared for analysis.
[0688] The preprocessed data is input into a machine learning model such as TensorFlow. This model is programmed to detect product anomalies based on a large amount of training data. Specifically, it learns the characteristics of defective products and determines whether or not a product's components are defective.
[0689] For example, if a manufactured smartphone case is cracked, a machine learning model can detect the anomaly in real time and immediately notify the product line manager. This real-time notification uses a web application framework such as Flask, and relevant information is sent to the manager's terminal each time an anomaly is detected.
[0690] An example of a prompt message used when a generative AI model is employed is as follows: "Using image data of the product captured by the smart camera, please use the AI model to detect surface anomalies and defects in real time."
[0691] In this way, the server enables the immediate identification and reporting of defective products during the manufacturing process, thereby improving the efficiency of quality control.
[0692] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0693] Step 1:
[0694] The server uses an image acquisition mechanism to obtain visualization data from products on the manufacturing line. The input is the manufactured product, and the output is high-resolution image data of the product. This image data is acquired in real time and used for rapid analysis.
[0695] Step 2:
[0696] The server transmits the acquired image data to a remote storage device for storage. The input is the high-resolution image data obtained in step 1, and the output is the image file stored on cloud storage. This storage operation makes the data accessible at any time.
[0697] Step 3:
[0698] The server uses a preprocessing module to perform standardization and denoising on the stored image data. The input is the original image on cloud storage, and the output is standardized, clean image data. This preprocessing is performed using OpenCV.
[0699] Step 4:
[0700] Using pre-processed image data, the server performs anomaly detection using a machine learning model. The input is the clean image data from step 3, and the output is the location and type of anomaly if detected. The trained model performs this using TensorFlow.
[0701] Step 5:
[0702] The server instantly notifies the administrator terminal based on the detected anomaly. The input is the anomaly information from step 4, and the output is a notification message detailing the anomaly. The notification is sent via a web application framework such as Flask.
[0703] This series of processes enables high-speed detection of defective products on the manufacturing line and appropriate reactions.
[0704] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0705] This invention combines an emotion engine with a system that analyzes the cause of death using medical image data to provide diagnostic information that takes into account the user's emotional state. In addition to acquiring, storing, preprocessing, analyzing, and generating reports on medical image data, this system includes an emotion engine that recognizes the user's emotions.
[0706] The server first receives medical image data from the image acquisition device and saves it to cloud storage. The saved images undergo preprocessing, including noise reduction and resolution adjustment, to prepare them for analysis by the AI model.
[0707] Next, the server inputs the pre-processed images into a deep learning model to extract features related to the cause of death. This process allows the model to determine the location of abnormalities and the likelihood of disease. Based on the analysis results, a detailed diagnostic report is generated.
[0708] The generated diagnostic report is provided to the user, and this is where the device utilizes an emotion engine. This emotion engine senses the user's facial expressions and tone of voice as they review the report and recognizes their emotions. Based on the recognized emotions, adjustments are made to how the report is presented and its content. For example, if the user is feeling anxious, additional explanations or supplementary information may be provided.
[0709] Users can view the adjusted reports through their devices and share information with other experts as needed. Furthermore, the emotional data recognized by the emotion engine is stored in cloud storage and used to improve future services.
[0710] This system goes beyond mere diagnosis, providing a new interface to improve the quality of medical services. It enables flexible responses tailored to the user's emotional state, improving the patient experience and delivering more personalized medical information.
[0711] The following describes the processing flow.
[0712] Step 1:
[0713] The server receives medical image data from the image acquisition device. This data is stored in DICOM format and organized and saved in cloud storage.
[0714] Step 2:
[0715] The server preprocesses the stored image data. Specifically, it removes noise and standardizes the image resolution. This process creates a uniform dataset suitable for analysis.
[0716] Step 3:
[0717] The server inputs pre-processed image data into a deep learning model. Based on knowledge learned from past training data, the model identifies abnormalities and lesions within the images. The resulting diagnostic results are then analyzed, and a report is generated.
[0718] Step 4:
[0719] The device displays the diagnostic report received from the server and simultaneously activates an emotion engine to recognize the user's emotions. The engine uses the camera and microphone to analyze the user's facial expressions and tone of voice to identify their emotions.
[0720] Step 5:
[0721] The device adjusts the report display based on the emotions it recognizes. For example, if the emotion engine detects user anxiety, the report will be adjusted to include additional explanations and reassuring information.
[0722] Step 6:
[0723] Users can review the adjusted report and, if necessary, share their diagnosis with other healthcare professionals. Additionally, emotional data recorded by the emotion engine is stored in the cloud for analysis.
[0724] Step 7:
[0725] The server will use the stored emotional data to improve future medical services. This will enable the provision of medical information more tailored to individual patients, aiming to improve patient satisfaction.
[0726] (Example 2)
[0727] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0728] Current medical imaging diagnostic systems have shortcomings, such as insufficient accuracy in diagnostic results and inadequate information provision tailored to the user's emotional state. Specifically, there is a need to improve the user experience because data comparison to improve the accuracy of medical image analysis and appropriate information provision tailored to the emotions of users receiving diagnostic results are not being carried out.
[0729] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0730] In this invention, the server includes means for recording medical image information acquired from an image acquisition means in a remote data storage means, means for performing data preprocessing on the recorded medical image information to standardize it, and means for performing diagnostic analysis using a machine learning model with the preprocessed medical image information. This makes it possible to recognize the user's emotional state and provide a diagnostic report that is adjusted according to that state, as well as to improve the accuracy of the analysis through comparison with past data.
[0731] "Image acquisition means" refers to devices and methods for collecting image data from medical imaging equipment.
[0732] "Remote data storage means" refers to servers or cloud storage that can store acquired data via the internet.
[0733] "Data preprocessing" refers to processes, including noise reduction and resolution adjustment, performed to improve the quality of acquired image data.
[0734] A "machine learning model" refers to an algorithm or system that extracts features from observed data and makes judgments or predictions.
[0735] "User equipment" refers to computers or mobile devices used by users to receive and view diagnostic reports.
[0736] "Means of recognizing emotional states" refers to technologies and devices that analyze a user's facial expressions and voice to determine their emotions.
[0737] A "diagnostic report" refers to a document or data that summarizes the results of medical image analysis and provides them to the user.
[0738] The system in this invention processes medical image information and provides a diagnostic report that takes into account the user's emotional state based on the analysis results. The embodiments of this system are described in detail below.
[0739] The server receives image data from the medical imaging device and records it in a remote data storage system. The received data is securely stored, for example, using a cloud-based storage service. Next, noise reduction is performed using the OpenCV library, and data preprocessing such as resolution adjustment is performed using the Python Imaging Library. This enables data analysis in a standardized format.
[0740] The server feeds pre-processed medical image data into machine learning models trained with TensorFlow or PyTorch to perform diagnostic analysis. The models utilize technologies such as Convolutional Neural Networks (CNNs) to identify abnormalities in the images and determine the likelihood of disease.
[0741] Based on the generated diagnostic analysis results, the server prepares to create a diagnostic report and send it to the user's device. Specifically, it uses visualization tools such as Matplotlib and Seaborn to graphically represent the diagnostic results and uses natural language generation technology to create a document that is easy for the user to understand.
[0742] When presenting a diagnostic report to the user, the device analyzes the user's facial expressions and voice using emotion recognition technology. This allows it to determine in real time whether the user is experiencing anxiety or doubts, and adjust the content and presentation method of the report accordingly. For emotion recognition, for example, Microsoft Azure APIs are used.
[0743] As a concrete example, consider a scenario where a user undergoes an MRI scan of their knee, and this system analyzes the results. The server performs noise reduction on the received image and begins analyzing for abnormalities in the knee area. As a result, a diagnostic report is generated stating, "No abnormalities were found in the knee." If the user is detected to be feeling anxious while viewing the report, the terminal displays a supplementary message saying, "There's no need to worry."
[0744] An example of a prompt message might be: "Analyze the MRI images of the knee, compile a diagnostic report indicating whether or not there are any abnormalities, and add information to alleviate the user's anxiety."
[0745] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0746] Step 1:
[0747] The server receives image data from a medical imaging device using an image acquisition method. The received data is received in a standardized format such as DICOM format. Since the input image data cannot be used directly for analysis, it is first recorded in a remote data storage system. During storage, the integrity of the data is checked, and if it is normal, the process proceeds to the next step.
[0748] Step 2:
[0749] The server performs data preprocessing on stored medical image information, including noise reduction and resolution adjustment. It takes the received image data as input and applies an OpenCV Gaussian filter to remove noise. Next, it adjusts the resolution using the Python Imaging Library and outputs standardized, high-quality image data. This process improves the accuracy of subsequent analyses.
[0750] Step 3:
[0751] The server feeds pre-processed image data into a machine learning model. This model is a generative AI model trained using TensorFlow, PyTorch, etc. The input is a pre-processed image, and the output is the anomaly detection result in the image. The model uses a Convolutional Neural Network (CNN) to extract features of the anomaly and make them ready for diagnosis.
[0752] Step 4:
[0753] The server generates a diagnostic report based on the results of a machine learning model's diagnostic analysis. The input is the analysis data, and the output is a diagnostic report in a user-friendly format. It uses tools like Matplotlib and Seaborn to graphically represent the data, and natural language generation technology to create the report content.
[0754] Step 5:
[0755] The device recognizes the user's emotions when providing the generated diagnostic report to the user. When the report is displayed, the built-in camera and microphone receive the user's facial expressions and voice as input, and the emotion engine analyzes their state. Using the Microsoft Azure API, the device obtains user emotion data (e.g., anxiety, reassurance) as output.
[0756] Step 6:
[0757] The device adjusts the content and presentation of the diagnostic report based on data obtained from the emotion engine. If the user is feeling anxious based on the emotional data input, a reassuring message such as "There's no need to worry" is added and displayed. This enables more personalized information delivery and improves the user experience.
[0758] Step 7:
[0759] Users review the diagnostic report provided through their device. They carefully examine the output report and, if necessary, request supplementary information or decide to share the information with other medical professionals. Throughout this process, the information is securely managed, allowing users to continue using the service with peace of mind.
[0760] (Application Example 2)
[0761] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0762] Traditional security systems simply record and play back video and audio data, but they have difficulty detecting abnormal situations or changes in people's emotional states in real time and prompting appropriate responses. Furthermore, they fail to provide reports that take user emotions into account, resulting in missed opportunities for effective countermeasures. There is a need to address these issues and provide a higher level of management and security.
[0763] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0764] In this invention, the server includes means for storing image data acquired from an image acquisition unit in network storage, means for preprocessing and standardizing the stored image data, and means for performing diagnostic analysis using a machine learning model with the preprocessed image data. This makes it possible to detect abnormal situations and changes in emotional states at the site in real time and provide information reports to user terminals. Furthermore, by analyzing the user's emotions and adjusting the report, it is possible to encourage appropriate responses.
[0765] The "image acquisition unit" refers to the devices and sensors used to collect video data within the system.
[0766] "Network storage" refers to an online storage system for storing and managing data via the internet.
[0767] "Preprocessing" refers to the preparatory steps for improving the accuracy of analysis by standardizing and denoising the acquired data.
[0768] A "machine learning model" refers to an algorithm that learns patterns from data and performs predictions and classifications.
[0769] "Diagnostic analysis" is the process of evaluating and determining specific events or conditions based on data.
[0770] An "information report" is a document or data that organizes analysis results and related information and provides it to users.
[0771] "User terminal" refers to a device used to display and operate information, and includes PCs, smartphones, and other similar devices.
[0772] "Emotional analysis" is the process of identifying and evaluating a person's emotional state based on video and audio data.
[0773] An "abnormal situation" refers to an unusual or irregular phenomenon or circumstance that is considered to require attention or action.
[0774] The system that realizes this invention consists of multiple servers, terminals, and users. First, the servers save the video data obtained from the image acquisition unit to network storage. The saved video data is preprocessed to remove noise and adjust the resolution. This preprocessing standardizes the data and makes it suitable for analysis. The preprocessed data is input into a machine learning model to detect abnormal situations and changes in emotional state in the video.
[0775] This analysis utilizes deep learning algorithms such as YOLO (You Only Look Once) and speech analysis using TensorFlow. The server generates an information report based on the analysis results and provides it to the user's device. This report also includes the analyzed emotional state. The device performs emotion analysis and flexibly adjusts the report content by sensing the user's facial expressions and voice nuances. OpenCV and Face API are used for emotion analysis, which enables the provision of additional information tailored to the user's emotions.
[0776] As a concrete example, in large commercial facilities, surveillance systems may use multiple cameras and microphones to perform real-time analysis and detect increases in anxiety or tension within the premises. When detected, an alert is sent to security staff, allowing for a swift response.
[0777] In this way, this system improves on-site safety and enables efficient management. An example of a prompt message used for this purpose is: "Provide real-time sentiment analysis data within the security area to detect situations where anxiety or tension is rising early. If an anomaly is detected, provide specific action guidelines as an alert."
[0778] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0779] Step 1:
[0780] The server receives video data from the image acquisition unit and stores it in network storage. Specifically, it collects video streams transmitted from cameras and surveillance sensors and saves each frame as digital data. The input is video data, and the output is the saved digital data.
[0781] Step 2:
[0782] The server performs preprocessing on the stored video data. This includes noise reduction and resolution adjustment, and aims to standardize the data. The input for preprocessing is the stored raw data, and the output is standardized data. Specifically, it applies filters to reduce noise and resizes frames.
[0783] Step 3:
[0784] The server inputs pre-processed data into a machine learning model, performs analysis, and detects abnormal situations and changes in emotional states. At this stage, algorithms such as YOLO are used to analyze the position and movement of people and objects. The input is pre-processed data, and the output is the detected features.
[0785] Step 4:
[0786] The server generates an information report based on the analysis results. The generated report includes identification of anomalies and an evaluation of detected emotional states. The input to the report is the analysis results, and the output is text information provided to the user. Specifically, this includes automatically generated comments and recommended actions.
[0787] Step 5:
[0788] The device receives the generated information report and analyzes the user's emotions. Using OpenCV and the Face API, it analyzes the user's facial expressions and voice tone to identify their emotional state. The input is data from the device's camera and microphone, and the output is parameters indicating the user's emotional state.
[0789] Step 6:
[0790] The device adjusts and presents an information report based on the recognized user's emotions. For example, if anxiety is detected, additional explanations or reassuring information are added. The input is the user's emotion parameters, and the output is the adjusted report. At this stage, the text is restructured and displayed on the screen.
[0791] Step 7:
[0792] Users review information reports provided through their devices and adjust their actions as needed. Based on the information presented, users understand the situation on-site and decide on corrective actions. The input is the adjusted report, and the output is the user's decision to take action. Specific examples include implementing safety measures and improvement plans tailored to the situation.
[0793] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0794] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0795] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0796] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0797] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0798] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0799] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0800] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0801] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0802] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0803] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0804] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0805] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0806] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0807] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0808] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0809] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0810] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0811] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0812] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0813] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0814] The following is further disclosed regarding the embodiments described above.
[0815] (Claim 1)
[0816] A means of saving medical image data acquired from an image acquisition device to cloud storage,
[0817] A means of preprocessing and standardizing stored medical image data,
[0818] A means for performing diagnostic analysis using an artificial intelligence model with pre-processed medical image data,
[0819] A means of generating a diagnostic report based on the diagnostic analysis results,
[0820] A means of providing the generated diagnostic report to the user terminal,
[0821] A system that includes this.
[0822] (Claim 2)
[0823] The system according to claim 1, comprising means for comparing previously stored medical image data with current image data to improve the accuracy of diagnostic analysis.
[0824] (Claim 3)
[0825] The system according to claim 1, comprising means for saving diagnostic analysis results on the cloud and making them available for future data analysis and research.
[0826] "Example 1"
[0827] (Claim 1)
[0828] A means for storing medical image data acquired from an image acquisition device in an information storage device,
[0829] A means of processing and standardizing information in stored medical image data,
[0830] A means for performing diagnostic analysis using a learning model with pre-processed medical image data,
[0831] A means for generating a diagnostic document based on the diagnostic analysis results,
[0832] A means of providing the generated diagnostic document to the user's terminal,
[0833] A means of encrypting and transmitting diagnostic documents and providing them securely to users,
[0834] A system that includes this.
[0835] (Claim 2)
[0836] The system according to claim 1, comprising means for comparing previously stored medical image data with current image data to improve the accuracy of diagnostic analysis.
[0837] (Claim 3)
[0838] The system according to claim 1, comprising means for storing diagnostic analysis results in an information storage device and making them available for future information analysis and research.
[0839] "Application Example 1"
[0840] (Claim 1)
[0841] A means for saving visualization data acquired from an image acquisition mechanism to a remote storage device,
[0842] A means of preprocessing and standardizing stored visualization data,
[0843] A means of performing analysis using a machine learning model with pre-processed visualization data,
[0844] A means of generating a report based on the analysis results,
[0845] A means of providing the generated report to the user's terminal,
[0846] A means for detecting defective products on a manufacturing line using a device that photographs the components of a product,
[0847] A means of detecting and notifying anomalies in real time,
[0848] A system that includes this.
[0849] (Claim 2)
[0850] The system according to claim 1, comprising means for comparing previously stored visualization data with current data to improve the accuracy of the analysis.
[0851] (Claim 3)
[0852] The system according to claim 1, further comprising means for saving analysis results on a remote storage device and making them available for future data analysis and research.
[0853] "Example 2 of combining an emotion engine"
[0854] (Claim 1)
[0855] A means for recording medical image information acquired from an image acquisition means to a remote data storage means,
[0856] A means of performing data preprocessing on recorded medical image information and standardizing it,
[0857] A means for performing diagnostic analysis using a machine learning model with pre-processed medical image information,
[0858] A means of generating a diagnostic report based on the diagnostic analysis results,
[0859] Means for providing the generated diagnostic report to the user's device,
[0860] A means of recognizing the user's emotional state and adjusting the content of the diagnostic report,
[0861] A system that includes this.
[0862] (Claim 2)
[0863] The system according to claim 1, comprising means for comparing previously recorded medical image information with current image information to improve the accuracy of diagnostic analysis.
[0864] (Claim 3)
[0865] The system according to claim 1, further comprising means for recording diagnostic analysis results in a remote data storage means and making them available for future data analysis and research.
[0866] "Application example 2 when combining with an emotional engine"
[0867] (Claim 1)
[0868] A means for saving image data acquired from the image acquisition unit to network storage,
[0869] A means of preprocessing and standardizing stored image data,
[0870] A means of performing diagnostic analysis using a machine learning model with preprocessed image data,
[0871] A means of generating an information report based on the diagnostic analysis results,
[0872] A means of providing the generated information report to the user's terminal,
[0873] A means of analyzing user emotions and adjusting information reports,
[0874] A means for analyzing data acquired from multiple recording and sound devices to detect and notify abnormal situations and emotional states,
[0875] A system that includes this.
[0876] (Claim 2)
[0877] The system according to claim 1, comprising means for comparing previously stored image data with current image data to improve the accuracy of diagnostic analysis.
[0878] (Claim 3)
[0879] The system according to claim 1, comprising means for storing diagnostic analysis results on a network and making them available for future data analysis and research. [Explanation of Symbols]
[0880] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of saving medical image data acquired from an image acquisition device to cloud storage, A means of preprocessing and standardizing stored medical image data, A means for performing diagnostic analysis using an artificial intelligence model with pre-processed medical image data, A means of generating a diagnostic report based on the diagnostic analysis results, A means of providing the generated diagnostic report to the user terminal, A system that includes this.
2. The system according to claim 1, comprising means for comparing previously stored medical image data with current image data to improve the accuracy of diagnostic analysis.
3. The system according to claim 1, comprising means for saving diagnostic analysis results on the cloud and making them available for future data analysis and research.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A