system

A system using image and text data analysis by a generative AI model provides quick and accurate medical diagnoses and referrals, addressing the challenge of accessing medical services for those without specialized knowledge.

JP2026035194APending Publication Date: 2026-03-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024138037
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2026-03-04

Smart Images

  • Figure 2026035194000001_ABST
    Figure 2026035194000001_ABST
Patent Text Reader

Abstract

To provide a system that improves access to medical services by enabling even users without specialized medical knowledge to receive early predictions of disease conditions and referrals to appropriate medical institutions. [Solution] A system including: a means for a user to input image data and text data relating to a medical condition or injury; a means for a terminal to package the image data and text data into a specified format and send it to a server; a means for the server to analyze the image data and infer the medical condition; a means for the server to analyze the text data and infer symptoms; a means for a generative AI model to integrate the analysis results and make an inferred diagnosis; a means for the server to select an appropriate medical institution based on the inferred diagnosis; a means for the server to send information about the inferred diagnosis and medical institution to the terminal; and a means for the terminal to display the information to the user.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, an increasing number of people are unable to receive necessary medical services due to busy schedules, geographical factors, or not knowing how to find an appropriate medical institution. In particular, for acute illnesses and injuries, early diagnosis and prompt treatment are required, but users without specialized medical knowledge find it difficult to self-diagnose and access appropriate medical institutions. Furthermore, with the advancement of a super-aging society, problems with diagnosis and appointments are arising as medical institutions exceed their capacity. Given this background, there is a demand for a fast and accurate non-face-to-face diagnostic support system. [Means for solving the problem]

[0005] The present invention is a system in which a user inputs image and text data related to a medical condition or injury, which the terminal packages and sends to a server. The server analyzes the image data to estimate the medical condition and the text data to estimate symptoms. These analysis results are integrated using a generative AI model to make a probable diagnosis. Furthermore, based on the estimated diagnosis, the server selects an appropriate medical institution in conjunction with the user's location information and sends the diagnosis results and information about the medical institution to the terminal. The terminal displays the received information to the user, helping them quickly access an appropriate medical institution. This system enables even users without specialized medical knowledge to receive early estimates of the medical condition and be referred to an appropriate medical institution, thereby improving access to medical services.

[0006] A "user" is an individual who uses the system to input information about a medical condition or injury and receive diagnosis results and information about a medical institution.

[0007] A "terminal" is a device that sends image and text data of medical conditions and injuries entered by the user to a server and displays the received diagnosis results and information about medical institutions. This applies to smartphones and computers.

[0008] The "server" is a computer system that receives image data and text data sent by the user, analyzes them, and makes a presumptive diagnosis and selects an appropriate medical institution.

[0009] "Image data" refers to data that a user visually records of the state of illness or injury, and is generally a photo file (e.g., JPEG, PNG).

[0010] "Text data" is character information that a user inputs to explain symptoms related to a medical condition or injury.

[0011] A "specified format" is a standard data format used to package, transmit, or receive data, and examples include JSON and XML.

[0012] A "generative AI model" is an artificial intelligence computational model that analyzes input image data and text data, integrates this information, and makes a presumptive diagnosis.

[0013] A "presumed diagnosis" is the assessment result of a disease or health condition predicted by a generative AI model.

[0014] A "medical institution" is a specialized facility for examining and treating illnesses and injuries, and includes hospitals and clinics.

[0015] "Location information" is information that indicates the user's current location, and is data that is used to select an appropriate medical institution within a specific area. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention is a system that allows users to input image and text data related to a medical condition or injury and receive a referral to an appropriate medical institution even if they do not have specialized medical knowledge. Specific embodiments of this system will be described below.

[0038] User input of images and text

[0039] Users can use a smartphone or computer to take pictures of their illness or injury and save the image data on the device. They can then enter their symptoms and thoughts about the condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text, "My hand is swollen and painful."

[0040] Sending data by the device

[0041] The device packages the image data and text data entered by the user into a specified format (for example, JSON or XML).The device then sends this data to a server via the Internet.At this time, the data is encrypted to ensure security.

[0042] Receipt and analysis of data by the server

[0043] The server receives and decodes the image and text data sent from the device. It then uses a generative AI model to analyze the received image data. Specifically, it uses algorithms such as convolutional neural networks to extract image features and estimate the condition. At the same time, the server analyzes the text data using a natural language processing (NLP) model to estimate symptoms.

[0044] Predictive diagnosis using generative AI models

[0045] The generative AI model on the server integrates the results of image analysis and text analysis to generate a probable diagnosis. This diagnosis includes a probable pathology and its reliability. For example, if image analysis detects a bone abnormality and text analysis identifies the symptom "swelling and pain," the probable diagnosis is "high probability of fracture."

[0046] Server selection of appropriate medical institution

[0047] Based on the estimated diagnosis, the server refers to the user's location information and selects the most appropriate medical institution. At this time, the server retrieves a list of affiliated medical institutions from the database and lists medical institutions close to the user's current location. The information on the selected medical institution includes location, contact information, available hours for consultation, etc.

[0048] Sending results from the server to the device

[0049] The server then sends the estimated diagnosis and information on the recommended medical institution to the device. The data is then re-encrypted and transferred to the device in a secure manner.

[0050] Displaying results on a terminal

[0051] The device decodes the received data and displays it in a user-friendly format. The user can then check the displayed estimated diagnosis and recommended medical institution information and make arrangements for a medical examination if necessary. For example, if there is a high possibility of a fracture, the user can call the orthopedic hospital contact number displayed on the device to make an appointment.

[0052] Specific examples

[0053] For example, if a user injures their hand, the following process occurs:

[0054] 1. A user takes a photo of a hand injury with their smartphone and types, "My hand is swollen and painful."

[0055] 2. The device packages the image data and text data in JSON format and sends it to the server.

[0056] 3. The server analyzes the image using an AI model and the text data using an NLP model.

[0057] 4. The server integrates the analysis results and generates a diagnosis of "high probability of fracture."

[0058] 5. The server selects a nearby orthopedic hospital based on the user's location information.

[0059] 6. The server sends the diagnosis results and hospital information to the terminal.

[0060] 7. The terminal displays the received information to the user, who then contacts the orthopedic hospital to arrange for a consultation.

[0061] The above is a specific embodiment of the present invention. This system makes it possible to quickly and accurately estimate the condition of a patient, even without specialized medical knowledge, and to assist in accessing an appropriate medical institution.

[0062] The processing flow will be explained below.

[0063] Step 1: The user takes an image of their condition or injury using a smartphone or computer, saves the image data on their device, and enters a description of their symptoms or condition into a text input form.

[0064] Step 2: The terminal converts the image data and text data entered by the user into a specified format (e.g., JSON or XML) and packages it.

[0065] Step 3: The device encrypts the packaged data and sends it to a server via the Internet.

[0066] Step 4: The server decodes the received image data and text data and prepares each for analysis.

[0067] Step 5: The server analyzes the image data using an image analysis module. Specifically, it uses a convolutional neural network (CNN) to extract image features and estimate the disease state.

[0068] Step 6: The server analyzes the text data using a natural language processing (NLP) module, extracting symptom-related keywords and context from the text and obtaining the analysis results.

[0069] Step 7: The generative AI model on the server integrates the results of image analysis and text analysis to make a probable diagnosis. For example, if image analysis indicates a suspected fracture and text analysis indicates pain or swelling, the model will generate a probable diagnosis of "high probability of fracture."

[0070] Step 8: Based on the estimated diagnosis, the server references the user's current location information and searches the database for the most suitable medical institution. Information on the target medical institutions is then listed.

[0071] Step 9: The server packages the estimated diagnosis and the information on the selected medical institution, encrypts it again, and transmits it to the terminal.

[0072] Step 10: The terminal decrypts the received data and displays it to the user, including the estimated diagnosis and detailed information (such as location, contact information, and consultation hours) of the recommended medical institution.

[0073] Step 11: The user checks the information about the medical institution displayed on the terminal and contacts the appropriate hospital to arrange for a medical examination.

[0074] Example 1

[0075] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0076] In today's world, it is difficult for users to quickly and accurately communicate information about their illness or injury to specialized medical institutions. It is also difficult for users without medical knowledge to select an appropriate medical institution. As a result, it can take a lot of time and effort to receive an appropriate diagnosis and treatment. To solve these problems, a system is needed that can effectively collect information about a user's illness or injury, analyze it quickly and accurately, and select an appropriate medical institution.

[0077] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0078] In this invention, the server includes: means for a user to input image data and text data related to a medical condition or injury; means for a terminal to package the image data and text data in a specified format and transmit the packaged image data and text data to the server; means for the terminal to encrypt the data; means for the server to receive and decrypt the image data and text data; means for the server to analyze the image data, extract features, and infer a medical condition; means for the server to analyze the text data, extract and infer symptoms; means for a generative AI model to integrate the analysis results and make a probable diagnosis; means for the server to select an appropriate medical institution based on the probable diagnosis; means for the server to transmit information about the probable diagnosis and medical facilities to the terminal; and means for the terminal to display the information to the user. This enables users to quickly and accurately understand their own medical condition or injury and to promptly visit an appropriate medical institution, even without specialized medical knowledge.

[0079] "Image data" is digital information that a user visually records regarding a medical condition or injury.

[0080] "Text data" is digital information in which a user records symptoms and thoughts about a medical condition or injury in written form.

[0081] A "terminal" is a device that a user uses to input data and send it to a server, and includes a smartphone or computer.

[0082] A "server" is a central processing device that receives data sent from a terminal, analyzes it, and returns the necessary information.

[0083] A "specified format" is a format in which data is arranged according to certain rules, and includes, for example, JSON and XML.

[0084] "Encryption" is a technique that uses a specific algorithm to convert data in order to transmit it securely.

[0085] "Decryption" is the technique of returning encrypted data to its original form.

[0086] "Features" are important attributes or patterns extracted from image data and are used to estimate the condition of a disease.

[0087] A "natural language processing (NLP) model" is an algorithm or method for analyzing text data and understanding human language.

[0088] A "generative AI model" is an artificial intelligence algorithm that integrates the results of image analysis and text analysis to make a presumptive diagnosis.

[0089] A "presumptive diagnosis" is the result of predicting the condition of a disease or injury based on analyzed data.

[0090] A "medical institution" is a facility where a user can receive medical examinations and treatment, and includes hospitals and clinics.

[0091] "Information" refers to data such as the diagnosis results and contact details and locations of medical institutions required by the user.

[0092] The present invention provides a system that allows a user to input image data and text data relating to a medical condition or injury, and easily refer the user to an appropriate medical institution. A specific embodiment of this system will be described below.

[0093] User data entry

[0094] Users use devices such as smartphones or computers to take pictures of their medical condition or injury and save them on the device. Next, they enter text data about their condition (e.g., "My hand is swollen and painful") into the device. At this stage, users do not need specialized medical knowledge and can enter data intuitively.

[0095] Device prepares to send data

[0096] The device packages the image and text data entered by the user into a specified format (e.g., JSON or XML), then encrypts the packaged data and transmits it securely to a server over the Internet.

[0097] Receipt and analysis of data by the server

[0098] The server receives the encrypted data sent from the device. After receiving the data, the server decrypts it and extracts the image data and text data. The server then analyzes the image data using a convolutional neural network (CNN) to extract image features. The server also analyzes the text data using a natural language processing (NLP) model to extract the described symptoms.

[0099] Predictive diagnosis using generative AI models

[0100] The generative AI model deployed on the server integrates the results of image analysis and text analysis. This model then uses the analysis results to generate a probable diagnosis, which includes a probable medical condition and its reliability. For example, by combining an image of a hand injury with the text data "your hand is swollen and painful," the model generates a probable diagnosis of "high probability of fracture."

[0101] Selection of medical institutions by the server

[0102] The server then references the user's location information based on the estimated diagnosis and selects an appropriate medical institution from a database. The information on the selected medical institution includes the location, contact information, and available consultation hours. The server then organizes this information, re-encrypts the data, and sends it to the device.

[0103] Displaying results on a terminal

[0104] The device receives and decrypts the encrypted data sent from the server. The decrypted data is displayed in a user-friendly format, allowing the user to refer to it and contact the appropriate medical institution and arrange for a medical examination. For example, it may display information such as "There is a high possibility of a fracture. The recommended medical institution is XX Orthopedic Hospital, contact number: XX-XX-XX."

[0105] Examples of concrete examples and prompts

[0106] As a concrete example, if a user injures their hand, the process is as follows:

[0107] 1. A user takes a photo of a hand injury with their smartphone and types, "My hand is swollen and painful."

[0108] 2. The device packages the image data and text data in JSON format and sends it to the server.

[0109] 3. The server analyzes the image using an AI model and the text data using an NLP model.

[0110] 4. The server integrates the analysis results and generates a diagnosis of "high probability of fracture."

[0111] 5. The server selects a nearby orthopedic hospital based on the user's location information.

[0112] 6. The server sends the diagnosis results and hospital information to the terminal.

[0113] 7. The terminal displays the received information to the user, who then contacts the orthopedic hospital to arrange for a consultation.

[0114] This system allows even users without specialized medical knowledge to quickly and accurately understand the patient's condition and refer them to the appropriate medical institution.

[0115] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0116] Step 1:

[0117] The user uses a smartphone or computer to take pictures of their condition or injury and save them on the device, then enters text about their condition (e.g., "My hand is swollen and painful") into the device.

[0118] Input: Image data and text data related to medical conditions and injuries

[0119] Output: Image data and text data stored on the device

[0120] Step 2:

[0121] The terminal acquires image data and text data input by the user.

[0122] Input: Image data and text data stored on the device

[0123] Output: Acquired image data and text data

[0124] Step 3:

[0125] The image data and text data acquired by the device are packaged in a specified format such as JSON.

[0126] Input: Acquired image data and text data

[0127] Output: Data packaged in the specified format

[0128] Step 4:

[0129] The device encrypts the packaged data and prepares it in a highly secure state.

[0130] Input: Data packaged in a specified format

[0131] Output: Encrypted data

[0132] Step 5:

[0133] The terminal transmits the encrypted data to a server via the Internet.

[0134] Input: Encrypted data

[0135] Output: Data sent to the server

[0136] Step 6:

[0137] The server receives the encrypted data sent from the terminal.

[0138] Input: Data sent over the internet

[0139] Output: Encrypted data received by the server

[0140] Step 7:

[0141] The server decrypts the received encrypted data and extracts the image data and text data.

[0142] Input: Received encrypted data

[0143] Output: Decoded image data, text data

[0144] Step 8:

[0145] The server analyzes the image data using a convolutional neural network (CNN) and extracts features.

[0146] Input: Decoded image data

[0147] Output: Image features

[0148] Step 9:

[0149] The server analyzes the text data using a natural language processing (NLP) model to extract symptoms.

[0150] Input: Decrypted text data

[0151] Output: Extracted symptoms

[0152] Step 10:

[0153] The server uses the generated AI model to integrate the results of image analysis and text analysis and make a presumptive diagnosis.

[0154] Input: Image features, extracted symptoms

[0155] Output: Estimated diagnosis result (condition and its reliability)

[0156] Step 11:

[0157] Based on the estimated diagnosis, the server refers to the user's location information and selects an appropriate medical institution.

[0158] Input: Estimated diagnosis result, user location information

[0159] Output: Information about the selected medical institution (location, contact information, available hours, etc.)

[0160] Step 12:

[0161] The server re-encrypts the diagnosis results and medical institution information and sends them to the terminal.

[0162] Input: Information on the selected medical institution, diagnosis results

[0163] Output: Encrypted data for transmission

[0164] Step 13:

[0165] The terminal receives the encrypted data sent from the server.

[0166] Input: Encrypted data to be sent

[0167] Output: Received encrypted data

[0168] Step 14:

[0169] The terminal decrypts the received encrypted data and displays it in a user-friendly format.

[0170] Input: Received encrypted data

[0171] Output: Diagnosis results and medical institution information displayed to the user

[0172] Through the above steps, the user can quickly and accurately understand the condition of their illness or injury and promptly seek medical attention at an appropriate medical institution.

[0173] (Application example 1)

[0174] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0175] In recent years, users have faced the challenge of finding an appropriate medical institution for their illness or injury quickly and accurately without specialized medical knowledge. Furthermore, there is a lack of systems that allow users to instantly make appointments and make payments based on diagnosis results. This creates a problem of inability to smoothly access and use medical services.

[0176] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0177] In this invention, the server includes means for selecting an appropriate medical institution based on a presumed diagnosis and transmitting that information, means for making a reservation at the recommended medical institution, and means for completing payment by electronic payment. This enables a user to quickly find an appropriate medical institution based on their condition or injury and seamlessly complete a series of processes from reservation to payment.

[0178] A "user" is someone who uses the system to input information about their medical condition or injury and wishes to receive recommendations or make appointments with medical institutions.

[0179] "Image data relating to a medical condition or injury" refers to photographs or image files taken by a user that show the state of a medical condition or injury.

[0180] "Text data" refers to character information entered by the user, including symptoms, impressions, and explanations of illnesses or injuries.

[0181] A "terminal" is a device, such as a smartphone or computer, that allows a user to access the system and input and send data.

[0182] A "server" is a computer system that receives data sent by users, analyzes it, selects medical institutions, and sends information.

[0183] A "specified format" is a format used to organize image data or text data into a certain format, such as JSON or XML.

[0184] "Encryption" is the process of converting data into a format that cannot be read by third parties in order to prevent data leakage during communication.

[0185] A "generative AI model" is a system that includes an artificial intelligence algorithm for analyzing and presumptively diagnosing medical conditions and symptoms using image and text data.

[0186] A "presumed diagnosis" is a provisional diagnosis derived from the results of the disease state and symptoms analyzed by the generative AI model.

[0187] "Selection of medical institution" refers to the procedure in which the server identifies an appropriate medical institution based on the estimated diagnosis result and taking into consideration the user's location information.

[0188] The "reservation means" is part of the system that allows the user to make an appointment with a recommended medical institution, and is a function that manages reservation information in cooperation with the server.

[0189] "Electronic payment" is a payment system conducted via the Internet, and refers to the use of online payment methods such as credit cards and electronic money.

[0190] This invention is a system that allows users to input image and text data about their medical condition or injury and receive referrals to appropriate medical institutions, even without specialized medical knowledge. This system is constructed using a terminal, a server, and a generative AI model.

[0191] First, the user takes a picture of their condition or injury using a smartphone or computer and saves the image data on the device. Next, the user enters their symptoms and thoughts about the condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text "My hand is swollen and painful."

[0192] The terminal packages the image data and text data entered by the user into a specified format (for example, JSON or XML). The terminal then sends this data to the server via the Internet. At this time, the data is encrypted to ensure security. For encryption, the pycryptodome library, for example, is used.

[0193] The server receives and decodes the image and text data sent from the device. The server uses a generative AI model to analyze the received image data. Specifically, it uses deep learning frameworks such as TENSORFLOW® and PyTorch to extract image features using algorithms such as convolutional neural networks (CNNs) and estimate the disease state. At the same time, the server uses a natural language processing (NLP) model to analyze the text data and estimate the symptoms.

[0194] The generative AI model combines the results of image analysis and text analysis to generate a probable diagnosis. This diagnosis includes a probable pathology and its reliability. For example, if image analysis detects a bone abnormality and text analysis identifies the symptom "swelling and pain," the probable diagnosis is "high probability of fracture."

[0195] The server selects the most appropriate medical institution based on the estimated diagnosis and the user's location information. At this time, the server retrieves a list of affiliated medical institutions from the database and lists the medical institutions closest to the user's current location. The information on the selected medical institution includes the location, contact information, and available hours for consultation.

[0196] The server then sends the estimated diagnosis and information on the recommended medical institution to the device. The data is then re-encrypted and transferred to the device in a secure manner.

[0197] The device decrypts the received data and displays it in a user-friendly format. The user can then check the displayed estimated diagnosis and information on recommended medical institutions, make an appointment with a medical institution within the app, and complete payment electronically. Payment can be made using APIs such as Stripe or PayPal.

[0198] Specific examples

[0199] For example, if a user injures their hand, the following process occurs:

[0200] 1. A user takes a photo of a hand injury with their smartphone and types, "My hand is swollen and painful."

[0201] 2. The device packages the image data and text data in JSON format, sends it to the server, and encrypts the data.

[0202] 3. The server analyzes the image using an AI model and the text data using an NLP model.

[0203] 4. The server integrates the analysis results and generates a diagnosis that "there is a high possibility of a fracture."

[0204] 5. The server selects a nearby orthopedic hospital based on the user's location information.

[0205] 6. The server sends the diagnosis results and medical institution information to the terminal.

[0206] 7. The terminal displays the received information to the user, who then contacts the orthopedic hospital within the app to make an appointment and completes payment electronically.

[0207] Prompt Sentence Examples

[0208] "Please provide a proper diagnosis and recommended medical care for the injury image taken by the user and the symptom text entered: 'My hand is swollen and painful.'"

[0209] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0210] Step 1:

[0211] The user takes an image of their condition or injury using a smartphone or computer and saves the image data on the device. The user then enters their symptoms and thoughts about the condition in text format. The input image data is a JPEG or PNG file, and the text data is a string of characters.

[0212] Step 2:

[0213] The device packages the image and text data entered by the user into the specified format. For example, this data is converted into JSON format. The packaged data is then encrypted to ensure security. The pycryptodome library is used for encryption, and the data is processed using the AES encryption algorithm.

[0214] Step 3:

[0215] The device sends the encrypted data to the server via the Internet in encrypted JSON format. The server receives the data using the HTTPS protocol to ensure secure communication.

[0216] Step 4:

[0217] The server decrypts the received data and converts it back into JSON image and text data, again using the pycryptodome library and the AES decryption algorithm, before formatting the decrypted data for analysis.

[0218] Step 5:

[0219] The server analyzes image data and text data. To analyze the image data, a deep learning framework such as TensorFlow or PyTorch is used to extract image features using a convolutional neural network (CNN). To analyze the text data, a natural language processing (NLP) model is used to infer symptoms from the text entered by the user. A generative AI model integrates these analysis results to make a probable diagnosis. For example, a CNN model detects bone abnormalities, and an NLP model analyzes the text "My hand is swollen and painful" and infers that there is a high possibility of a fracture.

[0220] Step 6:

[0221] The server selects the most suitable medical institution based on the estimated diagnosis and the user's location information. The server retrieves a list of affiliated medical institutions from an internal database and creates a list of medical institutions that can provide consultations. The server uses a geographic information system (GIS) to select the most suitable medical institution based on the user's location information.

[0222] Step 7:

[0223] The server encrypts the estimated diagnosis and information about the recommended medical institution and sends it back to the device. The data sent includes the estimated condition, reliability, name, location, contact information, and available hours of the recommended medical institution. The data is again AES encrypted and transmitted securely using the HTTPS protocol.

[0224] Step 8:

[0225] The device decrypts the received data and displays the diagnosis results and medical institution information so that the user can check them. Based on the displayed diagnosis results, the user can make an appointment with the recommended medical institution. The device application also has a function to contact the medical institution directly from the displayed information.

[0226] Step 9:

[0227] The user makes a reservation at a recommended medical institution from their device. They enter reservation information within the app and send a reservation request to the server. The server forwards the reservation information to the medical institution and confirms the reservation. The confirmed reservation information is sent to the user's device.

[0228] Step 10:

[0229] The user uses the terminal to make an electronic payment for a medical appointment. The terminal uses an API such as Stripe or PayPal to send the payment information to the server. The server processes the payment through a payment gateway and sends a confirmation to the user if the payment is successful.

[0230] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0231] This invention is a system that allows users to input image and text data related to their medical condition or injury and receive referrals to appropriate medical institutions, even without specialized medical knowledge. Furthermore, by combining this system with an emotion engine that recognizes the user's emotions, it is possible to make diagnoses and select medical institutions that take into account the user's emotional state. A specific embodiment of this system is described below.

[0232] User input of images and text

[0233] Users use their smartphones or computers to take pictures of their illnesses or injuries and save the image data on their devices. They then enter their symptoms and thoughts about their condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text, "My hand is swollen and painful."

[0234] Sending data by the device

[0235] The device converts the image data and text data entered by the user into a specified format (e.g., JSON or XML) and packages it. The device then encrypts the data and sends it to the server via the Internet.

[0236] Receipt and analysis of data by the server

[0237] The server receives and decodes the image and text data sent from the device. It then uses a generative AI model to analyze the received image data. Specifically, it uses algorithms such as convolutional neural networks (CNNs) to extract image features and estimate the disease state. At the same time, the server analyzes the text data using a natural language processing (NLP) model to estimate the symptoms.

[0238] Emotion engine recognizes emotional states

[0239] The server uses an emotion engine to recognize the user's emotional state using text data. The emotion engine analyzes keywords and context in the text to determine the user's emotional state, such as stress, anxiety, or pain. For example, if the text contains expressions such as "It hurts so much" or "I'm so anxious," the corresponding emotional state is recognized.

[0240] Predictive diagnosis using generative AI models

[0241] The generative AI model on the server combines the results of image analysis, text analysis, and emotion engine analysis to make a probable diagnosis. This diagnosis includes the estimated medical condition, its reliability, and the user's emotional state. For example, if image analysis suggests a fracture, text analysis identifies symptoms such as "swelling and pain," and the emotion engine recognizes "strong anxiety," the generative AI model will generate a probable fracture diagnosis.

[0242] Server selection of appropriate medical institution

[0243] Based on the estimated diagnosis, the server references the user's current location information and searches a database for the most suitable medical institution. The information on the medical institution includes location, contact information, and available consultation hours. The server also takes into account the user's emotional state. For example, if the server detects strong anxiety or stress, it will prioritize medical institutions that offer psychological counseling.

[0244] Sending results from the server to the device

[0245] The server packages the estimated diagnosis and information on the recommended medical institution, encrypts it again, and transmits it to the terminal.

[0246] Displaying results on a terminal

[0247] The device decodes the received data and displays it in a user-friendly format, including the estimated diagnosis and detailed information about the recommended medical institution (e.g., location, contact information, and opening hours). The user can then review the displayed information and contact the appropriate hospital to arrange an appointment.

[0248] Specific examples

[0249] For example, if a user has a hand injury, the following process occurs:

[0250] 1. A user takes a photo of a hand injury with their smartphone and enters, "My hand is swollen and painful, and I'm very anxious."

[0251] 2. The device packages the image data and text data in JSON format and sends it to the server.

[0252] 3. The server analyzes the images using an AI model and the text data using an NLP model and emotion engine.

[0253] 4. The server integrates the results of each analysis and generates a diagnosis of "high probability of fracture" and an emotional state indicating "high anxiety."

[0254] 5. The server selects nearby orthopedic hospitals and medical institutions that provide psychological counseling, taking into account the user's location information and emotional state.

[0255] 6. The server sends the diagnosis results and medical institution information to the terminal.

[0256] 7. The device displays the received information to the user, who then contacts the appropriate medical institution and arranges for an examination.

[0257] The above is a specific embodiment of the present invention. This system allows users to receive prompt and accurate diagnosis and access to medical institutions, taking into account not only their medical condition but also their emotional state.

[0258] The processing flow will be explained below.

[0259] Step 1:

[0260] The user takes an image of the medical condition or injury using a smartphone or computer, saves the image data on the device, and then enters a description of their symptoms or condition into a text input form. For example, they might enter "My hand is swollen and painful."

[0261] Step 2:

[0262] The device packages the image and text data entered by the user into a specified format (e.g., JSON or XML), encrypts the data, and prepares it for transmission.

[0263] Step 3:

[0264] The device sends the encrypted image data and text data to the server, and the data is transmitted securely over the Internet.

[0265] Step 4:

[0266] The server decrypts the received encrypted data and prepares it for analysis. The server passes image data to the image analysis module and text data to the natural language processing (NLP) module.

[0267] Step 5:

[0268] The server analyzes the image data using an image analysis module. Specifically, it uses a convolutional neural network (CNN) to extract image features and estimate the condition of the patient. For example, it can detect signs of a fracture from the image.

[0269] Step 6:

[0270] The server analyzes the text data using a natural language processing (NLP) module. It extracts symptom-related keywords and context from the text and analyzes the symptoms. For example, it extracts information such as "swelling" and "pain."

[0271] Step 7:

[0272] The server uses an emotion engine to recognize the user's emotional state from the text data. The emotion engine analyzes keywords and context in the text to determine the user's emotional state (e.g., anxiety, stress, pain level, etc.).

[0273] Step 8:

[0274] The generative AI model combines the results of image analysis, text analysis, and emotion engine analysis to produce a probable diagnosis, which includes an estimated medical condition, its confidence level, and the user's emotional state.

[0275] Step 9:

[0276] Based on the estimated diagnosis, the server references the user's current location and searches a database for the most suitable medical institution. The server also takes into account the user's emotional state when selecting an appropriate medical institution. For example, if strong anxiety or stress is detected, it will prioritize medical institutions that offer psychological counseling.

[0277] Step 10:

[0278] The server packages the estimated diagnosis results and information about the selected medical institution, encrypts them again, and sends them to the terminal.

[0279] Step 11:

[0280] The device decodes the received data and displays it in a user-friendly format, including the estimated diagnosis and detailed information about the recommended medical institution (e.g., location, contact information, consultation hours, etc.).

[0281] Step 12:

[0282] The user checks the medical institution information displayed on the device and contacts the appropriate hospital to arrange for a medical examination. For example, the user contacts a hospital in response to a message that reads, "There is a high possibility of a fracture. We recommend the following orthopedic hospital."

[0283] Example 2

[0284] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0285] Currently, many people have difficulty accurately and quickly communicating information about their illness or injury to medical institutions. In particular, it is difficult for users to accurately describe their symptoms and emotional state, which can delay the selection of an appropriate medical institution and diagnosis. To solve this problem, a system is needed that can effectively analyze user input data, perform a comprehensive diagnosis including emotional state, and quickly provide guidance to an appropriate medical institution.

[0286] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0287] In this invention, the server includes a means for analyzing image data input by a user to estimate a medical condition, a means for analyzing text data to estimate symptoms, and a means for recognizing the user's emotional state from the text data using an emotion engine. This enables a highly accurate inferential diagnosis to be made by comprehensively evaluating the image analysis results, text analysis results, and emotion analysis results. Furthermore, based on this inferential diagnosis, the server can quickly select an appropriate medical institution by referring to the user's location information and provide that information to the user, thereby enabling the provision of prompt and appropriate medical services.

[0288] "User" refers to an individual who uses the system to input information about a medical condition or injury.

[0289] "Terminal" refers to a device used by a user to input image and text data related to a medical condition or injury and send it to a server. Examples of such devices include smartphones and computers.

[0290] "Image data" refers to photographic data of illnesses or injuries that users take using their terminals and enter into the system.

[0291] "Text data" refers to textual information such as the status and impressions of a medical condition or injury that a user inputs into a device.

[0292] A "format" refers to the rules and forms for organizing and structuring data, such as JSON and XML.

[0293] A "server" is a central computer system for receiving and analyzing image data and text data sent from the terminals.

[0294] A "generative AI model" is a collection of algorithms and programs that run on a server and integrate the results of image analysis and text analysis to make a presumptive diagnosis.

[0295] A "presumed diagnosis" is the evaluation result of the disease state or symptoms calculated by the generative AI model, and is a preliminary assessment before an actual diagnosis is made.

[0296] An "emotion engine" is a software component that analyzes text data and recognizes the user's emotional state.

[0297] "Location information" is geographical data that indicates the user's current location, and is obtained from GPS data or the like.

[0298] "Medical institution" refers to a clinic, doctor's office, hospital, etc. that provides medical services to users.

[0299] "Packaging" is the process of combining multiple pieces of data into one format, and is done to maintain data consistency.

[0300] "Encryption" is the process of transforming data using a specific algorithm to protect the confidentiality of the data.

[0301] "Decryption" is the process of restoring encrypted data to its original form.

[0302] "Analysis" is the process of analyzing data and extracting meaningful information and patterns from it.

[0303] This invention is a system that allows users to input image and text data about their illness or injury and receive referrals to appropriate medical institutions, even without specialized medical knowledge. Furthermore, by combining this system with an emotion engine that recognizes the user's emotions, it is possible to make diagnoses and select medical institutions that take the user's emotional state into consideration.

[0304] Hardware and software used

[0305] The system consists of the following components:

[0306] User terminal: A device through which a user inputs and receives data, such as a smartphone or computer.

[0307] Server: A central computer system that analyzes data and selects medical institutions.

[0308] Generative AI models: Algorithms that perform data analysis, such as convolutional neural networks (CNNs) and natural language processing (NLP) models.

[0309] Emotion Engine: A software component for recognizing the emotional state of a user from their text data.

[0310] Processing flow

[0311] A specific embodiment for implementing this system will be described below.

[0312] User input of images and text

[0313] Users use their smartphones or computers to take pictures of their illnesses or injuries and save the image data on their devices. They then enter their symptoms and thoughts about their condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text, "My hand is swollen and painful."

[0314] Sending data by the device

[0315] The device converts the image data and text data entered by the user into a specified format (e.g., JSON format), packages it, encrypts it, and sends it to a server via the Internet.

[0316] Receipt and analysis of data by the server

[0317] The server receives and decodes the image and text data sent from the device. The server then uses a generative AI model to analyze the image data. Specifically, it uses a convolutional neural network (CNN) to extract image features and infer the disease state. At the same time, the server analyzes the text data with a natural language processing (NLP) model to infer the symptoms.

[0318] Emotion engine recognizes emotional states

[0319] The server uses an emotion engine to recognize the user's emotional state based on the text data. The emotion engine analyzes keywords and context within the text to determine the user's emotional state, such as stress, anxiety, or pain. For example, if the text contains expressions such as "It hurts so much" or "I'm so anxious," the corresponding emotional state is recognized.

[0320] Predictive diagnosis using generative AI models

[0321] The generative AI model on the server integrates the results of image analysis, text analysis, and emotion engine analysis to make a probable diagnosis. This diagnosis includes the estimated medical condition, its reliability, and the user's emotional state. For example, if image analysis suggests a fracture, text analysis identifies symptoms such as "swelling and pain," and the emotion engine recognizes "strong anxiety," the generative AI model will generate a probable fracture diagnosis.

[0322] Server selection of appropriate medical institution

[0323] The server then references the user's current location information based on the estimated diagnosis and searches a database for the most suitable medical institution. The information on the medical institution includes the location, contact information, and available consultation hours. The server also takes into account the user's emotional state. For example, if the server detects strong anxiety or stress, it will prioritize medical institutions that offer psychological counseling.

[0324] Sending results from the server to the device

[0325] The server packages the estimated diagnosis results and information on recommended medical institutions, encrypts them again, and sends them to the terminal.

[0326] Displaying results on a terminal

[0327] The device decodes the received data and displays it in a user-friendly format, including the estimated diagnosis and detailed information about the recommended medical institution (such as location, contact information, and opening hours). The user can then review the displayed information and contact the appropriate hospital to arrange an appointment.

[0328] Specific examples

[0329] For example, if a user has a hand injury, the following process occurs:

[0330] 1. A user takes a photo of a hand injury with their smartphone and enters, "My hand is swollen and painful, and I'm very anxious."

[0331] 2. The device packages the image data and text data in JSON format and sends it to the server.

[0332] 3. The server analyzes the images using an AI model and the text data using an NLP model and emotion engine.

[0333] 4. The server integrates the results of each analysis and generates a diagnosis of "high probability of fracture" and an emotional state indicating "high anxiety."

[0334] 5. The server selects nearby orthopedic hospitals and medical institutions that provide psychological counseling, taking into account the user's location information and emotional state.

[0335] 6. The server sends the diagnosis results and medical institution information to the terminal.

[0336] 7. The device displays the received information to the user, who then contacts the appropriate medical institution and arranges for an examination.

[0337] This process allows users to access medical care and receive a quick and accurate diagnosis that takes into account not only their medical condition but also their emotional state.

[0338] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0339] Step 1:

[0340] User input of images and text

[0341] Users use a smartphone or computer to take pictures of their condition or injury, save the image data on the device, and then enter detailed symptoms and impressions of their condition in text format.

[0342] Input: Image data and text data related to medical conditions and injuries

[0343] Output: Image data and text data stored on the device

[0344] Specific behavior:

[0345] The user activates the camera and takes a picture of the affected area.

[0346] After taking the photo, you upload the image through the device's application and enter details such as "My hand is swollen and painful, and I'm very anxious" in the text input field.

[0347] Step 2:

[0348] Sending data by the device

[0349] The device converts the image data and text data entered by the user into a specified format (e.g., JSON format) and packages it. The device then encrypts the data and sends it to the server via the Internet.

[0350] Input: User image and text data

[0351] Output: Encrypted packaging data

[0352] Specific behavior:

[0353] The device converts the image file and text data into JSON format.

[0354] The data is encrypted using the SSL / TLS protocol and sent to the server's API endpoint.

[0355] Step 3:

[0356] Receipt and analysis of data by the server

[0357] The server receives and decrypts the encrypted data sent from the device. Next, the server analyzes the received image data using a generative AI model (CNN algorithm) to extract image features. At the same time, the server analyzes the text data using a natural language processing (NLP) model to estimate symptoms.

[0358] Input: Encrypted packaging data

[0359] Output: Analyzed image features and text analysis results

[0360] Specific behavior:

[0361] The server reads the encrypted data from the SSL / TLS input buffer and decrypts it.

[0362] The server inputs the image data into a generative AI model, extracts features, and estimates the condition of the disease.

[0363] An NLP processing engine is used to analyze the text data and extract keywords and context related to the symptoms.

[0364] Step 4:

[0365] Emotion engine recognizes emotional states

[0366] The server analyzes the text data with an emotion engine to recognize the user's emotional state. The emotion engine analyzes keywords and context within the text to determine the user's stress, anxiety, pain, etc.

[0367] Input: Text data

[0368] Output: Data about the user's emotional state

[0369] Specific behavior:

[0370] The server inputs the text data into an emotion engine, which analyzes the keywords and context within the text.

[0371] The emotion engine analyzes expressions such as "It hurts so much" and "I'm so anxious" to determine the user's emotional state.

[0372] Step 5:

[0373] Predictive diagnosis using generative AI models

[0374] The generative AI model on the server integrates the results of image analysis, text analysis, and emotion engine analysis to produce a probable diagnosis, which includes the probable medical condition, its reliability, and the user's emotional state.

[0375] Input: Image analysis results, text analysis results, sentiment analysis results

[0376] Output: Estimated diagnosis result

[0377] Specific behavior:

[0378] The server integrates the results of image analysis, text analysis, and sentiment analysis.

[0379] The generative AI model estimates the condition based on the integrated data and generates a diagnosis such as "high probability of fracture."

[0380] Step 6:

[0381] Server selection of appropriate medical institution

[0382] Based on the estimated diagnosis, the server references the user's current location information and searches a database for the most suitable medical institution, including information on the location, contact details, and available hours for consultations.

[0383] Input: Estimated diagnosis result, user's current location information

[0384] Output: Appropriate medical institution information

[0385] Specific behavior:

[0386] The server obtains the user's current location information from GPS data and queries a location information database.

[0387] The server searches the medical institution database and generates a list of medical institutions that can provide consultations.

[0388] Step 7:

[0389] Sending results from the server to the device

[0390] The server packages the estimated diagnosis and information on the recommended medical institution, encrypts it again, and transmits it to the terminal.

[0391] Input: Estimated diagnosis result, medical institution information

[0392] Output: Encrypted diagnosis results and medical institution information

[0393] Specific behavior:

[0394] The server packages the diagnosis results and medical institution information in JSON format.

[0395] The server encrypts the packaged data and sends it to the user's terminal.

[0396] Step 8:

[0397] Displaying results on a terminal

[0398] The device decodes the received data and displays it in a user-friendly format, including a probable diagnosis and detailed information about recommended medical institutions.

[0399] Input: Encrypted diagnosis results and medical institution information

[0400] Output: Decrypted diagnosis and medical institution information

[0401] Specific behavior:

[0402] The device decodes the data it receives and parses the JSON format data to extract the diagnosis results and medical institution information.

[0403] The device updates the application UI and displays the diagnosis results and medical institution information to the user.

[0404] Through this detailed, step-by-step process, users can obtain an accurate and prompt diagnosis and access to the appropriate medical facility.

[0405] (Application example 2)

[0406] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0407] Conventional medical diagnosis systems have the problem that it is difficult for users without medical expertise to make an appropriate diagnosis and select a medical institution. Furthermore, there is a need for rapid medical diagnosis and guidance to medical institutions in emergency situations. In particular, a system that can provide appropriate information immediately is needed so that security staff can respond quickly on the scene.

[0408] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0409] In this invention, the server includes: a means for a user to input image data and text data related to a medical condition or injury; a means for a terminal to package the image data and text data into a specified format and transmit the packaged image data and text data to the server; a means for the server to analyze the image data and estimate a medical condition; a means for the server to analyze the text data and estimate symptoms; a generative AI model to integrate the analysis results and make an estimated diagnosis; a means for the server to select an appropriate medical institution based on the estimated diagnosis; a means for the server to transmit information about the estimated diagnosis and the medical institution to the terminal; a means for the terminal to display the information to the user; and a means for performing a rapid medical diagnosis on site using smart glasses. This enables a user (security staff) to receive a rapid and appropriate medical diagnosis on site and access an appropriate medical institution based on the results.

[0410] "Image data" is digital data entered by a user that contains visual information about a medical condition or injury.

[0411] "Text data" is digital data containing written information related to a medical condition or injury entered by a user.

[0412] A "user" is an individual who uses the system and who inputs information about a medical condition or injury.

[0413] A "terminal" is a device operated by a user, and is a device for packaging image data and text data and transmitting them to a server.

[0414] A "server" is a computer system that receives data sent from a terminal and performs analysis and inferential diagnosis.

[0415] A "generative AI model" is an artificial intelligence algorithm that integrates the analysis results of image data and text data to make a presumptive diagnosis.

[0416] "Format" refers to a data format for arranging image data and text data in a form that is easy for the server to process.

[0417] "Smart glasses" are wearable devices equipped with a camera function that users use in the field.

[0418] An "appropriate medical institution" is a medical facility that can provide optimal medical services to the user based on the estimated diagnosis results.

[0419] A "presumed diagnosis" is the result of a disease state inferred by a generative AI model based on an analysis of image data and text data.

[0420] "Display means" refers to a function that visually presents to the user the information that the terminal receives from the server.

[0421] "Emotional state" indicates the psychological and emotional state of the user, which is analyzed from text data or the like.

[0422] In order to implement the present invention, the following configurations and means are necessary.

[0423] Overall system overview

[0424] The user inputs image and text data of their condition or injury and sends the information to a server. The server analyzes the data, makes a probable diagnosis, selects an appropriate medical institution, and provides that information to the user. Furthermore, the use of smart glasses enables rapid medical diagnosis on the spot.

[0425] Hardware and Software

[0426] 1. Smart Glasses:

[0427] A wearable device with a camera function.

[0428] Supports on-site image capture and data entry.

[0429] 2. Terminal:

[0430] Personal computers and smartphones.

[0431] It supports the input of image data and text data and sends the data to the server.

[0432] 3. Server:

[0433] High performance computing systems.

[0434] Analyzes image and text data, performs presumptive diagnoses, and selects appropriate medical institutions.

[0435] The algorithms used are convolutional neural networks (CNN) and natural language processing (NLP).

[0436] Program processing

[0437] The device packages the image data and text data entered by the user into a specified format (e.g., JSON), encrypts it, and sends it to the server. The server decrypts the received data and performs the following processes.

[0438] 1. Analysis of imaging data:

[0439] The server uses a convolutional neural network (CNN) to extract features from the image data and estimate the disease state.

[0440] 2. Text data analysis:

[0441] Text data is analyzed using a natural language processing (NLP) model to estimate symptoms.

[0442] 3. Integrated analysis using generative AI models:

[0443] The generative AI model integrates the results of image analysis, text analysis, and the emotion engine analysis to make a presumptive diagnosis.

[0444] 4. Selection of appropriate medical facility:

[0445] Based on the estimated diagnosis results, the server selects the most appropriate medical institution, taking into consideration the user's location information and emotional state.

[0446] 5. Transmission and Display of Information:

[0447] The server transmits the generated estimated diagnosis and information about the medical institution to the terminal, which then displays them to the user.

[0448] Specific examples

[0449] For example, if security staff discover someone collapsed at a scene, they can take the following steps:

[0450] 1. Taking a photo of a fallen person using smart glasses.

[0451] 2. Enter the text data "Unconscious, not breathing."

[0452] 3. The device sends this data to the server.

[0453] 4. The server makes a probable diagnosis and generates information such as "possibility of cardiac arrest" and "nearest hospital with a cardiologist."

[0454] 5. The server sends the diagnosis results and information on recommended medical institutions to the terminal, which displays them to the user.

[0455] Prompt Sentence Examples

[0456] "Unconscious, not breathing"

[0457] "My hands are swollen and painful, and I'm very anxious"

[0458] By using these prompts, users can quickly and accurately access the appropriate medical institution.

[0459] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0460] Step 1:

[0461] The user uses the smart glasses to take images of their condition or injury and input text data about their symptoms. The user uses the smart glasses' camera function to take photos of their condition or injury, and then inputs text information such as "My hand is swollen and painful, and I feel very anxious" through the smart glasses' UI.

[0462] Input: Image and text data of medical conditions and injuries.

[0463] Output: Photographed image data and input text data.

[0464] Step 2:

[0465] The device packages the image and text data entered by the user into a specified format, encrypts it, and sends it to the server.The device then converts the data into a structured data format such as JSON and encrypts it using an encryption algorithm such as AES for security.The encrypted data is then sent to the server via the HTTPS protocol.

[0466] Input: Image data and text data.

[0467] Output: Encrypted and formatted data.

[0468] Step 3:

[0469] The server receives and decrypts the data sent from the device. The server receives the data via the HTTPS protocol and decrypts data encrypted with AES or other encryption protocols. It then analyzes the received JSON format data and separates the image data from the text data.

[0470] Input: Encrypted and formatted data.

[0471] Output: Decoded and separated image and text data.

[0472] Step 4:

[0473] The server analyzes the image data to estimate the condition of the patient. Specifically, it uses a convolutional neural network (CNN) to extract image features and estimate the condition. Based on the results of this analysis, a diagnosis such as "high possibility of fracture" is obtained.

[0474] Input: Image data.

[0475] Output: Estimated disease state based on image analysis.

[0476] Step 5:

[0477] The server analyzes the text data to infer symptoms. It uses a natural language processing (NLP) model to analyze keywords and context within the text data to infer symptoms. For example, from an expression such as "My hand is swollen and painful, and I feel very anxious," it extracts symptoms such as "swelling," "pain," and "anxiety."

[0478] Input: Text data.

[0479] Output: Symptom prediction results based on text analysis.

[0480] Step 6:

[0481] The generative AI model combines the results of image analysis, text analysis, and the emotion engine to generate a comprehensive diagnosis, such as "There is a high possibility of a fracture and the patient feels high anxiety."

[0482] Input: Image analysis results, text analysis results, emotion engine results.

[0483] Output: Integrated estimated diagnostic results.

[0484] Step 7:

[0485] The server selects an appropriate medical institution based on the user's location information and emotional state. Based on the estimated diagnosis, the server searches a database of medical institutions and selects the most suitable medical institution. For example, the server selects the "nearest orthopedic hospital" and "medical institution that offers psychological counseling" based on the user's location information and emotional state.

[0486] Input: Estimated diagnosis result, user location information, emotional state.

[0487] Output: Information on selected medical institutions.

[0488] Step 8:

[0489] The server then sends the generated estimated diagnosis and information about the medical institution to the terminal, where it re-encrypts the information and transmits it to the terminal using the HTTPS protocol.

[0490] Input: Estimated diagnosis result, medical institution information.

[0491] Output: Encrypted diagnosis results and medical institution information.

[0492] Step 9:

[0493] The device decrypts the information received from the server and displays it in a format that is easy for the user to understand. The device decrypts the encrypted data and displays information such as "High probability of fracture," "Nearest orthopedic hospital: XX Hospital," and "Psychological counseling: YY Clinic" through the UI.

[0494] Input: Encrypted diagnosis results and medical institution information.

[0495] Output: Presenting information in a user-friendly format.

[0496] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0497] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0498] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0499] [Second embodiment]

[0500] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0501] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0502] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0503] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0504] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0505] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0506] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0507] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0508] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0509] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0510] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0511] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0512] This invention is a system that allows users to input image and text data related to a medical condition or injury and receive a referral to an appropriate medical institution even if they do not have specialized medical knowledge. Specific embodiments of this system will be described below.

[0513] User input of images and text

[0514] Users can use a smartphone or computer to take pictures of their illness or injury and save the image data on the device. They can then enter their symptoms and thoughts about the condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text, "My hand is swollen and painful."

[0515] Sending data by the device

[0516] The device packages the image data and text data entered by the user into a specified format (for example, JSON or XML).The device then sends this data to a server via the Internet.At this time, the data is encrypted to ensure security.

[0517] Receipt and analysis of data by the server

[0518] The server receives and decodes the image and text data sent from the device. It then uses a generative AI model to analyze the received image data. Specifically, it uses algorithms such as convolutional neural networks to extract image features and estimate the condition. At the same time, the server analyzes the text data using a natural language processing (NLP) model to estimate symptoms.

[0519] Predictive diagnosis using generative AI models

[0520] The generative AI model on the server integrates the results of image analysis and text analysis to generate a probable diagnosis. This diagnosis includes a probable pathology and its reliability. For example, if image analysis detects a bone abnormality and text analysis identifies the symptom "swelling and pain," the probable diagnosis is "high probability of fracture."

[0521] Server selection of appropriate medical institution

[0522] Based on the estimated diagnosis, the server refers to the user's location information and selects the most appropriate medical institution. At this time, the server retrieves a list of affiliated medical institutions from the database and lists medical institutions close to the user's current location. The information on the selected medical institution includes location, contact information, available hours for consultation, etc.

[0523] Sending results from the server to the device

[0524] The server then sends the estimated diagnosis and information on the recommended medical institution to the device. The data is then re-encrypted and transferred to the device in a secure manner.

[0525] Displaying results on a terminal

[0526] The device decodes the received data and displays it in a user-friendly format. The user can then check the displayed estimated diagnosis and recommended medical institution information and make arrangements for a medical examination if necessary. For example, if there is a high possibility of a fracture, the user can call the orthopedic hospital contact number displayed on the device to make an appointment.

[0527] Specific examples

[0528] For example, if a user injures their hand, the following process occurs:

[0529] 1. A user takes a photo of a hand injury with their smartphone and types, "My hand is swollen and painful."

[0530] 2. The device packages the image data and text data in JSON format and sends it to the server.

[0531] 3. The server analyzes the image using an AI model and the text data using an NLP model.

[0532] 4. The server integrates the analysis results and generates a diagnosis of "high probability of fracture."

[0533] 5. The server selects a nearby orthopedic hospital based on the user's location information.

[0534] 6. The server sends the diagnosis results and hospital information to the terminal.

[0535] 7. The terminal displays the received information to the user, who then contacts the orthopedic hospital to arrange for a consultation.

[0536] The above is a specific embodiment of the present invention. This system makes it possible to quickly and accurately estimate the condition of a patient, even without specialized medical knowledge, and to assist in accessing an appropriate medical institution.

[0537] The processing flow will be explained below.

[0538] Step 1: The user takes an image of their condition or injury using a smartphone or computer, saves the image data on their device, and enters a description of their symptoms or condition into a text input form.

[0539] Step 2: The terminal converts the image data and text data entered by the user into a specified format (e.g., JSON or XML) and packages it.

[0540] Step 3: The device encrypts the packaged data and sends it to a server via the Internet.

[0541] Step 4: The server decodes the received image data and text data and prepares each for analysis.

[0542] Step 5: The server analyzes the image data using an image analysis module. Specifically, it uses a convolutional neural network (CNN) to extract image features and estimate the disease state.

[0543] Step 6: The server analyzes the text data using a natural language processing (NLP) module, extracting symptom-related keywords and context from the text and obtaining the analysis results.

[0544] Step 7: The generative AI model on the server integrates the results of image analysis and text analysis to make a probable diagnosis. For example, if image analysis indicates a suspected fracture and text analysis indicates pain or swelling, the model will generate a probable diagnosis of "high probability of fracture."

[0545] Step 8: Based on the estimated diagnosis, the server references the user's current location information and searches the database for the most suitable medical institution. Information on the target medical institutions is then listed.

[0546] Step 9: The server packages the estimated diagnosis and the information on the selected medical institution, encrypts it again, and transmits it to the terminal.

[0547] Step 10: The terminal decrypts the received data and displays it to the user, including the estimated diagnosis and detailed information (such as location, contact information, and consultation hours) of the recommended medical institution.

[0548] Step 11: The user checks the information about the medical institution displayed on the terminal and contacts the appropriate hospital to arrange for a medical examination.

[0549] Example 1

[0550] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0551] In today's world, it is difficult for users to quickly and accurately communicate information about their illness or injury to specialized medical institutions. It is also difficult for users without medical knowledge to select an appropriate medical institution. As a result, it can take a lot of time and effort to receive an appropriate diagnosis and treatment. To solve these problems, a system is needed that can effectively collect information about a user's illness or injury, analyze it quickly and accurately, and select an appropriate medical institution.

[0552] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0553] In this invention, the server includes: means for a user to input image data and text data related to a medical condition or injury; means for a terminal to package the image data and text data in a specified format and transmit the packaged image data and text data to the server; means for the terminal to encrypt the data; means for the server to receive and decrypt the image data and text data; means for the server to analyze the image data, extract features, and infer a medical condition; means for the server to analyze the text data, extract and infer symptoms; means for a generative AI model to integrate the analysis results and make a probable diagnosis; means for the server to select an appropriate medical institution based on the probable diagnosis; means for the server to transmit information about the probable diagnosis and medical facilities to the terminal; and means for the terminal to display the information to the user. This enables users to quickly and accurately understand their own medical condition or injury and to promptly visit an appropriate medical institution, even without specialized medical knowledge.

[0554] "Image data" is digital information that a user visually records regarding a medical condition or injury.

[0555] "Text data" is digital information in which a user records symptoms and thoughts about a medical condition or injury in written form.

[0556] A "terminal" is a device that a user uses to input data and send it to a server, and includes a smartphone or computer.

[0557] A "server" is a central processing device that receives data sent from a terminal, analyzes it, and returns the necessary information.

[0558] A "specified format" is a format in which data is arranged according to certain rules, and includes, for example, JSON and XML.

[0559] "Encryption" is a technique that uses a specific algorithm to convert data in order to transmit it securely.

[0560] "Decryption" is the technique of returning encrypted data to its original form.

[0561] "Features" are important attributes or patterns extracted from image data and are used to estimate the condition of a disease.

[0562] A "natural language processing (NLP) model" is an algorithm or method for analyzing text data and understanding human language.

[0563] A "generative AI model" is an artificial intelligence algorithm that integrates the results of image analysis and text analysis to make a presumptive diagnosis.

[0564] A "presumptive diagnosis" is the result of predicting the condition of a disease or injury based on analyzed data.

[0565] A "medical institution" is a facility where a user can receive medical examinations and treatment, and includes hospitals and clinics.

[0566] "Information" refers to data such as the diagnosis results and contact details and locations of medical institutions required by the user.

[0567] The present invention provides a system that allows a user to input image data and text data relating to a medical condition or injury, and easily refer the user to an appropriate medical institution. A specific embodiment of this system will be described below.

[0568] User data entry

[0569] Users use devices such as smartphones or computers to take pictures of their medical condition or injury and save them on the device. Next, they enter text data about their condition (e.g., "My hand is swollen and painful") into the device. At this stage, users do not need specialized medical knowledge and can enter data intuitively.

[0570] Device prepares to send data

[0571] The device packages the image and text data entered by the user into a specified format (e.g., JSON or XML), then encrypts the packaged data and transmits it securely to a server over the Internet.

[0572] Receipt and analysis of data by the server

[0573] The server receives the encrypted data sent from the device. After receiving the data, the server decrypts it and extracts the image data and text data. The server then analyzes the image data using a convolutional neural network (CNN) to extract image features. The server also analyzes the text data using a natural language processing (NLP) model to extract the described symptoms.

[0574] Predictive diagnosis using generative AI models

[0575] The generative AI model deployed on the server integrates the results of image analysis and text analysis. This model then uses the analysis results to generate a probable diagnosis, which includes a probable medical condition and its reliability. For example, by combining an image of a hand injury with the text data "your hand is swollen and painful," the model generates a probable diagnosis of "high probability of fracture."

[0576] Selection of medical institutions by the server

[0577] The server then references the user's location information based on the estimated diagnosis and selects an appropriate medical institution from a database. The information on the selected medical institution includes the location, contact information, and available consultation hours. The server then organizes this information, re-encrypts the data, and sends it to the device.

[0578] Displaying results on a terminal

[0579] The device receives and decrypts the encrypted data sent from the server. The decrypted data is displayed in a user-friendly format, allowing the user to refer to it and contact the appropriate medical institution and arrange for a medical examination. For example, it may display information such as "There is a high possibility of a fracture. The recommended medical institution is XX Orthopedic Hospital, contact number: XX-XX-XX."

[0580] Examples of concrete examples and prompts

[0581] As a concrete example, if a user injures their hand, the process is as follows:

[0582] 1. A user takes a photo of a hand injury with their smartphone and types, "My hand is swollen and painful."

[0583] 2. The device packages the image data and text data in JSON format and sends it to the server.

[0584] 3. The server analyzes the image using an AI model and the text data using an NLP model.

[0585] 4. The server integrates the analysis results and generates a diagnosis of "high probability of fracture."

[0586] 5. The server selects a nearby orthopedic hospital based on the user's location information.

[0587] 6. The server sends the diagnosis results and hospital information to the terminal.

[0588] 7. The terminal displays the received information to the user, who then contacts the orthopedic hospital to arrange for a consultation.

[0589] This system allows even users without specialized medical knowledge to quickly and accurately understand the patient's condition and refer them to the appropriate medical institution.

[0590] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0591] Step 1:

[0592] The user uses a smartphone or computer to take pictures of their condition or injury and save them on the device, then enters text about their condition (e.g., "My hand is swollen and painful") into the device.

[0593] Input: Image data and text data related to medical conditions and injuries

[0594] Output: Image data and text data stored on the device

[0595] Step 2:

[0596] The terminal acquires image data and text data input by the user.

[0597] Input: Image data and text data stored on the device

[0598] Output: Acquired image data and text data

[0599] Step 3:

[0600] The image data and text data acquired by the device are packaged in a specified format such as JSON.

[0601] Input: Acquired image data and text data

[0602] Output: Data packaged in the specified format

[0603] Step 4:

[0604] The device encrypts the packaged data and prepares it in a highly secure state.

[0605] Input: Data packaged in a specified format

[0606] Output: Encrypted data

[0607] Step 5:

[0608] The terminal transmits the encrypted data to a server via the Internet.

[0609] Input: Encrypted data

[0610] Output: Data sent to the server

[0611] Step 6:

[0612] The server receives the encrypted data sent from the terminal.

[0613] Input: Data sent over the internet

[0614] Output: Encrypted data received by the server

[0615] Step 7:

[0616] The server decrypts the received encrypted data and extracts the image data and text data.

[0617] Input: Received encrypted data

[0618] Output: Decoded image data, text data

[0619] Step 8:

[0620] The server analyzes the image data using a convolutional neural network (CNN) and extracts features.

[0621] Input: Decoded image data

[0622] Output: Image features

[0623] Step 9:

[0624] The server analyzes the text data using a natural language processing (NLP) model to extract symptoms.

[0625] Input: Decrypted text data

[0626] Output: Extracted symptoms

[0627] Step 10:

[0628] The server uses the generated AI model to integrate the results of image analysis and text analysis and make a presumptive diagnosis.

[0629] Input: Image features, extracted symptoms

[0630] Output: Estimated diagnosis result (condition and its reliability)

[0631] Step 11:

[0632] Based on the estimated diagnosis, the server refers to the user's location information and selects an appropriate medical institution.

[0633] Input: Estimated diagnosis result, user location information

[0634] Output: Information about the selected medical institution (location, contact information, available hours, etc.)

[0635] Step 12:

[0636] The server re-encrypts the diagnosis results and medical institution information and sends them to the terminal.

[0637] Input: Information on the selected medical institution, diagnosis results

[0638] Output: Encrypted data for transmission

[0639] Step 13:

[0640] The terminal receives the encrypted data sent from the server.

[0641] Input: Encrypted data to be sent

[0642] Output: Received encrypted data

[0643] Step 14:

[0644] The terminal decrypts the received encrypted data and displays it in a user-friendly format.

[0645] Input: Received encrypted data

[0646] Output: Diagnosis results and medical institution information displayed to the user

[0647] Through the above steps, the user can quickly and accurately understand the condition of their illness or injury and promptly seek medical attention at an appropriate medical institution.

[0648] (Application example 1)

[0649] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0650] In recent years, users have faced the challenge of finding an appropriate medical institution for their illness or injury quickly and accurately without specialized medical knowledge. Furthermore, there is a lack of systems that allow users to instantly make appointments and make payments based on diagnosis results. This creates a problem of inability to smoothly access and use medical services.

[0651] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0652] In this invention, the server includes means for selecting an appropriate medical institution based on a presumed diagnosis and transmitting that information, means for making a reservation at the recommended medical institution, and means for completing payment by electronic payment. This enables a user to quickly find an appropriate medical institution based on their condition or injury and seamlessly complete a series of processes from reservation to payment.

[0653] A "user" is someone who uses the system to input information about their medical condition or injury and wishes to receive recommendations or make appointments with medical institutions.

[0654] "Image data relating to a medical condition or injury" refers to photographs or image files taken by a user that show the state of a medical condition or injury.

[0655] "Text data" refers to character information entered by the user, including symptoms, impressions, and explanations of illnesses or injuries.

[0656] A "terminal" is a device, such as a smartphone or computer, that allows a user to access the system and input and send data.

[0657] A "server" is a computer system that receives data sent by users, analyzes it, selects medical institutions, and sends information.

[0658] A "specified format" is a format used to organize image data or text data into a certain format, such as JSON or XML.

[0659] "Encryption" is the process of converting data into a format that cannot be read by third parties in order to prevent data leakage during communication.

[0660] A "generative AI model" is a system that includes an artificial intelligence algorithm for analyzing and presumptively diagnosing medical conditions and symptoms using image and text data.

[0661] A "presumed diagnosis" is a provisional diagnosis derived from the results of the disease state and symptoms analyzed by the generative AI model.

[0662] "Selection of medical institution" refers to the procedure in which the server identifies an appropriate medical institution based on the estimated diagnosis result and taking into consideration the user's location information.

[0663] The "reservation means" is part of the system that allows the user to make an appointment with a recommended medical institution, and is a function that manages reservation information in cooperation with the server.

[0664] "Electronic payment" is a payment system conducted via the Internet, and refers to the use of online payment methods such as credit cards and electronic money.

[0665] This invention is a system that allows users to input image and text data about their medical condition or injury and receive referrals to appropriate medical institutions, even without specialized medical knowledge. This system is constructed using a terminal, a server, and a generative AI model.

[0666] First, the user takes a picture of their condition or injury using a smartphone or computer and saves the image data on the device. Next, the user enters their symptoms and thoughts about the condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text "My hand is swollen and painful."

[0667] The terminal packages the image data and text data entered by the user into a specified format (for example, JSON or XML). The terminal then sends this data to the server via the Internet. At this time, the data is encrypted to ensure security. For encryption, the pycryptodome library, for example, is used.

[0668] The server receives and decodes the image and text data sent from the device. It then uses a generative AI model to analyze the received image data. Specifically, it uses deep learning frameworks such as TensorFlow and PyTorch to extract image features using algorithms such as convolutional neural networks (CNNs) and estimate the condition of the patient. At the same time, the server uses a natural language processing (NLP) model to analyze the text data and estimate the symptoms.

[0669] The generative AI model combines the results of image analysis and text analysis to generate a probable diagnosis. This diagnosis includes a probable pathology and its reliability. For example, if image analysis detects a bone abnormality and text analysis identifies the symptom "swelling and pain," the probable diagnosis is "high probability of fracture."

[0670] The server selects the most appropriate medical institution based on the estimated diagnosis and the user's location information. At this time, the server retrieves a list of affiliated medical institutions from the database and lists the medical institutions closest to the user's current location. The information on the selected medical institution includes the location, contact information, and available hours for consultation.

[0671] The server then sends the estimated diagnosis and information on the recommended medical institution to the device. The data is then re-encrypted and transferred to the device in a secure manner.

[0672] The device decrypts the received data and displays it in a user-friendly format. The user can then check the displayed estimated diagnosis and information on recommended medical institutions, make an appointment with a medical institution within the app, and complete payment electronically. Payment can be made using APIs such as Stripe or PayPal.

[0673] Specific examples

[0674] For example, if a user injures their hand, the following process occurs:

[0675] 1. A user takes a photo of a hand injury with their smartphone and types, "My hand is swollen and painful."

[0676] 2. The device packages the image data and text data in JSON format, sends it to the server, and encrypts the data.

[0677] 3. The server analyzes the image using an AI model and the text data using an NLP model.

[0678] 4. The server integrates the analysis results and generates a diagnosis that "there is a high possibility of a fracture."

[0679] 5. The server selects a nearby orthopedic hospital based on the user's location information.

[0680] 6. The server sends the diagnosis results and medical institution information to the terminal.

[0681] 7. The terminal displays the received information to the user, who then contacts the orthopedic hospital within the app to make an appointment and completes payment electronically.

[0682] Prompt Sentence Examples

[0683] "Please provide a proper diagnosis and recommended medical care for the injury image taken by the user and the symptom text entered: 'My hand is swollen and painful.'"

[0684] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0685] Step 1:

[0686] The user takes an image of their condition or injury using a smartphone or computer and saves the image data on the device. The user then enters their symptoms and thoughts about the condition in text format. The input image data is a JPEG or PNG file, and the text data is a string of characters.

[0687] Step 2:

[0688] The device packages the image and text data entered by the user into the specified format. For example, this data is converted into JSON format. The packaged data is then encrypted to ensure security. The pycryptodome library is used for encryption, and the data is processed using the AES encryption algorithm.

[0689] Step 3:

[0690] The device sends the encrypted data to the server via the Internet in encrypted JSON format. The server receives the data using the HTTPS protocol to ensure secure communication.

[0691] Step 4:

[0692] The server decrypts the received data and converts it back into JSON image and text data, again using the pycryptodome library and the AES decryption algorithm, before formatting the decrypted data for analysis.

[0693] Step 5:

[0694] The server analyzes image data and text data. To analyze the image data, a deep learning framework such as TensorFlow or PyTorch is used to extract image features using a convolutional neural network (CNN). To analyze the text data, a natural language processing (NLP) model is used to infer symptoms from the text entered by the user. A generative AI model integrates these analysis results to make a probable diagnosis. For example, a CNN model detects bone abnormalities, and an NLP model analyzes the text "My hand is swollen and painful" and infers that there is a high possibility of a fracture.

[0695] Step 6:

[0696] The server selects the most suitable medical institution based on the estimated diagnosis and the user's location information. The server retrieves a list of affiliated medical institutions from an internal database and creates a list of medical institutions that can provide consultations. The server uses a geographic information system (GIS) to select the most suitable medical institution based on the user's location information.

[0697] Step 7:

[0698] The server encrypts the estimated diagnosis and information about the recommended medical institution and sends it back to the device. The data sent includes the estimated condition, reliability, name, location, contact information, and available hours of the recommended medical institution. The data is again AES encrypted and transmitted securely using the HTTPS protocol.

[0699] Step 8:

[0700] The device decrypts the received data and displays the diagnosis results and medical institution information so that the user can check them. Based on the displayed diagnosis results, the user can make an appointment with the recommended medical institution. The device application also has a function to contact the medical institution directly from the displayed information.

[0701] Step 9:

[0702] The user makes a reservation at a recommended medical institution from their device. They enter reservation information within the app and send a reservation request to the server. The server forwards the reservation information to the medical institution and confirms the reservation. The confirmed reservation information is sent to the user's device.

[0703] Step 10:

[0704] The user uses the terminal to make an electronic payment for a medical appointment. The terminal uses an API such as Stripe or PayPal to send the payment information to the server. The server processes the payment through a payment gateway and sends a confirmation to the user if the payment is successful.

[0705] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0706] This invention is a system that allows users to input image and text data related to their medical condition or injury and receive referrals to appropriate medical institutions, even without specialized medical knowledge. Furthermore, by combining this system with an emotion engine that recognizes the user's emotions, it is possible to make diagnoses and select medical institutions that take into account the user's emotional state. A specific embodiment of this system is described below.

[0707] User input of images and text

[0708] Users use their smartphones or computers to take pictures of their illnesses or injuries and save the image data on their devices. They then enter their symptoms and thoughts about their condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text, "My hand is swollen and painful."

[0709] Sending data by the device

[0710] The device converts the image data and text data entered by the user into a specified format (e.g., JSON or XML) and packages it. The device then encrypts the data and sends it to the server via the Internet.

[0711] Receipt and analysis of data by the server

[0712] The server receives and decodes the image and text data sent from the device. It then uses a generative AI model to analyze the received image data. Specifically, it uses algorithms such as convolutional neural networks (CNNs) to extract image features and estimate the disease state. At the same time, the server analyzes the text data using a natural language processing (NLP) model to estimate the symptoms.

[0713] Emotion engine recognizes emotional states

[0714] The server uses an emotion engine to recognize the user's emotional state using text data. The emotion engine analyzes keywords and context in the text to determine the user's emotional state, such as stress, anxiety, or pain. For example, if the text contains expressions such as "It hurts so much" or "I'm so anxious," the corresponding emotional state is recognized.

[0715] Predictive diagnosis using generative AI models

[0716] The generative AI model on the server combines the results of image analysis, text analysis, and emotion engine analysis to make a probable diagnosis. This diagnosis includes the estimated medical condition, its reliability, and the user's emotional state. For example, if image analysis suggests a fracture, text analysis identifies symptoms such as "swelling and pain," and the emotion engine recognizes "strong anxiety," the generative AI model will generate a probable fracture diagnosis.

[0717] Server selection of appropriate medical institution

[0718] Based on the estimated diagnosis, the server references the user's current location information and searches a database for the most suitable medical institution. The information on the medical institution includes location, contact information, and available consultation hours. The server also takes into account the user's emotional state. For example, if the server detects strong anxiety or stress, it will prioritize medical institutions that offer psychological counseling.

[0719] Sending results from the server to the device

[0720] The server packages the estimated diagnosis and information on the recommended medical institution, encrypts it again, and transmits it to the terminal.

[0721] Displaying results on a terminal

[0722] The device decodes the received data and displays it in a user-friendly format, including the estimated diagnosis and detailed information about the recommended medical institution (e.g., location, contact information, and opening hours). The user can then review the displayed information and contact the appropriate hospital to arrange an appointment.

[0723] Specific examples

[0724] For example, if a user has a hand injury, the following process occurs:

[0725] 1. A user takes a photo of a hand injury with their smartphone and enters, "My hand is swollen and painful, and I'm very anxious."

[0726] 2. The device packages the image data and text data in JSON format and sends it to the server.

[0727] 3. The server analyzes the images using an AI model and the text data using an NLP model and emotion engine.

[0728] 4. The server integrates the results of each analysis and generates a diagnosis of "high probability of fracture" and an emotional state indicating "high anxiety."

[0729] 5. The server selects nearby orthopedic hospitals and medical institutions that provide psychological counseling, taking into account the user's location information and emotional state.

[0730] 6. The server sends the diagnosis results and medical institution information to the terminal.

[0731] 7. The device displays the received information to the user, who then contacts the appropriate medical institution and arranges for an examination.

[0732] The above is a specific embodiment of the present invention. This system allows users to receive prompt and accurate diagnosis and access to medical institutions, taking into account not only their medical condition but also their emotional state.

[0733] The processing flow will be explained below.

[0734] Step 1:

[0735] The user takes an image of the medical condition or injury using a smartphone or computer, saves the image data on the device, and then enters a description of their symptoms or condition into a text input form. For example, they might enter "My hand is swollen and painful."

[0736] Step 2:

[0737] The device packages the image and text data entered by the user into a specified format (e.g., JSON or XML), encrypts the data, and prepares it for transmission.

[0738] Step 3:

[0739] The device sends the encrypted image data and text data to the server, and the data is transmitted securely over the Internet.

[0740] Step 4:

[0741] The server decrypts the received encrypted data and prepares it for analysis. The server passes image data to the image analysis module and text data to the natural language processing (NLP) module.

[0742] Step 5:

[0743] The server analyzes the image data using an image analysis module. Specifically, it uses a convolutional neural network (CNN) to extract image features and estimate the condition of the patient. For example, it can detect signs of a fracture from the image.

[0744] Step 6:

[0745] The server analyzes the text data using a natural language processing (NLP) module. It extracts symptom-related keywords and context from the text and analyzes the symptoms. For example, it extracts information such as "swelling" and "pain."

[0746] Step 7:

[0747] The server uses an emotion engine to recognize the user's emotional state from the text data. The emotion engine analyzes keywords and context in the text to determine the user's emotional state (e.g., anxiety, stress, pain level, etc.).

[0748] Step 8:

[0749] The generative AI model combines the results of image analysis, text analysis, and emotion engine analysis to produce a probable diagnosis, which includes an estimated medical condition, its confidence level, and the user's emotional state.

[0750] Step 9:

[0751] Based on the estimated diagnosis, the server references the user's current location and searches a database for the most suitable medical institution. The server also takes into account the user's emotional state when selecting an appropriate medical institution. For example, if strong anxiety or stress is detected, it will prioritize medical institutions that offer psychological counseling.

[0752] Step 10:

[0753] The server packages the estimated diagnosis results and information about the selected medical institution, encrypts them again, and sends them to the terminal.

[0754] Step 11:

[0755] The device decodes the received data and displays it in a user-friendly format, including the estimated diagnosis and detailed information about the recommended medical institution (e.g., location, contact information, consultation hours, etc.).

[0756] Step 12:

[0757] The user checks the medical institution information displayed on the device and contacts the appropriate hospital to arrange for a medical examination. For example, the user contacts a hospital in response to a message that reads, "There is a high possibility of a fracture. We recommend the following orthopedic hospital."

[0758] Example 2

[0759] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0760] Currently, many people have difficulty accurately and quickly communicating information about their illness or injury to medical institutions. In particular, it is difficult for users to accurately describe their symptoms and emotional state, which can delay the selection of an appropriate medical institution and diagnosis. To solve this problem, a system is needed that can effectively analyze user input data, perform a comprehensive diagnosis including emotional state, and quickly provide guidance to an appropriate medical institution.

[0761] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0762] In this invention, the server includes a means for analyzing image data input by a user to estimate a medical condition, a means for analyzing text data to estimate symptoms, and a means for recognizing the user's emotional state from the text data using an emotion engine. This enables a highly accurate inferential diagnosis to be made by comprehensively evaluating the image analysis results, text analysis results, and emotion analysis results. Furthermore, based on this inferential diagnosis, the server can quickly select an appropriate medical institution by referring to the user's location information and provide that information to the user, thereby enabling the provision of prompt and appropriate medical services.

[0763] "User" refers to an individual who uses the system to input information about a medical condition or injury.

[0764] "Terminal" refers to a device used by a user to input image and text data related to a medical condition or injury and send it to a server. Examples of such devices include smartphones and computers.

[0765] "Image data" refers to photographic data of illnesses or injuries that users take using their terminals and enter into the system.

[0766] "Text data" refers to textual information such as the status and impressions of a medical condition or injury that a user inputs into a device.

[0767] A "format" refers to the rules and forms for organizing and structuring data, such as JSON and XML.

[0768] A "server" is a central computer system for receiving and analyzing image data and text data sent from the terminals.

[0769] A "generative AI model" is a collection of algorithms and programs that run on a server and integrate the results of image analysis and text analysis to make a presumptive diagnosis.

[0770] A "presumed diagnosis" is the evaluation result of the disease state or symptoms calculated by the generative AI model, and is a preliminary assessment before an actual diagnosis is made.

[0771] An "emotion engine" is a software component that analyzes text data and recognizes the user's emotional state.

[0772] "Location information" is geographical data that indicates the user's current location, and is obtained from GPS data or the like.

[0773] "Medical institution" refers to a clinic, doctor's office, hospital, etc. that provides medical services to users.

[0774] "Packaging" is the process of combining multiple pieces of data into one format, and is done to maintain data consistency.

[0775] "Encryption" is the process of transforming data using a specific algorithm to protect the confidentiality of the data.

[0776] "Decryption" is the process of restoring encrypted data to its original form.

[0777] "Analysis" is the process of analyzing data and extracting meaningful information and patterns from it.

[0778] This invention is a system that allows users to input image and text data about their illness or injury and receive referrals to appropriate medical institutions, even without specialized medical knowledge. Furthermore, by combining this system with an emotion engine that recognizes the user's emotions, it is possible to make diagnoses and select medical institutions that take the user's emotional state into consideration.

[0779] Hardware and software used

[0780] The system consists of the following components:

[0781] User terminal: A device through which a user inputs and receives data, such as a smartphone or computer.

[0782] Server: A central computer system that analyzes data and selects medical institutions.

[0783] Generative AI models: Algorithms that perform data analysis, such as convolutional neural networks (CNNs) and natural language processing (NLP) models.

[0784] Emotion Engine: A software component for recognizing the emotional state of a user from their text data.

[0785] Processing flow

[0786] A specific embodiment for implementing this system will be described below.

[0787] User input of images and text

[0788] Users use their smartphones or computers to take pictures of their illnesses or injuries and save the image data on their devices. They then enter their symptoms and thoughts about their condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text, "My hand is swollen and painful."

[0789] Sending data by the device

[0790] The device converts the image data and text data entered by the user into a specified format (e.g., JSON format), packages it, encrypts it, and sends it to a server via the Internet.

[0791] Receipt and analysis of data by the server

[0792] The server receives and decodes the image and text data sent from the device. The server then uses a generative AI model to analyze the image data. Specifically, it uses a convolutional neural network (CNN) to extract image features and infer the disease state. At the same time, the server analyzes the text data with a natural language processing (NLP) model to infer the symptoms.

[0793] Emotion engine recognizes emotional states

[0794] The server uses an emotion engine to recognize the user's emotional state based on the text data. The emotion engine analyzes keywords and context within the text to determine the user's emotional state, such as stress, anxiety, or pain. For example, if the text contains expressions such as "It hurts so much" or "I'm so anxious," the corresponding emotional state is recognized.

[0795] Predictive diagnosis using generative AI models

[0796] The generative AI model on the server integrates the results of image analysis, text analysis, and emotion engine analysis to make a probable diagnosis. This diagnosis includes the estimated medical condition, its reliability, and the user's emotional state. For example, if image analysis suggests a fracture, text analysis identifies symptoms such as "swelling and pain," and the emotion engine recognizes "strong anxiety," the generative AI model will generate a probable fracture diagnosis.

[0797] Server selection of appropriate medical institution

[0798] The server then references the user's current location information based on the estimated diagnosis and searches a database for the most suitable medical institution. The information on the medical institution includes the location, contact information, and available consultation hours. The server also takes into account the user's emotional state. For example, if the server detects strong anxiety or stress, it will prioritize medical institutions that offer psychological counseling.

[0799] Sending results from the server to the device

[0800] The server packages the estimated diagnosis results and information on recommended medical institutions, encrypts them again, and sends them to the terminal.

[0801] Displaying results on a terminal

[0802] The device decodes the received data and displays it in a user-friendly format, including the estimated diagnosis and detailed information about the recommended medical institution (such as location, contact information, and opening hours). The user can then review the displayed information and contact the appropriate hospital to arrange an appointment.

[0803] Specific examples

[0804] For example, if a user has a hand injury, the following process occurs:

[0805] 1. A user takes a photo of a hand injury with their smartphone and enters, "My hand is swollen and painful, and I'm very anxious."

[0806] 2. The device packages the image data and text data in JSON format and sends it to the server.

[0807] 3. The server analyzes the images using an AI model and the text data using an NLP model and emotion engine.

[0808] 4. The server integrates the results of each analysis and generates a diagnosis of "high probability of fracture" and an emotional state indicating "high anxiety."

[0809] 5. The server selects nearby orthopedic hospitals and medical institutions that provide psychological counseling, taking into account the user's location information and emotional state.

[0810] 6. The server sends the diagnosis results and medical institution information to the terminal.

[0811] 7. The device displays the received information to the user, who then contacts the appropriate medical institution and arranges for an examination.

[0812] This process allows users to access medical care and receive a quick and accurate diagnosis that takes into account not only their medical condition but also their emotional state.

[0813] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0814] Step 1:

[0815] User input of images and text

[0816] Users use a smartphone or computer to take pictures of their condition or injury, save the image data on the device, and then enter detailed symptoms and impressions of their condition in text format.

[0817] Input: Image data and text data related to medical conditions and injuries

[0818] Output: Image data and text data stored on the device

[0819] Specific behavior:

[0820] The user activates the camera and takes a picture of the affected area.

[0821] After taking the photo, you upload the image through the device's application and enter details such as "My hand is swollen and painful, and I'm very anxious" in the text input field.

[0822] Step 2:

[0823] Sending data by the device

[0824] The device converts the image data and text data entered by the user into a specified format (e.g., JSON format) and packages it. The device then encrypts the data and sends it to the server via the Internet.

[0825] Input: User image and text data

[0826] Output: Encrypted packaging data

[0827] Specific behavior:

[0828] The device converts the image file and text data into JSON format.

[0829] The data is encrypted using the SSL / TLS protocol and sent to the server's API endpoint.

[0830] Step 3:

[0831] Receipt and analysis of data by the server

[0832] The server receives and decrypts the encrypted data sent from the device. Next, the server analyzes the received image data using a generative AI model (CNN algorithm) to extract image features. At the same time, the server analyzes the text data using a natural language processing (NLP) model to estimate symptoms.

[0833] Input: Encrypted packaging data

[0834] Output: Analyzed image features and text analysis results

[0835] Specific behavior:

[0836] The server reads the encrypted data from the SSL / TLS input buffer and decrypts it.

[0837] The server inputs the image data into a generative AI model, extracts features, and estimates the condition of the disease.

[0838] An NLP processing engine is used to analyze the text data and extract keywords and context related to the symptoms.

[0839] Step 4:

[0840] Emotion engine recognizes emotional states

[0841] The server analyzes the text data with an emotion engine to recognize the user's emotional state. The emotion engine analyzes keywords and context within the text to determine the user's stress, anxiety, pain, etc.

[0842] Input: Text data

[0843] Output: Data about the user's emotional state

[0844] Specific behavior:

[0845] The server inputs the text data into an emotion engine, which analyzes the keywords and context within the text.

[0846] The emotion engine analyzes expressions such as "It hurts so much" and "I'm so anxious" to determine the user's emotional state.

[0847] Step 5:

[0848] Predictive diagnosis using generative AI models

[0849] The generative AI model on the server integrates the results of image analysis, text analysis, and emotion engine analysis to produce a probable diagnosis, which includes the probable medical condition, its reliability, and the user's emotional state.

[0850] Input: Image analysis results, text analysis results, sentiment analysis results

[0851] Output: Estimated diagnosis result

[0852] Specific behavior:

[0853] The server integrates the results of image analysis, text analysis, and sentiment analysis.

[0854] The generative AI model estimates the condition based on the integrated data and generates a diagnosis such as "high probability of fracture."

[0855] Step 6:

[0856] Server selection of appropriate medical institution

[0857] Based on the estimated diagnosis, the server references the user's current location information and searches a database for the most suitable medical institution, including information on the location, contact details, and available hours for consultations.

[0858] Input: Estimated diagnosis result, user's current location information

[0859] Output: Appropriate medical institution information

[0860] Specific behavior:

[0861] The server obtains the user's current location information from GPS data and queries a location information database.

[0862] The server searches the medical institution database and generates a list of medical institutions that can provide consultations.

[0863] Step 7:

[0864] Sending results from the server to the device

[0865] The server packages the estimated diagnosis and information on the recommended medical institution, encrypts it again, and transmits it to the terminal.

[0866] Input: Estimated diagnosis result, medical institution information

[0867] Output: Encrypted diagnosis results and medical institution information

[0868] Specific behavior:

[0869] The server packages the diagnosis results and medical institution information in JSON format.

[0870] The server encrypts the packaged data and sends it to the user's terminal.

[0871] Step 8:

[0872] Displaying results on a terminal

[0873] The device decodes the received data and displays it in a user-friendly format, including a probable diagnosis and detailed information about recommended medical institutions.

[0874] Input: Encrypted diagnosis results and medical institution information

[0875] Output: Decrypted diagnosis and medical institution information

[0876] Specific behavior:

[0877] The device decodes the data it receives and parses the JSON format data to extract the diagnosis results and medical institution information.

[0878] The device updates the application UI and displays the diagnosis results and medical institution information to the user.

[0879] Through this detailed, step-by-step process, users can obtain an accurate and prompt diagnosis and access to the appropriate medical facility.

[0880] (Application example 2)

[0881] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0882] Conventional medical diagnosis systems have the problem that it is difficult for users without medical expertise to make an appropriate diagnosis and select a medical institution. Furthermore, there is a need for rapid medical diagnosis and guidance to medical institutions in emergency situations. In particular, a system that can provide appropriate information immediately is needed so that security staff can respond quickly on the scene.

[0883] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0884] In this invention, the server includes: a means for a user to input image data and text data related to a medical condition or injury; a means for a terminal to package the image data and text data into a specified format and transmit the packaged image data and text data to the server; a means for the server to analyze the image data and estimate a medical condition; a means for the server to analyze the text data and estimate symptoms; a generative AI model to integrate the analysis results and make an estimated diagnosis; a means for the server to select an appropriate medical institution based on the estimated diagnosis; a means for the server to transmit information about the estimated diagnosis and the medical institution to the terminal; a means for the terminal to display the information to the user; and a means for performing a rapid medical diagnosis on site using smart glasses. This enables a user (security staff) to receive a rapid and appropriate medical diagnosis on site and access an appropriate medical institution based on the results.

[0885] "Image data" is digital data entered by a user that contains visual information about a medical condition or injury.

[0886] "Text data" is digital data containing written information related to a medical condition or injury entered by a user.

[0887] A "user" is an individual who uses the system and who inputs information about a medical condition or injury.

[0888] A "terminal" is a device operated by a user, and is a device for packaging image data and text data and transmitting them to a server.

[0889] A "server" is a computer system that receives data sent from a terminal and performs analysis and inferential diagnosis.

[0890] A "generative AI model" is an artificial intelligence algorithm that integrates the analysis results of image data and text data to make a presumptive diagnosis.

[0891] "Format" refers to a data format for arranging image data and text data in a form that is easy for the server to process.

[0892] "Smart glasses" are wearable devices equipped with a camera function that users use in the field.

[0893] An "appropriate medical institution" is a medical facility that can provide optimal medical services to the user based on the estimated diagnosis results.

[0894] A "presumed diagnosis" is the result of a disease state inferred by a generative AI model based on an analysis of image data and text data.

[0895] "Display means" refers to a function that visually presents to the user the information that the terminal receives from the server.

[0896] "Emotional state" indicates the psychological and emotional state of the user, which is analyzed from text data or the like.

[0897] In order to implement the present invention, the following configurations and means are necessary.

[0898] Overall system overview

[0899] The user inputs image and text data of their condition or injury and sends the information to a server. The server analyzes the data, makes a probable diagnosis, selects an appropriate medical institution, and provides that information to the user. Furthermore, the use of smart glasses enables rapid medical diagnosis on the spot.

[0900] Hardware and Software

[0901] 1. Smart Glasses:

[0902] A wearable device with a camera function.

[0903] Supports on-site image capture and data entry.

[0904] 2. Terminal:

[0905] Personal computers and smartphones.

[0906] It supports the input of image data and text data and sends the data to the server.

[0907] 3. Server:

[0908] High performance computing systems.

[0909] Analyzes image and text data, performs presumptive diagnoses, and selects appropriate medical institutions.

[0910] The algorithms used are convolutional neural networks (CNN) and natural language processing (NLP).

[0911] Program processing

[0912] The device packages the image data and text data entered by the user into a specified format (e.g., JSON), encrypts it, and sends it to the server. The server decrypts the received data and performs the following processes.

[0913] 1. Analysis of imaging data:

[0914] The server uses a convolutional neural network (CNN) to extract features from the image data and estimate the disease state.

[0915] 2. Text data analysis:

[0916] Text data is analyzed using a natural language processing (NLP) model to estimate symptoms.

[0917] 3. Integrated analysis using generative AI models:

[0918] The generative AI model integrates the results of image analysis, text analysis, and the emotion engine analysis to make a presumptive diagnosis.

[0919] 4. Selection of appropriate medical facility:

[0920] Based on the estimated diagnosis results, the server selects the most appropriate medical institution, taking into consideration the user's location information and emotional state.

[0921] 5. Transmission and Display of Information:

[0922] The server transmits the generated estimated diagnosis and information about the medical institution to the terminal, which then displays them to the user.

[0923] Specific examples

[0924] For example, if security staff discover someone collapsed at a scene, they can take the following steps:

[0925] 1. Taking a photo of a fallen person using smart glasses.

[0926] 2. Enter the text data "Unconscious, not breathing."

[0927] 3. The device sends this data to the server.

[0928] 4. The server makes a probable diagnosis and generates information such as "possibility of cardiac arrest" and "nearest hospital with a cardiologist."

[0929] 5. The server sends the diagnosis results and information on recommended medical institutions to the terminal, which displays them to the user.

[0930] Prompt Sentence Examples

[0931] "Unconscious, not breathing"

[0932] "My hands are swollen and painful, and I'm very anxious"

[0933] By using these prompts, users can quickly and accurately access the appropriate medical institution.

[0934] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0935] Step 1:

[0936] The user uses the smart glasses to take images of their condition or injury and input text data about their symptoms. The user uses the smart glasses' camera function to take photos of their condition or injury, and then inputs text information such as "My hand is swollen and painful, and I feel very anxious" through the smart glasses' UI.

[0937] Input: Image and text data of medical conditions and injuries.

[0938] Output: Photographed image data and input text data.

[0939] Step 2:

[0940] The device packages the image and text data entered by the user into a specified format, encrypts it, and sends it to the server.The device then converts the data into a structured data format such as JSON and encrypts it using an encryption algorithm such as AES for security.The encrypted data is then sent to the server via the HTTPS protocol.

[0941] Input: Image data and text data.

[0942] Output: Encrypted and formatted data.

[0943] Step 3:

[0944] The server receives and decrypts the data sent from the device. The server receives the data via the HTTPS protocol and decrypts data encrypted with AES or other encryption protocols. It then analyzes the received JSON format data and separates the image data from the text data.

[0945] Input: Encrypted and formatted data.

[0946] Output: Decoded and separated image and text data.

[0947] Step 4:

[0948] The server analyzes the image data to estimate the condition of the patient. Specifically, it uses a convolutional neural network (CNN) to extract image features and estimate the condition. Based on the results of this analysis, a diagnosis such as "high possibility of fracture" is obtained.

[0949] Input: Image data.

[0950] Output: Estimated disease state based on image analysis.

[0951] Step 5:

[0952] The server analyzes the text data to infer symptoms. It uses a natural language processing (NLP) model to analyze keywords and context within the text data to infer symptoms. For example, from an expression such as "My hand is swollen and painful, and I feel very anxious," it extracts symptoms such as "swelling," "pain," and "anxiety."

[0953] Input: Text data.

[0954] Output: Symptom prediction results based on text analysis.

[0955] Step 6:

[0956] The generative AI model combines the results of image analysis, text analysis, and the emotion engine to generate a comprehensive diagnosis, such as "There is a high possibility of a fracture and the patient feels high anxiety."

[0957] Input: Image analysis results, text analysis results, emotion engine results.

[0958] Output: Integrated estimated diagnostic results.

[0959] Step 7:

[0960] The server selects an appropriate medical institution based on the user's location information and emotional state. Based on the estimated diagnosis, the server searches a database of medical institutions and selects the most suitable medical institution. For example, the server selects the "nearest orthopedic hospital" and "medical institution that offers psychological counseling" based on the user's location information and emotional state.

[0961] Input: Estimated diagnosis result, user location information, emotional state.

[0962] Output: Information on selected medical institutions.

[0963] Step 8:

[0964] The server then sends the generated estimated diagnosis and information about the medical institution to the terminal, where it re-encrypts the information and transmits it to the terminal using the HTTPS protocol.

[0965] Input: Estimated diagnosis result, medical institution information.

[0966] Output: Encrypted diagnosis results and medical institution information.

[0967] Step 9:

[0968] The device decrypts the information received from the server and displays it in a format that is easy for the user to understand. The device decrypts the encrypted data and displays information such as "High probability of fracture," "Nearest orthopedic hospital: XX Hospital," and "Psychological counseling: YY Clinic" through the UI.

[0969] Input: Encrypted diagnosis results and medical institution information.

[0970] Output: Presenting information in a user-friendly format.

[0971] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0972] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0973] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0974] [Third embodiment]

[0975] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0976] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0977] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0978] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0979] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0980] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0981] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0982] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0983] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0984] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0985] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0986] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0987] This invention is a system that allows users to input image and text data related to a medical condition or injury and receive a referral to an appropriate medical institution even if they do not have specialized medical knowledge. Specific embodiments of this system will be described below.

[0988] User input of images and text

[0989] Users can use a smartphone or computer to take pictures of their illness or injury and save the image data on the device. They can then enter their symptoms and thoughts about the condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text, "My hand is swollen and painful."

[0990] Sending data by the device

[0991] The device packages the image data and text data entered by the user into a specified format (for example, JSON or XML).The device then sends this data to a server via the Internet.At this time, the data is encrypted to ensure security.

[0992] Receipt and analysis of data by the server

[0993] The server receives and decodes the image and text data sent from the device. It then uses a generative AI model to analyze the received image data. Specifically, it uses algorithms such as convolutional neural networks to extract image features and estimate the condition. At the same time, the server analyzes the text data using a natural language processing (NLP) model to estimate symptoms.

[0994] Predictive diagnosis using generative AI models

[0995] The generative AI model on the server integrates the results of image analysis and text analysis to generate a probable diagnosis. This diagnosis includes a probable pathology and its reliability. For example, if image analysis detects a bone abnormality and text analysis identifies the symptom "swelling and pain," the probable diagnosis is "high probability of fracture."

[0996] Server selection of appropriate medical institution

[0997] Based on the estimated diagnosis, the server refers to the user's location information and selects the most appropriate medical institution. At this time, the server retrieves a list of affiliated medical institutions from the database and lists medical institutions close to the user's current location. The information on the selected medical institution includes location, contact information, available hours for consultation, etc.

[0998] Sending results from the server to the device

[0999] The server then sends the estimated diagnosis and information on the recommended medical institution to the device. The data is then re-encrypted and transferred to the device in a secure manner.

[1000] Displaying results on a terminal

[1001] The device decodes the received data and displays it in a user-friendly format. The user can then check the displayed estimated diagnosis and recommended medical institution information and make arrangements for a medical examination if necessary. For example, if there is a high possibility of a fracture, the user can call the orthopedic hospital contact number displayed on the device to make an appointment.

[1002] Specific examples

[1003] For example, if a user injures their hand, the following process occurs:

[1004] 1. A user takes a photo of a hand injury with their smartphone and types, "My hand is swollen and painful."

[1005] 2. The device packages the image data and text data in JSON format and sends it to the server.

[1006] 3. The server analyzes the image using an AI model and the text data using an NLP model.

[1007] 4. The server integrates the analysis results and generates a diagnosis of "high probability of fracture."

[1008] 5. The server selects a nearby orthopedic hospital based on the user's location information.

[1009] 6. The server sends the diagnosis results and hospital information to the terminal.

[1010] 7. The terminal displays the received information to the user, who then contacts the orthopedic hospital to arrange for a consultation.

[1011] The above is a specific embodiment of the present invention. This system makes it possible to quickly and accurately estimate the condition of a patient, even without specialized medical knowledge, and to assist in accessing an appropriate medical institution.

[1012] The processing flow will be explained below.

[1013] Step 1: The user takes an image of their condition or injury using a smartphone or computer, saves the image data on their device, and enters a description of their symptoms or condition into a text input form.

[1014] Step 2: The terminal converts the image data and text data entered by the user into a specified format (e.g., JSON or XML) and packages it.

[1015] Step 3: The device encrypts the packaged data and sends it to a server via the Internet.

[1016] Step 4: The server decodes the received image data and text data and prepares each for analysis.

[1017] Step 5: The server analyzes the image data using an image analysis module. Specifically, it uses a convolutional neural network (CNN) to extract image features and estimate the disease state.

[1018] Step 6: The server analyzes the text data using a natural language processing (NLP) module, extracting symptom-related keywords and context from the text and obtaining the analysis results.

[1019] Step 7: The generative AI model on the server integrates the results of image analysis and text analysis to make a probable diagnosis. For example, if image analysis indicates a suspected fracture and text analysis indicates pain or swelling, the model will generate a probable diagnosis of "high probability of fracture."

[1020] Step 8: Based on the estimated diagnosis, the server references the user's current location information and searches the database for the most suitable medical institution. Information on the target medical institutions is then listed.

[1021] Step 9: The server packages the estimated diagnosis and the information on the selected medical institution, encrypts it again, and transmits it to the terminal.

[1022] Step 10: The terminal decrypts the received data and displays it to the user, including the estimated diagnosis and detailed information (such as location, contact information, and consultation hours) of the recommended medical institution.

[1023] Step 11: The user checks the information about the medical institution displayed on the terminal and contacts the appropriate hospital to arrange for a medical examination.

[1024] Example 1

[1025] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1026] In today's world, it is difficult for users to quickly and accurately communicate information about their illness or injury to specialized medical institutions. It is also difficult for users without medical knowledge to select an appropriate medical institution. As a result, it can take a lot of time and effort to receive an appropriate diagnosis and treatment. To solve these problems, a system is needed that can effectively collect information about a user's illness or injury, analyze it quickly and accurately, and select an appropriate medical institution.

[1027] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1028] In this invention, the server includes: means for a user to input image data and text data related to a medical condition or injury; means for a terminal to package the image data and text data in a specified format and transmit the packaged image data and text data to the server; means for the terminal to encrypt the data; means for the server to receive and decrypt the image data and text data; means for the server to analyze the image data, extract features, and infer a medical condition; means for the server to analyze the text data, extract and infer symptoms; means for a generative AI model to integrate the analysis results and make a probable diagnosis; means for the server to select an appropriate medical institution based on the probable diagnosis; means for the server to transmit information about the probable diagnosis and medical facilities to the terminal; and means for the terminal to display the information to the user. This enables users to quickly and accurately understand their own medical condition or injury and to promptly visit an appropriate medical institution, even without specialized medical knowledge.

[1029] "Image data" is digital information that a user visually records regarding a medical condition or injury.

[1030] "Text data" is digital information in which a user records symptoms and thoughts about a medical condition or injury in written form.

[1031] A "terminal" is a device that a user uses to input data and send it to a server, and includes a smartphone or computer.

[1032] A "server" is a central processing device that receives data sent from a terminal, analyzes it, and returns the necessary information.

[1033] A "specified format" is a format in which data is arranged according to certain rules, and includes, for example, JSON and XML.

[1034] "Encryption" is a technique that uses a specific algorithm to convert data in order to transmit it securely.

[1035] "Decryption" is the technique of returning encrypted data to its original form.

[1036] "Features" are important attributes or patterns extracted from image data and are used to estimate the condition of a disease.

[1037] A "natural language processing (NLP) model" is an algorithm or method for analyzing text data and understanding human language.

[1038] A "generative AI model" is an artificial intelligence algorithm that integrates the results of image analysis and text analysis to make a presumptive diagnosis.

[1039] A "presumptive diagnosis" is the result of predicting the condition of a disease or injury based on analyzed data.

[1040] A "medical institution" is a facility where a user can receive medical examinations and treatment, and includes hospitals and clinics.

[1041] "Information" refers to data such as the diagnosis results and contact details and locations of medical institutions required by the user.

[1042] The present invention provides a system that allows a user to input image data and text data relating to a medical condition or injury, and easily refer the user to an appropriate medical institution. A specific embodiment of this system will be described below.

[1043] User data entry

[1044] Users use devices such as smartphones or computers to take pictures of their medical condition or injury and save them on the device. Next, they enter text data about their condition (e.g., "My hand is swollen and painful") into the device. At this stage, users do not need specialized medical knowledge and can enter data intuitively.

[1045] Device prepares to send data

[1046] The device packages the image and text data entered by the user into a specified format (e.g., JSON or XML), then encrypts the packaged data and transmits it securely to a server over the Internet.

[1047] Receipt and analysis of data by the server

[1048] The server receives the encrypted data sent from the device. After receiving the data, the server decrypts it and extracts the image data and text data. The server then analyzes the image data using a convolutional neural network (CNN) to extract image features. The server also analyzes the text data using a natural language processing (NLP) model to extract the described symptoms.

[1049] Predictive diagnosis using generative AI models

[1050] The generative AI model deployed on the server integrates the results of image analysis and text analysis. This model then uses the analysis results to generate a probable diagnosis, which includes a probable medical condition and its reliability. For example, by combining an image of a hand injury with the text data "your hand is swollen and painful," the model generates a probable diagnosis of "high probability of fracture."

[1051] Selection of medical institutions by the server

[1052] The server then references the user's location information based on the estimated diagnosis and selects an appropriate medical institution from a database. The information on the selected medical institution includes the location, contact information, and available consultation hours. The server then organizes this information, re-encrypts the data, and sends it to the device.

[1053] Displaying results on a terminal

[1054] The device receives and decrypts the encrypted data sent from the server. The decrypted data is displayed in a user-friendly format, allowing the user to refer to it and contact the appropriate medical institution and arrange for a medical examination. For example, it may display information such as "There is a high possibility of a fracture. The recommended medical institution is XX Orthopedic Hospital, contact number: XX-XX-XX."

[1055] Examples of concrete examples and prompts

[1056] As a concrete example, if a user injures their hand, the process is as follows:

[1057] 1. A user takes a photo of a hand injury with their smartphone and types, "My hand is swollen and painful."

[1058] 2. The device packages the image data and text data in JSON format and sends it to the server.

[1059] 3. The server analyzes the image using an AI model and the text data using an NLP model.

[1060] 4. The server integrates the analysis results and generates a diagnosis of "high probability of fracture."

[1061] 5. The server selects a nearby orthopedic hospital based on the user's location information.

[1062] 6. The server sends the diagnosis results and hospital information to the terminal.

[1063] 7. The terminal displays the received information to the user, who then contacts the orthopedic hospital to arrange for a consultation.

[1064] This system allows even users without specialized medical knowledge to quickly and accurately understand the patient's condition and refer them to the appropriate medical institution.

[1065] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1066] Step 1:

[1067] The user uses a smartphone or computer to take pictures of their condition or injury and save them on the device, then enters text about their condition (e.g., "My hand is swollen and painful") into the device.

[1068] Input: Image data and text data related to medical conditions and injuries

[1069] Output: Image data and text data stored on the device

[1070] Step 2:

[1071] The terminal acquires image data and text data input by the user.

[1072] Input: Image data and text data stored on the device

[1073] Output: Acquired image data and text data

[1074] Step 3:

[1075] The image data and text data acquired by the device are packaged in a specified format such as JSON.

[1076] Input: Acquired image data and text data

[1077] Output: Data packaged in the specified format

[1078] Step 4:

[1079] The device encrypts the packaged data and prepares it in a highly secure state.

[1080] Input: Data packaged in a specified format

[1081] Output: Encrypted data

[1082] Step 5:

[1083] The terminal transmits the encrypted data to a server via the Internet.

[1084] Input: Encrypted data

[1085] Output: Data sent to the server

[1086] Step 6:

[1087] The server receives the encrypted data sent from the terminal.

[1088] Input: Data sent over the internet

[1089] Output: Encrypted data received by the server

[1090] Step 7:

[1091] The server decrypts the received encrypted data and extracts the image data and text data.

[1092] Input: Received encrypted data

[1093] Output: Decoded image data, text data

[1094] Step 8:

[1095] The server analyzes the image data using a convolutional neural network (CNN) and extracts features.

[1096] Input: Decoded image data

[1097] Output: Image features

[1098] Step 9:

[1099] The server analyzes the text data using a natural language processing (NLP) model to extract symptoms.

[1100] Input: Decrypted text data

[1101] Output: Extracted symptoms

[1102] Step 10:

[1103] The server uses the generated AI model to integrate the results of image analysis and text analysis and make a presumptive diagnosis.

[1104] Input: Image features, extracted symptoms

[1105] Output: Estimated diagnosis result (condition and its reliability)

[1106] Step 11:

[1107] Based on the estimated diagnosis, the server refers to the user's location information and selects an appropriate medical institution.

[1108] Input: Estimated diagnosis result, user location information

[1109] Output: Information about the selected medical institution (location, contact information, available hours, etc.)

[1110] Step 12:

[1111] The server re-encrypts the diagnosis results and medical institution information and sends them to the terminal.

[1112] Input: Information on the selected medical institution, diagnosis results

[1113] Output: Encrypted data for transmission

[1114] Step 13:

[1115] The terminal receives the encrypted data sent from the server.

[1116] Input: Encrypted data to be sent

[1117] Output: Received encrypted data

[1118] Step 14:

[1119] The terminal decrypts the received encrypted data and displays it in a user-friendly format.

[1120] Input: Received encrypted data

[1121] Output: Diagnosis results and medical institution information displayed to the user

[1122] Through the above steps, the user can quickly and accurately understand the condition of their illness or injury and promptly seek medical attention at an appropriate medical institution.

[1123] (Application example 1)

[1124] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1125] In recent years, users have faced the challenge of finding an appropriate medical institution for their illness or injury quickly and accurately without specialized medical knowledge. Furthermore, there is a lack of systems that allow users to instantly make appointments and make payments based on diagnosis results. This creates a problem of inability to smoothly access and use medical services.

[1126] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1127] In this invention, the server includes means for selecting an appropriate medical institution based on a presumed diagnosis and transmitting that information, means for making a reservation at the recommended medical institution, and means for completing payment by electronic payment. This enables a user to quickly find an appropriate medical institution based on their condition or injury and seamlessly complete a series of processes from reservation to payment.

[1128] A "user" is someone who uses the system to input information about their medical condition or injury and wishes to receive recommendations or make appointments with medical institutions.

[1129] "Image data relating to a medical condition or injury" refers to photographs or image files taken by a user that show the state of a medical condition or injury.

[1130] "Text data" refers to character information entered by the user, including symptoms, impressions, and explanations of illnesses or injuries.

[1131] A "terminal" is a device, such as a smartphone or computer, that allows a user to access the system and input and send data.

[1132] A "server" is a computer system that receives data sent by users, analyzes it, selects medical institutions, and sends information.

[1133] A "specified format" is a format used to organize image data or text data into a certain format, such as JSON or XML.

[1134] "Encryption" is the process of converting data into a format that cannot be read by third parties in order to prevent data leakage during communication.

[1135] A "generative AI model" is a system that includes an artificial intelligence algorithm for analyzing and presumptively diagnosing medical conditions and symptoms using image and text data.

[1136] A "presumed diagnosis" is a provisional diagnosis derived from the results of the disease state and symptoms analyzed by the generative AI model.

[1137] "Selection of medical institution" refers to the procedure in which the server identifies an appropriate medical institution based on the estimated diagnosis result and taking into consideration the user's location information.

[1138] The "reservation means" is part of the system that allows the user to make an appointment with a recommended medical institution, and is a function that manages reservation information in cooperation with the server.

[1139] "Electronic payment" is a payment system conducted via the Internet, and refers to the use of online payment methods such as credit cards and electronic money.

[1140] This invention is a system that allows users to input image and text data about their medical condition or injury and receive referrals to appropriate medical institutions, even without specialized medical knowledge. This system is constructed using a terminal, a server, and a generative AI model.

[1141] First, the user takes a picture of their condition or injury using a smartphone or computer and saves the image data on the device. Next, the user enters their symptoms and thoughts about the condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text "My hand is swollen and painful."

[1142] The terminal packages the image data and text data entered by the user into a specified format (for example, JSON or XML). The terminal then sends this data to the server via the Internet. At this time, the data is encrypted to ensure security. For encryption, the pycryptodome library, for example, is used.

[1143] The server receives and decodes the image and text data sent from the device. It then uses a generative AI model to analyze the received image data. Specifically, it uses deep learning frameworks such as TensorFlow and PyTorch to extract image features using algorithms such as convolutional neural networks (CNNs) and estimate the condition of the patient. At the same time, the server uses a natural language processing (NLP) model to analyze the text data and estimate the symptoms.

[1144] The generative AI model combines the results of image analysis and text analysis to generate a probable diagnosis. This diagnosis includes a probable pathology and its reliability. For example, if image analysis detects a bone abnormality and text analysis identifies the symptom "swelling and pain," the probable diagnosis is "high probability of fracture."

[1145] The server selects the most appropriate medical institution based on the estimated diagnosis and the user's location information. At this time, the server retrieves a list of affiliated medical institutions from the database and lists the medical institutions closest to the user's current location. The information on the selected medical institution includes the location, contact information, and available hours for consultation.

[1146] The server then sends the estimated diagnosis and information on the recommended medical institution to the device. The data is then re-encrypted and transferred to the device in a secure manner.

[1147] The device decrypts the received data and displays it in a user-friendly format. The user can then check the displayed estimated diagnosis and information on recommended medical institutions, make an appointment with a medical institution within the app, and complete payment electronically. Payment can be made using APIs such as Stripe or PayPal.

[1148] Specific examples

[1149] For example, if a user injures their hand, the following process occurs:

[1150] 1. A user takes a photo of a hand injury with their smartphone and types, "My hand is swollen and painful."

[1151] 2. The device packages the image data and text data in JSON format, sends it to the server, and encrypts the data.

[1152] 3. The server analyzes the image using an AI model and the text data using an NLP model.

[1153] 4. The server integrates the analysis results and generates a diagnosis that "there is a high possibility of a fracture."

[1154] 5. The server selects a nearby orthopedic hospital based on the user's location information.

[1155] 6. The server sends the diagnosis results and medical institution information to the terminal.

[1156] 7. The terminal displays the received information to the user, who then contacts the orthopedic hospital within the app to make an appointment and completes payment electronically.

[1157] Prompt Sentence Examples

[1158] "Please provide a proper diagnosis and recommended medical care for the injury image taken by the user and the symptom text entered: 'My hand is swollen and painful.'"

[1159] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1160] Step 1:

[1161] The user takes an image of their condition or injury using a smartphone or computer and saves the image data on the device. The user then enters their symptoms and thoughts about the condition in text format. The input image data is a JPEG or PNG file, and the text data is a string of characters.

[1162] Step 2:

[1163] The device packages the image and text data entered by the user into the specified format. For example, this data is converted into JSON format. The packaged data is then encrypted to ensure security. The pycryptodome library is used for encryption, and the data is processed using the AES encryption algorithm.

[1164] Step 3:

[1165] The device sends the encrypted data to the server via the Internet in encrypted JSON format. The server receives the data using the HTTPS protocol to ensure secure communication.

[1166] Step 4:

[1167] The server decrypts the received data and converts it back into JSON image and text data, again using the pycryptodome library and the AES decryption algorithm, before formatting the decrypted data for analysis.

[1168] Step 5:

[1169] The server analyzes image data and text data. To analyze the image data, a deep learning framework such as TensorFlow or PyTorch is used to extract image features using a convolutional neural network (CNN). To analyze the text data, a natural language processing (NLP) model is used to infer symptoms from the text entered by the user. A generative AI model integrates these analysis results to make a probable diagnosis. For example, a CNN model detects bone abnormalities, and an NLP model analyzes the text "My hand is swollen and painful" and infers that there is a high possibility of a fracture.

[1170] Step 6:

[1171] The server selects the most suitable medical institution based on the estimated diagnosis and the user's location information. The server retrieves a list of affiliated medical institutions from an internal database and creates a list of medical institutions that can provide consultations. The server uses a geographic information system (GIS) to select the most suitable medical institution based on the user's location information.

[1172] Step 7:

[1173] The server encrypts the estimated diagnosis and information about the recommended medical institution and sends it back to the device. The data sent includes the estimated condition, reliability, name, location, contact information, and available hours of the recommended medical institution. The data is again AES encrypted and transmitted securely using the HTTPS protocol.

[1174] Step 8:

[1175] The device decrypts the received data and displays the diagnosis results and medical institution information so that the user can check them. Based on the displayed diagnosis results, the user can make an appointment with the recommended medical institution. The device application also has a function to contact the medical institution directly from the displayed information.

[1176] Step 9:

[1177] The user makes a reservation at a recommended medical institution from their device. They enter reservation information within the app and send a reservation request to the server. The server forwards the reservation information to the medical institution and confirms the reservation. The confirmed reservation information is sent to the user's device.

[1178] Step 10:

[1179] The user uses the terminal to make an electronic payment for a medical appointment. The terminal uses an API such as Stripe or PayPal to send the payment information to the server. The server processes the payment through a payment gateway and sends a confirmation to the user if the payment is successful.

[1180] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1181] This invention is a system that allows users to input image and text data related to their medical condition or injury and receive referrals to appropriate medical institutions, even without specialized medical knowledge. Furthermore, by combining this system with an emotion engine that recognizes the user's emotions, it is possible to make diagnoses and select medical institutions that take into account the user's emotional state. A specific embodiment of this system is described below.

[1182] User input of images and text

[1183] Users use their smartphones or computers to take pictures of their illnesses or injuries and save the image data on their devices. They then enter their symptoms and thoughts about their condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text, "My hand is swollen and painful."

[1184] Sending data by the device

[1185] The device converts the image data and text data entered by the user into a specified format (e.g., JSON or XML) and packages it. The device then encrypts the data and sends it to the server via the Internet.

[1186] Receipt and analysis of data by the server

[1187] The server receives and decodes the image and text data sent from the device. It then uses a generative AI model to analyze the received image data. Specifically, it uses algorithms such as convolutional neural networks (CNNs) to extract image features and estimate the disease state. At the same time, the server analyzes the text data using a natural language processing (NLP) model to estimate the symptoms.

[1188] Emotion engine recognizes emotional states

[1189] The server uses an emotion engine to recognize the user's emotional state using text data. The emotion engine analyzes keywords and context in the text to determine the user's emotional state, such as stress, anxiety, or pain. For example, if the text contains expressions such as "It hurts so much" or "I'm so anxious," the corresponding emotional state is recognized.

[1190] Predictive diagnosis using generative AI models

[1191] The generative AI model on the server combines the results of image analysis, text analysis, and emotion engine analysis to make a probable diagnosis. This diagnosis includes the estimated medical condition, its reliability, and the user's emotional state. For example, if image analysis suggests a fracture, text analysis identifies symptoms such as "swelling and pain," and the emotion engine recognizes "strong anxiety," the generative AI model will generate a probable fracture diagnosis.

[1192] Server selection of appropriate medical institution

[1193] Based on the estimated diagnosis, the server references the user's current location information and searches a database for the most suitable medical institution. The information on the medical institution includes location, contact information, and available consultation hours. The server also takes into account the user's emotional state. For example, if the server detects strong anxiety or stress, it will prioritize medical institutions that offer psychological counseling.

[1194] Sending results from the server to the device

[1195] The server packages the estimated diagnosis and information on the recommended medical institution, encrypts it again, and transmits it to the terminal.

[1196] Displaying results on a terminal

[1197] The device decodes the received data and displays it in a user-friendly format, including the estimated diagnosis and detailed information about the recommended medical institution (e.g., location, contact information, and opening hours). The user can then review the displayed information and contact the appropriate hospital to arrange an appointment.

[1198] Specific examples

[1199] For example, if a user has a hand injury, the following process occurs:

[1200] 1. A user takes a photo of a hand injury with their smartphone and enters, "My hand is swollen and painful, and I'm very anxious."

[1201] 2. The device packages the image data and text data in JSON format and sends it to the server.

[1202] 3. The server analyzes the images using an AI model and the text data using an NLP model and emotion engine.

[1203] 4. The server integrates the results of each analysis and generates a diagnosis of "high probability of fracture" and an emotional state indicating "high anxiety."

[1204] 5. The server selects nearby orthopedic hospitals and medical institutions that provide psychological counseling, taking into account the user's location information and emotional state.

[1205] 6. The server sends the diagnosis results and medical institution information to the terminal.

[1206] 7. The device displays the received information to the user, who then contacts the appropriate medical institution and arranges for an examination.

[1207] The above is a specific embodiment of the present invention. This system allows users to receive prompt and accurate diagnosis and access to medical institutions, taking into account not only their medical condition but also their emotional state.

[1208] The processing flow will be explained below.

[1209] Step 1:

[1210] The user takes an image of the medical condition or injury using a smartphone or computer, saves the image data on the device, and then enters a description of their symptoms or condition into a text input form. For example, they might enter "My hand is swollen and painful."

[1211] Step 2:

[1212] The device packages the image and text data entered by the user into a specified format (e.g., JSON or XML), encrypts the data, and prepares it for transmission.

[1213] Step 3:

[1214] The device sends the encrypted image data and text data to the server, and the data is transmitted securely over the Internet.

[1215] Step 4:

[1216] The server decrypts the received encrypted data and prepares it for analysis. The server passes image data to the image analysis module and text data to the natural language processing (NLP) module.

[1217] Step 5:

[1218] The server analyzes the image data using an image analysis module. Specifically, it uses a convolutional neural network (CNN) to extract image features and estimate the condition of the patient. For example, it can detect signs of a fracture from the image.

[1219] Step 6:

[1220] The server analyzes the text data using a natural language processing (NLP) module. It extracts symptom-related keywords and context from the text and analyzes the symptoms. For example, it extracts information such as "swelling" and "pain."

[1221] Step 7:

[1222] The server uses an emotion engine to recognize the user's emotional state from the text data. The emotion engine analyzes keywords and context in the text to determine the user's emotional state (e.g., anxiety, stress, pain level, etc.).

[1223] Step 8:

[1224] The generative AI model combines the results of image analysis, text analysis, and emotion engine analysis to produce a probable diagnosis, which includes an estimated medical condition, its confidence level, and the user's emotional state.

[1225] Step 9:

[1226] Based on the estimated diagnosis, the server references the user's current location and searches a database for the most suitable medical institution. The server also takes into account the user's emotional state when selecting an appropriate medical institution. For example, if strong anxiety or stress is detected, it will prioritize medical institutions that offer psychological counseling.

[1227] Step 10:

[1228] The server packages the estimated diagnosis results and information about the selected medical institution, encrypts them again, and sends them to the terminal.

[1229] Step 11:

[1230] The device decodes the received data and displays it in a user-friendly format, including the estimated diagnosis and detailed information about the recommended medical institution (e.g., location, contact information, consultation hours, etc.).

[1231] Step 12:

[1232] The user checks the medical institution information displayed on the device and contacts the appropriate hospital to arrange for a medical examination. For example, the user contacts a hospital in response to a message that reads, "There is a high possibility of a fracture. We recommend the following orthopedic hospital."

[1233] Example 2

[1234] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1235] Currently, many people have difficulty accurately and quickly communicating information about their illness or injury to medical institutions. In particular, it is difficult for users to accurately describe their symptoms and emotional state, which can delay the selection of an appropriate medical institution and diagnosis. To solve this problem, a system is needed that can effectively analyze user input data, perform a comprehensive diagnosis including emotional state, and quickly provide guidance to an appropriate medical institution.

[1236] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1237] In this invention, the server includes a means for analyzing image data input by a user to estimate a medical condition, a means for analyzing text data to estimate symptoms, and a means for recognizing the user's emotional state from the text data using an emotion engine. This enables a highly accurate inferential diagnosis to be made by comprehensively evaluating the image analysis results, text analysis results, and emotion analysis results. Furthermore, based on this inferential diagnosis, the server can quickly select an appropriate medical institution by referring to the user's location information and provide that information to the user, thereby enabling the provision of prompt and appropriate medical services.

[1238] "User" refers to an individual who uses the system to input information about a medical condition or injury.

[1239] "Terminal" refers to a device used by a user to input image and text data related to a medical condition or injury and send it to a server. Examples of such devices include smartphones and computers.

[1240] "Image data" refers to photographic data of illnesses or injuries that users take using their terminals and enter into the system.

[1241] "Text data" refers to textual information such as the status and impressions of a medical condition or injury that a user inputs into a device.

[1242] A "format" refers to the rules and forms for organizing and structuring data, such as JSON and XML.

[1243] A "server" is a central computer system for receiving and analyzing image data and text data sent from the terminals.

[1244] A "generative AI model" is a collection of algorithms and programs that run on a server and integrate the results of image analysis and text analysis to make a presumptive diagnosis.

[1245] A "presumed diagnosis" is the evaluation result of the disease state or symptoms calculated by the generative AI model, and is a preliminary assessment before an actual diagnosis is made.

[1246] An "emotion engine" is a software component that analyzes text data and recognizes the user's emotional state.

[1247] "Location information" is geographical data that indicates the user's current location, and is obtained from GPS data or the like.

[1248] "Medical institution" refers to a clinic, doctor's office, hospital, etc. that provides medical services to users.

[1249] "Packaging" is the process of combining multiple pieces of data into one format, and is done to maintain data consistency.

[1250] "Encryption" is the process of transforming data using a specific algorithm to protect the confidentiality of the data.

[1251] "Decryption" is the process of restoring encrypted data to its original form.

[1252] "Analysis" is the process of analyzing data and extracting meaningful information and patterns from it.

[1253] This invention is a system that allows users to input image and text data about their illness or injury and receive referrals to appropriate medical institutions, even without specialized medical knowledge. Furthermore, by combining this system with an emotion engine that recognizes the user's emotions, it is possible to make diagnoses and select medical institutions that take the user's emotional state into consideration.

[1254] Hardware and software used

[1255] The system consists of the following components:

[1256] User terminal: A device through which a user inputs and receives data, such as a smartphone or computer.

[1257] Server: A central computer system that analyzes data and selects medical institutions.

[1258] Generative AI models: Algorithms that perform data analysis, such as convolutional neural networks (CNNs) and natural language processing (NLP) models.

[1259] Emotion Engine: A software component for recognizing the emotional state of a user from their text data.

[1260] Processing flow

[1261] A specific embodiment for implementing this system will be described below.

[1262] User input of images and text

[1263] Users use their smartphones or computers to take pictures of their illnesses or injuries and save the image data on their devices. They then enter their symptoms and thoughts about their condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text, "My hand is swollen and painful."

[1264] Sending data by the device

[1265] The device converts the image data and text data entered by the user into a specified format (e.g., JSON format), packages it, encrypts it, and sends it to a server via the Internet.

[1266] Receipt and analysis of data by the server

[1267] The server receives and decodes the image and text data sent from the device. The server then uses a generative AI model to analyze the image data. Specifically, it uses a convolutional neural network (CNN) to extract image features and infer the disease state. At the same time, the server analyzes the text data with a natural language processing (NLP) model to infer the symptoms.

[1268] Emotion engine recognizes emotional states

[1269] The server uses an emotion engine to recognize the user's emotional state based on the text data. The emotion engine analyzes keywords and context within the text to determine the user's emotional state, such as stress, anxiety, or pain. For example, if the text contains expressions such as "It hurts so much" or "I'm so anxious," the corresponding emotional state is recognized.

[1270] Predictive diagnosis using generative AI models

[1271] The generative AI model on the server integrates the results of image analysis, text analysis, and emotion engine analysis to make a probable diagnosis. This diagnosis includes the estimated medical condition, its reliability, and the user's emotional state. For example, if image analysis suggests a fracture, text analysis identifies symptoms such as "swelling and pain," and the emotion engine recognizes "strong anxiety," the generative AI model will generate a probable fracture diagnosis.

[1272] Server selection of appropriate medical institution

[1273] The server then references the user's current location information based on the estimated diagnosis and searches a database for the most suitable medical institution. The information on the medical institution includes the location, contact information, and available consultation hours. The server also takes into account the user's emotional state. For example, if the server detects strong anxiety or stress, it will prioritize medical institutions that offer psychological counseling.

[1274] Sending results from the server to the device

[1275] The server packages the estimated diagnosis results and information on recommended medical institutions, encrypts them again, and sends them to the terminal.

[1276] Displaying results on a terminal

[1277] The device decodes the received data and displays it in a user-friendly format, including the estimated diagnosis and detailed information about the recommended medical institution (such as location, contact information, and opening hours). The user can then review the displayed information and contact the appropriate hospital to arrange an appointment.

[1278] Specific examples

[1279] For example, if a user has a hand injury, the following process occurs:

[1280] 1. A user takes a photo of a hand injury with their smartphone and enters, "My hand is swollen and painful, and I'm very anxious."

[1281] 2. The device packages the image data and text data in JSON format and sends it to the server.

[1282] 3. The server analyzes the images using an AI model and the text data using an NLP model and emotion engine.

[1283] 4. The server integrates the results of each analysis and generates a diagnosis of "high probability of fracture" and an emotional state indicating "high anxiety."

[1284] 5. The server selects nearby orthopedic hospitals and medical institutions that provide psychological counseling, taking into account the user's location information and emotional state.

[1285] 6. The server sends the diagnosis results and medical institution information to the terminal.

[1286] 7. The device displays the received information to the user, who then contacts the appropriate medical institution and arranges for an examination.

[1287] This process allows users to access medical care and receive a quick and accurate diagnosis that takes into account not only their medical condition but also their emotional state.

[1288] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1289] Step 1:

[1290] User input of images and text

[1291] Users use a smartphone or computer to take pictures of their condition or injury, save the image data on the device, and then enter detailed symptoms and impressions of their condition in text format.

[1292] Input: Image data and text data related to medical conditions and injuries

[1293] Output: Image data and text data stored on the device

[1294] Specific behavior:

[1295] The user activates the camera and takes a picture of the affected area.

[1296] After taking the photo, you upload the image through the device's application and enter details such as "My hand is swollen and painful, and I'm very anxious" in the text input field.

[1297] Step 2:

[1298] Sending data by the device

[1299] The device converts the image data and text data entered by the user into a specified format (e.g., JSON format) and packages it. The device then encrypts the data and sends it to the server via the Internet.

[1300] Input: User image and text data

[1301] Output: Encrypted packaging data

[1302] Specific behavior:

[1303] The device converts the image file and text data into JSON format.

[1304] The data is encrypted using the SSL / TLS protocol and sent to the server's API endpoint.

[1305] Step 3:

[1306] Receipt and analysis of data by the server

[1307] The server receives and decrypts the encrypted data sent from the device. Next, the server analyzes the received image data using a generative AI model (CNN algorithm) to extract image features. At the same time, the server analyzes the text data using a natural language processing (NLP) model to estimate symptoms.

[1308] Input: Encrypted packaging data

[1309] Output: Analyzed image features and text analysis results

[1310] Specific behavior:

[1311] The server reads the encrypted data from the SSL / TLS input buffer and decrypts it.

[1312] The server inputs the image data into a generative AI model, extracts features, and estimates the condition of the disease.

[1313] An NLP processing engine is used to analyze the text data and extract keywords and context related to the symptoms.

[1314] Step 4:

[1315] Emotion engine recognizes emotional states

[1316] The server analyzes the text data with an emotion engine to recognize the user's emotional state. The emotion engine analyzes keywords and context within the text to determine the user's stress, anxiety, pain, etc.

[1317] Input: Text data

[1318] Output: Data about the user's emotional state

[1319] Specific behavior:

[1320] The server inputs the text data into an emotion engine, which analyzes the keywords and context within the text.

[1321] The emotion engine analyzes expressions such as "It hurts so much" and "I'm so anxious" to determine the user's emotional state.

[1322] Step 5:

[1323] Predictive diagnosis using generative AI models

[1324] The generative AI model on the server integrates the results of image analysis, text analysis, and emotion engine analysis to produce a probable diagnosis, which includes the probable medical condition, its reliability, and the user's emotional state.

[1325] Input: Image analysis results, text analysis results, sentiment analysis results

[1326] Output: Estimated diagnosis result

[1327] Specific behavior:

[1328] The server integrates the results of image analysis, text analysis, and sentiment analysis.

[1329] The generative AI model estimates the condition based on the integrated data and generates a diagnosis such as "high probability of fracture."

[1330] Step 6:

[1331] Server selection of appropriate medical institution

[1332] Based on the estimated diagnosis, the server references the user's current location information and searches a database for the most suitable medical institution, including information on the location, contact details, and available hours for consultations.

[1333] Input: Estimated diagnosis result, user's current location information

[1334] Output: Appropriate medical institution information

[1335] Specific behavior:

[1336] The server obtains the user's current location information from GPS data and queries a location information database.

[1337] The server searches the medical institution database and generates a list of medical institutions that can provide consultations.

[1338] Step 7:

[1339] Sending results from the server to the device

[1340] The server packages the estimated diagnosis and information on the recommended medical institution, encrypts it again, and transmits it to the terminal.

[1341] Input: Estimated diagnosis result, medical institution information

[1342] Output: Encrypted diagnosis results and medical institution information

[1343] Specific behavior:

[1344] The server packages the diagnosis results and medical institution information in JSON format.

[1345] The server encrypts the packaged data and sends it to the user's terminal.

[1346] Step 8:

[1347] Displaying results on a terminal

[1348] The device decodes the received data and displays it in a user-friendly format, including a probable diagnosis and detailed information about recommended medical institutions.

[1349] Input: Encrypted diagnosis results and medical institution information

[1350] Output: Decrypted diagnosis and medical institution information

[1351] Specific behavior:

[1352] The device decodes the data it receives and parses the JSON format data to extract the diagnosis results and medical institution information.

[1353] The device updates the application UI and displays the diagnosis results and medical institution information to the user.

[1354] Through this detailed, step-by-step process, users can obtain an accurate and prompt diagnosis and access to the appropriate medical facility.

[1355] (Application example 2)

[1356] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1357] Conventional medical diagnosis systems have the problem that it is difficult for users without medical expertise to make an appropriate diagnosis and select a medical institution. Furthermore, there is a need for rapid medical diagnosis and guidance to medical institutions in emergency situations. In particular, a system that can provide appropriate information immediately is needed so that security staff can respond quickly on the scene.

[1358] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1359] In this invention, the server includes: a means for a user to input image data and text data related to a medical condition or injury; a means for a terminal to package the image data and text data into a specified format and transmit the packaged image data and text data to the server; a means for the server to analyze the image data and estimate a medical condition; a means for the server to analyze the text data and estimate symptoms; a generative AI model to integrate the analysis results and make an estimated diagnosis; a means for the server to select an appropriate medical institution based on the estimated diagnosis; a means for the server to transmit information about the estimated diagnosis and the medical institution to the terminal; a means for the terminal to display the information to the user; and a means for performing a rapid medical diagnosis on site using smart glasses. This enables a user (security staff) to receive a rapid and appropriate medical diagnosis on site and access an appropriate medical institution based on the results.

[1360] "Image data" is digital data entered by a user that contains visual information about a medical condition or injury.

[1361] "Text data" is digital data containing written information related to a medical condition or injury entered by a user.

[1362] A "user" is an individual who uses the system and who inputs information about a medical condition or injury.

[1363] A "terminal" is a device operated by a user, and is a device for packaging image data and text data and transmitting them to a server.

[1364] A "server" is a computer system that receives data sent from a terminal and performs analysis and inferential diagnosis.

[1365] A "generative AI model" is an artificial intelligence algorithm that integrates the analysis results of image data and text data to make a presumptive diagnosis.

[1366] "Format" refers to a data format for arranging image data and text data in a form that is easy for the server to process.

[1367] "Smart glasses" are wearable devices equipped with a camera function that users use in the field.

[1368] An "appropriate medical institution" is a medical facility that can provide optimal medical services to the user based on the estimated diagnosis results.

[1369] A "presumed diagnosis" is the result of a disease state inferred by a generative AI model based on an analysis of image data and text data.

[1370] "Display means" refers to a function that visually presents to the user the information that the terminal receives from the server.

[1371] "Emotional state" indicates the psychological and emotional state of the user, which is analyzed from text data or the like.

[1372] In order to implement the present invention, the following configurations and means are necessary.

[1373] Overall system overview

[1374] The user inputs image and text data of their condition or injury and sends the information to a server. The server analyzes the data, makes a probable diagnosis, selects an appropriate medical institution, and provides that information to the user. Furthermore, the use of smart glasses enables rapid medical diagnosis on the spot.

[1375] Hardware and Software

[1376] 1. Smart Glasses:

[1377] A wearable device with a camera function.

[1378] Supports on-site image capture and data entry.

[1379] 2. Terminal:

[1380] Personal computers and smartphones.

[1381] It supports the input of image data and text data and sends the data to the server.

[1382] 3. Server:

[1383] High performance computing systems.

[1384] Analyzes image and text data, performs presumptive diagnoses, and selects appropriate medical institutions.

[1385] The algorithms used are convolutional neural networks (CNN) and natural language processing (NLP).

[1386] Program processing

[1387] The device packages the image data and text data entered by the user into a specified format (e.g., JSON), encrypts it, and sends it to the server. The server decrypts the received data and performs the following processes.

[1388] 1. Analysis of imaging data:

[1389] The server uses a convolutional neural network (CNN) to extract features from the image data and estimate the disease state.

[1390] 2. Text data analysis:

[1391] Text data is analyzed using a natural language processing (NLP) model to estimate symptoms.

[1392] 3. Integrated analysis using generative AI models:

[1393] The generative AI model integrates the results of image analysis, text analysis, and the emotion engine analysis to make a presumptive diagnosis.

[1394] 4. Selection of appropriate medical facility:

[1395] Based on the estimated diagnosis results, the server selects the most appropriate medical institution, taking into consideration the user's location information and emotional state.

[1396] 5. Transmission and Display of Information:

[1397] The server transmits the generated estimated diagnosis and information about the medical institution to the terminal, which then displays them to the user.

[1398] Specific examples

[1399] For example, if security staff discover someone collapsed at a scene, they can take the following steps:

[1400] 1. Taking a photo of a fallen person using smart glasses.

[1401] 2. Enter the text data "Unconscious, not breathing."

[1402] 3. The device sends this data to the server.

[1403] 4. The server makes a probable diagnosis and generates information such as "possibility of cardiac arrest" and "nearest hospital with a cardiologist."

[1404] 5. The server sends the diagnosis results and information on recommended medical institutions to the terminal, which displays them to the user.

[1405] Prompt Sentence Examples

[1406] "Unconscious, not breathing"

[1407] "My hands are swollen and painful, and I'm very anxious"

[1408] By using these prompts, users can quickly and accurately access the appropriate medical institution.

[1409] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1410] Step 1:

[1411] The user uses the smart glasses to take images of their condition or injury and input text data about their symptoms. The user uses the smart glasses' camera function to take photos of their condition or injury, and then inputs text information such as "My hand is swollen and painful, and I feel very anxious" through the smart glasses' UI.

[1412] Input: Image and text data of medical conditions and injuries.

[1413] Output: Photographed image data and input text data.

[1414] Step 2:

[1415] The device packages the image and text data entered by the user into a specified format, encrypts it, and sends it to the server.The device then converts the data into a structured data format such as JSON and encrypts it using an encryption algorithm such as AES for security.The encrypted data is then sent to the server via the HTTPS protocol.

[1416] Input: Image data and text data.

[1417] Output: Encrypted and formatted data.

[1418] Step 3:

[1419] The server receives and decrypts the data sent from the device. The server receives the data via the HTTPS protocol and decrypts data encrypted with AES or other encryption protocols. It then analyzes the received JSON format data and separates the image data from the text data.

[1420] Input: Encrypted and formatted data.

[1421] Output: Decoded and separated image and text data.

[1422] Step 4:

[1423] The server analyzes the image data to estimate the condition of the patient. Specifically, it uses a convolutional neural network (CNN) to extract image features and estimate the condition. Based on the results of this analysis, a diagnosis such as "high possibility of fracture" is obtained.

[1424] Input: Image data.

[1425] Output: Estimated disease state based on image analysis.

[1426] Step 5:

[1427] The server analyzes the text data to infer symptoms. It uses a natural language processing (NLP) model to analyze keywords and context within the text data to infer symptoms. For example, from an expression such as "My hand is swollen and painful, and I feel very anxious," it extracts symptoms such as "swelling," "pain," and "anxiety."

[1428] Input: Text data.

[1429] Output: Symptom prediction results based on text analysis.

[1430] Step 6:

[1431] The generative AI model combines the results of image analysis, text analysis, and the emotion engine to generate a comprehensive diagnosis, such as "There is a high possibility of a fracture and the patient feels high anxiety."

[1432] Input: Image analysis results, text analysis results, emotion engine results.

[1433] Output: Integrated estimated diagnostic results.

[1434] Step 7:

[1435] The server selects an appropriate medical institution based on the user's location information and emotional state. Based on the estimated diagnosis, the server searches a database of medical institutions and selects the most suitable medical institution. For example, the server selects the "nearest orthopedic hospital" and "medical institution that offers psychological counseling" based on the user's location information and emotional state.

[1436] Input: Estimated diagnosis result, user location information, emotional state.

[1437] Output: Information on selected medical institutions.

[1438] Step 8:

[1439] The server then sends the generated estimated diagnosis and information about the medical institution to the terminal, where it re-encrypts the information and transmits it to the terminal using the HTTPS protocol.

[1440] Input: Estimated diagnosis result, medical institution information.

[1441] Output: Encrypted diagnosis results and medical institution information.

[1442] Step 9:

[1443] The device decrypts the information received from the server and displays it in a format that is easy for the user to understand. The device decrypts the encrypted data and displays information such as "High probability of fracture," "Nearest orthopedic hospital: XX Hospital," and "Psychological counseling: YY Clinic" through the UI.

[1444] Input: Encrypted diagnosis results and medical institution information.

[1445] Output: Presenting information in a user-friendly format.

[1446] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1447] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1448] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1449] [Fourth embodiment]

[1450] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1451] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1452] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1453] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1454] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1455] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1456] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1457] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1458] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1459] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1460] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1461] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1462] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1463] This invention is a system that allows users to input image and text data related to a medical condition or injury and receive a referral to an appropriate medical institution even if they do not have specialized medical knowledge. Specific embodiments of this system will be described below.

[1464] User input of images and text

[1465] Users can use a smartphone or computer to take pictures of their illness or injury and save the image data on the device. They can then enter their symptoms and thoughts about the condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text, "My hand is swollen and painful."

[1466] Sending data by the device

[1467] The device packages the image data and text data entered by the user into a specified format (for example, JSON or XML).The device then sends this data to a server via the Internet.At this time, the data is encrypted to ensure security.

[1468] Receipt and analysis of data by the server

[1469] The server receives and decodes the image and text data sent from the device. It then uses a generative AI model to analyze the received image data. Specifically, it uses algorithms such as convolutional neural networks to extract image features and estimate the condition. At the same time, the server analyzes the text data using a natural language processing (NLP) model to estimate symptoms.

[1470] Predictive diagnosis using generative AI models

[1471] The generative AI model on the server integrates the results of image analysis and text analysis to generate a probable diagnosis. This diagnosis includes a probable pathology and its reliability. For example, if image analysis detects a bone abnormality and text analysis identifies the symptom "swelling and pain," the probable diagnosis is "high probability of fracture."

[1472] Server selection of appropriate medical institution

[1473] Based on the estimated diagnosis, the server refers to the user's location information and selects the most appropriate medical institution. At this time, the server retrieves a list of affiliated medical institutions from the database and lists medical institutions close to the user's current location. The information on the selected medical institution includes location, contact information, available hours for consultation, etc.

[1474] Sending results from the server to the device

[1475] The server then sends the estimated diagnosis and information on the recommended medical institution to the device. The data is then re-encrypted and transferred to the device in a secure manner.

[1476] Displaying results on a terminal

[1477] The device decodes the received data and displays it in a user-friendly format. The user can then check the displayed estimated diagnosis and recommended medical institution information and make arrangements for a medical examination if necessary. For example, if there is a high possibility of a fracture, the user can call the orthopedic hospital contact number displayed on the device to make an appointment.

[1478] Specific examples

[1479] For example, if a user injures their hand, the following process occurs:

[1480] 1. A user takes a photo of a hand injury with their smartphone and types, "My hand is swollen and painful."

[1481] 2. The device packages the image data and text data in JSON format and sends it to the server.

[1482] 3. The server analyzes the image using an AI model and the text data using an NLP model.

[1483] 4. The server integrates the analysis results and generates a diagnosis of "high probability of fracture."

[1484] 5. The server selects a nearby orthopedic hospital based on the user's location information.

[1485] 6. The server sends the diagnosis results and hospital information to the terminal.

[1486] 7. The terminal displays the received information to the user, who then contacts the orthopedic hospital to arrange for a consultation.

[1487] The above is a specific embodiment of the present invention. This system makes it possible to quickly and accurately estimate the condition of a patient, even without specialized medical knowledge, and to assist in accessing an appropriate medical institution.

[1488] The processing flow will be explained below.

[1489] Step 1: The user takes an image of their condition or injury using a smartphone or computer, saves the image data on their device, and enters a description of their symptoms or condition into a text input form.

[1490] Step 2: The terminal converts the image data and text data entered by the user into a specified format (e.g., JSON or XML) and packages it.

[1491] Step 3: The device encrypts the packaged data and sends it to a server via the Internet.

[1492] Step 4: The server decodes the received image data and text data and prepares each for analysis.

[1493] Step 5: The server analyzes the image data using an image analysis module. Specifically, it uses a convolutional neural network (CNN) to extract image features and estimate the disease state.

[1494] Step 6: The server analyzes the text data using a natural language processing (NLP) module, extracting symptom-related keywords and context from the text and obtaining the analysis results.

[1495] Step 7: The generative AI model on the server integrates the results of image analysis and text analysis to make a probable diagnosis. For example, if image analysis indicates a suspected fracture and text analysis indicates pain or swelling, the model will generate a probable diagnosis of "high probability of fracture."

[1496] Step 8: Based on the estimated diagnosis, the server references the user's current location information and searches the database for the most suitable medical institution. Information on the target medical institutions is then listed.

[1497] Step 9: The server packages the estimated diagnosis and the information on the selected medical institution, encrypts it again, and transmits it to the terminal.

[1498] Step 10: The terminal decrypts the received data and displays it to the user, including the estimated diagnosis and detailed information (such as location, contact information, and consultation hours) of the recommended medical institution.

[1499] Step 11: The user checks the information about the medical institution displayed on the terminal and contacts the appropriate hospital to arrange for a medical examination.

[1500] Example 1

[1501] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1502] In today's world, it is difficult for users to quickly and accurately communicate information about their illness or injury to specialized medical institutions. It is also difficult for users without medical knowledge to select an appropriate medical institution. As a result, it can take a lot of time and effort to receive an appropriate diagnosis and treatment. To solve these problems, a system is needed that can effectively collect information about a user's illness or injury, analyze it quickly and accurately, and select an appropriate medical institution.

[1503] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1504] In this invention, the server includes: means for a user to input image data and text data related to a medical condition or injury; means for a terminal to package the image data and text data in a specified format and transmit the packaged image data and text data to the server; means for the terminal to encrypt the data; means for the server to receive and decrypt the image data and text data; means for the server to analyze the image data, extract features, and infer a medical condition; means for the server to analyze the text data, extract and infer symptoms; means for a generative AI model to integrate the analysis results and make a probable diagnosis; means for the server to select an appropriate medical institution based on the probable diagnosis; means for the server to transmit information about the probable diagnosis and medical facilities to the terminal; and means for the terminal to display the information to the user. This enables users to quickly and accurately understand their own medical condition or injury and to promptly visit an appropriate medical institution, even without specialized medical knowledge.

[1505] "Image data" is digital information that a user visually records regarding a medical condition or injury.

[1506] "Text data" is digital information in which a user records symptoms and thoughts about a medical condition or injury in written form.

[1507] A "terminal" is a device that a user uses to input data and send it to a server, and includes a smartphone or computer.

[1508] A "server" is a central processing device that receives data sent from a terminal, analyzes it, and returns the necessary information.

[1509] A "specified format" is a format in which data is arranged according to certain rules, and includes, for example, JSON and XML.

[1510] "Encryption" is a technique that uses a specific algorithm to convert data in order to transmit it securely.

[1511] "Decryption" is the technique of returning encrypted data to its original form.

[1512] "Features" are important attributes or patterns extracted from image data and are used to estimate the condition of a disease.

[1513] A "natural language processing (NLP) model" is an algorithm or method for analyzing text data and understanding human language.

[1514] A "generative AI model" is an artificial intelligence algorithm that integrates the results of image analysis and text analysis to make a presumptive diagnosis.

[1515] A "presumptive diagnosis" is the result of predicting the condition of a disease or injury based on analyzed data.

[1516] A "medical institution" is a facility where a user can receive medical examinations and treatment, and includes hospitals and clinics.

[1517] "Information" refers to data such as the diagnosis results and contact details and locations of medical institutions required by the user.

[1518] The present invention provides a system that allows a user to input image data and text data relating to a medical condition or injury, and easily refer the user to an appropriate medical institution. A specific embodiment of this system will be described below.

[1519] User data entry

[1520] Users use devices such as smartphones or computers to take pictures of their medical condition or injury and save them on the device. Next, they enter text data about their condition (e.g., "My hand is swollen and painful") into the device. At this stage, users do not need specialized medical knowledge and can enter data intuitively.

[1521] Device prepares to send data

[1522] The device packages the image and text data entered by the user into a specified format (e.g., JSON or XML), then encrypts the packaged data and transmits it securely to a server over the Internet.

[1523] Receipt and analysis of data by the server

[1524] The server receives the encrypted data sent from the device. After receiving the data, the server decrypts it and extracts the image data and text data. The server then analyzes the image data using a convolutional neural network (CNN) to extract image features. The server also analyzes the text data using a natural language processing (NLP) model to extract the described symptoms.

[1525] Predictive diagnosis using generative AI models

[1526] The generative AI model deployed on the server integrates the results of image analysis and text analysis. This model then uses the analysis results to generate a probable diagnosis, which includes a probable medical condition and its reliability. For example, by combining an image of a hand injury with the text data "your hand is swollen and painful," the model generates a probable diagnosis of "high probability of fracture."

[1527] Selection of medical institutions by the server

[1528] The server then references the user's location information based on the estimated diagnosis and selects an appropriate medical institution from a database. The information on the selected medical institution includes the location, contact information, and available consultation hours. The server then organizes this information, re-encrypts the data, and sends it to the device.

[1529] Displaying results on a terminal

[1530] The device receives and decrypts the encrypted data sent from the server. The decrypted data is displayed in a user-friendly format, allowing the user to refer to it and contact the appropriate medical institution and arrange for a medical examination. For example, it may display information such as "There is a high possibility of a fracture. The recommended medical institution is XX Orthopedic Hospital, contact number: XX-XX-XX."

[1531] Examples of concrete examples and prompts

[1532] As a concrete example, if a user injures their hand, the process is as follows:

[1533] 1. A user takes a photo of a hand injury with their smartphone and types, "My hand is swollen and painful."

[1534] 2. The device packages the image data and text data in JSON format and sends it to the server.

[1535] 3. The server analyzes the image using an AI model and the text data using an NLP model.

[1536] 4. The server integrates the analysis results and generates a diagnosis of "high probability of fracture."

[1537] 5. The server selects a nearby orthopedic hospital based on the user's location information.

[1538] 6. The server sends the diagnosis results and hospital information to the terminal.

[1539] 7. The terminal displays the received information to the user, who then contacts the orthopedic hospital to arrange for a consultation.

[1540] This system allows even users without specialized medical knowledge to quickly and accurately understand the patient's condition and refer them to the appropriate medical institution.

[1541] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1542] Step 1:

[1543] The user uses a smartphone or computer to take pictures of their condition or injury and save them on the device, then enters text about their condition (e.g., "My hand is swollen and painful") into the device.

[1544] Input: Image data and text data related to medical conditions and injuries

[1545] Output: Image data and text data stored on the device

[1546] Step 2:

[1547] The terminal acquires image data and text data input by the user.

[1548] Input: Image data and text data stored on the device

[1549] Output: Acquired image data and text data

[1550] Step 3:

[1551] The image data and text data acquired by the device are packaged in a specified format such as JSON.

[1552] Input: Acquired image data and text data

[1553] Output: Data packaged in the specified format

[1554] Step 4:

[1555] The device encrypts the packaged data and prepares it in a highly secure state.

[1556] Input: Data packaged in a specified format

[1557] Output: Encrypted data

[1558] Step 5:

[1559] The terminal transmits the encrypted data to a server via the Internet.

[1560] Input: Encrypted data

[1561] Output: Data sent to the server

[1562] Step 6:

[1563] The server receives the encrypted data sent from the terminal.

[1564] Input: Data sent over the internet

[1565] Output: Encrypted data received by the server

[1566] Step 7:

[1567] The server decrypts the received encrypted data and extracts the image data and text data.

[1568] Input: Received encrypted data

[1569] Output: Decoded image data, text data

[1570] Step 8:

[1571] The server analyzes the image data using a convolutional neural network (CNN) and extracts features.

[1572] Input: Decoded image data

[1573] Output: Image features

[1574] Step 9:

[1575] The server analyzes the text data using a natural language processing (NLP) model to extract symptoms.

[1576] Input: Decrypted text data

[1577] Output: Extracted symptoms

[1578] Step 10:

[1579] The server uses the generated AI model to integrate the results of image analysis and text analysis and make a presumptive diagnosis.

[1580] Input: Image features, extracted symptoms

[1581] Output: Estimated diagnosis result (condition and its reliability)

[1582] Step 11:

[1583] Based on the estimated diagnosis, the server refers to the user's location information and selects an appropriate medical institution.

[1584] Input: Estimated diagnosis result, user location information

[1585] Output: Information about the selected medical institution (location, contact information, available hours, etc.)

[1586] Step 12:

[1587] The server re-encrypts the diagnosis results and medical institution information and sends them to the terminal.

[1588] Input: Information on the selected medical institution, diagnosis results

[1589] Output: Encrypted data for transmission

[1590] Step 13:

[1591] The terminal receives the encrypted data sent from the server.

[1592] Input: Encrypted data to be sent

[1593] Output: Received encrypted data

[1594] Step 14:

[1595] The terminal decrypts the received encrypted data and displays it in a user-friendly format.

[1596] Input: Received encrypted data

[1597] Output: Diagnosis results and medical institution information displayed to the user

[1598] Through the above steps, the user can quickly and accurately understand the condition of their illness or injury and promptly seek medical attention at an appropriate medical institution.

[1599] (Application example 1)

[1600] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1601] In recent years, users have faced the challenge of finding an appropriate medical institution for their illness or injury quickly and accurately without specialized medical knowledge. Furthermore, there is a lack of systems that allow users to instantly make appointments and make payments based on diagnosis results. This creates a problem of inability to smoothly access and use medical services.

[1602] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1603] In this invention, the server includes means for selecting an appropriate medical institution based on a presumed diagnosis and transmitting that information, means for making a reservation at the recommended medical institution, and means for completing payment by electronic payment. This enables a user to quickly find an appropriate medical institution based on their condition or injury and seamlessly complete a series of processes from reservation to payment.

[1604] A "user" is someone who uses the system to input information about their medical condition or injury and wishes to receive recommendations or make appointments with medical institutions.

[1605] "Image data relating to a medical condition or injury" refers to photographs or image files taken by a user that show the state of a medical condition or injury.

[1606] "Text data" refers to character information entered by the user, including symptoms, impressions, and explanations of illnesses or injuries.

[1607] A "terminal" is a device, such as a smartphone or computer, that allows a user to access the system and input and send data.

[1608] A "server" is a computer system that receives data sent by users, analyzes it, selects medical institutions, and sends information.

[1609] A "specified format" is a format used to organize image data or text data into a certain format, such as JSON or XML.

[1610] "Encryption" is the process of converting data into a format that cannot be read by third parties in order to prevent data leakage during communication.

[1611] A "generative AI model" is a system that includes an artificial intelligence algorithm for analyzing and presumptively diagnosing medical conditions and symptoms using image and text data.

[1612] A "presumed diagnosis" is a provisional diagnosis derived from the results of the disease state and symptoms analyzed by the generative AI model.

[1613] "Selection of medical institution" refers to the procedure in which the server identifies an appropriate medical institution based on the estimated diagnosis result and taking into consideration the user's location information.

[1614] The "reservation means" is part of the system that allows the user to make an appointment with a recommended medical institution, and is a function that manages reservation information in cooperation with the server.

[1615] "Electronic payment" is a payment system conducted via the Internet, and refers to the use of online payment methods such as credit cards and electronic money.

[1616] This invention is a system that allows users to input image and text data about their medical condition or injury and receive referrals to appropriate medical institutions, even without specialized medical knowledge. This system is constructed using a terminal, a server, and a generative AI model.

[1617] First, the user takes a picture of their condition or injury using a smartphone or computer and saves the image data on the device. Next, the user enters their symptoms and thoughts about the condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text "My hand is swollen and painful."

[1618] The terminal packages the image data and text data entered by the user into a specified format (for example, JSON or XML). The terminal then sends this data to the server via the Internet. At this time, the data is encrypted to ensure security. For encryption, the pycryptodome library, for example, is used.

[1619] The server receives and decodes the image and text data sent from the device. It then uses a generative AI model to analyze the received image data. Specifically, it uses deep learning frameworks such as TensorFlow and PyTorch to extract image features using algorithms such as convolutional neural networks (CNNs) and estimate the condition of the patient. At the same time, the server uses a natural language processing (NLP) model to analyze the text data and estimate the symptoms.

[1620] The generative AI model combines the results of image analysis and text analysis to generate a probable diagnosis. This diagnosis includes a probable pathology and its reliability. For example, if image analysis detects a bone abnormality and text analysis identifies the symptom "swelling and pain," the probable diagnosis is "high probability of fracture."

[1621] The server selects the most appropriate medical institution based on the estimated diagnosis and the user's location information. At this time, the server retrieves a list of affiliated medical institutions from the database and lists the medical institutions closest to the user's current location. The information on the selected medical institution includes the location, contact information, and available hours for consultation.

[1622] The server then sends the estimated diagnosis and information on the recommended medical institution to the device. The data is then re-encrypted and transferred to the device in a secure manner.

[1623] The device decrypts the received data and displays it in a user-friendly format. The user can then check the displayed estimated diagnosis and information on recommended medical institutions, make an appointment with a medical institution within the app, and complete payment electronically. Payment can be made using APIs such as Stripe or PayPal.

[1624] Specific examples

[1625] For example, if a user injures their hand, the following process occurs:

[1626] 1. A user takes a photo of a hand injury with their smartphone and types, "My hand is swollen and painful."

[1627] 2. The device packages the image data and text data in JSON format, sends it to the server, and encrypts the data.

[1628] 3. The server analyzes the image using an AI model and the text data using an NLP model.

[1629] 4. The server integrates the analysis results and generates a diagnosis that "there is a high possibility of a fracture."

[1630] 5. The server selects a nearby orthopedic hospital based on the user's location information.

[1631] 6. The server sends the diagnosis results and medical institution information to the terminal.

[1632] 7. The terminal displays the received information to the user, who then contacts the orthopedic hospital within the app to make an appointment and completes payment electronically.

[1633] Prompt Sentence Examples

[1634] "Please provide a proper diagnosis and recommended medical care for the injury image taken by the user and the symptom text entered: 'My hand is swollen and painful.'"

[1635] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1636] Step 1:

[1637] The user takes an image of their condition or injury using a smartphone or computer and saves the image data on the device. The user then enters their symptoms and thoughts about the condition in text format. The input image data is a JPEG or PNG file, and the text data is a string of characters.

[1638] Step 2:

[1639] The device packages the image and text data entered by the user into the specified format. For example, this data is converted into JSON format. The packaged data is then encrypted to ensure security. The pycryptodome library is used for encryption, and the data is processed using the AES encryption algorithm.

[1640] Step 3:

[1641] The device sends the encrypted data to the server via the Internet in encrypted JSON format. The server receives the data using the HTTPS protocol to ensure secure communication.

[1642] Step 4:

[1643] The server decrypts the received data and converts it back into JSON image and text data, again using the pycryptodome library and the AES decryption algorithm, before formatting the decrypted data for analysis.

[1644] Step 5:

[1645] The server analyzes image data and text data. To analyze the image data, a deep learning framework such as TensorFlow or PyTorch is used to extract image features using a convolutional neural network (CNN). To analyze the text data, a natural language processing (NLP) model is used to infer symptoms from the text entered by the user. A generative AI model integrates these analysis results to make a probable diagnosis. For example, a CNN model detects bone abnormalities, and an NLP model analyzes the text "My hand is swollen and painful" and infers that there is a high possibility of a fracture.

[1646] Step 6:

[1647] The server selects the most suitable medical institution based on the estimated diagnosis and the user's location information. The server retrieves a list of affiliated medical institutions from an internal database and creates a list of medical institutions that can provide consultations. The server uses a geographic information system (GIS) to select the most suitable medical institution based on the user's location information.

[1648] Step 7:

[1649] The server encrypts the estimated diagnosis and information about the recommended medical institution and sends it back to the device. The data sent includes the estimated condition, reliability, name, location, contact information, and available hours of the recommended medical institution. The data is again AES encrypted and transmitted securely using the HTTPS protocol.

[1650] Step 8:

[1651] The device decrypts the received data and displays the diagnosis results and medical institution information so that the user can check them. Based on the displayed diagnosis results, the user can make an appointment with the recommended medical institution. The device application also has a function to contact the medical institution directly from the displayed information.

[1652] Step 9:

[1653] The user makes a reservation at a recommended medical institution from their device. They enter reservation information within the app and send a reservation request to the server. The server forwards the reservation information to the medical institution and confirms the reservation. The confirmed reservation information is sent to the user's device.

[1654] Step 10:

[1655] The user uses the terminal to make an electronic payment for a medical appointment. The terminal uses an API such as Stripe or PayPal to send the payment information to the server. The server processes the payment through a payment gateway and sends a confirmation to the user if the payment is successful.

[1656] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1657] This invention is a system that allows users to input image and text data related to their medical condition or injury and receive referrals to appropriate medical institutions, even without specialized medical knowledge. Furthermore, by combining this system with an emotion engine that recognizes the user's emotions, it is possible to make diagnoses and select medical institutions that take into account the user's emotional state. A specific embodiment of this system is described below.

[1658] User input of images and text

[1659] Users use their smartphones or computers to take pictures of their illnesses or injuries and save the image data on their devices. They then enter their symptoms and thoughts about their condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text, "My hand is swollen and painful."

[1660] Sending data by the device

[1661] The device converts the image data and text data entered by the user into a specified format (e.g., JSON or XML) and packages it. The device then encrypts the data and sends it to the server via the Internet.

[1662] Receipt and analysis of data by the server

[1663] The server receives and decodes the image and text data sent from the device. It then uses a generative AI model to analyze the received image data. Specifically, it uses algorithms such as convolutional neural networks (CNNs) to extract image features and estimate the disease state. At the same time, the server analyzes the text data using a natural language processing (NLP) model to estimate the symptoms.

[1664] Emotion engine recognizes emotional states

[1665] The server uses an emotion engine to recognize the user's emotional state using text data. The emotion engine analyzes keywords and context in the text to determine the user's emotional state, such as stress, anxiety, or pain. For example, if the text contains expressions such as "It hurts so much" or "I'm so anxious," the corresponding emotional state is recognized.

[1666] Predictive diagnosis using generative AI models

[1667] The generative AI model on the server combines the results of image analysis, text analysis, and emotion engine analysis to make a probable diagnosis. This diagnosis includes the estimated medical condition, its reliability, and the user's emotional state. For example, if image analysis suggests a fracture, text analysis identifies symptoms such as "swelling and pain," and the emotion engine recognizes "strong anxiety," the generative AI model will generate a probable fracture diagnosis.

[1668] Server selection of appropriate medical institution

[1669] Based on the estimated diagnosis, the server references the user's current location information and searches a database for the most suitable medical institution. The information on the medical institution includes location, contact information, and available consultation hours. The server also takes into account the user's emotional state. For example, if the server detects strong anxiety or stress, it will prioritize medical institutions that offer psychological counseling.

[1670] Sending results from the server to the device

[1671] The server packages the estimated diagnosis and information on the recommended medical institution, encrypts it again, and transmits it to the terminal.

[1672] Displaying results on a terminal

[1673] The device decodes the received data and displays it in a user-friendly format, including the estimated diagnosis and detailed information about the recommended medical institution (e.g., location, contact information, and opening hours). The user can then review the displayed information and contact the appropriate hospital to arrange an appointment.

[1674] Specific examples

[1675] For example, if a user has a hand injury, the following process occurs:

[1676] 1. A user takes a photo of a hand injury with their smartphone and enters, "My hand is swollen and painful, and I'm very anxious."

[1677] 2. The device packages the image data and text data in JSON format and sends it to the server.

[1678] 3. The server analyzes the images using an AI model and the text data using an NLP model and emotion engine.

[1679] 4. The server integrates the results of each analysis and generates a diagnosis of "high probability of fracture" and an emotional state indicating "high anxiety."

[1680] 5. The server selects nearby orthopedic hospitals and medical institutions that provide psychological counseling, taking into account the user's location information and emotional state.

[1681] 6. The server sends the diagnosis results and medical institution information to the terminal.

[1682] 7. The device displays the received information to the user, who then contacts the appropriate medical institution and arranges for an examination.

[1683] The above is a specific embodiment of the present invention. This system allows users to receive prompt and accurate diagnosis and access to medical institutions, taking into account not only their medical condition but also their emotional state.

[1684] The processing flow will be explained below.

[1685] Step 1:

[1686] The user takes an image of the medical condition or injury using a smartphone or computer, saves the image data on the device, and then enters a description of their symptoms or condition into a text input form. For example, they might enter "My hand is swollen and painful."

[1687] Step 2:

[1688] The device packages the image and text data entered by the user into a specified format (e.g., JSON or XML), encrypts the data, and prepares it for transmission.

[1689] Step 3:

[1690] The device sends the encrypted image data and text data to the server, and the data is transmitted securely over the Internet.

[1691] Step 4:

[1692] The server decrypts the received encrypted data and prepares it for analysis. The server passes image data to the image analysis module and text data to the natural language processing (NLP) module.

[1693] Step 5:

[1694] The server analyzes the image data using an image analysis module. Specifically, it uses a convolutional neural network (CNN) to extract image features and estimate the condition of the patient. For example, it can detect signs of a fracture from the image.

[1695] Step 6:

[1696] The server analyzes the text data using a natural language processing (NLP) module. It extracts symptom-related keywords and context from the text and analyzes the symptoms. For example, it extracts information such as "swelling" and "pain."

[1697] Step 7:

[1698] The server uses an emotion engine to recognize the user's emotional state from the text data. The emotion engine analyzes keywords and context in the text to determine the user's emotional state (e.g., anxiety, stress, pain level, etc.).

[1699] Step 8:

[1700] The generative AI model combines the results of image analysis, text analysis, and emotion engine analysis to produce a probable diagnosis, which includes an estimated medical condition, its confidence level, and the user's emotional state.

[1701] Step 9:

[1702] Based on the estimated diagnosis, the server references the user's current location and searches a database for the most suitable medical institution. The server also takes into account the user's emotional state when selecting an appropriate medical institution. For example, if strong anxiety or stress is detected, it will prioritize medical institutions that offer psychological counseling.

[1703] Step 10:

[1704] The server packages the estimated diagnosis results and information about the selected medical institution, encrypts them again, and sends them to the terminal.

[1705] Step 11:

[1706] The device decodes the received data and displays it in a user-friendly format, including the estimated diagnosis and detailed information about the recommended medical institution (e.g., location, contact information, consultation hours, etc.).

[1707] Step 12:

[1708] The user checks the medical institution information displayed on the device and contacts the appropriate hospital to arrange for a medical examination. For example, the user contacts a hospital in response to a message that reads, "There is a high possibility of a fracture. We recommend the following orthopedic hospital."

[1709] Example 2

[1710] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1711] Currently, many people have difficulty accurately and quickly communicating information about their illness or injury to medical institutions. In particular, it is difficult for users to accurately describe their symptoms and emotional state, which can delay the selection of an appropriate medical institution and diagnosis. To solve this problem, a system is needed that can effectively analyze user input data, perform a comprehensive diagnosis including emotional state, and quickly provide guidance to an appropriate medical institution.

[1712] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1713] In this invention, the server includes a means for analyzing image data input by a user to estimate a medical condition, a means for analyzing text data to estimate symptoms, and a means for recognizing the user's emotional state from the text data using an emotion engine. This enables a highly accurate inferential diagnosis to be made by comprehensively evaluating the image analysis results, text analysis results, and emotion analysis results. Furthermore, based on this inferential diagnosis, the server can quickly select an appropriate medical institution by referring to the user's location information and provide that information to the user, thereby enabling the provision of prompt and appropriate medical services.

[1714] "User" refers to an individual who uses the system to input information about a medical condition or injury.

[1715] "Terminal" refers to a device used by a user to input image and text data related to a medical condition or injury and send it to a server. Examples of such devices include smartphones and computers.

[1716] "Image data" refers to photographic data of illnesses or injuries that users take using their terminals and enter into the system.

[1717] "Text data" refers to textual information such as the status and impressions of a medical condition or injury that a user inputs into a device.

[1718] A "format" refers to the rules and forms for organizing and structuring data, such as JSON and XML.

[1719] A "server" is a central computer system for receiving and analyzing image data and text data sent from the terminals.

[1720] A "generative AI model" is a collection of algorithms and programs that run on a server and integrate the results of image analysis and text analysis to make a presumptive diagnosis.

[1721] A "presumed diagnosis" is the evaluation result of the disease state or symptoms calculated by the generative AI model, and is a preliminary assessment before an actual diagnosis is made.

[1722] An "emotion engine" is a software component that analyzes text data and recognizes the user's emotional state.

[1723] "Location information" is geographical data that indicates the user's current location, and is obtained from GPS data or the like.

[1724] "Medical institution" refers to a clinic, doctor's office, hospital, etc. that provides medical services to users.

[1725] "Packaging" is the process of combining multiple pieces of data into one format, and is done to maintain data consistency.

[1726] "Encryption" is the process of transforming data using a specific algorithm to protect the confidentiality of the data.

[1727] "Decryption" is the process of restoring encrypted data to its original form.

[1728] "Analysis" is the process of analyzing data and extracting meaningful information and patterns from it.

[1729] This invention is a system that allows users to input image and text data about their illness or injury and receive referrals to appropriate medical institutions, even without specialized medical knowledge. Furthermore, by combining this system with an emotion engine that recognizes the user's emotions, it is possible to make diagnoses and select medical institutions that take the user's emotional state into consideration.

[1730] Hardware and software used

[1731] The system consists of the following components:

[1732] User terminal: A device through which a user inputs and receives data, such as a smartphone or computer.

[1733] Server: A central computer system that analyzes data and selects medical institutions.

[1734] Generative AI models: Algorithms that perform data analysis, such as convolutional neural networks (CNNs) and natural language processing (NLP) models.

[1735] Emotion Engine: A software component for recognizing the emotional state of a user from their text data.

[1736] Processing flow

[1737] A specific embodiment for implementing this system will be described below.

[1738] User input of images and text

[1739] Users use their smartphones or computers to take pictures of their illnesses or injuries and save the image data on their devices. They then enter their symptoms and thoughts about their condition in text format. For example, if a user injures their hand, they can take a photo of the injury and enter the text, "My hand is swollen and painful."

[1740] Sending data by the device

[1741] The device converts the image data and text data entered by the user into a specified format (e.g., JSON format), packages it, encrypts it, and sends it to a server via the Internet.

[1742] Receipt and analysis of data by the server

[1743] The server receives and decodes the image and text data sent from the device. The server then uses a generative AI model to analyze the image data. Specifically, it uses a convolutional neural network (CNN) to extract image features and infer the disease state. At the same time, the server analyzes the text data with a natural language processing (NLP) model to infer the symptoms.

[1744] Emotion engine recognizes emotional states

[1745] The server uses an emotion engine to recognize the user's emotional state based on the text data. The emotion engine analyzes keywords and context within the text to determine the user's emotional state, such as stress, anxiety, or pain. For example, if the text contains expressions such as "It hurts so much" or "I'm so anxious," the corresponding emotional state is recognized.

[1746] Predictive diagnosis using generative AI models

[1747] The generative AI model on the server integrates the results of image analysis, text analysis, and emotion engine analysis to make a probable diagnosis. This diagnosis includes the estimated medical condition, its reliability, and the user's emotional state. For example, if image analysis suggests a fracture, text analysis identifies symptoms such as "swelling and pain," and the emotion engine recognizes "strong anxiety," the generative AI model will generate a probable fracture diagnosis.

[1748] Server selection of appropriate medical institution

[1749] The server then references the user's current location information based on the estimated diagnosis and searches a database for the most suitable medical institution. The information on the medical institution includes the location, contact information, and available consultation hours. The server also takes into account the user's emotional state. For example, if the server detects strong anxiety or stress, it will prioritize medical institutions that offer psychological counseling.

[1750] Sending results from the server to the device

[1751] The server packages the estimated diagnosis results and information on recommended medical institutions, encrypts them again, and sends them to the terminal.

[1752] Displaying results on a terminal

[1753] The device decodes the received data and displays it in a user-friendly format, including the estimated diagnosis and detailed information about the recommended medical institution (such as location, contact information, and opening hours). The user can then review the displayed information and contact the appropriate hospital to arrange an appointment.

[1754] Specific examples

[1755] For example, if a user has a hand injury, the following process occurs:

[1756] 1. A user takes a photo of a hand injury with their smartphone and enters, "My hand is swollen and painful, and I'm very anxious."

[1757] 2. The device packages the image data and text data in JSON format and sends it to the server.

[1758] 3. The server analyzes the images using an AI model and the text data using an NLP model and emotion engine.

[1759] 4. The server integrates the results of each analysis and generates a diagnosis of "high probability of fracture" and an emotional state indicating "high anxiety."

[1760] 5. The server selects nearby orthopedic hospitals and medical institutions that provide psychological counseling, taking into account the user's location information and emotional state.

[1761] 6. The server sends the diagnosis results and medical institution information to the terminal.

[1762] 7. The device displays the received information to the user, who then contacts the appropriate medical institution and arranges for an examination.

[1763] This process allows users to access medical care and receive a quick and accurate diagnosis that takes into account not only their medical condition but also their emotional state.

[1764] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1765] Step 1:

[1766] User input of images and text

[1767] Users use a smartphone or computer to take pictures of their condition or injury, save the image data on the device, and then enter detailed symptoms and impressions of their condition in text format.

[1768] Input: Image data and text data related to medical conditions and injuries

[1769] Output: Image data and text data stored on the device

[1770] Specific behavior:

[1771] The user activates the camera and takes a picture of the affected area.

[1772] After taking the photo, you upload the image through the device's application and enter details such as "My hand is swollen and painful, and I'm very anxious" in the text input field.

[1773] Step 2:

[1774] Sending data by the device

[1775] The device converts the image data and text data entered by the user into a specified format (e.g., JSON format) and packages it. The device then encrypts the data and sends it to the server via the Internet.

[1776] Input: User image and text data

[1777] Output: Encrypted packaging data

[1778] Specific behavior:

[1779] The device converts the image file and text data into JSON format.

[1780] The data is encrypted using the SSL / TLS protocol and sent to the server's API endpoint.

[1781] Step 3:

[1782] Receipt and analysis of data by the server

[1783] The server receives and decrypts the encrypted data sent from the device. Next, the server analyzes the received image data using a generative AI model (CNN algorithm) to extract image features. At the same time, the server analyzes the text data using a natural language processing (NLP) model to estimate symptoms.

[1784] Input: Encrypted packaging data

[1785] Output: Analyzed image features and text analysis results

[1786] Specific behavior:

[1787] The server reads the encrypted data from the SSL / TLS input buffer and decrypts it.

[1788] The server inputs the image data into a generative AI model, extracts features, and estimates the condition of the disease.

[1789] An NLP processing engine is used to analyze the text data and extract keywords and context related to the symptoms.

[1790] Step 4:

[1791] Emotion engine recognizes emotional states

[1792] The server analyzes the text data with an emotion engine to recognize the user's emotional state. The emotion engine analyzes keywords and context within the text to determine the user's stress, anxiety, pain, etc.

[1793] Input: Text data

[1794] Output: Data about the user's emotional state

[1795] Specific behavior:

[1796] The server inputs the text data into an emotion engine, which analyzes the keywords and context within the text.

[1797] The emotion engine analyzes expressions such as "It hurts so much" and "I'm so anxious" to determine the user's emotional state.

[1798] Step 5:

[1799] Predictive diagnosis using generative AI models

[1800] The generative AI model on the server integrates the results of image analysis, text analysis, and emotion engine analysis to produce a probable diagnosis, which includes the probable medical condition, its reliability, and the user's emotional state.

[1801] Input: Image analysis results, text analysis results, sentiment analysis results

[1802] Output: Estimated diagnosis result

[1803] Specific behavior:

[1804] The server integrates the results of image analysis, text analysis, and sentiment analysis.

[1805] The generative AI model estimates the condition based on the integrated data and generates a diagnosis such as "high probability of fracture."

[1806] Step 6:

[1807] Server selection of appropriate medical institution

[1808] Based on the estimated diagnosis, the server references the user's current location information and searches a database for the most suitable medical institution, including information on the location, contact details, and available hours for consultations.

[1809] Input: Estimated diagnosis result, user's current location information

[1810] Output: Appropriate medical institution information

[1811] Specific behavior:

[1812] The server obtains the user's current location information from GPS data and queries a location information database.

[1813] The server searches the medical institution database and generates a list of medical institutions that can provide consultations.

[1814] Step 7:

[1815] Sending results from the server to the device

[1816] The server packages the estimated diagnosis and information on the recommended medical institution, encrypts it again, and transmits it to the terminal.

[1817] Input: Estimated diagnosis result, medical institution information

[1818] Output: Encrypted diagnosis results and medical institution information

[1819] Specific behavior:

[1820] The server packages the diagnosis results and medical institution information in JSON format.

[1821] The server encrypts the packaged data and sends it to the user's terminal.

[1822] Step 8:

[1823] Displaying results on a terminal

[1824] The device decodes the received data and displays it in a user-friendly format, including a probable diagnosis and detailed information about recommended medical institutions.

[1825] Input: Encrypted diagnosis results and medical institution information

[1826] Output: Decrypted diagnosis and medical institution information

[1827] Specific behavior:

[1828] The device decodes the data it receives and parses the JSON format data to extract the diagnosis results and medical institution information.

[1829] The device updates the application UI and displays the diagnosis results and medical institution information to the user.

[1830] Through this detailed, step-by-step process, users can obtain an accurate and prompt diagnosis and access to the appropriate medical facility.

[1831] (Application example 2)

[1832] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1833] Conventional medical diagnosis systems have the problem that it is difficult for users without medical expertise to make an appropriate diagnosis and select a medical institution. Furthermore, there is a need for rapid medical diagnosis and guidance to medical institutions in emergency situations. In particular, a system that can provide appropriate information immediately is needed so that security staff can respond quickly on the scene.

[1834] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1835] In this invention, the server includes: a means for a user to input image data and text data related to a medical condition or injury; a means for a terminal to package the image data and text data into a specified format and transmit the packaged image data and text data to the server; a means for the server to analyze the image data and estimate a medical condition; a means for the server to analyze the text data and estimate symptoms; a generative AI model to integrate the analysis results and make an estimated diagnosis; a means for the server to select an appropriate medical institution based on the estimated diagnosis; a means for the server to transmit information about the estimated diagnosis and the medical institution to the terminal; a means for the terminal to display the information to the user; and a means for performing a rapid medical diagnosis on site using smart glasses. This enables a user (security staff) to receive a rapid and appropriate medical diagnosis on site and access an appropriate medical institution based on the results.

[1836] "Image data" is digital data entered by a user that contains visual information about a medical condition or injury.

[1837] "Text data" is digital data containing written information related to a medical condition or injury entered by a user.

[1838] A "user" is an individual who uses the system and who inputs information about a medical condition or injury.

[1839] A "terminal" is a device operated by a user, and is a device for packaging image data and text data and transmitting them to a server.

[1840] A "server" is a computer system that receives data sent from a terminal and performs analysis and inferential diagnosis.

[1841] A "generative AI model" is an artificial intelligence algorithm that integrates the analysis results of image data and text data to make a presumptive diagnosis.

[1842] "Format" refers to a data format for arranging image data and text data in a form that is easy for the server to process.

[1843] "Smart glasses" are wearable devices equipped with a camera function that users use in the field.

[1844] An "appropriate medical institution" is a medical facility that can provide optimal medical services to the user based on the estimated diagnosis results.

[1845] A "presumed diagnosis" is the result of a disease state inferred by a generative AI model based on an analysis of image data and text data.

[1846] "Display means" refers to a function that visually presents to the user the information that the terminal receives from the server.

[1847] "Emotional state" indicates the psychological and emotional state of the user, which is analyzed from text data or the like.

[1848] In order to implement the present invention, the following configurations and means are necessary.

[1849] Overall system overview

[1850] The user inputs image and text data of their condition or injury and sends the information to a server. The server analyzes the data, makes a probable diagnosis, selects an appropriate medical institution, and provides that information to the user. Furthermore, the use of smart glasses enables rapid medical diagnosis on the spot.

[1851] Hardware and Software

[1852] 1. Smart Glasses:

[1853] A wearable device with a camera function.

[1854] Supports on-site image capture and data entry.

[1855] 2. Terminal:

[1856] Personal computers and smartphones.

[1857] It supports the input of image data and text data and sends the data to the server.

[1858] 3. Server:

[1859] High performance computing systems.

[1860] Analyzes image and text data, performs presumptive diagnoses, and selects appropriate medical institutions.

[1861] The algorithms used are convolutional neural networks (CNN) and natural language processing (NLP).

[1862] Program processing

[1863] The device packages the image data and text data entered by the user into a specified format (e.g., JSON), encrypts it, and sends it to the server. The server decrypts the received data and performs the following processes.

[1864] 1. Analysis of imaging data:

[1865] The server uses a convolutional neural network (CNN) to extract features from the image data and estimate the disease state.

[1866] 2. Text data analysis:

[1867] Text data is analyzed using a natural language processing (NLP) model to estimate symptoms.

[1868] 3. Integrated analysis using generative AI models:

[1869] The generative AI model integrates the results of image analysis, text analysis, and the emotion engine analysis to make a presumptive diagnosis.

[1870] 4. Selection of appropriate medical facility:

[1871] Based on the estimated diagnosis results, the server selects the most appropriate medical institution, taking into consideration the user's location information and emotional state.

[1872] 5. Transmission and Display of Information:

[1873] The server transmits the generated estimated diagnosis and information about the medical institution to the terminal, which then displays them to the user.

[1874] Specific examples

[1875] For example, if security staff discover someone collapsed at a scene, they can take the following steps:

[1876] 1. Taking a photo of a fallen person using smart glasses.

[1877] 2. Enter the text data "Unconscious, not breathing."

[1878] 3. The device sends this data to the server.

[1879] 4. The server makes a probable diagnosis and generates information such as "possibility of cardiac arrest" and "nearest hospital with a cardiologist."

[1880] 5. The server sends the diagnosis results and information on recommended medical institutions to the terminal, which displays them to the user.

[1881] Prompt Sentence Examples

[1882] "Unconscious, not breathing"

[1883] "My hands are swollen and painful, and I'm very anxious"

[1884] By using these prompts, users can quickly and accurately access the appropriate medical institution.

[1885] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1886] Step 1:

[1887] The user uses the smart glasses to take images of their condition or injury and input text data about their symptoms. The user uses the smart glasses' camera function to take photos of their condition or injury, and then inputs text information such as "My hand is swollen and painful, and I feel very anxious" through the smart glasses' UI.

[1888] Input: Image and text data of medical conditions and injuries.

[1889] Output: Photographed image data and input text data.

[1890] Step 2:

[1891] The device packages the image and text data entered by the user into a specified format, encrypts it, and sends it to the server.The device then converts the data into a structured data format such as JSON and encrypts it using an encryption algorithm such as AES for security.The encrypted data is then sent to the server via the HTTPS protocol.

[1892] Input: Image data and text data.

[1893] Output: Encrypted and formatted data.

[1894] Step 3:

[1895] The server receives and decrypts the data sent from the device. The server receives the data via the HTTPS protocol and decrypts data encrypted with AES or other encryption protocols. It then analyzes the received JSON format data and separates the image data from the text data.

[1896] Input: Encrypted and formatted data.

[1897] Output: Decoded and separated image and text data.

[1898] Step 4:

[1899] The server analyzes the image data to estimate the condition of the patient. Specifically, it uses a convolutional neural network (CNN) to extract image features and estimate the condition. Based on the results of this analysis, a diagnosis such as "high possibility of fracture" is obtained.

[1900] Input: Image data.

[1901] Output: Estimated disease state based on image analysis.

[1902] Step 5:

[1903] The server analyzes the text data to infer symptoms. It uses a natural language processing (NLP) model to analyze keywords and context within the text data to infer symptoms. For example, from an expression such as "My hand is swollen and painful, and I feel very anxious," it extracts symptoms such as "swelling," "pain," and "anxiety."

[1904] Input: Text data.

[1905] Output: Symptom prediction results based on text analysis.

[1906] Step 6:

[1907] The generative AI model combines the results of image analysis, text analysis, and the emotion engine to generate a comprehensive diagnosis, such as "There is a high possibility of a fracture and the patient feels high anxiety."

[1908] Input: Image analysis results, text analysis results, emotion engine results.

[1909] Output: Integrated estimated diagnostic results.

[1910] Step 7:

[1911] The server selects an appropriate medical institution based on the user's location information and emotional state. Based on the estimated diagnosis, the server searches a database of medical institutions and selects the most suitable medical institution. For example, the server selects the "nearest orthopedic hospital" and "medical institution that offers psychological counseling" based on the user's location information and emotional state.

[1912] Input: Estimated diagnosis result, user location information, emotional state.

[1913] Output: Information on selected medical institutions.

[1914] Step 8:

[1915] The server then sends the generated estimated diagnosis and information about the medical institution to the terminal, where it re-encrypts the information and transmits it to the terminal using the HTTPS protocol.

[1916] Input: Estimated diagnosis result, medical institution information.

[1917] Output: Encrypted diagnosis results and medical institution information.

[1918] Step 9:

[1919] The device decrypts the information received from the server and displays it in a format that is easy for the user to understand. The device decrypts the encrypted data and displays information such as "High probability of fracture," "Nearest orthopedic hospital: XX Hospital," and "Psychological counseling: YY Clinic" through the UI.

[1920] Input: Encrypted diagnosis results and medical institution information.

[1921] Output: Presenting information in a user-friendly format.

[1922] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1923] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1924] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1925] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's ...

Claims

1. a means for a user to input image and text data relating to a medical condition or injury; a means for packaging the image data and text data in a designated format in the terminal and transmitting the packaged data to a server; A server analyzes the image data and estimates the condition of the patient; A server analyzes the text data and estimates symptoms; A means for a generative AI model to integrate the analysis results and perform a presumptive diagnosis; A means for the server to select an appropriate medical institution based on the estimated diagnosis; a server that transmits the estimated diagnosis and information on the medical institution to a terminal; means for the terminal to display said information to a user; A system including:

2. The system of claim 1, wherein the generative AI model integrates the results of image analysis and text analysis to make a presumptive diagnosis.

3. 2. The system according to claim 1, wherein the server selects an appropriate medical institution based on the user's location information and transmits the information to the terminal.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A