Information processing system
Patent Information
- Application Number
- CN202610242478.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-02-28
- Publication Date
- 2026-09-22
AI Technical Summary
此外,所述处理器还被配置为:基于所述解析结果生成用于指示推荐适当医疗机构的提示词,并生成用于向用户输出的推荐消息,以便将疾病类型及重症度信息与医疗机构资源信息关联,向用户推送个性化的就医推荐,完成从智能问诊、风险评估到就医引导和用药支持的一体化技术方案,从而有效解决现有技术中疾病诊断不准确、重症度评估不足以及与医疗资源联动不充分的问题
[0005]进一步地,本发明的系统中,所述处理器还被配置为:与邻近的医疗机构联动,选择适当的药品,并生成用于指示所述生成式人工智能模型对所述药品的提供进行控制的提示词,从而在完成疾病诊断与重症度判断后,能够基于标准化的提示词与医疗机构或药房系统进行交互,实现药品的自动选定与提供,提高线上健康服务与线下药品供应之间的协同效率。此外,所述处理器还被配置为:基于所述解析结果生成用于指示推荐适当医疗机构的提示词,并生成用于向用户输出的推荐消息,以便将疾病类型及重症度信息与医疗机构资源信息关联,向用户推送个性化的就医推荐,完成从智能问诊、风险评估到就医引导和用药支持的一体化技术方案,从而有效解决现有技术中疾病诊断不准确、重症度评估不足以及与医疗资源联动不充分的问题。
Smart Images

Figure CN122800178A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to an information processing system. Background Technology
[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.
[0003] The main technical challenge this invention addresses is that existing AI-based health consultation or online diagnosis systems generally suffer from the following problems: First, these systems often only perform simple matching or rule-based judgments on user-input health information, making it difficult to provide timely and accurate possible disease names and lacking the ability to comprehensively analyze complex symptoms. Second, after providing a suspected disease name, the system often fails to systematically assess the severity of the disease using authoritative databases, resulting in an inability to provide users with tiered risk warnings and medical advice. Third, existing systems lack close integration with offline medical services, lacking standardized interactive instructions and control logic, making it difficult to coordinate with nearby medical institutions for drug selection and provision, and also difficult to recommend suitable medical institutions to users based on specific analysis results. Therefore, there is an urgent need for a system that can: automatically generate prompts for disease diagnosis and severity assessment based on user health status information using generative AI models; and, based on this, coordinate with medical institutions for drug provision and generate medical institution recommendation information, thereby improving the intelligence and reliability of the connection between health consultation and medical services. Summary of the Invention
[0004] To address the aforementioned technical challenges, this invention provides an information processing system comprising a processor configured to: receive health-related information from a user; parse the received information using a generative artificial intelligence (AI) model and generate prompt words instructing the AI model to parse the information to diagnose possible disease names; query a database to determine the severity of the diagnosed disease name and generate prompt words instructing the AI model to evaluate the severity based on the database. Through this structure, this invention utilizes the natural language understanding and generation capabilities of a generative AI model to convert the user's unstructured health information into standardized prompt words that the model can use to perform disease diagnosis and severity assessment, thereby improving the accuracy of disease inference and the precision of severity assessment.
[0005] Furthermore, in the system of the present invention, the processor is also configured to: collaborate with nearby medical institutions, select appropriate medications, and generate prompts to instruct the generative artificial intelligence model to control the provision of the medications. This allows the processor to interact with medical institutions or pharmacy systems based on standardized prompts after disease diagnosis and severity assessment, enabling automatic medication selection and provision, and improving the collaborative efficiency between online health services and offline medication supply. In addition, the processor is also configured to: generate prompts to recommend appropriate medical institutions based on the analysis results, and generate recommendation messages to be output to the user. This associates disease type and severity information with medical institution resource information, pushing personalized medical recommendations to the user, thus completing an integrated technical solution from intelligent consultation and risk assessment to medical guidance and medication support. This effectively solves the problems of inaccurate disease diagnosis, insufficient severity assessment, and inadequate collaboration with medical resources in existing technologies.
[0006] "System" refers to an overall device or platform consisting of one or more hardware components and / or software modules, used to perform processing functions such as receiving user health information, calling generative artificial intelligence models, accessing databases, linking with medical institutions, and generating output information. It can be deployed on local devices, servers, or in cloud environments.
[0007] "Processor" refers to a hardware and / or virtual computing unit used to execute computer program instructions, process input data, and perform functions such as receiving user information, generating prompts, accessing databases, and controlling external system linkages, including but not limited to a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a combination thereof.
[0008] "User" refers to an entity that uses the system and provides health status-related information to the system. This entity can be an individual patient or their agent, including any human individual who interacts with the system through a terminal device.
[0009] "Information related to health status" refers to various data that can reflect the user's current or past physical and mental health status, including but not limited to symptom descriptions, body temperature, heart rate, past medical history, allergy history, medication history, physical examination results, and other subjective or objective information related to disease diagnosis.
[0010] "Generative AI models" refer to AI models that can automatically generate text, vectors, or other forms of output based on input data. In particular, models that can parse user health information based on prompts and generate content related to disease diagnosis and severity assessment include, but are not limited to, large language models, multimodal generative models, or their variants.
[0011] "Prompt words" refer to instructional text or structured data generated by the processor and input into the generative artificial intelligence model. They are used to explicitly specify the parsing goals, diagnostic tasks, or control logic of the generative artificial intelligence model, thereby guiding the model to perform operations such as diagnosing disease names, assessing the severity of illness, controlling drug provision, or recommending medical institutions based on the user's health information.
[0012] "Analysis" refers to the process by which a generative artificial intelligence model, after receiving prompts and information related to a user's health status, understands, extracts features, performs semantic analysis, and makes inferences on that information in order to output results corresponding to disease diagnosis, severity assessment, or other medical-related tasks.
[0013] "Possible disease names" refer to one or more suspected disease names inferred by the generative artificial intelligence model after analyzing information related to the user's health status. These names represent the most likely disease type corresponding to the user's current health status, but are not guaranteed to be the final diagnosis.
[0014] A database is a collection of structured or semi-structured data used to store and retrieve data from a processor. This includes medical knowledge or experience data related to disease names, severity indicators, treatment guidelines, risk grading rules, etc., and can be a relational database, a non-relational database, or a distributed data storage system.
[0015] "Severity" refers to an indicator or classification result used to characterize the severity, urgency, and potential risk level of a disease or health condition. It can take the form of numerical scores, level labels (such as mild, moderate, severe), or a combination of both.
[0016] "Assessing the severity of illness" refers to the process of analyzing, scoring, or classifying the risk level of a diagnosed disease based on disease-related knowledge and rules stored in a database, thereby arriving at quantitative or qualitative conclusions regarding the severity and urgency of the current disease.
[0017] "Nearby medical institutions" refers to medical service units that are geographically located in the user's location or within a preset range and have the ability to provide medical services or supply medicines, including but not limited to hospitals, clinics, community health service centers, and cooperative medical institutions that have established data or business linkages with the system.
[0018] "Linkage" refers to the process by which the processor interacts with external medical institutions, pharmacy systems, etc., through a predetermined communication interface or protocol, including operations such as sending prescription information, receiving review results, checking drug inventory, or making appointments.
[0019] "Appropriate medicines" refer to medicines that are selected from multiple options based on the name and severity of the diagnosed disease, combined with the user's individual characteristics and contraindications, and that have a high degree of match in terms of efficacy, safety and indications.
[0020] "Medicine provision" refers to the process of supplying selected appropriate medicines to users in the form of prescriptions, preparation, or dispensing through collaboration with medical institutions or pharmacy systems. This includes generating prescription information and arranging medicine dispensing or delivery.
[0021] "Analysis results" refers to the output content obtained by the generative artificial intelligence model after receiving prompt words and analyzing information related to the user's health status. This content includes at least the possible disease name and / or the corresponding severity assessment, and may also include additional information such as explanations of related symptoms and risk warnings.
[0022] "Recommending appropriate medical institutions" refers to the process of screening and ranking multiple medical institutions based on information such as disease type and severity contained in the analysis results, combined with geographical location, hospital resources and specialty settings, in order to determine one or more medical institutions that are more suitable for the user to seek medical treatment.
[0023] "Recommended message" refers to the message content generated by the processor and output to the user to prompt or guide the user to select a medical institution. This content may include information such as the name, address, department type, consultation suggestions and related contact information of the recommended medical institution. Attached Figure Description
[0024] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.
[0025] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0026] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.
[0027] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0028] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.
[0029] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.
[0030] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.
[0031] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.
[0032] Figure 9 This represents an emotion map that maps multiple emotions.
[0033] Figure 10 This represents an emotion map that maps multiple emotions.
[0034] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.
[0035] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.
[0036] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.
[0037] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation
[0038] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.
[0039] First, let me explain the terminology used in the following instructions.
[0040] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0041] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.
[0042] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.
[0043] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.
[0044] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.
[0045] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0046] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.
[0047] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0048] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.
[0049] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.
[0050] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0051] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0052] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0053] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0054] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0055] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0056] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0057] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."
[0058] In existing health consultation and preliminary diagnosis systems, servers typically perform simple keyword matching or rule-based judgments on user-inputted health status information. Servers struggle to fully understand the contextual information and implicit semantics within natural language descriptions, leading to inaccurate disease candidate results and poor scalability. When utilizing generative AI models, servers often use fixed prompts to invoke the model, making it difficult to dynamically adjust prompts based on model output quality and user input characteristics. This results in unstable and poorly structured generated results, hindering automated severity assessment and resource recommendation in subsequent data processing. Furthermore, the lack of a unified data structure and processing flow when associating unstructured text output from generative AI models with structured data in medical databases makes it difficult for servers to perform efficient joint calculations of disease candidate information, disease classification information, and disease severity information, and to promptly return actionable guidance information to user terminals. Therefore, improving the server-side processes for acquiring, preprocessing, generating prompts, invoking generative AI models, and subsequent structured parsing and severity calculation of health status information—to enhance the accuracy and usability of disease candidate information, improve overall computational efficiency, and improve system response quality—has become an unresolved computer technology problem.
[0059] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.
[0060] In this invention, the server includes a device for acquiring health status-related information from a user terminal, a device for converting the health status-related information into structured information, a device for generating prompts for a generative artificial intelligence model based on the health status-related information and the structured information, and acquiring disease candidate information through the generative artificial intelligence model, a device for parsing the disease candidate information and extracting disease classification information, a device for calculating disease severity information based on the disease classification information and the health status-related information reference information storage device, a device for generating diagnostic result information and behavioral guidance information based on the disease candidate information and the disease severity information, and a device for storing the diagnostic result information and... The device for sending behavioral guidance information to the user terminal; wherein, the device for generating prompt statements for generative artificial intelligence models is configured to dynamically update the prompt statements based on the content of the health status-related information and the response format of the generative artificial intelligence model, and to repeatedly control the query processing for the generative artificial intelligence model to improve the accuracy of the disease candidate information; the device for generating diagnostic result information and behavioral guidance information is configured to extract candidates of medical service providers and medical-related items based on the disease severity information and the diagnostic result information, referencing a medical-related resource information storage device, and generate recommendation information messages containing the candidates as the behavioral guidance information. Thus, on the server side, a unified data structure and processing flow can be used for preprocessing of health status-related information in natural language form, adaptive generation of prompt statements, invocation of generative artificial intelligence models, structured parsing of results, and joint calculation with the database, achieving high-precision generation of disease candidate information and disease severity information, improving the usability and controllability of generative artificial intelligence model output in medical assistance scenarios, reducing server processing overhead, and improving the overall system response efficiency and stability, thereby substantially improving the health status analysis and decision support capabilities based on generative artificial intelligence models in the field of computer technology.
[0061] "System" refers to an integrated computer implementation consisting of one or more processing devices, storage devices, and communication devices, used to perform health status information processing, model invocation, and result output.
[0062] "User terminal" refers to an electronic device operated by a user for inputting health status-related information and receiving diagnostic results and behavioral guidance information, including but not limited to smartphones, tablets, personal computers, or wearable devices.
[0063] "Health status related information" refers to data input by users through user terminals that reflects the user's current or past physical condition, including but not limited to symptom descriptions, physiological parameters, past disease information, and medication use.
[0064] "Structured information" refers to machine-processable data generated based on health status-related information and organized according to a predetermined data format, including fielded, tagged, or encoded information, for subsequent calculation and retrieval.
[0065] "Generative artificial intelligence models" refer to artificial intelligence models trained based on machine learning algorithms that can automatically generate text output based on input prompts. These models are used to perform natural language analysis on health status-related information and generate candidate disease information.
[0066] "Prompt statements" refer to natural language text or semi-structured text constructed by the server to invoke generative artificial intelligence models, which are used to provide task descriptions, input constraints, and output format requirements to the generative artificial intelligence models.
[0067] "Disease candidate information" refers to the result data set output by a generative artificial intelligence model after receiving prompts and related inputs, which contains one or more possible disease names and their related descriptions.
[0068] "Disease classification information" refers to structured information extracted and summarized from disease candidate information, used to classify or label diseases, including disease type, onset location, etiology category or risk level, etc.
[0069] "Information storage device" refers to storage resources used to electronically store and manage data related to health status analysis, including database systems, file storage systems, or cloud storage services.
[0070] "Disease severity information" refers to indicator information calculated based on disease classification information and health status-related information, used to characterize the severity or risk level of a disease, including classification results, risk coefficients, or criticality markers.
[0071] "Diagnosis result information" refers to a data set generated by the server after integrating disease candidate information, disease severity information, and related rules, which is used to represent a comprehensive judgment result on the user's health status.
[0072] "Behavioral guidance information" refers to the suggestions generated by the server based on diagnostic results and disease severity information to guide users in taking subsequent actions, including medical advice, self-observation advice, or medication-related advice.
[0073] "Medical-related resource information storage device" refers to a storage device that stores data related to medical service resources, including databases or other data storage systems for medical service providers, pharmaceuticals, service accessibility, and cost information.
[0074] "Medical service providers" refers to organizations or institutions that can provide medical-related services to users, including medical institutions, telemedicine platforms, or other entities with medical service functions.
[0075] "Medical-related items" refer to physical resources related to the prevention, mitigation or treatment of diseases, including medicines, medical devices, health care products and other items used for health management.
[0076] This invention will be described in detail focusing on the collaborative processing of the server, terminal, and user. It will also provide a detailed description of how the system implemented under this invention is achieved, covering aspects such as hardware configuration, software modules, data structures, algorithm flow, and specific application methods of the generative artificial intelligence model. This embodiment targets health status-related information, but the technical concept of this invention can also be extended to other natural language health consultation scenarios.
[0077] I. Overall Hardware and Software Composition of the System In this invention, the server operates as the core computing node.
[0078] The hardware used in servers includes a central processing unit (CPU), main memory (RAM), persistent storage (solid-state drive), network interface card, and optional graphics processing unit (GPU). Servers establish communication connections with multiple terminals via a local area network (LAN) or the Internet.
[0079] The basic software used by servers includes: Servers use operating systems, such as Linux kernel-based server operating systems.
[0080] The server uses middleware and backend frameworks, such as Python-based web frameworks (like Flask), Java-based web frameworks (like Spring), or JavaScript-based runtime environments (like Node.js).
[0081] The server uses a database management system, such as a relational database (like MySQL or PostgreSQL) or a document-oriented database (like MongoDB), as an information storage device and a storage device for medical-related resource information.
[0082] The server uses a generative artificial intelligence model interface, such as a cloud interface system based on a large-scale pre-trained language model (a generative artificial intelligence model similar to GPT-4), to call the model via an HTTPS API.
[0083] In this invention, the terminal is used as a user interaction node.
[0084] The hardware used in the terminal includes mobile terminals (such as smartphones), tablet computing devices, or desktop computing devices, and the terminal has a touch screen or keyboard input device, a display, a processor, and local storage.
[0085] The underlying software used by the terminal includes mobile operating systems (such as Android and iOS), desktop operating systems (such as Windows), and front-end applications, such as native mobile applications or browser front-end applications. Front-end applications can be implemented using HTML, CSS, JavaScript, or mobile application frameworks (such as React Native and Flutter frameworks).
[0086] In this invention, users operate a terminal to input health status-related information in natural language and browse diagnostic results and behavioral guidance information returned by the server on the terminal.
[0087] II. Server-side module composition and program implementation 1. Health Information Receiving and Structured Module The server in this invention includes a health information receiving module.
[0088] The server provides a set of network interfaces through a web application framework to receive health status-related information from the terminal.
[0089] After receiving a request message from the terminal containing a natural language description of symptoms, body temperature value, and past medical information, the server creates a corresponding request object in memory. This object is organized in key-value pairs, such as fields like "symptom text", "body temperature value", and "illness text".
[0090] The server in this invention includes a structured information generation module.
[0091] The server preprocesses natural language fields using algorithms such as word segmentation, syntactic segmentation, and regular expression parsing.
[0092] The server uses a symptom dictionary or medical terminology table to map key symptoms in user input to standardized labels such as “headache”, “cough”, and “high fever”.
[0093] The server parses the body temperature string into a floating-point number and adds a unit marker field.
[0094] The server matches past disease texts with an internal coding table, mapping texts such as "hypertension" to standard disease codes.
[0095] The server thus generates a structured information object, which uses a fixed data structure (such as a record or document containing multiple fields) so that subsequent algorithm modules can process it efficiently.
[0096] Through this structured processing, the server transforms unstructured natural language into a unified data representation that can be reused by multiple algorithms, thereby reducing the computational overhead of subsequent repeated parsing and improving overall processing efficiency.
[0097] 2. Prompt Statement Generation and Dynamic Adjustment Module The server in this invention includes a prompt statement generation module.
[0098] The server constructs prompts for a generative artificial intelligence model based on health status information and structured information. The server maintains multiple prompt templates in memory, which include fixed description sections and insertable variable sections.
[0099] In one example, the server generates the following prompt: "The user's symptoms are: headache and cough. Body temperature is 37.8 degrees Celsius. Pre-existing condition is: hypertension. Based on the above information, please use your medical knowledge to diagnose the possible diseases. List at least two possible diseases and sort them from most likely to least likely. Provide a one-sentence description for each disease." In another example, the server generates the following prompt: "The user's symptoms are: persistent cough and sore throat. Body temperature: 39.2 degrees Celsius. The user has no known illnesses. Please diagnose possible diseases based on this information, and briefly describe the typical symptoms and approximate severity of each disease (preliminary judgment of mild / moderate / severe). Output only medically relevant content." In this invention, the server dynamically adjusts the prompt statements.
[0100] The server adjusts the constraints of the prompt statements based on the quality of past responses from the generative artificial intelligence model (such as the degree of structure and whether there are missing items) and the complexity of the current input (number of symptoms, number of diseases, etc.). For example, it may add output format specifications, limit the number of characters, or require a list of items to be returned.
[0101] By dynamically adjusting the parameters, the server optimizes the contextual conditions input to the generative AI model without modifying the model's internal parameters. This makes the model's natural language output more consistent with the predetermined structure, thereby reducing the burden of post-processing on the server side and improving computational efficiency.
[0102] 3. Description of Generative Artificial Intelligence Model Invocation and Internal Structure The server in this invention includes a generative artificial intelligence model invocation module.
[0103] The server constructs a request for the cloud-based generative artificial intelligence model service using an HTTP client library. The request includes control information such as the model name, prompts, temperature parameters, and maximum generation length parameters.
[0104] In one example, the server sets the temperature parameter to a lower value to make the model output more deterministic; when more diverse output is needed, the server can set the temperature parameter slightly higher to obtain a variety of disease candidates.
[0105] In one implementation, the generative artificial intelligence model employs a multi-layer Transformer neural network structure based on a self-attention mechanism.
[0106] Generative artificial intelligence models include embedding layers, which map input characters or tokenized words to a high-dimensional vector space; Generative artificial intelligence models consist of several encoder-decoder stacked layers, each of which includes a self-attention sub-layer, a feedforward fully connected sub-layer, and a normalization layer; Generative artificial intelligence models use a self-attention mechanism to calculate the relevance weights between positions in the input sequence in order to capture long-distance dependencies and contextual semantic relationships; Generative AI models use large-scale corpora during the training phase and update network weights by minimizing the cross-entropy loss function. The weight update algorithm uses stochastic gradient descent with kinetic terms or adaptive learning rate optimization algorithm. Generative AI models can be fine-tuned based on medical corpora during pre-training, and the server only performs forward propagation calculations during the inference phase without updating the weights.
[0107] In this invention's system, the server does not retrain the model's internal weights online. Instead, it fully utilizes the model's existing semantic representation and reasoning capabilities through carefully designed prompts and structured preprocessing. This approach delegates complex semantic reasoning to neural networks with highly nonlinear expressive capabilities, while employing rule-based data flow management and structured analysis on the server side to optimize the allocation of computing resources and tasks.
[0108] 4. Model Output Parsing and Disease Classification Information Extraction Module The server in this invention includes a model output parsing module.
[0109] After receiving the text returned by the generative artificial intelligence model, the server uses methods such as string splitting, regular expression matching, and template matching to parse the names and corresponding descriptions of each disease.
[0110] The server splits the text into several candidate disease records according to the output format agreed upon in the prompt statement (such as a numbered list or semicolon-separated format).
[0111] The server assigns a unique identifier to each record and stores fields such as disease name, descriptive text, and preliminary severity description in its internal data structure.
[0112] The server in this invention includes a disease classification information extraction module.
[0113] The server obtains the standard code and category of a disease (such as respiratory diseases, circulatory system diseases, etc.) by matching the disease name with the main disease table in the information storage device.
[0114] The server uses rule mapping or lightweight classification models to extract keywords such as "mild," "moderate," "severe," or "high-risk" from the model text description as structured attributes.
[0115] The server thus generates a disease classification information object, which includes fields such as disease code, disease category, preliminary severity level label, and possible danger signal markers.
[0116] 5. Interaction module for calculating disease severity and storing information The server in this invention includes a module for calculating disease severity information.
[0117] The server accesses a pre-stored table of basic severity rules for diseases in the information storage device. This table contains the basic severity, common age groups, common complications, etc. for each disease.
[0118] The server takes disease classification information, the user's structured body temperature value, and disease information as input, and performs a severity calculation according to a predefined algorithm.
[0119] In one example, the server uses the following computational logic: If the server detects a body temperature field value greater than or equal to 38.5 degrees Celsius, it will raise the severity level of the corresponding disease by one level. If the server detects that the disease information contains codes related to hypertension or cardiovascular disease, it will add a risk factor to diseases related to the cerebrovascular or cardiovascular system. If the server detects keywords such as "sudden onset," "severe," or "confusion" in the model description text, it will directly mark the relevant disease as a high-risk category.
[0120] Through this non-conventional processing flow that combines rules with structured information, the server can correct and finely classify model outputs without relying entirely on a single model output to make a final judgment. This reduces the uncertainty brought about by a single natural language output, thereby improving the credibility and interpretability of severity assessment.
[0121] After the server completes the calculation, it stores the disease severity information along with the disease candidate information and disease classification information in an information storage device, forming a searchable record of all previous diagnoses. This centralized storage facilitates the optimization of subsequent statistical analysis and model invocation strategies, and improves data management.
[0122] 6. Diagnostic Result Information and Behavioral Guidance Information Generation Module The server in this invention includes a diagnostic result generation module.
[0123] The server ranks the disease candidates according to a preset priority rule based on the combined probability of multiple disease candidate information, the severity level, and the combination of symptoms currently entered by the user, and selects one or more main recommended diseases.
[0124] The server uses template generation technology to populate structured data into natural language interpretation templates, generating diagnostic result information text that includes summaries of primary disease, secondary disease, and severity level.
[0125] The server in this invention includes a behavior guidance information generation module.
[0126] The server accesses the medical-related resource information storage device to query the inventory or availability information of medical service providers and medical-related items related to the current user's geographical location or network area.
[0127] The server combines information on the severity of the disease to select the appropriate level of medical service provider. For example, it recommends general outpatient clinics in mild cases and emergency or specialized medical institutions in moderate or high-risk cases.
[0128] When generating behavioral guidance information, the server may include the following: whether immediate medical attention is required, the recommended type of department to consult, whether home observation is suitable, and warning signs to be aware of.
[0129] In this way, the server connects the results of multi-source structured data computation within the computer with real-world medical service resources, enabling the invention to go beyond abstract data processing and have a direct technical impact on actual medical resource usage and user behavior.
[0130] 7. Terminal display and user interaction methods The terminal in this invention includes a user interface display module.
[0131] After receiving diagnostic results and behavioral guidance information from the server, the terminal displays the candidate disease names and their corresponding severity levels in a list format on the screen.
[0132] The terminal can use different colors or icons to represent different levels of severity of illness in order to enhance users' perception of the importance of the information.
[0133] In one embodiment, the terminal also displays an overall summary paragraph generated by the server, allowing users to quickly understand the overall assessment of their current health status.
[0134] In this invention, users browse the terminal interface and, based on behavioral guidance information, choose whether to immediately go to a medical service provider, whether to re-enter more detailed symptoms, and whether to save the record. This human-computer interaction process relies on the server's efficient computing and data processing capabilities, but does not change the technical characteristics of the server side.
[0135] III. Technical Effects and Improvements in Computer Technology The server achieves the following technical effects through a combination of processes: generating structured information, dynamically adjusting prompts, collaborating with generative artificial intelligence models, and calculating severity rules. By performing structured preprocessing on natural language input, the server reduces the complexity of parsing the raw text each time the generative AI model is invoked, thereby reducing the uncertainty of the model input and improving the reliability of the output.
[0136] By dynamically updating and repeating prompts, the server enables the generative AI model to converge to a more structured output that better meets system requirements across multiple iterations, thereby reducing post-processing complexity and improving the overall computational efficiency of the system.
[0137] By combining the model output with the rule data in the information storage device, the server can correct and refine the model output, avoiding complete reliance on the "black box" judgment of the model. This unconventional method of mixing human rules with neural network results helps to improve diagnostic accuracy and reduce errors.
[0138] By associating diagnostic results and behavioral guidance information with information about healthcare providers and healthcare-related items, the server enables the computer's internal processing results to directly support real-world healthcare resource allocation decisions, thereby achieving overall optimization of communication and data flow (such as reducing unnecessary remote repetitive consultation requests and lowering network load).
[0139] Through the above-described embodiments, this invention does not simply automate the thought process of human doctors, but rather achieves improved processing speed, enhanced result accuracy, and optimized data management and communication load through coordinated improvements to multiple internal computer technologies, such as data structure, model calling strategy, prompt statement design, and database joint computation. Thus, it provides a novel implementation method for health status analysis and decision support based on generative artificial intelligence models in the field of computer technology.
[0140] use Figure 11 The processing procedure is explained.
[0141] Step 1: Users input their health status information through the terminal.
[0142] In the graphical interface of the terminal, users can enter natural language text (such as "headache, cough") in the symptom input box, enter a value (such as "37.8") in the body temperature input box, and enter or select disease information (such as "hypertension") in the past disease input box.
[0143] Input: Symptom text, body temperature text, and illness text entered by the user on the terminal interface.
[0144] Output: A set of raw user input data objects stored in the terminal memory.
[0145] Step 2: The terminal performs format validation and preprocessing on the raw input.
[0146] The terminal checks if the symptom text is empty; if so, it prompts the user to add it on the interface. The terminal uses a local script to determine if the body temperature text is in a parsable numerical format; if an illegal character is detected, it prompts the user to re-enter it. The terminal removes leading and trailing spaces from each field, merges duplicate delimiters, and encapsulates each field into a key-value pair data structure.
[0147] Input: Raw user input data object (symptoms, body temperature, illness in text form).
[0148] Output: A standardized input data object that has been validated and preprocessed (still in the terminal memory).
[0149] Step 3: The terminal sends standardized input data to the server.
[0150] The terminal constructs an HTTPS request containing standardized input data, sets the Content-Type to application / json, encodes fields such as symptoms, body temperature, and illness status into JSON format, and targets the server's publicly exposed health analysis interface address. The terminal sends the request through the network module and registers a callback locally to wait for the server's response.
[0151] Input: Standardized input data object.
[0152] Output: The network request message sent to the server, and the terminal being in a state of waiting for the server's response.
[0153] Step 4: The server receives and parses the request data sent by the terminal.
[0154] The server listens on a predefined port through a web application framework. Upon receiving an HTTPS POST request from the terminal, it reads a JSON string from the request body. The server then calls a JSON parsing library to convert the JSON string into an internal data structure (such as a dictionary or object) and extracts the symptom text, temperature value text, and illness text fields from it.
[0155] Input: A JSON-formatted network request message from the terminal.
[0156] Output: An internal server request object containing symptom text, body temperature text, and illness text.
[0157] Step 5: The server converts health status-related information into structured information.
[0158] The server performs word segmentation and keyword extraction on the symptom text, comparing words such as "headache" and "cough" with the internal symptom dictionary and mapping them to standard symptom tags; the server converts the body temperature text into floating-point number form and adds unit markers (such as degrees Celsius); the server matches the disease text with the previous disease coding table to obtain the corresponding standard disease code; the server combines these results into a structured information object.
[0159] Input: Original symptom text, body temperature text, and illness text from the server's internal request object.
[0160] Output: A structured information object containing fields such as standard symptom labels, body temperature values and units, and disease codes.
[0161] Step 6: The server generates prompts for generative artificial intelligence models.
[0162] The server reads the raw text information and structured information, embeds them into a predefined prompt template, and generates a complete natural language description through string concatenation or template rendering. The server explicitly requires the model to output a list of disease names and brief descriptions in the prompt statement, and specifies the output order or format.
[0163] Input: Raw health status related information (text) and structured information objects.
[0164] Output: A text message containing symptoms, body temperature, description of illness, and task requirements.
[0165] Step 7: The server dynamically adjusts the prompts based on historical performance and current input.
[0166] The server queries the information storage device for past model output quality records related to the current user session (such as whether the disease list is missing or whether the format is not standardized), and counts the number of symptoms currently input, the degree of abnormal body temperature, and the complexity of the disease. Based on these indicators, the server adds or modifies the constraints of the prompt statement, such as adding "Please output as a numbered list" or "Please keep it within 200 characters". The server generates the revised final prompt statement.
[0167] Input: Initial prompt text, historical call quality metrics, and current input complexity metrics.
[0168] Output: The final prompt text, dynamically adjusted and formatted.
[0169] Step 8: The server invokes a generative artificial intelligence model to obtain disease candidate information.
[0170] The server constructs a request using an HTTP client library, places the final prompt statement in the request body, sets the model name, temperature parameters, maximum generation length, etc., and sends the request to the generative artificial intelligence model server. After receiving the response, the server extracts the text content generated by the model from the JSON response. This text contains multiple possible diseases and their corresponding descriptions.
[0171] Input: Final prompt text and model call parameters.
[0172] Output: Model output text containing natural language descriptions of disease candidates.
[0173] Step 9: The server parses the model output and extracts disease candidate information and disease classification information.
[0174] The server performs line segmentation or number segmentation on the model's output text, identifying disease names and descriptions line by line. The server removes serial numbers and irrelevant characters through string cleaning, and matches the disease names with the disease coding table to obtain standard disease codes and disease categories. The server encapsulates the name, description, code, and category of each disease into a list of disease candidate information and disease classification information.
[0175] Input: Natural language text output by the model.
[0176] Output: A list containing multiple disease candidate entries, each record including disease classification information such as disease name, description, standard code, and disease category.
[0177] Step 10: The server calculates disease severity information by referring to the information storage device.
[0178] For each disease candidate, the server retrieves the basic severity data and related rules for the corresponding disease from the information storage device; the server combines the basic severity with structured body temperature values, disease codes, and symptom tags, and executes predefined rules or weighted calculation algorithms to generate the final severity level (such as mild, moderate, or severe) and numerical risk score; the server generates a corresponding disease severity information object for each disease candidate.
[0179] Input: List of disease candidate information and disease classification information, structured information objects, and severity rule data in the information storage device.
[0180] Output: A list of disease severity information generated for each disease candidate.
[0181] Step 11: The server generates diagnostic results based on the severity of the illness and the diagnostic results.
[0182] The server sorts all disease candidates based on severity level and model output order, marking the most likely and severe disease as the primary candidate and other diseases as secondary candidates. The server constructs a diagnostic result data structure, which includes the main disease name, several candidate diseases, corresponding severity level and brief description, and converts this information into user-friendly natural language description text.
[0183] Input: List of candidate disease information and list of disease severity information.
[0184] Output: Structured diagnostic results data objects and natural language text of the diagnostic results.
[0185] Step 12: The server generates behavioral guidance information and associates it with medical-related resources.
[0186] Based on the severity level of the primary disease, the server retrieves appropriate levels of medical service providers (such as general outpatient clinics, specialist outpatient clinics, and emergency rooms) and commonly used medical-related items (such as commonly used drugs and assistive devices) from the medical-related resource information storage device. The server generates behavioral guidance text based on the risk level and resource accessibility, such as "It is recommended to seek medical treatment within 24 hours" or "If the following symptoms occur, please seek emergency treatment immediately," and embeds the recommended medical service provider types or resource lists into the text.
[0187] Input: Diagnostic result data objects, disease severity information, and resource data in medical-related resource information storage devices.
[0188] Output: Behavior guidance information data object and corresponding natural language behavior guidance text.
[0189] Step 13: The server returns diagnostic results and behavioral guidance information to the terminal.
[0190] The server encapsulates the diagnostic result data object and the behavior guidance data object into a JSON response, and includes a natural language description field in the response; the server sends the response to the terminal address that initiated the request via HTTPS; the server records the core content of this response and the processing time in the log for subsequent performance analysis and fault diagnosis.
[0191] Input: Diagnostic result data object and behavioral guidance data object.
[0192] Output: A network response message containing diagnostic results and behavioral guidance sent to the terminal.
[0193] Step 14: The terminal parses the server response and displays the results.
[0194] After receiving the server's response, the terminal reads the response content from the network buffer and uses a JSON parsing function to generate a local data object. The terminal extracts the disease name, severity level, diagnosis description, and behavioral guidance text, and presents them on the interface in the form of lists and paragraphs. The terminal can use highlighting or color marking to mark key diseases according to the severity level to increase users' attention to high-risk information.
[0195] Input: The JSON response message returned by the server.
[0196] Output: Diagnostic results and behavioral guidance information displayed on the terminal interface, as well as subsequent operation options available to the user (such as re-entering, saving records, etc.).
[0197] Step 15: Users make subsequent decisions based on the information displayed on the terminal.
[0198] Users can read the disease candidates, severity levels, and behavioral guidelines on the terminal. Based on prompts such as "mild," "moderate," and "severe," users can decide whether to seek medical attention immediately, observe changes in symptoms, or re-enter more detailed symptom information. If users need further consultation, they can select the entry point on the terminal to re-enter or contact the medical service provider.
[0199] Input: Diagnostic results and behavioral guidance information displayed on the terminal.
[0200] Output: Medical decisions or subsequent interactive instructions made by the user based on the system output. These instructions are then input into the system processing flow again in the next round of interaction.
[0201] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0202] In the healthcare field, with the rapid increase in information processing volume such as remote consultations and online medical payments, existing computer processing workflows based on fixed rules or simple classification models typically only perform rough matching of user symptoms or single disease determination. It is difficult to complete the following series of processes uniformly and efficiently within the same computer system: performing structured preprocessing of multi-source health data, automatically generating prompts suitable for generative artificial intelligence models, making refined disease and severity determinations based on the generated results, and further closely integrating the determination results with the automatic selection strategy of medical expense payment schemes.
[0203] In existing technologies, servers mostly use generative AI models only as "question-answering interfaces" or "auxiliary consultation tools." Their inputs are mostly fixed natural language questions written manually or directly by the front end. The server side lacks a systematic mechanism for constructing prompts for generative AI models, leading to the following technical problems: (1) The server cannot automatically generate prompts optimized for specific consultation scenarios based on the structured features of user symptom information, physiological measurement information and past information. The reasoning effect of generative artificial intelligence models depends on human experience, and the stability and repeatability are poor. (2) The server lacks a technical solution to automatically parse the disease name and severity from the natural language diagnostic results output by the generative artificial intelligence model and link it with the cost information storage department. As a result, the subsequent medical expense payment plan still needs to be manually interpreted and configured, the calculation process is fragmented, and the overall processing efficiency is low. (3) When designing medical expense payment schemes, servers often use business logic that is independent of the diagnosis process. They cannot dynamically associate the severity of the disease with payment conditions (one-time payment or installment payment, etc.) in a single automated processing link, making it difficult to provide differentiated and personalized payment schemes for users with different health risks.
[0204] Therefore, an improved computer implementation scheme is needed. On the server side, through unified preprocessing and feature extraction of health status data, prompts adapted to generative artificial intelligence models are automatically generated. The diagnostic results output by this model drive the automatic selection and generation of payment schemes. This improves the computer's ability to integrate online consultation and cost decision-making at the algorithm and system structure levels, enhances the stability of diagnostic reasoning, the automation of payment scheme generation, and the overall processing efficiency of the system.
[0205] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.
[0206] In this invention, the server includes a device for receiving health-related information from a user terminal and generating structured data; a device for preprocessing the structured data and extracting features including symptom information, physiological measurement information, and past information, and generating prompt statements based on the features to be input into a generative artificial intelligence model; a device for inputting the prompt statements into the generative artificial intelligence model to obtain diagnostic result information including possible disease names and the severity corresponding to the disease names; a device for extracting the disease names and the severity from the diagnostic result information and determining a medical expense payment scheme suitable for the user by referring to multiple medical expense payment conditions pre-defined in a expense information storage unit corresponding to different severity levels; and a device for sending the diagnostic result information and the medical expense payment scheme to the user terminal. This enables an integrated computer processing flow on the server side, from structured processing of health data, automatic construction of prompt statements, generative artificial intelligence model diagnosis, to automatic generation of payment schemes based on severity, improving the controllability and stability of the input and output of the generative artificial intelligence model, reducing human intervention, and enhancing the automation level of online medical consultation and cost decision-making and the overall system processing performance.
[0207] "User terminal" refers to an information processing device operated by a user to input health status-related information and receive diagnostic results and medical expense payment plans returned by the server, including but not limited to smartphones, tablets, desktop computers, or other electronic devices that can communicate with the server via a network.
[0208] "Health status related information" refers to data that reflects a user's current or historical health status, including but not limited to symptom descriptions, physiological measurement information, past medical history, medication use, and other personal information related to health assessment.
[0209] "Structured data" refers to a data set generated by a server based on health status-related information sent by a user terminal, organized according to a predetermined data format and fields. This allows various types of information to be represented by fixed fields or labels, facilitating subsequent retrieval, calculation, or modeling processing by subsequent programs.
[0210] "Preprocessing" refers to the data processing operations performed by the server on the structured data before it is used for subsequent model inference, such as format regularization, missing value handling, standardization, word segmentation, encoding, or feature extraction, in order to improve the effectiveness and stability of subsequent data analysis or model calculation.
[0211] "Symptom information" refers to health status description information that reflects the user's subjective feelings or external manifestations, including but not limited to symptom-related data such as headache, fever, cough, chest tightness, and fatigue, which are entered through natural language or preset options.
[0212] "Physiological measurement information" refers to numerical or graded data that can quantify the body's state, obtained through measurement equipment or user input, including but not limited to body temperature, heart rate, blood pressure, blood oxygen saturation, weight, and related physiological indicators.
[0213] "Past information" refers to historical data related to a user's past health and medical treatment, including but not limited to past medical history, chronic disease history, allergy history, surgical history, long-term medication history, and other past medical records related to the current diagnosis.
[0214] "Features" refer to numerical or symbolic data that the server extracts or calculates from health status-related information during preprocessing to represent a user's health status. These include, but are not limited to, encoded symptom vectors, standardized values of physiological measurements, disease history codes, and other feature data that can be used as input to the model.
[0215] "Generative artificial intelligence models" refer to computational models built on machine learning or deep learning techniques that can automatically generate output content such as text based on input prompts, including but not limited to language models or multimodal generative models used for question answering, dialogue, text generation, and reasoning analysis.
[0216] "Prompt statements" refer to natural language text constructed by the server based on structured data and features and input into the generative artificial intelligence model. They are used to explicitly provide the model with diagnostic tasks, constraints, and output requirements, thereby guiding the model to generate diagnostic results information related to the disease name and severity.
[0217] "Diagnostic result information" refers to the result data related to the user's health status assessment obtained by the generative artificial intelligence model based on the prompt statements and after being parsed by the server. This includes, but is not limited to, possible disease names, the severity of each disease, risk warnings, and related suggestions.
[0218] "Disease name" refers to the medical or generic name used to identify a specific disease or health condition, indicating the specific disease category or symptom type that a user may have.
[0219] "Severity" refers to the evaluation result used to indicate the degree of impact of a disease on a user's health or the risk level of disease development, including but not limited to mild, moderate, severe or other predetermined level classifications.
[0220] The “cost information storage unit” refers to a data storage unit used to store information related to medical costs. It can be a database, a memory, or other data storage structure, used to store various pre-set medical cost payment conditions and their parameters based on the type and severity of the disease.
[0221] "Medical expense payment conditions" refer to the rules for settling medical expenses applicable under different diagnostic results and severity, including but not limited to whether installment payments are allowed, the number of installments, the amount per installment, preferential conditions, payment period, and other parameters related to the payment method.
[0222] "Medical expense payment plan" refers to a specific payment arrangement plan generated by the server for a specific user based on the medical expense payment conditions stored in the diagnostic result information and expense information storage department. This includes, but is not limited to, a one-time payment plan, an installment payment plan, and a combination plan that includes various fees and payment rules.
[0223] "One-time payment conditions" refer to payment rules where medical expenses are settled in full at a specified time in a single payment manner, including but not limited to parameter settings such as total amount, payment deadline, and payment channel.
[0224] "Installment payment terms" refer to payment rules that allow medical expenses to be paid in multiple installments over a period of time, including but not limited to the number of installments, the amount payable in each installment, the down payment ratio, the repayment cycle, and the setting of related fees.
[0225] In one embodiment of the invention, the system includes a server, a terminal, and a communication network connecting the two. The server is configured to operate in a data center or cloud computing environment, and the terminal is configured to be an information processing device carried or operated by a user, such as a smartphone, tablet, or personal computer.
[0226] In terms of hardware, a server may include a general-purpose processor, memory, a network interface, and an optional graphics processing unit. In terms of software, a server may run an operating system, such as a Unix-like operating system; an application service framework, such as a web service framework implemented using a scripting language; a machine learning framework, such as a deep learning framework for building and inferring neural network models; and a database management system for storing cost information and diagnostic records.
[0227] The terminal may include a processor, touchscreen display, wireless communication module, memory, and optional sensors (such as body temperature sensor, heart rate sensor, etc.) in terms of hardware. In terms of software, the terminal may run a mobile operating system and communicate with the server through a browser or dedicated application. Users provide health status-related information to the server through the terminal's touch interface or voice input interface, including symptom text, physiological measurements, and historical information.
[0228] After receiving health status information from the terminal, the server organizes this information into a data structure with fixed fields in its memory. For example, the server can set character fields for symptom text, floating-point fields for physiological measurements such as body temperature and blood pressure, and enumeration or encoding fields for past diseases. The server further establishes structured data objects internally, uniformly mapping different types of health information to a set of structured data fields for unified processing by subsequent algorithm modules.
[0229] The server performs preprocessing operations on the structured data. Specifically, it uses a Chinese word segmentation tool to segment the symptom text, breaking the sentence down into a sequence of words; it uses a stop word list to remove frequently used words that are not significant for diagnosis; and it uses a word vector model or an encoder-based language model to vectorize the retained words, converting the text into multi-dimensional numerical vectors. For physiological measurements such as body temperature, heart rate, and blood pressure, the server performs type conversion and range checks, and uses normalization or standardization formulas to map values of different dimensions to a unified numerical range to improve the numerical stability of the neural network computation process. For past disease information, the server assigns a unique code to each disease, converting it into a sparse vector or one-hot encoded form as part of the modeling features.
[0230] After completing the preprocessing and feature extraction described above, the server combines these numerical and symbolic features into a feature set. Based on this feature set, the server constructs prompts for the generative artificial intelligence model. The server pre-stores multiple prompt templates in its memory; each template includes a prompt task description, an input data summary, and output format requirements. The server selects a template that matches the medical diagnostic task and replaces the placeholders in the template with a natural language summary of the current user's symptoms, physiological measurements, and past information, thereby generating the complete prompt text.
[0231] For example, the server can generate the following prompt: "The user's symptoms are: headache, slight fever, and cough. The body temperature is 37.6 degrees Celsius, and the past medical history is hypertension. Based on this information, please analyze the possible disease names (you can list 1-3), and give the severity (mild / moderate / severe) and a brief description for each disease. Output the structured results in Simplified Chinese." Alternatively, the server can generate the following prompt: "The user's symptoms are: chest tightness and mild shortness of breath. The body temperature is 36.8 degrees Celsius, and the past medical history is hypertension. Please guess the possible disease name, assess the severity of the disease (mild, moderate, or severe), give a brief reason, and output it in the form of an item." The server inputs the generated prompts into the generative AI model. In one embodiment, the generative AI model can employ a multi-layer neural network structure based on a self-attention mechanism, such as a language model containing multiple encoders and decoders. During the model inference phase, the server does not perform weight updates, only forward propagation. During inference, the server first segments and tokenizes the prompts, then converts them into discrete token sequences, and finally converts them into vector sequences through an embedding layer. The server then sequentially performs matrix multiplication, non-linear activation, and normalization operations on the vector sequences in multiple attention sub-layers and feedforward sub-layers to obtain a contextual representation. In the output phase, the server samples word-by-word or character-by-character based on probability distributions to generate the diagnostic result text.
[0232] During model training, the server can utilize pre-collected and labeled datasets of consultation texts, electronic medical record summaries, and diagnostic labels to optimize model parameters through supervised learning. During training, the server employs loss functions, such as cross-entropy loss, to measure the difference between the model-generated sequence and the target sequence. The server uses stochastic gradient descent or its variants to update model weights. Data augmentation techniques can be employed, such as synonym replacement, random masking, or order perturbation of symptom text, to improve the model's generalization ability across different representations. Through this training process, the model's internal parameters form a statistical representation of the relationship between symptom combinations and disease categories and severity, enabling the model to automatically output relatively stable and medically accurate diagnostic text based on prompts during deployment.
[0233] After receiving the diagnostic result text output by the generative artificial intelligence model, the server parses the text in its natural language parsing module. The server can use regular expressions, keyword matching, sequence labeling models, and other methods to locate disease name phrases and severity phrases in the diagnostic text. The server maps the extracted disease names to a standard disease coding library to eliminate ambiguity caused by synonyms; the server converts severity text (e.g., "mild," "moderate," "severe") into predefined level identifiers, forming a structured diagnostic result information object.
[0234] The server pre-stores medical expense payment conditions corresponding to different severity levels in the expense information storage department, including one-time payment options and multiple installment payment options. The expense information storage department can use a relational database or a key-value store system. The server configures different parameter combinations for different severity levels, such as the total cost estimation range, the maximum number of allowed installments, the minimum payment amount per installment, and preferential conditions. When determining a medical expense payment plan, the server queries the expense information storage department based on the severity field in the diagnosis result information object to obtain a set of payment conditions matching that severity level, and generates a specific payment plan according to preset rules (e.g., prioritizing plans with lower total costs or plans based on the user's historical preferences).
[0235] In one example, when the severity is mild, the server selects a one-time payment option from the cost information storage unit, determines the total cost and the amount of the single payment, and generates a first payment plan. When the severity is moderate or severe, the server selects an installment payment option from the cost information storage unit, determines the number of installments, the amount per installment, and the installment period, and generates a second payment plan. The server can also generate multiple candidate plans simultaneously and mark the recommendation level in the results. The terminal displays the candidate plans and the recommended plans to the user on the display interface.
[0236] The server encapsulates structured diagnostic results and medical expense payment plans into output information and sends it to the terminal. Upon receiving this information, the terminal displays the disease name, severity, brief recommendations, and payment plan details on its screen using a user interface. Users can browse different payment plans on the terminal interface and select and confirm a plan via touch. After the user makes a selection, the terminal sends the result back to the server, which updates the corresponding record in its database and can further interact with external payment processing systems to generate payment links or payment instructions, thereby triggering the actual fund settlement process in the real world.
[0237] Through the aforementioned structure and processing methods, the system of this invention not only achieves automated diagnosis and cost decision-making for health status data, but also brings about multiple improvements at the computer technology level. The server, through a unified structured data model and feature extraction process, transforms the originally scattered text and numerical data into a unified set of features, thereby making the input of the generative artificial intelligence model more standardized, reducing output fluctuations caused by inconsistent prompt statements, and improving the stability of diagnostic results. By centrally constructing prompt statements on the server side, rather than allowing free input from the terminal or manually, the content and format of the prompt statements are more consistent with the distribution during model training, reducing distribution bias in the inference stage and significantly improving the accuracy of the model in generating disease names and severity levels.
[0238] Furthermore, the server establishes a direct data link between the output of the generative AI model and the cost information storage unit, enabling the diagnostic results to be automatically parsed and utilized by the machine, forming an automatic mapping link from "model output text" to "payment scheme parameters." This link, through rule-based parsing and encoding, reduces the number of times humans need to participate in result interpretation and cost configuration, and reduces intermediate human intervention nodes from the perspective of data flow within the computer system, thereby reducing communication rounds and processing latency, and improving overall processing speed.
[0239] During the model training phase, the server employs an error feedback and weight update mechanism to continuously adjust internal parameters, enabling the system to maintain high diagnostic accuracy even when faced with diverse symptom combinations and complex historical information. The server expands the training sample space through data augmentation and other techniques, improving the model's robustness to rare symptom combinations and variable expressions, and reducing diagnostic errors caused by differences in language expression. Compared to expert systems based on fixed rules, this invention establishes finer-grained, higher-dimensional feature associations through the parameter learning mechanism of neural networks, enabling the capture of complex nonlinear relationships and achieving higher diagnostic accuracy and more flexible payment scheme inference capabilities.
[0240] This invention can also have several alternative implementations. The server can employ different types of generative artificial intelligence models, such as language models based on recurrent neural networks, or models based on a hybrid convolutional and attention structure. As long as the model can generate diagnostic text containing the disease name and severity based on the prompt, it can be incorporated into the technical solution of this invention. The server can also employ a multi-level rule system in the cost information storage unit, dynamically adjusting payment conditions based on additional features such as disease type, past disease combinations, and age, rather than being limited to a single severity dimension. In one implementation, the server can use the output of the generative artificial intelligence model as the upper-level inference result, and then combine it with a lightweight discriminative model or rule engine to verify severity or filter unreasonable results, thereby further improving the reliability of the overall system output.
[0241] By tightly integrating structured feature extraction, automatic construction of prompts, generative artificial intelligence model inference, and automatic generation of cost plans into a data processing architecture within the same server, this invention achieves a novel computer implementation method, making online consultation and medical cost decision-making a single, optimizable algorithmic process. This process is executed internally with specific data structures and algorithmic sequences, and its behavior is progressively optimized through training and parameter tuning. This results in quantifiable technical improvements in processing speed, diagnostic accuracy, data management consistency, and communication load, rather than simply automating the traditional manual decision-making process.
[0242] use Figure 12 The processing procedure is explained.
[0243] Step 1: Users enter health status information on the terminal and submit the information to the server.
[0244] Input: Information related to the user's health status (including symptom text, physiological measurements, past information, etc.).
[0245] Output: A health status-related information data packet generated by the terminal and sent to the server.
[0246] The terminal displays symptom input boxes, body temperature input boxes, and a list of past medical conditions on the interface. Users input symptoms such as "headache, slight fever, cough" via keyboard or touch, enter a body temperature of "37.6", and select "hypertension" as a past medical condition. The terminal encapsulates these inputs into request data with fixed fields and sends it to the server via a network communication protocol.
[0247] Step 2: The server receives health status information sent by the terminal and generates structured data.
[0248] Input: A data packet containing health status-related information sent by the terminal.
[0249] Output: A structured health data object stored in the server's memory or database.
[0250] The server receives data packets from the network interface, parses the packet format, and reads the symptom string, body temperature field, and past medical history field. The server checks for missing fields; if any are missing, it sets a default value or returns an error message. If all fields are present, the server creates a data structure in memory containing multiple key-value pairs, mapping the symptom text to a text field, converting the body temperature value to a floating-point number and storing it in a numeric field, mapping past medical history items to a predefined encoding, and storing all this information as a structured data object.
[0251] Step 3: The server preprocesses the structured data and extracts features.
[0252] Input: Structured health data object.
[0253] Output: Includes symptom vectors, standardized values of physiological measurements, and a set of features encoded by previous diseases.
[0254] The server first performs Chinese word segmentation on the symptom text field, breaking long sentences down into a list of words. It then uses a stop word list to remove words without diagnostic significance, inputting the remaining words into a word vector model or encoder. Each word is converted into a numerical vector, and these vectors are combined into a symptom vector representation through averaging or concatenation. The server performs range validation and type conversion on physiological measurements such as body temperature, then uses a normalization formula to calculate standardized values; for example, converting 37.6 degrees Celsius into a numerical representation between 0 and 1. The server also maps past medical conditions such as "hypertension" to a set of predefined disease codes and generates corresponding unique thermal encoding vectors. Finally, the server combines the symptom vectors, standardized physiological measurement values, and past medical condition codes into a unified set of features, providing input for subsequent generation of prompts and model inference.
[0255] Step 4: The server generates prompts based on feature values and constructs natural language text.
[0256] Input: A set of features (including symptom vectors, standardized values of physiological measurements, and codes of past diseases) and raw text information.
[0257] Output: Natural language text prompts for generative artificial intelligence models.
[0258] The server generates a readable summary based on the original content contained in the feature set. For example, it reconstructs the symptom list into "headache, slight fever, cough", restores the standardized body temperature value to "body temperature is 37.6 degrees Celsius", and converts the past disease code back to the name "hypertension". The server reads a preset prompt statement template from the memory, such as "The user's symptoms are: {symptoms}. Body temperature is {body temperature} degrees Celsius, and past medical history is {medical history}. Please analyze the possible disease names (list 1-3) based on this information, and give the severity (mild / moderate / severe) and brief description of each disease. Output the structured results in simplified Chinese." The server replaces the placeholders in the template with the current summary content to form a complete prompt statement text, and temporarily stores the prompt statement in memory, ready to be used as input for the generative artificial intelligence model.
[0259] Step 5: The server will input the prompt statement into the generative artificial intelligence model and obtain the diagnostic result text.
[0260] Input: The server-generated prompt text in natural language.
[0261] Output: Natural language text of the diagnostic results output by the generative artificial intelligence model.
[0262] The server calls the generative AI model inference interface, sending the prompt as input to the model inference module. Internally, the model performs word segmentation and tokenization on the prompt, converting the text into a tokenized sequence and mapping it to a vector sequence through an embedding layer. Subsequently, the model performs matrix multiplication, attention weight calculation, and non-linear activation in a multi-layer neural network structure to generate a contextual representation. The decoding part generates diagnostic text containing "possible disease names" and "severity" word by word based on probability distribution. The server waits for the inference process to complete and receives the complete diagnostic result text from the model interface, such as "Based on the information provided by the user, the most likely disease is a mild upper respiratory tract infection (mild cold), with a mild severity. It is recommended to rest more, drink more water, and monitor blood pressure changes." Step 6: The server parses structured diagnostic result information from the diagnostic result text.
[0263] Input: Natural language text of the diagnostic results output by the generative artificial intelligence model.
[0264] Output: A structured diagnostic results object containing the disease name, severity, and additional suggestions.
[0265] The server uses a natural language processing module to perform pattern matching and field extraction on the diagnostic result text. The server locates the disease name phrase "mild upper respiratory tract infection (mild cold)" through keyword search and regular expressions, and maps it to an internal disease coding table. The server searches for severity keywords such as "mild," "moderate," and "severe," extracting "mild" as the severity field value. The server can also store the remaining suggestions, such as "rest more, drink more water, and monitor blood pressure changes," as explanatory text in the suggestions field. The server integrates the disease code, severity level, and suggestions into a structured diagnostic result object for use by the subsequent cost calculation module.
[0266] Step 7: The server determines the medical expense payment plan based on the structured diagnostic results and the cost information storage unit.
[0267] Input: Structured diagnostic results objects (including disease codes and severity) and payment terms data in the cost information storage department.
[0268] Output: A medical expense payment scheme object containing payment method, amount, and installment parameters.
[0269] The server retrieves the cost information storage unit based on the severity field in the diagnosis result object, reading a set of payment conditions corresponding to "mild," such as a one-time payment condition and candidate installment payment conditions. The server selects the payment plan type based on internal strategies (e.g., prioritizing one-time payment for lower severity). In the case of mild severity, the plan type is set to "First Payment Plan," i.e., one-time payment, and the total cost is calculated according to the cost rules. The server encapsulates the calculated total cost value, payment method identifier, and explanatory text into a payment plan object. If the severity is moderate or severe, the server looks up installment payment conditions, determines the number of installments, the amount per installment, and the payment cycle, generating a "Second Payment Plan" object.
[0270] Step 8: The server packages the diagnosis results and the medical expense payment plan and sends them to the terminal.
[0271] Input: Structured diagnostic results objects and medical expense payment scheme objects.
[0272] Output: Response data sent to the terminal, including diagnostic information and payment scheme information.
[0273] The server combines the disease name, severity, and recommended description fields with the payment plan's payment method, cost, and installment period fields into a unified response data structure. The server sends this response data to the terminal via a network interface. Upon receiving the response, the terminal parses the content and displays information such as "Preliminary diagnosis: Mild upper respiratory tract infection (mild cold), severity: mild. Recommended payment plan: One-time payment of 800 yuan." on the display interface.
[0274] Step 9: Users can view diagnostic results and payment options on the terminal and select or confirm a payment option.
[0275] Input: Diagnostic results and payment scheme information displayed on the terminal interface.
[0276] Output: User's instructions for selecting a payment method.
[0277] After reading the disease name, severity, and cost plan details returned by the server via the terminal, the user clicks the "Accept One-Time Payment Plan" button on the terminal interface, or selects one of the multiple candidate plans. The user's input is converted into a payment plan confirmation instruction internally by the terminal, which then encapsulates this instruction along with the identifier of the selected plan into a data packet, ready to send it back to the server.
[0278] Step 10: The terminal sends the user's payment plan selection to the server and completes the plan confirmation process.
[0279] Input: The user's payment plan selection instruction.
[0280] Output: Payment scheme confirmation data sent to the server.
[0281] After the user completes their selection, the terminal encapsulates the selected scheme type, identifier, and user identifier into confirmation request data and sends it to the server via a network protocol. Upon receiving this confirmation data, the server updates the payment scheme status of the corresponding user record in its database to "confirmed" and can trigger subsequent payment instruction generation or interaction with external payment systems as needed.
[0282] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.
[0283] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."
[0284] In existing technologies, computer-aided medical diagnostic systems typically rely on single-modality data analysis, such as rule matching based solely on structured questionnaire data or classification of lesion images solely based on image recognition models. Firstly, this single-modality approach struggles to integrate patient image information with consultation text information in a timely and accurate manner, resulting in insufficient accuracy and robustness of diagnostic results. Secondly, the logic used in traditional systems to generate diagnostic results and explanations is often based on fixed rules or simple templates, failing to adaptively generate personalized and interpretable textual descriptions according to the specific case context, thus reducing user trust in the system's output. Thirdly, existing systems often implement image analysis, consultation text parsing and severity assessment, medical institution recommendations, and medical supply suggestions in a fragmented manner, lacking a unified computing architecture and data flow mechanism. This leads to poor system scalability, high maintenance costs, and difficulty in uniformly optimizing and tuning the performance of each processing step.
[0285] Furthermore, in the process of applying generative artificial intelligence models to medical auxiliary diagnosis, traditional solutions often directly input user descriptions or simple keywords, lacking the organic integration and context construction of image features, structured consultation data, and candidate disease information. This results in unstable generation results, inconsistent context, and may even output content that does not match the actual symptoms, making it difficult to meet the requirements of medical scenarios with high safety and consistency requirements.
[0286] Therefore, an improved solution for computer implementation is needed, providing a technical solution from the perspectives of system architecture and data processing: a solution that can automatically complete image preprocessing, deep learning feature extraction, diagnostic text structuring, candidate disease fusion calculation on the server side, and further construct prompt statements adapted to generative artificial intelligence models, so as to achieve controllable and structured management of the input and output process of generative artificial intelligence models, thereby improving the overall computational performance of multimodal medical auxiliary diagnostic systems in terms of accuracy, scalability, and interpretability.
[0287] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.
[0288] In this invention, the server includes means for receiving images of biological parts from a user terminal and converting the images into a predetermined image format upon receipt and acquiring the images through a communication channel; means for performing image processing on the received images to perform pixel value normalization, noise removal, and pixel array transformation to generate input data for image analysis; means for using the input data for image analysis to extract image features using a machine learning model through image recognition and matching them with reference data in a data storage device storing symptom information to infer multiple candidate diseases; and means for acquiring text data of consultation content input by the user and generating structured data containing symptoms, location of occurrence, duration, accompanying symptoms, and medical history from the text data using natural language processing technology. The device comprises: an apparatus for fusing the structured data with the plurality of candidate diseases to calculate the probability of each candidate disease and generate diagnostic context information containing the fusion result; an apparatus for generating prompt statements based on the diagnostic context information and inputting the prompt statements and the diagnostic context information into the generative artificial intelligence model to generate diagnostic results and explanatory text about the candidate diseases; an apparatus for generating prompt statements related to severity assessment for each disease contained in the diagnostic results with reference to a data storage device storing severity information and inputting them into the generative artificial intelligence model to generate explanatory text about severity assessment; and an apparatus for sending diagnostic information obtained based on the diagnostic results and the severity assessment to a user terminal. This allows for the formation of a unified data pathway within the server, encompassing image preprocessing, deep learning feature extraction, diagnostic text structuring, candidate disease fusion calculation, and the construction and reasoning of generative AI model prompts. This enables the generative AI model to generate information in a controlled manner upon receiving structured diagnostic context information and prompts, thereby improving the computational accuracy and stability of multimodal medical auxiliary diagnosis, enhancing the system's ability to handle complex input scenarios, and improving the interpretability and consistency of diagnostic results and explanatory texts. Ultimately, this achieves performance optimization and structural improvement of the computer implementation itself.
[0289] "User terminal" refers to a computing device operated by a user for collecting, inputting and sending data related to health status, including but not limited to smartphones, tablets, personal computers and smart terminals with communication functions.
[0290] "Biological part images" refer to digital image data that represents the state of local areas of the human body or other organisms, including but not limited to static images of parts such as skin, mucous membranes, limbs or organ surfaces.
[0291] "Preset image format" refers to the digital image encoding format that the system pre-sets and is suitable for subsequent image processing and model input, including but not limited to JPEG format, PNG format, and other standard formats that meet the requirements of the image processing interface.
[0292] "Communication channel" refers to a wired or wireless communication path used to transmit data between a user terminal and a server, including but not limited to network connections based on Internet Protocol, mobile communication networks, or local area network connections.
[0293] Image processing refers to the process of performing a series of operations or algorithms on input image data to improve image quality or transform the image representation, including but not limited to scaling, cropping, normalization, noise reduction, and color space conversion.
[0294] "Pixel value normalization" refers to the process of linearly or non-linearly transforming the numerical range of each pixel in an image to make it fall into a predetermined numerical range, so as to adapt to the calculation requirements of subsequent algorithms or models.
[0295] "Noise removal" refers to the process of suppressing or eliminating random interference signals generated during the imaging and transmission processes in an image through filtering or other algorithms.
[0296] "Pixel array transformation" refers to the process of changing the arrangement of image pixels in space or channels, including but not limited to operations such as size scaling, cropping, channel rearrangement, and dimensional expansion.
[0297] "Input data for image analysis" refers to standardized data structures generated after image processing that meet the input requirements of image recognition models, usually represented in tensor or matrix form.
[0298] "Machine learning models for image recognition" refers to computational models trained by learning from a large amount of labeled image data, which can automatically extract features from images and perform classification or matching tasks, including but not limited to convolutional neural network models and other deep learning models.
[0299] "Image feature quantity" refers to a high-dimensional numerical vector or feature representation extracted from an input image by a machine learning model for image recognition, used to characterize information such as lesion morphology, color, and texture in the image.
[0300] "Data storage device" refers to a storage system used to store and manage data related to symptoms, diseases, medical institutions, medical supplies, and severity of illness, including but not limited to database systems, file storage systems, and combinations thereof.
[0301] "Symptom information" refers to various manifestation data related to the disease state, including but not limited to subjective symptom descriptions, objective physical signs and characteristics, and attribute information related to the lesion site.
[0302] "Reference data" refers to a standardized set of information pre-stored in a data storage device for comparison or matching with the data to be analyzed, including but not limited to feature vectors and corresponding disease labels of known case samples.
[0303] "Candidate disease" refers to a disease category that is inferred from the analysis of image features and / or other diagnostic data and has a certain probability of occurring in the current case.
[0304] "Text data of consultation content" refers to string data of medical information such as symptoms, course of disease, and past medical history that users input through natural language.
[0305] "Natural Language Processing" refers to a set of computational methods and algorithms used to process natural language text, such as word segmentation, part-of-speech tagging, entity recognition, syntactic analysis, and semantic understanding.
[0306] "Structured data" refers to data that is organized into a format that is easy to store and retrieve, after the original text or other unstructured data has been parsed and extracted, according to predetermined fields and data types.
[0307] "Diagnostic contextual information" refers to comprehensive contextual data generated by fusing candidate diseases with structured consultation data. This data includes disease probability, key symptom factors, and background information, and is used to drive subsequent diagnostic reasoning and generative model invocation.
[0308] "Generative AI models" refer to AI models that can automatically generate text content based on input prompts and contextual information, including but not limited to large-scale language models based on deep learning.
[0309] "Prompt statements" refer to input text constructed to guide generative artificial intelligence models to perform specific tasks. They typically include task descriptions, contextual information, and output requirements.
[0310] "Diagnosis results" refers to a set of output information about possible diseases and their corresponding explanations, obtained based on image features, structured medical data, and generative artificial intelligence model reasoning.
[0311] "Explanatory text" refers to natural language text generated by generative artificial intelligence models based on prompts, used to explain diagnostic results, severity assessments, or recommendations.
[0312] "Severity information" refers to data used to characterize the severity, urgency, or risk level of various diseases, including but not limited to risk classification and recommended time limit for seeking medical treatment.
[0313] "Severity assessment" refers to the assessment of the severity and urgency of the current disease state by combining diagnostic results with severity information.
[0314] "Medical service institution information" refers to a collection of data related to medical institutions, including but not limited to attribute information such as institution name and type, departments, geographical location, capacity to receive patients, and business hours.
[0315] "Recommended medical service institutions" refers to one or more medical institutions selected from the medical service institution information based on diagnostic results and severity assessment, which are suitable to provide diagnosis and treatment services to the current user.
[0316] "Medical supplies information" refers to attribute data about items related to symptom management or treatment, such as medicines, medical devices, and dressings, including but not limited to indications, dosage, contraindications, and precautions.
[0317] "Medical supply candidates" refers to one or more medical supply options that are selected as suitable for self-management or use in conjunction with telemedicine, given a diagnostic result and severity assessment.
[0318] "Precautions for use" refers to warnings, limitations, and operating recommendations related to the use, safety, and adverse reactions of medical products.
[0319] "Consultation Conditions with Medical Personnel" refers to the description of the triggering conditions that indicate to users in what situations they need to consult a physician or other medical personnel in a timely manner, such as changes in symptoms, medication reactions, or disease progression.
[0320] In the embodiments of this invention, the server, terminal, and user each perform different technical functions. The system of this invention can be deployed in an information processing environment including a server computer, user terminal equipment, and network communication infrastructure. The server can employ computer hardware with a multi-core central processing unit and a graphics processing unit, and the graphics processing unit can employ a graphics processor that supports general-purpose computing; the terminal can be an electronic device with a camera and network communication capabilities, such as a smartphone, tablet computer, or personal computer.
[0321] In one implementation, the server uses an operating system and server framework software to implement network interfaces and task scheduling. The server can run service programs based on a general-purpose operating system and utilize network server software (such as a general architecture based on a combination of reverse proxy and application server) to receive and distribute requests from terminals. Internally, the server uses programming languages to run image processing libraries (such as OpenCV and Pillow) and numerical computing libraries (such as NumPy), and loads machine learning models for image recognition and generative artificial intelligence models based on deep learning frameworks (such as TensorFlow or PyTorch).
[0322] In one embodiment, the terminal uses the camera interface and image encoding library provided by the terminal operating system to convert images of biological parts captured by the user into a predetermined image format, such as JPEG or PNG. The terminal then packages the image file and consultation text into a multi-part form or a request body containing binary data and text fields via a network communication library (e.g., a common HTTP client library), and sends it to the server via the HTTPS protocol. After receiving the diagnostic information returned by the server, the terminal invokes the user interface rendering module to display the structured text and labels on the screen.
[0323] In one implementation, the user triggers the image acquisition function via the terminal to take pictures of the skin, mucous membranes, or other biological sites, and then enters descriptions of symptoms, disease course, accompanying symptoms, and past medical history in the text input area of the terminal application. After confirming that the input is correct, the user triggers a data transmission operation through the terminal interface to initiate a request to the server.
[0324] In one embodiment, the server performs image preprocessing on the image data uploaded by the terminal. The server uses an image processing library to decode the received binary image data into a multi-dimensional pixel matrix and resizes the image to meet the input size requirements of the convolutional neural network model, such as 224×224 or other fixed sizes. The server performs pixel value normalization, linearly mapping the pixel values of each channel from the range of 0–255 to the interval of 0–1 or -1–1. The server further reduces noise and compression artifacts using spatial filtering algorithms (such as median filtering or Gaussian filtering), thereby improving the signal-to-noise ratio of the subsequent model during the feature extraction stage.
[0325] In a certain implementation, the server loads a convolutional neural network model based on a deep learning framework. This model may include multiple convolutional layers, batch normalization layers, non-linear activation layers, pooling layers, and fully connected layers. The server performs forward propagation computation on the image tensor: the convolutional layers perform local linear transformations and weighted summations on the input feature map through sliding convolution kernels; the batch normalization layers normalize the mean and variance of intermediate features; the non-linear activation layers (e.g., ReLU) introduce non-linearity; the pooling layers (e.g., max pooling or average pooling) downsample the spatial dimension; and finally, based on the high-level feature map, it is mapped to the feature vector space or class probability space through fully connected layers.
[0326] In one implementation, the server does not directly use the final classification output. Instead, it extracts feature vectors from a high-dimensional hidden layer as image features, such as real-valued vectors of length 512 or 1024. The server compares this feature vector with a pre-stored set of reference features in the data storage device. For this, the server can use vector similarity calculation algorithms, such as cosine similarity or Euclidean distance, and sort the similarity results of all reference samples to obtain the similarity or confidence scores for multiple candidate diseases. The server records these candidate diseases and their probability values as image parsing results and stores them in association with the current request identifier.
[0327] In another implementation, the server performs natural language processing on the text data of the user's consultation content. The server uses a word segmentation and tagging module to segment and tag the text, and uses an entity recognition model or rule matching module to identify symptom words, body part words, time expressions, negation words, and expressions related to past medical history. The server can use a sequence labeling-based neural network model, such as a bidirectional recurrent neural network combined with a conditional random field or a sequence labeling model based on a Transformer structure, to label each position in the text sequence and output the label category for each word. The server merges adjacent words with the same label category to form entities and generates structured data based on preset field mappings, such as fields like "symptoms," "location," "duration," "accompanying symptoms," "fever," and "similar past medical history."
[0328] In one implementation, the server fuses candidate diseases and their confidence scores obtained from image parsing with structured medical history data. The server can employ a combination of rule-based and statistical model-based approaches. On one hand, the server uses a rule engine to downweight candidates with obvious conflicts; for example, if a candidate disease is typically accompanied by high fever but the structured data clearly does not show fever, the server lowers the overall score for that candidate disease. On the other hand, the server uses models such as multilayer perceptrons or gradient-boosting decision trees to concatenate image features with structured medical history features into a joint feature vector, which is then input into the fusion model to obtain the overall probability of each candidate disease. The server reorders the candidate disease list based on the overall probability and generates a diagnostic contextual information data structure containing disease names, probability values, and key supporting features.
[0329] In one implementation, the server constructs prompts for a generative artificial intelligence model based on diagnostic context information. The server maintains a text template containing a task description, a summary of image analysis results, a summary of consultation information, and output format requirements. The server reads candidate disease names, probabilities, and key symptom elements from the diagnostic context information and inserts them into the template to generate complete prompts. For example, the server can generate the following prompts: "You are a generative artificial intelligence model for dermatological auxiliary diagnosis."
[0330] Here is a user's skin image and consultation information: I. Image analysis results (provided by a convolutional neural network model, for reference only): - Candidate disease 1: Contact dermatitis, probability 0.62 - Candidate disease 2: Urticaria, probability 0.25 - Candidate disease 3: Eczema, probability 0.10 II. Consultation Information (input by the user and structured): - Lesion site: Right forearm - Main symptom: small red rashes - Duration: 3 days - Accompanying symptoms: Mild itching, no fever - Previous history: There has been no similar situation before. Based on the image analysis results and consultation information, please provide the three most likely disease names, along with a brief explanation of the reasoning and approximate probability (high / medium / low) for each disease name, and whether to recommend immediate in-person consultation and the recommended department / clinic. In another implementation, the server can also construct separate prompts for severity assessment, embedding candidate diseases along with severity information (such as acute severe illness markers and potential complication risk levels): "The following are candidate diseases and their approximate probabilities inferred from images and consultation information. Please evaluate the severity of each disease (mild / moderate / severe) according to common clinical severity criteria, and indicate which symptom changes require emergency medical attention." When the server invokes a generative AI model in one implementation, it loads a language model based on a multi-layered self-attention structure using a deep learning framework. This model is pre-trained and fine-tuned on a large-scale general corpus and a selected medical-related corpus. The server sets decoding parameters for the model, such as temperature, sampling strategy, and maximum generation length, to control the stability and controllability of the output. The server inputs the constructed prompts and diagnostic context information into the model interface and receives the diagnostic description and suggestion text generated by the model.
[0331] In one implementation, the server combines the output of a generative artificial intelligence model with internal structured data to form a final diagnostic information record. The server performs security filtering and consistency checks on the output text, such as detecting statements that clearly conflict with the structured data and whether it contains prohibited content. When significant inconsistencies are detected, the server can correct them by regenerating the data and labeling uncertainties, thus creating a safer and more reliable output for the user.
[0332] In one implementation, the server queries a data storage device containing information on medical service institutions. Based on the disease type and severity assessment in the diagnosis results, it uses an algorithm based on geographic location and department matching to select recommended medical service institutions. The server calculates a comprehensive score based on indicators such as the distance between the institution and the current user's geographic location, departmental suitability, and the ability to treat critically ill patients, and selects several institutions as recommendations. The server can construct the following prompt statements for the generative artificial intelligence model to generate natural language recommendation descriptions: "Based on the currently estimated disease type and severity, as well as the patient's location, please generate a text that recommends suitable types of medical institutions and the best time to seek medical attention to the user, and explains the reasons for the recommendation." Similarly, in another implementation, the server accesses medical supply information data, filters suitable self-management medical supply candidates based on diagnostic results and severity assessments, and constructs prompts including usage precautions and consultation conditions for medical personnel. A generative artificial intelligence model then generates natural language instructions. The server packages this information along with basic diagnostic instructions into response data and sends it to the terminal.
[0333] In one implementation, the terminal receives structured response data from the server and displays the main disease names, probabilities, severity assessments, recommended medical service providers, and self-management suggestions in a partitioned layout through a user interface module. The terminal can further support users clicking on a candidate disease to expand and view detailed explanations and precautions provided by the generative artificial intelligence model.
[0334] This invention, through meticulously designed data structures and modular divisions in multiple embodiments, achieves streamlined processing of image feature extraction, text structuring, multimodal fusion, and generative AI model invocation. Internally, the server uses a unified request identifier to manage each diagnostic request, linking the image preprocessing module, feature extraction module, text parsing module, fusion calculation module, prompt generation module, and generative AI model inference module to form a monitorable computational path. The server extracts high-dimensional features at the image level using convolutional neural networks, constructs structured fields at the text level using sequence labeling and entity extraction methods, and introduces multimodal joint features and rule corrections at the fusion level. This ensures that the generative AI model receives diagnostic contextual information that has already undergone algorithmic filtering and structured processing, rather than raw, chaotic input. This specific data flow and module combination reduces the model's sensitivity to irrelevant information, lowers the probability of generating erroneous or inconsistent content, and thus improves the overall accuracy and stability of the diagnostic results.
[0335] This invention improves processing speed by implementing unified feature caching and intermediate result reuse on the server side, thereby reducing redundant computations in some implementations. For example, when generating different styles of explanatory text through multiple iterations, the server reuses existing image features and structured diagnostic data, only reconstructing prompts and calling generative artificial intelligence models, thus avoiding costly convolutional neural network computations again. Regarding network transmission, this invention reduces the transmission load of the original high-resolution image by performing preliminary image compression and format unification on the terminal side, thereby reducing the risk of network congestion and accelerating response speed.
[0336] Through the synergistic effect of the above-mentioned technical features, this invention enables the server to perform more than just simple automation of manual operations in multimodal medical auxiliary diagnosis scenarios. Instead, it introduces specific algorithm structures, data structures, and processing sequences into the internal details of image parsing, text understanding, multimodal fusion, and generative artificial intelligence model invocation, thereby achieving a comprehensive improvement in computational accuracy, processing speed, and system interpretability.
[0337] use Figure 13 The processing procedure is explained.
[0338] Step 1: Users use the terminal to collect and input diagnosis and treatment-related data.
[0339] Users take images of the diseased biological parts using the terminal's camera and enter descriptions of symptoms, duration of illness, accompanying symptoms, and past medical history in the text input area of the terminal interface.
[0340] Input: Real-world images of the lesion site and the user's natural language consultation information.
[0341] Output: The raw image data object (e.g., bitmap object) cached internally by the terminal and the raw consultation text string.
[0342] Step 2: The terminal performs local preprocessing and packaging of the acquired images and text.
[0343] The terminal calls the image encoding interface provided by the operating system to compress and encode the raw image object into a predefined image format (such as JPEG or PNG) file data, while setting the compression quality to control the file size. The terminal combines the user-input consultation text with the session identifier into a string or key-value pair structure. The terminal uses a network library to construct an HTTP request, appending the image as binary data or a multipart form field to the request body, and writing the consultation text as a text field into the request body.
[0344] Input: Original image data object and original consultation text string.
[0345] Output: A network request message containing image binary data and consultation text fields.
[0346] Step 3: The terminal sends the request to the server through a secure network channel.
[0347] The client establishes an encrypted connection with the server using the HTTPS protocol and sends a constructed HTTP POST request to the server's specified interface address. The client appends necessary authentication information, such as a token or user identifier, to the request header so that the server can perform identity verification and access control.
[0348] Input: The completed HTTP request message.
[0349] Output: The request data stream transmitted over the network to the server.
[0350] Step 4: The server receives and parses the requests sent by the terminal.
[0351] The server uses web server software and application frameworks to listen for specific interface paths and receive HTTPS requests from terminals. The server decrypts and parses the requests, extracting the image binary data and consultation text fields from the request body, and generating a unique request identifier for each request. The server validates parameters such as image size, format, and text length; if any parameters are found to be inconsistent with preset conditions, an error response is generated.
[0352] Input: Encrypted data stream received by the network layer.
[0353] Output: The original image binary data, the original consultation text string, and the corresponding request identifier from within the server.
[0354] Step 5: The server performs standardized preprocessing on the received images.
[0355] The server uses an image processing library to decode the image's binary data into a pixel matrix and resizes the image to fit the input size requirements of the convolutional neural network model. The server then applies filtering algorithms to remove noise, suppressing random noise and some compression artifacts. Next, the server normalizes the pixel values, mapping the original 0–255 values to the range required by the model (e.g., 0–1). Finally, the server converts the image matrix into a tensor format for the deep learning framework and adds batch dimensions.
[0356] Input: Raw image binary data.
[0357] Output: A preprocessed image tensor conforming to the input format of the image recognition model.
[0358] Step 6: The server uses an image recognition model to extract image features and infer candidate diseases.
[0359] The server loads a pre-trained convolutional neural network model and performs forward inference computation using the pre-processed image tensor as input. The server extracts high-dimensional feature vectors from the network's high-level feature maps or the penultimate layer, using these as the image features of the current image. The server calculates the similarity between this feature vector and a pre-stored set of reference feature vectors in the data storage device, sorts them according to the similarity scores, selects the disease labels corresponding to the most similar reference samples as candidate diseases, and generates a probability or confidence value for each candidate disease.
[0360] Input: Preprocessed image tensor.
[0361] Output: A data structure containing image parsing results for multiple candidate diseases and their corresponding confidence levels.
[0362] Step 7: The server performs natural language preprocessing and structuring on the consultation text.
[0363] The server receives the raw consultation text string and first performs text cleaning to remove extra spaces, special characters, and meaningless punctuation. The server then calls the natural language processing module to perform word segmentation and part-of-speech tagging, and uses entity recognition algorithms to identify key segments related to symptoms, location, time, negation expressions, and past medical history. Based on preset field mappings, the server fills the identified entities into a structured data structure, such as a JSON object or database record, forming a standardized set of consultation elements.
[0364] Input: The original consultation text string.
[0365] Output: Structured consultation data containing fields such as symptoms, location, duration, accompanying symptoms, and past medical history.
[0366] Step 8: The server integrates image analysis results with structured medical consultation data to calculate the overall disease probability.
[0367] The server combines candidate diseases and their confidence scores obtained from image parsing with various fields from the structured medical history data to form a joint feature vector. The server can use a multilayer perceptron or other classification model to perform forward computation on these joint features to obtain the overall probability of each candidate disease. The server can also downweight candidate diseases that clearly do not match the medical history information according to predefined rules; for example, when the structured data indicates no fever, the server lowers the score of diseases that are usually accompanied by high fever. The server ultimately generates a list of candidate diseases sorted by overall probability and records key symptom features used to support the judgment, forming diagnostic contextual information data.
[0368] Input: Image parsing result data structure and structured medical consultation data.
[0369] Output: A diagnostic contextual information data structure containing candidate diseases, overall probabilities, and key symptom elements.
[0370] Step 9: The server constructs prompts for a generative artificial intelligence model based on diagnostic context information.
[0371] The server reads candidate disease names, probability values, symptom elements, and disease course information from the diagnostic context and populates this information into a pre-designed text template. The server specifies the model role, task objective, input information, and expected output format within the template to construct a semantically complete prompt. For example, the server generates a text segment including "Image analysis results (candidate diseases and probabilities)," "Medical information (location, symptoms, time, accompanying symptoms, past medical history)," and "Please provide the most likely disease and reasons." The server stores the constructed prompt in memory and associates it with a request identifier.
[0372] Input: Diagnostic context information data structure.
[0373] Output: Diagnostic prompt text for generative artificial intelligence models.
[0374] Step 10: The server constructs supplementary prompts for severity assessment and treatment recommendations.
[0375] The server retrieves a list of candidate diseases from the diagnostic context information and queries the data storage device for the severity information and complication risk levels corresponding to these diseases. Based on this information, the server inserts the disease name and a reference severity description into another text template, generating prompts for a guided generative AI model to assess the severity and urgency of medical attention, such as asking the model to output "mild / moderate / severe" and whether immediate medical attention is required. The server can also simultaneously construct prompts for recommending medical service providers or medical supplies, embedding information such as geographical location, department type, available resources, and self-care conditions.
[0376] Input: List of candidate diseases and their corresponding severity information.
[0377] Output: Supplementary prompts for severity assessment and recommendations.
[0378] Step 11: The server submits prompts to the generative artificial intelligence model and receives the generated results.
[0379] The server calls the generative AI model service interface, inputting diagnostic and supplementary prompts into the model in a predetermined order or combination, while setting generation parameters such as temperature and maximum length. The generative AI model performs self-attention computation and decoding on the server or dedicated computing node, outputting natural language text about possible diseases, severity assessments, medical advice, and self-management suggestions. The server receives the generated text segments from the interface, parses and formats them, separating the parts related to disease conclusions, explanations, severity levels, and recommendations.
[0380] Input: Diagnostic prompts and supplementary prompts.
[0381] Output: Diagnostic description text, severity description text, and recommendation text returned by the generative artificial intelligence model.
[0382] Step 12: The server performs technical verification and integration of the generated results.
[0383] The server performs content checks on the generated text, verifying whether the disease names match the candidate disease set in the diagnostic context information, and detecting any descriptions that clearly contradict the structured consultation data. The server marks paragraphs that violate preset rules and, if necessary, modifies the prompts and re-requests generation, or adds uncertainty prompts to the output. The server integrates the final confirmed diagnosis, severity assessment, recommended medical service provider information, and medical supply instructions into a unified response data structure for easy parsing and display by the terminal.
[0384] Input: Raw generated text and diagnostic context information.
[0385] Output: A validated and integrated diagnostic information response data structure.
[0386] Step 13: The server sends diagnostic information to the terminal.
[0387] The server uses the application framework to serialize the diagnostic information data structure into a response message and returns this response message to the requesting terminal via the HTTPS protocol. The server appends a request identifier and timestamp to the response for subsequent tracking and logging. The server also records the processing path and key performance indicators of this request in the logging system.
[0388] Input: Diagnostic information response data structure.
[0389] Output: Diagnostic information response message transmitted to the terminal via the network.
[0390] Step 14: The terminal receives and parses the diagnostic results returned by the server.
[0391] The terminal obtains the server's response message from the network layer, decodes and parses it into JSON, mapping the disease list, probability, severity assessment, recommended medical service institutions, and medical supply descriptions into internal data objects. Based on a predefined interface layout, the terminal fills the key results into the corresponding interface components, such as list items, labels, and text areas.
[0392] Input: The diagnostic information response message returned by the server.
[0393] Output: A set of diagnostic information data objects suitable for interface display.
[0394] Step 15: The terminal displays diagnostic results to the user and provides an interactive interface.
[0395] The terminal displays multiple candidate diseases and their probability level labels on the screen, showcasing textual explanations, severity classifications, and treatment recommendations provided by a generative artificial intelligence model. For each candidate disease, the terminal offers an interactive "expand details" control, allowing users to view more specific explanatory text and precautions. The terminal can also provide jump links to guide users to view the locations of recommended medical service institutions or initiate further consultations.
[0396] Input: A set of diagnostic information data objects.
[0397] Output: A visual representation of the diagnostic results on the terminal display, along with the status of user-operable interactive controls.
[0398] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0399] The following problems exist in existing computer-based online consultation and health management technologies: First, although the server can receive health status information input by the user, it usually only performs simple keyword matching or fixed rule judgment. It cannot dynamically construct prompts that adapt to generative artificial intelligence models according to different scenarios, resulting in unstable and uncontrollable generation results, making it difficult to guarantee the consistency and interpretability of diagnostic assistance results.
[0400] Second, traditional systems often mix multiple steps such as "disease inference", "severity assessment" and "generation of user-readable instructions" into a single call. They lack phased prompts and multi-round interaction mechanisms, which makes it impossible for the server to manage and reuse the intermediate results of each stage in a structured manner. It is also not conducive to fine-tuning the output in complex situations, thus limiting the reliability of generative artificial intelligence models in medical scenarios.
[0401] Third, existing technologies typically do not incorporate multi-source data such as emotional information and location information as structured features within the computer into the prompt statement construction process of generative artificial intelligence models. As a result, when servers perform drug recommendations and medical institution recommendations, they cannot fully reflect the user's urgency, psychological state, and geographical constraints, making it difficult to generate truly personalized recommendation results that meet actual constraints.
[0402] Fourth, the server side lacks the systematic processing capability for the "prompt statement - model output - structured data" link, and lacks a unified arrangement of the calling order and data dependencies between different business sub-tasks (disease diagnosis, severity assessment, drug recommendation, medical institution recommendation). As a result, only the final text result is written to the database, rather than structured diagnostic data that can be directly used by subsequent computer programs, which reduces the scalability and maintainability of the system.
[0403] Therefore, it is necessary to provide a new computer implementation scheme by introducing a phased prompt generation and multi-round call mechanism for generative artificial intelligence models on the server side, and tightly coupling medical information sources, emotional information, location information and diagnostic process, thereby improving the accuracy and personalization of diagnosis, while improving the computer technology performance of the server in terms of prompt construction, result parsing and structured data management.
[0404] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.
[0405] In this invention, the server includes: a component for receiving health-related information from a user terminal and performing natural language processing to extract symptom elements and past elements; a component for automatically generating a first type of prompt statement for inputting into a generative artificial intelligence model based on the extracted elements and original health information to obtain a diagnostic result containing possible disease names and their reasons; a component for obtaining the severity of each disease by referring to an information source storing medical information for the disease names in the diagnostic result and generating a second type of prompt statement for overall severity evaluation so that the generative artificial intelligence model outputs the overall severity; a component for generating a third type of prompt statement for user-oriented explanatory information based on the diagnostic result and the overall severity so that the generative artificial intelligence model outputs user-readable explanatory information; and optionally, a component for generating prompt statements for drug recommendations and medical institution recommendations based on the diagnostic result, severity, emotional information, and location information, calling the generative artificial intelligence model to obtain corresponding recommendation results, and sending the explanatory information and recommendation results to the terminal. This transforms generative artificial intelligence models from simple, single-instance black-box calls into multi-round collaborative computation processes orchestrated in stages by the server and finely controlled by prompt statements. This enables the server to generate, manage, and reuse structured diagnostic and recommendation data within the computer, improving the consistency and interpretability of diagnostic assistance results. At the same time, without changing the terminal hardware, it significantly improves the computer technology performance of online consultation and health management systems in terms of prompt statement construction, model call control, result parsing, and personalized recommendations.
[0406] "User terminal" refers to an electronic device used by users to input health status information, emotional information and location information, and to communicate data with a server, including but not limited to smartphones, tablet computing devices or desktop computing devices.
[0407] "Health status related information" refers to a set of data sent by the user terminal to reflect the user's current physical and mental state, including symptom information, past information, allergy information, lifestyle information, and subjective descriptions related to health.
[0408] "Symptom information" refers to text or structured data used in health-related information to characterize a user's current perceived or perceived abnormal state, such as fever, cough, pain level and duration.
[0409] "Past information" refers to data related to a user's past health status, including past medical history, past treatment history, past medication history, and other historical medical information that may affect the current diagnosis.
[0410] Natural Language Processing (NLP) refers to the computational processing by which a server performs word segmentation, part-of-speech tagging, entity recognition, and syntactic analysis on text data provided in natural language form in order to extract structured semantic elements.
[0411] "Symptom elements" refer to one or more symptom features extracted from symptom information through natural language processing and represented in a structured form, which are used for constructing subsequent prompts and inferring diseases.
[0412] "Generative artificial intelligence models" refer to artificial intelligence models trained based on machine learning algorithms that can automatically generate corresponding output text or structured results based on input prompts. These models are used to perform tasks such as disease inference, severity assessment, explanatory information generation, and recommendation information generation.
[0413] "Prompt statements" refer to text data used as input to generative artificial intelligence models. This text data describes task instructions, context information, and constraints in natural language form, guiding the generative artificial intelligence model to output results in the expected format and content.
[0414] "Diagnosis results" refers to the output information obtained by processing prompts containing symptom elements and past information based on a generative artificial intelligence model. This output information includes at least one or more possible disease names and the reason for their generation.
[0415] "Possible disease name" refers to the identifier of a disease that the generative artificial intelligence model considers to be consistent with the user's health status in the diagnostic results, and is used to indicate a candidate diagnosis.
[0416] "Information source" refers to a storage system accessible by a server for storing medical information, drug information, or medical institution information, including local databases, external databases, or remote information service systems.
[0417] "Medical information" refers to a collection of data stored in information sources that is related to disease definitions, clinical manifestations, severity classifications, and treatment strategies, and is used to support auxiliary analysis of diagnostic results.
[0418] "Severity" refers to the level of severity of a disease or overall health status, usually expressed in predefined levels such as mild, moderate, and severe.
[0419] "Overall Severity" refers to an overall validation indicator that reflects the user's current health risk, obtained by comprehensively considering the severity information of multiple possible disease names in the diagnostic results.
[0420] "Explanatory information" refers to user-oriented natural language text generated based on the diagnosis results and overall severity of the illness, used to explain the health status and recommended measures to users in an easy-to-understand manner.
[0421] "Emotional information" refers to data used to represent a user's current emotional state, including the categories and intensity of emotions such as tension, anxiety, fear, and calmness obtained through voice, text, or image analysis.
[0422] "Drug information" refers to data stored in information sources that relates to drug names, ingredients, indications, adverse reactions, dosage and administration, and contraindications, and is used to support the selection and recommendation of candidate drugs.
[0423] "Candidate drugs" refers to one or more drugs selected from drug information based on diagnostic results, severity of illness, and emotional information that may be suitable for the user's current health status.
[0424] "Drug recommendation results" refer to the recommendation information output by a generative artificial intelligence model based on prompts containing candidate drugs and relevant constraints, used to guide the provision or purchase of drugs, including a list of recommended drugs and the reasons for it.
[0425] "Drug delivery processing" refers to the computer processing flow used to prepare or distribute drugs, which is executed through system interaction with the drug provider based on drug recommendations.
[0426] "Drug purchase processing" refers to the computerized process of completing drug transactions through electronic payment or order management based on drug recommendations.
[0427] "Location information" refers to data used to represent the geographical location of a user terminal or user, including but not limited to latitude and longitude, administrative regions, or location signals provided by location services.
[0428] "Medical institution information" refers to a collection of data stored in information sources that is related to the name, address, department setup, service capacity, and opening hours of medical institutions.
[0429] "Candidate medical institutions" refers to one or more medical institutions that may be suitable for the user to seek medical treatment, selected from medical institution information based on diagnostic results, severity of illness, and location information.
[0430] "Recommended medical institutions" refers to medical institutions that are ultimately selected by a generative artificial intelligence model or server from among the candidate medical institutions based on predetermined judgment criteria and recommended to users.
[0431] "Recommendation information" refers to output data that includes recommended medical institutions and the reasons for the recommendations, used to explain the basis for the recommendations to users and assist them in making medical decisions.
[0432] In one embodiment of the present invention, the server serves as the core computing platform, and the terminal serves as the user interface and data acquisition device. The user inputs health-related data through the terminal and receives diagnostic and recommendation results from the server. This embodiment focuses on describing the internal module division, data structure, algorithm flow, and collaborative method with the generative artificial intelligence model and prompt statements of the server, thereby supporting the system functions described in the claims.
[0433] I. Overall System Composition The server includes a processor, a memory, a network interface, and a program stored in the memory. When the server's processor executes the program, it implements various functional modules of the present invention, including but not limited to: a data receiving module, a natural language processing module, a feature extraction module, a prompt statement generation module, a generative artificial intelligence invocation module, a medical information retrieval module, a severity assessment module, an emotion analysis and processing module, a drug recommendation module, a medical institution recommendation module, a result interpretation generation module, and a data storage and logging module.
[0434] The terminal includes a processor, memory, display unit, input unit (touchscreen, keyboard), audio acquisition unit (microphone), image acquisition unit (camera), positioning unit (positioning module), and network communication unit. When the terminal executes the application, it collects and encapsulates the user's symptom text, voice, image, and location data into structured data, sends it to the server through the network interface, and receives diagnostic results, explanatory information, drug recommendations, and medical institution recommendations returned by the server, which are then displayed on the display unit.
[0435] Users can input health-related information through the terminal, including symptom information, past history, allergy information, lifestyle habits information, and optional emotional descriptions. If necessary, they can also take images of lesions such as skin through the camera.
[0436] II. Core Software Modules and Data Processing Performed by the Server 1. Natural Language Processing and Feature Extraction The server uses a natural language processing library (such as a general-purpose natural language processing framework) to implement its natural language processing module. The server takes the symptom text sent by the user through the terminal as input and performs the following data processing: The server performs word segmentation, part-of-speech tagging, and named entity recognition on the text, extracting symptom words (such as "fever," "headache," and "cough"), time expressions (such as "3 days" and "one week"), intensity expressions (such as "mild" and "severe"), and negative words (such as "no chest pain" and "no difficulty breathing"). The server stores these extracted elements as structured features, for example, in the form of key-value pairs or feature vectors: "symptom = fever, duration = 3 days, highest body temperature = 38.5 degrees Celsius," etc.
[0437] The server uses a feature extraction module to construct a unified feature object by combining these structured features with the user's past information and allergy information. This feature object serves as input to the subsequent prompt generation module and the generative artificial intelligence model invocation module, thereby avoiding direct reliance on the original free text, reducing noise and ambiguity, and improving the consistency and controllability of the model input.
[0438] Through the above processing, the server converts unstructured natural language into structured symptom elements and prior elements within the computer, achieving dimensionality reduction and semantic abstraction of the original data, thereby reducing the complexity of constructing subsequent prompt statements and improving overall computational efficiency.
[0439] 2. Structure and Invocation Method of Generative Artificial Intelligence Models In this embodiment, the server uses a generative artificial intelligence model as its text generation and inference engine. The server can employ a multi-layered encoder-decoder architecture based on a self-attention mechanism, such as a deep neural network model containing several self-attention layers and feedforward layers. The server pre-trains or fine-tunes this model using a large-scale medical corpus, with training objectives including next-word prediction and instruction following tasks.
[0440] During the training phase, the server uses the cross-entropy loss function as the error function to measure the difference between the actual output and the model's predicted output in the training samples, and updates the model weights through backpropagation and gradient descent. The server can use optimization algorithms such as adaptive gradient optimization to set parameters such as the learning rate and weight decay. During training, the server employs data augmentation methods, such as synonym replacement, template expansion, and noise injection, to increase the diversity of training samples, thereby improving the model's robustness to various user expressions.
[0441] During the inference phase, the server takes the prompt statements as the model input sequence and uses beam search or temperature sampling strategies to generate the output text. The server controls the stability and security of the output by setting hyperparameters such as generation temperature, maximum length, and prohibition of certain inappropriate words. These specific generation parameters can be configured separately for different tasks (diagnosis, severity assessment, explanatory information generation, drug recommendation, and medical institution recommendation), thereby achieving the technical effect of multiple tasks sharing the same generative artificial intelligence model within the computer without interfering with each other.
[0442] 3. Automatic generation and phased invocation of prompt statements The server uses a prompt statement generation module to automatically construct multiple types of prompt statements based on different tasks. The server embeds symptom elements, past medical history, and allergy information into predefined query templates to form the first type of prompt statement for diagnostic tasks. For example, the server can generate the following sample prompt statement: "User symptoms: Fever for three days, body temperature 38.5 degrees Celsius, cough with yellow phlegm, chest tightness."
[0443] User's past medical history: Asthma.
[0444] User's history of drug allergies: allergic to penicillin.
[0445] Based on the information above, please list the three most likely diseases, sorting them from most likely to least likely, and provide a brief explanation for each disease. The server inputs this type of prompt statement into a generative artificial intelligence model to obtain a diagnosis result containing one or more possible disease names and their rationales. Based on the diagnosis result, the server then retrieves the severity definitions of each disease from medical information sources, such as the criteria for classifying mild, moderate, and severe. The server then uses this severity information to construct a second type of prompt statement, such as: List of possible diseases: 1. Acute bronchitis 2. Pneumonia 3. Upper respiratory tract infection Current symptoms: fever of 38.5 degrees Celsius, cough with yellow phlegm, chest tightness, but no difficulty breathing or chest pain.
[0446] Based on the above illness and symptoms, please provide an overall severity assessment (choose from mild, moderate, or severe), and indicate whether immediate emergency medical attention is required. By employing this phased prompting design, the server logically decouples the diagnostic results from the severity assessment, while technically enabling the structured storage and reuse of intermediate model outputs. These structured intermediate results can be used for log tracking, model tuning, and subsequent tasks (such as drug recommendations and medical institution recommendations), thereby significantly improving system scalability and maintainability.
[0447] The server also uses the third type of prompt statement generation module to automatically generate user-oriented explanation requests based on the diagnosis results and overall severity, such as: The structured data of the diagnostic results are as follows: - Illness: Upper respiratory tract infection; Severity: Mild. - Main symptoms: fever, cough, yellow phlegm; - Recommended treatment: Rest at home, observe changes in symptoms, take medication as prescribed, and seek medical attention if symptoms do not improve within three days.
[0448] Please address the needs of ordinary users, using simple and reassuring language, and explain the current situation and offer suggestions in no more than 150 words. The server thus obtains user-readable, structured, and logically supported explanatory information. Because the prompts are automatically generated by the server program, rather than being manually written each time, this automated, template-based prompt mechanism represents a new internal computer control method. Combined with generative artificial intelligence models, it enables controllable management of complex, multi-turn reasoning processes.
[0449] 4. Fusion processing of emotional information and multi-source data In this embodiment, the server incorporates the user's emotional state as an additional feature into the diagnosis and recommendation process through an emotion analysis module. The server can analyze emotional text or voice sent by the user via the terminal, extracting the emotion category and intensity. The server combines this emotional information with symptom elements, past information, location, and other features to form an extended feature object, which is used to explicitly prompt in the prompting statements, for example: "User's physical symptoms: chest pain, palpitations, mild shortness of breath."
[0450] User sentiment: According to the analysis, the current sentiment is intense fear and high anxiety.
[0451] Please consider that emotions may indicate a potential risk of serious illness, comprehensively assess the overall severity of the condition, and determine whether immediate travel to the emergency room or calling for emergency medical assistance is necessary. By explicitly writing emotional information into prompts, the server enables generative AI models to exhibit multimodal fusion capabilities during computation, unlike traditional rule-based systems. The server can also internally set independent rules, such as "when the overall severity of illness is moderate but the emotional state is highly anxious, increase the medical recommendation level." This hybrid approach combines numerical threshold judgment with generative AI reasoning, offering high flexibility and interpretability, thereby improving decision-making accuracy at the computer technology level.
[0452] 5. Technical Implementation of Drug Recommendation and Medical Institution Recommendation In the drug recommendation module, the server takes the diagnosis result, severity of illness, emotional information, and allergy information as input, accesses the drug information source, and filters candidate drugs. The server uses algorithms such as keyword matching, indication checking, and ingredient interaction checking to filter and rank the candidate drugs. Then, the server embeds these candidate drugs and their constraints into a drug recommendation prompt, for example: Diagnosis: Mild upper respiratory tract infection.
[0453] User's history of drug allergies: allergic to penicillin.
[0454] User sentiment: Mild anxiety.
[0455] Based on the information above, please recommend 2-3 suitable medications, including the medication name, purpose, dosage, and precautions. Please avoid recommending medications containing penicillin. The server extracts structured information (drug name, dosage, frequency) from the drug recommendations generated by the generative AI model, and cross-validates it with local rules (maximum dosage limits, age limits, etc.) before outputting it to the terminal. This combination of "model generation + rule validation" differs from traditional manual prescriptions or simple rule engines. It can reduce the risk of errors while ensuring flexibility, and reduce the systematic errors caused by inconsistent output from a technical perspective.
[0456] The server also utilizes multi-source information in the medical institution recommendation module. Using location and map services, the server selects candidate medical institutions from the available information based on the user's location and desired department type, and sorts them according to distance, specialty matching, and emergency capabilities. The server then constructs recommendation prompts for medical institutions, such as: "Diagnosis: Suspected acute cardiovascular event, severity: moderate to severe."
[0457] User location: A certain district in a certain city.
[0458] Optional medical institutions: 1. Medical institution A, located 2 kilometers away, has a 24-hour cardiology emergency department; 2. Medical institution B, located 3.5 kilometers away, has a cardiology outpatient department but no emergency department.
[0459] Based on the urgency of the situation, recommend the best hospital to the user and explain the reasons for the recommendation in concise language. The server thus obtains recommended medical institutions and their reasons for recommendation, and sends this information to the terminal for navigation and appointment booking. By centrally managing prompts and data sources on the server side, the system reduces the overhead of multiple requests and API calls, which helps improve overall response speed and reduce network communication load.
[0460] III. Explanation of Technical Effects and Causal Relationship The server achieves the following technical effects through the aforementioned structured feature extraction, phased prompt statement generation, multi-round generative artificial intelligence invocation, and hybrid rule validation: 1. Improved accuracy During the diagnostic process, the server first extracts symptom and historical data elements before constructing prompts, thereby reducing noise in the original text and making the input information for the generative AI model more structured and complete. Through phased calls, the server can optimize the disease list, severity level, and user instructions separately, avoiding interference caused by processing multiple tasks simultaneously in a single call, thus improving overall diagnostic accuracy.
[0461] 2. Improved processing speed and computational efficiency The server breaks down the complex reasoning process into several modules and multiple rounds of calls. Through structured caching of intermediate results, it enables the reuse of some results and incremental computation. For example, the server can re-make recommendations only for newly added sentiment information or location changes without changing the basic diagnosis, instead of regenerating all diagnostic results. This modular design reduces unnecessary redundant calculations and improves server resource utilization.
[0462] 3. Improved Data Management and Maintainability The server internally uses a unified data structure to store prompt statements, model outputs, and structured diagnostic results, making logging and tracking analysis simple and transparent. Developers can use these structured records to statistically analyze the model's output distribution under different symptom combinations and then tailor prompt statement templates and model parameters accordingly.
[0463] 4. Unconventional decision-making rules and unique processing methods The server employs a hybrid processing strategy across multiple stages: "generative AI model inference + fixed rule verification + multi-source feature fusion." This combination differs from simple human decision automation or traditional rule engines. For example, in drug recommendation and disease severity assessment, the server considers both the textual interpretation of the model output and filters information using predefined safety thresholds and disabling rules from the information sources, thus forming a unique decision path within the machine. This decision path is explicitly implemented through software modules, representing a novel data processing method within computers.
[0464] 5. Reduced communication load and response time The server integrates multiple pieces of information (such as diagnosis, severity, emotion, and location) with a single prompt, obtaining a comprehensive recommendation result in a single call to a generative AI model, reducing the number of round trips between requests and external model services. The server also reduces frequent access to external services by caching medical and pharmaceutical information in a local database, thereby lowering network communication load and shortening average response time.
[0465] IV. Multiple Implementation Methods and Alternative Solutions In different implementations, the server can employ different generative artificial intelligence model structures. For example, the server can use a sequence generation model with a decoder-only structure, or a multimodal model with an encoder-decoder structure, to support mixed text and image input. In the sentiment analysis part, the server can use only a text sentiment analysis module, or combine it with a speech sentiment recognition model, achieving different balances between accuracy and cost.
[0466] The terminal can be a mobile terminal, desktop device, or embedded medical device in different implementations. Users can input health status-related information through various methods such as touch, voice, or images. The server can deploy generative artificial intelligence models locally or perform model calculations through remote inference services, depending on the network environment and computing resources.
[0467] As can be seen from the above description of various implementation forms, this invention is not limited to a specific hardware platform or a single model architecture. Instead, by adopting a unified prompt statement generation and multi-round inference control mechanism on the server side, it integrates multi-source health data, emotion data, and location information into the same calculation process. This enables structured calling and collaborative calculation of generative artificial intelligence models at the computer technology level, thereby achieving the technical effects of improving diagnostic accuracy, accelerating response speed, enhancing data management capabilities, and reducing communication and computing burden.
[0468] use Figure 14 The processing procedure is explained.
[0469] Step 1: Users input their health status information through the terminal.
[0470] Users can manually input text data such as symptom descriptions, past medical history, and allergy information in the application interface on the terminal, and can also choose to input symptoms by voice or take images of the lesion site by camera.
[0471] Input: Text data (symptoms, past medical history, allergy information, etc.) entered or collected by the user on the terminal, optional voice data, and image data.
[0472] Output: A structured request data packet generated internally by the terminal (containing text fields, audio / video file references, timestamps, and user identifiers).
[0473] Based on the user's actions, the terminal packages the above data into a request object and prepares to send it to the server over the network.
[0474] Step 2: The terminal sends a health status-related request to the server.
[0475] The terminal uses a network communication module and a secure communication protocol to send the data packet generated in step 1 to the preset interface address of the server.
[0476] Input: Request data packets (health status information and media data) from within the terminal.
[0477] Output: HTTP / HTTPS requests transmitted to the server over the network.
[0478] Before sending, the terminal encrypts the request data and attaches a user identifier and a session identifier to the request header so that the server can authenticate and track the session.
[0479] Step 3: The server receives and parses health status-related information.
[0480] The server receives requests from the terminal at the designated interface, verifies the data format and user identity, and parses the text fields and media file references in the request body.
[0481] Input: An HTTP / HTTPS request (containing JSON text and media data) sent from the terminal.
[0482] Output: The server's raw text string, media data handle, and session context object.
[0483] The server stores the parsed raw text and session information into an in-memory structure and records a new session log entry in the database to provide a correlation identifier for subsequent processing.
[0484] Step 4: The server performs natural language processing on the symptom text and extracts features.
[0485] The server calls the natural language processing module to perform word segmentation, part-of-speech tagging, entity recognition, and dependency parsing on the original text, extracting symptom words, time expressions, degree expressions, and negation structures.
[0486] Input: Raw text string (containing symptom description, past medical history, allergy information).
[0487] Output: Structured feature objects (list of symptom elements, list of past elements, list of allergy elements).
[0488] The server performs syntactic analysis on each sentence, converting phrases such as "fever for three days", "no chest pain", and "mild cough" into standardized fields, such as "symptom=fever, duration=3 days; symptom=cough, severity=mild; chest_pain=false", and stores them in a feature object.
[0489] Step 5: The server constructs initial diagnostic prompts.
[0490] Based on the feature objects and original text from step 4, the server uses a template engine to generate the first type of prompt statements for the diagnostic task, populating the natural language template with symptom elements, past information, and allergy information.
[0491] Input: Structured feature objects (symptom features, past features, allergy features) and raw text.
[0492] Output: First-class prompt text for generative artificial intelligence models.
[0493] The server inserts explicit instructions and output requirements into the template, such as requiring a list of several possible disease names and reasons, thereby ensuring that the output of the generative artificial intelligence model has a fixed structure and is parseable.
[0494] Step 6: The server invokes a generative artificial intelligence model to perform disease candidate diagnoses.
[0495] The server sends the prompt statement generated in step 5 to the generative artificial intelligence model interface, sets parameters such as model name, maximum output length, and temperature, and waits for the model to return diagnostic text.
[0496] Input: The text of the first type of prompt statement.
[0497] Output: Diagnostic text results containing multiple possible disease names and corresponding reasons.
[0498] The server interacts with the model service internally via HTTP or other remote procedure call protocols, and records the call parameters and response time for subsequent performance optimization.
[0499] Step 7: The server parses the diagnostic text and generates a list of diseases.
[0500] The server performs regular expression matching or structured parsing on the diagnostic text returned in step 6 to extract the disease name, sorting, and brief reason, generating a unified disease list structure.
[0501] Input: Diagnostic text returned by the generative artificial intelligence model.
[0502] Output: A structured list of diseases (each entry includes the disease name, sort number, and reason for diagnosis).
[0503] The server handles errors such as failed parsing or incorrectly formatted text by re-calling the model or reverting to an alternative template to ensure the integrity of inputs in subsequent processes.
[0504] Step 8: The server queries the baseline information on the severity of each disease from medical information sources.
[0505] Based on the list of disease names generated in step 7, the server accesses local or remote medical information sources to retrieve the standard severity grading rules and key warning signs for each disease.
[0506] Input: A structured list of diseases (a set of disease names).
[0507] Output: Baseline information on the severity of each disease (such as classification rules and typical severe symptoms).
[0508] The server integrates the search results into a mapping structure as supplementary reference data for severity assessment, thereby reducing reliance on the single judgment of generative artificial intelligence models.
[0509] Step 9: The server constructs prompts for severity assessment.
[0510] The server combines the disease list and current symptom elements obtained in step 7 with the baseline information on severity obtained in step 8 to generate a second type of prompt statement, requesting the generative artificial intelligence model to conduct a unified assessment of the overall severity.
[0511] Input: Disease list, symptom elements, and baseline information on severity.
[0512] Output: Text of the second type of severe illness assessment prompt for generative artificial intelligence models.
[0513] The server limits the output format in the prompt statement to "choose one of mild, moderate or severe, and give a judgment on whether emergency treatment is required", so that the generated results are easier for the program to parse and use for decision-making.
[0514] Step 10: The server calls a generative artificial intelligence model to obtain the overall severity level.
[0515] The server inputs the prompt from step 9 into the generative artificial intelligence model and receives the severity level and urgency advice text returned by the model.
[0516] Input: Text of the prompt statement for the second category of severe illness assessment.
[0517] Output: Text results including overall severity level (mild / moderate / severe) and emergency room recommendations.
[0518] The server extracts keywords from the model output and converts content such as "severe severity = moderate severity, it is recommended to seek outpatient treatment as soon as possible" into structured fields for subsequent logical judgment.
[0519] Step 11: The server integrates diagnostic results and severity information.
[0520] The server combines the disease list from step 7 with the overall severity and emergency recommendations from step 10 to form a complete diagnostic summary data structure.
[0521] Input: Structured list of diseases, overall severity information, emergency department recommendations.
[0522] Output: Comprehensive diagnostic summary object (including disease candidates, severity labels, and emergency markers).
[0523] The server writes this comprehensive object into the database as the core record of the session diagnostic results, providing a unified data source for subsequent drug recommendations and medical institution recommendations.
[0524] Step 12: The server analyzes and processes sentiment information (if any).
[0525] Users can input text to describe their emotions on the terminal, or transmit emotion-related data through voice and images; the server performs emotion classification model inference on this data to obtain the user's current emotion label and intensity.
[0526] Input: Emotion-related raw data (text, voice, or image).
[0527] Output: Structured emotional information (emotional category, such as anxiety, fear, and emotional intensity).
[0528] The server associates and stores emotion information with diagnostic summary objects, and inserts it as an additional feature in the construction of subsequent prompt statements to reflect the impact of the user's psychological state on decision-making.
[0529] Step 13: The server constructs user-defined prompts.
[0530] Based on the comprehensive diagnostic summary and emotional information, the server generates a third type of prompt statement and requests a generative artificial intelligence model to generate easy-to-understand explanatory text so that users can understand their current health status and recommended measures.
[0531] Input: Comprehensive diagnostic summary of the subject and emotional information.
[0532] Output: Text of third-class user instruction prompts for generative artificial intelligence models.
[0533] The server provides a structured data summary in the prompt and makes explicit requirements on tone, length, and style, such as "soothing tone, about 150 words, suitable for non-professionals to read".
[0534] Step 14: The server invokes a generative artificial intelligence model to generate user description information.
[0535] The server inputs the third type of prompt statement into the generative artificial intelligence model to obtain a natural language explanatory text that explains the diagnosis results, severity of the illness, and recommended actions.
[0536] Input: Text of the third type of user instruction prompt.
[0537] Output: User-oriented instructional text.
[0538] The server performs security filtering and sensitive word checks on the explanatory text to ensure that the output content complies with medical communication standards, and saves the approved explanatory text along with the session records.
[0539] Step 15: The server performs drug candidate screening (if drug recommendations are needed).
[0540] Based on a comprehensive diagnostic summary, emotional information, and allergy information, the server filters out a set of drugs from drug information sources that match the candidate diseases and do not contain contraindicated ingredients.
[0541] Input: Comprehensive diagnostic summary of the subject, emotional information, allergy information, and drug information source data.
[0542] Output: A list of candidate drugs (including drug name, ingredients, indications, contraindications, and usual dosage).
[0543] The server sorts candidate drugs according to indication matching, risk level, and emotional state requirements (such as whether sedation is required), forming a preliminary candidate set of recommended drugs.
[0544] Step 16: The server constructs a prompt statement for recommending medicines.
[0545] The server inserts a list of candidate drugs and diagnostic summaries into a prompt template specifically for drug recommendations, and requests a generative artificial intelligence model to generate the final recommended drugs and descriptions.
[0546] Input: List of candidate drugs, comprehensive diagnostic summary of the subject, and emotional information.
[0547] Output: Text of drug recommendation suggestions for generative artificial intelligence models.
[0548] The server requires the model to clearly state the purpose, recommended dosage, and precautions for each recommended drug in the prompt statement, and emphasizes avoiding conflicts with ingredients that the user is known to be allergic to or contraindicated.
[0549] Step 17: The server invokes a generative artificial intelligence model and analyzes the drug recommendation results.
[0550] The server sends the prompt from step 16 to the generative artificial intelligence model. After receiving the text containing the recommended medicine and its reasons, it parses out fields such as medicine name, dosage, and frequency of use.
[0551] Input: Text of a drug recommendation prompt.
[0552] Output: Structured drug recommendation results (a list of drug entries, including name, dosage, frequency, and reason).
[0553] The server performs rule verification on the parsed drugs again, filtering out recommendations with unreasonable dosages or conflicts with basic rules to ensure the security and consistency of the output results.
[0554] Step 18: The server performs candidate screening of medical institutions (if medical recommendations are needed).
[0555] Based on the comprehensive diagnostic summary and user location information, the server filters out candidate medical institutions with relevant departments that are within a reasonable distance from the medical institution information source.
[0556] Input: Comprehensive diagnostic summary object, user location information, and medical institution information source data.
[0557] Output: List of candidate medical institutions (including name, address, specialty information, emergency capacity, and distance).
[0558] The server sorts candidate medical institutions based on departmental suitability, whether they have emergency services capabilities, and distance.
[0559] Step 19: The server is configured with the recommended prompt statements for medical institutions.
[0560] The server integrates the list of candidate medical institutions, diagnosis results, severity information, and location information to generate medical institution recommendation prompts, and requests the generative artificial intelligence model for selection and interpretation.
[0561] Input: List of candidate medical institutions, comprehensive diagnostic summary objects, user location information.
[0562] Output: Recommendation prompts for medical institutions based on generative artificial intelligence models.
[0563] The server prompts the model to prioritize recommending medical institutions with emergency capabilities or specialty advantages based on the urgency level, and outputs a brief reason for the recommendation.
[0564] Step 20: The server invokes a generative artificial intelligence model and analyzes the recommendations from medical institutions.
[0565] The server sends the prompt from step 19 to the generative artificial intelligence model to obtain the text containing the name of the recommended medical institution and the reason for the recommendation, and parses it into structured recommendation information.
[0566] Input: The text of the medical institution's recommended prompt.
[0567] Output: Structured medical institution recommendation results (including recommended medical institutions, ranking, and reasons).
[0568] The server performs consistency checks on the recommendation results, such as confirming that the recommended medical institutions actually exist and are in the candidate list, to avoid invalid or out-of-service institutions.
[0569] Step 21: The server generates the final response data packet and sends it to the terminal.
[0570] The server integrates user instructions, drug recommendations, medical institution recommendations, and necessary emergency prompts into a unified response object, which is then sent to the terminal via a network interface.
[0571] Input: User description text, structured drug recommendation results, structured medical institution recommendation results, emergency room prompts.
[0572] Output: The response data packet (JSON or equivalent structured response) sent to the terminal.
[0573] The server de-identifies sensitive information before sending it and records response summaries and timestamps for logging and auditing.
[0574] Step 22: The terminal receives and displays the diagnostic and recommendation results returned by the server.
[0575] After receiving the server's response, the terminal parses the JSON or structured response and maps the diagnostic summary, user instructions, drug recommendations, and medical institution recommendations to the interface components for display.
[0576] Input: The response data packet returned by the server.
[0577] Output: The diagnostic results interface displayed on the terminal screen (text description, drug recommendation list, medical institution list, etc.).
[0578] The terminal uses different colors and reminder methods (such as pop-ups and icon highlighting) based on the severity and emergency status of the condition to guide users to pay attention to serious warning information.
[0579] Step 23: Users can then perform further actions based on the information provided.
[0580] Users can read diagnostic instructions, view recommended medications and medical institutions on the terminal interface, and choose whether to purchase medication, make an appointment with a medical institution, or go to the emergency room immediately based on the prompts.
[0581] Input: Diagnostic and recommendation information displayed on the terminal.
[0582] Output: The user's action selection (e.g., clicking to buy medicine, clicking to make an appointment, ignoring suggestions, etc.).
[0583] The user's operation instructions are captured by the terminal and packaged into a new request, which is then sent to the server to trigger subsequent payment processing, appointment processing, or session termination processes, thereby completing the entire closed loop.
[0584] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0585] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0586] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0587] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0588] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0589] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0590] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0591] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0592] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0593] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0594] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0595] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0596] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0597] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0598] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0599] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0600] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0601] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0602] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0603] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0604] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0605] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0606] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0607] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0608] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0609] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0610] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0611] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0612] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0613] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0614] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0615] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0616] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0617] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0618] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0619] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0620] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0621] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0622] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0623] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0624] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0625] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0626] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0627] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0628] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0629] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0630] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0631] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0632] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0633] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0634] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0635] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0636] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).
[0637] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0638] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0639] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0640] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0641] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0642] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0643] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0644] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0645] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0646] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0647] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0648] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0649] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0650] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0651] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0652] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0653] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0654] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0655] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0656] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0657] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0658] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0659] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0660] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0661] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0662] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0663] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0664] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0665] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0666] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0667] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0668] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0669] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0670] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0671] In addition, the following notes are provided in response to the above explanation.
[0672] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for obtaining health status-related information from a user terminal; A device for converting the health status-related information into structured information; An apparatus for generating prompt statements for a generative artificial intelligence model based on the health status-related information and the structured information, and for obtaining disease candidate information through the generative artificial intelligence model; A device for parsing the disease candidate information and extracting disease classification information; A device for calculating disease severity information by referring to an information storage device based on the disease classification information and the health status-related information; A device for generating diagnostic results and behavioral guidance information based on the disease candidate information and the disease severity information; A device for sending the diagnostic results information and the behavioral guidance information to the user terminal.
[0673] (Note 2) The information processing system according to Appendix 1 is characterized in that, The apparatus for generating prompt statements for a generative artificial intelligence model is configured to dynamically update the prompt statements based on the content of the health status-related information and the response format of the generative artificial intelligence model, and to repeatedly control the query processing for the generative artificial intelligence model in order to improve the accuracy of the disease candidate information.
[0674] (Note 3) The information processing system according to Appendix 1 is characterized in that, The device for generating diagnostic result information and behavioral guidance information is configured to extract candidates of medical service providers and medical-related items based on the disease severity information and the diagnostic result information, with reference to a medical-related resource information storage device, and generate the behavioral guidance information as a recommendation message containing the candidates.
[0675] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: An apparatus for receiving health status-related information from a user terminal and generating structured data containing such health status-related information; An apparatus for preprocessing the structured data, extracting features including symptom information, physiological measurement information and past information, and generating prompt statements for input to a generative artificial intelligence model based on the features. A device for inputting the prompt statement into the generative artificial intelligence model to obtain diagnostic result information including possible disease names and the severity corresponding to the disease names; An apparatus for extracting the disease name and severity from the diagnostic result information, referring to a cost information storage unit that is equipped with multiple medical expense payment conditions corresponding to different severity levels, and determining a medical expense payment scheme applicable to the user based on the severity level; A device for sending the diagnostic results and the medical expense payment plan as output information to the user terminal.
[0676] (Note 2) According to the information processing system described in Appendix 1, the device for generating the prompt statement is configured to generate natural language text as input to the generative artificial intelligence model, the natural language text including symptom information, physiological measurement information and a summary of past information, an indication for inferring possible disease names and the severity of the disease, and an output request for an evaluation item for selecting a medical expense payment scheme based on the severity of the disease.
[0677] (Note 3) According to the information processing system described in Appendix 1, the device for determining the medical expense payment scheme is configured to: determine the medical expense payment scheme as a first payment scheme including a one-time payment condition when the severity is low, and determine the medical expense payment scheme as a second payment scheme including an installment payment condition when the severity is high, and notify the user terminal by selecting at least one of the first payment scheme and the second payment scheme based on the diagnostic result information.
[0678] Example 2 (Note 1) An information processing system, characterized in that it comprises: An apparatus for receiving images of biological parts from a user terminal, converting the images into a predetermined image format upon receipt, and acquiring the images through a communication channel; A means for performing image processing on the received image to perform pixel value normalization, noise removal, and pixel array transformation to generate input data for image analysis; An apparatus for using the input data for image analysis to extract image features through an image recognition machine learning model and match them with reference data in a data storage device storing symptom information, thereby estimating multiple candidate diseases. An apparatus for acquiring text data of consultation content input by a user, and generating structured data containing symptoms, location of occurrence, duration of occurrence, accompanying symptoms, and medical history from the text data using natural language processing technology; An apparatus for fusing the structured data with the plurality of candidate diseases, calculating the probability of each candidate disease, and generating diagnostic contextual information containing the fusion results; An apparatus for generating prompt statements based on the diagnostic context information and inputting them into a generative artificial intelligence model, and for inputting the prompt statements and the diagnostic context information into the generative artificial intelligence model to generate diagnostic results and explanatory text about the candidate disease; An apparatus for generating prompt statements related to the severity assessment of each disease referenced in the diagnostic results and inputting them into the generative artificial intelligence model to generate explanatory text about the severity assessment. A device for sending diagnostic information based on the diagnostic results and the severity assessment to a user terminal.
[0679] (Note 2) According to the information processing system described in Appendix 1, the system is configured to: based on the diagnostic results and the severity assessment, refer to a data storage device storing information on medical service institutions, determine recommended medical service institutions including the treatment department, geographical location, and urgency of the visit, generate a prompt statement to prompt the recommended medical service institutions, and input the prompt statement into the generative artificial intelligence model to generate a recommendation message.
[0680] (Note 3) According to the information processing system described in Appendix 1, the system is configured to: based on the diagnostic results and the severity assessment, refer to a data storage device storing medical supplies information, select medical supply candidates suitable for symptom self-management or telemedicine, generate prompt statements for generating explanatory text containing usage precautions for the candidates and conditions for consulting medical personnel, and input the prompt statements into the generative artificial intelligence model to generate the explanatory text.
[0681] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for receiving health status-related information from a user's terminal; An apparatus for performing natural language processing on symptom information and past information contained in the health status-related information to extract symptom elements. A device for generating prompt statements for input to a generative artificial intelligence model based on the symptom elements and information related to the health status; A device for inputting the prompt statement into the generative artificial intelligence model to obtain a diagnostic result containing possible disease names and their reasons; A device for obtaining the severity of each disease name by referring to an information source storing medical information, based on the disease names contained in the diagnostic results. An apparatus for generating prompt statements related to the severity assessment based on the diagnostic results and the severity, and inputting the prompt statements into the generative artificial intelligence model so that the generative artificial intelligence model can evaluate the overall severity. An apparatus for generating prompt statements to provide information to a user based on the diagnostic results and the overall severity of the illness, and inputting the prompt statements into the generative artificial intelligence model to obtain user-oriented explanatory information; A device for sending the explanatory information to the terminal.
[0682] (Note 2) The information processing system according to Appendix 1 is characterized in that, The system further includes: a device for selecting candidate drugs from an information source storing drug information based on the diagnosis result, the severity of the illness, and emotional information representing the user's emotional state; generating a prompt statement for a generative artificial intelligence model to generate recommendation content related to the candidate drugs; inputting the prompt statement into the generative artificial intelligence model to obtain drug recommendation results; and performing drug delivery processing or drug purchase processing according to the drug recommendation results.
[0683] (Note 3) The information processing system according to Appendix 1 is characterized in that, The system further includes: a device for extracting candidate medical institutions from an information source storing medical institution information based on the diagnosis result, the severity of the illness, the user's location information, and the emotional information; generating a prompt statement containing a judgment criterion for selecting a recommended medical institution from the candidate medical institutions; inputting the prompt statement into the generative artificial intelligence model to obtain recommendation information containing the recommended medical institution and the reasons for the recommendation; and sending the recommendation information to the terminal.
Claims
1. An information processing system, characterized in that, include: processor; The processor is configured as follows: Receive information from users regarding their health status; The received information is parsed using a generative artificial intelligence model, and prompt words are generated to instruct the generative artificial intelligence model to parse the information in order to diagnose possible disease names; The database is queried to determine the severity of the diagnosed disease name, and prompts are generated to instruct the generative artificial intelligence model to evaluate the severity based on the database.
2. The information processing system according to claim 1, characterized in that, The processor is also configured to: collaborate with nearby medical institutions, select appropriate medications, and generate prompts to instruct the generative artificial intelligence model to control the provision of the medications.
3. The information processing system according to claim 1, characterized in that, The processor is also configured to: generate prompts based on the parsing results to indicate recommendations for appropriate medical institutions, and generate recommendation messages to be output to the user.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A