Unexplained pneumonia risk case identification and information acquisition method
By combining large language models and RAG technology with MCP technology to extract key information from electronic medical records, and by combining risk case identification rule algorithms and standardized data collection with doctor-side plugins, the problems of accuracy and incomplete information collection in the identification of risk cases of pneumonia of unknown cause in existing technologies have been solved, achieving efficient and accurate case identification and information collection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI MUNICIPAL CENT FOR DISEASE CONTROL & PREVENTION
- Filing Date
- 2026-01-08
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies for identifying cases of pneumonia of unknown origin using electronic medical record data are inaccurate, lack comprehensive consideration of clinical information, and lack rapid multi-source data integration and information collection standards, leading to missed diagnoses, misdiagnoses, and information omissions.
Clinical signs and symptoms are extracted using a large language model and retrieval enhancement generation (RAG) technology. The latest data is obtained by combining model context protocol (MCP) technology. Probability scores are calculated through majority voting and risk case identification rule algorithms. Information collection is standardized through a doctor-side plugin.
It improved the accuracy of identifying high-risk cases of pneumonia of unknown origin, ensured the completeness and accuracy of information collection, and enhanced the data support capabilities for epidemic prevention and control.
Smart Images

Figure CN121885189A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a technology for predicting disease risk using electronic medical record data, and particularly to a method for identifying and collecting information on cases of pneumonia of unknown cause. Background Technology
[0002] With the rapid development of medical informatization, hospitals have accumulated a large amount of electronic medical record data. This data contains a wealth of patient information. In pneumonia prevention and control, the rapid and accurate identification of cases of pneumonia of unknown origin is crucial for epidemic monitoring, prevention and control, and patient treatment. With the development of artificial intelligence technology, using electronic medical record data for disease risk prediction has become a research hotspot; however, this field still faces many challenges in identifying pneumonia risk cases.
[0003] Existing technologies have involved unstructured data (text medical records, examination reports, etc.) processed by NLP (Natural Language Processing) technology, or directly relying on structured data (laboratory reports, medication records, ICD diagnostic codes, etc.) to identify pneumonia cases through simple rule matching methods. For example, they can search based on keywords in the medical records (such as "solidification", "solidification shadow" or "infiltration"), identify cases by matching the International Classification of Diseases (ICD) pneumonia codes, or use statistical models to analyze relevant feature data in the medical records.
[0004] (1) The research mainly focuses on pneumonia cases identified from a clinical perspective, while there are relatively few studies on the early identification of pneumonia of unknown cause risk from a public health perspective.
[0005] (2) The application technology is relatively simple and it is difficult to comprehensively consider various complex clinical information. It is insufficient to mine the hidden and deep-seated disease-related information in electronic medical records, such as the doctor's subjective description in the medical record and the patient's self-feeling. These information may contain important diagnostic clues, but traditional methods are difficult to extract effective information from them for comprehensive judgment, which may affect the accuracy of identification. When faced with complex diseases and atypical symptoms, it is easy to miss or misdiagnose.
[0006] (3) Existing related technologies lack effective means to quickly integrate multi-source data into the model when processing data, and cannot analyze and judge based on the latest information in a timely manner, which limits their application effect in actual medical scenarios.
[0007] (4) After identifying high-risk cases, there is a lack of unified standards and effective reminders for the key information of the cases that need to be collected. Furthermore, due to differences in doctors' autonomy and clinical experience, key information may be missed, which is not conducive to the subsequent judgment of the pneumonia type of the case and the assessment of the epidemic risk.
[0008] Background references: Jones G, Amoah J, Klein EY, et al. Development of an Electronic Algorithm to Identify in Real Time Adults Hospitalized WithSuspected Community-Acquired Pneumonia[J]. Open Forum Infectious Diseases, 2021. DOI: 10.1093 / ofid / ofab291. Summary of the Invention To address the issues of low accuracy and difficulty in practical application of electronic medical record data for identifying pneumonia risk cases, a method for identifying and collecting information on pneumonia of unknown cause risk cases is proposed.
[0009] The technical solution of this invention is as follows: A method for identifying and collecting information on high-risk cases of pneumonia of unknown cause, including: Step 1: Clinical sign or symptom extractor workflow based on large language model: Step 1.1: The input module receives prompts and electronic medical record data; at the same time, it establishes a connection with the electronic medical record data source through MCP technology, obtains the latest data related to the patient in real time, and integrates this data with existing electronic medical record data to form more comprehensive input data for use by subsequent modules; Step 1.2: The RAG module starts. This module retrieves relevant text information from the integrated multi-source electronic medical record data, generates text blocks suitable for processing by a large language model, and appends them to the input prompts. Step 1.3: The large language model module receives the processed input data prompts, runs it multiple times, and each time identifies clinical signs or symptoms based on the input information and generates corresponding text output; Step 1.4: The results processing module collects the output results of multiple runs of the large language model and uses majority voting to determine the final clinical signs or symptoms identification results for use by subsequent modules; Step 2: Algorithm flow for rule-based model of risk cases of pneumonia of unknown cause: Step 2.1: The data receiving module acquires the recognition results output by the clinical sign or symptom extractor based on the large language model; Step 2.2: The algorithm calculation module dynamically analyzes and calculates the acquired clinical signs, symptoms and laboratory test data according to the pre-set risk case identification rules; the algorithm will calculate the probability score of patients being identified as risk cases of pneumonia of unknown cause according to different classification scenarios based on the weight and combination relationship of each sign, and calculate and refresh the score results in real time according to the update of electronic medical record data. Step 2.3: The result output module compares the calculated probability score with a preset threshold; if the score exceeds the threshold, the patient is determined to be a high-risk case of pneumonia of unknown cause, and the corresponding judgment result is output; if the score does not exceed the threshold, the patient is determined to be a non-risk case. Step 3: Standardized Collection Process for Risk Case Information of Integrated Medical Treatment and Prevention: Step 3.1: When the algorithm of the rule model for identifying risk cases of pneumonia of unknown cause determines that a patient is a risk case, the reminder module set in the doctor's terminal plugin is activated to send a reminder to the doctor; Step 3.2: Following the prompts, the doctor collects patient characteristic information through the doctor's terminal plugin according to the data collection standard requirements. The collection results are normalized and standardized, and the background automatically matches and connects with existing data in the hospital management information system. After the doctor saves and uploads the data, it is transmitted to the disease control system platform, and an early warning reminder is sent. The staff of the disease control agency will then complete the form after conducting further epidemiological investigation. Step 3.3: The collected data is verified by the data verification module to ensure the accuracy and completeness of the data; Step 3.4: The validated data is stored in the risk case database as a dataset for subsequent analysis.
[0010] Furthermore, in step 1.1, the prompt message is: "You are a clinician with extensive experience in handling pneumonia cases; your task is to identify the following abnormal clinical signs and symptoms: [clinical signs or symptoms]; think step by step and provide your response in the following JSON format: {[clinical signs or symptoms]: ["Yes" or "No", "brief reason"]}; medical records: [RAG context]". The electronic medical record data includes various types of data such as patient basic information, symptom descriptions, examination and test reports, and medical records.
[0011] Furthermore, in step 2.2, the different classification scenarios include: severe pneumonia cases and severe pneumonia-prone cases in adults, severe pneumonia cases and severe pneumonia-prone cases in children, critical pneumonia cases, clustered pneumonia cases, and cases at risk of pneumonia caused by emerging rare pathogens.
[0012] Furthermore, in step 3.2, the collected patient characteristic information includes: confirmation of basic information, epidemiological history, dynamic changes in symptoms and test results, and recording of expert consultation opinions; the epidemiological investigation includes: recent travel history, contact history, activity trajectory, and contact information.
[0013] Furthermore, in step 3.3, the verification methods include data format checking, data range checking, and logical relationship checking; if there are problems with the data, it should be promptly reported to the doctor for re-collection or correction.
[0014] Furthermore, in step 3.1, the algorithm for the rule-based model for identifying risk cases of pneumonia of unknown cause is as follows: Feature label recognition based on regular expressions: Based on the standard data interface of collected medical information, a series of regular expression rules are designed and applied to quickly extract risk features from EMR data fields. These rules cover key risk indicators including fever, pneumonia imaging, shortness of breath, respiratory rate, arterial blood gas, and pathogen detection. Through pattern matching of regular expressions, cases that meet the risk definition can be quickly and accurately identified from the daily aggregated EMR data based on unstructured text fields including chief complaint, present medical history, and examination reports.
[0015] Furthermore, in step three, the plugins and disease control system platform enable authorized access to electronic medical records (EMR) and intelligent prompts across multiple scenarios at the doctor's workstation, supporting data exchange and closed-loop management between medical and preventive care. Its technical framework includes: Lightweight embedded deployment, based on WebSocket communication, requires no modification to the core business system. Relying on unified identity authentication and access management, it enables "one person, one file" access based on permissions, ensuring data security and privacy compliance. The contextualized intelligent prompt engine, with its built-in rule engine and knowledge base, covers various public health business areas. Real-time triggering prompts for business scenarios improve the efficiency of diagnosis and treatment and public health work. The system adopts an "event-driven + rule-matching + real-time response" triggering model to ensure accurate and timely prompts, including triggering mechanisms for operational behavior, inspection and check alerts, and cross-domain collaborative triggering.
[0016] Furthermore, in step one, the large language model sets the temperature parameter and processes the results of multiple runs. The temperature parameter is set to 0.3 to minimize illusions and maintain the consistency of text generation, ensuring the reliability of the recognition results. Multiple runs and taking the majority of results further improve the accuracy of recognition and assist in improving accuracy.
[0017] The beneficial effects of this invention are as follows: (1) Higher accuracy: Through in-depth mining of multi-source data of electronic medical records and artificial intelligence model algorithms based on large models, it is possible to make full use of various patient information and avoid the negative situation of keywords through semantic recognition. Compared with the simple keyword matching and primary statistical models of existing related technologies, it greatly improves the accuracy of identifying cases of pneumonia of unknown cause and reduces missed diagnosis and misdiagnosis.
[0018] (2) Efficiency improvement: It realizes automated identification, can quickly process a large amount of electronic medical record data, and give identification results in a short time, which saves valuable time for epidemic prevention and control. Manual screening and existing simple technologies are less efficient when processing large-scale data.
[0019] (3) Standardize information collection: The doctor-side plugin reminds doctors to collect information and provides standardized collection requirements to ensure the completeness and accuracy of the collected information. This helps to quickly determine the type of pneumonia in cases and potential epidemic risks, and provides strong data support for epidemic prevention and control. Attached Figure Description
[0020] Figure 1 This is a general example diagram illustrating a technique for identifying pneumonia risk cases using electronic medical record data. Detailed Implementation
[0021] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0022] Terminology Explanation: 1. Cases of pneumonia of unknown cause at risk: A case is considered at risk if it meets one of the following three criteria: 1) Cases of severe (critical) pneumonia with a high suspicion of infection, rapid disease progression, and ineffective empirical treatment, and which are ruled out as being caused by common respiratory infectious disease pathogens through routine diagnosis and treatment, or cannot be explained by common pathogen infection. 2) Within two weeks, two or more cases of pneumonia with a suspected epidemiological history were found, and infection by common respiratory infectious disease pathogens was ruled out; 3) Suspected new or rare pathogens or novel pathogens were detected in pneumonia case specimens.
[0023] 2. Large Language Model (LLM): A deep learning-based language processing model that, after being trained on a large amount of text, can understand and generate natural language. In this invention, it is used to extract clinical signs and symptoms from electronic medical records.
[0024] 3. Retrieval-Enhanced Generation (RAG) Technology: This technology integrates information retrieval and language generation techniques. By retrieving relevant information, it assists the language model in generating more accurate text, thus solving the problem of language model limitations due to the extremely large input length of clinical records.
[0025] 4. MCP (Model Context Protocol) technology: A model context protocol technology that integrates feature and semantic information from multiple data sources to provide richer and more accurate background knowledge for the model by constructing a unified context framework. In this invention, MCP technology is used to establish a connection between the LLM and the data source, enabling it to obtain the latest data in real time, thereby improving service capabilities and response accuracy.
[0026] 5. Temperature parameter: A parameter that controls the randomness of output during text generation in a large language model. The lower the value, the more stable and deterministic the output. This invention sets a temperature parameter to balance the accuracy and stability of text generation.
[0027] 6. JSON format: A lightweight data exchange format that facilitates data storage, transmission, and parsing. This invention is used for the output of a large language model clinical sign or symptom extractor.
[0028] 7. Rule-based model algorithm for identifying high-risk cases of pneumonia of unknown origin: A rule-based algorithm that calculates the probability of a patient being a high-risk case of pneumonia of unknown origin based on structured clinical signs and symptoms.
[0029] 8. Doctor-side clinic risk alert and information collection program: A small application installed on the front end of the electronic medical record system used by doctors. It is used to remind doctors to collect risk case information according to established rules and to carry out data collection work in accordance with standard specifications.
[0030] A method for identifying and collecting information on high-risk cases of pneumonia of unknown cause, including: 1. Steps of a clinical sign or symptom extractor based on a large language model: Step 1: The input module receives prompts and electronic medical record (EMR) data. The prompt is: "You are a clinician with extensive experience in handling pneumonia cases. Your task is to identify the following abnormal clinical signs and symptoms: [clinical signs or symptoms]. Think step by step and provide your response in the following JSON format: {[clinical signs or symptoms]: ["Yes" or "No", "brief reason (original medical record data, etc.)"]}. Medical records: [RAG context]". The EMR data includes various types of data such as basic patient information, symptom descriptions, examination and test reports, and medical records. Simultaneously, a connection is established with the EMR data source through MCP technology to obtain the latest data related to the patient in real time, such as recent test results and new medical records. This data is then integrated with existing EMR data to form more comprehensive input data for subsequent modules.
[0031] Step 2: Since clinical records may exceed the predefined input length of the large language model, the RAG module is activated. This module retrieves relevant text information from the integrated multi-source electronic medical record data, generates text blocks suitable for processing by the large language model, and appends them to the input prompts. During this process, MCP technology leverages its connection to external resources to provide the RAG module with a more precise search scope and data support, helping it to more efficiently filter and extract key information, thus improving search efficiency and accuracy.
[0032] Step 3: The large language model module receives the processed input prompts, runs it three times, and each time identifies clinical signs or symptoms based on the input information and generates corresponding JSON format text output.
[0033] Step 4: The results processing module collects the output results of multiple runs of the large language model and uses majority voting to determine the final clinical signs or symptoms identification results for use by subsequent modules.
[0034] 2. Algorithm steps for rule-based model of identifying risk cases of pneumonia of unknown cause: Step 1: The data receiving module obtains the recognition results output by the clinical sign or symptom extractor based on the large language model, which is JSON data containing the presence or absence of various clinical signs or symptoms and the corresponding reasons.
[0035] Step 2: The algorithm calculation module dynamically analyzes and calculates the acquired clinical signs, symptoms, and laboratory test data according to pre-set risk case identification rules. For example, if multiple signs related to pneumonia of unknown origin, such as fever, shortness of breath, and abnormal lung imaging, appear simultaneously, the algorithm will calculate the probability score of the patient being identified as a risk case of pneumonia of unknown origin based on the weight and combination relationship of each sign, categorizing the risk cases of pneumonia of unknown origin into different classifications such as severe and severe pneumonia cases in adults, severe and severe pneumonia cases in children, critical pneumonia cases, clustered pneumonia cases, and pneumonia cases with emerging rare pathogens. The algorithm also updates the score results in real time based on the updates to the electronic medical record data.
[0036] Step 3: The results output module compares the calculated probability score with a preset threshold. If the score exceeds the threshold, the patient is determined to be a high-risk case of pneumonia of unknown cause, and the corresponding judgment result is output; if the score does not exceed the threshold, the patient is determined to be a non-risk case.
[0037] 3. Standardized steps for collecting information on risk cases involving the integration of medical treatment and disease prevention: Step 1: When the algorithm of the rule model for identifying high-risk cases of pneumonia of unknown cause determines that a patient is a high-risk case, the reminder module (doctor-side plugin) is activated to send a reminder to the doctor.
[0038] Step 2: Following the prompts, the doctor uses the doctor-side plugin to collect patient characteristic information according to the data collection standards. Address information, workplace name, and test results are normalized and standardized. Existing data from the HIS (Hospital Information System) is automatically matched and integrated in the background. This information includes confirmation of basic information, epidemiological history, dynamic changes in symptoms and test results (such as respiratory-related clinical manifestations like fever, cough, and sore throat; medication use; blood routine tests, chest CT scan follow-up results, etc.), and records expert consultation opinions. After the doctor saves and uploads the data, it is transmitted to the disease control system platform, and a warning SMS reminder is sent via SMS or email. Disease control personnel then conduct further epidemiological investigations and complete the necessary forms, including recent travel history, contact history, activity trajectory, and contact information.
[0039] Step 3: The collected data is verified by the data validation module to ensure its accuracy and completeness. Validation methods include data format checking, data range checking, and logical relationship checking. If any problems are found with the data, it is promptly reported to the doctor for re-collection or correction.
[0040] Step 4: The validated data is stored in a professional high-risk case database as a high-quality dataset for subsequent clinical diagnosis, epidemic prevention and control analysis, etc.
[0041] Data collection standards require: Epidemiological history: including accurate records of activity trajectory within 14 days prior to the onset of illness, specifying the cities or regions visited; detailed records of contact history with confirmed or suspected cases, including the time, manner, and symptoms of both parties at the time of contact; records of crowded places visited and the duration of stay.
[0042] Dynamic changes in symptoms: Record the onset time of fever, fever pattern (continuous fever, remittent fever, etc.), highest body temperature and trend of body temperature changes; onset time, frequency (e.g., number of coughs per day), and nature (dry cough, sputum; if sputum is coughed up, record the color, characteristics, and amount of sputum); degree of dyspnea (mild, moderate, severe, judged based on the patient's subjective feelings and objective manifestations) and time of onset.
[0043] Examination and test results: Track the dynamic changes of indicators such as white blood cell count, lymphocyte count, and neutrophil count in routine blood tests; record the time and imaging manifestations of chest imaging examinations (changes in the location, shape, and size of lung shadows); monitor changes in inflammatory markers such as C-reactive protein and procalcitonin.
[0044] For competitors to achieve the same goal, the fundamental processes that are difficult to bypass include: extracting clinical signs and symptoms based on a large language model combined with RAG and MCP technologies; calculating risk scores and identifying high-risk cases using a professionally considered public health risk case identification rule model algorithm; and transmitting data in clinical and disease control application scenarios through specialized plug-ins and information platform systems, while reminding staff to collect risk case information in a standardized manner. Among these, the most crucial steps are the information extraction using a large language model combined with RAG and MCP technologies, and the calculation of risk scores using a case identification rule model algorithm defined from a public health risk perspective. These two steps are key to accurately identifying high-risk cases and are difficult to replace.
[0045] This invention extracts clinical signs and symptoms by combining a large language model with RAG and MCP technologies: RAG technology solves the input length problem, MCP technology provides real-time multi-source data support for LLM, and the large language model uses its powerful text understanding and generation capabilities to accurately extract key information from electronic medical records. This is the key difference between this invention and traditional methods.
[0046] The core step involves using a specific risk case identification rule model algorithm to calculate risk scores and determine risky cases. This step calculates the probability score of a risky case based on preset calculation rules, combined with clinical signs, symptom data, and laboratory test results. This step directly determines the accuracy of the identification results. Through scientific and reasonable algorithm design, multiple factors can be comprehensively considered to accurately assess the patient's risk level. The specific risk case identification rule model algorithm is as follows: Feature Tag Recognition Based on Regular Expressions: Based on the standard data interface for medical information collected by the national pre-processing software, a series of regular expression rules were designed and applied to quickly extract risk features from over 600 EMR data fields. These rules cover key risk indicators such as fever, pneumonia imaging, shortness of breath, respiratory rate, arterial blood gas, and pathogen detection (e.g., the chief complaint field contains "fever" and does not contain the words "no fever," "no fever," or "no fever," "respiratory rate ≥30 breaths / min," "oxygen saturation ≤93%," etc.). Through pattern matching of regular expressions, cases meeting the risk definition can be quickly and accurately identified from millions of EMR data points collected daily, based on unstructured text fields such as chief complaint, present medical history, and examination reports.
[0047] A dedicated plugin alerts doctors and standardizes the collection of information for high-risk cases. This ensures that key information is collected promptly and systematically after identifying high-risk cases, and uses normalization techniques to address the difficulty in standardizing the identification and processing of address information and test results, providing strong support for subsequent diagnosis and prevention efforts. The specific plugin and information platform system are as follows: The outpatient intelligent plugin, relying on the city-level provincial integrated platform, enables authorized access to electronic medical records (EMR) and intelligent prompts in multiple scenarios at the doctor's workstation, supporting the interoperability and closed-loop management of medical and preventive data. Its main technical framework includes: 1. Lightweight embedded deployment, based on WebSocket communication, requires no modification to the core business system, has strong adaptability and fast deployment, reduces the connection cost for medical institutions, and relies on unified identity authentication and permission management to realize "one person, one file" access according to permissions, ensuring data security and privacy compliance; 2. Contextualized intelligent prompt engine with built-in rule engine and knowledge base, covering various public health business areas, and real-time triggered prompts for business scenarios to improve the efficiency of diagnosis and treatment and public health work; 3. Adopt an "event-driven + rule matching + real-time response" triggering mode to ensure accurate and timely prompts, including triggering mechanisms such as operation behavior triggering, inspection and check warnings, inspection and check warnings, and cross-domain collaborative triggering.
[0048] The auxiliary steps include setting the temperature parameters of the language model, processing the results of multiple runs, acquiring data through the data receiving module, outputting data through the result output module, verifying data through the data verification module, and storing data through the data storage module.
[0049] Setting the temperature parameter of the large language model and processing the results of multiple runs: The temperature parameter is set to 0.3 to minimize illusions and maintain the consistency of text generation, ensuring the reliability of recognition results; multiple runs and taking the majority of results further improve the accuracy of recognition, but multiple runs of the model are not necessary for every data processing, and are an auxiliary operation to improve accuracy.
[0050] Data receiving module and result output module: The data receiving module is responsible for acquiring the data required for identification and providing support for core calculations; the result output module formats the calculation results for easy subsequent use, but they themselves do not participate in the core risk assessment calculation process.
[0051] Data verification module and data storage module: The data verification module is used to ensure the quality of the collected data. Although it does not directly participate in the information collection process, it is crucial to the reliability of the data. The data storage module is responsible for storing the verified data in an orderly manner to facilitate subsequent querying and use, but it does not involve information collection or key processing itself.
[0052] The invention adds a step to process the input length using RAG technology: Compared to simply using a large language model for recognition (assuming the basic process), this invention adds a step to process the input length using RAG technology. This is because clinical records often exceed the predefined input length of a large language model. Without processing, the large language model cannot fully acquire the information, which will seriously affect the accuracy of the recognition results.
[0053] Adding the step of connecting to external data sources using MCP technology: Compared to the process without MCP technology, this step is added to address the issue of LLMs being unable to access external data sources, enabling LLMs to obtain the latest data in real time, improving their service capabilities and response accuracy, and providing more comprehensive and timely data support for the extraction of clinical signs and symptoms.
[0054] Adding a majority vote operation for multiple runs of the large language model reduces the uncertainty of model output, improves recognition reliability, and avoids incorrect judgments caused by random errors in a single run.
[0055] The invention adds a multi-indicator comprehensive calculation step: Compared with the basic process of simply judging the risk of cases based on a single symptom, this invention adds a multi-indicator comprehensive calculation step. Risk assessment of pneumonia of unknown cause requires comprehensive consideration of multiple clinical signs and symptoms, as well as laboratory test results; judgment based on a single symptom cannot take into account the comprehensive risk.
[0056] Add a comparison step with preset thresholds: Define clear judgment criteria to make the identification results more objective and operable. Update risk scores dynamically by tracking the latest diagnosis and treatment activities in real time through updates to HIS data (such as test results).
[0057] The invention adds standardized information collection steps: Compared to the basic process of only identifying high-risk cases without standardized information collection, this invention adds reminder collection, data verification, and data storage steps. Reminder collection ensures that doctors do not miss any information collection tasks; data verification ensures the usability of the collected data and avoids erroneous data misleading subsequent work; data storage facilitates long-term tracking and analysis of high-risk case data, providing data support for the formulation of epidemic prevention and control strategies.
[0058] (1) Algorithm Alternatives: In terms of model algorithms, besides building artificial intelligence models based on large models, other types of machine learning models can be tried, such as support vector machines and random forests. These models may have unique advantages when processing certain types of data. Compared with algorithms based on large models, support vector machines may be more computationally efficient when processing small samples and nonlinear data, but they have higher requirements for data feature engineering, and their generalization ability when processing large-scale complex electronic medical record data may not be as good as algorithms based on large models. Random forests have better noise resistance and interpretability, but they may not be as good as deep learning algorithms based on large models in mining complex relationships in electronic medical record data.
[0059] (2) Alternative data collection solutions: In the data collection stage, in addition to obtaining electronic medical record data from the hospital information system, data from patients' wearable devices, such as physiological data like body temperature and heart rate recorded by smart bracelets, can also be considered. Compared with relying solely on electronic medical record data, adding wearable device data can provide patients with more real-time and continuous health information, but this requires solving the problems of data synchronization and integration, and may involve challenges such as patient privacy protection.
[0060] Key technical points: 1. Application of large language model based on RAG and MCP technologies: RAG technology is used to solve the problem of clinical record input length, and MCP technology provides real-time multi-source data for large language model. Combined with the powerful text processing capabilities of large language model, it can accurately extract clinical signs and symptoms, which is the key to improving the recognition accuracy of this invention and is of the highest importance.
[0061] 2. Algorithm for identifying risk cases of pneumonia of unknown cause: Through scientific and reasonable algorithm design, the probability score of risk cases is calculated by comprehensively considering multiple signs, providing an objective basis for judging risk cases, and its importance ranks second.
[0062] 3. Doctor-side plugin and standardized information collection: The plugin reminds doctors and standardizes the information collection process and standards, and transmits the data to the risk pneumonia module of the disease control information platform. Through medical and prevention collaboration, it ensures the efficient collection of professional data on risk cases, providing data support for subsequent diagnosis and prevention. Its importance ranks third.
[0063] The above-described embodiments are merely one implementation of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention should be determined by the appended claims.
Claims
1. A method for identifying and collecting information on high-risk cases of pneumonia of unknown cause, characterized in that, include: Step 1: Clinical sign or symptom extractor workflow based on large language model: Step 1.1: The input module receives prompts and electronic medical record data; at the same time, it establishes a connection with the electronic medical record data source through MCP technology, obtains the latest data related to the patient in real time, and integrates this data with existing electronic medical record data to form more comprehensive input data for use by subsequent modules; Step 1.2: The RAG module starts. This module retrieves relevant text information from the integrated multi-source electronic medical record data, generates text blocks suitable for processing by a large language model, and appends them to the input prompts. Step 1.3: The large language model module receives the processed input data prompts, runs it multiple times, and each time identifies clinical signs or symptoms based on the input information and generates corresponding text output; Step 1.4: The results processing module collects the output results of multiple runs of the large language model and uses majority voting to determine the final clinical signs or symptoms identification results for use by subsequent modules; Step 2: Algorithm flow for rule-based model of risk cases of pneumonia of unknown cause: Step 2.1: The data receiving module acquires the recognition results output by the clinical sign or symptom extractor based on the large language model; Step 2.2: The algorithm calculation module dynamically analyzes and calculates the acquired clinical signs, symptoms and laboratory test data according to the pre-set risk case identification rules; the algorithm will calculate the probability score of patients being identified as risk cases of pneumonia of unknown cause according to different classification scenarios based on the weight and combination relationship of each sign, and calculate and refresh the score results in real time according to the update of electronic medical record data. Step 2.3: The result output module compares the calculated probability score with a preset threshold. If the score exceeds the threshold, the patient is determined to be a case of pneumonia of unknown cause at risk, and the corresponding judgment result is output. If the score does not exceed the threshold, it is determined to be a non-risk case; Step 3: Standardized Collection Process for Risk Case Information of Integrated Medical Treatment and Prevention: Step 3.1: When the algorithm of the rule model for identifying risk cases of pneumonia of unknown cause determines that a patient is a risk case, the reminder module set in the doctor's terminal plugin is activated to send a reminder to the doctor; Step 3.2: Following the prompts, the doctor collects patient characteristic information through the doctor's terminal plugin according to the data collection standard requirements. The collection results are normalized and standardized, and the background automatically matches and connects with existing data in the hospital management information system. After the doctor saves and uploads the data, it is transmitted to the disease control system platform, and an early warning reminder is sent. The staff of the disease control agency will then complete the form after conducting further epidemiological investigation. Step 3.3: The collected data is verified by the data verification module to ensure the accuracy and completeness of the data; Step 3.4: The validated data is stored in the risk case database as a dataset for subsequent analysis.
2. The method for identifying and collecting information on high-risk cases of pneumonia of unknown cause according to claim 1, characterized in that, In step 1.1, the prompt message is: "You are a clinician with extensive experience in handling pneumonia cases; your task is to identify the following abnormal clinical signs and symptoms: [clinical signs or symptoms]; think step by step and provide your response in the following JSON format: {[clinical signs or symptoms]: ["Yes" or "No", "brief reason"]}; medical records: [RAG context]". The electronic medical record data includes various types of data such as basic patient information, symptom descriptions, examination and test reports, and medical records.
3. The method for identifying and collecting information on high-risk cases of pneumonia of unknown cause according to claim 1, characterized in that, In step 2.2, the different classification scenarios include: severe pneumonia cases and severe pneumonia-prone cases in adults, severe pneumonia cases and severe pneumonia-prone cases in children, critical pneumonia cases, clustered pneumonia cases, and cases at risk of pneumonia caused by emerging rare pathogens.
4. The method for identifying and collecting information on high-risk cases of pneumonia of unknown cause according to claim 1, characterized in that, In step 3.2, the patient characteristic information collected includes: confirmation of basic information, epidemiological history, dynamic changes in symptoms and test results, and recording of expert consultation opinions; the epidemiological investigation includes: recent travel history, contact history, activity trajectory, and contact information.
5. The method for identifying and collecting information on high-risk cases of pneumonia of unknown cause according to claim 1, characterized in that, In step 3.3, the verification methods include data format checking, data range checking, and logical relationship checking; if there are problems with the data, it should be promptly reported to the doctor for re-collection or correction.
6. The method for identifying and collecting information on high-risk cases of pneumonia of unknown cause according to claim 1, characterized in that, In step 3.1, the algorithm for the rule-based model for identifying cases of pneumonia of unknown origin is as follows: Feature label recognition based on regular expressions: Based on the standard data interface of collected medical information, a series of regular expression rules are designed and applied to quickly extract risk features from EMR data fields; These rules cover key risk indicators including fever, pneumonia imaging, shortness of breath, respiratory rate, arterial blood gas, and pathogen detection; through pattern matching of regular expressions, cases that meet the risk definition can be quickly and accurately identified from the daily aggregated EMR data based on unstructured text fields including chief complaint, present medical history, and examination reports.
7. The method for identifying and collecting information on high-risk cases of pneumonia of unknown cause according to claim 1, characterized in that, In step three, the plugins and disease control system platform enable authorized access to electronic medical records (EMR) and intelligent prompts across multiple scenarios at the doctor's workstation, supporting data exchange and closed-loop management between medical and preventive healthcare. Its technical framework includes: Lightweight embedded deployment, based on WebSocket communication, requires no modification to the core business system. Relying on unified identity authentication and access management, it enables "one person, one file" access based on permissions, ensuring data security and privacy compliance. The contextualized intelligent prompt engine, with its built-in rule engine and knowledge base, covers various public health business areas. Real-time triggering prompts for business scenarios improve the efficiency of diagnosis and treatment and public health work. The system adopts an "event-driven + rule-matching + real-time response" triggering model to ensure accurate and timely prompts, including triggering mechanisms for operation behavior, inspection and check warnings, and cross-domain collaborative triggering.
8. The method for identifying and collecting information on high-risk cases of pneumonia of unknown cause according to claim 1, characterized in that, In step one, the large language model sets the temperature parameter and processes the results of multiple runs. The temperature parameter is set to 0.3 to minimize illusions and maintain the consistency of text generation, ensuring the reliability of the recognition results. Multiple runs and the majority result are taken to further improve the accuracy of recognition and assist in improving accuracy.