A method and system for retrieving auxiliary diagnostic information based on a respiratory knowledge base

CN122817293APending Publication Date: 2026-09-25JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611278320.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

现有的AI医疗诊断系统,大多依赖于大量数据的机器学习,存在“黑盒”决策、缺乏可解释性、对罕见病的辅助诊断能力弱等问题

Benefits of technology

本发明能提供可解释的辅助诊断信息,通过提供诊断信息依据和可解释性分析,增强辅助诊断信息的准确、高效和可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817293A_ABST
    Figure CN122817293A_ABST
Patent Text Reader

Abstract

The application provides a kind of based on respiratory knowledge base auxiliary diagnosis information retrieval method and system, belong to artificial intelligence and respiratory system disease cross technical field. Including the following steps: the respiratory system related data of patient is collected;Different data types collected are preprocessed using different preprocessing methods;Based on the respiratory system related data of patient, relevant diagnostic information is retrieved in medical knowledge base;The relevant entries in medical knowledge base are preferentially retrieved in the retrieval process, if relevant entries are not found in medical knowledge base, then call transformer large model to assist in reasoning of diagnostic information, otherwise return the auxiliary diagnostic information of medical knowledge base;The application provides the basis and explainability analysis of auxiliary diagnostic information. The application improves the accuracy and reliability of auxiliary diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of artificial intelligence and respiratory diseases, specifically relating to an auxiliary diagnostic information retrieval method and system based on a respiratory knowledge base. Background Technology

[0002] Chronic obstructive pulmonary disease (COPD) is a common chronic airway disease with a high global prevalence and a significant impact on patients' quality of life. Traditional COPD diagnosis relies on pulmonary function tests, imaging studies, and clinical symptoms, which suffers from diagnostic delays, high misdiagnosis rates, and insufficient individualized treatment. Particularly in cold regions, long-term exposure to low temperatures is closely associated with COPD incidence, but current diagnostic methods struggle to effectively integrate environmental exposure factors, leading to poor treatment outcomes. Therefore, providing an accurate and reliable auxiliary diagnostic method is urgently needed.

[0003] The prior art represented by the aforementioned documents has at least the following unresolved technical problems or defects: Most existing AI-based medical diagnostic systems rely on machine learning with massive amounts of data, resulting in problems such as "black box" decision-making, lack of interpretability, and weak ability to assist in the diagnosis of rare diseases. These systems often fail to effectively utilize existing medical knowledge and guidelines, leading to low accuracy and reliability in assisted diagnosis. Summary of the Invention

[0004] The purpose of this invention is to provide a method for retrieving auxiliary diagnostic information based on a respiratory knowledge base, thereby improving the accuracy and reliability of auxiliary diagnosis.

[0005] To address the aforementioned technical problems, this invention provides a method for retrieving auxiliary diagnostic information based on a respiratory knowledge base, comprising the following steps: Step S1: Data Acquisition: Collect respiratory system-related data from the patient; Step S2: Data preprocessing: Different preprocessing methods are used to preprocess different types of collected data; Step S3: Assisted Diagnostic Information Retrieval: Based on the patient's respiratory system-related data, retrieve relevant diagnostic information from the medical knowledge base; the retrieval process prioritizes searching for relevant entries in the medical knowledge base. If no relevant entries are found in the medical knowledge base, the transformer large model is invoked to perform assisted diagnostic information reasoning; otherwise, the assisted diagnostic information from the medical knowledge base is returned. The aforementioned transformer large model supports both structured and unstructured medical knowledge bases. It is trained based on massive amounts of data and possesses an understanding of the pathological mechanisms, diagnostic criteria, and treatment strategies of respiratory diseases. It decomposes problems through CoT chain reasoning and ToT tree reasoning, and can generate verifiable reasoning paths based on the patient's respiratory system-related data through the RAG tracing mechanism combined with the medical knowledge base to perform auxiliary diagnostic information reasoning. Otherwise, it returns auxiliary diagnostic information from the medical knowledge base. The auxiliary diagnostic information retrieval includes the following steps: Query construction: Transforming patients' respiratory system-related data into query statements; Query decomposition: The query statement is broken down into several simpler query statements; Retrieval Execution: Based on the decomposed query statement, the query is executed in the medical knowledge base. A knowledge graph is constructed based on the medical knowledge base. During retrieval, a hierarchical traversal is performed, starting from the root node and traversing the knowledge graph layer by layer. The node and relationship with the highest similarity to the query are selected and the query results are returned. Multiple query results are obtained for each decomposed query statement. Result fusion: Merge multiple query results based on similarity to obtain a fused result set of multiple query statements; Result sorting: Sort the query results in the fused result set and select the results most relevant to the query as the retrieved auxiliary diagnostic information.

[0006] Preferably, the data types include image data, text data, and numerical data.

[0007] Preferably, the preprocessing of the image data includes at least one of the following steps, selected and combined according to image quality and diagnostic objectives: Format conversion: Converting image data from different formats into a unified format; Image denoising: using filtering algorithms to remove noise from an image; Median filtering: For each pixel in an image, replace it with the median value of its neighboring pixels; Gaussian filtering: Smooths the image by applying a Gaussian kernel and performing convolution. Image enhancement: Adjusts the contrast and brightness of an image to improve its clarity; Histogram equalization: Adjusts the histogram of an image to make the image contrast more uniform; CLAHE: Performs local histogram equalization on an image to improve local contrast. Image segmentation: dividing an image into different regions; Thresholding segmentation: Dividing an image into different regions based on the grayscale values ​​of pixels; Region growing: Starting from the seed point, adjacent pixels are merged into the same region based on pixel similarity; Deep learning segmentation: Image segmentation using deep learning models; Feature extraction: Extracting features from an image; Gray-level co-occurrence moments: used to extract texture features from an image; HOG: Used to extract edge features from an image; Deep learning features: using deep learning models to extract features from images.

[0008] Preferably, the preprocessing of the text data includes the following steps: Text cleaning: Remove special characters, HTML tags, and punctuation marks from text; Word segmentation: dividing text into multiple words; Remove stop words: Remove words that have been stopped from the text; Text vectorization: Converting text into a vector representation.

[0009] Preferably, the preprocessing of the text data further includes at least one of the following steps: Part-of-speech tagging: Tagging the part of speech of each word; Named entity recognition: Identifying named entities in text data.

[0010] Preferably, the preprocessing of the numerical data includes the following steps: Missing value handling: Fill or delete missing values ​​in numerical data, using the mean, median, or interpolation to fill missing values; Outlier handling: Detecting and handling outliers in numerical data, using box plots or Z-scores to detect outliers; Data standardization: scaling data to the same range.

[0011] Preferably, in the step of assisting in diagnostic information retrieval, the knowledge base retrieval includes knowledge graph retrieval, which first undergoes preprocessing, including the following steps: Embedding vectorization: Convert keywords and entities in the knowledge graph, as well as the relationships between entities, into vector representations to generate embedding vectors, including keyword vectors and entity vectors; Similarity calculation: Calculate the similarity between keyword vectors and entity vectors in the knowledge graph.

[0012] Preferably, the similarity is expressed as cosine similarity: similarity(A, B) = (A·B) / (||A||) ||B||).

[0013] Preferably, the auxiliary diagnostic information retrieval includes the following steps: Query construction: Transforming patients' respiratory system-related data into query statements; Query decomposition: The query statement is broken down into several simpler query statements; Retrieval Execution: The decomposed query is executed in the medical knowledge base. During retrieval, a hierarchical traversal is performed, starting from the root node and traversing the knowledge graph layer by layer. The node and relationship with the highest similarity to the query are selected and the query results are returned. Multiple query results are obtained for each decomposed query statement. Result fusion: Merge multiple query results based on similarity to obtain a fused result set of multiple query statements; Result sorting: Sort the query results in the fused result set and select the most relevant results as the retrieved auxiliary diagnostic information.

[0014] Preferably, in the step of assisting in the diagnostic information retrieval, the query construction step includes the following steps: using a pre-trained NER model to identify entities from patient data, using a pre-trained RE model to extract the relationships between entities, and generating a query statement based on the entities and relationships; In the steps of assisting in the diagnostic information retrieval, the query decomposition step includes the following steps: using a syntax analyzer to parse the query statement, identify the query target and constraints, and decompose the query according to the query target and constraints.

[0015] Preferably, in the step of assisting in the diagnostic information retrieval, the retrieval execution step includes the following steps: for retrieval of structured medical knowledge bases, performing a query in the RDF triple database using the SPARQL query language; for retrieval of unstructured medical knowledge bases, performing a query using a vector similarity search algorithm. The result fusion step includes the following steps: Calculate the correlation and cluster the results based on the correlation, filtering out duplicate or irrelevant results.

[0016] The present invention also provides an auxiliary diagnostic information retrieval system based on a respiratory knowledge base, which uses the above-mentioned auxiliary diagnostic information retrieval method based on a respiratory knowledge base, including a data acquisition module, a data preprocessing module, and a knowledge base retrieval module; The data acquisition module is used to collect respiratory system-related data from patients; The data preprocessing module is used to preprocess different data types using different preprocessing methods, such as cleaning, standardizing, and format conversion of the collected data. The knowledge base retrieval module is used to retrieve relevant diagnostic information from the medical knowledge base based on the patient's respiratory system-related data. The retrieval process prioritizes searching for relevant entries in the medical knowledge base. If no relevant entries are found in the medical knowledge base, the large model is invoked to perform auxiliary diagnostic information reasoning, and the knowledge base retrieval process described in step S1 is executed. Otherwise, the auxiliary diagnostic information from the medical knowledge base is returned.

[0017] Compared with the prior art, the present invention has the following beneficial effects: This invention can provide interpretable auxiliary diagnostic information, and enhances the accuracy, efficiency and reliability of auxiliary diagnostic information by providing diagnostic information basis and interpretability analysis. Attached Figure Description

[0018] Figure 1 This is a flowchart of an auxiliary diagnostic information retrieval method based on a respiratory knowledge base, according to an embodiment of the present invention.

[0019] Figure 2 This is a schematic diagram of the sorting process according to an embodiment of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] To better understand the purpose and function of this invention, the invention will be described in further detail below with reference to the accompanying drawings.

[0022] Example 1 like Figure 1 As shown, the auxiliary diagnostic information retrieval method based on a respiratory knowledge base of the present invention will be described in detail below with reference to a specific embodiment of the present invention.

[0023] This invention provides a method for retrieving auxiliary diagnostic information based on a respiratory knowledge base, comprising the following steps: Step S1: Data Acquisition: Collect respiratory system-related data from the patient; Step S2: Data preprocessing: Different preprocessing methods are used to preprocess different types of collected data. After verification, step S3 is executed. Step S3: Assisted Diagnostic Information Retrieval: Based on the patient's respiratory system-related data, retrieve relevant diagnostic information from the medical knowledge base; the retrieval process prioritizes searching for relevant entries in the medical knowledge base. If no relevant entries are found in the medical knowledge base, the transformer large model is invoked to perform assisted diagnostic information reasoning; otherwise, the assisted diagnostic information from the medical knowledge base is returned. The aforementioned large-scale transformer model supports both structured and unstructured medical knowledge bases. Trained on massive amounts of data, it possesses an understanding of the pathological mechanisms, diagnostic criteria, and treatment strategies of respiratory diseases. Through CoT chain reasoning and ToT tree reasoning, it decomposes problems and can generate verifiable reasoning paths based on the patient's respiratory system-related data through the RAG tracing mechanism combined with the medical knowledge base. This allows for auxiliary diagnostic information reasoning; otherwise, it returns auxiliary diagnostic information from the medical knowledge base. Finally, it generates an interpretable report and outputs it to the doctor. The auxiliary diagnostic information retrieval includes the following steps: Query construction: Transforming patients' respiratory system-related data into query statements; Query decomposition: The query statement is broken down into several simpler query statements; Retrieval Execution: Based on the decomposed query statement, the query is executed in the medical knowledge base. A knowledge graph is constructed based on the medical knowledge base. During retrieval, a hierarchical traversal is performed, starting from the root node and traversing the knowledge graph layer by layer. The node and relationship with the highest similarity to the query are selected and the query results are returned. Multiple query results are obtained for each decomposed query statement. Result fusion: Merge multiple query results based on similarity to obtain a fused result set of multiple query statements; Result sorting: Sort the query results in the fused result set and select the results most relevant to the query as the retrieved auxiliary diagnostic information.

[0024] In this embodiment, the data types include image data, text data, and numerical data.

[0025] In this embodiment, the preprocessing of the image data includes at least one of the following steps, selected and combined according to image quality and diagnostic objectives: Format conversion: Converting image data from different formats into a unified format; Image denoising: using filtering algorithms to remove noise from an image; Median filtering: For each pixel in an image, replace it with the median value of its neighboring pixels; Gaussian filtering: Smooths the image by applying a Gaussian kernel and performing convolution. Image enhancement: Adjusts the contrast and brightness of an image to improve its clarity; Histogram equalization: Adjusts the histogram of an image to make the image contrast more uniform; CLAHE: Performs local histogram equalization on an image to improve local contrast. Image segmentation: dividing an image into different regions; Thresholding segmentation: Dividing an image into different regions based on the grayscale values ​​of pixels; Region growing: Starting from the seed point, adjacent pixels are merged into the same region based on pixel similarity; Deep learning segmentation: Image segmentation using deep learning models; Feature extraction: Extracting features from an image; Gray-level co-occurrence moments: used to extract texture features from an image; HOG: Used to extract edge features from an image; Deep learning features: using deep learning models to extract features from images.

[0026] In this embodiment, the preprocessing of the text data includes the following steps: Text cleaning: Remove special characters, HTML tags, and punctuation marks from text; Word segmentation: dividing text into multiple words; Remove stop words: Remove words that have been stopped from the text; Text vectorization: converting text into a vector representation; Part-of-speech tagging: Tagging the part of speech of each word; Named entity recognition: Identifying named entities in text data; Text cleaning, word segmentation, stop word removal, and text vectorization are mandatory steps, while the remaining steps are optional in this embodiment.

[0027] In this embodiment, the preprocessing of the numerical data includes the following steps: Missing value handling: Fill or delete missing values ​​in numerical data, using the mean, median, or interpolation to fill missing values; Outlier handling: Detect and handle outliers in numerical data. Use box plots or Z-scores to detect outliers. For example, if the CT slice thickness is >1mm, issue an underestimation warning. Data standardization: scaling data to the same range.

[0028] In this embodiment, the step of retrieving auxiliary diagnostic information includes knowledge base retrieval, which involves preprocessing and includes the following steps: Embedding vectorization: Convert keywords and entities in the knowledge graph, as well as the relationships between entities, into vector representations to generate embedding vectors, including keyword vectors and entity vectors; Similarity calculation: Calculate the similarity between keyword vectors and entity vectors in the knowledge graph.

[0029] In this embodiment, the similarity is calculated using cosine similarity: similarity(A, B) = (A·B) / (||A|| ||B||).

[0030] In this embodiment, the auxiliary diagnostic information retrieval includes the following steps: Query construction: Transforming patients' respiratory system-related data into query statements; Query decomposition: The query statement is broken down into several simpler query statements; Retrieval Execution: The decomposed query is executed in the medical knowledge base. During retrieval, a hierarchical traversal is performed, starting from the root node and traversing the knowledge graph layer by layer. The node and relationship with the highest similarity to the query are selected and the query results are returned. Multiple query results are obtained for each decomposed query statement. Result fusion: Merge multiple query results based on similarity to obtain a fused result set of multiple query statements; Result sorting: Sort the query results in the fused result set and select the most relevant results as the retrieved auxiliary diagnostic information.

[0031] In this embodiment, the query construction step in the auxiliary diagnostic information retrieval step includes the following steps: using a pre-trained NER model to identify entities from patient data, using a pre-trained RE model to extract the relationships between entities, and generating query statements based on entities and relationships; In the steps of assisting in the diagnostic information retrieval, the query decomposition step includes the following steps: using a syntax analyzer to parse the query statement, identify the query target and constraints, and decompose the query according to the query target and constraints.

[0032] In this embodiment, the retrieval execution step of the auxiliary diagnostic information retrieval step includes the following steps: for structured medical knowledge base retrieval, the SPARQL query language is used to perform a query in the RDF triple database; for unstructured medical knowledge base retrieval, the vector similarity search algorithm is used to perform a query. The result fusion step includes the following steps: Calculate the correlation and cluster the results based on the correlation, filtering out duplicate or irrelevant results.

[0033] Example 2 According to a specific embodiment of the present invention, a respiratory knowledge base-based auxiliary diagnostic information retrieval system of the present invention will be described in detail below, wherein the respiratory knowledge base-based auxiliary diagnostic information retrieval method used can be used alone.

[0034] like Figure 1 As shown, an auxiliary diagnostic information retrieval system based on a respiratory knowledge base utilizes the aforementioned respiratory knowledge base retrieval method. The system includes a data acquisition module, a data preprocessing module, and a knowledge base retrieval module. The data acquisition module collects respiratory system data such as patients' chest CT images, blood gas analysis results, and pathology reports. The data preprocessing module uses different preprocessing methods to preprocess different data types, such as cleaning, standardizing, and format conversion. The knowledge base retrieval module retrieves relevant diagnostic information from a medical knowledge base based on the patient's respiratory system-related data. The retrieval process prioritizes searching for relevant entries in the medical knowledge base. If no relevant entry is found, a large model is invoked for auxiliary diagnostic information reasoning, and the knowledge base retrieval process described in step S1 is executed; otherwise, auxiliary diagnostic information from the medical knowledge base is returned. The medical knowledge base includes the "Respiratory Disease Diagnosis and Treatment Knowledge Base 2025," the Global Initiative for Chronic Obstructive Pulmonary Disease (GOLD 2025 Guidelines), and the Chinese Medical Association's Diagnostic Criteria for Respiratory Failure, among others.

[0035] I. The data acquisition module will be described in detail below: 1. The types of data input into the retrieval system include: Image data: The image data input into the retrieval system includes medical images such as CT scans, X-rays, and MRIs. Formats: DICOM, JPEG, PNG, etc.

[0036] Text data: Text data input into the retrieval system includes medical history records, physical examination reports, laboratory test reports, etc. Formats: TXT, PDF, Word, etc.

[0037] Numerical data: Numerical data input into the retrieval system includes laboratory test results, vital signs, age, gender, etc. Format: CSV, Excel, etc.

[0038] 2. Channels for inputting data into the retrieval system include: Manual upload: Doctors or patients manually upload data.

[0039] Interface integration: It interfaces with Hospital Information System (HIS), Picture Archiving and Communication System (PACS), and Laboratory Information System (LIS) to automatically acquire data.

[0040] II. The data preprocessing module will be described in detail below: The data preprocessing module can preprocess image data, text data, and numerical data. The preprocessing process for each type of data is described below: 1. Image data preprocessing includes the following steps: (1) Format conversion: Convert image data of different formats into a unified format (e.g., DICOM).

[0041] (2) Image denoising: Use filtering algorithms to remove noise from the image.

[0042] (3) Median filtering: For each pixel in the image, replace it with the median of its neighboring pixels. The formula for median filtering is: output(x, y) = median{neighbor(x, y)}.

[0043] (4) Gaussian filtering: The image is smoothed by convolving it with a Gaussian kernel. The formula for Gaussian filtering is: output(x, y) = ΣΣ G(u, v) input(x - u, y - v), where G(u, v) = (1 / (2πσ²)) exp(-(u² + v²) / (2σ²)) (5) Image enhancement: Adjust the contrast and brightness of the image to improve the image clarity.

[0044] (6) Histogram equalization: Adjusts the histogram of an image to make the image contrast more uniform. The formula for histogram equalization is: s = round(L... (cdf(r) - cdf_min) / (cdf_max - cdf_min)), where cdf(r) is the cumulative distribution function and L is the number of gray levels.

[0045] (7) CLAHE (Contrast Limited Adaptive Histogram Equalization): Performs local histogram equalization on an image to improve the local contrast of the image.

[0046] (8) Image segmentation: dividing an image into different regions, such as the lungs, lesions, etc.

[0047] (9) Threshold segmentation: The image is segmented into different regions based on the gray value of the pixels.

[0048] (10) Region growing: Starting from the seed point, adjacent pixels are merged into the same region based on the similarity of pixels.

[0049] (11) Deep learning segmentation: Image segmentation using deep learning models (e.g., U-Net).

[0050] (12) Feature extraction: Extract features from the image, such as emphysema index, lesion size, lesion density, etc.

[0051] (13) Gray co-occurrence moment (GLCM): used to extract texture features of an image.

[0052] (14) HOG (Histogram of Oriented Gradients): used to extract edge features of an image.

[0053] (15) Deep learning features: using deep learning models to extract features from images.

[0054] The above preprocessing steps can be selectively combined according to actual needs. For example: Low-noise, high-quality imagery: may only require format conversion and histogram equalization.

[0055] Detection of complex lesions may require a combination of denoising, CLAHE enhancement, deep learning segmentation, and feature extraction.

[0056] Specific feature analysis: For example, emphysema detection may focus on gray-level co-occurrence moments (texture features), while lung nodule identification may rely on HOG (edge ​​features).

[0057] The order directly affects the processing results. Table 1 shows the recommended typical order of steps and the reasons.

[0058] Table 1 Recommended processing steps and reasons

[0059] An incorrect preprocessing order presents a potential problem: Segmentation followed by denoising: The segmentation result may contain regions with noise interference, leading to inaccurate segmentation boundaries.

[0060] Enhance first, then denoise: Enhancement may amplify noise, making the denoising effect worse.

[0061] Skip format conversion: Images of different formats may fail to be processed due to inconsistent metadata.

[0062] The impact of different preprocessing steps on the retrieval process: (1) Data quality and retrieval accuracy High-quality preprocessing (such as denoising + CLAHE) can improve the accuracy of feature extraction. For example, in emphysema detection, clear texture features (gray-level co-occurrence moments) can more reliably reflect lung structural damage. In nodule recognition, edge features extracted by HOG are more prominent in the enhanced image, improving segmentation accuracy.

[0063] Low-quality preprocessing (such as over-filtering) can lead to feature loss. For example, excessive smoothing by Gaussian filtering may blur nodule boundaries, causing feature extraction to fail.

[0064] (2) Feature diversity and retrieval comprehensiveness Multi-step combinations (such as CLAHE + gray-level co-occurrence moments + HOG) can extract multi-dimensional features, covering different pathological features: COPD diagnosis: Combining texture (grayscale co-occurrence moments) and edge (HOG) features allows for the simultaneous analysis of lung structural damage and airway narrowing.

[0065] Lung fibrosis detection: CLAHE-enhanced contrast combined with deep learning features can more comprehensively capture the heterogeneity of fibrotic regions.

[0066] (3) Search efficiency and computational cost Simplifying the process (such as only format conversion and splitting) can reduce computational resource consumption and is suitable for scenarios with high real-time requirements.

[0067] Complex combinations of steps (such as deep learning segmentation + multi-feature extraction) may increase computation time, but may improve diagnostic reliability.

[0068] Examples of pretreatment step selection in practical respiratory medicine applications: Scenario 1: COPD Diagnosis Preprocessing options: Format conversion → CLAHE → Median filtering → Gray-level co-occurrence moment extraction.

[0069] Cause: Emphysema in COPD manifests as low-density areas and abnormal texture. CLAHE enhances contrast, and gray-level co-occurrence moments can quantify texture heterogeneity.

[0070] Scenario 2: Classification of benign and malignant pulmonary nodules Preprocessing options: Format conversion → Gaussian filtering → HOG extraction → Deep learning segmentation.

[0071] Reasons: Gaussian filtering reduces noise interference, HOG captures nodule edge features, and deep learning segmentation accurately locates the nodule region.

[0072] Scenario 3: Pulmonary fibrosis assessment Preprocessing options: format conversion → histogram equalization → region growing → deep learning feature extraction.

[0073] Reason: Histogram equalization improves the contrast of fibrotic regions, region growing segments the lesion area, and deep learning features capture fibrosis patterns.

[0074] In practical applications, the preprocessing steps can be flexibly combined according to the image quality and diagnostic objectives.

[0075] When combining preprocessing steps, the processing order needs to be optimized, following the logic of "format conversion → denoising → enhancement → segmentation → feature extraction" to avoid performance degradation caused by reverse order.

[0076] Impact on effect: High-quality preprocessing improves feature accuracy, and multi-step combination enhances the comprehensiveness of retrieval, but the computational cost needs to be balanced.

[0077] 2. Text data preprocessing includes the following steps: (1) Text cleaning: remove special characters, HTML tags, punctuation marks and the like from the text.

[0078] (2) Word segmentation: split the text into different words. Available tools for word segmentation include: Jieba word segmentation: a commonly used Chinese word segmentation tool; spaCy word segmentation: a commonly used English word segmentation tool.

[0079] (3) Stop word removal: remove meaningless words in the text, for example, "de", "le", "shi" and the like.

[0080] (4) Part-of-speech tagging: tag the part of speech of each word, for example, nouns, verbs, adjectives and the like.

[0081] (5) Named Entity Recognition (NER): recognize named entities in text data, for example, diseases, drugs, symptoms and the like.

[0082] (6) Text vectorization: convert text into vector representation.

[0083] (7) TF-IDF: calculate the weight of a word. The formula of TF-IDF is: TF-IDF(t, d, D) = TF(t, d) IDF(t, D) (8) Word2Vec: used to generate vector representations of words. It trains word vectors using a neural network, and maps words to a low-dimensional vector space.

[0084] (9) BERT: used to generate vector representations of text.

[0085] Not all of the above preprocessing steps will be adopted, which specifically depends on the characteristics of the knowledge base, retrieval objectives and computing resources.

[0086] The core steps, which are also mandatory steps, include: text cleaning, word segmentation, stop word removal, and text vectorization (e.g., TF-IDF / Word2Vec / BERT).

[0087] Optional steps: Part-of-speech tagging / named entity recognition: adopted when it is necessary to distinguish entity types (e.g., disease names, drug names) or analyze grammatical structures.

[0088] TF-IDF vs. Word2Vec vs. BERT: select according to requirements: TF-IDF is suitable for simple scenarios, Word2Vec captures semantics, and BERT takes into account both semantics and context, but has higher computational cost.

[0089] The execution order of steps directly affects the processing effect. The typical processing flow is: Text cleaning (removing noise such as special symbols and stop words); Word segmentation (segmenting continuous text into meaningful units); Stop word removal (filtering meaningless high-frequency words such as "de" and "shi" in Chinese); Part-of-speech tagging / named entity recognition (if structured analysis is required); Text vectorization (converting text into numerical representations) Potential problems of different processing orders: Word segmentation first then cleaning: incorrectly segmented special characters may be retained (e.g., "COPD-1" is incorrectly segmented).

[0090] Vectorization first then cleaning: noise data will contaminate vector representations and reduce retrieval accuracy.

[0091] Skipping stop word removal: may lead to vector dimension explosion and affect efficiency.

[0092] Influence of preprocessing steps on retrieval Word segmentation accuracy: directly affects the quality of vectorization. For example, if "FEV1" is incorrectly segmented into "FEV" and "1" in medical text, key information will be lost.

[0093] Selection of vectorization methods: TF-IDF: suitable for keyword matching, but cannot handle synonyms (e.g., "asthma" and "bronchial asthma").

[0094] Word2Vec: can capture semantic similarity, but has insufficient handling of polysemy (e.g., the different meanings of "lung" in "pulmonary infection" and "pulmonary function").

[0095] BERT: can dynamically generate context-related vectors, which is suitable for complex semantic retrieval, but requires a large amount of computing resources.

[0096] In summary, the selection and order of preprocessing steps should be tailored to the specific scenario. For example, when searching for respiratory diseases (such as COPD), medical terminology (such as "FEV1" and "small airway lesions") should be retained, and BERT may be used to capture contextual semantics. A well-designed preprocessing workflow can significantly improve search results.

[0097] 3. Numerical data preprocessing includes the following steps: (1) Missing value handling: Missing values ​​in numerical data can be filled or deleted. Missing values ​​can be filled by the mean, the median, or by interpolation.

[0098] (2) Outlier handling: Detecting and handling outliers in numerical data. Outliers can be detected using box plots or Z-scores.

[0099] (3) Data standardization: scaling the data to the same range. For example, using Min-Max standardization (the formula for Min-Max standardization: x' = (x - min) / (max - min)) to scale the data to the range of 0 to 1, or using Z-score standardization (the formula for Z-score standardization: x' = (x - μ) / σ) to scale the data to the range of 0 mean and 1 standard deviation.

[0100] Typically, all numerical data preprocessing steps are used, but the specific steps depend on the data characteristics and requirements: Handling missing values: Respiratory physiological data (such as FEV1, FVC) often have missing values, which need to be filled or deleted.

[0101] Handling of outliers: Physiological indicators (such as blood oxygen saturation) may have abnormal values ​​due to measurement errors or extreme cases, which need to be detected and handled.

[0102] Data standardization: Different indicators (such as vital capacity in L vs. blood oxygen percentage) have large differences in dimensions, and standardization can eliminate the influence.

[0103] There is a clear order to the preprocessing steps; the recommended workflow is as follows: 1. Handling missing values ​​→ 2. Handling outliers → 3. Data standardization Different orders can have unexpected effects, as follows: Standardize before processing missing values: Standardization may mask missing values ​​(e.g., filling with 0), leading to biases in subsequent analyses. For example, the mean of missing values ​​after standardization may be calculated incorrectly, affecting the accuracy of the filling.

[0104] Handle outliers before handling missing values: Outlier handling may alter the data distribution, leading to inaccurate mean / median values ​​when filling in missing values. For example, after removing outliers, the mean of the remaining data may deviate from the true value.

[0105] Disordered processing: Disorganized data processing logic can lead to unreliable results. For example, standardized outliers may be misclassified as normal values.

[0106] The impact of different preprocessing steps on the retrieval process is as follows: (1) Handling missing values: Mean imputation may underestimate data volatility and affect the robustness of search results.

[0107] Interpolation: Preserves data trends, but has high computational complexity and is suitable for time series data (such as respiratory waveforms).

[0108] (2) Outlier handling: Deletion method: May result in data loss and affect the scope of search coverage.

[0109] Replacement method: may introduce bias and needs to be combined with domain knowledge (such as whether extreme values ​​in respiratory data are true pathological manifestations).

[0110] (3) Data standardization: Z-score: Suitable for normally distributed data, but may not be applicable to long-tailed distributions (such as lung function indicators).

[0111] Min-Max: Preserves the original distribution, but is sensitive to outliers and should be used after outlier handling.

[0112] In summary, the order and methods of preprocessing directly affect data quality in respiratory knowledge base retrieval. For example, when processing pulmonary function data, missing values ​​should be filled in first (e.g., using the patient's historical mean), then outliers should be detected (e.g., whether oxygen saturation <90% indicates a true pathological condition), and finally, standardization should be performed (e.g., unifying FEV1 and FVC to the 0-1 range). A well-designed preprocessing workflow can improve the accuracy and reliability of retrieval.

[0113] Data acquisition and preprocessing are crucial steps in a retrieval system. Appropriate preprocessing methods can improve data quality and reliability, providing a foundation for subsequent retrieval and analysis.

[0114] III. The following is a detailed introduction to the knowledge base retrieval module: The knowledge base retrieval module is used to retrieve relevant diagnostic information from the medical knowledge base based on the patient's respiratory system-related data. The retrieval process prioritizes searching for relevant entries in the medical knowledge base. If no relevant entry is found, a large model is invoked to perform auxiliary diagnostic information reasoning, and the knowledge base retrieval process described in step S1 is executed; otherwise, auxiliary diagnostic information from the medical knowledge base is returned. The medical knowledge base includes the "Respiratory Disease Diagnosis and Treatment Knowledge Base 2025," the Global Initiative for Chronic Obstructive Pulmonary Disease (GOLD 2025 Guidelines), and the Chinese Medical Association's Diagnostic Criteria for Respiratory Failure, among others.

[0115] 1. Knowledge base construction and organization Before detailing the retrieval process, let's briefly explain the construction of the knowledge base. This system uses a hybrid knowledge base, which includes: A structured medical knowledge base: built on ontology, containing entities such as diseases, symptoms, drugs, and examinations, as well as their relationships. It is stored in RDF triples (Resource Description Framework).

[0116] Unstructured medical knowledge base: Contains textual data such as medical literature, guidelines (GOLD2025, GINA, ATS / ERS), disease diagnostic criteria (ICD-11), imaging feature databases, pathological feature databases, drug treatment plan databases, clinical trial data, and expert experience. It is stored using a vector database, and text is converted into vector representations using embedding technology.

[0117] Hybrid knowledge base: This combines structured and unstructured medical knowledge bases to achieve knowledge fusion and reasoning.

[0118] Knowledge Representation: Knowledge is organized in the form of a knowledge graph. Nodes represent entities (diseases, symptoms, imaging features, drugs, etc.), and edges represent relationships between entities (e.g., a disease "causes" symptoms, and imaging features "hint" at a disease).

[0119] Knowledge graph structure: It adopts a multi-level structure, organizing knowledge from macro to micro levels. For example, knowledge can be divided into the following three layers: Top level: Disease classification system (e.g., respiratory diseases, cardiovascular diseases).

[0120] Intermediate layer: Specific diseases (e.g., COPD, asthma, pneumonia).

[0121] The underlying layer includes the disease's symptoms, imaging features, pathological features, and treatment plans.

[0122] 2. Forced Search Protocol and Keyword Construction Forced retrieval protocol: When retrieving auxiliary diagnostic information, the retrieval of the knowledge base is enforced first to ensure that all diagnostic conclusions are adequately supported by knowledge.

[0123] Keyword construction for search: Keyword extraction based on patient data: Extracting keywords from patients' imaging reports, laboratory test results, and medical history information.

[0124] Keyword expansion: Expand keywords using synonyms, near-synonyms, and hyponyms to improve search coverage.

[0125] Keyword weighting: Different weights are assigned to keywords based on their importance. For example, keywords related to imaging features have higher weights, while keywords related to medical history have lower weights.

[0126] 3. Implementation steps for hierarchical traversal of a knowledge graph Root node localization: Based on the patient's initial symptoms or imaging characteristics, locate the root node of the knowledge graph (e.g., respiratory diseases).

[0127] Hierarchical traversal: Starting from the root node, traverse the knowledge graph level by level.

[0128] Breadth-first traversal: Prioritizes traversing nodes at the same level to expand the search scope.

[0129] Depth-first traversal: Deeply traverses the next level of nodes to refine the search results.

[0130] Relationship filtering: Based on the patient's characteristics, filter out irrelevant nodes and relationships.

[0131] Similarity matching: Calculate the similarity between patient features and knowledge graph nodes, and select the most matching node.

[0132] Results sorting: Sort the search results according to similarity, and return the most matching node first.

[0133] 4. Knowledge Graph Retrieval Process First, perform the following embedding vectorization preprocessing: convert keywords and entities in the knowledge graph, as well as the relationships between entities, into vector representations. Models such as Word2Vec, GloVe, and BERT are used to generate embedding vectors, including keyword vectors and entity vectors.

[0134] Then, similarity calculation is performed: the similarity between the keyword vector and the entity vector in the knowledge graph is calculated, preferably using cosine similarity: similarity(A,B)=(A·B) / (||A|| ||B||); Then perform a hierarchical traversal: starting from the root node, traverse the knowledge graph layer by layer, and select the nodes and relationships with the highest similarity.

[0135] Knowledge fusion: The retrieved knowledge is integrated to form a complete knowledge chain.

[0136] In theory, all steps can be used, but in practical applications, adjustments need to be made according to specific requirements: Vectorization of embeddings is essential and forms the basis of knowledge graph retrieval.

[0137] Similarity calculation: This is mandatory and is used to match keywords with entities.

[0138] Hierarchical traversal: Optional, used when the knowledge graph has a complex hierarchical structure and requires deep retrieval.

[0139] Knowledge fusion: Optional, used when it is necessary to integrate knowledge from multiple sources to form a complete chain.

[0140] The above steps have a strict order; the recommended process is as follows: Embedding vectorization → Similarity calculation → Hierarchical traversal → Knowledge fusion.

[0141] Different execution orders can have unexpected effects: Hierarchical traversal followed by vectorization: This method fails to calculate the similarity between keywords and entities, leading to inaccurate search results. For example, in the respiratory knowledge graph, if "COPD" is not vectorized, direct traversal may miss key nodes.

[0142] First, knowledge is fused, then vectorized: The fused knowledge chain lacks vector representation and cannot be matched with keywords. For example, after fusing "asthma-bronchitis-emphysema", if it is not vectorized, the relevant entities cannot be retrieved.

[0143] Calculating similarity before hierarchical traversal may miss deep nodes with high similarity, leading to incomplete search results. For example, when searching for "pulmonary fibrosis," if deep relationships are not traversed, its association with "interstitial lung disease" may be overlooked.

[0144] The impact of different preprocessing steps on the retrieval process is as follows: Vectorization of Embedding: Quality impact: Using low-quality vectors (such as models not trained for the respiratory medicine field) may lead to misclassification of similarity between "pneumonia" and "tuberculosis".

[0145] Method selection: Embedding (such as models trained on lung function data) can improve retrieval accuracy.

[0146] After similarity calculation, the results need to be sorted and the best match selected. At this point, a filtering process is required based on the set threshold. If the threshold is too low, irrelevant entities may be retrieved (such as mismatch between "lung" and "alveolar"); if it is too high, key associations may be missed (such as between "asthma" and "airway hyperresponsiveness").

[0147] Level traversal: Depth control: If the traversal is too shallow, the "COPD-chronic bronchitis-emphysema" link may be missed; if it is too deep, noise may be introduced (such as the "emphysema-lung cancer" link not being directly related).

[0148] Knowledge integration: Strategy selection: If simple splicing is used, it may lead to redundancy (such as the repetition of "asthma-bronchospasm-airway stenosis"); if semantic reasoning is used, more accurate links may be generated (such as "asthma-inflammatory factors-airway remodeling").

[0149] The following example illustrates this concept in a respiratory medicine setting: Suppose the search keyword is "pulmonary fibrosis": Embedding Vectorization: Converts "pulmonary fibrosis" and entities in the knowledge graph (such as "interstitial lung disease" and "collagen deposition") into vectors.

[0150] Similarity calculation: Calculate the cosine similarity between "pulmonary fibrosis" and "interstitial lung disease" (e.g., 0.92), and the similarity between "pulmonary emphysema" and "pulmonary emphysema" (e.g., 0.35), and sort and filter them.

[0151] Hierarchical traversal: Starting from the root node "Respiratory diseases", we delve deeper layer by layer, selecting "Interstitial lung disease" and its child node "Collagen deposition" with the highest similarity.

[0152] Knowledge integration: Integrating "pulmonary fibrosis - interstitial lung disease - collagen deposition - decreased lung function" to form a complete chain.

[0153] 5. Knowledge Base Retrieval Process The knowledge base retrieval process consists of the following steps: (1) Query construction: Transform the patient's respiratory system-related data into query statements.

[0154] (2) Query decomposition: Decompose complex query statements into multiple simple queries.

[0155] (3) Retrieval execution: Execute the decomposed query in the medical knowledge base to obtain multiple query results; (4) Result fusion: Multiple query results are fused based on similarity to obtain a fused result set of multiple query statements; (5) Result sorting: Sort the query results in the fusion result set and select the most relevant results as the retrieved auxiliary diagnostic information.

[0156] In the respiratory knowledge base-based auxiliary diagnostic information retrieval method, the multiple results in the result fusion stage originate from the independent retrieval results of multiple simple query statements after query decomposition. The following is a detailed analysis: I. Sources of Multiple Results Query decomposition granularity: The original query (e.g., "differential diagnosis of decreased FEV1 in COPD patients") is decomposed into multiple subqueries (e.g., "pathological mechanism of COPD", "common causes of decreased FEV1", "imaging features of small airway lesions"). Each subquery is retrieved from a specific module in the knowledge base (e.g., disease database, examination database, literature database).

[0157] Diversity of search results: The results returned by each subquery may include: Structured data: such as diagnostic criteria (COPD classification in the GOLD guidelines) and examination indicators (FEV1 / FVC ratio).

[0158] Unstructured data: such as literature abstracts and treatment recommendations in expert consensus.

[0159] Related information: such as the pathological association between COPD and small airway lesions, and a list of differential diagnoses for decreased FEV1.

[0160] Redundancy and conflicting results: For example, the subquery "pathological mechanisms of COPD" might return "small airway inflammation" and "alveolar structural damage," while the subquery "causes of decreased FEV1" might return "small airway obstruction" and "chest deformity." These results partially overlap or conflict in content (e.g., whether they emphasize the core of small airway lesions).

[0161] II. Implementation Methods of Result Fusion 1. Rule-based fusion strategy Prioritization: Results are weighted based on authority (e.g., guidelines > expert consensus > general literature) or the credibility of the data source (e.g., PubMed literature > blog posts).

[0162] Conflict resolution: Logical rules: If multiple results point to the same cause (e.g., "small airway disease"), they are merged into a unified conclusion; if the results are contradictory (e.g., "FEV1 decline is caused by airway obstruction" vs. "FEV1 decline is caused by decreased lung elasticity"), a confidence threshold is introduced (e.g., a conclusion is retained when >80% of the literature supports it).

[0163] Medical knowledge graph: Validate the consistency of results using predefined causal relationships (such as "COPD → small airway inflammation → decrease in FEV1").

[0164] 2. Machine Learning-Based Fusion Strategy Feature extraction: Extract features from search results (such as keyword frequency, publication year, and citation count).

[0165] Model training: Using labeled data (such as clinical cases with known correct answers) to train a classification model to predict the relevance or confidence of the results.

[0166] Dynamic weighting: Adjust the weights of different subquery results based on user needs (e.g., "early diagnosis" vs. "treatment plan").

[0167] 3. Graph-based fusion strategy Construct a network of connections: Entities (such as “COPD”, “FEV1”, and “small airway lesions”) in the search results are used as nodes, and relationships (such as “cause-symptom” and “examination-diagnosis”) are used as edges to form a knowledge graph.

[0168] Path analysis: Identify key findings using shortest path algorithms (such as Dijkstra's algorithm) or community detection algorithms (such as Louvain's algorithm). For example, if "small airway disease" is mentioned in multiple paths, its importance is reinforced.

[0169] 4. Weighted average method Simple weighting: Assign weights based on the complexity or number of results of the subquery (e.g., the subquery "pathological mechanism" has a weight of 0.4, and the subquery "examination indicators" has a weight of 0.3).

[0170] Confidence weighting: The results are weighted by combining the confidence level of the source (e.g., 0.9 for guideline entries and 0.6 for general literature).

[0171] III. Optimization of the Fusion Results Deduplication and standardization: Combine duplicate diagnostic recommendations (e.g., "inhaled bronchodilator" and "β2 receptor agonist"). Standardize terminology (e.g., map "FEV1" and "forced expiratory volume" to the same indicator).

[0172] Contextual association: Integrate the scattered results into a coherent logical chain (e.g., "COPD → small airway inflammation → decreased FEV1 → inhalation therapy").

[0173] Supplement missing information (e.g., if the search results do not mention "lung function assessment", supplement relevant guidelines from the knowledge base).

[0174] User needs adaptation: Adjust the output format based on the user's identity (e.g., doctor vs. patient). For example, doctors may need detailed pathological mechanisms, while patients may need simplified treatment suggestions.

[0175] Dynamically filter irrelevant results (such as excluding complications that are not related to the current problem).

[0176] IV. Technical Implementation Path of Result Ranking 1. Selection of sorting algorithm Vector space model ranking: Converts search results and query statements into vectors (such as TF-IDF, Word2Vec), and calculates relevance scores using cosine similarity. Deep learning ranking models: Use pre-trained models such as BERT and RoBERTa to perform semantic matching on query-result pairs, and output probability values ​​as the ranking criteria. Ensemble ranking framework: Employs gradient boosting decision tree models such as LambdaMART, integrating multiple features (such as relevance, authority, and timeliness) for ranking. 2. Feature Engineering Design Table 2 shows the feature engineering design table.

[0177] Table 2 Feature Engineering Design Table

[0178] 3. The sorting process is as follows: Figure 2 As shown.

[0179] Feature extraction is performed on the fusion result set, mainly extracting features such as relevance, authority, and timeliness; then, a ranking model is used to calculate the comprehensive score for model prediction; next, online learning and dynamic adjustments are made based on user feedback (click rate, dwell time); finally, the results are ranked and the top N results are output.

[0180] V. Special Considerations in the Field of Respiratory Medicine 1. Disease spectrum characteristics Differentiating between acute and chronic conditions: For acute respiratory failure, emergency treatment guidelines should be prioritized; for chronic obstructive pulmonary disease (COPD), long-term management plans should be emphasized. Handling overlapping symptoms: Differential diagnosis of pneumonia and pulmonary tuberculosis requires increased weighting of specific content in the ranking process. 2. Priority of Evidence Levels Table 3 is a table of evidence level priorities.

[0181] Table 3. Priority Table of Evidence Levels

[0182] 3. Personalized sorting strategy Patient profile matching: The sorting weight is adjusted based on characteristics such as age, smoking history, and allergy history; for example, lung cancer screening-related content is displayed first for smoking patients.

[0183] Treatment phase adaptation: Diagnostic phase: Increase the weighting of content related to differential diagnosis; Treatment phase: Prioritize displaying the latest drug clinical trial data; VI. Analysis of Typical Application Scenarios Scenario: Pneumonia diagnosis search Basic ranking: Sort by relevance score (e.g., "diagnostic criteria for pneumonia" ranked in the top 3). Dynamic adjustment: If a user enters "childhood pneumonia", the weight of pediatric guidelines will be automatically increased; if the search time is in winter, the ranking weight of "influenza-related pneumonia" will be increased.

[0184] Results presentation: The first position shows the latest version of the "Guidelines for the Diagnosis and Treatment of Pneumonia"; the second position shows high-quality RCT studies from the past 3 years; and the third position shows the key points of differential diagnosis from authoritative experts.

[0185] VII. Technical Challenges and Solutions Table 4 presents the solutions proposed by this invention to address the technical challenges.

[0186] Table 4. Technical Challenges and Solutions

[0187] VIII. Verification and Optimization Methods Offline assessment: Ranking quality was assessed using the NDCG (Normalized Discounted Cumulative Gain) index, and a dedicated test set for respiratory medicine (containing 500+ typical cases) was constructed.

[0188] Online A / B testing: Compare user click-through rates, diagnostic accuracy, and other metrics of different ranking strategies, and dynamically optimize ranking parameters using a multi-armed slot machine algorithm.

[0189] This structured design ensures both the versatility of the retrieval system and the specific needs of the respiratory medicine field. In practical deployment, a three-tier architecture of "basic sorting + domain adaptation + personalized adjustments" is recommended, continuously improving the sorting performance through ongoing user feedback and data iteration.

[0190] Example 3 The following detailed description of the respiratory knowledge base-based auxiliary diagnostic information retrieval system of the present invention is based on a specific embodiment. Where details are not exhaustive, the method described in Embodiment 2 shall apply.

[0191] I. Data Collection 1. Impact data (images / videos) Data types: chest X-rays, CT scan images, bronchoscopy videos.

[0192] Example: CT scan of the lungs of a COPD patient shows areas of emphysema (black low-density shadows) and thickening of the bronchial walls.

[0193] Lung cancer X-ray: Lung nodule shadow (diameter >3cm) with spiculation.

[0194] Data sources: hospital imaging databases (such as PACS systems) and publicly available medical imaging datasets (such as NIH ChestX-ray14).

[0195] 2. Text data Data type: Medical record: Patient's chief complaint (e.g., "chronic cough with sputum for 10 years"), past medical history (e.g., "smoking history for 30 years").

[0196] Medical literature: Diagnostic criteria in the Global Guidelines for the Diagnosis and Treatment of COPD (2023 edition).

[0197] Patient's symptom description: self-reported on social media or forums (e.g., "Recently, I have chest tightness and shortness of breath after activity").

[0198] Data sources: Electronic Health Records (EHR), PubMed literature, and online patient communities (such as DXY.cn).

[0199] 3. Numerical data Data type: Lung function tests: FEV1 / FVC ratio (e.g., 0.65), FEV1 predicted percentage (e.g., 50%).

[0200] Blood parameters: oxygen saturation (SpO2: 92%), C-reactive protein (CRP: 80 mg / L).

[0201] Imaging quantification: Percentage of emphysema volume on CT scan (e.g., 40%).

[0202] Data sources: hospital laboratory systems, wearable devices (such as blood oxygen monitoring watches).

[0203] II. Pretreatment Steps 1. Text Data Preprocessing Word segmentation and standardization: Use medical NLP tools (such as BioBERT) to unify "chronic obstructive pulmonary disease" as "COPD (ICD-10: J44.9)".

[0204] Remove stop words (e.g., "de", "le") and retain key symptoms (e.g., "cough", "hemoptysis").

[0205] Entity Recognition: Recognize diseases (e.g., "asthma"), drugs (e.g., "salmeterol"), and examination methods (e.g., "pulmonary function test").

[0206] Relation Extraction: Extract "COPD → Symptom: Chronic cough" and "Salmeterol → Indication: COPD" from literatures.

[0207] 2. Numerical Data Preprocessing Missing Value Processing: For missing FEV1 values, mean imputation is adopted (e.g., the mean FEV1 of COPD patients is 60%).

[0208] Outlier Detection: The IQR method is used to eliminate abnormal records with CRP values > 200 mg / L.

[0209] Normalization: Map the FEV1 / FVC ratio to the interval [0,1] (e.g., 0.65 → 0.65).

[0210] 3. Image Data Preprocessing Image Cleaning: Remove blurred or underexposed CT images.

[0211] Feature Extraction: The ResNet-50 model is used to extract texture features of lung CT (e.g., Gabor filter response).

[0212] Annotation: Annotate emphysema regions (label: Emphysema) and nodule positions (coordinates: x=120, y=150).

[0213] III. Construction of Knowledge Graph 1. Definition of Entities and Relations Entity Types: Diseases (COPD, asthma), symptoms (cough, hemoptysis), examinations (pulmonary function test), drugs (salmeterol).

[0214] Relation Types: Disease → Symptom (e.g., COPD → Chronic cough) Drug → Indication (e.g., Salmeterol → COPD) Examination → Index (e.g., Pulmonary function test → FEV1) 2. Example of Knowledge Graph (COPD) - [:typical symptoms] -> (chronic cough) (COPD) - [:diagnostic criteria] -> (FEV1 / FVC < 0.7) (Sameterol) - [:indications] -> (COPD) (Lung function test) - [:indicators] -> (FEV1) 3. Knowledge Integration Conflict resolution: If two studies do not describe the diagnostic criteria for COPD in a consistent manner (e.g., FEV1 / FVC <0.7 vs. <0.65), the latest guidelines (e.g., GOLD 2023) should be adopted first.

[0215] Data source weight: Authoritative guidelines (such as GOLD) have a weight of 1.0, while patient forums have a weight of 0.3.

[0216] IV. Search Process 1. User Question Analysis Question: "What are the diagnostic criteria for COPD?" Analysis results: Entity: COPD Intent: To query diagnostic criteria.

[0217] 2. Query Construction Cypher query statement: MATCH (d:Disease {name: "COPD"})-[:diagnostic criteria]->(criteria) RETURNcriteria 3. Results Retrieval and Fusion Knowledge graph results: FEV1 / FVC < 0.7 (from GOLD 2023).

[0218] Supplementary data: Recent studies (such as the 2023 Lancet Respir Med) suggest that "FEV1 / FVC < 0.65 is more sensitive".

[0219] 4. Result Sorting and Output Sorting rules: Weighting: Guidelines (1.0) > High-impact factor journals (0.8) > Patient forums (0.3).

[0220] Output example: COPD diagnostic criteria (GOLD 2023): Lung function test: FEV1 / FVC < 0.7 (after ruling out other diseases).

[0221] Symptoms: Chronic cough with sputum, lasting ≥3 months / year, or ≥2 years.

[0222] Imaging findings: Chest CT scan showed emphysema or thickening of the bronchial walls.

[0223] This system integrates multi-source heterogeneous data (images, text, numerical values) to construct a knowledge graph of respiratory diseases, and combines medical guidelines and the latest research to achieve accurate retrieval of diagnostic criteria.

[0224] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A method for retrieving auxiliary diagnostic information based on a respiratory knowledge base, characterized in that, Includes the following steps: Step S1: Data Acquisition: Collect respiratory system-related data from the patient; Step S2: Data preprocessing: Different preprocessing methods are used to preprocess different types of collected data; Step S3: Assisted Diagnostic Information Retrieval: Based on the patient's respiratory system-related data, retrieve relevant diagnostic information from the medical knowledge base; the retrieval process prioritizes searching for relevant entries in the medical knowledge base. If no relevant entries are found in the medical knowledge base, the transformer large model is invoked to perform assisted diagnostic information reasoning; otherwise, the assisted diagnostic information from the medical knowledge base is returned. The aforementioned transformer large model supports both structured and unstructured medical knowledge bases. It is trained based on massive amounts of data and possesses an understanding of the pathological mechanisms, diagnostic criteria, and treatment strategies of respiratory diseases. It decomposes problems through CoT chain reasoning and ToT tree reasoning, and can generate verifiable reasoning paths based on the patient's respiratory system-related data through the RAG tracing mechanism combined with the medical knowledge base, thereby assisting in the reasoning of diagnostic information. The auxiliary diagnostic information retrieval includes the following steps: Query construction: Transforming patients' respiratory system-related data into query statements; Query decomposition: The query statement is broken down into several simpler query statements; Retrieval Execution: Based on the decomposed query statement, the query is executed in the medical knowledge base. A knowledge graph is constructed based on the medical knowledge base. During retrieval, a hierarchical traversal is performed, starting from the root node and traversing the knowledge graph layer by layer. The node and relationship with the highest similarity to the query are selected and the query results are returned. Multiple query results are obtained for each decomposed query statement. Result fusion: Merge multiple query results based on similarity to obtain a fused result set of multiple query statements; Result sorting: Sort the query results in the fused result set and select the results most relevant to the query as the retrieved auxiliary diagnostic information.

2. The auxiliary diagnostic information retrieval method based on a respiratory knowledge base according to claim 1, characterized in that, Data types include image data, text data, and numerical data.

3. The auxiliary diagnostic information retrieval method based on a respiratory knowledge base according to claim 2, characterized in that, The preprocessing of the image data includes at least one of the following steps, selected and combined based on image quality and diagnostic objectives: Format conversion: Converting image data from different formats into a unified format; Image denoising: using filtering algorithms to remove noise from an image; Median filtering: For each pixel in an image, replace it with the median value of its neighboring pixels; Gaussian filtering: Smooths the image by applying a Gaussian kernel and performing convolution. Image enhancement: Adjusts the contrast and brightness of an image to improve its clarity; Histogram equalization: Adjusts the histogram of an image to make the image contrast more uniform; CLAHE: Performs local histogram equalization on an image to improve local contrast. Image segmentation: dividing an image into different regions; Thresholding segmentation: Dividing an image into different regions based on the grayscale values ​​of pixels; Region growing: Starting from the seed point, adjacent pixels are merged into the same region based on pixel similarity; Deep learning segmentation: Image segmentation using deep learning models; Feature extraction: Extracting features from an image; Gray-level co-occurrence moments: used to extract texture features from an image; HOG: Used to extract edge features from an image; Deep learning features: using deep learning models to extract features from images.

4. The auxiliary diagnostic information retrieval method based on a respiratory knowledge base according to claim 2, characterized in that, The preprocessing of the text data includes the following steps: Text cleaning: Remove special characters, HTML tags, and punctuation marks from text; Word segmentation: dividing text into multiple words; Remove stop words: Remove words that have been stopped from the text; Text vectorization: Converting text into a vector representation.

5. The auxiliary diagnostic information retrieval method based on a respiratory knowledge base according to claim 4, characterized in that, The preprocessing of the text data also includes at least one of the following steps: Part-of-speech tagging: Tagging the part of speech of each word; Named entity recognition: Identifying named entities in text data; The preprocessing of the numerical data includes the following steps: Missing value handling: Fill or delete missing values ​​in numerical data, using the mean, median, or interpolation to fill missing values; Outlier handling: Detecting and handling outliers in numerical data, using box plots or Z-scores to detect outliers; Data standardization: scaling data to the same range.

6. The auxiliary diagnostic information retrieval method based on a respiratory knowledge base according to claim 1, characterized in that, The step of retrieving auxiliary diagnostic information first involves preprocessing, which includes the following steps: Embedding vectorization: Convert keywords and entities in the knowledge graph, as well as the relationships between entities, into vector representations to generate embedding vectors, including keyword vectors and entity vectors; Similarity calculation: Calculate the similarity between keyword vectors and entity vectors in the knowledge graph.

7. The auxiliary diagnostic information retrieval method based on a respiratory knowledge base according to claim 1, characterized in that, The query construction includes the following steps: using a pre-trained NER model to identify entities from patient data, using a pre-trained RE model to extract relationships between entities, and generating query statements based on entities and relationships.

8. The auxiliary diagnostic information retrieval method based on a respiratory knowledge base according to claim 1, characterized in that, The query decomposition includes the following steps: parsing the query statement using a syntax analyzer, identifying the query target and constraints, and decomposing the query according to the query target and constraints.

9. The auxiliary diagnostic information retrieval method based on a respiratory knowledge base according to claim 1, characterized in that, The retrieval process includes the following steps: for retrieval of structured medical knowledge bases, queries are performed in the RDF triple database using the SPARQL query language; for retrieval of unstructured medical knowledge bases, queries are performed using a vector similarity search algorithm. The result fusion step includes the following steps: Calculate the correlation and cluster the results based on the correlation, filtering out duplicate or irrelevant results.

10. A respiratory knowledge base-based auxiliary diagnostic information retrieval system, characterized in that, The auxiliary diagnostic information retrieval method based on a respiratory knowledge base according to any one of claims 1-9 includes a data acquisition module, a data preprocessing module, and a knowledge base retrieval module; The data acquisition module is used to collect respiratory system-related data from patients; The data preprocessing module is used to preprocess different types of collected data using different preprocessing methods; The knowledge base retrieval module is used to retrieve relevant diagnostic information from the medical knowledge base based on the patient's respiratory system-related data. The retrieval process prioritizes searching for relevant entries in the medical knowledge base. If no relevant entries are found in the medical knowledge base, the large model is invoked to perform auxiliary diagnostic information reasoning, and the knowledge base retrieval process described in step S1 is executed. Otherwise, the auxiliary diagnostic information from the medical knowledge base is returned.