Rapid filing method and system based on semantic extraction
Through fast file building methods and systems based on semantic extraction, feature extraction, difference detection and knowledge graph construction of multi-source medical data is solved, and the problems of low efficiency and insufficient accuracy of medical data processing in the existing technology are achieved, and efficient and accurate data utilization and diagnostic support are achieved.
Patent Information
- Application Number
- CN202510175654.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively process multi-source heterogeneous medical data, especially in the case of inconsistent data quality, inconsistent formats, and redundant information, resulting in low data utilization and insufficient diagnostic accuracy and therapeutic decision support capabilities.
A rapid file building method and system based on semantic extraction is adopted to generate complete lesion descriptions and disease analysis to improve data utilization and diagnostic accuracy through technical means such as feature extraction, difference detection, knowledge graph construction and automated typesetting of standardized data.
It realizes efficient collection, cleaning and utilization of multi-source medical data, improves data quality and utilization, enhances diagnostic accuracy and treatment decision support capabilities, and reduces the need for manual review.
Smart Images

Figure CN120148720A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semantic extraction, and particularly to a fast filing method and system based on semantic extraction. Background Art
[0002] With the continuous progress of modern medical technology, the medical data generated in the process of clinical diagnosis and treatment has shown explosive growth. This data includes various forms such as medical record texts, medical images, test reports, treatment records, etc. However, while the diversity and complexity of this data improve the accuracy of medical diagnosis and the ability to support treatment decisions, they also bring a series of problems such as inconsistent data quality, non-uniform formats, and information redundancy. High-quality and standardized medical data is crucial for the scientificity and accuracy of disease diagnosis, treatment decision-making, and health management. Especially in the medical scenarios driven by artificial intelligence, the quality of data directly affects the performance and credibility of intelligent diagnosis systems. Therefore, how to efficiently collect, clean, and utilize multi-source heterogeneous medical data has become an important problem to be solved urgently.
[0003] In the Chinese invention patent with the application publication number of CN116821199A, a traditional Chinese medicine case data extraction system is disclosed, including a metadata module, a corpus, a data collection module, a data verification module, a data decomposition module, a case filing module, and a query module; the metadata module is used to set entity classes, dictionaries, and semantic relationships and perform maintenance; the corpus is used to form semi-structured documents according to the imported literature; the data collection module collects data for existing cases; the data verification module verifies the data collected by the data collection module. By adopting the method of image information collection, the collection rate of case text information is improved, and before extracting data, a standardized metadata module and corpus are cited to standardize the collected data, improving the analysis efficiency of case data and ensuring the standardization of case filing.
[0004] Since the sources of medical data are extensive and the formats are diverse, including unstructured texts, medical images, and structured test data, etc., these data are often accompanied by problems such as information loss, noise interference, and non-standard data formats during the collection process, resulting in uneven data quality. In this case, the diversity and semantic ambiguity of medical terms make the processing of text data complicated. Especially in cross-domain semantic analysis, spelling correction, and term standardization, the prior art lacks a systematic solution. More importantly, in the fusion analysis of multi-modal data, it is difficult to achieve unified representation and processing of semantic conflicts or logical inconsistencies between text features and image features through traditional methods, further increasing the difficulty of data mining and analysis.
[0005] Therefore, the present invention provides a fast filing method and system based on semantic extraction. Summary of the Invention
[0006] (1) Technical problems to be solved
[0007] In view of the deficiencies of the prior art, the present invention provides a fast filing method and system based on semantic extraction. By respectively extracting features from the normalized data, a complete lesion description is generated; the text features and image features after screening are subjected to difference detection, and the difference degree is constructed from the obtained difference dimension data. According to the relationship between the difference degree and the difference threshold, corresponding difference removal response strategies are adopted; based on knowledge and data, inference rules are constructed and then a disease knowledge graph is constructed, and the extracted patient features are matched with the rules in the knowledge graph to obtain disease analysis results; structured data is collected to draw a symptom score trend graph, the key elements of the event are extracted from the patient records, and an electronic medical record document generated by an automated typesetting tool; the unstructured text is converted into structured data to improve the utilization rate of the data and facilitate further analysis; thus, the technical problems raised in the background art are solved.
[0008] (2) Technical solutions
[0009] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0010] A fast filing method based on semantic extraction includes collecting multi-source disease data and then performing quality detection. The quality degree Qto is constructed from the obtained quality detection data. If the quality degree does not exceed the expectation, the multi-source disease data is normalized and the normalized data is obtained;
[0011] Feature extraction is respectively performed on the normalized data, key features are screened according to the closeness centrality, and the text features and image features among the remaining features after screening are combined to generate a complete lesion description;
[0012] Difference detection is performed on the text features and image features after screening, and the difference degree Cyo is constructed from the obtained difference dimension data. According to the relationship between the difference degree Cyo and the difference threshold, corresponding difference removal response strategies are adopted, and the terms in the lesion description are mapped to standard terms;
[0013] Based on knowledge and data, inference rules are constructed and then a disease knowledge graph is constructed, and the extracted patient features are matched with the rules in the knowledge graph to obtain disease analysis results;
[0014] Structured data is collected to draw a symptom score trend graph, the key elements of the event are extracted from the patient records, and an electronic medical record document generated by an automated typesetting tool is encrypted and corresponding access permissions are set; among them, the update frequency f(t) of the electronic medical record document is constrained, and the constraint method is as follows:
[0015]
[0016] Where: r i is the risk level of the i-th risk node, w i (t) is the weighting factor related to the i-th risk node, s(t) is the symptom severity at time t, and λ max is the upper limit of the maximum update frequency.
[0017] Furthermore, multi-source data collection is carried out on the condition data, including medical record texts, medical images, and test data;
[0018] The trained data quality evaluation model is used to evaluate the quality of the condition data obtained from each data source, and the corresponding data quality scores are output. The quality degree Qto is generated from the data quality scores. If the obtained quality degree Qto does not exceed the preset quality threshold, a data processing instruction is sent to the outside.
[0019] Furthermore, after receiving the data processing instruction, for cleaning the non-medical related content in the text data using a medical NLP tool, the key statements after cleaning are obtained; taking the extracted key statements as input, the trained spelling correction model is used to correct spelling mistakes and replace synonyms, and missing value filling is performed after unifying the units in the test data. After combination, the corresponding standardized text is obtained;
[0020] The trained defect detection and recognition model is used to detect defects in the image data, and the corresponding defect features are obtained; according to the defect features, the corresponding defect optimization strategy is matched by the image defect optimization library; the defect optimization strategy is executed to optimize the image, and the image resolution and pixel pitch are standardized to obtain the standardized medical image.
[0021] Furthermore, taking the standardized text as input, the trained NLP model is used to extract key features to obtain text disease features; taking the standardized image as input, the trained deep learning segmentation model is used to automatically segment the lesion area and extract image disease features such as shape, area, and boundary;
[0022] Similarity analysis is performed between different disease features of the same type respectively to obtain the similarity between two different disease features;
[0023] Closeness centrality analysis is performed on the obtained several similarity data, and the disease features are labeled with the obtained closeness centrality. The disease features with closeness centrality exceeding the centrality threshold are used as the remaining features.
[0024] Furthermore, the screened text features and image features are mapped to the same semantic space, and logical conflict detection is performed to obtain the corresponding difference dimension data. The difference degree Cyo is constructed from the difference dimension data,
[0025] If the difference degree Cyo is lower than the difference threshold, correct the difference data according to the pre-set correction rules;
[0026] If the difference degree Cyo is not lower than the difference threshold, send an alarm instruction to the outside and prompt for manual review. The alarm information includes detailed conflict information, difference analysis reports, conflict descriptions, and the degree of difference.
[0027] Furthermore, the method for constructing the difference degree Cyo from the difference dimension data is as follows:
[0028]
[0029] In the formula: Cyo represents the multi-dimensional difference degree in the time interval t 1 ,t 2 inside, w S (t), w L (t) is a time-dependent weight function, is the change rate of the size of the lesion at time t, ||L A (t) - L B (t)|| is the spatial position difference at time t, N A (t) and N B (t) are the property vectors of the lesion at time t.
[0030] Furthermore, construct a medical ontology for defining terms, classifications, and semantic relationships in the medical field. After using a word segmentation tool to split the text into individual words or phrases, perform stop word filtering to remove irrelevant vocabulary, and use the medical ontology or term library to map the terms in the lesion description to standard terms.
[0031] Furthermore, collect relevant domain knowledge and data, and transform the knowledge into structured information; for each disease or reasoning scenario, clarify the input variables and logical relationships, and represent the reasoning rules in the form of a decision table and define rule variables;
[0032] Based on knowledge and data, deduce the logical relationship between input variables and outputs, extract reasoning rules; and define priorities according to the importance or applicability of the rules, and assign probability weights to conditions when the rules involve uncertainty; construct a disease knowledge graph based on the reasoning rules.
[0033] Furthermore, use NLP technology to analyze the patient's chief complaint and extract keywords as disease characteristics, extract characteristics from laboratory tests or imaging results as examination characteristics, extract patient background characteristics from the patient's basic information, and combine them to obtain patient characteristics;
[0034] Match the extracted patient features with the rules in the knowledge graph, evaluate according to the importance of each condition in the rules, and use the inference engine to combine the knowledge graph and the extracted features to finally deduce possible analysis results.
[0035] Furthermore, extract time points and corresponding events from the patient's chief complaint, examination reports, and treatment records through natural language processing technology, automatically record them in the form of a time series, and establish a structured timeline data model;
[0036] Record the appearance of the patient's symptoms, examination results, and treatment intervention information in chronological order, store them as structured data, and automatically update them according to the gradually input diagnosis and treatment data.
[0037] Furthermore, collect the required structured data, including specific descriptions, time points, index values, and event types; use the trained disease evaluation model to evaluate symptoms with the structured data as input, obtain symptom scores, and when the symptom scores are higher than the preset symptom thresholds, send an alarm instruction to the outside and use the corresponding time nodes as risk nodes; draw a symptom score trend chart based on the distribution of time points and annotate it with structured data.
[0038] A fast medical record filing system based on semantic extraction, including a data specification unit that performs quality inspection after collecting multi-source disease data, constructs a quality degree Qto from the obtained quality inspection data, and if the quality degree does not exceed the expectation, performs normalization processing on the multi-source disease data and obtains the normalized data;
[0039] A feature screening unit that extracts features from the normalized data respectively, screens key features according to the closeness to the center, combines the text features and image features among the remaining features after screening, and generates a complete lesion description;
[0040] A difference detection unit that performs difference detection on the screened text features and image features, constructs a difference degree Cyo from the obtained difference dimension data, and adopts corresponding de-difference response strategies according to the relationship between the difference degree Cyo and the difference threshold, and maps the terms in the lesion description to standard terms;
[0041] A rule construction unit that constructs an inference rule based on knowledge and data and then constructs a disease knowledge graph, matches the extracted patient features with the rules in the knowledge graph, and obtains disease analysis results;
[0042] A document generation unit that draws a symptom score trend chart with the collected structured data, extracts the key elements of the events from the patient records, and generates an electronic medical record document using an automated typesetting tool, encrypts the patient's electronic medical record document and sets corresponding access permissions.
[0043] (III) Beneficial effects
[0044] The present invention provides a fast filing method and system based on semantic extraction, having the following beneficial effects:
[0045] 1. By generating a quality score and a quality degree Qto, low-quality data can be real-time warned and processed during the data collection stage, reducing diagnostic deviations caused by incomplete or incorrect data; the full-process automated operation from data cleaning to error correction can greatly reduce the need for manual review, while ensuring the improvement of efficiency and accuracy, unifying the units and filling in missing values for the inspection data, ensuring the consistency and integrity of numerical data, and providing more reliable numerical inputs for model training and result inference.
[0046] 2. By using a defect detection model to automatically detect defects in medical images, the abnormal areas in the images can be accurately located. Combining with an image optimization strategy library and automatically matching the best optimization scheme can specifically improve the quality of the image data.
[0047] 3. Extracting key information from standardized texts based on the trained NLP model can quickly identify disease-related features, avoid missing key diagnostic information, comprehensively improve the efficiency and accuracy of text feature extraction. The feature extraction process under multi-modal input greatly reduces the steps of manual participation, improving efficiency and consistency; combining text and image disease characteristics to form a comprehensive lesion description can generate a more data-driven disease description.
[0048] 4. Automatically detecting semantic conflicts or data logic inconsistencies between the two, reducing the subjectivity and error rate of manual verification; through the construction of differential dimension data and the calculation of the difference degree, the differences between text and images in lesion descriptions can be quantitatively analyzed, providing a basis for quantitative monitoring of lesion development.
[0049] 5. Standardizing medical terms and performing term disambiguation through a context analysis model can effectively eliminate semantic ambiguity, ensuring the unity and accuracy of term understanding; through the context analysis model, medical terms with complex semantics can be further accurately parsed, greatly improving the model's processing ability for complex texts and semantic understanding level.
[0050] 6. Extracting key information from the patient's chief complaint and examination through NLP technology, combined with rule matching in the knowledge graph, can quickly perform a likelihood assessment. The disease feature assessment mechanism based on weight calculation can quantify the disease risk.
[0051] 7. Encrypting the patient's file through AES or RSA to ensure the security of sensitive data. At the same time, setting access permissions to avoid data leakage. Combining with an automatic typesetting tool and a structured data model, the generated electronic medical record not only improves the standardization of the document, but also greatly saves the doctor's recording time and improves the efficiency of medical documentation work. Brief Description of the Drawings
[0052] Figure 1 It is a schematic flowchart of the fast filing method based on semantic extraction of the present invention;
[0053] Figure 2 It is a schematic structural diagram of the fast filing system based on semantic extraction of the present invention. Detailed Embodiments
[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] Please refer to Figure 1 , the present invention provides a fast filing method based on semantic extraction, including,
[0056] Step 1: After collecting multi-source disease data, perform quality detection, construct a quality degree Qto from the obtained quality detection data. If the quality degree does not exceed the expectation, perform normalization processing on the multi-source disease data and obtain the normalized data;
[0057] The said Step 1 includes the following contents:
[0058] Step 101: Conduct multi-source data collection on disease data, including medical record texts, medical images, and test data, etc.; train a machine learning algorithm with the labeled sample data to obtain a trained data quality evaluation model;
[0059] Use the trained data quality evaluation model to evaluate the quality of the disease data obtained from each data source, output the corresponding data quality score, and generate the quality degree Qto from the data quality score in the following way:
[0060]
[0061] Weight coefficient: 0 ≤ S 1 ≤ 1, 0 ≤ S 2 ≤ 1, 0 ≤ S 3 ≤ 1 and S 3 + S 2 + S 1 = 1, the weight coefficient can be obtained by referring to the entropy value method; Gs i is the quality score of the medical record text in the i-th stage, Gs a is the corresponding mean value, Ts i is the quality score of the medical image in the i-th stage, Ts a is the corresponding mean value, Jsi is the quality score of the inspection data for the i-th stage, Js a is the corresponding mean value;
[0062] According to historical data and the management expectations for data quality, a quality threshold is set in advance; if the obtained quality degree Qto does not exceed the preset quality threshold, it indicates that the current data quality is poor, and it is necessary to process the collected multi-source data and send a data processing instruction to the outside;
[0063] When in use, by collecting multi-source data for the condition data (such as medical records text, medical images, inspection data, etc.) and combining with the data quality evaluation model, the accuracy, integrity, and reliability of the data can be systematically evaluated to ensure high-quality data is input into the subsequent analysis process and improve data quality and credibility: by generating a quality score and a quality degree Qto, and combining with the setting of the quality threshold, low-quality data can be real-time warned and processed during the data collection stage, reducing diagnostic deviations caused by incomplete or incorrect data;
[0064] Step 102, after receiving the data processing instruction, for cleaning non-medical related content in the text data using a medical NLP tool to obtain the cleaned key sentences; for example, using a word segmentation tool to cut the text into independent words, using a medical stop word list to filter out irrelevant words (such as non-medical information like "de", "le", "shi", etc.), and using a domain-specific dictionary (such as SNOMED, ICD coding table) to retain medical-related keywords; using the TF-IDF (Term Frequency - Inverse Document Frequency) method to calculate the importance of each word to complete the extraction of key sentences;
[0065] The trained spelling correction model is obtained by training the labeled sample data with a machine learning algorithm combined with a custom medical term dictionary; using the extracted key sentences as input, the trained spelling correction model is used to correct spelling mistakes and replace synonyms, such as replacing "lump" with "nodule", and after unifying the units in the inspection data, the missing values of the numerical data are predicted using a machine learning model and then filled, and the corresponding standardized text is obtained after combination;
[0066] When in use, by cleaning non-medical related content in the text using a medical NLP tool and combining with a medical stop word list and a domain dictionary (such as SNOMED, ICD coding table), the effective key information related to diagnosis and treatment can be maximally extracted, and noise can be eliminated from the massive data; through the spelling correction model and TF-IDF weight calculation, the standardization of medical terms is further realized, ensuring semantic consistency between different medical text data and improving the generalization ability of the model;
[0067] Full-process automated operations from data cleaning to error correction significantly reduce the need for manual review while ensuring improved efficiency and accuracy. Standardize units and fill in missing values for inspection data to ensure the consistency and integrity of numerical data, providing a more reliable numerical input for model training and result inference;
[0068] Step 103: Use the trained defect detection and recognition model to perform defect detection on the image data to obtain corresponding defect features; Based on the correspondence between the defect features and the image optimization strategy, match the corresponding defect optimization strategy from the pre-constructed image defect optimization library; Execute the defect optimization strategy to optimize the image, and standardize the image resolution and pixel pitch to obtain a normalized medical image;
[0069] When in use, combine the content in Steps 101 to 103:
[0070] Automated defect detection of medical images through the defect detection model can accurately locate abnormal areas in the image, reducing the doctor's image review burden. Combining with the image optimization strategy library to automatically match the best optimization plan (such as adjusting resolution, pixel pitch, etc.) can specifically improve the quality of image data.
[0071] Step Two: Extract features from the standardized data respectively, screen the key features based on the proximity to the center, and combine the text features and image features among the remaining features after screening to generate a complete lesion description;
[0072] The said Step Two includes the following content:
[0073] Step 201: Use the trained NLP model to extract key features with the standardized text as the input to obtain text disease features; Use the trained deep learning segmentation model to automatically segment the lesion area with the standardized image as the input to extract image disease features such as shape, area, and boundary;
[0074] When in use, extracting key information (such as symptoms, time, location, etc.) from the standardized text based on the trained NLP model can quickly identify disease-related features, avoid missing key diagnostic information, comprehensively improve the efficiency and accuracy of text feature extraction, and the feature extraction process under multi-modal input significantly reduces the steps of manual participation, improving efficiency and consistency.
[0075] Step 202: Conduct similarity analysis among different disease characteristics of the same type to obtain the similarity between two different disease characteristics. Perform closeness centrality analysis on the obtained similarity data to obtain the corresponding closeness centrality. Label the disease characteristics with the closeness centrality. Use the disease characteristics with closeness centrality exceeding the centrality threshold as the remaining characteristics. Combine the text characteristics and image characteristics in the remaining characteristics after screening to generate a complete lesion description.
[0076] During use, combine the content in Steps 201 and 202:
[0077] The similarity analysis and feature centrality screening mechanism effectively remove redundant or superfluous feature information, making the remaining features more accurate and important, further reducing the analysis dimension and computational overhead. Combining text and image disease characteristics to form a comprehensive lesion description helps provide more comprehensive disease information. Based on the results of feature centrality and similarity analysis, a more data-driven disease description can be generated to support further clinical decision-making.
[0078] Step 3: Conduct difference detection on the screened text characteristics and image characteristics, construct a difference degree Cyo from the obtained difference dimension data, and adopt corresponding difference removal response strategies according to the relationship between the difference degree Cyo and the difference threshold, and map the terms in the lesion description to standard terms.
[0079] The above Step 3 includes the following content:
[0080] Step 301: Map the screened text characteristics and image characteristics to the same semantic space, conduct logical conflict detection, compare the similarity of their feature values or descriptive semantics, obtain the corresponding difference dimension data, including lesion size, location, nature, development trend, etc., and construct the difference degree Cyo from the difference dimension data in the following way:
[0081]
[0082] In the formula: Cyo represents the multi-dimensional difference degree in the time interval t 1 , t 2 within, w S (t), w L (t) is a time-dependent weight function that can be dynamically adjusted according to the importance of each dimension in different time periods. is the change rate of the lesion size at time t, ||L A (t) - L B (t)|| is the spatial position difference at time t, N A (t) and N B (t) are the property vectors of lesion A and lesion B at time t;
[0083] Set a difference threshold in advance according to historical data and the management expectations of data differences.
[0084] If the obtained difference degree Cyo is lower than the difference threshold, correct the difference data according to the pre-set correction rules; for example, if the image shows a slight shadow but the text is normal, the conclusion is corrected to "normal".
[0085] If the obtained difference degree Cyo is not lower than the difference threshold, send an alarm instruction to the outside and prompt for manual review. The alarm information includes detailed conflict information, difference analysis reports, conflict descriptions, and the degree of difference, etc.
[0086] During use, after mapping text and image features to the same semantic space, automatically detect semantic conflicts or data logic inconsistencies between the two, reducing the subjectivity and error rate of manual verification; through the construction of difference dimension data and the calculation of the difference degree, the differences between text and image in lesion descriptions can be quantitatively analyzed (such as lesion size, location, nature, etc.), providing a basis for quantitative monitoring of lesion development. Through the setting of the difference degree threshold and correction rules, slightly contradictory data can be automatically adjusted or corrected, and an alarm instruction can be issued for situations with a higher difference degree, promptly prompting for manual review to ensure the scientificity and accuracy of the diagnosis conclusion.
[0087] Step 302: Construct a medical ontology for defining terms, classifications, and semantic relationships in the medical field, specifically including concepts, terms, relationships, and attributes, etc.; after using a word segmentation tool to split the text into individual words or phrases, perform stop word filtering to remove irrelevant vocabulary, and use the medical ontology or term library to map the terms in the lesion description to standard terms; if a term corresponds to multiple meanings, use a context analysis model (such as BioBERT, SciBERT) to extract context information related to the term (such as "lung" or "thyroid"), and then determine the specific meaning to achieve the effect of term disambiguation.
[0088] During use, combine the content in Steps 301 and 302:
[0089] Standardize medical terms using a medical ontology and term library (such as SNOMED or UMLS), and perform term disambiguation through a context analysis model, which can effectively eliminate semantic ambiguity and ensure the unity and accuracy of term understanding; through the context analysis model, medical terms with complex semantics can be further accurately parsed, greatly improving the model's processing ability and semantic understanding level for complex texts.
[0090] Step Four: After constructing inference rules based on knowledge and data, construct a disease knowledge graph, match the extracted patient features with the rules in the knowledge graph, and obtain disease analysis results.
[0091] Step 4 includes the following content:
[0092] Step 401: Before starting to build the inference rules, clarify the goals to be achieved and the scope of reasoning, including problem description, reasoning scope, and target output, etc.;
[0093] Collect relevant domain knowledge and data. The knowledge sources can be domain experts, literature and guidelines, and data records, etc.; transform the knowledge into structured information; for each disease or reasoning scenario, clarify the following elements: Input variables: The input data on which the inference rules depend, such as symptoms, examination results, and basic patient information; Output variables: The conclusions or action suggestions to be drawn by the inference rules; Logical relationship: The causal relationship or conditional dependence between the input and the output;
[0094] The inference rules are represented in the form of a decision table and define rule variables, including determining the value range of the variables: discrete values (such as "with / without symptoms") or continuous values (such as "body temperature > 38°C"), and at the same time, the variables need to be observable and measurable;
[0095] Based on knowledge and data, deduce the logical relationship between the input variables and the output, and extract clear inference rules; and define priorities according to the importance or applicability of the rules. For example, the inference rules for tuberculosis take precedence over those for the common cold; when the rules involve uncertainty, probability weights need to be assigned to the conditions;
[0096] When in use, after clarifying the inference goals and logical relationships, it can make the inference process more standardized and targeted, greatly improving the reliability of the inference engine in diagnosis and decision-making. Defining the inference rules in the form of a decision table can clearly present the input, output, and their logical relationships, facilitating the update, optimization, and expansion of the rules. By assigning probability weights to the conditions to handle uncertainty problems, for example, combining the patient's disease characteristics and probability reasoning, it can flexibly handle complex medical scenarios.
[0097] Step 402: Build a disease knowledge graph based on the inference rules, which mainly includes the following steps:
[0098] Knowledge acquisition: Collect medical knowledge from multiple data sources, including medical literature (such as PubMed), medical guidelines (such as SNOMED-CT), electronic health records (EHR), etc. The data types include entities such as diseases, symptoms, drugs, and examination methods, as well as the relationships between them (such as "causes", "manifested as").
[0099] Data cleaning and normalization: Clean and normalize the collected data. Remove redundant and incorrect information, and map the text to standard medical terms (such as UMLS or SNOMED). For example, "pneumonia" and "lung infection" can be normalized to the same term.
[0100] Graph database modeling: Represent knowledge as a graph structure, where entities are nodes and relationships are edges. For example, the "manifested as" relationship between the "tuberculosis" node and the "cough" node. Such a visualized map structure facilitates storage and reasoning;
[0101] Knowledge storage: Utilize graph databases (such as Neo4j or GraphDB) to store the knowledge graph. These tools support complex relationship queries and are suitable for processing large-scale medical data.
[0102] Data supplementation and reasoning support: Represent nodes and relationships in the graph through embedding models (such as Node2Vec or TransE) to provide support for subsequent reasoning.
[0103] When in use, the knowledge graph comprehensively integrates heterogeneous medical data (such as diseases, symptoms, drugs, examination methods, etc.) in the form of a graph structure, which can provide a unified data interface for large-scale medical research and clinical practice. Based on the storage architecture of the knowledge graph, it can support real-time knowledge updates. By combining embedding models (such as Node2Vec) to represent the nodes and relationships of the knowledge graph, the logical inference efficiency and accuracy of the subsequent inference engine can be optimized;
[0104] Step 403: After using NLP technology to analyze the patient's chief complaint, extract keywords as disease characteristics, extract characteristics from laboratory tests or imaging results as examination characteristics, extract the patient's background characteristics from the patient's basic information, and combine them to obtain the patient's characteristics;
[0105] Match the extracted patient characteristics with the rules in the knowledge graph, and evaluate according to the importance of each condition in the rules. For example, the weight of cough is 80%, the weight of hemoptysis is 50%, and the weight of chest X-ray shadow is 90%. The calculation result shows that the possibility of having tuberculosis is 36%. Use an inference engine (such as Neo4j or Drools) to combine the knowledge graph and the extracted characteristics to finally deduce the possible analysis results;
[0106] When in use, combine the content in Steps 401 to 403:
[0107] Extract key information from the patient's chief complaint and examination through NLP technology, and combine it with the rule matching in the knowledge graph to quickly conduct a possibility assessment. Based on the disease characteristic assessment mechanism calculated by weight, the disease risk (such as the percentage of possibility) can be quantified, significantly improving the scientific nature and quantification ability of diagnosis;
[0108] Step Five: Draw a symptom score trend chart with the collected structured data, extract the key elements of the event from the patient's records, and generate an electronic medical record document using an automated typesetting tool. Encrypt the patient's electronic medical record document and set corresponding access permissions;
[0109] The above Step Five includes the following content:
[0110] Step 501: Extract time points and corresponding events from the patient's chief complaint, examination reports, and treatment records through natural language processing (NLP) technology, automatically record them in the form of a time series, and establish a structured timeline data model; record information such as the onset of the patient's symptoms, examination results, and treatment interventions in chronological order, store them as structured data, and automatically update them as the medical data is gradually input;
[0111] When in use, systematically record the whole process of the patient from the initial diagnosis to the treatment, provide intuitive and complete data support for the analysis of disease progression, and the automatic update mechanism ensures the timeliness of the medical record data. For example, it can quickly reflect the changes in the condition after treatment, facilitating doctors to adjust the treatment strategy in a timely manner;
[0112] Step 502: Collect the required structured data, including specific descriptions, time points, the occurrence dates of each event or record; index values, quantitative data or category information reflecting the dynamic condition; event types, such as symptoms, examinations, and treatments;
[0113] Train a machine learning algorithm with the labeled sample data to obtain a trained disease evaluation model;
[0114] Use the trained disease evaluation model to conduct symptom assessment with the structured data as the input, obtain the symptom score. When the symptom score is higher than the preset symptom threshold, send an alarm instruction to the outside and use the corresponding time node as a risk node; draw a symptom score trend chart based on the distribution of time points and label it with structured data;
[0115] When in use, by drawing the symptom score trend chart, it can visually display the changes in the patient's condition, quickly identify the critical time points of treatment effects or disease deterioration, convert unstructured text into structured data, improve the utilization rate of data, and facilitate further analysis or modeling by machine learning models;
[0116] Step 503: Extract the key elements of the event from the patient record, including time points, event types (such as symptoms, examinations, treatments), and event descriptions, and generate an electronic medical record document using an automatic typesetting tool. Encrypt the patient file using AES or RSA and set the corresponding access permissions;
[0117] Restrict the update frequency of the electronic medical record document in the current stage according to the distribution of risk nodes and the symptom risk level. The restriction method is as follows:
[0118]
[0119] Where: f(t) is the update frequency of the electronic medical record at time t, r i$r_i$ represents the risk level of the $i$-th risk node, indicating the severity of the patient's condition when the node is triggered. i $r_i\in[0,1]$, where $r_i = 1$ represents the highest risk; i $w_i(t)$ is the weighting factor associated with the $i$-th risk node, indicating the weighted impact on this node at time point $t$. $s(t)$ is the severity of symptoms at time $t$, and $s(t)\in[0,1]$, where $s(t)=1$ represents the most severe symptoms. $\lambda$ i is the upper limit of the maximum update frequency; max
[0120] Update the electronic medical record document on the update nodes that meet the constraints, so as to adapt its update frequency to the current actual situation and reduce ineffective updates;
[0121] When in use, combine the content in steps 501 to 503:
[0122] Encrypt the patient file through AES or RSA to ensure the security of sensitive data. At the same time, set access permissions to avoid data leakage. Combining an automatic typesetting tool and a structured data model, the generated electronic medical record not only improves the standardization of the document, but also greatly saves the doctor's recording time and improves the efficiency of medical documentation work.
[0123] Please refer to Figure 2 , the present invention provides a fast medical record filing system based on semantic extraction, including
[0124] A data specification unit, which performs quality detection after collecting multi-source disease data, constructs a quality degree $Q_{to}$ from the obtained quality detection data. If the quality degree does not exceed the expectation, normalize the multi-source disease data and obtain the normalized data;
[0125] A feature screening unit, which respectively extracts features from the normalized data, screens the key features according to the proximity centrality, and combines the text features and image features in the remaining features after screening to generate a complete lesion description;
[0126] A difference detection unit, which performs difference detection on the screened text features and image features, constructs a difference degree $C_{yo}$ from the obtained difference dimension data, and adopts corresponding difference elimination response strategies according to the relationship between the difference degree $C_{yo}$ and the difference threshold, and maps the terms in the lesion description to standard terms;
[0127] A rule construction unit, which constructs an inference rule based on knowledge and data and then constructs a disease knowledge graph, matches the extracted patient features with the rules in the knowledge graph, and obtains the disease analysis result;
[0128] A document generation unit that collects structured data to draw a symptom score trend chart, extracts key elements of events from patient records, and generates an electronic medical record document using an automated typesetting tool, encrypts the patient's electronic medical record document and sets corresponding access permissions.
[0129] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0130] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0131] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only for some logical function divisions, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings, direct couplings, or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0132] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0133] As described above, only the specific implementation manners of this application are provided, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A rapid archiving method based on semantic extraction, characterized by: include, After collecting multi-source disease data, quality inspection is performed, and the quality degree Qto is constructed based on the acquired quality inspection data. If the quality degree does not exceed expectations, the multi-source disease data is standardized and the standardized data is obtained; The normalized data were subjected to feature extraction, key features were screened based on the proximity to the center, and the text features of the remaining features after screening were combined with the image features to generate a complete lesion description; Perform difference detection on the screened text features and image features, and construct the difference degree Cyo from the obtained difference dimension data. According to the relationship between the difference degree Cyo and the difference threshold, adopt the corresponding de-difference response strategy to map the terms in the lesion description to standard terms. After building inference rules based on knowledge and data, a disease knowledge graph is constructed, and the extracted patient features are matched with the rules in the knowledge graph to obtain disease analysis results; The symptom score trend chart is drawn by collecting structured data, extracting the key elements of the event from the patient record, and using the electronic medical record document generated by the automatic typesetting tool to encrypt the patient's electronic medical record document and set the corresponding access rights; the update frequency f(t) of the electronic medical record document is constrained as follows: Where: r i is the risk level of the ith risk node, w i (t) is the weighting factor associated with the i-th risk node, s(t) is the severity of the symptoms at time t, and λ max The upper limit of the maximum update frequency.
2. The rapid archiving method based on semantic extraction according to claim 1 is characterized in that: Collect multi-source data on medical conditions, including medical records, medical images, and test data; The trained data quality evaluation model is used to evaluate the quality of the medical data obtained from various data sources, and the corresponding data quality score is output. The quality degree Qto is generated from the data quality score. If the obtained quality degree Qto does not exceed the preset quality threshold, a data processing instruction is issued to the outside.
3. The rapid archiving method based on semantic extraction according to claim 2 is characterized in that: After receiving the data processing instruction, the non-medical related content in the text data is cleaned using the medical NLP tool to obtain the cleaned key sentences; the extracted key sentences are used as input, the spelling error correction model after training is used to correct spelling errors and synonym replacement, and the missing values are filled after the unit of the test data is unified, and the corresponding standardized text is obtained after combination; Use the trained defect detection and recognition model to perform defect detection on the image data and obtain the corresponding defect features; According to the defect characteristics, the corresponding defect optimization strategy is matched from the image defect optimization library; Execute defect optimization strategies to optimize images, standardize image resolution and pixel spacing, and obtain standardized medical images.
4. The rapid archiving method based on semantic extraction according to claim 3 is characterized in that: Taking normalized text as input, the trained NLP model is used to extract key features and obtain text disease characteristics; taking normalized images as input, the trained deep learning segmentation model is used to automatically segment the lesion area and extract image disease characteristics; Perform similarity analysis on different disease characteristics of the same type to obtain the similarity between two different disease characteristics; A proximity centrality analysis is performed on the obtained similarity data, and the disease features are marked with the obtained proximity centrality, and the disease features whose proximity centrality exceeds the centrality threshold are taken as the remaining features.
5. The rapid archiving method based on semantic extraction according to claim 4 is characterized in that: Map the filtered text features and image features to the same semantic space, perform logical conflict detection, obtain the corresponding difference dimension data, and construct the difference degree Cyo from the difference dimension data; If the difference degree Cyo is lower than the difference threshold, the difference data is corrected according to the preset correction rules; If the difference Cyo is not lower than the difference threshold, an alarm command is issued to the outside and a manual review is prompted. The alarm information contains detailed conflict information and difference analysis report, conflict description and difference degree.
6. The rapid archiving method based on semantic extraction according to claim 5 is characterized in that: The way to construct the difference degree Cyo from the difference dimension data is as follows: Where: Cyo represents the multi-dimensional difference degree in the time interval [t1, t2], w S (t),w L (t) is the time-dependent weight function, is the rate of change of the size of the lesion at time t, ||L A (t)-L B (t)|| is the spatial position difference at time t, N A (t) and N B (t) is the property vector of the lesion at time t.
7. The rapid archiving method based on semantic extraction according to claim 6 is characterized in that: Construct a medical ontology to define the terminology, classification, and semantic relationships in the medical field. Use a word segmentation tool to split the text into individual words or phrases, perform stop word filtering to remove irrelevant words, and use a medical ontology or term library to map the terms in the lesion description to standard terms.
8. The rapid archiving method based on semantic extraction according to claim 7 is characterized in that: Collect relevant domain knowledge and data, and transform knowledge into structured information; for each disease or reasoning scenario, clearly define input variables and logical relationships, express reasoning rules in the form of decision tables, and define rule variables; Based on knowledge and data, the logical relationship between input variables and outputs is derived and reasoning rules are extracted. Priorities are defined based on the importance or applicability of the rules, and probability weights are assigned to conditions when the rules involve uncertainty. A disease knowledge graph is constructed based on the reasoning rules.
9. The rapid archiving method based on semantic extraction according to claim 8 is characterized in that: Use NLP technology to analyze the patient's chief complaint and extract keywords as disease features, extract features from laboratory tests or imaging results as examination features, extract patient background features from the patient's basic information, and combine them to obtain patient features; The extracted patient features are matched with the rules in the knowledge graph, and the importance of each condition in the rule is evaluated. The reasoning engine is used to combine the knowledge graph and the extracted features to finally derive possible analysis results.
10. The rapid archiving method based on semantic extraction according to claim 9 is characterized in that: Through natural language processing technology, time points and corresponding events are extracted from patient complaints, examination reports and treatment records, and automatically recorded in the form of time series to establish a structured timeline data model; The patient's symptom onset, examination results and treatment intervention information are recorded in chronological order and stored as structured data, which is automatically updated based on the gradual input of diagnosis and treatment data.
11. The rapid archiving method based on semantic extraction according to claim 10, characterized in that: Collect the required structured data, including specific descriptions, time points, indicator values, and event types; use the structured data as input, use the trained disease evaluation model to perform symptom assessment, and obtain the symptom score. If the symptom score is higher than the preset symptom threshold, issue an alarm command to the outside and use the corresponding time node as a risk node; draw a symptom score trend chart based on the distribution of time points and annotate it with structured data.
12. Rapid archiving system based on semantic extraction, characterized by: include, The data standardization unit collects multi-source disease data and performs quality inspection. The quality degree Qto is constructed based on the acquired quality inspection data. If the quality degree does not exceed the expectation, the multi-source disease data is standardized and the standardized data is obtained. The feature screening unit extracts features from the normalized data, screens key features based on the proximity to the center, combines the text features and image features in the remaining features after screening, and generates a complete lesion description; The difference detection unit performs difference detection on the screened text features and image features, and constructs the difference degree Cyo from the acquired difference dimension data. According to the relationship between the difference degree Cyo and the difference threshold, a corresponding de-difference response strategy is adopted to map the terms in the lesion description into standard terms. The rule construction unit builds the disease knowledge graph after building the inference rules based on knowledge and data, matches the extracted patient features with the rules in the knowledge graph, and obtains the disease analysis results; The document generation unit collects structured data to draw symptom score trend charts, extracts key elements of events from patient records, and uses electronic medical record documents generated by automated typesetting tools to encrypt patients' electronic medical record documents and set corresponding access rights.
Citation Information
Patent Citations
Traditional Chinese medicine case data extraction system and method
CN116821199A
Cited By
RCT literature labeling method and system based on model consensus
CN120429756A