Structured extraction method and system for multi-format medical record text based on configuration rule
Patent Information
- Application Number
- CN202610952633.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-06-30
AI Technical Summary
[0007]鉴于上述的分析,本发明实施例旨在提供一种基于配置化规则的多格式病历文本结构化提取方法及系统,用以解决现有病历文本结构化提取方法在面对多格式、低质量数据时提取鲁棒性差、规则维护困难、缺乏可信度保障的问题
[0018]与现有技术相比,本发明至少可实现如下有益效果之一:
Smart Images

Figure CN122472026B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical information technology and data processing technology, and in particular to a method and system for structured extraction of multi-format medical record text based on configuration rules. Background Technology
[0002] In current medical informatics practices, the automated structured extraction of medical record text data faces significant challenges, primarily due to the following bottlenecks: The data sources are heterogeneous in format and inconsistent in quality: they include various formats such as plain text clinical records, standard / non-standard XML, and non-standard HTML, and a large amount of historical data has problems such as missing tags, character encoding errors, and non-standard writing (such as writing variations and arbitrary spaces), which makes the general parser ineffective.
[0003] Rigid rules and maintenance difficulties: Traditional extraction methods based on hard coding or fixed templates require the development of dedicated code for each data format variant. When business rules change or new data sources need to be integrated, the modification, testing, and deployment cycles are long, costly, and extremely inflexible.
[0004] Poor robustness in extraction: Most existing methods rely on precise keywords or regular expression matching, which have poor tolerance for noise in the text (such as OCR errors, synonyms, and formatting tweaks), leading to omissions or mis-extractions of key fields.
[0005] Lack of credibility assurance and optimization loop: The extraction process is mostly a "black box", and the results lack effective verification; moreover, when there are problems with the extraction quality, it is difficult to quickly locate the root cause (whether it is a data problem, a rule problem or a model problem), and it is impossible to form a data-driven system self-optimization.
[0006] Therefore, there is an urgent need for a multi-format medical record text structure extraction method and system based on configuration rules, which can achieve intelligent tolerance of data defects, dynamic adaptation to diverse formats, and ensure the credibility of results in medical record text structure extraction. Summary of the Invention
[0007] Based on the above analysis, the embodiments of the present invention aim to provide a method and system for structured extraction of multi-format medical record text based on configurable rules, in order to solve the problems of poor robustness, difficulty in rule maintenance, and lack of credibility guarantee in existing medical record text structured extraction methods when facing multi-format and low-quality data.
[0008] On one hand, embodiments of the present invention provide a method for structured extraction of multi-format medical record text based on configuration rules, including: Step S1: Format identification is performed on the original medical record data, which includes several original document data. Based on the format identifier and / or character encoding range in each original document data, the data format corresponding to each original document data is determined. The original medical record data is repaired and mapped based on the data format to generate corresponding standard structure medical record data. The data format includes markup language format. Based on the data format, the original medical record data is repaired and mapped to generate corresponding standard structured medical record data, including: The original document data in the markup language format is subjected to a first fault-tolerant parsing to generate a corresponding primary parse tree and error records. Based on the primary parse tree and the error records, the original document data is subjected to a first intelligent repair to generate a corresponding well-structured document object. The first fault-tolerant parsing includes feature tag recognition and stack parsing of the original document data. The first intelligent repair includes tag completion and structure normalization. The well-structured document object is mapped and transformed to generate the corresponding standard structured medical record data; Step S2: Based on the extraction requirements, configure rules using a preset rule base to generate declarative rules; wherein, the declarative rules include anchor rules and extraction rules, the anchor rules include precise anchor rules and / or fuzzy fault-tolerant anchor rules; the preset rule base includes a keyword and variant knowledge base; Based on the extraction requirements, rules are configured using a pre-defined rule base to generate declarative rules, including: Based on the target words for extraction requirements and the keyword and variant knowledge base, configure and generate the anchor point rules; Based on the information extraction range of the extraction requirements, configure and generate the extraction rules; Step S3: Based on the precise anchor point rule and / or the fuzzy fault-tolerant anchor point rule, determine the extraction starting point position in the standard structured medical record data. Extract information from the extraction starting point position according to the extraction rule to obtain candidate field data and the corresponding data spectrum. Perform multi-level verification on the candidate field data to obtain the structured data package corresponding to the extraction requirement. The structured data package includes target field data, the data spectrum, and a verification report. The multi-level verification includes one or more of rule logic verification, standard terminology database verification, and lightweight NLP semantic verification.
[0009] Furthermore, it also includes: Step S4: Perform root cause analysis on the execution results of step S3, generate optimization recommendations, and update the declarative rules and / or the preset rule base according to the optimization recommendations; wherein, the preset rule base also includes a verification knowledge base and a verification model.
[0010] Furthermore, the data format also includes plain text format; repairing and mapping the original medical record data based on the data format to generate corresponding standard structured medical record data also includes: The original document data in the plain text format is subjected to a second fault-tolerant parsing and a second intelligent repair to generate a structured text object; The structured text object is mapped and transformed to generate the corresponding standard structured medical record data.
[0011] Further, the format recognition includes first format recognition and second format recognition; for original medical record data comprising several original document data, format recognition is performed based on the format identifier and / or character encoding range in each original document data to determine the data format corresponding to each original document data, including: For each of the original document data, the first format identification is performed based on the format identifier and / or character encoding range of the original document data. If the first format identification is successful, the data format corresponding to the original document data is determined based on the identification result of the first format identification. The first format identification includes identifier matching based on the format identifier of the original medical record data and encoding analysis based on the character encoding range of the original medical record data. If the first format recognition fails, then the original document data is subjected to data format feature extraction based on the format identifier of the original document data, and the data format features are input into a lightweight classification model for the second format recognition to determine the data format corresponding to the original document data.
[0012] Furthermore, the first fault-tolerant parsing includes syntax analysis and stack parsing; the first intelligent repair includes tag completion and structure normalization. The original document data in the markup language format undergoes a first fault-tolerant parsing process to generate a corresponding primary parse tree and error records. Based on the primary parse tree and the error records, the original document data undergoes a first intelligent repair process to generate a well-structured document object, including: A tolerant scanner is used to perform syntactic analysis on the original document data in the markup language format to identify the feature tags in the original document data; Based on the feature tags, the original document data is parsed using the stack parsing to generate the corresponding primary parse tree and error records; Based on the primary parse tree and the error records, the original document data is processed by tag completion and structural normalization to generate the well-structured document object.
[0013] Furthermore, the second fault-tolerant parsing includes encoding detection and error correction, and text segmentation; the second intelligent repair includes paragraph recombination and structural simulation. The original document data in the plain text format undergoes a second fault-tolerant parsing and a second intelligent repair to generate a structured text object, including: The original document data in the plain text format is subjected to encoding detection and error correction to generate a Unicode character stream; The Unicode character stream is divided into text segments based on newline characters to generate a sequence of text lines. The paragraph recombination is performed on the text line sequence to generate logical paragraphs; The structure of the logical paragraphs is simulated to generate the structured text object.
[0014] Furthermore, the declarative rules also include multi-level verification rules; The candidate field data is validated at multiple levels to obtain the structured data package corresponding to the extraction requirements, including: According to the multi-level verification rules, the candidate field data is subjected to multi-level verification, and the candidate field data that passes the multi-level verification is standardized to obtain the target field data and the corresponding verification report.
[0015] Furthermore, the fuzzy fault-tolerant anchor point rule includes fuzzy anchor points, anchoring distances, and anchoring ranges; Based on the aforementioned fuzzy fault-tolerant anchor point rule, the starting point position for extraction is determined in the standard structured medical record data, including: Based on the anchoring range, candidate nodes are determined from all text nodes of the standard structure medical record data, and the text distance between the fuzzy anchor point and the candidate node is calculated. If the text distance is not less than the anchoring distance, then the position of the candidate node in the standard structure medical record data is determined as the extraction starting point position.
[0016] Furthermore, based on the precise anchor point rule and the fuzzy fault-tolerant anchor point rule, the extraction starting point position is determined in the standard structured medical record data, including: Based on the precise anchor point rules, anchor points are located in the standard structured medical record data to determine the first starting point position; Based on the fuzzy fault-tolerant anchor point rule, anchor point positioning is performed on the standard structured medical record data to determine the second starting point position and the corresponding confidence level; The extraction starting point position is determined based on the first starting point position, the second starting point position, and the corresponding confidence level.
[0017] On the other hand, embodiments of the present invention provide a multi-format medical record text structure extraction system based on configuration rules, including: An adaptive repair engine is used to identify the format of original medical record data, which includes several original document data, based on the format identifier and / or character encoding range in each original document data, determine the data format corresponding to each original document data, repair and map the original medical record data based on the data format, and generate corresponding standard structured medical record data; wherein, the data format includes markup language format; Based on the data format, the original medical record data is repaired and mapped to generate corresponding standard structured medical record data, including: The original document data in the markup language format is subjected to a first fault-tolerant parsing to generate a corresponding primary parse tree and error records. Based on the primary parse tree and the error records, the original document data is subjected to a first intelligent repair to generate a corresponding well-structured document object. The first fault-tolerant parsing includes feature tag recognition and stack parsing of the original document data. The first intelligent repair includes tag completion and structure normalization. The well-structured document object is mapped and transformed to generate the corresponding standard structured medical record data; The intelligent rule configuration center is used to configure rules and generate declarative rules based on extraction requirements using a preset rule library; wherein, the declarative rules include anchor rules and extraction rules, the anchor rules include precise anchor rules and / or fuzzy fault-tolerant anchor rules; the preset rule library includes a keyword and variant knowledge base; Based on the extraction requirements, rules are configured using a pre-defined rule base to generate declarative rules, including: Based on the target words for extraction requirements and the keyword and variant knowledge base, configure and generate the anchor point rules; Based on the information extraction range of the extraction requirements, configure and generate the extraction rules; An enhanced rule execution engine is validated to determine the extraction start point position in the standard structured medical record data based on the precise anchor point rules and / or the fuzzy fault-tolerant anchor point rules. Information is extracted from the extraction start point position according to the extraction rules to obtain candidate field data and the corresponding data hierarchy. Multi-level validation is performed on the candidate field data to obtain a structured data package corresponding to the extraction requirements. The structured data package includes target field data, the data hierarchy, and a validation report. The multi-level validation includes one or more of rule logic validation, standard terminology database validation, and lightweight NLP semantic validation.
[0018] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects: 1. By recognizing the format of the original medical record data, repairing and mapping it based on different data formats, the system transforms medical record documents in multiple formats such as text, XML, and HTML into queryable standard structured medical record data. Based on the configuration of the preset rule base, declarative rules are generated, and anchor rules are integrated to extract information and perform multi-level verification on the standard structured medical record data. This achieves highly adaptable and accurate structured extraction of medical record documents in multiple formats such as text, XML, and HTML, and simultaneously obtains data genealogy and verification reports, providing credibility assurance for the extracted target field data.
[0019] 2. By incorporating fuzzy fault-tolerant anchor point rules into declarative rules, we can effectively address issues such as writing variations, non-standard writing, and OCR noise in real-world medical record texts. This improves the robustness and coverage of information extraction for non-standardized and noisy text, ensuring that the starting point for extraction can still be reliably found even when medical record data is not written in a standard way. This solves the problem of failure caused by minor format differences in traditional rule-based methods.
[0020] 3. By performing root cause analysis on the results of information extraction and multi-level verification, optimization recommendations are generated, and declarative rules and preset rule bases are updated to achieve continuous evolution of information extraction capabilities and multi-level verification capabilities, enabling the multi-format medical record text structure extraction method of this application to have optimization closed-loop capabilities.
[0021] 4. By constructing an adaptive repair engine with non-standard data repair capabilities, an intelligent rule configuration center supporting fuzzy fault-tolerant anchor point rules, and a verification-enhanced rule execution engine integrating multi-level verification and capable of obtaining data lineages and verification reports, the system achieves efficient, accurate, and reliable processing of multi-format heterogeneous medical records. Furthermore, by constructing a quality governance platform with root cause analysis and optimization recommendations, the intelligent rule configuration center and verification-enhanced rule execution engine are optimized and updated, enabling the multi-format medical record text structure extraction system to possess continuous evolution capabilities.
[0022] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description
[0023] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. Figure 1This is a flowchart illustrating a method for structured extraction of multi-format medical record text based on configuration rules in an embodiment of the present invention. Figure 2 This is a schematic diagram of the main modules of a multi-format medical record text structure extraction system based on configuration rules in an embodiment of the present invention. Detailed Implementation
[0024] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0025] A specific embodiment of the present invention discloses a method for structured extraction of multi-format medical record text based on configuration rules, such as... Figure 1 As shown, it includes: Step S1: For the original medical record data including several original document data, the format is identified according to the format identifier and / or character encoding range in each original document data, the data format corresponding to each original document data is determined, and the original medical record data is repaired and mapped based on the data format to generate corresponding standard structure medical record data; wherein, the data format includes markup language format; Based on the data format, the original medical record data is repaired and mapped to generate corresponding standard structured medical record data, including: The original document data in the markup language format is subjected to a first fault-tolerant parsing to generate a corresponding primary parse tree and error records. Based on the primary parse tree and the error records, the original document data is subjected to a first intelligent repair to generate a corresponding well-structured document object. The first fault-tolerant parsing includes the identification of feature tags and stack parsing of the original document data. The first intelligent repair includes tag completion and structure normalization. The well-structured document object is mapped and transformed to generate the corresponding standard structured medical record data.
[0026] In this embodiment, the original medical record data includes several original document data. Different original document data may be in different data formats, such as text, XML, and HTML. The format of each original document data is identified, and the corresponding repair method is determined based on the data format of the original document data. The original document data is repaired, and the repaired original document data is mapped to generate queryable standard structure medical record data. The data format includes markup language format. Step S1 specifically includes steps S11-S12.
[0027] Step S11: For the original medical record data including several original document data, perform format identification based on the format identifier and / or character encoding range in each original document data, and determine the data format corresponding to each original document data.
[0028] For original medical record data comprising several original document data, format identification is performed based on the format identifier and / or character encoding range in each original document data to determine the data format corresponding to each original document data. Format identification includes first format identification and second format identification.
[0029] In this embodiment, format recognition is performed through steps S111-S112 to determine the data format corresponding to each original document data. It should be noted that in this embodiment, the original document data includes the data stream of the original medical record document stored in code form.
[0030] Step S111: For each of the original document data, perform the first format recognition based on the format identifier and / or the character encoding range of the original document data. If the first format recognition is successful, determine the data format corresponding to the original document data based on the recognition result of the first format recognition. The first format recognition includes identifier matching based on the format identifier of the original medical record data and encoding analysis based on the character encoding range of the original medical record data.
[0031] For each original document data, a first format identification is performed based on the format identifier and character encoding range of the original document data. If the first format identification is successful, the data format corresponding to the original document data is determined based on the identification result. The first format identification includes identifier matching based on the format identifier of the original medical record data and encoding analysis based on the character encoding range of the original medical record data. In this embodiment, the format identifier includes a signature-type format identifier, and the first format identification is specifically performed through steps a-b.
[0032] Step a: Quickly match the data format of the original document data based on the format identifier. If the match is successful, determine the data format of the original document data based on the matching result.
[0033] In this embodiment, the data format of the original document data is quickly matched based on the signature class format identifier. If the identifier matches successfully, the data format of the original document data is determined according to the matching result. For data streams in standardized markup language formats, a signature class format identifier is typically set at the front end of the data stream, such as a magic number or other forms of file signature. For example, standard XML files typically use signature class format identifiers.<?xml version="1.0"?> At the beginning, HTML5 documents begin with<!DOCTYPE html> In the beginning, then "<?xml version="1.0"?> "This is the signature class format identifier corresponding to the XML format, "<!DOCTYPE html> "This is the signature class format identifier corresponding to the HTML5 format. The signature class format identifier of the original document data is compared and matched with the predefined identifier to complete the identifier matching of the original medical record data.
[0034] Specifically, for each original document data, the first byte of the original document data is read as the corresponding signature class format identifier, and matched with the predefined identifiers in the predefined format signature library. If the match is successful, the data format of the original document data is determined according to the matched predefined identifier (i.e., the matching result). If the match fails, proceed to step b. In this embodiment, the markup language format includes XML format and HTML format.
[0035] The length of the header bytes is determined based on a predefined format signature library. For example, if all predefined identifiers in the predefined format signature library have lengths of 512 bytes and 1024 bytes, then the length of the header bytes will also have lengths of 512 bytes and 1024 bytes. For each original document data, 512 bytes and 1024 bytes starting from the first byte are read as the two header bytes of the original document data. For each header byte, it is compared and matched with each predefined identifier of the same byte length in the predefined format signature library. For example, the original document data has two header bytes, header byte A (512 bytes) and header byte B (1024 bytes). The format signature library includes 6 predefined identifiers: predefined identifier 1 (512 bytes), predefined identifier 2 (512 bytes), predefined identifier 3 (512 bytes), predefined identifier 4 (1024 bytes), predefined identifier 5 (1024 bytes), and predefined identifier 6 (1024 bytes). Then, header byte A is compared and matched with predefined identifier 1, predefined identifier 2, and predefined identifier 3 in sequence, and header byte B is compared and matched with predefined identifier 4, predefined identifier 5, and predefined identifier 6 in sequence. If any comparison is successful, such as header byte A and predefined identifier 1 having the same characters for each byte, then predefined identifier 1 is considered a matched predefined identifier, and predefined identifier 1 is used as the data format of the original document data.
[0036] By quickly matching the data format of each original document data using signature-based format identifiers, documents in standardized markup language formats can be efficiently identified.
[0037] Step b: Perform encoding analysis on the data format of the original document data based on the character encoding range. If the encoding analysis is successful, determine the data format of the original document data based on the encoding analysis results.
[0038] In this embodiment, the original document data that failed the fast matching in step a is subjected to encoding analysis. The character encoding range of the original document data is statistically analyzed. If it meets the encoding range of a certain data format, the encoding analysis is considered to have passed, and the data format that meets the criteria (i.e., the encoding analysis result) is determined as the data format of the original document data. If it does not meet the encoding range of any data format, the first format recognition is considered to have failed, and the process proceeds to step S112. For example, if the character encoding range corresponding to the original document data conforms to the byte rules corresponding to the UTF-8 format "00~7F, C2~F4 at the beginning + followed by 80~BF" and there are no illegal bytes (e.g., C0 / C1 / F5~FF), then the original document data is considered to meet the plain text format, that is, the encoding analysis of the original document data has passed, the encoding analysis result is plain text format, and the data format of the original document data is determined to be plain text format.
[0039] Further, the original document data in plain text format determined in step b is checked according to encoding conventions. If the check passes, the determined data format is retained; if the check fails, the first format recognition is considered to have failed, and the process proceeds to step S112. For example, a document in regular UTF-8 format does not contain a large number of non-text control characters (e.g., 00–1F, 7F). If the proportion of 00–1F and 7F in all characters of the original document data reaches the character proportion threshold (e.g., 0.1%), it is determined that the check has failed.
[0040] The first format recognition method can effectively identify raw document data in standardized markup language format and plain text format.
[0041] Step S112: If the first format recognition fails, then the original document data is extracted according to the format identifier of the original document data, and the data format features are input into a lightweight classification model for the second format recognition to determine the data format corresponding to the original document data.
[0042] In this embodiment, the format identifier also includes a symbolic format identifier. For original document data that fails to identify the first format, feature label recognition and statistical analysis are performed based on the symbolic format identifier of the original document data to extract data format features. The extracted data format features are then input into a lightweight classification model for classification, and the classification probability corresponding to each data format is output, thus completing the second format identification of the original document data. The data format with the highest classification probability is taken as the data format corresponding to the original document data. For example, the lightweight classification model outputs three probability values, which represent the classification probabilities corresponding to XML format, HTML format, and plain text format, respectively. For a certain original document data, the three probability values are 80%, 15%, and 5%, respectively. Therefore, the data format of the original document data is XML format in the markup language format.
[0043] The data format characteristics include the density of key symbols, nesting density, key tags, structural density, and key heading density corresponding to symbol-based format identifiers. Key symbol density refers to the frequency of occurrence of key symbol-based format identifiers in the original document data; for example, the percentage of characters in angle brackets "<>" out of the total number of characters in the original document data. Nesting density refers to the frequency of occurrence of symbol-based format identifiers that demonstrate hierarchical nesting in the original document data; for example, " the proportion of characters that appear in pairs of " and "< / div", " and "; key tag density refers to the occurrence frequency of key tags in the original document data, wherein key tags refer to symbolic format identifiers that reflect the data flow format, such as , , <clinicaldocument>Complete tags, and <html、html、<body、body、body> Incomplete tags, and the total percentage of all complete and incomplete tags in the total number of characters in the original document data are used as the key tag density; structural density refers to the frequency of occurrence of formatting symbols such as line breaks and indentation in the original document data; key heading density refers to the frequency of occurrence of formatting symbols representing key headings in the original document data, such as key field headings unique to structured or semi-structured original document data, such as "Chief Complaint:" and "Diagnosis:". In this embodiment, the structured or semi-structured original document data is divided into plain text format.
[0044] In this embodiment, the lightweight classification model uses a rule-based decision tree and is trained using historical raw document data. In other embodiments, Naive Bayes or other existing technologies may also be used; no limitation is made here.
[0045] This embodiment addresses the situation where real-world original medical record data formats are mixed, labels are missing, labels are incomplete, or file signatures are incorrect by using multi-level, progressive format recognition. All original document data are divided into two major data formats: markup language format and plain text format (TEXT). Furthermore, the markup language format includes Extensible Markup Language and its variants (XML format) and Hypertext Markup Language and its variants (HTML format).
[0046] Step S12: Repair and map the original medical record data based on the data format to generate corresponding standard structure medical record data.
[0047] Based on the data format, the original medical record data is repaired and mapped to generate corresponding standard structure medical record data. This includes repairing and mapping the original medical record data in markup language format to generate corresponding standard structure medical record data. Step S12 includes steps S121-S122.
[0048] Step S121: Perform a first fault-tolerant parsing on the original document data in the markup language format to generate a corresponding primary parse tree and error record. Based on the primary parse tree and the error record, perform a first intelligent repair on the original document data to generate a corresponding well-structured document object. The first fault-tolerant parsing includes the identification of feature tags and stack parsing of the original document data. The first intelligent repair includes tag completion and structure normalization.
[0049] The original document data in markup language format is subjected to first fault-tolerant parsing to generate a corresponding primary parse tree and error record. Based on the primary parse tree and error record, the original document data is subjected to first intelligent repair to generate a corresponding well-structured document object, including steps S1211-S1213.
[0050] Step S1211: parsing the raw document data in the markup language format by using a lenient scanner, and identifying the feature tags in the raw document data.
[0051] Parsing the raw document data in the markup language format by using a lenient scanner, and identifying feature tags in the raw document data. Specifically, the raw document data in the markup language format is scanned by the lenient scanner to identify " <tag> ”、"< / tag> ", " <tag / > " and other feature tags in the document. When a malformed feature tag (e.g., "<tag", an unclosed tag) is encountered, error reporting and scanning stopping are not performed immediately; instead, the malformed feature tag is corrected according to a predefined correction rule, for example, inserting a missing ">" to close the feature tag, completing minimal correction on the raw document data to continue scanning, taking both original complete feature tags and corrected feature tags as the feature tags included in the raw document data, and recording the correction process.
[0052] Further, during the parsing of the raw document data in the markup language format by using the lenient scanner, the method further includes minimal correction of feature attributes. Specifically, the raw document data in the markup language format is scanned by the lenient scanner to identify feature attributes such as "attr ="valu"" in the document. When a malformed feature attribute (e.g., "attr ="value, missing quotation mark) is encountered, error reporting and scanning stopping are not performed immediately; instead, the malformed feature attribute is corrected according to a predefined correction rule, for example, inserting a missing quotation mark to close the feature attribute, completing minimal correction on the feature attributes of the raw document data, so that both original complete feature attributes and corrected feature attributes in the raw document data are identified as the feature attributes included in the raw document data in a subsequent processing process.
[0053] Step S1212: performing the stack parsing on the raw document data based on the feature tags, and generating a corresponding primary parse tree and an error record.
[0054] Performing stack parsing on the raw document data based on the feature tags identified by the lenient scanner, and generating a corresponding primary parse tree and an error record. Specifically, a stack-based parsing algorithm is adopted: when a start tag is encountered, the tag is pushed onto the stack; when an end tag is encountered, it is not required to strictly match the tag at the top of the stack, and the parsing algorithm searches upward in the stack for a tag with the same name to perform matching and closing, and automatically records unmatched feature tags in the middle as nodes to be repaired. A corresponding primary parse tree and a corresponding error record are generated according to the parsed matching and closing structure. The stack parsing adopted in this embodiment can tolerate a large number of tag nesting errors.
[0055] The error records include nodes that need to be repaired, corrections to feature labels in step S1211, such as unmatched feature labels and corrected unclosed label positions.
[0056] For example, a section of the original document data is as follows: <main> <section> <txt> Text content 1 < / txt> < / section> <txt> Text content 2 < / txt> <section> <txt> Text content 3 < / txt> < / section> < / main> This data segment contains 13 feature labels, including 6 start labels: 1 " <main>"、2" <section>", 3" <txt> 7 closing tags: 3 ";< / txt> ", 3"< / section> "、1"< / main> ".", among which, the stack corresponding to "text content 1"[ <main>, <section> , <txt>"]", top label of the stack <txt> The closing tag after "text content 1"< / txt> Match the top label of the stack. <txt>", which meets the strict matching requirement; while for two feature labels with the same name" <section> "and 3"< / section> "Cannot be strictly matched, the second one"< / txt> < / txt> < / section> "With the first" <section> Perform a matching closure, closing the first "< / section> "As an unmatched closing tag, the third "" is considered a node that needs to be repaired, along with the second "". <section>"Strict matching closure. Furthermore," <main> "and"< / main> "The second one" <txt> "With the second"< / txt> "The third one" <txt> "With the third"< / txt> "Strict matching closure. Based on the parsed matching closure structure, generate the corresponding primary parse tree and the corresponding error record (the first one)."< / section> The example of a primary parse tree is as follows: main ├─section │└─text→Text content 1 │└─text→Text content 2 └─section └─text→Text content 3 For example, if a segment of the original document data contains unmatched feature tags "" and "", then these two feature tags will be treated as the corresponding error records.
[0057] Through the tolerant scanner's syntax analysis and stack parsing, even in the event of syntax errors such as tag mismatches and unclosed attribute values, a primary parse tree containing the text content of the original document data is still constructed, and corresponding error records are generated.
[0058] Step S1213: Based on the primary parse tree and the error record, perform tag completion and structure normalization on the original document data to generate the well-structured document object.
[0059] Based on the initial parse tree and error records, the original document data is labeled and normalized to generate a well-formed document, which is a well-formed document object.
[0060] First, tag completion is performed based on context awareness to close the tags. For unmatched feature tags in erroneous records, the correct closing tags are intelligently inserted or corrected based on the feature tag signature, the feature tag position in the original document data, and the semantics of adjacent feature tags. Based on the aforementioned example, for the first example in step S1212, in the second " <txt>"Insert before" <section> ", making the original first "< / section> "Compared to the original first one" <section> "Matching closure, the original second"< / section> "with the newly inserted" <section>"Match closure; for the second example, correct "" to "".
[0061] Secondly, based on the data format and preset nesting specifications, unreasonable nesting is reorganized. For example, the preset nesting specifications corresponding to HTML format include " No nesting If the original document data contains Nested Then, the unreasonable nesting will be reorganized, for example, in " "Insert before" ", compared with the original "Matching, will match the original" "Close early; will" "and" "Promoted one level; in" "Insert after" ", compared with the original "Matching will make the original" "Complete the closure."
[0062] For the original document data in markup language format, the process involves parsing and repairing it through steps S1211-S1213. Without relying on perfect syntax, the process aims to restore the document's structured information as much as possible and output a well-structured document object with basically correct syntax and clean structure, including the repaired parse tree and the text content of the original document data.
[0063] Step S122: Map and transform the well-structured document object to generate the corresponding standard structure medical record data.
[0064] Well-structured document objects are mapped and transformed to generate corresponding standard structured medical record data. In this embodiment, a Unified Document Model (UDM) is used to map and transform well-structured document objects to generate corresponding standard structured medical record data. The Unified Document Model includes a UDM Schema (Unified Data Model Structure Definition) and a writing adapter. A pre-defined, general UDM Schema defines each level of the well-structured document object, representing "document," "chapter," "section," "paragraph," "body text," etc. For example, the writing adapter for HTML format will... <h1>The tag is converted to a "Heading 1" node in the UDM. The tags are converted into "paragraph" nodes. By writing an adapter, the well-structured document object is converted into the corresponding standard structured medical record data, including the text content of the original medical record data and metadata.
[0065] Taking a well-structured document object corresponding to the original HTML document data as an example, the parse tree (DOM tree) nodes and text content of the well-structured document object are mapped one by one to the corresponding nodes of the UDM according to the corresponding semantics and hierarchical relationships. Each UDM node includes the corresponding text content, as well as the metadata of the original document data, such as the source XPath (path language) and line numbers.
[0066] Furthermore, the data format also includes plain text format, and step S12 also includes steps S123-S124.
[0067] Step S123: Perform second fault-tolerant parsing and second intelligent repair on the original document data in the plain text format to generate a structured text object.
[0068] In this embodiment, the second fault-tolerant parsing includes encoding detection and error correction, and text segmentation; the second intelligent repair includes paragraph recombination and structural simulation, and step S123 includes steps S1231-S1234.
[0069] Step S1231: Perform encoding detection and error correction on the original document data in the plain text format to generate a Unicode character stream.
[0070] Encoding detection and error correction are performed on raw document data in plain text format to generate a Unicode character stream. Statistical character distribution analysis and general encoding detection libraries (such as chardet) are used to identify possible multi-byte encodings (such as GBK, UTF-8). For parts that fail to decode, a heuristic replacement or ignoring strategy is employed to ensure a consistent Unicode character stream output.
[0071] Step S1232: Divide the Unicode character stream into text segments based on newline characters to generate a text line sequence.
[0072] The Unicode character stream is divided into text lines based on newline characters to generate a text line sequence. The Unicode character stream is initially split into lines based on newline characters (such as "\n" and "\r\n"). Irregular newlines (such as multiple consecutive newlines or space newlines) are merged or cleaned up to form a regular text line sequence with uniform encoding and organized by line.
[0073] By using encoding detection and error correction and text segmentation, this method solves problems such as encoding chaos and irregular delimiters in raw document data in plain text format. It reliably converts the raw document data from raw byte stream to character stream and identifies basic text blocks (i.e., text line sequences) for secondary intelligent repair.
[0074] Step S1233: Perform the paragraph recombination on the text line sequence to generate logical paragraphs.
[0075] The text line sequence is reorganized into paragraphs to generate logical paragraphs. Specifically, regular expressions and keyword matching are used to identify common medical record field titles (e.g., "Chief Complaint:", "Past Medical History:"), combined with indentation, line numbering (e.g., "1.", "2."), and special punctuation (e.g., "..."). Visual cues such as "" are used to merge consecutive lines of text into logical paragraphs.
[0076] Step S1234: Perform structural simulation on the logical paragraph to generate the structured text object.
[0077] The logical paragraphs are structurally simulated to generate structured text objects. Specifically, based on indentation levels and semantic relationships (e.g., "present illness history" should include multiple symptom description paragraphs), a lightweight tree-like hierarchical structure is constructed, simulating XML-like parent-child node relationships to generate structured text objects while preserving the text content of the original document data. Furthermore, the font format (e.g., bold mode) of the titles in the original medical record document can be scanned and identified to assist in the construction of the tree-like hierarchical structure. Steps S1233-S1234 complete the analysis of the text line sequence, identifying and reconstructing logical structures such as chapters, sections, paragraphs, and lists, generating structured text objects.
[0078] Through second fault-tolerant parsing and second intelligent repair, the unstructured original document data is parsed and the implicit document logical structure is reconstructed, outputting a structured text object with logical hierarchy (e.g., chapters, sections, paragraphs).
[0079] Step S124: Map and transform the structured text object to generate the corresponding standard structured medical record data.
[0080] The structured text objects are mapped and transformed to generate corresponding standard structured medical record data. In this embodiment, a Unified Document Model (UDM) is used to map and transform the structured text objects to generate corresponding standard structured medical record data. The Unified Document Model includes a UDM Schema (Unified Data Model Structure Definition) and an adapter. A pre-defined, general UDM Schema defines each level of the structured text object as representing a "document," "chapter," "section," "paragraph," or "body text." The adapter converts the structured text objects into corresponding standard structured medical record data, including the original medical record data's text content and metadata. The specific principle is the same as step S122 and will not be repeated here.
[0081] By identifying the format of the original medical record data, different fault-tolerant parsing and intelligent repair methods are adopted for original document data with different data formats. This allows for greater tolerance of format defects at the data level, and the repaired data is converted into standard structured medical record data, providing a stable, consistent, clean, and queryable data source for information extraction in the subsequent step S3.
[0082] Step S2: Based on the extraction requirements, configure rules using a preset rule base to generate declarative rules; wherein, the declarative rules include anchor rules and extraction rules, the anchor rules include precise anchor rules and / or fuzzy fault-tolerant anchor rules; the preset rule base includes a keyword and variant knowledge base; Based on the extraction requirements, rules are configured using a pre-defined rule base to generate declarative rules, including: Based on the target words for extraction requirements and the keyword and variant knowledge base, configure and generate the anchor point rules; Based on the information extraction range of the extraction requirements, the extraction rules are configured and generated.
[0083] Based on the extraction requirements, rules are configured using a preset rule base to generate declarative rules corresponding to the extraction requirements. In this embodiment, the preset rule base includes a keyword and variant knowledge base, and step S2 specifically includes steps S21-S22.
[0084] Step S21: Configure and generate the anchor point rules based on the target words for extraction requirements and the keyword and variant knowledge base.
[0085] In this embodiment, the keyword and variant knowledge base is a structured thesaurus built by statistically analyzing common spellings, variants, and variant patterns of key fields in historical medical record data. It includes multiple preset keyword groups; for example, the chief complaint-keyword group includes the keyword "chief complaint," and corresponding variant reference words "chief-complaint," "chief complaint," and the text distance between each variant reference word and the keyword. In this embodiment, the text distance is determined based on edit distance. In other embodiments, the text distance can also be determined based on Jaccard similarity, cosine similarity, or other existing technologies.
[0086] In this embodiment, based on a keyword and variant knowledge base, anchor rules corresponding to extraction requirements are defined declaratively through a visual interface. These rules include precise anchor rules and / or fuzzy, error-tolerant anchor rules, used to determine the starting point position for extraction in standard structured medical record data. For example, if precise positioning is desired, precise anchor points and anchor ranges are set to generate precise anchor rules; if fuzzy positioning is desired, fuzzy anchor points, anchor distances, and anchor ranges are set to generate fuzzy, error-tolerant anchor rules.
[0087] In this embodiment, the extraction requirement includes extraction target words, which indicate the fields from which corresponding data needs to be extracted. For example, extracting the content corresponding to "chief complaint" and "discharge diagnosis," where "chief complaint" and "discharge diagnosis" are the extraction target words corresponding to this extraction requirement, which can be configured by the user in the visual interface. It should be noted that for an extraction requirement, different anchor rules can be selected for different extraction target words, and multiple anchor rules can also be set for the same extraction target word at the same time.
[0088] For example, the extraction requirement is to extract content corresponding to "chief complaint" and "discharge diagnosis," including the two target keywords "chief complaint" and "discharge diagnosis." For "chief complaint," if fuzzy positioning is desired, select "Fuzzy Tolerant Anchor Point" in the visualization interface, and set the fuzzy anchor point, anchor distance, and anchor range for "chief complaint" to generate fuzzy tolerance anchor point rules. For example, set the fuzzy anchor point to "chief complaint," the anchor distance to 0.85, and the anchor range to the title field. For "discharge diagnosis," if precise positioning is desired, select "Precise Anchor Point" in the visualization interface, and set the precise anchor point and anchor range for "discharge diagnosis" to generate precise anchor point rules. For example, set the precise anchor point to "discharge diagnosis" and the anchor range to the body text field.
[0089] For standard medical record data with well-defined formats and stable structures, precise anchor rules can be used with exact matching techniques such as regular expressions and XPath for accurate positioning and high execution efficiency. For variable and non-standard field titles in the real world (such as "Chief Complaint" and "Past Medical History"), fuzzy error-tolerant anchor rules can be used for positioning, effectively handling writing variations, extra spaces, special punctuation, and minor OCR errors, serving as a supplement and enhancement to precise anchor rules. The combination of precise anchor rules and fuzzy error-tolerant anchor rules greatly improves the robustness of the method in this embodiment for positioning non-standardized and noisy text, ensuring that the starting point for extraction can still be reliably found and anchor positioning completed even when the writing is non-standard.
[0090] Step S22: Configure and generate the extraction rules according to the information extraction range of the extraction requirements.
[0091] Based on the information extraction scope required, extraction rules are configured to define the extraction scope after determining the starting point position, and to extract data within this scope as candidate field data. The information extraction scope indicates the desired range of data to be extracted. Rules are configured based on this scope to generate extraction rules. For example, the information extraction scope could be the paragraph containing the starting point, with the extraction rule configured as "Apply a regular expression to extract all text from the segment marker preceding the starting point to the segment marker following the starting point"; or the information extraction scope could be the text at the same level as the starting point, with the extraction rule configured as "Apply a regular expression to extract all text from the starting point until the next heading at the same level as the starting point"; or the information extraction scope could be the adjacent sentence, with the extraction rule configured as "Apply a regular expression to extract all text from the starting point until the next sentence break". For example, the anchor range is the title field, and the extraction rule is "apply a regular expression to extract all text starting from the extraction start position until the next title at the same level as the extraction start position"; or the extraction rule is "configure a regular expression to capture the content after the colon," for example, "chief complaint: headache for 3 days," if "chief complaint" is the extraction start position, then "headache for 3 days" is extracted as candidate field data. In this embodiment, the extraction rules corresponding to the extraction requirements are defined in a declarative manner through a visual interface.
[0092] Improving the accuracy of information extraction content boundaries through anchor point rules can simplify extraction rules, reduce the data processing complexity of the information extraction process, and enhance the accuracy and noise resistance of information extraction by combining anchor point rules with extraction rules.
[0093] Furthermore, the preset rule base also includes a verification knowledge base and a verification model, the declarative rules also include multi-level verification rules, and step S2 also includes step S23.
[0094] Step S23: Based on the extraction requirements, the verification knowledge base, and the verification model, configure and generate multi-level verification rules.
[0095] Based on the extraction requirements, validation knowledge base, and validation model, multi-level validation rules are configured and generated. In this embodiment, the validation knowledge base includes a standard terminology database, built based on ICD-10, LOINC, and a drug dictionary; the validation model includes a lightweight NLP semantic model, built based on existing NLP semantic models and trained using the standard terminology database; the multi-level validation rules include setting one or more of the following for the extraction requirements: rule logic validation, standard terminology database validation, and lightweight NLP semantic validation. For example, the multi-level validation rules for the extraction requirements might be rule logic validation and standard terminology database validation of the candidate field data corresponding to the extraction requirements. In this embodiment, the multi-level validation rules for the extraction requirements are defined declaratively through a visual interface.
[0096] The process includes: rule-based logic verification, which checks the numerical range, format, and correlation of the extracted candidate field data; standard terminology database verification, which maps the mandatory or suggested fields in the extracted candidate field data to the standard medical terminology system based on the standard terminology database, where mandatory and suggested fields can be pre-set in the standard terminology database, mandatory fields are replaced and the replacement is recorded in the verification assurance, and suggested fields are retained and the suggestions are recorded in the verification report; and lightweight NLP semantic verification, which quickly scores the semantic rationality of the extracted candidate field data, intuitively reflecting the credibility of each candidate field data, such as determining whether a piece of text belongs to "symptom description" or whether a value is within the physiologically reasonable range.
[0097] In this embodiment, the anchor rules, extraction rules, and multi-level verification rules in the rule configuration results are converted into computer-readable language as declarative rules.
[0098] In other embodiments, based on the rule configuration results completed by the visual interface, a simple and standardized rule format is preset using URLD (Unified Rule Description Language) with JSON as the carrier. The rule configuration results corresponding to the extraction requirements are converted into declarative rules expressed in URLD, so that users who make extraction requirements only need to declare their extraction intent in the visual interface, such as what fields to extract, without requiring users to write specific parsing code for different formats such as XML, HTML, and text.
[0099] Step S3: Based on the precise anchor point rule and / or the fuzzy fault-tolerant anchor point rule, determine the extraction starting point position in the standard structured medical record data. Extract information from the extraction starting point position according to the extraction rule to obtain candidate field data and the corresponding data spectrum. Perform multi-level verification on the candidate field data to obtain the structured data package corresponding to the extraction requirement. The structured data package includes target field data, the data spectrum, and a verification report. The multi-level verification includes one or more of rule logic verification, standard terminology database verification, and lightweight NLP semantic verification.
[0100] Based on the anchor rules, extraction rules, and multi-level validation rules of the declarative rules, information extraction and multi-level validation are performed on the standard structured medical record data to obtain the structured data package corresponding to the extraction requirements. The structured data package includes target field data, data genealogy, and validation report. Multi-level validation includes one or more of the following: rule logic validation, standard terminology database validation, and lightweight NLP semantic validation. Step S3 includes steps S31-S33.
[0101] Step S31: Determine the starting point position for extraction in the standard structured medical record data based on the precise anchor point rule and / or the fuzzy fault-tolerant anchor point rule.
[0102] In this embodiment, step S31 can be executed using any one of steps S311, S312, or S313.
[0103] Step S311: Determine the starting point position for extraction in the standard structured medical record data based on the precise anchor point rule.
[0104] In this embodiment, the precise anchor point rule includes precise anchor points and anchoring range. Based on the precise anchor points and anchoring range, anchor points are located in the standard structured medical record data to determine the starting point position for extraction.
[0105] First, candidate nodes are determined from all text nodes in the standard structured medical record data based on the anchoring range.
[0106] In this embodiment, text nodes refer to UDM nodes corresponding to levels such as "document", "chapter", "section", "paragraph", and "body text". For example, if the anchoring range is the title field, then "document", "chapter", "section", and "paragraph" are determined as candidate nodes, and a precise search for anchor points is performed in the text content of document title, chapter title, section title, and paragraph title.
[0107] Secondly, if a precise anchor point exists among the candidate nodes, the position of the precise anchor point in the standard structured medical record data will be used as the starting point for extraction.
[0108] A strict matching search is performed on the text content of candidate nodes. If a precise anchor point exists, its position is used as the starting point for extraction. This strict matching search for precise anchor points from the text content can be performed using existing keyword extraction tools; no restrictions are placed here. For example, if the precise anchor point is "chief complaint," and "chief complaint" exists in the paragraph title, then the position where "chief complaint" appears in that paragraph title is used as the starting point for extraction.
[0109] Furthermore, if no precise anchor point exists among the candidate nodes, the starting point position is set to empty. In this embodiment, if the starting point position is empty, the process returns to step S2 to reconfigure the rules and generate new declarative rules.
[0110] Step S312: Determine the starting point position for extraction in the standard structured medical record data based on the fuzzy fault-tolerant anchor point rule.
[0111] In this embodiment, the fuzzy fault-tolerant anchor rule includes a fuzzy anchor, an anchoring distance, and an anchoring range. The anchor positioning process corresponding to the fuzzy fault-tolerant anchor rule includes: taking the keywords of the preset keyword group corresponding to the fuzzy anchor and all variant reference words as search target words, and performing fuzzy matching search in the text content of the candidate nodes of the standard structure medical record data determined according to the anchoring range, specifically including steps S3121-S3122.
[0112] Step S3121: Based on the anchoring range, determine candidate nodes from all text nodes of the standard structure medical record data, and calculate the text distance between the fuzzy anchor point and the candidate node.
[0113] First, candidate nodes are determined from all text nodes in the standard structured medical record data based on the anchoring range.
[0114] The principle is the same as that of determining candidate nodes in step S311, so it will not be repeated here.
[0115] Secondly, calculate the text distance between the fuzzy anchor point and the candidate node.
[0116] Using the keywords of the preset keyword group corresponding to the fuzzy anchor point and all variant reference words as the search target words, a sliding search is performed in the text content of the candidate nodes, and the text distance between the word segment and the search target word at each step is calculated.
[0117] For example, the sliding step size is 0.5 times the minimum number of characters of the search target word to 2 times the maximum number of characters. Based on the above example, the keywords "chief complaint", variant reference words "chief-complaint" and "chief complaint" are used as search target words. A sliding search is performed on the text content of document titles, chapter titles, section titles and paragraph titles in standard structured medical record data. First, the number of characters corresponding to the keyword "chief complaint" is used as the sliding step size. The text distance between the word segment corresponding to each step size in the title field and each search target word is calculated.
[0118] If a word in the title field has a text distance of not less than 0.85 from any search target word, then the position where that word appears will be used as the starting point for extraction.
[0119] Step S3122: If the text distance is not less than the anchoring distance, then the position of the candidate node in the standard structure medical record data is determined as the extraction starting point position.
[0120] If a candidate node has a text distance to a word segment that is not less than the anchor distance, then the position where the word appears in the candidate node in the standard structure medical record data is determined as the starting point for extraction.
[0121] In other embodiments, the anchor point localization process corresponding to the fuzzy fault-tolerant anchor point rule includes: filtering from a preset keyword group corresponding to the fuzzy anchor point based on the anchoring distance; using the filtered keywords and variant reference words as search target words; and performing a strict matching search in the text content of candidate nodes of the standard structured medical record data determined according to the anchoring range. Based on the aforementioned example, the keyword "chief complaint" and variant reference words in the chief complaint-keyword group whose text distance from "chief complaint" is not less than 0.85 are all used as search target words, and the search is performed in the text content of the title field of the standard structured medical record data. If a search target word exists, the position of the search target word in the standard structured medical record data is determined as the extraction starting point position.
[0122] Step S313: Determine the starting point position for extraction in the standard structured medical record data based on the precise anchor point rule and the fuzzy fault-tolerant anchor point rule.
[0123] In this embodiment, step S313 is executed according to steps S3131-S3133.
[0124] Step S3131: Based on the precise anchor point rules, anchor point positioning is performed on the standard structure medical record data to determine the first starting point position.
[0125] Based on the precise anchor point corresponding to the precise anchor point rule, anchor point positioning is performed in the standard structure medical record data to determine the first starting point position. The principle is the same as step S311, and will not be repeated here.
[0126] Step S3132: Based on the fuzzy fault-tolerant anchor point rule, anchor point positioning is performed on the standard structure medical record data to determine the second starting point position and the corresponding confidence level.
[0127] Anchor point location is performed on standard structured medical record data based on fuzzy fault-tolerant anchor point rules to determine the second starting point position. The principle is the same as step S312, and will not be repeated here. In this embodiment, the text distance corresponding to the candidate node as the second starting point position is used as the confidence level corresponding to the second starting point position.
[0128] Step S3133: Determine the extraction starting point position based on the first starting point position, the second starting point position, and the corresponding confidence level.
[0129] The extraction starting point position is determined based on the first starting point position, the second starting point position, and the corresponding confidence level. In this embodiment, both the first starting point position and the second starting point position are used as the extraction starting point position.
[0130] In other embodiments, if the first starting position is not empty, only the first starting position is determined as the extraction starting position; or, based on the set first anchoring requirement (e.g., anchoring 3 extraction starting positions), the first and second starting positions required by the first anchoring requirement are selected as extraction starting positions according to their confidence levels from high to low; or, based on the set second anchoring requirement (e.g., confidence level exceeding 0.85), the first and second starting positions with confidence levels reaching the second anchoring requirement are selected as extraction starting positions. The confidence level corresponding to the first starting position is 1 by default.
[0131] In other embodiments, step S313 is performed according to steps S3134-S3137.
[0132] Step S3134: Based on the precise anchor point rule, perform precise positioning on the standard structure medical record data to determine the location of the third starting point.
[0133] Based on the precise anchor point corresponding to the precise anchor point rule, anchor point positioning is performed in the standard structured medical record data to determine the location of the third starting point. The principle is the same as step S311, and will not be repeated here.
[0134] Step S3135: Determine whether the precise anchor point rule has been successfully anchored.
[0135] In this embodiment, if the third starting position is not empty, the precise anchor rule is considered to have been successfully anchored. In other embodiments, a third anchoring requirement can be set when configuring the precise anchor rule. If the third starting position meets the third anchoring requirement, the precise anchor rule is considered to have been successfully anchored. For example, the third anchoring requirement is to match at least two third starting positions. If there is only one third starting position, the precise anchor rule is considered to have failed to anchor; if the third starting position includes two starting positions, the precise anchor rule is considered to have been successfully anchored.
[0136] Step S3136: If the precise anchor point rule is successfully anchored, then the third starting point position is taken as the extraction starting point position.
[0137] If the precise anchor point rule is successfully anchored, the third starting point position will be used as the extraction starting point position.
[0138] Step S3137: If the precise anchor point rule fails to anchor, then anchor point positioning is performed on the standard structure medical record data based on the fuzzy fault-tolerant anchor point rule to determine the starting point position for extraction.
[0139] Based on the fuzzy fault-tolerant anchor point rule, anchor point positioning is performed in standard structured medical record data to determine the starting point position for extraction. The principle is the same as step S312, and will not be repeated here.
[0140] Step S32: Extract information from the starting point position according to the extraction rules to obtain candidate field data and the corresponding data spectrum.
[0141] For standard structured medical record data, starting from the extraction start point, information is extracted from the standard structured medical record data according to the extraction rules to obtain candidate field data and the corresponding data hierarchy. For example, regular expressions are used to extract the text content after the colon at the extraction start point as candidate field data, and the meta-information corresponding to the text node at the extraction start point is used as a component of the data hierarchy corresponding to the candidate field data. For example, the data hierarchy includes the position of the candidate field data, the anchor point rules used, the extraction rules used, and the original text in the original document file corresponding to the original document data.
[0142] Step S33: Perform multi-level validation on the candidate field data to obtain the structured data package corresponding to the extraction requirements.
[0143] The candidate field data is subjected to multi-level validation to obtain the structured data package corresponding to the extraction requirements. In this embodiment, the candidate field data is subjected to multi-level validation according to the multi-level validation rules, and the candidate field data that passes the multi-level validation is standardized to obtain the target field data and the corresponding validation report.
[0144] Based on multi-level validation rules, the extracted candidate field data undergoes real-time multi-level validation. The candidate field data that passes multi-level validation is then standardized to obtain the target field data and the corresponding validation report. During the multi-level validation process, any validation failure generates a quality event with an error level (e.g., error, warning, suspicious), triggers corresponding actions (e.g., marking, calling backup rules, interrupting the process, forcing field replacement, suggesting fields), and is recorded in the validation report.
[0145] Rule-based logic validation verifies the extracted candidate field data for numerical range, format, and correlation. For example, it reads the data of adjacent text nodes in the standard structured medical record data for a candidate field and performs correlation verification. For instance, if the candidate field data is diastolic blood pressure, it extracts the systolic blood pressure value from the adjacent text node and verifies whether the diastolic blood pressure value is less than the systolic blood pressure value. The standard structured medical record data obtained through step S1 has a standardized tree structure, making field location more direct and efficient.
[0146] Standard terminology database validation maps mandatory or suggested fields in the extracted candidate field data to the standard medical terminology system based on the standard terminology database. For example, terminology mapping is performed based on diagnostic results and ICD-10 encoding. The context of candidate field data in standard structured medical record data is read. For the abbreviation "CAP" in the target field data, it is checked whether the parent node of the target field data is a discharge diagnosis to eliminate ambiguity, achieve more accurate mapping, and obtain more accurate structured data packets.
[0147] Lightweight NLP semantic validation quickly scores the semantic reasonableness of candidate field data, intuitively reflecting the credibility of each candidate field. For example, to verify whether a piece of text (candidate field data) belongs to "symptom description", it reads the context of the candidate field data in standard structured medical record data and the corresponding text format (e.g., bold headings) in the original medical record document file of the standard structured medical record data, judges the narrative coherence and semantic reasonableness, and determines the score.
[0148] In this embodiment, all multi-level verifications are assumed to pass by default, and the verification failures in the multi-level verification are recorded in the verification report. In other embodiments, verification pass conditions can be set, such as no "error" level quality events occurring, or the total number of quality events not exceeding 10. Meeting these verification pass conditions determines that the multi-level verification has passed.
[0149] For candidate field data that passes multi-level validation, data standardization processing is performed, such as string cleaning and date formatting, to obtain target field data and corresponding validation reports. Furthermore, the replacement records for forced fields and date formatting are updated in the data genealogy to record the complete source of each target field's data.
[0150] In this embodiment, the structured data packet adopts a layered JSON format, including a business data layer and a metadata layer. The business data layer stores target field data, such as "chief complaint: headache for 3 days" and "discharge diagnosis: J06.9 acute upper respiratory tract infection", which can serve as a clean and standardized data source that can be directly consumed by downstream clinical business coefficients (such as electronic medical record archiving and clinical research analysis). The metadata layer stores data lineage and verification reports, which are strongly associated with the business data layer, and records the quality status and extraction process of each target field data, so as to achieve complete traceability of the processing process.
[0151] Furthermore, prior to step S31, the following steps are also included: Step S30: Compile the declarative rules to generate data extraction instructions.
[0152] The declarative rules generated in step S2 are pre-compiled into high-performance internal instructions, namely data extraction instructions. For example, regular expression patterns in the rules are pre-compiled into reusable objects, and XPath expressions are optimized and cached, transforming "declarative intents" into "efficient execution instructions." Furthermore, the data extraction instructions are hot-deployed, supporting real-time dynamic updates without requiring service restarts, ensuring business continuity. Steps S31-S33 are executed according to the data extraction instructions, improving extraction performance.
[0153] Furthermore, all execution processes in step S3 are recorded and used for iterative optimization of declarative rules and preset rule bases. The method for extracting structured medical record text in multiple formats also includes: Step S4: Perform root cause analysis on the execution results of step S3, generate optimized recommendations, and update the declarative rules and / or the preset rule base according to the optimized recommendations; wherein the preset rule base also includes a verification knowledge base and a verification model.
[0154] The execution status of step S3 is recorded, and root cause analysis is performed on the execution results based on the recorded execution status to drive iterative optimization of declarative rules and preset rule base, specifically including steps S41-S43.
[0155] Step S41: Record the number of times step S3 is executed and the execution performance indicators.
[0156] Record the number of times step S3 is executed and the execution performance metrics. In this embodiment, the execution performance metrics include the processing success rate, the single-class verification failure rate, and the rule execution performance.
[0157] The success rate is the ratio of the number of records in the structured data package whose overall quality score (quality_score) for the target field data is higher than the passing threshold to the number of times step S3 is executed. The overall quality score for each target field data is calculated using a weighted scoring method. A base score (e.g., 80 points) is obtained for passing rule logic verification, and a base score is added (e.g., an addition ratio of 1.2) for passing standard terminology library verification. The score of lightweight NLP semantic verification is used as a correction coefficient.
[0158] The single-class validation failure rate is the ratio of the number of validation failures triggered by rule logic validation, standard terminology library validation, and lightweight NLP semantic validation to the number of times that type of validation is executed. It is used to determine the bottleneck in the overall quality score.
[0159] Rule execution performance, including average rule compilation time, average execution time, cache hit rate, etc., is used to evaluate and optimize the execution efficiency of multi-level verification.
[0160] Step S42: Based on the number of executions, the execution performance indicators, the data spectrum, and the verification report, perform root cause analysis on the execution results of step S3 to generate optimization recommendations; wherein, the optimization recommendations include at least one of rule configuration optimization recommendations and rule base optimization recommendations.
[0161] Based on the number of executions, performance metrics, data hierarchy, and validation report, root cause analysis is performed on the execution results of step S3 to generate optimization recommendations; among them, optimization recommendations include at least one of rule configuration optimization recommendations and rule base optimization recommendations.
[0162] This embodiment employs a Root Cause Analysis Engine (RCA Engine) to perform root cause analysis. It automatically aggregates and mines patterns from processing failures (where the overall quality score corresponding to the target field data is not higher than the acceptable threshold), validation failures, and quality events to determine the causes of failure and generate optimization recommendations. In this embodiment, the optimization recommendations include at least one of rule configuration optimization recommendations or rule base optimization recommendations.
[0163] For example, through association rule mining, it was discovered that "when the data source = Hospital A and the field = 'chief complaint', the failure rate of standard terminology database validation increases significantly." This generates a rule base optimization recommendation, adding new medical terms to the annotated terminology database of the validation knowledge base. For "frequent extraction failures due to writing variations," the reasons for failure are usually: the preset keyword group in the fuzzy tolerance anchor configuration does not cover the new variation, or the anchor distance is set too high. After clustering the genealogy information of such failure cases, the root cause analysis engine extracts the new writing variation reference words that caused the matching failure, such as "main complaint," and automatically generates optimization recommendations, including rule configuration optimization recommendations: "Add the variation reference word 'main complaint' to the chief complaint-keyword group" and "It is recommended to fine-tune the anchor distance from 0.85 to 0.82." The rule base optimization recommendation is: "Add the variation reference word 'main complaint' to the chief complaint-keyword group." For a large number of term mapping failures, a message will be displayed: "The diagnostic field contains the unincluded term 'XX syndrome'," generating optimization recommendations, including a rule base optimization recommendation: "It is recommended to update the ICD-10 dictionary."
[0164] Step S43: Update the declarative rules and / or the preset rule base according to the optimization recommendations.
[0165] Update the declarative rules and / or preset rule base according to the optimization recommendations, including one or more of steps S431 and S432.
[0166] Step S431: Update the declarative rules according to the rule configuration optimization recommendation.
[0167] Based on the rule configuration optimization recommendations, the declarative rules corresponding to the current extraction requirements are updated. Specifically, based on the rule configuration optimization recommendations, new variant reference words are added to the keyword and variant knowledge base of the preset rule base, or the anchor distance is adjusted, and declarative rules are regenerated. Based on the new declarative rules, the data is recompiled and hot-deployed to extract information and perform multi-level verification on the standard structured medical record data, resulting in a new structured data package.
[0168] Step S432: Update the preset rule base according to the rule base optimization recommendations.
[0169] The preset rule base is updated based on rule base optimization recommendations. New variant reference terms are added to the keyword and variant knowledge base of the preset rule base, or new medical terms are added to the validation knowledge base of the preset rule base.
[0170] Furthermore, manual review and collaboration can be added. Data with an overall quality score below the threshold or key validation failures can be transferred to the manual review queue. After the reviewers confirm or correct the results, the confirmed or corrected results (such as confirming a new variant reference word or correcting a mapping relationship) are used as high-quality feedback samples to optimize the keyword and variant knowledge base, validation knowledge base, and validation model of the preset rule base. For example, high-quality feedback samples can be used as training samples to train and update a lightweight NLP semantic model.
[0171] Through root cause analysis and optimized recommendations, the declarative rules for current extraction needs are updated, forming a data-driven closed loop of "processing-verification-analysis (root cause identification)-optimized recommendations (specific optimization actions)-rule configuration optimization-hot deployment update of declarative rules-reprocessing". The keyword and variant knowledge base, verification knowledge base, and verification model of the preset rule base are updated to achieve continuous evolution of information extraction capabilities and multi-level verification capabilities.
[0172] This invention provides a multi-format medical record text structure extraction system based on configuration rules, such as... Figure 2 As shown, it includes: An adaptive repair engine is used to identify the format of the original medical record data, determine the corresponding data format, repair and map the original medical record data based on the data format, and generate corresponding standard structure medical record data. The intelligent rule configuration center is used to configure rules and generate declarative rules based on the extracted requirements and a preset rule library; wherein, the declarative rules include anchor rules, and the anchor rules include precise anchor rules and / or fuzzy fault-tolerant anchor rules; An enhanced rule execution engine is used to extract information and perform multi-level verification on the standard structured medical record data according to the declarative rules, so as to obtain a structured data package corresponding to the extraction requirements; wherein, the structured data package includes target field data, data genealogy and verification report.
[0173] The above-described method and system embodiments are based on the same principles, and their related aspects can be referenced from each other to achieve the same technical effects. For specific implementation processes, please refer to the foregoing embodiments, which will not be repeated here.
[0174] Furthermore, embodiments of the present invention provide another multi-format medical record text structure extraction system based on configurable rules, including an adaptive repair engine, an intelligent rule configuration center, a verification-enhanced rule execution engine, and a quality governance platform.
[0175] (1) Adaptive Repair Engine An adaptive repair engine is used to identify the format of raw medical record data comprising several original document data, determine the data format corresponding to each of the original document data, and repair and map the raw medical record data based on the data format to generate corresponding standard structured medical record data. The format identification includes a first format identification; the first format identification includes identifier matching based on the format identifier of the raw medical record data and encoding analysis based on the character encoding range of the raw medical record data; the data format includes a markup language format.
[0176] Based on the data format, the original medical record data is repaired and mapped to generate corresponding standard structured medical record data, including: The original document data in the markup language format is subjected to a first fault-tolerant parsing to generate a corresponding primary parse tree and error records. Based on the primary parse tree and the error records, the original document data is subjected to a first intelligent repair to generate a corresponding well-structured document object. The first fault-tolerant parsing includes feature tag recognition and stack parsing of the original document data. The first intelligent repair includes tag completion and structure normalization. The well-structured document object is mapped and transformed to generate the corresponding standard structured medical record data.
[0177] In this embodiment, the adaptive repair engine includes an intelligent format detection and triage module, a progressive flowing repair pipeline, and a standard structure medical record data generation module.
[0178] (1.1) Intelligent format detection and traffic splitting module The intelligent format detection and routing module is used to identify the format of each original document data and determine the data format corresponding to each original document data; the data format includes markup language format and plain text format.
[0179] The intelligent format detection and routing module performs format recognition and method implementation steps S11 based on the same principle. Related aspects can be mutually referenced, and the same technical effects can be achieved. For specific implementation details, please refer to the aforementioned embodiments, which will not be repeated here.
[0180] The intelligent format detection and diversion module is also used to divert each raw document data to the corresponding progressive streaming repair pipeline based on the data format.
[0181] (1.2) Progressive flow repair pipeline In this embodiment, the progressive streaming repair pipeline includes an XML / HTML processing pipeline and a plain text processing pipeline. If the original document data corresponds to a markup language format, the original document data is diverted to the XML / HTML processing pipeline; if the original document data corresponds to a plain text format, the original document data is diverted to the plain text processing pipeline.
[0182] The XML / HTML processing pipeline is used to perform first-level fault-tolerant parsing and first-level intelligent repair on the raw document data in markup language format, generating well-structured document objects.
[0183] This method is based on the same principle as step S121 in the embodiment of the method, and the relevant parts can be borrowed from each other to achieve the same technical effect. For the specific implementation process, please refer to the foregoing embodiment, which will not be repeated here.
[0184] The plain text processing pipeline is used to perform secondary fault-tolerant parsing and secondary intelligent repair on raw document data in plain text format, generating structured text objects.
[0185] This method is based on the same principle as step S123 in the embodiment of the method, and the relevant parts can be borrowed from each other and can achieve the same technical effect. For the specific implementation process, please refer to the foregoing embodiment, which will not be repeated here.
[0186] (1.3) Unified Document Model Module The Unified Document Model module is used to map and transform well-structured document objects and structured text objects to generate corresponding standard structured medical record data.
[0187] Steps S122 and S124 of the method embodiment are based on the same principle, and their related aspects can be mutually referenced, achieving the same technical effect. For specific implementation process, please refer to the foregoing embodiments, which will not be repeated here.
[0188] (2) Intelligent rule configuration center The intelligent rule configuration center is used to configure rules and generate declarative rules based on extraction requirements using a preset rule library; wherein, the declarative rules include anchor rules and extraction rules, the anchor rules include precise anchor rules and / or fuzzy fault-tolerant anchor rules; the preset rule library includes a keyword and variant knowledge base; Based on the extraction requirements, rules are configured using a pre-defined rule base to generate declarative rules, including: Based on the target words for extraction requirements and the keyword and variant knowledge base, configure and generate the anchor point rules; Based on the information extraction range of the extraction requirements, the extraction rules are configured and generated.
[0189] In this embodiment, the preset rule base is stored in the intelligent rule configuration center. The construction of the preset rule base, as well as the configuration of rules for extraction requirements and the generation of declarative rules, are based on the same principle as step S3 in the method embodiment. The relevant parts can be referenced from each other and can achieve the same technical effect. For the specific implementation process, please refer to the foregoing embodiments, which will not be repeated here.
[0190] In this embodiment, a unified rules description language (URLD) that is independent of format is used to achieve complete decoupling of business logic (requirement extraction) from program code and underlying data format.
[0191] (3) Verify the enhanced rule execution engine An enhanced rule execution engine is validated to determine the extraction start point position in the standard structured medical record data based on the precise anchor point rules and / or the fuzzy fault-tolerant anchor point rules. Information is extracted from the extraction start point position according to the extraction rules to obtain candidate field data and the corresponding data hierarchy. Multi-level validation is performed on the candidate field data to obtain a structured data package corresponding to the extraction requirements. The structured data package includes target field data, the data hierarchy, and a validation report. The multi-level validation includes one or more of rule logic validation, standard terminology database validation, and lightweight NLP semantic validation.
[0192] The enhanced rule execution engine is used to efficiently and accurately run the declarative rules generated by the intelligent rule configuration center on standard structured medical record data. It performs precise positioning and / or fuzzy positioning, precise extraction and real-time multi-level verification, and generates a structured data package including target field data, data genealogy and verification report, realizing the whole process from data positioning to trusted output.
[0193] Furthermore, the enhanced rule execution engine includes a two-layer fault-tolerant anchor module, an extractor, and a real-time verification and transformation pipeline.
[0194] (3.1) Double-layer fault-tolerant anchor point module The dual-layer fault-tolerant anchor point module is used to locate anchor points in the standard structured medical record data based on the precise anchor point rules and / or the fuzzy fault-tolerant anchor point rules, and to determine the starting point position for extraction.
[0195] This method is based on the same principle as step S31 in the method embodiment, and the relevant parts can be borrowed from each other to achieve the same technical effect. For the specific implementation process, please refer to the foregoing embodiments, which will not be repeated here.
[0196] Furthermore, the dual-layer fault-tolerant anchor point module includes a precise anchor point layer, a fuzzy fault-tolerant anchor point layer, and a fusion layer. The precise anchor point layer is used to locate anchor points in standard structured medical record data based on precise anchor point rules, the fuzzy fault-tolerant anchor point layer is used to locate anchor points in standard structured medical record data based on fuzzy fault-tolerant anchor point rules, and the fusion layer is used to determine the extraction starting point position based on the positioning results of the precise anchor point layer and / or the fuzzy fault-tolerant anchor point layer.
[0197] (3.2) Extractor The extractor is used to extract information from standard structured medical record data based on the starting point and extraction rules, and to obtain candidate field data and corresponding data genealogy.
[0198] This method is based on the same principle as step S32 in the embodiment of the method, and the relevant parts can be borrowed from each other to achieve the same technical effect. For the specific implementation process, please refer to the foregoing embodiment, which will not be repeated here.
[0199] (3.3) Real-time verification and conversion pipeline The real-time validation and transformation pipeline is used to perform multi-level validation on candidate field data according to multi-level validation rules, and to standardize the candidate field data that passes multi-level validation to obtain a structured data package corresponding to the extraction requirements. The structured data package includes target field data, data genealogy, and validation report.
[0200] This method is based on the same principle as step S33 in the embodiment of the method, and the relevant parts can be borrowed from each other to achieve the same technical effect. For the specific implementation process, please refer to the foregoing embodiment, which will not be repeated here.
[0201] Furthermore, the enhanced rule execution engine also includes a rule compilation and hot deployment module, which compiles the declarative rules generated by the intelligent rule configuration center into data extraction instructions and hot-deploys these instructions within the enhanced rule execution engine. This rule compilation and hot deployment capability enables the optimization and updating of declarative rules to be completed within minutes through configuration, eliminating the need for development and deployment, thus greatly improving business agility.
[0202] In this embodiment, the enhanced rule execution engine is designed as a stateless service that can be horizontally scaled. Through the collaboration of a central rule repository and a distributed cache, consistency and high performance are ensured during concurrent processing of massive amounts of medical records, meeting enterprise-level big data processing needs.
[0203] (4) Quality governance platform The quality governance platform is used to perform root cause analysis on the execution results of the verification-enhanced rule execution engine, generate optimization recommendations, and update the preset rule base of the declarative rules and / or intelligent rules configuration center based on the optimization recommendations; the preset rule base also includes a verification knowledge base and a verification model.
[0204] This method is based on the same principle as step S4 in the embodiment of the method, and the relevant parts can be borrowed from each other to achieve the same technical effect. For the specific implementation process, please refer to the foregoing embodiment, which will not be repeated here.
[0205] This invention provides an example of the execution flow of a multi-format medical record text structure extraction system based on configurable rules, taking as an example a scenario where discharge summaries from different channels, including exported HTML web pages with non-standard formats, standard XML, and plain text records, require the extraction of the "chief complaint" and "discharge diagnosis" fields.
[0206] Step 1: Repair the original medical record data.
[0207] The original HTML tag "chief complaint" in the original medical record data was mismatched. The adaptive repair engine automatically corrected it and generated standard structured medical record data.
[0208] Step 2: Configure rules in the intelligent rule configuration center to generate declarative rules.
[0209] Regarding the extraction of the target "chief complaint": Anchor point rule settings: Enable "fuzzy fault-tolerant anchor points", select main complaint - keyword group, set the anchor distance to 0.85, select "edit distance" for the anchor distance algorithm, and set the anchor range to the title field.
[0210] Extraction rule settings: Configure regular expressions to capture the content after the colon.
[0211] Multi-level validation rule settings: Add "Lightweight NLP Semantic Validation" and link a lightweight symptom description classification model.
[0212] For extracting the target "discharge diagnosis": Anchor point rule settings: Enable "Precise Anchor Point", set the precise anchor point to "Discharge Diagnosis", and set the anchor range to the text field.
[0213] Extraction rule settings: Configure regular expressions to capture the content after the colon.
[0214] Multi-level validation rule settings: Add "Standard Terminology Library Validation" and set it to ICD-10 mapping.
[0215] Step 3: Verify the enhanced rule execution engine's execution information extraction and multi-level verification, and output a structured data packet.
[0216] Target field data 1: The enhanced rule execution engine successfully identified "[chief complaint]" in the standard structure medical record data corresponding to the HTML format by calculating the edit distance, extracted the text "headache for 3 days" (i.e. candidate target data), and obtained a high comprehensive quality score through lightweight NLP semantic verification.
[0217] Target field data 2: Successfully identified "chief complaint:" in the standard structured medical record data corresponding to plain text format, extracted the text "cough with fever for 2 days" (i.e. candidate target data), and obtained a high comprehensive quality score through lightweight NLP semantic verification.
[0218] Target field data 3: Successfully identify "discharge diagnosis" in the standard structure medical record data corresponding to the XML format, extract the diagnosis text (i.e., candidate target data), perform standard terminology database verification, and record the ICD-10 encoding mapping results.
[0219] The output is a structured data packet, including the target field data corresponding to "chief complaint" and "discharge diagnosis", and records the pass status and details of each channel verification (e.g., the ICD-10 encoding mapping result of the diagnosis).
[0220] Step 4: The quality governance platform performs root cause analysis and optimizes declarative rules and preset rule base.
[0221] The quality governance platform's monitoring revealed a sudden drop in the fuzzy matching success rate of the "Main Complaint" field in data from a certain source. Analysis indicated that a new title variant, "Main Request:", had appeared in this batch of data. It was recommended to copy and modify the rules for this source, adding the variant reference term "Main Request" to the preset rule base and generating a new declarative rule. After the new declarative rule was hot-deployed, the data extraction success rate of this source immediately recovered.
[0222] Based on the above examples, the structured extraction system for formatted medical records proposed in this embodiment not only achieves efficient and accurate extraction of medical records in multiple formats, but also demonstrates strong practicality, adaptability and sustainable evolution capabilities through its built-in intelligent fault tolerance, real-time verification and data-driven optimization loop.
[0223] In summary, the multi-format medical record text structure extraction method and system based on configuration rules according to embodiments of the present invention has at least one of the following beneficial effects: 1. By repairing the original medical record data, medical record documents in multiple formats such as text, XML, and HTML are transformed into queryable standard structured medical record data; declarative rules are generated based on the configuration of the preset rule base, and anchor point rules are integrated to extract information and perform multi-level verification on the standard structured medical record data. This achieves highly adaptable and accurate structured extraction of medical record documents in multiple formats such as text, XML, and HTML, and simultaneously obtains data genealogy and verification reports, providing credibility assurance for the extracted target field data.
[0224] 2. By incorporating fuzzy fault-tolerant anchor point rules into declarative rules, we can effectively address issues such as writing variations, non-standard writing, and OCR noise in real-world medical record texts. This improves the robustness and coverage of information extraction for non-standardized and noisy text, ensuring that the starting point for extraction can still be reliably found even when medical record data is not written in a standard way. This solves the problem of failure caused by minor format differences in traditional rule-based methods.
[0225] 3. By performing root cause analysis on the results of information extraction and multi-level verification, optimization recommendations are generated, and declarative rules and preset rule bases are updated to achieve continuous evolution of information extraction capabilities and multi-level verification capabilities, enabling the multi-format medical record text structure extraction method of this application to have optimization closed-loop capabilities.
[0226] 4. By constructing an adaptive repair engine with non-standard data repair capabilities, an intelligent rule configuration center supporting fuzzy fault-tolerant anchor point rules, and a verification-enhanced rule execution engine integrating multi-level verification and capable of obtaining data lineages and verification reports, the system achieves efficient, accurate, and reliable processing of multi-format heterogeneous medical records. Furthermore, by constructing a quality governance platform with root cause analysis and optimization recommendations, the intelligent rule configuration center and verification-enhanced rule execution engine are optimized and updated, enabling the multi-format medical record text structure extraction system to possess continuous evolution capabilities.
[0227] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0228] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. < / h1> < / section> < / txt> < / main> < / clinicaldocument>
Claims
1. A method for structured extraction of multi-format medical record text based on configuration rules, characterized in that, include: Step S1: For the original medical record data including several original document data, the format is identified according to the format identifier and / or character encoding range in each original document data, the data format corresponding to each original document data is determined, and the original medical record data is repaired and mapped based on the data format to generate corresponding standard structure medical record data; wherein, the data format includes markup language format; Based on the data format, the original medical record data is repaired and mapped to generate corresponding standard structured medical record data, including: The original document data in the markup language format is subjected to a first fault-tolerant parsing to generate a corresponding primary parse tree and error records. Based on the primary parse tree and the error records, the original document data is subjected to a first intelligent repair to generate a corresponding well-structured document object. The first fault-tolerant parsing includes the identification of feature tags and stack parsing of the original document data. The first intelligent repair includes tag completion and structure normalization. The well-structured document object is mapped and transformed using a unified document model to generate the corresponding standard structured medical record data. The original document data in the markup language format undergoes a first fault-tolerant parsing process to generate a corresponding primary parse tree and error records. Based on the primary parse tree and the error records, the original document data undergoes a first intelligent repair process to generate a well-structured document object, including: A tolerant scanner is used to perform syntactic analysis on the original document data in the markup language format to identify the feature tags in the original document data; Based on the feature tags, the original document data is parsed using the stack parsing to generate the corresponding primary parse tree and error records; Based on the primary parse tree and the error record, the original document data is processed by tag completion and structural normalization to generate the well-structured document object; Step S2: Based on the extraction requirements, configure rules using a preset rule base to generate declarative rules; wherein, the declarative rules include anchor rules and extraction rules, the anchor rules include precise anchor rules and / or fuzzy fault-tolerant anchor rules; the preset rule base includes a keyword and variant knowledge base; Based on the extraction requirements, rules are configured using a pre-defined rule base to generate declarative rules, including: Based on the target words for extraction requirements and the keyword and variant knowledge base, configure and generate the anchor point rules; Based on the information extraction range of the extraction requirements, configure and generate the extraction rules; Configure the rules for generating the anchor points, including: Set precise anchor points and anchoring ranges to generate the precise anchor point rules; and / or set fuzzy anchor points, anchoring distances, and anchoring ranges to generate the fuzzy fault-tolerant anchor point rules; Step S3: Based on the precise anchor point rule and / or the fuzzy fault-tolerant anchor point rule, determine the extraction starting point position in the standard structured medical record data. Extract information from the extraction starting point position according to the extraction rule to obtain candidate field data and corresponding data lineage. Perform multi-level verification on the candidate field data to obtain the structured data package corresponding to the extraction requirement. The structured data package includes target field data, the data lineage, and a verification report. The multi-level verification includes one or more of rule logic verification, standard terminology base verification, and lightweight NLP semantic verification.
2. The method according to claim 1, characterized in that, Also includes: Step S4: Perform root cause analysis on the execution results of step S3, generate optimization recommendations, and update the declarative rules and / or the preset rule base according to the optimization recommendations; wherein, the preset rule base also includes a verification knowledge base and a verification model.
3. The method according to claim 1, characterized in that, The data format also includes plain text format; Based on the data format, the original medical record data is repaired and mapped to generate corresponding standard structured medical record data, which also includes: The original document data in the plain text format is subjected to a second fault-tolerant parsing and a second intelligent repair to generate a structured text object; The structured text object is mapped and transformed to generate the corresponding standard structured medical record data.
4. The method according to claim 1, characterized in that, The format recognition includes first format recognition and second format recognition; For original medical record data comprising several original document data, format identification is performed based on the format identifier and / or character encoding range in each original document data to determine the data format corresponding to each original document data, including: For each of the original document data, the first format identification is performed based on the format identifier and / or the character encoding range of the original document data. If the first format identification is successful, the data format corresponding to the original document data is determined based on the identification result of the first format identification. The first format identification includes identifier matching based on the format identifier of the original medical record data and encoding analysis based on the character encoding range of the original medical record data. If the first format recognition fails, then the original document data is subjected to data format feature extraction based on the format identifier of the original document data, and the data format features are input into a lightweight classification model for the second format recognition to determine the data format corresponding to the original document data.
5. The method according to claim 3, characterized in that, The second fault-tolerant parsing includes code detection and error correction, and text segmentation; the second intelligent repair includes paragraph recombination and structural simulation. The original document data in the plain text format undergoes a second fault-tolerant parsing and a second intelligent repair to generate a structured text object, including: The original document data in the plain text format is subjected to encoding detection and error correction to generate a Unicode character stream; The Unicode character stream is divided into text segments based on newline characters to generate a sequence of text lines. The paragraph recombination is performed on the text line sequence to generate logical paragraphs; The structure of the logical paragraphs is simulated to generate the structured text object.
6. The method according to claim 1, characterized in that, The declarative rules also include multi-level validation rules; The candidate field data is validated at multiple levels to obtain the structured data package corresponding to the extraction requirements, including: According to the multi-level verification rules, the candidate field data is subjected to multi-level verification, and the candidate field data that passes the multi-level verification is standardized to obtain the target field data and the corresponding verification report.
7. The method according to claim 1, characterized in that, The fuzzy fault-tolerant anchor point rule includes fuzzy anchor point, anchoring distance, and anchoring range; Based on the aforementioned fuzzy fault-tolerant anchor point rule, the starting point position for extraction is determined in the standard structured medical record data, including: Based on the anchoring range, candidate nodes are determined from all text nodes of the standard structure medical record data, and the text distance between the fuzzy anchor point and the candidate node is calculated. If the text distance is not less than the anchoring distance, then the position of the candidate node in the standard structure medical record data is determined as the extraction starting point position.
8. The method according to claim 1, characterized in that, Based on the precise anchor point rules and the fuzzy fault-tolerant anchor point rules, the starting point position for extraction is determined in the standard structured medical record data, including: Based on the precise anchor point rules, anchor points are located in the standard structured medical record data to determine the first starting point position; Based on the fuzzy fault-tolerant anchor point rule, anchor point positioning is performed on the standard structured medical record data to determine the second starting point position and the corresponding confidence level; The extraction starting point position is determined based on the first starting point position, the second starting point position, and the corresponding confidence level.
9. A multi-format medical record text structure extraction system based on configuration rules, characterized in that, include: An adaptive repair engine is used to identify the format of original medical record data, which includes several original document data, based on the format identifier and / or character encoding range in each original document data, determine the data format corresponding to each original document data, repair and map the original medical record data based on the data format, and generate corresponding standard structured medical record data; wherein, the data format includes markup language format; Based on the data format, the original medical record data is repaired and mapped to generate corresponding standard structured medical record data, including: The original document data in the markup language format is subjected to a first fault-tolerant parsing to generate a corresponding primary parse tree and error records. Based on the primary parse tree and the error records, the original document data is subjected to a first intelligent repair to generate a corresponding well-structured document object. The first fault-tolerant parsing includes feature tag recognition and stack parsing of the original document data. The first intelligent repair includes tag completion and structure normalization. The well-structured document object is mapped and transformed using a unified document model to generate the corresponding standard structured medical record data. The original document data in the markup language format undergoes a first fault-tolerant parsing process to generate a corresponding primary parse tree and error records. Based on the primary parse tree and the error records, the original document data undergoes a first intelligent repair process to generate a well-structured document object, including: A tolerant scanner is used to perform syntactic analysis on the original document data in the markup language format to identify the feature tags in the original document data; Based on the feature tags, the original document data is parsed using the stack parsing to generate the corresponding primary parse tree and error records; Based on the primary parse tree and the error record, the original document data is processed by tag completion and structural normalization to generate the well-structured document object; The intelligent rule configuration center is used to configure rules and generate declarative rules based on extraction requirements using a preset rule library; wherein, the declarative rules include anchor rules and extraction rules, the anchor rules include precise anchor rules and / or fuzzy fault-tolerant anchor rules; the preset rule library includes a keyword and variant knowledge base; Based on the extraction requirements, rules are configured using a pre-defined rule base to generate declarative rules, including: Based on the target words for extraction requirements and the keyword and variant knowledge base, configure and generate the anchor point rules; Based on the information extraction range of the extraction requirements, configure and generate the extraction rules; Configure the rules for generating the anchor points, including: Set precise anchor points and anchoring ranges to generate the precise anchor point rules; and / or set fuzzy anchor points, anchoring distances, and anchoring ranges to generate the fuzzy fault-tolerant anchor point rules; An enhanced rule execution engine is validated to determine the extraction start point position in the standard structured medical record data based on the precise anchor point rules and / or the fuzzy fault-tolerant anchor point rules. Information is extracted from the extraction start point position according to the extraction rules to obtain candidate field data and corresponding data lineages. Multi-level validation is performed on the candidate field data to obtain a structured data package corresponding to the extraction requirements. The structured data package includes target field data, the data lineage, and a validation report. The multi-level validation includes one or more of rule logic validation, standard terminology database validation, and lightweight NLP semantic validation.
Citation Information
Patent Citations
A medical record document standardize processing system and method
CN109408635A
Electronic medical record structured method and system and related equipment
CN111352987A