Medical text structured intelligent processing system and method based on local large model
Patent Information
- Application Number
- CN202610756308.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2046-05-29
AI Technical Summary
具体解决静态术语映射导致的语义失真问题、语义掩码技术造成的信息碎片化问题和云端部署带来的数据安全与合规风险问题
[0025]本发明的有益效果在于:本发明通过本地轻量化部署彻底规避了传统医疗 NLP模型依赖云端传输带来的隐私泄露风险,完全能够符合医疗数据监管条件。相比于现有技术中粗暴的语义掩码方案,本系统引入医学逻辑规则库,利用人工智能算法剥离噪声,精准识别并修正诊疗记录中的逻辑冲突。这种兼顾无损降噪与合规审计的技术闭环填补了行业在医疗数据监管与逻辑校验领域的空白。具体效果如下:
Smart Images

Figure CN122287613B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical data processing and natural language processing technology, and relates to a medical text structured intelligent processing system and method based on a local large model. Background Technology
[0002] With the rapid development of medical informatization, the structured and standardized processing of clinical free text has become a core prerequisite for unlocking the value of medical big data. Existing medical text processing technologies mainly include the following: (1) Standardized mapping technology based on static terminology database: Early technologies, such as the UMLS system published by Olivier Bodenreider et al. in the paper "The Unified Medical Language System (UMLS): integrating biomedical terminology", mainly relied on building a large standard terminology database for semantic mapping. However, this strongly coupled mapping scheme not only faces the dilemma of extremely high cost of dynamic updating and maintenance of the terminology database, but also easily destroys the subtle semantics of the original text by forcibly replacing personalized descriptions with standard terms, resulting in poor adaptability.
[0003] (2) Privacy Protection Technology Based on Semantic Masking: In the paper "Automatic de-identification of textual documents in the electronic health record: a review of recent research", Stephane M Meystre et al. proposed an automatic de-identification framework based on semantic categories and discussed semantic masking and de-privacy technology. However, semantic masking technology has an irreconcilable inherent contradiction: excessive masking leads to fragmentation of key clinical information, while insufficient masking leaves the risk of re-identification. More importantly, semantic masking is a destructive privacy protection method and cannot support the complete information restoration required for downstream clinical decision-making, which contradicts the data quality requirements of real-world research.
[0004] (3) Cloud-based NLP technology: Most current medical NLP models rely heavily on cloud computing power. Data is at great risk of leakage during cross-domain transmission and lacks supporting localized traceability and anti-tampering mechanisms, making it difficult to meet the stringent requirements of medical supervision for data security and compliance throughout the entire process.
[0005] Existing processing methods mostly focus on surface-level text correction and simple transformation, failing to address the core dimension of medical logical consistency. This often results in inherent logical conflicts between symptoms, examination results, and medication plans in the processed text, leading to a severe lack of data credibility. Furthermore, current medical text processing is often limited to large cloud-based models, neglecting the local confidentiality requirements of medical information and the limitations of computing power. Summary of the Invention
[0006] In view of this, the purpose of this invention is to provide a medical text structuring intelligent processing system and method based on a local large-scale model. Through a locally deployed lightweight large-scale model, it achieves accurate noise reduction without the need for semantic masking, and identifies and corrects text conflicts based on medical logic. This allows for the complete preservation of core clinical information while constructing a localized processing closed loop with secure traceability capabilities, filling the gaps in existing technologies regarding logical review and compliance control. Specifically, it addresses the semantic distortion problem caused by static terminology mapping, the information fragmentation problem caused by semantic masking technology, and the data security and compliance risks brought about by cloud deployment.
[0007] To achieve the above objectives, the present invention provides the following technical solution: Solution 1: A method for intelligent structural processing of medical text based on a local large model, specifically including the following steps: S1: Lightweight model adaptation in intranet environment: In a physically isolated medical intranet environment, the model size is compressed by performing weight quantization on mainstream open-source medical models.
[0008] S2: Medical semantic parsing and positive sign identification: A semantic parsing and positive screening mechanism based on medical context enhancement is used to identify positive medical entities; the semantic parsing and positive screening mechanism based on medical context enhancement uses the hidden layer representation of a large model to calculate the semantic information entropy of each word, identify key medical entities and their attribute states in the text, and automatically filter negative descriptions through existence determination, retaining only positive signs.
[0009] S3: Structured representation of positive signs: Convert the positive medical entities identified in step S2 into standardized structural outputs.
[0010] S4: Use a logical self-consistency evaluation function to quantify the inherent rationality of medical records.
[0011] S5: Localized security audit mechanism, specifically including: For each processed medical text, the system will generate a block containing the original text hash, processing operator characteristics, operator signature and timestamp, to ensure that the hard requirements of medical data traceability and compliance supervision are met locally, making the data controllable and traceable.
[0012] Furthermore, in step S1, low-rank adaptive (LoRA) fine-tuning technology is used to compress the model size so that it can be adapted to the existing general-purpose servers or edge computing devices of medical institutions.
[0013] Furthermore, in step S2, the semantic parsing and positive screening mechanism based on medical context enhancement specifically includes: defining medical entities. Existence state These correspond to the three cases of existence, non-existence, and not mentioned, respectively; the degree of definition is modified. The system quantifies the severity of the description; it calculates terms by constructing a medical context-aware attention mechanism. Importance weight in the medical context for:
[0014] in, ∈[0,1] is the balance factor, which is determined by maximizing the accuracy of positive signs extraction on the validation set; TF-IDF refers to word frequency-inverse document frequency, a statistical method used to evaluate the importance of a word in a text or corpus. The local deployment of this system requires that this statistic be a static parameter pre-calculated based on the local historical desensitized medical record corpus of medical institutions, and updated regularly during the accumulation of the corpus. Indicates terms In medical corpus The core purpose of semantic information entropy in context is to leverage the sensitivity of large models to medical contexts to identify words with high information content in specific diagnostic and treatment contexts, thereby defining the importance of clinical semantics. Its calculation method is as follows:
[0015] in, For medical context categories (such as symptom descriptions, examination results, and medical history statements); Terms output for large models Belongs to context The posterior probability is calculated from the last hidden state through the Softmax layer; This represents the total number of medical context categories. A negative trigger word detection and positive screening mechanism is adopted; the set of negative trigger words N = {none, none, not, deny, not seen, disappear, alleviate, exclude, reject} is defined, and a negative scope resolution function is constructed:
[0016] in, Negation trigger word For target terms The probability of determining the scope. For the detected first k A negative trigger word, and These are the hidden representations of the negation word and the target word, respectively. Encode the relative positions of the two words in the sentence. For learnable parameter vectors, For the Sigmoid function; when When the value is greater than the threshold, the term is determined. Its existence state within the negation scope This entity will be filtered out and will not participate in subsequent structured output.
[0017] Furthermore, in step S2, based on the definitions of degree adverbs in the 5th edition of the *Modern Chinese Dictionary* and clinical physician opinions, the system constructs a degree adverb hierarchical mapping table for degree modifier identification. For example, D={(mild, 0.2), (mild, 0.3), (moderate, 0.5), (obvious, 0.7), (significant, 0.8), (large, 0.85), (serious, 0.9), (extreme, 1.0)}, etc., and determines the degree value through fuzzy matching and semantic similarity calculation.
[0018] in, Adjectives that modify the target entity For semantic similarity based on BioBERT encoding, As a preset level value, For the j-th degree adverb entry in the degree adverb hierarchy mapping table, This is a pre-defined hierarchy mapping table for medical degree adverbs. If no explicit degree modifier is specified, the default value is used. =0.5.
[0019] Furthermore, in step S3, the structured output format is defined as a quadruple:
[0020] The system uses standardized medical entity mapping to link extracted entities to a standard terminology database (such as ICD-10 and SNOMED CT), and the mapping score is used for mapping. Calculated using cosine similarity:
[0021] in, To extract the semantic vector of an entity, Vector representation of standard terminology.
[0022] Furthermore, in step S4, a logical consistency evaluation function is used to quantify the internal rationality of the medical records, and its logical conflict risk score is used. The calculation process is defined as follows:
[0023] in, It is the k-th type of medical logical association matrix. The total number of categories in the medical logical association matrix. These are the importance weights for the corresponding dimensions; if the calculated score... If the data falls below a certain threshold, a logical conflict warning will be triggered, prompting manual review and thus improving data credibility.
[0024] Option 2: A medical text structuring intelligent processing system based on a local large model. This system relies on medical logical self-consistency rules to achieve noise reduction, conflict verification, and standardization of free medical text. Specifically, it includes: The lightweight model adaptation module for intranet environments compresses the size of mainstream open-source medical models by performing weight quantization on them in physically isolated medical intranet environments. The medical semantic parsing and positive sign recognition module uses a medical context-enhanced semantic parsing and positive screening mechanism to identify positive medical entities. The positive sign structured representation module converts the identified positive medical entities into standardized structural outputs. The dynamic logic verification module uses a logic self-consistency evaluation function to quantify the inherent rationality of medical records; The localized security audit module for medical compliance adopts a localized security audit mechanism to ensure that the hard requirements of medical data traceability and compliance supervision are met locally, making the data controllable and traceable.
[0025] The beneficial effects of this invention are as follows: This invention completely avoids the privacy leakage risks associated with traditional medical NLP models relying on cloud transmission through local lightweight deployment, fully complying with medical data regulatory requirements. Compared to the crude semantic masking schemes in existing technologies, this system introduces a medical logic rule base, utilizing artificial intelligence algorithms to remove noise and accurately identify and correct logical conflicts in medical records. This closed-loop technology, which balances lossless noise reduction and compliant auditing, fills a gap in the industry's medical data supervision and logic verification fields. Specific effects are as follows: (1) To address the semantic distortion problem caused by static terminology mapping, this invention aims to provide a semantic parsing method based on a large model. By introducing a medical context-enhanced attention mechanism and a degree modification recognition algorithm, the system can achieve accurate entity extraction and standardized mapping while preserving clinical descriptions, achieving a balance between semantic integrity and terminology standardization, and supporting the demand for high-quality clinical data in real-world research.
[0026] (2) To address the information fragmentation problem caused by semantic masking technology, this invention aims to construct a data processing mode that combines lossless noise reduction with logical self-consistency verification. By deploying a local model to avoid cross-domain data transmission, and combining the medical logic rule base to perform conflict detection and correction on the extracted content, the system can achieve privacy compliance without relying on semantic masking, ensuring the continuity, authenticity and availability of medical records for clinical decision-making.
[0027] (3) In response to the data security and compliance risks brought about by cloud deployment, this invention aims to achieve end-to-end localized processing in a physically isolated environment. By fine-tuning the open-source medical model locally, the system is adapted to general servers and even edge computing devices in medical institutions. At the hardware level, the data outflow path is cut off, while an anti-tampering audit chain is established to meet the requirements of medical industry standards.
[0028] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0029] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 Flowchart of a medical free text noise reduction mapping system; Figure 2 Flowchart for the medical semantic parsing and positive sign recognition module; Figure 3 This is a flowchart illustrating the complete closed-loop process of data processing and compliance auditing within a medical intranet environment. Detailed Implementation
[0030] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0031] Example 1: Please see Figures 1-3 This embodiment provides a medical text structuring intelligent processing system based on a local large-scale model. It constructs a closed-loop processing system covering the entire process from local deployment to logical verification and security tracing. The core of its technical solution lies in the precise extraction of structured elements through deep integration of deterministic medical logic rules and a probabilistic lightweight large-scale model. This overcomes the difficulties of traditional structuring methods in accurately extracting required fields, negative semantics such as presence / absence, and degree modifiers. The system mainly consists of four core modules: a lightweight model adaptation module for intranet environments, a medical semantic parsing and positive sign recognition module, a positive sign structured representation module, a dynamic logic verification module, and a localized security audit module for medical compliance.
[0032] The system first compresses the model size by performing weight quantization on mainstream open-source medical models within a physically isolated medical intranet environment. This weight quantization process includes mapping the original floating-point precision model weights to low-bit integer representations using INT8 or INT4 linear quantization. Calculate the quantized value, where For the original weights, Scaling factor This is the zero-point offset, which is dequantized during inference. The calculations are then performed using approximate floating-point values. Building upon weight quantization, a model pruning strategy is employed to remove redundant attention heads and fully connected layer parameters. This is combined with low-rank adaptive (LoRA) fine-tuning to freeze the main parameters of the original model and train only the low-rank matrix. and (Right now ,in This refers to the weight matrix that is frozen and not involved in fine-tuning in the original pre-trained large model. This is the updated weight matrix used for inference or forward propagation after LoRA fine-tuning, ensuring the compressed model size is suitable for the memory and computing power constraints of general-purpose servers or edge computing devices in medical institutions. This completely cuts off the data outflow path at the physical level, ensuring all diagnostic and treatment data is processed within the intranet environment.
[0033] In the medical semantic parsing and positive sign recognition module, the system addresses the problem of extracting symptom-presence-degree triples from free clinical text by designing a semantic parsing and positive sign filtering mechanism based on medical context enhancement. This mechanism utilizes the hidden layer representation of a large model to calculate the semantic information entropy of each word, identifies key medical entities and their attribute states in the text, and automatically filters negative descriptions through existence determination, retaining only positive signs for structured output. Specifically, it defines medical entity e... i Existence state S(e) i )∈{1,0,-1}, corresponding to the three cases of existence, non-existence, and not mentioned, respectively; the degree modifier D(e) is defined. i The term t ∈ [0,1] quantifies the severity of the description. The system constructs a medical context-aware attention mechanism to calculate the severity of the term t. i Importance weight in the medical context for:
[0034] in, ∈[0,1] is a balance factor, determined by maximizing the accuracy of positive sign extraction on the validation set. TF-IDF refers to Term Frequency-Inverse Document Frequency, a statistical method used to evaluate the importance of a word in a text or corpus. The local deployment of this system requires that this statistic be a static parameter pre-calculated based on the local historical de-identified medical record corpus of medical institutions, and updated periodically during the corpus accumulation process. Indicates term t i In medical corpus The core purpose of semantic information entropy in context is to leverage the sensitivity of large models to medical contexts to identify words with high information content in specific diagnostic and treatment contexts, thereby defining the importance of clinical semantics. Its calculation method is as follows:
[0035] in, For medical context categories (such as symptom descriptions, examination results, and medical history statements). Terms output for large models Belongs to context The posterior probability is calculated from the last hidden state through the Softmax layer. Addressing the key challenge of negation semantic recognition, the system introduces a negation trigger word detection and positive screening mechanism. The set of negation trigger words N = {none, none, not, deny, not seen, disappeared, alleviated, excluded, rejected} is defined, and a negation scope parsing function is constructed:
[0036] in, For the detected first k A negative trigger word, and These are the hidden representations of the negation word and the target word, respectively. Encode the relative positions of the two words in the sentence. For learnable parameter vectors, For the Sigmoid function. When When the value is greater than the threshold, the term is determined. Within the negation domain, its existence state S(e) i If the expression is 0, the entity will be filtered out and will not participate in subsequent structured output. For degree modifier recognition, the system constructs a degree adverb hierarchical mapping table, such as D={(slight, 0.2), (mild, 0.3), (moderate, 0.5), (obvious, 0.7), (significant, 0.8), (heavy, 0.85), (severe, 0.9), (extreme, 1.0)}, etc., and determines the degree value through fuzzy matching and semantic similarity calculation.
[0037] in, Adjectives that modify the target entity For semantic similarity based on BioBERT encoding, This is the preset level value. If no explicit level modifier is specified, the default value will be used. =0.5.
[0038] Based on the above analysis results, the structured representation generation module converts positive medical entities (i.e., S(ei)=1) into standardized structured output. The structured output format is defined as a quadruple:
[0039] The system uses medical entity standardization mapping to link extracted entities to a standard terminology database (such as ICD-10 and SNOMED CT), and the mapping score is calculated using cosine similarity.
[0040] in, To extract the semantic vector of an entity, Vector representation of standard terminology.
[0041] The dynamic logic verification module addresses the lack of logical verification in existing technologies. This module will process the previously extracted standard medical entities. , The abstraction is performed and projected into a multidimensional tensor space to generate corresponding entity feature vectors. and The system relies on a pre-built, rigorous medical logic rule base to construct a medical logic association matrix covering multiple complications, medication contraindications, and other dimensions. It then performs deep-level interaction and self-consistency calculations on these originally isolated entity features. By designing a highly sensitive logical self-consistency assessment method, the system quantifies the internal rationality of medical records, with its core logical conflict risk scoring... The calculation process is rigorously defined as follows:
[0042] in, It is the k-th type of medical logical association matrix. These are the importance weights for the corresponding dimensions, and K is the total number of categories in the association matrix. The logical consistency assessment method is as follows: if the calculated score... If the data falls below the threshold, a logical conflict warning is triggered, prompting manual review to improve data credibility.
[0043] Finally, the system establishes a security audit mechanism in the local environment. For each processed medical text, the system generates a block containing the original text hash, processing operator characteristics, operator signature, and timestamp, ensuring that the hard requirements for medical data traceability and compliance supervision are met locally, making the data controllable and traceable.
[0044] The audit mechanism consists of three steps. First, a unique hash value is generated for the original medical record text to ensure that the original cannot be replaced. Second, the system generates an operation log, recording the operator's account, operation time, and processing method. Finally, the operation log is stored in a chain structure, with each new block containing the hash value of the previous block to ensure that the data cannot be deleted from a single point of failure. When regulatory verification is required, the system can first determine the compliance of the medical records from raw data to structured data, compare the hash values to confirm the reliability of the data, and trace the specific responsible party if problems occur.
[0045] Example 2: Please see Figures 1-3 This embodiment provides a method for intelligent structural processing of medical text based on a local large model, specifically including the following steps: Step 1: A single NVIDIA RTX 4090 graphics card (24GB VRAM) was configured on the local server. The Qwen2-7B-Instruct model, quantized to 4-bit, was used as the base model. Adaptive fine-tuning was performed on a dataset of 50,000 anonymized Chinese electronic medical records using LoRA technology to ensure a balance between model accuracy and size. The LoRA configuration parameters were: rank r=16, alpha=128, dropout=0.05, and learning rate 2×10⁻⁶. -4 Batch size 4, training epochs 20, early stopping mechanism set. Fine-tuning target is medical entity recognition and attribute extraction task, loss function is:
[0046] in, Cross-entropy loss for entity recognition, For existence classification loss, For the regression loss of the degree value, this embodiment sets λ1=0.5 and λ2=0.3 as balancing factors.
[0047] Step 2: Based on expert consensus, clinical practice guidelines, and the National Essential Medicines List, a correlation matrix covering symptoms, diagnosis, and contraindications was constructed. When the system receives free clinical text, the following structured workflow is executed: An example is: "The patient was found to have mild wheezing during the examination, but no pleural effusion." 1) Input text segmentation and encoding: Generate context-dependent hidden layer representation H∈RL×d using the fine-tuned Qwen2-7B, where L is the sequence length and d=4096 is the hidden layer dimension.
[0048] 2) Identify medical entities and types: Secondly, based on CRF layer decoding, wheezing → Symptom, location index [7,9]; pleural effusion → Symptom, location index [13,16].
[0049] 3) Existence status determination and positive screening: For "wheezing": check its preceding modifier "present", there is no negative trigger word, determine S(wheezing)=1 (existence), and enter the structured process; For "pleural effusion": check the negative trigger word "absent", calculate the negative scope Scope(absent, pleural effusion)=0.89>0.5, determine S(pleural effusion)=0 (non-existent), the entity is filtered out and not output.
[0050] 4) Degree modification identification for positive entities: The modifier "slight" precedes "wheezing". Query the degree mapping table, simbert(slight, slight) = 1.0, and the judgment D(wheezing) = 0.2.
[0051] 5. Standardize the mapping of positive entities: "Wheezing sound" is mapped to SNOMED CT code "886096002|Wheezing|", with a cosine similarity C=0.94.
[0052] 6. Structured Output: {Structured Result: [{Diagnosis: Wheezing, Standard Code: SNOMED CT: 886096002, Severity: Mild, Severity Value: 0.2, Confidence: 0.94}], Filtered Entity: [{Entity: Pleural Effusion, Filtering Reason: Negative State Recognition, Confidence: 0.91}], Original Text: The patient was found to have mild wheezing during examination, but no pleural effusion.} Step 3: Based on authoritative treatment guidelines, the system pre-constructs a medical logical association matrix covering diseases and medication contraindications. After entity extraction, assuming the system identifies the core entity pair "Diagnosis: Peptic ulcer" and "Medication: Enteric-coated aspirin," it then calculates the matrix outer product using a large model encoding method, and compares it with the k-th type of medical logical association matrix. Perform inner product operation and output. Because aspirin's pharmacological irritation of the gastric mucosa is logically incompatible with ulcer diagnosis, the output compliance score will fall below the safety baseline. The system then intercepts the structured archiving process of this medical record and triggers a tamper-proof level logical conflict warning for clinicians to review.
[0053] Step 4: After the above noise reduction, mapping, and self-consistency verification are all successful, the system automatically generates an irreversible operation block at the local underlying layer. This block tightly encapsulates the ultrafast hash value of the original medical record text, the characteristic signature of the processing operator, and the operation timestamp accurate to milliseconds. This step not only outputs high-quality structured medical data to the hospital's central database, but also solidifies a highly tamper-proof secure traceability chain at extremely low hardware costs, ensuring that the entire data processing flow fully complies with the stringent requirements of medical data compliance and regulation.
[0054] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for intelligent structural processing of medical text based on a local large model, characterized in that, The method includes the following steps: S1: In a physically isolated medical intranet environment, the size of the open-source medical model is compressed by performing weight quantization on the large model. S2: A semantic parsing and positive screening mechanism based on medical context enhancement is used to identify positive medical entities; the semantic parsing and positive screening mechanism based on medical context enhancement uses the hidden layer representation of a large model to calculate the semantic information entropy of each word, identify key medical entities and their attribute states in the text, and automatically filter negative descriptions through existence determination, retaining only positive signs. The semantic parsing and positive screening mechanism based on medical context enhancement specifically includes: defining medical entities. Existence state These correspond to the three cases of existence, non-existence, and not mentioned, respectively; the degree of definition is modified. The severity of the description is quantified; an attention mechanism for medical context awareness is constructed to calculate the term level. Importance weight in the medical context for: in, ∈[0,1] is the balance factor, which is determined by maximizing the accuracy of positive feature extraction on the validation set; TF-IDF refers to term frequency-inverse document frequency, a statistical method used to evaluate the importance of a word in a text or corpus; Indicates terms In medical corpus The semantic information entropy in the context is calculated as follows: in, For medical context categories; Terms output for large models Belongs to context The posterior probability; This represents the total number of medical context categories. A negative trigger word detection and positive screening mechanism is adopted; a set N of negative trigger words is defined, and a negative scope resolution function is constructed: in, Negation trigger word For target terms The probability of determining the scope. For the detected first k A negative trigger word, and These are the hidden representations of the negation word and the target word, respectively. Encode the relative positions of the two words in the sentence. For learnable parameter vectors, For the Sigmoid function; when When the value is greater than the threshold, the term is determined. Its existence state within the negation scope This entity will be filtered. For degree modifier recognition, a degree adverb hierarchical mapping table is constructed, and the degree value is determined through fuzzy matching and semantic similarity calculation. in, Adjectives that modify the target entity For semantic similarity based on BioBERT encoding, As a preset level value, For the j-th degree adverb entry in the degree adverb hierarchy mapping table, This is a pre-defined hierarchical mapping table for medical degree adverbs; S3: Convert the positive medical entities identified in step S2 into standardized structural output; Define the structured output format as quadruples: Through standardized mapping of medical entities, the extracted entities are linked to a standard terminology database, and the mapping score is calculated. Calculated using cosine similarity: in, To extract the semantic vector of an entity, Vector representation of standard terminology; After converting positive medical entities into standardized structural outputs, a logical self-consistency evaluation function is used to quantify the internal rationality of medical records, including a logical conflict risk score. The calculation process is defined as follows: in, It is the k-th type of medical logical association matrix. The total number of categories in the medical logical association matrix. These are the importance weights for the corresponding dimensions; if the calculated score... If the value is below the threshold, a logical conflict warning will be triggered, prompting manual review. After the structure output is completed, a localized security audit mechanism is also included. Specifically, for each processed medical text, a block containing the original text hash, processing operator characteristics, operator signature and timestamp is generated to ensure that the hard requirements of medical data traceability and compliance supervision are met locally, making the data controllable and traceable.
2. The medical text structuring intelligent processing method according to claim 1, characterized in that, In step S1, low-rank adaptive fine-tuning technology is used to compress the model size so that it can be adapted to the existing general-purpose servers or edge computing devices of medical institutions.
3. A system for implementing the intelligent processing method for structuring medical text as described in claim 1 or 2, characterized in that, The system includes: The lightweight model adaptation module for intranet environments compresses the size of large open-source medical models by performing weight quantization on them in physically isolated medical intranet environments. The medical semantic parsing and positive sign recognition module uses a medical context-enhanced semantic parsing and positive screening mechanism to identify positive medical entities. The positive sign structured representation module converts the identified positive medical entities into standardized structural outputs. The dynamic logic verification module uses a logic self-consistency evaluation function to quantify the inherent rationality of medical records; The localized security audit module for medical compliance adopts a localized security audit mechanism to ensure that the hard requirements of medical data traceability and compliance supervision are met locally, making the data controllable and traceable.
4. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the intelligent processing method for structuring medical text as described in claim 1 or 2.
Citation Information
Patent Citations
File anti-desensitization self-learning recognition system and method based on information entropy
CN120580704A
Anxiety state real-time evaluation method based on multi-modal fusion
CN120973949A
Non-intrusive electronic medical record structured generation, traceability verification and automatic filling method and system based on localized large model
CN121812039A