Unstructured text entity recognition and relation extraction method under small sample learning
By preprocessing unstructured text and building a model using a meta-learning algorithm, the accuracy problem of entity and relationship extraction in small sample scenarios is solved, and rapid adaptation and efficient recognition are achieved in fields such as medicine and law.
Patent Information
- Application Number
- CN202510928021.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies have difficulty accurately identifying entities and relationships in unstructured text in small sample scenarios, especially in fields such as medicine and law when labeled data is scarce. The model has poor generalization ability and the accuracy and recall rate of entity and relationship extraction are low.
By preprocessing unstructured text, using pre-trained language models to obtain contextual semantic features, combining meta-learning algorithms to build models, extracting entities and relationships, and optimizing the model through data enhancement and regularization, it is suitable for small sample scenarios.
It achieves rapid knowledge transfer in small sample scenarios, adapts to new domain texts, reduces false positives and missed positives, improves the accuracy of entity boundary recognition and semantic relationship analysis, and maintains long-term stable high-quality services.
Smart Images

Figure CN120633635A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and specifically to a method for unstructured text entity recognition and relationship extraction under small sample learning. Background Art
[0002] In the field of natural language processing, entity and relation extraction is a key technology for extracting valuable information from text; however, current entity and relation extraction methods face many challenges when faced with small sample scenarios; traditional methods usually require a large amount of labeled data for model training. When labeled data is scarce, the model's generalization ability is poor, making it difficult to accurately identify entities and relations in new fields or rare types.
[0003] After searching, the Chinese patent number CN118332106A discloses a Chinese entity relationship extraction method based on additional relational information. It has the characteristics of incorporating relational information into triple extraction through potential relation prediction by using Chinese BERT, and labeling entities through potential relation prediction. In practical applications, many fields (such as medical, legal, and financial) often face the problem of insufficient labeled data. In addition, the diverse forms and complex structures of unstructured texts further increase the difficulty of entity and relationship extraction. Some existing small-sample learning methods are not comprehensive enough in extracting text features when processing unstructured texts, and lack the comprehensive utilization of semantic and structural information in the text, resulting in low accuracy and recall of entity and relationship extraction. Therefore, it is urgent to propose an unstructured text entity and relationship extraction method suitable for small-sample learning to improve the entity and relationship extraction performance in small-sample scenarios. Summary of the Invention
[0004] To achieve the above objectives, the present invention is implemented through the following technical solutions:
[0005] Step 1: Preprocess the unstructured text, including word segmentation, part-of-speech tagging, and syntactic analysis, to convert the text into a structured word vector representation; use a pre-trained language model to encode the text and obtain the contextual semantic features of the text;
[0006] Step 2: Based on the contextual semantic features of the text, a meta-learning model is constructed using a meta-learning algorithm. This model is trained with a small number of labeled samples and a large number of unlabeled samples to learn the common knowledge between different tasks.
[0007] Step 3: Based on the meta-learning model, the word vectors in the text are processed to identify entities, including boundaries and types. Based on the identified entities, the semantic relationships between entities are analyzed, the relationship types between entities are extracted, and the dependency relationships between entities and relationships are optimized using dynamic programming methods.
[0008] Step 4: During the model training process, add regularization terms, perform data augmentation on a small number of labeled samples based on the optimized dependencies between entities and relationships, evaluate the model using the validation set, and adjust the model parameters and structure based on the evaluation results to optimize the model.
[0009] Furthermore, the process of preprocessing the unstructured text is as follows:
[0010] Word segmentation: Use domain adaptation tools to segment text according to semantic boundaries and load custom dictionaries;
[0011] Part-of-speech tagging: Use statistical learning models to assign part-of-speech tags to words;
[0012] Syntactic analysis: Use dependency parser to build grammatical dependencies and extract key syntactic structures;
[0013] Generate structured word vectors: Use a pre-trained model to map words into dense vectors, and fine-tune the model for domain words to make the vector representation fit the domain semantic distribution.
[0014] Furthermore, the process of encoding the text is as follows:
[0015] First, the position encoding algorithm is introduced to give words position features, allowing the model to perceive the temporal structure of the text; then the pre-trained language model is used for encoding, and long-distance dependencies are captured through self-attention. The input is input according to the model specifications, and the hidden layer representation is output through a multi-layer encoder; then the outputs of different levels are integrated, the bottom layer captures grammatical features, and the upper layer obtains semantic features, which are fused through weighting or splicing; finally, for small sample scenarios, a small amount of domain labeled data is used to fine-tune the model; multiple features are spliced, the sequence is standardized, and the feature vector is processed before output.
[0016] Furthermore, the process of building the meta-learning model is as follows:
[0017] Split meta-tasks, abstract entity and relationship extraction tasks into multiple meta-tasks, sample across domains, and capture task commonalities; using the improved prototype network as the framework, map text features to the semantic space through the feature encoder, prototype generator, and distance measurement module.
[0018] Furthermore, the process of learning common knowledge between different tasks is:
[0019] By clustering cross-task semantic prototypes, we cluster meta-task prototype vectors to find common semantic patterns across domains. By summarizing meta-rules, we incorporate the encoding into regularization terms into training to guide the model to learn interpretable extraction patterns.
[0020] Furthermore, the process of identifying the entity is:
[0021] The word vector is input into the meta-learning model encoder, attention is used to enhance features, and the entity location is found in combination with syntax; sliding window detection is used, BIOES is used to mark boundaries, and multiple classifiers are integrated to predict types; finally, CRF is used to constrain the label sequence and facilitate cross-task knowledge transfer.
[0022] Furthermore, the process of extracting the relationship type between entities is as follows:
[0023] The identified entities are grouped into candidate pairs and context features are extracted. The features are input into the meta-learning model relationship classifier, and the relationship type is predicted based on the similarity with the relationship prototype, and multi-label classification is used to process the relationship situation. The domain relationship graph is used to filter the prediction and optimize the extraction by focusing on key related words through self-attention.
[0024] Furthermore, the process of optimizing the dependency between entities and relationships is as follows:
[0025] Define entity and relationship compatibility rules and utilize transitive constraints; use dynamic programming to calculate transition probabilities based on entity positions and use the Viterbi algorithm to find the optimal path; iteratively optimize parameters, verify consistency based on domain knowledge, and correct contradictory predictions.
[0026] Furthermore, the process of performing data enhancement on a small number of labeled samples is as follows:
[0027] Rule-based text transformation uses domain dictionaries to replace synonyms and restructure sentences; based on the generative model, samples are generated based on entity and relationship types, and MLM masks fill in the text; through adversarial training and adding perturbations, the model learns the essential semantic features.
[0028] Furthermore, the process of optimizing the model is as follows:
[0029] Balance complexity and expressiveness by adjusting the number of neural network layers; optimize feature interaction by improving the attention mechanism; integrate multiple models to reduce variance; incrementally train new data and compress the model through knowledge distillation.
[0030] The unstructured text entity recognition and relationship extraction method based on small sample learning provided by the present invention has the following beneficial effects:
[0031] (1) The present invention constructs meta-tasks through cross-domain task sampling, and the model learns the common features of entity and relationship extraction in different fields; when faced with small sample new tasks or new domain texts, it can quickly transfer knowledge without the need for retraining with a large amount of labeled data; for example, when migrating from medical text extraction to legal text, it can quickly adapt, avoid duplication of work, save resources, and effectively solve the difficulties of traditional methods in cross-domain applications, and easily cope with diverse text scenarios.
[0032] (2) The present invention comprehensively utilizes a variety of technical means, such as combining syntactic structure and attention mechanism to enhance entity feature recognition, and utilizing relational semantic graph and dynamic programming to optimize relation extraction; it can accurately locate entity boundaries, determine types, and accurately analyze semantic relationships between entities, thus significantly reducing the rates of misjudgment and missed judgments, and the extraction results are highly consistent with the true semantics of the text and the logic of domain knowledge.
[0033] (3) The present invention uses data enhancement strategies to expand labeled samples, combines regularization to prevent overfitting, and dynamically adjusts model parameters and structures based on evaluation. As new labeled data increases, knowledge can be updated through incremental training. The model can also be optimized through knowledge distillation. It always evolves in sync with actual application needs, continuously improves extraction capabilities and efficiency, and ensures long-term, stable, and high-quality services. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION
[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0036] Example
[0037] See also Figure 1 , an embodiment of the present application provides a method for unstructured text entity recognition and relationship extraction under small sample learning, the method comprising:
[0038] Step 1: Preprocess the unstructured text, including word segmentation, part-of-speech tagging, and syntactic analysis, to convert the text into a structured word vector representation; use a pre-trained language model to encode the text and obtain the contextual semantic features of the text;
[0039] Preprocessing of unstructured text:
[0040] Word segmentation:
[0041] To segment semantic units, use domain-adapted word segmentation tools, such as the PubMedBERT word segmenter in the medical field and THULAC in the legal field, to split continuous text into word sequences according to semantic boundaries.
[0042] For example, “The patient visited the doctor due to headache on July 1, 2025” is split into “The patient / on / July 1, 2025 / due to / headache / visited the doctor.”
[0043] For special field terms, such as the medical term "coronary atherosclerosis", custom dictionaries are loaded to ensure word segmentation accuracy and avoid segmentation errors, such as misclassifying "echocardiogram" as "ultrasound / cardiogram".
[0044] Part-of-speech tagging:
[0045] By semantically anchoring grammatical functions, statistical learning-based part-of-speech tagging models, such as HMM and CRF, are used to assign a part-of-speech tag to each word, clarifying its grammatical role. For example, "headache" (noun, n), "see a doctor" (verb, v), and "July 1, 2025" (time word, t).
[0046] The annotation model is trained in conjunction with the domain corpus to improve the annotation accuracy of professional vocabulary, such as annotating "statute of limitations" in legal texts as a noun phrase.
[0047] Syntactic analysis:
[0048] Hierarchical parsing of text structure, through dependency syntax analyzers, such as spaCy's SyntaxSuggester, to build grammatical dependencies between words and generate a dependency tree.
[0049] For example, in “the doctor prescribes a prescription”, “doctor” (subject, nsubj) → “prescribe” (predicate, root) → “prescription” (object, dobj).
[0050] Extract key syntactic structures, such as subject, predicate, object, attributive, adverbial, and complement, to provide structural clues for subsequent entity recognition. For example, entity objects often appear in the object position.
[0051] Structured word vector representation:
[0052] Static word vectors are generated by mapping the numerical value of text. Using a pre-trained general word vector model, words are mapped into dense vectors, typically with dimensions of 100-300. For example, the word vectors for "headache" and "dizziness" are close in semantic space, reflecting semantic relevance.
[0053] For domain-specific vocabulary (such as "atrial fibrillation" and "blockchain"), the word vector model is fine-tuned through the domain corpus to make the vector representation of professional vocabulary closer to the domain semantic distribution.
[0054] Encode the text:
[0055] By explicitly embedding temporal information and introducing a positional encoding algorithm, we leverage the Transformer's sine-cosine encoding to add positional features to each word, encoding its order within the sentence. The position of the first word in a sentence is encoded as [0.1, 0.2, ...], and the position of the second word as [0.3, 0.4, ...], ensuring that the model can perceive the temporal structure of the text.
[0056] Deep extraction of contextual semantics:
[0057] By constructing dynamic semantic representations, a pre-trained language model is used to encode text. The model uses a self-attention mechanism to capture long-range dependencies between words. For example, in the sentence "patients experience adverse reactions after taking medication," the model can associate the causal relationship between "drug" and "adverse reactions," generating dynamic word vectors that incorporate contextual semantics.
[0058] Its input format follows the model specification and outputs the hidden layer representation of each word through a multi-layer Transformer encoder, such as BERT's 12-layer output, with a dimension of 768 per layer and multi-level feature fusion.
[0059] Multi-level feature fusion:
[0060] Integrate the outputs of different levels of the pre-trained model. The bottom hidden layers, such as layers 1-4, capture grammatical features such as parts of speech and syntactic structure; the high-level hidden layers, such as layers 8-12, capture semantic features such as entity relationships and sentiment tendencies.
[0061] A weighted fusion strategy is adopted to assign weights to each layer according to task requirements; or a splicing strategy is used to splice multi-layer outputs into longer vectors to generate contextual features containing multi-dimensional semantic information, and perform domain adaptation fine-tuning.
[0062] Domain adaptation fine-tuning:
[0063] For small sample learning scenarios, a small amount of domain-labeled data is used to fine-tune the pre-trained model so that the model parameters adapt to the semantic distribution of the specific domain.
[0064] For example, in medical texts, the fine-tuned model can more accurately identify the semantic features of professional entities such as "pulmonary embolism" and "myocardial infarction".
[0065] Feature integration and output:
[0066] By generating a unified structured representation, static word vectors, positional encodings, and contextual features output by the pre-trained model are spliced by dimension to form a unified word-level representation vector.
[0067] For example, the final representation of each word is [static word vector (300 dimensions) + position encoding (100 dimensions) + BERT high-level features (768 dimensions)], with a total dimension of 1168 dimensions.
[0068] By normalizing sequence features, long texts are truncated or padded, for example, to a fixed sequence length of 512, ensuring consistent lengths for the text sequences fed into the model and facilitating batch computation. Layer-wise normalization or standardization of feature vectors prevents the impact of numerical discrepancies on subsequent model training. This ensures that the final output, structured word vector representation, incorporates both the basic semantics and positional information of words and the dynamic semantic features of the context, providing high-quality input features for the entity and relationship extraction tasks of subsequent meta-learning models.
[0069] Step 2: Based on the contextual semantic features of the text, a meta-learning model is constructed using a meta-learning algorithm. This model is trained with a small number of labeled samples and a large number of unlabeled samples to learn the common knowledge between different tasks.
[0070] Building a meta-learning model:
[0071] Split the meta-task and abstract the entity and relationship extraction tasks into multiple meta-tasks, each of which contains several support sets and query sets.
[0072] For example, in the medical field, the meta-task can be defined as "extracting the 'disease-symptom' relationship from pneumonia-related texts", the support set contains 5 labeled samples, and the query set contains 10 unlabeled samples.
[0073] By sampling cross-domain tasks, we can capture commonality among tasks. Meta-tasks cover entity and relationship types in different fields, such as healthcare, law, and finance.
[0074] For example, tasks such as "drug-side effects" (medical) and "contract-breach of contract clauses" (legal) are included at the same time, forcing the model to learn cross-domain extraction patterns.
[0075] The meta-learning architecture uses an improved prototype network as the basic framework. Through a feature encoder, text features composed of a pre-trained language model and a multi-layer neural network are mapped into a semantic embedding space. A prototype generator is used to calculate prototype vectors for each category, using entity or relationship types, as a "template" for the category semantics. Based on a distance metric module, the semantic distance between the query sample and the prototype is calculated using Euclidean distance or cosine similarity to achieve few-shot classification.
[0076] Training with a small number of labeled samples and a large number of unlabeled samples:
[0077] For each support set sample in each meta-task, the feature encoder is used to process it to obtain the semantic embedding vector xi of each sample. According to the category c to which the sample belongs, all semantic embedding vectors under the same category c are aggregated to calculate the prototype vector pc of the category. The semantic embedding vectors xi of all samples in the support set Sc belonging to category c are processed by the feature encoder function f(xi) and then divided by the number of samples in category c, Nc.
[0078] For example, in the "symptom entity recognition" task, the support set contains labeled samples such as "headache" and "fever", and the prototype vector psymptom is the mean of these sample features, representing the semantic prototype of the "symptom" category;
[0079] Query set prediction and loss calculation:
[0080] For a sample xq in the query set, we first calculate the distance d(xq,pc) between it and each prototype vector. These distances are then converted into probabilities of belonging to different categories using the softmax function. This probability indicates the likelihood that a given sample xq belongs to category c. To optimize the model parameters, we use the cross-entropy loss function. When calculating the loss, we first traverse all samples in the query set. For each sample, we calculate the difference between the predicted and true results based on its true category label yq,c and the model's predicted category probability P(c|xq). Finally, we sum the differences across all samples in the query set to obtain the loss value for the entire query set. Q represents the query set, and yq,c is the true category label for sample xq.
[0081] Meta-optimization strategies:
[0082] Using MAML’s gradient update method, the model parameters θ are iteratively adjusted through meta-training: For each meta-task, the loss L is calculated using the support set i (θ), and calculate the gradient ▽ θ L i (θ), we get the task-specific parameters, which are:
[0083] θ′=θ-α▽ θ L i (θ)
[0084] Where α is the learning rate within the task. Use θ′ to calculate the meta-loss L on the query set. meta (θ′), back-propagation updates the global parameters θ, so that the model can quickly adapt to new tasks and perform focused learning of key features.
[0085] Focused learning of key features:
[0086] By enhancing the attention mechanism and adding multiple layers of self-attention modules to the feature encoder, we can capture semantic dependencies within the text. For example, in the sentence "The patient developed gastrointestinal bleeding after taking aspirin," the self-attention mechanism can strengthen the semantic connection between "aspirin" and "gastrointestinal bleeding" while reducing the interference caused by irrelevant words such as "patient" and "take."
[0087] Task-related attention dynamically assigns feature weights across different meta-tasks and designs a task embedding vector, such as one-hot encoding to represent the task type. The attention mechanism calculates the correlation between features and tasks, thereby determining the weight of each feature for a specific task.
[0088] Unlabeled data:
[0089] Using a semi-supervised meta-learning strategy, pseudo-label generation and screening, and the current meta-learning model, pseudo-labels are generated for large amounts of unlabeled data. Reliable samples are then screened based on confidence. For example, results with a predicted probability greater than 0.9 are considered highly reliable and added to the support set. For example, in legal text processing, the model generates pseudo-labels for the "party-rights" relationship for unlabeled contract clauses, then selects high-confidence samples for meta-task training.
[0090] Perform perturbations on unlabeled data, such as replacing words or adjusting word order, and require the model to output the same prediction results when processing samples before and after the perturbation, which can improve the generalization ability of the model.
[0091] Explicit extraction of common knowledge:
[0092] Through cross-task semantic prototype clustering, the prototype vectors of all meta-tasks are clustered to discover common semantic patterns across domains.
[0093] For example, the "drug" entity prototype in the medical field and the "securities" entity prototype in the financial field can be clustered into the "entity-object" class, extracting the common knowledge that "noun entities usually serve as relationship subjects."
[0094] Step 3: Based on the meta-learning model, the word vectors in the text are processed to identify entities, including boundaries and types. Based on the identified entities, the semantic relationships between entities are analyzed, the relationship types between entities are extracted, and the dependency relationships between entities and relationships are optimized using dynamic programming methods.
[0095] Entity Recognition:
[0096] Perform feature enhancement and representation. Input the word vectors obtained in Step 1 into the encoder of the meta-learning model, and strengthen entity-related features through the attention mechanism. For example, in medical texts, focus on the context information of keywords such as "patient", "drug", "symptom", etc., and suppress the interference of redundant words (such as "de", "le").
[0097] Combine syntactic structure information (such as the subject and object positions in the dependency tree), and give priority to the grammatical positions where entities may appear.
[0098] Boundary recognition and classification:
[0099] Among them, for sliding window detection, traverse the text using a sliding window, and calculate the probability that the text within the window belongs to an entity through the meta-learning model.
[0100] For example, set the window size to 5 - 7 words. When "acute myocardial infarction" is detected, the model determines that this window is a disease entity.
[0101] Use the BIOES annotation system to perform boundary annotation on the identified entities, distinguishing the start (B), inside (I), end (E), single-word entity (S) and non-entity (O) of the entity. For example, "The patient sought medical treatment due to hypertension" is annotated as "patient (S-person) due to (O) hypertension (S-disease) sought medical treatment (O)".
[0102] [[ID=**18**]]Integrate through multiple classifiers. Use the classifier of the meta-learning model to predict the entity type, and combine domain knowledge constraints. For example, medical entity types are limited to "disease", "drug", "symptom", etc., to improve the classification accuracy.
[0103] Entity recognition optimization:
[0104] Through CRF sequence constraints, introduce a conditional random field to constrain the entity label sequence, and consider the transition probability between labels. For example, after B-disease, only I-disease or E-disease can follow, to eliminate unreasonable boundary predictions.
[0105] Through cross-task knowledge transfer, utilize the cross-domain common knowledge learned by the meta-learning model in Step 2 to enhance the recognition ability of rare entities. For example, transfer the recognition pattern of "stock code" in the financial field to the recognition of "gene coding" in the medical field;
[0106] Extract the relationship types between entities:
[0107] Through the modeling and classification of semantic associations, perform entity pair construction and feature extraction, and combine the identified entities in pairs to form candidate relationship pairs. For example, in the text "Aspirin treats headache", generate the candidate pair (Aspirin, headache).
[0108] Note: There seems to be a small error in the original text. In the description of "through multiple classifiers for integration" in , it should be "through multiple classifiers for integration" instead of "through multiple classifiers for integrated", and this has been corrected in the translation.Contextual features are used to extract features such as text fragments between entity pairs (such as "treatment"), the relative position of entities in sentences, and syntactic paths, such as the shortest path connecting two entities in the dependency tree.
[0109] Through relationship classification and type determination, the entity pair features are input into the meta-learning model's relationship classifier. The prototype network calculates the similarity with various relationship prototypes to predict the relationship type. For example, (aspirin, headache) has the highest similarity with the "drug-treatment-symptom" prototype and is therefore determined to be a treatment relationship.
[0110] It supports the existence of multiple relationship types for entity pairs (such as "causality" and "treatment"), outputs probability distribution through a multi-label classifier, and takes the categories above the threshold as valid relationships.
[0111] Relationship extraction optimization:
[0112] By constraining the semantic graph, we can construct a domain-specific relationship graph, such as the "drug-target-disease" network in the medical field. We can then use the prior knowledge in the graph to filter out unreasonable relationship predictions. For example, if the graph shows a therapeutic relationship between "aspirin" and "headache," the credibility of that prediction will be enhanced.
[0113] By focusing on key related words between entity pairs, such as "lead to" and "affect", the interference of irrelevant context is suppressed.
[0114] Joint reasoning of entities and relations:
[0115] Define compatibility rules between entity types and relationship types. For example, the "drug" entity can only be associated with relationships such as "treatment" and "side effect", and the "symptom" entity cannot be the initiator of a relationship.
[0116] Transitivity constraint: Utilize the transitivity of relationships. For example, if A is a subclass of B and B has a relationship with C, then A and C may have the same relationship, thus optimizing the inference results.
[0117] Through dynamic programming, each entity position in the text is treated as a state, with the state value representing the possible combinations of entity types and relationship types at that position. Based on entity-relationship constraints and contextual information, the state transition probability and transition function are calculated. For example, if the current entity is "disease" and the next entity is "drug," the probability of transitioning to the "treatment" relationship increases. The optimal path search uses the Viterbi algorithm to find the globally optimal entity-relationship combination path and maximize the joint probability.
[0118] Through iterative optimization and consistency checking, we first perform preliminary entity and relationship extraction, then adjust the dynamic programming parameters based on the extraction results, and iteratively optimize. For example, if the first extraction finds that the relationship between "patient" and "symptom" appears frequently, we adjust the state transition weights to strengthen this relationship.
[0119] Through consistency checking, we check whether the extraction results meet the domain knowledge constraints, such as time sequence and logical consistency, and correct contradictory predictions. For example, if "disease appears after treatment" is extracted, but the disease occurred before treatment, it will be corrected to "misdiagnosis";
[0120] Step 4: During model training, add regularization terms. Based on the optimized dependencies between entities and relationships, perform data augmentation on a small number of labeled samples. Use the validation set to evaluate the model. Based on the evaluation results, adjust the model parameters and structure to optimize the model.
[0121] Regularization processing:
[0122] By suppressing the generalization enhancement of overfitting, L1 / L2 regularization constraints are performed, and L1 or L2 regularization terms are added to the model loss function to constrain the complexity of the model parameters.
[0123] For example, L2 regularization penalizes large weight parameters, allowing the model to learn smoother decision boundaries and reduce sensitivity to training data noise.
[0124] Applied to the encoder and classifier parameters of the meta-learning model to prevent the model from overfitting to specific feature patterns on small samples.
[0125] During model training, we randomly "drop" some neuron outputs to force the model to learn more robust feature representations. For example, setting a 50% dropout rate in the fully connected layers of a meta-learning model randomly ignores the contributions of half of the neurons at each iteration. Combined with batch normalization, this further stabilizes the training process and accelerates convergence.
[0126] Terminate training early when validation set performance stops improving to avoid model overfitting. For example, if the validation set F1 value shows no significant improvement for 10 consecutive training epochs, stop training and save the current optimal model parameters.
[0127] Data enhancement: Through the value expansion of limited samples, rule-based text transformation, and synonym replacement using domain dictionaries for entity and relationship descriptions.
[0128] For example, replace “patient” with “disease” and “treatment” with “heal” to generate new samples with equivalent semantics but different expressions.
[0129] Change the text structure by converting active sentences into passive sentences, splitting long sentences, etc. For example, reorganize "The doctor prescribed aspirin" into "Aspirin was prescribed by the doctor" while keeping the core entities and relationships unchanged.
[0130] Sample expansion based on generative model:
[0131] Through conditional generative adversarial networks, realistic text samples are generated based on entity types and relation types.
[0132] For example, input the condition "drug-treatment-disease" to generate new text of the type "ibuprofen relieves arthritis pain".
[0133] Based on the masked language model, the original text is randomly token-masked, and the padded new text is generated through the pre-trained language model.
[0134] For example, complete "aspirin treatment [MASK]" to "aspirin treatment for headaches."
[0135] Robustness is enhanced through adversarial training, which adds carefully designed perturbations to the input text, such as small shifts in the word vector space, and requires the model to maintain stable predictions for the perturbed samples.
[0136] For example, adding Gaussian noise to word vectors forces the model to learn more essential semantic features.
[0137] Model evaluation and validation:
[0138] Through multi-dimensional performance measurement and evaluation indicator selection, we calculate precision, recall, and F1 value to evaluate the accuracy of entity boundary and type prediction. We calculate the above indicators for each relationship type separately, paying attention to macro-average and micro-average, and balancing the evaluation weights of rare and common relationships.
[0139] Validation set design:
[0140] Stratified sampling is used to ensure that the validation set contains representative samples of various entities and relationships. For example, in the medical field, the proportion of relationship types such as "disease-symptom" and "drug-side effect" in the validation set is consistent with the real data distribution. By designing cross-domain validation sets, the model's ability to generalize to unseen domains is tested. For example, a legal text validation set is used to evaluate a model trained in the medical field to test the cross-domain transfer effect of meta-learning.
[0141] Construct a confusion matrix to analyze misclassification patterns, such as which entity types are easily confused, such as "disease" and "symptoms." Visualize the model's attention distribution to check whether the text areas the model focuses on are consistent with human intuition. For example, in relation extraction, check whether the model correctly focuses on relation indicators such as "cause" and "treat."
[0142] Perform grid search or random search on key hyperparameters such as learning rate, batch size, and dropout rate. For example, try different values of learning rate in the range of 0.001-0.01 and select the configuration that performs best on the validation set.
[0143] Model structure optimization:
[0144] By adjusting the number of layers, increasing or decreasing the number of neural network layers in the meta-learning model, we can balance model complexity and expressiveness. For example, in the medical field, increasing the number of Transformer layers in the encoder may enhance understanding of professional terminology.
[0145] Improve the attention mechanism, adjust the number of self-attention and cross-attention heads, and optimize feature interaction. For example, in relation extraction, increase the number of cross-entity attention heads to strengthen the capture of inter-entity relationship features. Train multiple models with different initializations or hyperparameter configurations and combine the prediction results through voting or weighted averaging. For example, ensemble five models with similar performance to reduce the prediction variance of a single model.
[0146] When new labeled data is acquired, the model is updated using incremental training. For example, every time 100 labeled medical samples are added, the model parameters are fine-tuned while retaining the original knowledge.
[0147] Distill the knowledge of a complex model (the teacher model) into a lightweight model (the student model) to improve inference efficiency. For example, knowledge distillation can be used to compress the BERT-base model into DistilBERT, maintaining similar performance while reducing computing resource consumption.
[0148] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution.
[0149] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment as needed.
[0150] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. Unstructured text entity recognition and relationship extraction method under small sample learning, characterized by: The method includes: Step 1: Preprocess the unstructured text, including word segmentation, part-of-speech tagging, and syntactic analysis, to convert the text into a structured word vector representation; use a pre-trained language model to encode the text and obtain the contextual semantic features of the text; Step 2: Based on the contextual semantic features of the text, a meta-learning model is constructed using a meta-learning algorithm. This model is trained with a small number of labeled samples and a large number of unlabeled samples to learn the common knowledge between different tasks. Step 3: Based on the meta-learning model, the word vectors in the text are processed to identify entities, including boundaries and types. Based on the identified entities, the semantic relationships between entities are analyzed, the relationship types between entities are extracted, and the dependency relationships between entities and relationships are optimized using dynamic programming methods. Step 4: During the model training process, add regularization terms, perform data augmentation on a small number of labeled samples based on the optimized dependencies between entities and relationships, evaluate the model using the validation set, and adjust the model parameters and structure based on the evaluation results to optimize the model.
2. The unstructured text entity recognition and relationship extraction method under small sample learning according to claim 1 is characterized in that The process of preprocessing unstructured text is as follows: Word segmentation: Use domain adaptation tools to segment text according to semantic boundaries and load custom dictionaries; Part-of-speech tagging: Use statistical learning models to assign part-of-speech tags to words; Syntactic analysis: Use dependency parser to build grammatical dependencies and extract key syntactic structures; Generate structured word vectors: Use a pre-trained model to map words into dense vectors, and fine-tune the model for domain words to make the vector representation fit the domain semantic distribution.
3. The unstructured text entity recognition and relationship extraction method under small sample learning according to claim 2 is characterized in that: The process of encoding the text is as follows: First, a positional encoding algorithm is introduced to assign positional features to words, allowing the model to perceive the temporal structure of the text. Then, a pre-trained language model is used for encoding, and self-attention is used to capture long-range dependencies. The input is input according to the model specifications, and a multi-layer encoder outputs the hidden layer representation. Then, the outputs of different levels are integrated, with the bottom layer capturing grammatical features and the upper layer acquiring semantic features, which are then fused through weighting or splicing. Finally, for small sample scenarios, the model is fine-tuned using a small amount of domain-labeled data. Multiple features are spliced together, the sequence is standardized, and the feature vector is processed before output.
4. The unstructured text entity recognition and relationship extraction method under small sample learning according to claim 1 is characterized in that The process of building the meta-learning model is as follows: Split meta-tasks, abstract entity and relationship extraction tasks into multiple meta-tasks, sample across domains, and capture task commonalities; Using the improved prototype network as the framework, text features are mapped to the semantic space through the feature encoder, prototype generator, and distance measurement module.
5. The unstructured text entity recognition and relationship extraction method under small sample learning according to claim 4 is characterized in that: The process of learning common knowledge between different tasks is: By clustering cross-task semantic prototypes, we cluster meta-task prototype vectors to find common semantic patterns across domains. By summarizing meta-rules, we incorporate the encoding into regularization terms into training to guide the model to learn interpretable extraction patterns.
6. The unstructured text entity recognition and relationship extraction method under small sample learning according to claim 1 is characterized in that The process of identifying entities is as follows: The word vector is input into the meta-learning model encoder, attention is used to enhance features, and the entity location is found in combination with syntax; sliding window detection is used, BIOES is used to mark boundaries, and multiple classifiers are integrated to predict types; finally, CRF is used to constrain the label sequence and cross-task knowledge transfer is used.
7. The method for unstructured text entity recognition and relationship extraction under small sample learning according to claim 6 is characterized in that: The process of extracting the relationship type between entities is as follows: Combine the identified entities into candidate pairs and extract contextual features; The features are input into the meta-learning model relationship classifier, and the relationship type is predicted based on the similarity with the relationship prototype, and multi-label classification is used to process the relationship situation; the domain relationship graph is used to filter the prediction, and the key related words are focused on by self-attention to optimize the extraction.
8. The method for unstructured text entity recognition and relationship extraction under small sample learning according to claim 1, characterized in that: The process of optimizing the dependencies between entities and relationships is as follows: Define entity and relationship compatibility rules and use transitive constraints; use dynamic programming to calculate transition probabilities based on entity positions and use the Viterbi algorithm to find the optimal path; Iteratively optimize parameters, verify consistency based on domain knowledge, and correct conflicting predictions.
9. The method for unstructured text entity recognition and relationship extraction under small sample learning according to claim 1, characterized in that: The process of data augmentation for a small number of labeled samples is as follows: Rule-based text transformation uses domain dictionaries to replace synonyms and restructure sentences. Based on the generative model, samples are generated based on entity and relationship types, and MLM masks fill in the text. Through adversarial training and adding perturbations, the model learns the essential semantic features.
10. The method for unstructured text entity recognition and relationship extraction under small sample learning according to claim 9, characterized in that: The process of optimizing the model is as follows: Balance complexity and expressiveness by adjusting the number of neural network layers; optimize feature interaction by improving the attention mechanism; integrate multiple models to reduce variance; incrementally train new data and compress the model through knowledge distillation.
Citation Information
Patent Citations
Chinese entity relation extraction method based on additional relation information
CN118332106A
Cited By
Forest fire knowledge modeling method based on named entity recognition and relation extraction
CN121524352A
Long text official document key information extraction agent method based on large model
CN121542411A