Entity relationship extraction research based on semi-supervised learning and curriculum learning

By combining course learning with semi-supervised learning and using iterative training with labeled and unlabeled data, we solved the problem of scarcity of labeled data in entity relationship extraction and improved the performance and accuracy of the model in low-resource environments.

CN120654691APending Publication Date: 2025-09-16NANJING TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510712354.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing entity relationship extraction methods mainly rely on supervised learning, which requires a large amount of labeled data, resulting in high labeling costs and difficulty in effective training in small sample or low-resource environments. In particular, the performance is insufficient when labeled data is scarce in professional fields.

Method used

A curriculum learning strategy is adopted to divide the labeled data into buckets of different difficulty levels. Pseudo-labels are generated through adaptive dynamic thresholds. Combined with semi-supervised learning methods, the model is gradually optimized. Iterative training is performed using a small amount of labeled data and a large amount of unlabeled data to form a curriculum-guided semi-supervised entity relationship joint extraction framework.

Benefits of technology

It significantly improves the generalization ability and performance of the model under limited data, effectively solves the problem of scarce labeled data, and improves the accuracy and efficiency of entity relationship extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654691A_ABST
    Figure CN120654691A_ABST
Patent Text Reader

Abstract

An entity relationship extraction method based on semi-supervised learning and curriculum learning is characterized in that a given sentence S is equal to {e1, e2,..., en}, and the target of entity relationship joint extraction is to comprehensively extract all entities and relationships thereof, so that all possible triples in a (s, r, o) form are identified; an overall framework of the entity relationship extraction method based on semi-supervised learning and curriculum learning comprises the following five steps: (1) based on an entity and relationship overlapping condition, dividing initial labeled data into data sets with different difficulties; (2) training a teacher network model by using the course classification data set; (3) screening a high-confidence pseudo-tag triple through a semi-supervised learning strategy; (4) adding the pseudo label as a real label into the original training set, and training a student network by adopting extended data of course classification; (5) taking the trained student network as a new teacher network for iterative optimization; and carrying out relation extraction by using the finally constructed large-scale entity relation extraction model with better performance. The model provided by the invention effectively solves the problem of insufficient label data in the entity relationship extraction task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and specifically is an entity relationship extraction method based on semi-supervised learning (SSL) and curriculum learning (CL). Background Art

[0002] Entity relationship extraction is a core task in natural language processing (NLP), which aims to identify entities and their relationships from unstructured text [1]. Specifically, the goal of entity relationship extraction is to automatically identify "entity-relationship-entity" triplets from text, representing the entities in the text and the relationships between them. In the text "Washington is the capital of the United States", the goal of the entity relationship extraction task is to identify the triple (Washington, capital, United States). This technology has important applications in knowledge graph construction [2], intelligent question-answering systems [3], machine translation [4] and other fields.

[0003] Current entity relationship extraction methods, especially deep learning neural network models based on Transformer[5], are mainstream. These methods utilize the self-attention mechanism of the Transformer architecture to effectively capture long-range dependencies in text, thereby improving the accuracy of entity relationship extraction. BERT[6] (Bidirectional Encoder Representations from Transformers), a pre-trained language model, has become an important tool in this field. These models learn rich semantic information on large-scale text data through pre-training tasks, and are then able to efficiently perform downstream tasks such as entity recognition and relationship classification.

[0004] Although existing entity relationship extraction methods have achieved significant success in many tasks, most methods still rely on supervised learning, which requires a large amount of labeled datasets [7] for effective model training. In supervised learning, training data must be manually labeled to identify the entities in the text and their relationships. This process is not only time-consuming and labor-intensive, but also extremely costly, especially when large-scale data is required. Therefore, the acquisition of labeled data has become a bottleneck for model training. In particular, in many professional fields, relevant labeled data is often scarce, making it difficult for existing models to be effectively trained in small sample or low-resource environments.

[0005] To address the problem of data scarcity, we proposed semi-supervised learning strategies[8][9] and curriculum learning strategies

[10]

[11] . Semi-supervised learning significantly reduces the reliance on large-scale labeled datasets by jointly utilizing a small amount of labeled data and a large amount of unlabeled data, thereby improving the performance of the model in limited data scenarios. The curriculum learning strategy enables the model to first consolidate the initial representation through basic tasks, and then gradually process more complex data associations, ultimately improving the generalization ability of the model. Summary of the Invention

[0006] In order to solve the above problems, the present invention proposes a curriculum learning guided semi-supervised entity relationship joint extraction framework. The framework demonstrates high generalization ability and is compatible with existing entity relationship extraction methods. It effectively solves the challenge of model optimization when labeled data is limited. Specifically, the framework adopts a curriculum learning strategy to divide the initial labeled data into different "buckets" according to the overlap of entities and relationships. These data buckets are processed by the model in turn to establish an initial teacher model. Next, an adaptive dynamic threshold is applied to generate pseudo labels for unlabeled data, thereby realizing iterative updates of the training dataset. The expanded training set is further optimized through curriculum learning and cyclic optimization. In each iteration, the retrained model becomes the teacher model, relabels the unlabeled data, and gradually incorporates all instances into the training set.

[0007] The detailed construction steps of the present invention are:

[0008] (1) Using a curriculum learning strategy, a traditional teacher network is trained with a small amount of labeled data;

[0009] (2) Screening high-confidence pseudo-label triplets through a semi-supervised learning strategy;

[0010] (3) Add the pseudo labels as true labels to the original training set;

[0011] (4) Train the student network using the curriculum learning strategy with the expanded data;

[0012] (5) The student network is trained as a new teacher network to participate in the next round of iteration to improve performance.

[0013] The main contributions of the present invention are as follows:

[0014] 1. This paper proposes a novel curriculum-guided semi-supervised learning framework for entity relationship extraction. This framework demonstrates compatibility with general entity relationship extraction methods by leveraging limited labeled data and large-scale unlabeled data. This framework innovatively combines semi-supervised learning with curriculum learning to address the scarcity of labeled data.

[0015] 2. The framework implements a difficulty quantification method based on entity overlap features to optimize the use of labeled data. This method establishes a progressive optimization strategy that gradually transitions from simple examples to more complex ones. By introducing a dynamic course scheduling mechanism, the framework effectively reduces overfitting in the early training stages while improving overall model performance.

[0016] 3. Experimental evaluations using three common entity relation extraction models on three benchmark datasets validate the robustness and effectiveness of the proposed framework. Empirical results show that the proposed framework achieves significantly better performance in the task of joint entity relation extraction compared to traditional supervised learning methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is the overall framework diagram of the method (model) proposed in this invention.

[0018] Figure 2 This is the course learning framework diagram of the present invention.

[0019] Figure 3 It is a radar chart on the single entity overlap dataset and the entity pair overlap dataset.

[0020] Figure 4 It is a visualization of the subject in the triple. DETAILED DESCRIPTION

[0021] The following is a further explanation of this solution with the help of the accompanying figures and specific implementation methods. The first part summarizes the entire framework method; the second part details the implementation of the curriculum learning strategy; the third part introduces the selection of pseudo-labels, including the implementation of supervised learning and semi-supervised learning; the fourth part covers the experimental details, including the dataset used, the selection of baseline models, the presentation of results, ablation experiments and the visualization analysis of results; the fifth part provides concluding remarks and future directions.

[0022] 1. Project Overview

[0023] Given a sentence S = {e1, e2, ..., e n}, the goal of entity-relation joint extraction is to comprehensively extract all entities and their relations, thereby identifying all potential triples (s, r, o). This study proposes a novel entity-relation extraction framework that organically combines self-training and curriculum learning to overcome the limitations of data scarcity. The framework effectively selects high-confidence pseudo-label samples from a large amount of unlabeled data, significantly improving the performance of relation extraction. Experimental results show that the framework has strong generalization capabilities in different extraction architectures, especially improving performance in low-resource environments. This integrated framework not only optimizes the utilization of unlabeled data, but also establishes an effective paradigm for resource-efficient relation extraction.

[0024] Figure 1 We present our proposed overall framework, which consists of five main steps: (1) In the initial stage, the teacher network performs curriculum learning on labeled data, where samples are gradually organized according to difficulty to promote the model to gradually adapt to complexity; (2) The trained teacher network generates pseudo-labels for unlabeled data, and then retains high-confidence triplets through semi-supervised screening; (3) An enhanced training set is constructed by combining pseudo-labeled data with the original labeled data; (4) The curriculum learning strategy is adopted again to train the student network using the expanded training set; (5) The trained student network iteratively plays the role of teacher for subsequent pseudo-label generation and model optimization.

[0025] 2. System Model

[0026] 2.1 Course Difficulty Classification

[0027] The first task of this paper is to design a course level suitable for the entity relationship extraction task. The key is to identify features that can make it easier for the model to extract entities and relationships from the data. Considering that there are overlapping entities in most text data, the frequency of these overlaps is proposed as an indicator to quantify the difficulty of the data. The underlying assumption is that a dataset with a high degree of entity overlap will bring greater complexity to the extraction task. The proposed course learning difficulty estimator is as follows Figure 3 Specifically, for each labeled data sample (x i ,y i ), where y i ={(s i , r i , o i )|i=1,2,...,n}. Therefore, the difficulty level of the triple is d s (y i ) is defined as:

[0028] d s (y i )=αN un (y i )+βN ro (y i )+γN eo (y i ),

[0029] Among them, N un (y i ) represents the number of unique entity pairs in the labeled data, N ro (y i ) represents the number of entity pairs with single entity overlap, N eo (y i ) represents the number of (s, o) entity pairs that overlap completely. α, β, and γ represent Nun (y i ), N ro (y i ) and N eo (y i ), which is used to quantify the impact of entity overlap on data complexity. In general, triples are classified according to their overlap complexity: triples with no entity or relationship overlap are the simplest, followed by triples with single entity overlap, and finally triples with entity pair overlap are considered the most complex.

[0030] A progressive training scheduler is used to structure the entire training process. Specifically, the training set D is initially divided into several buckets according to the complexity and difficulty of the training samples, denoted as {D1, ...D T Each bucket contains training examples with similar difficulty scores, enabling the model to most effectively learn hierarchical knowledge at different difficulty levels by gradually increasing training complexity. Training begins with the simplest bucket, D1, initially focusing on relatively simple, easy-to-learn examples to ensure the model acquires a solid foundation of knowledge, enabling good generalization capabilities.

[0031] As training progresses, after a predetermined number of epochs or when the model reaches convergence, the next bucket (containing more challenging examples) is incorporated into the current training set. This incremental addition increases the complexity of the training set, prompting the model to continuously improve its performance as the training difficulty increases. Once all buckets have been incorporated and utilized, the training process continues for more epochs to further fine-tune the model.

[0032] 2.2 Supervised Learning

[0033] The main goal of this framework is to transform natural language sentences into high-dimensional numerical vectors, thereby facilitating the efficient processing of text data by machine learning models. To achieve this goal, we utilize BERT (Bidirectional Encoder Representations from Transformers), a state-of-the-art model based on the Transformer architecture, to extract key information and generate context-aware embedding vectors. BERT is derived from the Transformer architecture and adopts a self-attention mechanism to train the encoder-decoder model. It uses scaled dot product attention and multi-head attention by computing the dot product between the query and the key and then obtaining the attention weights through softmax. This enables the model to simultaneously focus on information from different subspaces and capture local and global dependencies. The self-attention layer connects all network positions, enabling the embedding vector to distinguish important semantic information and extract valuable insights. This mechanism is crucial for establishing long-range dependencies and revealing hidden patterns in the input sequence, which makes the Transformer perform well in natural language processing tasks. In particular, the self-attention mechanism is defined by the following formula:

[0034]

[0035] Where Q, K, and V represent query, key, and value matrices respectively, and d k is the dimension of the key vector. These matrices are obtained by linearly transforming the input X:

[0036] Q=XW Q ,

[0037] K=XW K ,

[0038] V=XW V ,

[0039] Among them, W Q 、W K and W V is a learnable weight matrix. The dot product between the query and the key is calculated and passed Scaled to prevent large values ​​from causing gradient instability. A softmax function is then applied to obtain attention weights, which are used to weight the value vector V. This enables the model to simultaneously focus on information from different representation subspaces.

[0040] Through the self-attention mechanism, the relationship between each word vector and other word vectors in the sentence is calculated to generate the context representation of each word. Specifically, for each word vector e i , whose context indicates h i It is calculated by the following formula:

[0041]

[0042] Among them, α ij Expressing word e i and word e j The attention weight between V Represents the value projection matrix in the self-attention mechanism. Attention weight α ij It is calculated by the softmax function:

[0043]

[0044] Let entity e i and entity e j The corresponding word vector sequences are and Their relationship categories are predicted through linear transformation and Softmax function:

[0045]

[0046] in, The feature cascade representing the entity representation, W r is the learnable weight matrix, b r is the bias term of the affine transformation. Based on the above processing, the sentence is converted into a high-dimensional vector, which is convenient for subsequent entity and relationship extraction tasks.

[0047] Generally speaking, most entity relationship extraction models usually consist of four key components: input preparation and text encoding, entity recognition, entity pair extraction, and relationship extraction. First, contextualized word representations are generated through a pre-trained model (such as BERT). Subsequently, the entity boundaries are identified and entities are extracted through the feature extraction module, and each entity is assigned a corresponding category label. Next, all possible entity pairs are systematically extracted from the identified entities. Finally, for each pair of entities, their contextual representations are concatenated and input into a classifier to predict the relationship between them. During the training process, these models optimize the network parameters through backpropagation, gradually improving the ability to detect entities and their relationships. This process follows the principle of entity-relationship co-optimization, and triples are extracted through probability scoring. It should be noted that although different entity-relationship extraction models may differ in architectural implementation, they generally follow these basic processes. The joint probability p(s, r, o|X) of triple extraction can be expressed as:

[0048] p(s,r,o|X)=p(s|X)·p(o|s,X)·p(r|s,o,X),

[0049] Here, p(s|X) represents the probability of extracting entity s given input X, p(o|s,X) represents the probability of extracting entity o given X and entity s, and p(r|s,o,X) represents the probability of extracting relation r given X, entity s, and entity o. This model jointly performs entity recognition and relation extraction in a unified framework by leveraging contextual information X, extracting entities and their corresponding relations.

[0050] Given a small labeled dataset D l and a large unlabeled dataset D u , the goal of supervised learning is to train a baseline model f θ , is called the teacher model. This is achieved by labeling the dataset D l This is achieved by training a standard entity-relation extraction model on , using a supervised learning strategy:

[0051]

[0052] in, Represents the binary cross entropy loss function used to calculate label loss. iRepresents the actual label, usually in the form of one-hot encoding, indicating the true category of the sample. i represents the probability of the i-th label.

[0053] 2.3. Semi-supervised learning

[0054] For each unlabeled sample x u ∈D u , the data is processed by the teacher model to generate a predicted output. Since the model usually outputs a distribution representing the probability of each potential label, pseudo labels can be generated by selecting the label with the highest probability The details are as follows:

[0055]

[0056] in, Indicates that given an input x u and model parameters θ, the posterior probability of label y, the pseudo label is assigned the label with the highest probability.

[0057] In general pseudo-labeling methods, the strategy is usually to directly select the category with the highest predicted probability as the pseudo-label. However, in practical applications, the model may show low confidence on all categories, resulting in relatively low predicted probability values. In this case, if the label with the highest probability is still forcibly selected as the pseudo-label, a large number of noisy labels may be introduced. To solve this problem, a confidence threshold mechanism is proposed. If the model's prediction confidence for the unlabeled sample in round t exceeds The prediction is then assigned as the pseudo label for the sample. This selection criterion generates the pseudo label dataset used in round t:

[0058]

[0059] To maximize the utility of unlabeled data, a triplet threshold is used. This threshold reflects the model's average confidence on the unlabeled dataset and is dynamically adjusted as training progresses. This adaptive mechanism helps to select unlabeled samples more and more precisely while ensuring that incorrect pseudo-labels are excluded in a timely manner. The threshold update takes into account the current model's prediction confidence and the global data distribution, allowing for more precise control over whether samples are included in the training set throughout the training process. The threshold is defined as follows:

[0060]

[0061] in, represents the global threshold at the tth iteration, is the initial threshold (usually set to a higher value in the first round to ensure that high-confidence pseudo labels are used for model training), and λ is a smoothing factor (ranging from 0 to 1) used to control the previous threshold Current threshold The influence weight of . Represents the sum of the maximum confidence of all samples in the current batch.

[0062] The selected high-confidence pseudo-label samples enhance the training set and are actually treated as true labels. Therefore, the model not only learns from the original labeled data, but also utilizes the supervisory signal from a large amount of unlabeled data. Specifically, for each high-confidence pseudo-label sample x u , pseudo labels is regarded as its true label and is included in the existing annotation set D l , thus forming an expanded training set

[0063]

[0064] Specifically, the difficulty evaluator reclassifies the expanded dataset and organizes the data from simple to complex for subsequent training iterations. This integration enables the model to integrate knowledge from pseudo-labeled data while strengthening learning from the original dataset. Therefore, this approach enhances the learning and generalization capabilities of the model. In the t+1th iteration, the model is trained using both supervised learning and semi-supervised learning strategies, and the total loss function is Defined as:

[0065]

[0066] In the tth iteration, the model learns from both the labeled dataset and the pseudo-labeled dataset. The learning process consists of two parts: supervised learning and semi-supervised learning. In the supervised learning part, the loss function quantifies the difference between the model prediction and the true label, ensuring that the model accurately captures the distribution of the labeled data. In the unsupervised learning part, the model is trained on the unlabeled dataset D. u Then apply the screening mechanism to select high confidence predictions and generate a set of pseudo-label data D u′ This process enhances the training signal by identifying reliable predictions and utilizing unlabeled data to improve the model’s generalization capabilities.

[0067] 2.3. Overall framework approach

[0068] Algorithm 1 presents an overview of the proposed framework. Initially, a difficulty estimator is used to partition the labeled dataset into buckets of different difficulty levels. This process ensures that the model is trained sequentially, starting with the simplest samples and gradually processing more complex samples. After training on the labeled data, a semi-supervised learning method is used to generate pseudo labels for the unlabeled data. The model makes triplet predictions for these unlabeled data and assigns a confidence score to each triplet. If the confidence score exceeds a predefined threshold Tt, the corresponding triplet is considered a valid pseudo label and is added to the training set. On the contrary, if the confidence score is below Predictions for will be discarded. This prevents the introduction of noisy or low-quality pseudo-labels, thereby preserving the integrity of the training data and reducing the potential negative impact on model performance.

[0069] The proposed framework fundamentally differs from traditional supervised learning-based entity relationship extraction methods. Unlike traditional supervised methods that require a large number of manually annotated samples, this framework adopts an iterative self-training paradigm to select high-confidence pseudo-labels. This architecture combines semi-supervised learning (which effectively utilizes unlabeled data) with a curriculum learning strategy (which optimally utilizes annotated examples). This synergistic integration enables the method to maintain strong performance even in data-scarce scenarios.

[0070] Algorithm 1 is based on a curriculum-guided semi-supervised entity relationship extraction framework.

[0071]

[0072] 3 Experimental design

[0073] 3.1 Dataset

[0074] In the entity relationship extraction experiments, three widely used public benchmark datasets are introduced to evaluate the performance. The detailed statistics of these three datasets are as follows:

[0075] Table 1: Datasets used

[0076]

[0077] NYT dataset

[12] : This dataset is a distantly supervised single-sentence corpus for relation extraction, derived from articles in The New York Times. This benchmark dataset contains over 50,000 entity-relation pairs, covering multiple fields such as geography, politics, and economics, and is one of the most widely used resources in relation extraction research.

[0078] WeNLG dataset

[13] : This dataset is a manually annotated corpus originally designed for natural language generation tasks. The dataset is characterized by high-quality annotations and a wide coverage of historical, scientific, and cultural fields. It contains linguistically diverse but concise sentence samples with rich relation triples.

[0079] CoNLL dataset

[14] : This dataset is a widely used dataset derived from news articles, containing approximately 800 sentences with carefully annotated entities and relations. Known for its annotation accuracy and rich relation diversity, this dataset remains an important resource for promoting the development of relation extraction methods.

[0080] We split the labeled data from all three datasets into training sets at a ratio of 50% and 30%, respectively, with the remaining 50% and 70% used as unlabeled training data. To avoid the bias associated with traditional random sampling—which can produce non-representative datasets and impair model performance—we employed a curriculum-based stratified sampling strategy: we first categorized the dataset by course difficulty level and then sampled proportionally within each difficulty level. This approach ensured a balanced representation of the training set across both the feature space and the difficulty distribution.

[0081] 3.2 Baseline Model

[0082] In the biomedical relation extraction task, we conducted a series of experiments to verify the effectiveness and practicality of our method and compared it with some strong baseline methods. The following are some baseline methods used for comparison:

[0083] 1. Att-Bilstm: Att-Bilstm

[15] combines the attention mechanism with the Long Short-Term Memory (LSTM) network. Bilstm captures long-range dependencies, while the attention mechanism highlights words related to the relationship, helping to extract key information. It performs well in tasks dealing with simple relationships and effectively captures contextual information.

[0084] 2. CasRel: CasRel

[16] uses a cascaded binary annotation framework for end-to-end entity relationship extraction. The model first identifies entities in the text and then classifies the relationship between each pair of entities, effectively solving the problem of relationship overlap. With its simple and easy-to-implement architecture, CasRel is particularly suitable for handling complex relationship extraction tasks.

[0085] 3. OneRel: OneRel

[17] is an entity relation extraction model based on a unified annotation framework. It integrates entity recognition and relation extraction into a single step through a unified module. The model has a simple architecture and demonstrates high efficiency in training and inference, making it particularly suitable for tasks that require a close integration of entity and relation extraction. OneRel excels in jointly extracting entities and relations.

[0086] 3.3 Evaluation Metrics

[0087] In this experiment, we evaluated our proposed framework on the three public benchmark datasets mentioned above and compared its performance with the original model method. Precision P, recall R, and micro-average F1 were used as indicators to evaluate model performance. The introduced indicators are defined as follows:

[0088]

[0089] Among them, TP, FP and FN represent the number of true positives, false positives and false negatives, respectively.

[0090] 3.4 Experimental Results

[0091] In this study, we employed two evaluation methods: exact matching and partial matching. Exact matching considers a triple (s, r, o) correctly predicted only if all its components—the subject (s), the relation (r), and the object (o)—are accurately predicted. In contrast, partial matching considers a triple correctly predicted as long as the subject (s) and the object (o) are correctly identified, regardless of whether the relation (r) between them is correct. Our experimental results and analysis primarily discuss the performance of exact matching.

[0092] Table 2: Experimental results on the NYT dataset. The improved results are marked in bold.

[0093]

[0094] Table 3: Experimental results on the WeNLG dataset. The improved results are marked in bold.

[0095]

[0096] Table 4: Experimental results on the CoNLL dataset. The improved results are marked in bold.

[0097]

[0098] The results in Tables 2, 3, and 4 show that our framework significantly improves model performance compared to the baseline model alone. Under the condition of exact matching with only 50% labeled data, our framework significantly improves the performance of three entity relationship extraction methods (Att-Bilstm, CasRel, and OneRel): on the NYT dataset, CasRel's F1 score increases from 79.8% to 83.7%; on the WebNLG dataset, CasRel and OneRel's F1 scores increase from 85.3% and 76.2% to 87.6% and 77.9%, respectively; and on the CoNLL dataset, Att-Bilstm, CasRel, and OneRel's recall increases by 4.0%, 3.0%, and 2.9%, respectively, while their F1 scores also continue to improve. These results fully demonstrate the framework's effectiveness in improving entity relationship extraction performance with limited labeled data.

[0099] When the proportion of labeled data is reduced to 30%, experimental results show that our framework maintains a competitive advantage over Att-Bilstm, CasRel, and OneRel on three datasets (NYT, WebNLG, and CoNLL): On the NYT dataset, the framework outperforms the baseline methods in both recall (R) and F1 value, with Att-Bilstm's F1 value increasing from 61.7% to 62.5%, CasRel significantly increasing from 75.9% to 80.0%, and OneRel significantly increasing from 82. 3% to 82.9%; on the WebNLG dataset, the framework significantly improved the recall rate and F1 value, especially Att-Bilstm (F1: 65.0%→66.4%) and CasRel (F1: 79.9%→82.4%); on the CoNLL dataset, the framework significantly improved the F1 value of each method - Att-Bilstm (20.7%→28.9%), CasRel (27.6%→29.5%) and OneRel (34.9%→40.4%).

[0100] Under the partial matching setting, the proposed framework shows significant improvements at different annotated data ratios (50%, 30%). When using 50% annotated data, the F1 scores of all three datasets show a consistent improvement, with the CasRel method showing the most significant improvement on the NYT dataset (from 80.6% to 84.0%). Even when the annotation ratio is reduced to 30%, the framework still maintains an F1 score that is superior to the baseline model. These results fully demonstrate the effectiveness of the framework in low-resource scenarios with limited annotated data.

[0101] 4 Performance Analysis and Comparison

[0102] In the following discussion, we analyze various factors that influence the model's performance in the above experiments, including the curriculum learning strategy and the course difficulty level of the data. In addition, we also analyze the decision boundary characteristics between different entities through vector visualization.

[0103] 4.1 Model performance under the course learning strategy:

[0104] To explore the impact of the curriculum learning strategy on entity relationship extraction performance, this study designed an ablation experiment: First, the model was trained using a standard semi-supervised learning method without the curriculum learning strategy, and its performance on various datasets was recorded. Subsequently, the curriculum learning strategy was introduced and its performance improvement was evaluated. As shown in Table 5, the model's performance on the WebNLG dataset with a 50% annotation ratio was as follows when the curriculum learning strategy was not used.

[0105] Table 5: Model performance on data with different noise ratios

[0106]

[0107] As shown in Table 5, the curriculum learning strategy has different effects on different models: in the Att-Bilstm model, the precision (P) increased from 82.4% to 83.7%, but the F1 value remained basically unchanged; in contrast, the precision (P) of the CasRel model increased from 87.9% to 88.1%, the recall (R) increased from 87.0% to 87.2%, and the F1 value increased from 87.4% to 87.6% accordingly; and the OneRel model achieved significant improvements in precision (79.6% → 80.7%), recall (74.5% → 75.4%), and F1 value (76.9% → 77.9%).

[0108] 4.2 Performance under higher course difficulty data:

[0109] Experimental results show that while the proposed framework performs well on all three datasets, the distribution of class samples within these datasets varies significantly. This imbalance can affect the model's training process and final performance, especially in the case of class imbalance. To more comprehensively evaluate the effectiveness of the framework, we conducted extended experiments using the CasRel model as a baseline: on WebNLG datasets with annotation ratios of 30% and 50%, based on the aforementioned curriculum learning strategy, we specifically tested the model's performance on data with higher curriculum difficulty (including single entity overlap and entity pair overlap). The specific performance is shown in Tables 6 and 7.

[0110] Table 6: Performance of single entity overlapping categories

[0111]

[0112] Table 7: Performance of entity pairs with overlapping categories

[0113]

[0114] Experimental results show that in the single entity overlap task, the CSJE framework proposed in this paper achieves stable improvements over the baseline model in all evaluation metrics (precision, recall, and F1 value), regardless of whether the proportion of labeled data is 30% or 50%. More importantly, in the more challenging entity pair overlap scenario, our framework successfully addresses the key flaws of the baseline CasRel model: although CasRel achieved an F1 value of 81.5% with 30% labeled data, its recall rate was only 72.7%, indicating that the model is insufficient in detecting low-frequency entity pairs. By incorporating semi-supervised self-training into the curriculum learning framework, CSJE not only improves the F1 value to 91.3% (an increase of 9.8 percentage points), but also significantly improves the recall performance. These findings strongly demonstrate that the proposed framework can effectively solve the problem of complex entity pair overlap in scenarios with scarce annotations.

[0115] 4.3 Entity Visualization

[0116] The visualization results clearly show that the framework can effectively represent the entity distribution in the vector space. Figure 4 As shown, the projection results show obvious entity clustering characteristics - entities with similar semantic properties are close to each other in space. Specifically: Figure (a) shows the clustering distribution of specific entity types (such as people's names); Figure (b) shows the clustering characteristics of location entities (such as parks and valleys); Figure (c) shows the clustering of country / region entities (such as Afghanistan and Canada). These findings confirm that the framework is able to capture the semantic similarity between entities and present it intuitively through dimensionality reduction projection. Such visualization results can provide practical and effective guidance for downstream tasks.

[0117] 5. Conclusion

[0118] This paper proposes a curriculum-guided semi-supervised learning framework, CSJE, to address the challenge of model optimization under data-scarce conditions. This framework improves existing supervised entity relationship extraction methods through two major innovations: (1) a curriculum-based training strategy enhances model robustness; and (2) a confidence-based self-training mechanism generates reliable pseudo-labels from unlabeled data. The synergistic integration of curriculum learning and semi-supervised learning maximizes the utility of both labeled and unlabeled data, demonstrating significantly superior extraction performance compared to traditional methods that rely solely on labeled data.

[0119] Experimental results demonstrate that the CSJE framework achieves significant performance improvements when applied to a variety of mainstream models on three benchmark datasets (NYT, WebNLG, and CoNLL). Notably, the framework maintains excellent recall and F1 scores even in scenarios with limited labeled data. Furthermore, extensive experiments validate the framework's high compatibility with existing entity relationship extraction paradigms, providing new insights for advancing research in semi-supervised information extraction.

[0120] References:

[0121] [1] L.

[0122] [2] H. Luo, Y. Yang, T. Yao, Y. Guo, Z. Tang, W. Zhang, S. Peng, K. Wan, M. Song, W. Lin, et al., Text2nkg: Fine-grained n-ary relation extraction for n-ary relational knowledge graph construction, Advances in Neural Information Processing Systems 37 (2024) 27417-27439.

[0123] [3]X.Li, F.Yin, Z.Sun,

[0124] [4]Z.Wang,J.Yahg,T.Li,L.Chai,J.Liu,Y.Mo,J.Bai,Z.Li,Multilingualentity and relation extraction from unified to language-specific training,in:Proceedings of the 2023 International Conference on Electronics,Computers andCommunication Technology,2023,pp.98-105.

[0125] [5]Z.Yang,J.-K.Lee,Bert for entity-relation extraction in biomedicaltexts,in:Proceedings of the 28th International Conference on ComputationalLinguistics,2020,pp.1223-1231.

[0126] [6]M.V.Koroteev,Berr:A review of applications in natural languageprocessing and understanding,Journal of Computer Science and Tcchnology 36(4)(2021)764-784.

[0127] [7]R.Xiao,L.Feng,K.Tang,J.Zhao,Y.Li,G.Chen,H.Wang,Targetedrepresentation alignment for open-world semi-supervised learning,in:Proceedings of the 2024IEEE / CVF Conference on Computer Vision and PatternRecognition,2024,pp.23072-23082.

[0128] [8]JLGarrido-Labrador,A.Serrano-Mamolar,J.Maudes-Raedo,JJRodr′ιguez,C.Garc′ιa-Osorio,Ensemble methods and semi-supervised learning forinformation fusion:A review and future research directions,Information Fusion107(2024)102310.

[0129] [9]Y.Wang,H.Jian,J.Zhuang,H.Guo,Y.Leng,Sslmm:Semi-supervised learningwith missing modalities for multimodal sentiment analysis,Information Fusion120(2025)103058.

[0130]

[10] Y.Gu,S.Zheng,Z.Xu,Q.Yin,L.Li,J.Li,An efficient curriculum learning-based strategy for molecular graph learning.,Brieffngs inBioinformatics 23(2022)bbac099-bbac099.

[0131]

[11] X.Wang,Y.Chen,W.Zhu,A Survey on Curriculum Learning,IEEETransactions on Pattern Analysis & Machine Intelligence 44(2022)4555-4576.

[0132]

[12] S.Riedel,L.Yao,A.McCallum,Modeling relations and their mentionswithout labeled text,in:Proceedings of the 2010European conference on MachineIearning and knowledge discovery in databases:Part III,2010,pp.148-163.

[0133]

[13] C.Gardent,A.Shimorina,S.Narayan,L.Perez-Beltrachini,Creatingtraining corpora for nlg micro-planning,in:55th Annual Meeting of theAssociation for Computational Linguistics,ACL 2017,Association forComputational Linguistics(ACL),2017,pp.179-188.

[0134]

[14] E.F.Tjong Kim Sang,F.De Meulder,Introduction to the conll-2003shared task:language-independent named entity recognition,in:Proceedings ofthe seventh conference on Natural language learning at HLTNAACL 2003-Volume4,2003,pp.142-147.

[0135]

[15] P.Zhou,W.Shi,J.Tian,Z.Qi,B.Li,H.Hao,B.Xu,Attention-basedbidirectional long short-term memory networks for relation classification,in:Proceedings of the 54th Annual Meeting of the Association for ComputationalLinguistics,2016,pp.207-212.

[0136]

[16] Z.Wei,J.Su,Y.Wang,Y.Tian,Y.Chang,A novel cascade binary taggingframework for relational triple extraction,in:Proceedings of the 58th AnnualMeeting of the Association for Computational Linguistics,2020,pp.1476-1488.

[0137]

[17] Y.M.Shang,H.Huang,X.L.Mao,Onerel:Joint entity and relationextraction with one module in one step,in:Proceedings of the 36th AAAIConference on Artificial Intelligence,Association for the Advancement ofArtificial Intelligence,2022,pp.11285-11293。

Claims

1. A method for entity relationship extraction based on semi-supervised learning and curriculum learning. Given a sentence S = {e1, e2, ..., e n The goal of entity-relationship joint extraction is to comprehensively extract all entities and their relationships, thereby identifying all possible triples of the form (s, r, o). The specific steps are as follows: 1) Based on the overlap of entities and relations, the initial labeled data is divided into data sets of different difficulty levels. The number of initial data sets is relatively small; 2) Train a teacher network model using the course classification dataset; 3) Screening high-confidence pseudo-label triplets through a semi-supervised learning strategy; 4) Pseudo-labels are added to the original training set as true annotations, and the student network is trained using the extended data of course classification; 5) The trained student network is used as the new teacher network for iterative optimization to build a large-scale entity relationship extraction model with better performance. In step 1), given a labeled data (x i ,y i ), y i ={(s i , r i , o i )|i=1,2,...,n}, where, Course classification difficulty d(y i ) is defined as: d s (the i )=αN un (the i ) + βN ro (the i ) + γN eo (the i ), Among them, N un (y i ) represents the number of unique entity pairs in the labeled data, N ro (y i ) represents the number of entity pairs with single entity overlap, N eo (y i ) represents the number of (s, o) entity pairs that overlap completely. α, β, and γ represent N un (y i ), N ro (y i ) and N eo (y i ) is used to quantify the impact of entity overlap on data complexity. In step 2), based on the above course classification data, a teacher network model f is trained in the order of increasing course difficulty. θ , this process is based on supervised learning strategy, specifically: Assume that x represents the input data, y i Represents the actual label (usually in the form of one-hot encoding), which is used to identify the true category of the sample, p i Then it represents the predicted probability of the i-th label: Assume f θ (x) represents the trained teacher model, where θ represents the parameters of the network f; The cross entropy loss function of the supervised training process is expressed as: Among them, the loss function contains the sum of the losses of the three parts (s, r, o), which can also be expressed as: By minimizing the supervision function, we can obtain a teacher network with relatively good performance f θ (x). In step 3), after obtaining the teacher network, θ (x), we need to predict pseudo labels for unlabeled data. For the unlabeled dataset D u Each sample x in u , the teacher model will make predictions and generate a probability distribution output, which is as follows: Among them, pseudo labels The corresponding label with the highest probability will be assigned. In practical applications, the model may have low prediction confidence for all categories, resulting in an overall small output probability value. To solve this problem, this paper proposes to introduce a confidence threshold mechanism, specifically: in, represents the set of pseudo labels selected in round t, and S represents the confidence threshold used by the model when screening pseudo labels in the T-th iteration. In step 4), the high-confidence pseudo-label samples screened out will be added to the training set and used as true labels. This process belongs to the semi-supervised learning process. Specifically, for each high-confidence pseudo-label sample, its pseudo-label will be regarded as the true label and merged into the original annotation set D l Thus, the expanded training set is formed: In step 5), we add the pseudo-labeled data with high confidence from step 4) to the initial training set as labeled data. At this time, we comprehensively consider the current model prediction confidence and the global data distribution to achieve precise control of the access of training set samples. The threshold update formula is expressed as: in, represents the global threshold at the tth iteration, is the initial threshold (usually set to a higher value in the first round to ensure This framework is fundamentally different from traditional supervised learning-based entity relationship extraction methods. Unlike traditional supervised methods that require large numbers of manually labeled samples, this framework employs an iterative self-training paradigm to select high-confidence pseudo-labels. Its architecture deeply integrates semi-supervised learning (which efficiently utilizes unlabeled data) with curriculum learning strategies (which optimize the use of labeled samples). This synergistic integration enables the model to maintain excellent performance even in data-scarce scenarios.