Reliable low-resource entity analysis method based on rule enhancement
By matching MD-guided data augmentation and model training, the problem of insufficient interpretability and accuracy of entity resolution methods is solved, especially under low resource conditions, the performance of entity resolution systems is significantly improved.
Patent Information
- Application Number
- CN202510614297.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-22
AI Technical Summary
The existing learning-based entity analytics lack interpretability, the cost of labeling training data is high, and the unbalanced training data leads to insufficient model accuracy, especially in low-resource scenarios.
采用匹配依赖MD指导数据增强,通过合成高质量的伪标签样本,生成满足匹配依赖MD的正样本和违反MD的负样本,结合最小化损失函数训练模型参数,提升模型在低资源条件下的准确性和可解释性。
Improve the accuracy of entity resolution under low resource conditions, while giving model interpretability, alleviating the problem of sample category imbalance, generating trusted pseudo-labels, and improving the F1 score performance of multiple data sets.
Smart Images

Figure CN120524952A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of entity resolution technology, and more specifically, to a reliable low-resource entity resolution method based on rule enhancement. Background Art
[0002] Entity Resolution (ER) is the task of determining whether two data entries refer to the same real-world entity. Figure 1 As shown in the figure, in two relational tables containing a large number of tuples describing products, the goal of entity resolution is to identify all pairs of tuples that refer to the same product (e.g., t1 and t4). In some scenarios, entity resolution is also known as entity linking, deduplication, or record linking. Entity resolution is crucial for data cleaning and integration, and has a profound impact on data-intensive downstream applications such as e-commerce, censuses, and corporate recruitment. As an important and enduring problem in the database field, entity resolution has been studied for decades.
[0003] Entity resolution solutions can be mainly divided into learning-based solutions and rule-based solutions. Although rule-based solutions are inherently interpretable, their accuracy is often insufficient compared to learning-based solutions, which makes the latter more popular in the research of this task. However, learning-based solutions have the following disadvantages:
[0004] a. Machine learning solutions, such as deep learning, are black boxes, and their predictions lack interpretability.
[0005] b. Labeled training data requires a large amount of manual labeling, which is very expensive and limits the scale of data annotation and the capabilities of the model.
[0006] c. Since most tuple pairs are unmatched, even with a large number of labeled training samples, the number of positive samples (i.e., matching tuple pairs) is far less than the number of negative samples (i.e., unmatched tuple pairs). The imbalance of training data further hinders the training of high-quality ER models.
[0007] Therefore, it is necessary to provide a reliable low-resource entity resolution method based on rule enhancement. Summary of the Invention
[0008] The purpose of the present invention is to provide a reliable low-resource entity resolution method based on rule enhancement to overcome the defects of the prior art.
[0009] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0010] A reliable low-resource entity resolution method based on rule enhancement includes the following steps:
[0011] S1. Extract matching dependency MD from a given dataset;
[0012] S2, using extraction matching dependent MD to enhance the training sample to obtain the enhanced training sample;
[0013] S3. Based on the enhanced training samples, the model is trained with the goal of minimizing the loss function L to obtain the optimal model parameters.
[0014] Furthermore, the step S1 specifically includes the following steps:
[0015] S11. Merge the given relational data tables R and R′ into a unified relational data table R″, where each record in R″ represents a tuple pair, marked as matched (1) or unmatched (0);
[0016] S12. Generate candidate MDs by identifying attribute-level dependencies for the relational data table R″.
[0017] Furthermore, the step S12 specifically includes:
[0018] For each attribute pair (A i ,A j ), calculate the similarity score based on their respective embedding vectors;
[0019] Based on the similarity score, by applying a threshold vector σ=[σ1,σ2,…,σ m ] to construct a candidate MD, and the candidate MD meets the following requirements: {Sim(t i· A1,t j· A1)≥σ1∧…∧Sim(t i· A m ,t j· A m )≥σ m}→Match, which means that if the similarity of a pair of tuples to the corresponding attribute embedding vectors reaches a threshold, the two tuples are considered to be matched.
[0020] Furthermore, the step S2 specifically includes the following steps:
[0021] S21, using labeled training sample set S train Synthesize new samples from the attribute values of each tuple in ;
[0022] S22, using matching dependency MD to assign pseudo labels to the new samples, and screen to obtain a set S of synthetic samples syn ;
[0023] S23, the set of synthetic samples S syn Compared with the existing labeled training sample set S train Merge to get the enhanced training sample S hybrid , as an enhanced training set for training entity resolution models.
[0024] Furthermore, the step S21 synthesizes each tuple pair When any attribute value A i ∈A, randomly select two values from all possible values in the tuple pair, introduce a perturbation to them by adding, deleting or changing characters, and then assign the perturbed values to A of the synthetic left tuple respectively i The value of the attribute and the synthesized right tuple A i The value of the attribute
[0025] Furthermore, the pseudo label assigned to the new sample in step S22 using the matching dependency MD is expressed as: Where Φ represents the set of matching dependency MDs, φ represents a matching dependency MD, and I(·) is an indicator function, which takes the value of 1 when the independent variable is true and 0 otherwise.
[0026] Furthermore, the loss function is minimized in step S3 The formula for training the model for the target is: where θ * represents the optimal model parameters.
[0027] Compared with the prior art, the advantages of the present invention are:
[0028] Compared with other data augmentation schemes, this invention adopts matching-dependent MD to guide and constrain the data augmentation process. The positive samples generated by this method all meet the constraints of matching-dependent MD, and the negative samples all violate the constraints of matching-dependent MD. Therefore, the pseudo labels of the synthesized samples are interpretable and easy for human users to believe.
[0029] Compared with other machine learning-based entity resolution systems, the present invention can use matching dependency MD to explain the model's prediction results when performing inference prediction after model training is completed, bringing explainability to the prediction results, which is not available in most existing technologies.
[0030] Compared with other data augmentation methods, this invention uses matching-dependent MD to synthesize data, which can generate both positive and negative samples. According to actual application requirements, it can effectively alleviate the problem of sample category imbalance during model training.
[0031] Compared with other entity resolution systems based on machine learning, the present invention has higher classification accuracy in low-resource scenarios with limited labeled training samples, avoiding a large amount of manpower for labeling data without causing a significant drop in accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0033] Figure 1 It is an example relational table used for entity resolution in the prior art.
[0034] Figure 2 This is a schematic diagram of the principle of the reliable low-resource entity resolution method based on rule enhancement of the present invention.
[0035] Figure 3 It is the task performance of different entity resolution methods of the present invention without / with the present invention. DETAILED DESCRIPTION
[0036] The preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more precise definition of the protection scope of the present invention.
[0037] See Figure 2 As shown, this embodiment discloses a reliable low-resource entity resolution method based on rule enhancement, comprising the following steps:
[0038] Step S1: Extract matching dependency MD from a given dataset.
[0039] In this embodiment, step S1 specifically includes the following steps:
[0040] Step S11: merge the given relational data tables R and R′ into a unified relational data table R″. Each record in R″ represents a tuple pair, marked as matched (1) or unmatched (0). This preprocessing ensures the structuring of the input data for use in subsequent steps, thereby achieving effective similarity evaluation and matching dependency MD.
[0041] Step S12: Generate candidate MDs by identifying attribute-level dependencies in the relational data table R″.
[0042] For each attribute pair (A i ,Aj ), calculate the similarity score based on their respective embedding vectors. Taking cosine similarity as an example, its similarity function is defined as Sim(t i· A m ,t j· A m )=cos(embedding(t i· A m ),embedding(t j· A m )).
[0043] Based on the similarity score, by applying a threshold vector σ=[σ1,σ2,…,σ m ] to construct a candidate MD, and the candidate MD meets the following requirements: {Sim(t i· A1,t j· A1)≥σ1∧…∧Sim(t i· A m ,t j· A m )≥σ m}→Match, which means that if the similarity of a pair of tuples to the corresponding attribute embedding vectors reaches a threshold, the two tuples are considered to be matched.
[0044] Only matching dependency MDs with high support and confidence are retained for further use. These metrics ensure that the selected matching dependency MDs are both generalizable and accurate in defining matches. At the end of this process, a set of high-quality matching dependency MDs is obtained, which serve as interpretable rules for data augmentation and model training.
[0045] Step S2: Enhance the training samples using extraction matching dependent MD to obtain enhanced training samples.
[0046] In this embodiment, step S2 specifically includes the following steps:
[0047] Step S21: Use the labeled training sample set S train The attribute values of each tuple in are used to synthesize new samples.
[0048] Specifically, when synthesizing each tuple pair When any attribute value A i ∈A, randomly select two values from all possible values in the tuple pair, introduce a perturbation to them by adding, deleting or changing characters, and then assign the perturbed values to A of the synthetic left tuple respectively i The value of the attribute and the synthesized right tuple A i The value of the attribute
[0049] Step S22: Use matching dependency MD to assign pseudo labels to the new samples, and filter to obtain a set S of synthetic samples. syn .
[0050] Specifically, the pseudo label assigned to the new sample using the matching dependency MD in step S22 is expressed as: Where Φ represents the set of matching dependency MDs, φ represents a matching dependency MD, and I(·) is an indicator function that takes the value 1 when the independent variable is true and 0 otherwise. This means that if there exists a matching dependency MD in the set of matching dependency MDs such that the similarity of the embedding vectors of the corresponding attribute values of a tuple pair satisfies the respective thresholds of the matching dependency MD, then the pseudo label of this tuple pair is 1, otherwise it is 0. It is worth noting that our data augmentation based on matching dependency MDs can generate both positive and negative samples, thereby providing simultaneous enhancement of high-quality positive and negative samples for low-resource ER tasks.
[0051] Step S23: The set of synthetic samples S syn Compared with the existing labeled training sample set S train Merge to get the enhanced training sample S hybrid , as an enhanced training set for training entity resolution models.
[0052] Specifically, according to the specific application requirements (such as the ratio of positive samples to negative samples), the synthesized tuple pairs are selected. For negative samples, random selection is performed, and for positive samples, tuple pairs that meet the MD constraints of high support and high confidence are given priority. After selection, the synthetic sample set S is obtained. syn , the sample set finally provided for ER model training is S hybrid =S syn ∪S train .
[0053] In this embodiment, MD-based data augmentation has two advantages. First, the pseudo-labels of synthetic samples are annotated based on the high-support, high-confidence MD mined from the data, ensuring the high confidence of the pseudo-labels. Second, the MD used for annotation can be regarded as a first-order logic formula, making the rules for annotating synthetic samples clear and easy to understand for users, and more convincing to users.
[0054] Step S3: Minimize the loss function based on the enhanced training samples In order to train the model and obtain the optimal model parameters, the present invention introduces rules into the entity resolution method based on machine learning and artificial intelligence. To implement this entity resolution method, it is necessary to train the artificial intelligence model to optimize the model parameters.
[0055] Among them, minimizing the loss function The formula for training the model for the target is: Where θ represents the optimal model parameters.
[0056] The following shows the task performance of the present invention compared with other existing entity resolution methods on the public entity resolution dataset. Figure 3 As shown, the F1 score is used as an indicator. AG, DA, DS, WA, and AB represent the Amazon-Google dataset, the DBLP-ACM dataset, the DBLP-Scholar dataset, the Walmart-Amazon dataset, and the Abt-Buy dataset, respectively. For each method, their task performance is tested without and with the present invention. It can be seen that whether it is other existing entity resolution methods or methods that only use pre-trained language models, the average F1 scores on multiple datasets are improved after combining the present invention, proving the effectiveness of the present invention in entity resolution tasks.
[0057] The present invention proposes to use matching-dependent MD for data enhancement in entity resolution tasks, which can effectively constrain the data enhancement process and generate highly reliable synthetic samples, thereby significantly improving the task performance of the entity resolution system in low-resource scenarios and alleviating the problem of sample category imbalance during model training. It also combines matching-dependent MD with the entity resolution method based on machine learning, which not only improves the accuracy but also gives the entity resolution system a certain degree of interpretability, alleviating the black box problem of the machine learning method to a certain extent, so that the entity resolution system has both the high accuracy based on the machine learning method and the interpretability based on the rule-based method.
[0058] Although the embodiments of the present invention are described in conjunction with the accompanying drawings, the patent owner may make various changes or modifications within the scope of the appended claims. As long as they do not exceed the scope of protection described in the claims of the present invention, they should be within the scope of protection of the present invention.
Claims
1. A reliable low-resource entity resolution method based on rule enhancement, characterized in that: The following steps are involved: S1. Extract matching dependency MD from a given dataset; S2, using extraction matching dependent MD to enhance the training sample to obtain the enhanced training sample; S3, based on the enhanced training samples, to minimize the loss function Train the model for the target and obtain the optimal model parameters.
2. The reliable low-resource entity resolution method based on rule enhancement according to claim 1 is characterized in that: The step S1 specifically includes the following steps: S11. Merge the given relational data tables R and R′ into a unified relational data table R″, where each record in R″ represents a tuple pair, marked as matched (1) or unmatched (0); S12. Generate candidate MDs by identifying attribute-level dependencies for the relational data table R″.
3. The reliable low-resource entity resolution method based on rule enhancement according to claim 1 is characterized in that: The step S12 specifically includes: For each attribute pair (A i ,A j ), calculate the similarity score based on their respective embedding vectors; Based on the similarity score, by applying a threshold vector σ=[σ1,σ2,…,σ m ] to construct a candidate MD, and the candidate MD meets the following requirements: It means that if the similarity of the corresponding attribute embedding vectors of a pair of tuples reaches a threshold, the two tuples are considered to be matched.
4. The reliable low-resource entity resolution method based on rule enhancement according to claim 1 is characterized in that: The step S2 specifically includes the following steps: S21, using labeled training sample set S train Synthesize new samples from the attribute values of each tuple in ; S22, using matching dependency MD to assign pseudo labels to the new samples, and screen to obtain a set S of synthetic samples syn ; S23, the set of synthetic samples S syn Compared with the existing labeled training sample set S train Merge to get the enhanced training sample S hybrid , as an enhanced training set for training entity resolution models.
5. The reliable low-resource entity resolution method based on rule enhancement according to claim 4 is characterized in that: The step S21 is to synthesize each tuple pair When any attribute value A i ∈A, randomly select two values from all possible values in the tuple pair, introduce a perturbation to them by adding, deleting or changing characters, and then assign the perturbed values to A of the synthetic left tuple respectively i The value of the attribute and the synthesized right tuple A i The value of the attribute 6. The reliable low-resource entity resolution method based on rule enhancement according to claim 4 is characterized in that: The pseudo label assigned to the new sample in step S22 using the matching dependency MD is expressed as: Where Φ represents the set of matching dependency MDs, φ represents a matching dependency MD, and I(·) is an indicator function, which takes the value of 1 when the independent variable is true and 0 otherwise.
7. The reliable low-resource entity resolution method based on rule enhancement according to claim 4 is characterized in that: In step S3, the loss function is minimized The formula for training the model for the target is: where θ * represents the optimal model parameters.