Reliable low-resource entity analysis method based on rule enhancement

By matching MD-guided data augmentation and model training, the problem of insufficient interpretability and accuracy of entity resolution methods is solved, especially under low resource conditions, the performance of entity resolution systems is significantly improved.

CN120524952APending Publication Date: 2025-08-22NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510614297.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The existing learning-based entity analytics lack interpretability, the cost of labeling training data is high, and the unbalanced training data leads to insufficient model accuracy, especially in low-resource scenarios.

Method used

采用匹配依赖MD指导数据增强,通过合成高质量的伪标签样本,生成满足匹配依赖MD的正样本和违反MD的负样本,结合最小化损失函数训练模型参数,提升模型在低资源条件下的准确性和可解释性。

Benefits of technology

Improve the accuracy of entity resolution under low resource conditions, while giving model interpretability, alleviating the problem of sample category imbalance, generating trusted pseudo-labels, and improving the F1 score performance of multiple data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120524952A_ABST
    Figure CN120524952A_ABST
Patent Text Reader

Abstract

The invention discloses a reliable low-resource entity analysis method based on rule enhancement. The reliable low-resource entity analysis method comprises the following steps: S1, extracting a matching dependency MD from a given data set; s2, enhancing the training sample by adopting extraction matching dependence MD to obtain an enhanced training sample; and S3, according to the enhanced training sample, training the model by taking a minimization loss function # imgabs0 # as a target to obtain an optimal model parameter. According to the method, the matching dependence MD guidance and constraint data enhancement process is adopted, the generated positive samples meet the constraint of the matching dependence MD, and the negative samples violate the constraint of the matching dependence MD, so that the pseudo labels of the synthesized samples have interpretability and are easy to be trusted by human users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of entity resolution technology, and more specifically, to a reliable low-resource entity resolution method based on rule enhancement. Background Art

[0002] Entity Resolution (ER) is the task of determining whether two data entries refer to the same real-world entity. Figure 1 As shown in the figure, in two relational tables containing a large number of tuples describing products, the goal of entity resolution is to identify all pairs of tuples that refer to the same product (e.g., t1 and t4). In some scenarios, entity resolution is also known as entity linking, deduplication, or record linking. Entity resolution is crucial for data cleaning and integration, and has a profound impact on data-intensive downstream applications such as e-commerce, censuses, and corporate recruitment. As an important and enduring problem in the database field, entity resolution has been studied for decades.

[0003] Entity resolution solutions can be mainly divided into learning-based solutions and rule-based solutions. Although rule-based solutions are inherently interpretable, their accuracy is often insufficient compared to learning-based solutions, which makes the latter more popular in the research of this task. However, learning-based solutions have the following disadvantages:

[0004] a. Machine learning solutions, such as deep learning, are black boxes, and their predictions lack interpretability.

[0005] b. Labeled training data requires a large amount of manual labeling, which is very expensive and limits the scale of data annotation and the capabilities of the model.

[0006] c. Since most tuple pairs are unmatched, even with a large number of labeled training samples, the number of positive samples (i.e., matching tuple pairs) is far less than the number of negative samples (i.e., unmatched tuple pairs). The imbalance of training data further hinders the training of high-quality ER models.

[0007] Therefore, it is necessary to provide a reliable low-resource entity resolution method based on rule enhancement. Summary of the Invention

[0008] The purpose of the present invention is to provide a reliable low-resource entity resolution method based on rule enhancement to overcome the defects of the prior art.

[0009] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0010] A reliable low-resource entity resolution method based on rule enhancement includes the following steps:

[0011] S1. Extract matching dependency MD from a given dataset;

[0012] S2, using extraction matching dependent MD to enhance the training sample to obtain the enhanced training sample;

[0013] S3. Based on the enhanced training samples, the model is trained with the goal of minimizing the loss function L to obtain the optimal model parameters.

[0014] Furthermore, the step S1 specifically includes the following steps:

[0015] S11. Merge the given relational data tables R and R′ into a unified relational data table R″, where each record in R″ represents a tuple pair, marked as matched (1) or unmatched (0);

[0016] S12. Generate candidate MDs by identifying attribute-level dependencies for the relational data table R″.

[0017] Furthermore, the step S12 specifically includes:

[0018] For each attribute pair (A i ,A j ), calculate the similarity score based on their respective embedding vectors;

[0019] Based on the similarity score, by applying a threshold vector σ=[σ1,σ2,…,σ m ] to construct a candidate MD, and the candidate MD meets the following requirements: {Sim(t i· A1,t j· A1)≥σ1∧…∧Sim(t i· A m ,t j· A m )≥σ m}→Match, which means that if the similarity of a pair of tuples to the corresponding attribute embedding vectors reaches a threshold, the two tuples are considered to be matched.

[0020] Furthermore, the step S2 specifically includes the following steps:

[0021] S21, using labeled training sample set S train Synthesize new samples from the attribute values ​​of each tuple in ;

[0022] S22, using matching dependency MD to assign pseudo labels to the new samples, and screen to obtain a set S of synthetic samples syn ;

[0023] S23, the set of synthetic samples S syn Compared with the existing labeled training sample set S train Merge to get the enhanced training sample S hybrid , as an enhanced training set for training entity resolution models.

[0024] Furthermore, the step S21 synthesizes each tuple pair When any attribute value A i ∈A, randomly select two values ​​from all possible values ​​in the tuple pair, introduce a perturbation to them by adding, deleting or changing characters, and then assign the perturbed values ​​to A of the synthetic left tuple respectively i The value of the attribute and the synthesized right tuple A i The value of the attribute

[0025] Furthermore, the pseudo label assigned to the new sample in step S22 using the matching dependency MD is expressed as: Where Φ represents the set of matching dependency MDs, φ represents a matching dependency MD, and I(·) is an indicator function, which takes the value of 1 when the independent variable is true and 0 otherwise.

[0026] Furthermore, the loss function is minimized in step S3 The formula for training the model for the target is: where θ * represents the optimal model parameters.

[0027] Compared with the prior art, the advantages of the present invention are:

[0028] Compared with other data augmentation schemes, this invention adopts matching-dependent MD to guide and constrain the data augmentation process. The positive samples generated by this method all meet the constraints of matching-dependent MD, and the negative samples all violate the constraints of matching-dependent MD. Therefore, the pseudo labels of the synthesized samples are interpretable and easy for human users to believe.

[0029] Compared with other machine learning-based entity resolution systems, the present invention can use matching dependency MD to explain the model's prediction results when performing inference prediction after model training is completed, bringing explainability to the prediction results, which is not available in most existing technologies.

[0030] Compared with other data augmentation methods, this invention uses matching-dependent MD to synthesize data, which can generate both positive and negative samples. According to actual application requirements, it can effectively alleviate the problem of sample category imbalance during model training.

[0031] Compared with other entity resolution systems based on machine learning, the present invention has higher classification accuracy in low-resource scenarios with limited labeled training samples, avoiding a large amount of manpower for labeling data without causing a significant drop in accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0033] Figure 1 It is an example relational table used for entity resolution in the prior art.

[0034] Figure 2 This is a schematic diagram of the principle of the reliable low-resource entity resolution method based on rule enhancement of the present invention.

[0035] Figure 3 It is the task performance of different entity resolution methods of the present invention without / with the present invention. DETAILED DESCRIPTION

[0036] The preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more precise definition of the protection scope of the present invention.

[0037] See Figure 2 As shown, this embodiment discloses a reliable low-resource entity resolution method based on rule enhancement, comprising the following steps:

[0038] Step S1: Extract matching dependency MD from a given dataset.

[0039] In this embodiment, step S1 specifically includes the following steps:

[0040] Step S11: merge the given relational data tables R and R′ into a unified relational data table R″. Each record in R″ represents a tuple pair, marked as matched (1) or unmatched (0). This preprocessing ensures the structuring of the input data for use in subsequent steps, thereby achieving effective similarity evaluation and matching dependency MD.

[0041] Step S12: Generate candidate MDs by identifying attribute-level dependencies in the relational data table R″.

[0042] For each attribute pair (A i ,Aj ), calculate the similarity score based on their respective embedding vectors. Taking cosine similarity as an example, its similarity function is defined as Sim(t i· A m ,t j· A m )=cos(embedding(t i· A m ),embedding(t j· A m )).

[0043] Based on the similarity score, by applying a threshold vector σ=[σ1,σ2,…,σ m ] to construct a candidate MD, and the candidate MD meets the following requirements: {Sim(t i· A1,t j· A1)≥σ1∧…∧Sim(t i· A m ,t j· A m )≥σ m}→Match, which means that if the similarity of a pair of tuples to the corresponding attribute embedding vectors reaches a threshold, the two tuples are considered to be matched.

[0044] Only matching dependency MDs with high support and confidence are retained for further use. These metrics ensure that the selected matching dependency MDs are both generalizable and accurate in defining matches. At the end of this process, a set of high-quality matching dependency MDs is obtained, which serve as interpretable rules for data augmentation and model training.

[0045] Step S2: Enhance the training samples using extraction matching dependent MD to obtain enhanced training samples.

[0046] In this embodiment, step S2 specifically includes the following steps:

[0047] Step S21: Use the labeled training sample set S train The attribute values ​​of each tuple in are used to synthesize new samples.

[0048] Specifically, when synthesizing each tuple pair When any attribute value A i ∈A, randomly select two values ​​from all possible values ​​in the tuple pair, introduce a perturbation to them by adding, deleting or changing characters, and then assign the perturbed values ​​to A of the synthetic left tuple respectively i The value of the attribute and the synthesized right tuple A i The value of the attribute

[0049] Step S22: Use matching dependency MD to assign pseudo labels to the new samples, and filter to obtain a set S of synthetic samples. syn .

[0050] Specifically, the pseudo label assigned to the new sample using the matching dependency MD in step S22 is expressed as: Where Φ represents the set of matching dependency MDs, φ represents a matching dependency MD, and I(·) is an indicator function that takes the value 1 when the independent variable is true and 0 otherwise. This means that if there exists a matching dependency MD in the set of matching dependency MDs such that the similarity of the embedding vectors of the corresponding attribute values ​​of a tuple pair satisfies the respective thresholds of the matching dependency MD, then the pseudo label of this tuple pair is 1, otherwise it is 0. It is worth noting that our data augmentation based on matching dependency MDs can generate both positive and negative samples, thereby providing simultaneous enhancement of high-quality positive and negative samples for low-resource ER tasks.

[0051] Step S23: The set of synthetic samples S syn Compared with the existing labeled training sample set S train Merge to get the enhanced training sample S hybrid , as an enhanced training set for training entity resolution models.

[0052] Specifically, according to the specific application requirements (such as the ratio of positive samples to negative samples), the synthesized tuple pairs are selected. For negative samples, random selection is performed, and for positive samples, tuple pairs that meet the MD constraints of high support and high confidence are given priority. After selection, the synthetic sample set S is obtained. syn , the sample set finally provided for ER model training is S hybrid =S syn ∪S train .

[0053] In this embodiment, MD-based data augmentation has two advantages. First, the pseudo-labels of synthetic samples are annotated based on the high-support, high-confidence MD mined from the data, ensuring the high confidence of the pseudo-labels. Second, the MD used for annotation can be regarded as a first-order logic formula, making the rules for annotating synthetic samples clear and easy to understand for users, and more convincing to users.

[0054] Step S3: Minimize the loss function based on the enhanced training samples In order to train the model and obtain the optimal model parameters, the present invention introduces rules into the entity resolution method based on machine learning and artificial intelligence. To implement this entity resolution method, it is necessary to train the artificial intelligence model to optimize the model parameters.

[0055] Among them, minimizing the loss function The formula for training the model for the target is: Where θ represents the optimal model parameters.

[0056] The following shows the task performance of the present invention compared with other existing entity resolution methods on the public entity resolution dataset. Figure 3 As shown, the F1 score is used as an indicator. AG, DA, DS, WA, and AB represent the Amazon-Google dataset, the DBLP-ACM dataset, the DBLP-Scholar dataset, the Walmart-Amazon dataset, and the Abt-Buy dataset, respectively. For each method, their task performance is tested without and with the present invention. It can be seen that whether it is other existing entity resolution methods or methods that only use pre-trained language models, the average F1 scores on multiple datasets are improved after combining the present invention, proving the effectiveness of the present invention in entity resolution tasks.

[0057] The present invention proposes to use matching-dependent MD for data enhancement in entity resolution tasks, which can effectively constrain the data enhancement process and generate highly reliable synthetic samples, thereby significantly improving the task performance of the entity resolution system in low-resource scenarios and alleviating the problem of sample category imbalance during model training. It also combines matching-dependent MD with the entity resolution method based on machine learning, which not only improves the accuracy but also gives the entity resolution system a certain degree of interpretability, alleviating the black box problem of the machine learning method to a certain extent, so that the entity resolution system has both the high accuracy based on the machine learning method and the interpretability based on the rule-based method.

[0058] Although the embodiments of the present invention are described in conjunction with the accompanying drawings, the patent owner may make various changes or modifications within the scope of the appended claims. As long as they do not exceed the scope of protection described in the claims of the present invention, they should be within the scope of protection of the present invention.

Claims

1. A reliable low-resource entity resolution method based on rule enhancement, characterized in that: The following steps are involved: S1. Extract matching dependency MD from a given dataset; S2, using extraction matching dependent MD to enhance the training sample to obtain the enhanced training sample; S3, based on the enhanced training samples, to minimize the loss function Train the model for the target and obtain the optimal model parameters.

2. The reliable low-resource entity resolution method based on rule enhancement according to claim 1 is characterized in that: The step S1 specifically includes the following steps: S11. Merge the given relational data tables R and R′ into a unified relational data table R″, where each record in R″ represents a tuple pair, marked as matched (1) or unmatched (0); S12. Generate candidate MDs by identifying attribute-level dependencies for the relational data table R″.

3. The reliable low-resource entity resolution method based on rule enhancement according to claim 1 is characterized in that: The step S12 specifically includes: For each attribute pair (A i ,A j ), calculate the similarity score based on their respective embedding vectors; Based on the similarity score, by applying a threshold vector σ=[σ1,σ2,…,σ m ] to construct a candidate MD, and the candidate MD meets the following requirements: It means that if the similarity of the corresponding attribute embedding vectors of a pair of tuples reaches a threshold, the two tuples are considered to be matched.

4. The reliable low-resource entity resolution method based on rule enhancement according to claim 1 is characterized in that: The step S2 specifically includes the following steps: S21, using labeled training sample set S train Synthesize new samples from the attribute values ​​of each tuple in ; S22, using matching dependency MD to assign pseudo labels to the new samples, and screen to obtain a set S of synthetic samples syn ; S23, the set of synthetic samples S syn Compared with the existing labeled training sample set S train Merge to get the enhanced training sample S hybrid , as an enhanced training set for training entity resolution models.

5. The reliable low-resource entity resolution method based on rule enhancement according to claim 4 is characterized in that: The step S21 is to synthesize each tuple pair When any attribute value A i ∈A, randomly select two values ​​from all possible values ​​in the tuple pair, introduce a perturbation to them by adding, deleting or changing characters, and then assign the perturbed values ​​to A of the synthetic left tuple respectively i The value of the attribute and the synthesized right tuple A i The value of the attribute 6. The reliable low-resource entity resolution method based on rule enhancement according to claim 4 is characterized in that: The pseudo label assigned to the new sample in step S22 using the matching dependency MD is expressed as: Where Φ represents the set of matching dependency MDs, φ represents a matching dependency MD, and I(·) is an indicator function, which takes the value of 1 when the independent variable is true and 0 otherwise.

7. The reliable low-resource entity resolution method based on rule enhancement according to claim 4 is characterized in that: In step S3, the loss function is minimized The formula for training the model for the target is: where θ * represents the optimal model parameters.