Comparison learning entity matching method and system based on similar samples
By constructing a comparative learning method for similar samples, the problem of difficulty in capturing subtle differences between entities in the prior art is solved, and more accurate entity matching and stronger generalization capabilities are achieved, especially in the recognition of similar entities.
Patent Information
- Application Number
- CN202510539497.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-26
AI Technical Summary
Existing entity matching methods are difficult to effectively capture key subtle differences between similar but mismatched entities, resulting in poor misjudgment and generalization, especially in the case of unbalanced data categories.
Using a comparison learning method based on similar samples, by constructing positive entity-to-sample, negative entity-to-sample and similar entity-to-sample, using the contrast learning mechanism to train the model, learn the similarity and differences between entities, introduce similar samples to pay attention to subtle differences, and amplify these differences through interactive attention mechanisms, set the best threshold to improve prediction accuracy.
More accurate entity matching is achieved, which can effectively distinguish similar samples, improve the generalization ability and prediction accuracy of the model on the data set, and especially perform better than the baseline method in the recognition of similar entities.
Smart Images

Figure CN120541536A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data mining, and in particular to a method and system for entity matching based on comparative learning of similar samples. Background Art
[0002] With the advent of the big data era, data sources are constantly enriching, and the amount of data from different sources is exploding. Data from different sources often contain heterogeneous, redundant, and fragmented information about the same entity. Therefore, how to construct more accurate and comprehensive entity information from massive amounts of data has attracted widespread attention. Entity matching aims to find the same entity representing the real world in data from different sources. It is crucial for reducing data redundancy and improving data quality, and is a key foundational technology for data fusion. Currently, it is widely used in areas such as user profiling, information retrieval, and cross-platform user alignment, and has important practical significance.
[0003] Early approaches primarily included rule-based, crowdsourcing-based, and machine learning-based methods. These methods relied on manual intervention, were unsuitable for datasets with noisy data or missing values, and exhibited poor generalization. With the development of deep learning, methods based on similarity feature comparison, such as Seq2SeqMatcher, DITTO, JointMatcher, and EMTransformer, modeled the entity matching task as a binary classification problem based on a cross-entropy loss by training deep learning models end-to-end. Embedding techniques were used to convert the two entities to be matched into vectors. Entity similarity features were then calculated using different vector interaction strategies, and these features were then mapped to match probabilities. If the probability was greater than a set threshold (e.g., 0.5), the two entities were considered a match; otherwise, they were considered mismatched. However, methods based on similarity feature comparison struggled to effectively capture important subtle differences, leading to the misclassification of similar but non-matching entities as matches. Because similar entities are often highly similar, or even nearly identical, except for a few minor differences, existing methods failed to adequately address these subtle differences, often resulting in high overall similarity scores and inaccurate predictions. For example, the "Corsair Vengeance LPX Black 32GB (2x16GB) DDR4 PC4-17000 2133MHz Dual Channel Kit"@en 2x16GB|CMK32GX4M2A2133C13 Novatech"@en" and the "Corsair Vengeance LPX Black 32GB (2x16GB) DDR4 PC4-21300 2666MHz Dual Channel Kit"@en 2x16GB|CMK32GX4M2A2666C16 Novatech"@en" are two different memory modules manufactured by Corsair. Aside from the operating frequencies (17000 vs. 21300), data transfer speeds (2133 vs. 2666), and corresponding model numbers, all other information is identical, but they are not the same entity.
[0004] In addition, existing methods use a cross-entropy loss model to classify prediction results into matching (positive class) and mismatching (negative class) decision boundaries, categorizing similar but mismatching entities and dissimilar and mismatching entities into the negative class, ignoring the differences between these two types of entities. Furthermore, because similar entities are very close to each other in the feature space, it is difficult for the model to learn an accurate decision boundary, making it difficult to effectively distinguish similar entities. For a given entity, there are also differences between different entities that do not match it. Furthermore, when data categories are unbalanced, predictions tend to favor a larger number of categories, resulting in poor recognition of minority categories. Effectively learning sample similarity and difference features so that the model can capture and detect key subtle differences to distinguish similar samples is a challenge. Summary of the Invention
[0005] To this end, the present invention provides a comparative learning entity matching method and system based on similar samples to solve the problem that existing entity matching is difficult to learn similar but not matching key subtle differences between entities, resulting in unsatisfactory entity matching accuracy.
[0006] According to the design scheme provided by the present invention, on the one hand, a comparative learning entity matching method based on similar samples is provided, comprising:
[0007] Obtain the entity data set to be matched and serialize the entities to be matched in the data set;
[0008] The serialized representation of the entity to be matched is input into the entity matching model, and the entity matching model is used to obtain the matching results of each entity pair in the entity data set to be matched. The entity matching model is based on positive entity pair samples, negative entity pair samples and similar entity pair samples and is trained using a contrastive learning mechanism to enable the model to learn the similarities and differences between different entities, wherein the similar entity pair samples are entity sample data whose entity pairs are similar but not matched.
[0009] As the comparative learning entity matching method based on similar samples of the present invention, further, the entities to be matched in the data set are serialized and represented, including:
[0010] For the attributes in the entity and the attribute values corresponding to each attribute, the entity is serialized and converted by adding marking tags to obtain the serialized representation corresponding to the entity. The marking tags include attribute tags and attribute value tags.
[0011] As the comparative learning entity matching method based on similar samples of the present invention, further, the construction process of positive entity pair samples, negative entity pair samples and similar entity pair samples includes:
[0012] For entity sample raw data with positive and negative sample labels, an embedding vector is generated for each entity to represent the entity's semantic features. The entity sample raw data includes entity pairs from different data sources and annotated with positive and negative sample labels, where the positive sample label is used to indicate a matching entity pair and the negative sample label is used to indicate an unmatched entity pair.
[0013] Calculating the similarity between different entities in the original data based on the embedding vectors, and constructing a similarity matrix using the similarity, wherein each element in the similarity matrix represents the similarity value between corresponding entity pairs;
[0014] The average similarity value corresponding to the positive sample label in the original data is calculated based on the similarity value of the entity pair in the similarity matrix, and the similarity threshold is set based on the average value. The entity pairs with similarity values greater than the similarity threshold but without positive sample labels are regarded as similar entity pair samples.
[0015] As the comparative learning entity matching method based on similar samples of the present invention, the construction process of positive entity pair samples, negative entity pair samples and similar entity pair samples further includes:
[0016] Calculate the average similarity value corresponding to the negative sample label in the original data based on the similarity value of the entity pair in the similarity matrix, and set the negative sample similarity threshold based on the similarity average value;
[0017] For each entity in the original data, exclude the corresponding entity pairs with positive sample labels in the original data, select entity pairs with similarity values less than the negative sample similarity threshold as negative sample entity pairs based on the entity pair similarity values in the similarity matrix, and add them to the negative entity pair sample.
[0018] As the comparative learning entity matching method based on similar samples of the present invention, the model is further trained using the comparative learning mechanism, including:
[0019] Construct a contrastive learning loss function based on the cosine similarity of entity pairs;
[0020] Based on the contrastive learning loss function and using positive entity pair samples, negative entity pair samples and similar entity pair samples to train the entity matching model, the model learns the similarities and differences between positive entity pair samples, negative entity pair samples and similar entity pair samples, and determines the optimal threshold of the model prediction output, so as to maximize the probability of the model prediction output accuracy using the optimal threshold.
[0021] As the contrastive learning entity matching method based on similar samples of the present invention, further, the contrastive learning loss function is expressed as Among them, cos() represents cosine similarity, label(j, k) and label(x, y) represent entity pairs e respectively. j 、ek and entity pair e x 、e y The sample labels of both, λ is a hyperparameter.
[0022] As the comparative learning entity matching method based on similar samples of the present invention, further, the entity matching model is used to predict and output the matching results of each entity pair, including:
[0023] Using the pre-trained language model to obtain the entity serialization representation embedding vector of the entity pair to be matched, and merging the entity serialization representation embedding vectors of the entity pair to be matched to obtain a merged embedding vector;
[0024] Merge the embedding vectors of the entity serialization representations of the entity pairs to be matched respectively and perform vector fusion to obtain the fusion vector of the entity pairs to be matched;
[0025] The attention distribution matrix of the fusion vector is obtained based on the entity interaction direction. The entity interaction attention expression is obtained using the attention distribution matrix, and the expression is fused with the merged embedding vector to obtain the semantically relevant perceptual embedding of the entity to be matched;
[0026] Extract the similarity of the perceptual embeddings of the entity pairs to be matched, and output the matching results of the entity pairs to be matched based on the similarity prediction.
[0027] On the other hand, the present invention also provides a comparative learning entity matching system based on similar samples, comprising: a sequence conversion module and an entity matching module, wherein:
[0028] The sequence conversion module is used to obtain the entity data set to be matched and serialize the entities to be matched in the data set;
[0029] The entity matching module is used to input the serialized representation of the entity to be matched into the entity matching model, and use the entity matching model to obtain the matching results of each entity pair in the entity data set to be matched. The entity matching model is based on positive entity pair samples, negative entity pair samples and similar entity pair samples and uses a contrastive learning mechanism to train the model so that the model learns the similarities and differences between different entities, wherein the similar entity pair samples are entity sample data with similar but unmatched entity pairs.
[0030] Beneficial effects of the present invention:
[0031] This method uses similar but mismatched entities as similar samples, providing a more comprehensive, high-quality set of comparison samples for the contrastive learning process. This method accurately learns the similarities and differences between entity pairs—positive, negative, and similar—through contrastive learning. It also leverages a mutual attention mechanism to amplify subtle differences between similar entities, paying extra attention to similar samples and achieving more accurate entity matching. Further experimental validation on public datasets demonstrates that this approach outperforms baseline methods on datasets with relatively small overall differences, effectively distinguishing similar samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 Schematic diagram of the entity matching process of comparative learning based on similar samples in an embodiment;
[0033] Figure 2 Schematic diagram of the entity matching algorithm framework based on similar samples in the embodiment;
[0034] Figure 3 This is a comparison diagram of similar sample matching in the embodiment;
[0035] Figure 4 Schematic diagram of experimental results under different similarity thresholds in the embodiment;
[0036] Figure 5 Schematic diagram of the ablation experiment results in the embodiment. DETAILED DESCRIPTION
[0037] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention is further described in detail below with reference to the accompanying drawings and technical solutions.
[0038] Entity matching aims to determine whether entity information from different sources refers to the same object in the real world, which is crucial for downstream applications such as data fusion and information retrieval. Given a set T of entity pairs to be matched, for each entity pair (e i ,e j )∈T, each entity e i They are all composed of key-value pairs {(att1,val1),(att2,val2),…,(att k ,val k )}, where att k and val k are the kth attribute name and attribute value of the entity respectively. Output: a set N of entity pair matching results, each of which represents the current entity pair (e i ,e j ) refer to the same entity in the real world.
[0039] Existing methods based on comparison aggregation and pre-trained models both model the entity matching task as a binary classification problem, mapping entity pair matching or mismatching to positive (1) or negative (0). This causes the model to classify similar but mismatching entities and dissimilar and mismatching entities as negative classes for learning, causing the model to ignore the differences between these two entities during the learning process. For a single entity, there may be many mismatching entities, and these different entities also have differences. However, existing methods do not consider the inherent differences between these mismatching entities.
[0040] In entity matching tasks, subtle semantic differences between entities can significantly impact matching results. Existing approaches based on comparative aggregation and pre-trained models struggle to accurately capture these subtle semantic differences when the entities are similar overall, and struggle to distinguish between similar but mismatched entities.
[0041] For this purpose, the embodiments of the present invention, see Figure 1 As shown, a comparative learning entity matching method based on similar samples is provided, including:
[0042] S101: Obtain a dataset of entities to be matched, and perform serialization representation on the entities to be matched in the dataset.
[0043] Specifically, the serialized representation of the entities to be matched in the dataset can be designed to include:
[0044] For the attributes in the entity and the attribute values corresponding to each attribute, the entity is serialized and converted by adding marking tags to obtain the serialized representation corresponding to the entity. The marking tags include attribute tags and attribute value tags.
[0045] For the entity to be matched e i and e j , the following serialization conversion formula can be used to achieve sequence conversion,
[0046] serialize(x)=[COL]attr1 [VAL]val1 ...[COL]attr k [VAL]val k (1)
[0047] x is an entity object with multiple attributes attr1 through attrk. Each attribute has a corresponding value, val1, or valk. [COL] and [VAL] are the corresponding attribute name and value tags. For example, if entity x has an object with the attributes name and age, and the value of name is John and the value of age is 20, the result of serialize(x) can be expressed as [COL]name[VAL]John[COL]age[VAL]20.
[0048] S102. Input the serialized representation of the entity to be matched into the entity matching model, and use the entity matching model to obtain the matching results of each entity pair in the entity data set to be matched. The entity matching model is based on positive entity pair samples, negative entity pair samples and similar entity pair samples and is trained using a contrastive learning mechanism to enable the model to learn the similarities and differences between different entities, wherein the similar entity pair samples are entity sample data whose entity pairs are similar but not matched.
[0049] The construction process of positive entity pair samples, negative entity pair samples, and similar entity pair samples includes:
[0050] For entity sample raw data with positive and negative sample labels, an embedding vector is generated for each entity to represent the entity's semantic features. The entity sample raw data includes entity pairs from different data sources and annotated with positive and negative sample labels, where the positive sample label is used to indicate a matching entity pair and the negative sample label is used to indicate an unmatched entity pair.
[0051] Calculating the similarity between different entities in the original data based on the embedding vectors, and constructing a similarity matrix using the similarity, wherein each element in the similarity matrix represents the similarity value between corresponding entity pairs;
[0052] The average similarity value corresponding to the positive sample label in the original data is calculated based on the similarity value of the entity pair in the similarity matrix, and the similarity threshold is set based on the average value. The entity pairs with similarity values greater than the similarity threshold but without positive sample labels are regarded as similar entity pair samples.
[0053] A key challenge in training contrastive learning models is how to construct high-quality positive and negative samples. In model training, incorrect labels will mislead the judgment of the entity matching model and lead to incorrect results, so high-quality training samples are crucial. Because in the entity matching task, changing only one number may cause the originally matched entities to become mismatched, even if the rest of their information is the same and they are very similar overall. For example, "AMD Ryzen Threadripper 1950X Box" and "AMD Ryzen Threadripper 1900X Box" represent two different products. In practice, most entities that are similar to each other do not refer to the same entity, and it is often difficult for the model to learn useful difference information about such entities. Therefore, in the embodiment of this case, it is considered to construct such entities separately as a type of sample, namely "similar samples", to provide effective information for subsequent multi-label contrastive learning. Similar samples refer to a type of sample in the contrastive learning process, in which the two entities are similar overall but do not match.
[0054] The automatic sample pair generation strategy aims to more effectively utilize training data and identify "similar samples" for additional attention in subsequent comparative learning. It consists of two components: ① a similar sample construction method and ② a negative sample construction method. Leveraging the powerful semantic expression capabilities of the pre-trained language model (RoBERTa, a BERT variant), we generate an embedding vector for each entity, representing its semantic information.
[0055] Different similarity functions (such as cosine distance, Euclidean distance, and Manhattan distance) are commonly used to quantify the semantic similarity between entities in a dataset. Cosine similarity effectively ignores differences in text length and focuses solely on their content structure, ensuring that the similarity measure is primarily based on the key information or themes contained, rather than being overly influenced by word count. Therefore, cosine distance was chosen as the similarity function. An entity similarity matrix is constructed between all T entities. Similar but mismatched "similar samples" are constructed using similarity functions and generative algorithms, and negative samples are additionally constructed to fully utilize the training data.
[0056] Given a list containing entities to be matched (e i ,e j ), first for each entity e i Generate embedding vector V(e i ) represents the semantic features of the entity. i ,e j ) to calculate the similarity and obtain a similarity matrix M, where each element M ij Represents entity e i and entity e j According to the generation algorithm, some M entities are selected from the unmatched entity pairs. ij Entities with high similarity will correspond to entity e i and e j Constructed into similar sample pairs.
[0057] The goal of similar sample generation is to find entity pairs with a high matching probability in the original negative samples and construct them as similar samples. A common approach is to consider the semantic similarity between entities and set a threshold. Similar label generation provides additional constraints compared to similarity-based methods to ensure high quality similar labels.
[0058] Specifically, for (e i ,e j ), e i 、e j From different data sources A and B. For each entity e i 、e j , and obtain its embedding vector V(e i )V(e j), capturing the semantic features of entities. Calculate the similarity between entities from different sources to build a similarity matrix M. Each element M in M ij Represents entity e i and entity e j The similarity value between them. Alignment matrix M n×n Each row represents entity e i and all entities e in data source B j Similarity score. For each row element M i,j Sort the similarity list and select the top 3 entities with the highest similarity. Since the matching entity pair information is known (supervised), first calculate the average similarity of all positive sample pairs Then, a threshold t is set. According to the sorted similarity list, among the entities whose similarity is greater than the threshold, the positive samples are excluded according to the annotation set, and ej and ei are constructed as similar sample pairs.
[0059] In the process of model training, the selection of negative samples is crucial. If the negative samples are too simple and easy to distinguish, the model may fail to fully learn complex features, resulting in poor performance on new data. According to the entity similarity matrix M, since the similarity sample generation process is too simple for the entity e i , its element negative samples may be constructed as similar samples due to their high similarity, which may lead to entity e i There is no negative sample pair involved in the subsequent training process. Therefore, negative samples are constructed by generating negative labels to ensure that each entity e i There are negative samples. We can intuitively get the entity e i Different entities with similarity less than the threshold are constructed as new negative sample pairs.
[0060] Specifically, the average similarity value corresponding to the negative sample label in the original data can be calculated based on the similarity value of the entity pairs in the similarity matrix, and the negative sample similarity threshold can be set based on the similarity average value; for each entity in the original data, the corresponding entity pairs with positive sample labels in the original data are excluded, and the entity pairs with similarity values less than the negative sample similarity threshold are selected as negative sample entity pairs based on the similarity values of the entity pairs in the similarity matrix, and added to the negative entity pair sample.
[0061] By calculating the average similarity of all negative sample pairs As the threshold for selecting the similarity of negative samples, for each entity e i , traverse M i,j Sort the similarity in descending order, exclude positive samples according to the annotation set, and select samples less than the threshold The entity e'j as e i Negative samples are constructed as negative sample pairs, and negative sample pairs are constructed for entities with fewer negative samples. iIf it matches e'j, then entity e'k is randomly selected as a negative sample of ei.
[0062] Among them, the construction algorithm of positive entity pair samples, negative entity pair samples and similar entity pair samples is shown in Algorithm 1.
[0063]
[0064] The description of each symbol in the algorithm is shown in Table 1.
[0065] Table 1 Symbols
[0066]
[0067]
[0068] Using three types of high-quality samples, for each pair of entity feature vectors E i ′、E j ′, through contrastive learning training, for each entity, the distance to its positive sample in the vector space is less than the distance to the similar sample and less than the distance to the negative sample, so that different types of samples are distributed in a more reasonable sample space, thereby learning and distinguishing the differences between different types of samples, especially effectively distinguishing similar samples from positive samples.
[0069] Among them, the contrastive learning mechanism is used to train the model, which can be designed to include:
[0070] Construct a contrastive learning loss function based on the cosine similarity of entity pairs;
[0071] Based on the contrastive learning loss function and using positive entity pair samples, negative entity pair samples and similar entity pair samples to train the entity matching model, the model learns the similarities and differences between positive entity pair samples, negative entity pair samples and similar entity pair samples, and determines the optimal threshold of the model prediction output, so as to maximize the probability of the model prediction output accuracy using the optimal threshold.
[0072] like Figure 2 As shown in the figure, in the SSEM framework of the contrastive learning entity matching algorithm, this solution screens similar samples in the data and constructs three types of sample data sets: positive samples, negative samples, and similar samples; focuses on the main similar fragments between samples, fully extracts the subtle difference features between entities, learns a more reasonable distribution in the vector space through the improved contrast loss, and judges the entity matching results through the cosine similarity distance metric.
[0073] Compared to the cross-entropy loss, which typically treats all mismatched samples as similar, contrastive learning introduces a similarity metric to distinguish different mismatched samples. Contrastive learning models not only learn the similarities between matching samples, but also the differences between mismatched samples, thereby capturing richer inter-sample relationships.
[0074] It can better cope with the diversity and noise between samples, thereby improving the generalization ability of the model.
[0075] Traditional contrastive learning typically only considers the relationship between positive and negative samples. However, contrastive learning methods, in addition to distinguishing between positive and negative samples, also introduce similar samples, controlling the distance between positive samples to be smaller than the distance between similar samples. This further optimizes the learning process and enhances the expressive power of the embedding space. For an entity, the distance between its positive sample in vector space is less than that between its similar samples, and less than that between its negative samples. This allows learning the similarities and differences between different entities, enabling better differentiation between similar entities. The loss function is as follows:
[0076]
[0077] Here, cos() represents cosine similarity, and label() represents the sample label. A dataset with three labels, "match," "similar," and "mismatch," is constructed. The similarity between "matching" sentences is greater than that between two "similar" sentences, and the similarity between two "similar" sentences is greater than that between two "mismatching" sentences. For example, when j = x, if three pairs of entities have label(j,k) > label(j,y) > label(j,z), then for entity j, the distance z in vector space is greater than y > k. λ > 0 is a hyperparameter. This allows the model to robustly learn similarities and differences between entities, making it easier to distinguish between similar samples, especially the subtle differences that are crucial in matching, thereby improving model generalization.
[0078] The algorithm framework utilizes three types of samples for comparative learning: positive, similar, and negative. The entity matching task ultimately involves determining two types of results: matches and mismatches. Therefore, a threshold is set. Entity pairs with similarity greater than the threshold are considered matches, while those with similarity less than the threshold are considered mismatches. By continuously iteratively learning the optimal threshold, the model's prediction accuracy is maximized.
[0079] The Powell method is a direct search method that does not rely on the first-order or second-order derivative information of the objective function. Instead, it approaches the optimal solution through iterative search. In this embodiment, the Powell method can be used to determine the optimal threshold. The specific steps can be summarized as follows:
[0080] 1. Initialization: Select the initial point T0 and the initial direction set D0. Usually, the standard basis vectors can be selected as the initial direction set.
[0081] 2. Linear search: For each direction di∈Dk, perform a linear search to find the step size αi such that L(Tk+αidi) is minimized.
[0082] Update the current point to Tk(i)=Tk+αidi.
[0083] 3. Update the direction set: Calculate the new direction dk+1=Tk(n)-Tk and use it to replace the first direction in the direction set, that is, Dk+1={dk+1,d1,d2,…,dn-1}.
[0084] 4. Update the iteration point: use Tk(n) as the starting point of the next iteration, that is, Tk+1=Tk(n).
[0085] 5. Check the termination condition: If the termination condition is met (such as the function value change is less than the tolerance or the maximum number of iterations is reached),
[0086] Then stop the iteration and output the current point as the optimal solution; otherwise, continue to the next iteration.
[0087] The Powell algorithm iterates to find an optimal threshold. Predictions greater than the threshold are considered matches, while those less than the threshold are mismatches. The predicted results are compared with the true labels to maximize the proportion of samples with correct predictions. The resulting threshold can be used to adjust the prediction probability to improve the accuracy of the matching result classification.
[0088] Specifically, the entity matching model is used to predict and output the matching results of each entity pair, including:
[0089] Using the pre-trained language model to obtain the entity serialization representation embedding vector of the entity pair to be matched, and merging the entity serialization representation embedding vectors of the entity pair to be matched to obtain a merged embedding vector;
[0090] Merge the embedding vectors of the entity serialization representations of the entity pairs to be matched respectively and perform vector fusion to obtain the fusion vector of the entity pairs to be matched;
[0091] The attention distribution matrix of the fusion vector is obtained based on the entity interaction direction. The entity interaction attention expression is obtained using the attention distribution matrix, and the expression is fused with the merged embedding vector to obtain the semantically relevant perceptual embedding of the entity to be matched;
[0092] Extract the similarity of the perceptual embeddings of the entity pairs to be matched, and output the matching results of the entity pairs to be matched based on the similarity prediction.
[0093] For the serialized representation of the entity pair to be matched, input it into the pre-trained language model and obtain e i 、e j The embedding vector E i 、E j , so that the overall semantic attention of each entity is only distributed in its own word vector itself, which makes it easier to capture the semantic core words of the entity. Then use formula (2) to convert the entity e i and e j Merge and serialize them into a whole, and get E after pre-training model t , to capture deeper semantic and similarity features between entities. where l i and l j represents the total length of the text embedded vectorized by the pre-trained model, and d represents the dimension of the embedding. Among them, l t Indicates serialize(e i ,e j ) The total length of the text after . t =l i +l j -1. The sequence merging method is as follows:
[0094] serialize(i,j)=[CLS]serialize(i)[SEP]serialize(j)[SEP] (2)
[0095] In calculating E t The Transformer's self-attention mechanism can cause semantic interactions between entities, potentially leading to information crucial for making matching decisions being assigned a lower weight by the self-attention mechanism and thus overlooked by the model. Therefore, we considered using both embeddings simultaneously to more meticulously utilize the representation of each sentence and flexibly fuse features. Furthermore, we used cross-attention to interact with the two embeddings, focusing more on key and similar segments within entity pairs. This allows the model to prioritize important words while minimizing the need for attention to other words.
[0096] In this embodiment, first, i and E t [CLS] and E j and E t [CLS] performs vector fusion. t [CLS]∈R d , representing the [CLS] token in E t Embedded in. Due to E t [CLS] contains the overall semantics of the entity pair, so this fused vector can directly benefit e i and ej The interactive attention between them. By fusing E i and E t The representation is used to calculate the fusion representation E i ′:
[0097] E′ i =E i +E′ t [CLS]
[0098] where E′ t [CLS]=repeat(E t [CLS],l i ). The repetition operation here is used to expand E t The dimension of [CLS] and E i Similarly, the fusion representation E′ can be obtained v .
[0099] For e i and e j Attention to the interaction between the two directions (e i →e j and e i ←e j ) compare entities and obtain two comparison matrices as matching information, and then execute e i to e j Attention and e j and e i Attention. i to e j Note that E′ is calculated v The attention distribution matrix A is as follows:
[0100]
[0101] in Specifically, repeat(B j ,l j ) means B j The elements of l are repeated j times and expand the dimension. j Then merge with A and add back to E i On, we get the expression C u :
[0102] C u =E u +ATE′ v
[0103] Similarly, through the interactive attention between v and u, we can get the expression C v Finally, by adding Cu and C v Respectively with E t Fusion, we get the relevant perceptual embeddings under two overall semantics:
[0104]
[0105]
[0106] Get the embedded vector after interactive attention and And through maximum pooling, similar features are extracted, focusing on important matching information between entities and amplifying small differences between similar entities.
[0107] Furthermore, based on the above method, an embodiment of the present invention also provides a comparative learning entity matching system based on similar samples, comprising: a sequence conversion module and an entity matching module, wherein:
[0108] The sequence conversion module is used to obtain the entity data set to be matched and serialize the entities to be matched in the data set;
[0109] The entity matching module is used to input the serialized representation of the entity to be matched into the entity matching model, and use the entity matching model to obtain the matching results of each entity pair in the entity data set to be matched. The entity matching model is based on positive entity pair samples, negative entity pair samples and similar entity pair samples and uses a contrastive learning mechanism to train the model so that the model learns the similarities and differences between different entities, wherein the similar entity pair samples are entity sample data with similar but unmatched entity pairs.
[0110] To verify the effectiveness of this solution, the following is a further explanation based on experimental data:
[0111] Experimental analysis was conducted on the WDC public dataset. The WDC datasets primarily consist of four different product categories: Computers, Watches, Cameras, and Shoes. These datasets contain 26 million product descriptions collected from e-commerce websites. These datasets are widely used for entity matching tasks. Furthermore, each dataset is divided into different sizes: Small, Medium, Large, and xLarge. The training, validation, and test sets of each dataset follow a 3:1:1 ratio. Summary statistics of the datasets are shown in Table 2.
[0112] Table 2 Statistics of WDC dataset
[0113]
[0114] The experiments were all implemented in PyTorch 1.8.0 and hugging face transformers, and were performed on a device equipped with an Nvidia V100 GPU.
[0115] 1. Ditto: Fine-tuning the pre-trained LM through three optimizations (i.e., domain knowledge, TF-IDF summary, and data augmentation).
[0116] 2. JointMatcher: Adds relevance-aware encoder and digit-aware encoder to focus on important fragments in entities and better capture digit matching information.
[0117] 3. R-SupCon: Using contrastive learning, first pre-training with supervised contrastive loss, and then fine-tuning based on cross-entropy.
[0118] We use the evaluation metric commonly used in previous entity matching methods: F1 Score. F1 score is a commonly used evaluation metric in machine learning, especially in imbalanced classification problems. It combines the precision (Precision) and recall (Recall) of the model, and can more comprehensively evaluate the performance of the model. In binary classification problems, the predicted category is usually called a positive example (Positive) or a negative example (Negative). Among them, the number of samples that are actually positive and predicted as positive are true positive examples (True Positive, TP), the number of samples that are actually negative but mistakenly predicted as positive are false positive examples (False Positive, FP), the number of samples that are actually negative and predicted as negative are true negative examples (True Negative, TN), and the number of samples that are actually positive but mistakenly predicted as negative are false negative examples (False Negative, FN). This leads to the following evaluation metrics:
[0119] (1)Precision:
[0120]
[0121] Precision, also known as the accuracy rate, measures the proportion of samples predicted as positive that are actually positive. It reflects the reliability of the model's positive predictions. A higher precision rate means a higher proportion of samples predicted as positive that are actually positive, and fewer false positives (FPs).
[0122] (2) Recall:
[0123]
[0124] Recall, also known as recall, measures the proportion of samples predicted as positive among samples that are actually positive. It reflects the model's ability to find all positive examples. A higher recall means the model is more likely to find all positive examples and has fewer false negatives (FN).
[0125] (3) F1 Score:
[0126]
[0127] The results of this solution are compared with Ditto, Jointmatcher, and R-SupCon, as shown in Table 3, which shows the comparison results under different datasets.
[0128] Table 3 Comparison of F1 value results between this solution and existing methods
[0129]
[0130]
[0131] Size represents the different sizes of each dataset. As can be seen in the table, the proposed algorithm, SSEM, performs well on all four datasets. Compared to the results of Ditto and JointMatcher, SSEM outperforms both methods in all cases. Compared to R-SupCon, SSEM also achieves better results in most cases across all four datasets. Specifically, the maximum F1 value increases on datasets of different sizes for the four product categories are: 0.16, 0.05, 0.17, and 1.08.
[0132] In order to verify the effectiveness of similar entity matching, the similarity threshold when constructing the algorithm based on the sample All sample pairs in the data set with similarity above the threshold are screened for testing. Figure 3 The results show the number of mispredictions for similar entities by different methods. While the R-SupCon method lags behind the SSEM method in overall F1 score, the SSEM method achieves the fewest misclassifications in the similar entity test results. This indicates that the model is able to distinguish similar samples to a certain extent. Comparative learning between three types of samples and focusing on similar segments within entity pairs can improve entity matching performance.
[0133] The similarity threshold determines the quantity and quality of similar samples, and the learning of similar samples has an important impact on the performance of the model. According to the average positive sample similarity x, the similarity threshold is set to x-0.1, x-0.05, and x. The experimental results are shown in Figure 2. Figure 4As shown in the figure, the experimental results show that as the threshold increases, the algorithm's F1 value gradually increases. When the similarity threshold is below x-0.1, the model's F1 value remains essentially unchanged or decreases, with only a slight increase on the shoes dataset. Therefore, the model's attribute similarity threshold can be set to x-0.1.
[0134] To verify the effectiveness of this solution, we conducted ablation experiments. Specifically, we compared the overall approach with the removal of some modules. For ease of presentation, we use the terms EN and SG to represent the fine-grained perceptual encoder module and the automatic sample module, respectively. Specifically, EM-EN represents the experimental results when the EN module is removed and only the similar sample generation module (SG) is used; EM-SG represents the experimental results when the SG module is removed and only the fine-grained perceptual encoder is used to encourage the model to focus on similar fragment information. The detailed comparison information is shown in Table 4.
[0135] Table 4 Ablation analysis results
[0136]
[0137]
[0138] Table 4 shows that removing the SG module causes a decrease in F1 scores on datasets of varying sizes, with a particularly significant drop on the Small dataset. This is because the Small dataset provides limited training samples, while the SG module adds additional similar and negative samples, enabling the model to learn effective feature information. As the data size increases and the number of training samples increases, adding the SG module improves the F1 scores on the Large and XLarge datasets, but the increase is less pronounced than on the Small dataset, demonstrating that the SG module is more effective on small datasets.
[0139] Depend on Figure 5 As shown in the figure, the experimental results show that when the data scale is small, the sample generation module significantly improves the F1 value on such datasets, and achieves significant improvements on all four types of datasets; in addition, it can be seen that regardless of the size of the dataset, the fine-grained perceptual encoder improves the F1 value and effectively identifies key similar fragments in entity pairs.
[0140] The above experimental results show that compared with the Ditto and jointMatcher methods, the F1 values of this solution in four different product data sets are improved by 2.25, 0.22, 1.40, and 3.58 respectively. It can achieve relatively accurate entity matching and has good application prospects in user profiling, information retrieval, cross-platform user alignment and other fields.
[0141] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0142] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0143] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.
[0144] Those skilled in the art will appreciate that all or part of the steps in the above method can be performed by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disk. Alternatively, all or part of the steps in the above embodiment can be implemented using one or more integrated circuits. Accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or software functional modules. The present invention is not limited to any specific combination of hardware and software.
[0145] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A comparative learning entity matching method based on similar samples, characterized in that: Include: Obtain the entity data set to be matched and serialize the entities to be matched in the data set; The serialized representation of the entity to be matched is input into the entity matching model, and the entity matching model is used to obtain the matching results of each entity pair in the entity data set to be matched. The entity matching model is based on positive entity pair samples, negative entity pair samples and similar entity pair samples and is trained using a contrastive learning mechanism to enable the model to learn the similarities and differences between different entities, wherein the similar entity pair samples are entity sample data whose entity pairs are similar but not matched.
2. The entity matching method based on comparative learning of similar samples according to claim 1 is characterized in that: Serialize the entities to be matched in the dataset, including: For the attributes in the entity and the attribute values corresponding to each attribute, the entity is serialized and converted by adding marking tags to obtain the serialized representation corresponding to the entity. The marking tags include attribute tags and attribute value tags.
3. The entity matching method based on comparative learning of similar samples according to claim 1, characterized in that: The construction process of positive entity pair samples, negative entity pair samples, and similar entity pair samples includes: For entity sample raw data with positive and negative sample labels, an embedding vector is generated for each entity to represent the entity's semantic features. The entity sample raw data includes entity pairs from different data sources and annotated with positive and negative sample labels, where the positive sample label is used to indicate a matching entity pair and the negative sample label is used to indicate an unmatched entity pair. Calculating the similarity between different entities in the original data based on the embedding vectors, and constructing a similarity matrix using the similarity, wherein each element in the similarity matrix represents the similarity value between corresponding entity pairs; The average similarity value corresponding to the positive sample label in the original data is calculated based on the similarity value of the entity pair in the similarity matrix, and the similarity threshold is set based on the average value. The entity pairs with similarity values greater than the similarity threshold but without positive sample labels are regarded as similar entity pair samples.
4. The entity matching method based on comparative learning of similar samples according to claim 3 is characterized in that: The construction process of positive entity pair samples, negative entity pair samples, and similar entity pair samples also includes: Calculate the average similarity value corresponding to the negative sample label in the original data based on the similarity value of the entity pair in the similarity matrix, and set the negative sample similarity threshold based on the similarity average value; For each entity in the original data, exclude the corresponding entity pairs with positive sample labels in the original data, select entity pairs with similarity values less than the negative sample similarity threshold as negative sample entity pairs based on the entity pair similarity values in the similarity matrix, and add them to the negative entity pair sample.
5. The entity matching method based on comparative learning of similar samples according to claim 1, characterized in that: The model is trained using contrastive learning, including: Construct a contrastive learning loss function based on the cosine similarity of entity pairs; Based on the contrastive learning loss function and using positive entity pair samples, negative entity pair samples and similar entity pair samples to train the entity matching model, the model learns the similarities and differences between positive entity pair samples, negative entity pair samples and similar entity pair samples, and determines the optimal threshold of the model prediction output, so as to maximize the probability of the model prediction output accuracy using the optimal threshold.
6. The entity matching method based on comparative learning of similar samples according to claim 5, characterized in that: The contrastive learning loss function is expressed as Among them, cos() represents cosine similarity, label(j, k) and label(x, y) represent entity pairs e respectively. j 、e k and entity pair e x 、e y The sample labels of both, λ is a hyperparameter.
7. The entity matching method based on comparative learning of similar samples according to claim 1, characterized in that: Use the entity matching model to predict and output the matching results of each entity pair, including: Using the pre-trained language model to obtain the entity serialization representation embedding vector of the entity pair to be matched, and merging the entity serialization representation embedding vectors of the entity pair to be matched to obtain a merged embedding vector; Merge the embedding vectors of the entity serialization representations of the entity pairs to be matched respectively and perform vector fusion to obtain the fusion vector of the entity pairs to be matched; The attention distribution matrix of the fusion vector is obtained based on the entity interaction direction. The entity interaction attention expression is obtained using the attention distribution matrix, and the expression is fused with the merged embedding vector to obtain the semantically relevant perceptual embedding of the entity to be matched; Extract the similarity of the perceptual embeddings of the entity pairs to be matched, and output the matching results of the entity pairs to be matched based on the similarity prediction.
8. A comparative learning entity matching system based on similar samples, characterized in that: Contains: sequence conversion module and entity matching module, among which, The sequence conversion module is used to obtain the entity data set to be matched and serialize the entities to be matched in the data set; The entity matching module is used to input the serialized representation of the entity to be matched into the entity matching model, and use the entity matching model to obtain the matching results of each entity pair in the entity data set to be matched. The entity matching model is based on positive entity pair samples, negative entity pair samples and similar entity pair samples and uses a contrastive learning mechanism to train the model so that the model learns the similarities and differences between different entities, wherein the similar entity pair samples are entity sample data with similar but unmatched entity pairs.
9. An electronic device, characterized in that: include: at least one processor, and a memory coupled to the at least one processor; The memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 can be implemented.