Generalized entity matching method based on self-supervised learning and attribute perception

By adopting self-supervised learning and attribute perception methods in entity matching, label instances are automatically generated and multiple entity information are integrated, and the problem of relying on manual annotation and ignoring structural information in the prior art is solved, and more efficient and accurate entity matching is achieved.

CN120030474APending Publication Date: 2025-05-23NORTHEASTERN UNIV CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510102253.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing entity matching methods rely on a large number of manual annotations, consume labor, and ignore the structural information of entity data, resulting in poor matching performance.

Method used

A generalized entity matching method based on self-supervised learning and attribute perception is proposed. Tag instances are automatically generated through pseudo-label generation strategy, entity name, structure information, attribute information and relationship information are integrated, and the impact of structural information on matching results is improved through the sharing weight mechanism.

Benefits of technology

Reduce the dependence of manual tagging, make full use of information in entity data, improve the accuracy and efficiency of entity matching, and is better than existing baseline models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030474A_ABST
    Figure CN120030474A_ABST
Patent Text Reader

Abstract

The invention provides a generalized entity matching method based on self-supervised learning and attribute perception, and belongs to the technical field of natural language processing, and the method comprises the steps: employing a pseudo tag generation strategy through a self-supervised learning method, and obtaining positive and negative training samples of entity matching; a shared weight mechanism is adopted to improve the embedding learning process of an entity, that is, attribute type embedding and attribute value embedding corresponding to attribute types are cascaded, so that attribute types obtained by attribute value sharing learning pay attention to weight scores, and the obtained attribute type embedding and attribute value embedding are used for performing information aggregation in a cascading manner, so as to improve the embedding learning efficiency of the entity. And the entity overall feature information is further integrated to complement each other, so that the entity alignment accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a generalized entity matching method based on self-supervised learning and attribute perception. Background Art

[0002] Entity matching is an important task in the field of data integration. Data integration is an important concept in the field of data science and artificial intelligence. It involves extracting valuable information from data of different sources, formats and types and integrating them into a unified data set. With the continuous development of the field of data integration, entity matching technology is also constantly improving. From the early algorithms based on statistical learning, to the application of deep learning technology, to the introduction of attention mechanism, the performance and accuracy of entity matching technology have been significantly improved.

[0003] Entity matching, also known as record linking, entity resolution or object recognition, is also called duplicate detection in a single data source. In 1946, Dunn et al. defined the task of entity matching as finding records pointing to the same entity in multiple data sources. Entity matching mainly includes entity matching within a data source and entity matching between multiple data sources. Matching within the same data source is mainly aimed at situations where there are duplicate entity records in the data source; matching between data sources mainly matches entities between multiple data sources. This level of matching can better associate entities between different data sources, laying the foundation for more comprehensive mining of user behavior and user characteristics.

[0004] The current mainstream entity matching method is the entity matching method based on deep learning, which is mainly divided into the entity matching method based on large language models and the entity matching method based on graph neural networks. The entity matching method based on large language models serializes the descriptions of the relevant entities and regards them as sentences. It uses a pre-trained language model fine-tuned for the entity matching task to model the entity matching task as a sentence pair classification problem. The method based on graph neural networks obtains the embedded representation of the entity by recursively aggregating the neighbors of the entity, which can better capture the structural information between nodes.

[0005] At present, the entity matching problem is limited to matching between two homogeneous tables. However, in the real scenario of data science, matching mostly exists between heterogeneous data sets. At present, the entity matching method based on large language models requires a large number of manually annotated training data sets, which is too labor-intensive. In addition, the commonly used pre-trained large language models aim to predict missing words in sentences and predict the next sentence based on the current input sentence, while the goal of fine-tuning is to classify sentence pairs. Therefore, there is a difference in goals between the two, which will make the model used for entity matching tasks perform poorly. Furthermore, the current entity matching method based on large language models treats the data describing entities completely as sentences, ignoring the structure of the data, resulting in the lack of matching information. The method based on graph neural network obtains the embedded representation of the entity by recursively aggregating the neighbors of the entity, which can better capture the structural information between nodes. However, most of the current entity matching methods based on graph neural networks only focus on the single structural information in entity matching to obtain the embedded representation of the entity. The relationship information, name information, and attribute information in the entity data are not fully utilized, and the interaction between information is ignored, which affects the embedded representation of the entity and leads to poor matching performance. At the same time, some graph-based methods and random walk strategies are not scalable on large data sets and are time-consuming and material-intensive.

[0006] In addition, the use of unsupervised learning methods to replace the current supervised learning methods, although reducing the use of manually labeled labels, also creates new problems. That is, unsupervised learning has no labels to guide the learning process, which makes it more difficult to identify the accuracy of the model or obtain high-quality clustering or dimensions. Noise data and outliers can significantly distort the data structure and affect the learning results. Since the data is not labeled, large errors and defects may occur in the prediction, resulting in inaccurate results. Unsupervised learning may require a lot of computing resources, especially when processing large data sets. Summary of the invention

[0007] In order to solve the problem of over-reliance on manually labeled labels and loss of entity structure data in the entity matching process, this application proposes a generalized entity matching method based on self-supervised learning and attribute perception, specifically, an entity matching method that can get rid of the reliance on manually labeled labels, integrates entity names, structure information, attribute information, and relationship information in entity data, and uses pseudo-label generation strategies to automatically generate label instances. At the same time, it considers the various information carried by the entity itself to optimize and improve the embedded representation of the entity. In addition, the influence of structural information on the matching results is improved through a shared weight mechanism, so that entities can achieve a one-to-one stable match and improve the accuracy of entity matching.

[0008] In the first aspect, the present application proposes a generalized entity matching method based on self-supervised learning and attribute perception, comprising:

[0009] Through the self-supervised learning method, a pseudo-label generation strategy is adopted to obtain positive and negative training samples for entity matching;

[0010] The positive and negative training samples of entity matching are divided into training set and test set using a preset ratio;

[0011] The language model is trained using the training set to obtain a pre-trained language model;

[0012] Randomly select the first attribute type and the second attribute type in the test set;

[0013] Inputting the first attribute type and the second attribute type into a pre-trained language model to obtain M first attribute type embeddings and M second attribute type embeddings;

[0014] Embed the M first attribute types and embed the attribute values ​​corresponding to the first attribute types, and perform cascading to obtain a first cascade sequence; embed the M second attribute types and embed the attribute values ​​corresponding to the second attribute types, and perform cascading to obtain a second cascade sequence;

[0015] Inserting the marker into the starting position of the first cascade sequence to obtain a first marker sequence, and inserting the marker into the starting position of the second cascade sequence to obtain a second marker sequence;

[0016] The output of the pre-trained language model, the first tag sequence, and the second tag sequence are fused and input into the fully connected layer;

[0017] The output of the fully connected layer is input into the Softmax layer, and the output of the Softmax layer is used as the generalized entity matching result.

[0018] The method uses a pseudo-label generation strategy through a self-supervised learning method to obtain positive and negative training samples for entity matching, including:

[0019] Use existing structured datasets in entity matching tasks to build benchmark datasets;

[0020] Performing an initial entity embedding representation on the benchmark data set to obtain a sequence of the benchmark data set;

[0021] The sequence of the benchmark data set is divided into a first entity set and a second entity set, and the cosine similarity is used to calculate the most similar entity and the second most similar entity of each entity in the first entity set in the second entity set, and the cosine similarity is used to calculate the most similar entity and the second most similar entity of each entity in the second entity set in the first entity set;

[0022] Generate a positive label generation strategy and a negative label generation strategy based on the cosine similarity calculation results of the first entity set and the second entity set;

[0023] Positive label generation strategy and negative label generation strategy are adopted to obtain positive and negative training samples for entity matching.

[0024] The positive label generation strategy includes:

[0025] In the first entity set T and the second entity set T′, for each entity e i ∈T, let e′ j and e′ k is the second entity set T′ with e i The most similar and second most similar entities, for each entity e′ j ∈T′, let e l and e u is the first entity set T that is related to e′ j For the most similar and second most similar entities, the positive label generation strategy must meet the following two conditions:

[0026] (1)e i and e′ j The most similar vectors to each other, that is, e i =e l ;

[0027] (2) For each entity in the first entity set T and the second entity set T′, the margin between the most similar and second similar vectors of the entity is greater than the set threshold θ.

[0028] The negative label generation strategy is calculated as follows:

[0029] h vhn =(1-μ)h neg +μh anc ,μ∈(0,1)

[0030] Among them, h neg For entity e i The embedding vector of any instance in the same batch, μ is the defined hyperparameter, h anc For entity e i The embedding representation of h vhn Embedding vectors of negative samples generated using the negative label generation strategy.

[0031] The step of inputting the first attribute type and the second attribute type into a pre-trained language model to obtain M first attribute type embeddings and M second attribute type embeddings includes:

[0032] The first attribute type is input into the pre-trained language model for embedding operation, and M first attribute type embeddings A=(a 1 ,a 2 …a m );

[0033] The second attribute type is input into the pre-trained language model for embedding operation to obtain M second attribute type embeddings A'=(a' 1 ,a′ 2 …a′ m ).

[0034] The M first attribute type embeddings and the attribute value embeddings corresponding to the first attribute type are cascaded to obtain a first cascade sequence, which is calculated as follows:

[0035]

[0036] Among them, a i is the i-th attribute type embedding of the first attribute type, v i is the attribute value embedding corresponding to the first attribute type, α i is the attention weight of the i-th attribute of the first attribute class, e type The sum of the first attribute type, e value is the sum of the first attribute value, and e is the first cascade sequence;

[0037] Embed the M second attribute types and concatenate the attribute values ​​corresponding to the second attribute types to obtain a second concatenated sequence, which is calculated as follows:

[0038]

[0039] Among them, a′ i is the i-th attribute type embedding of the second attribute type, v′ i is the attribute value embedding corresponding to the second attribute type, α′ i is the attention weight of the i-th attribute of the second attribute class, e type′ The sum result of the second attribute type, e value′ is the sum of the second attribute value, and e′ is the second cascade sequence.

[0040] The output of the pre-trained language model, the first tag sequence, and the second tag sequence are fused and input into the fully connected layer. The calculation formula is as follows:

[0041] C e =S(e)+E [CLS] ; C e′ =S(e′)+E [CLS]

[0042]

[0043] Among them, E [CLS] is the output of the pre-trained language model, S(e) is the first tag sequence, S(e′) is the second tag sequence, σ is an activation function, is the attribute feature embedding representation of the fused entity, The first joint entity C e With the second joint entity C e′ cascade.

[0044] In a second aspect, the present application proposes an electronic device comprising: one or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the generalized entity matching method based on self-supervised learning and attribute perception.

[0045] In a third aspect, the present application proposes a computer-readable storage medium storing executable instructions, which, when executed, enable a processor to execute the generalized entity matching method based on self-supervised learning and attribute perception.

[0046] In a fourth aspect, the present application proposes a computer program product, including a computer program or instructions, which, when executed by a processor, implements the generalized entity matching method based on self-supervised learning and attribute perception.

[0047] Beneficial effects:

[0048] This application proposes a generalized entity matching method based on self-supervised learning and attribute perception, which makes full use of the information carried by the entity data itself to improve the matching effect. Through the self-supervised learning method, a positive and negative label generation strategy is proposed to generate training instances, and a shared weight mechanism is proposed to improve the entity embedding learning process. Specifically, the attribute value shares the learned attribute type attention weight score, and the attribute type embedding and attribute value embedding are used to aggregate information in a cascade manner. Then the overall feature information of the entity is further integrated to complement each other and improve the accuracy of entity alignment. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 A flow chart of a generalized entity matching method based on self-supervised learning and attribute perception in an embodiment of the present invention;

[0050] Figure 2 A positive and negative label generation strategy diagram based on self-supervised learning according to an embodiment of the present invention;

[0051] Figure 3 The shared attention mechanism of an embodiment of the present invention calculates a weight graph. DETAILED DESCRIPTION

[0052] The specific implementation of the present application is further described in detail below in conjunction with the drawings and examples.

[0053] Embodiment 1:

[0054] This embodiment proposes a generalized entity matching method based on self-supervised learning and attribute perception, such as Figure 1 As shown, including:

[0055] Step S1: Using a self-supervised learning method and a pseudo-label generation strategy, positive and negative training samples for entity matching are obtained;

[0056] Step S2: Divide the positive and negative training samples of entity matching into a training set and a test set using a preset ratio;

[0057] Step S3: training the language model using the training set to obtain a pre-trained language model;

[0058] Step S4: randomly select the first attribute type and the second attribute type in the test set;

[0059] Step S5: input the first attribute type and the second attribute type into a pre-trained language model to obtain M first attribute type embeddings and M second attribute type embeddings;

[0060] Step S6: embedding the M first attribute types and the attribute values ​​corresponding to the first attribute types, and cascading them to obtain a first cascade sequence; embedding the M second attribute types and the attribute values ​​corresponding to the second attribute types, and cascading them to obtain a second cascade sequence;

[0061] Step S7: inserting the marker into the starting position of the first cascade sequence to obtain a first marker sequence, and inserting the marker into the starting position of the second cascade sequence to obtain a second marker sequence;

[0062] Step S8: fusing the output of the pre-trained language model, the first tag sequence, and the second tag sequence and inputting them into the fully connected layer;

[0063] Step S9: Input the output of the fully connected layer to the Softmax layer, and use the output of the Softmax layer as the generalized entity matching result.

[0064] In step S1, the self-supervised learning method adopts a pseudo-label generation strategy to obtain positive and negative training samples for entity matching, such as Figure 2 As shown, including:

[0065] Step S1.1: Use the existing structured datasets in the entity matching task to build a benchmark dataset;

[0066] In this embodiment, the already contained data sets that are accurately annotated by humans can be utilized and the expensive cost of secondary manual labeling can be avoided. Since they were originally proposed to match relational data, only necessary conversions need to be performed to build equivalent semi-structured or unstructured tables without changing the meaning of the records. Two heterogeneous structured data sets and semi-structured data sets are obtained from public data sets, including the attribute types, attribute values, entity name information, etc. of the entities. Specifically, the main dependent resources for creating the benchmark data sets include the Magellan database, the DeepMatcher data set, and the Web Data Commons collection (WDC) data set;

[0067] Two resources are introduced from the Magellan repository. They consist of several smaller datasets. The table schemas in different datasets are also different, and each dataset is stored in a CSV-like format. Here, the task of matching structured datasets with semi-structured datasets is taken as an example.

[0068] Step S1.2: Performing an initial entity embedding representation on the benchmark dataset to obtain a sequence of the benchmark dataset;

[0069] In this embodiment, the description information of the entities in the benchmark data set contains rich semantic information. For a structured table, the data entry with the nth attribute can be represented as e={attr i ,val i} i∈[1,n] ,attr i is the name of the i-th attribute, val i is the attribute value of the i-th attribute, and the serialized representation is:

[0070] serialize(e):[COL]attr 1 [VAL]val 1 ...[COL]attr n [VAL]val n (1)

[0071] Among them, [COL] and [VAL] are two special tags, representing the attribute name and the starting field of the corresponding attribute value respectively. serialize(e) is the serialized representation. For example, given an entity in a structured table, it is serialized as:

[0072] [COL]title[VAL]efficient similarity...[COL]authors[VAL]renald...[COL]venue[VAL]

[0073] SIGMOD[COL]year[VAL]2003(2)

[0074] Semi-structured tables are serialized in the same way as above. The differences are: (i) for nested attributes, [COL] and [VAL] tags should be recursively added along the attribute name and the value in each nested layer; (ii) for attributes whose content is a list, the elements in the list are concatenated into a string separated by commas. For example, given a semi-structured entity, it is serialized as:

[0075] [COL]title[VAL]efficient similarity...[COL]year[VAL]2003[COL]authors[VAL]

[0076] ronald fagin ravi kumar d.sivakumar(3)

[0077] Since the text data itself is a sentence, no serialization operation is required.

[0078] Step S1.3: Divide the sequence of the benchmark data set into a first entity set and a second entity set, use cosine similarity to calculate the most similar entity and the second most similar entity in the second entity set for each entity in the first entity set, and use cosine similarity to calculate the most similar entity and the second most similar entity in the first entity set for each entity in the second entity set;

[0079] In this embodiment, the cosine similarity calculation is performed using the initial embedding representation of the entity to obtain the similarity results between entities from two different data sets, which are calculated as follows:

[0080] θ=cosine(e i ,e′ j ) (4)

[0081] Among them, cosine(·) is the cosine similarity calculation representation of two entity embeddings, e i and e′ j Represent two entities from entity set T and entity set T′ respectively; i and e′ j Represents the embedding vector of two entities in space;

[0082] The cosine similarity results of step S1.3 are used to obtain the most similar and second most similar entity representations for each entity, and a positive label generation strategy is introduced;

[0083] Step S1.4: Generate a positive label generation strategy and a negative label generation strategy through the cosine similarity calculation results of the first entity set and the second entity set;

[0084] Step S1.5: Adopt the positive label generation strategy and the negative label generation strategy to obtain positive and negative training samples for entity matching.

[0085] The positive label generation strategy includes:

[0086] In the first entity set T and the second entity set T′, for each entity e i ∈T, let e′ j and e′ k is the second entity set T′ with e i The most similar and second most similar entities, for each entity e′ j ∈T′, let e l and e u is the first entity set T that is related to e′ j For the most similar and second most similar entities, the positive label generation strategy must meet the following two conditions:

[0087] (1)e i and e′ j The most similar vectors to each other, that is, e i =e l ;

[0088] (2) For each entity in the first entity set T and the second entity set T′, the margin between the most similar and the second similar vectors of the entity is greater than the set threshold θ, that is, δ 1 =Sim(e i -e′ j )-Sim(e i -e′ k ) and δ 2 =Sim(e′ j -e l )-Sim(e′ j -e u ) are all greater than θ; where Sim(·) is equivalent to cosine(·);

[0089] The principle of the virtual strict negative label generation strategy is as follows. For each instance, the conventional negative label construction method is to use other instances in the same batch as its negative examples. However, such negative sample pairs are easily distinguished from positive sample pairs. For example, if we generate a negative tuple pair ("Apple Inc.", "Apple") for a positive tuple pair ("Apple Inc.", "Google"), then it is obvious that "Google" and "Apple Inc." are not equivalent. This type of negative tuple pair has low information content and therefore contributes little to the embedding training process. Ideally, we expect to generate effective negative entity pairs. Strict entity pairs, that is, those instances that are more similar to the original instance but do not match, can effectively solve the problem of the above simple negative instances being easy to distinguish. However, the problem with using too strict entity pairs is that the model can easily mistake strict negative instances for positive instances. In this regard, a compromise idea is proposed to generate virtual strict negative samples. The negative label generation strategy is calculated as follows:

[0090] h vhn =(1-μ)h neg +μh anc ,μ∈(0,1) (5)

[0091] Among them, h neg For entity e i The embedding vector of any instance in the same batch, μ is the defined hyperparameter, h anc For entity e i The embedding representation of h vhn Embedding vectors of negative samples generated using the negative label generation strategy.

[0092] In steps S2 to S4, the training samples obtained in step S1 are divided into a training set and a test set in a ratio of 3:1. Figure 3As shown, the matching stage adopts a shared weight mechanism, the principle of which is that a large proportion of "noise" attributes in the data description of the entity contribute little to entity matching. For example, the entity "Titanic" in DBPedia has the attribute "IMDB ID" to record its ID in the IMDB website, but the "IMDB ID" attribute does not exist in "Titanic" in YAGO. Here, "IMDB ID" in DbPedia can be regarded as a noise attribute, which is not important for matching the entity "Titanic". In order to learn the importance of different attributes to entities, it is proposed to use attribute type embedding to obtain attention weights and make attribute types and attribute values ​​share the same attention weight mechanism. Specifically explained as follows, for example, in the attribute triple (Tom, age, "12"), "age" is the attribute type of the entity "Tom", and "12" is the value of "age"; the language model is trained using the training set to obtain a pre-trained language model, and the public BERT (Bidirectional Encoder Representations from Transformers, bidirectional transformer model) large language model used in this embodiment. Take the first attribute type and the second attribute type in the test set at will.

[0093] In step S5, the first attribute type and the second attribute type are input into a pre-trained language model to obtain M first attribute type embeddings and M second attribute type embeddings, including:

[0094] The first attribute type is input into the pre-trained language model for embedding operation, and M first attribute type embeddings A=(a 1 ,a 2 …a m );

[0095] In this embodiment, the pre-trained language model is used to embed the M attribute types of entity e, and A=(a 1 ,a 2 …a m ), calculate the attention weight α of the i attributes of the first attribute class i :

[0096] α i =softmax(A T ω a a i ) (6)

[0097] Among them, ω a represent Here, the attention weight is obtained by embedding the attribute type, and the attribute value weight is consistent with the attribute type weight.

[0098] Similarly, the second attribute type is input into the pre-trained language model for embedding operation, and M second attribute type embeddings A′=(a′ 1 ,a′ 2 …a′ m ). Calculate the attention weight α′ of the i attributes of the second attribute class i :

[0099] α′ i =softmax(A′ T ω′ a a′ i ) (7)

[0100] Among them, ω′ a represent The weight matrix of .

[0101] In step S6, the M first attribute type embeddings and the attribute value embeddings corresponding to the first attribute type are cascaded to obtain a first cascade sequence, which is calculated as follows:

[0102]

[0103] Among them, a i is the i-th attribute type embedding of the first attribute type, v i is the attribute value embedding corresponding to the first attribute type, α i is the attention weight of the i-th attribute of the first attribute class, e type The sum of the first attribute type, e value is the sum of the first attribute value, and e is the first cascade sequence;

[0104] Embed the M second attribute types and concatenate the attribute values ​​corresponding to the second attribute types to obtain a second concatenated sequence, which is calculated as follows:

[0105]

[0106] Among them, a′ i is the i-th attribute type embedding of the second attribute type, v′ i is the attribute value embedding corresponding to the second attribute type, α′ i is the attention weight of the i-th attribute of the second attribute class, e type′ The sum result of the second attribute type, e value′ is the sum of the second attribute value, and e′ is the second cascade sequence.

[0107] In step S7, the marker is inserted into the starting position of the first cascade sequence to obtain a first marker sequence, and the marker is inserted into the starting position of the second cascade sequence to obtain a second marker sequence;

[0108] In this embodiment, for the method based on the pre-trained language model, the tag [CLS] is further inserted in front of the sequence. This tag is a special tag required for BERT to encode the sequence pair into a feature vector, which will be fed into the fully connected layer for classification. The final serialization is represented as:

[0109] S(e)=[CLS]serialize(e)[SEP] serialize(e')[SEP] (12)

[0110] In step S8, the output of the pre-trained language model, the first tag sequence, and the second tag sequence are fused and input into the fully connected layer. The calculation formula is as follows:

[0111] C e =S(e)+E [CLS] ; C e′ =S(e′)+E [CLS] (13)

[0112]

[0113] Among them, E [CLS] is the output of the pre-trained language model, S(e) is the first tag sequence, S(e′) is the second tag sequence, σ is an activation function, is the attribute feature embedding representation of the fused entity, The first joint entity C e With the second joint entity C e′ cascade.

[0114] It is understandable that embedding two data items separately may lose cross-entity information (e.g., the difference between two product ID spans), so the combination of records e and e′ should be taken as input to obtain the joint context embedding. [CLS] Between S(e′) and E [CLS] The fusion operations are performed between them. [CLS] It contains the overall semantics of the record pair, so the fusion operation can ensure that the interactive attention between e and e′ will not be lost.

[0115] A variant of the cross entropy loss L1 is used as the target training function; the classification signal E is learned by feeding the input into a multi-layer Transformer encoder. [CLS] ;

[0116]

[0117] Among them, the logical value d * The sentence features E of two entities [CLS] Generate, c is the dimension of the entity embedding representation,

[0118] In order to better jointly learn the embedding representation of entities and relations, we should consider how to minimize the semantic distance between matching entities. At the input, the pre-trained language model assigns an initial embedding to each token of the sentence S(e), denoted as e. The special symbol [CLS] at the front of each sentence also has an initial embedding, denoted as E [CLS] The embedding of each token will be updated after executing the multi-layer Transformer encoder. Max-pooling is applied here to obtain a fixed-length embedding e for representing entity e. Specifically, max-pooling generates a fixed-length embedding by selecting the maximum value in each dimension among all the embedding tokens in the tuple. Cosine embedding loss is used As the target training function, it aims to minimize the semantic distance. (positive label set And maximize the distance between unmatched tuples, the negative label set )

[0119]

[0120] Among them, μ is the edge hyperparameter artificially set for generating negative label instances, and cosine(·) is the cosine distance metric;

[0121] After training, we get an entity embedding representation based on structural information. The input data is mapped to a new space through a linear transformation through the Linear layer to integrate the features extracted by the previous training layer and provide a decision basis for classification or regression tasks. According to formula (14), we calculate

[0122] In step S9, the output of the fully connected layer is input to the Softmax layer, and the output of the Softmax layer is used as the generalized entity matching result.

[0123] In this embodiment, the feature embedding obtained in step S8 is input into the Softmax layer to convert the output of the neural network into a probability distribution, and the output value ranges from 0 to 1, and the sum of all output values ​​is 1. This means that each output value can be regarded as a probability value, and during the training process, the Softmax layer is used in conjunction with the loss function to simplify the gradient calculation, making the gradient descent algorithm more stable;

[0124] The entity matching task belongs to a binary classification problem. In the binary classification problem, the Softmax function degenerates into the Logistic function (also called the Sigmoid function), which is used to output the probability of belonging to one of the two categories. pm is defined as:

[0125]

[0126] The comparison results between this embodiment and the existing baseline model are as follows:

[0127] Table 1 Comparison of the effects of this embodiment and the existing baseline model;

[0128]

[0129]

[0130] As can be seen from Table 1, precision, recall and F1 score are selected as evaluation indicators and expressed in percentage (%). The present invention uses 60% of pseudo-label entity pairs as training data, and the remaining 30% is used for testing. The method of the present invention is evaluated on six benchmark data sets of generalized entity matching, which are heterogeneous table structure, homogeneous semi-structure, heterogeneous semi-structure, semi-structure-structure, semi-structure-text-Watch, and semi-structure-text-Camera. Among them, the semi-structure related data set contains 20,000 heterogeneous entity pairs. The experimental results show that the method of the present invention is superior to the existing baseline model.

[0131] Compared with the existing methods, the method proposed in this embodiment uses a positive and negative label generation strategy to generate pseudo labels, solves the high cost problem caused by manual labeling, fully considers the structural information, relationship information, and attribute information of entities in entity matching, and uses a shared weight mechanism to help entities capture key matching information, reduce the interference of noise information in entity information, and effectively optimize the embedded representation of entities. At the same time, the global entity matching loss function strategy is adopted to improve the training accuracy of the large language model and improve the accuracy of entity matching. The experimental results are shown in the following table, and the accuracy (Precision), recall (Recall) and F1 score are selected as evaluation indicators and expressed in percentage (%). The present invention uses 60% of pseudo-label entity pairs as training data, and the remaining 30% is used for testing. The method of the present invention evaluates the method on six benchmark data sets of generalized entity matching, which are heterogeneous table structure, homogeneous semi-structure, heterogeneous semi-structure, semi-structure-structure, semi-structure-text-Watch, and semi-structure-text-Camera. Among them, the semi-structure related data set contains 20,000 heterogeneous entity pairs. The experimental results show that the method of the present invention is better than the existing baseline model.

[0132] Embodiment 2:

[0133] This embodiment proposes an electronic device, including: one or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the generalized entity matching method based on self-supervised learning and attribute perception.

[0134] The electronic device can be a mobile phone, a computer or a tablet computer, etc., including a memory and a processor, wherein a computer program is stored on the memory, and when the computer program is executed by the processor, a generalized entity matching method based on self-supervised learning and attribute perception as described in the embodiment is implemented. It can be understood that the electronic device can also include an input / output (I / O) interface and a communication component.

[0135] The processor is used to execute all or part of the steps in the generalized entity matching method based on self-supervised learning and attribute perception as described in the above embodiment. The memory is used to store various types of data, which may include instructions of any application or method in the electronic device, as well as data related to the application.

[0136] The processor can be an application specific integrated circuit (ASIC), a digital signal processor (DSP), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor or other electronic components, and is used to execute a generalized entity matching method based on self-supervised learning and attribute perception described in the above embodiment.

[0137] Embodiment 3:

[0138] This embodiment provides a computer-readable storage medium storing executable instructions. When the instructions are executed and implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0139] The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of a generalized entity matching method based on self-supervised learning and attribute perception described in various embodiments of the present application.

[0140] The aforementioned storage media include: flash memory, hard disk, multimedia card, card-type memory (for example, SD (Secure Digital Memory Card) or DX (Memory Data Register, MDR abbreviation, memory data register) memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, CD, server, APP (Application, abbreviation of application software) application store and other media that can store program verification codes, on which computer programs are stored. When the computer program is executed by the processor, the various steps of the above-mentioned generalized entity matching method based on self-supervised learning and attribute perception can be implemented.

[0141] Embodiment 4:

[0142] This embodiment proposes a computer program product, including a computer program or instructions, which, when executed by a processor, implements the generalized entity matching method based on self-supervised learning and attribute perception.

[0143] Based on such understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a computer program product.

[0144] The various embodiments in the present application are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

[0145] The protection scope of the present application is not limited to the above-mentioned embodiments. Obviously, those skilled in the art can make various changes and modifications to the present disclosure without departing from the scope and spirit of the present disclosure. If these changes and modifications fall within the scope of the claims of the present disclosure and their equivalents, the intention of the present disclosure also includes these changes and modifications.

Claims

1. A generalized entity matching method based on self-supervised learning and attribute perception, characterized in that: include: Through the self-supervised learning method, a pseudo-label generation strategy is adopted to obtain positive and negative training samples for entity matching; The positive and negative training samples of entity matching are divided into training set and test set using a preset ratio; Use the training set to train the language model to obtain a pre-trained language model; Randomly select the first attribute type and the second attribute type in the test set; Inputting the first attribute type and the second attribute type into a pre-trained language model to obtain M first attribute type embeddings and M second attribute type embeddings; Embed the M first attribute types and embed the attribute values ​​corresponding to the first attribute types, and perform cascading to obtain a first cascade sequence; embed the M second attribute types and embed the attribute values ​​corresponding to the second attribute types, and perform cascading to obtain a second cascade sequence; Inserting the marker into the starting position of the first cascade sequence to obtain a first marker sequence, and inserting the marker into the starting position of the second cascade sequence to obtain a second marker sequence; The output of the pre-trained language model, the first tag sequence, and the second tag sequence are fused and input into the fully connected layer; The output of the fully connected layer is input into the Softmax layer, and the output of the Softmax layer is used as the generalized entity matching result.

2. According to claim 1, a generalized entity matching method based on self-supervised learning and attribute perception is characterized in that: The method uses a pseudo-label generation strategy through a self-supervised learning method to obtain positive and negative training samples for entity matching, including: Use existing structured datasets in entity matching tasks to build benchmark datasets; Performing an initial entity embedding representation on the benchmark data set to obtain a sequence of the benchmark data set; The sequence of the benchmark data set is divided into a first entity set and a second entity set, and the cosine similarity is used to calculate the most similar entity and the second most similar entity of each entity in the first entity set in the second entity set, and the cosine similarity is used to calculate the most similar entity and the second most similar entity of each entity in the second entity set in the first entity set; Generate a positive label generation strategy and a negative label generation strategy based on the cosine similarity calculation results of the first entity set and the second entity set; Positive label generation strategy and negative label generation strategy are adopted to obtain positive and negative training samples for entity matching.

3. A generalized entity matching method based on self-supervised learning and attribute perception according to claim 2, characterized in that: The positive label generation strategy includes: In the first entity set T and the second entity set T′, for each entity e i ∈T, let e j ′ and e′ k is the second entity set T′ with e i The most similar and second most similar entities, for each entity e j ′∈T′, let e l and e u is the first entity set T with e j For the two most similar and second most similar entities, the positive label generation strategy must meet the following two conditions: (1)e i and e j ′ are the most similar vectors to each other, that is, e i =e l ; (2) For each entity in the first entity set T and the second entity set T′, the margin between the most similar and second similar vectors of the entity is greater than the set threshold θ.

4. The generalized entity matching method based on self-supervised learning and attribute perception according to claim 2, characterized in that: The negative label generation strategy is calculated as follows: h vhn =(1-μ)h neg +μh anc ,μ∈(0,1) Among them, h neg For entity e i The embedding vector of any instance in the same batch, μ is the defined hyperparameter, h anc For entity e i The embedding representation of h vhn Embedding vectors of negative samples generated using the negative label generation strategy.

5. The generalized entity matching method based on self-supervised learning and attribute perception according to claim 1, characterized in that: The step of inputting the first attribute type and the second attribute type into a pre-trained language model to obtain M first attribute type embeddings and M second attribute type embeddings includes: The first attribute type is input into the pre-trained language model for embedding operation, and M first attribute type embeddings A = (a1, a2…a m ); The second attribute type is input into the pre-trained language model for embedding operation, and M second attribute type embeddings A'=(a'1, a'2…a' m ).

6. A generalized entity matching method based on self-supervised learning and attribute perception according to claim 1, characterized in that: The M first attribute type embeddings and the attribute value embeddings corresponding to the first attribute type are cascaded to obtain a first cascade sequence, which is calculated as follows: Among them, a i is the i-th attribute type embedding of the first attribute type, v i is the attribute value embedding corresponding to the first attribute type, α i is the attention weight of the i-th attribute of the first attribute class, e type The sum of the first attribute type, e value is the sum of the first attribute value, and e is the first cascade sequence; Embed the M second attribute types and concatenate the attribute values ​​corresponding to the second attribute types to obtain a second concatenated sequence, which is calculated as follows: Among them, a′ i is the i-th attribute type embedding of the second attribute type, v′ i is the attribute value embedding corresponding to the second attribute type, α′ i is the attention weight of the i-th attribute of the second attribute class, e type′ The sum result of the second attribute type, e value′ is the sum of the second attribute value, and e′ is the second cascade sequence.

7. The generalized entity matching method based on self-supervised learning and attribute perception according to claim 1, characterized in that: The output of the pre-trained language model, the first tag sequence, and the second tag sequence are fused and input into the fully connected layer. The calculation formula is as follows: C e =S(e)+E [CLS] ;C e′ =S(e′)+E [CLS] Among them, E [CLS] is the output of the pre-trained language model, S(e) is the first tag sequence, S(e′) is the second tag sequence, σ is an activation function, is the attribute feature embedding representation of the fused entity, The first joint entity C e With the second joint entity C e′ cascade.

8. An electronic device, characterized in that: include: One or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute a generalized entity matching method based on self-supervised learning and attribute perception as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: It stores executable instructions, which, when executed, enable a processor to execute a generalized entity matching method based on self-supervised learning and attribute perception as described in any one of claims 1 to 7.

10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, a generalized entity matching method based on self-supervised learning and attribute perception as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Planar geographic entity self-supervised matching method and device and medium

    CN122112660A