An entity matching method, system, electronic device and storage medium
By constructing undirected and powerless graphs and random walk sampling to obtain samples, and continuing to pre-train the BERT model in combination with entity and attribute context, the problem of failing to fully utilize attribute context in the prior art is solved and the accuracy of entity matching is improved.
Patent Information
- Application Number
- CN202111220113.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-20
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-10-20
AI Technical Summary
The prior art only considers the entity context in entity matching, and fails to fully utilize the attribute context, thereby affecting the matching accuracy.
By using the attributes in the entity collection as nodes to build an undirected and unrighteous graph, perform random walk sampling, obtain samples and build a corpus, continue to pre-train the BERT model, and introduce the attribute context for entity matching.
Improve the accuracy of entity matching, and improve the entity matching effect based on pre-trained language models by combining entity and attribute context.
Smart Images

Figure CN113961714B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to an entity matching method, system, electronic device, and storage medium. Background Art
[0002] Currently, a large amount of data is being generated in various fields, such as e-commerce, social networking, transportation, catering, and so on. This data contains a large amount of valuable information that can help enterprises improve operational efficiency and enhance the user experience. However, in the era of big data, there is a huge challenge in making better use of this data, namely multi-source data integration. Since each enterprise, or even each department of the same enterprise, will establish independent databases according to its own needs, and there may be redundant information between these databases. Therefore, integrating multiple databases from different sources and in different forms to provide a unified data view has important value.
[0003] There is an important problem in the field of data integration, called entity matching or entity resolution. The goal of entity matching is to determine whether two entities in a database refer to the same thing in the real world. An entity is data that describes things in the real world. A single entity e can be regarded as a set of key-value pairs, denoted as e = {attr i , val i} 1≤i≤m , where m is the number of attributes in the entity, attr i is the attribute name, and val i is the attribute value. Given two entity sets D and D', the entity matching task is to determine whether any pair of entities (e, e') refers to the same thing in the real world, where e ∈ D and e' ∈ D'. For example: Given two entities, namely Entity 1 (Name: Zhang San, Age: 30, Address: Chaoyang District, Beijing, Occupation: Programmer) and Entity 2 (Name: Zhang San, Age: 31, Address: Haidian District, Beijing, Occupation: Programmer). Then, are Entity 1 and Entity 2 referring to the same person? This is the problem faced by entity matching.
[0004] The following methods are mainly used in the prior art to solve the above problems:
[0005] The first method is the technology based on pre-trained word vectors; among them, the general idea of the pre-trained word vectors based on entity serialization is to convert the entity into text, and then use word2vec to obtain the pre-trained word vectors. Specifically: (1) Given an entity set D, serialize all entities in D into text; the specific serialization method is serialize(e) = val1 val 2 …val 3 That is, all the attribute values of the entity are concatenated. (2) Use all the serialized entities to construct the corpus T; (3) Use word2vec to train word vectors on the corpus T; The second method is a technique based on a pre-trained language model. Specifically, this method converts the entity pair (e, e') into a text pair (s, s'), and then uses a text matching model to perform entity matching, where s = serialzie(e) and s' = serialize(e').
[0006] The prior art has at least the following defects: Among them, there is a disadvantage in the pre-trained word vectors based on entity serialization, that is, only the context within a single entity is considered during training. As Figure 1 shown, the gender of "Entity 1" is "male". When training the word vector of the word "male", not only the context of "Entity 1" should be considered, but also the context of the attribute "gender" should be considered. Currently, the best-performing entity matching method is based on a pre-trained language model, which mainly benefits from the powerful expressive ability of the pre-trained language model. However, this method also only considers the entity context and does not consider the attribute context. Summary of the Invention
[0007] In view of the technical problem that only the entity context is considered and the attribute context is not considered during the training of the above-mentioned pre-trained language model, the present invention proposes an entity matching method, system, electronic device and storage medium.
[0008] In a first aspect, an embodiment of the present application provides an entity matching method, including:
[0009] Obtain a target entity set by merging a first entity set and a second entity set, and construct an undirected unweighted graph with the attributes of each entity in the target entity set as nodes;
[0010] Perform random walk sampling in the undirected unweighted graph, convert the obtained path into text, and construct a corpus using the text;
[0011] Use the corpus to continue pre-train the BERT model;
[0012] Construct a text matching model according to the pre-trained BERT model, convert the entity matching corpus annotated in the corpus into text matching corpus, and train the text matching model using the text matching corpus;
[0013] Use the trained text matching model to match the entity to be matched.
[0014] The above entity matching method, wherein the construction of the undirected unweighted graph includes:
[0015] Select entities from the target entity set without replacement, and use the attributes of the selected entities as nodes to construct a complete graph;
[0016] Merge the complete graph into an initialized empty graph to obtain a preliminary undirected unweighted graph;
[0017] Traverse the target entity set, iteratively update the preliminary undirected unweighted graph, and obtain the final undirected unweighted graph.
[0018] The above entity matching method, wherein the construction of the corpus includes:
[0019] Perform random walk sampling in the final undirected unweighted graph;
[0020] Convert the path obtained by each random walk sampling into text, and repeat random walk sampling until the corpus is constructed.
[0021] The above entity matching method, wherein the random walk sampling specifically includes: when the random walk sampling reaches between the set minimum sampling length and the set maximum sampling length, each sampling terminates with a set probability, and when the random walk sampling reaches the set maximum sampling length, the sampling terminates.
[0022] The above entity matching method, wherein the continued pre-training includes: continuing to pre-train the BERT model using the pre-training task MLM based on the constructed corpus.
[0023] The above entity matching method, wherein the matching of the entity to be matched includes: converting the entity to be matched into text to be matched, inputting the text to be matched into the text matching model, and obtaining an entity matching result.
[0024] In a second aspect, an embodiment of the present application provides an entity matching system, including:
[0025] Graph generation module: Obtain a target entity set by merging a first entity set and a second entity set, and use the attributes of each entity in the target entity set as nodes to construct an undirected unweighted graph;
[0026] Graph sampling module: Perform random walk sampling in the undirected unweighted graph, convert the path obtained after sampling into text, and construct a corpus using the text;
[0027] Continued pre-training module: Use the corpus to continue pre-train the BERT model;
[0028] Training module: Construct a text matching model based on the pre-trained BERT model, convert the entity matching corpus marked in the corpus into a text matching corpus, and use the text matching corpus to train the text matching model;
[0029] Entity matching module: Use the trained text matching model to match the entities to be matched.
[0030] The above entity matching system, wherein the graph generation module includes:
[0031] Complete graph construction unit: Select entities from the target entity set without replacement, and use the attributes of the selected entities as nodes to construct a complete graph;
[0032] Initial undirected unweighted graph obtaining unit: Merge the complete graph into an initialized empty graph to obtain an initial undirected unweighted graph;
[0033] Final undirected unweighted graph obtaining unit: Traverse the target entity set, iteratively update the initial undirected unweighted graph, and obtain a final undirected unweighted graph.
[0034] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the entity matching method as described in the first aspect above.
[0035] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the entity matching method as described in the first aspect above.
[0036] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0037] 1. Use the attributes of all entities in the entity set as nodes to construct an undirected unweighted graph. Since the edges in the graph have neither direction nor weight, it is convenient to uniformly and randomly sample the nodes in the graph later, so that the obtained samples have a certain probability of containing the same attributes of different entities;
[0038] 2. The samples obtained by random sampling contain "entity context" and "attribute context", so that the attribute context can be incorporated into the pre-trained language model, which can further improve the entity matching method based on the pre-trained language model and improve the matching accuracy. The accuracy of data capabilities is improved. Description of the Drawings
[0039] Figure 1 It is an example diagram of entity context and attribute context in the entity set provided by the present invention;
[0040] Figure 2 Schematic diagram of steps of an entity matching method provided by the present invention;
[0041] Figure 3 Based on the Figure 2 Flow chart of step S1 in
[0042] Figure 4 Based on the Figure 2 Flow chart of step S2 in
[0043] Figure 5 Example diagram of an undirected unweighted graph generated by the present invention;
[0044] Figure 6 Example diagram of single sampling provided by the present invention;
[0045] Figure 7 Schematic diagram of the structure of a text matching model provided by the present invention;
[0046] Figure 8 Schematic diagram of the process of an embodiment of an entity matching method provided by the present invention;
[0047] Figure 9 Framework diagram of an entity matching system provided by the present invention;
[0048] Figure 10 Framework diagram of a computer device according to an embodiment of the present application.
[0049] Among them, the reference numerals are:
[0050] 1. Graph generation module; 11. Complete graph construction unit; 12. Initial undirected unweighted graph acquisition unit; 13. Final undirected unweighted graph acquisition unit; 2. Graph sampling module; 21. Random sampling unit; 22. Path conversion unit; 3. Continue pre-training module; 4. Training module; 5. Entity matching module; 81. Processor; 82. Memory; 83. Communication interface; 80. Bus. Detailed implementation manners
[0051] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided in the present application without creative efforts shall fall within the scope of protection of the present application.
[0052] Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present application. For those of ordinary skill in the art, without creative efforts, the present application can also be applied to other similar scenarios based on these drawings. In addition, it can also be understood that although the efforts made in such a development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing, or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as the content disclosed in the present application being insufficient.
[0053] In the present application, the mention of "embodiment" means that the specific features, structures, or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.
[0054] Unless otherwise defined, the technical terms or scientific terms involved in the present application should have the ordinary meaning understood by those of ordinary skill in the technical field to which the present application belongs. The terms "a", "an", "one", "the" and similar words involved in the present application do not indicate a limitation in quantity and can represent a singular or plural number. The terms "including", "comprising", "having" and any variations thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or units, but may also include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products or devices. The terms "connected", "coupled" and similar words involved in the present application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The term "plurality" involved in the present application refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the preceding and following associated objects. The terms "first", "second", "third", etc. involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0055] The present invention will be described in detail with reference to the embodiments shown in the accompanying drawings. It should be noted, however, that these embodiments are not intended to limit the present invention, and any equivalent transformation or substitution in terms of function, method, or structure made by those of ordinary skill in the art based on these embodiments shall fall within the protection scope of the present invention.
[0056] Before elaborating on the various embodiments of the present invention in detail, an overview of the core inventive concept of the present invention is provided and will be elaborated in detail through the following several embodiments.
[0057] The present invention transforms the entity set into an undirected unweighted graph, samples the graph through random walks, thereby obtaining samples, constructs a corpus using the obtained samples, and continues to pre-train the BERT model based on the corpus, improving the entity matching method based on the pre-trained language model and enhancing the effect of entity matching.
[0058] The basic concepts and symbols in the specific embodiments are as follows:
[0059] Let and be two entity sets and have the same attributes {A 1 , A 2 , …, A m}, where m is the number of attributes;
[0060] For any entity e i on the attribute A j , its value is denoted as e i [A j ;
[0061] The goal of entity matching is to determine whether any entity pair (e, e') refers to the same thing in reality, where e ∈ D and e' ∈ D'.
[0062] Embodiment 1:
[0063] As Figure 2 , Figure 8 shows, this embodiment discloses a specific implementation of an entity matching method (hereinafter referred to as "the method").
[0064] Specifically, the method disclosed in this embodiment mainly includes the following steps:
[0065] Step S1: Obtain the target entity set by merging the first entity set and the second entity set, and construct an undirected unweighted graph with the attributes of each entity in the target entity set as nodes;
[0066] Among them, as Figure 3 shows, step S1 specifically includes the following content:
[0067] Step S11: Select entities from the target entity set without replacement, and use the attributes of the selected entities as nodes to construct a complete graph;
[0068] Specifically, the target entity set is obtained by merging the first entity set D and the second entity set D'; That is From the target entity set Select entity e without replacement i , that is Take all attributes e i of entity e i [A 1 ,…,e i [A m as nodes to construct a complete graph G i .
[0069] Step S12: Merge the complete graph into an initialized empty graph to obtain a preliminary undirected unweighted graph;
[0070] Specifically, merge the complete graph G i into an initialized empty graph G. Specifically, add the nodes that do not exist in G in G to G, and add the corresponding edges to G as well; i
[0071] Step S13: Traverse the target entity set and iteratively update the preliminary undirected unweighted graph to obtain the final undirected unweighted graph.
[0072] Specifically, if then go to Step S11; otherwise, return the constructed undirected unweighted graph G; for example Figure 5 Figure 5 shown, the undirected unweighted graph is composed of the "name", "age", "gender", and "address" of entity 1 and entity 2 as nodes.
[0073] Step S2: Perform random walk sampling in the undirected unweighted graph, convert the sampled path into text, and use the text to construct a corpus;
[0074] Among them, as Figure 4 shown, Step S2 specifically includes the following contents:
[0075] Step S21: Perform random walk sampling in the final undirected unweighted graph;
[0076] Specifically, when the random walk sampling reaches between the set minimum sampling length and the set maximum sampling length, each sampling terminates with a set probability, and when the random walk sampling reaches the set maximum sampling length, the sampling terminates. In this embodiment, a node is sampled from graph G in a uniform sampling manner, called n 0; centered around n 0 , uniformly sample its adjacent nodes to obtain node n 1 ; centered around n 1 , uniformly sample its adjacent nodes to obtain node n 2 ; and so on; when sampling reaches , each sampling terminates with a manually specified probability p, where l 1 is the minimum sampling length set manually; when the sampling exceeds , the sampling directly terminates, where l 2 is the maximum sampling length set manually.
[0077] Step S22: Convert the paths obtained by each random walk sampling into text, and repeat the random walk sampling until the corpus construction is completed.
[0078] Specifically, convert the sampled path n 0 , n 1 ,... into text; for example Figure 6 shown, Figure 6 is an example of a single sampling, the starting node is "Zhang San", the entire path is indicated by the dotted arrow, and the finally sampled text is "Zhang San male Li Si Chaoyang District". Repeat Step S21 until enough text is sampled to form the corpus T.
[0079] Step S3: Use the corpus to continue pre-training the BERT model;
[0080] Specifically, only use the pre-training task MLM (Masked Language Model) to continue pre-training the BERT model using the corpus T. Among them, BERT is a pre-training model proposed by Google in 2018, which is mainly composed of multiple Transformer encoders stacked together. This pre-training language model can be used for various downstream natural language processing tasks, such as: text classification, named entity recognition, text matching and other tasks. The specific application method is: add some structures on the basis of BERT to adapt to the specific task, for example, in the named entity recognition task, CRF can be added after BERT; then, use the specific task data to fine-tune this model.
[0081] Step S4: Construct a text matching model based on the pre-trained BERT model, convert the entity matching corpus marked in the corpus into a text matching corpus, and use the text matching corpus to train the text matching model.
[0082] Among them, the BERT-based text matching model regards the text matching task as a text pair classification task. Specifically, given the matching text t 1 and t2 , use special characters [CLS] and [SEP] to splice two texts together. For example, t 1 is "Text matching is so simple", t 2 is "Text matching is so difficult". Then, the finally spliced text is "[CLS]Text matching is so simple[SEP]Text matching is so difficult[SEP]". Input this text into the text matching model, and then the model predicts 1 for matching and 0 for non-matching.
[0083] The text matching model based on BERT mainly includes: BERT encoding layer and classification layer. The BERT encoding layer is responsible for converting the input text into a vector representation; the classification layer is responsible for judging whether the texts match based on the output of the BERT layer, and it is mainly composed of a fully connected layer with an input of 2 and a Softmax layer. As Figure 7 shown.
[0084] Step S5: Use the trained text matching model to match the entity to be matched;
[0085] Specifically, convert the entity to be matched into a text to be matched, input the text to be matched into the above text matching model, and obtain the entity matching result.
[0086] Among them, the core of using the text matching model to complete entity matching is to convert the entity matching corpus into a text matching corpus. Specifically, given a pair of entities to be matched:
[0087] Entity 1 (Name: Zhang San, Age: 30, Address: Chaoyang District, Beijing, Occupation: Programmer)
[0088] Entity 2 (Name: Zhang San, Age: 31, Address: Haidian District, Beijing, Occupation: Programmer).
[0089] Then Entity 1 can be converted into the text: Zhang San 30 Chaoyang District, Beijing Programmer; Entity 2 is converted into the text: Zhang San 31 Haidian District, Beijing Programmer. Input the converted text to be matched into the text matching model, and the output result is "1 (matching)" or "0 (non-matching)", so that the text matching model can be used to complete entity matching.
[0090] The entity matching method proposed by the present invention constructs an undirected unweighted graph with the attributes of each entity as nodes, and after random walk sampling, the obtained samples will contain the same attributes of different entities with a certain probability. Therefore, this method introduces "attribute context" into the pre-trained language model, thereby improving the effect of the pre-trained language model in entity matching tasks and improving the matching accuracy.
[0091] Example Two:
[0092] Combined with an entity matching method disclosed in Embodiment 1, this embodiment discloses a specific implementation example of an entity matching system (hereinafter referred to as "the system").
[0093] Referring to Figure 9 as shown, the system includes:
[0094] Graph generation module 1: Obtain a target entity set by merging a first entity set and a second entity set, and construct an undirected unweighted graph with the attributes of each entity in the target entity set as nodes;
[0095] Specifically, the graph generation module 1 includes:
[0096] Complete graph construction unit 11: Select entities from the target entity set without replacement, and construct a complete graph with the attributes of the selected entities as nodes;
[0097] Initial undirected unweighted graph obtaining unit 12: Merge the above complete graph into an initialized empty graph to obtain an initial undirected unweighted graph;
[0098] Final undirected unweighted graph obtaining unit 13: Traverse the target entity set, and perform iterative update on the initial undirected unweighted graph to obtain the final undirected unweighted graph.
[0099] Graph sampling module 2: Perform random walk sampling in the above undirected unweighted graph, convert the path obtained after sampling into text, and construct a corpus using this text;
[0100] Specifically, the graph sampling module 2 includes:
[0101] Random sampling unit 21: Perform random walk sampling in the final undirected unweighted graph;
[0102] Path conversion unit 22: Convert the path obtained by each random walk sampling into text, and repeat random walk sampling until the corpus is constructed.
[0103] Continuous pre-training module 3: Use the above corpus to perform continuous pre-training on the BERT model;
[0104] Training module 4: Construct a text matching model according to the pre-trained BERT model, convert the labeled entity matching corpus in the corpus into text matching corpus, and use the text matching corpus to train the text matching model;
[0105] Entity matching module 5: Use the trained text matching model to match the entities to be matched.
[0106] Hereinafter, the entity matching system proposed by the present invention will be further described in detail in combination with specific embodiments. It includes:
[0107] I. Graph Generation Module
[0108] The goal of this module is to construct a unified undirected unweighted graph G from entity sets D and D'. The specific steps are as follows:
[0109] (1) Initialize an empty graph G;
[0110] (2) Merge entity sets D and D', that is
[0111] (3) Randomly select an entity e without replacement from , that is i ;
[0112] (4) Take all the attributes of entity e i e i [A 1 , …, e i [A m as nodes to construct a complete graph G i ;
[0113] (5) Merge G i into graph G. Specifically: Add the nodes in G i that do not exist in G to G, and add the corresponding edges to G;
[0114] (6) If , then go to step (3); otherwise, return the constructed undirected unweighted graph G; as shown in Figure 5 for example.
[0115] II. Graph Sampling Module
[0116] The goal of this module is to sample G through random walk to construct a corpus for pre-training. Taking a single sampling as an example, the working principle of this module is introduced as follows:
[0117] (1) Sample a node from graph G in a uniform sampling manner, called n 0 ;
[0118] (2) Taking n 0 as the center, sample its adjacent nodes uniformly to obtain node n 1 ;
[0119] (3) Taking n 1 as the center, sample its adjacent nodes uniformly to obtain node n 2 ;
[0120] (4) And so on; when sampling reaches , each sampling terminates with a manually specified probability p, where l 1is the minimum sampling length set manually;
[0121] (5). When the sampling exceeds , the sampling is directly terminated, where l 2 is the maximum sampling length set manually;
[0122] (6). Convert the sampled path n 0 , n 1 , … into text;
[0123] Repeat the above steps until enough text is sampled and form a corpus T;
[0124] Figure 6 is an example of a single sampling, with the starting node being "Zhang San", and the entire path being indicated by a dashed arrow.
[0125] III. Continuing pre-training module
[0126] This module uses the generated corpus T to perform continuing pre-training on BERT, and only uses the pre-training task MLM (Masked Language Model).
[0127] IV. Entity matching module
[0128] The goal of this module is to complete the entity matching task. Specifically, it includes the following steps:
[0129] (1) Use the model BERT output by the "continuing pre-training module" to construct a text matching model M;
[0130] (2) Convert the labeled entity matching corpus into a text matching corpus;
[0131] (3) Use the text matching corpus to train the model M;
[0132] For the technical solutions of the same parts in an entity matching system disclosed in this embodiment and an entity matching method disclosed in Embodiment 1, please refer to what is described in Embodiment 1 and will not be elaborated here.
[0133] Embodiment 3:
[0134] Combined with Figure 10 as shown, this embodiment discloses a specific implementation of a computer device. The computer device may include a processor 81 and a memory 82 storing computer program instructions.
[0135] Specifically, the above-mentioned processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits implementing the embodiments of the present application.
[0136] Among them, the memory 82 may include a mass storage for data or instructions. By way of example and not limitation, the memory 82 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 82 may include removable or non-removable (or fixed) media. Where appropriate, the memory 82 may be internal or external to the data processing device. In a particular embodiment, the memory 82 is non-volatile memory. In a particular embodiment, the memory 82 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable read-only memory (EAROM), or a flash memory, or a combination of two or more of these. Where appropriate, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended date out dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0137] The memory 82 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81.
[0138] The processor 81 reads and executes the computer program instructions stored in the memory 82 to implement any one of the entity matching methods in the above embodiments.
[0139] In some of the embodiments, the computer device may further include a communication interface 83 and a bus 80. Among them, as Figure 10 shown, the processor 81, the memory 82, and the communication interface 83 are connected through the bus 80 and complete communication with each other.
[0140] The communication interface 83 is used to implement communication between the various modules, devices, units, and / or devices in the embodiments of the present application. The communication port 83 can also implement data communication with other components such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations, etc.
[0141] Bus 80 includes hardware, software, or both, and couples components of a computer device to each other. Bus 80 includes, but is not limited to, at least one of the following: Data Bus, Address Bus, Control Bus, Expansion Bus, Local Bus. By way of example and not limitation, Bus 80 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable bus or a combination of two or more of these. Where appropriate, Bus 80 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.
[0142] In addition, in combination with the entity matching method in the above embodiments, an embodiment of the present application can be implemented by providing a computer-readable storage medium. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by a processor, any one of the entity matching methods in the above embodiments is implemented.
[0143] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0144] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. An entity matching method, characterized in that, it includes: Obtaining a target entity set by merging a first entity set and a second entity set, and respectively using the attributes of each entity in the target entity set as nodes to construct an undirected unweighted graph; Performing random walk sampling in the undirected unweighted graph, converting the paths obtained after sampling into texts, and constructing a corpus using the texts; Continuing to pre-train the BERT model using the corpus; Constructing a text matching model according to the pre-trained BERT model, converting the entity matching corpus labeled in the corpus into text matching corpus, and training the text matching model using the text matching corpus; Using the trained text matching model to match the entities to be matched; wherein, the constructing of the undirected unweighted graph includes: Selecting entities without replacement from the target entity set, and using the attributes of the selected entities as nodes to construct a complete graph; Merging the complete graph into an initialized empty graph to obtain a preliminary undirected unweighted graph; Traversing the target entity set, and iteratively updating the preliminary undirected unweighted graph to obtain the final undirected unweighted graph; wherein, the constructing of the corpus includes: Performing random walk sampling in the final undirected unweighted graph; Converting the paths obtained by each random walk sampling into texts, and repeating the random walk sampling until the corpus is constructed.
2. The entity matching method according to claim 1, characterized in that, the random walk sampling specifically includes: when the random walk sampling reaches between the set minimum sampling length and the set maximum sampling length, each sampling terminates with a set probability, and when the random walk sampling reaches the set maximum sampling length, the sampling terminates.
3. The entity matching method according to claim 1, characterized in that, the continuing pre-training includes: continuing to pre-train the BERT model using the pre-training task MLM based on the constructed corpus.
4. The entity matching method according to claim 1, characterized in that, the matching of the entities to be matched includes: converting the entities to be matched into texts to be matched, inputting the texts to be matched into the text matching model, and obtaining an entity matching result.
5. An entity matching system, characterized in that, it includes: A graph generation module: obtaining a target entity set by merging a first entity set and a second entity set, and respectively using the attributes of each entity in the target entity set as nodes to construct an undirected unweighted graph; A graph sampling module: performing random walk sampling in the undirected unweighted graph, converting the paths obtained after sampling into texts, and constructing a corpus using the texts; A continuing pre-training module: using the corpus to continue pre-training the BERT model; A training module: constructing a text matching model according to the pre-trained BERT model, converting the entity matching corpus labeled in the corpus into text matching corpus, and training the text matching model using the text matching corpus; An entity matching module: using the trained text matching model to match the entities to be matched; wherein, the graph generation module includes: Complete graph construction unit: Entities are selected without replacement from the set of target entities, and the attributes of the selected entities are used as nodes to construct a complete graph; Initial undirected unweighted graph acquisition unit: The complete graph is merged into an initially empty graph to obtain an initial undirected unweighted graph; Final undirected unweighted graph acquisition unit: The target entity set is traversed, and the initial undirected unweighted graph is iteratively updated to obtain a final undirected unweighted graph; Among them, the graph sampling module includes: Random walk sampling is performed in the final undirected unweighted graph; The path obtained by each random walk sampling is converted into text, and random walk sampling is repeated until the corpus construction is completed.
6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, when the processor executes the computer program, the entity matching method described in any one of claims 1 to 4 is implemented.
7. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the program is executed by a processor, the entity matching method described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Author name disambiguation method based on heterogeneous graph convolutional neural network embedding
CN110516146A
Knowledge representation learning method and device, equipment and storage medium
CN111475658A