Data Processing Method, Apparatus, Device, and Medium
Through the entity matching model of cross-attribute symbol comparison and error correction of symbol comparison, the impact of dirty data on entity matching accuracy is solved, the accuracy of deduplication and fusion processing of the data set is improved, and the data quality is improved.
Patent Information
- Application Number
- CN202210203266.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-03
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-03-03
AI Technical Summary
Existing entity matching techniques rely on the data quality of entity records, resulting in low accuracy of entity matching in the presence of dirty data, which in turn affects the quality of the data set.
Through the entity matching model, symbol comparison across attributes and error correction processing are performed, attributes in entity records are ignored for global comparison, information misalignment and errors are corrected, and the accuracy of entity matching is improved.
Improve the accuracy of entity matching, thereby improving the accuracy of data deduplication and fusion processing, and improving the quality of the data set.
Smart Images

Figure CN114595293B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technologies, and in particular, to a data processing method, apparatus, device, and medium. Background Art
[0002] Entity Matching, also known as Entity Alignment or Entity Resolution, aims to determine whether entity descriptions from different sources refer to the same entity or object in the real world. Among them, an entity description, also known as an Entity Record, is a structured object containing multiple attributes.
[0003] In the field of data processing, entity matching is applied to reduce redundant entity records and / or obtain more accurately described entity records to improve the quality of the dataset where the entity records are located. Currently, entity matching can be modeled as a binary classification task. In this task, the attribute values of two entity records are compared to obtain a comparison vector, and then based on the comparison vector, the classification result of the pair of two entity records is obtained, that is, the two entity records "match" or "do not match".
[0004] However, the accuracy of entity matching in the above manner is overly dependent on the data quality of entity records. In actual situations, most entity records have varying degrees of dirty data, resulting in low accuracy of entity matching in the above manner, and further resulting in low data quality of the dataset after entity matching processing. Summary of the Invention
[0005] Multiple aspects of this application provide a data processing method, apparatus, device, and medium to solve the problem of low data quality of the dataset after entity matching processing.
[0006] In a first aspect, an embodiment of this application provides a data processing method, including: obtaining a first dataset, where the first dataset contains multiple entity records; performing entity matching processing on the multiple entity records through an entity matching model to obtain a matching result, where the entity matching processing includes cross-attribute symbol comparison and error correction of the comparison result of the symbol comparison; and performing deduplication processing and / or fusion processing on the first dataset according to the matching result to obtain a second dataset.
[0007] In a second aspect, an embodiment of the present application provides a data processing device, including: an acquisition unit configured to acquire a first data set, where the first data set includes a plurality of entity records; an entity matching unit configured to perform entity matching processing on the plurality of entity records through an entity matching model to obtain a matching result, where the entity matching processing includes cross-attribute symbol comparison and error correction of the comparison result of the symbol comparison; and a data set processing unit configured to perform duplicate removal processing and / or fusion processing on the first data set according to the matching result to obtain a second data set.
[0008] In a third aspect, an embodiment of the present application provides a cloud server, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the cloud server can execute the data processing method provided in the first aspect.
[0009] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it is the data processing method provided in the first aspect.
[0010] In an embodiment of the present application, a first data set including a plurality of entity records is acquired; entity matching processing is performed on the plurality of entity records in the first data set through an entity matching model to obtain a matching result, where the entity matching processing includes cross-attribute symbol comparison and error correction of the comparison result of the symbol comparison; and duplicate removal processing and / or fusion processing is performed on the first data set according to the matching result to obtain a second data set. Thus, through cross-attribute symbol comparison, the entity structure reflected by the attributes in the entity records is ignored, and global comparison of symbols in the entity record pairs is achieved. By performing error correction on the comparison result of the symbol comparison, the adverse impact of dirty data on the accuracy of entity matching is solved, the accuracy of entity matching is effectively improved, and further the accuracy of duplicate removal processing and / or fusion processing on the data set is improved, and the quality of the processed data is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0012] Figure 1 is a schematic diagram of an application scenario provided according to an embodiment of the present application;
[0013] Figure 2 is a schematic flowchart of the data processing method provided by an embodiment of the present application;
[0014] Figure 3aStructural schematic of the entity matching model provided by the embodiments of the present application Figure 1 ;
[0015] Figure 3b Flow schematic of entity matching processing in the data processing method provided by the embodiments of the present application;
[0016] Figure 3c Structural schematic of the entity matching model provided by the embodiments of the present application Figure 2 ;
[0017] Figure 4 Structural example diagram of the entity matching model provided by the embodiments of the present application;
[0018] Figure 5 Block diagram of the structure of the data processing device 500 provided by an embodiment of the present application;
[0019] Figure 6 Structural schematic of a cloud server provided by an exemplary embodiment of the present application. Detailed implementation manners
[0020] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0021] First, some terms in the embodiments of the present application are explained:
[0022] Entity: An entity refers to an object that actually exists in the real world. For example, an actually existing commodity belongs to an entity, and an entity can also be called an object.
[0023] Entity matching: Entity matching refers to matching entity descriptions from different sources to determine whether the entity descriptions from different sources point to the same entity or object in the real world. If the entity descriptions from different sources point to the same entity or object in the real world, then the entity descriptions from different sources pointing to the same entity or object can be fused (i.e., information complementarity) to obtain a more accurate and more detailed entity description of the entity or object, or data deduplication can be performed on the entity descriptions from different sources pointing to the same entity or object to reduce data redundancy and improve data quality. In addition, a knowledge graph (such as a commodity knowledge graph) can be constructed based on the fused and / or deduplicated entity descriptions. The nodes in the knowledge graph are entities, and the edges between the nodes in the knowledge graph are entity descriptions and the relationships between entities.
[0024] Entity record: That is, entity description, which refers to a structured object containing multiple attributes. In the structured object, each attribute stores a corresponding attribute value, and the attribute value consists of several symbols. Among them, an entity record can be expressed as <Attribute 1, Attribute value of Attribute 1>, <Attribute 2, Attribute value of Attribute 2>, <Attribute 3, Attribute value of Attribute 3>,....
[0025] Dirty entity: An entity in an entity record that has one or more data quality problems such as information misalignment, information loss, and information error. Among them, information misalignment means that the attribute value of a certain attribute appears in the attribute value of another attribute, information loss means an attribute with a missing attribute value, and information error means an incorrect attribute value corresponding to an attribute.
[0026] Table 1
[0027] Title Category Brand Flavor Specification Barcode A-brand vitamin drink with nationwide free shipping Lime-flavored functional drink Brand B 600 ml
[0028] As an example, Table 1 gives the structural schematic of the entity record of a certain beverage product. As shown in Table 1, the first row lists multiple attributes of the beverage product: title, category, brand, flavor, specification, and barcode. The second row shows the attribute values of these multiple attributes. The attribute values include several symbols. For example, in the attribute value "600 milliliters", "600" is one symbol and "milliliters" is one symbol. Among them, the attribute value "lime flavor" of "flavor" appears in the attribute value of "category", resulting in information misalignment; the attribute value of "brand" is "Brand B", but from the title, it can be concluded that the correct attribute value of "brand" is "Brand A", resulting in information error; the attribute value of "barcode" is missing, resulting in information loss.
[0029] Entity record pair: It contains two entity records to be matched. The two entity records in the entity record pair may point to the same entity or object, or they may point to different entities or objects.
[0030] Business knowledge graph: A semantic network description that depicts the attributes of commodity entities and the relationships between commodity entities. Among them, in the commodity knowledge graph, nodes represent commodity entities or concepts (such as TV), and edges are composed of entity attributes or relationships. In the field of commodity knowledge graph, one entity record pair includes two commodity entity records, and each commodity entity record is a structured description of a commodity information. For example, Table 1 is a commodity entity record.
[0031] In related technologies, the accuracy of entity matching models relies on the high data quality of the input entity records. However, in practical applications, the data sources of entity records are diverse, the data quality varies, and there will be varying degrees of dirty data. In particular, entity records extracted from unstructured / semi-structured web pages (for example, product entity information extracted from e-commerce platforms, user description information extracted from social media, etc.) are prone to information dislocation, information missing, and information errors, which have a significant impact on the performance of the entity matching model, as follows:
[0032] (1) Information misalignment and missing information often lead to structural information damage in entity records, causing the entity matching model to mistakenly focus on some unimportant noise information or pay insufficient attention to some important information with high recognition;
[0033] (2) Information errors can easily lead to semantic noise, causing inconsistent or even contradictory semantic expressions in the same entity record, misleading the entity model to make incorrect judgments.
[0034] Therefore, how to effectively identify and eliminate the interference caused by dirty data to the entity matching model is a key challenge faced by the entity matching task in practical application scenarios.
[0035] To alleviate the interference of dirty data in entity records on the entity matching model, there are two solutions:
[0036] (1) Conduct symbol comparison across attributes;
[0037] (2) Introducing data enhancement in the entity matching model.
[0038] Among them, improving the symbol comparison within the attribute to a symbol comparison across attributes, that is, a global comparison on the entity record, can alleviate the impact of attribute value misalignment, that is, information misalignment, on entity matching to a certain extent, but it cannot solve the impact of missing information and information errors on entity matching. Introducing data enhancement in the entity matching model can train a more robust entity matching model. However, the core idea of current data enhancement is to force the model to learn more comprehensive matching features by generating lower-quality training data, so that it is not easy to overfit some simple features. In essence, this type of data enhancement method cannot correct problems such as structural damage and semantic noise caused by dirty data to entity records.
[0039] Embodiments of the present application propose a data processing method, apparatus, device, and medium to better solve the problem that the above-mentioned dirty data leads to low accuracy of entity matching. In the embodiments of the present application, a first data set is obtained, and the first data set contains multiple entity records; entity matching processing is performed on the multiple entity records through an entity matching model to obtain a matching result. The entity matching processing includes cross-attribute symbol comparison and error correction of the comparison result of the symbol comparison. Among them, the cross-attribute symbol comparison, that is, the global comparison of symbols in the entity record pair, ignores the attributes in the entity record during the comparison process, solves the adverse effect of information misalignment on entity matching, and corrects the comparison result of the symbol comparison to further solve the adverse effect of dirty data on entity matching, improving the accuracy of entity matching based on the entity matching model; finally, according to the matching result, deduplication processing and / or fusion processing are performed on the first data set to obtain a second data set. Thus, by improving the accuracy of entity matching, the accuracy of data deduplication processing and / or fusion processing is improved, and the quality of the processed data is improved.
[0040] The technical solution provided by the embodiments of the present application can be applied to the entity information processing scenario. In particular, it can be applied to the scenario of processing commodity entity information in the construction of a commodity knowledge graph. In this scenario, entity information of commodity entities can be collected under authorization to obtain entity records of commodity entities, and entity matching is performed on different entity records to determine entity records belonging to the same commodity entity, thereby improving the accuracy of the construction of the commodity knowledge graph.
[0041] Exemplarily, Figure 1 FIG. is a schematic diagram of an application scenario provided according to an embodiment of the present application. As Figure 1 shown, the device designed in this application scenario includes a client and a data processing device ( Figure 1 taking the data processing device as a server as an example). The client can send a first data set containing multiple entity records to the data processing device. The data processing device performs entity matching processing on the entity records in the first data set, and performs deduplication processing and / or fusion processing on the first data set based on the matching result to obtain a second data set to improve the data quality of the second data set. The data processing device can return the second data set to the client or store the second data set.
[0042] Taking the entity records A and B included in the entity record pair as an example, in the data processing device, the entity matching process includes: performing cross-attribute symbol comparison on the entity records A and B, correcting the symbol comparison result, and obtaining the matching result of the entity records A and B, that is, the result of whether the entity records A and B belong to the same entity. If the entity records A and B belong to the same entity, duplicate removal processing and / or fusion processing can be performed on the entity records A and B to remove duplicate entity records in the first dataset and / or obtain a more accurate entity record of the entity to which the entity records A and B belong.
[0043] Next, in combination with the Figure 1 application scenario shown above, the technical solution of the present application will be described in detail through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0044] It should be noted that the execution subject of the embodiments of the present application can be an electronic device, and the electronic device can be a terminal or a server. Among them, the terminal can be a personal digital assistant (PDA) device, a handheld device with wireless communication function (such as a smart phone, a tablet computer), a computing device (such as a personal computer (PC)), a vehicle-mounted device, a wearable device (such as a smart watch, a smart bracelet), a smart home device (such as a smart display device), etc. Among them, the server can be a single server, or a server cluster, can be a distributed server, or a centralized server, and can also be a cloud server.
[0045] Refer to Figure 2 , Figure 2 which is a schematic flowchart of the data processing method provided by the embodiments of the present application. As Figure 2 shown, the data processing method includes:
[0046] S201. Obtain a first dataset, where the first dataset contains multiple entity records
[0047] Among them, in the first dataset, each entity record is a structured object containing multiple attributes, and the entity record also contains the attribute values corresponding to the attributes. An attribute value can include one or more symbols. Specifically, reference can be made to the example of the entity record given in Table 1 above, which will not be repeated here.
[0048] In this embodiment, the first data set collected in advance can be obtained from a database; alternatively, with authorization, entity records can be collected on the network to obtain the first data set. For example, after obtaining the authorization of an e-commerce platform, entity records of products can be collected on the e-commerce platform; alternatively, the first data set input by a user or sent by other devices can be obtained.
[0049] S202. Perform entity matching processing on multiple entity records through an entity matching model to obtain a matching result. The entity matching processing includes cross-attribute symbol comparison and error correction of the comparison results of the symbol comparison.
[0050] In the first data set, two entity records may point to the same entity or object in the real world, or may belong to the same entity or object pointing to a real event. Two entity records pointing to the same entity or object may have exactly the same data or complementary data. For example, in two entity records collected from different e-commerce platforms, one entity record contains the price of a product, while the other entity record does not. Therefore, there are problems of data redundancy in the first data set and inaccurate data in the entity records. Therefore, it is necessary to perform entity matching processing on the first data set to determine whether two entity records in the first data set point to the same entity or object, and then, based on the matching result, perform deduplication processing and / or fusion processing on the first data set to improve the data quality of the first data set.
[0051] In this embodiment, in the entity matching processing of the first data set, an entity matching model can be used to perform entity matching processing on multiple entity records in the first data set to obtain a matching result. Among them, the matching result may include whether two entity records in the first data set point to the same entity or object; alternatively, the matching result may include all entity records in the first data set that point to the same entity or object. For example, in the matching result, all entity records pointing to entity a include entity record A and entity record B, and all entity records pointing to entity b include entity record C, entity record D, and entity record E.
[0052] The process of performing entity matching processing on multiple entity records using an entity matching model may include:
[0053] First, the attributes in the entity record to be matched can be ignored, and cross-attribute symbol comparison can be performed on the entity record to be matched. That is, the symbols under the attributes in the entity record can be compared with the symbols under the same and different attributes in another entity record for entity matching, and the comparison result of the symbol comparison can be obtained. For example, the symbol "Global Free Shipping A Brand Vitamins" under "Title" in the entity record shown in Table 1 can be compared with the symbols under multiple attributes such as "Title", "Category", and "Brand" in another entity record. Compared with only performing symbol comparison within the same attribute, cross-attribute symbol comparison is equivalent to global symbol comparison, which solves the adverse effects of information misalignment on entity matching and improves the accuracy of entity matching.
[0054] Next, in the entity matching model, after obtaining the comparison result of the symbol comparison, the comparison result of the symbol comparison can be corrected to solve the adverse effects of information loss and information error on entity matching, and further solve the adverse effects of information misalignment on entity matching, thereby improving the accuracy of entity matching.
[0055] S203. According to the matching result, perform deduplication processing and / or fusion processing on the first data set to obtain a second data set.
[0056] Among them, the second data set can be the first data set subjected to deduplication processing and / or fusion processing, including multiple entity records; or, the second data set can be a knowledge graph obtained based on the first data set subjected to deduplication processing and / or fusion processing. Here, the process of constructing a knowledge graph based on multiple entity records will not be elaborated and limited.
[0057] In this embodiment, after obtaining the matching result, duplicate entity records can be deleted from the entity records that match in the first data set (i.e., entity records pointing to the same entity or object) (for example, only one entity record is retained among the matching entity records), so as to achieve deduplication processing of the first data set and reduce data redundancy in the first data set; and / or, the entity records that match in the first data set can be merged into one entity record to improve the data completeness and accuracy of the entity records, and at the same time, data redundancy in the first data set can also be reduced. In this way, a second data set with higher data quality is obtained.
[0058] In the embodiment of the present application, in the process of entity matching of multiple entity records in the first data set by using the entity matching model, cross-attribute symbol comparison and correction of the comparison result of the symbol comparison are adopted to solve the adverse effects of dirty data on entity matching, improve the accuracy of entity matching, and further improve the accuracy of deduplication processing and / or fusion processing of the first data set, and improve the data quality of the processed data.
[0059] In some embodiments, error correction is performed on the comparison results of symbol comparison, including: performing entity structure reconstruction and semantic noise reduction on the comparison results of symbol comparison. Among them, entity structure reconstruction is beneficial to solving the adverse effects of attribute errors, that is, information misalignment, on entity matching, and semantic noise reduction is beneficial to solving the adverse effects of attribute value errors, that is, information errors and information missing, on entity matching. Thus, through these two error correction methods of entity structure reconstruction and semantic noise reduction, the accuracy of entity matching is improved.
[0060] Reference Figure 3a , Figure 3a is a schematic diagram of the structure of the entity matching model provided by the embodiments of the present application Figure 1 . As Figure 3a shown, the entity matching model includes a comparison network layer, a structure reconstruction network layer, a semantic noise reduction network layer, and an aggregation network layer. Among them:
[0061] In the entity matching model: the comparison network layer is used to compare each symbol in the symbol sequence corresponding to the entity record in the entity record pair, and obtain an initial comparison vector corresponding to each symbol in the symbol sequence; the structure reconstruction network layer is used to perform entity structure reconstruction on the initial comparison vector corresponding to each symbol to obtain an intermediate comparison vector corresponding to each symbol; the semantic noise reduction network layer is used to perform semantic noise reduction on the intermediate comparison vector corresponding to each symbol to obtain an enhanced comparison vector corresponding to each symbol; the aggregation network layer is used to perform aggregation processing on the enhanced comparison vectors corresponding to each symbol to obtain the matching result of the entity record pair, that is, the result of whether the two entity records in the entity record pair match successfully or fail to match.
[0062] Based on Figure 3a the model structure shown, Figure 3b is a schematic diagram of the process of entity matching processing in the data processing method provided by the embodiments of the present application Figure 1 . As Figure 3b shown, the process of entity matching processing includes:
[0063] S301. In the entity record pair, according to the attribute values corresponding to multiple attributes in the entity record, determine the symbol sequence corresponding to the entity record.
[0064] Among them, multiple entity records in the first dataset can form at least one entity record pair, and the entity record pair includes two entity records. Performing entity matching processing on multiple entity records in the first dataset includes performing entity matching processing on one or more entity record pairs formed by multiple entity records in the first dataset, and the matching result obtained by performing entity matching processing on multiple entity records in the first dataset may include the matching result of the entity record pair.
[0065] In this embodiment, in the entity record pair, for each entity record, the attribute values corresponding to multiple attributes in the entity record can be obtained, and according to the symbols in the multiple attribute values, a symbol sequence corresponding to the entity record can be obtained. For example, based on the entity record shown in Table 1, the symbol sequence "Free shipping nationwide, A-brand vitamin drink, lime flavor, functional drink, B-brand, 600 ml" can be obtained. Thus, each entity record in the entity record pair can be represented as a symbol sequence without attributes, that is, a symbol sequence without the structural information of the entity record, which helps to realize the global comparison of symbols in the subsequent process, without being limited to the symbol comparison under the same attribute, and further helps to solve the problem of low accuracy of entity matching caused by information dislocation.
[0066] Among them, the entity records in the entity record pair can be respectively referred to as the first entity record and the second entity record. Therefore, the attribute values corresponding to multiple attributes in the first entity record can be obtained, and according to the attribute values corresponding to the multiple attributes, the first symbol sequence corresponding to the first entity record can be determined; the attribute values corresponding to multiple attributes in the second entity record can be obtained, and according to the attribute values corresponding to the multiple attributes, the second symbol sequence corresponding to the second entity record can be determined.
[0067] S302. Through the comparison network layer, each symbol in the symbol sequence is compared to obtain the initial comparison vector corresponding to each symbol in the symbol sequence.
[0068] In this embodiment, as Figure 3a shown, the symbol sequences corresponding to the entity records in the entity record pair can be input into the comparison network layer. In the comparison network layer, symbol-level comparison is performed on the symbol sequences corresponding to different entity records to obtain the initial comparison vector corresponding to each symbol in the symbol sequence. That is, each symbol in the first symbol sequence corresponding to the first entity record is compared with each symbol in the second symbol sequence corresponding to the second entity record to obtain the initial comparison vector corresponding to each symbol in the first symbol sequence and the initial comparison vector corresponding to each symbol in the second symbol sequence, realizing the global comparison of symbols.
[0069] Among them, the data form of the initial comparison vector corresponding to the symbol is a vector, that is, the initial comparison vector corresponding to the symbol is the initial comparison vector corresponding to the symbol. The element value in the initial comparison vector is the value obtained by comparing the symbol with the corresponding symbol in another symbol sequence. For example, in the initial comparison vector corresponding to the first symbol in the first symbol sequence, the first element value is the value obtained by comparing the first symbol in the first symbol sequence with the first symbol in the second symbol sequence, and the second element value is the value obtained by comparing the first symbol in the first symbol sequence with the second symbol in the second symbol sequence, and so on.
[0070] Optionally, the value obtained by comparing symbol with symbol can be a value reflecting the similarity degree between symbols. The larger the value, the higher the similarity degree, such as the similarity between symbols; or, it can also be a value reflecting the difference degree between symbols. The larger the value, the higher the difference degree, such as the difference between symbols.
[0071] S303. Through the structure reconstruction network layer, perform entity structure reconstruction on the initial comparison vector to obtain the intermediate comparison vectors corresponding to each symbol in the symbol sequence.
[0072] Among them, the intermediate comparison vector is the initial comparison vector after volume structure reconstruction.
[0073] In this embodiment, as Figure 3a shown, each symbol sequence and the initial comparison vectors corresponding to each symbol in the symbol sequence can be input into the structure reconstruction network layer. In the structure reconstruction network layer, the attributes corresponding to each symbol in the symbol sequence can be corrected. Since the attributes corresponding to the symbols reflect the entity structure, the corrected attributes corresponding to each symbol can be added to the initial comparison vectors corresponding to each symbol to achieve entity structure reconstruction of the initial comparison vectors corresponding to each symbol.
[0074] Specifically, the attributes corresponding to each symbol in the first symbol sequence can be corrected to obtain the corrected attributes corresponding to each symbol in the first symbol sequence, and the corrected attributes corresponding to each symbol in the first symbol sequence are added to the initial comparison vectors corresponding to each symbol in the first symbol sequence to obtain the intermediate comparison vectors corresponding to each symbol in the first symbol sequence; similarly, the attributes corresponding to each symbol in the second symbol sequence can be corrected to obtain the corrected attributes corresponding to each symbol in the second symbol sequence, and the corrected attributes corresponding to each symbol in the second symbol sequence are added to the initial comparison vectors corresponding to each symbol in the second symbol sequence to obtain the intermediate comparison vectors corresponding to each symbol in the second symbol sequence.
[0075] Thus, by correcting the attributes corresponding to all symbols in the structure reconstruction network layer, that is, correcting the entity structures corresponding to all symbols, and adding the correct entity structures to the initial comparison vectors, the adverse effect of attribute errors on the accuracy of entity matching is reduced, that is, the adverse effect of information misalignment on the accuracy of entity matching is reduced.
[0076] In addition, the importance of different attributes is different. Adding the correct entity structure to the initial comparison vector is equivalent to adding reference information reflecting the importance of symbols to the initial comparison vector, enabling the subsequent aggregation layer to distinguish the importance of different symbols when aggregating multiple comparison vectors, paying more attention to the comparison vectors of symbols corresponding to whether the corresponding entity structures match more discriminative attributes, and obtaining a more reliable entity matching result.
[0077] S304. Through the semantic noise reduction network layer, perform semantic noise reduction on the intermediate comparison vectors to obtain enhanced comparison vectors corresponding to each symbol in the symbol sequence.
[0078] Among them, the enhanced comparison vector is the intermediate comparison vector after semantic noise reduction.
[0079] In this embodiment, as Figure 3a shown, each symbol sequence and the intermediate comparison vectors corresponding to each symbol in the symbol sequence can be input into the semantic noise reduction network layer. In the semantic noise reduction network layer, the semantic correctness of each symbol in the symbol sequence can be detected to obtain the correct rate of each symbol in the symbol sequence; according to the correct rate of each symbol in the symbol sequence, the intermediate comparison vectors of each symbol are adjusted to obtain the enhanced comparison vectors corresponding to each symbol.
[0080] Specifically, in the semantic noise reduction network layer, the semantic correctness of each symbol in the first symbol sequence can be detected to obtain the correct rate of each symbol in the first symbol sequence, and according to the correct rate of each symbol in the first symbol sequence, the intermediate comparison vectors corresponding to each symbol in the first symbol sequence are adjusted to obtain the enhanced comparison vectors corresponding to each symbol in the first symbol sequence; the semantic correctness of each symbol in the second symbol sequence can be detected to obtain the correct rate of each symbol in the second symbol sequence, and according to the correct rate of each symbol in the second symbol sequence, the intermediate comparison vectors corresponding to each symbol in the second symbol sequence are adjusted to obtain the enhanced comparison vectors corresponding to each symbol in the second symbol sequence.
[0081] Thus, by detecting the semantic correctness of all symbols in the semantic noise reduction network layer and adjusting the intermediate comparison vectors corresponding to the symbols according to the detected semantic correctness, the influence of symbol errors (i.e., attribute value errors) on the accuracy of entity matching is reduced, that is, the influence of information errors and information loss on the accuracy of entity matching is reduced.
[0082] S305. Through the aggregation network layer, perform aggregation processing on the enhanced comparison vectors to obtain the matching result of the entity record pair.
[0083] In this embodiment, as Figure 3aAs shown, the enhanced contrast vectors corresponding to each symbol can be input into the aggregation network layer. In the aggregation network layer, the enhanced contrast vectors corresponding to each symbol are aggregated to obtain the matching result of the entity record pair. Among them, the matching result of the entity record pair can be that the first entity record and the second entity record match successfully, indicating that the first entity record and the second entity record belong to the same entity; or, the matching result of the entity record pair can be that the first entity record and the second entity record match fails, indicating that the first entity record and the second entity record belong to different entities; or, the matching result of the entity record pair can be the probability value that the first entity record and the second entity record belong to the same entity, and it can be further determined whether the first entity record and the second entity record belong to the same entity according to this probability value.
[0084] Among them, the aggregation operation of the enhanced contrast vectors corresponding to each symbol in the aggregation network layer, for example, is an operation such as vector weighting, summing, or averaging, and is not limited here.
[0085] In the embodiment of the present application, in the entity matching model, a contrast network layer for global symbol contrast is introduced to reduce the adverse impact of attribute value misalignment, that is, information error, on entity matching. A structure reconstruction network layer and a semantic noise reduction network layer for self-correction are introduced to reduce the adverse impact of dirty data problems such as attribute errors, attribute value errors, and attribute value missing on entity matching, so as to ensure the accuracy of entity matching even when there is dirty data in the entity record.
[0086] Next, multiple network structure diagrams of the entity matching model are provided, and feasible solutions for entity structure reconstruction and semantic denoising are provided. Since the first symbol sequence and the second symbol sequence go through the same processing process in entity structure reconstruction and semantic denoising, the subsequent processing processes of the first symbol sequence and the second symbol sequence will not be described separately.
[0087] Refer to Figure 3c , Figure 3c which is the structural schematic of the entity matching model provided by the embodiment of the present application Figure 2 .
[0088] (1) Optional Model Structure 1
[0089] An optional model structure: As Figure 3c shown, on the basis of the entity matching model shown in Figure 3a , the structure reconstruction network layer may include a first language model and a first context information extraction network. Among them:
[0090] The first language model is used to perform feature encoding on the symbol sequence in the input structure reconstruction network layer to obtain a first encoded representation corresponding to the symbol sequence; the first context information extraction network is used to extract context information from the first encoded representation to obtain a first context representation corresponding to the symbol sequence. Among them, the first language model can adopt a pre-trained language model. During the training process of the entity matching model, there is no need to train the first language model.
[0091] At this time, a possible implementation of S303 includes: determining a first context representation corresponding to the symbol sequence through the first language model and the first context information extraction network; predicting the attributes of each symbol in the symbol sequence according to the first context representation corresponding to the symbol sequence to obtain an attribute prediction vector corresponding to each symbol in the symbol sequence; and reconstructing the entity structure of each symbol in the symbol sequence according to the attribute prediction vector corresponding to each symbol in the symbol sequence to obtain an intermediate comparison vector corresponding to each symbol in the symbol sequence.
[0092] In the contrast network layer, global symbol contrast is performed by ignoring the entity structure. However, considering that the entity structure is very important for the entity matching task, in order to improve the accuracy of entity matching, it is still necessary to apply the entity structure to entity matching. Due to problems such as misalignment, missing, and error of attribute values, the entity structure in the entity record of a dirty entity is often damaged, resulting in insufficient effectiveness and reliability of the entity matching result. Therefore, during the process of applying the entity structure to entity matching, it is necessary to correct the entity structure to reduce the adverse impact of the damaged entity structure on the accuracy of entity matching.
[0093] In this implementation, to solve these problems, considering that the symbols in the symbol sequence can provide discriminative context information and the context information can reflect the correct entity structure, a structure reconstruction network layer is introduced into the entity matching model, and a first language model and a first context information extraction network are introduced into the structure reconstruction network layer. The context information in the symbol sequence is extracted through the first language model and the first context information extraction network to obtain a first context representation of the symbol sequence. Among them, the first context representation of the symbol sequence includes the first context representation of each symbol in the symbol sequence.
[0094] Specifically, the symbol sequence can be input into the first language model, and the first language model performs feature encoding on the symbol sequence to obtain a first encoded representation corresponding to the symbol sequence; the first encoded representation corresponding to the symbol sequence is input into the first context information extraction network, and the first context information extraction network extracts context information from the first encoded representation corresponding to the symbol sequence to obtain a first context representation corresponding to the symbol sequence.
[0095] After that, for each symbol in the symbol sequence, the following operations can be performed respectively: Based on the first context representation of the symbol, the correct attributes of the symbol can be predicted to obtain the probabilities of the symbol belonging to multiple attributes in the entity record corresponding to the symbol sequence; Based on the probabilities of the symbol belonging to multiple attributes, the attribute prediction vector corresponding to the symbol can be obtained; By fusing the attribute prediction vector of the symbol with the initial comparison vector of the symbol, the entity structure reconstruction of the initial comparison vector of the symbol can be realized, and the correct entity structure of the symbol is successfully introduced into the initial comparison vector of the symbol. Among them, the elements in the attribute prediction vector correspond one by one to multiple attributes, and the element is the probability that the correct attribute corresponding to the symbol is the attribute corresponding to the element.
[0096] Thus, by predicting the correct attributes of each symbol in the symbol sequence by using the context information of the symbol sequence, the entity structure in the entity record can be corrected, and then the problem of entity structure damage can be solved, and the correct entity structure is introduced in entity matching. Among them, the attribute prediction process has nothing to do with the original entity structure in the entity record, that is, the predicted attribute corresponding to the symbol has nothing to do with the attribute corresponding to the symbol in the entity record, and each symbol may be assigned the correct attribute through attribute prediction.
[0097] Optionally, the first language model adopts the BERT model. Among them, BERT is a pre-trained language model, and its full name is the Bidirectional Encoder Representation from Transformers model. Thus, the accuracy of feature encoding the symbol sequence by the first language model is improved by using the BERT model.
[0098] Optionally, the first context information extraction network adopts a Long Short-Term Memory (LSTM) network to improve the accuracy of context information extraction for the symbol sequence through the LSTM network.
[0099] (2) Optional model structure two
[0100] Another optional model structure: As Figure 3c shown, the structure reconstruction network layer may further include a first linear network layer and a first classification network layer. Among them, the first linear network layer and the first classification network layer are used to predict the probabilities of each symbol in the symbol sequence belonging to multiple attributes based on the first context representation corresponding to the symbol sequence.
[0101] At this time, in the process of predicting the attributes of each symbol in the symbol sequence according to the first context representation to obtain the attribute prediction vector corresponding to each symbol in the symbol sequence, a possible implementation method includes: successively processing the first context representation corresponding to the symbol sequence through the first linear network layer and the first classification network layer to obtain the probabilities of each symbol in the symbol sequence belonging to multiple attributes; and obtaining the attribute prediction vector corresponding to each symbol in the symbol sequence according to the probabilities of each symbol in the symbol sequence belonging to multiple attributes. Among them, the multiple attributes refer to multiple attributes in the entity record corresponding to the symbol sequence.
[0102] In this implementation method, for each symbol in the symbol sequence, the following operations can be performed: input the first context representation corresponding to the symbol into the first linear network layer to obtain the output data of the first linear network layer, and input the output data of the first linear network layer into the first classification network layer to obtain the output data of the first classification network layer. This output data includes the probabilities of the symbol belonging to multiple attributes, and the attribute corresponding to the maximum probability is most likely the predicted attribute of the symbol; based on the probabilities of the symbol belonging to multiple attributes, determine the attribute prediction vector corresponding to the symbol.
[0103] For example, the symbol sequence corresponding to the entity record in Table 1 is "National Free Shipping, A Brand Vitamin Drink, Lime Flavor, Functional Drink, B Brand, 600 ml". The probabilities of each symbol in this symbol sequence belonging to "Title", "Category", "Brand", "Flavor", "Specification", "Barcode" can be determined through the first linear network layer and the first classification network layer. If the probability of the symbol belonging to "Title" is the largest, then determine the predicted attribute of the symbol as "Title".
[0104] Optionally, the first classification network layer can adopt a softmax layer. Thus, the attribute prediction of the symbol is realized through the softmax layer, improving the accuracy of predicting the attributes of the symbol.
[0105] As an example, the symbol sequence is {T 11 ,..., T 1i ,..., T 1m}, and the first context representation of this symbol sequence obtained through the first language model and the first context information extraction network is {r 11 ,..., r 1i ,..., r 1m}, where m represents the symbol sequence, T 1i represents the i-th symbol in the symbol sequence, and r 1i represents the first context representation corresponding to the i-th symbol in the symbol sequence. The formula for processing the first context representation of the symbol sequence through the first linear network layer and the first classification network layer can be expressed as:
[0106] v 1i= softmax(δ(r 1i W + b))
[0107] Wherein, W and b are network parameters to be learned in the first linear network layer, which can be learned during the training process of the entity matching model. t is the total number of attributes in the entity record. softmax() represents the first classification network layer, δ() represents the first linear network layer, d2 represents the size of the hidden layer in the first context information extraction network, and v 1i is the attribute prediction vector corresponding to the i-th symbol in the symbol sequence, which includes the probabilities that the i-th symbol belongs to t attributes.
[0108] In some embodiments, based on the entity matching model shown in Figure 3c , during the process of reconstructing the entity structure for the initial comparison vectors of each symbol in the symbol sequence according to the attribute prediction vectors corresponding to each symbol in the symbol sequence, a possible implementation includes: concatenating the initial comparison vectors of each symbol in the symbol sequence with the attribute prediction vectors corresponding to each symbol in the symbol sequence to obtain the intermediate comparison vectors of each symbol in the symbol sequence. Thus, the fusion of the initial comparison vector of the symbol and the attribute prediction vector of the symbol is achieved through data concatenation, and the entity structure is more completely introduced into the comparison vector.
[0109] In this implementation, optionally, for each symbol in the symbol sequence: the structure vector corresponding to the symbol can be determined according to the attribute vector and the attribute prediction vector corresponding to the symbol; the initial comparison vector of the symbol is concatenated with the structure vector corresponding to the symbol to obtain the intermediate comparison vector of the symbol. Among them, the attribute vector is a vector parameter that is randomly initialized and adjusted through training during the training process of the entity matching model, and each row vector of it is a context representation of an attribute.
[0110] Further, the calculation formula of the structure vector corresponding to the symbol can be expressed as:
[0111] g 1i = v 1i P
[0112] Wherein, P is the attribute vector,, d3 is the dimension of the context representation of each attribute, and g 1i is the structure vector corresponding to the i-th symbol in the symbol sequence.
[0113] Further, the formula for concatenating the initial comparison vector of the symbol with the structure vector corresponding to the symbol can be expressed as:
[0114] h′ 1i = [h 1i ; g 1i
[0115] where h 1i represents the initial comparison vector of the i-th symbol in the symbol sequence, and h' 1i represents the intermediate comparison vector of the i-th symbol in the symbol sequence.
[0116] (3) Optional model structure three
[0117] Another optional model structure: As Figure 3c shown, based on the entity matching model shown in Figure 3a the semantic noise reduction network layer may include a second language model and a second context information extraction network. Among them:
[0118] The second language model is used to perform feature encoding on the symbol sequence input to the semantic noise reduction network layer to obtain a second encoded representation corresponding to the symbol sequence; the second context information extraction network is used to extract context information from the second encoded representation to obtain a second context representation corresponding to the symbol sequence. Among them, the second language model can adopt a pre-trained language model, and during the training process of the entity matching model, there is no need to train the second language model.
[0119] At this time, a possible implementation of S304 includes: determining the second context representation corresponding to the symbol sequence through the second language model and the second context information extraction network; performing correctness analysis on each symbol in the symbol sequence according to the second context representation corresponding to the symbol sequence to obtain the confidence of each symbol in the symbol sequence; performing semantic noise reduction on the intermediate comparison vectors of each symbol in the symbol sequence according to the confidence of each symbol in the symbol sequence to obtain enhanced comparison vectors corresponding to each symbol in the symbol sequence. Thus, by using the language model and the context information extraction network, the accuracy of analyzing the correctness of symbols is improved, the accuracy of the confidence reflecting the correctness of symbols is improved, and further the noise reduction effect of semantic noise reduction based on the confidence of symbols is improved.
[0120] The semantic noise caused by incorrect attribute values will have a greater impact on the accuracy of entity matching. Therefore, incorrect attribute values can be analyzed in the entity record and then removed or weakened to achieve semantic noise reduction, and the analysis of incorrect attribute values can be achieved by performing correctness analysis on the symbols in the symbol sequence. Therefore, an entity matching model with the ability to analyze the correctness of symbols in the symbol sequence can effectively solve the adverse impact of semantic noise on entity matching.
[0121] In this implementation, considering that adjacent symbols in the same symbol sequence are usually correlated, the correctness of symbols can be analyzed based on the context information of the symbols. Therefore, a semantic noise reduction network layer is introduced into the entity matching model, and a second language model and a second context information extraction network are introduced into the semantic noise reduction network layer. Through the second language model and the second context information extraction network, the context information in the symbol sequence is extracted to obtain the second context representation of the symbol sequence. Among them, the second context representation of the symbol sequence includes the second context representations of the symbols in the symbol sequence.
[0122] Specifically, the symbol sequence can be input into the second language model, and the second language model performs feature encoding on the symbol sequence to obtain the second encoded representation corresponding to the symbol sequence; the second encoded representation corresponding to the symbol sequence is input into the second context information extraction network, and the second context information extraction network extracts the context information from the second encoded representation corresponding to the symbol sequence to obtain the second context representation corresponding to the symbol sequence.
[0123] After that, for each symbol in the symbol sequence, the following operations can be performed respectively: the correctness of the symbol can be analyzed based on the second context representation of the symbol to obtain the confidence of the symbol, where the confidence of the symbol is the correctness score of the symbol, and the higher the confidence, the more likely the symbol is correct; based on the confidence of the symbol, the intermediate comparison vector of the symbol is adjusted to achieve semantic noise reduction of the intermediate comparison vector of the symbol.
[0124] Thus, by using the context information of the symbol sequence to predict the correctness of each symbol in the symbol sequence, the attribute values in the entity record are corrected, thereby solving the adverse effect of semantic noise caused by misaligned attribute values on entity matching.
[0125] Optionally, the second language model uses the BERT model to improve the accuracy of feature encoding of the symbol sequence.
[0126] Optionally, the first context information extraction network uses the LSTM network to improve the accuracy of context information extraction of the symbol sequence through the LSTM network.
[0127] (4) Optional model structure four
[0128] Another optional model structure: As Figure 3c shown, the semantic noise reduction network layer further includes a second linear network layer and a second classification network layer. Among them, the second linear network layer and the second classification network layer are used to analyze the correctness of each symbol in the symbol sequence based on the second context representation corresponding to the symbol sequence to obtain the confidence of each symbol in the symbol sequence.
[0129] At this time, in the process of analyzing the correctness of each symbol in the symbol sequence according to the second context representation to obtain the confidence of each symbol in the symbol sequence, a possible implementation method includes: sequentially processing the second context representation through the second linear network layer and the second classification network layer to obtain the confidence of each symbol in the symbol sequence.
[0130] In this implementation method, for each symbol in the symbol sequence, the following operations can be performed: input the second context representation corresponding to the symbol into the second linear network layer to obtain the output data of the second linear network layer; input the output data into the second classification network layer to obtain the output data of the second classification network layer, and the output data includes the confidence of the symbol.
[0131] Optionally, the process of determining the confidence of the symbol can be expressed as:
[0132] α 1i =P(z=1|T 1i )
[0133] where z = 1 indicates that the i-th symbol T in the symbol sequence 1i is the correct symbol in the entity record corresponding to the symbol sequence, and α 1i represents the confidence that T 1i is the correct symbol in the entity record corresponding to the symbol sequence, and P() represents the semantic noise reduction network layer.
[0134] In some embodiments, based on Figure 3c the entity matching model shown, in the process of performing semantic noise reduction on the intermediate comparison vectors of each symbol in the symbol sequence according to the confidence of each symbol in the symbol sequence to obtain the enhanced comparison vectors of each symbol in the symbol sequence, a possible implementation method includes: weighting the confidence of each symbol in the symbol sequence and the intermediate comparison vectors of each symbol in the symbol sequence to obtain the enhanced comparison vectors of each symbol in the symbol sequence. Thus, by weighting the confidence of the symbol and the intermediate comparison vector of the symbol, a lower weight value is assigned to the symbol with a lower confidence, realizing the soft deletion of the intermediate comparison vector corresponding to the symbol with more semantic noise and improving the semantic denoising effect.
[0135] In this implementation method, for each symbol in the symbol sequence: weighting the confidence of the symbol and the intermediate comparison vector of the symbol to obtain the enhanced comparison vector of the symbol.
[0136] Optionally, the calculation formula of the enhanced comparison vector of the symbol can be expressed as:
[0137] ″h 1i =α 1i h′ 1i
[0138] where h″1i The representation symbol T 1i of the enhanced contrast vector,
[0139] In addition to the above soft deletion strategy, the enhanced contrast vectors corresponding to symbols with a confidence level less than the threshold can also be set to 0.
[0140] Based on any of the foregoing embodiments, optionally, the contrast model adopts a BERT model to implement symbol contrast in the symbol sequence through the BERT model, improve the contrast effect, and reduce the training burden of the entity matching model.
[0141] Based on any of the foregoing embodiments, training data can be used to perform end-to-end training on the entity matching model. Among them, the training data includes multiple entity matching pairs and the true labels of multiple entity matching pairs. The true label of an entity matching pair is matching or not matching. For example, a label of 1 indicates matching, and a label of 0 indicates non-matching. During the training process, the entity matching model is adjusted based on the error between the matching result of the entity matching pair output by the entity matching model and the true label of the entity matching pair.
[0142] Referring to Figure 4 , Figure 4 which is a structural example diagram of the entity matching model provided by the embodiments of the present application.
[0143] As Figure 4 shown, the entity matching model includes an input layer, a contrast network layer, a structure reconstruction network layer, a semantic noise reduction network layer, an aggregation network layer, and a classification network layer connected to the aggregation network layer.
[0144] As Figure 4 shown, during the process of matching entity record e1 and entity record e2 through the entity matching model:
[0145] First, through the input layer, the symbol sequence {T 11 ,..., T 1i ,..., T 1m} corresponding to entity record e1 and the symbol sequence {T 21 , …, T 2i , …, T 2m} corresponding to entity record e2 are input into the contrast network layer.
[0146] Next, in the contrast network layer, symbol contrast that is independent of the structure (i.e., independent of attributes and cross-attributes) is performed on the two symbol sequences to obtain the initial contrast vectors corresponding to each symbol in the two symbol sequences.
[0147] Next, in the structure reconstruction network layer, determine the probabilities that each symbol in the two symbol sequences belongs to multiple attributes. Based on the probabilities that the symbols belong to multiple attributes, obtain the attribute prediction vectors corresponding to the symbols. Based on the attribute prediction vectors corresponding to the symbols, perform entity structure reconstruction on the initial contrast vectors corresponding to the symbols to obtain the intermediate contrast vectors corresponding to the symbols. Among them, Figure 4 Attr1 and Attr k in z represent the first attribute, the k-th attribute, and the z-th attribute respectively.
[0148] Next, in the semantic noise reduction network layer, perform correctness analysis on each symbol in the two symbol sequences to obtain the confidence levels of each symbol, and weight the confidence levels of each symbol with the intermediate contrast vectors corresponding to each symbol to obtain the enhanced contrast vectors corresponding to each symbol. Among them, in the left frame of the semantic noise reduction network layer shown in Figure 4 , 0.3, 1.0, 0.7, and 0.9 are the confidence levels of multiple symbols in the symbol sequence corresponding to the entity record e1; in the right frame of the semantic noise reduction network layer shown in Figure 4 , 1.0, 0.1, 0.6, and 0.8 are the confidence levels of multiple symbols in the symbol sequence corresponding to the entity record e2.
[0149] Finally, in the aggregation network layer, perform entity layer aggregation on the enhanced contrast vectors corresponding to multiple symbols. For example, aggregate the enhanced contrast vectors corresponding to each symbol in the symbol sequence corresponding to the entity record e1 and the enhanced contrast vectors corresponding to each symbol in the symbol sequence corresponding to the entity record e2 into a vector, and input this vector into the softmax layer to obtain the matching result between the entity record e and the entity record e2, that is, whether the entity record e1 and the entity record e2 belong to the same entity.
[0150] Refer to Figure 5 , Figure 5 which is the structural block diagram of the data processing device 500 provided by the embodiment of the present application. The data processing device 500 provided by the embodiment of the present application includes: an acquisition unit 501, an entity matching unit 502, and a data set processing unit 503, where:
[0151] The acquisition unit 501 is used to acquire a first data set, and the first data set contains multiple entity records;
[0152] The entity matching unit 502 is used to perform entity matching processing on multiple entity records through an entity matching model to obtain a matching result, and the entity matching processing includes cross-attribute symbol comparison and error correction of the comparison results of the symbol comparison;
[0153] The dataset processing unit 503 is configured to perform deduplication processing and / or fusion processing on the first dataset according to the matching result to obtain the second dataset.
[0154] In an alternative embodiment, the entity matching model includes a contrast network layer, a structure reconstruction network layer, a semantic noise reduction layer, and an aggregation network layer. Multiple entity records form at least one entity record pair. The matching result includes the matching result of the entity record pair. The entity matching unit 502 is specifically configured to: in the entity record pair, determine a symbol sequence corresponding to the entity record according to the attribute values corresponding to multiple attributes in the entity record; through the contrast network layer, compare each symbol in the symbol sequence to obtain an initial contrast vector corresponding to each symbol in the symbol sequence; through the structure reconstruction network layer, perform entity structure reconstruction on the initial contrast vector to obtain an intermediate contrast vector corresponding to each symbol in the symbol sequence; through the semantic noise reduction network layer, perform semantic noise reduction on the intermediate contrast vector to obtain an enhanced contrast vector corresponding to each symbol in the symbol sequence; through the aggregation network layer, perform aggregation processing on the enhanced contrast vector to obtain the matching result of the entity record pair.
[0155] In an alternative embodiment, the structure reconstruction network layer includes a first language model and a first context information extraction network. The entity matching unit 502 is specifically configured to: through the first language model and the first context information extraction network, determine a first context representation corresponding to the symbol sequence; according to the first context representation, predict the attributes of each symbol in the symbol sequence to obtain an attribute prediction vector corresponding to each symbol in the symbol sequence; according to the attribute prediction vectors corresponding to each symbol in the symbol sequence, perform entity structure reconstruction on the initial contrast vectors of each symbol in the symbol sequence to obtain an intermediate contrast vector.
[0156] In an alternative embodiment, the structure reconstruction network layer further includes a first linear network layer and a first classification network layer. The entity matching unit 502 is specifically configured to: sequentially process the first context representation through the first linear network layer and the first classification network layer to obtain the probabilities that each symbol in the symbol sequence belongs to multiple attributes; according to the probabilities that each symbol in the symbol sequence belongs to multiple attributes, obtain an attribute prediction vector corresponding to each symbol in the symbol sequence.
[0157] In an alternative embodiment, the entity matching unit 502 is specifically configured to: splice the initial contrast vectors of each symbol in the symbol sequence with the attribute prediction vectors corresponding to each symbol in the symbol sequence to obtain an intermediate contrast vector for each symbol in the symbol sequence.
[0158] In an alternative embodiment, the semantic noise reduction network layer includes a second language model and a second context information extraction network entity matching unit 502 is specifically configured to: determine a second context representation corresponding to the symbol sequence through the second language model and the second context information extraction network; perform correctness analysis on each symbol in the symbol sequence according to the second context representation to obtain the confidence of each symbol in the symbol sequence; perform semantic noise reduction on the intermediate comparison vectors of each symbol in the symbol sequence according to the confidence of each symbol in the symbol sequence to obtain enhanced comparison vectors of each symbol in the symbol sequence.
[0159] In an alternative embodiment, the semantic noise reduction network layer further includes a second linear network layer and a second classification network layer, and the entity matching unit 502 is specifically configured to: process the second context representation sequentially through the second linear network layer and the second classification network layer to obtain the confidence of each symbol in the symbol sequence.
[0160] In an alternative embodiment, the entity matching unit 502 is specifically configured to: weight the confidence of each symbol in the symbol sequence and the intermediate comparison vectors of each symbol in the symbol sequence to obtain enhanced comparison vectors of each symbol in the symbol sequence.
[0161] The data processing device provided by the embodiments of the present application is used to execute the technical solutions in the corresponding foregoing method embodiments, and its implementation principles and technical effects are similar and will not be elaborated herein.
[0162] Among them, the technical solutions provided by the embodiments of the present application can be implemented on a cloud server.
[0163] Figure 6 FIG. is a schematic structural diagram of a cloud server provided by an exemplary embodiment of the present application. The cloud server is used to run a data processing method, and is used to perform entity matching processing on a data set including multiple entity records, and perform deduplication processing and / or fusion processing based on the matching results. As Figure 6 shown, the cloud server includes: a memory 63 and a processor 64.
[0164] The memory 63 is used to store computer programs and can be configured to store various other data to support operations on the cloud server. The memory 63 may be an Object Storage Service (OSS).
[0165] The memory 63 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0166] A processor 64, coupled to a memory 63, is configured to execute a computer program in the memory 63 for: obtaining a first data set including a plurality of entity records; performing entity matching processing on the plurality of entity records through an entity matching model to obtain a matching result, where the entity matching processing includes cross-attribute symbol comparison and error correction of the comparison result of the symbol comparison; and performing duplicate removal processing and / or fusion processing on the first data set according to the matching result to obtain a second data set.
[0167] In an alternative embodiment, the entity matching model includes a comparison network layer, a structure reconstruction network layer, a semantic noise reduction layer, and an aggregation network layer. The plurality of entity records form at least one entity record pair, and the matching result includes the matching result of the entity record pair. When the processor 64 performs entity matching processing on the plurality of entity records through the entity matching model to obtain a matching result, it is specifically configured to: in the entity record pair, determine a symbol sequence corresponding to the entity record according to the attribute values corresponding to multiple attributes in the entity record; compare each symbol in the symbol sequence through the comparison network layer in the entity matching model to obtain an initial comparison vector corresponding to each symbol in the symbol sequence; perform entity structure reconstruction on the initial comparison vector through the structure reconstruction network layer to obtain an intermediate comparison vector corresponding to each symbol in the symbol sequence; perform semantic noise reduction on the intermediate comparison vector through the semantic noise reduction network layer to obtain an enhanced comparison vector corresponding to each symbol in the symbol sequence; and perform aggregation processing on the enhanced comparison vector through the aggregation network layer to obtain the matching result of the entity record pair.
[0168] In an alternative embodiment, the structure reconstruction network layer includes a first language model and a first context information extraction network. When the processor 64 performs entity structure reconstruction on the initial comparison vector through the structure reconstruction network layer to obtain an intermediate comparison vector corresponding to each symbol in the symbol sequence, it specifically includes: determining a first context representation corresponding to the symbol sequence through the first language model and the context information extraction network; predicting the attributes of each symbol in the symbol sequence according to the first context representation to obtain an attribute prediction vector corresponding to each symbol in the symbol sequence; and performing entity structure reconstruction on the initial comparison vector of each symbol in the symbol sequence according to the attribute prediction vector corresponding to each symbol in the symbol sequence to obtain the intermediate comparison vector.
[0169] In an alternative embodiment, the structure reconstruction network layer further includes a first linear network layer and a first classification network layer. When the processor 64 predicts the attributes of each symbol in the symbol sequence based on the first context representation to obtain the attribute prediction vectors corresponding to the symbols in the symbol sequence, it specifically includes: successively processing the first context representation through the first linear network layer and the first classification network layer to obtain the probabilities of each symbol in the symbol sequence belonging to multiple attributes; and obtaining the attribute prediction vectors corresponding to the symbols in the symbol sequence according to the probabilities of each symbol in the symbol sequence belonging to multiple attributes.
[0170] In an alternative embodiment, when the processor 64 reconstructs the entity structure of the initial comparison vectors of the symbols in the symbol sequence based on the attribute prediction vectors corresponding to the symbols in the symbol sequence to obtain the intermediate comparison vectors, it specifically includes: concatenating the initial comparison vectors of the symbols in the symbol sequence with the attribute prediction vectors corresponding to the symbols in the symbol sequence to obtain the intermediate comparison vectors of the symbols in the symbol sequence.
[0171] In an alternative embodiment, the semantic noise reduction network layer includes a second language model and a second context information extraction network. When the processor 64 performs semantic noise reduction on the intermediate comparison vectors through the semantic noise reduction network layer to obtain the enhanced comparison vectors, it specifically includes: determining the second context representation corresponding to the symbol sequence through the second language model and the second context information extraction network; performing correctness analysis on each symbol in the symbol sequence according to the second context representation to obtain the confidence levels of the symbols in the symbol sequence; and performing semantic noise reduction on the intermediate comparison vectors of the symbols in the symbol sequence according to the confidence levels of the symbols in the symbol sequence to obtain the enhanced comparison vectors of the symbols in the symbol sequence.
[0172] In an alternative embodiment, the semantic noise reduction network layer further includes a second linear network layer and a second classification network layer. When the processor 64 performs correctness analysis on each symbol in the symbol sequence according to the second context representation to obtain the confidence levels of the symbols in the symbol sequence, it specifically includes: successively processing the second context representation through the second linear network layer and the second classification network layer to obtain the confidence levels of the symbols in the symbol sequence.
[0173] In an alternative embodiment, when the processor 64 performs semantic noise reduction on the intermediate comparison vectors of the symbols in the symbol sequence according to the confidence levels of the symbols in the symbol sequence to obtain the enhanced comparison vectors of the symbols in the symbol sequence, it specifically includes: weighting the confidence levels of the symbols in the symbol sequence and the intermediate comparison vectors of the symbols in the symbol sequence to obtain the enhanced comparison vectors of the symbols in the symbol sequence.
[0174] Further, as Figure 6 shown, the cloud server further includes other components such as a firewall 61, a load balancer 62, a communication component 65, and a power supply component 66.Figure 6 Only some components are schematically shown, which does not mean that the cloud server only includes Figure 6 the components shown.
[0175] The cloud server provided by the embodiment of the present application converts an entity record into a corresponding symbol sequence; through the comparison network layer in the entity matching model, each symbol in the symbol sequence is compared to obtain an initial comparison vector corresponding to each symbol; through the structure reconstruction network layer in the entity matching model, entity structure reconstruction is performed on the initial comparison vectors corresponding to each symbol to obtain an intermediate comparison vector corresponding to each symbol; through the semantic noise reduction network layer in the entity matching model, semantic noise reduction is performed on the intermediate comparison vectors of each symbol to obtain an enhanced comparison vector corresponding to each symbol; through the aggregation network layer in the entity matching model, aggregation processing is performed on the enhanced comparison vectors to obtain a matching result of the entity record pair. Thus, global comparison of symbols is realized on the basis of ignoring the entity structure in the entity record, and error correction of structural information errors and symbol errors in the entity record is realized through two error correction modules, namely the structure reconstruction network layer and the semantic noise reduction network layer, thereby solving the adverse effect of dirty data on the accuracy of entity matching and effectively improving the accuracy of entity matching..
[0176] Correspondingly, the embodiment of the present application further provides a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to implement the steps in the above method embodiments (such as Figure 2 , Figure 3b the method shown).
[0177] Correspondingly, the embodiment of the present application further provides a computer program product, including a computer program / instructions, which when executed by a processor causes the processor to implement the steps in the above method embodiments (such as Figure 2 , Figure 3b the method shown).
[0178] The above Figure 6 The communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0179] The aboveFigure 6 The power supply component in [it] provides power for various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.
[0180] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0181] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0182] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0183] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0184] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0185] The memory may include non-permanent memory in the form of computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0186] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media, such as modulated data signals and carrier waves.
[0187] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0188] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A data processing method, characterized in that, Including: Obtain a first data set, where the first data set contains multiple entity records; Perform entity matching processing on the multiple entity records through an entity matching model to obtain a matching result. The entity matching processing includes cross-attribute symbol comparison and error correction of the comparison result of the symbol comparison. The error correction of the comparison result of the symbol comparison includes: performing entity structure reconstruction and semantic noise reduction on the comparison result of the symbol comparison; the entity matching model includes a comparison network layer, a structure reconstruction network layer, a semantic noise reduction network layer, and an aggregation network layer; According to the matching result, perform deduplication processing and / or fusion processing on the first data set to obtain a second data set; The multiple entity records form at least one entity record pair, and the matching result includes the matching result of the entity record pair. The step of performing entity matching processing on the multiple entity records through the entity matching model to obtain a matching result includes: In the entity record pair, determine a symbol sequence corresponding to the entity record according to the attribute values corresponding to multiple attributes in the entity record; Through the comparison network layer, compare each symbol in the symbol sequence to obtain an initial comparison vector corresponding to each symbol in the symbol sequence; Through the structure reconstruction network layer, correct the attributes corresponding to each symbol in the symbol sequence to obtain the corrected attributes corresponding to each symbol in the symbol sequence, and add the corrected attributes corresponding to each symbol in the symbol sequence to the initial comparison vector corresponding to each symbol in the symbol sequence to achieve entity structure reconstruction of the initial comparison vector corresponding to each symbol; Through the semantic noise reduction network layer, detect the semantic correctness of each symbol in the symbol sequence to obtain the confidence level of each symbol in the symbol sequence; according to the confidence level of each symbol in the symbol sequence, perform semantic noise reduction on the intermediate comparison vector of each symbol to obtain an enhanced comparison vector corresponding to each symbol; Through the aggregation network layer, perform aggregation processing on the enhanced comparison vectors to obtain the matching result of the entity record pair.
2. The data processing method according to claim 1, wherein The structure reconstruction network layer includes a first language model and a first context information extraction network. The step of performing entity structure reconstruction on the initial comparison vector through the structure reconstruction network layer to obtain an intermediate comparison vector corresponding to each symbol in the symbol sequence includes: Through the first language model and the first context information extraction network, determine a first context representation corresponding to the symbol sequence; According to the first context representation, predict the attributes of each symbol in the symbol sequence to obtain an attribute prediction vector corresponding to each symbol in the symbol sequence; According to the attribute prediction vector corresponding to each symbol in the symbol sequence, perform entity structure reconstruction on the initial comparison vector of each symbol in the symbol sequence to obtain the intermediate comparison vector.
3. The data processing method according to claim 2, wherein The structure reconstruction network layer further includes a first linear network layer and a first classification network layer. The step of predicting the attributes of each symbol in the symbol sequence according to the first context representation to obtain an attribute prediction vector corresponding to each symbol in the symbol sequence includes: Process the first context representation through the first linear network layer and the first classification network layer in sequence to obtain the probabilities that each symbol in the symbol sequence belongs to the multiple attributes; Obtain the attribute prediction vector corresponding to each symbol in the symbol sequence according to the probabilities that each symbol in the symbol sequence belongs to the multiple attributes.
4. The data processing method according to claim 2, wherein The entity structure reconstruction of the initial contrast vector of each symbol in the symbol sequence according to the attribute prediction vector corresponding to each symbol in the symbol sequence to obtain the intermediate contrast vector includes: Concatenate the initial contrast vector of each symbol in the symbol sequence with the attribute prediction vector corresponding to each symbol in the symbol sequence to obtain the intermediate contrast vector of each symbol in the symbol sequence.
5. The data processing method according to any one of claims 1 to 4, characterized in that, The semantic noise reduction network layer includes a second language model and a second context information extraction network. The semantic noise reduction of the intermediate contrast vector through the semantic noise reduction network layer to obtain the enhanced contrast vector includes: Determine the second context representation corresponding to the symbol sequence through the second language model and the second context information extraction network; Perform correctness analysis on each symbol in the symbol sequence according to the second context representation to obtain the confidence of each symbol in the symbol sequence; Perform semantic noise reduction on the intermediate contrast vector of each symbol in the symbol sequence according to the confidence of each symbol in the symbol sequence to obtain the enhanced contrast vector of each symbol in the symbol sequence.
6. The data processing method according to claim 5, characterized in that The semantic noise reduction network layer further includes a second linear network layer and a second classification network layer. The performing correctness analysis on each symbol in the symbol sequence according to the second context representation to obtain the confidence of each symbol in the symbol sequence includes: Process the second context representation through the second linear network layer and the second classification network layer in sequence to obtain the confidence of each symbol in the symbol sequence.
7. A data processing device, characterized in that, Includes: An acquisition unit for acquiring a first data set, where the first data set contains multiple entity records; An entity matching unit for performing entity matching processing on the multiple entity records through an entity matching model to obtain a matching result. The entity matching processing includes cross-attribute symbol comparison and error correction of the comparison result of the symbol comparison. The error correction of the comparison result of the symbol comparison includes: performing entity structure reconstruction and semantic noise reduction on the comparison result of the symbol comparison. The entity matching model includes a comparison network layer, a structure reconstruction network layer, a semantic noise reduction network layer, and an aggregation network layer; A data set processing unit for performing deduplication processing and / or fusion processing on the first data set according to the matching result to obtain a second data set; The multiple entity records form at least one entity record pair, and the matching result includes the matching result of the entity record pair. The entity matching unit is specifically used for: In the entity record pair, determine the symbol sequence corresponding to the entity record according to the attribute values corresponding to the multiple attributes in the entity record; Compare each symbol in the symbol sequence through the comparison network layer to obtain the initial contrast vector corresponding to each symbol in the symbol sequence; By reconstructing the network layer of the structure, correcting the attributes corresponding to each symbol in the symbol sequence, obtaining the corrected attributes corresponding to each symbol in the symbol sequence, and adding the corrected attributes corresponding to each symbol in the symbol sequence to the initial comparison vector corresponding to each symbol in the symbol sequence, the entity structure reconstruction of the initial comparison vector corresponding to each symbol is realized; By the semantic noise reduction network layer, detecting the semantic correctness of each symbol in the symbol sequence to obtain the confidence of each symbol in the symbol sequence; according to the confidence of each symbol in the symbol sequence, performing semantic noise reduction on the intermediate comparison vector of each symbol to obtain the enhanced comparison vector corresponding to each symbol; Through the aggregation network layer, performing an aggregation process on the enhanced comparison vector to obtain the matching result of the entity record pair.
8. A cloud server, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the cloud server can execute the data processing method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the data processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Entity pair matching method and device in database, electronic equipment and storage medium
CN113569554A