Data cleaning method and apparatus for erroneously matched entity, device and medium

By performing abnormal detection of data tuples and segmentation correction of mixed entities, the problem of information loss in error matching entity cleaning is solved, and an efficient and accurate data cleaning process is achieved.

WO2025107392A1PCT designated stage expired Publication Date: 2025-05-30SHENZHEN INST OF COMPUTING SCI

Patent Information

Application Number
PCT/CN2023/140142
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-21
Filing Date
2023-12-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art can easily lead to information loss when cleaning the wrong matching entity, seriously affecting the performance of downstream applications.

Method used

By performing exception detection on the acquired data tuples, identify entity tuples of mixed entity types, and construct split entities based on exception properties for correction, ensuring that information is not lost during the data cleaning process.

Benefits of technology

Effectively identify and split wrong matching entities, avoid information loss, improve the accuracy and reliability of data cleaning, and thus improve the performance of downstream applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023140142_30052025_PF_FP_ABST
    Figure CN2023140142_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application is applicable to the technical field of data cleaning, and particularly relates to a data cleaning method and apparatus for an erroneously matched entity, a device and a medium. The method comprises: performing anomaly detection on an acquired data tuple to obtain an entity tuple represented as anomalous and anomalous attributes of the entity tuple; determining the entity type of the entity tuple, and if it is determined that the entity type of the entity tuple is a hybrid entity type, determining that the entity tuple is a hybrid entity; on the basis of each anomalous attribute of the hybrid entity and a corresponding attribute value, constructing a segmented entity corresponding to each anomalous attribute; and respectively performing anomaly correction on each segmented entity to obtain a corrected segmented entity, and determining that all corrected segmented entities are results of cleaning the hybrid entity. The hybrid entity is found and the segmented entity is constructed for each anomalous attribute of the hybrid entity, so as to respectively correct and clean the segmented entities, thereby preventing data from being discarded.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, equipment and medium for cleaning data of mismatched entities

[0001] This application is based on the Chinese invention application with application number 202311572735.7 filed on November 21, 2023, and entitled “Method, device, equipment and medium for cleaning data of incorrectly matched entities”, and claims its priority. Technical Field

[0002] The present application is applicable to the field of data cleaning technology, and in particular relates to a method, apparatus, device and medium for cleaning data of mismatched entities. Background Art

[0003] As data sources increase, data quality issues such as missing values, outliers, and conflicting attributes often arise. Missing and outlier data can be caused by mechanical issues (such as hardware failure) or privacy concerns. Such data requires cleaning before it can be used downstream. In addition to the aforementioned reasons, conflicting attributes can also arise from incorrect data fusion (i.e., entity resolution). For example, two people named Zhang San, one with Shenzhen as their city and the other with Guangxi as their province, may have their information mistakenly merged. For conflicting attributes caused by incorrect data fusion, during the data cleaning process, if one attribute (such as Shenzhen as their city) is assumed to be correct and the other attribute (such as changing the province from Guangxi to Guangdong) is corrected, while the information for one entity (such as Zhang San in Shenzhen) is successfully corrected, the information for the other entity (i.e., another Zhang San in Guangxi) is lost. Therefore, the existing cleaning process for incorrectly matched entities can result in information loss, which can severely impact the performance of downstream applications using this data.

[0004] Therefore, how to identify and split mismatched entities to facilitate data cleaning while avoiding data loss has become an urgent problem to be solved.

[0005] Summary of the Invention

[0006] The embodiments of the present application provide a method, apparatus, device and medium for cleaning data of mismatched entities to solve the problem of how to identify, split and process mismatched entities so as to facilitate data cleaning while avoiding data loss.

[0007] A method for cleaning data of mismatched entities, the method comprising:

[0008] Performing anomaly detection on the acquired data tuples to obtain entity tuples characterized as abnormal and abnormal attributes in the entity tuples, wherein the data tuples include at least one of the entity tuples;

[0009] detecting an entity type of the entity tuple, and if it is detected that the entity type of the entity tuple is a mixed entity type, determining that the entity tuple is a mixed entity, wherein the entity tuple corresponding to the mixed entity type represents more than one object;

[0010] According to each abnormal attribute and the corresponding attribute value in the hybrid entity, a segmentation entity corresponding to each abnormal attribute is constructed, wherein each segmentation entity includes an abnormal attribute and a corresponding attribute value;

[0011] Anomalies are corrected for each segmented entity to obtain corrected segmented entities, and all corrected segmented entities are determined to be the result of cleaning the mixed entity.

[0012] A device for cleaning data of an erroneously matched entity, the cleaning device comprising:

[0013] an anomaly detection module, configured to perform anomaly detection on the acquired data tuples to obtain entity tuples characterized as abnormal and abnormal attributes in the entity tuples, wherein the data tuples include at least one entity tuple;

[0014] a type detection module, configured to detect an entity type of the entity tuple, and determine that the entity tuple is a hybrid entity if it is detected that the entity type of the entity tuple is a hybrid entity type, wherein the entity tuple corresponding to the hybrid entity type represents more than one object;

[0015] An entity segmentation module is used to construct a segmentation entity corresponding to each abnormal attribute and the corresponding attribute value in the hybrid entity, wherein each segmentation entity includes an abnormal attribute and a corresponding attribute value;

[0016] The entity cleaning module is used to perform abnormality correction on each segmented entity to obtain a corrected segmented entity, and determine that all corrected segmented entities are the results of cleaning the mixed entity.

[0017] A computer device includes a memory, a processor, and a readable storage medium stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the readable storage medium:

[0018] Performing anomaly detection on the acquired data tuples to obtain entity tuples characterized as abnormal and abnormal attributes in the entity tuples, wherein the data tuples include at least one of the entity tuples;

[0019] detecting an entity type of the entity tuple, and if it is detected that the entity type of the entity tuple is a mixed entity type, determining that the entity tuple is a mixed entity, wherein the entity tuple corresponding to the mixed entity type represents more than one object;

[0020] According to each abnormal attribute and the corresponding attribute value in the hybrid entity, a segmentation entity corresponding to each abnormal attribute is constructed, wherein each segmentation entity includes an abnormal attribute and a corresponding attribute value;

[0021] Anomalies are corrected for each segmented entity to obtain corrected segmented entities, and all corrected segmented entities are determined to be the result of cleaning the mixed entity.

[0022] One or more computer-readable storage media storing computer-readable instructions, wherein the computer-readable instructions, when executed by one or more processors, cause the one or more processors to perform the following steps:

[0023] Performing anomaly detection on the acquired data tuples to obtain entity tuples characterized as abnormal and abnormal attributes in the entity tuples, wherein the data tuples include at least one of the entity tuples;

[0024] detecting an entity type of the entity tuple, and if it is detected that the entity type of the entity tuple is a mixed entity type, determining that the entity tuple is a mixed entity, wherein the entity tuple corresponding to the mixed entity type represents more than one object;

[0025] According to each abnormal attribute and the corresponding attribute value in the hybrid entity, a segmentation entity corresponding to each abnormal attribute is constructed, wherein each segmentation entity includes an abnormal attribute and a corresponding attribute value;

[0026] Anomalies are corrected for each segmented entity to obtain corrected segmented entities, and all corrected segmented entities are determined to be the result of cleaning the mixed entity.

[0027] The present application performs anomaly detection on the acquired data tuple to obtain an entity tuple characterized as abnormal and an abnormal attribute in the entity tuple, detects the entity type of the entity tuple, and if it is detected that the entity type of the entity tuple is a mixed entity type, determines that the entity tuple is a mixed entity, and constructs a split entity corresponding to each abnormal attribute in the mixed entity according to each abnormal attribute and the corresponding attribute value, performs anomaly correction on each split entity respectively, obtains a corrected split entity, and determines that all corrected split entities are the results of cleaning the mixed entity. By finding the mixed entity and constructing a split entity for each abnormal attribute in the mixed entity, the split entities can be corrected and cleaned respectively to avoid data discard. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0029] FIG1 is a schematic diagram of an application environment of a method for cleaning data of mismatched entities provided in Example 1 of the present application;

[0030] FIG2 is a flow chart of a method for cleaning data of mismatched entities provided in Example 2 of the present application;

[0031] FIG3 is a flow chart of a method for cleaning data of mismatched entities provided in Example 3 of the present application;

[0032] FIG4 is a flow chart of a method for cleaning data of mismatched entities provided in a fourth embodiment of the present application;

[0033] FIG5 is a flow chart of a method for cleaning data of mismatched entities provided in Example 5 of the present application;

[0034] FIG6 is a schematic diagram of a cleaning method according to a fifth embodiment of the present invention;

[0035] FIG7 is a flow chart of a method for cleaning data of an incorrectly matched entity provided in Example 6 of the present application;

[0036] FIG8 is a flow chart of a method for cleaning data of an incorrectly matched entity provided in Example 7 of the present application;

[0037] FIG9 is a schematic structural diagram of a device for cleaning data of an incorrectly matched entity provided in Example 8 of the present application;

[0038] FIG10 is a schematic structural diagram of a computer device provided in Example 9 of the present application. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0040] In order to illustrate the technical solution of the present application, specific embodiments are provided below.

[0041] A method for cleaning data of an incorrectly matched entity provided in the first embodiment of the present application can be applied in an application environment such as that shown in FIG1 , wherein a server communicates with a client, and the server is used to carry the data cleaning method to provide data cleaning services. The client can request data cleaning services from the server by providing corresponding data tuples, thereby obtaining cleaned data. Of course, the server can also clean the data tuples stored in itself. The client includes but is not limited to PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, personal digital assistants (PDAs), and other devices. The computer device corresponding to the server can be implemented using an independent server or a server cluster consisting of multiple servers.

[0042] Referring to FIG2 , which is a flow chart of a method for cleaning data of an incorrectly matched entity provided in Example 2 of the present application, the method for cleaning data of an incorrectly matched entity is applied to the server in FIG1 , and the user corresponding to the client sends the data to be cleaned (i.e., data tuples) to the server to trigger the server to perform the cleaning task. As shown in FIG2 , the method for cleaning data of an incorrectly matched entity may include the following steps:

[0043] Step S201 : performing anomaly detection on the acquired data tuple to obtain entity tuples characterized as abnormal and abnormal attributes in the entity tuples.

[0044] In this embodiment, a tuple is a row or a column in a relational table in a database. If a row is represented as a tuple, the row of data corresponds to an entity. Each column in the row of data represents the attributes of the entity. The entity tuple is a row of data in the relational table. The data tuple includes at least one entity tuple. Specifically, the data tuple can be a relational table.

[0045] In the process of collecting and organizing data, the attributes of two different entities may be mistakenly merged into one entity. For example, there are two entities named A. The gender attribute of the first entity A is "male (that is, the attribute value of the gender attribute)", and the age attribute of the second entity A is "25 years old (the attribute value of the age attribute)". After the two entities are mistakenly merged, the entity A is obtained, with the gender attribute of "male" and the age attribute of "25 years old". Obviously, this is an incorrect fusion result and a manifestation of entity anomaly.

[0046] Furthermore, during data collection and organization, the attribute values ​​of an entity may conflict. For example, entity B may have a gender attribute of "male" and a medical treatment attribute of "gynecology." Since men are not advised to seek medical treatment at gynecology clinics, the attributes of entity B may be incorrect, which is a manifestation of an entity anomaly. Of course, other entity anomalies include instances where an attribute is empty or exceeds a limit.

[0047] For data tuples, the anomaly detection method can be used to filter out abnormal entity tuples and also determine the abnormal attributes corresponding to the abnormal entity tuples. The object of anomaly detection can be an attribute value. If there is an abnormality in the attribute value, the corresponding attribute is also an abnormal attribute, and the corresponding entity tuple is the abnormal entity tuple.

[0048] By analyzing the characteristics of the attributes, we can obtain the corresponding detection rules. These detection rules can be used to automatically detect anomalies. The rules can be constructed based on logical conditions or machine learning models. Of course, other forms of rules can also be used to construct rules. It is necessary to ensure that the constructed rules can automatically detect anomalies.

[0049] Step S202 : detecting the entity type of the entity tuple. If it is detected that the entity type of the entity tuple is a mixed entity type, determining that the entity tuple is a mixed entity.

[0050] In this embodiment, for the above-mentioned abnormal entity tuples, there are single entities and mixed entities, where the mixed entity is a mixed entity type, and the entity tuple corresponding to the mixed entity type represents more than one object, and the single entity is a single entity type, and the entity tuple corresponding to the single entity type represents only one object.

[0051] An object can refer to an object, item, person, place, etc. in the real world. For example, for an entity tuple, the corresponding entity in the real world is city M, and the city M is a unique object, then the entity tuple is a single entity. For another example, for an entity tuple, the corresponding entity is represented by a name, and there are multiple person objects corresponding to the name in the real world, then the entity tuple is a mixed entity.

[0052] Step S203: constructing a segmentation entity corresponding to each abnormal attribute according to each abnormal attribute and the corresponding attribute value in the mixed entity.

[0053] In this embodiment, when performing anomaly detection on the hybrid entity in the above step S201, all abnormal attributes that are abnormal in the hybrid entity are determined, wherein the characteristics of the abnormal attributes include conflict anomalies, empty set anomalies, error anomalies, etc. For the hybrid entity, since the hybrid entity corresponds to more than one object, the abnormal attribute characterized by conflict anomalies should not belong to one object. Therefore, based on the abnormal attribute characterized by conflict anomalies, the hybrid entity can be segmented into segmented entities representing different objects, that is, one segmented entity corresponds to one object.

[0054] Among them, each split entity contains an abnormal attribute and a corresponding attribute value, and the attributes in each split entity correspond one-to-one to the attributes of the mixed entity, that is, the mixed entity contains attribute 1, attribute 2 and attribute 3, then each split entity also includes attribute 1, attribute 2 and attribute 3. If attribute 2 and attribute 3 are abnormal attributes with conflicting anomalies, they can be divided into two split entities. The attribute value corresponding to attribute 2 in the mixed entity is used as the attribute value of attribute 2 of one split entity, and the attribute value corresponding to attribute 3 in the mixed entity is used as the attribute value of attribute 3 of another split entity. Attributes that do not have corresponding attribute values ​​in each split entity can be left blank, or a random or preset value can be written.

[0055] For example, for a mixed entity, the entity name is Name1, which corresponds to multiple objects in reality. The corresponding gender attribute in its entity tuple is "male", and the corresponding medical attribute is "gynecology". Based on the principle of conflict, the gender attribute and the medical attribute are conflicting attributes. Based on this, the entity name Name1 can be divided into two split entities, such as the first split entity and the second split entity. Among them, the name of the first split entity is Name1, the gender attribute is "male", and the medical attribute is empty. The name of the second split entity is also Name1, the gender attribute is empty, and the medical attribute is "gynecology".

[0056] In step S204 , abnormality correction is performed on each segmented entity to obtain a corrected segmented entity, and all corrected segmented entities are determined to be the result of cleaning the mixed entity.

[0057] In this embodiment, after entity segmentation in the above-mentioned step S203, the obtained segmented entity is an entity of a single entity type. However, due to the segmentation, there are empty or incorrect attribute values ​​in the segmented entity. Therefore, it is necessary to perform abnormal correction on each segmented entity, wherein the correction is for each abnormal attribute in the segmented entity, and the correction may include operations such as error correction and missing completion.

[0058] All attributes in the corrected segmented entities have been corrected and missing attributes have been completed. All corrected segmented entities are the result of cleaning the mixed entities.

[0059] Anomaly correction may refer to replacing the original erroneous value with the corresponding correct value or filling in the empty data. Among them, anomaly correction may use corresponding rules, models, trusted knowledge graphs, and trusted source data to match the original abnormal data with the correct value.

[0060] For example, for the first split entity and the second split entity split above, the name of the first split entity is Name1, the gender attribute is "male", and the medical attribute is empty. The name of the second split entity is also Name1, the gender attribute is empty, and the medical attribute is "gynecology". At this time, for the first split entity, the two attribute values ​​of Name1 and "male" are used to match the medical information as "andrology" from the trusted knowledge graph. Therefore, "andrology" can be filled in the medical attribute of the first split entity as its attribute value. For the second split entity, the attribute value of "gynecology" is used. Using the preset rules, it can be determined that the gender information is "female". Therefore, "female" can be filled in the gender attribute of the second split entity as its attribute value.

[0061] The embodiment of the present application performs anomaly detection on the acquired data tuple to obtain an entity tuple characterized as abnormal and an abnormal attribute in the entity tuple, and detects the entity type of the entity tuple. If it is detected that the entity type of the entity tuple is a mixed entity type, the entity tuple is determined to be a mixed entity. According to each abnormal attribute and the corresponding attribute value in the mixed entity, a split entity corresponding to each abnormal attribute is constructed, and the abnormality is corrected for each split entity to obtain a corrected split entity. All corrected split entities are determined to be the result of cleaning the mixed entity. By finding the mixed entity and constructing a split entity for each abnormal attribute in the mixed entity, the split entities are corrected and cleaned respectively to avoid data discard.

[0062] See Figure 3, which is a flow chart of a method for cleaning data of mismatched entities provided in Example 3 of the present application. As shown in Figure 3, the entity type of the entity tuple is detected in step S202, and the following steps may also be included:

[0063] Step S301: Obtain a preset knowledge graph, and map and extract the objects of the entity tuple according to the preset knowledge graph to obtain an object set of the entity tuple.

[0064] In this embodiment, the preset knowledge graph may refer to graph data containing some known data and the relationships between these data. Based on the known data and the relationships between these data, the entity objects corresponding to the entity tuple can be retrieved, thereby obtaining the object set corresponding to the entity tuple. Mapping extraction requires a portion of the attribute information in the entity tuple, specifically attribute information that describes the entity tuple's characteristics in the real world, such as the name of a country.

[0065] Taking the country name as an example, the country name of entity A is Z. Using the country name Z, it can be mapped to the country corresponding to the country name from the preset knowledge graph. The country is currently a unique object. Therefore, there are no other objects for the entity A. For example, taking the name of a person as an example, the name of entity A is M. Using the name M, it can be mapped to 10 people from the preset knowledge graph, that is, the entity A corresponds to 10 objects. In actual use, in order to improve the accuracy, when mapping and extracting objects in the entity tuple, the name can be combined with at least one conflicting attribute to determine the corresponding object, such as the name M and the gender attribute, thereby further limiting the characteristics of the object and being able to match 3 objects, which is more accurate than the aforementioned 10 objects.

[0066] Step S302 : detecting whether the object set contains at least two objects. If it is detected that the object set contains at least two objects, determining that the entity type of the entity tuple is a mixed entity type.

[0067] In this embodiment, if the object set contains two or more objects, it means that the entity tuple has multiple entities in the real world. Therefore, the entity tuple is a mixed entity, that is, the corresponding entity type is a mixed entity type.

[0068] Of course, after detecting whether the object set contains at least two objects, if it is detected that the object set contains one object, the entity type of the entity tuple is determined to be a single entity type.

[0069] The embodiment of the present application provides a method for distinguishing the types of abnormal attributes based on a knowledge graph, which can effectively distinguish between mixed entity categories and single entity categories, so as to implement different cleaning processes for entity tuples of different categories, making the obtained cleaning results more accurate.

[0070] See Figure 4, which is a flow chart of a method for cleaning data of mismatched entities provided in Example 4 of the present application. As shown in Figure 4, in the above step S201, anomaly detection is performed on the acquired data tuple to obtain entity tuples characterized as abnormal and abnormal attributes in the entity tuples, and the following steps may also be included:

[0071] Step S401: Acquire data tuples and preset anomaly detection rules.

[0072] In this embodiment, the preset anomaly detection rules are rules constructed based on the analysis of association relationships, relationships between conditions, etc., and can specifically include logic-based rules and model-based rules. The model can be a machine learning model, a neural network model, etc.

[0073] Logic-based rules can use rule discovery technology to analyze the relationship between data or conditions from existing data. They can also be rules manually specified according to needs. The rules can be expressed in the form of logical predicates. The form of the logical predicate is X->Y, which means that if condition X is met, then condition Y must be met. Conversely, if condition X is met but Y is not, this situation does not meet the rule, and the attribute corresponding to Y is an abnormal attribute. The use of logical predicates can also be combined with trusted source data Γ to better verify and match conditions X and Y.

[0074] The rules based on the machine learning model can be embodied in the form of a model predicate in the form of M(X,Y). Its function is to judge the correlation between attributes X and Y. If the correlation is lower than a certain threshold and attribute X is credible, then the current entity tuple is considered to not satisfy the rule, and the attribute corresponding to attribute Y is an abnormal attribute. The model predicate may not be combined with the source data Γ, but it may be necessary to rely on the source data Γ in the process of discovering the model predicate.

[0075] Step S402 : For any entity tuple in the data tuple, use a preset anomaly detection rule to perform an anomaly detection on each attribute of any entity tuple, and obtain an abnormal attribute whose detection result is abnormal.

[0076] Step S403: If any entity tuple has an abnormal attribute, then the entity tuple is determined to be an entity tuple representing an abnormality, and the abnormal attribute belonging to the entity tuple is determined from all abnormal attributes.

[0077] In this embodiment, anomaly detection includes the above-mentioned conflict anomalies, empty set anomalies, error anomalies, etc. Of course, it can also include other anomalies. If an entity tuple contains abnormal attributes, it can be determined as an abnormal entity tuple, otherwise it is a normal entity tuple.

[0078] If an entity tuple is determined to be abnormal, all abnormal attributes in it need to be aggregated to form an abnormal attribute set. Specifically, the abnormalities can be classified according to different categories. A conflicting abnormality corresponds to a conflicting attribute set, and so on. In subsequent use, the corresponding abnormal attribute set can be directly called according to different classifications. For example, in step S203 of the above embodiment 2, the conflicting attribute set can be directly called.

[0079] In the embodiment of the present application, the abnormal attributes of the entity tuple can be detected based on rules, which can cover a wider range of abnormal data and perform comprehensive detection, which helps to quickly and accurately realize abnormality detection, thereby facilitating subsequent cleaning operations.

[0080] 5 is a flow chart of a method for cleaning data of mismatched entities provided in a fifth embodiment of the present application. As shown in FIG5 , the above step S204 of correcting the abnormality of each segmented entity to obtain the corrected segmented entity may include the following steps:

[0081] Step S501 : for any segmented entity in each segmented entity, detecting whether the attribute value of the segmented entity corresponding to the abnormal attribute is empty.

[0082] Step S502: If it is detected that the attribute value of the segmented entity corresponding to the abnormal attribute is not empty, the correct attribute value is matched according to the segmented entity and the corresponding abnormal attribute in combination with a preset matching rule and / or matching model.

[0083] Step S503: Using the correct attribute value to replace and correct the attribute value of the segmented entity corresponding to the abnormal attribute, to obtain a corrected segmented entity.

[0084] In this embodiment, exception correction is divided into replacement correction and completion correction, among which, replacement correction means that there is an original attribute value under the corresponding exception attribute, and a matching attribute value needs to be used to replace the original attribute value. Completion correction means that the corresponding attribute value under the corresponding exception attribute is empty, and a matching attribute value needs to be used to fill the empty exception attribute.

[0085] Among them, the preset matching rule and the logic-based rule in step S401 of the above-mentioned embodiment 4 are rules of the same nature. Of course, the two can be the same rule. The preset matching model and the model in the model-based rule in step S401 of the above-mentioned embodiment 4 are models of the same nature. Of course, the two can be the same model.

[0086] It should be noted that the above logic-based rules and model-based rules can realize anomaly detection and can also match the correct attribute value for each abnormal attribute during the anomaly detection process.

[0087] Optionally, after detecting whether the attribute value of the segmented entity corresponding to the abnormal attribute is empty, the following is further included:

[0088] If it is detected that the attribute value of the abnormal attribute corresponding to the segmented entity is empty, the preset completion rules, completion knowledge graph and / or completion model are used to complete and correct the abnormal attribute corresponding to the segmented entity to obtain the corrected segmented entity.

[0089] Among them, the completion rule is a rule of the same nature as the above-mentioned logic-based rule, the completion model is a model of the same nature as the model in the above-mentioned model-based rule, and the completion knowledge graph is a knowledge graph of the same nature as the above-mentioned preset knowledge graph. Specifically, the rules, models, and knowledge graphs defined in this embodiment can be the same as the rules, models, knowledge graphs, etc. in the aforementioned embodiments. Therefore, the process of anomaly detection and subsequent matching of attribute values ​​can be achieved based on rules, models, and knowledge graphs. Of course, these rules, models, and knowledge graphs can be self-learned through continuous data accumulation in the process of executing the method of this application to make themselves more accurate.

[0090] As shown in Figure 6, a schematic diagram of the framework of a cleaning method provided in Example 5 of the present application is provided, wherein, for a data tuple containing N entity tuples, an abnormal entity tuple is obtained by anomaly detection, and then a mixed entity of a mixed entity type is obtained by type detection (a single entity of a single entity type can also be obtained), and then the corresponding entity tuple enters the process of splitting the entity (for mixed entities, the single entity is not split), and the entity after the split is corrected (the single entity is directly corrected), and finally, the missing values ​​of the split and corrected entities are completed, and finally an entity tuple with no abnormal attributes and no missing attributes is obtained, that is, the cleaning of the entity tuple is achieved. In this figure, the knowledge graph, rules, models, etc. can refer to the records in the specific embodiments, which act on the operation processes such as anomaly and type detection, splitting and correcting entities, and missing value completion. In addition, users can also participate in each operation process, that is, user verification is performed to improve the accuracy of cleaning.

[0091] This embodiment can experiment with the above framework and obtain an effective cleaning method that integrates logic-based rules, knowledge graphs, and machine learning model-based rules. The accuracy of the present application in distinguishing attribute error categories (whether the entity tuple needs to be segmented or only needs to be corrected), the accuracy of attribute assignment and correction, and the accuracy of missing value completion are improved by 31.8%, 8.3%, and 39.5% respectively compared to existing error correction methods. In addition, compared to using only rule-based methods or only machine learning model-based methods, the accuracy can be improved by 35.5% and 30.3% respectively; at the same time, the execution efficiency of the method of the present application has been improved to a certain extent, such as processing 1,057,217 entity tuples in only 1,481 seconds.

[0092] See Figure 7, which is a flow chart of a method for cleaning data of mismatched entities provided in Example 6 of the present application. As shown in Figure 7, after mapping and extracting the objects of the entity tuple in step S301 to obtain the object set of the entity tuple, the following steps may be included:

[0093] Step S701 : detecting whether the object set contains one object. If it is detected that the object set contains one object, determining that the entity type of the entity tuple is a single entity type.

[0094] Step S702 : determining that the entity tuple is a single entity, cleaning the single entity to obtain a cleaned single entity, wherein the entity tuple corresponding to the single entity type is represented as an object.

[0095] In this embodiment, as for the single entity type, it has been explained in the above embodiment and will not be repeated here.

[0096] See Figure 8, which is a flow chart of a method for cleaning data of an incorrectly matched entity provided in Example 7 of the present application. As shown in Figure 8, the cleaning of a single entity in step S702 to obtain a cleaned single entity may include the following steps:

[0097] Step S801: Set the attribute value of each abnormal attribute in a single entity to empty.

[0098] Step S802: For any abnormal attribute whose attribute value is empty, a corresponding matching attribute value is matched according to the single entity and the abnormal attribute.

[0099] Step S803: Fill in the abnormal attributes with matching attribute values ​​to obtain a cleaned single entity.

[0100] In this embodiment, the exceptions in a single entity do not include conflict exceptions, which are generally empty set exceptions and error exceptions. For error exceptions, the attribute value of the exception attribute can be left empty and matched and filled together with the exception attribute of the empty set exception to obtain a cleaned single entity.

[0101] Of course, the empty set exception of a single entity can be matched and filled together with the empty set exception of the above mixed entity, that is, after splitting the above mixed entity into split entities, the split entity is cleaned as a single entity.

[0102] In the process of matching single entities and abnormal attributes to corresponding matching attribute values, the aforementioned rules, models, and knowledge graphs can also be used for matching, corresponding to logical completion, knowledge graph completion, and machine learning model completion, respectively. The idea behind logical completion is to use logic-based rules to complete missing attributes. For example, if the city of work is Shenzhen, then the province of work is Guangdong. Knowledge graph completion uses known attributes in the entity tuple to match entities in the knowledge graph G, and then uses the information in G to complete missing attribute values ​​in the entity tuple. For example, based on partial information about a movie, a corresponding movie match can be found in the Internet Movie Database (IMDb), and the missing movie information can be completed using information from IMDb. The input for machine learning model completion is: a pre-trained model M, known attribute values ​​t[A] in entity tuple t, the missing attribute B of t, and a set of possible values ​​for attribute B, Cand(B). For each possible value val in Cand(B), model M calculates the correlation strength score between t[A] and val, and ultimately uses the value val that achieves the maximum correlation score to complete t[B].

[0103] It should be known that among the above three completion strategies, the first two (i.e., logical completion and knowledge graph completion) are preferred because they have higher credibility. When the first two strategies fail, machine learning model completion can be used to fill in missing values.

[0104] Corresponding to the method for cleaning the data of mismatched entities in the above embodiment, FIG9 shows a structural block diagram of the device for cleaning the data of mismatched entities provided in the eighth embodiment of the present application. The above cleaning device is applied to the server in FIG1 . The user corresponding to the client sends the data to be cleaned (i.e., data tuples) to the server to trigger the server to perform the cleaning task. For ease of explanation, only the part relevant to the embodiment of the present application is shown.

[0105] Referring to FIG9 , the cleaning device comprises:

[0106] Anomaly detection module 91 is used to perform anomaly detection on the acquired data tuples to obtain entity tuples characterized as abnormal and abnormal attributes in the entity tuples, wherein the data tuples include at least one entity tuple;

[0107] a type detection module 92 for detecting an entity type of the entity tuple, and determining that the entity tuple is a mixed entity if the entity type of the entity tuple is detected to be a mixed entity type, wherein the entity tuple corresponding to the mixed entity type represents more than one type of object;

[0108] An entity segmentation module 93 is configured to construct a segmentation entity corresponding to each abnormal attribute according to each abnormal attribute and the corresponding attribute value in the mixed entity, wherein each segmentation entity includes an abnormal attribute and a corresponding attribute value;

[0109] The mixed entity cleaning module 94 is configured to perform abnormality correction on each segmented entity to obtain a corrected segmented entity, and determine that all corrected segmented entities are the results of cleaning the mixed entity.

[0110] Optionally, the type detection module 92 includes:

[0111] A knowledge graph mapping unit is used to obtain a preset knowledge graph, and map and extract the objects of the entity tuple according to the preset knowledge graph to obtain an object set of the entity tuple;

[0112] The type detection unit is configured to detect whether the object set contains at least two objects, and if it is detected that the object set contains at least two objects, determine that the entity type of the entity tuple is a mixed entity type.

[0113] Optionally, the anomaly detection module 91 includes:

[0114] An acquisition unit, used to acquire data tuples and preset anomaly detection rules;

[0115] An anomaly detection unit is used to perform an anomaly detection on each attribute of any entity tuple in the data tuple using a preset anomaly detection rule, and obtain an anomaly attribute whose detection result is abnormal;

[0116] The abnormality determination unit is used to determine that any entity tuple is an entity tuple representing an abnormality if an abnormal attribute exists in any entity tuple, and to determine the abnormal attribute belonging to the entity tuple from all abnormal attributes.

[0117] Optionally, the entity cleaning module 94 includes:

[0118] An attribute value detection unit, configured to detect, for each segmented entity, whether an attribute value of an abnormal attribute corresponding to the segmented entity is empty;

[0119] A first attribute value matching unit is configured to match a correct attribute value based on the segmented entity and the corresponding abnormal attribute in combination with a preset matching rule and / or matching model if it is detected that the attribute value of the segmented entity corresponding to the abnormal attribute is not empty;

[0120] The attribute value replacement unit is used to replace and correct the attribute value of the segmented entity corresponding to the abnormal attribute with the correct attribute value to obtain a corrected segmented entity.

[0121] Optionally, the entity cleaning module 94 further includes:

[0122] The attribute value completion unit is used to detect whether the attribute value of the abnormal attribute corresponding to the segmented entity is empty. If it is detected that the attribute value of the abnormal attribute corresponding to the segmented entity is empty, the preset completion rules, completion knowledge graph and / or completion model are used to complete and correct the abnormal attribute corresponding to the segmented entity to obtain the corrected segmented entity.

[0123] Optionally, the cleaning device further comprises:

[0124] A single entity type determination module is used to determine that the entity type of the entity tuple is a single entity type if it is detected that the object set contains one object after mapping and extracting the objects of the entity tuple to obtain the object set of the entity tuple;

[0125] The single entity cleaning module is used to determine that the entity tuple is a single entity, clean the single entity, and obtain the cleaned single entity, wherein the entity tuple corresponding to the single entity type is represented as an object.

[0126] Optionally, a single entity cleaning module includes:

[0127] An attribute value preprocessing unit, used for setting the attribute value of each abnormal attribute in a single entity to empty;

[0128] A second attribute value matching unit is used to match any abnormal attribute whose attribute value is empty to a corresponding matching attribute value according to the single entity and the abnormal attribute;

[0129] The attribute value filling unit is used to fill the abnormal attributes with matching attribute values ​​to obtain a cleaned single entity.

[0130] It should be noted that the information interaction, execution process and other contents between the above modules are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0131] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as shown in FIG10 . The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a readable storage medium, and a database. The internal memory provides an environment for the operation of the operating system and the readable storage medium in the non-volatile storage medium. The database of the computer device is used to store user original data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the readable storage medium is executed by the processor, a method for cleaning data of an incorrectly matched entity is implemented.

[0132] In one embodiment, a computer device is provided, comprising a memory, a processor, and a readable storage medium stored in the memory and executable by the processor. When the processor executes the readable storage medium, the steps of the method for cleaning data of mismatched entities in the above-described embodiment are implemented, such as steps S201-S204 shown in FIG. 2 , or the steps shown in FIG. 3 through FIG. 8 . To avoid repetition, these steps are not described here. Alternatively, when the processor executes the readable storage medium, the functions of the various modules / units in the embodiment of the user data processing device are implemented, such as the functions of the anomaly detection module 91, type detection module 92, entity segmentation module 93, and mixed entity cleaning module 94 shown in FIG. 9 . To avoid repetition, these steps are not described here.

[0133] In one embodiment, one or more readable storage media storing computer-readable instructions are provided. The computer-readable storage media stores computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors implement the steps of the method for cleaning data of mismatched entities in the above-mentioned embodiment, such as steps S201-S204 shown in FIG2 , or the steps shown in FIG3 to FIG8 . To avoid repetition, they are not described here. Alternatively, when the processor executes the readable storage medium, the functions of each module / unit in this embodiment of the user data processing device are implemented, such as the functions of the anomaly detection module 91, type detection module 92, entity segmentation module 93 and mixed entity cleaning module 94 shown in FIG9 . To avoid repetition, they are not described here. The readable storage medium in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.

[0134] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a readable storage medium, and the readable storage medium can be stored in a non-volatile computer-readable storage medium. When the readable storage medium is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0135] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0136] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for cleaning mis-matched entity data, wherein, the cleaning method includes: Performing anomaly detection on the obtained data tuples to obtain entity tuples characterized as anomalies and the anomalous attributes in the entity tuples, wherein the data tuples include at least one of the entity tuples; Detecting the entity type of the entity tuples, and if the detected entity type of the entity tuples is a mixed entity type, determining the entity tuples as mixed entities, wherein the entity tuples corresponding to the mixed entity type represent more than one object; Constructing a segmented entity corresponding to each anomalous attribute according to each anomalous attribute and the corresponding attribute value in the mixed entity, wherein each segmented entity includes an anomalous attribute and the corresponding attribute value; Performing anomaly correction on each segmented entity respectively to obtain corrected segmented entities, and determining all the corrected segmented entities as the result of cleaning the mixed entity.

2. The cleaning method according to claim 1, wherein, the detecting the entity type of the entity tuples includes: Obtaining a preset knowledge graph, and according to the preset knowledge graph, performing mapping extraction on the objects of the entity tuples to obtain an object set of the entity tuples; Detecting whether the object set contains at least two objects, and if it is detected that the object set contains at least two objects, determining that the entity type of the entity tuples is a mixed entity type.

3. The cleaning method according to claim 1, wherein, the performing anomaly detection on the obtained data tuples to obtain entity tuples characterized as anomalies and the anomalous attributes in the entity tuples includes: Obtaining data tuples and preset anomaly detection rules; For any one entity tuple in the data tuples, using the preset anomaly detection rules to perform anomaly detection on each attribute of the any one entity tuple to obtain anomalous attributes with a detection result of anomaly; If there are the anomalous attributes in the any one entity tuple, determining the any one entity tuple as an entity tuple characterized as an anomaly, and determining the anomalous attributes belonging to the entity tuple from all the anomalous attributes.

4. The cleaning method according to any one of claims 1 to 3, wherein, the performing anomaly correction on each segmented entity respectively to obtain corrected segmented entities includes: For any one segmented entity in each segmented entity, detecting whether the attribute value corresponding to the anomalous attribute of the segmented entity is empty; If it is detected that the attribute value corresponding to the anomalous attribute of the segmented entity is not empty, then according to the segmented entity and the corresponding anomalous attribute, combining preset matching rules and / or a matching model, matching the correct attribute value; Using the correct attribute value to replace and correct the attribute value corresponding to the anomalous attribute of the segmented entity to obtain a corrected segmented entity.

5. The cleaning method according to claim 4, wherein, after the detecting whether the attribute value corresponding to the anomalous attribute of the segmented entity is empty, it further includes: If it is detected that the attribute value of the abnormal attribute corresponding to the segmented entity is empty, the preset completion rule, completion knowledge graph, and / or completion model are used to complete and correct the abnormal attribute corresponding to the segmented entity, and the corrected segmented entity is obtained.

6. The cleaning method according to claim 2, wherein, after mapping and extracting the object of the entity tuple to obtain the object set of the entity tuple, it further includes: detecting whether there is one object in the object set, and if it is detected that there is one object in the object set, determining that the entity type of the entity tuple is a single entity type; determining that the entity tuple is a single entity, and cleaning the single entity to obtain the cleaned single entity, wherein the entity tuple corresponding to the single entity type represents one type of object.

7. The cleaning method according to claim 6, wherein, the cleaning the single entity to obtain the cleaned single entity includes: setting the attribute value of each abnormal attribute in the single entity to be empty; for any abnormal attribute with an empty attribute value, matching the corresponding matching attribute value according to the single entity and the abnormal attribute; using the matching attribute value to fill the abnormal attribute to obtain the cleaned single entity.

8. A cleaning device for data of mis-matched entities, wherein, the cleaning device includes: an anomaly detection module, configured to perform anomaly detection on the acquired data tuple to obtain an entity tuple characterized as abnormal and the abnormal attributes in the entity tuple, wherein the data tuple includes at least one of the entity tuples; a type detection module, configured to detect the entity type of the entity tuple, and if it is detected that the entity type of the entity tuple is a mixed entity type, determining that the entity tuple is a mixed entity, wherein the entity tuple corresponding to the mixed entity type represents more than one type of object; an entity segmentation module, configured to construct a segmented entity corresponding to each abnormal attribute according to each abnormal attribute and the corresponding attribute value in the mixed entity, wherein each segmented entity includes an abnormal attribute and the corresponding attribute value; a mixed entity cleaning module, configured to perform anomaly correction on each segmented entity respectively to obtain the corrected segmented entity, and determining all the corrected segmented entities as the result of cleaning the mixed entity.

9. A computer device, including a memory, a processor, and a readable storage medium stored in the memory and executable on the processor, wherein, when the processor executes the readable storage medium, the following steps are implemented: performing anomaly detection on the acquired data tuple to obtain an entity tuple characterized as abnormal and the abnormal attributes in the entity tuple, wherein the data tuple includes at least one of the entity tuples; detecting the entity type of the entity tuple, and if it is detected that the entity type of the entity tuple is a mixed entity type, determining that the entity tuple is a mixed entity, wherein the entity tuple corresponding to the mixed entity type represents more than one type of object; Construct a segmentation entity corresponding to each abnormal attribute according to each abnormal attribute and the corresponding attribute value in the mixed entity, where each segmentation entity contains an abnormal attribute and the corresponding attribute value; Perform anomaly correction on each segmentation entity respectively to obtain a corrected segmentation entity, and determine that all the corrected segmentation entities are the result of cleaning the mixed entity.

10. The computer device according to claim 9, wherein, the detection of the entity type of the entity tuple includes: Obtain a preset knowledge graph, and perform mapping extraction on the object of the entity tuple according to the preset knowledge graph to obtain an object set of the entity tuple; Detect whether the object set contains at least two objects. If it is detected that the object set contains at least two objects, determine that the entity type of the entity tuple is a mixed entity type.

11. The computer device according to claim 9, wherein, the anomaly detection of the obtained data tuple to obtain an entity tuple characterized as abnormal and the abnormal attributes in the entity tuple includes: Obtain a data tuple and a preset anomaly detection rule; For any entity tuple in the data tuple, use the preset anomaly detection rule to perform anomaly detection on each attribute of the any entity tuple to obtain an abnormal attribute with a detection result of abnormality; If there is the abnormal attribute in the any entity tuple, determine that the any entity tuple is an entity tuple characterized as abnormal, and determine the abnormal attributes belonging to the entity tuple from all the abnormal attributes.

12. The computer device according to any one of claims 9 to 11, wherein, the performing anomaly correction on each segmentation entity respectively to obtain a corrected segmentation entity includes: For any segmentation entity in each segmentation entity, detect whether the attribute value of the abnormal attribute corresponding to the segmentation entity is empty; If it is detected that the attribute value of the abnormal attribute corresponding to the segmentation entity is not empty, match the correct attribute value according to the segmentation entity and the corresponding abnormal attribute, in combination with a preset matching rule and / or matching model; Use the correct attribute value to replace and correct the attribute value of the abnormal attribute corresponding to the segmentation entity to obtain a corrected segmentation entity.

13. The computer device according to claim 12, wherein, after the detection of whether the attribute value of the abnormal attribute corresponding to the segmentation entity is empty, it further includes: If it is detected that the attribute value of the abnormal attribute corresponding to the segmentation entity is empty, use a preset completion rule, completion knowledge graph and / or completion model to perform completion correction on the abnormal attribute corresponding to the segmentation entity to obtain the corrected segmentation entity.

14. The computer device according to claim 10, wherein, after the performing mapping extraction on the object of the entity tuple to obtain the object set of the entity tuple, it further includes: Detect whether there is one object in the object set. If it is detected that the object set contains one object, determine that the entity type of the entity tuple is a single entity type; Determine that the entity tuple is a single entity, and clean the single entity to obtain a cleaned single entity, where the entity tuple corresponding to the single entity type represents an object.

15. The computer device according to claim 14, wherein, the cleaning the single entity to obtain a cleaned single entity includes: setting the attribute value of each abnormal attribute in the single entity to be empty; for any abnormal attribute with an empty attribute value, matching a corresponding matching attribute value according to the single entity and the abnormal attribute; using the matching attribute value to fill the abnormal attribute to obtain a cleaned single entity.

16. One or more readable storage media storing computer-readable instructions, the computer-readable storage media storing computer-readable instructions, wherein, when the computer-readable instructions are executed by one or more processors, the one or more processors are caused to perform the following steps: perform anomaly detection on the obtained data tuples to obtain entity tuples characterized as anomalies and the abnormal attributes in the entity tuples, where the data tuples include at least one of the entity tuples; detect the entity type of the entity tuple, and if it is detected that the entity type of the entity tuple is a mixed entity type, determine that the entity tuple is a mixed entity, where the entity tuple corresponding to the mixed entity type represents more than one object; construct a split entity corresponding to each abnormal attribute according to each abnormal attribute and the corresponding attribute value in the mixed entity, where each split entity includes an abnormal attribute and the corresponding attribute value; perform anomaly correction on each split entity respectively to obtain a corrected split entity, and determine that all the corrected split entities are the result of cleaning the mixed entity.

17. The readable storage medium according to claim 16, wherein, the detecting the entity type of the entity tuple includes: obtain a preset knowledge graph, and perform mapping extraction on the objects of the entity tuple according to the preset knowledge graph to obtain an object set of the entity tuple; detect whether the object set contains at least two objects, and if it is detected that the object set contains at least two objects, determine that the entity type of the entity tuple is a mixed entity type.

18. The readable storage medium according to claim 16, wherein, the performing anomaly detection on the obtained data tuples to obtain entity tuples characterized as anomalies and the abnormal attributes in the entity tuples includes: obtain data tuples and preset anomaly detection rules; for any entity tuple in the data tuples, use the preset anomaly detection rules to perform anomaly detection on each attribute of the any entity tuple to obtain abnormal attributes with a detection result of anomaly; if there are abnormal attributes in the any entity tuple, determine that the any entity tuple is an entity tuple characterized as abnormal, and determine the abnormal attributes belonging to the entity tuple from all the abnormal attributes.

19. The readable storage medium according to any one of claims 16 to 18, wherein, the performing anomaly correction on each split entity respectively to obtain a corrected split entity includes: For any one of the segmented entities in each segmented entity, detect whether the attribute value corresponding to the abnormal attribute of the segmented entity is empty; If it is detected that the attribute value corresponding to the abnormal attribute of the segmented entity is not empty, then according to the segmented entity and the corresponding abnormal attribute, combined with a preset matching rule and / or matching model, match the correct attribute value; Use the correct attribute value to replace and correct the attribute value corresponding to the abnormal attribute of the segmented entity to obtain the corrected segmented entity.

20. The readable storage medium according to claim 19, wherein, after detecting whether the attribute value corresponding to the abnormal attribute of the segmented entity is empty, further comprising: If it is detected that the attribute value corresponding to the abnormal attribute of the segmented entity is empty, then use a preset completion rule, completion knowledge graph and / or completion model to complete and correct the abnormal attribute corresponding to the segmented entity to obtain the corrected segmented entity.

Citation Information

Patent Citations

  • Entity identification method and system

    CN108491373A

  • Name correction method, device, electronic equipment and storage medium

    CN113947073A

  • Processing method for identity discrimination and data self-complementation of multi-source primary and secondary entities

    CN114969041A

  • Methods and systems for data cleaning

    US20160004743A1

Cited By

  • Cross-domain entity identity matching and information fusion method and system

    CN121278657A

  • Cross-domain entity identity matching and information fusion methods and systems

    CN121278657B