Automatic entity splitting method and apparatus, device, and medium
By splitting the properties in the entity tuple and deriving the attribute value of the attributes in the entity tuple using a preset logical rule group, the accuracy problem of automated entity splitting in the prior art is solved, and high-accuracy data processing is achieved.
Patent Information
- Application Number
- PCT/CN2023/137742
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-21
- Filing Date
- 2023-12-11
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to automatically and accurately split the wrongly merged entity tuples, making it difficult to guarantee data accuracy.
By obtaining the entity tuple to be split and the preset logical rule group, using the logical rule group to split the attributes in the entity tuple, perform attribute value derivation and conflict resolution, and determine the final attribute value, thereby achieving automated entity splitting.
It realizes automatic accurate splitting of entities, improving the use effect of data in scenarios with high accuracy requirements.
Smart Images

Figure CN2023137742_30052025_PF_FP_ABST
Abstract
Description
Automated entity splitting method, device, equipment and medium
[0001] This application is based on the Chinese invention application with application number 202311563603.8 filed on November 21, 2023, entitled “Automated entity splitting method, device, equipment and medium”, and claims priority. Technical Field
[0002] The present application is applicable to the field of data processing technology, and in particular relates to an automated entity splitting method, device, equipment and medium. Background Art
[0003] Currently, one of the most studied topics in data quality is entity resolution. The entity resolution problem is to identify tuples belonging to the same entity and merge these tuples into one tuple. Entity resolution is also known as record linking, deduplication, merging / cleaning, and record matching. It has been a routine operation in many applications. Through machine learning or logical rules, a lot of research has been done on entity resolution.
[0004] A related problem involves splitting merged tuples consisting of mismatched entities. In practice, different entities may be mistakenly merged into the same tuple. For example, in user-contributed projects, mismatched merged tuples can be observed due to users collaborating on data independently. In third-party or public data, the data has been processed by other parties, and since the third-party data processing methods may not be correct, different entities may also be mistakenly matched to the same tuple. In data lineage, data cleaning always starts from a checkpoint. Mismatches in entity tuples may be caused by historical operations and are difficult to detect. Entity splitting can effectively break down erroneous data. Entity splitting is the opposite of entity resolution. For example, role-based splitting can split person attributes by their roles. However, in most applications, manual entity splitting can only be performed by moving attributes one by one. However, since the merging method used during entity resolution is unknown, it is impossible to identify the key points of the mismerge based on the merging method, making it impossible to accurately split the entities.
[0005] Therefore, how to automatically split entities and obtain accurate splitting results in order to cope with scenarios with high data accuracy requirements has become an urgent problem to be solved.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide an automated entity splitting method, apparatus, device, and medium to solve the problem of how to automatically split entities and obtain accurate splitting results, so as to cope with scenarios with high data accuracy requirements.
[0008] An automated entity splitting method, the automated entity splitting method comprising:
[0009] Obtaining an entity tuple to be split and a preset logical rule group, and using the logical rule group to split attributes in the entity tuple that represent the same entity into split tuples, to obtain at least one split tuple;
[0010] Using the logical rule group to deduce attribute values for all attributes in each split tuple, obtaining all attribute value results for each attribute in each split tuple;
[0011] According to all attribute value results of each attribute in each split tuple, final attribute values of all attributes in each split tuple are determined.
[0012] An automated entity splitting device, comprising:
[0013] A tuple splitting module is used to obtain an entity tuple to be split and a preset logical rule group, and use the logical rule group to split the attributes of the entity tuple that represent the same entity into a split tuple to obtain at least one split tuple;
[0014] An attribute derivation module is used to derive attribute values for all attributes in each split tuple using the logical rule group to obtain all attribute value results for each attribute in each split tuple;
[0015] The attribute determination module is used to determine the final attribute values of all attributes in each split tuple according to all attribute value results under each attribute in each split tuple.
[0016] A computer device includes a memory, a processor, and a readable storage medium stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the readable storage medium:
[0017] Obtaining an entity tuple to be split and a preset logical rule group, and using the logical rule group to split attributes in the entity tuple that represent the same entity into split tuples, to obtain at least one split tuple;
[0018] Using the logical rule group to deduce attribute values for all attributes in each split tuple, obtaining all attribute value results for each attribute in each split tuple;
[0019] According to all attribute value results of each attribute in each split tuple, final attribute values of all attributes in each split tuple are determined.
[0020] One or more computer-readable storage media storing computer-readable instructions, wherein the computer-readable instructions, when executed by one or more processors, cause the one or more processors to perform the following steps:
[0021] Obtaining an entity tuple to be split and a preset logical rule group, and using the logical rule group to split attributes in the entity tuple that represent the same entity into split tuples, to obtain at least one split tuple;
[0022] Using the logical rule group to deduce attribute values for all attributes in each split tuple, obtaining all attribute value results for each attribute in each split tuple;
[0023] According to all attribute value results of each attribute in each split tuple, final attribute values of all attributes in each split tuple are determined.
[0024] The present application obtains an entity tuple to be split and a preset logical rule group, uses the logical rule group to split the attributes of the entity tuple that represent the same entity into a split tuple, and obtains at least one split tuple, uses the logical rule group to deduce attribute values for all attributes in each split tuple, and obtains all attribute value results under each attribute in each split tuple, and determines the final attribute values of all attributes in each split tuple based on all attribute value results under each attribute in each split tuple, thereby realizing automatic splitting of the entity, and deriving and verifying the attribute values after the splitting, thereby obtaining a more accurate splitting result, which is helpful for using data in scenarios with higher accuracy requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0026] FIG1 is a schematic diagram of an application environment of an automated entity splitting method provided in Example 1 of the present application;
[0027] FIG2 is a flow chart of an automated entity splitting method provided in Example 2 of the present application;
[0028] FIG3 is a flow chart of an automated entity splitting method provided in Example 3 of the present application;
[0029] FIG4 is a flow chart of an automated entity splitting method provided in Example 4 of the present application;
[0030] FIG5 is a flow chart of an automated entity splitting method provided in Example 5 of the present application;
[0031] FIG6 is a schematic diagram of a framework of an automated entity splitting method provided in Example 6 of the present application;
[0032] FIG7 is a flow chart of an automated entity splitting method provided in Example 7 of the present application;
[0033] FIG8 is a flow chart of an automated entity splitting method provided in Example 8 of the present application;
[0034] FIG9 is a schematic structural diagram of an automated entity splitting device provided in Example 9 of the present application;
[0035] FIG10 is a schematic structural diagram of a computer device provided in Example 10 of the present application. DETAILED DESCRIPTION
[0036] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0037] In order to illustrate the technical solution of the present application, specific embodiments are provided below.
[0038] An automated entity splitting method provided in the first embodiment of the present application can be applied in an application environment such as that shown in FIG1 , wherein a server communicates with a client, and the server is used to carry the automated entity splitting method to provide data splitting services. The client can request data splitting services from the server by providing corresponding entity tuples, thereby obtaining split tuples. Of course, the server can also split the entity tuples stored in itself. The client includes but is not limited to PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud-based computer devices, personal digital assistants (PDAs), and other devices. The computer device corresponding to the server can be implemented using an independent server or a server cluster consisting of multiple servers.
[0039] Referring to Figure 2, which is a flow chart of an automated entity splitting method provided in Example 2 of the present application, the automated entity splitting method is applied to the server in Figure 1. The user corresponding to the client sends the entity tuple to be split to the server, triggering the server to perform the splitting task. As shown in Figure 2, the automated entity splitting method may include the following steps:
[0040] Step S201 : obtaining an entity tuple to be split and a preset logic rule group, and using the logic rule group to split attributes in the entity tuple that represent the same entity into a split tuple, thereby obtaining at least one split tuple.
[0041] In this embodiment, a tuple is a row or a column in a relational table in a database. If a row is represented by a tuple, the row of data corresponds to an entity. Each column in the row of data represents the attributes of the entity. The entity tuple is a row of data in the relational table.
[0042] In the process of collecting and organizing data, the attributes of two different entities may be mistakenly merged into one entity. For example, there are two entities named A. The gender attribute of the first entity A is "male (that is, the attribute value of the gender attribute)", and the age attribute of the second entity A is "25 years old (the attribute value of the age attribute)". After the two entities are mistakenly merged, the entity A is obtained, with the gender attribute of "male" and the age attribute of "25 years old". Obviously, this is an incorrect fusion result.
[0043] It can be understood that the attribute values of the attributes in the incorrectly merged entity tuple represent more than one object. The object may refer to an object, item, person, place, etc. in the real world. For example, for an entity tuple, the corresponding entity is represented by a name. The entity name is Name1, which corresponds to multiple objects in reality. The corresponding gender attribute in its entity tuple is "male", and the corresponding medical attribute is "gynecology". Based on the principle of conflict, the gender attribute and the medical attribute are conflicting attributes. Based on this, the entity name Name1 can be divided into two split entities, such as the first split entity and the second split entity, where the name of the first split entity is Name1, the gender attribute is "male", and the medical attribute is empty. The name of the second split entity is also Name1, the gender attribute is empty, and the medical attribute is "gynecology".
[0044] After splitting, for the same entity, its attributes should be classified into a split tuple, and the split tuple can only include the attributes split from the entity tuple. For example, if the entity tuple contains attribute 1, attribute 2, and attribute 3, attribute 1 and attribute 2 are split into a split tuple 1, and attribute 3 is split into an attribute tuple 2. Of course, the attributes in the split tuple can be set to be equivalent to the attributes in the entity tuple. For attributes without attribute values, they can be left blank, that is, if the entity tuple contains attribute 1, attribute 2, and attribute 3, the split tuple should also contain attribute 1, attribute 2, and attribute 3.
[0045] The preset logical rule group includes at least one logical rule, wherein a logical rule φ consists of a rule condition X and a result Y. If X holds, then Y is also considered to hold. In other words, φ represents the result Y obtained through logical deduction based on condition X. The logical rules in the preset logical rule group can be obtained offline, or in real time before executing the method of this embodiment.
[0046] Logical rules can be used to determine whether an entity tuple contains the aforementioned incorrect fusions, thereby determining whether the entity tuple needs to be split. Incorrectly fused entities require splitting. Furthermore, logical rules can be used to assign values to the attributes of an entity tuple. This means that logical rules can be used to match specific attributes to possible values, which are then used for assignment. In a database, a logical rule φ can be applied to different data. h represents an application of a logical rule φ.
[0047] The logic rules in this embodiment can be obtained by manual design or by using an automated rule discovery method. Specifically, the condition X and result Y of the logic rule can be enumerated manually or by an automated discovery method.
[0048] Step S202: deduce attribute values for all attributes in each split tuple using a logical rule group to obtain all attribute value results for each attribute in each split tuple.
[0049] In this embodiment, for the split tuples obtained by the above splitting, the attributes therein may correspond to incorrect attribute values or be empty. In order to ensure the correctness and validity of the split tuples, it is necessary to deduce the attributes in the split tuples to obtain accurate attribute values.
[0050] For any attribute in any split tuple, use the logical rule group to deduce the attribute value of the attribute. If the logical rule group includes more than one logical rule, it is necessary to use each logical rule to deduce the attribute value of the attribute separately in an iterative manner. Each logical rule corresponds to an attribute value result after the attribute value is deduced, and then there is at least one attribute value result under this attribute.
[0051] The attribute value derivation process is essentially a process of repairing the attribute value of the attribute, that is, applying logical rules to repair it. The repair result is the above attribute value result. These attribute value results can be stored in a set In , each attribute value result can be represented in the form of (t[A], c), where (t[A], c) indicates that the attribute value of attribute A of a split tuple t is c.
[0052] Using a logical rule φ for deduction, it can be characterized as:
[0053] where h is the application of the logical rule φ, The set of attribute value results that existed before the logical rule φ was derived. In order to add a new set of attribute value results, the application of the logic rule φ needs to meet the following conditions: φ is a valid logic rule, and There is a new attribute value result derived by h, that is, the application of φ in A new (t[A], c) is added, where a valid logical rule may refer to a condition X in the logical rule that has passed the verification of the verification data. The logical rule that has passed the verification data is deduced. As long as the logical rule and the verification data are correct, the deduction result is also correct.
[0054] Step S203: Determine the final attribute values of all attributes in each split tuple according to all attribute value results under each attribute in each split tuple.
[0055] In this embodiment, since there may be multiple attribute value results obtained through the above derivation, and for an attribute in a split tuple, the attribute value needs to be uniquely determined, it is necessary to analyze the attribute value results obtained through the above derivation to obtain the final attribute value of each attribute.
[0056] That is, if there are two attribute value results for an attribute, and there is a conflict between the two attribute value results, then a corresponding conflict resolution method needs to be used to obtain the exact attribute value of the attribute.
[0057] At this point, it should be considered that the above step S202 can deduce all attribute value sets of t[A], which are recorded as attribute value set C. The final attribute value of t[A] needs to be determined based on the attribute value set C. Here, all attribute values in the attribute value set C are screened to obtain the final attribute value. The screening can be based on the calculation and comparison of data correlation, for example, the calculation of the correlation between the attribute values in the attribute value set C and other attributes of the split tuple t.
[0058] The embodiment of the present application obtains an entity tuple to be split and a preset logical rule group, uses the logical rule group to split the attributes of the entity tuple that represent the same entity into a split tuple, and obtains at least one split tuple, uses the logical rule group to deduce attribute values for all attributes in each split tuple, and obtains all attribute value results under each attribute in each split tuple, and determines the final attribute values of all attributes in each split tuple based on all attribute value results under each attribute in each split tuple, thereby realizing automated splitting of the entity, and deriving and verifying the attribute values after the splitting, thereby obtaining a more accurate splitting result, which is helpful for using data in scenarios with higher accuracy requirements.
[0059] See Figure 3, which is a flow chart of an automated entity splitting method provided in Example 3 of the present application. As shown in Figure 3, in step S201, the attributes representing the same entity in the entity tuple are split into split tuples using a logical rule group to obtain at least one split tuple. The following steps may also be included:
[0060] Step S301 : Using a logical rule group, determine whether the attribute value of the first attribute and the attribute value of the second attribute in the entity tuple represent the same entity.
[0061] Step S302: If the attribute value of the first attribute and the attribute value of the second attribute represent the same entity, it is determined that the attribute value of the first attribute and the attribute value of the second attribute both belong to a split tuple.
[0062] Step S303: If the attribute value of the first attribute and the attribute value of the second attribute represent different entities, it is determined that the attribute value of the first attribute belongs to one split tuple and the attribute value of the second attribute belongs to another split tuple.
[0063] In this embodiment, the first attribute and the second attribute are not the same attribute, and both are arbitrary attributes in the entity tuple. That is, to determine whether the attributes in the entity tuple represent the same entity, it is necessary to determine whether the attribute values of the two attributes should refer to the same entity. When the attribute values of the two attributes are determined by the logical rules, it is determined that the two attributes do not satisfy the logical rules, then the two attributes are determined to represent different entities. For example, the above-mentioned gender attribute is "male" and the medical attribute is "gynecology". The attribute values between the gender attribute and the medical attribute are determined to not belong to the same entity. The specific corresponding logical rules are: male corresponds to andrology, and female corresponds to gynecology. Therefore, the male and gynecology in the entity tuple do not belong to the same entity.
[0064] If two attributes represent the same entity, their values are split into a single split tuple. If the two attributes represent different entities, the value of one attribute is split into one split tuple, and the value of the other attribute is split into another split tuple. In this case, the split tuple can include the attributes that were split into it, as well as other attributes in the entity tuple. However, these other attributes do not have attribute values and can be obtained through subsequent derivation.
[0065] See Figure 4, which is a flow chart of an automated entity splitting method provided in Example 4 of the present application. As shown in Figure 4, the above step S203 determines the final attribute values of all attributes in each split tuple based on all attribute value results under each attribute in each split tuple, and may also include the following steps:
[0066] Step S401: for any split tuple, the attributes in the split tuple with attribute value conflicts are taken as attributes to be resolved.
[0067] In this embodiment, during the attribute value derivation process in step S202, for an attribute, if attribute value derivation is performed using one logical rule to obtain a first attribute value result, and then attribute value derivation is performed using another logical rule to obtain a second attribute value result, whether there is a conflict between the first attribute value result and the second attribute value result is determined. If there is a conflict, the attribute is a conflict-resolved attribute. Of course, conflict resolution is not required for this attribute only if all attribute value results for the attribute do not conflict with each other.
[0068] For a logical rule, after deducing the attribute value of an attribute, two attribute value results may be obtained. You can also refer to the above process and take the attribute as the research object. As long as there is a conflict in the attribute value results under the attribute, it is an attribute to be resolved.
[0069] The above-mentioned conflict may refer to an essential difference or a large difference between two attribute values. The method for determining the conflict may be set according to requirements. For example, if the two attribute values are numerical, they are considered to be different or the difference between them is greater than a threshold, and they are determined to be in conflict. If the two attribute values are character types, it is necessary to judge the character similarity or semantic similarity, and determine whether there is a conflict based on the similarity. For example, "andrology" and "gynecology" are two attribute values with essential semantic differences, and it can be determined that the two are in conflict.
[0070] Step S402: Obtain a verified attribute set of the split tuple.
[0071] In this embodiment, the verified attribute set includes verified attributes, and the corresponding attribute values of the verified attributes are confirmed to belong to the corresponding split entity. For example, for split entity 1, it includes attribute 1, attribute 2, and attribute 3, among which the attribute values corresponding to attribute 1 and attribute 2 are confirmed to belong to split entity 1. In this case, attribute 1 and attribute 2 are the verified attributes of the split tuple.
[0072] The verified attributes can be some attributes and corresponding attribute values set according to requirements. Of course, since the entity tuple is the data obtained through entity parsing, although there are fusion errors, its attribute values may be correct. Therefore, the attributes split from the entity tuple can also be used as the verified attributes in the split tuple. In addition, after the attributes of the split tuple are subsequently deduced, the derived attributes without conflict can also be used as the verified attributes in the split tuple.
[0073] Step S403: All attribute value results in the attribute to be resolved are screened according to the attribute values of all attributes in the verified attribute set, and the filtered attribute value results are used as the final attribute value of the attribute to be resolved.
[0074] In this embodiment, the attributes in the verified attribute set can indicate some characteristics of the entity corresponding to the split tuple to a certain extent. Therefore, the attribute value of the attribute to be solved should be related to the partial characteristics. Therefore, based on this, all attribute value results of the attribute to be solved can be screened to obtain the related attribute value as the final attribute value.
[0075] For a split tuple, when using the preset logical rule group to derive the attribute value of each attribute, since the preset logical rule group needs to derive each attribute one by one, if the currently derived attribute is the first derived attribute in the split tuple, and the first derived attribute is the attribute to be resolved, the preset verified attribute or the attribute split from the entity tuple can be used to filter all the attribute value results in the attribute to be resolved to obtain the final attribute value.
[0076] If the currently derived attribute is not the first derived attribute in the split tuple, and the attribute is the attribute to be resolved, you can use the preset verified attributes, the attributes split from the entity tuple, or the attributes with the final attribute value determined to filter all the attribute value results in the attribute to be resolved to obtain the final attribute value.
[0077] In an embodiment of the present application, the attribute value of the verified attribute can be used to screen the conflicting attribute value results to resolve the attribute value conflict problem under the attribute, thereby determining the final attribute value of the attribute and ensuring the uniqueness of the split tuple.
[0078] 5 is a flowchart of an automated entity splitting method provided in a fifth embodiment of the present application. As shown in FIG5 , obtaining a verified attribute set of a split tuple in step S402 may include the following steps:
[0079] Step S501: Attributes in the split tuple that do not have attribute value result conflicts and attributes to be resolved in the split tuple whose final attribute values have been determined are regarded as verified attributes.
[0080] Step S502: Obtain a verified attribute set based on the verified attributes.
[0081] In this embodiment, for a split tuple, if there is no conflict in the attribute value result corresponding to the attribute after performing attribute value derivation, then the attribute is not a conflicting attribute, and the corresponding attribute value result is its final attribute value, which can be used as the verified attribute of the split tuple. In addition, if the attribute is a conflicting attribute, but after conflict resolution (i.e., step S403 above), a final attribute value is obtained, it can also be used as the verified attribute of the split tuple.
[0082] In an embodiment of the present application, a verified attribute set is constructed based on the verified attributes. The verified attribute set can be continuously updated in the subsequent process of deducing the attributes of the split tuple, making the attributes in the set richer and helping to improve the accuracy of conflict attribute resolution.
[0083] 6 is a flowchart of an automated entity splitting method provided in Example 6 of the present application. As shown in FIG6 , the above step S403 filters all attribute value results of the attributes to be resolved based on the attribute values of all attributes in the verified attribute set, and obtains the filtered attribute value results as the final attribute value of the attribute to be resolved, which may include the following steps:
[0084] Step S601: Use the pre-trained association analysis model to calculate the association strength between each attribute value result in the attribute to be resolved and the attribute values of all attributes in the verified attribute set, and obtain the association strength score corresponding to each attribute value result in the attribute to be resolved.
[0085] In this embodiment, the pre-trained association analysis model can be a model based on a neural network architecture, which can perform correlation analysis on the input data and the target data and output a correlation strength score, wherein the input data is an attribute value result in the attribute to be solved, and the target data is the attribute value of all attributes in the verified attribute set.
[0086] Step S602 : Filter the association strength scores of all attribute value results in the attribute to be resolved and obtain the attribute value result with the highest association strength score as the target result, and determine the target result as the final attribute value of the attribute to be resolved.
[0087] In this embodiment, after the calculation of the above-mentioned step S601, the association strength score corresponding to each attribute value result in the attribute to be resolved is obtained, and the highest association strength score is determined from all the association strength scores. The highest association strength score is used as the target result, which indicates that the corresponding attribute value has the strongest correlation with the attributes in the verified attribute set. Therefore, the reliability of the target result being the correct attribute value is the highest, that is, the target result is used as the final attribute value of the attribute to be resolved.
[0088] For example, a conflict is found in the attribute value of t[A] in the split tuple t, and the conflict is resolved by the following strategy based on attribute relevance:
[0089] Assume that the set of all attribute values of t[A] that can be deduced through logical rules is C, and assume that the set of all verified attributes of t is By fine-tuning a large-scale language model (such as BERT), a model M for judging attribute relevance can be trained. Specifically, given a candidate attribute value result c in C, The candidate attribute value result c and all verified attributes will be output The association strength of t[A] is taken as the final attribute value of t[A].
[0090] See Figure 7, which is a flow chart of an automated entity splitting method provided in Example 7 of the present application. As shown in Figure 7, the logical rule group includes at least two logical rules. Step S202 uses the logical rule group to derive attribute values for all attributes in each split tuple to obtain all attribute value results for each attribute in each split tuple. The following steps may be included:
[0091] Step S701: Acquire verification data, verify each logic rule according to the verification data, and obtain the verified logic rules.
[0092] In this embodiment, the logical rule group contains more than two logical rules. Before using each logical rule to chase and execute each attribute, it is necessary to verify whether the logical rule is valid. At this time, verification data is needed. The verification data can be pre-set data for verifying condition X in the logical rule. The deduction results after each derivation can also be collected in the subsequent derivation process to update the verification data.
[0093] Step S702: for any split tuple, any attribute in the split tuple is used as the current derivation attribute, and each verified logical rule is called in turn to derive the attribute value of the current derivation attribute to obtain all attribute value results under the current derivation attribute.
[0094] In this embodiment, each logical rule is used for pursuit execution to obtain a pursuit sequence for the current derivation attribute, that is, the logical rules are {φ1, φ2, ..., φn}, and the corresponding logical rules are executed from 1 to n in sequence, and finally the attribute value result set of the current derivation attribute is obtained.
[0095] Step S703: traverse all attributes in the split tuple and all split tuples to obtain all attribute value results under each attribute in each split tuple.
[0096] In this embodiment, for a split tuple, each attribute needs to be processed through step S702 to obtain an attribute value result set. Thus, the chasing result for the split tuple is a chasing sequence from the derivation of the first attribute to the derivation of the last attribute, as follows:
[0097] in, The result set of attribute values derived for the first attribute. The resulting set of attribute values for the last attribute derivation. The chasing process ends when no logical rules can be applied anymore.
[0098] See Figure 8, which is a flow chart of an automated entity splitting method provided in Example 8 of the present application. As shown in Figure 8, the above-mentioned step S702 of deriving attribute values for the current derived attribute to obtain all attribute value results under the current derived attribute may include the following steps:
[0099] Step S801: derive an attribute value for a current derivation attribute using a first logic rule to obtain a first attribute value result.
[0100] Step S802: Use the second logic rule to derive the attribute value of the current derivation attribute to obtain a second attribute value result.
[0101] Step S803: Using a preset data structure, store the attribute value results after attribute value deduction for each verified logical rule.
[0102] In this embodiment, the first logic rule is any logic rule among the verified logic rules, and the second logic rule is a logic rule among the verified logic rules that is executed after the first logic rule is executed.
[0103] Among them, if it is detected that the first attribute value result is the same as the second attribute value result, one of the first attribute value result and the second attribute value result is retained as the attribute value result of the current derived attribute; if it is detected that the first attribute value result and the second attribute value result are different, the first attribute value result and the second attribute value result are retained as the attribute value results of the current derived attribute; that is, after using a logical rule to derive the attribute value, it is necessary to compare with the attribute value result that has been derived before to determine whether it is the same as the attribute value result that has been derived before. If it is the same, only one attribute value result needs to be retained. If it is different, both attribute value results need to be retained. In the subsequent judgment of whether it is a conflicting attribute to be resolved, if an attribute has two or more attribute value results, it is a conflicting attribute to be resolved. If an attribute has only one attribute value result, it is not a conflicting attribute.
[0104] If it is detected that the use of the second logical rule depends on the attribute value result of the first logical rule, the first attribute value result and the second logical rule in the preset data structure are called to perform attribute value derivation on the current derivation attribute to obtain the second attribute value result. Because enumerating the application of logical rules is relatively difficult, and the application of a logical rule may depend on the results of the application of other logical rules used in previous attribute value derivation, a dedicated data structure is designed to record attribute value results that can only be temporarily stored during the application of logical rules. This allows for efficient enumeration and reuse of logical rule applications based on the attribute value results stored in this data structure.
[0105] In addition, for the verification of logical rules, there is no need to check again the application of logical rules that have been used without exception. Based on the attribute value results and corresponding logical rules stored in the data structure, only the application of affected or unverified logical rules can be checked.
[0106] This embodiment is applied to a data set of the real Internet Movie Database (IMDb), and the overall splitting accuracy can reach 92%. In addition, when the number of tuples in the data set reaches 1,057,217, the running time of the method of the present application is only 1481 seconds. When the scaling factor of the data set changes from 20% to 100%, maintaining a preset data structure to record temporary attribute value results can effectively enumerate and reuse the application of logical rules, improve the efficiency of calculations, and run 13.5 times faster than the method of simply enumerating all application results. When more logical rules are used to perform derivation, the method of the present application requires more time. Even so, compared with the method of simply enumerating all application results, the running speed is still 11.8 times faster.
[0107] Corresponding to the automated entity splitting method described in the preceding embodiment, FIG9 shows a block diagram of an automated entity splitting apparatus according to a ninth embodiment of the present application. This automated entity splitting apparatus is applied to the server shown in FIG1 . The user corresponding to the client sends the entity tuple to be split to the server, triggering the server to execute the splitting task. For ease of illustration, only the portion relevant to the present embodiment is shown.
[0108] Referring to FIG9 , the automated entity splitting device includes:
[0109] The tuple splitting module 91 is configured to obtain an entity tuple to be split and a preset logical rule group, and to use the logical rule group to split the attributes of the entity tuple that represent the same entity into split tuples, thereby obtaining at least one split tuple.
[0110] An attribute derivation module 92 is configured to derive attribute values for all attributes in each split tuple using a logical rule group, and obtain all attribute value results for each attribute in each split tuple;
[0111] The attribute determination module 93 is used to determine the final attribute values of all attributes in each split tuple according to all attribute value results of each attribute in each split tuple.
[0112] Optionally, the tuple splitting module 91 includes:
[0113] an attribute value determination unit, configured to determine, using a logical rule group, whether an attribute value of a first attribute and an attribute value of a second attribute in the entity tuple represent the same entity, wherein the first attribute and the second attribute are not the same attribute and are both arbitrary attributes in the entity tuple;
[0114] a first splitting unit, configured to determine that the attribute value of the first attribute and the attribute value of the second attribute both belong to a split tuple if the attribute value of the first attribute and the attribute value of the second attribute represent the same entity;
[0115] The second splitting unit is configured to determine that the attribute value of the first attribute belongs to one split tuple and the attribute value of the second attribute belongs to another split tuple if the attribute value of the first attribute and the attribute value of the second attribute represent different entities.
[0116] Optionally, the attribute determination module 93 includes:
[0117] a conflicting attribute determining unit, configured to, for any split tuple, take the attributes in the split tuple that have conflicting attribute values as attributes to be resolved;
[0118] An attribute set determining unit, configured to obtain a verified attribute set of the split tuple, wherein the verified attribute set includes attributes that have passed verification;
[0119] The final attribute value determination unit is used to screen all attribute value results in the attribute to be resolved according to the attribute values of all attributes in the verified attribute set, and obtain the screened attribute value result as the final attribute value of the attribute to be resolved.
[0120] Optionally, the verified attribute determination unit includes:
[0121] The attribute determination subunit is used to treat the attributes in the split tuple that do not have attribute value conflict and the unresolved attributes in the split tuple whose final attribute values have been determined as verified attributes;
[0122] The attribute set determination subunit is used to obtain a verified attribute set based on the verified attributes.
[0123] Optionally, the final attribute value determination unit includes:
[0124] The association calculation subunit is used to use the pre-trained association analysis model to calculate the association strength between each attribute value result in the attribute to be solved and the attribute values of all attributes in the verified attribute set, and obtain the association strength score corresponding to each attribute value result in the attribute to be solved;
[0125] The attribute value determination subunit is used to screen the attribute value result with the highest association strength score from the association strength scores of all attribute value results in the attribute to be solved as the target result, and determine the target result as the final attribute value of the attribute to be solved.
[0126] Optionally, the logical rule group includes at least two logical rules, and the attribute derivation module 92 includes:
[0127] A rule verification unit is used to obtain verification data, verify each logical rule according to the verification data, and obtain the logical rules that pass the verification;
[0128] The attribute derivation unit, for any split tuple, takes any attribute in the split tuple as the current derivation attribute, calls each verified logical rule in turn, derives the attribute value of the current derivation attribute, and obtains all attribute value results under the current derivation attribute;
[0129] The loop execution unit is used to traverse all attributes in the split tuple and all split tuples to obtain all attribute value results under each attribute in each split tuple.
[0130] Optionally, the attribute derivation unit includes:
[0131] A first attribute derivation subunit is configured to derive an attribute value of a current derivation attribute using a first logic rule to obtain a first attribute value result, wherein the first logic rule is any logic rule among the verified logic rules;
[0132] A second attribute derivation subunit is configured to derive an attribute value for the current derivation attribute using a second logic rule to obtain a second attribute value result, wherein the second logic rule is a logic rule after the first logic rule is executed among the verified logic rules;
[0133] The attribute value storage subunit is used to store the attribute value results derived from the attribute value of each verified logical rule using a preset data structure;
[0134] If it is detected that the first attribute value result and the second attribute value result are the same, then one of the first attribute value result and the second attribute value result is retained as the attribute value result of the current attribute derivation; if it is detected that the first attribute value result and the second attribute value result are different, then the first attribute value result and the second attribute value result are retained as the attribute value result of the current attribute derivation;
[0135] If it is detected that the use of the second logic rule depends on the attribute value result of the first logic rule, the first attribute value result and the second logic rule in the preset data structure are called to derive the attribute value of the current derivation attribute to obtain the second attribute value result.
[0136] It should be noted that the information interaction, execution process and other contents between the above modules are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0137] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as shown in FIG10 . The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a readable storage medium, and a database. The internal memory provides an environment for the operation of the operating system and the readable storage medium in the non-volatile storage medium. The database of the computer device is used to store user original data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the readable storage medium is executed by the processor, an automated entity splitting method is implemented.
[0138] In one embodiment, a computer device is provided, comprising a memory, a processor, and a readable storage medium stored in the memory and executable on the processor. When the processor executes the readable storage medium, the steps of the automated entity splitting method in the above-described embodiment are implemented, such as steps S201-S203 shown in FIG2 , or the steps shown in FIG3 to FIG8 . To avoid repetition, these steps are not described here. Alternatively, when the processor executes the readable storage medium, the functions of the modules / units in the embodiment of the user data processing device are implemented, such as the functions of the tuple splitting module 91, the attribute derivation module 92, and the attribute determination module 93 shown in FIG9 . To avoid repetition, these steps are not described here.
[0139] In one embodiment, one or more readable storage media storing computer-readable instructions are provided, and the computer-readable storage media stores computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors implement the steps of the automated entity splitting method in the above-mentioned embodiment, such as steps S201-S203 shown in FIG2 , or the steps shown in FIG3 to FIG8 . To avoid repetition, they are not described here. Alternatively, when the processor executes the readable storage medium, the functions of each module / unit in this embodiment of the user data processing device are implemented, such as the functions of the tuple splitting module 91, the attribute derivation module 92 and the attribute determination module 93 shown in FIG9 . To avoid repetition, they are not described here. The readable storage medium in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.
[0140] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a readable storage medium, and the readable storage medium can be stored in a non-volatile computer-readable storage medium. When the readable storage medium is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0141] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0142] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. An automated entity splitting method, wherein, the automated entity splitting method includes: Obtaining an entity tuple to be split and a preset logical rule group, and using the logical rule group to split the attributes representing the same entity in the entity tuple into a split tuple, obtaining at least one split tuple; Using the logical rule group to perform attribute value derivation on all attributes in each split tuple, obtaining all attribute value results for each attribute in each split tuple; Determining the final attribute values of all attributes in each split tuple according to all the attribute value results for each attribute in each split tuple.
2. The automated entity splitting method according to claim 1, wherein, the step of using the logical rule group to split the attributes representing the same entity in the entity tuple into a split tuple, obtaining at least one split tuple, includes: Using the logical rule group to determine whether the attribute values of the first attribute and the second attribute in the entity tuple represent the same entity, wherein the first attribute and the second attribute are not the same attribute, and both are arbitrary attributes in the entity tuple; If the attribute values of the first attribute and the second attribute represent the same entity, determining that the attribute values of the first attribute and the second attribute both belong to a split tuple; If the attribute values of the first attribute and the second attribute represent different entities, determining that the attribute value of the first attribute belongs to a split tuple, and the attribute value of the second attribute belongs to another split tuple.
3. The automated entity splitting method according to claim 1, wherein, the step of determining the final attribute values of all attributes in each split tuple according to all the attribute value results for each attribute in each split tuple, includes: For any split tuple, taking the attributes with conflicting attribute value results in the split tuple as attributes to be resolved; Obtaining the verified attribute set of the split tuple, wherein the verified attribute set includes attributes that have passed verification; According to the attribute values of all attributes in the verified attribute set, screening all the attribute value results of the attributes to be resolved, and obtaining the screened attribute value results as the final attribute values of the attributes to be resolved.
4. The automated entity splitting method according to claim 3, wherein, the step of obtaining the verified attribute set of the split tuple includes: Taking the attributes without conflicting attribute value results in the split tuple, and the attributes to be resolved for which the final attribute values have been determined in the split tuple as attributes that have passed verification; Obtaining the verified attribute set according to the attributes that have passed verification.
5. The automated entity splitting method according to claim 3, wherein, the step of screening all the attribute value results of the attributes to be resolved according to the attribute values of all attributes in the verified attribute set, and obtaining the screened attribute value results as the final attribute values of the attributes to be resolved, includes: Using a pre-trained association analysis model, calculate the association strength between each attribute value result in the to-be-solved attributes and the attribute values of all attributes in the verified attribute set, and obtain the association strength score corresponding to each attribute value result in the to-be-solved attributes; Screen from the association strength scores of all attribute value results in the to-be-solved attributes to obtain the attribute value result with the highest association strength score as the target result, and determine the target result as the final attribute value of the to-be-solved attributes.
6. The automated entity splitting method according to claim 1, wherein, The logic rule group includes at least two logic rules. Using the logic rule group to perform attribute value derivation on all attributes in each split tuple to obtain all attribute value results under each attribute in each split tuple includes: Obtain verification data, and verify each logic rule according to the verification data to obtain the logic rules that pass the verification; For any split tuple, use any attribute in the split tuple as the current derivation attribute, and sequentially call each verified logic rule to perform attribute value derivation on the current derivation attribute to obtain the All attribute value results; Traverse all attributes in the split tuple and all split tuples to obtain all attribute value results under each attribute in each split tuple.
7. The automated entity splitting method according to claim 6, wherein, Performing attribute value derivation on the current derivation attribute to obtain all attribute value results under the current derivation attribute includes: Use the first logic rule to perform attribute value derivation on the current derivation attribute to obtain the first attribute value result, and the first logic rule is any logic rule in the verified logic rules; Use the second logic rule to perform attribute value derivation on the current derivation attribute to obtain the second attribute value result, and the second logic rule is the logic rule in the verified logic rules after the execution of the first logic rule; Use a preset data structure to store the attribute value results after attribute value derivation by each verified logic rule; Wherein, if it is detected that the first attribute result is the same as the second attribute value result, retain one of the first attribute value result and the second attribute value result as the attribute value result of the current derivation attribute. If it is detected that the first attribute value result is different from the second attribute value result, retain the first attribute value result and the second attribute value result as the attribute value result of the current derivation attribute; If it is detected that the use of the second logic rule depends on the attribute value result of the first logic rule, call the first attribute value result and the second logic rule in the preset data structure to perform attribute value derivation on the current derivation attribute to obtain the second attribute value result.
8. An automated entity splitting device, wherein, The automated entity splitting device includes: A tuple splitting module, configured to obtain an entity tuple to be split and a preset logical rule set, and use the logical rule set to split the attributes representing the same entity in the entity tuple into a split tuple, so as to obtain at least one split tuple; An attribute derivation module, configured to use the logical rule set to perform attribute value derivation on all attributes in each split tuple, so as to obtain all attribute value results for each attribute in each split tuple; An attribute determination module, configured to determine the final attribute value of all attributes in each split tuple according to all attribute value results for each attribute in each split tuple.
9. A computer device, including a memory, a processor, and a readable storage medium stored in the memory and operable on the processor, wherein, when the processor executes the readable storage medium, the following steps are implemented: Obtain an entity tuple to be split and a preset logical rule set, and use the logical rule set to split the attributes representing the same entity in the entity tuple into a split tuple, so as to obtain at least one split tuple; Use the logical rule set to perform attribute value derivation on all attributes in each split tuple, so as to obtain all attribute value results for each attribute in each split tuple; Determine the final attribute value of all attributes in each split tuple according to all attribute value results for each attribute in each split tuple.
10. The computer device according to claim 9, wherein, the step of using the logical rule set to split the attributes representing the same entity in the entity tuple into a split tuple, so as to obtain at least one split tuple, includes: Use the logical rule set to determine whether the attribute values of the first attribute and the second attribute in the entity tuple represent the same entity, wherein the first attribute and the second attribute are not the same attribute, and both are arbitrary attributes in the entity tuple; If the attribute values of the first attribute and the second attribute represent the same entity, it is determined that the attribute values of the first attribute and the second attribute both belong to a split tuple; If the attribute values of the first attribute and the second attribute represent different entities, it is determined that the attribute value of the first attribute belongs to a split tuple, and the attribute value of the second attribute belongs to another split tuple.
11. The computer device according to claim 9, wherein, the step of determining the final attribute value of all attributes in each split tuple according to all attribute value results for each attribute in each split tuple, includes: For any split tuple, use the attributes with conflicting attribute value results in the split tuple as the attributes to be resolved; Obtain the verified attribute set of the split tuple, wherein the verified attribute set includes the attributes that have passed verification; According to the attribute values of all attributes in the verified attribute set, screen all attribute value results of the attributes to be resolved, and obtain the screened attribute value results as the final attribute values of the attributes to be resolved.
12. The computer device according to claim 11, wherein, the step of obtaining the verified attribute set of the split tuple includes: Take the attributes in the split tuple that do not have conflicting attribute value results, and the to-be-solved attributes in the split tuple for which the final attribute values have been determined as the attributes that have passed verification. Obtain a set of verified attributes based on the attributes that have passed verification.
13. The computer device according to claim 11, wherein, The screening of all the attribute value results in the to-be-solved attributes according to the attribute values of all the attributes in the set of verified attributes to obtain the screened attribute value results as the final attribute values of the to-be-solved attributes includes: Using a pre-trained association analysis model to calculate the association strength between each attribute value result in the to-be-solved attributes and the attribute values of all the attributes in the set of verified attributes, and obtaining the association strength score corresponding to each attribute value result in the to-be-solved attributes; Screen from the association strength scores of all the attribute value results in the to-be-solved attributes to obtain the attribute value result with the highest association strength score as the target result, and determine the target result as the final attribute value of the to-be-solved attributes.
14. The computer device according to claim 9, wherein, The logical rule group includes at least two logical rules. The derivation of the attribute value results for all the attributes in each split tuple using the logical rule group includes: Obtain verification data, and verify each logical rule according to the verification data to obtain the verified logical rules; For any one split tuple, take any attribute in the split tuple as the current derivation attribute, and sequentially call each verified logical rule to perform attribute value derivation on the current derivation attribute to obtain all the attribute value results under the current derivation attribute; Traverse all the attributes in the split tuple and all the split tuples to obtain all the attribute value results under each attribute in each split tuple.
15. The computer device according to claim 14, wherein, The derivation of the attribute value results for the current derivation attribute to obtain all the attribute value results under the current derivation attribute includes: Use the first logical rule to perform attribute value derivation on the current derivation attribute to obtain the first attribute value result, and the first logical rule is any one of the verified logical rules; Use the second logical rule to perform attribute value derivation on the current derivation attribute to obtain the second attribute value result, and the second logical rule is the logical rule after the execution of the first logical rule among the verified logical rules; Use a preset data structure to store the attribute value results after the attribute value derivation by each verified logical rule; wherein, if it is detected that the first attribute result is the same as the second attribute value result, then retain one of the first attribute value result and the second attribute value result as the attribute value result of the current derivation attribute, and if it is detected that the first attribute value result is different from the second attribute value result, then retain the first attribute value result and the second attribute value result as the attribute value result of the current derivation attribute; If it is detected that the use of the second logical rule depends on the attribute value result of the first logical rule, then the first attribute value result and the second logical rule in the preset data structure are called to perform attribute value derivation on the current derived attribute to obtain the second attribute value result.
16. One or more readable storage media storing computer-readable instructions, the computer-readable storage media storing computer-readable instructions, wherein, when the computer-readable instructions are executed by one or more processors, the one or more processors are caused to perform the following steps: Obtain an entity tuple to be split and a preset logical rule group, and use the logical rule group to split the attributes representing the same entity in the entity tuple into a split tuple to obtain at least one split tuple; Use the logical rule group to perform attribute value derivation on all attributes in each split tuple to obtain all attribute value results for each attribute in each split tuple; Determine the final attribute values of all attributes in each split tuple according to all attribute value results for each attribute in each split tuple.
17. The readable storage media according to claim 16, wherein, the using the logical rule group to split the attributes representing the same entity in the entity tuple into a split tuple to obtain at least one split tuple includes: Use the logical rule group to determine whether the attribute values of the first attribute and the second attribute in the entity tuple represent the same entity, wherein the first attribute and the second attribute are not the same attribute and are both arbitrary attributes in the entity tuple; If the attribute values of the first attribute and the second attribute represent the same entity, then determine that the attribute values of the first attribute and the second attribute both belong to a split tuple; If the attribute values of the first attribute and the second attribute represent different entities, then determine that the attribute value of the first attribute belongs to a split tuple and the attribute value of the second attribute belongs to another split tuple.
18. The readable storage media according to claim 16, wherein, the determining the final attribute values of all attributes in each split tuple according to all attribute value results for each attribute in each split tuple includes: For any split tuple, use the attributes with conflicting attribute value results in the split tuple as the attributes to be resolved; Obtain the verified attribute set of the split tuple, wherein the verified attribute set includes the attributes that have passed verification; According to the attribute values of all attributes in the verified attribute set, screen all attribute value results of the attributes to be resolved to obtain the screened attribute value results as the final attribute values of the attributes to be resolved.
19. The readable storage media according to claim 18, wherein, the obtaining the verified attribute set of the split tuple includes: Use the attributes without conflicting attribute value results in the split tuple and the attributes to be resolved for which the final attribute values have been determined in the split tuple as the attributes that have passed verification; Obtain a set of verified attributes according to the verified attributes.
20. The readable storage medium according to claim 18, wherein filtering all the attribute value results of the to-be-solved attributes according to the attribute values of all the attributes in the set of verified attributes, and obtaining the filtered attribute value results as the final attribute values of the to-be-solved attributes, including: using a pre-trained association analysis model to calculate the association strength between each attribute value result of the to-be-solved attributes and the attribute values of all the attributes in the set of verified attributes, and obtaining the association strength score corresponding to each attribute value result of the to-be-solved attributes; screening from the association strength scores of all the attribute value results of the to-be-solved attributes to obtain the attribute value result with the highest association strength score as the target result, and determining the target result as the final attribute value of the to-be-solved attributes.
Citation Information
Patent Citations
Data cleaning system based on internet information
CN104268216A
Rule script generation method and device, computer equipment and medium
CN115545006A
Entity edge building method, device and equipment of knowledge graph and medium
CN116629359A
Processing method and device for triple of knowledge graph and computer equipment
CN117033648A
Method and apparatus for optimizing data while preserving provenance information for the data
US20080126399A1