Computer-implemented method, computer program, and computer system

The method enhances MDM systems by processing unstructured data to improve record matching accuracy, leveraging both structured and unstructured data for a more unified customer data view.

JP7812597B2Active Publication Date: 2026-02-10INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022052207
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-29
Filing Date
2022-03-28
Publication Date
2026-02-10
Estimated Expiration
2042-03-28

AI Technical Summary

Technical Problem

Existing Master Data Management (MDM) systems face challenges in effectively matching and linking customer data from different sources to create a unified view, as they primarily rely on structured data and lack efficient methods for incorporating and leveraging unstructured data for improved accuracy.

Method used

A computer-implemented method that processes unstructured data objects to identify attribute values, compares these values across records to determine similarity, and assigns weighted contributions based on structured and unstructured data matches to enhance record matching accuracy.

Benefits of technology

Improves the accuracy of record matching by leveraging unstructured data, allowing for better identification of duplicate records and creating a more comprehensive and unified customer data view.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007812597000001
    Figure 0007812597000001
  • Figure 0007812597000002
    Figure 0007812597000002
  • Figure 0007812597000003
    Figure 0007812597000003
Patent Text Reader

Abstract

To meet a continuous need to improve data matching to data in master data management systems.SOLUTION: A computer-implemented method comprises processing unstructured objects of each of multiple records of a database for identifying a set of one or more values of attributes in the unstructured objects of the record. The sets of unstructured attribute values of two records of the database may be compared for determining a similarity level between the two sets. It may be determined whether the two records represent the same entity based on the comparison result.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates generally to the field of digital computer systems, and more particularly to methods for record matching in database systems. [Background technology]

[0002] Enterprise data matching deals with matching and linking customer data received from different sources to create a single version of the truth. Solutions based on Master Data Management (MDM) address enterprise data and perform data indexing, matching, and linking. Summary of the Invention [Problem to be solved by the invention]

[0003] Master data management systems may provide access to this data, however there is a continuing need for improved data matching to the data within the master data management systems. [Means for solving the problem]

[0004] Various embodiments provide methods, computer systems, and computer program products as described by the subject matter of the independent claims. Advantageous embodiments are described in the dependent claims. Where embodiments of the invention are not mutually exclusive, they may be freely combined with one another. In one aspect, the invention relates to a computer-implemented method for record matching in a database system, wherein a record represents an entity and the record is associated with one or more unstructured data objects. The method comprises the steps of: processing the unstructured object of each record of a plurality of records of a database (e.g., of the database system) to identify a set of one or more values ​​of attributes, hereinafter referred to as unstructured attribute values, in the unstructured object of each record; comparing the sets of unstructured attribute values ​​of two records of the database to determine a level of similarity between the two sets; and determining whether the two records represent the same entity based on a result of the comparison.

[0005] In another aspect, the invention relates to a computer program product comprising a computer-readable storage medium having computer-readable program code embodied thereon, the computer-readable program code being configured to implement all of the steps of the method of the preceding embodiment. In another aspect, the invention relates to a computer system for record matching, wherein a record represents an entity and the record is associated with one or more unstructured data objects. The computer system is configured to process the unstructured object of each record of a plurality of records of a database to identify a set of one or more values ​​of attributes, hereinafter referred to as unstructured attribute values, in the unstructured object of each record, compare the sets of unstructured attribute values ​​of two records of the database to determine a level of similarity between the two sets, and determine whether the two records represent the same entity based on a result of the comparison. [Brief explanation of the drawings]

[0006] In the following, embodiments of the invention will be described in more detail, by way of example only, and with reference to the following drawings, in which:

[0007] [Figure 1] 1 is a block diagram of a database device according to an example of the present subject matter.

[0008] [Figure 2] 1 is a flowchart of a method for record matching in a database system according to an example of the present subject matter.

[0009] [Figure 3] 1 is a flowchart of a method for record matching in a database system according to an example of the present subject matter.

[0010] [Figure 4] 1 is a flowchart of a method for record matching in a database system according to an example of the present subject matter.

[0011] [Figure 5A] 1 is a flowchart of a method for comparing two records according to an example of the present subject matter.

[0012] [Figure 5B] FIG. 1 illustrates the resulting set of records and unstructured attribute values ​​associated with an unstructured object.

[0013] [Figure 5C] FIG. 1 illustrates the resulting set of records and unstructured attribute values ​​associated with an unstructured object.

[0014] [Figure 5D] FIG. 10 illustrates the results of a comparison between two sets of unstructured attribute values.

[0015] [Figure 6]FIG. 1 is a diagram representing a computerized system adapted to implement steps of one or more methods as encompassed by the present subject matter. DETAILED DESCRIPTION OF THE INVENTION

[0016] The description of various embodiments of the present invention is presented for illustrative purposes and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been selected to best explain the principles of the embodiments, practical applications of or technical improvements to the technology found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0017] Service provider computer systems typically include storage for storing information related to customers and services. Some of this information may be provided by customers when they fill out registration forms, service request forms, or other forms, such as contracts, banking transactions, certificates, etc. As a result, unstructured objects may be stored in these systems. Unstructured objects may be objects that contain attribute values ​​in an unstructured format. These attributes are named unstructured attributes to distinguish them from attributes of records, which are named structured attributes. Unstructured objects may allow attributes to be associated with corresponding attribute values. Unstructured objects may be files, documents, or objects that contain free-form text or embedded values. Examples of unstructured objects may include word processing documents (e.g., Microsoft Word documents in their native format), Adobe Acrobat documents, emails, image files, video files, audio files, and other files in the native format associated with the software application that created them. Furthermore, such computer systems may store customer and service data in a structured format in the form of database records. As a result, the same entity, such as a person, may be associated with information in different formats in the system. For example, a record representing a particular person may be associated with unstructured information about that particular person. A data record or record is a set of related data items, such as a particular user's name, date of birth, and class. A record represents an entity, which may be a user, object, or concept about which information is stored in the record. The terms "data record" and "record" are used interchangeably.

[0018] Such a computer system may thus comprise records having attributes, termed structured attributes, and unstructured objects having values ​​for unstructured attributes. The values ​​of unstructured attributes may be the complete value of the unstructured attribute or a portion of this complete value. For example, the value "Street" may be the value of the unstructured attribute "Address," but the value "Street" may only be a portion of the address; other values, such as the city name, may also comprise the complete or entire value of the address.

[0019] Most of the operations performed in these computer systems can involve matching records. Matching records involves comparing structured attribute values ​​of records. Matched records (records that can be merged) are records that represent the same entity. The level of match between two records indicates the similarity of the attribute values ​​of the two records.

[0020] The present subject matter can be advantageous because it can improve the record matching process by leveraging existing unstructured objects, such as documents. The present subject matter can gain insights for decision making from unstructured documents. This can be particularly advantageous because many industrial systems can store large amounts of unstructured information attached to master data. For example, in insurance, there are many insurance policies attached to customer records. In manufacturing and utility industries, there are repair and maintenance manuals attached to product records.

[0021] According to one embodiment, the method further comprises evaluating one or more occurrence characteristics for each identified unstructured attribute value, wherein the occurrence characteristics for the identified unstructured attribute value in the unstructured object of a particular record include either a frequency of occurrence of the particular unstructured attribute value in the unstructured object of the particular record (termed a first occurrence characteristic) or an indication of other identified unstructured attribute values ​​for the particular record that co-occur with the particular unstructured attribute value in the unstructured object (termed a second occurrence characteristic), and wherein comparing the two sets of unstructured attribute values ​​comprises comparing the evaluated occurrence characteristics of the unstructured attribute values ​​of one of the two sets with the evaluated occurrence characteristics of the unstructured attribute values ​​of the other of the two sets.

[0022] For example, if the same value exists in two sets and has the same frequency of occurrence in the two sets, this may provide a strong indication of the similarity of the records representing the two sets. Similarly, if the same value exists in two sets and the same value belongs to the same co-occurrence in the two sets, this may provide a strong indication of the similarity of the records representing the two sets. Thus, the occurrence property may further improve the accuracy of the comparison of the two sets, which in turn may improve the accuracy of the record matching process.

[0023] According to one embodiment, records have values ​​for attributes, hereinafter referred to as structured attributes, and determining whether the two records represent the same entity comprises assigning an initial contribution weight to each structured attribute of the plurality of structured attributes; selecting unstructured attribute values ​​that are similar and present in both two sets based on the results of the comparison; if a structured attribute value of the plurality of structured attribute values ​​does not match any of the selected unstructured attributes, replacing the contribution weight of the structured attribute with a weight indicative of the similarity between the two sets; if a structured attribute value of the plurality of structured attribute values ​​fully or partially matches a selected unstructured attribute, increasing the contribution weight of the structured attribute; and comparing the two records using the contribution weights.

[0024] This embodiment may allow for weighted matching rules that assign a weight (e.g., an integer weight) to each structured attribute of the records being compared. For each structured attribute, the associated contribution weight may be multiplied by the similarity score of the two records, and the scores may be summed. If the sum is greater than or equal to a threshold, the two records being compared are considered a match. This embodiment may be particularly advantageous when comparing a large number of attributes without having a single attribute that differs in the two records resulting in a mismatch.

[0025] According to one embodiment, the method further comprises running an aggregation algorithm to aggregate values ​​of the selected unstructured attribute values ​​that constitute the complete value of the respective unstructured attribute, resulting in zero or more aggregate values, and wherein the comparison with the structured attribute values ​​is performed using the processed selected unstructured attribute values.

[0026] For example, the selected unstructured attribute values ​​may include values ​​v1, v2, ..., and vr. Each of the values ​​is a value for a respective unstructured attribute. However, some values ​​may not be the complete value for the respective unstructured attribute. For example, v1 may be a first name value and v2 may be a last name value, both of which are values ​​for the unstructured attribute "Full Name," but v1 and v2 are not the complete values. The aggregate of v1 and v2 may be the complete value for the unstructured attribute "Full Name." Each selected unstructured attribute value may be processed to determine whether it is a complete value for the respective unstructured attribute, and if it is not a complete value, other non-complete values ​​to aggregate with it may be determined. Determining which values ​​can be aggregated together may be performed using a second occurrence characteristic evaluated for the selected unstructured attribute values; for example, values ​​may be aggregated if they co-occur within the same sentence or paragraph and are non-complete values ​​for the same unstructured attribute. Alternatively or additionally, according to one embodiment, the aggregating step includes grouping the unstructured attribute values ​​of each of the sets into a plurality of groups based on categories of the unstructured attributes, and the aggregation is performed for values ​​that belong to the same group, and also based on a second occurrence characteristic of the group.

[0027] According to one embodiment, comparing the two records includes comparing the values ​​of the structured attributes of the two records, resulting in individual match scores for each structured attribute of the records, and combining the individual match scores using the contribution weights. A match score may be, for example, a value in the range 0 to 100, which represents the degree to which two values ​​are similar. A value of 100 indicates that the two values ​​are identical, and a value of 0 indicates no similarity.

[0028] According to one embodiment, if the two records represent the same entity, then the two records are merged into a single record; otherwise, the two records remain separate.

[0029] FIG. 1 illustrates an exemplary computer system 100. The computer system 100 may be configured to perform, for example, master data management or data warehousing, or both; for example, the computer system 100 may enable a deduplication system. The computer system 100 includes a data integration system 101 and one or more client systems or data sources 105. The client systems 105 may include computer systems (e.g., computer systems such as those described with reference to FIG. 6). The client systems 105 may communicate with the data integration system 101 via a network connection, including, for example, a wireless local area network (WLAN) connection, a wide area network (WAN) connection, a local area network (LAN) connection, the Internet, or a combination thereof. The data integration system 101 may control access (e.g., read and write access) to a database or repository 103, which includes structured records 107 and is therefore referred to herein as a structured repository. The data integration system 101 may control access (eg, read and write access) to another repository 110, which is referred to herein as an unstructured repository because it contains unstructured objects 111.

[0030] As shown in FIG. 1 , each structured record 107 stored in the structured repository 103 may have values ​​for a set of attributes a_1...a_N (N≧1), such as a name attribute. While this example is described with respect to a small number of attributes, more or fewer attributes may be used. Each record 107 may represent an entity, such as a person. The data records 107 stored in the central repository 103 may be received from client systems 105 and processed (e.g., converted to a unified structure) by the data integration system 101 before being stored in the central repository 103. For example, the records received from the client systems 105 may have a structure that differs from the structure of the records stored in the central repository 103. In another example, the data integration system 101 may import data records from the central repository 103 from client systems 105 using one or more Extract-Transform-Load (ETL) batch processes, or via HyperText Transport Protocol ("HTTP") communications or other types of data exchange.

[0031] Unstructured object 111 may include, for example, a scanned document or form. Unstructured object 111 may be received, for example, from client system 105. As indicated by the dashed lines, each entity or record in structured repository 103 may be associated with one or more unstructured objects in unstructured repository 110. For example, each record R_i of at least some of records 107 in structured repository 103 may be associated with m_i unstructured objects “OB”_1, “OB”_2, ... “OB”_(m_i) in unstructured repository 110, where m_i≧1. For example, an employee who has a record in structured repository 103 describing their name, social media, etc., may also have their employment contract scanned and stored in unstructured repository 110; for example, unstructured object 111 may be provided by a CMIS-enabled system. For example, in an MDM system such as IBM MDM, an OOTB connector to a content management system such as Filenet may enable access to unstructured objects. For example, a master data record in the MDM system may be associated with a unique resource locator that allows for discovery of the associated unstructured document in the content management system based on a standard such as CMIS.

[0032] The data integration system 101 may be configured to process the records 107 and unstructured objects 111 using one or more algorithms, such as algorithm 120, that implement at least a portion of the present methods. For example, the data integration system 101 may process the data records 107 and unstructured objects 111 using algorithm 120 to identify duplicate records in the structured repository 103. Although shown as separate components, the repository 103 or repository 110, or both, may be part of the data integration system 101 in other examples.

[0033] In one example, algorithm 120 may include a matching engine 121 that matches records. Algorithm 120 further includes a token bag comparator 122 that compares sets or bags of unstructured attribute values ​​associated with the records to be compared. Algorithm 120 further includes a token bag manager 123 that may manage the bags determined by token extractor 124.

[0034] Figure 2 is a flowchart of a method of record matching in a system according to one example of the present subject matter. For illustrative purposes, the method illustrated in Figure 2 may be implemented in the system shown in Figure 1, but is not limited to this implementation. The method of Figure 2 may be performed by, for example, data integration system 101.

[0035] In step 201, the unstructured objects for each record of at least a portion of the records of database 103 may be processed to identify a set of one or more values ​​for unstructured attributes in the unstructured object for each record. Each record of the at least a portion of the records may be associated with one or more unstructured objects.

[0036] In one example, at least a portion of the records may include all of the records in database 103. For each record in the database, it may be determined whether the record is associated with one or more unstructured objects. The associated unstructured objects may then be processed to identify values ​​for unstructured attributes. If the record is not associated with any unstructured objects, the next unprocessed record may be processed, and so on, until all records have been processed. Processing all of the records in the database may be advantageous because it may prepare all information in advance that may be ready for immediate use at a later stage.

[0037] In one example, at least some of the records may include a subset of the records of database 103. For each record in the subset of records, it may be determined whether the record is associated with one or more unstructured objects. The associated unstructured objects may then be processed to identify values ​​for the unstructured attributes. The subset of records may, for example, include only those records that a user needs to process. This example may be advantageous because it allows for on-demand processing of records, which may save resources that would otherwise be required by processing records whose results would not be used.

[0038] Identifying values ​​of unstructured attributes in an unstructured object may be performed, for example, by parsing the object and performing data mining analysis to identify values ​​of the attributes.

[0039] Thus, processing the unstructured objects of each record R_i in step 201 may result in a bag or collection (termed "Bag"_i) of values ​​for unstructured attributes b_1, b_2...b_(M_i), where M_i≧1. The unstructured attributes b_1, b_2...b_(M_i) of each collection "Bag"_i may or may not include attributes from the structured attributes a_1...a_N. For example, a record representing a student may include structured attributes such as "Student ID," "Class," "Age," "Name," etc., while the documents 111 associated with that student may include values ​​for different attributes such as "Address," or values ​​for the same attribute such as "Name," or both. The unstructured attributes of one set 'Bag'_i may or may not contain the unstructured attributes of another set, i.e., they may or may not share unstructured attributes with another set 'Bag'_j; for example, two student records may be associated with completely different documents, one associated with an insurance policy document while the other student record is associated with a resume, resulting in different identified unstructured attributes for the two students.

[0040] Each unstructured attribute may have at least one value in a respective set 'Bag'_i. At least one value may contain duplicate values. For example, the record of employee "X" may be associated with a document that has been processed to identify values ​​for unstructured attributes, resulting in a bag ('Bag'_x) of identified values ​​for the unstructured attributes "Car Model," "Phone Number," and "Address." Bag 'Bag'_x may contain multiple values ​​for the attribute "Car Model," e.g., person "X" has several cars listed in the document. Bag 'Bag'_x may contain five duplicate values ​​for the attribute "Phone Number," because the same number appears in several documents for employee "X." In other words, the set associated with person "X"'s record has three unstructured attributes, but may contain multiple values ​​for each unstructured attribute.

[0041] The sets of unstructured attribute values ​​obtained in step 201 may be used to determine whether records 107 are duplicates. For example, in step 203, two sets 'Bag'_i and 'Bag'_j of two records R_i and R_j, respectively, may be compared to determine the level of similarity between the two sets 'Bag'_i and 'Bag'_j. That is, the values ​​of unstructured attributes b_1, b_2...b_(M_i) may be compared with the values ​​of unstructured attributes b_1, b_2...b_(M_j). In one example, the comparison may be a pairwise comparison between all possible pairs of values ​​of the two sets, or may be a pairwise comparison between pairs of values ​​of the same unstructured attribute. The comparison may result in individual similarity scores, which are combined to obtain a similarity score between the compared sets. In another example, the Jaccard Similarity algorithm may be used to compare sets 'Bag'_i and 'Bag'_j. Figure 3 provides an example implementation of the comparison step 203. The comparison as described herein is performed between two records, but is not limited to such; more than two records may be compared by comparing their respective sets as described using the example of two records.

[0042] Therefore, the result of the comparison in step 203 may be used to determine whether the two records represent the same entity in step 205. The similarity level between two sets 'Bag'_i and 'Bag'_j may indicate the similarity between two records R_i and R_j, respectively, e.g., if two bags are very similar, this indicates that the two records represent the same entity.

[0043] Figure 3 is a flowchart of a method for comparing records according to one example of the present subject matter. For illustrative purposes, the method described in Figure 3 may be implemented in the system shown in Figure 1, but is not limited to this implementation. The method of Figure 3 may be performed, for example, by data integration system 101. The method of Figure 3 provides an example implementation of comparison step 203 of Figure 2. For example, records 107 to be compared may be associated with respective bags or sets of unstructured attribute values, as described with reference to Figure 2.

[0044] In step 301, one or more occurrence properties may be evaluated for each identified unstructured attribute value. For example, each set of unstructured attribute values, 'Bag'_i, may be processed to evaluate an occurrence property for each unstructured attribute value in that set. That is, the occurrence property for each value of attribute b_1 may be evaluated, the occurrence property for each value of attribute b_2 may be evaluated, and so on.

[0045] In a first example, the occurrence property may be the frequency of occurrence of an unstructured attribute value in the unstructured object of a particular record. In this case, the frequency of occurrence of the value of attribute b_1 in bag 'Bag'_i of record R_i may be determined. Similarly, the frequency of occurrence of the value of attribute b_2 in bag 'Bag'_i of record R_i may be determined, and so on. Following the example of employee X's record, the value of the unstructured attribute "phone number" has a frequency of occurrence of 5 because this value appeared 5 times in the documents associated with employee X.

[0046] In a second example, the occurrence property of each value in set 'Bag'_i of record R_i may be an indication of other values ​​in the same set 'Bag'_i that co-occur with that value in the unstructured object. For example, for each record R_i, the values ​​of the unstructured attributes in each set 'Bag'_i may be processed to identify how often and which attribute values ​​are mentioned together in the same sentence or paragraph.

[0047] Therefore, a comparison of two records R_i and R_j may be performed by comparing the evaluated occurrence characteristics of the unstructured attribute values ​​of one set 'Bag'_i with the evaluated occurrence characteristics of the unstructured attribute values ​​of the other set 'Bag'_j in step 303. For example, if the same values ​​are present in the two sets 'Bag'_i and 'Bag'_j, and those same values ​​have the same frequency of occurrence in the two sets, this may provide a strong indication of the similarity of records R_i and R_j.

[0048] 4 is a flowchart of a method for comparing records according to an example of the present subject matter. For illustrative purposes, the method illustrated in FIG. 4 may be implemented in the system shown in FIG. 1, but is not limited to this implementation. The method of FIG. 4 may be performed, for example, by data integration system 101. For example, the method of FIG. 4 may compare two records R_i and R_j.

[0049] In step 401, an initial contribution weight may be assigned to each structured attribute a_1...a_N. For example, to compare two employee records, the attribute "employee ID" may be assigned a higher weight than the attribute "name" because two employees may have the same name but may be unlikely to have the same employee ID. Therefore, the employee ID may advantageously have a higher contribution in determining a match. For example, an integer weight may be assigned to each structured attribute a_1...a_N of the records R_i and R_j being compared.

[0050] In step 403, unstructured attribute values ​​that are similar and that are present in two sets 'Bag'_i and 'Bag'_j of two records R_i and R_j may be selected. This may be done, for example, by taking the intersection of the two sets 'Bag'_i and 'Bag'_j. Figures 5C-5D provide an exemplary implementation of step 403. Step 403 may result in a set referenced by 'Bag'_i ∩ 'Bag'_j that contains the selected unstructured attribute values.

[0051] In step 405, the initial contribution weights may be adapted or adjusted. This adaptation may be performed by comparing the unstructured attribute values ​​in the intersection subset 'Bag'_i∩'Bag'_j with the values ​​of the structured attributes a_1...a_N in the two records R_i and R_j. For example, if a structured attribute value of the plurality of structured attribute values ​​does not match any of the selected unstructured attributes, the contribution weight of that structured attribute may be replaced with a weight indicating the similarity between the two sets. If a structured attribute value of the plurality of structured attribute values ​​fully or partially matches a selected unstructured attribute, the contribution weight of that structured attribute may be increased by a predefined value.

[0052] In step 407, two records R_i and R_j may be compared using the contribution weights. Pairs of values ​​for each structural attribute a_1...a_N may be compared, resulting in N similarity scores. The N similarity scores may be combined using a weighted sum, for example, by multiplying each similarity score by an adapted contribution weight and summing the scores. The resulting scores may be compared to a threshold to determine whether the two records R_i and R_j are duplicate or non-duplicate records.

[0053] FIG. 5A is a flowchart of a method for comparing records according to an example of the present subject matter. For illustrative purposes, the method described in FIG. 5A may be implemented in the system shown in FIG. 1, but is not limited to this implementation. The method of FIG. 5A may be performed, for example, by data integration system 101. FIGS. 5B and 5C illustrate two records (e.g., MDM records) R_1 and R_2 to be compared. As also shown in FIGS. 5B and 5C, record R_1 is associated with a set of unstructured objects 'OB'_1, 'OB'_2...'OB'_(m_1), and record R_2 is associated with a set of unstructured objects 'OB'_1, 'OB'_2...'OB'_(m_2). For example, record R_1 may be associated with 14 documents, while R_2 may be associated with 17 documents. Two records R_1 and R_2 represent people, such as employees. Two records R_1 and R_2 have values ​​for structured attributes such as "Name," "Address," "Date of Birth ("DOB")," "Gender," "Marital Status," and "SSN." Each of the structured attributes may be assigned a contribution weight as follows: Name weight: Medium, Address weight: Medium, DOB weight: High, Gender weight: Very Low, Marital Status weight: High, SSN weight: Very High. The values ​​High, Medium, and Very High may be represented by respective integers that can be used to perform a weighted sum.

[0054] In step 501, documents associated with each of two records may be processed to identify values ​​for unstructured attributes. This may result in one set of unstructured values ​​(also called an entity token bag) "Bag"_1 for record R_1, as shown in FIG. 5B, and one set of unstructured values ​​"Bag"_2 for record R_2, as shown in FIG. 5C. Step 501 may enable parsing the associated unstructured content for the MDM records using an entity detection module that detects people's names, addresses, personal confidential information, and other entities of interest. As shown in FIGS. 5B and 5C, each of the two sets "Bag"_1 and "Bag"_2 includes values ​​such as "John" for the unstructured attribute "Name" and "United States" for the unstructured attribute "Country."

[0055] Each of the values ​​in the two sets 'Bag'_1 and 'Bag'_2 may be associated with an occurrence property. This is shown in Figures 5B and 5C, where each value is associated with its frequency of occurrence. For example, the value "Street" appeared four times in the 14 documents associated with record R_1, while this value appeared three times in the 17 documents associated with record R_2. For example, all extracted values ​​may be stored in a so-called entity token bag along with their frequency and entity association score (indicating how often and which entities are mentioned together in the same sentence, paragraph, etc.). The entity association score may be the second occurrence property defined herein.

[0056] In step 503, the two sets 'Bag'_1 and 'Bag'_2 may be compared to calculate a similarity score for the entire bag. For example, similar bags may have many identical values ​​and very similar entity association scores. Therefore, an intersection set or bag may be determined in step 503. The intersection bag may be determined by including in the intersection bag only values ​​that are present in all bags 'Bag'_1 and 'Bag'_2. The resulting intersection bag 'Bag'_1 ∩ 'Bag'_2 is shown in FIG. 5D. As shown in FIG. 5D, values ​​in regular font occur with the same frequency in all bags 'Bag'_1 and 'Bag'_2. Values ​​in italic font occur with different frequencies in all bags 'Bag'_1 and 'Bag'_2. Values ​​in bold font occur in both bags 'Bag'_1 and 'Bag'_2, but not in all structured record attributes. For example, the intersection value 'Baker' is not part of record R_1, and therefore it is in bold font.

[0057] In step 505, the intersection bag may be used to adjust the weights assigned to the structural attributes of records R_1 and R_2.

[0058] For example, if the intersection bags have the same value of the unstructured attribute IATT with the same frequency, but this value does not match the value of the structured attribute ATT corresponding to the unstructured attribute IATT, the weight of the structured attribute ATT may be replaced with the token bag weight. These types of values ​​(e.g., values ​​of IATT) are written in bold font, which indicates that the match based on the structured attribute ATT may be incorrect.

[0059] If the intersection bag has the same value of IATT with the same frequency for something that partially or completely matches a structured attribute, the weight of that structured attribute ATT may be increased. These types of values ​​(e.g., IATT values) are written in normal font, indicating that matching based on that attribute may be strengthened.

[0060] If the intersection token bags have the same value for something that only partially matches a structured attribute (e.g., address or name), but with a different frequency, the weight of that structured attribute may be increased. These types of values ​​(e.g., IATT values) are written in italic font, indicating that matching based on that attribute may allow for a course correct partial match.

[0061] In step 507, it may be determined whether the two records represent the same entity based on a comparison of the two records R_1 and R_2 using the adjusted weights.

[0062] FIG. 6 depicts a general computerized system 600 adapted to implement at least some of the steps of the methods included in the present disclosure.

[0063] It will be understood that the methods described herein are at least partially non-interactive and automated via a computerized system, such as a server or embedded system. However, in exemplary embodiments, the methods described herein can be implemented in a (partially) interactive system. These methods can also be implemented in software 612, 622 (including firmware 622), hardware (processor) 605, or a combination thereof. In exemplary embodiments, the methods described herein are implemented in software as executable programs and executed by a special-purpose or general-purpose digital computer, such as a personal computer, workstation, minicomputer, or mainframe computer. Thus, the most common system 600 includes a general-purpose computer 601.

[0064] In an exemplary embodiment, with respect to the hardware architecture, as shown in FIG. 5 , a computer 601 includes a processor 605, a memory (main memory) 610 coupled to a memory controller 615, and one or more input or output (I / O) devices (or peripherals) 10, 645 communicatively coupled via a local input / output controller 635. The input / output controller 635 can be, but is not limited to, one or more buses or other wired or wireless connections, as known in the art. The input / output controller 635 may include additional elements, such as controllers, buffers (caches), drivers, repeaters, and receivers, to enable communication, but these are omitted for brevity. Furthermore, the local interface may include address, control, or data connections, or a combination thereof, to enable appropriate communication between the aforementioned components. As described herein, the I / O devices 10, 645 may generally include any generalized cryptographic or smart card known in the art.

[0065] Processor 605 is a hardware device that executes software stored, among other things, in memory 610. Processor 605 can be any custom or commercially available processor, a central processing unit (CPU), a coprocessor among several processors associated with computer 601, a semiconductor-based microprocessor (in the form of a microchip or chipset), or generally any device that executes software instructions.

[0066] The memory 610 can include any one or combination of volatile memory elements (e.g., random access memory (RAM, e.g., DRAM, SRAM, SDRAM, etc.)) and non-volatile memory elements (e.g., ROM, erasable programmable read-only memory (EPROM), electronically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM)). It should be noted that the memory 610 can have a distributed architecture, where various components are located remotely from each other but can be accessed by the processor 605.

[0067] The software in memory 610 may include one or more separate programs, each of which includes an ordered list of executable instructions that implement logical functions, particularly functions involved in embodiments of the present invention. In the example of Figure 5, the software in memory 610 includes instructions 612, for example, instructions for managing a database, such as a database management system.

[0068] The software in memory 610 will typically also include a suitable operating system (OS) 611. The OS 611 essentially controls the execution of other computer programs, such as software 612, which may implement methods as described herein.

[0069] The methods described herein may be in the form of a source program 612, an executable program 612 (object code), a script, or any other entity comprising a set of instructions 612 to be executed. In the case of a source program, the program must be translated via a compiler, assembler, interpreter, etc., which may or may not be included in memory 610, in order to operate properly in conjunction with the OS 611. Furthermore, the methods may be written as an object-oriented programming language with classes of data and methods, or a procedural programming language with routines, subroutines, or functions, or a combination thereof.

[0070] In an exemplary embodiment, a conventional keyboard 650 and mouse 655 may be coupled to the input / output controller 635. Other output devices, such as I / O devices 645, may include input devices, such as, but not limited to, printers, scanners, microphones, etc. Finally, I / O devices 10, 645 may further include devices that communicate both input and output, such as, but not limited to, network interface cards (NICs) or modulators / demodulators (for accessing other files, devices, systems, or networks), radio frequency (RF) or other transceivers, telephone interfaces, bridges, routers, etc. I / O devices 10, 645 may be any generalized cryptographic card or smart card known in the art. System 600 may further include a display controller 625 coupled to a display 630. In an exemplary embodiment, system 600 may further include a network interface that couples to a network 665. Network 665 may be an IP-based network for communication between computer 601 and any external servers, clients, etc. via a broadband connection. Network 665 transmits and receives data between computer 601 and external systems 30, which may be involved in performing some or all of the method steps discussed herein. In an exemplary embodiment, network 665 may be a managed IP network managed by a service provider. Network 665 may be implemented wirelessly using wireless protocols and technologies such as WiFi, WiMax, etc. Network 665 may also be a packet-switched network such as a local area network, wide area network, metropolitan area network, the Internet network, or other similar types of network environments.The network 665 may be a fixed wireless network, a wireless local area network (WLAN), a wireless wide area network (WWAN), a personal area network (PAN), a virtual private network (VPN), an intranet, or other suitable network system, and includes equipment for receiving and transmitting signals.

[0071] If computer 601 is a PC, workstation, intelligent device, etc., the software in memory 610 may further include a basic input / output system (BIOS) 622. The BIOS is a set of basic software routines that initializes and tests hardware at startup, starts OS 611, and supports the transfer of data between hardware devices. The BIOS is stored in ROM so that the BIOS can be executed when computer 601 is booted.

[0072] When computer 601 is in operation, processor 605 is configured to execute software 612 stored in memory 610 to communicate data to memory 610 and to generally control the operation of computer 601 in accordance with the software. The methods and OS 611 described herein are read by processor 605, possibly buffered within processor 605, and then executed, in whole or in part, but typically the latter.

[0073] 6, the methods can be stored on any computer-readable medium, such as storage 620, for use by or in connection with any computer-related system or method. Storage 620 can include disk storage, such as HDD storage.

[0074] The present invention may be a system, method, or computer program product, or combination thereof, at any possible level of technical detail of integration. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions that cause a processor to perform aspects of the present invention.

[0075] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or ridge structures in grooves that record instructions, and any suitable combination of the foregoing. Computer-readable storage medium, as used herein, is not to be construed as a transitory signal per se, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electrical signal transmitted through a wire.

[0076] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions to a computer-readable storage medium in the respective computing / processing device for storage.

[0077] The computer-readable program instructions for carrying out the operations of the present invention may be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for an integrated circuit, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk®, C++, etc., and procedural programming languages ​​such as the “C” programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer readable program instructions to personalize the electronic circuitry by utilizing state information of the computer readable program instructions to perform aspects of the present invention.

[0078] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0079] These computer-readable program instructions may be provided to a computer processor or other programmable data processing apparatus to produce a machine, whereby the instructions, executed by the computer processor or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowcharts or block diagrams, or both. These computer-readable program instructions may also be stored on a computer-readable storage medium, whereby the instructions can instruct a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, such that the computer-readable storage medium having the instructions stored thereon comprises an article of manufacture including instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts or block diagrams, or combination thereof.

[0080] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be executed on the computer, other programmable apparatus, or other device to generate a computer-implemented process, whereby the instructions executing on the computer, other programmable apparatus, or other device implement the functions / operations specified in one or more blocks of the flowcharts or block diagrams, or a combination thereof.

[0081] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions that implement the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may actually be implemented as a single step, or may be executed concurrently, substantially concurrently, partially, or fully overlapping in time, or the blocks may even be executed in the reverse order depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or operations or executes a combination of dedicated hardware and computer instructions.

Claims

1. 1. A computer-implemented method for record matching in a database system, wherein a record represents an entity, the record being associated with one or more unstructured data objects, the computer-implemented method comprising: processing the unstructured data object for each record of a plurality of records of a database to identify a set of one or more values ​​of attributes, hereafter referred to as unstructured attribute values, in the unstructured data object for each record; comparing the sets of unstructured attribute values ​​of two records in the database to determine a level of similarity between the two compared sets; determining whether the two records represent the same entity based on a result of the comparison; A computer-implemented method comprising:

2. The method further comprises evaluating, for each identified unstructured attribute value, one or more occurrence characteristics, wherein the occurrence characteristics for a particular unstructured attribute value identified in the unstructured data object of a particular record include:

2. The computer-implemented method of claim 1, wherein the comparing step includes comparing the two sets of unstructured attribute values ​​to one of a frequency of occurrence of the particular unstructured attribute value in the unstructured data object for the particular record and an indication of other identified unstructured attribute values ​​for the particular record that co-occur with the particular unstructured attribute value in the unstructured data object, and wherein comparing the two sets of unstructured attribute values ​​comprises comparing the evaluated occurrence characteristics of the unstructured attribute values ​​of one of the two sets to the evaluated occurrence characteristics of the unstructured attribute values ​​of the other of the two sets.

3. 3. The computer-implemented method of claim 1, further comprising grouping the unstructured attribute values ​​of each of the plurality of sets into a plurality of groups based on their category, and wherein the comparison between two of the sets is performed by comparing groups of the same category.

4. The records have values ​​for attributes hereafter referred to as structural attributes, and the step of determining whether the two records represent the same entity comprises: assigning an initial contribution weight to each of said plurality of structured attributes; selecting unstructured attribute values ​​that are present in both of said sets based on the results of said comparison; responsive to determining that a structured attribute value of a plurality of structured attribute values ​​does not match any of the selected unstructured attributes, replacing the initial contribution weight of the structured attribute with a weight indicative of similarity between two of the sets; increasing the initial contribution weight of a structured attribute in response to determining that a structured attribute value of the plurality of structured attribute values ​​fully or partially matches a selected unstructured attribute; comparing the two records using the initial contribution weights; The computer-implemented method of claim 1 , comprising:

5. The computer-implemented method of claim 4 , wherein the selecting step comprises intersecting two of the sets to yield an intersection subset.

6. 6. The computer-implemented method of claim 4, wherein the unstructured attribute values ​​are the complete value of the respective unstructured attribute or a portion of the complete value of the unstructured attribute, and wherein the selecting step further comprises running an aggregation algorithm to aggregate values ​​of the selected unstructured attribute values ​​that make up the complete value of the respective unstructured attribute, resulting in zero or more aggregate values, and wherein the comparison with the plurality of structured attribute values ​​is performed using the processed selected unstructured attribute values.

7. 7. The computer-implemented method of claim 6, wherein the aggregating comprises grouping the unstructured attribute values ​​of each of the plurality of sets into a plurality of groups based on their categories, and the aggregation is performed on values ​​that belong to the same group.

8. 8. The computer-implemented method of claim 6 or 7, wherein each of the selected unstructured attribute values ​​is present with the same frequency of occurrence in each of the two sets.

9. 9. The computer-implemented method of claim 6, wherein comparing the two records comprises: comparing the values ​​of the structured attributes of the two records, resulting in individual match scores for each structured attribute of the records; combining the individual match scores using the initial contribution weights; and comparing the combined score to a predefined threshold.

10. The computer-implemented method of claim 1 , further comprising merging the two records into a single record, wherein the two records represent the same entity.

11. The computer-implemented method of claim 1 , performed in response to receiving a respective request to match the record.

12. 12. The computer-implemented method of claim 1, further comprising repeating the step of comparing the two sets of unstructured attribute values ​​to compare further records of the database until all records of the database have been compared.

13. 13. The computer-implemented method of claim 1, wherein the method is performed by a master data management (MDM) system, the records to be compared are MDM records, and the identification of the set of one or more values ​​of attributes is performed by an entity detection module of the master data management system.

14. The computer-implemented method of claim 1 , wherein the unstructured data object corresponds to a document.

15. The computer-implemented method of claim 14 , wherein the unstructured data object corresponds to a scanned document.

16. 16. The computer-implemented method of claim 1, further comprising providing information to a person associated with the compared records indicating the unstructured data objects associated with the two records.

17. The computer-implemented method of claim 1 performed in response to storing the compared records.

18. The computer-implemented method of claim 5 , wherein the intersection includes the selected unstructured attribute values.

19. 1. A computer program for record matching in a database system, wherein a record represents an entity, the record being associated with one or more unstructured data objects, the computer program comprising: The processor processing the unstructured data object for each record of a plurality of records of a database to identify a set of one or more values ​​of attributes, hereinafter referred to as unstructured attribute values, in the unstructured data object for each record; comparing the sets of unstructured attribute values ​​of two records in the database to determine a level of similarity between the two compared sets; determining whether the two records represent the same entity based on the results of the comparison; 1. A computer program comprising:

20. 1. A computer system for record matching, wherein a record represents an entity, the record being associated with one or more unstructured data objects, the computer system comprising: one or more computer processors; one or more computer-readable storage media; and program instructions stored on the one or more computer-readable storage media, the program instructions comprising: instructions for processing the unstructured data object for each record of a database to identify a set of one or more values ​​of attributes, hereafter referred to as unstructured attribute values, in the unstructured data object for each record; instructions for comparing the sets of unstructured attribute values ​​of two records in the database to determine a level of similarity between the two compared sets; instructions for determining whether the two records represent the same entity based on a result of the comparison; A computer system comprising:

Citation Information

Patent Citations

  • Method and device of database-file cooperation

    JP2001282593A

  • Cooperative master data management

    JP2005537595A

  • Code conversion system, code conversion method, code correspondence relationship information generation method and computer program

    JP2008250861A

  • Supplementing Structured Information About Entities With Information From Unstructured Data Sources

    US20130325881A1

  • Automatic entity resolution with rules detection and generation system

    US20180137150A1