Automatic detection of matching fields in entity resolution systems
By using matching functions and scoring mechanisms in the entity parsing system to automatically select and optimize matching fields, the problem of low matching efficiency caused by reliance on manual selection in existing technologies is solved, achieving more efficient and accurate data matching.
Patent Information
- Application Number
- CN202180048991.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-14
- Filing Date
- 2021-07-13
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2041-07-13
AI Technical Summary
Existing technologies struggle to efficiently select the correct data matching fields in entity parsing systems. They rely on users' domain expertise, and the system cannot automatically optimize the selection of matching fields, resulting in low matching efficiency.
The computer-implemented method utilizes matching functions and scoring mechanisms to automatically evaluate the usefulness of potential matching fields, select and optimize matching fields, and incorporates a feedback mechanism based on reference data to automatically recommend and adjust matching fields to improve matching accuracy.
It improves the accuracy and efficiency of data matching in the entity parsing system, reduces the occurrence of false affirmations and false negations, and optimizes the performance of the matching process.
Smart Images

Figure CN115803727B_ABST
Abstract
Description
Background Technology
[0001] This disclosure generally relates to the field of master data management, and more specifically to identifying data matching fields / attributes used in matching and linking entity resolution systems.
[0002] Typically, master data management is used to define and manage an organization's critical data. Master data management can provide processes for collecting, matching, merging, and distributing organizational data to allow for consistency, accuracy, and control over the use and maintenance of that data.
[0003] Entity resolution systems provide useful tools within master data management. They allow organizations to connect heterogeneous data sources to provide an understanding of possible entity matches and non-obvious relationships within and across different data sets. Summary of the Invention
[0004] According to an aspect of the invention, there exists a computer-implemented method, computer program product, and / or system that performs the following operations (not necessarily in the following order): obtaining a plurality of payload attribute fields associated with payload data; determining one or more potential matching fields from the plurality of payload attribute fields; determining a matching function for each of the one or more potential matching fields; determining an attribute score for each of the one or more potential matching fields based at least in part on the matching function; obtaining a score list for a reference dataset; determining the correlation between the attribute score of each of the potential matching fields and the score list of the reference dataset; selecting one or more new matching fields from the one or more potential matching fields based at least in part on the correlation between the attribute score and the score list of the reference dataset; selecting one or more attribute fields from the selected new matching fields for matching the payload data; and providing the one or more attribute fields for matching for matching data in an entity resolution system. Attached Figure Description
[0005] Figure 1 This is a block diagram view of a first embodiment of the system according to the present disclosure;
[0006] Figure 2 This is a flowchart illustrating a method of the first embodiment, which is at least partially performed by the system of the first embodiment;
[0007] Figure 3 This is a block diagram illustrating the machine logic (e.g., software) portion of the system of the first embodiment;
[0008] Figure 4 This is a functional block diagram illustrating another example embodiment capable of performing the methods according to this disclosure; and
[0009] Figure 5 Example properties of data records according to embodiments of this disclosure are shown. Detailed Implementation
[0010] According to aspects of this disclosure, systems and methods can be provided to allow the automatic detection and / or recommendation of new matching fields associated with payload data for matching and linking data in an entity parsing system. In particular, the systems and methods of this disclosure can provide methods for evaluating potential new matching fields / attributes based on payload data that has not yet been matched (e.g., payload fields / attributes), and determining the usefulness of the new matching fields, for example, by comparing them with reference data and determining their impact on the number of false positives and / or false negatives.
[0011] Typically, master data management is used to define and manage an organization's critical data. Master data management can provide processes for collecting, matching, merging, and distributing organizational data to allow for consistency, accuracy, and control over the use and maintenance of that data. Entity resolution systems provide useful tools within master data management and can allow organizations to connect heterogeneous data sources to provide an understanding of possible entity matches and non-obvious relationships within and across different data sets. Entity resolution can allow determining when references to real-world entities (e.g., within payload data records) refer to the same entity or different entities.
[0012] Master data management solutions typically involve matching and linking data as core capabilities. Finding duplicate matches in a given population usually involves a significant number of comparisons (e.g., n² comparisons), but with grouping in an index, the number of comparisons can be limited to the selected candidate set.
[0013] However, even with good candidate selection criteria, there are often too many possibilities to choose actual fields for matching from all available attributes (e.g., possibly hundreds) within the payload dataset. Typically, the person implementing the system will select a portion of the input records (e.g., a list of attributes) that forms the basis for comparisons (matching fields) during the initial phase of implementation. The key to optimal quality matching and performance is the ability to select the correct set of fields for grouping indexes and matching. Often, making such a selection relies on the user's domain expertise, and the system does not assist the user in selecting the correct set of fields to perform the matching. This disclosure provides a mechanism that uses feedback, leveraging comparisons with known good matching pairs, to automatically identify good matching fields from unmatched data (from payload attributes).
[0014] In some embodiments, after the initial deployment configuration of the entity resolution system is completed, the system can evaluate payload data matching fields, for example, by comparing the current matching results with pre-defined reference results. The system can then perform optimization by evaluating additional potential matching fields and determining the usefulness of these potential matching fields by examining the impact of each potential field on false positives and false negatives. The system can then automatically qualify or recommend one or more additional matching attribute fields to use with the corresponding weights of the matching attribute fields and new scoring thresholds.
[0015] This Detailed Description provides the following sub-sections: hardware and software environment; (multiple) example embodiments; further comments and / or embodiments; and limitations.
[0016] Hardware and software environment
[0017] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the invention.
[0018] Computer-readable storage media can be tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or recessed structures with instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0019] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device, or downloaded via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network) to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the respective computing / processing device.
[0020] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (e.g., Smalltalk, C++, etc.) and conventional procedural programming languages (e.g., the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, to perform aspects of this invention, electronic circuits, including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute computer-readable program instructions to personalize the electronic circuits by utilizing state information from the computer-readable program instructions.
[0021] This document describes aspects of the invention with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0022] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes an article of writing comprising instructions for implementing aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0023] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus or other device, perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0024] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions mentioned in the blocks may occur in a non-linear order as shown in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0025] Embodiments of possible hardware and software environments for the software and / or methods according to the present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a functional block diagram illustrating various parts of an exemplary networked computer system 100, which may include: a server subsystem 102; client subsystems 104, 106, 108, 110, 112; a communication network 114; a server computer 200; a communication unit 202; a processor group 204; an input / output (I / O) interface set 206; a memory device 208; a persistent storage device 210; a display device 212; an external device set 214; a random access memory (RAM) device 230; a cache memory device 232; and a program 300.
[0026] Subsystem 102 represents various computer subsystems in this invention in several aspects. Therefore, several parts of subsystem 102 will now be discussed in the following paragraphs.
[0027] Subsystem 102 may be a laptop computer, tablet computer, netbook computer, personal computer (PC), desktop computer, personal digital assistant (PDA), smartphone, or any programmable electronic device capable of communicating with the client subsystem via network 114. Program 300 is a collection of machine-readable instructions and / or data for creating, managing, and controlling certain software functions, which will be discussed in detail in the Example Implementations section of the Detailed Implementation section below. As an example, program 300 may include an attribute detector, a matching attribute recommender, etc.
[0028] Subsystem 102 is capable of communicating with other computer subsystems via network 114. Network 114 may be, for example, a local area network (LAN), a wide area network (WAN) such as the Internet, or a combination of both, and may include wired, wireless, or fiber optic connections. Typically, network 114 may be any combination of connections and protocols that support communication between server and client subsystems.
[0029] Subsystem 102 is shown as a block diagram with multiple double-headed arrows. These double-headed arrows (without separate reference numerals) represent a communication structure that provides communication between the various components of subsystem 102. This communication structure can be implemented using any architecture designed to transfer data and / or control information between processors (such as microprocessors, communication and network processors, etc.), system memory, peripheral devices, and any other hardware components within the system. For example, the communication structure can be implemented at least partially using one or more buses.
[0030] The multiple memory devices 208 and the multiple persistent storage devices 210 are computer-readable storage media. Typically, the multiple memory devices 208 may include any suitable volatile or non-volatile computer-readable storage media. It should also be noted that now and / or in the near future: (i) the external device set 214 may provide some or all of the memory for subsystem 102; and / or (ii) devices external to subsystem 102 may provide memory for subsystem 102.
[0031] Program 300 is stored in persistent storage devices 210 for access and / or execution by one or more respective computer processors in processor group 204, typically via one or more memories of memory devices 208. The persistent storage devices 210: (i) are at least more persistent than signals in transit; (ii) store programs (including their soft logic and / or data) on tangible media (such as magnetic domains or optical domains); and (iii) are much less persistent than permanent storage. Alternatively, data storage may be more persistent and / or permanent than the type of storage provided by the persistent storage devices 210.
[0032] Program 300 may include machine-readable and executable instructions and / or substantial data (i.e., the type of data stored in the database). For example, program 300 may include machine-readable and executable instructions to provide execution of methods as further disclosed herein. In this particular embodiment, persistent storage device 210 includes a magnetic hard disk drive. Among some possible variations, persistent storage device 210 may include a solid-state hard disk drive, a semiconductor storage device, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), flash memory, or any other computer-readable storage medium capable of storing program instructions or digital information.
[0033] The media used by the persistent storage device(s) 210 may also be removable. For example, a removable hard disk drive may be used with the persistent storage device(s) 210. Other examples include optical discs and disks, thumb drives and smart cards, which are inserted into the drive for transfer to another computer-readable storage medium that is also part of the persistent storage device 210.
[0034] In these examples, communication unit 202 provides communication with other data processing systems or devices outside of subsystem 102. In these examples, communication unit 202 includes one or more network interface cards. Communication unit 202 can provide communication by using one or both of physical and wireless communication links. Any software modules discussed herein can be downloaded to persistent storage devices (such as persistent storage devices 210) via a communication unit (such as communication unit 202).
[0035] I / O interface set 206 allows input and output of data to other devices that can be locally connected to the server computer 200 in a manner that enables data communication. For example, I / O interface set 206 provides connectivity to external device set 214. External device set 214 typically includes devices such as keyboards, keypads, touchscreens, and / or other suitable input devices. External device set 214 may also include portable computer-readable storage media, such as, for example, thumb drives, portable optical discs or disks, and memory cards. Software and data (e.g., program 300) used to implement embodiments of the invention can be stored on such portable computer-readable storage media. In these embodiments, the associated software may (or may not) be loaded, in whole or in part, onto persistent storage device 210 via I / O interface set 206. I / O interface set 206 is also connected to display device 212 for data communication.
[0036] Display device 212 provides a mechanism for displaying data to a user and may be, for example, a computer monitor or a smartphone display screen.
[0037] The programs described herein are identified based on applications that implement them in specific embodiments of the invention. However, it should be understood that any particular program terminology used herein is for convenience only, and therefore the invention should not be limited to use only in any particular application identified and / or implied by such terminology.
[0038] Various embodiments of the invention have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Numerous modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or improvements to existing technologies in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0039] (Multiple) Example Implementations
[0040] Figure 2 A flowchart 250 is shown depicting a computer-implemented method according to the present invention. Figure 3 A program 300 is shown for performing at least some of the method operations in flowchart 250. Regarding... Figure 2 One or more flowchart blocks may be identified by dashed lines and indicate optional steps that may be included, but are not required in the depicted embodiments. Extensive reference will now be made to the process in the following paragraphs. Figure 2 (For the method operation box) and Figure 3 (Regarding the software framework) This paper discusses the method and the associated software.
[0041] like Figure 2 As shown, in some embodiments, the operation for automatically identifying new data matching fields begins at operation S252, where the attribute classifier module 320, etc., obtains multiple payload attribute fields associated with the payload data to be processed in the entity parsing / master data management system. Processing proceeds to operation S254, where the attribute classifier module 320 can determine one or more potential matching fields from the multiple payload attribute fields associated with the payload data.
[0042] For example, such as Figure 5 As shown, the attributes of the input payload data record 500 can be divided into three categories: grouping field 502, existing matching field 504, and remaining payload field 506. In some embodiments, the grouping field 502 may include one or more index fields 508a, and may be used to classify data records into groups and / or index records. The matching field 504 may include one or more existing matching fields 510a, which may be used to identify matching entities within a data record. The remaining payload field 506 may include one or more new potential matching fields 512a, which may be evaluated by the system to limit and / or recommend new matching fields to be used in the entity resolution system. Figure 5 Legend 501 provides an indication of which attribute type is present in the payload data, such as index field 508a, existing match field 510a, and potential match field 512a.
[0043] In some embodiments, the attribute classifier module 320 may determine the data category of each of the payload attribute fields and make the determination of potential matching fields partially based on the data category of the attribute. In some embodiments, the determination of potential matching fields may be based on whether the attribute includes transaction information (e.g., information describing a single transaction rather than an entity).
[0044] The process proceeds to operation S256, where the attribute scorer module 325 and others determine an appropriate matching / comparison function for the potential matching field. For example, the matching / comparison function may include an exact match of the attribute field, a partial match of a string in the attribute field, a match of certain characters in the attribute field, etc. The process proceeds to operation S258, where the attribute scorer module 325 and others calculate a score for the potential matching field based on the selected matching / comparison function. In some embodiments, the attribute scorer module 325 may determine a matching score threshold to be used for comparing potential matching fields and / or an initial weight to be applied to the potential matching fields.
[0045] The process proceeds to operation S260, where the attribute scorer module 325, etc., obtains a score list for the reference dataset. In some embodiments, the score list for the reference dataset can provide expected matching results and the associated ratio of false positives to false negatives. In operation S262, the attribute scorer module 325, etc., determines the relevance of the attribute score of the potential matching field to the score list of the reference dataset. For example, the attribute scorer module 325 can provide a feedback mechanism that utilizes a comparison of the results of the potential matching field with known good matches in the results of the reference dataset.
[0046] The process proceeds to operation S264, where the attribute selector module 330, etc., selects one or more new matching fields from (a plurality of) potential matching fields based at least in part on the relevance of the attribute scores to the score list of the reference dataset. For example, in some embodiments, the attribute selector module 330, etc., may sort the potential matching fields based on descending order of relevance and select a limited number of top entries as new matching fields. In some embodiments, the attribute selector module 330, the attribute scorer module 325, etc., may update the matching threshold to account for the increase in the overall score based on including the additional matching fields.
[0047] In some embodiments, the process optionally proceeds to operation S266, where the weight generator module 335, etc., can determine the optimal weight for each of the selected new matching fields. For example, in some embodiments, the system can determine the optimal weight to use when scoring the new matching fields based on the optimal ratio of false positives and / or false negatives achieved when using the new matching fields during the matching process of reference data, etc. (e.g., adjusted from the initial weight of the payload attribute).
[0048] The process proceeds to operation S268, where the matching attribute limiting module 340, etc., selects one or more new attribute fields from the selected new matching fields to match the payload data. In some embodiments, for example, the selection of one or more new attribute fields for matching the payload is based at least in part on a threshold ratio for false positives and / or false negatives. In operation S270, the matching attribute limiting module 340, etc., limits and / or provides one or more new attribute fields for matching, and optionally, provides associated optimal weights for the new attribute fields used to match data in the entity resolution system.
[0049] Figure 4 A functional block diagram 400 is shown that enables the execution of another example embodiment of the method according to this disclosure. For example... Figure 4As shown, the attributes 402 associated with the data to be processed by the entity resolution system can be grouped into grouping / indexing attributes 404, guided comparison / matching attributes 406, and remaining payload attributes 408. The system can then process the payload attributes 408 to determine potential new matching attributes that can be used by the entity resolution system, for example, as shown in the pre-matching process 410.
[0050] The pre-matching process 410 may begin by determining the data category 412 for each of the payload attributes 408. The system can then determine the appropriate matching function 414 to be applied with the payload attributes(s)(s)408. The system can increase the matching threshold 416 to accommodate new attributes being included during the matching process. The system can also determine the initial weights 418 to be used for scoring the payload attributes 408.
[0051] Then, along with the guiding comparison / matching attribute 406, the processed payload attribute 408, the updated matching threshold 416, and the initial weights 418 of the payload attribute scores can be provided, so that the matching process 420 can be performed using (multiple) reference datasets.
[0052] The updated attributes used for matching can be used to match records in the reference dataset and are provided to determine the false positive and / or false negative ratios. Based at least in part on the matching results 422 and the false positive and / or false negative ratios, the system can determine the optimal weight 424 based on the best achievable false positive / false negative ratio for each of the potential new matching attributes (e.g., payload attributes). The payload attributes and their respective false positive / false negative ratios can be analyzed by the system 426. The system can then, for example, limit and / or recommend one or more payload attributes 428 as new matching attributes based on which payload attributes achieve false positive / false negative ratios within a defined threshold or limit. These new matching attributes can then be used in the entity parsing system to process the payload data, for example, to provide data matching and linking capabilities.
[0053] Further comments and / or examples
[0054] Some embodiments of the present invention can provide adjustments to the matching threshold used when analyzing scores for new potential matching attributes. For example, as a new matching field is added, the total score can increase, and the matching threshold can be adjusted to account for higher scores and avoid too many false positives during the matching process. In some embodiments, the matching threshold adjustment can begin by determining the score range of the matching function selected for the payload attribute (potential matching field), for example, a score range of 0 to N, and then the current threshold used for matching can be updated based on a factor of the score range (e.g., 0.9, 0.8, 0.7, 0.6, etc.). For example, the current threshold can be adjusted by a factor of 0.8N such that the updated threshold T used for matching... NEW =T OLD +0.8N.
[0055] Some embodiments of the present invention may provide for determining an optimal weight for each potential new matching field, for example, for adjusting the match score of the new matching field. In some embodiments, an initial weight (K) (e.g., initial weight K = 1) may be determined for the potential matching field before it is used in the matching process. The initial weight K may be used to determine the match score of the potential matching field using a reference dataset, and the matching field may be used to determine the total number of false positives and / or false negatives. In response to too many false positives / false negatives in the matching process, the initial weight may be adjusted by a defined amount. In one example, in response to too many false positives, the weight K may be reduced. For example, in some embodiments, the weight K may be reduced in steps of 0.1 when the optimal weight is determined. The scoring process of the matching field using the reference dataset may then be repeated, and the weight K may be further adjusted in the case of too many false positives and / or false negatives. This process of adjusting the weight K and repeating the scoring process may be repeated iteratively until an optimal combination of false positives / false negatives is achieved. In response to achieving the optimal combination of false positives / false negatives, the final adjusted weight may be set as the optimal weight for the matching field (e.g., an acceptable ratio of false positives and / or false negatives). Then, this optimal weight can be used to refine the total score of the matching field. For example, the total score S can be equal to the old score S0 plus the weighted value of the new score S1 (S = S0 + K * S1).
[0056] limited
[0057] The term "invention" should not be considered an absolute indication that the subject matter described by the term "invention" is covered by the filed claims or by claims that may eventually be published after the patent application; while the term "invention" is used to help the reader get a general sense that the disclosure herein is potentially new, this understanding is experimental and provisional as indicated by the use of the term "invention," and this understanding may change during the patent examination process as relevant information develops and the claims are potentially modified.
[0058] Examples: See the definition of "invention" - similar note applies to the term "example".
[0059] And / or: includes or; for example, A, B "and / or" C means that at least one of A, B, or C is true and applicable.
[0060] Include / contain / comprise: Unless otherwise expressly stated, this means "includes but is not necessarily limited to".
[0061] Data communication: any kind of data communication scheme known now or developed in the future, including wireless communication, wired communication and communication routing with wireless and wired components; data communication is not limited to: (i) direct data communication; (ii) indirect data communication; and / or (iii) data communication in which the format, packetization state, medium, encryption state and / or protocol remain constant throughout the data communication process.
[0062] Receive / Provide / Send / Input / Output: Unless otherwise expressly specified, these terms should not be considered to imply: (i) any particular degree of directness regarding the relationship between their object and subject; and / or (ii) the absence of any intermediate component, action, and / or thing between their object and subject.
[0063] Module / Submodule: Any collection of hardware, firmware, and / or software that operates to perform a certain function, regardless of whether the module is: (i) in a single local proximity; (ii) distributed over a wide area; (iii) in a single proximity within a large segment of software code; (iv) located within a single segment of software code; (v) located in a single storage device, memory, or medium; (vi) mechanically connected; (vii) electrically connected; and / or (viii) connected in a data communication manner.
[0064] Computer: Any device with significant data processing and / or machine-readable instruction reading capabilities, including but not limited to: desktop computers, mainframe computers, laptop computers, field-programmable gate array (FPGA) based devices, smartphones, personal digital assistants (PDAs), portable or embedded computers, embedded device type computers, and application-specific integrated circuit (ASIC) based devices.
Claims
1. A computer-implemented method comprising: obtaining a plurality of payload attribute fields associated with payload data; determining one or more potential match fields from the plurality of payload attribute fields; determining a match function for each of the one or more potential match fields; determining an attribute score for each of the one or more potential match fields based at least in part on the match function; obtaining a score list for a reference data set; determining a correlation of the attribute score for each of the potential match fields to the score list for the reference data set; selecting one or more new match fields from the one or more potential match fields based at least in part on the correlation of the attribute score to the score list for the reference data set; selecting one or more attribute fields for matching against the payload data from the one or more new match fields; and providing the one or more attribute fields for matching for use in matching data in an entity resolution system.
2. The computer-implemented method of claim 1, wherein determining the one or more potential match fields from the plurality of payload attribute fields is based in part on a data category of each of the plurality of payload attribute fields.
3. The computer-implemented method of claim 1, wherein selecting the one or more new match fields from the one or more potential match fields comprises ordering the potential match fields in a descending order of correlation and selecting a defined number of top entries as new match fields.
4. The computer-implemented method of claim 1, wherein the score list for the reference data set provides an expected match result and associated false positive and false negative rates.
5. The computer-implemented method of claim 1, further comprising determining an optimal weight for each of the selected new match fields, wherein determining the optimal weight for each of the selected new match fields comprises: determining an initial weight for one of the selected new match fields; performing a scoring process for the selected new match field based on the reference data set and the initial weight for the selected new match field; determining a total number of false positives and false negatives for the reference data set using the selected new match field; determining a new weight for the selected new match field based on a false positive and false negative rate; repeating the scoring and weight adjustment for the selected new match field until an acceptable false positive and false negative rate is achieved; determining a final adjusted weight as the optimal weight for the selected new match field; and providing the optimal weight for the selected new match field along with a corresponding one or more attribute fields for matching.
6. A computer-implemented method comprising: obtaining a plurality of payload attribute fields associated with payload data; determining one or more potential match fields from the plurality of payload attribute fields; determining a match function for each of the one or more potential match fields; determining an attribute score for each of the one or more potential match fields based at least in part on the match function; obtaining a score list for a reference data set; determining a correlation of the attribute score for each of the potential match fields to the score list for the reference data set; selecting one or more new match fields from the one or more potential match fields based at least in part on the correlation of the attribute score to the score list for the reference data set; selecting one or more attribute fields for matching against the payload data from the one or more new match fields; and providing the one or more attribute fields for matching for use in matching data in an entity resolution system.
7. The computer-implemented method of claim 6, wherein determining the one or more potential match fields from the plurality of payload attribute fields is based in part on a data category of each of the plurality of payload attribute fields.
8. The computer-implemented method of claim 6, wherein selecting the one or more new match fields from the one or more potential match fields comprises ordering the potential match fields in a descending order of correlation and selecting a defined number of top entries as new match fields.
9. The computer-implemented method of claim 6, wherein the score list for the reference data set provides an expected match result and associated false positive and false negative rates.
10. The computer-implemented method of claim 6, further comprising determining an optimal weight for each of the selected new match fields, wherein determining the optimal weight for each of the selected new match fields comprises: determining an initial weight for one of the selected new match fields; performing a scoring process for the selected new match field based on the reference data set and the initial weight for the selected new match field; determining a total number of false positives and false negatives for the reference data set using the selected new match field; determining a new weight for the selected new match field based on a false positive and false negative rate; repeating the scoring and weight adjustment for the selected new match field until an acceptable false positive and false negative rate is achieved; determining a final adjusted weight as the optimal weight for the selected new match field; and providing the optimal weight for the selected new match field along with a corresponding one or more attribute fields for matching.
6. The computer-implemented method of claim 1, wherein selecting one or more attribute fields for matching against the payload data from the selected new matching fields is based at least in part on a threshold ratio for false positives and false negatives.
7. The computer-implemented method of claim 1, further comprising: determining a score range for the matching function for each of the potential matching fields; determining an updated threshold for matching by adjusting a current threshold for matching based on a factor of the score range; and providing the updated threshold for matching for use in selecting one or more new matching fields from the one or more potential matching fields, wherein the updated threshold is adjusted for an increase in total score based on including additional matching fields.
8. A computer program product comprising a computer readable storage medium having stored thereon: program instructions programmed to obtain a plurality of payload attribute fields associated with payload data; program instructions programmed to determine one or more potential matching fields from the plurality of payload attribute fields; program instructions programmed to determine a matching function for each of the one or more potential matching fields; program instructions programmed to determine an attribute score for each of the one or more potential matching fields based at least in part on the matching function; program instructions programmed to obtain a list of scores for a reference data set; program instructions programmed to determine a correlation of the attribute score for each of the potential matching fields to the list of scores for the reference data set; program instructions programmed to select one or more new matching fields from the one or more potential matching fields based at least in part on the correlation of the attribute score to the list of scores for the reference data set; program instructions programmed to select one or more attribute fields for matching against the payload data from the selected new matching fields; and program instructions programmed to provide the one or more attribute fields for matching for use in matching data in an entity resolution system.
9. The computer program product of claim 8, wherein determining one or more potential matching fields from the plurality of payload attribute fields is based in part on a data category of each of the plurality of payload attribute fields.
10. The computer program product of claim 8, wherein selecting the one or more new matching fields from the one or more potential matching fields comprises ordering the potential matching fields in a descending order of correlation and selecting a limited number of top entries as new matching fields.
11. The computer program product of claim 8, wherein the list of scores for the reference data set provides an expected match result and an associated ratio of false positives and false negatives.
12. The computer program product of claim 8, wherein the computer- readable storage medium has further stored thereon: program instructions programmed to determine an initial weight for a selected one of the new matching fields; program instructions programmed to perform a scoring process for the selected new matching field based on the reference data set and the determined weight for the selected new matching field; program instructions programmed to determine a total number of false positives and false negatives for the reference data set using the selected new matching field; program instructions programmed to determine a new weight for the selected new matching field based on a ratio of false positives and false negatives; program instructions programmed to repeat the scoring and weight adjustment until an acceptable ratio of false positives and false negatives is achieved; program instructions programmed to determine a final adjusted weight as an optimal weight for the selected new matching field; and program instructions programmed to provide the optimal weight for the selected new matching field along with a corresponding one or more attribute fields for matching.
13. The computer program product of claim 8, wherein the computer- readable storage medium has further stored thereon: program instructions programmed to determine, for each of the potential matching fields, a score range for the matching function; program instructions programmed to determine an updated threshold for matching by adjusting a current threshold for matching based on a factor of the score range; and program instructions programmed to provide the updated threshold for matching for use in selecting one or more new matching fields from the one or more potential matching fields, wherein the updated threshold accounts for an increase in overall score based on including additional matching fields.
14. The computer program product of claim 8, wherein selecting one or more attribute fields for matching against the payload data from the selected new matching fields is based at least in part on a threshold ratio of false positives and false negatives.
15. A computer system comprising: a set of processors; and a computer-readable storage medium; wherein: the set of processors is structured, positioned, connected, and / or programmed to execute program instructions stored on the computer-readable storage medium; and the stored program instructions comprise: program instructions programmed to obtain a plurality of payload attribute fields associated with payload data; program instructions programmed to determine one or more potential matching fields from the plurality of payload attribute fields; program instructions programmed to determine a matching function for each of the one or more potential matching fields; program instructions programmed to determine an attribute score for each of the one or more potential matching fields based at least in part on the matching function; program instructions programmed to obtain a score list for a reference data set; program instructions programmed to determine an initial weight for a selected one of the new matching fields; program instructions programmed to perform a scoring process for the selected new matching field based on the reference data set and the determined weight for the selected new matching field; program instructions programmed to determine a total number of false positives and false negatives for the reference data set using the selected new matching field; program instructions programmed to determine a new weight for the selected new matching field based on a ratio of false positives and false negatives; program instructions programmed to repeat the scoring and weight adjustment until an acceptable ratio of false positives and false negatives is achieved; program instructions programmed to determine a final adjusted weight as an optimal weight for the selected new matching field; and program instructions programmed to provide the optimal weight for the selected new matching field along with a corresponding one or more attribute fields for matching. program instructions programmed to determine a relevance of the attribute score for each of the potential match fields to the list of scores for a reference data set; program instructions programmed to select one or more new match fields from the one or more potential match fields based at least in part on the relevance of the attribute score to the list of scores for a reference data set; program instructions programmed to select one or more attribute fields from the selected new match fields for use in matching against the payload data; and program instructions programmed to provide the one or more attribute fields for use in matching for use in matching data in an entity resolution system.
16. The computer system of claim 15, wherein determining one or more potential match fields from the plurality of payload attribute fields is based in part on a data category of each of the plurality of payload attribute fields.
17. The computer system of claim 15, wherein selecting the one or more new match fields from the one or more potential match fields comprises ordering the potential match fields in a descending order of relevance, and selecting a limited number of top entries as new match fields.
18. The computer system of claim 15, wherein the stored program instructions further comprise: program instructions programmed to determine an initial weight for one of the selected new match fields; program instructions programmed to perform a scoring process for the selected new match field based on the reference data set and the determined weight for the selected new match field; program instructions programmed to determine a total number of false positives and false negatives for the reference data set using the selected new match field; program instructions programmed to determine a new weight for the selected new match field based on a ratio of false positives and false negatives; program instructions programmed to repeat the scoring and weight adjustment until an acceptable ratio of false positives and false negatives is achieved; program instructions programmed to determine a final adjusted weight as an optimal weight for the selected new match field; and program instructions programmed to provide the optimal weight for the selected new match field along with the corresponding one or more attribute fields for use in matching.
19. The computer system of claim 15, wherein the stored program instructions further comprise: program instructions programmed to determine a score range for the matching function for each of the potential match fields; program instructions programmed to determine an updated threshold for matching by adjusting a current threshold for matching based on a factor of the score range; and program instructions programmed to provide the updated threshold for matching for use in selecting one or more new match fields from the one or more potential match fields.
20. The computer system of claim 15, wherein selecting one or more attribute fields for matching against the payload data from the selected new matching field is based at least in part on a threshold ratio for false positives and false negatives.
Citation Information
Patent Citations
A method and apparatus for outputting information
CN109271556A
Methods and systems for linking data records from disparate databases
CN110709826A