Identifying transformations in data architecture using anti-factual interpretation of entity matching
By using a graph neural network (GNN) model and actionable tracing techniques to generate a counterfactual explanation and transformation list, the problem of high computational resource consumption in probabilistic data matching engines is solved, and efficient data matching optimization is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2026-03-27
AI Technical Summary
Existing probabilistic data matching engines consume high computational resources and are inefficient during the data matching process, making it difficult to effectively identify and optimize data transformations to achieve entity matching.
By training a graph neural network (GNN) model and using an actionable tracing technique, an ordered list of counterfactual interpretations and transformations is generated, optimizing the data matching process and reducing computational resource consumption.
The counterfactual explanations and ordered transformation lists generated by the GNN model significantly reduce the computational resource consumption in the data matching process, and improve the efficiency and accuracy of data matching.
Smart Images

Figure CN121753034A_ABST
Abstract
Description
Background Technology
[0001] This disclosure generally relates to systems and methods for using entity matching, and more specifically, to computer-implemented methods, computer systems, and computer program products for generating and using counterfactual interpretations of entity matching.
[0002] In probabilistic data matching engines (such as Match 360), data transformations are performed to improve matching. Such transformations can also be used in other parts of the data architecture. Given that entity matching is a core component of the data architecture, it can be leveraged to improve other downstream tasks.
[0003] Explaining why two entities do not match, or what should be done to make two records match, can help identify data transformations. This way of explaining entity matching is called counterfactual interpretation. Summary of the Invention
[0004] In one embodiment, a system and method are provided that can generate counterfactual interpretations of entity matching, generate an ordered list of transformations, identify transformations to data from the counterfactual interpretations of entity matching, and identify an ordered list of transformations to be applied to the data.
[0005] In one embodiment, the computer-implemented method and computer program product can be configured to train a graph neural network (GNN) model based on the output of a probabilistic matching engine to perform entity matching. Counterfactual interpretations can be determined for non-matches of entities. A list of data transformations can be identified using the GNN model through actionable recourse. In some embodiments, the list of data transformations can be sorted using the GNN model based on computational overhead and improvements in entity matching.
[0006] In another embodiment, a system includes a processor, a data bus coupled to the processor, a memory coupled to the data bus, and a computer-usable medium containing computer program code including instructions executable by the processor. The instructions executable by the processor can train a graph neural network (GNN) model based on the output of a probabilistic matching engine to perform entity matching and determine counterfactual interpretations for non-matches of entities. The GNN model can be used to identify a list of data transformations through actionable tracing.
[0007] These and other features will become apparent from the following detailed description of its illustrative embodiments, which is read in conjunction with the accompanying drawings. Attached Figure Description
[0008] The accompanying drawings are illustrative embodiments. They do not illustrate all embodiments. Other embodiments may be used additionally or alternatively. Details that may be obvious or unnecessary may be omitted to save space or for more efficient illustration. Some embodiments may be practiced using additional components or steps and / or without all components or steps shown. When the same reference numerals appear in different drawings, they refer to the same or similar components or steps.
[0009] Figure 1 Data entries from two data sources are shown, and various data transformations performed on them, consistent with the illustrative embodiment, are illustrated to achieve data entry matching.
[0010] Figure 2 A flowchart illustrating the identification of counterfactual interpretations, the identification of data transformations, and the creation of an ordered list of transformations, consistent with the illustrated and illustrative embodiments, is provided.
[0011] Figure 3 A flowchart illustrating a process consistent with the illustrative embodiments is shown; and
[0012] Figure 4 This is a functional block diagram illustration of a computer hardware platform, consistent with the illustrative embodiments, that can be used to implement methods for identifying counterfactual interpretations, identifying transformations of data, and creating an ordered list of transformations. Detailed Implementation
[0013] In the following detailed description, many specific details are illustrated by way of examples to provide a thorough understanding of the relevant teachings. However, it should be apparent, however, that these teachings can be practiced without these details. In other cases, well-known methods, processes, components, and / or circuits have already been described at a higher level without details to avoid unnecessarily obscuring aspects of these teachings.
[0014] As described in more detail below, aspects of this disclosure provide systems and methods for generating and using counterfactual explanations of entity matching. The generated counterfactual explanations can be used to generate an ordered list of transformations using a graph neural network (GNN) model, and to identify transformations to the data using actionable tracing. The generated counterfactual explanations can also be used to identify an ordered list of transformations to be applied to the data using actionable group tracing.
[0015] According to one aspect of this disclosure, a computer-implemented method, system, and computer program product are provided for improving entity matching in a probabilistic matching engine that performs several functions, including training a graph neural network (GNN) model based on the output of the probabilistic matching engine to perform entity matching, and determining counterfactual interpretations for non-matches of entities. The GNN model can be used to identify a list of data transformations through actionable tracing. Using a GNN model reduces computational resources compared to testing data transformations on a probabilistic matching engine.
[0016] In embodiments that can be combined with the foregoing embodiments, the list of data transformations can be sorted based on computational overhead and the estimated improvement in entity matching. The sorting of data transformations can provide those that offer improved entity matching while minimizing computational overhead.
[0017] In embodiments that can be combined with the foregoing embodiments, a GNN model is used to perform the sorting of the data transformation list. Compared to determining the order of data transformations on a probabilistic matching engine, using a GNN model reduces computational resources. In embodiments that can be combined with the foregoing embodiments, the computational cost can be determined using the feature cost (R) value, as discussed in more detail below.
[0018] In embodiments that can be combined with the foregoing embodiments, one or more data transformations can be deployed on the probabilistic matching engine to improve entity matching. Such deployment can result in improved data matching within the probabilistic matching engine. In embodiments that can be combined with the foregoing embodiments, an upper limit can be set on the number of data transformations in the sorted list of data transformations. Such an upper limit can limit how many data transformations can be performed on the probabilistic matching engine at a time.
[0019] In embodiments that can be combined with the foregoing embodiments, user feedback can be received to approve or reject one or more data transformations from a list of data transformations. The availability of user feedback and data management ensures that unwanted data transformations are rejected.
[0020] In embodiments that can be combined with the foregoing embodiments, rules can be generated based on a list of data transformations. These rules can be incorporated to improve matching in a probabilistic matching engine.
[0021] Although the operation / function descriptions presented in this paper are understandable to human thought, they are not abstract concepts of operations / functions separate from their computational implementation. Rather, operations / functions represent specifications for appropriately configured computing devices. As discussed in detail below, the operation / function language will be read in its appropriate technical context, i.e., as a concrete specification of the physical implementation.
[0022] Therefore, one or more methods discussed in this paper can provide a procedural model for identifying counterfactual interpretations, identifying transformations of data, and creating an ordered list of transformations. This can have the technical effect of significantly reducing the computational resources and overhead required to provide an ordered list of transformations using a GNN model, without needing to test the transformations of the data within the probabilistic matching engine.
[0023] It should be understood that the aspects taught herein are beyond the capabilities of human thought. It should also be understood that the various embodiments of the subject matter disclosed herein may include information that could not be manually obtained by an entity such as a human user. For example, the type, quantity, and / or kind of information included in the performance of the discussions herein may be more complex than information that could reasonably be manually processed by a human user.
[0024] refer to Figure 1 The first data source 100 may include columns for record id, customer name, DOB (date of birth), and other attributes relevant to identifying the individual. The second data source 102 may include columns for record id, first name, last name, DOB, and other attributes relevant to identifying the individual. Traditional probabilistic data matching engines may include a normalizer applied by default to the data. For example, the transformation indicated by arrow 106 may be a standard transformation that transforms an individual into a known equivalent of other attributes relevant to identifying the individual. Furthermore, the transformation indicated by arrow 104 may be a standard transformation where a split customer name can be equivalent to a combined first and last name. Learned transformations, such as those indicated by arrow 110, can also be used, where such transformations can be ignored or corrected to create entity matches.
[0025] Using the system and method of this disclosure, if a record shown in the first data source 100 is found to be mismatched with a record shown in the second data source 102, such a mismatch can be identified, and a counterfactual explanation suggesting that the last two digits of the birth year have been swapped can be provided. Furthermore, the system and method of this disclosure can identify what transformations can be used to create matches in a probabilistic data matching engine and the computational overhead for such transformations, as defined below. Finally, the system and method of this disclosure can then provide an ordered list of transformations that can optimize the matching of records in the probabilistic data matching engine.
[0026] refer to Figure 2The system 200, which learns counterfactual explanations about entity matching, can be provided with input data 202 from one or more data sources. A probabilistic matching engine 204, such as Match 360, can analyze the system's master data 206 and determine matches within the data 208. Non-matches can be output by the graph neural network and analyzed to determine why a match did not occur as a counterfactual explanation 210. For example, if the names and dates of birth of two data entries match in the data, but the probabilistic matching engine does not identify the two data entries as a match, the system 200 can determine that this is because some attributes related to identifying the individuals are different, such as different addresses.
[0027] Then, system 200 can determine, via actionable tracing 214, what transformations 212 can be used to match two data entries. For example, system 200 can determine that both the city name and the street name must be changed to create a match. GNN model predictions can be performed on the samples to determine whether certain transformations might lead to a match. This can be helpful because, as described below, the tracing overhead in probabilistic matching engines is very expensive. Furthermore, system 200 can receive human feedback before applying transformations. This use of data management is not available in traditional probabilistic matching engines.
[0028] Generally, for an operable trail, the trail rule r is a tuple r = (c, c') embedded in an if-then structure, i.e., "if c, then c'", where c and c' are conjunctions of predicates in the form of "feature-value". This is an operator (e.g., age ≥ 50). The pursuit set S is a set of unordered pursuit rules, i.e., S = {(c1, c′1), (c2, c′2)……(cL, c′L)}. The two-level pursuit set R is a hierarchical model with multiple pursuit sets, each embedded within an outer if-then structure. R = {(q1, c11, c′11), (q1, c12, c′12)……(q2, c21, c′21)……}, where qi corresponds to the subgroup descriptor. The two-level pursuit set can be used to provide pursuit to an instance x as follows: if x exactly satisfies one of the rules i in R, i.e., x satisfies qi ∧ ci, then its pursuit is c′. I If x does not satisfy any of the rules in R, then R cannot provide recourse to x; or if x satisfies more than one of the rules in R, then its recourse is given by the rule with the highest probability of providing correct recourse. It should be noted that this probability can be calculated directly from the data.
[0029] According to various aspects of the invention, the operable tracing 214 can be used for an entity matching task f: X → Y, with a target result y. Y, and instance x X, such that f(x) != y (Result not satisfied, no match). The counterfactual interpretation attempts to provide a perturbation vector 'a' that flips the prediction into an entity match, i.e., f(x+a) = y. These interpretations can be used alone, or applied to similar data after user approval, or learned as rules. Interpretations can also be used to learn data asset-specific transformations (operations). To achieve this, in f(x+a) = y Under the conditions, , where A is the set of feasible operations (i.e., perturbation vectors), and c is the cost function that measures the effort required to perform operation a.
[0030] The cost 218 (sometimes referred to in this paper as the computational overhead of operational tracing) can be determined by the feature cost (R), which can be defined as the sum of the costs of each feature present in c and whose value changes from c to c′, computed over all triples (q, c, c′) of R. Furthermore, the feature cost (R) can be determined as the sum of the magnitudes of the changes in feature values from c to c′, computed over all triples of R. The cost associated with each feature can be understood by comparing the input with pairwise features. For this purpose, the Bradley-Terry model can be used, which states that if p i,j If feature I is less operable (more difficult to change) than feature j, then this probability can be calculated as p. i,j – e βi / (e βi + e βj ), where β i and β j These correspond to the costs of features i and j, respectively. It should be noted that p... i,j It can be calculated directly from the paired comparisons obtained from the survey experts, i.e., p i,j = (Number of comparisons i>j) / Total number of comparisons i, j). This can be achieved by learning the feature β. i and β j The cost of obtaining features is estimated using the maximum a posteriori probability (MAP).
[0031] Identifying individual transformations may not be particularly useful; therefore, a list of transformations is typically provided as an ordered list 220 of transformations. System 200 can also create an ordered list 220 of transformations that can optimize matching in the probabilistic matching engine 204. The ordered list 220 of transformations can be learned via actionable group tracing, where multiple transformation operations can be modeled to determine transformations that optimize matching in the probabilistic matching engine 204 while minimizing cost 218. System 200 allows optional human review 222 and can generate rules 224, which can then be applied to the probabilistic matching engine 204 for future input data 202. The permutations used to generate counterfactual interpretations can provide a predetermined upper limit on the number of transformations to be applied to the data. The computational cost of testing the ordered list of transformations can be saved by attempting to trace the cost and identify the order first on the GNN model rather than in the probabilistic matching engine.
[0032] The counterfactual interpretation of multiple instances can be determined as follows. Given a set of N instances, , making A group and parameters Find the operation It is the optimal solution to the following problem:
[0033] .
[0034] Counterfactual Explanation Trees (CETs) can be used as follows. Given a set of... , It is the optimal solution to the following optimization problem:
[0035] .
[0036] Although the first learning goal Evaluation targeting x The average inefficiency of the assigned operation a = h(x) That is, cost c(a | x) and loss l 01 (f(x + a), +1), but the second term This is the total number of leaves (i.e., operations) included in h. This can be achieved by adjusting the parameters. This allows for a trade-off between the effectiveness of operations assigned by CET h and the interpretability of h. It should be noted that the CET framework can be applied to any classifier f and cost function c in existing counterfactual explanation methods.
[0037] Exemplary process
[0038] It may now be helpful to consider a more advanced discussion of the example procedure. For this purpose, Figure 3 An illustrative process 300 is presented related to methods for identifying counterfactual interpretations, identifying transformations of data, and creating an ordered list of transformations. Process 300 is shown as a collection of boxes in a logic flowchart, representing a series of operations that can be implemented in hardware, software, or a combination thereof. In the context of software, a box represents a computer-executable instruction that performs the operation when executed by one or more processors. Typically, computer-executable instructions may include routines, programs, objects, components, data structures, etc., that perform functions or implement abstract data types. In each process, the order in which operations are described is not intended to be construed as limiting, and any number of described boxes can be combined and / or executed in parallel in any order to implement the process.
[0039] refer to Figure 3 Box 302 of process 300 may include the operation of training a GNN model based on the output of the probabilistic matching engine. As described above, a GNN can be used to perform entity matching. While the probabilistic matching engine is optimized for scale, the trained GNN model can be optimized for fast and computationally inexpensive inference.
[0040] At box 304, GNN interpretability techniques can be used to find explanations for non-matching data (counterfactual explanations). At box 306, based on metrics from the probabilistic matching engine, the data administrator (using their domain knowledge) can optionally decide whether data transformations should be attempted to increase matching on the data assets. At box 308, actionable trails can be used to identify transformations. Typically, all possible data transformations can be identified, each of which is an actionable trail, as described above.
[0041] While the operable trace provides a list of transformations, at box 310, the operable group trace can be used to identify subsets of transformations, also known as an ordered list of transformations, where the GNN model can be used to run different orders and different subsets of the list of transformations to sort the transformations based on computational overhead and improvements in the estimated matches when deployed on a probabilistic matching engine.
[0042] Example computing platform
[0043] Various aspects of this disclosure are described by way of text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of a computer program product (CPP). With respect to any flowchart, depending on the technology involved, operations may be performed in a different order than that shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in a manner that at least partially overlaps in time.
[0044] Computer Program Product Embodiment (“CPP Embodiment” or “CPP”) is a term used in this disclosure to describe any set of one or more storage media (also referred to as “media”) collectively included in a set of one or more storage devices that collectively include machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A “storage device” is any tangible device capable of holding and storing instructions used by a computer processor. Without limitation, a computer-readable storage medium can be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), optical disc read-only memory (CD-ROM), digital universal disc (DVD), memory sticks, floppy disks, mechanical encoding devices (such as punched cards or pits / platforms formed in the main surface of the disk), or any suitable combination of the foregoing. As used herein, the term computer-readable storage medium should not be construed as storage in the form of transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, optical pulses through fiber optic cables, electrical signals transmitted through wires, and / or other transmission media. As those skilled in the art will understand, data typically moves at certain incidental points in time during the normal operation of the storage device (such as during access, defragmentation, or garbage collection), but this does not render the storage device transient, as the data is not transient at the time of storage.
[0045] refer to Figure 4 The computing environment 400 includes examples of environments for executing at least some computer code relating to the execution of the methods of the present invention, including counterfactual interpretation generation and usage engine block 500. In addition to block 500, the computing environment 400 includes, for example, a computer 401, a wide area network (WAN) 402, an end-user equipment (EUD) 403, a remote server 404, a public cloud 405, and a private cloud 406. In this embodiment, the computer 401 includes a processor set 410 (including processing circuitry 420 and a cache 421), a communication infrastructure 411, volatile memory 412, persistent storage device 413 (including an operating system 422 and block 500, as described above), a peripheral device set 414 (including a user interface (UI) device set 423, a storage device 424, and an Internet of Things (IoT) sensor set 425), and a network module 415. The remote server 404 includes a remote database 430. Public cloud 405 includes gateway 440, cloud orchestration module 441, host physical machine set 442, virtual machine set 443, and container set 444.
[0046] Computer 401 can take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device now known or developed in the future capable of running programs, accessing networks, or querying databases (e.g., remote database 430). As is well understood in the field of computer technology, and depending on the technology, the execution of computer-implemented methods can be distributed among multiple computers and / or multiple locations. On the other hand, in this presentation of computing environment 400, the detailed discussion focuses on a single computer, specifically computer 401, to keep the presentation as simple as possible. Computer 401 can reside in the cloud, even if it is... Figure 4 The computer 401 is not shown in the cloud. On the other hand, except to any extent that can be definitively indicated, the computer 401 is not required to be in the cloud.
[0047] Processor set 410 includes one or more computer processors of any type now known or to be developed in the future. Processing circuitry 420 may be distributed across multiple packages, for example, multiple coordinated integrated circuit chips. Processing circuitry 420 may implement multiple processor threads and / or multiple processor cores. Cache 421 is memory located within the processor chip package(s) and is typically used for data or code that should be readily accessible by the threads or cores running on processor set 410. Cache memory is typically organized into multiple levels based on its relative proximity to the processing circuitry. Alternatively, some or all of the cache used in the processor set may be located “off-chip.” In some computing environments, processor set 410 may be designed to work with qubits and perform quantum computing.
[0048] Computer-readable program instructions are typically loaded onto computer 401 to cause processor group 410 of computer 401 to perform a series of operational steps to implement a computer-implemented method, such that the instructions thus executed instantiate the method specified in the flowcharts and / or descriptive descriptions of the computer-implemented method included in this document (collectively, the “method of the invention”). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 421 and other storage media discussed below. The program instructions and associated data are accessed by processor group 410 to control and direct the execution of the method of the invention. In computing environment 400, at least some of the instructions for performing the method of the invention may be stored in persistent storage device 413 in block 500.
[0049] Communication structure 411 is a signal transmission path that allows various components of computer 401 to communicate with each other. Typically, this structure consists of switches and conductive paths, such as switches and conductive paths forming buses, bridges, physical input / output ports, etc. Other types of signal communication paths can be used, such as fiber optic communication paths and / or wireless communication paths.
[0050] Volatile memory 412 is any type of volatile memory now known or to be developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 412 is characterized by random access, but this is not necessary unless explicitly stated otherwise. In computer 401, volatile memory 412 is located in a single package and is internal to computer 401; however, alternatively or additionally, volatile memory may be distributed across multiple packages and / or located externally relative to computer 401.
[0051] The persistent storage device 413 is any form of non-volatile storage device for a computer that is now known or to be developed in the future. The non-volatility of this storage device means that the stored data is maintained regardless of whether power is supplied to the computer 401 and / or directly to the persistent storage device 413. The persistent storage device 413 may be a read-only memory (ROM), but typically at least a portion of the persistent storage device allows data to be written, deleted, and rewritten. Some common forms of persistent storage devices include hard disks and solid-state storage devices. The operating system 422 may take several forms, such as various known proprietary operating systems or open-source portable operating system interface type operating systems employing a kernel. The code included in box 500 typically includes at least some of the computer code involved in performing the methods of the present invention.
[0052] Peripheral device set 414 includes a collection of peripheral devices for computer 401. Data communication connections between peripheral devices and other components of computer 401 can be implemented in various ways, such as Bluetooth connectivity, near field communication (NFC) connectivity, connections via cables (such as Universal Serial Bus (USB) type cables), plug-in connections (e.g., Secure Digital (SD) cards), connections via local area communication networks, and even connections via wide area networks such as the Internet. In various embodiments, UI device set 423 may include components such as displays, speakers, microphones, wearable devices (e.g., goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Storage device 424 is an external storage device, such as an external hard drive, or a pluggable storage device, such as an SD card. Storage device 424 can be persistent and / or volatile. In some embodiments, storage device 424 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments requiring computer 401 to have substantial storage (e.g., where computer 401 locally stores and manages a large database), this storage can be provided by peripheral storage devices designed for storing very large amounts of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 425 consists of sensors that can be used in IoT applications. For example, one sensor could be a thermometer, while another could be a motion detector.
[0053] Network module 415 is a collection of computer software, hardware, and firmware that allows computer 401 to communicate with other computers via WAN 402. Network module 415 may include hardware such as a modem or Wi-Fi transceiver, software for packing and / or unpacking data for transmission over a communication network, and / or web browser software for transmitting data over the Internet. In some embodiments, the network control and network forwarding functions of network module 415 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing Software-Defined Networking (SDN), the control and forwarding functions of network module 415 are performed on physically separate devices, such that the control function manages several different network hardware devices. Computer-readable program instructions for performing the methods of the present invention can typically be downloaded to computer 401 from an external computer or external storage device via a network adapter card or network interface included in network module 415.
[0054] A WAN 402 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances using any technology now known or to be developed in the future for transmitting computer data. In some embodiments, a WAN 402 may be replaced and / or supplemented by a local area network (LAN) designed to transmit data between devices located in a local area, such as a Wi-Fi network. WANs and / or LANs typically include computer hardware such as copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.
[0055] End User Equipment (EUD) 403 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating computer 401) and can take any of the forms discussed above in conjunction with computer 401. EUD 403 typically receives helpful and useful data from the operation of computer 401. For example, assuming computer 401 is designed to provide recommendations to the end user, these recommendations are typically transmitted from network module 415 of computer 401 to EUD 403 via WAN 402. In this way, EUD 403 can display or otherwise present the recommendations to the end user. In some embodiments, EUD 403 can be a client device, such as a thin client, a thick client, a mainframe computer, a desktop computer, etc.
[0056] Remote server 404 is any computer system that provides at least some data and / or functionality to computer 401. Remote server 404 can be controlled and used by the same entity operating computer 401. Remote server 404 represents (multiple) machines that collect and store helpful and useful data for use by other computers, such as computer 401. For example, if computer 401 is designed and programmed to provide recommendations based on historical data, that historical data can be provided to computer 401 from a remote database 430 of remote server 404.
[0057] Public cloud 405 is any computer system available to multiple entities, providing on-demand availability of computer system resources and / or other computing capabilities, particularly data storage (cloud storage) and computing power, without direct active management by the user. Cloud computing typically leverages resource sharing to achieve consistency and economies of scale. Direct and active management of the computing resources of public cloud 405 is performed by the computer hardware and / or software of cloud orchestration module 441. The computing resources provided by public cloud 405 are typically implemented by virtual computing environments running on various computers constituting host physical set 442, which is a universe of physical computers in and / or available to public cloud 405. Virtual computing environments typically take the form of virtual machines from virtual machine set 443 and / or containers from container set 444. It should be understood that these VCEs can be stored as images and can be transferred between various physical machine hosts as images or after the VCEs are instantiated. Cloud orchestration module 441 manages the transfer and storage of images, deploys new instantiations of VCEs, and manages active instantiation of VCE deployments. Gateway 440 is a collection of computer software, hardware, and firmware that allows public cloud 405 to communicate via WAN 402.
[0058] Now, we will provide some further explanation of Virtualized Computing Environments (VCEs). A VCE can be stored as an "image." A new active instance of a VCE can be instantiated from an image. Two familiar types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows multiple isolated user-space instances, called containers, to exist. From the perspective of the programs running within them, these isolated user-space instances typically behave like real computers. Computer programs running on a regular operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running within a container can only use the contents of the container and the devices allocated to the container; this is a feature known as containerization.
[0059] Private cloud 406 is similar to public cloud 405, except that computing resources are only available for use by a single enterprise. While private cloud 406 is depicted as communicating with WAN 402, in other embodiments, private cloud can be completely disconnected from the internet and accessible only via a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types), typically implemented by different vendors. Each of the multiple clouds remains a separate and discrete entity, but the larger hybrid cloud architecture is bound together by standardization or proprietary technology that enables orchestration, management, and / or data / application portability across the multiple component clouds. In this embodiment, public cloud 405 and private cloud 406 are both part of a larger hybrid cloud.
[0060] in conclusion
[0061] The various embodiments described in this teaching are for illustrative purposes and are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles of the embodiments, their practical application, or technical improvements to technologies found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.
[0062] While what is considered the best state and / or other examples have been described above, it should be understood that various modifications may be made therein, and the subject matter disclosed herein can be implemented in various forms and examples, and the teachings can be applied to many applications, only some of which have been described herein. The appended claims are intended to claim protection for any and all applications, modifications, and variations that fall within the true scope of this teaching.
[0063] The components, steps, features, purposes, benefits, and advantages discussed herein are merely illustrative. None of them, and the discussions relating to them, are intended to limit the scope of protection. While various advantages have been discussed herein, it should be understood that not all embodiments are necessarily intended to include all advantages. Unless otherwise stated, all measurements, values, ratings, locations, sizes, dimensions, and other specifications set forth in this specification (including in the appended claims) are approximate and not precise. They are intended to have a reasonable range consistent with the functions they pertain to and the conventions of the art to which they belong.
[0064] Many other embodiments are also envisioned. These include embodiments with fewer, additional, and / or different components, steps, features, objects, benefits, and advantages. These also include embodiments in which the components and / or steps are arranged and / or ordered differently.
[0065] This document describes aspects of the disclosure with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0066] These computer-readable program instructions can be provided to a processor of a suitably configured computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / operations specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to function in a certain way; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / operations specified in one or more blocks of the flowchart and / or block diagram.
[0067] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / operations specified in one or more boxes of a flowchart and / or block diagram.
[0068] The calling flows, flowcharts, and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a portion of a module, segment, or instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions mentioned in the blocks may not occur in the order shown in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0069] While the foregoing has been described in conjunction with exemplary embodiments, it should be understood that the term "exemplary" means only as an example, and not the best or optimal. Nothing else stated or described above is intended or should be construed as causing any component, step, feature, object, benefit, advantage, or equivalent to be offered to the public, whether or not it is recited in the claims.
[0070] It should be understood that the terms and expressions used herein have the general meaning consistent with those in the respective fields of investigation and research, unless otherwise specified herein. Relational terms such as "first" and "second" may be used only to distinguish one entity or operation from another, without necessarily requiring or implying any actual such relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but may also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further constraints, an element beginning with "a" or "an" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes that element.
[0071] This abstract of disclosure is provided to allow the reader to quickly determine the nature of this technical disclosure. It should be understood at the time of submission that it is not intended to interpret or limit the scope or meaning of the claims. Furthermore, as can be seen in the foregoing detailed description, various features are combined in various embodiments for the purpose of simplifying this disclosure. The method of this disclosure should not be construed as reflecting an intention to have more features than expressly recited in each claim of the claimed embodiments. Rather, as reflected in the following claims, the subject matter of the invention lies in fewer than all features of a single disclosed embodiment. Therefore, the following claims are incorporated herein by reference, wherein each claim is independently claimed as a separate subject matter.
Claims
1. A computer-implemented method for improving entity matching, comprising: Train a graph neural network (GNN) model based on the output of the probability matching engine to perform entity matching; For non-matches between entities, a counterfactual interpretation can be determined; as well as The GNN model is used to identify a list of data transformations through actionable tracing.
2. The computer-implemented method according to claim 1 further includes: Based on improvements in computational overhead and estimated entity matching, the data transformation list is sorted.
3. The computer-implemented method of claim 2, wherein the GNN model is used to perform the sorting of the data transformation list.
4. The computer-implemented method of claim 2 further includes determining a feature overhead (R) value to calculate the computational overhead.
5. The computer-implemented method according to claim 2, further comprising: Deploy one or more data transformations from a sorted list of data transformations on the probabilistic matching engine to improve entity matching.
6. The computer-implemented method according to claim 2, further comprising: Set an upper limit on the number of data transformations in the sorted list of the data transformations.
7. The computer-implemented method according to any one of the preceding claims further includes receiving user input to approve or reject one or more data transformations in the data transformation list.
8. The computer-implemented method according to any one of the preceding claims further includes generating rules based on the data transformation list.
9. A system comprising: processor; The data bus coupled to the processor; Memory coupled to the data bus; as well as A computer-usable medium containing computer program code, the computer program code including instructions that are executable by the processor and configured to: Train a graph neural network (GNN) model based on the output of the probability matching engine to perform entity matching; For non-matches of entities, determine counterfactual interpretations; and The GNN model is used to identify a list of transformable data through actionable tracing.
10. The system according to claim 9, wherein, The instructions are also configured to sort the list of data transformations based on computational overhead and an improvement in estimated entity matching.
11. The system according to claim 10, wherein, The GNN model is used to perform the sorting of the data transformation list.
12. The system according to claim 10, wherein, The instructions are also configured to deploy one or more data transformations from a sorted list of data transformations on the probabilistic matching engine to improve entity matching.
13. The system according to claim 10, wherein, The instruction is also configured to set an upper limit on the number of data transformations in the sorted list of the data transformations.
14. The system according to any one of claims 9 to 13, wherein, The instruction is also configured to receive user input to approve or reject one or more data transformations in the data transformation list.
15. The system according to any one of claims 9 to 14, wherein, The instruction is also configured to generate rules based on the data transformation list.
16. A computer program product for improving matching in a probabilistic matching engine, the computer program product comprising a computer-readable storage medium having program instructions embodied therein, the program instructions being executable by a computer to cause the computer to: The graph neural network (GNN) model is trained based on the output of the probabilistic matching engine to perform entity matching; Determine the counterfactual interpretation for the mismatched entity; and The GNN model is used to identify a list of transformable data through actionable tracing.
17. The computer program product according to claim 16, wherein, The instructions are also configured to cause the computer to sort the list of data transformations based on computational overhead and an estimated improvement in entity matching.
18. The computer program product of claim 17, wherein the GNN model is used to perform the sorting of the data transformation list.
19. The computer program product according to claim 17, wherein, The instructions are also configured to cause the computer to deploy one or more data transformations from a sorted list of data transformations on the probabilistic matching engine to improve entity matching.
20. The computer program product according to any one of claims 16 to 19, wherein, The instructions are also configured to cause the computer to receive user input to approve or reject one or more data transformations in the data transformation list.