Advanced Analysis of Identity Graph Data Structures

An automated sandbox environment in entity resolution systems assesses data source changes in identity graphs, optimizing entity resolution by identifying active individuals and evaluating graph impacts, thus enhancing efficiency and accuracy.

JP7880326B2Active Publication Date: 2026-06-25LIVERAMP
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
LIVERAMP
Filing Date
2021-08-11
Publication Date
2026-06-25

AI Technical Summary

Technical Problem

Entity resolution systems face challenges in maintaining the accuracy and efficiency of identity graphs due to the dynamic nature of data sources, which require frequent evaluation and manual processes are impractical for large datasets, necessitating an automated and computer-feasible solution.

Method used

An automated sandbox environment is used to test combinations of candidate data sources, analyzing the impact on identity graphs through person, person + touchpoint, and activity value processes, identifying changes such as addition, removal, merging, or splitting of entities, and categorizing individuals as 'active' or 'inactive' based on client usage patterns.

Benefits of technology

Provides a comprehensive, automated analysis of identity graph evolution, enabling efficient evaluation of data source impacts while maintaining graph quality and client relevance, reducing manual effort and computational resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007880326000001
    Figure 0007880326000001
  • Figure 0007880326000002
    Figure 0007880326000002
  • Figure 0007880326000003
    Figure 0007880326000003
Patent Text Reader

Abstract

The environment measures the value of data sources as input to the identity graph in terms of the impact of their inclusion or removal. Combinations of candidate sources are sent to the sandbox environment to produce the desired output. The person process, person + touchpoint process, and activity value process are executed. Results include whether a person was added or removed, whether a person created a point of failure, and whether a person was merged or split. The output provides an analysis of the evolution of the identity graph within the entity resolution system based on the selection of data sets used to build the graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 070,911, filed Aug. 27, 2020, entitled "System and Method for Evolutionary Analysis of Identity Graph". Such application is hereby incorporated by reference in its entirety.

Background Art

[0002] Entity resolution systems are used to determine whether data related to real-world entities actually refers to the same entity or different entities. These systems may, for example, be used to determine whether different items of data related to a person actually relate to the same real-world person. This type of entity resolution system must overcome many complex issues, such as people using different names or nicknames in different contexts, name or address changes, different people with the same name, and so on. Entity resolution systems often use identity graphs to track data related to entities. An identity graph (or more generally, a data graph) is a data structure that links together data related to the same entity. For example, an identity graph may be formed from a set of nodes, each containing an item of data about a given entity, with edges connecting those nodes together if they relate to the same entity. Various types of data sources may be used to build and maintain identity graphs. The available data sources for a given area of ​​an entity may change over time, with new data sources becoming available or old data sources becoming unavailable. Therefore, the identity graph may be updated periodically or even continuously. The accuracy of the entity resolution system directly depends on the accuracy of the identity graph used to support the system; therefore, the data sources used to build and maintain the identity graph must be carefully selected.

[0003] The impact of a set of data sources on the evolutionary expansion of the identity graph within an entity resolution system can change throughout the system's lifespan. In entity resolution systems relating to individuals, data sources that were once valuable in terms of the unique scope of personally identifiable information (PII) that expresses a definition of a person may no longer provide such information as a particular PII rapidly spreads through many different data sources. Similarly, the quality of PII may degrade over time due to intentional or unintentional obfuscation, abbreviation, or copying errors concerning a particular PII. Existing data sources should be periodically re-evaluated to manage the costs associated with the data sources incorporated into the system and to maintain a continuous level of quality within the system. Furthermore, if a set of existing data sources needs to be removed for contractual or other reasons, it may be beneficial to determine whether the loss of this set of sources should be mitigated to maintain the quality of the system, and if so, which aspects of the identity graph require mitigation.

[0004] The situations described above may require an in-depth analysis of the sequence of changes to the data graph for the data sources involved and other associated sources. For example, if a candidate data source is intended as the final replacement for one or more existing sources, it may be advantageous to first determine how the removal of the existing sources might affect the identity graph. This would require starting with the existing graph, then removing all sources that are expected to be replaced. The candidate source would then be added to this updated version, and the impact of adding the new source would be evaluated. Finally, the original data graph would be compared to the completely modified graph to determine the overall differences.

[0005] The data graphs that form the foundation of a business entity resolution system are enormous, accommodating tens of billions to hundreds of billions of records and hundreds of millions to billions of individuals. Therefore, evaluations using the entire identity graph in a manual comparison process, as in the example above, require such large computing resources that a fully contextualized evaluation of the calculated results is not feasible. In addition, given the vast number of potential data sources and their constantly changing nature, performing the manual process described above to evaluate various choices is no longer practical. Thus, there is a need for systems and methods to perform this function in an automated manner, while also operating within a computer-feasible framework within a business-relevant timeframe.

[0006] The references mentioned in this section on technical background do not constitute prior art to the present invention. [Overview of the Initiative] [Means for solving the problem]

[0007] This invention relates to an automated environment capable of measuring the values ​​of individual sources or subsets of sources in terms of their actual impact on the underlying identity graph and in terms of direct comparisons between other sources. In a given implementation, a sandbox environment is created in which various combinations of candidate sources can be tested and the results determined. Person processes, person + touchpoint processes, and activity value processes may be executed as subcomponents of the system. The results include whether, as a result of the changes, persons (or persons + touchpoints) were added or removed in the sandbox combination, whether persons (or persons + touchpoints) created fault points, and whether persons were merged or split. The environment's output provides an analysis of the evolution of the identity graph within the entity resolution system, based on the selection of data sets used to build the graph.

[0008] These, as well as other, features, purposes, and advantages of the present invention will be better understood by considering the following detailed description in conjunction with the drawings described below. [Brief explanation of the drawing]

[0009] [Figure 1] This is an overall process flow diagram for an embodiment of the present invention. [Figure 2] This is a flowchart illustrating the process involving individuals in an embodiment of the present invention. [Figure 3] This is a flowchart of the person + touchpoint process for an embodiment of the present invention. [Figure 4] This is a flowchart of the activity value process for an embodiment of the present invention. [Modes for carrying out the invention]

[0010] Before the present invention is described in further detail, it should be understood that the scope of the invention is limited solely by the claims, and that the invention is not limited to the specific embodiments described, and that the terms used in describing the specific embodiments are for the purpose of describing those specific embodiments only and are not intended to limit them.

[0011] Embodiments of the present invention may now be described with reference to the accompanying drawings beginning with Figure 1. The first component of the present invention is the configuration of a “sandbox” test storage area 10 which will be used for the analysis of a specified data source. If only one sandbox 10 is desired, its geolocation is identified. For example, if the data to be interpreted has applicability across the entire United States, the selection of geolocation should endeavor to include major patterns of normalized cultural, socioeconomic, and ethnic diversity as much as the entire United States. To constitute a high-density subset of individuals expected for a geolocation, the sandbox should contain all personally identifiable information (PII) records for each person included. Selected individuals are chosen from those for whom the data graph shows recent evidence that the person has a strong association with that geolocation. One type of association is a postal link to a geolocation, such as a household containing individuals whose addresses are within that geolocation. Another type is a digital association, where at least one of a person’s telephone numbers has an area code associated with that geolocation and has evidence of recent use or activity. Once Sandbox 10 is configured, the associated resulting identity graph (resulting identity graph subset) for this subset is saved and represents an initial baseline from which a sequence of adjustments is made regarding the addition or removal of further data files.

[0012] The next component is a process that takes the identity graph and the names of the data sources 12 to be added or removed as input. This process then constructs the people from the input graph, along with input modifications, using a person formation process for the entire identity graph. In the case of adding a set of data sources 12, all of the data is added to the sandbox 10. This is necessary because some of the new data may reflect different geolocation information about the people in the sandbox 10. In the case of removing a set of data, only the PII records that contributed to the baseline graph by that set alone will be removed from the sandbox 10.

[0013] When the data in Sandbox 10 is modified, the same process used to construct the entire graph is used to create a merged identity graph in order to form people from Sandbox 10. Once people are formed, the modified process of the process that links the entire graph calculates persistent identifiers or links for both the formed people and their PII records. Persistence in this context means that any PII records or people that have not been changed during the person formation process will retain the same identifier used in the baseline, and any brand new PII record will acquire a new unique identifier, as well as a newly formed person whose defining PII comes solely from the new data. These identifiers can take any desired form, such as an alphanumeric string. In cases where people in the input data graph are changed only by the introduction of a new PII record, the baseline identifier is persisted. In cases where people in the input data graph are merged together, identifier assignments are made to minimize the changes that will be visible when using the matching service for a particular set of data, such as when a person in the graph splits into many different people, or when people in the graph lose some of their defining PII records. The process to achieve this requires a relevance assessment and matching request for each of the PII records involved. For example, in a case where a person is split into different people (because data previously identified as relating to one person is determined to actually relate to multiple people), the identifier of the original person is assigned to the new person whose data is most recent and who has the most matching hits for the defining PII record.

[0014] As new individuals are formed in a persistent manner and assigned identifiers, this modified sandbox data graph is saved to sandbox 10. If further modifications are needed (as described earlier), this identity graph may be used as input to this component in an iterative manner.

[0015] The next component of the present invention takes all the identity graphs configured in a desired modification sequence and calculates the difference between any pair of data sets. By default, data graphs are paired based on the linear order of configuration from the previous component, but any pair of data graphs may be compared by this component. In the example in Figure 1, there are two candidate sources A and B, and a candidate data source D to be removed. Thus, for comparison with the existing graphs, various combinations are calculated in sandbox 10, including adding only data source A, adding only data source B, removing only data source D, adding both data source A and data source B, adding data source B in combination with removing data source D, adding both data source A and data source B in combination with removing data source D, and so on, completing all possible combinations.

[0016] The difference calculated to explain the evolutionary impact on the graph represents the radical changes to the graph caused by the modification. One such change is the creation of new individuals from new data (which only occurs when new data is added). This difference indicates that some of the data provided by the newly added source is clearly different from what exists in the input data graph. However, because the input data graph is confined to a specific geolocation, only new individuals with postal, digital, or other touchpoint instances that directly link this geolocation to the individual are meaningful. A second change is the complete deletion of all existing PII records for individuals in the input data graph. This can occur when the modification is the removal of a set of data sources, and if it occurs, each instance is meaningful for the evolution of the input data graph. Subsequently, one or more individuals in the input data graph may be combined into a single individual by either the deletion or addition of data sources. This behavior (consolidation) is meaningful for the evolution of the input data graph because, regardless of how the consolidation occurred, its impact extends to individuals in the original input graph. The same applies to division, that is, when a single person splits into two or more different people.

[0017] The differences discussed so far relate to actual person formation, but a further general developmental effect that can be obtained concerns whether the actual PII record and the corresponding person have supporting data sources. Any PII record with only one contributing source is a "point of failure" record in the data graph, as already mentioned, because the removal of that contributing source can cause significant changes in the data graph. Therefore, it is important to identify the PII records that, when a set of data sources is removed from the data graph, have become such "point of failure" records rather than simply disappearing. Moving from the PII record level to the person level (i.e., disparate sets of PII records), a person becomes a "point of failure" person if the removal of a set of data sources creates a person for whom all defining PII records are "point of failure" records. This concept of a "point of failure" person must be extended to cases where not all defining PII records are "point of failure" records. This occurs when all records contain the PII that many, if not all, of the users or clients of the entity resolution system have as their definition of that person. If those records are removed in the future, even if the person could still exist in the data graph, clients would no longer be able to access or find that person. For example, person P1 has three PII records with multiple data sources supporting the represented PII, and one PII record that is a "point of failure". All clients who obtain this person as a result of a matching service obtain it only through the PII in the "point of failure" record. If the record is removed, the person will still be retained, but no clients will be able to access that person through the remaining three PII records.

[0018] Figure 2 illustrates the person process 20 described above. Using the standard source person record 21 and the modified person source record 23, the various processes applied are to check in step 25 whether people have been added or removed, to check for point of failure reduction in step 26, to check for consolidation in step 27, to count the added touchpoints in step 28, and to check in step 29 whether people have been split into multiple records. The partial results from each of these steps in the partial person process result 31 are merged in the person process merge 24 to create the person process result 22. Figure 3 similarly illustrates the person + touchpoint process 30. Using the standard source person + touchpoint record 36 and the modified source person + touchpoint record 33, the various processes applied are to check for person + touchpoints to be added or removed in step 35, and to check for point of failure reduction in step 37. The partial results from these two steps in the Partial Person + Touchpoint Process Result 38 are merged in the Person + Touchpoint Process Merge 34 to create the Person + Touchpoint Process Result 32.

[0019] The process then divides the computed data into two sets. The first (and primary) set is the difference containing the most sought-after individuals for a particular purpose, referred herein as “active” individuals. The second category is the complement of the first category, referred herein as “inactive” individuals. The concept of “active” is often based primarily on the residual log of the entity resolution system’s matching service, which provides information about which individuals were returned from the matching service and the specific PII records that produced the actual match. Although client input is not logged, this information provides a clear signal of which PIIs in the identity graph are responsible for the success of each match. Different perspectives exist on the definition of an “active” individual, and in many contexts, there is a desire to have a sequence of definitions that measure different degrees or types of activity. The present invention also considers any such user-defined sequence using the data available to the system in various embodiments. However, at least one of the definitions chosen to be used includes a temporal interpretation of client use of the resolution system’s matching service.

[0020] To calculate the set of active individuals, in some embodiments, the most recent time window is selected, with a width of at least six months. This width is calculated based on the historical usage patterns of the majority of clients in the system. For example, if the majority of clients use the match service monthly to quarterly, a six-month window will generate a signal of typical usage. In other situations, a larger window, such as twelve months, may be used. Using the time signal of client match log values, the count of the number of job units per client per PII record is the criterion for match. A job unit is either a single batch job from a single client, or a set of match calls for a transaction by a typical client that is temporally dense (appearing within clearly defined start and end times). A single PII record may be "hit" multiple times by the match service within a job unit, which can lead to an artificially distorted interpretation of the count. Therefore, for each client, a PII record "hit" per job unit is counted only once. In cases where it is desired that the concept of "active" be defined differently for different types of clients (such as financial institutions or retail businesses), the resulting signal is broken down into an appropriate number of sub-signals.

[0021] For each sub-signal, one interpretation of an "active" person is represented in terms of several patterns of time signals from the matching service results log. These patterns may include, but are not limited to, the relative relevance of most non-zero counts, whether the signal has increased or decreased from the oldest past time to the present, and the amount of monthly fluctuation (primary difference). For example, when a person changes their postal address or phone number, these changes are rarely propagated simultaneously to all of that person's financial and retail accounts. It often takes several months for the change to reach all of those accounts (if at all). In these cases, this new PII begins to appear slowly in the signal with a very small count, but over time, this signal will begin to exhibit a clear pattern of increasing counts. The magnitude of the count may be ignored, as it is this increasing count behavior that clearly indicates this new PII that is important to the client of the resolution system. Similarly, some companies purchase "prospect" files of potential new customers, and these files are often run through the system's matching service to see if any of the people in the file are already customers. Since such prospect files are not run at a constant frequency, these instances can be identified in the signal by numerous fluctuations, the differences being far greater than the usual and expected disturbances. This type of signal cannot indicate the interest of a known client (customer) and is therefore often not considered an "active" person.

[0022] When an active person is identified, the previously calculated differences between identity graphs are separated into those containing at least one active person and those not containing any active people. The evolutionary impact of the differences in the latter set has a significantly lower probability of altering the system's data graph in terms of affecting the system's clients compared to the former. Therefore, separating the differences helps in interpreting the results to consider the overall impact in a more expressive and justifiable way.

[0023] Figure 4 provides an overall view of this activity value process 40. The standard source 41 and the modified source 43 are used as inputs to check the record activity count process 45. The activity value result 42 is the output of this sub-process. Then, as shown in Figure 1, at the merge step 14, the person process result 22, the person + touch point result 32, and the activity value result 42 can be combined to produce an overall result 16 for the entire process.

[0024] The overall result 16 provides a count of the differences for each distinct type and is presented for every two or more counts. The following is an exemplary result where a single data source has been removed from the initial data graph of the sandbox 10. [5404267,[2571398,306,15],[3799,311,151],[190771,23105,20310],[209069,19,2]] The first value indicates that there were a total of 5.4M removed PII records when the PII records were contributed by only this one source. The next 3-tuple represents the difference from point to point for individuals who lose some, but not all, of their PII records. The first value (2.57M) shows the total number of individuals in the sandbox data graph where this occurred. The next two values ​​represent the count for two different definitions of "active" individuals, with the first value being less restrictive than the second. Subsequently, the next 3-tuple represents the same kind of count for individuals who lost all of their PII records, followed by 3-tuples for individuals who were split into two or more individuals, and finally 3-tuples for individuals who were merged with another individual. It should be noted that the effect of merging when data is removed may appear fragmentary, and this case is often overlooked. However, a PII record for an individual may be a definitive PII record that separates two or more strongly related subsets of PII records, and removing it removes enough context to continue the splitting of those subsets.

[0025] These steps interpret a single set of source files as a unit, independent of other sets that are the target (as described below, by deliberately ordering the sets and analyzing different permutations through the described process repeatedly for the same set, some relationship between multiple sets of source files can be inferred). Very often, the usage context starts with a (large) set of source files, and the question to answer is which subset of the entire set is the "good" subset that, when added to or removed from the entity resolution identity graph, enhances the resulting resolution and / or minimizes the negative impact on the resolution. From this larger perspective rather than the direct impact on person formation, the intent is to determine the impact on the resolution ability for each person with respect to the presented touchpoint instances that define the person, namely, postal addresses, email addresses, and phone numbers. If a person can have multiple PII records contributed by many data sources but there are no instances of a particular touchpoint type (no phone number, no email, etc.), there is no ability for a user of the resolution system to access that person through the matching service that uses that touchpoint type.

[0026] In another variation, the present invention addresses the "point of failure" problem not from the perspective of a specific PII record, but rather from the perspective of a minimal subset of source files whose removal would remove all instances of a given touchpoint type about a person. Below, we will use email addresses to illustrate the process, but the invention also applies to other touchpoint types such as phone numbers, postal addresses, IP addresses, and others. A source file (rather than a person in the identity graph) is a "point of failure" if, removing all PII records for which this file is the sole contributor from the data graph results in a person who had an email address before removal but does not after. Removing a source file often removes several email addresses about a person, but such removal of email addresses is not necessarily detrimental to either the development of the data graph or the current state of the client experience using the matching service. In fact, historically, early email addresses provided included a large number of "generated" or fake email addresses, but clients have never used those emails as PII about customers. Removing such email addresses can significantly improve the personification of data graphs. However, removing all email addresses for a particular person is far more likely to have a negative impact on the graph and the user experience using matching services.

[0027] The concept of a data source "point of failure" extends not only to a single source file but also to a subset of source files. Thus, in various embodiments, the present invention calculates the number of people in an input identity graph whose entire email address is lost. The input to this component is the input graph defined above and a set of data sources having PII records that will be considered for potential removal from the identity graph. Each element of the set of data sources may be either a single data source or a set of data sources (all of which either remain in the graph or all must be removed and are therefore treated as one).

[0028] As previously mentioned, both client and developmental impacts of any loss of information should be considered in relation to the previously defined concept of an "active" person. Again, this invention considers any sequence of definitions of the degree of "being active." The inputs are the previously defined input identity graph, the set of touchpoint types to be considered in the analysis, the sequence of definitions of an "active" person, and the set of source files to be considered for potential removal from the data graph. The types of calculations and outputs are described below. 1. About each input touchpoint type 1a. For each combination of subsets of sources: The count of people in the input data graph, which lost all instances of its own input touchpoint type due to combination removal but not in smaller subsets of combinations, is calculated for all people, as well as for each person included in the input definition of an "active" person, and 2. Possible output data formats include groupings based on all combinations of a single source file entry in the input, as well as lists sorted based on counts.

[0029] The results from these two main components ("person"-based differences and "source"-based differences) provide a rich, multidimensional view of the key areas of impact on proposed changes in the underlying data that form the identity graph of the resolution system. Often, a very narrow view drives proposals such as increasing the number of emails and other digital touchpoints to broaden the scope of matching services. However, each expected improvement comes at a cost in terms of some degree of negative impact. Decisions to make such changes have significantly altered the parameters and context that define the overall value and improvement concepts. Therefore, this invention is designed to provide a rich overview of the representation of these two key dimensions of data graph development.

[0030] The systems and methods described herein may be implemented in various embodiments by any combination of hardware and software. For example, in one embodiment, the system and method may be implemented by a computer system or a collection of computer systems, each including one or more processors that execute program instructions stored on a computer-readable storage medium coupled to the processor. The program instructions can implement the functionality described herein. Various systems and methods, as illustrated in the drawings and described herein, represent exemplary implementations. The order of steps in the methods may be changed, and various elements may be added to, modified, or omitted from the system.

[0031] The computing systems or computing devices described herein may be implemented using hardware components of cloud computing systems or non-cloud computing systems. A computer system may include, but is not limited to, any type of device, including, commodity servers, personal computer systems, desktop computers, laptop or notebook computers, mainframe computer systems, handheld computers, workstations, network computers, consumer devices, application servers, storage devices, mobile phones, or any type of computing node or device in general. A computing system may include one or more processors (any of which may include a number of processing cores, which may be single-threaded or multi-threaded) coupled to system memory via an input / output (I / O) interface. A computer system may further include a network interface coupled to the I / O interface.

[0032] In various embodiments, the computer system may be a single-processor system containing one processor or a multi-processor system containing multiple processors. The processor may be any suitable processor capable of executing computing instructions. For example, in various embodiments, the processor may be a general-purpose or embedded processor implementing one of a variety of instruction set architectures. In a multi-processor system, each processor may, but not necessarily, typically implement the same instruction set. The computer system also includes one or more network communication devices (e.g., network interfaces) for communicating with other systems and / or components on a communication network such as a local area network, a wide area network, or the Internet. For example, a client application running on a computing device may use the network interface to communicate with a server application running on a single server or a cluster of servers implementing one or more of the components of the system described herein in a cloud computing or non-cloud computing environment, as implemented in various subsystems. In another example, an instance of a server application running on a computer system can communicate with other instances of the application that may be implemented on other computer systems using a network interface.

[0033] A computing device also includes one or more persistent storage devices and / or one or more I / O devices. In various embodiments, a persistent storage device may correspond to a disk drive, tape drive, solid-state memory, other mass storage devices, or any other persistent storage device. A computer system (or a distributed application or operating system running on it) may, as desired, store instructions and / or data in a persistent storage device and, as needed, retrieve the stored instructions and / or data. For example, in some embodiments, persistent storage may include solid-state drives attached to its server nodes. Multiple computer systems may share the same persistent storage device, or they may share a pool of persistent storage devices, in which devices may represent the same or different storage technologies.

[0034] A computer system includes one or more system memories that can store code / instructions and data accessible by the processor. System memory may include multiple levels of memory and memory caches, for example, in a system designed to exchange information within memory based on access speed. Interleaving and exchange can extend to persistent storage in virtual memory implementations. Techniques used to implement memory may include, for example, static random-access memory (RAM), dynamic RAM, read-only memory (ROM), non-volatile memory, or flash-type memory. Similar to persistent storage, multiple computer systems may share the same system memory or a pool of system memory. System memory or memory may contain program instructions executable by the processor to implement the routines described herein. In various embodiments, program instructions may be encoded in binary, assembly language, any interpreted language such as Java®, a compiled language such as C / C++, or any combination thereof; the specific languages ​​listed herein are illustrative only. In some embodiments, program instructions can implement a number of separate clients, server nodes, and / or other components.

[0035] In some implementations, program instructions may include instructions that are executable to implement an operating system, which may be any of the various operating systems such as UNIX®, LINUX, MacOS®, or Microsoft Windows®. Some or all of the program instructions may be provided as a computer program product or software that includes a non-temporary computer-readable storage medium storing the instructions, and the instructions may be used to program a computer system (or other electronic device) and carry out processes according to various implementations. The non-temporary computer-readable storage medium may include any mechanism for storing information in a machine-readable format (e.g., software). Generally speaking, a non-temporary computer-accessible medium may include magnetic or optical media, computer-readable storage media such as disks or DVDs / CD-ROMs coupled to a computer system via an I / O interface, or memory media. The non-temporary computer-readable storage medium may also include any volatile or non-volatile media, such as RAM or ROM, which may be included in some embodiments of the computer system, as system memory or another type of memory. In other implementations, program instructions may be communicated using optical, acoustic, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.) transmitted over a network and / or a communication medium such as a wired or wireless link, which may be implemented via a network interface. The network interface may be used to interface with other devices, which may include other computer systems or any type of external electronic device.In general, system memory, persistent storage, and / or remote storage accessible on other devices over a network can store data blocks, copies of data blocks, metadata associated with data blocks and / or their state, database configuration information, and / or any other information that can be used to implement the routines described herein.

[0036] In certain implementations, an I / O interface can coordinate I / O traffic between the processor, system memory, and any peripheral devices in a system, including via a network interface or other peripheral interfaces. In some embodiments, the I / O interface can perform any necessary protocol, timing, or other data transformation to convert data signals from one component (e.g., system memory) into a format suitable for use by another component (e.g., the processor). In some embodiments, the I / O interface can include support for devices attached via various types of peripheral buses, such as the Peripheral Component Interconnect (PCI) bus standard or a variation of the Universal Serial Bus (USB) standard. Also, in some embodiments, some or all of the functionality of the I / O interface, such as interface to system memory, may be directly integrated into the processor.

[0037] The network interface can enable data exchange between a computer system and other devices attached to the network, such as other computer systems (which may implement one or more storage system server nodes, primary nodes, read-only nodes, and / or clients of the database system described herein). In addition, the I / O interface can enable communication between the computer system and various I / O devices and / or remote storage. Input / output devices may, in some embodiments, include one or more display terminals, keyboards, keypads, touchpads, scanning devices, voice or optical recognition devices, or any other devices suitable for inputting or retrieving data by one or more computer systems. These may be directly connected to a specific computer system, or broadly connected to a number of computer systems in a cloud computing environment, or to other systems involving a number of computer systems. The number of input / output devices may exist in communication with a computer system, or may be distributed across various nodes of a distributed system including a computer system. The user interfaces described herein may be visualized to the user using various types of display screen technologies. In some implementations, input may be received through a display using touchscreen technology, and in other implementations, input may be received through a keyboard, mouse, touchpad, or other input technology, or any combination thereof.

[0038] In some embodiments, similar input / output devices may be separate from the computer system and can interact with one or more nodes of a distributed system, including the computer system, via wired or wireless connections, such as a network interface. The network interface can typically support one or more wireless networking protocols (e.g., Wi-Fi / IEEE 802.11, or another wireless networking standard). The network interface can support communication over any suitable wired or wireless general data network, such as other types of Ethernet® networks. Additionally, the network interface can support communication over telecommunications / telephone networks, such as analog voice networks or digital fiber optic networks, over storage area networks, such as Fibre Channel storage area networks (SANs), or over any other suitable type of network and / or protocol.

[0039] Any of the distributed system embodiments described herein, or any of their components, may be implemented as one or more network-based services in a cloud computing environment. For example, read-write nodes and / or read-only nodes in the database hierarchy of a database system may present database services and / or other types of data storage services using the distributed storage systems described herein to clients as network-based services. In some embodiments, network-based services may be implemented by software and / or hardware systems designed to support interoperable machine-to-machine interaction over a network. Web services may have interfaces described in a machine-readable format, such as Web Services Description Language (WSDL). Other systems may interact with network-based services in a manner defined by the description of the network-based service's interface. For example, a network-based service may define a variety of actions that other systems can invoke, and may define specific application programming interfaces (APIs) that other systems may expect to adapt to when requesting these actions.

[0040] In various embodiments, network-based services may be requested or invoked through the use of messages containing parameters and / or data associated with network-based service requests. Such messages may be formatted according to a specific markup language, such as XML, and / or encapsulated using a protocol, such as the Simple Object Access Protocol (SOAP). To fulfill a network-based service request, a client of the network-based service may assemble a message containing the request and transmit that message to an addressable endpoint corresponding to the web service (e.g., a Uniform Resource Locator (URL)) using an internet-based application layer transport protocol, such as the Hypertext Transfer Protocol (HTTP). In some embodiments, network-based services may be implemented using Representational State Transfer (REST) ​​techniques rather than message-based techniques. For example, a network-based service implemented according to REST techniques may be invoked through parameters contained within HTTP methods such as PUT, GET, or DELETE.

[0041] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art in which the invention pertains. While any methods and materials similar to or equivalent to those described herein may also be used in the practice or experimentation of the invention, only a limited number of exemplary methods and materials are described herein. It will be apparent to those skilled in the art that many more modifications are possible without departing from the concept of the invention as described herein.

[0042] All terms used herein should be interpreted in the most broadly conceivable way, consistent with the context. In particular, the terms “comprises” and “comprising” should be interpreted non-exclusively when referring to elements, components, or steps, indicating that the referred elements, components, or steps may exist, be used, or be combined with other elements, components, or steps not explicitly mentioned. Where grouping is used herein, all individual members of a group, as well as all conceivable combinations and subcombinations of groups, are intended to be individually included in this disclosure. Where a scope is specified herein, all subscopes within that scope, and all distinct points within that scope, are intended to be individually included in this disclosure. All references cited herein are incorporated herein by reference, insofar as they do not conflict with the disclosure herein.

[0043] Although the present invention has been described with reference to certain preferred and alternative embodiments, these embodiments are intended to be illustrative only and are not intended to limit the scope of the invention to the entirety described in the appended claims.

Claims

1. A system for performing evolutionary analysis of data structures, An identity graph stored on one or more storage devices, wherein the identity graph includes records computerically represented on the storage device by nodes containing items of personally identifiable information (PII) about a person and edges connecting nodes related to the same person, and each node of the record includes a geolocation associated with the person, A geographically constrained sandbox storage area communicating with the identity graph, configured to provide an isolated test area for modifying the nodes and edges of the identity graph, One or more hardware processors communicating with the aforementioned one or more storage devices Includes, When the one or more storage devices are run by the one or more hardware processors, Creating a subset of the identity graph that consists only of the records containing the same geolocation, and storing the subset of the identity graph in the sandbox storage area. Adding at least one candidate data source to the aforementioned sandbox storage area, To produce at least one modified sandbox data graph, the personification process is used to combine a subset of the identity graph with the at least one candidate data source, To maintain continuity with the subset of the identity graph, the calculation of persistent identifiers for both persons and PII records in the at least one modified sandbox data graph, comprising: maintaining the same identifiers for unmodified PII records and persons; generating new identifiers for new PII records and new persons; and resolving identifier conflicts when existing persons are merged or split. Output a result set that identifies changes to records between a subset of the identity graph and the at least one modified sandbox data graph. Instructions that cause one or more hardware processors to perform an action including the above Remember this, The aforementioned modification includes at least one of the following: adding a person, removing a person, creating a point of obstruction, merging a person, or splitting a person. A system that enables improvement of the identity graph by providing quantifiable metrics for evaluating candidate data sources before integration into or removal from the identity graph, in a computer-efficient manner that reduces computing resource requirements compared to performing analysis on the complete identity graph.

2. The system according to claim 1, wherein the subset of the identity graph consists only of records relating to persons having postal connections to the same geolocation.

3. The system according to claim 2, wherein the subset of the identity graph consists only of records of individuals who are members of a household in the same geolocation.

4. The system according to claim 1, wherein the subset of the identity graph consists only of records relating to persons who have telephone numbers having area codes corresponding to the same geolocation.

5. The system according to claim 1, wherein the subset of the identity graph further comprises only records of persons who have had recent activity on telephone numbers having area codes corresponding to the same geolocation.

6. The system according to claim 1, wherein the at least one candidate data source includes data that will be removed from the subset of the identity graph.

7. The system according to claim 1, wherein the at least one candidate data source includes data that will be added to a subset of the identity graph.

8. The system according to claim 1, wherein the one or more storage devices further store instructions causing the one or more hardware processors to compute identifiers for persons in the at least one modified sandbox data graph when executed by the one or more hardware processors.

9. The system according to claim 8, wherein the identifier for a person in the at least one modified sandbox data graph includes a new identifier for a person present in the at least one modified sandbox data graph, rather than in a subset of the identity graph.

10. The system according to claim 8, wherein the identifier for a person in the at least one modified sandbox data graph includes a unified identifier for a person who is merged in the at least one modified sandbox data graph but was separate in a subset of the identity graph.

11. The system according to claim 1, wherein the at least one modified sandbox data graph comprises a plurality of modified sandbox data graphs.

12. The system according to claim 1, wherein the one or more storage devices further store instructions causing the one or more hardware processors to combine the subset of the identity graph with the at least one candidate data source in order to produce at least one modified sandbox data graph by performing a person process on the subset of the identity graph when executed by the one or more hardware processors.

13. The system according to claim 12, wherein the person process includes checking for persons to be added or removed.

14. The system according to claim 13, wherein the person process includes checking for the reduction of person fault points.

15. The system according to claim 14, wherein the aforementioned personnel process includes checking for integration.

16. The system according to claim 15, wherein the person process includes a process for counting added touchpoints.

17. The system according to claim 16, wherein the person process includes a process for checking the divided records.

18. The system according to claim 1, wherein the one or more storage devices further store instructions causing the one or more hardware processors to combine the subset of the identity graph with the at least one candidate data source in order to produce at least one modified sandbox data graph by performing a person + touchpoint process in the subset of the identity graph when executed by the one or more hardware processors.

19. The system according to claim 18, wherein the person-touchpoint process includes checking for people-touchpoints to be added or removed.

20. The system according to claim 19, wherein the person-to-touchpoint process includes checking for the reduction of person-to-touchpoint fault points.

21. The system according to claim 1, wherein the one or more storage devices further store instructions causing the one or more hardware processors to combine the subset of the identity graph and the at least one candidate data source to produce at least one modified sandbox data graph by performing activity processes in the subset of the identity graph to identify an active person, when executed by the one or more hardware processors.

22. A method for performing evolutionary analysis on a data structure using one or more hardware processors, A step of creating a geographically constrained subset of an identity graph containing multiple person records, wherein each person record contains multiple touchpoints related to the person, and the geographically constrained subset of the identity graph consists only of the person records that contain the same geolocation. The steps include storing a subset of the identity graph in a sandbox test storage area configured to provide an isolated environment for testing data modifications, A step of adding at least one candidate data source to the sandbox test storage area, wherein the at least one candidate data source includes multiple records, each containing multiple touchpoints related to a person. The steps include combining a subset of the identity graph and the at least one candidate data source using a personification process to produce at least one modified sandbox data graph, A step of calculating a persistent identifier for each of the plurality of person records in the at least one modified sandbox data graph in order to maintain continuity with the subset of the identity graph, the step of calculating a persistent identifier which includes maintaining the same identifier for person records that have not been changed, generating a new identifier for new person records, and resolving identifier conflicts when person records are merged or split. The steps include: outputting a result set that identifies specific changes to person records between a subset of the identity graph and the at least one modified sandbox data graph; Includes, The aforementioned modification includes at least one of the following: adding a person, removing a person, creating a point of obstruction, merging a person, or splitting a person. A method that enables optimization of data source selection for the identity graph by providing quantifiable metrics for evaluating the impact of candidate data sources in a computer-efficient manner using a reduced test environment, based on the result set.

23. The method according to claim 22, wherein the at least one candidate data source includes data that will be removed from the subset of the identity graph.

24. The method according to claim 22, wherein the at least one candidate data source includes data that will be added to the subset of the identity graph.

25. The method according to claim 22, further comprising the step of calculating an identifier for a person in the at least one modified sandbox data graph.

26. The method according to claim 25, wherein the identifier for a person in the at least one modified sandbox data graph includes a new identifier for a person present in the at least one modified sandbox data graph, rather than in a subset of the identity graph.

27. The method according to claim 25, wherein the identifier for a person in the at least one modified sandbox data graph includes a unified identifier for a person who is merged in the at least one modified sandbox data graph but was separate in a subset of the identity graph.

28. The method according to claim 22, wherein the step of outputting a result set that identifies specific changes to person records between a subset of the identity graph and the at least one modified sandbox data graph includes the step of performing a person + touchpoint process in the at least one modified sandbox data graph, wherein the person + touchpoint process includes one or both of the steps of checking for people to be added or removed and checking for the reduction of fault points among the people.

29. The method of claim 28, wherein the step of outputting a result set that identifies specific changes to person records between a subset of the identity graph and the at least one modified sandbox data graph includes the step of performing an activity process in the at least one modified sandbox data graph, the activity process including the step of identifying active people in the at least one modified sandbox data graph.

Citation Information

Patent Citations

  • System, method and computer program product for database change management

    US10268709B1

  • Branchable graph databases

    US20170212945A1

  • Profile enrichment

    US20170316380A1

  • Method and System for Distributed Processing in a Messaging Platform

    US20180121269A1

  • Data management for netcentric computing systems

    US7403946B1