Massive heterogeneous data ingestion and user resolution

By employing batch indexing and lazy data interpretation techniques, the system addresses the issues of flexibility and accuracy when processing heterogeneous data in credit data systems. This enables fast and accurate credit data processing and analysis, thereby improving the system's efficiency and reliability.

CN116205724BActive Publication Date: 2026-04-21EXPERIAN INFORMATION SOLUTIONS INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
EXPERIAN INFORMATION SOLUTIONS INC
Filing Date
2018-01-31
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing credit data systems suffer from insufficient flexibility, adaptability, accuracy, reliability, and interoperability when processing large-scale heterogeneous data, resulting in low data integration efficiency, delayed credit report generation, and difficulty in providing accurate credit information in real time.

Method used

We employ a batch indexing process and a lazy data interpretation method, generate reverse personal identifiers (reverse PIDs) through hash functions, and perform minimal processing during the data ingestion phase to reduce the ETL process. We also utilize domain classification and domain vocabulary annotation techniques to achieve efficient association and labeling of heterogeneous data.

Benefits of technology

It improves the processing efficiency and accuracy of the credit data system, reduces the formation of defects, can quickly generate accurate credit reports and credit profiles, supports real-time data analysis, and reduces system complexity and storage requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116205724B_ABST
    Figure CN116205724B_ABST
Patent Text Reader

Abstract

This disclosure relates to systems and methods for efficiently organizing data association, attributes, annotation, and interpretation of large-scale heterogeneous data. Incoming data is received and identification information (“information”) is extracted. A multi-dimensional reduction function is applied to the information, and the information is grouped into sets of similar information based on the function results. Filtering rules are applied to the sets to exclude mismatched information. The sets are then merged into information groups based on whether they contain at least one common piece of information. Common links can be associated with information within a group. If the incoming data includes identification information associated with a common link, the incoming data is assigned a common link. In some embodiments, the incoming data is not modified but assigned to a domain.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese invention patent application 2018800088771 (invention title: large-scale heterogeneous data ingestion and user parsing), filed on January 31, 2018. Technical Field

[0002] This disclosure relates to a system and related methods for efficiently organizing and relating large-scale heterogeneous data elements to users, including data association, attributes, annotations, and explanations. This system and methods can be implemented to provide real-time access to users' historical data elements, which has not been previously achieved. Background Technology

[0003] Credit events can be collected, compiled, and analyzed to provide an individual's creditworthiness in the form of a credit report. A credit report typically includes multiple credit attributes, such as credit score, credit account information, and other information related to the user's financial worth. For example, a credit score is important because it establishes the necessary level of trust between transacting entities. Financial institutions, such as lenders, credit card providers, banks, car dealerships, and brokers, can conduct business transactions more securely based on credit scores. Summary of the Invention

[0004] Systems and methods related to data association, attributes, annotation, and interpretation systems and related approaches for efficiently organizing large-scale heterogeneous data are disclosed.

[0005] One general aspect includes a computer system for determining an account holder identifier for collected event information, the computer system comprising: one or more hardware computer processors; and one or more storage devices configured to store software instructions configured to be executed by the one or more hardware computer processors to cause the computer system to: receive multiple event messages associated with corresponding multiple events from multiple data sources; for each event message: access a data store including an association between the data source and an identifier parameter, the identifier parameter including at least an indication of one or more identifiers included in the event message from the corresponding data source; determine, at least based on the identifier parameter of the data source of the event message, the identifier included in the event message as indicated in the accessed data store; extract the identifier from the event message at least based on the corresponding identifier parameter, wherein the combination of identifiers includes a unique identifier associated with a unique user; access a multiple hash functions, each hash function associated with the combination of identifiers; and for For each unique identifier, multiple hash values ​​are calculated by evaluating multiple hash functions; based on whether the unique identifiers share a common hash value calculated using a common hash function, the unique identifiers are selectively grouped into sets of unique identifiers associated with the common hash value; for each set of unique identifiers: one or more matching rules are applied, the one or more matching rules including criteria for comparing unique identifiers within the set; unique identifiers satisfying the one or more matching rules are determined as unique identifier matching sets; unique identifier matching sets, each including at least one common unique identifier, are merged to provide one or more merged sets that do not have a common unique identifier with other merged sets; for each merged set: a reverse personal identifier is determined; the reverse personal identifier is associated with each of the unique identifiers in the merged set; for each unique identifier: an event message associated with at least one combination of identifiers associated with the unique identifier is identified, and the reverse personal identifier is associated with the identified event message. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.

[0006] Implementations may include one or more of the following features: A computer system wherein the hash function includes at least: a first hash function that evaluates a first combination of at least a portion of a first identifier and at least a portion of a second identifier extracted from event information; and a second hash function that evaluates a second combination of at least a portion of a first identifier and at least a portion of a third identifier extracted from event information. A computer system wherein the first hash function is selected based on an identifier type of one or more of the first identifier or the second identifier. A computer system wherein the first identifier is a user's Social Security number, the second identifier is a user's last name, and the first combination is a concatenation of fewer than all digits of the Social Security number and fewer than all characters of the user's last name. A computer system wherein a first event set includes a plurality of events associated with a first hash value, and a second event set includes a plurality of events each associated with a second hash value. A computer system wherein the identifier is selected from: first name, last name, initial of middle name, middle name, date of birth, Social Security number, taxpayer ID, or country ID. A computer system wherein the computer system generates a reverse mapping that associates the reverse personal identifier with each of the remaining unique identifiers in the merged set, and stores the mapping in a data store. The computer system further includes: assigning a reverse personal identifier to each of a plurality of event information including the remaining unique identifier, based on a reverse personal identifier assigned to the remaining unique identifier. The computer system further includes a hash function comprising a position-sensitive hash algorithm. The computer system further includes one or more matching rules comprising one or more identifier resolution rules that compare u in one or more sets with account holder information in an external database or CRM system to identify a match for one or more matching rules. The computer system further includes identifier resolution rules comprising criteria indicating the matching standard between account holder information and identifiers. The computer system further includes merging sets comprising, for each of the one or more sets, repeating the following process: pairing each unique identifier in the set with another unique identifier in the set to create a unique identifier pair; determining a common unique identifier in the pair; and, in response to determining the common unique identifier, grouping non-common unique identifiers from the pair having common unique identifiers until the list of unique identifiers contained within the resulting group is mutually exclusive between the resulting groups. The computer system further includes determining a common unique identifier in the pair by classifying the unique identifiers in the pair. Implementations of the described techniques may include hardware, methods or processes, or computer software on a computer-accessible medium.

[0007] Another general aspect includes a computer system comprising: one or more hardware computer processors and one or more storage devices configured to store software instructions configured to be executed by the one or more hardware computer processors to cause the computer system to: receive a plurality of events from one or more data sources, wherein at least some of the events have a heterogeneous structure; store the heterogeneous events for access by an external process; for each data source: identify a domain at least in part based on a data structure or data from the data source; access a vocabulary associated with the identified domain; and for each event information: determine whether the event matches some or all of the vocabulary; associate the event with a corresponding domain or vocabulary; and associate one or more tags with portions of the event based on the determined domain. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.

[0008] Implementations may include one or more of the following features. A computer system includes software instructions, when executed by one or more hardware processors, configured to cause the computer system to: receive a request for user-associated information in a first domain; execute one or more domain resolvers configured to identify user-associated events having one or more tags associated with the first domain; and provide at least some of the identified events to the requesting entity. The computer system further includes at least some of the identified events comprising only the portion of the identified events associated with one or more tags associated with the first domain. Implementations of the described techniques may include hardware, methods or processes, or computer software on a computer-accessible medium.

[0009] Another general aspect includes a computerized method comprising, through a computing system having one or more computer processors: receiving multiple event messages from one or more data sources, the multiple event messages having heterogeneous data structures; determining a domain in each of the one or more data sources, at least in part based on the data sources, data structures associated with the data sources, or one or more of the event messages from the data sources; accessing a domain dictionary associated with the determined domain, the domain dictionary including domain vocabulary, domain syntax, and / or annotation criteria; annotating one or more portions of the event messages from the determined domain using the domain vocabulary based on the annotation criteria; receiving a request for the event messages or data included in the event messages; interpreting the event messages based on one or more annotated portions of the event messages; and providing the requested data based on the interpretation. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method. Attached Figure Description

[0010] Certain embodiments will now be described with reference to the following accompanying drawings. Throughout the drawings, reference numerals may be used repeatedly to indicate the correspondence between referenced elements. The drawings are provided to illustrate exemplary embodiments described herein and not to limit the scope of this disclosure or the claims.

[0011] Figure 1A An exemplary credit data system according to some embodiments of the present disclosure is shown.

[0012] Figure 1B Exemplary methods for generating, processing, and storing credit data according to some embodiments are illustrated.

[0013] Figure 2A An example of processing a heterogeneous set of events in a sequence according to some embodiments is shown.

[0014] Figure 2B An exemplary credit data system that interacts with various applications or services according to some embodiments is shown.

[0015] Figure 3 An exemplary credit data system architecture for simultaneously creating credit status and credit associations for analysis, according to some embodiments, is shown.

[0016] Figure 4 An exemplary batch indexing process in this embodiment is illustrated, including identifier stripping, identifier matching, and identifier stamping.

[0017] Figure 5 An example of identifier stripping according to some embodiments is shown.

[0018] Figure 6 An exemplary process for reducing data dimensionality using a hash algorithm is shown according to some embodiments.

[0019] Figure 7 An exemplary identifier resolution process according to some embodiments is shown.

[0020] Figure 8 An exemplary collection merging process according to some embodiments is illustrated.

[0021] Figure 9 Examples of associating a reverse personal identifier (“reverse PID”) with a unique identifier according to some embodiments are shown.

[0022] Figure 10 An example of a reverse PID for credit event stamping is shown according to some embodiments.

[0023] Figure 11A-11D An exemplary implementation of the sample identifier matching process is shown.

[0024] Figure 12This is a flowchart of an exemplary method for efficiently organizing large-scale heterogeneous data according to some embodiments.

[0025] Figures 13A-13C An exemplary data model is shown, illustrating the probability of defects associated with data as it flows from data ingestion to data consumption.

[0026] Figure 14 Various types of data sources that can provide heterogeneous event information about individuals are illustrated, and these data can be accessed and analyzed in various embodiments.

[0027] Figure 15 Exemplary domains and their associated vocabulary are shown according to some embodiments.

[0028] Figure 16 Exemplary systems and processes, according to some embodiments, are shown for labeling event information and then using the labeled event information to provide data understanding.

[0029] Figure 17 This is a flowchart of an exemplary method for interpreting incoming data to minimize the impact of defects in the system, according to some embodiments. Detailed Implementation

[0030] This disclosure provides various architectures and embodiments of systems and methods related to data association, attribute, annotation, and interpretation systems and related methods for efficiently organizing large-scale heterogeneous data. The disclosed systems and methods can be implemented to provide credit data based on an intelligent and efficient credit data architecture.

[0031] More accurate and reliable credit-related information can further enhance the confidence level of entities reviewing such information. For example, providing accurate and reliable credit statements, cash flow statements, balance statements, credit scores, or other credit attributes can more accurately depict an individual's creditworthiness. Ideally, collecting all credit-related information relevant to an individual and updating their credit attributes each time such information is collected would provide this more accurate and reliable credit attribute. However, there are very practical technical challenges in obtaining more timely, accurate, and reliable credit attributes. The same or similar challenges can also apply to other types of data collection, storage, and analysis. For example, systems may also be hampered by the timely parsing of large amounts of event data related to specific individuals and their association with travel-related events, crime-related events, education-related events, etc. Therefore, any discussion of technical issues and solutions in the context of credit-related information in this paper is equally applicable to other types of information.

[0032] One technological challenge involves processing a massive volume of credit events that need to be collected, analyzed, stored, and made accessible to requesting entities. For example, if there are 40 million people, and each person has 20 accounts (e.g., bank accounts, mortgages, car rentals, credit cards), then there are 800 million accounts continuously generating credit events. By modest assumptions, if each credit event contains 1,000 bytes of data, a large volume of raw credit events over 12 months could be approximately 10 gigabytes or more. If some internal guidelines or external regulations require archiving credit events for 5 years, the data volume could approach 50 gigabytes. The increasing trend of digital transactions driven by population growth and the rise of digital transaction applications further complicates this challenge. In traditional data collection models, data collection and analysis are viewed as different steps in a cross-process, which may not meet the needs for rapid analysis, presentation, and reporting.

[0033] Another technical challenge involves processing event data in various formats. Events can be received from a variety of entities, such as lenders, credit card providers, banks, car dealerships, brokers, etc. Typically, credit events provided by entities have proprietary data structures or schemas. The collected data is usually stored in databases (such as relational databases), which, while having the advantage of organizing standard data structures in a structured way, may be inadequate when collecting data with heterogeneous structures. Furthermore, such databases may require resource-intensive processes involving extract, transform, and load (ETL) operations. ETL operations often also require significant programming effort to merge data structures from new data sources.

[0034] Even if the collected data is successfully transformed into a schema that conforms to the database provided by the database, the schema is often too rigid to accommodate the information. With the continuous emergence of new data sources with completely different data structures, schema expansion can quickly become a daunting task. Therefore, database managers make decisions based on either (1) pruning additional information that might become important at some point (essentially pruning to fit square data into a circular schema), or (2) completely discarding incompatible information if it is known that future analysis will be inaccurate. Neither approach is ideal, as both lead to incompleteness or inaccuracy.

[0035] Beyond the challenges of data collection, there are also technical challenges related to analysis. For example, such systems can be extremely slow in generating credit reports for individuals. The system searches through gigabytes of data (each year) for records matching the requested individual to generate a credit statement. Such systems could take days or weeks to calculate the credit statements for 40 million people. Not only do delayed statements fail to reflect an individual's current status, but it also indicates that a significant amount of computational resources are tied to the task of generating these statements. This is undesirable for mechanisms that detect fraud based on credit data, as the data on credit reports may be days old by the time it is provided to the user. Furthermore, even if a fraudulent transaction has been resolved, it can take days, weeks, or more to update the change on the credit report. Therefore, it is no exaggeration to say that the credit statements generated by these reporting systems can be misleading in terms of the true creditworthiness of the individuals they reflect.

[0036] Delayed access to results is not the only challenge in analytics. Often, an individual's personally identifiable information is inaccurate or outdated. For example, someone might use "101 Main Street" for a credit card but "101 Main St." for her mortgage account, or, quite commonly, change their phone number. A credit event from one financial institution might use a more recent phone number, while a credit event from another financial institution might use an expired one. This irregularity and outdated personally identifiable information present unique challenges for data analysts, such as accurately resolving user credit events from multiple sources based on mismatched personally identifiable information between those events.

[0037] Credit data storage and analysis systems can implement data models where a rigorous ETL process is placed near data ingestion to standardize incoming data. This ETL process involves reconstruction, transformation, and interpretation. As will be described, early interpretation can mean introducing defects into the data stream early, and extending the lifespan of each defect before data consumption, providing ample opportunity for its propagation. Furthermore, because such systems update the ETL process for each new incoming data with a new data structure, significant software and engineering work is spent merging the new incoming data. Ultimately, the marginal effort to maintain upstream interpretation can overwhelm such systems. Additionally, the ETL process can transform the original data or create essentially similar copies of it. If defects are discovered during interpretation after the original data has been converted to a standard format, significant information loss may occur. Alternatively, if the original event data is essentially copied, there is wasted storage space and a severe impact on the processing capacity of large datasets. In various implementations of credit data systems, one or more of the following technical problems or challenges may be encountered:

[0038] • Data integration methods, such as data warehouses and data marts, attempt to extract meaningful data items from incoming data and transform them into standardized target data structures;

[0039] • As the number of data sources increases, the size and complexity of the software required to transform data from various types of sources also increase;

[0040] • The marginal effort required to merge new data sources becomes increasingly large because merging new data sources and formats requires modifications to existing software;

[0041] • Merging new data sources and types may require modifications to the target data structure, necessitating the conversion of existing data from one format to another;

[0042] • The complexity of software modifications and data transformations can lead to defects. If defects persist unnoticed for an extended period, significant effort and costs must be incurred to eliminate their impact through further software modifications and data transformations, creating a continuous cycle.

[0043] • These data integration methods may have high leverage because they attempt to interpret and transform data closer to the point of ingestion.

[0044] Therefore, such credit data systems (and other high-volume data analytics systems) are technically challenged, at least in terms of lack of flexibility, adaptability, accuracy, reliability, interoperability, defect management, and storage optimization.

[0045] definition

[0046] To facilitate understanding of the systems and methods discussed herein, several terms are defined below. The terms defined below, as well as other terms used herein, should be interpreted to include the definitions provided, the common and conventional meanings of the terms, and / or any other implied meanings of the individual terms. Therefore, the following definitions do not limit the meaning of these terms, but are merely illustrative.

[0047] The terms “user,” “individual,” “consumer,” and “customer” should be interpreted to include individual persons and user groups, such as, for example, married couples or family partners, organizations, groups, and business entities. Furthermore, these terms may be used interchangeably. In some embodiments, these terms refer to the user’s computing device, rather than the actual human operator of the computing device, or the actual human operator of the computing device.

[0048] Personally identifiable information (also referred to herein as "PII") includes any information about a user that can be used alone to uniquely identify a particular user to a third party. According to embodiments, and depending on the combination of user data that may be provided to third parties, a PII may include first and / or last name, middle name, address, email address, Social Security number, IP address, passport number, vehicle registration number, credit card number, date of birth, and / or home / work / mobile phone number. In some embodiments, user IDs that are very difficult to associate with a particular user may still be considered a PII, such as if these IDs are unique to the corresponding user. For example, a user's Facebook numeric ID may be considered a PII for both Facebook and third parties.

[0049] User input (also known as "input") generally refers to any type of input provided by a user, intended to be received and / or stored by one or more computing devices to update displayed data and / or the manner in which the data is displayed. Non-limiting examples of such user input include keyboard input, mouse input, digital pen input, voice input, finger touch input (e.g., via a touch-sensitive display), gesture input (e.g., hand movement, finger movement, arm movement, movement of any other limb, and / or body movement), etc.

[0050] Credit data typically refers to user data collected and maintained by one or more credit bureaus (e.g., Experian, TransUnion, and Equifax), such as data that affects a consumer's creditworthiness. Credit data can include transaction or status data, including but not limited to credit inquiries, mortgage payments, loan status, bank accounts, daily transactions, number of credit cards, utility payments, etc. Depending on the implementation method (and possibly the regulations of the region where credit data is stored and / or accessed), some or all of the credit data may be subject to regulatory requirements restricting the sharing of credit data with requesting entities, such as those based on the Fair Credit Reporting Act (FCRA) and / or other similar federal regulations in the United States. As used herein, "regulated data" generally refers to credit data as an example of such regulated data. However, regulated data can include other types of data, such as HIPPA-regulated medical data. Credit data can describe every user data item associated with a user, such as account balances, account transactions, or any combination of user data items.

[0051] Credit documents and credit reports typically refer to collections of credit data associated with a user, such as those that can be provided to the user, to a requesting entity that has been authorized by the user to access the user's credit data, or to a requesting entity that has the permitted purpose of accessing the user's credit data without the user's authorization (e.g., under FCRA).

[0052] A credit event (or “event”) generally refers to information associated with an event reported by an institution (including banks, credit card providers, or other financial institutions) to one or more credit bureaus and / or the credit data systems discussed herein. A credit event may include, for example, information associated with payments, purchases, bill payment due dates, bank transactions, credit inquiries, and / or any other event that may be reported to a credit bureau. Typically, a credit event is associated with a single user. For example, a credit event may be details of a specific transaction, such as a purchase of a specific product (e.g., Target, $12.53, grocery store, etc.), or it may be information associated with a credit limit (e.g., Citi credit card, balance $458, minimum payment $29, credit limit $1000, etc.). Generally, a credit event is associated with one or more unique identifiers, where each unique identifier includes one or more unique identifiers associated with a specific user (e.g., a consumer). For example, each identifier may include all or part of a user’s PII, such as username, physical address, Social Security number (“SSN”), bank account identifier, email address, telephone number, country ID (e.g., passport or driver’s license), etc.

[0053] A reversed PID refers to a unique identifier assigned to a specific user to form a one-to-one relationship. A reversed PID can be associated with a user's identifier, such as a specific PII (e.g., SSN "555-55-5555") or a combination of identifiers (e.g., the name "John Smith" and the address "100 Connecticut Ave"), to form a one-to-many relationship (between the PID and each of the multiple combinations of identifiers associated with the user). A specific reversed PID can be associated with event data when the event data includes identifiers or combinations of identifiers associated with it (referred to herein as a "stamp"). Therefore, the system can identify event data associated with a specific user based on multiple combinations of user identifiers included in the event data, using the reversed PID and its associated identification information.

[0054] Credit Data System

[0055] When determining whether to lend to a user, allow them to open an account, or lease from a user, and / or when making decisions about many other relationships or transactions that may take credit value into account, entities (such as lenders, credit card providers, banks, car dealers, brokers, etc.) typically request and consider credit data associated with the user. Entities requesting credit data (which may include requests for credit reports or credit scores) may submit credit inquiries to credit bureaus or credit card resellers. Credit reports or credit scores can be determined based at least on the analysis and calculation of credit data associated with the user's bank accounts, daily transactions, number of credit cards, loan history, etc. Furthermore, previous inquiries from other entities can also influence a user's credit report or credit score.

[0056] Entities (e.g., financial institutions) may also want to obtain users' up-to-date credit data (e.g., credit scores and / or credit reports) to better decide whether to lend to them. However, there can be significant delays in generating new credit reports or credit scores. In some cases, credit bureaus may only update a user's credit report or credit score once a month. As mentioned above, the reason for these significant delays may be that credit bureaus need to collect, analyze, and compute large amounts of data to generate credit reports or credit scores. The process of collecting credit data (such as a user's credit score) that may affect a user's credit value from credit events is generally referred to as "data ingestion" in this document. Credit data systems can achieve data ingestion using system-to-system horizontal data flow, such as by using batch ETL processes (e.g., as briefly discussed above).

[0057] In an ETL data ingestion system, credit events associated with multiple users can be transferred from different data sources to a database (online system), such as one or more relational databases. The online system can extract, transform, and load raw data associated with different users from these different data sources. The online system can then normalize, edit, and write the raw data across multiple tables in the first relational database. When the online system inserts data into the database, it must match the credit data with identification data about the consumer to link the data to the correct consumer record. When new data arrives, the online system needs to repeat this process and update multiple tables in the first relational database. Because incoming data (such as names, addresses, etc.) often contains errors, does not conform to established data structures, is incomplete, and / or has other data quality or integrity issues, new data may trigger a re-evaluation of some previously determined data links. In this case, the online system may unlink the credit data and relink it to new and / or historical consumer records.

[0058] In some cases, certain event data should be excluded from credit data storage, such as if errors are detected in the data files provided by the data source, or if there are defects in the credit data system software that may have misprocessed historical data. For example, a non-intelligent credit data system storing data in date format MM / DD / YYYY might accept incoming data from a data source using date format DD / MM / YY, potentially introducing errors into the calculation of a user's credit value. Alternatively, such data might cause the credit data system to reject it entirely, leading to incomplete and / or inaccurate calculations of the user's credit value. Worse still, if erroneous data has already been consumed by the credit data system to measure a user's (albeit inaccurate) credit value, the system may need to address the complex situation of not only excluding the erroneous data but also mitigating all its effects. Failure to do so could leave the online database in an inconsistent or inaccurate state.

[0059] This incremental processing logic makes the data ingestion process complex, error-prone, and slow. In an ETL implementation, the online system can send data to a batch processing system that includes a second database. The batch processing system can then extract, transform, and load data associated with a user's credit attributes to generate credit scores and analytics reports for promotional and account review purposes. Because extracting, transforming, and loading data into the batch processing system takes time, credit scores and analytics reports may lag behind the online system by hours or even days. In cases where user identification data is being updated, the lagging batch processing system may continue to reflect outdated and potentially inaccurate user identification data, potentially breaking the link between incoming credit data and user data and providing inaccurate credit data until the link is corrected and transmitted to the batch processing system.

[0060] Overview of Improved Credit Data Systems

[0061] This disclosure describes a faster and more efficient credit data system designed to address the aforementioned technical problems. The credit data system can sequentially process heterogeneous sets of events, simultaneously create credit status and credit attributes for analysis, perform batch indexing processes, and / or create credit profiles in real time by merging credit status and real-time events; each process will be described in further detail below.

[0062] The batch indexing process effectively "clusters" unique identifiers by first reducing the dimensionality of the original credit events, identifying false positives, and then providing a complete set of empirically verified unique identifiers that can be associated with a user. This allows for more efficient large-scale association of credit events with the correct users. By using the process combination of this invention in a specific order, the credit data system solves the specific problem of effectively identifying credit events belonging to a particular user with exponential efficiency. Furthermore, assigning inverse PIDs enables a new and more efficient data configuration, allowing the credit data system to provide requested credit data about a user at exponentially faster speeds. The improved credit data system can generate various analyses (such as credit reports) of a user's activity and status based on the latest credit events associated with that user.

[0063] Credit data systems enable lazy data interpretation, where the system does not alter heterogeneous incoming data from multiple data sources, but rather annotates or labels the data without performing ETL processes. By performing minimal processing only near data ingestion, the credit system minimizes the software size and complexity around data ingestion, significantly reducing defect formation and management issues. Furthermore, by omitting ETL processing and preserving data in its original heterogeneous form, the system can accept any type of data without losing valuable information. Domain classification and domain vocabulary annotation provide new data structures that allow for late-stage interpretation components such as parsers. Late-stage parsers reduce the overall defect impact on the system and make parsers easier to add and adapt, thus improving existing systems.

[0064] While this document discusses some embodiments of credit data systems or other similar naming systems with reference to various features and advantages, any of the features and advantages discussed may be combined or separated within the additional constraints of credit data systems.

[0065] Figure 1A An exemplary credit data system 102 of this disclosure is shown, which can be implemented by a credit bureau or an authorized agent of a credit bureau. Figure 1AIn this system, credit data system 102 receives credit events 122A-122C associated with different users 120A-120C. Credit data system 102 may include components such as an indexing engine 104, an identification engine 106, an event caching engine 108, a classification engine 110, and / or a credit data storage 112. As will be described in further detail, credit data system 102 can efficiently match specific credit events to appropriate corresponding users. Credit data system 102 may store credit events 122A-122C, credit data 114, and / or the associations between different users and credit events 122A-122C or credit data 114 in credit data storage 112, which may be a credit database of a credit bureau. In some embodiments, the credit database may be distributed across multiple databases and / or multiple credit data storage 112. Therefore, the credit data ingestion and storage processes, components, architectures, etc., discussed herein can be used to replace most existing credit data storage systems, such as batch processing systems. In response to receiving a credit inquiry request from an external entity 116 (e.g., a financial institution, a lender, a potential landholder, etc.), the credit data system 102 can quickly generate any requested credit data 118 (e.g., a specific transaction, a credit report, a credit score, a custom credit attribute of a specific requesting entity, etc.) based on the target user’s updated credit event data.

[0066] Furthermore, the credit data system can implement a batch indexing process. Introducing a batch indexing process eliminates the need for ETL on data from different credit events to conform to the requirements of a specific database or data structure, thus reducing or even eliminating bottlenecks related to ETL for credit events. As will be described in further detail in this application, the batch indexing process utilizes the components of the credit data system 102: indexing engine 104, identification engine 106, event cache engine 108, classification engine 110, and / or credit data storage 112. Indexing engine 104 can assign hash values ​​to unique identifiers (see...). Figure 4-10(To be described in further detail), so as to facilitate "clustering" of similar unique identifiers. Identification engine 106 can apply matching rules to resolve any issues related to the "clustered" unique identifiers, thereby generating a subset containing only empirically verified unique identifiers associated with the user. Classification engine 110 can merge the subsets into groups of unique identifiers associated with the same user. Event caching engine 108 can generate reverse personal identifiers ("reverse PIDs") and associate each unique identifier in the group with a reverse PID. Credit data system 102 can store the association between the reverse PID and the unique identifier as a reverse PID mapping in credit data storage 112 or any other accessible data storage. Then, using the reverse PID mapping, credit data system 102 can stamp the user associated with a credit event containing any unique identifier in the group. Credit data system 102 can store stamp associations 140 related to credit events 122A-122N belonging to the user in a flat file or database. (See reference...) Figure 4-10 Each component and its internal workings are described in further detail.

[0067] Invariant handling of heterogeneous credit events

[0068] Figure 1B Exemplary generation, processes, and storage of heterogeneous credit events are illustrated according to some embodiments. User 120 transacts with one or more business entities 124A-124N (such as merchants). Transactions may include purchases, sales, borrowing, loans, etc., and transactions may generate credit events. For example, user 120 uses a credit card to purchase an item, generating credit transaction data, which is collected by financial institutions 126A-126B (such as VISA, MasterCard, American Express, banks, mortgage lenders, etc.). Financial institutions 126A-B may share these transactions as credit events 122A-122N with credit data storage 112.

[0069] Each credit event 122A-122N may contain one or more unique identifiers that associate the credit event 122A-122N with the specific user 120 that generated the credit event 122A-122N. Unique identifiers may include various user identification information, such as name (first name, middle name, last name, and / or full name), address, Social Security number (“SSN”), bank account information, email address, telephone number, country ID (passport or driver's license), etc. Unique identifiers may also include partial names, partial addresses, partial telephone numbers, partial country IDs, etc. When financial institutions 126A-126B provide credit events 122A-122N for collection and analysis by a credit data system, the credit event can typically be identified as associated with a specific user through a combination of user identification information. For example, multiple people may share the same first and last name (consider “James Smith”), so first and last names may be too inclusive of other users' credit events. However, combinations of user identification information, such as full name plus telephone number, can achieve identification that meets the required needs. While each financial institution 126 may provide credit events 122A-122N in different formats, a credit event may include user identification information or a combination of user identification information, which can be used to associate the credit event with a specific user. This user identification information or combination of user identification information forms a unique identifier for the user. Therefore, multiple unique identifiers can be associated with a specific user.

[0070] The credit data system can operate using heterogeneous credit events 122A-122N, which have different data structures and provide distinct unique identifiers. For example, a credit event from a mortgage financial institution may include an SSN and a country ID, while a credit event from VISA may include a name and address, but not an SSN or country ID. Instead of performing ETL on credit events 122A-122N to standardize them for storage in the credit data store 112, the credit data system can perform a batch indexing process (as described in the following reference). Figure 4-10 (Detailed description) A reverse PID is generated for the set of unique identifiers that may be associated with user 120. The reverse PID can be assigned to credit events 122A-122N.

[0071] As will be described in further detail, the batch indexing process reduces or eliminates the significant computational resource overhead associated with heterogeneous ETL formats, significantly reducing processing overhead. Furthermore, the beneficial effect of assigning inverse PIDs to credit events is that once the correct inverse PID is assigned to a credit event, the credit data system 102 no longer needs to manage credit events based on the unique identifier contained within the credit event. In other words, once the credit data system 102 has identified the user associated with a credit event, it no longer needs to perform a search operation to find a unique identifier in credit events 122A-122N, but simply looks up the inverse PID assigned to the user in credit events 122A-122N. For example, in response to receiving a credit data request 118 from an external entity 116 (such as a financial institution, lender, potential landholder, etc.), a credit data system with a batch indexing process can use the user's inverse PID to quickly compile a list of credit events for user 120 and can provide any requested credit data 114 almost instantly.

[0072] Example of sequential processing of heterogeneous event sets

[0073] Figure 2A-2B An example of sequential processing of a heterogeneous set of events is illustrated. The credit data system can receive raw credit events from a high-throughput data source 202 via a high-throughput ingestion process 204. The credit data system can then store the raw credit events in a data store 206. The credit data system can perform a high-throughput cleaning process 208 on the raw credit events. The credit data system can then generate normalized clean events and store them in a data store 210. The credit data system can perform a high-throughput identifier resolution and keying process 212. The credit data system can store the identified events with keying in a data store 214. The identified credit events can then be categorized in process 216 and stored in an event set data store 218.

[0074] The credit data system can also generate local views in process 220. In process 220, the credit data system can load a set of user events associated with the user (identified events in data store 214, which have optionally been classified by classification process 216) from event set data store 218 into memory in process 222. The system can then calculate attribute 224, score model 226, and generate nested local views 228. The credit data system can then store the attribute calculations in analysis data (columnar) store 230. The analysis data can be used in application 234 to generate credit scores for the user. The nested local views can be stored in credit status (KV container) data store 232. Data in the credit status data store can be used in data manager application process 236 and credit query service 238.

[0075] During sequential processing, credit events can remain in the same state as they are sent to the credit data system by financial institutions. Financial data can also remain unchanged.

[0076] Example of creating credit status and credit attributes for analysis.

[0077] Figure 3 An exemplary credit data structure for simultaneously creating credit status and credit attributes for analysis is illustrated. The data structure can be virtually divided into three interaction layers: a batch processing layer 302, a service layer 320, and a speed layer 340. In the batch processing layer 302, a high-throughput data source 304 sends raw credit events to a data storage 310 via a high-throughput ingestion process 306. The credit data system can correct and PID-stamp the raw credit events 312 and store the corrected credit events in the data storage 314. The credit data system can then pre-compute 316 the corrected credit events associated with each user to generate a credit status and store each user's credit status in the data storage 322. The credit data system can store all credit attributes associated with each user in the data storage 324. The credit attributes associated with a user can then be accessed by various credit applications 326.

[0078] In velocity layer 340, various high-frequency data sources 342 can send new credit events to the credit data system via a high-frequency ingestion process 344. The credit data system can perform a low-latency correction process 348 and then store the new credit events associated with each user in data storage 350. New credit events associated with a user may cause a change in the user's credit status. The new credit status can be stored in data storage 328. The credit data system can then execute a credit profile lookup service process 330 to look up watermarks, thereby finding the stored credit status associated with the user. In some embodiments, the event caching engine is configured to allow the inclusion of recent credit events, even those not yet recorded in the user's complete credit status, in the credit attributes provided to third-party requesters. For example, when event data is added to the credit data storage (e.g., which may take hours or even days to complete), events stored in the new credit event data storage 350 can store the most recent credit events and be accessed upon receiving a credit query. Therefore, the requested report / rating can include credit events within milliseconds of receiving an event from a creditor.

[0079] The credit data system can use various local applications 332 to calculate credit scores or generate credit reports for users based on new credit status. Additionally, the credit data system can send instructions to the high-frequency ingestion process 344 via the high-frequency message channel 352. New credit events can then be sent again by the high-frequency ingestion process 344 to the file writer process 346. The credit data system can then store the new credit events in an event batch 308. Finally, the new credit events can be stored in the data storage 310 via the high-throughput ingestion process 306.

[0080] Credit data systems can store credit events in their raw form, generate credit status based on these events, and calculate attributes for users. When financial institutions send new credit events or detect errors in existing credit events, the credit data system can perform a credit profile lookup service to change the credit status or merge the credit status with real-time events. The credit data system can generate updated credit profiles based on the updated credit status.

[0081] Simultaneous creation of credit status and credit attributes allows for monitoring changes in a user's credit status and updating credit attributes upon detection of a change. Changes in a user's credit status can be caused by new credit events or errors detected in existing credit events. Credit events can remain at least partially identical because the credit data system does not extract, transform, or load data into a database. If the credit data system subsequently detects an invalid event, it can simply exclude the invalid event from future creations. Therefore, with the help of the credit data system, real-time reporting of events can be reflected in the user profile within minutes.

[0082] Example of a batch indexing process

[0083] Figure 4 A bulk indexing process according to some embodiments is illustrated, which includes the following steps: identifier stripping 402, identifier matching 410, and identifier stamping 440. The bulk indexing process can be a particularly powerful process for identifying and grouping different unique identifiers of a user (e.g., a credit event from VISA with an expired phone number can be grouped with a credit event from American Express with a newer phone number). One advantage of grouping different identifiers is that the user's credit data can be accurate and complete. The bulk indexing process can make credit data systems more efficient and responsive.

[0084] The identifier stripping process 402 extracts identifier fields (e.g., SSN, country ID, phone number, email, etc.) from credit events. The credit data system can segment 404 credit events by different financial institutions (e.g., credit card providers or lenders) and / or accounts. The credit data system can then extract 406 identifier fields from the segmented credit events without modifying the credit events. The identifier stripping process 402 may include a specialized extraction process for each different credit event format provided by different financial institutions. In some embodiments, the identifier stripping process 402 may perform a deduplication process 406 to remove identical or substantially similar identifier fields before generating a unique identifier (which may be a combination of identifier fields) associated with the credit event. This process will refer to... Figure 5 Further detailed description.

[0085] In the identifier matching process 410, the credit data system can perform a process to reduce the dimensionality of the unique identifiers determined in the identifier stripping process 402. For example, the location-sensitive hashing process 412 can be such a process. Depending on the design of the hashing process, the location-sensitive hashing process can calculate hash values ​​(e.g., identifier hash value 414) that increase or decrease the probability of collision based on the similarity of the original hash keys (e.g., unique identifier 408). For example, a well-designed hashing process can take different but similar unique identifiers, such as “John Smith, 1983 / 08 / 24, 92833-2983” and “Jonathan Smith, 1983 / 08 / 24, 92833” (full name, birthday, and ZIP code), and summarize the different but similar unique identifiers into the same hash value. Based on the shared public hash value, the two unique identifiers may be grouped into a set as potential matching unique identifiers associated with the user (details of the hash-based grouping process will refer to...). Figure 6 (Further details)

[0086] However, because hash functions can lead to unexpected collisions, the set based on hash values ​​may contain false positives (e.g., incorrectly associating credit events unrelated to the user with the user. For example, one of John's unique identifiers has the same hash value as one of Jane's unique identifiers, and after hash association, they might be grouped into the same set of unique identifiers associated with Jane). The credit data system can apply matching rule application 416 to the set of unique identifiers to remove false positive unique identifiers from the set. Various matching rules can be designed to optimize the probability of detecting false positives. An exemplary matching rule could be "exact match for country ID only," which removes unique identifiers with unarchived country IDs from the set of unique identifiers associated with the user. Another matching rule could be "minimum match for both name and ZIP code," where the minimum can be determined by comparing the calculated score of the match between the name and ZIP code with a minimum threshold score. Once false positives are removed from each set, the resulting subset 418 of matched identifiers contains only verified unique identifiers.

[0087] In some embodiments, matching rules can be designed considering the confidence level of each user identifier. For example, a driver's license number from a vehicle registration authority (VNA) can be associated with a high confidence level, and it may only be necessary to check the driver's license number to obtain an exact match. On the other hand, ZIP codes have a lower confidence level. Furthermore, matching rules can be designed to consider the history of associations for a particular record. If the record comes from a long-established bank account, the matching rules may not require rigorous scrutiny. On the other hand, if the record comes from a newly opened account, more stringent matching rules may be needed to remove false positives (e.g., identifying records in the set that may be associated with another user). This process will refer to... Figure 7 Further details: Matching rules can be applied to some or all of a set. Similarly, some or all of a matching rule can be applied to a single set.

[0088] Then, the subset of unique identifiers 418 can be merged with other subsets containing other unique identifiers of the user. Each subset 418 contains only the unique identifiers used to correctly identify the user. However, because the dimensionality reduction process may introduce false negative classes, it cannot be guaranteed that subsets 418 will be summarized into the same hash value. Therefore, some unique identifiers associated with a user may be placed in different subsets 418 when grouped based on hash values. Through the set merging process 420, when the subsets have common unique identifiers, the credit data system can group the two subsets into a single group (e.g., matching identifiers 422) that contains all unique identifiers associated with a specific user.

[0089] The credit data system can then assign reverse PIDs to each unique identifier in the merged group. Based on these assignments, the credit data system can then create a 424-to-426 reverse PID mapping, where each reverse PID is associated with multiple unique identifiers in the group associated with a specific user. This process will refer to... Figure 9 Further detailed description.

[0090] In the exemplary identification stamping process 440, the reverse PID mapping 426 can be used to stamp the segmented credit events 404 to generate PID-stamped credit events 430. In some embodiments, reverse PID stamping leaves the credit events associated with the reverse PID unchanged. This process will refer to Figure 10 Further detailed description.

[0091] Example of label stripping

[0092] Figure 5 An example of an identifier stripping process according to some embodiments is shown. In some embodiments, the credit data system may “correct” heterogeneous credit events 510 received from various financial institutions (e.g., e1, e2, e3, e4, e5, ...). “Correction” can be considered a process of fixing obvious quality problems. For example, a street address may be “100 Main Street” or “100 Main St.”. The credit data system may identify obvious quality problems such as the absence of spaces between the street address number and the street name, and / or modify “St.” to read as “Street”, and vice versa. The correction process may intelligently fix some identified quality problems but not others. For example, while the address described above may be corrected, correcting the user's name may be undesirable. Truncating, replacing, or otherwise modifying the user's name may cause problems more serious than retaining the information as a whole. Therefore, in some embodiments, the credit data system may selectively correct credit events 502.

[0093] The credit data system can segment credit events 504 by different financial institutions and / or accounts. The credit data system can extract the identifier field from credit events 406, and can optionally perform a deduplication process to remove redundant identifier fields. The credit data system can then generate a unique identifier based on the extracted identifier field. The identifier stripping process begins with credit event 510 and extracts the unique identifier 512. Figure 5 In the example, credit events e1, e2, e3, e4, e5, ... 510 may contain records r1, r2, r3, r4, ... 512. Each record may contain some or all of a unique identifier.

[0094] Figure 5Describe the beneficial effects of the identifier stripping process. If there are 40 million people, each with 20 accounts, generating credit events over 10 years (each event occupying 1000 bytes), there are approximately 96 gigabytes of credit event data. On the other hand, when there are the same number of people with the same number of accounts, the identifier attributes of the credit events only occupy approximately 3.2 gigabytes. If the stripped unique identifier 408 (which contains 1 / 30th of the credit event data) can be used to establish a correct association between credit events and specific users, the credit data system significantly narrows the scope of data that needs to be analyzed to be associated with specific users. Therefore, the credit data system has significantly reduced the computational overhead of the subsequent identifier matching process.

[0095] Identifier matching: Examples of position-sensitive hashes

[0096] Figure 6 An exemplary process for reducing data dimensionality using a hash algorithm, according to some embodiments, is illustrated. Records containing unique identifiers (r1-r16) from the identifier stripping process are listed in rows, and different hash functions (h1-hk) are listed in columns. The table shown in rows and columns is for illustrative purposes only, and the process can be implemented in any reasonably applicable manner. Furthermore, the collision rate in the figure (i.e., applying the hash function to different records yields the same hash value) does not reflect the likelihood of collisions occurring when genuine credit events are involved.

[0097] Multiple hash functions (e.g., h1 602, h5 604, etc.) can be applied to each record (e.g., r1-r16) to generate hash values ​​(e.g., h1' 606, h5' 608, h1' 610, h1' 612, etc.). Here, each row-column combination represents applying the hash function in the column to the record in the row to generate the hash value of that row-column combination. For example, hash function h1 602 applied to unique identifier r2 620 produces hash value h1' 610.

[0098] In some embodiments, each hash function can be designed to control the probability of collision for a given record. For example, h1602 could be a hash function for finding similar names by colliding with other records that have similar names. On the other hand, h5604 could be a hash function for SSNs, where the probability of collision is lower than that of hash function h1 for finding similar names. Various hash functions can be designed to better control the probability of collision. One of the beneficial effects of the disclosed credit data system is its ability to replace or complement various hash functions. The credit data system does not require a specific type of hash function, but rather allows users (e.g., data engineers) to easily interact with different hash functions to experiment and design to improve the overall system. This advantage can be very significant. For example, when a data engineer wants to migrate the credit data system to another country that uses a different character set (such as Chinese or Korean), the data engineer can replace a hash function for the English alphabet with one that yields better results for Chinese or Korean characters. Furthermore, in cases where country ID formats differ, such as South Korea's SSN using 12 digits while the United States' SSN uses 9 digits, a hash function more suitable for 12 digits can be used instead of a 9-digit hash function.

[0099] Although Figure 6 The diagram shows unmodified records r1-r16, but some embodiments may preprocess the records to obtain modified records more suitable for a given hash function. For example, the first name in a record can be concatenated with the last name to form a temporary record for use by a hash function specifically designed for such modified records. Another example could be truncating a 9-digit SSN number to the last 4 digits before applying the hash function. Similarly, users can modify records to better control the likelihood and outcome of collisions.

[0100] Figure 6 The diagram illustrates a hash function h1 that generates two distinct hash values, h1' 606 and h1" 612. Records {r1, r2, r3, r4, and r5} are associated with hash value h1' 606, while records {r12, r13, r14, r15, and r16} are associated with hash value h1" 612. Records can be grouped into sets based on their association with specific hash values. For example, the diagram shows a set 630 of hash values ​​h1' and a set 632 of hash values ​​h1" containing related records. Similarly, Figure 6 Six groups of records are identified and represented based on the common hash values ​​associated with them. For example, the hash values ​​h1' 606 and h5' 608 shown for record r1, each record can be associated with multiple hash values, each corresponding to a hash function.

[0101] For reference Figure 4As described, records with a common hash value can be grouped (“clustered”) into a set. For example, records {r1, r2, r3, r4, r5} share a common hash value h1' and are grouped into set 630. Similarly, records {r2, r7, r15} share a common hash value h4' and are grouped into set 632. As shown in both groups, some records (e.g., r2) can be grouped into multiple sets, while some records are grouped into a single set.

[0102] This hash-based grouping can be an ultra-fast process, requiring minimal computational resources. Hash functions have low operational complexity and can compute hash values ​​for large amounts of data in a short time. By grouping similar records together into sets, the process of identifying which records are associated with a specific user is greatly simplified. In a sense, the scope of all credit events that need to be associated with a user is narrowed down to just the records in those sets.

[0103] However, as referenced Figure 4 As briefly mentioned, using hash functions and resulting hash values ​​to group records may not be ideal, as it may contain false positives. In some embodiments, the resulting sets may carry "potential matches," but these sets may contain records that have not yet been rigorously verified to be associated with a user. For example, a set 630 of records with a specific hash value h1', namely {r1, r2, r3, r4, r5}, may contain records that are included in set 630 not because they have similar unique identifiers but because they have common hash values.

[0104] The credit data system then uses a rigorous identifier resolution process (“matching rule application”) to remove such false positives from each set.

[0105] Identifier Matching: Examples of Matching Rules

[0106] Figure 7 An example of an identifier resolution process according to some embodiments is shown. (Refer to...) Figure 6Following the described grouping process, the credit data system can apply one or more identifier resolution rules (“matching rules”) to the record set to remove false positive records from the set. Various matching rules can be designed to optimize the likelihood of detecting false positives. An exemplary matching rule could be “exact match for country ID only,” which removes records from the set of potentially matching records associated with a user that were assigned to the same set due to having the same hash value but were found to have different country IDs when checked by the matching rule. Matching rules can be based on exact matches or similar matches. For example, matching rules could also include “exact match for country ID,” “minimum match for country ID and last name,” and “exact match for country ID and similar match for last name.”

[0107] In some embodiments, a matching rule may calculate one or more confidence scores and compare them to one or more associated thresholds. For example, a matching rule “minimum match of both name and ZIP code” may have a threshold score to determine the minimum match, and the matching rule may discard records whose calculated scores are below the threshold. Matching rules may examine a record’s identifier (e.g., name, country ID, age, date of birth, etc.), format, length, or other characteristics and / or attributes of the record. Some examples include:

[0108] • Content: Reject unless the country ID is an exact match.

[0109] • Content: If there is a minimum match for both the country ID and the last name, accept.

[0110] • Content: Accept if there is an exact match for the country ID and a similar match for the surname.

[0111] • Format: If the user identification information (such as SSN) does not contain 9 digits, reject.

[0112] • Length: If the length of the user identification information does not match the length of the associated archived user identification information, reject the application.

[0113] • Content, format, and length: If the driver's license does not begin with "CA" and is not followed by X digits, it will be rejected.

[0114] The matching rules can also be any other combination of these criteria.

[0115] The resulting subset 418, after applying the matching rules, contains either the same or fewer records compared to the original set. Figure 7 It shows Figure 6The original sets (e.g., 702 and 704) after the hash value grouping process, and the resulting subsets (e.g., 712 and 714) after applying the matching rules. For example, in their respective orders, the sets associated with h1', h2', h3', h4', h1", and h2' initially contained 5, 6, 4, 3, 5, and 3 records, respectively. After applying the matching rules, the resulting subsets contained 3, 3, 2, 2, 2, and 2 records, all of which were previously included in the original sets. Using the matching rules increases the confidence that all remaining records are associated with the user.

[0116] Identifier matching: Examples of set merging

[0117] Figure 8 An exemplary set merging process according to some embodiments is illustrated. As discussed regarding existing systems, users sometimes change their personally identifiable information. An example is provided whereby a user may not have updated the phone number associated with their mortgagor. If the user updates their phone number with a credit card provider such as VISA, the credit events reported by the mortgagor and VISA will contain different phone numbers, while other information will be the same. This irregularity presents a unique challenge to data analysts because although both credit events should be associated with a specific user, the associated unique identifiers may differ, so the hash function may not group them into the same set. When records containing unique identifiers are not grouped into the same set, matching rules cannot correct for false negative classes (records that should have been placed in the same set but were not). Therefore, it is necessary to identify such irregular records generated by the same user and correctly associate these records with the user. The set merging process provides an efficient solution to this problem.

[0118] exist Figure 7 Following the matching process, each subset of results contains records that can be associated with the user with high confidence. Figure 8 There are 6 such subsets. The first subset 802 contains {r1, r3, r5}, and the second subset 804 contains {r3, r5, r15}. These two subsets may have become separate subsets because none of the hash functions produced a common hash value.

[0119] Further examination of the first and second subsets reveals that both subsets contain at least one common record r3. Since each subset is associated with a unique user, all records within the same subset can also be associated with the same unique user. Logically, if two different subsets associated with a unique user have at least one common record, then both different subsets should be associated with that unique user, and these two different subsets can be merged into a single group containing all records from both subsets. Therefore, based on the common record r3, the first subset 802 and the second subset 804 are combined to produce an extended group containing the records of the two subsets (i.e., {r1, r3, r5, r15}) after the set merging process. Similarly, based on the common record r15, another subset 808 containing {r2, r15} can be merged into this extended group to form another extended group 820 containing {r1, r2, r3, r5, r15}. Similarly, based on the other subsets 806 and 810, another group 822 containing {r10, r12, r16} can be formed. After the set merging process is completed, the records of all resulting groups will be mutually exclusive. Each merged group can contain all records containing a unique identifier associated with the user.

[0120] Exemplary set merging process

[0121] The above-mentioned set merging can be achieved using various methods. When the number of records is in the millions or even billions, the speed of merging sets can be critical. Here, an efficient grouping method is described.

[0122] The grouping algorithm first simplifies each set to a second-order relation (i.e., pairs). Then, the algorithm groups the second-order relations by the leftmost record. Next, the algorithm reverses or rotates the second-order relations to produce additional pairs. Then, the algorithm again groups the second-order relations by the leftmost record. Similarly, the algorithm repeats these processes until all subsets are merged into the final group. Each final group can be associated with a user.

[0123] For the purpose of explanation, Figure 7 The subsets in the algorithm are applied after the matching rules. These subsets are:

[0124] {r1, r3, r5}, {r3, r5, r15}, {r10, r12}, {r2, r15}, {r12, r16}, and {r1, r3}.

[0125] Starting with a subset, generate record pairs from the subset (i.e., reduce each group to a second-order relation). For example, the first subset containing {r1, r3, r5} can generate pairs:

[0126] (r1, r3)

[0127] (r3, r5)

[0128] (r1, r5)

[0129] The second subset containing {r3, r5, r15} can generate pairs:

[0130] (r3, r5)

[0131] (r5, r15)

[0132] (r3, r15)

[0133] The third subset containing {r10, r12} can generate pairs:

[0134] (r10, r12)

[0135] The fourth subset containing {r2, r15} can generate pairs:

[0136] (r2, r15)

[0137] The fifth subset containing {r12, r16} can generate pairs:

[0138] (r12, r16)

[0139] The sixth subset containing {r1, r3} can generate pairs:

[0140] (r1, r3)

[0141] The example merge process can list all pairs. Since duplicates do not contain any additional information, they have been removed:

[0142] (r1, r3)

[0143] (r3, r5)

[0144] (r1, r5)

[0145] (r5, r15)

[0146] (r3, r15)

[0147] (r10, r12)

[0148] (r2, r15)

[0149] (r12, r16)

[0150] Rotate or reverse each pair:

[0151] (r1, r3)

[0152] (r3, r1)

[0153] (r3, r5)

[0154] (r5, r3)

[0155] (r1, r5)

[0156] (r5, r1)

[0157] (r5, r15)

[0158] (r15, r5)

[0159] (r3, r15)

[0160] (r15, r3)

[0161] (r10, r12)

[0162] (r12, r10)

[0163] (r2, r15)

[0164] (r15, r2)

[0165] (r12, r16)

[0166] (r16, r12)

[0167] Group by the first record, where the first record is common to all pairs:

[0168] {r1, r3, r5}

[0169] {r3, r1, r5, r15}

[0170] {r5, r3, r1, r15} - Repeat

[0171] {r15, r5, r3, r2}

[0172] {r10, r12}

[0173] {r12, r10, r16}

[0174] {r2, r15}

[0175] {r16, r12}

[0176] Another round of pair generation. Duplicates are not shown:

[0177] (r1, r3)

[0178] (r3, r5)

[0179] (r1, r5)

[0180] (r3, r15)

[0181] (r1, r15)

[0182] (r5, r15)

[0183] (r15, r5)

[0184] (r15, r3)

[0185] (r15, r2)

[0186] (r5, r2)

[0187] (r3, r2)

[0188] (r10, r12)

[0189] (r12, r10)

[0190] (r12, r16)

[0191] (r10, r16)

[0192] (r2, r15)

[0193] (r16, r12)

[0194] Rotate or reverse each pair: duplicates not shown:

[0195] (r1, r3)

[0196] (r3, r5)

[0197] (r1, r5)

[0198] (r3, r15)

[0199] (r1, r15)

[0200] (r15, r1)

[0201] (r5, r15)

[0202] (r5, r3)

[0203] (r5, r1)

[0204] (r3, r1)

[0205] (r15, r5)

[0206] (r15, r3)

[0207] (r15, r2)

[0208] (r5, r2)

[0209] (r2, r5)

[0210] (r3, r2)

[0211] (r2, r3)

[0212] (r10, r12)

[0213] (r12, r10)

[0214] (r12, r16)

[0215] (r10, r16)

[0216] (r16, r10)

[0217] (r2, r15)

[0218] (r16, r12)

[0219] Group by the leftmost record, where the first record is common to all pairs:

[0220] {r1, r3, r5, r15}

[0221] {r2, r3, r5, r15}

[0222] {r3, r1, r2, r5, r15}

[0223] {r5, r1, r2, r3, r15} - Repeat

[0224] {r10, r12, r16}

[0225] {r12, r10, r16} - Repeat

[0226] {r15, r1, r2, r3, r5} - Repeat

[0227] {r16, r10, r12} - Repeat

[0228] Another round of pair generation. Duplicates are not shown:

[0229] (r1, r3)

[0230] (r1, r5)

[0231] (r1, r15)

[0232] (r2, r3)

[0233] (r2, r5)

[0234] (r2, r15)

[0235] (r3, r5)

[0236] (r3, r15)

[0237] (r5, r15)

[0238] (r1, r2)

[0239] (r2, r1)

[0240] (r10, r12)

[0241] (r10, r16)

[0242] (r12, r16)

[0243] Rotate or reverse each pair. Duplicates are not shown.

[0244] (r1, r3)

[0245] (r3, r1)

[0246] (r1, r5)

[0247] (r5, r1)

[0248] (r1, r15)

[0249] (r16, r1)

[0250] (r2, r3)

[0251] (r3, r2)

[0252] (r2, r5)

[0253] (r5, r2)

[0254] (r2, r15)

[0255] (r15, r2)

[0256] (r3, r5)

[0257] (r5, r3)

[0258] (r3, r15)

[0259] (r15, r3)

[0260] (r5, r15)

[0261] (r15, r5)

[0262] (r1, r2)

[0263] (r2, r1)

[0264] (r10, r12)

[0265] (r12, r10)

[0266] (r10, r16)

[0267] (r16, r10)

[0268] (r12, r16)

[0269] (r16, r12)

[0270] Group by the leftmost record, where the first record is common to all pairs:

[0271] {r1, r2, r3, r5, r15}

[0272] {r3, r1, r2, r5, r15} - Repeat

[0273] {r5, r1, r2, r3, r15} - Repeat

[0274] {r10, r12, r16}

[0275] {r12, r10, r16} - Repeat

[0276] {r16, r10, r12} - Repeat

[0277] By repeating the exemplary process: (1) creating pairs, (2) rotating or reversing each pair, and (3) grouping by the leftmost record, the subsets are merged into... Figure 8 The result sets shown are {r1, r2, r3, r5, r15} and {r10, r12, r16}.

[0278] Example of creating an event's reverse PID and identifier stamp

[0279] Figure 9 An exemplary process for associating a reversed PID with an identifier, according to some embodiments, is illustrated. For each final group associated with a user, the credit data system can assign a reversed PID. The reversed PIDs can be generated sequentially by the credit data system. Figure 9Two final groups are provided: group 902 contains {r1, r2, r3, r5, r15}, and group 904 contains {r10, r12, r16}. An inverse PID p1 is assigned to the first group, and an inverse PID p2 is assigned to the second group. Each inverse PID is associated with all records contained within the assigned group.

[0280] The credit data system can create a reverse PID mapping 426, which contains the association between records and their reverse PIDs. The reverse PID mapping 426 can be stored as a flat file or in a structured database. Once the reverse PID mapping is generated, the credit data system can incrementally update the mapping 426. (See reference...) Figure 8 Each group represents a set of all records associated with a specific user (along with a unique identifier contained within each record). Therefore, whenever two records have the same inverse PID, the credit data system can determine that these records are associated with a specific user, regardless of any differences in the records. Inverse PIDs can be used to stamp credit events.

[0281] Figure 10 An example of the identification stamping process is shown. The credit data process can access event 404 segmented by lender and / or account and reverse PID mapping 426, and provide them as input to the stamping process 428 to generate PID stamping event 430 based on one or more unique identifiers contained in the associated record. The stamped credit event 430 can be stored in a data store.

[0282] From grouping similar records into potential matches using hash functions to merging sets to stamping credit events with inverse PIDs, credit data systems maximize grouping. Grouping narrows the scope of analyzed credit events and facilitates faster access to them in the future. By using intelligent grouping instead of computationally intensive searches, credit data systems improve efficiency by several orders of magnitude. For example, using inverse PIDs to retrieve credit events associated with a user and generate credit reports is 100x more efficient.

[0283] For ease of disclosure, Figure 11A-11D It shows data with specific details. Figures 6-8 An exemplary identifier matching process. Figure 11A An exemplary process for reducing data dimensionality using hashing algorithms is provided, applied to specific values ​​in tabular form. Figure 11AThe leftmost column 11102 of the table lists records r1-r16 contained within the credit event. For example, record r1 could be {"John Smith", "111-22-3443", "06 / 10 / 1970", "100 Connecticut Ave", "Washington DC", "20036"}, record r2 could be {"Jonah Smith", "221-11-4343", "06 / 10 / 1984", "100 Connecticut Ave", "YourTown DC", "20036"}, and so on.

[0284] These records contain user identification information (e.g., record r1 654 contains user identification information "John Smith" (name), "111-22-3443" (SSN), "06 / 10 / 1970" (birthday), "100 Connecticut Ave" (street address), "YourTown DC" (city and state), and "20036" (ZIP code). Extracting user identification information from credit events ( Figure 4 (406), and optionally deduplication. User identification information can provide a unique identifier, either alone or in combination, that can associate records and associated credit events with a specific user. As shown in the figure, a record may include a unique identifier.

[0285] Various financial institutions can provide more or less different user identification information. For example, VISA can only provide first and last names (see, for example, r1), while American Express can provide a middle name in addition to first and last names (see, for example, r15). Some financial institutions can provide credit events where one or more user identification information is missing in total, such as not providing a driver's license number (see, for example, r1-r16 do not include a driver's license number).

[0286] Although there is no limit to how many hash functions can be applied to a record, Figure 11AThree exemplary hash functions, h1 11104, h2 11106, and h3 11108, are shown. As mentioned above, each hash function can be designed for (i.e., increasing or decreasing the collision rate) different personal identifiers or combinations of personal identifiers. Additionally, although not required, the personal identifier can be preprocessed to generate a hash key to achieve the purpose of each hash function. For example, hash function h1 11104 uses the preprocessed hash key “SSN number, user's last name, month of birth, and day of birth”. Record r1 can be preprocessed to provide the hash key “21Smith0610”. Using the preprocessing of h1 11104, records r2, r3, r4, and r5 will also provide the same hash key “21Smith0610”. However, for hash function h1 11104, record r14 will provide a different hash key “47Smith0610”. Different hash keys may yield different hash values. For example, the same hash key "21Smith0610" for r1, r2, r3, r4, and r5 yields "KN00NKL", while the hash key "47Smith0610" yields some other hash values. Therefore, records sharing the same hash value "KN00NKL" (i.e., r1, r2, r3, r4, and r5) are grouped as potential matches according to the hash function h1 11104.

[0287] Hash function h2 11106 uses a different preprocessing, namely "SSN, birth month, birth day". Records r3, r5, and r15, based on the preprocessing of h2 11106, produce the hash key "111-22-34340610". Using hash function h2 11106, the hash key is calculated as "VB556NB". However, hash functions can lead to unexpected collisions (in other words, false positives). Unexpected collisions result in unexpected records in the potential matching set. For example, based on the preprocessing of hash function h2 11106, record r14 obtains the hash key "766-87-16420610", which is different from the hash key "111-22-34340610" associated with r3, r5, and r15, but is still calculated as the same hash value "VB556NB". Therefore, when records are associated based on the same hash value generated by a shared hash function, the potential set of records belonging to a certain user may unintentionally include records belonging to different users. As mentioned above, and will also be used... Figure 11B The specific examples illustrate that matching rules can help parse the identifiers of false positive records in each set.

[0288] Each hash function can produce more than one set of potential matching records. For example, Figure 11AThe example illustrates how hash functions compute two sets of hash values, “VB556NB” and “NH1772TT”. Each set of hash values ​​represents a set of potential matching records. According to this example, hash function h2 11106 produces the hash value “VB556NB” with a set of potential matching records {r3, r5, r14, r16}, and the hash value “NH1772TT” with a set of potential matching records {r8, r9, r10, r12}.

[0289] Figure 11B The set of potential matching records 11202, 11204, 11206, and 11208 based on the common hash value is shown. Figure 11A The potential matching record set 11202 associated with the hash value “KN00NKL” includes {r1, r2, r3, r4, r5}. Similarly, the potential matching record set 11204 associated with the hash value “VB556NB” includes {r3, r5, r14, r16}. The potential matching record set 11206 associated with the hash value “NH1772TT” includes {r8, r9, r10, r12}. Similarly, the potential matching record set 11208 associated with the hash value “BBGT77TG” includes {r12, r13, r14, r15, r16}.

[0290] Each set may contain false positives. For example, although the potential matching record set 11202 associated with the hash value “VB556NB” includes {r1, r2, r3, r4, r5}, r2 and r4 do not appear to belong to the set of records that should be associated with John (Frederick) Smith because r2 has a different “SSN and year of birth”, and r4 has a different “name, SSN, year of birth, address, city, state, and ZIP code”. Determining whether r1, r3, or r5 are false positives is difficult because the SSN and year of birth vary only slightly (two digits in the SSN are reversed or the birth years differ by only one year). Therefore, records r2 and r4 are likely false positives, while r1, r3, and r5 are true positives. Similarly, other sets may contain both true and false positives.

[0291] Figure 11C This demonstrates parsing using one or more matching rules. Figure 11B The identifier in the set (i.e., removing such false positives). See reference. Figure 7Various matching rules are disclosed. For example, applying a rule such as "exact match for last name, at most two SSNs reversed and birth years differing by less than 2 years" can successfully remove possible false positives from set 11302, thus providing a subset containing only {r1, r3, r5}. In some embodiments, records in the set may be compared with the user's archived data (e.g., verified user identification information). In some embodiments, records in the set may be compared with each other to first determine personal identifiers that are highly likely to be true positives, and then a matching rule is applied to the determined personal identifiers.

[0292] In some embodiments, matching rules can calculate a confidence score and compare it to a threshold to accept or reject records in a set. For example, set 11304 with the hash value “VB556NB” can use a rule that calculates a character match score for a name. Record r14 has the full name “Eric Frederick”, which best matches “John Frederick Smith” and / or 9 out of 18 characters of “John Smith Frederick” among other records in set 11304. Therefore, a score of 50% can be calculated and compared with a minimum match threshold, such as 70%, and the credit data system can reject r14 from set 11304. Other matching rules can be designed and applied to sets 11302, 11304, 11306, and 11308 to remove rejected records and generate subsets. In some embodiments, some or all of these matching rules can be applied to different sets 11302, 11304, 11306, and 11308. Figure 11C The diagram shows a subset containing {r1, r3, r5}, {r3, r5, r15}, {r10, r12}, and {r12, r16}.

[0293] Figure 11D It shows the Figure 11C The subsets 11302, 11304, 11306, and 11308 identified in the dataset are used to apply set merging rules, thereby providing the merged groups 11402 and 11404. Figure 11C Each subset 11302, 11304, 11306, and 11308 contains records that can be associated with the user with high confidence. After applying the matching rules, Figure 11C Four such subsets are provided. The first subset 11302 contains {r1, r3, r5}, and the second subset 11304 contains {r3, r5, r15}.

[0294] Further examination of the first and second subsets reveals that both subsets contain at least one common record r3. Since each subset is associated with a unique user, all records within the same subset can also be associated with the same unique user. Logically, if two different subsets associated with a unique user have at least one common record, then both different subsets should be associated with that unique user, and these two different subsets can be merged into a single group containing all records from both subsets. Therefore, based on the common record r3, the first subset 11302 and the second subset 11304 are combined to produce a group 11402 containing records from both subsets (i.e., {r1, r3, r5, r15}) after the set merging process. Similarly, based on the other subsets 11306 and 11308, another group 11404 containing {r10, r12, r16} can be formed. After the set merging process is complete, the records in all resulting groups will be mutually exclusive. Each merged group can contain all records containing a unique identifier associated with a user.

[0295] When about Figure 8 When the described algorithm is applied to the original subset:

[0296] {r1, r3, r5}, {r3, r5, r15}, {r10, r12}, and {r12, r16}

[0297] Starting with a subset, generate record pairs from the subset (i.e., reduce each group to a second-order relation). For example, the first subset containing {r3, r5, r15} can generate pairs:

[0298] (r1, r3)

[0299] (r1, r5)

[0300] (r3, r5)

[0301] The second subset containing {r3, r5, r15} can generate pairs:

[0302] (r3, r5)

[0303] (r3, r15)

[0304] (r5, r15)

[0305] The third subset containing {r10, r12} can generate pairs:

[0306] (r10, r12)

[0307] The fourth subset containing {r12, r16} can generate pairs:

[0308] (r12, r16)

[0309] The example merge process can list all pairs. Since duplicates do not contain any additional information, they have been removed:

[0310] (r1, r3)

[0311] (r1, r5)

[0312] (r3, r5)

[0313] (r3, r15)

[0314] (r5, r15)

[0315] (r10, r12)

[0316] (r12, r16)

[0317] Rotate or reverse each pair:

[0318] (r1, r3)

[0319] (r3, r1)

[0320] (r1, r5)

[0321] (r5, r1)

[0322] (r3, r5)

[0323] (r5, r3)

[0324] (r3, r15)

[0325] (r15, r3)

[0326] (r5, r15)

[0327] (r15, r5)

[0328] (r10, r12)

[0329] (r12, r10)

[0330] (r12, r16)

[0331] (r16, r12)

[0332] Group by the first record, where the first record is common to all pairs:

[0333] {r1, r3, r5}

[0334] {r3, r1, r5, r15}

[0335] {r5, r1, r3, r15}

[0336] {r10, r12}

[0337] {r12, r10, r16}

[0338] {r15, r3, r5}

[0339] {r16, r12}

[0340] Another round of pair generation. Duplicates are not shown:

[0341] (r3, r1)

[0342] (r3, r5)

[0343] (r3, r15)

[0344] (r1, r5)

[0345] (r1, r15)

[0346] (r5, r15)

[0347] (r10, r12)

[0348] (r16, r12)

[0349] (r10, r16)

[0350] Rotate or reverse each pair: duplicates not shown:

[0351] (r3, r1)

[0352] (r1, r3)

[0353] (r3, r5)

[0354] (r5, r3)

[0355] (r3, r15)

[0356] (r15, r3)

[0357] (r1, r5)

[0358] (r5, r1)

[0359] (r1, r15)

[0360] (r15, r1)

[0361] (r5, r15)

[0362] (r15, r5)

[0363] (r10, r12)

[0364] (r12, r10)

[0365] (r16, r12)

[0366] (r12, r16)

[0367] (r10, r16)

[0368] (r16, r10)

[0369] Group by the leftmost record, where the first record is common to all pairs:

[0370] {r1, r3, r5, r15}

[0371] {r3, r1, r5, r15} - Repeat

[0372] {r5, r1, r3, r15} - Repeat

[0373] {r10, r12, r16}

[0374] {r12, r10, r16} - Repeat

[0375] {r15, r1, r3, r5} - Repeat

[0376] {r16, r12, r10} - Repeat

[0377] After applying the set merging algorithm, the two groups {r1, r3, r5, r15} and {r10, r12, r16} each have records remaining that are mutually exclusive.

[0378] Figure 12 This is a flowchart 1200 illustrating a method for efficiently organizing large-scale heterogeneous data. The method shown is implemented by a computing system, which may be a credit data system. Method 1200 begins at block 1202, where the computing system receives multiple event information from one or more data sources. The event information data source may be a financial institution. In some embodiments, the event information may have a heterogeneous data structure between event information from the same financial institution and / or across multiple financial institutions. The event information contains at least one personally identifiable information (“identification field” or “identifier”) that associates the event information with an account holder associated with the account that generated the credit event. For example, credit event information (or simply “credit event”) may contain one or more identification fields that associate the credit event with a specific user who generated the credit event by executing a credit transaction.

[0379] A computer system can access multiple event information by directly accessing a storage device or data store that contains existing event information from a data source, or by obtaining event information in real time via a network.

[0380] In box 1204, the computer system can extract the identifying fields of the account holder included in the event information. Identifying field extraction can involve formatting, transformation, matching, syntax parsing, etc. Identifying fields can include SSN, name, address, ZIP code, phone number, email address, or anything that can be used individually or in combination to attribute event information to an account holder. For example, a name and address may be sufficient to identify an account holder. An SSN can also be used to identify an account holder. If the number of event messages is in the billions and received from many data sources using heterogeneous formats, some accounts may not provide certain identifying fields, and some identifying fields may contain mis-entered or incorrect information. Therefore, when processing large amounts of event information, it is important to consider combinations of identifying fields. For example, distinguishing account holders solely based on SSNs may lead to incorrect identification of the associated account holder when an SSN is mis-entered. By relying on other available identifying fields, such as name and address, an intelligent computer system can correctly attribute event information to the same user. Combinations of identifying fields can form unique identifiers used to attribute event information to the user associated with the event.

[0381] In box 1206, the computer system may optionally deduplicate unique identifiers to remove duplicates. For example, an event message, when extracted, may provide "John Smith," "555-55-5555" (SSN), "jsmith@email.com" (email address), and "333-3333-3333" (phone number). Another event message may also provide "John Smith," "555-55-5555" (SSN), "jsmith@email.com" (email address), and "333-3333-3333" (phone number). Since the unique identifiers of the two event messages are the same, they can be deduplicated. Because one of the unique identifiers is removed, the operation at box 1208 only applies to unique identifiers that are not duplicated.

[0382] In box 1208, a computer system can reduce the dimensionality of a unique identifier using a multi-dimensionality reduction process. The goal of this box is to “cluster” unique identifiers based on some similarity contained within them. An exemplary process that can be used to reduce the dimensionality of a unique identifier based on the contained similarity can be a position-sensitive hash function. The computer system can provide multiple such dimensionality reduction processes, each targeting one aspect of the similarity contained in the unique identifier, to provide multiple “clusters” of similar (and potentially belonging to a particular user) unique identifiers. When using a position-sensitive hash function, a unique identifier is associated with a hash value, where each applied hash function generates a hash value for a given unique identifier. Therefore, each unique identifier can be associated with the hash value of each hash function.

[0383] In box 1210, the computer system groups uniquely identified components into sets, at least in part based on the results of a dimensionality reduction function that have common values. The components are then grouped into sets. Figure 6 Describe it in detail at an abstract level and in Figure 11B Detailed description using specific sample values. For example, regarding... Figure 6 and Figure 11B As described, the result set contains potential matches and may also contain false positives.

[0384] At box 1212, the computer system applies one or more matching rules with criteria for removing false positives to each set of unique identifiers. After the matching rules are applied to remove false positives, the set may become a subset of its previous set before the matching rules were applied, including only verified unique identifiers.

[0385] In box 1214, the computer system merges subsets to obtain a set of unique identifiers. The set merging process includes identifying common unique identifiers within the subsets, and merging the subsets containing at least one common unique identifier when the computer system finds that identifier. Set merging is... Figure 8 Describe it in detail at an abstract level and in Figure 11D The text describes this in detail using specific sample values. Furthermore, examples of efficient methods for set merging are disclosed above. After set merging, the merged groups consist of mutually exclusive unique identifiers.

[0386] In box 1216, the computer system provides a unique reverse PID for each group. In a sense, the process recognizes that each group represents a unique account holder. In box 1218, the computer system assigns the reverse PID provided for each group to all unique identifiers contained within each associated group. In a sense, the process recognizes that each unique identifier found in the event information can identify the event information as belonging to a specific account holder associated with the reverse PID.

[0387] At box 1220, the computer system examines the event information to find a unique identifier, and when the unique identifier is found, it stamps the event information with the inverse PID associated with that unique identifier.

[0388] Ingestion and consumption of heterogeneous data sets (HDC)

[0389] When a system collects and analyzes large amounts of heterogeneous data, some incoming data may contain or cause "defects." Defects can be broadly defined as any factor that leads to software modifications or data transformation. For example, some financial institutions reporting credit events may provide non-standardized data, requiring significant ETL processing as part of data ingestion. During the ETL process, some defects may be introduced. One example could be telephone numbers using the format "(###)###-####" instead of "###.###.####". Another example is the contrast between European date formats and US date formats. Yet another example could be a defect introduced due to the adoption of daylight saving time. Therefore, these defects may be introduced due to software errors or a lack of design universality in the ETL process. Sometimes, human error can also be a factor leading to certain format defects. Therefore, existing systems are not adequately prepared to address the formation and handling of defects, leaving room for improvement.

[0390] Existing data integration methods, such as data warehouses and data marts, attempt to extract meaningful data items from incoming data and transform them into a standardized target data structure. Typically, as the number of data sources providing heterogeneous data grows, so does the scale and complexity of the software and engineering work required to transform or otherwise process the ever-growing collections of heterogeneous data. These systemic and human demands can grow to a point where the marginal effort to modify and maintain existing systems can introduce further defects. For example, merging new data sources and formats may require modifying the data structure of an existing system, which sometimes necessitates converting existing data from an older format to a new one. This conversion process introduces new defects. If these defects persist unnoticed for a long time, significant effort and cost must be incurred to eliminate their impact through further software modifications and data transformations. Ironically, such software modifications and data transformations can lead to defects.

[0391] The credit data system described in this article addresses the defect management problem by implementing what can be called "lazy interpretation" of the data, which will be referenced below. Figures 13A-13C The defect model is described in further detail.

[0392] Defect Model

[0393] Figure 13AThis is a general defect model 13100, which illustrates the probability of defects associated with data as it flows from data ingestion to data consumption (i.e., from left to right) across multiple system states. The system can have an associated “defect surface” 13102, which can be defined as the probability distribution of a given software component having defects based on its functional scope and design complexity. The height of the defect surface 13102 reflects the defect probability P(D) combined with the functional scope and design complexity. In other words, the height of the defect surface 13102 will be higher when the software's functional scope and design complexity are high, and lower when the software's functional scope and design complexity are low. The defect surface 13102 is mostly flat, indicating that the software's functional scope and design complexity do not change between states.

[0394] Figure 13A The concept of "defect leverage" is also illustrated. Defect leverage can be defined as the amount (or distance) by which a downstream software component may be affected by a given defect. Defect 13104, located near data ingestion, is farther downstream, therefore its defect leverage is greater than that of defect 13106, located near data consumption. Based on the defect probability and defect leverage, the defect torque can be calculated, which can be defined as:

[0395] Defect torque = Defect probability * Defect leverage effect.

[0396] Defect torque can be understood as the potential impact of a defect on the system. The integral sum of defect torques can quantify the expected value of the defect quantity in the system. Therefore, we aim to minimize the sum of defect torques.

[0397] Figure 13B A defect surface model 13200 of a system using an ETL process is shown. Reconstruction, transformation, and normalization (all of which can be part of the ETL process) are performed during early data ingestion. Furthermore, interpretation also occurs early ingestion to assist the ETL process. Understanding collection, as part of analysis and reporting, occurs at the end of the data stream, close to data consumption.

[0398] As mentioned above, the complexity of the ETL process increases when dealing with heterogeneous data sources. Therefore, Figure 13B The defect surface 13202 is shown to be higher near data ingestion (indicating higher functional range and software complexity) and lower near data consumption. The system exhibits the highest defect surface 13202 at the point of highest defect leverage (near data ingestion) and the lowest defect surface 13202 at the point of lowest defect leverage (near data consumption).

[0399] This type of high-to-low defect surface 13202 poses a problem when considering defect torque. Defect torque is defined as the product of defect probability and defect leverage, where the integral of the defect torque quantifies the expected value of the defect quantity in the system. In this existing system, because high values ​​are multiplied by high values ​​and low values ​​by low values, the integral sum of the products can be quite large. Therefore, the expected value of the defect quantity can also be quite large.

[0400] Figure 13C A defect surface model of a credit data system is illustrated. Unlike existing systems, the credit data system does not perform ETL processes (e.g., reconstruction, transformation, standardization, recording, etc.), but its processing can be limited to verification, correction (e.g., performing quality control), and matching / linking incoming data. The verification, correction, and matching / linking processes are not as complex as the software components of ETL processes and have a lower probability of defects. Therefore, Figure 13C The defect surface 13302 of the credit data system is shown to be lower near data ingestion and higher near data consumption. Correspondingly, the system exhibits the lowest defect surface 13202 at the point of highest defect leverage (near data ingestion) and the highest defect surface 13202 at the point of lowest defect leverage (near data consumption).

[0401] This type of low-to-high defect surface 13302 is highly advantageous when considering defect moments. In a credit data system, because low defect probability multiplies by high defect leverage, and high defect probability multiplies by low defect leverage, the integral sum of the products can be much smaller than in existing systems. Therefore, the credit data system improves defect management regarding data ingestion and data consumption.

[0402] Lazy interpretation of data

[0403] The "lazy interpretation" system does not interpret the incoming data near the data ingestion (e.g., Figure 13B Instead of using the data model 13200 shown for conventional systems, the interpretation is delayed as late as possible in the pipeline from data to understanding in order to minimize the overall defect torque. Figure 13C An exemplary defect model 13300 of such an inertial interpretation system according to one embodiment is shown.

[0404] Lazy interpretation systems can accept any type of event data, such as data from data sources with various data types, formats, structures, and meanings. For example, Figure 14Various types of event data related to anchor entity 1402 (shown as a specific user in this example) are illustrated. An anchor entity can be any other entity from which event data is provided for parsing. For example, an anchor entity can be a specific user, and various data sources can provide heterogeneous data events associated with that specific user, such as vehicle loan records 1404, mortgage records 1406, credit card records 1408, utility records 1410, DMV records 1412, court records 1414, tax records 1416, employment records 1418, etc.

[0405] In some embodiments, when accessing new event data, the system identifies only the minimum information required to bind the data to the correct anchor entity. For example, the anchor entity could be a specific user, and the minimum information required to bind new data to that user could be identifying information such as name, country ID, or address. When new data is received, the system can look up the minimum set of identifying information for a specific user within that data and bind that data to one or more user-associated tags (e.g., a reverse PID is an example of a user-associated tag when the anchor entity is a user associated with a credit event). For a given set of data, the lazy interpretation system can subsequently use the tags to identify the correct anchor entity. The process of tagging can be... Figure 13C The matching / linking process in the system. In some embodiments, the matching / linking process does not change the incoming data or data structure.

[0406] The tagging / matching / linking process can be similar to cataloging books. For example, based on the International Standard Book Number (“ISBN”), title, and / or author, librarians can place books in the correct sections and on the correct shelves. The content or plot of the book is not necessary in the cataloging process. Similarly, based on minimal information identifying the anchor entity, vehicle loan record 1404 can be associated with a specific anchor entity. In some embodiments, each record and / or data source can be associated with a domain (see reference). Figure 15 (Further description). For example, vehicle loan record 1404 or vehicle loan data source can be associated with the "Vehicle Loan Domain", credit card record 1408 or credit card data source can be associated with the "Credit Domain", and mortgage record 1406 or mortgage data source can be associated with the "Mortgage Domain".

[0407] In some embodiments, the lazy interpretation system may include an Anchored Entity Resolution (AER) process that corrects the labels bound to previously received data to associate them with a best-known anchored entity. The best-known anchored entity may change dynamically based on information contained in newly incoming data, such as analysis of previously received data or improvements based on the resolution of the anchored entity itself. In some embodiments, anchored entity resolution may update previously bound labels. The anchored entity resolution process may run periodically or continuously in the background or foreground, and may be automatically triggered when a predetermined event occurs and / or initiated by a system monitor, requesting entity, or other user.

[0408] Lazy interpretation systems limit the probability of defects to the interpretation and processing of identified information. By eliminating the ETL process of traditional systems, lazy interpretation systems reduce the software and engineering work required to transform or otherwise address the ever-increasing size and complexity of heterogeneous data sets. Figure 13C As shown, the upstream state of the defect surface 13302 near the data consumption is lower, thereby reducing the defect torque.

[0409] Domain dictionary and vocabulary

[0410] A lazy interpretation system may include one or more parsers for interpreting data. Figure 13C ,13304). Interpretation components of existing systems ( Figure 13B Unlike data ingestion, the interpreter component (e.g., the "parser") in a lazy interpreter system is located closer to data consumption. Figure 13C (13304). The resolver can be associated with domains such as credit domain 1502, utility domain 1504 and / or mortgage domain 1506.

[0411] A lazy interpretation system can associate incoming data or a data source with one or more domains. For example, credit card record 1408 or its data source might already be associated with the "Credit Domain". Each domain includes a dictionary containing the vocabulary for that domain. Figure 15The diagram illustrates the domains and their associated vocabularies. For example, the credit domain 1502 might have associated dictionaries including the terms "@credit_limit", "@current_balance", and "@past_due_balance". Similarly, the utility domain 1504 might have associated dictionaries including the terms "@current_balance" and "@pass_due_balance". As shown, terms can be repeated across different domains, such as "@current_balance" and "@past_due_balance". However, each domain has its own set of interpretation rules and parser associated with that specific domain, enabling appropriate interpretation of the same term in one domain based on the domain of each record, unlike the interpretation of the same term in another domain.

[0412] Based on a dictionary and its vocabulary, one or more parsers examine the content of records and tag fields or values ​​using matching terms. The parsing process can be analogous to scanning a book to identify / interpret relevant content. Similar to scanning a history book for content related to "George Washington" and tagging descriptions of George Washington's birthplace, birthday, age, etc., with "@George_Washington," the credit parser 1508 can scan records from a credit data source or within a credit domain and identify / interpret content that may be related to credit limits, tagging the identified / interpreted content with the "@Credit_Limit" label. Figure 16 An example is shown using @Credit_Limit to tag the identified content. Similarly, the utility domain resolver 1510 can scan records (such as utility invoices) from utility data sources or utility domains and identify content that may be related to overdue balances and tag the identified content with the "@past_due_balance" label.

[0413] Once tagged, including Figure 13C Downstream components for consistency checks, understanding, and / or reporting can use the vocabulary of the recorded domains to analyze the contents of the records. In some embodiments, downstream components (e.g., any understanding calculation component 1512) can interpret records from one or more domains for their use. For example, a mortgage scoring component can look up “@Credit_Limit” in data from the credit domain before determining the creditworthiness of a potential borrower.

[0414] Advantageously, lazy interpretation has the beneficial effect of reducing the impact of defects. The interpretation performed by the parser above, such as... Figure 13CAs shown in 13304, the interpretation is closer to data consumption than that provided by existing systems. Therefore, the shortcomings in the lazy interpretation system have limited leverage, and thus their impact is reduced.

[0415] Another beneficial effect of lazy interpretation systems is that the system does not require modification of the original or existing heterogeneous event data. Instead of ETL processing that standardizes data for storage and interpretation, the system tags and defers interpretation to the parser. If one or more parsers are found to introduce defects into the domain, data engineers can simply update those domain parsers. Because the original or existing event data remains unchanged, re-executing the parser can quickly eliminate defects without data loss. Furthermore, in some embodiments, because no data is copied throughout the data stream, data engineers can correct, delete, or exclude any data without updating other databases.

[0416] Therefore, lazy interpretation systems do not require an ETL process for data ingestion, thus allowing for the rapid and low-cost introduction of new data sources.

[0417] Figure 16 An exemplary process 1600 is illustrated, according to some embodiments, using lazy interpretation of some sample content. Domain dictionary 1602 may include domain vocabulary 1604 and domain syntax 1606. Domain vocabulary 1604 may include annotations (e.g., references...) Figure 15 The keyword definition of the described tagging data. Domain vocabulary 1604 may include "primitive words" and "compound words". In some embodiments, primitive words are tags that are directly associated with (or "annotations") some parts of heterogeneous data. For example, a lazy interpretation system tags some parts of incoming data 1610 with @CreditLimit and @Balance. Compound words are synthesized from one or more primitive words or other variables using domain syntax 1606. An example of domain syntax 1606 could be "The average balance of N records is equal to the sum of the balances of each account divided by N", which can be represented in domain syntax 1606 with two primitive words @Balance as "@AverageBalance[n]=Sum(@Balance) / n").

[0418] Domain dictionary 1602 may also include predefined source templates 1608 for heterogeneous data sources. Source template 1608 acts as a lens to expose important fields. For example, a simple exemplary source template might be "For incoming data 1610 from a VISA data source, the 6th data field is @CreditLimit, and the 7th data field is @Balance". Annotator 1612 can use one or more such source templates 1608 to tag incoming data in label / annotation fields to generate annotated data 1614. In some embodiments, machine learning models and / or other artificial intelligence can be used to supplement or replace source template 1608 in identifying and exposing important fields.

[0419] The lazy interpretation system may also include one or more domain resolvers 1616. Domain resolver 1616 may use annotations / tags and rules embedded in its software to present fully annotated data to the application. In some embodiments, in addition to or instead of the annotations / tags provided by annotation adder 1612, the domain resolver may provide annotations / tags to generate fully annotated data. Domain resolver 1616 may reference domain dictionary 1602 when presenting fully annotated data to the application or in its own annotations / tags.

[0420] Credit score calculation application 1618 and understanding calculation application 1620 are provided as exemplary applications that can use fully annotated data. Credit score calculation application 1618 can calculate credit scores (or other scores) for one or more users based on the annotated data and provide them to the requesting entity. Similarly, understanding calculation application 1620 can provide analytics or reports, including balance statements, cash flow statements, spending habits, possible savings tips, etc. In some embodiments, various applications, including credit score calculation application 1618 and understanding calculation application 1620, can use fully annotated data combined with a reverse PID from a batch indexing process to quickly identify all annotated records belonging to a specific user and generate reports or analytics related to that user.

[0421] Figure 17 This is a flowchart 1700 illustrating a method for interpreting incoming data to minimize the impact of defects in a system, according to some embodiments. According to the embodiments, Figure 17 The method may include fewer or more boxes, and these boxes may be executed in an order different from the order shown.

[0422] Beginning with box 1702, the interpreter system (e.g., one or more components of a credit data system discussed elsewhere in this document) receives multiple event messages from one or more data sources (see [link to document]). Figure 14The data source can be a mortgagor, credit card provider, utility company, vehicle dealership providing vehicle loan records, DMV, court, IRS, employer, bank, or any other information source that can be associated with the entity requiring entity resolution. In some embodiments, the data source provides multiple event information in heterogeneous data formats or structures.

[0423] In box 1704, the lazy interpretation system determines the category or type (also referred to herein as a “domain”) of information associated with the data source. The domain of the data source can be determined based on information provided by the data source. In some embodiments, the system is able to determine (or confirm, where the data source provides domain information) the associated domain by examining the data structure of the data source. In some embodiments, event information may include hints indicating the domain of a particular data source, and the system is able to determine the domain of the data source based on these hints. For example, if the event information (or most of the event information) includes the terms “water” or “gas”, the system may automatically determine that the data source should be associated with a utility domain.

[0424] In box 1706, the system accesses the domain dictionary for the determined domain. The domain dictionary may include a domain vocabulary, domain grammar, and / or annotation standards, examples of which are referenced above. Figure 16 describe.

[0425] In box 1708, the system annotates event information from identified domains using a domain dictionary. For example, based on annotation criteria, the system evaluates the event information and identifies one or more portions of the available domain vocabulary for annotation. Figure 16 An exemplary event information 1610 prior to annotation is shown, followed by an annotated event information 1614 with annotations associated with that event information. In some embodiments, event information is updated only by domain annotations (such as those in the exemplary annotated event information 1614) and otherwise remains unchanged. In some embodiments, once event information is annotated, it remains undisturbed until the system receives a data request for the event information, such as information associated with a specific annotation (e.g., a request for @Creditlimit data for event information might be to calculate a consumer's total credit limit across multiple accounts, which could be included in a credit report or similar consumer risk analysis report).

[0426] In box 1710, the system receives a data request for event information. The request can be for event information (e.g., all event information including a specific comment or combination of comments) or for specific data included within the event information (e.g., a portion of event information specifically associated with a comment). For example, see reference... Figure 16The annotated event information 1614 allows requests to be made for the entire annotated credit event information or only for the @Balance data within the credit event information. Data requests can originate from another component of the system, such as a score calculation application or an understanding calculation application, or from another requesting entity, such as a third party.

[0427] In box 1712, the system uses one or more domain resolvers to analyze the event information to identify the requested information. (See reference...) Figure 16 As described, a domain resolver can use a domain dictionary to interpret event information. For example, a domain resolver can use a domain vocabulary to look up one or more primitive words. The domain resolver can then use a domain grammar to determine compound words based on the one or more primitive words. In some embodiments, a domain resolver can request another domain resolver to provide the necessary data for its interpretation. For example, a mortgage domain resolver can request @credit_score from a credit domain resolver when generating its compound words, based on the domain grammar of the required credit score. In box 1714, the system provides the requested data to the requesting application or requesting entity.

[0428] Additional Examples

[0429] It should be understood that not all objectives or advantages may be achieved according to any particular embodiment described herein. Therefore, for example, some embodiments may be configured to operate in a manner that achieves or optimizes one or more advantages taught herein, without necessarily achieving other objectives or advantages taught or suggested herein.

[0430] All processes described herein can be implemented and fully automated by software code modules executed by a computing system comprising one or more computers or processors. In some embodiments, at least some processes can be implemented individually or in combination using virtualization technologies such as cloud computing, application containerization, or Lambda architecture. The code modules can be stored on any type of non-transitory computer-readable medium or other computer storage device. Some or all of the methods can be implemented in dedicated computer hardware.

[0431] Based on this disclosure, many other variations besides those described herein will be apparent. For example, depending on the embodiment, certain actions, events, or functions of any algorithm described herein may be performed in a different order, or may be added together, combined, or omitted (e.g., not all described actions or events are necessary for the practice of the algorithm). Furthermore, in some embodiments, actions or events may be performed concurrently, for example, through multithreaded processing, interrupt handling, or multiple processors or processor cores, or on other parallel architectures, rather than sequentially. Additionally, different tasks or processes may be performed by different machines and / or computing systems capable of working together.

[0432] Unless otherwise specifically stated, conditional language, such as in particular “can,” “could,” “might,” or “may,” is understood in the context to generally convey certain features, elements, and / or processes included in some embodiments but not in others. Therefore, such conditional language is not generally intended to imply that features, elements, and / or processes are necessary in any way for one or more embodiments, or that one or more embodiments must include logic for determining, with or without user input or prompting, whether such features, elements, and / or processes are included in or will be performed in any particular embodiment.

[0433] Unless otherwise specifically stated, disjunctive language such as the phrase “at least one of X, Y, or Z” is understood in the context to generally indicate that an item, term, etc., can be X, Y, or Z or any combination thereof (e.g., X, Y, and / or Z). Therefore, such disjunctive language is generally not intended and should not imply that some embodiments require the presence of at least one X, at least one Y, or at least one Z.

[0434] Any process description, element, or block described herein and / or shown in the flowcharts in the accompanying drawings should be understood as potentially representing a code module, segment, or portion comprising one or more executable instructions for implementing a particular logical function or element in the process. Alternative implementations are included within the scope of the embodiments described herein, wherein elements or functions are omitted or performed in a different order than that shown or discussed (including substantially simultaneously or in reverse order), depending on the functionality involved, as will be understood by those skilled in the art.

[0435] Unless otherwise explicitly stated, articles such as “a” or “an” should generally be interpreted as including one or more of the described items. Therefore, phrases such as “devices configured to…” are intended to include one or more of the described devices. Such one or more enumerated devices may also be collectively configured to perform the stated descriptions. For example, “processors configured to perform descriptions A, B, and C” could include a first processor configured to perform description A, which works in conjunction with a second processor configured to perform descriptions B and C.

[0436] It should be emphasized that many variations and modifications can be made to the above embodiments, and their elements should be understood as one of other acceptable examples. All such modifications and variations are intended to be included within the scope of this disclosure.

Claims

1. A computer system for determining an account holder identifier for collected event information, the computer system comprising: One or more hardware computer processors; as well as One or more storage devices configured to store software instructions configured to be executed by the one or more hardware computer processors to cause the computer system to: Receive multiple event information associated with multiple corresponding events from multiple data sources; For each event message: Access includes data storage associated with a data source and an identifier parameter, the identifier parameter including at least an indication of one or more identifiers included in event information from the corresponding data source; Based at least on the identifier parameter of the data source of the event information, determine the identifier included in the event information as indicated in the accessed data storage; and Identifiers are extracted from the event information based at least on corresponding identifier parameters, wherein the combination of identifiers includes a unique identifier associated with a unique user. Access multiple hash functions, each of which is associated with a combination of identifiers; For each unique identifier, multiple hash values ​​are calculated by evaluating the multiple hash functions; Based on whether the unique identifiers share the common hash value calculated using the common hash function, the unique identifiers are selectively grouped into a set of unique identifiers associated with the common hash value; For each set of unique identifiers: Apply one or more matching rules, said one or more matching rules including criteria for comparing unique identifiers within said set; and The unique identifiers that satisfy one or more matching rules are determined as the unique identifier matching set; By repeatedly creating record pairs from each unique identifier matching set, reversing each record pair, and grouping by the leftmost record, wherein the leftmost record is common between the record pairs, until the unique identifier matching sets are merged, wherein each merged set is associated with a user, thereby merging unique identifier matching sets that each include at least one common unique identifier to provide one or more merged sets that do not have a common unique identifier with other merged sets. For each merged set: Determine the reverse personal identifier; and Associate the reverse personal identifier with each of the unique identifiers in the merged set to create a reverse personal identifier mapping; For each unique identifier, use the reverse personal identifier mapping: Identify event information associated with at least one of the combinations of identifiers associated with the unique identifier; and The reverse personal identifier is associated with the identified event information, wherein each reverse personal identifier is associated with multiple unique identifiers in the merged set, and the merged set is associated with the unique user.

2. The computer system as claimed in claim 1, wherein, The hash function includes at least: A first hash function evaluates a first combination of at least a portion of a first identifier and at least a portion of a second identifier extracted from the event information; and The second hash function evaluates a second combination of at least a portion of the first identifier and at least a portion of the third identifier extracted from the event information.

3. The computer system as described in claim 2, wherein, The first hash function is selected based on the identifier type of one or more of the first identifier or the second identifier.

4. The computer system as described in claim 2, wherein, The first identifier is the Social Security number of the unique user, the second identifier is the last name of the unique user, and the first combination is a concatenation of fewer than all the digits of the Social Security number and fewer than all the characters of the last name of the unique user.

5. The computer system as described in claim 2, wherein, The first event set includes multiple events associated with the first hash value, and the second event set includes multiple events, each associated with the second hash value.

6. The computer system as claimed in claim 1, wherein, The identifier is selected from: first name, last name, first letter of middle name, middle name, date of birth, social security number, taxpayer ID, or country ID.

7. The computer system as claimed in claim 1, wherein, The computer system generates a reverse mapping that associates the reverse personal identifier with each of the remaining unique identifiers in the merged set, and stores the mapping in a data storage.

8. The computer system of claim 7, further comprising: Based on the reverse personal identifier assigned to the remaining unique identifier, the reverse personal identifier is assigned to each of the plurality of event information including the remaining unique identifier.

9. The computer system as claimed in claim 1, wherein, The hash function includes position-sensitive hashing.

10. The computer system as claimed in claim 1, wherein, The one or more matching rules include one or more identifier resolution rules, which compare u in one or more sets with account holder information in an external database or CRM system to identify a match for the one or more matching rules.

11. The computer system of claim 10, wherein, The identifier resolution rules include criteria that indicate the matching standards between the account holder information and the identifier.

12. The computer system as claimed in claim 1, wherein, Creating a record pair from each unique identifier matching set involves pairing each unique identifier in the set with another unique identifier in the set to create a unique identifier pair; The records are grouped by the leftmost record, which is common to the record pairs, including grouping non-common unique identifiers from unique identifier pairs that have the common unique identifier, until the unique identifier lists contained in the result group are mutually exclusive between the result groups.

13. The computer system of claim 12, wherein, The process also includes classifying the unique identifiers in the unique identifier pair.

14. The computer system as claimed in claim 1, wherein, The multiple event information includes information about at least one credit event.

Citation Information

Patent Citations

  • Information search system, information management device, information search method, information management method, and recording medium

    US20120109990A1

  • User similarity groups for on-line marketing

    US20150287091A1