Risk assessment techniques based on dynamic data selection
By generating data samples and calculating metrics for each data source, the system addresses the challenge of selecting high-quality and non-redundant data sources, enhancing machine learning accuracy and reducing costs.
Patent Information
- Application Number
- PCT/US2024/037767
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-12
- Publication Date
- 2026-01-15
AI Technical Summary
Existing systems face challenges in selecting high-quality and non-redundant data sources for machine learning, leading to inaccurate predictions and increased costs due to redundant and low-quality data from multiple external vendors.
A system that generates data samples from each data source, including random, fraud, and synthetic samples to calculate metrics such as coverage, fraud detection, and trust, and selects the best data source based on these metrics to improve data quality and reduce redundancy.
This approach enhances the accuracy of machine learning models by selecting robust data sources, reducing redundant data, and improving predictive power while minimizing costs.
Smart Images

Figure IMGF000018_0001 
Figure IMGF000018_0002 
Figure IMGF000019_0001
Abstract
Description
RISK ASSESSMENT TECHNIQUES BASED ON DYNAMIC DATA SELECTIONTECHNICAL FIELD
[0001] The present disclosure relates generally to controlling interactions between computing systems. More specifically, but not by way of limitation, this disclosure relates to systems and methods for evaluating and selecting data sets based on one or more data source scoring metrics.BACKGROUND
[0002] Various systems rely on data from a number of sources to make operational decisions or determine risk. For example, a system can use internal data, as well as data from a number of external vendors or sources in determining risk for a target entity. But the internal data and external data from multiple sources can overlap or be redundant. Additionally, different external data sources can provide differing coverage and differing quality. This leads entities managing systems relying on such external data to acquire redundant, and possibly low quality, data.SUMMARY
[0003] Various aspects of the present disclosure provide systems and methods for data source selection. The system can receive a data selection request from a remote computing device and data from a set of data sources. In some aspects, for each data source in a set of data sources, the system can: generate a random data sample, a fraud data sample, and a synthetic data sample; generate a set of metrics based on the random data sample, the fraud data sample, and the synthetic data sample; and based on the set of metrics, determine a data source score. In some aspects, the random data sample can include a subset of data from the respective data source, the fraud data sample can include a set of labelled data from the respective data source; and the synthetic data sample can include randomized data from the respective data source. In some aspects, the system can select, a data source from the set of data sources based at least in part on the data source score. The system can transmit, to a remote computing device, a responsive message including at least an indication of the selected data source.
[0004] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification, any or all drawings, and each claim.
[0005] The foregoing, together with other features and examples, will become more apparent upon referring to the following specification, claims, and accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS
[0006] FIG. 1 is a block diagram depicting an example of an operating environment in which a computing system can be used to evaluate data from a number of sources according to some aspects of the present disclosure.
[0007] FIG. 2 is a block diagram depicting a system for generating a risk assessment associated with a target entity according to some aspects of the present disclosure.
[0008] FIG. 3 is a flow chart illustrating a method for generating a risk assessment associated with a target entity according to some aspects of the present disclosure.
[0009] FIG. 4 is a block diagram depicting an example of a computing device, which can be used to implement the embodiments described herein according to some aspects of the present disclosure.DETAILED DESCRIPTION OF THE INVENTION
[0010] Systems and methods described herein relate to selecting data sets based on one or more data source scoring metrics. The accuracy of machine learning outcomes and predictions relies, in part, on the fidelity and quality of the data on which a machine learning model is trained. Thus, systems reliant on machine learning outcomes are vulnerable to inaccuracy and poor predictive power if the machine learning model is not trained on high-quality and robust data. Accordingly, systems and methods described herein enable dynamic selection of a data set based on one or more data source scoring metrics. For example, a data source can be evaluated or scored based on coverage, risk, verification, affiliation, trust, demographic consistency, concurrence with other data sources, fraud detection, consistency, and sensitivity. In some examples, one or more of these metrics can be combined to determine a data source effectiveness score that can be used to evaluate a data set provided by a data source.
[0011] Certain aspects described herein for selecting data sets based on one or more data source scoring metrics can address issues associated with selecting an external data source or data vendor. For example, systems and methods disclosed herein can determine the coverage of a data source to determine an amount of overlap with available internal data. This metric can be used to avoid acquiring redundant data. In another example, a fraud score can measure a data vendor’s ability to correctly identify fraudulent data in a data set. This metric can be used to select a data source from a vendor that accurately flags fraudulent data. Such a metric may be important, for example, in applications relying on a machine learning model to be accurately trained to identify fraudulent activity. By providing these and other metrics, the systems and methods described herein can provide an explorable set of data source scores that facilitate more data selection, thereby improving outcomes of machine learning models trained on the selected data. This can improve an entity’s ability to, for example, prevent fraudulent activities and enhance security in online environments and of online interactions by relying on machine learning models trained to accurately identify risk.
[0012] In some examples, a computing system can receive a data selection request from a remote computing device. The computing system can also access data from a set of data sources. The set of data sources may contain data accessible to the remote computing device, or may contain data provided by an entity managing the remote computing device. The request can include, in some examples, a set of thresholds associated with available data source scoring metrics. For example, depending on the intended application of the selected data set, the remote computing device may define different thresholds for each available data source scoring metric. In some examples, the request can specify the set of data sources from which a data source is to be selected. In other examples, the computing system may determine a selected data source from among a predetermined set of data sources.
[0013] The computing system can generate data samples for each data source of the set of data sources. The data samples can include: a data sample including a subset of data from a data source; a fraud data sample including a set of labelled data from the data source; and a synthetic data sample including randomized data from the data source. The data sample can be a randomly selected subset of data from the data source. The fraud data sample can be a randomly selected subset of data from the data source, where the fraud data has been analyzed by the computing system or the remote computing device, and fraudulent records have been labelled as fraudulentor potentially fraudulent. The synthetic data sample can be, in some cases, based on the data sample, or it can be based on a randomly selected subset of data from the data source. The synthetic data can be generated by randomizing the subset of data.
[0014] The three generated data sets can be used, for example, to generate a set of metrics for the respective data source. The metrics can include: a coverage score; a verification score; an affiliation score; a risk score; a trust score; a fraud score; a concurrence score; a demographic score; a sensitivity score; and a consistency score. In some examples, one or more of these scores can be combined to determine a data source score for the respective data source.
[0015] Disclosed systems and methods can use the data source score to select a data source from the set of data sources. In some examples, data sources of the set of data sources can be ranked from highest to lowest based on the associated data source score, and the data source having the highest data source score can be selected. In another example, a subset of data sources having data source scores above a score threshold. In some aspects, the computing system can transmit one or more of the metrics and the data source score associated with each data source to the remote computing device. Thus, a user of the remote computing device can explore the metrics associated with each data source to compare the data sources. In some aspects, when the data selection request message includes a set of metric thresholds, the computing system can select a data source based on the set of metric thresholds. For example, the remote computing device can specify different thresholds for each metric. A data source having a highest number of metrics exceeding the thresholds can be selected.
[0016] In some examples, the computing system can use the selected data source to determine a risk indicator for a target entity. The data selection request message can include a target entity identifier, such that, in addition to selecting a data source, the computing system can train a risk assessment model on the data source and use the trained risk assessment model to generate a risk indicator for the target entity. The system can then transmit the risk indicator to the remote computing device. The risk indicator can be used to control access of the target entity to an interactive computing environment. For example, the risk indicator can be included in a responsive message to the request for evaluating the target entity such that the responsive message can be used to allow, challenge, or deny access to the target entity. For example, if the risk indicator isbelow a predefined threshold, a request by the target entity to access the interactive computing environment may be automatically denied or flagged for manual review.
[0017] Certain aspects described herein, which can include generating one or more metrics and a data source score and providing a responsive message including the data source score, can improve at least the technical field of machine learning. For example, machine learning outcomes depend on the training data used to train the machine learning model. Thus, a model trained on more robust and unbiased data is likely to be more accurate than a model trained on a small or inaccurate data set. Further, disclosed systems and methods enable selection of a data source that is robust and that has the characteristics or qualities required for its intended use. This combination of features provides a robust system for evaluating data sources to reduce redundant data and improve data quality. As an example, a user can set a high fraud score threshold such that the selected data source is the most accurate of the data sources in flagging fraudulent records, thus making it more well suited for use in training a machine learning model to identify fraudulent transactions. In another example, an entity can reduce costs associated with purchasing a vendor’s data set by determining how much overlap one data set has with another. This can also improve system efficiency by reducing the need for preprocessing operations to handle duplicate data or merge data sets.
[0018] Certain aspects described herein, which can further include generating one or more risk indicators associated with target entities and providing a responsive message using the risk indicator, can improve at least the technical fields of controlling interactions between computing environments, access control for a computing environment, or a combination thereof. For instance, by generating and transmitting the responsive message, the computing system can cause access to a computing system to be controlled more accurately by training a machine learning model on a robust data set, thereby improving machine learning outcomes and predictive power. The risk indicator may be used to better predict whether the target entity requesting access is legitimate, and using the risk indicator may yield fewer malicious interactions than if the responsive message is not used. Further, the risk assessment computing system leverages distinctive components of the risk indicator to create a robust and easily implemented framework.
[0019] These illustrative examples are given to introduce the reader to the general subject matter discussed here and are not intended to limit the scope of the disclosed concepts. Thefollowing sections describe various additional features and examples with reference to the drawings in which like numerals indicate like elements, and directional descriptions are used to describe the illustrative examples but, like the illustrative examples, should not be used to limit the present disclosure.Operating Environment Example for Selecting a Data Source and Generating a Risk Indicator associated with a Target Entity
[0020] Referring now to the drawings, FIG. 1 is a block diagram depicting an example of an operating environment in which a risk assessment computing system can be used to provide a risk assessment associated with a target entity according to some aspects of the present disclosure. FIG. 1 depicts examples of hardware components of a computing system 102, according to some aspects. The computing system 102 can be a specialized computing system that may be used for processing large amounts of data using a large number of computer processing cycles. In other examples, the computing system 102 may be or include a general-purpose computing system. The computing system 102 can include a server 104 for performing data source selection (e.g., analyzing data sources and determining data source metrics) with respect to a set of data sources. In some examples, the server 104 can further perform a risk assessment (e.g., predicting future risk associated with the target entity, predicting the legitimacy of the target entity, etc.) with respect to a target entity, such as a target individual or a user computing device.
[0021] The server 104 can include one or more processing devices that can execute program code, such as a data selection application 106 and risk assessment application 108. The program code can be stored on a non-transitory computer-readable medium or other suitable medium. The data selection application 106 can include one or more modules or components executing software code to complete one or more steps for determining a data source score or one or more metrics associated with a data set from a given data source. For example, the data selection application 106 can include a data sampling module 110 and a data source scoring module 112. The data sampling module 110 can access a data source and generate one or more data samples based on data stored in the data source (e.g., external data sources 114). These data samples can include a randomly selected data sample, a fraud data sample (e.g., a data sample with fraudulent records identified and labelled), and a synthetic data sample (e.g., a scrambled or otherwise randomized set of data from the data source). The data source scoring module 112 can determine, based on thedata samples, a set of metrics associated with each of data sources 114. The metrics can be used to determine a data source score, which can be used in selecting a data source from the data sources 114.
[0022] The server 104 can also include a risk assessment application 108. The risk assessment application 108 can access data from the selected data source and train a risk assessment model on the data from the selected data source. The risk assessment model can be trained to determine a risk indicator associated with a target entity. The target entity, for example, can be specified in the request for data selection and can be an entity requesting access to an interactive computing environment 124.
[0023] In some examples, the server 104 can perform risk assessment operations or access control operations for validating or otherwise authenticating a target entity, for example using other suitable modules, models, components, etc. of the server 104. The server 104 can receive data associated with the target entity from a selected external data source of external data sources 114 (e.g., the database selected based on the data source score determined by the data selection application 106), as well as internal data 118 from data repository 116, or any suitable combination thereof. In some aspects, the risk assessment application 108 can authenticate or deny a request for an interaction involving the target entity by generating a risk indicator using the target entity data retrieved from the selected data source of the external data sources 114 and the data repository 116.
[0024] In some aspects, the target entity data can be determined or stored in one or more network-attached storage units on which various repositories, databases, or other structures are stored. An example of these data structures can include the data repository 116. In some examples, a combination of internal data sets 118 and data sets from one or more selected external data sources 114 can be used to train a risk assessment machine-learning models. The risk assessment machine-learning model can be trained to generate a risk indicator for use in determining whether to allow a target entity to access a secure or protected resource.
[0025] Network-attached storage units may store a variety of different types of data organized in a variety of different ways and from a variety of different sources. For example, the network- attached storage unit may include storage other than primary storage located within the server 104 that is directly accessible by processors located therein. In some aspects, the network-attachedstorage unit may include secondary, tertiary, or auxiliary storage, such as large hard drives, servers, and virtual memory, among other types of suitable storage. Storage devices may include portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing and containing data. A machine-readable storage medium or computer-readable storage medium may include a non-transitory medium in which data can be stored and that does not include carrier waves or transitory electronic signals. Examples of a non-transitory medium may include, for example, a magnetic disk or tape, optical storage media such as a compact disk or digital versatile disk, flash memory, memory devices, or other suitable media.
[0026] Furthermore, the computing system 102 can communicate with various other computing systems. The other computing systems can include user computing systems 120, such as smartphones, personal computers, etc., client computing systems 122, and other suitable computing systems. In one example, the client computing system 122 can transmit a data selection request to the server 104. The request can specify one or more of external data sources 114 to evaluate. The data selection application 106 can generate data samples from each data source 114 and generate metrics and a data source score for each data source of data sources 114. The data selection application 106 can determine a selection of a data source from among data sources 114 and transmit an identification of the selected data source to the client computing system 122. The client computing system 122 can then use the selected data source to train a machine learning model on robust and non-redundant data. In some aspects, the server 104 can transmit the one or more metrics and data source scores to the client computing system 122 such that the client computing system 122 can display a graphical user interface (GUI) enabling a user to explore the metrics and data source scores to evaluate each data source 114.
[0027] In another example, user computing systems 120 may transmit, such as in response to receiving input from the target entity, requests for accessing the interactive computing environment 124 to the client computing systems 122. In response, the client computing systems 122 can send authentication queries to the server 104, and the server 104 can receive data associated with the target entity from an internal data source and a selected external data source and generate a risk indicator associated with the target entity. While FIG. 1 illustrates that the computing system 102 and the client computing systems 122 are separate systems, the computing system 102 and the client computing systems 122 can be one system. For example, the computing system 102 can be a part of the client computing systems 122, or vice versa. In other examples,the risk assessment application 108 may be hosted on a separate server, or a separate system, and may be in communication with the computing system 102 to receive an indication of the selected data source or data sources, which in turn can be used to provide training data for training a risk assessment model.
[0028] As illustrated in FIG. 1, the computing system 102 may interact with the client computing systems 122, the user computing systems 120, or a combination thereof via one or more public data networks 126 to facilitate interactions between users of the user computing systems 120 and the interactive computing environment 126. For example, the computing system 102 can facilitate the client computing systems 122 providing a user interface to the user computing system 120 for receiving various data from the user. The computing system 102 can transmit validated risk assessment data, for example similarity-preserving hashes, comparisons or scores determined therefrom, etc., to the client computing systems 122 for providing, challenging, or rejecting, etc. access of the target entity to the interactive computing environment 124. In some examples, the computing system 102 can additionally communicate with third-party systems to receive risk assessment data, entity data, and the like, through the public data network 126. In some examples, the third-party systems can provide real-time (e.g., streamed) data about the target entity, historical data about the target entity, etc. to the computing system 102.
[0029] In another example, the client computing systems 122 can request a risk indicator determined by the risk assessment application 108 for a target entity. Prior to determining the risk indicator, the data selection application 106 can determine which data source of data sources 114 to use as the source of training data to train a risk assessment model for determining the risk indicator. In another example, the client computing system 122 can request a data source selection from the computing system 102, such that the computing system 102 determines, based on one or more metrics, a selected data source to be used by the client computing system 122 in training a model. The metrics on which the data sources are evaluated can be compared to a set of default thresholds, or a threshold associated with each metric can be received by the computing system 102 from the client computing system 122.
[0030] Each client computing system 122 may include one or more devices such as individual servers or groups of servers operating in a distributed manner. In some examples, a client computing system 122 can include any computing device or group of computing devices operatedby a seller, lender, or other suitable entity that can provide products or services. The client computing system 122 can include one or more server devices. The one or more server devices can include or can otherwise access one or more non-transitory computer-readable media.
[0031] The client computing system 122 can further include one or more processing devices and an interface capable of displaying a GUI to a user of the client computing system 122. For example, in response to a request for a data source selection from the computing system 102, the client computing system 122 can receive a set of metrics and data source scores associated with the data sources 114. The client computing system 122 can be configured to display a GUI through which a user can explore the metrics associated with each data source 114. In some examples, the GUI can include one or more interactive elements or data visualizations associated with the one or more metrics and data source scores, such that a user of the client computing system 122 can visually compare the data sources 114.
[0032] The client computing system 122 can further include one or more processing devices that can be capable of providing an interactive computing environment 124, such as a user interface, etc., that can perform various operations. The interactive computing environment 124 can include executable instructions stored in one or more non-transitory computer-readable media. The instructions providing the interactive computing environment 124 can configure one or more processing devices to perform the various operations. In some aspects, the executable instructions for the interactive computing environment 124 can include instructions that provide one or more graphical interfaces. The graphical interfaces can be used by a user computing system 120 to access various functions of the interactive computing environment 124. For instance, the interactive computing environment 124 may transmit data to and receive data, such as via the graphical interface, from a user computing system 120 to shift between different states of the interactive computing environment 124, where the different states allow one or more electronic interactions between the user computing system 120 and the client computing system 122 to be performed.
[0033] In some examples, the client computing system 122 may include other computing resources associated therewith (e.g., not shown in FIG. 1), such as server computers hosting and managing virtual machine instances for providing cloud computing services, server computers hosting and managing online storage resources for users, server computers for providing database services, and others. The interaction between the user computing system 120, the client computingsystem 122, and the computing system 102, or any suitable sub-combination thereof may be performed through graphical user interfaces, such as the user interface, presented by the computing system 102, the client computing system 122, other suitable computing systems of the computing environment 100, or any suitable combination thereof. The graphical user interfaces can be presented to the user computing system 120 or the client computing system 122. Application programming interface (API) calls, web service calls, or other suitable techniques can be used to facilitate interaction between any suitable combination or sub-combination of the client computing system 122, the user computing system 120, and the computing system 102.
[0034] A user computing system 120 can include any computing device or other communication device that can be operated by a user or entity, such as the user entity, which may include a consumer or a customer. The user computing system 120 can include one or more computing devices such as laptops, smartphones, and other personal computing devices. A user computing system 120 can include executable instructions stored in one or more non-transitory computer-readable media. The user computing system 120 can additionally include one or more processing devices configured to execute program code to perform various operations. In various examples, the user computing system 120 can allow a user to access certain online services or other suitable products, services, or computing resources from a target entity, such as the client computing system 122, to engage in mobile commerce with the client computing system 122, to obtain controlled access to electronic content, such as the interactive computing environment 124, hosted by the client computing system 122, etc.
[0035] In some examples, the user or a target entity can use the user computing system 120 to engage in an electronic interaction with the client computing system 122 via the interactive computing environment 124. The computing system 102 can receive a request, for example from the user computing system 122, to access the interactive computing environment 124 and can use target entity data or any other suitable data or signals determined therefrom, to determine whether to provide access, to challenge the request, to deny the request, etc. The computing system 102 can, for example, train a machine learning model to determine whether to provide access, to challenge the request, or to deny the request, etc. by determining a risk indicator. The machine learning model can be trained on internal data sets 118 as well as on data from one or more of data sources 114, as selected by the data selection application 106. In some aspects, the client computing system 122 can include a risk assessment application, such that the client computingsystem 122 requests a data source selection from the data selection application 106, and trains a risk assessment model using data from the selected data source or data sources.
[0036] In some examples, the generation of a risk indicator associated with a target entity by the computer system 102 or by the client computing system 122 can be triggered by an interaction between the user computing system 120 and the client computing system 122. An electronic interaction between the user computing system 120 and the client computing system 122 can include, for example, the user computing system 120 being used to request a financial loan or other suitable services or products from the client computing system 122, and so on. An electronic interaction between the user computing system 120 and the client computing system 122 can also include, for example, one or more queries for a set of sensitive or otherwise controlled data, accessing online financial services provided via the interactive computing environment 124, submitting an online credit card application or other digital application to the client computing system 122 via the interactive computing environment 124, operating an electronic tool within the interactive computing environment 124 (e.g., a content-modification feature, an applicationprocessing feature, etc.), etc.
[0037] In some aspects, an interactive computing environment 124 implemented through the client computing system 122 can be used to provide access to various online functions. As a simplified example, a user interface or other interactive computing environment 124 provided by the client computing system 122 can include electronic functions for requesting computing resources, online storage resources, network resources, database resources, or other types of resources. In another example, a website or other interactive computing environment 124 provided by the client computing system 122 can include electronic functions for obtaining one or more financial services, such as an asset report, management tools, credit card application and transaction management workflows, electronic fund transfers, etc.
[0038] A user computing system 120 can be used to request access to the interactive computing environment 124 provided by the client computing system 124. The client computing system 124 can submit a request, such as in response to a request made by the user computing system 120 to access the interactive computing environment 124, for risk assessment to the computing system 102 and can selectively grant or deny access to various electronic functions based on risk assessment performed by the computing system 102 using a selected data source or data sources.In another example, the client computing environment can request a data source selection from the computing system 102 and can use data from the selected data sources to train a risk assessment model stored and executed by the client computing system 122. Based on the request, or continuously or substantially contemporaneously, the computing system 102 or the client computing system 122 can determine one or more risk signals or risk indicators for data associated with the target entity, which may submit or may have submitted the request via the user computing system 120. The risk signals or risk indicators can be determined using a model trained on data from a selected data source or data sources, as well as internal data. Based on a risk indicator determined from the risk assessment application 108, the computing system 102, the client computing system 122, or a combination thereof can determine whether to grant the access request of the user computing system 120 to certain features of the interactive computing environment 124. The computing system 102, the client computing system 122, or a combination thereof can use the risk indicator for other suitable purposes such as identifying a manipulated identity, controlling a real-world interaction, and the like.
[0039] In a simplified example, the system illustrated in FIG. 1 can configure the server 104 to be used for controlling access to the interactive computing environment 124. The server 104 can retrieve data associated with the target entity in response to a request to access the interactive computing environment 124. The data may, for example, be retrieved based on identity information (e.g., information collected by the client computing system 122 via a user interface provided to the user computing system 120) provided by the client computing system 122 or received via other suitable computing systems. The server 104 can retrieve the data associated with the target entity from one or more data sources 114, as determined by the data selection application 106. The data sources 114 can store, for example, historical data, transaction data, financial data, and the like. The server 104 can determine a risk indicator associated with the target entity by training a risk assessment model on data from the data source(s) selected by the data selection application 106. The server 104 can transmit the risk indicator, or any inference derived therefrom, to the client computing system 122 for use in controlling access to the interactive computing environment 124. In another example, the server 104 can transmit an indication of the selected data sources to the client computing system 122, such that the client computing system 122 can train a risk assessment model on data from the selected data sources and can determine a risk indicator using the trained risk assessment model.
[0040] The risk indicator associated with the target entity, or any suitable score or comparison determined therefrom, can be used, for example by the computing system 102, the client computing system 122, etc., to determine whether the risk associated with the target entity accessing a good or a service provided by the client computing system 122 using exceeds a threshold, thereby granting, challenging, or denying access by the target entity to the interactive computing environment 124. For example, if the computing system 102 or the client computing system 122 determines that the risk indicator indicates that risk associated with the identity element is lower than a threshold value, then the client computing system 122 associated with the service provider can generate or otherwise provide access permission to the user computing system 120 that requested the access. The access permission can include, for example, cryptographic keys used to generate valid access credentials or decryption keys used to decrypt access credentials. The client computing system 122 can also allocate resources to the target entity and provide a dedicated web address for the allocated resources to the user computing system 120, for example, by adding the user computing system 120 in the access permission. With the obtained access credentials or the dedicated web address, the user computing system 120 can establish a secure network connection to the interactive computing environment 124 hosted by the client computing system 122 and access the resources via invoking API calls, web service calls, HTTP requests, other suitable mechanisms or techniques, etc.
[0041] In some examples, the computing system 102 may determine whether to grant, challenge, or deny the access request made by the user computing system 120 for accessing the interactive computing environment 124. For example, based on the risk indicator associated with the target entity, the computing system 102 or the client computing system 122 can determine that the target entity is a legitimate entity that made the access request and may authenticate the request. In other examples, the computing system 102 or the client computing system 122 can challenge or deny the access attempt if the computing system 102 or the client computing system 122 determines that the target entity may not be a legitimate entity.
[0042] In some examples, the risk indicator used to determine access to the interactive computing environment 124 may be determined at least in part based on output from one or more machine-learning models. For example, a machine-learning model may be trained on data from one or more data sources selected from the data sources 114 by the data selection application 106. The data selection application 106 can determine one or more metrics and data source scoresassociated with each of the data sources 114. The metrics and data source scores can be used to compare the data sources 114 to determine and select one or more data sources containing data with characteristics suited to the desired application of the data, i.e., using the data to train a machine-learning model to determine a risk indicator. The thresholds or characteristics that determine which data source is selected may vary depending upon the desired use of the data from the selected data source.
[0043] Each communication within the computing environment 100 may occur over one or more data networks, such as a public data network 126, a network 128 such as a private data network, or some combination thereof. A data network may include one or more of a variety of different types of networks, including a wireless network, a wired network, or a combination of a wired and wireless network. Examples of suitable networks include the Internet, a personal area network, a local area network (“LAN”), a wide area network (“WAN”), or a wireless local area network (“WLAN”). A wireless network may include a wireless interface or a combination of wireless interfaces. A wired network may include a wired interface. The wired or wireless networks may be implemented using routers, access points, bridges, gateways, or the like, to connect devices in the data network.
[0044] The number of devices illustrated in FIG. 1 is provided for illustrative purposes. Different numbers of devices may be used. For example, while certain devices or systems are shown as single devices in FIG. 1 , multiple devices may instead be used to implement these devices or systems. Similarly, devices or systems that are shown as separate may be instead implemented in a signal device or system.Architecture for Implementing a System for Selecting a Data Source
[0045] FIG. 2 is a block diagram depicting an example environment 200 for selecting a data source according to some aspects of the present disclosure. The environment 200 can include components as described above with reference to FIG. 1. In some examples, computing system 102 may be the same as the client computing system 122. Other implementations or architectures, however, are possible.
[0046] The environment 200 can include the computing system 122 and data sources 114 including a Data Source 1, Data Source 2, ... , and Data Source N. Each data source can store different data sets. In some examples, the data sources 114 can store overlapping or related datasets. In some examples, the data stored in data sources 114 can be transaction data or financial data associated with a set of identities. The data can include, for example, personally identifiable information (PII) such as names, addresses, phone numbers, Social Security Numbers (SSNs), email addresses, dates of birth (DOBs), and the like. Each data source 114 may be provided by a different vendor, such that different vendors may maintain the data sources differently (e.g., with differing levels of identity verification, security, or fraud detection). Additionally, the data sources 114 can contain redundant data, data overlapping with other data sources of data sources 114, or data overlapping or redundant with the internal data sets 118.
[0047] The computing system 102 can receive a request for selecting a data source. In some examples, the request may include a set or subset of available data sources (e.g., data sources 114 or a subset thereof). In some aspects, the request can specify a set of thresholds associated with a set of metrics or a threshold associated with a data source score. The thresholds can be set depending on the intended use of the data from the selected data set. For example, where the client computing system 122 includes preprocessing functionality, the threshold for a data source containing duplicate data from another source or an internal data source may be lower than if the client computing system 122 does not include preprocessing or deduplicating functionality. In another example, data source coverage may be important to reduce cost associated with purchasing multiple data sets from multiple data sources and thus the coverage threshold may be higher than if data coverage were not important.
[0048] For each data source in the data sources 114, the data sampling module 110 can generate one or more data samples. The minimum sample size for each sample can be determined using methods to determine the sample size required for statistically significant results. The data sampling module 110 can generate a random data sample of records randomly selected from each data source. The random data sample can be used to measure the characteristics of each data source. The data sampling module 110 can also generate a fraud data sample. The fraud data sample can include records associated with known fraud cases. For example, records in a data source can be reviewed or analyzed to identify fraudulent records. These records can be compiled into the fraud data sample. The fraud data sample can be used to evaluate how a vendor of a data source (e.g., a vendor of Data Source 1, Data Source 2, ... , or Data Source N) differentiates identities associated with fraudulent activities. The data sampling module 110 can also generate a synthetic data sample. The synthetic data sample can be generated by randomizing a data samplefrom the data source. In some instances, machine learning can be used to randomize PII data, or to generate synthetic data from a set of PII data. The synthetic data sample can be used to mimic fraud patterns to test a vendor’s ability to identify fraud and their susceptibility to false trust. In some aspects, the synthetic data can be generated by fuzzifying the random data sample.
[0049] The data scoring module 112 can use the data samples generated by the data sampling module 110 to determine a set of metrics for each data source. The metrics can include: a coverage score; a verification score; an affiliation score; a risk score; a trust score; a fraud score; a concurrence score; a demographic score; a sensitivity score; and a consistency score.
[0050] The coverage score can reflect, for example, the percentage of the overall population that is covered by the data source. This can be, for example, the percentage of the overall population of identities that have records with enough features to be used in identity authentication. Coverage can be determined using Equation 1.Equation 1where Nsurveyis the data sample size, Ftis the number of features used by the vendor in identity authentication, and i is the feature from 1 to a, or the total number of features. Coverage can be calculated based on the random data sample generated by the data sampling module 110.
[0051] The verification score can indicate how many transactions of the data sample can be verified using a particular PII element. In other words, the verification score can measure how many of the identities in the random data sample are verified based on a PII element (e.g., phone number, email address, SSN, etc.). The verification score can be determined using Equation 2:Equation 2where Vtare the verification features used in identity authentication (e.g., the PII elements used to verify an identity in the data source) from i = 1 to i = y, where y is the total number of verification features.
[0052] The affiliation score measures the percentage of identities in the data sample that are affiliated with one or more PII elements. For example, the affiliation score can measure the number of identities with which a particular PII element has been successfully affiliated. Successful affiliation may be, for example, determined based on an affiliation score of the PII element to the identity being above a threshold. The affiliation score can be determined using Equation 3 :Equation 3where Atare the affiliation features used in affiliating an identity in the data source and where 8 is the total number of affiliation features.
[0053] The risk score can measure a data source vendor’s ability to identify risk. For example, the risk score can indicate a data source vendor’s ability to identify a risky identity or risky record. Risk associated with an identity can be measured based on a number of features associated with a PII element. For example, for a phone number, features indicating risk can be a toll-free area code, account tenure, a voice-over-IP (VOIP) number. In another example, for an email address, features indicating risk can include account tenure, affiliation with a risky domain, etc. The risk score can be determined using Equation 4:Equation 4where Rtare the features used to identify risk and f> is the total number of features. In some examples, an identity may satisfy the requirement for being counted as a risk based on a risk score determined for the identity based on a feature and compared to a risk threshold.
[0054] The trust score can be built off the number of identities in the data sample that are verified, affiliated, and associated with no risk, thereby measuring a data source vendor’s ability to determine whether or not to trust an identity. In some aspects, the data sample can include indications of whether or not an identity has been verified and authenticated, and whether or not the identity is associated with a risk. In some examples, this data is determined by the data sourcevendor. In other examples, this data is determined by the computing system 102 based on an analysis of the data sample. The trust score can be determined using Equation 5:Equation 5
[0055] In some examples, trust can also be based on the number of identities in the sample that have a risk score below a risk threshold. The parameters used to determine the trust score can be varied based on the data’s intended use.
[0056] The fraud score can be based on a trust rate determined by comparing the data sample, with the fraud data sample. For example, if an identity is in the fraud sample, and therefore associated with fraudulent activity, the identity should be identified as fraudulent in the data sample. The number of identities that are identified as fraud in the fraud sample that are not identified as fraud in the data sample are counted. The trust rate should be 0 if the data source vendor correctly identified all fraud in the data sample. Thus, the lower the trust score, the more fraud in the data sample was correctly identified.
[0057] The concurrence score can be used to compare data sources or data source vendors. Thus, the data source scoring module 112 can determine, for each pair of data sources (e.g., Data Source 1 and Data Source 2, Data Source 1 and Data Source N, Data Source 2 and Data Source N, etc.) a measure of agreement between the data sources. In some examples, the concurrence score can be a summation of the pairwise correlations between each pair of data sources involving a particular data source. The concurrence can be determined by first computing a trust rate for each data source vendor, as described above. The correctly identified fraud in each data source’s data sample can be compared to determine if the data sources missed the same fraudulent identities in their respective data samples. The pair of data sources may have higher concurrence score if the respective data source vendors missed the same fraudulent identities and a lower concurrence score if the respective data source vendors missed different fraudulent identities.
[0058] The demographic score can measure whether the data in the data sample, for a particular demographic group, can be used to reach a particular outcome. For example, the data sample can be separated into age groups (e.g., demographic groups) such that the data in each age group is used to determine an outcome. The outcome can be compared with an expected outcomeor an average outcome for the age group. The demographic score can be a count of the number of demographic groups for which the determined outcome is not equal to, or is outside one or more standard deviations, of the expected outcome for that demographic group.
[0059] As an example of demographic score, a data sample can include financial data associated with a set of identities, where the data can be used to calculate a credit score for each identity. The data sample can be split into subsets of data based on identity age. For each subset of data, the data scoring module 112 can determine an average credit score. For each subset of data, the average credit score can be compared to an expected average credit score for the age group corresponding to the particular subset. If the average credit score for the age group differs from the expected average credit score for the age group, the demographic score can be increased by one. Thus, a lower demographic score indicates that the outcomes based on the data sample agree with the expected outcomes.
[0060] The sensitivity score can measure the level at which a data source vendor penalizes their trust rate when real data has been altered, masked, or permutated. The sensitivity score can be determined based on the synthetic data set. The sensitivity score can be determined as the ratio of trust rate in the synthetic population over the trust rate in a control group, reversed,, and scaled from 0 to 100. Thus, the highest sensitivity score can be 100, while a low sensitivity of 0 can indicate that the trust rate based on the synthetic data sample is the same as the trust rate in the control population, implying that the data source vendor could not distinguish between data sets.
[0061] The consistency score can measure the ability of a data source vendor to handle duplicate records and reach the same results. For example, if an identity is handled twice within a short timeframe A(e.g., to recalculate an outcome for the identity) the outcome should be the same both times it is calculated. Thus, the consistency score can indicate a percentage of nondiscrepancy in decisions (e.g., outcome calculations) made by the data source vendor. The consistency score can be determined based on Equation 6:Equation 6where NDoEis the data sample size and a is the number of features used to authenticate an identity in the data sample.
[0062] In some aspects, the data source scoring module 112 can combine one or more of the above-described scores to determine a data source score. The data source score can, in some cases, be a composite score reflective of the quality and robustness of the data source. As an example, the data source module 112 can combine the coverage score, verification score, affiliation score, risk score, trust score, and fraud score for each of the data sources 114 to determine a data source score for each of the databases 114. The data source scores provide a measure by which the data sources can be compared. In some examples, the data source score can be based on differently weighted metrics. For example, for a particular data application, the fraud score may be weighted more heavily than the verification score, etc.
[0063] In some aspects, the computing system 102 can further generate a responsive message including the set of metrics (e.g., a coverage score; a verification score; an affiliation score; a risk score; a trust score; a fraud score; a concurrence score; a demographic score; a sensitivity score; and a consistency score), as well as the data source score for each evaluated data source. The responsive message can cause the client computing system 122 to display a GUI through which a user can view and compare the data sources based on the available scores and metrics. For example, the GUI can include interactive data visualizations to facilitate a comparison of the data sources such that a user is able to select a data source having data best suited for its intended use.
[0064] In some aspects, the computing system 102 can automatically select a data source or data sources for use based on comparison of one or more metrics or the data source scores with a set of thresholds. The thresholds can, in some examples, be specified by the client computing system 122 in the data source selection request, or can be a set of default predetermined thresholds. In some examples, the client computing system 122 can receive the set of metrics and data source scores and compare the metrics and scores with a set of thresholds.
[0065] Thus, the computing system 102 can evaluate the data sources 114 to determine comparable metrics and data source scores. Accordingly, Data Source 1 can be compared to Data Source 2, and so on. The comparable metrics and data source scores enable a user or system to evaluate the data included in the data source, as well as the data source vendor, to determine a bestfit data set for an intended application. This enables systems to access, from a data source vendor, a robust and accurate data set.Techniques for Selecting a Data Source
[0066] FIG. 3 is a flow chart illustrating an example of a process 300 for selecting a data source according to some aspects of the present disclosure. In some examples, the operations of the process 300, or any subset thereof, may be performed by the computing system 102 via the server 104, or by the client computing system 122, but other suitable systems, devices, or subsets or combinations thereof may perform one or more operations described with respect to the process 300. For illustrative purposes, the process 300 is described with reference to certain examples depicted in the figures. Other implementations are possible.
[0067] At block 302, the process 300 involves receiving a data selection request from the client computing system 122. The data selection request can indicate a set of data sources to be evaluated. In some examples, the data selection request can be initiated in response to an access request from the user computing system 120 to access the interactive computing environment 124. The data selection request can be used to select a data source for training data to train a machine learning model to determine a risk indicator associated with a target entity.
[0068] At block 304, the process 300 involves, for each data source, generating a random data sample, a fraud data sample, and a synthetic data sample. The random data sample can be a sample of a minimum sample size randomly selected from the data source. The fraud data sample can be a data sample including identities that have previously been determined to be fraudulent (e.g., a set of records labelled as being fraudulent). The synthetic data sample can include randomized data from the data source. In some examples, the synthetic data sample can be based on the random data sample.
[0069] At block 306, the process 300 involves, for each data source, generating a set of metrics based on the random data sample, the fraud data sample, and the synthetic data sample associated with the respective data source. As discussed above, the set of metrics can include: a coverage score; a verification score; an affiliation score; a risk score; a trust score; a fraud score; a concurrence score; a demographic score; a sensitivity score; and a consistency score. The data required to determine the set of metrics (e.g., affiliation data, verification data, trust rate, etc.) can be determined by the computing system 102, the client computing system 122, or the data source 1vendor associated with the respective data source. For example, the data selection application 106 can include data preprocessing for each data source to determine the additional information required to calculate the set of metrics. In another example, the client computing system 122 or a data source vendor can calculate the additional information and provide the additional information to the computing system 102.
[0070] At block 308, the process 300 involves determining, for each data source and based on the set of metrics associated with the data source, a data source score. In some examples, the data source score can be a weighted sum of one or more of the set of metrics. In another embodiment, the data source score can be normalized to a particular scale, to facilitate comparison between the data sources.
[0071] At block 310, the process 300 involves selecting, based at least in part on the data source score, a data source from the set of data sources. In some examples, the data sources can be ranked by data source score from high to low and the data source with the highest data source score can be selected. In other examples, data sources having data source scores above a threshold value may be identified and transmitted, with their associated metrics, to the client computing system 122 for further comparison.
[0072] At block 312, the process 300 involves transmitting, to a remote computing device (e.g., the client computing device 122), a responsive message comprising at least an indication of the selected data source. In some examples, the computing system 102 can further receive a request for determining a risk indicator associated with a target entity, where the risk indicator is determined by a machine learning model trained on data from the selected data source. For example, the risk indicator can be used in controlling an interaction involving a target entity or access of the target entity to a restricted system. In another example, the computing system 102 can provide the set of metrics and data source scores associated with the set of data sources to the client computing system 122 in a GUI configured to allow a user of the client computing system 122 to explore and compare the data sources prior to selecting a data source.
[0073] Systems and methods described herein provide advantages over traditional methods for selecting data sources (e.g., based on a singular metric or by comparing data sources using different parameters). For example, disclosed systems and methods allow for dynamic selection of a data source tailored to a specific application based on comparable metrics and scores. The metrics andthresholds can be modified based on the desired data source properties, such that a data source selected for an application is the best fit for that application from a set of data sources. This further can reduce required data preprocessing, thereby improving computing efficiency.
[0074] Because machine learning outcomes depend on the data used to train the machine learning model, the selection of a data source for training data can impact machine learning outcomes and predictive power. Thus, by comparing data sources, disclosed systems and methods can enable selection of an appropriate data source to provide training data for a machine learning model. In risk assessment applications, this can yield more accurate risk predictions, thereby improving system security.Example of Computing System
[0075] Any suitable computing system or group of computing systems can be used to perform the operations for the techniques described herein. For example, FIG. 4 is a block diagram depicting an example of a computing device 400, which can be used to implement the server 104. The computing device 400 can include various devices for communicating with other devices in the computing environment 100, as described with respect to FIG. 1. The computing device 400 can include various devices for performing one or more operations, such as risk assessment operations, described above with respect to FIGs. 1-3.
[0076] The computing device 400 can include a processor 402 that can be communicatively coupled to a memory 404. The processor 402 can execute computer-executable program code stored in the memory 404, can access information stored in the memory 404, or both. Program code may include machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, among others.
[0077] Examples of a processor 402 can include a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or any other suitable processing device. The processor 402 can include any suitable number of processing devices,including one. The processor 402 can include or communicate with a memory 404. The memory 404 can store program code that, when executed by the processor 402, causes the processor 402 to perform the operations described herein.
[0078] The memory 404 can include any suitable non-transitory computer-readable medium. The computer-readable medium can include any electronic, optical, magnetic, or other storage device capable of providing a processor with computer-readable program code or other program code. Non-limiting examples of a computer-readable medium can include a magnetic disk, memory chip, optical storage, flash memory, storage class memory, ROM, RAM, an ASIC, magnetic storage, or any other medium from which a computer processor can read and execute program code. The program code may include processor-specific program code generated by a compiler or an interpreter from code written in any suitable computer-programming language. Examples of suitable programming language can include Hadoop, C, C++, C#, Visual Basic, Java, Python, Perl, JavaScript, ActionScript, etc.
[0079] The computing device 400 may also include a number of external or internal devices such as input or output devices. For example, the computing device 400 is illustrated with an input / output interface 408 that can receive input from input devices or provide output to output devices. A bus 406 can also be included in the computing device 400. The bus 406 can communicatively couple one or more components of the computing device 400.
[0080] The computing device 400 can execute program code 414 that can include data selection application 106 and risk assessment application 108. The program code 414 for the risk assessment application 106 may be resident in any suitable computer-readable medium and may be executed on any suitable processing device. For example, and as illustrated in FIG. 4, the program code 414 for the data selection application 106 and the risk assessment application 108 can reside in the memory 404 at the computing device 400 along with the program data 416 associated with the program code 414. Executing the data selection application 106 or the risk assessment application 108 can configure the processor 402 to perform at least a portion of the operations described herein.
[0081] In some aspects, the computing device 400 can include one or more output devices. One example of an output device can be or include the network interface device 410 illustrated in FIG. 4. A network interface device 410 can include any device or group of devices suitable forestablishing a wired or wireless data connection to one or more data networks described herein. Non-limiting examples of the network interface device 410 can include an Ethernet network adapter, a modem, etc.
[0082] Another example of an output device can include the presentation device 412 depicted in FIG. 4. A presentation device 412 can include any device or group of devices suitable for providing visual, auditory, or other suitable sensory output. Non-limiting examples of the presentation device 412 can include a touchscreen, a monitor, a speaker, a separate mobile computing device, etc. In some aspects, the presentation device 412 can include a remote clientcomputing device that communicates with the computing device 400 using one or more data networks described herein. In other aspects, the presentation device 412 can be omitted.
[0083] The foregoing description of some examples has been presented only for the purpose of illustration and description and is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Numerous modifications and adaptations thereof will be apparent to those skilled in the art without departing from the spirit and scope of the disclosure.
Claims
ClaimsWhat is claimed is:
1. A system comprising: a processor; and a non-transitory computer-readable medium comprising instructions that are executable by the processor for causing the processor to perform operations comprising: receiving a data selection request from a remote computing device indicating a set of data sources; for each data source of the set of data sources: generating a random data sample, a fraud data sample, and a synthetic data sample, wherein: the random data sample comprises a subset of data from the respective data source; the fraud data sample comprises a set of labelled data from the respective data source; and the synthetic data sample comprises randomized data from the respective data source, generating a set of metrics based on the random data sample, the fraud data sample, and the synthetic data sample, and based on the set of metrics, determining a data source score; selecting, a data source from the set of data sources based at least in part on the data source score; and transmitting, to the remote computing device, a responsive message comprising at least an indication of the selected data source.
2. The system of claim 1, wherein the set of metrics comprises a coverage score based on a data sample of the respective data source, and wherein the coverage score comprises a percentage of a population represented by the data sample of the respective data source.
3. The system of claim 1, wherein the set of metrics comprises a risk score based on a data sample of the respective data source, and wherein the risk score is generated based on a percentage of identities in the data sample of the respective data source that are associated with at least one risk factor.
4. The system of claim 1 , wherein the set of metrics comprises a verification score based on a data sample of the respective data source, and wherein the verification score is generated based on a number of verification features used to authenticate an identity in the data sample of the respective data source.
5. The system of claim 1, wherein the set of metrics comprises an affiliation score based on a data sample of the respective data source, and wherein the affiliation score is generated based on a number of affiliation features used to authenticate an identity in the data sample of the respective data source.
6. The system of claim 1, wherein the set of metrics comprise a trust score based on a data sample of the respective data source, and wherein the trust score is generated based on a percentage of identities in the respective data sample that are verified and affiliated.
7. The system of claim 1 , wherein the operations further comprise: determining a demographic score based on a random data sample of the respective data source, and wherein the demographic score is determined by: determining an average outcome for a set of identities in the random data sample, wherein the set of identities belong to an age group, and comparing the average outcome to a threshold outcome associated with the age group; and transmitting, to the remote computing device, the demographic score associated with the selected data source.
8. A method comprising: receiving, at a computing system, a data selection request from a remote computing device indicating a set of data sources; for each data source of the set of data sources: generating a random data sample, a fraud data sample, and a synthetic data sample, wherein: the random data sample comprises a subset of data from the respective data source; the fraud data sample comprises a set of labelled data from the respective data source; and the synthetic data sample comprises randomized data from the respective data source, generating a set of metrics based on the random data sample, the fraud data sample, and the synthetic data sample, and based on the set of metrics, determining a data source score; selecting, by the computing system, a data source from the set of data sources based at least in part on the data source score; and transmitting, from the computing system to the remote computing device, a responsive message comprising at least an indication of the selected data source.
9. The method of claim 8, wherein the method further comprises: determining a set of concurrence scores, wherein each concurrence score is determined for a pair of data sources in the set of data sources by: generating a trust rate for each demographic group present in each respective data sample, and for each pair of data sources, comparing the trust rate associated with each data source of the pair of data sources; and transmitting, to the remote computing device, a set of concurrence scores associated with the selected data source.
10. The method of claim 8, wherein set of metrics comprises a fraud score, and wherein the fraud score is generated based on a comparison of identities in each respective random data sample with identities in the fraud data sample.
11. The method of claim 8, wherein the set of metrics comprises a consistency score, and wherein the consistency score is generated based on repeatability of an outcome calculated for an entity in the random data sample.
12. The method of claim 8, wherein the set of metrics comprises a sensitivity score, and wherein the sensitivity score is generated based on a comparison of a first trust rate determined from a random data sample of each respective data source with a second trust rate determined from a synthetic data sample of each respective data source.
13. The method of claim 8, wherein the method further comprises: generating, by the computing device, a risk score using data from the selected data source; and transmitting, by the computing device to the remote computing device, the risk score.
14. The method of claim 8, wherein the responsive message is configured to display, via an interface of the remote computing device, a graphical user interface (GUI) comprising the set of metrics and the data source score associated with each data source of the set of data sources.
15. A non-transitory computer-readable medium comprising instructions that are executable by a processor for causing the processor to perform operations comprising: receiving, at a computing system, a data selection request from a remote computing device indicating a set of data sources; for each data source of the set of data sources: generating a random data sample, a fraud data sample, and a synthetic data sample, wherein:the random data sample comprises a subset of data from the respective data source; the fraud data sample comprises a set of labelled data from the respective data source; and the synthetic data sample comprises randomized data from the respective data source, generating a set of metrics based on the random data sample, the fraud data sample, and the synthetic data sample, and based on the set of metrics, determining a data source score; selecting, by the computing system, a data source from the set of data sources based at least in part on the data source score; and transmitting, from the computing system to the remote computing device, a responsive message comprising at least an indication of the selected data source.
16. The non-transitory computer-readable medium of claim 15, wherein the set of metrics comprises a coverage score based on a data sample of the respective data source, and wherein the coverage score comprises a percentage of a population represented by the data sample of the respective data source.
17. The non-transitory computer-readable medium of claim 15, wherein the set of metrics comprises a risk score based on a data sample of the respective data source, and wherein the risk score is generated based on a percentage of identities in the data sample of the respective data source that are associated with at least one risk factor.
18. The non-transitory computer-readable medium of claim 15, wherein the set of metrics comprises a verification score based on a data sample of the respective data source, and wherein the verification score is generated based on a number of verification features used to authenticate an identity in the data sample of the respective data source.
19. The non-transitory computer-readable medium of claim 15, wherein the set of metrics comprises an affiliation score based on a data sample of the respective data source, andwherein the affiliation score is generated based on a number of affiliation features used to authenticate an identity in the data sample of the respective data source.
20. The non-transitory computer-readable medium of claim 15, wherein the operations further comprise: generating a risk score using data from the selected data source; and transmitting, to the remote computing device, the risk score.