Machine-learning techniques with large graphs
By densifying semi-labeled graphs with synthetic edges and reducing graph size, the method enhances the capability of machine-learning models to predict risk indicators for secure access control, addressing challenges of sparsity and explosion in large graph data processing.
Patent Information
- Application Number
- PCT/US2024/061747
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-03
AI Technical Summary
Machine-learning models struggle to effectively process and train on large, semi-labeled graph data due to issues such as graph sparsity and explosion, limiting their ability to accurately predict risk indicators for entities accessing secured resources.
The method involves generating synthetic edges to densify semi-labeled graphs, using language models to connect isolated nodes and reduce graph size through mapping similar nodes, enabling training on large graph representations.
This approach improves the performance and accuracy of machine-learning models in predicting risk indicators by efficiently handling large graph data, facilitating secure access control to computing environments.
Smart Images

Figure US2024061747_03072025_PF_FP_ABST
Abstract
Description
MACHINE-LEARNING TECHNIQUES WITH LARGE GRAPHSCross- References to Related Applications
[0001] This application claims the benefit of U.S. Provisional Patent Application no. 63 / 615,908, filed December 29, 2023, the full disclosure of which is incorporated herein by reference in its entirety for all purposes.Technical Field
[0002] The present disclosure relates generally to artificial intelligence. More specifically, but not by way of limitation, this disclosure relates to machine-learning models for predicting risk and controlling access to secured resources.Background
[0003] Graph processing can be used to analyze large data sets and generate predictions based on that data. A graph may include a number of connected and unconnected nodes, where nodes are connected to each other by edges. Graph processing has a number of applications in data science such as label propagation, graph partitioning, node classification, risk prediction, and so on. Machine-learning models, including neural networks, can be trained on graph data to identify relationships within large data sets.Summary
[0004] Various aspects of the present disclosure provide systems and methods for training a machine-learning model for risk assessment and outcome prediction using large graph representations. The machine-learning model is trained to compute a risk indicator from predictor variables for a target entity. In some aspects, the trained machine-learning model may generate, for the target entity, the risk indicator based on the machine-learning model trained on a similarity graph having nodes, edges, and synthetic edges. The machine-learning model may further transmit, to a remote computing device, a responsive message including at least the risk indicator for use in controlling access of the target entity to one or more interactive computing environments.
[0005] The training of the machine-learning model further involves generating training data. The training data is generated by a process that includes receiving a data set including transaction data. In some aspects, the process includes generating a similarity graph having a set of nodes including a subset of nodes connected by edges, where each node represents an entity with which a transaction occurred. In some aspects, the process can further includeidentifying one or more isolated nodes. An isolated node can be a node that is not connected to a node of the set of nodes by an edge. In some aspects, for each of the one or more isolated nodes, the process may include creating a synthetic edge between the isolated node and the node of the subset of nodes based on a similarity between the isolated node and the node of the subset of nodes.
[0006] In some aspects, the machine-learning model can be used to predict risk indicators. For example, a risk assessment query for a target entity can be received from a remote computing device. In response to the assessment query, an output risk indicator for the target entity can be computed by applying the machine-learning model to predictor variables associated with the target entity. A responsive message including at least the output risk indicator can be transmitted to the remote computing device.
[0007] Additional aspects of the present disclosure provide systems and methods for training a machine-learning model for risk assessment and outcome prediction using bi-partite graph representations. The machine-learning model is trained to compute a risk indicator from predictor variables for a target entity. In some aspects, the trained machine-learning model may generate, for the target entity, the risk indicator based on the machine-learning model trained on a bi-partite similarity graph. The machine-learning model may further transmit, to a remote computing device, a responsive message including at least the risk indicator for use in controlling access of the target entity to one or more interactive computing environments.
[0008] The training of the machine-learning model further involves generating training data. The training data includes a bi-partite similarity graph including a first set of nodes and a second set of nodes. The bi-partite similarity graph can include a first set of edges connects nodes of the first set of nodes and represents similarities between the connected nodes of the first set of nodes. The bi-partite similarity graph can also include a second set of edges connects nodes of the first set of nodes to nodes of the second set of nodes and represents relationships between the respective connected nodes. The bi-partite similarity graph can further include a third set of edges connects nodes of the second set of nodes and represents similarities between the connected nodes of the second set of nodes.
[0009] In some aspects, the machine-learning model can be used to predict risk indicators. For example, a risk assessment query for a target entity can be received from a remote computing device. In response to the assessment query, an output risk indicator for the target entity can be computed by applying the machine-learning model to predictor variables associated with the target entity. A responsive message including at least the output risk indicator can be transmitted to the remote computing device.
[0010] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification, any or all drawings, and each claim.
[0011] The foregoing, together with other features and examples, will become more apparent upon referring to the following specification, claims, and accompanying drawings.Brief Description of the Drawings
[0012] FIG. 1 is a block diagram depicting an example of a computing environment in which a machine-learning model can be trained and applied in a risk assessment application according to certain aspects of the present disclosure.
[0013] FIG. 2 is a flow chart depicting an example of a process for utilizing a machinelearning model to generate risk indicators for a target entity based on predictor variables associated with the target entity according to certain aspects of the present disclosure.
[0014] FIG. 3 is a flow chart depicting an example of a process for generating training data for a machine-learning model according to certain aspects of the present disclosure.
[0015] FIG. 4 is a diagram depicting an example of label propagation in graph data according to certain aspects of the present disclosure.
[0016] FIG. 5 is a diagram depicting an example of a heterogeneous network according to certain aspects of the present disclosure.
[0017] FIG. 6 is a flow chart depicting another example of a process for generating training data for a machine-learning model according to certain aspects of the present disclosure.
[0018] FIG. 7 is a diagram depicting an example of merging transactions to mitigate graph explosion when generating training data for a machine-learning model according to certain aspects of the present disclosure.
[0019] FIG. 8 is a block diagram depicting an example of a computing system suitable for implementing aspects of the techniques and technologies presented herein.Detailed Description
[0020] Machine-learning techniques can provide powerful predictive insights. In one example, complex data can be used to predict a risk associated with an entity accessing asecured resource such as a computing environment. To make such predictions, a machinelearning model may be trained on large and complex data sets. For example, a machinelearning model, such as a graph neural network (GNN), can make determinations based on a graph generated from a large set of data points. Disclosed examples provide systems and methods for training a machine-learning model using large graph representations. Further, disclosed examples facilitate the handling of large graphs, which is often difficult due to the scale of such graphs.
[0021] In some cases, machine-learning algorithms may be limited in their ability to determine patterns directly from graph data if the data from which the graph was generated (e.g., a large set of tabularized data) lacks labels or lacks complete labels. In semi-supervised learning, machine-learning algorithms infer the labels of unlabeled data from their direct and indirect relationships with labelled data. Certain aspects described herein for training a machine-learning model for risk assessment or other outcome predictions based on semilabelled graph data can address one or more issues identified above.
[0022] In one example, a machine-learning algorithm can categorize millions or billions of semi-labeled data points to yield a large graph of millions of nodes and billions of edges. Processing and storing such a graph are infeasible with the computing and storage resources available to many organizations. Further, as additional data points are collected, the generated graph will continue to grow in size over time. Disclosed examples describe the construction of a graph through graph densification and mapping, which overcomes graph sparsity (i.e., where not all nodes of the graph are related to some part of the graph and are isolated) and graph explosion (i.e., growth of the live graph).
[0023] To overcome the above-described problems associated with training machinelearning models using large sets of tabularized data for applications in risk assessment, aspects of the disclosure provide for processes to generate synthetic edges to densify a semi-labelled graph. Aspects of the disclosure can be used to represent large data sets, including tabularized data sets, as large graphs, which enables a machine-learning model to be trained using this data, thereby improving the machine-learning model outcomes and accuracy. For example, semilabelled data may be received by a computing system. From this data, the computing system can generate a similarity graph where each node represents a keyword associated with data in the data set. Each node may also be associated with a label and a frequency. The edges of the similarity graph may connect related nodes. For example, each edge of the similarity graph may be based on a threshold degree of similarity between the connected nodes. The similarity can be measured, for example, by comparing the keywords or other data associated with eachnode to determine a measure of relatedness. The computing system can further determine one or more clusters of nodes where a node’s inclusion in a cluster is based on its distance from a centroid. Each cluster may represent a category of nodes.
[0024] The computing system can also identify isolated nodes that are not connected via an edge to another node of the graph. These isolated nodes can make it difficult for a machinelearning model, such as a GNN, to perform information propagation over a node’s neighborhood (e.g., the nodes connected to a particular node). To overcome this difficulty, a language model can be used to extract node representations and use these representations to compute similarities between an isolated node and the other nodes of the graph. In some examples, a tunable hyperparameter can be used to filter the most similar pairs of nodes such that synthetic edges can be created between the isolated nodes and a respective most similar node where the similarity is above a threshold associated with the hyperparameter. The hyperparameter can be based on, for example, a desired graph density or degree of connectedness.
[0025] Further, the computing system can mitigate graph explosion by mapping similar nodes onto each other to reduce memory usage and reduce complexity. In such a case, one node may be used to represent the nodes of a particular category. In another example, one node may be used to represent a subset of nodes having a particular level of similarity between each other. This level of similarity may be determined, for example, based on a desired density of the resulting graph.
[0026] Once graph densification is complete and synthetic edges are created to connect isolated nodes to the graph, the graph may be used to train a machine-learning model (e.g., a GNN). The trained machine-learning model can be used to generate a risk indicator for controlling access of a target entity to an interactive computing environment. The above examples can be used to determine a risk indicator for a target entity based on incomplete or semi-labelled data. In other examples, an adjacency matrix can be created from the generated graph. This adjacency matrix can be used by the GNN to determine node embeddings for the generated graph, which can be used for further graph analysis.
[0027] Accordingly, certain aspects of the disclosed systems and methods improve the performance of machine-learning models, which were previously unable to be trained on large sets of data. Disclosed systems and methods enable the generation of a large graph representation of the tabularized data, such that a machine-learning model can be trained using the tabularized data set. In some aspects, disclosed systems and methods describe tuning of graph density via one or more hyperparameters to control the size and complexity of the graphrepresentation, thereby facilitating use of a large graph representation to train a machinelearning model efficiently.
[0028] Further, semi-labelled graph data can be used to train a machine-learning model. The semi-labeled graph may be supplemented, as described below, through label propagation. Training a machine-learning model on the supplemented graph can improve machine-learning outcomes as the machine-learning model does not have to infer information about unlabeled nodes.
[0029] These illustrative examples are given to introduce the reader to the general subject matter discussed here and are not intended to limit the scope of the disclosed concepts. The following sections describe various additional features and examples with reference to the drawings in which like numerals indicate like elements, and directional descriptions are used to describe the illustrative examples but, like the illustrative examples, should not be used to limit the present disclosure.
[0030] Referring now to the drawings, FIG. l is a block diagram depicting an example of an operating environment 100 in which a risk assessment computing system 130 builds and trains a machine-learning model that can be used to predict risk indicators based on predictor variables. FIG. 1 depicts examples of hardware components of a risk assessment computing system 130, according to some aspects. The risk assessment computing system 130 is a specialized computing system that may be used for processing large amounts of data using a large number of computer processing cycles. The risk assessment computing system 130 can include a model training server 110 for building and training a machine-learning model 120. The risk assessment computing system 130 can further include a risk assessment server 118 for performing a risk assessment for given predictor variables 124 using the trained machinelearning model 120.
[0031] The model training server 110 can include one or more processing devices that execute program code, such as a model training application 112. The program code is stored on a non-transitory computer-readable medium. The model training application 112 can execute one or more processes to train a machine-learning model for predicting risk indicators based on predictor variables 124.
[0032] In some examples, the machine-learning model training samples 126 can be generated from graph data 142 associated with various entities, such as users or organizations. The graph data 142 can include attributes of each of the entities. For example, the graph data 142 can include a graph generated based on a semi-labeled data set. The graph data for each entity can be represented as a similarity graph having nodes and edges. In some scenarios, thegraph data 142 includes a graph generated from a large-scale data set, such as 200 million records. The graph data 142 can also be stored in the risk data repository 122.
[0033] To generate the training samples 126, the model training server 110 can execute a graph processing application 140 that can perform label propagation on graph data from the graph data 142. For example, the graph data 142 can include a graph having a number of unlabeled isolated nodes where each node represents a transaction. An unlabeled node can represent a transaction where transaction data (e.g., data indicating a transacting entity) is not available. The graph processing application 140 can be configured to determine, for each pair of nodes in a graph, a similarity between the nodes based on keywords or metadata associated with each node. The similar nodes may be connected by an edge based on whether the respective similarity between each pair of nodes is above a predetermined threshold. The graph processing application 140 may also identify isolated nodes, which are not connected to any other node of the graph. A language model may be used to determine similarities between an isolated node and each connected node of the graph. For example, the isolated node can be compared with a cluster centroid to determine a cluster of nodes with which the isolated node shares similarity. The isolated node may be connected via a synthetic edge to a node of the cluster with which it has the greatest calculated similarity. The generated graph can be stored as the model training samples 126. In some aspects, the label associated with the node to which the isolated node is connected may be propagated along the newly created synthetic edge and applied to the isolated node.
[0034] In some aspects, the model training application 112 can build and train a machinelearning model 120 utilizing model training samples 126. The model training samples 126 can include multiple training vectors consisting of training predictor variables and training risk indicator outputs corresponding to the training vectors. The model training samples 126 can be stored in one or more network-attached storage units on which various repositories, databases, or other structures are stored. Examples of these data structures are the risk data repository 122.
[0035] Network-attached storage units may store a variety of different types of data organized in a variety of different ways and from a variety of different sources. For example, the network-attached storage unit may include storage other than primary storage located within the model training server 110 that is directly accessible by processors located therein. In some aspects, the network-attached storage unit may include secondary, tertiary, or auxiliary storage, such as large hard drives, servers, virtual memory, among other types. Storage devices may include portable or non-portable storage devices, optical storage devices, and various othermediums capable of storing and containing data. A machine-readable storage medium or computer-readable storage medium may include a non-transitory medium in which data can be stored and that does not include carrier waves or transitory electronic signals. Examples of a non-transitory medium may include, for example, a magnetic disk or tape, optical storage media such as a compact disk or digital versatile disk, flash memory, memory, or memory devices.
[0036] The risk assessment server 118 can include one or more processing devices that execute program code, such as a risk assessment application 114. The program code is stored on a non-transitory computer-readable medium. The risk assessment application 114 can execute one or more processes to utilize the machine-learning model 120 trained by the model training application 112 to predict risk indicators based on input predictor variables 124.
[0037] The output of the trained machine-learning model 120 can be used to modify a data structure in the memory or a data storage device. For example, the predicted risk indicator and / or the explanation codes can be utilized to reorganize, flag, or otherwise change the predictor variables 124 involved in the prediction by the machine-learning model 120. For instance, predictor variables 124 stored in the risk data repository 122 can be attached with flags indicating their respective amount of impact on the risk indicator. Different flags can be utilized for different predictor variables 124 to indicate different levels of impacts. Additionally, or alternatively, the locations of the predictor variables 124 in the storage, such as the risk data repository 122, can be changed so that the predictor variables 124 or groups of predictor variables 124 are ordered, ascendingly or descendingly, according to their respective amounts of impact on the risk indicator.
[0038] By modifying the predictor variables 124 in this way, a more coherent data structure can be established which enables the data to be searched more easily. In addition, further analysis of the machine-learning model 120 and the outputs of the machine-learning model 120 can be performed more efficiently. For instance, predictor variables 124 having the most impact on the risk indicator can be retrieved and identified more quickly based on the flags and / or their locations in the risk data repository 122. Further, updating the machine-learning model 120, such as re-training the machine-learning model 120 based on new values of the predictor variables 124, can be performed more efficiently especially when computing resources are limited. For example, updating or retraining the machine-learning model 120 can be performed by incorporating new values of the predictor variables 124 having the most impact on the output risk indicator based on the attached flags without utilizing new values of all the predictor variables 124.
[0039] Furthermore, the risk assessment computing system 130 can communicate with various other computing systems, such as client computing systems 104. For example, client computing systems 104 may send risk assessment queries to the risk assessment server 118 for risk assessment, or may send signals to the risk assessment server 118 that control or otherwise influence different aspects of the risk assessment computing system 130. The client computing systems 104 may also interact with user computing systems 106 via one or more public data networks 108 to facilitate interactions between users of the user computing systems 106 and interactive computing environments provided by the client computing systems 104.
[0040] Each client computing system 104 may include one or more third-party devices, such as individual servers or groups of servers operating in a distributed manner. A client computing system 104 can include any computing device or group of computing devices operated by a seller, lender, or other providers of products or services. The client computing system 104 can include one or more server devices. The one or more server devices can include or can otherwise access one or more non-transitory computer-readable media. The client computing system 104 can also execute instructions that provide an interactive computing environment accessible to user computing systems 106. Examples of the interactive computing environment include a mobile application specific to a particular client computing system 104, a web-based application accessible via a mobile device, etc. The executable instructions are stored in one or more non-transitory computer-readable media.
[0041] The client computing system 104 can further include one or more processing devices that are capable of providing the interactive computing environment to perform operations described herein. The interactive computing environment can include executable instructions stored in one or more non-transitory computer-readable media. The instructions providing the interactive computing environment can configure one or more processing devices to perform operations described herein. In some aspects, the executable instructions for the interactive computing environment can include instructions that provide one or more graphical interfaces. The graphical interfaces are used by a user computing system 106 to access various functions of the interactive computing environment. For instance, the interactive computing environment may transmit data to and receive data from a user computing system 106 to shift between different states of the interactive computing environment, where the different states allow one or more electronics transactions between the user computing system 106 and the client computing system 104 to be performed.
[0042] In some examples, a client computing system 104 may have other computing resources associated therewith (not shown in FIG. 1), such as server computers hosting andmanaging virtual machine instances for providing cloud computing services, server computers hosting and managing online storage resources for users, server computers for providing database services, and others. The interaction between the user computing system 106 and the client computing system 104 may be performed through graphical user interfaces presented by the client computing system 104 to the user computing system 106, or through an application programming interface (API) calls or web service calls.
[0043] A user computing system 106 can include any computing device or other communication device operated by a user, such as a consumer or a customer. The user computing system 106 can include one or more computing devices, such as laptops, smartphones, and other personal computing devices. A user computing system 106 can include executable instructions stored in one or more non-transitory computer-readable media. The user computing system 106 can also include one or more processing devices that are capable of executing program code to perform operations described herein. In various examples, the user computing system 106 can allow a user to access certain online services from a client computing system 104 or other computing resources, to engage in mobile commerce with a client computing system 104, to obtain controlled access to electronic content hosted by the client computing system 104, etc.
[0044] For instance, the user can use the user computing system 106 to engage in an electronic transaction with a client computing system 104 via an interactive computing environment. An electronic transaction between the user computing system 106 and the client computing system 104 can include, for example, the user computing system 106 being used to request online storage resources managed by the client computing system 104, acquire cloud computing resources (e.g., virtual machine instances), and so on. An electronic transaction between the user computing system 106 and the client computing system 104 can also include, for example, query a set of sensitive or other controlled data, access online financial services provided via the interactive computing environment, submit an online credit card application or other digital application to the client computing system 104 via the interactive computing environment, operating an electronic tool within an interactive computing environment hosted by the client computing system (e.g., a content-modification feature, an application-processing feature, etc.).
[0045] In some aspects, an interactive computing environment implemented through a client computing system 104 can be used to provide access to various online functions. As a simplified example, a website or other interactive computing environment provided by an online resource provider can include electronic functions for requesting computing resources,online storage resources, network resources, database resources, or other types of resources. In another example, a website or other interactive computing environment provided by a financial institution can include electronic functions for obtaining one or more financial services, such as loan application and management tools, credit card application and transaction management workflows, electronic fund transfers, etc. A user computing system 106 can be used to request access to the interactive computing environment provided by the client computing system 104, which can selectively grant or deny access to various electronic functions. Based on the request, the client computing system 104 can collect data associated with the user and communicate with the risk assessment server 118 for risk assessment. Based on the risk indicator predicted by the risk assessment server 118, the client computing system 104 can determine whether to grant the access request of the user computing system 106 to certain features of the interactive computing environment.
[0046] In a simplified example, the system depicted in FIG. 1 can configure a machinelearning model to be used for accurately determining risk indicators, such as credit scores, and using predictor variables and determining adverse action codes or other explanation codes for the predictor variables. A predictor variable can be any variable predictive of risk that is associated with an entity. Any suitable predictor variable that is authorized for use by an appropriate legal or regulatory framework may be used.
[0047] Examples of predictor variables used for predicting the risk associated with an entity accessing online resources include, but are not limited to, variables indicating the demographic characteristics of the entity (e.g., name of the entity, the network or physical address of the company, the identification of the company, the revenue of the company), variables indicative of prior actions or transactions involving the entity (e.g., past requests of online resources submitted by the entity, the amount of online resource currently held by the entity, and so on.), variables indicative of one or more behavioral traits of an entity (e.g., the timeliness of the entity releasing the online resources), etc. Similarly, examples of predictor variables used for predicting the risk associated with an entity accessing services provided by a financial institute include, but are not limited to, indicative of one or more demographic characteristics of an entity (e.g., age, gender, income, etc.), variables indicative of prior actions or transactions involving the entity (e.g., information that can be obtained from credit files or records, financial records, consumer records, or other data about the activities or characteristics of the entity), variables indicative of one or more behavioral traits of an entity, etc.
[0048] In some aspects, predictor variables can be extracted from a labelled graph. For example, predictor variables can be derived from keywords and metadata associated with eachnode of the graph and from the relationships between the nodes of the graph as indicated by edges and synthetic edges between the nodes.
[0049] The predicted risk indicator can be used by the service provider to determine the risk associated with the entity accessing a service provided by the service provider, thereby granting or denying access by the entity to an interactive computing environment implementing the service. For example, if the service provider determines that the predicted risk indicator is lower than a threshold risk indicator value, then the client computing system 104 associated with the service provider can generate or otherwise provide access permission to the user computing system 106 that requested the access. The access permission can include, for example, cryptographic keys used to generate valid access credentials or decryption keys used to decrypt access credentials. The client computing system 104 associated with the service provider can also allocate resources to the user and provide a dedicated web address for the allocated resources to the user computing system 106, for example, by adding it in the access permission. With the obtained access credentials and / or the dedicated web address, the user computing system 106 can establish a secure network connection to the computing environment hosted by the client computing system 104 and access the resources via invoking API calls, web service calls, HTTP requests, or other proper mechanisms.
[0050] Each communication within the operating environment 100 may occur over one or more data networks, such as a public data network 108, a network 116 such as a private data network, or some combination thereof. A data network may include one or more of a variety of different types of networks, including a wireless network, a wired network, or a combination of a wired and wireless network. Examples of suitable networks include the Internet, a personal area network, a local area network (“LAN”), a wide area network (“WAN”), or a wireless local area network (“WLAN”). A wireless network may include a wireless interface or a combination of wireless interfaces. A wired network may include a wired interface. The wired or wireless networks may be implemented using routers, access points, bridges, gateways, or the like, to connect devices in the data network.
[0051] The number of devices depicted in FIG. 1 is provided for illustrative purposes. Different numbers of devices may be used. For example, while certain devices or systems are shown as single devices in FIG. 1, multiple devices may instead be used to implement these devices or systems. Similarly, devices or systems that are shown as separate, such as the model training server 110 and the risk assessment server 118, may be instead implemented in a signal device or system.
[0052] FIG. 2 is a flow chart depicting an example of a process 200 for using a machinelearning model to generate risk indicators for a target entity based on predictor variables associated with the target entity. One or more computing devices (e.g., the risk assessment server 118) implement operations depicted in FIG. 2 by executing suitable program code (e.g., the risk assessment application 114). For illustrative purposes, the process 200 is described with reference to certain examples depicted in the figures. Other implementations, however, are possible.
[0053] At block 202, the process 200 involves receiving a risk assessment query for a target entity from a remote computing device, such as a computing device associated with the target entity requesting the risk assessment. The risk assessment query can also be received by the risk assessment server 118 from a remote computing device associated with an entity authorized to request risk assessment of the target entity.
[0054] At operation 204, the process 200 involves accessing a machine-learning model trained to generate risk indicator values based on input predictor variables or other data suitable for assessing risks associated with an entity, such as graph data or other predictive data generated based on graph data. Examples of predictor variables can include data associated with an entity that describes prior actions or transactions involving the entity (e.g., information that can be obtained from credit files or records, financial records, consumer records, or other data about the activities or characteristics of the entity), behavioral traits of the entity, demographic traits of the entity, or any other traits that may be used to predict risks associated with the entity. In some aspects, predictor variables can be obtained from credit files, financial records, consumer records, etc. In other aspects, predictor variables can be generated based on analysis of graph data. For example, predictor variables could be generated based on analysis of related transaction records stored as a graph. The risk indicator can indicate a level of risk associated with the entity, such as a credit score of the entity.
[0055] The machine-learning model can be constructed and trained based on training samples including training predictor variables and training risk indicator outputs. Additional details regarding generating training data to train the machine-learning model will be presented below with regard to FIGS. 3 and 4.
[0056] At operation 206, the process 200 involves applying the machine-learning model to generate a risk indicator for the target entity specified in the risk assessment query. Predictor variables associated with the target entity can be used as inputs to the machine-learning model. The predictor variables associated with the target entity can be obtained from a predictor variable database configured to store predictor variables associated with various entities. Theoutput of the machine-learning model would include the risk indicator for the target entity based on its current predictor variables.
[0057] At operation 208, the process 200 involves generating and transmitting a response to the risk assessment query. The response can include the risk indicator generated using the machine-learning model. The risk indicator can be used for one or more operations that involve performing an operation with respect to the target entity based on a predicted risk associated with the target entity. In one example, the risk indicator can be used to control access to one or more interactive computing environments by the target entity. As discussed above with regard to FIG. 1, the risk assessment computing system 130 can communicate with client computing systems 104, which may send risk assessment queries to the risk assessment server 118 to request risk assessment. The client computing systems 104 may be associated with technological providers, such as cloud computing providers, online storage providers, or financial institutions such as banks, credit unions, credit-card companies, insurance companies, or other types of organizations. The client computing systems 104 may be implemented to provide interactive computing environments for customers to access various services offered by these service providers. Customers can utilize user computing systems 106 to access the interactive computing environments thereby accessing the services provided by these providers.
[0058] For example, a customer can submit a request to access the interactive computing environment using a user computing system 106. Based on the request, the client computing system 104 can generate and submit a risk assessment query for the customer to the risk assessment server 118. The risk assessment query can include, for example, an identity of the customer and other information associated with the customer that can be utilized to generate predictor variables. The risk assessment server 118 can perform a risk assessment based on predictor variables generated for the customer and return the predicted risk indicator to the client computing system 104.
[0059] Based on the received risk indicator, the client computing system 104 can determine whether to grant the customer access to the interactive computing environment. If the client computing system 104 determines that the level of risk associated with the customer accessing the interactive computing environment and the associated technical or financial service is too high, the client computing system 104 can deny access by the customer to the interactive computing environment. Conversely, if the client computing system 104 determines that the level of risk associated with the customer is acceptable, the client computing system 104 can grant access to the interactive computing environment by the customer and the customer wouldbe able to use the various services provided by the service providers. For example, with the granted access, the customer can utilize the user computing system 106 to access clouding computing resources, online storage resources, web pages or other user interfaces provided by the client computing system 104 to execute applications, store data, query data, submit an online digital application, operate electronic tools, or perform various other operations within the interactive computing environment hosted by the client computing system 104.
[0060] In other examples, the machine-learning model can also be used to generate adverse action codes or other explanation codes for the predictor variables. Adverse action code can indicate an effect or an amount of impact that a predictor variable has or a group of predictor variables have on the value of the risk indicator, such as credit score (e.g., the relative negative impact of the predictor variable(s) on a risk indicator such as the credit score). In some aspects, the risk assessment application uses the machine-learning model to provide adverse action codes that are compliant with regulations, business policies, or other criteria used to generate risk evaluations. Examples of regulations to which the machine-learning model conforms and other legal requirements include the Equal Credit Opportunity Act (“ECOA”), Regulation B, and reporting requirements associated with ECOA, the Fair Credit Reporting Act (“FCRA”), the Dodd-Frank Act, and the Office of the Comptroller of the Currency (“OCC”).
[0061] In some implementations, the explanation codes can be generated for a subset of the predictor variables that have the highest impact on the risk indicator. For example, the risk assessment application 114 can determine the rank of each predictor variable based on the impact of the predictor variable on the risk indicator. A subset of the predictor variables including a certain number of highest-ranked predictor variables can be selected and explanation codes can be generated for the selected predictor variables. The risk assessment application 114 may provide recommendations to a target entity based on the generated explanation codes. The recommendations may indicate one or more actions that the target entity can take to improve the risk indicator (e.g., improve a credit score).
[0062] In some aspects, process 200 can include training the machine-learning model on bi-partite similarity graph data. A bi-partite similarity graph can include nodes representing receiving systems and nodes representing transmitting systems. A first set of edges can connect the nodes representing receiving systems and a second set of edges can connect nodes representing transmitting systems with the respective nodes representing the receiving systems to which they have transmitted data. Thus, the machine-learning model is trained on a graph containing transaction data, as well as data associated with additional information about the transmitting systems and the associated transactions with receiving systems.
[0063] Referring now to FIG. 3, a flow chart depicting an example of a process 300 for propagating information in semi-labeled graph data is presented. One or more computing devices (e.g., the model training server 110) implement operations depicted in FIG. 3 by executing suitable program code (e.g., the graph processing application 140). For illustrative purposes, the process 300 is described with reference to certain examples depicted in the figures. Other implementations, however, are possible.
[0064] At block 302, the process 300 involves the graph processing application 140 receiving a data set including one or more keywords. For example, the graph processing application 140 can receive raw transaction data in which each row represents a transaction. Each transaction record can include one or more keywords, such as an entity name, transaction amount, transaction type, etc. As used herein, a transaction can be an interaction between two transacting entities. For example, a transaction can be a merchant-consumer exchange of money for goods. In another example, a transaction can include providing a resource (e.g., data storage) to a computing system. Each transaction record may be associated with metadata such as a transaction amount, an entity location, and a processing provider, among others.
[0065] At block 304, the process 300 involves the graph processing application 140 generating a similarity graph from the received data set. The similarity graph may have a set of nodes, e.g., representing each transacting entity, and edges connecting a subset of nodes. The similarity may be based on a pre-trained language model, such as a bidirectional encoder representations and transformers (BERT) language model, comparing representations of the respective keywords associated with each node to each other. The similarity can be, for example, a percentage or other quantitative measure of similarity between the representations of the keywords. Edges may be created between nodes having a similarity above a given similarity level or above a threshold similarity.
[0066] In some examples, the threshold similarity is based on a hyperparameter associated with a desired graph density or connectedness. For example, the threshold similarity required to create an edge may be lower to create a denser graph, e.g., nodes having lower similarity with each other may be connected, resulting in more connections between nodes and thereby creating a denser graph. In other examples, a desire for greater accuracy of the connections may yield a sparser graph with fewer connections between nodes having a stronger similarity to each other. Accordingly, the hyperparameter may be tuned to determine the threshold similarity thereby controlling the number of connections made in the graph.
[0067] In some aspects, generating the similarity graph can involve the graph processing application 140 extracting, from each node of the set of nodes, a representation of the nodebased on a pre-trained language model. The representation may be, for example, an embedding or feature vector. The graph processing application 140 can compute a pairwise similarity for each pair of nodes in the set of nodes based on the extracted representations. Edges can be created between each pair of nodes having a pairwise similarity greater than a threshold level of similarity. In some examples, each node may be connected to zero other nodes, one other node, or more than one other node based on the node’s similarity with each of the other nodes of the similarity graph.
[0068] In some examples, one or more categories may be determined for various nodes of the similarity graph. A node’s category may be the same as, or related to, its label. The graph processing application 140 can determine a category for each node of the subset of connected nodes based on the label associated with each node. In some examples, a category of nodes may be those nodes within a threshold distance from a centroid that is determined to be a centroid of a cluster of nodes. Those nodes within the threshold distance of the centroid can be associated with the category of the centroid.
[0069] At block 306, the process 300 involves the graph processing application 140 identifying isolated nodes. Isolated nodes of the similarity graph are those which are not connected to any other node of the set of nodes of the similarity graph. In some examples, this can occur when a node is not or cannot be associated with a label from which a similarity to other nodes or a category can be determined. Isolated nodes hinder the machine-learning model’s ability to propagate information and to determine relationships between nodes of the similarity graph.
[0070] At block 308, the process 300 involves the graph processing application 140 creating a synthetic edge between the isolated node and a node of the subset of nodes (i.e., the connected nodes) based on a similarity. For example, the graph processing application 140 may determine one or more clusters of related, labelled nodes. For each cluster, the graph processing application 140 can determine a similarity between the centroid of the cluster and an isolated node. Then, for the cluster having the greatest similarity to the isolated node, a pairwise similarity between the isolated node and each node of the cluster can be determined. As described above, each similarity can be determined using a node representation and a language model to determine similarities between nodes based on, for example, keywords or other metadata associated with each node and contained in the transaction data. In certain examples, the similarity between a pair of nodes may be based on embeddings generated for the similarity graph.
[0071] In an example, for a given isolated node, the graph processing application 140 may determine a similarity for between the isolated node and each other node of the similarity graph. These similarities can be compared to determine the node with which the isolated node shares the greatest similarity. In some embodiments, the isolated nodes may be connected via a synthetic edge to any nodes with which it has a similarity above the similarity threshold. Accordingly, in some examples, an isolated node may be connected via synthetic edge to more than one node of the similarity graph.
[0072] The graph processing application 140 can further propagate labels from the labelled nodes, along the created synthetic edges, to the previously isolated nodes. Accordingly, the received semi-labelled data can be used to create a labelled similarity graph.
[0073] In some examples, categories may be determined for the nodes of the similarity graph once synthetic edges to isolated nodes are generated. To mitigate graph explosion (e.g., to stop unrestricted growth of the graph), the graph processing application 140 can generate a reduced size similarity graph by mapping nodes associated with the same category onto a single node. In some examples, the mapping may be based on nodes having a similarity above a particular similarity threshold with those nodes to which they are connected. Accordingly, connected nodes that have a similarity above the similarity threshold may be mapped onto a single node associated with a category or with a label of the mapped nodes.
[0074] The training samples can be generated based on the generated similarity graph. In some examples, the training samples can be generated based on the reduced size similarity graph. In turn, the model training application 112 may use the training samples to train the machine-learning model 120. In some examples the machine-learning model 120 may be a GNN or a graph convolutional network (GCN). The risk assessment application 114 may use the trained machine-learning model to determine values of predictor variables. The predictor variables can be based, in part, on metadata associated with the nodes of the similarity graph. In another example, predictor variables can be determined by training the machine-learning model 120 on an adjacency matrix generated from the similarity graph. In one example, a GNN trained on a graph of transaction data can predict a category into which a customer falls or may predict particular customer behavior based on the relationships between transactions of the similarity graph.
[0075] The risk assessment application 114 can generate a risk indicator based on the predictor variables. As discussed above, the risk indicator can determine whether or not a target entity is granted access to a secure resource like an interactive computing environment. For example, a risk indicator may be a binary value (e.g., 1 or 0) indicative of whether or not toallow the target entity to access the interactive computing environment. In another example, the risk indicator can be a value and can be compared with a predetermined threshold. If the risk indicator is below the threshold the entity can be denied access to the interactive computing environment.
[0076] FIG. 4 illustrates a diagram depicting an example of a semi-labelled graph. In the illustrated example, each node represents a computing system associated with a data transaction made by the target entity. For example, the target entity may be a networking device configured to route data to one or more computing systems. As shown in FIG. 4, an exemplary similarity graph 400 may include labelled nodes 402A, 402B, 402C. As described above, graph processing application 140 may receive transaction data and analyze the data to generate a similarity graph. For example, the similarity graph may be based on a comparison of keywords (“System A,” “System B,” and “System C”) associated with each of the nodes 402A, 402B, 402C using a language model. In some examples, the label, “Data Type A,” may be based on a mapping of each system to a data type label associated with a data type handled by the system. In other examples, the label may be metadata associated with the received transaction data. Based on the determination that the respective pairs of the nodes 402A, 402B, and 402C have a similarity above a certain threshold, edges representative of the respective system similarities may be drawn.
[0077] In some examples, the nodes 402A, 402B, and 402C may form a cluster 406 associated with the label “Data Type A.” The cluster can be determined by the graph processing application 140 based on, for example, a distance of each of the nodes 402A, 402B, and 402C from a centroid 408. In other examples, nodes sharing the same label can be considered a cluster associated with the particular data type involved in the transaction between the target entity and each system (e.g., “Data Type A”). In another example, the label may be determined by applying a pre-trained language model to data (e.g., keywords or metadata) associated with each node.
[0078] The data used to generate similarity graph 400 may include a node 404 associated with unlabeled transactions associated with a System X. Once the similarity graph is generated for the labelled data, graph processing application 140 may determine that node 404 is an isolated node and is not connected to any other node of the similarity graph 400. The graph processing application 140 may compare a representation of the node 404 with the centroid 408 of cluster 406 and the centroids of other clusters (not shown) of the similarity graph 400. Once the graph processing application 140 identifies a most similar cluster (e.g., cluster 406) and then determines a similarity for each of the nodes 402A, 402B, and 402C, to determine amost similar node of the cluster 406 with the node 404 and each of the nodes of the similarity graph 400. Based on this determination, a synthetic edge can be constructed between the node 404 and the node 402B. For example, node 404 may have a similarity with node 402B that is greater than its similarity with the nodes 402 A and 402C. In another example, the similarity between the node 404 and the node 402B may be above a predetermined threshold based on a desired density or connectedness of the similarity graph 400.
[0079] The generation of the synthetic edge between the node 404 and the node 402B facilitates a number of operations. For example, the label “Data Type A” may now be propagated along the synthetic edge to the node 404. Further, by connecting the isolated node 404 with the node 402B, the graph processing application 140 can perform additional operations on the similarity graph 400, such as graph partitioning, label propagation, and node classification. Accordingly, the similarity graph 400 can be optimized to yield a more accurately trained machine-learning model.
[0080] Additionally, in some examples, to overcome graph explosion, nodes having the same label can be mapped onto a single node. Using the similarity graph 400 as an example, the nodes 402A, 402B, and 402C all are labelled “Data Type A.” Thus, these three nodes can be mapped onto a single node, e.g., the node 402B, and the node 402B can represent all three nodes. Further, once the similarity edge is constructed between the node 404 and the node 402B and the label “Data Type A” is propagated to the node 404, the node 404 can also be mapped onto the node 402B. Accordingly, this method of mapping similarly labelled nodes avoids graph explosion, thereby reducing the amount of storage and processing power required to handle the similarity graph 400.
[0081] In another example, the graph processing application 140 can generate graph embeddings for each cluster of the similarity graph 400. In some examples, the graph embedding for each cluster may be a centroid embedding determined based on the centroid 408 of the cluster 406. The graph processing application 140 can then identify a similar cluster associated with the isolated node 404.
[0082] Once the similarity graph 400 including the synthetic edge is generated, it can be used by the machine-learning model training application 112 as part of the training process for training a machine-learning model to determine a risk indicator. The trained machine-learning model can be used by risk assessment application 114 to determine the risk indicator for the target entity based on the associated predictor variables 124.
[0083] In some examples, the transaction data can be used to generate a heterogeneous, or bi-partite, graph in which the nodes represent either a receiving system or a transmitting system. An example of such a graph 500 is depicted in FIG. 5.
[0084] Graph 500 includes transmitting systems 502A and 502B, as well as a system 402B (i.e., a receiving system) and unlabeled system 404. A transmitting system can be, for example, a system that transmits or otherwise communicates data to another system, referred to herein as a receiving system. In graph 500, there are two types of edges: edges that represent data transactions (e.g., edge 504); and edges that represent similarity between receiving systems (e.g., synthetic similarity edge 506). Each edge representing a transaction (e.g., edge 504) can include transaction metadata such as a frequency of transactions and median size of the transaction. Each transaction edge can include other transaction data or metadata, such as a period of time in which the transaction occurred, or other statistical data associated with transactions between the transmitting system and receiving system connected by the edge. This edge data can be used to infer similarities between transmitting systems and to classify the transmitting systems. For example, graph 500 includes an edge 508 representing an inferred similarity between transmitting system A and transmitting system B.
[0085] In some aspects, the edge 506 can be determined as described above with reference to FIG. 2. Edge 508 can be generated based on a similarity between the transmitting system A and the transmitting system B. For example, a similarity can be determined based on transaction data associated with each transmitting system’s respective transactions with the receiving system B. If the similarity is greater than a predetermined classifier threshold, an edge can be created between the transmitting systems A and B.
[0086] Graph 500 can enable a system (e.g., risk assessment computing system 130), to generate risk predictions based on transmitting system similarity by using a machine-learning model trained on heterogeneous graph data.
[0087] FIG. 6 illustrates a flow chart depicting an example of a process 600 for determining similarities between transmitting systems. One or more computing devices (e.g., the model training server 110) implement operations depicted in FIG. 6 by executing suitable program code (e.g., the graph processing application 140). For illustrative purposes, the process 600 is described with reference to certain examples depicted in the figures. Other implementations, however, are possible.
[0088] At block 602, the process 600 can include receiving a risk assessment query and a similarity graph (e.g., similarity graph 400). For example, process 600 can be performed by the graph processing application 140 to further improve the accuracy of the machine-learningmodel 120 by providing information from which the machine-learning model 120 can predict future transactions, classify transmitting systems, and predict risk.
[0089] In some aspects, in response to receiving the risk assessment query, the graph processing application 140 can generate a similarity graph (e.g., similarity graph 400). In other aspects, the graph processing application 140 can receive the risk assessment query and retrieve (e.g., from the risk data repository 122) a previously generated similarity graph. The risk assessment query can include, for example, a target entity, which may be a transmitting system.
[0090] At block 604, the process 600 can include adding nodes to the similarity graph, where the added nodes represent transmitting systems. For example, the graph processing application 140 can add nodes representing transmitting systems to a similarity graph containing nodes representing receiving systems (e.g., graph 400) to create the bipartite similarity graph. The data and metadata associated with each transmitting system can be included in a tabularized data set from which the similarity graph was generated.
[0091] At block 606, the process 600 can include generating edges between transmitting systems and receiving systems where each edge represents a transaction (e.g., a transfer of data from the transmitting system to the receiving system). In some aspects, to mitigate graph explosion, each edge can represent a number of merged transactions between the transmitting system and respective receiving system. Thus, each edge can represent multiple transactions between the involved parties to reduce the size of the generated bi-partite similarity graph.
[0092] Thus, the bi-partite similarity graph can include two sets of nodes: a first set of nodes representing receiving systems connected by edges representing similarities between the receiving systems; and a second set of nodes representing transmitting systems that are connected to receiving nodes based on transactions with respective receiving systems. In some aspects, the graph processing application 140 can further generate edges between transmitting systems indicating an inferred similarity between the connected transmitting systems. For example, the graph processing application 140 can determine a frequency of transactions between a transmitting system and a receiving system, as well as a median size of the transactions. The frequencies and medians for each transmitting system transacting with a particular receiving system can be compared and analyzed to extract transaction behavior of the transmitting entities. The transaction behavior can be used, for example, to classify transmitting systems and infer similarities between transmitting systems.
[0093] The generated bi-partite similarity graph can be used to train a machine learning model as described above with reference to FIG. 2. Thus, the bi-partite similarity graph can be used to train a machine learning model to output a risk indicator associated with a target entity.
[0094] FIG. 7 is an illustration of an example bi-partite similarity graph 700 generated using the process 600. For example, the bi-partite similarity graph 700 can be generated based on transaction data and can be used to train a machine-learning model to predict risk for a target entity.
[0095] Graph 700 can include a first set of nodes 702, representing receiving systems. The graph processing application 140 can use natural language processing to determine a cluster 704 includes receiving systems 702A, 702B, and 702C based on a similarity between these systems. In some aspects, one or more of the edges 706 between the receiving systems 702A, 702B, and 702C can be a synthetic edge generated by the graph processing application 140 as described above. The graph processing applciation 140 can determine similarities between the receiving systems to generate clusters of similar receiving systems (e.g., cluster 704) based on natural language processing of transaction data as discussed above.
[0096] The bi-partite similarity graph 700 can further include a second set of nodes 708 representing transmitting systems. Edge 710A can represent one or more transactions between transmitting system 708A and receiving system 702B. Edge 710B can represent one or more transactions between transmitting system 708B and receiving system 702B. Edge 710C can represent one or more transactions between transmitting system 708B and receiving system 702C.
[0097] In some aspects, the graph processing application 140 can generate one or more metrics associated with the transactions represented by edges 710A, 710B, and 710C. Metrics can include statistics such as a frequency of transactions and median size of transactions. In some aspects, the median transaction size is used to reduce the influence of outliers on analyses of the transactions. Based on these metrics, the graph processing application 140 can extract transaction behaviors of the transmitting systems 708A and 708B. The extracted transaction behaviors can be used to determine a similarity between the transmitting systems 708A and 708B, which can be represented by edge 712.
[0098] The bi-partite similarity graph 700 can be used to train a machine-learning model (e.g., machine-learning model 120). The information represented in the bi-partite similarity graph 700 enables the machine-learning model to classify systems, predict future transactions, and predict risk. By generating the bi-partite similarity graph 700, the risk assessment computing system 130 can train a machine-learning model on data sets that are typically too cumbersome to feasibly train a machine-learning model.
[0099] Any suitable computing system or group of computing systems can be used to perform the operations for the machine-learning operations described herein. For example,FIG. 8 is a block diagram depicting an example of a computing device 800, which can be used to implement the risk assessment server 118 or the model training server 110. The computing device 800 can include various devices for communicating with other devices in the operating environment 100, as described with respect to FIG. 1. The computing device 800 can include various devices for performing one or more transformation operations described above with respect to FIGS. 1-7.
[0100] The computing device 800 can include a processor 802 that is communicatively coupled to a memory 804. The processor 802 executes computer-executable program code stored in the memory 804, accesses information stored in the memory 804, or both. Program code may include machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, among others.
[0101] Examples of a processor 802 include a microprocessor, an application-specific integrated circuit, a field-programmable gate array, or any other suitable processing device. The processor 802 can include any number of processing devices, including one. The processor 802 can include or communicate with a memory 804. The memory 804 stores program code that, when executed by the processor 802, causes the processor to perform the operations described in this disclosure.
[0102] The memory 804 can include any suitable non-transitory computer-readable medium. The computer-readable medium can include any electronic, optical, magnetic, or other storage device capable of providing a processor with computer-readable program code or other program code. Non-limiting examples of a computer-readable medium include a magnetic disk, memory chip, optical storage, flash memory, storage class memory, ROM, RAM, an ASIC, magnetic storage, or any other medium from which a computer processor can read and execute program code. The program code may include processor-specific program code generated by a compiler or an interpreter from code written in any suitable computerprogramming language. Examples of suitable programming language include Hadoop, C, C++, C#, Visual Basic, Java, Python, Perl, JavaScript, ActionScript, etc.
[0103] The computing device 800 may also include a number of external or internal devices such as input or output devices. For example, the computing device 800 is shown with aninput / output interface 808 that can receive input from input devices or provide output to output devices. A bus 806 can also be included in the computing device 800. The bus 806 can communicatively couple one or more components of the computing device 800.
[0104] The computing device 800 can execute program code 814 that includes the risk assessment application 114, the model training application 112, and / or the graph processing application 140. The program code 814 for the risk assessment application 114 and / or the model training application 112 may be resident in any suitable computer-readable medium and may be executed on any suitable processing device. For example, as depicted in FIG. 8, the program code 814 for the risk assessment application 114, the model training application 112, and / or the graph processing application 140 can reside in the memory 804 at the computing device 800 along with the program data 516 associated with the program code 514, such as the predictor variables 124, the model training samples 126, and / or the graph data 142. Executing the risk assessment application 114, the model training application 112, or the graph processing application 140 can configure the processor 802 to perform the operations described herein.
[0105] In some aspects, the computing device 800 can include one or more output devices. One example of an output device is the network interface device 810 depicted in FIG. 8. A network interface device 810 can include any device or group of devices suitable for establishing a wired or wireless data connection to one or more data networks described herein. Non-limiting examples of the network interface device 810 include an Ethernet network adapter, a modem, etc.
[0106] Another example of an output device is the presentation device 812 depicted in FIG. 8. A presentation device 812 can include any device or group of devices suitable for providing visual, auditory, or other suitable sensory output. Non-limiting examples of the presentation device 812 include a touchscreen, a monitor, a speaker, a separate mobile computing device, etc. In some aspects, the presentation device 812 can include a remote client-computing device that communicates with the computing device 800 using one or more data networks described herein. In other aspects, the presentation device 812 can be omitted.
[0107] The foregoing description of some examples has been presented only for the purpose of illustration and description and is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Numerous modifications and adaptations thereof will be apparent to those skilled in the art without departing from the spirit and scope of the disclosure.
Claims
Claims1. A method that includes one or more processing devices performing operations comprising: determining, using a machine-learning model trained using a training process, a risk indicator for a target entity from predictor variables associated with the target entity, wherein the machine-learning model is trained based on a bi-partite similarity graph comprising: a first set of nodes and a second set of nodes, wherein: a first set of edges connects nodes of the first set of nodes and represents similarities between the connected nodes of the first set of nodes, a second set of edges connects nodes of the first set of nodes to nodes of the second set of nodes and represents relationships between the respective connected nodes, and a third set of edges connects nodes of the second set of nodes and represents similarities between the connected nodes of the second set of nodes; and generating, for the target entity, the risk indicator based on the machine-learning model trained on the bi-partite similarity graph; and transmitting, to a remote computing device, a responsive message comprising at least the risk indicator for use in controlling access of the target entity to one or more interactive computing environments.
2. The method of claim 1, wherein the method further comprises: generating the bi-partite similarity graph by: extracting, from each node of the first set of nodes, a representation of each node based on a pre-trained language model; computing a pairwise similarity for each pair of nodes of the first set of nodes based on the representations of each node; and creating the first set of edges between each pair of nodes having a pairwise similarity greater than a threshold level of similarity.
3. The method of claim 2, wherein the threshold level of similarity is based on a hyperparameter configured to control a density of the bi-partite similarity graph.
4. The method of claim 1, wherein the method further comprises: generating the bi-partite similarity graph by: generating an edge between a first node of the second set of nodes and a second node of the second set of nodes based on a second similarity being above a classifier threshold, wherein the second similarity is based on a first relationship between the first node of the second set of nodes and a second node of the first set of nodes and a second relationship between the second node of the second set of nodes and the second node of the first set of nodes, and generating a reduced size graph based on the bi-partite similarity graph by merging the first node of the second set of nodes and the second node of the second set of nodes.
5. The method of claim 4, wherein the second similarity is determined based on a frequency and a median quantity associated with each of the first relationship and the second relationship.
6. The method of claim 4, wherein the machine-learning model is trained using the reduced size graph.
7. The method of claim 1, wherein the method further comprises: generating the bi-partite similarity graph by: identifying an isolated node of the first set of nodes, wherein the isolated node is not connected to another node of the first set of nodes; determining a similarity between the isolated node and each other node of the first set of nodes; and creating a synthetic edge between the isolated node and each other node of the first set of nodes where the similarity is greater than a synthetic edge similarity threshold.
8. The method of claim 7, wherein generating the bi-partite similarity graph further comprises: propagating a label associated with a connected node to the isolated node along the synthetic edge.
9. The method of claim 1, wherein the machine-learning model is trained using an adjacency matrix associated with the bi-partite similarity graph.
10. The method of claim 1, wherein the machine-learning model comprises a graph neural network (GNN).
11. A system comprising: a processing device; and a memory device in which instructions executable by the processing device are stored for causing the processing device to: determine, using a machine-learning model trained using a training process, a risk indicator for a target entity from predictor variables associated with the target entity, wherein the machine-learning model is trained based on a bi-partite similarity graph comprising: a first set of nodes and a second set of nodes, wherein: a first set of edges connects nodes of the first set of nodes and represents similarities between the connected nodes of the first set of nodes, a second set of edges connects nodes of the first set of nodes to nodes of the second set of nodes and represents relationships between the respective connected nodes, and a third set of edges connects nodes of the second set of nodes and represents similarities between the connected nodes of the second set of nodes; and generating, for the target entity, the risk indicator based on the machine-learning model trained on the bi-partite similarity graph; and transmitting, to a remote computing device, a responsive message comprising at least the risk indicator for use in controlling access of the target entity to one or more interactive computing environments.
12. The system of claim 11, wherein the instructions executable by the processing device further cause the processing device to: generating the bi-partite similarity graph by: extracting, from each node of the first set of nodes, a representation of each node based on a pre-trained language model; computing a pairwise similarity for each pair of nodes of the first set of nodes based on the representations of each node; and creating the first set of edges between each pair of nodes having a pairwise similarity greater than a threshold level of similarity.
13. The system of claim 12, wherein the threshold level of similarity is based on a hyperparameter configured to control a density of the bi-partite similarity graph.
14. The system of claim 11, wherein the instructions executable by the processing device further cause the processing device to: generate the bi-partite similarity graph by: generating an edge between a first node of the second set of nodes and a second node of the second set of nodes based on a second similarity being above a classifier threshold, wherein the second similarity is based on a first relationship between the first node of the second set of nodes and a second node of the first set of nodes and a second relationship between the second node of the second set of nodes and the second node of the first set of nodes, and generating a reduced size graph based on the bi-partite similarity graph by merging the first node of the second set of nodes and the second node of the second set of nodes.
15. The system of claim 14, wherein the second similarity is determined based on a frequency and a median quantity associated with each of the first relationship and the second relationship.
16. A non-transitory computer-readable storage medium having program code that is executable by a processor device to cause a computing device to perform operations, the operations comprising:determining, using a machine-learning model trained using a training process, a risk indicator for a target entity from predictor variables associated with the target entity, wherein the machine-learning model is trained based on a bi-partite similarity graph comprising: a first set of nodes and a second set of nodes, wherein: a first set of edges connects nodes of the first set of nodes and represents similarities between the connected nodes of the first set of nodes, a second set of edges connects nodes of the first set of nodes to nodes of the second set of nodes and represents relationships between the respective connected nodes, and a third set of edges connects nodes of the second set of nodes and represents similarities between the connected nodes of the second set of nodes; and generating, for the target entity, the risk indicator based on the machine-learning model trained on the bi-partite similarity graph comprising nodes, edges, and the synthetic edges; and transmitting, to a remote computing device, a responsive message comprising at least the risk indicator for use in controlling access of the target entity to one or more interactive computing environments.
17. The non-transitory computer-readable storage medium of claim 16, wherein the operations further comprise generating the bi-partite similarity graph by: extracting, from each node of the first set of nodes, a representation of each node based on a pre-trained language model; computing a pairwise similarity for each pair of nodes of the first set of nodes based on the representations of each node; and creating the first set of edges between each pair of nodes having a pairwise similarity greater than a threshold level of similarity.
18. The non-transitory computer-readable storage medium of claim 17, wherein the threshold level of similarity is based on a hyperparameter configured to control a density of the bi-partite similarity graph.
19. The non-transitory computer-readable storage medium of claim 16, wherein the operations further comprise generating the bi-partite similarity graph further by: generating an edge between a first node of the second set of nodes and a second node of the second set of nodes based on a second similarity being above a classifier threshold, wherein the second similarity is based on a first relationship between the first node of the second set of nodes and a second node of the first set of nodes and a second relationship between the second node of the second set of nodes and the second node of the first set of nodes, and generating a reduced size graph based on the bi-partite similarity graph by merging the first node of the second set of nodes and the second node of the second set of nodes.
20. The non-transitory computer-readable storage medium of claim 19, wherein the second similarity is determined based on a frequency and a median quantity associated with each of the first relationship and the second relationship.
Citation Information
Patent Citations
System and method for structure learning for graph neural networks
US20220101103A1
Guided exploration for conversational business intelligence
US20220277031A1
Entity Tag Association Prediction Method, Device, and Computer Readable Storage Medium
US20240419942A1
Entity tag association prediction method and device and computer readable storage medium
WO2023093205A1