Target permutation testing to generate computationally efficient models

Target permutation testing addresses the lack of transparency in deep learning models by evaluating feature significance through permuted data, enhancing interpretability and efficiency by reducing complexity and mitigating overfitting.

WO2026030452A1PCT designated stage Publication Date: 2026-02-05EQUIFAX INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/039889
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-30
Filing Date
2025-07-30
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing machine learning models, particularly deep learning models, lack transparency in evaluating feature importance, leading to computational inefficiencies, overfitting, and reduced generalization capability due to the presence of unimportant features.

Method used

A target permutation testing method that permutes the target variable to evaluate feature significance by training models on permuted data, allowing for the identification and removal of irrelevant features, thereby enhancing model interpretability and computational efficiency.

Benefits of technology

The method improves model interpretability and efficiency by reducing model complexity, mitigating overfitting, and improving generalization, making deep learning models more transparent and trustworthy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025039889_05022026_PF_FP_ABST
    Figure US2025039889_05022026_PF_FP_ABST
Patent Text Reader

Abstract

A computing device can train a model on original data and determine first characteristics of the trained model. The computing device can permute a target variable of the original data by randomly shuffling the target variable to obtain a set of permutations of the original data. The computing device can train a permuted model for each permutation of the set of permutations to generate a set of permuted models, determine second characteristics for each permuted model of the set of permuted models, and compare the second characteristics for each permuted model and the first characteristics of the trained model to determine relevance of a set of features associated with the first characteristics to an output of the trained model. The computing device can adjust the trained model to include only a subset of features that are relevant to the output of the trained model.
Need to check novelty before this filing date? Find Prior Art

Description

TARGET PERMUTATION TESTING TO GENERATE COMPUTATIONALLY EFFICIENT MODELSCross-Reference to Related Applications

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 677,168, filed July 30, 2024, and titled TARGET PERMUTATION TESTING TO GENERATE COMPUTATIONALLY EFFICIENT MODELS, the entirety of which is incorporated herein by reference.Technical Field

[0002] The present disclosure relates generally to permutation testing for modeling. More specifically, but not by way of limitation, this disclosure relates to target permutation testing to identify and select relevant features for use in models to improve computational efficiency.Background

[0003] Statistical modeling has long been foundational to data-driven decision making. Approaches like linear / logistic regression and other generalized linear models are tools for understanding relationships in data. These techniques admit clear, quantifiable metrics that clarify the significance of a model’s features to guide developers in developing more concise, yet more robust and interpretable, models. As the size and complexity of both data and models have grown, the range of data science has expanded beyond the bounds of traditional statistics and the approach to data science has shifted from a “data modeling” culture to an “algorithmic modeling” culture. Consequently, prediction has achieved a perceived importance over inference, and so-called ‘black box’ modeling approaches (e.g., deep learning) have become more prominent. But these models lack transparent mechanisms to evaluate the importance and contributions of individual features, complicating the interpretative process. The opacity of the modeling approaches relative to traditional linear models is due to high capacity and flexibility of the modeling approaches. That is, the approaches are useful in modeling the intricate, complex relationships intrinsic to modern data problems. In a truly black box model where the relationship between the inputs and outputs is completely inscrutable, compliance with various trustworthiness and explainability regulations may be impossible.

[0004] This opacity is not only a hindrance to trustworthiness and explainability, but it also complicates model validation and understanding and restricts the data scientists and machine learning engineers tasked with designing, implementing, and overseeing such systems. Furthermore, the presence of unimportant features is significantly more detrimental to the performance of neural network models than other machine learning models. Therefore, using a meaningful and concise set of inputs may be important for optimizing performance of neural networks while avoiding unnecessary computational costs, mitigating overfitting, and improving generalization capability.Summary

[0005] Various aspects of the present disclosure provide systems and methods for target permutation testing to identify and select relevant features for use in models to improve computational efficiency. In one example, a computer-implemented method includes training, by a processor, a generalized linear model on a set of original data. The method further includes extracting, by the processor, a first set of coefficients from the trained generalized linear model. Additionally, the method includes permuting, by the processor, a target variable of the set of original data by randomly shuffling the target variable to obtain a set of permutations of the set of original data. The method also includes training, by the processor and using a set of training settings and a set of hyperparameters of the trained generalized linear model, a permuted model for each permutation of the set of permutations to generate a set of permuted models. Further, the method includes extracting, by the processor, a second set of coefficients for each permuted model of the set of permuted models. Furthermore, the method includes comparing, by the processor, the second set of coefficients for each permuted model and the first set of coefficients of the trained generalized linear model to determine relevance of a set of features associated with the first set of coefficients to an output of the trained generalized linear model. Moreover, the method includes adjusting, by the processor, the trained generalized linear model to include only a subset of features of the set of features that are determined to be relevant to the output of the trained generalized linear model.

[0006] In an additional example, system includes a processing device and a memory device in which instructions executable by the processing device are stored for causing the processing device to perform operations. The operations include training a neural network model on a set of original data and determining a first set of gradients of the trained neuralnetwork model. The operations further include permuting a target variable of the set of original data by randomly shuffling the target variable to obtain a set of permutations of the set of original data and training, using a set of training settings of the trained neural network model, a permuted model for each permutation of the set of permutations to generate a set of permuted models. Additionally, the operations include determining a second set of gradients for each permuted model of the set of permuted models and comparing the second set of gradients or second test statistics computed from the second set of gradients for each permuted model and the first set of gradients or first test statistics computed from the first set of gradients of the trained neural network model to determine relevance of a set of features associated with the set of gradients to an output of the trained neural network model. Further, the operations include adjusting the trained neural network model to include only a subset of features of the set of features that are determined to be relevant to the output of the trained neural network model.

[0007] In an additional example, a non-transitory computer-readable storage medium has program code that is executable by a processor device to cause the processor device to perform operations. The operations include training a machine-learning model on a set of original data and determining a first set of characteristics of the trained machine-learning model. Additionally, the operations include permuting a target variable of the set of original data by randomly shuffling the target variable to obtain a set of permutations of the set of original data and training, using a set of training settings of the machine-learning model, a permuted model for each permutation of the set of permutations to generate a set of permuted models. Further, the operations include determining a second set of characteristics for each permuted model of the set of permuted models and comparing the second set of characteristics for each permuted model and the first set of characteristics of the trained machine-learning model to determine relevance of a set of features associated with the first set of characteristics to an output of the trained machine-learning model. Furthermore, the operations include adjusting the trained machine-learning model to include only a subset of features of the set of features that are determined to be relevant to an output of the trained machine-learning model.

[0008] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimedsubject matter. The subject matter should be understood by reference to appropriate portions of the entire specification, any or all drawings, and each claim.

[0009] The foregoing, together with other features and examples, will become more apparent upon referring to the following specification, claims, and accompanying drawings.Brief Description of the Drawings

[0010] FIG. 1 is a block diagram depicting an example of a computing environment in which permutation testing can occur, according to certain aspects of the present disclosure.

[0011] FIG. 2 is a flow chart depicting an example of a process for performing permutation testing with a generalized linear model, according to certain aspects of the present disclosure.

[0012] FIG. 3 is a diagram depicting an example of the permutation testing with the generalized linear model, according to certain aspects of the present disclosure.

[0013] FIG. 4 is a flow chart depicting an example of a process of performing permutation testing with a neural network, according to certain aspects of the present disclosure.

[0014] FIG. 5 is a diagram depicting an example of the permutation testing with the neural network, according to certain aspects of the present disclosure.

[0015] FIG. 6 is a block diagram depicting an example of a computing system suitable for implementing aspects of the permutation testing, according to certain aspects of the present disclosure.Detailed Description

[0016] Certain aspects described herein are provided for target permutation testing to identify and select relevant features for use in models to improve computational efficiency. In some examples, the models may be used in a risk assessment computing system. The risk assessment computing system, in response to receiving a risk assessment query for an entity, can access a model designed to generate a risk indicator for the entity based on input predictor variables associated with the entity. The model may also be used to output other features beyond risk indicators associated with an entity. In an example, the model may use several model features to generate the risk indicator or other output in response to a query. The risk assessment computing system can transmit a response to the riskassessment query for use by a remote computing system in controlling access to one or more interactive computing environments. The response can include the risk indicator and, in some examples, explanatory data associated with the model generating the risk indicator.

[0017] The models described in this disclosure may include global, feature-based explanations for differentiable models, which may be useful for gleaning insights at the population-level rather than the instance-level. In the focus on analyzing the input-output behavior of models in terms of their features, the disclosure relates to a broad category of feature importance methodologies. While the present approach is associated with a category of permutation-based feature importance methodologies, unlike many other techniques in this category, the approach does not fall under a more general umbrella of removal-based explanations, which seek to simulate feature removal to infer feature importance. Rather, the present approach employs permutations to break the relationship between the target and the input covariates in their entirety.

[0018] In short, the present approach works as follows. First, a machine-learning model is fit on a given dataset and either the coefficients (e.g., in the case of a generalized linear model) or the gradients at each data point (e.g., in the case of a neural network model) are extracted. As used herein, the terms generalized linear model and neural network model may generally be described as machine-learning models. In the neural network model case, the absolute values of the gradients are aggregated to calculate a test statistic - essentially a global measure of feature importance. Next, many copies of the dataset are created where the target variable has been randomly permuted and models are fitted to each of the copies. As before, the coefficients / gradients are extracted, and, if applicable, test statistics are calculated. Since any relationships between the predictors and targets in these models are false by design, these results give empirical null distributions by which the probability of observing values of original coefficients / gradients / statistics that are as extreme or greater can be evaluated under the assumption that no relationship exists between the inputs and the target.

[0019] By permuting the target variable, the present approach places no requirements on the correlational structure of the predictors or the targets (in the case of multiple target variables). Additionally, the present approach evaluates the significance of all features jointly, obviating the need to conduct a separate test for each feature of interest (albeit at the expense of being able to isolate a single feature’s conditional contribution). The presentdisclosure explains the feature importance quantification method, its theoretical underpinnings, and its practical application in neural networks, specifically. Accordingly, the method enhances model interpretability and enables model conciseness by identifying the most relevant features in the input space. The disclosure provides a robust technique for feature importance in deep learning.

[0020] Certain aspects described herein can provide a model that can be computationally efficient. For instance, the testing of feature importance may result in a model that includes fewer, but more relevant features. As a result, the model generated or optimized from the target permutation testing method can be operate using fewer computing resources than other models. For example, a reduction in model complexity by removing features from the model may result in faster training times and lower computational demands. Further, models trained with fewer, but more significant features may generalize better to unseen data. This may occur because the model learns to capture the most impactful patterns without or with minimal noise or redundancy that often accompanies large feature sets. Further, reducing the number of features can help in mitigating overfitting, which is a challenge associated with complex models like deep neural networks. By focusing on significant features, models are less likely to learn noise from the training data, which improves model performance on real-world data.

[0021] The permutation test methodology described herein includes a unique approach to evaluating feature significance in neural networks and other predictive models. The described techniques offers several innovative aspects that enhance the value of models in the context of machine learning and deep learning. For example, unlike other permutation tests that shuffle individual features, this method shuffles the target variable. This approach tests the model’s sensitivity to the overall structure and relationships within the data. By disrupting the association between the features and the target while keeping the features themselves intact, it provides a robust measure of how model performance degrades when the underlying data structure is altered. This can be particularly revealing about the model’s dependency on actual versus spurious correlations. With enough parameters, a neural network can represent any complicated and nonlinear associations between its features and targets. Accordingly, this method can potentially be a universal test to identify any types of correlations regardless of their complexity.

[0022] Further, the methodology evaluates the importance of all features simultaneously rather than focusing on one at a time. This global perspective is valuable because it considers the interdependencies among features, which are often overlooked in individual feature assessments. By permuting the target variable, it assesses the collective influence of the features on model predictions, providing a more comprehensive understanding of feature relevance. Additionally, the method only requires running the test once for the entire model, whereas other approaches require a separate test for each variable. This makes the method described herein less computationally intensive than other methods.

[0023] By assessing the significance of features in a holistic manner, this approach can help identify and eliminate unnecessary or noisy features that may lead to overfitting. This is particularly useful in neural networks, where overfitting can significantly degrade model generalizability. Moreover, the ability to quantify and explain the importance of features in deep learning models addresses the “black-box” nature of deep learning methods. Quantifying and explaining the importance of features may enhance the transparency and trustworthiness of the models. Better understanding of what drives model predictions can aid in more informed decision-making and increase the acceptance of these models in critical applications.

[0024] Furthermore, in some statistical models, such as logistic and linear regression, correlated features can disrupt the modeling process, particularly when one of the features has no significant impact on the target variable. This may make it challenging to identify features that do not influence the outcome. But the present method can identify these unimportant features through significance testing. The combination of these factors positions the permutation test method as a significant advancement in the field of machine learning, particularly in enhancing the interpretability and efficiency of deep learning models. This approach not only furthers the understanding of complex models but also aligns with the growing need for transparent, accountable, and efficient Al systems in various sectors.

[0025] These illustrative examples are given to introduce the reader to the general subject matter discussed here and are not intended to limit the scope of the disclosed concepts. The following sections describe various additional features and examples with reference to the drawings in which like numerals indicate like elements, and directionaldescriptions are used to describe the illustrative examples but, like the illustrative examples, should not be used to limit the present disclosure.Operating Environment Example for Model Operations

[0026] Referring now to the drawings, FIG. 1 is a block diagram depicting an example of an operating environment 100 in which a risk assessment computing system 130 can build and train a model that can be utilized to predict risk indicators based on predictor variables. FIG. 1 depicts examples of hardware components of a risk assessment computing system 130, according to some aspects. The risk assessment computing system 130 can be a specialized computing system that may be used for processing large amounts of data using a large number of computer processing cycles. The risk assessment computing system 130 can include a model training server 110 for building and training a model 120. The risk assessment computing system 130 can further include a risk assessment server 118 for performing a risk assessment for given predictor variables 124 using the model 120. While FIG. 1 is described with respect to the risk assessment computing system 130, it may be appreciated that the risk assessment computing system 130 can be used in applications such as transportation, healthcare, finance, education, criminal justice, and any other applications where decision based control is desired.

[0027] The model training server 110 can include one or more processing devices that execute program code, such as a model training application 112. The program code can be stored on a non-transitory computer-readable medium. The model training application 112 can execute one or more processes to train and optimize a model for predicting risk indicators based on predictor variables 124.

[0028] In some aspects, the model training application 112 can build and train a model 120 utilizing an initial training dataset 126. The initial training dataset 126 can include training records with each training record having multiple predictor variables. The model training application 112 can additionally train the model 120 utilizing additional training dataset(s) 128, which may be received after the model 120 is initially trained. The additional training dataset(s) 128 can each include additional predictor variables. The initial training dataset 126 and the additional training dataset(s) 128 can be stored in one or more network- attached storage units on which various repositories, databases, or other structures are stored. Examples of these data structures are the risk data repository 122. In one or more examples, the model training application 112 can further optimize the model120 using the permutation testing described below to remove model features that are determine to not be relevant to an output of the model 120.

[0029] Network- attached storage units may store a variety of different types of data organized in a variety of different ways and from a variety of different sources. For example, the network- attached storage unit may include storage other than primary storage located within the model training server 110 that is directly accessible by processors located therein. In some aspects, the network-attached storage unit may include secondary, tertiary, or auxiliary storage, such as large hard drives, servers, virtual memory, among other types. Storage devices may include portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing and containing data. A machine- readable storage medium or computer-readable storage medium may include a non- transitory medium in which data can be stored and that does not include carrier waves or transitory electronic signals. Examples of a non-transitory medium may include, for example, a magnetic disk or tape, optical storage media such as a compact disk or digital versatile disk, flash memory, memory, or memory devices.

[0030] The risk assessment server 118 can include one or more processing devices that execute program code, such as a risk assessment application 114. The program code can be stored on a non-transitory computer-readable medium. The risk assessment application 114 can execute one or more processes to utilize the model 120 trained by the model training application 112 to predict risk indicators based on input predictor variables 124. In some examples, the model 120 can also be used to generate explanatory data for the predictor variables, which can indicate an effect or an amount of impact that one or more predictor variables have on the risk indicator.

[0031] Further, the risk assessment computing system 130 can communicate with various other computing systems, such as client computing systems 104. For example, client computing systems 104 may send risk assessment queries to the risk assessment server 118 for risk assessment or may send signals to the risk assessment server 118 that control or otherwise influences different aspects of the risk assessment computing system 130. The client computing systems 104 may also interact with user computing systems 106 via one or more public data networks 108 to facilitate interactions between users of the user computing systems 106 and interactive computing environments provided by the client computing systems 104.

[0032] Each client computing system 104 may include one or more third-party devices, such as individual servers or groups of servers operating in a distributed manner. A client computing system 104 can include any computing device or group of computing devices operated by a seller, lender, or other providers of products or services. The client computing system 104 can include one or more server devices. The one or more server devices can include or can otherwise access one or more non-transitory computer-readable media. The client computing system 104 can also execute instructions that provide an interactive computing environment accessible to user computing systems 106. Examples of the interactive computing environment include a mobile application specific to a particular client computing system 104, a web-based application accessible via a mobile device, etc. The executable instructions are stored in one or more non-transitory computer- readable media.

[0033] The client computing system 104 can further include one or more processing devices that are capable of providing the interactive computing environment to perform operations described herein. The interactive computing environment can include executable instructions stored in one or more non-transitory computer-readable media. The instructions providing the interactive computing environment can configure one or more processing devices to perform operations described herein. In some aspects, the executable instructions for the interactive computing environment can include instructions that provide one or more graphical interfaces. The graphical interfaces are used by a user computing system 106 to access various functions of the interactive computing environment. For instance, the interactive computing environment may transmit data to and receive data from a user computing system 106 to shift between different states of the interactive computing environment, where the different states allow one or more electronics transactions between the user computing system 106 and the client computing system 104 to be performed.

[0034] In some examples, a client computing system 104 may have other computing resources associated therewith (not shown in FIG. 1), such as server computers hosting and managing virtual machine instances for providing cloud computing services, server computers hosting and managing online storage resources for users, server computers for providing database services, and others. The interaction between the user computing system 106 and the client computing system 104 may be performed through graphical userinterfaces presented by the client computing system 104 to the user computing system 106, or through an application programming interface (API) calls or web service calls.

[0035] A user computing system 106 can include any computing device or other communication device operated by a user, such as a consumer or a customer. The user computing system 106 can include one or more computing devices, such as laptops, smartphones, and other personal computing devices. A user computing system 106 can include executable instructions stored in one or more non-transitory computer-readable media. The user computing system 106 can also include one or more processing devices that are capable of executing program code to perform operations described herein. In various examples, the user computing system 106 can allow a user to access certain online services from a client computing system 104 or other computing resources, to engage in mobile commerce with a client computing system 104, to obtain controlled access to electronic content hosted by the client computing system 104, etc.

[0036] For instance, the user can use the user computing system 106 to engage in an electronic transaction with a client computing system 104 via an interactive computing environment. An electronic transaction between the user computing system 106 and the client computing system 104 can include, for example, the user computing system 106 being used to request online storage resources managed by the client computing system 104, acquire cloud computing resources (e.g., virtual machine instances), and so on. An electronic transaction between the user computing system 106 and the client computing system 104 can also include, for example, query a set of sensitive or other controlled data, access online financial services provided via the interactive computing environment, submit an online credit card application or other digital application to the client computing system 104 via the interactive computing environment, operating an electronic tool within an interactive computing environment hosted by the client computing system (e.g., a content-modification feature, an application-processing feature, etc.).

[0037] In some aspects, an interactive computing environment implemented through a client computing system 104 can be used to provide access to various online functions. As a simplified example, a website or other interactive computing environment provided by an online resource provider can include electronic functions for requesting computing resources, online storage resources, network resources, database resources, or other types of resources. In another example, a website or other interactive computing environmentprovided by a financial institution can include electronic functions for obtaining one or more financial services, such as loan application and management tools, credit card application and transaction management workflows, electronic fund transfers, etc. A user computing system 106 can be used to request access to the interactive computing environment provided by the client computing system 104, which can selectively grant or deny access to various electronic functions. Based on the request, the client computing system 104 can collect data associated with the user and communicate with the risk assessment server 118 for risk assessment. Based on the risk indicator predicted by the risk assessment server 118, the client computing system 104 can determine whether to grant the access request of the user computing system 106 to certain features of the interactive computing environment.

[0038] In a simplified example, the system depicted in FIG. 1 can generate a model to be used for accurately determining risk indicators, such as credit scores. The model may be generated through target permutation testing of features used in the model. The permutation tests rely on comparisons of a metric measured from collected data to that as if the data had randomly occurred. The present approach permutes the target in a model which both relieves the independence requirement and enables evaluation of all features simultaneously. The techniques focus on models of which associations between inputs and outputs can be summarized numerically, for example, generalized linear models or neural networks.

[0039] The predicted risk indicator determined by the model can be utilized by the service provider to determine the risk associated with the entity accessing a service provided by the service provider, thereby granting or denying access by the entity to an interactive computing environment implementing the service. For example, if the service provider determines that the predicted risk indicator is lower than a threshold risk indicator value, then the client computing system 104 associated with the service provider can generate or otherwise provide access permission to the user computing system 106 that requested the access. The access permission can include, for example, cryptographic keys used to generate valid access credentials or decryption keys used to decrypt access credentials. The client computing system 104 associated with the service provider can also allocate resources to the user and provide a dedicated web address for the allocated resources to the user computing system 106, for example, by adding it in the accesspermission. With the obtained access credentials and / or the dedicated web address, the user computing system 106 can establish a secure network connection to the computing environment hosted by the client computing system 104 and access the resources via invoking API calls, web service calls, HTTP requests, or other proper mechanisms.

[0040] Each communication within the operating environment 100 may occur over one or more data networks, such as a public data network 108, a network 116 such as a private data network, or some combination thereof. A data network may include one or more of a variety of different types of networks, including a wireless network, a wired network, or a combination of a wired and wireless network. Examples of suitable networks include the Internet, a personal area network, a local area network ("LAN"), a wide area network ("WAN"), or a wireless local area network ("WLAN"). A wireless network may include a wireless interface or a combination of wireless interfaces. A wired network may include a wired interface. The wired or wireless networks may be implemented using routers, access points, bridges, gateways, or the like, to connect devices in the data network.

[0041] The number of devices depicted in FIG. 1 is provided for illustrative purposes. Different numbers of devices may be used. For example, while certain devices or systems are shown as single devices in FIG. 1, multiple devices may instead be used to implement these devices or systems. Similarly, devices or systems that are shown as separate, such as the model training server 110 and the risk assessment server 118, may be instead implemented in a signal device or system.Examples of Operations Improving Generalized Linear Modelins

[0042] To assess the significance of features in generalized linear models, henceforth dubbed linear models, an empirical permutation test methodology is adopted. To set up the problem, let X, Y) be the original dataset with A being the input matrix of k feature vectors {Xi, Xz , Xk}, and Y , the target. Without loss of generality, Y may be assumed to be a vector since, with multi-output models, the test can be performed for each individual output and coefficient set. The goal then is to determine how significantly each feature Xj influences the prediction Y of a model by computing an empirical p-value.

[0043] In short, the test may permute Y and compare the model coefficients between the original data and permuted data. The empirical p-value for each feature may then measure the probability that the absolute value of the original coefficient is less extreme than the absolute value of the coefficients on the permuted data.

[0044] FIG. 2 is a flow chart depicting an example of a process 200 for performing permutation testing with a generalized linear model, according to certain aspects of the present disclosure. One or more computing devices (e.g., the model training server 110) implement operations depicted in FIG. 2 by executing suitable program code (e.g., the model training application 112). For illustrative purposes, the process 200 is described with reference to certain examples depicted in the figures. Other implementations, however, are possible. Further, FIG. 3 is a diagram 300 depicting an example of the permutation testing with the generalized linear model of the process 200, according to certain aspects of the present disclosure. The process 200 is described below with respect to features depicted in the diagram 300.

[0045] At block 202, the process 200 involves training a generalized linear model. The generalized linear model Y = <J (W • X + b~) is fit on the original data (X, T), where IV = w1, w2, ... , wk) is a coefficient vector, b is a bias scalar, and cr( ) is a link function for the model. The linear model is fitted on (X, Y ) only in this step. The training of block 202 is depicted at SI in the diagram 300 of FIG. 3, where original data 302 is used as the training data to train a generalized linear model 304. The result of the training is represented by S2 in the diagram 300, which provides the trained linear model 306.

[0046] At block 204, the process 200 involves extracting coefficients of the trained linear model 306. The extracted coefficients represent the importance and influence of each feature Xton the predicted target Y. In the trained linear model 306, the extracted coefficients are included in the coefficient vector W.

[0047] At block 206, the process 200 involves permuting a target variable Y. Permuting the target variable Y may be performed by randomly shuffling the variable N times to obtain N permutations (X, Y^^),j G {1 ... n . In all permutations, predictor variables X remain constant. This alteration reduces the correlations between all predictor variables and the target variable to random. Permuting the target variable is represented as S3 in the diagram 300 of FIG. 3, and permuted data 308 is the result of permuting the target variable.

[0048] At block 208, the process 200 involves training a generalized linear model on the permuted data 308. For each permutation (X, K^), a model Y^ = cr(lV^7-)• X + 7-*) is trained. All models trained on the permuted data use the same training settingsand hyperparameters as the original model trained at block 202. For simplicity, models trained on permuted data may be referred to as permuted models. The training of block 208 is depicted at S4 in the diagram 300 of FIG. 3, where the permuted data 308 is used as the training data to train a generalized linear model 310. The result of the training is represented by S5 in the diagram 300, which provides permuted models 312 after the training process is repeated N times for each permutation of the permuted data 308.

[0049] At block 210, the process 200 involves extracting coefficients from each of the permuted models. In other words, for each permuted model 312, the coefficient vector represents the correlation of X and arandomly shuffled Y.

[0050] At block 212, the process 200 involves comparing coefficients from the trained linear model 306 and the permuted models 312. For example, each coefficientof the permuted models 312 to the coefficient wtof the trained linear model 306. A comparison indicating that| > | wt| means that the correlation between Xi and a randomly created Y is more extreme than that in the collected data. The empirical p-value for feature X is then computed as:with n( ) being a count function. In the diagram 300, the comparison between the coefficients is represented by S6 to produce the p-value 314 for feature Xi.

[0051] At block 214, the process 200 involves adjusting the trained model. In an example, the trained model 306 may be adjusted to include only the features of the set of features determined to be relevant to the outcome of the trained linear model based on the empirical p-value.

[0052] The metric pt calculated using Equation 1 indicates the proportion of permutations where the permuted coefficients are more extreme than the original coefficients. In other words, the metric measures how likely a randomly generated dataset yields a more significant association between X and Y than in the collected data. Lower empirical p-values mean the X to Y association is more significant as it is less likely to occur randomly. In some examples, a threshold of 0.05 may be used for the p-value toscreen out less significant features, though other thresholds may also be used. The resulting model with only features determined to be relevant (i.e., those with a p-value below at or below a p-value threshold) may be implemented with the model 120 in the risk assessment server 118 to perform risk assessment operations.Examples of Operations Improving Neural Network Modeling

[0053] With a neural network, the modeled correlation between an input feature Xtand the prediction Y becomes much more complicated than that of a linear model due to the large number of coefficients stacked in nonlinear layers. Therefore, instead of examining the extremeness of any coefficient in W, that metric is evaluated for the partial gradient of Y by X (i.e., dY / dXi). Furthermore, because a neural network generates a separate value for each instance in data, the final test statistic is aggregated by using the mean absolute function:(Equation 2)where s is the data size of X, T). Using a neural network to approximate the function between inputs and targets, this method can also act as a universal test to reveal much more complicated associations compared to linear models.

[0054] Algorithmically, the test for neural networks follows a similar process to that of generalized linear models. More specifically, one neural network is trained on the original data, and one for each of the N permuted data. Then, the significance of a feature is measured by the rate that the test statistic of a feature is less extreme than that from random. Because different network architectures have different learning capacity, models in the original data and all permuted ones have the same hyperparameters. Below, the permutation test for a neural network is described.

[0055] FIG. 4 is a flow chart depicting an example of a process 400 for performing permutation testing with a neural network model, according to certain aspects of the present disclosure. One or more computing devices (e.g., the model training server 110) implement operations depicted in FIG. 4 by executing suitable program code (e.g., the model training application 112). For illustrative purposes, the process 400 is described with reference to certain examples depicted in the figures. Other implementations, however, are possible. Further, FIG. 5 is a diagram 500 depicting an example of the permutation testing with the neural network model of the process 400, according to certain aspects of the presentdisclosure. The process 400 is described below with respect to features depicted in the diagram 500.

[0056] At block 402, the process 400 involves training a neural network model. Given a dataset (X, Y) with X being the features, and Y the target, the neural network can be trained to estimate Y = 0(A), where 0( ) is the function that is parameterized by the model. The neural network may be finetuned prior to a significance test. Training techniques such as early-stopping and drop-out can also be utilized. The training of block 402 is depicted at SI in the diagram 500 of FIG. 5, where original data 502 is used as the training data to train a neural network model 504. The result of the training is represented by S2 in the diagram 500, which provides the trained neural network model 506.

[0057] At block 404, the process 400 involves determining gradients of the trained neural network model 506. The gradient of the model output Y with respect to each input feature A, dY / dX, is computed using back-propagation. These gradients provide a summarized quantitative measure of how much changes in each input feature X impact the output. Then, the test statistics TZfor X can be calculated using Equation 2.

[0058] At block 406, the process 400 involves permuting a target variable Y. Permuting the target variable Y may be performed by randomly shuffling the variable N times to obtain N permutations (X, Y^^J E {1 ... N}. In all permutations, predictor variables A remain constant. This alteration reduces the correlations between all predictor variables and the target variable to random. Permuting the target variable is represented as S3 in the diagram 500 of FIG. 5, and permuted data 508 is the result of permuting the target variable.

[0059] At block 408, the process 400 involves training a neural network model on the permuted data 508. For each permutation (X, K(-7')), a model 0( / )() is trained. All models trained on the permuted data use the same architecture, training settings (e.g., learning rate, momentum, early stopping criteria, dropout rates, etc.) as the model 0( ) trained at block 402. For simplicity, models trained on permuted data may be referred to as permuted neural network models. The training of block 408 is depicted at S4 in the diagram 500 of FIG. 5, where the permuted data 508 is used as the training data to train a neural network model 510. The result of the training is represented by S5 in thediagram 500, which provides permuted models 512 after the training process is repeated N times for each permutation of the permuted data 508.

[0060] At block 410, the process 400 involves calculating a partial gradient from each of the permuted models 512. In other words, for each neural network 0w, the partial gradient of the permuted target Y is calculated with respect to feature X, (i.e., dY^ / dX ). Then, the test statistics T '1for Xtcan be obtained using Equation 2.

[0061] At block 412, the process 400 involves comparing gradients from the trained neural network model 506 and the permuted gradients from the permuted models 512. For example, the empirical p-value for each feature with respect to the current neural network architecture is computed as:In other words, the significance measurement ofX, is evaluated by the empirical probability that its contribution to the prediction is less extreme than had the model been trained from randomly occurring data. In the diagram 500, the comparison between the coefficients is represented by S6 to produce the p-value 514 for feature .

[0062] At block 414, the process 400 involves adjusting the trained neural network model. In an example, the trained neural network model 506 may be adjusted to include only the features of the set of features determined to be relevant to the outcome of the trained neural network based on the empirical p-value.

[0063] As in generalized linear models discussed above with respect to FIGS. 2 and 3, a feature may be more significant to a neural network if the features impact on the prediction is less likely to happen by randomness. This significance measurement could be different across neural network architectures, even when they are trained from the same data. As divergences in model hyperparameters yield variable learning capacities, a more complicated neural network may find a feature more important whereas a shallower one does not. Therefore, this test may be performed after hyperparameter tuning is completed. Further, due to the increase in complexity in transformer-based models, opting for the use of a shallower neural network during the permutation test may reduce the demanded time and resources. The permutation test still provides useful insights and helpsreduce the models’ input space significantly while maintaining or even improving performances.

[0064] Unlike traditional statistical and tree-based models, neural networks often show improved performance when pruned of features deemed statistically insignificant. The presence of superfluous or noisy features in neural networks can lead to overfitting and reduce the generalizability of the model. The presently described techniques demonstrate that generalized linear models and neural network models can operate more efficiently with fewer, but more relevant, features. This has significant implications for model training and optimization: By employing fewer features, models can reduce complexity, which often results in faster training times and lower computational demands. This is particularly advantageous in environments where resources are constrained or when deploying models to edge devices. Further, models trained with fewer, significant features may generalize better to unseen data. This is because the model may learn to capture the most impactful patterns without or with minimal noise or redundancy, which often accompanies large feature sets. Additionally, reducing the number of features can help in mitigating overfitting, which is a challenge associated with complex models like deep neural networks. By focusing on significant features, models are less likely to learn noise from the training data, improving their performance on test or real-world data.

[0065] Example of Computing System for Machine-Learning Operations

[0066] Any suitable computing system or group of computing systems can be used to perform the operations for the machine-learning operations described herein. For example, FIG. 6 is a block diagram depicting an example of a computing device 600, which can be used to implement the risk assessment server 118 or the model training server 110. The computing device 600 can include various devices for communicating with other devices in the operating environment 100, as described with respect to FIG. 1. The computing device 600 can include various devices for performing one or more operations described above with respect to FIGS. 1-5.

[0067] The computing device 600 can include a processor 602 that is communicatively coupled to a memory 604. The processor 602 executes computer-executable program code stored in the memory 604, accesses information stored in the memory 604, or both. Program code may include machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, asoftware package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, among others.

[0068] Examples of a processor 602 include a microprocessor, an application-specific integrated circuit, a field-programmable gate array, or any other suitable processing device. The processor 602 can include any number of processing devices, including one. The processor 602 can include or communicate with a memory 604. The memory 604 stores program code that, when executed by the processor 602, causes the processor to perform the operations described in this disclosure.

[0069] The memory 604 can include any suitable non-transitory computer-readable medium. The computer-readable medium can include any electronic, optical, magnetic, or other storage device capable of providing a processor with computer-readable program code or other program code. Non-limiting examples of a computer-readable medium include a magnetic disk, memory chip, optical storage, flash memory, storage class memory, ROM, RAM, an ASIC, magnetic storage, or any other medium from which a computer processor can read and execute program code. The program code may include processor-specific program code generated by a compiler or an interpreter from code written in any suitable computer-programming language. Examples of suitable programming language include Hadoop, C, C++, C#, Visual Basic, Java, Python, Perl, JavaScript, ActionScript, etc.

[0070] The computing device 600 may also include a number of external or internal devices such as input or output devices. For example, the computing device 600 is shown with an input / output interface 608 that can receive input from input devices or provide output to output devices. A bus 606 can also be included in the computing device 600. The bus 606 can communicatively couple one or more components of the computing device 600.

[0071] The computing device 600 can execute program code 614 that includes the risk assessment application 114 and / or the model training application 112. The program code 614 for the risk assessment application 114 and / or the model training application 112 maybe resident in any suitable computer-readable medium and may be executed on any suitable processing device. For example, as depicted in FIG. 6, the program code 614 for the risk assessment application 114 and / or the model training application 112 can reside in the memory 604 at the computing device 600 along with the program data 616 associated with the program code 614, such as the predictor variables 124 and / or the initial training dataset 126. Executing the risk assessment application 114 or the model training application 112 can configure the processor 602 to perform the operations described herein.

[0072] In some aspects, the computing device 600 can include one or more output devices. One example of an output device is the network interface device 610 depicted in FIG. 6. A network interface device 610 can include any device or group of devices suitable for establishing a wired or wireless data connection to one or more data networks described herein. Non-limiting examples of the network interface device 610 include an Ethernet network adapter, a modem, etc.

[0073] Another example of an output device is the presentation device 612 depicted in FIG. 6. A presentation device 612 can include any device or group of devices suitable for providing visual, auditory, or other suitable sensory output. Non-limiting examples of the presentation device 612 include a touchscreen, a monitor, a speaker, a separate mobile computing device, etc. In some aspects, the presentation device 612 can include a remote client-computing device that communicates with the computing device 600 using one or more data networks described herein. In other aspects, the presentation device 612 can be omitted.

[0074] The foregoing description of some examples has been presented only for the purpose of illustration and description and is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Numerous modifications and adaptations thereof will be apparent to those skilled in the art without departing from the spirit and scope of the disclosure.

Claims

ClaimsWhat is claimed is:

1. A computer-implemented method comprising: training, by a processor, a generalized linear model on a set of original data; extracting, by the processor, a first set of coefficients from the trained generalized linear model; permuting, by the processor, a target variable of the set of original data by randomly shuffling the target variable to obtain a set of permutations of the set of original data; training, by the processor and using a set of training settings and a set of hyperparameters of the trained generalized linear model, a permuted model for each permutation of the set of permutations to generate a set of permuted models; extracting, by the processor, a second set of coefficients for each permuted model of the set of permuted models; comparing, by the processor, the second set of coefficients for each permuted model and the first set of coefficients of the trained generalized linear model to determine relevance of a set of features associated with the first set of coefficients to an output of the trained generalized linear model; and adjusting, by the processor, the trained generalized linear model to include only a subset of features of the set of features that are determined to be relevant to the output of the trained generalized linear model.

2. The method of claim 1, further comprising: determining a risk indicator of a target based on the adjusted, trained generalize linear model; and transmitting a risk indicator to a third-party, the risk indicator used to control access to a secure computing environment.

3. The method of claim 1, wherein comparing the second set of coefficients for each permuted model and the first set of coefficients from the trained generalized linear model comprises calculating an empirical p-value for each feature of the set of features to determine how likely each permuted model of the set of permuted models is to yield a moresignificant association between each feature of the set of features and the target variable than the trained generalized linear model.

4. The method of claim 3, wherein each feature of the subset of features comprises an empirical p-value below a predetermined threshold p-value.

5. The method of claim 1, wherein the adjusted, trained generalized linear model uses fewer computational resources than the trained generalized linear model to generate an output.

6. The method of claim 1, wherein the relevance of each feature of the set of features is determined simultaneously with the remaining features of the set of features.

7. The method of claim 1, wherein the target variable comprises a multi-output vector.

8. A system comprising: a processing device; and a memory device in which instructions executable by the processing device are stored for causing the processing device to: train a neural network model on a set of original data; determine a first set of gradients of the trained neural network model; permute a target variable of the set of original data by randomly shuffling the target variable to obtain a set of permutations of the set of original data; train, using a set of training settings of the trained neural network model, a permuted model for each permutation of the set of permutations to generate a set of permuted models; determine a second set of gradients for each permuted model of the set of permuted models; compare the second set of gradients or second test statistics computed from the second set of gradients for each permuted model and the first set of gradients or first test statistics computed from the first set of gradients of the trained neuralnetwork model to determine relevance of a set of features associated with the set of gradients to an output of the trained neural network model; and adjust the trained neural network model to include only a subset of features of the set of features that are determined to be relevant to the output of the trained neural network model.

9. The system of claim 8, wherein the instructions executable by the processing device are stored for further causing the processing device to: determine a risk indicator of a target based on the adjusted, trained neural network model; and transmit a risk indicator to a third-party, the risk indicator used to control access to a secure computing environment.

10. The system of claim 8, wherein comparing the second set of gradients or the second test statistics for the set of permuted models and the first set of gradients or the first test statistics from the trained neural network model comprises calculating an empirical p-value for each feature of the set of features to determine how likely each permuted model of the set of permuted models is to yield a more significant association between each feature of the set of features and the target variable than the trained neural network model.

11. The system of claim 10, wherein each feature of the subset of features comprises an empirical p-value below a predetermined threshold p-value.

12. The system of claim 10, wherein the empirical p-value of a feature of the set of features is calculated using a test statistic that represents a mean absolute function of the gradient of the feature.

13. The system of claim 8, wherein the adjusted, trained neural network model uses fewer computational resources than the trained neural network model to generate an output.

14. The system of claim 8, wherein the relevance of each feature of the set of features is determined simultaneously with the remaining features of the set of features.

15. A non-transitory computer-readable storage medium having program code that is executable by a processor device to cause the processor device to: train a machine-learning model on a set of original data; determine a first set of characteristics of the trained machine-learning model; permute a target variable of the set of original data by randomly shuffling the target variable to obtain a set of permutations of the set of original data; train, using a set of training settings of the machine-learning model, a permuted model for each permutation of the set of permutations to generate a set of permuted models; determine a second set of characteristics for each permuted model of the set of permuted models; compare the second set of characteristics for each permuted model and the first set of characteristics of the trained machine-learning model to determine relevance of a set of features associated with the first set of characteristics to an output of the trained machinelearning model; and adjust the trained machine-learning model to include only a subset of features of the set of features that are determined to be relevant to an output of the trained machine-learning model.

16. The non-transitory computer-readable storage medium of claim 15, wherein comparing the second set of characteristics for each permuted model and the first set of characteristics of the trained machine-learning model comprises calculating an empirical p- value for each feature of the set of features to determine how likely each permuted model of the set of permuted models is to yield a more significant association between each feature of the set of features and the target variable than the trained machine-learning model.

17. The non-transitory computer-readable storage medium of claim 16, wherein each feature of the subset of features comprises an empirical p-value below a predetermined threshold p-value.

18. The non-transitory computer-readable storage medium of claim 15, wherein the machine-learning model comprises a generalized linear model or a neural network model.

19. The non-transitory computer-readable storage medium of claim 15, wherein the program code is further executable by the processor device to cause the processor device to: determine a risk indicator of a target based on the adjusted, trained machine-learning model; and transmit a risk indicator to a third-party, the risk indicator used to control access by an entity to a secure computing environment.

20. The non-transitory computer-readable storage medium of claim 15, wherein the first set of characteristics and the second set of characteristics comprise model coefficients or model gradients.

Citation Information

Patent Citations

  • Analysis and display of cybersecurity risks for enterprise data

    US20160012235A1

  • AU2021313070A1