Systems and methods for automated transfer learning using domain disentanglement

Automatically searching for hyperparameters of nuisance censorship mode and method through Autotransfer Framework solves the problem that deep neural networks are difficult to achieve robust cross-domain learning when dealing with nuisance factors, and realizes seamless decoupling of nuisance factors and automatic optimization of DNN architecture.

JP7672588B2Active Publication Date: 2025-05-07MITSUBISHI ELECTRIC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024555572
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-02-01
Filing Date
2022-09-30
Publication Date
2025-05-07
Estimated Expiration
2042-09-30

AI Technical Summary

Technical Problem

The prior art requires manual determination of the connectivity and architecture of functional blocks when designing deep neural networks (DNNs), resulting in a large number of trials and errors required for architecture optimization, and it is difficult to achieve robust cross-domain learning when dealing with nuisance factors such as noise, interference, deviations and domain shifts.

Method used

An automated machine learning method, called Autotransfer Framework, is proposed to achieve seamless decoupling of nuisance factors by searching for hyperparameters of different nuisance censorship patterns and methods. The framework uses classification and continuous collaborative search spaces to adjust the degree of domain release, and finds the best balance between task distinction and nuisance invariant features through multiple censoring methods and hyperparameter adjustments.

Benefits of technology

Through the Autotransfer Framework, it is possible to efficiently search for DNN model architectures suitable for nuisance, reducing the dependence of manual design, improving the robustness and efficiency of cross-domain learning, and reducing the sensitivity to nuisance factors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672588000006
    Figure 0007672588000006
  • Figure 0007672588000007
    Figure 0007672588000007
  • Figure 0007672588000008
    Figure 0007672588000008
Patent Text Reader

Abstract

A system and method for automated construction of artificial neural network architectures are provided. The system includes a set of interfaces and data links configured to receive and transmit signals. The signals include a dataset of training data, validation data, and test data. The signals include a set of random factors in a multidimensional signal. Some of the random factors are associated with task labels for identifying and nuisance variations. The system further includes a set of reconfigurable deep neural network (DNN) blocks, a set of memory banks for storing hyperparameters, trainable variables, intermediate neuron signals, and provisional calculated values ​​including forward pass signals and backward pass gradients. The system further includes at least one processor connected to the interface and the memory banks, the at least one processor configured to present the signals and datasets to the reconfigurable DNN blocks. The at least one processor is configured to explore hyperparameters of regularization modules, pre-processing methods, and post-processing methods to achieve Bayesian inference robust to nuisances such that the reconfigurable DNN blocks are transferable to new datasets with domain shifts.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an automated training system for artificial neural networks, and more particularly to an automated transfer learning and domain adaptation system for artificial neural networks using nuisance factor disentanglement. [Background technology]

[0002] Great progress in deep learning methods based on deep neural networks (DNNs) has solved various problems in data processing, including media signal processing for video, audio, and images, physical data processing for radio waves, electric pulses, and light beams, and physiological data processing for heart rate, temperature, and blood pressure. For example, DNNs have enabled more practical design of human-machine interfaces (HMIs) through the analysis of users' biosignals such as electroencephalograms (EEG) and electromyograms (EMG). However, such biosignals are highly variable depending on the biological state of each subject as well as the imperfections of the measurement sensors and the inconsistencies of the experimental setup. Thus, frequent calibration is often required in typical HMI systems. In addition to HMI systems, data analysis often encounters many nuisance factors such as noise, interference, bias, domain shift, etc. Therefore, deep learning that is robust to those nuisance factors across different dataset domains is required. Summary of the Invention [Problem to be solved by the invention]

[0003] Towards the solution of this problem, nuisance-invariant methods that employ adversarial training, such as Adversarial Conditional Variational AutoEncoder (A-CVAE), have emerged to reduce domain calibration to realize cross-domain generalization deep learning, such as subject-invariant HMI systems. Compared to standard DNN classifiers / regressors, integrating additional functional blocks, such as an encoder, a nuisance-conditional decoder, and an adversarial network, provides superior nuisance-invariant performance, since domain generalization can be obtained without new domain data. The DNN structure can potentially be extended with more functional blocks and more latent layers. However, most works rely on human design to determine the block connectivity and architecture of DNNs. Specifically, DNN methods are often handcrafted by experts who use human insight to design data models. Methods to optimize the architecture of DNNs require a trial-and-error approach. A new framework of automated machine learning (AutoML) has been proposed to automatically explore different DNN architectures. Automating hyperparameter and architecture search in the context of AutoML can facilitate DNN design suitable for processing nuisance-invariant data. In addition to DNN architectures, there are many approaches to stabilize DNN training behavior by regularizing trainable parameters, such as adversarial disentanglement and L2 / L1 norm regularization.

[0004] Learning data representations that capture task-relevant features but are invariant to nuisance variations remains a key challenge in machine learning. VAE introduced a variational Bayesian inference method incorporating an auto-coupled architecture where generative and inferential models can be jointly trained. This method was extended with CVAE, which introduces conditioning variables that can be used to represent nuisance variations, and regularized VAE, which considers disentangling nuisance variables from the latent representation. The concept of adversarial learning was considered in Generative Adversarial Networks (GANs) and has been adopted in a myriad of applications. The contemporaneous discovery of Adversarially Learned Inference (ALI) and Bidirectional GANs (BiGANs) proposed an adversarial approach towards training autoencoders. Adversarial training has also been combined with VAE to regularize and disentangle the latent representation such that robust learning against nuisances is achieved. Searching for DNN models using hyperparameter optimization has been thoroughly studied in a related framework called AutoML. These automated methods include architecture search, learning rule design, and augmented search. Most work uses evolutionary optimization or reinforcement learning frameworks to tune hyperparameters or build network architectures from preselected building blocks. Recent work on AutoML-Zero considers augmentation to eliminate human knowledge and insight for a fully automated design from scratch.

[0005] However, AutoML requires a lot of search time to find the best hyperparameters due to the explosion of the search space. In addition, without a good justification, most of the search space of link connections will be meaningless. In order to develop a system for the automated construction of neural networks with justification, a method called AutoBayes was proposed. The AutoBayes method explores different Bayesian graphs to represent the inherent graphical relationships between variable data for the generative model, and then builds the most plausible inference graph to connect the encoder, decoder, classifier, regressor, adversary, and domain estimator. Using the so-called Bayes-Ball algorithm, the most compact inference graph for a given Bayesian graph can be automatically constructed, and some factors are identified as independent variables from the domain factors that should be truncated by the adversarial block. Adversarial truncation to disentangle nuisance factors from the feature space is verified to be effective for domain generalization in pre-shot transfer learning and domain adaptation in post-shot transfer learning.

[0006] However, adversarial training requires careful selection of hyperparameters because too strong truncation impairs the primary task performance due to insufficient weighting of the primary objective function. Also, adversarial truncation is not the only regularization approach to promote independence from nuisance variables in the feature space. For example, minimizing the mutual information between nuisances and features can be achieved by a mutual information gradient estimator (MIGE). Similarly, there are different such truncation approaches and scoring methods to consider. Due to the so-called no-free-lunch theorem, there is no single method that can universally achieve the best performance across different problems and datasets. Exploring domain disentanglement approaches requires time / resource-intensive trial and error to find the best solution. Therefore, there is a need to efficiently identify the best truncation approach that depends on the specific problem for transfer learning robust to nuisances. [Means for solving the problem]

[0007] The present invention provides a way to design machine learning models such that nuisance factors are seamlessly disentangled by exploring various hyperparameters of truncation modes and methods for transfer learning that is robust to domain shifts across pre-shot and post-shot phases. The present invention enables AutoML to efficiently search potential transfer learning modules, and hence we call it the auto-transfer framework. One embodiment uses categorical and continuous joint search spaces between different truncation modes and methods with truncation hyperparameters to adjust the level of domain disentanglement. The truncation modes include, but are not limited to, marginal distributions, conditional distributions, and complementary distributions to control the mode of disentanglement. The truncation methods encourage the features in the machine learning models to be independent of the nuisance parameters, so that feature extraction that is robust to nuisances is achieved. However, because too strong truncation generally degrades downstream task performance, auto-transfer adjusts the hyperparameters to search for the best tradeoff between task-discriminative features and nuisance-invariant features. Truncation methods include, but are not limited to, adversarial networks, mutual information gradient estimation (MIGE), pairwise discrepancy, and Wasserstein distance.

[0008] The invention provides ways to tune their hyperparameters under AutoML frameworks such as Bayesian optimization, reinforcement learning, and heuristic optimization. Yet another embodiment explores different preprocessing mechanisms including domain-robust data augmentation, filter banks, and wavelet kernels to improve nuisance-robust inference across many different data formats such as time series signals, spectrograms, cepstrum, and other tensors. Another embodiment uses variational sampling for semi-supervised settings where nuisance factors are not fully available for training. Another embodiment provides ways to convert one data structure to another with mismatched dimensionality by using tensor projection with optimal transport methods and independent component mapping with common spatial patterns to enable heterogeneous transfer learning. An embodiment implements an ensemble method to explore stacking protocols through cross-validation to reuse multiple explored models at once. In addition to pre-shot transfer learning when zero data is available in the target domain during the training phase, the present invention also provides post-shot transfer learning when some data is available in the target domain during the training or fine-tuning phase, such as zero-shot learning (all data in the target domain are unlabeled), one-shot learning, and few-shot learning. Hypernetwork adaptation provides a way to automatically generate auxiliary models that directly control the parameters of the basic inference model by analyzing the consistent evolution behavior in the hypothesis post-shot learning phase. Post-shot learning includes, but is not limited to, successive unpacking and fine-tuning with perturbation minimization from source domain to target domain with or without pseudo-labeling.

[0009] The present disclosure relates to a system and method for automated construction of artificial neural networks through the exploration of different truncation modules and preprocessing methods. Specifically, the system of the present invention introduces an automated transfer learning framework called auto-transfer that explores different disentanglement approaches for inference models linking classifier, encoder, decoder and estimator blocks to optimize nuisance-invariant machine learning pipelines. In one embodiment, the framework is applied to a set of physiological datasets, where we have access to subject and class labels during training, and provide an analysis of its capabilities for subject transfer learning with / without variational modeling and adversarial training. The framework can be effectively utilized in semi-supervised multi-class classification, multidimensional regression and data reconstruction tasks for various dataset formats such as media and electrical signals as well as biological signals.

[0010] Some embodiments of the present disclosure are based on the recognition that a new concept called AutoBayes explores various different Bayesian graph models to facilitate the search for the best inference strategy suitable for a nuisance-robust HMI system. Using the Bayes-Ball algorithm, our method can automatically build plausible link connections between the classifier, encoder, decoder, nuisance estimator and adversarial DNN blocks. We observed a huge performance gap between the best and worst graph models, implying that using a deterministic model without graph exploration may suffer from poor classification results. In addition, the best model for a physiological dataset does not necessarily work best for different data, which prompts us to use AutoBayes for adaptive model generation given a target dataset. One embodiment extends the macro-level AutoBayes framework to integrate micro-level AutoML to optimize the hyperparameters of each DNN block. The present invention is based on the recognition that some nodes in a Bayesian graph are marginally or conditionally independent from other nodes. Our inventive auto-transfer framework further explores various truncation modes and methods to promote such independence at specific hidden nodes of the DNN model to enhance the AutoBayes framework.

[0011] Our invention allows AutoML to efficiently search for potential architectures that have solid theoretical reasons to consider them. Our method is based on the realization that the dataset is hypothetically modeled using a directed Bayesian graph, and we therefore call it the AutoBayesian method. One embodiment uses a Bayesian graph search with different factorization orders of the joint probability distribution. Our invention also provides a method for creating compact architectures with pruning links based on conditional independence derived from the Bayes-Ball algorithm through the Bayesian graph hypothesis. Yet another method can optimize the inference graph with different factorization orders of the likelihood, which allows for automatically constructing a combined generative graph and an inference graph. It achieves natural architectures based on VAE with / without conditional links. Yet another embodiment uses domain disentanglement with an auxiliary network attached to latent variables that are independent of the nuisance parameters, thereby achieving feature extraction that is robust to nuisances. Yet another case uses intentionally redundant graphs with conditional grafting to facilitate feature extraction that is robust to nuisances. Yet another embodiment uses an ensemble graph that combines estimates and disentanglement methods of multiple different Bayesian graphs to improve performance. For example, Wasserstein distance can be used instead of divergence to measure independence scores. One embodiment uses a dynamic attention network to realize the ensemble method. Also, cycle consistency of VAE and model consistency between different inference graphs are both addressed. Another embodiment uses a graph neural network to exploit the geometric information of the data, and the pruning strategy is assisted by belief propagation between Bayesian graphs to verify associations.

[0012] The system provides a systematic automated framework approach to search for the best inference graph model associated with a Bayesian graph model well suited to reproduce the training dataset. The proposed system automatically formulates various distinct Bayesian graphs by factorizing the joint probability distribution with respect to the data, class labels, subject identification (ID), and unique latent representations. Given a Bayesian graph, several meaningful inference graphs are generated through a Bayes-Ball algorithm to prune redundant links and achieve high accuracy of estimation. To promote robustness against nuisance parameters such as subject ID, the explored Bayesian graphs can provide justification for using domain disentanglement with / without variational modeling. As one of the embodiments, Auto-Transfer with Auto-Bayes can achieve superior performance across various physiological datasets for cross-subject, cross-session, and cross-device transfer learning.

[0013] In the system of the present invention, a variety of different truncation methods for transfer learning are considered, for example, for classification of biosignal data. The system is established to address the difficulty of transfer learning for biosignals, known as the problem of "negative transfer", whereby a naive attempt to combine datasets from multiple subjects or sessions may paradoxically degrade model performance due to domain differences in response statistics. The method of the present invention tackles such subject transfer problems by training the model to be invariant to changes in nuisance variables representing subject identifiers. Specifically, the method automatically explores several established approaches to build a set of good approaches based on mutual information estimation and generative modeling. For example, the method is enabled on real datasets such as various electroencephalography (EEG), electromyography (EMG), and electrocorticography (ECoG) datasets, showing that these methods can improve generalization to unfamiliar subjects. Some embodiments also explore ensembling strategies to combine a set of these good approaches into a single meta-model to gain additional performance. Further exploration of these methods via hyper-parameter tuning can yield additional generalization improvements. For some embodiments, the systems and methods can be combined with existing test-time online adaptation techniques from zero-shot and few-shot learning frameworks to achieve even better subject transfer performance.

[0014] An important approach to the transfer learning problem is to truncate the encoder model such that it learns a representation that is useful for the task while containing minimal information about the changes in nuisance variables that will change as part of our transfer learning setup. Specifically, we consider a dataset composed of high-dimensional data (e.g., raw EEG inputs) with task-relevant labels (e.g., EEG task categories) and nuisance labels (e.g., subject ID or poster ID). Intuitively, we seek to learn a representation that captures only the variation that is relevant to the task. The motivation behind this approach is related to information bottleneck methods, but with an important difference. While information bottleneck methods and their variational variants seek to learn useful, compressed representations from a supervised dataset without additional information about the nuisance variation, we explicitly use additional nuisance labels to draw conclusions about the types of variation in the data that should not affect the output of our model. Many transfer learning settings will have such nuisance labels readily available; and intuitively, the model should benefit from this additional source of supervision. The system can provide unobvious benefits to learning subject-invariant representations by exploring various regularization modules for domain disentanglement.

[0015] Also, according to some embodiments of the present invention, a system for automated construction of artificial neural network architectures is provided. In this case, the system may include a set of interfaces and data links configured to transmit and receive signals. The signals include a data set of training data, validation data, and test data, and the signals include a set of random variable factors in a multidimensional signal X, some of the random variable factors being associated with a task label Y and nuisance variations S for identification. The system further includes a set of memory banks for storing a set of reconfigurable DNN blocks, each of the reconfigurable DNN blocks being configured with a main task pipeline module for identifying the task label Y from the multidimensional signal X, and a set of auxiliary regularization modules for adjusting the disentanglement between a plurality of latent variables Z and the nuisance variations S. The memory banks further include hyperparameters, trainable variables, intermediate neuron signals, and preliminary calculated values ​​including forward pass signals and backward pass gradients. The system further includes at least one processor connected to the interfaces and memory banks and configured to present the signals and the data set to the reconfigurable DNN blocks. The at least one processor is configured to perform a search across a set of graphical models, a set of pre-shot regularization methods, a set of pre-processing methods, a set of post-processing methods, and a set of post-shot adaptation methods to reconfigure the reconfigurable DNN block such that the task prediction is insensitive to the nuisance variation S by modifying hyper-parameters in the memory bank.

[0016] Further, some embodiments of the present invention provide a computer-implemented method for automated construction of artificial neural network architectures. The computer-implemented method may include providing a dataset of training data, validation data, and test data. The dataset includes a set of random variable factors in a multidimensional signal X, some of which are associated with a task label Y and nuisance variation S for identification. The computer-implemented method may further include configuring a set of reconfigurable DNN blocks for identifying the task label Y from the multidimensional signal X. The set of reconfigurable DNN blocks includes a set of auxiliary regularization modules for adjusting the disentanglement between a plurality of latent variables Z and the nuisance variation S. The computer-implemented method may further include training the set of reconfigurable DNN blocks via stochastic gradient optimization on the training data such that the task prediction is accurate, and exploring the set of auxiliary regularization modules on the validation data to search for the best hyperparameters such that the task prediction is insensitive to the nuisance variation S.

[0017] The accompanying drawings, which are included to provide a further understanding of the invention, illustrate embodiments of the invention and, together with the description, explain the principles of the invention. [Brief description of the drawings]

[0018] [Figure 1A] FIG. 1 illustrates an inference method for classifying Y given data X under latency Z and semi-labeled nuisances S, according to an embodiment of the present disclosure. [Figure 1B] FIG. 1 illustrates an inference method for classifying Y given data X under latency Z and semi-labeled nuisances S, according to an embodiment of the present disclosure. [Figure 1C] FIG. 1 illustrates an inference method for classifying Y given data X under latency Z and semi-labeled nuisances S, according to an embodiment of the present disclosure. [Figure 2A]1 illustrates an example Bayesian graph model and an interference model for a particular factorization, in accordance with some embodiments of the present disclosure. [Figure 2B] 1 illustrates an example Bayesian graph model and an interference model for a particular factorization, in accordance with some embodiments of the present disclosure. [Figure 2C] 1 illustrates an example Bayesian graph model and an interference model for a particular factorization, in accordance with some embodiments of the present disclosure. [Figure 3A] FIG. 1 illustrates an example Bayesian graph model for a data-generating model under automated exploration, in accordance with some embodiments of the present disclosure. [Figure 3B] FIG. 1 illustrates an example Bayesian graph model for a data-generating model under automated exploration, in accordance with some embodiments of the present disclosure. [Figure 3C] FIG. 1 illustrates an example Bayesian graph model for a data-generating model under automated exploration, in accordance with some embodiments of the present disclosure. [Figure 3D] FIG. 1 illustrates an example Bayesian graph model for a data-generating model under automated exploration, in accordance with some embodiments of the present disclosure. [Figure 3E] FIG. 1 illustrates an example Bayesian graph model for a data-generating model under automated exploration, in accordance with some embodiments of the present disclosure. [Figure 3F] FIG. 1 illustrates an example Bayesian graph model for a data-generating model under automated exploration, in accordance with some embodiments of the present disclosure. [Figure 3G] FIG. 1 illustrates an example Bayesian graph model for a data-generating model under automated exploration, in accordance with some embodiments of the present disclosure. [Figure 3H] FIG. 1 illustrates an example Bayesian graph model for a data-generating model under automated exploration, in accordance with some embodiments of the present disclosure. [Figure 3I] FIG. 1 illustrates an example Bayesian graph model for a data-generating model under automated exploration, in accordance with some embodiments of the present disclosure. [Figure 3J] FIG. 1 illustrates an example Bayesian graph model for a data-generating model under automated exploration, in accordance with some embodiments of the present disclosure. [Figure 3K] FIG. 1 illustrates an example Bayesian graph model for a data-generating model under automated exploration, in accordance with some embodiments of the present disclosure. [Figure 4A] FIG. 2 illustrates an example inference factor graph model associated with a particular generative model, in accordance with some embodiments of the present disclosure. [Figure 4B] FIG. 2 illustrates an example inference factor graph model associated with a particular generative model, in accordance with some embodiments of the present disclosure. [Figure 4C] FIG. 2 illustrates an example inference factor graph model associated with a particular generative model, in accordance with some embodiments of the present disclosure. [Figure 4D] FIG. 2 illustrates an example inference factor graph model associated with a particular generative model, in accordance with some embodiments of the present disclosure. [Figure 4E] FIG. 2 illustrates an example inference factor graph model associated with a particular generative model, in accordance with some embodiments of the present disclosure. [Figure 4F] FIG. 2 illustrates an example inference factor graph model associated with a particular generative model, in accordance with some embodiments of the present disclosure. [Figure 4G] FIG. 2 illustrates an example inference factor graph model associated with a particular generative model, in accordance with some embodiments of the present disclosure. [Figure 4H] FIG. 2 illustrates an example inference factor graph model associated with a particular generative model, in accordance with some embodiments of the present disclosure. [Figure 4I] FIG. 2 illustrates an example inference factor graph model associated with a particular generative model, in accordance with some embodiments of the present disclosure. [Figure 4J] FIG. 2 illustrates an example inference factor graph model associated with a particular generative model, in accordance with some embodiments of the present disclosure. [Figure 4K]FIG. 2 illustrates an example inference factor graph model associated with a particular generative model, in accordance with some embodiments of the present disclosure. [Figure 4L] FIG. 2 illustrates an example inference factor graph model associated with a particular generative model, in accordance with some embodiments of the present disclosure. [Figure 5A] FIG. 1 illustrates one of the ten basic rules of the Bayes-Ball algorithm with shaded conditional nodes as condition factors, according to an embodiment of the present disclosure. [Figure 5B] FIG. 1 illustrates one of the ten basic rules of the Bayes-Ball algorithm with shaded conditional nodes as condition factors, according to an embodiment of the present disclosure. [Figure 5C] FIG. 1 illustrates one of the ten basic rules of the Bayes-Ball algorithm with shaded conditional nodes as condition factors, according to an embodiment of the present disclosure. [Figure 5D] FIG. 1 illustrates one of the ten basic rules of the Bayes-Ball algorithm with shaded conditional nodes as condition factors, according to an embodiment of the present disclosure. [Figure 5E] FIG. 1 illustrates one of the ten basic rules of the Bayes-Ball algorithm with shaded conditional nodes as condition factors, according to an embodiment of the present disclosure. [Figure 5F] FIG. 1 illustrates one of the ten basic rules of the Bayes-Ball algorithm with shaded conditional nodes as condition factors, according to an embodiment of the present disclosure. [Figure 5G] FIG. 1 illustrates one of the ten basic rules of the Bayes-Ball algorithm with shaded conditional nodes as condition factors, according to an embodiment of the present disclosure. [Figure 5H] FIG. 1 illustrates one of the ten basic rules of the Bayes-Ball algorithm with shaded conditional nodes as condition factors, according to an embodiment of the present disclosure. [Figure 5I]FIG. 1 illustrates one of the ten basic rules of the Bayes-Ball algorithm with shaded conditional nodes as condition factors, according to an embodiment of the present disclosure. [Figure 5J] FIG. 1 illustrates one of the ten basic rules of the Bayes-Ball algorithm with shaded conditional nodes as condition factors, according to an embodiment of the present disclosure. [Figure 6] FIG. 1 illustrates an example algorithm illustrating the overall procedure of the AutoBayes algorithm for searching model architectures, according to an embodiment of the present disclosure. [Figure 7A] FIG. 1 illustrates an exemplary algorithm illustrating the overall procedure of subset selection for pairwise score estimation based on Bernoulli and Clique criteria, according to an embodiment of the present disclosure. [Figure 7B] FIG. 1 illustrates an exemplary algorithm illustrating the overall procedure of subset selection for pairwise score estimation based on Bernoulli and Clique criteria, according to an embodiment of the present disclosure. [Figure 8] FIG. 1 illustrates an example model for predicting task label Y from data X in the main pipeline of encoder f and decoder g, where latent factor Z is regularized by a set of truncation methods to disentangle nuisance factor S, according to an embodiment of the present disclosure. [Figure 9A] FIG. 13 illustrates example pseudocode describing the adversarial truncation method in bounded truncation, conditional truncation, and complementary truncation modes in accordance with an embodiment of the present disclosure. [Figure 9B] FIG. 13 illustrates example pseudocode describing the adversarial truncation method in bounded truncation, conditional truncation, and complementary truncation modes in accordance with an embodiment of the present disclosure. [Figure 9C] FIG. 13 illustrates example pseudocode describing the adversarial truncation method in bounded truncation, conditional truncation, and complementary truncation modes in accordance with an embodiment of the present disclosure. [Figure 10A]FIG. 13 illustrates example pseudocode describing a mutual information gradient estimation (MIGE)-based truncation method in marginal, conditional, and complementary truncation modes according to an embodiment of the present disclosure. [Figure 10B] FIG. 13 illustrates example pseudocode describing a mutual information gradient estimation (MIGE)-based truncation method in marginal, conditional, and complementary truncation modes according to an embodiment of the present disclosure. [Figure 10C] FIG. 13 illustrates example pseudocode describing a mutual information gradient estimation (MIGE)-based truncation method in marginal, conditional, and complementary truncation modes according to an embodiment of the present disclosure. [Figure 11A] FIG. 1 illustrates example pseudocode describing maximum mean discrepancy (MMD)-based truncation methods in marginal, conditional, and complementary truncation modes, according to an embodiment of the present disclosure. [Figure 11B] FIG. 13 illustrates example pseudocode describing a maximum mean discrepancy (MMD)-based truncation method in marginal, conditional, and complementary truncation modes, according to an embodiment of the present disclosure. [Figure 11C] FIG. 13 illustrates example pseudocode describing a maximum mean discrepancy (MMD)-based truncation method in marginal, conditional, and complementary truncation modes, according to an embodiment of the present disclosure. [Figure 12A] FIG. 13 illustrates an example pseudocode describing a pairwise MMD-based truncation method in bounded, conditional, and complementary truncation modes according to an embodiment of the present disclosure. [Figure 12B] FIG. 13 illustrates an example pseudocode describing a pairwise MMD-based truncation method in bounded, conditional, and complementary truncation modes according to an embodiment of the present disclosure. [Figure 12C]FIG. 13 illustrates an example pseudocode describing a pairwise MMD-based truncation method in bounded, conditional, and complementary truncation modes according to an embodiment of the present disclosure. [Figure 13A] FIG. 1 illustrates an example pseudocode describing a boundary equilibrium generative adversarial network (BEGAN) discriminator-based truncation method in marginal, conditional, and complementary truncation modes, according to an embodiment of the present disclosure. [Figure 13B] FIG. 1 illustrates an example pseudocode describing a Bounded Balanced Generative Adversarial Network (BEGAN) discriminator-based truncation method in marginal, conditional, and complementary truncation modes, according to an embodiment of the present disclosure. [Figure 13C] FIG. 1 illustrates an example pseudocode describing a Bounded Balanced Generative Adversarial Network (BEGAN) discriminator-based truncation method in marginal, conditional, and complementary truncation modes, according to an embodiment of the present disclosure. [Figure 14A] FIG. 2 illustrates an exemplary set of post-processing modules used in the post-shot adaptation stage, according to an embodiment of the present disclosure. [Figure 14B] FIG. 2 illustrates an exemplary set of post-processing modules used in the post-shot adaptation stage, according to an embodiment of the present disclosure. [Figure 14C] FIG. 2 illustrates an exemplary set of post-processing modules used in the post-shot adaptation stage, according to an embodiment of the present disclosure. [Figure 14D] FIG. 2 illustrates an exemplary set of post-processing modules used in the post-shot adaptation stage, according to an embodiment of the present disclosure. [Figure 15A] FIG. 2 illustrates an exemplary set of pre-processing modules, according to an embodiment of the present disclosure. [Figure 15B] FIG. 2 illustrates an exemplary set of pre-processing modules, according to an embodiment of the present disclosure. [Figure 15C] FIG. 2 illustrates an exemplary set of pre-processing modules, according to an embodiment of the present disclosure. [Figure 15D] FIG. 2 illustrates an exemplary set of pre-processing modules, according to an embodiment of the present disclosure. [Figure 15E] FIG. 2 illustrates an exemplary set of pre-processing modules, according to an embodiment of the present disclosure. [Figure 15F] FIG. 2 illustrates an exemplary set of pre-processing modules, according to an embodiment of the present disclosure. [Figure 16] FIG. 1 is a schematic diagram of a system configured with a processor, a memory, and an interface in accordance with an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0019] Various embodiments of the present invention are described below with reference to the drawings. Note that the drawings are not drawn to scale, and elements of similar structure or function are represented by the same reference numerals throughout the drawings. Furthermore, the drawings are merely intended to facilitate the description of specific embodiments of the present invention. The drawings are not intended as an exhaustive description of the present invention or as limitations on the scope of the present invention. In addition, aspects described in connection with a particular embodiment of the present invention are not necessarily limited to that embodiment, and may be practiced in any other embodiment of the present invention.

[0020] FIG. 1A shows an example schematic diagram of an artificial intelligence (AI) model that provides inference to identify a task label Y from observed data X. The task label is a categorical identification number or a non-categorical continuous value. For categorical inference, the AI ​​model performs a classification task, while for non-categorical inference, the model performs a regression task. The task label is a scalar value or a vector of multiple values. The observed data is Media data such as images, photographs, movies, text, characters, voice, music, audio and voice; Physical data such as radio waves, optical signals, electrical pulses, temperature, pressure, acceleration, velocity, vibration, mass, moisture, and force; Physiological data such as heart rate, blood pressure, electroencephalogram, electromyogram, electrocardiogram, mechanomyogram, electrooculogram, galvanic skin response, magnetoencephalogram, and electrocorticography;

[0036] It is a tensor format having at least one axis for representing many signal and sensor data including, but not limited to:

[0021] For example, the AI ​​model predicts emotions from a user's electroencephalogram measurements, where the data is a three-axis tensor representing a spatiotemporal spectrogram from a multi-channel sensor over the measurement time. All available data signals with X and Y pairs are bundled together as a whole batch of data sets for training the AI ​​model, which are called training data or training data sets for supervised learning. For some embodiments, as a semi-supervised setting, there is no task label Y for part of the training data set.

[0022] The AI ​​model may be realized by a reconfigurable deep neural network (DNN) model whose architecture is specified by a set of hyperparameters. The set of hyperparameters includes, but is not limited to, the number of hidden nodes, the number of hidden layers, the type of activation function, graph edge connectivity, and cell combinations. Reconfigurable DNN architectures are typically based on multi-layer perceptrons that use cell combinations such as fully connected, convolutional, recurrent, pooling, and normalization layers with many trainable parameters such as affine transformation weights and biases. The types of activation functions include, but are not limited to, sigmoid, hard sigmoid, log sigmoid, tanh, hard tanh, softmax, soft shrink, hard shrink, tanh shrink, rectified linear unit, soft sign, exponential linear unit, sigmoid linear unit, mish, hard swish, and soft plus. The graph edge connectivity includes, but is not limited to, skip addition, skip concatenation, skip product, branching, and looping. For example, residual networks use skip connections from one hidden layer to another, which allows for stable learning of deeper layers.

[0023] The DNN model is trained on a training dataset to minimize or maximize an objective function by gradient methods such as stochastic gradient descent, adaptive momentum gradient, root mean square propagation, adaptive gradient, adaptive delta, adaptive max, elastic backpropagation, and weighted adaptive momentum. For some embodiments, the training dataset is split into multiple sub-batches for local gradient updates. For some embodiments, a portion of the training dataset is set aside for a validation dataset to evaluate the performance of the trained DNN model. In some embodiments, the validation dataset from the training dataset is rotated for cross-validation. The manners for splitting the training data into sub-batches for cross-validation include, but are not limited to, random sampling, weighted random sampling, one session set aside, one subject set aside, and one region set aside. Typically, the data distribution for each sub-batch is not the same due to domain shift.

[0024] Gradient-based optimization algorithms have several hyperparameters, such as learning rate and weight decay. The learning rate is an important parameter to choose and can be automatically adjusted by several scheduling methods, such as step function, exponential function, trigonometric function, and flat adaptive decay. Non-gradient optimization such as evolutionary strategies, genetic algorithms, differential evolution, and Nelder-Mead can also be used. Objective functions include, but are not limited to, L1 loss, mean squared error loss, cross entropy loss, connectionist time classification loss, negative log-likelihood loss, Kullback-Leibler divergence loss, margin ranking loss, hinge loss, and Huber loss.

[0025] Standard AI models that do not have guidance about hidden nodes may suffer from local minimum trapping due to over-parameterized DNN architectures to solve task problems. To stabilize training convergence, several regularization methods are used. For example, L1 / L2 norm is used to regularize affine transformation weights. Batch normalization and dropout methods are also widely used as general regularization methods to prevent overfitting. Other regularization methods include but are not limited to drop connect, drop block, drop pass, shake drop, spatial drop, zone out, probabilistic depth, probabilistic width, spectral normalization, and shake shake. However, those well-known regularization methods do not exploit the underlying data distribution. Most datasets have certain probabilistic relationships between X and Y and many nuisance factors S that disturb task prediction performance. For example, physiological datasets such as EEG signals are highly dependent on the subjects' mental states and measurement conditions as such nuisance factors S. The nuisance variations include a set of subject identification information, session number, biological state, environmental state, sensor state, position, orientation, sampling rate, time, and sensitivity. For yet another example, an electromagnetic dataset such as a Wi-Fi signal is subject to indoor environment, surrounding users, interference, and hardware imperfections. The present disclosure provides a way to efficiently regularize DNN blocks by considering nuisance factors such that the AI ​​model does not feel the domain shift caused by the change of the nuisance factors. Auxiliary Regularization Module

[0026] A DNN model can be decomposed into an encoder part and a classifier part (or a regressor part for regression tasks), where the encoder part extracts a feature vector from data X as a latent variable Z, and the classifier part predicts a task label Y from the latent variable Z. For example, the latent variable Z is a vector of hidden nodes in the middle layer of a DNN model. An exemplary pipeline of an AI model composed of an encoder block and a classifier block is shown in FIG. 1B. Here, the encoder is given X to generate Z, and the decoder is given Y to predict Z.

[0027] In addition to the main pipeline of the encoder and classifier, FIG. 1B shows a schematic diagram illustrating an exemplary DNN block with an additional auxiliary regularization module for regularizing the latent variables Z. Specifically, a decoder DNN block is attached to the latent variables Z to reconstruct the original data X with additional conditional information of the nuisance variations S. This conditional decoder can facilitate disentanglement of the nuisance domain information S from the latent variables Z. For example, the nuisance domain variations S include subject identification information (ID), measurement session ID, noise level, subject height / weight / age information, etc. for a physiological dataset. By disentangling those nuisance factors S from the latent variables (Z), the present invention can realize a subject-invariant generic human-machine interface without long calibration sessions. The auxiliary DNN block, called the decoder, is trained to minimize another loss function, such as the mean squared error loss or the Gaussian negative log-likelihood loss, to reconstruct X from Z.

[0028] For some embodiments, the latent variation Z is a set of multiple latent factors Z1, Z2, ..., Z L Each of these is further decomposed into a set of nuisance factors Z1, Z2, …, Z LIn addition, some nuisance factors are partially known or unknown depending on the dataset. For known labels of nuisance factors, the DNN block can be trained in a supervised manner, while it requires a semi-supervised manner for unlabeled nuisance factors. For the semi-supervised case, pseudo-labeling based on variational sampling over all potential labels of nuisance factors is used for some embodiments, for example based on the so-called Gumbel softmax reparameterization trick. For example, some of the data in the dataset does not have subject age information, while the rest of the data has age information to be used for supervised regularization.

[0029] The DNN block in FIG. 1B has another auxiliary regularization module attached to the latent variable Z to estimate the nuisance variation S. This regularization DNN block is used to further facilitate the disentanglement of nuisance factors to be robust, and is often called an adversarial network. Because the regularization DNN block is trained to minimize a loss function to estimate S from Z, while the main pipeline DNN block is trained to maximize the loss function to censor the nuisance information. The adversarial block is trained alternately with associated hyperparameters including adversarial coefficients, adversarial learning rate, adversarial alternation interval, and architecture specifications.

[0030] The graphical model in FIG. 1B is known as Adversarial Conditional Variational Autoencoder (A-CVAE) for Unsupervised Feature Extraction for Downstream Task Classifier. The graphical model of A-CVAE has various graph nodes and graph edges to represent the connectivity across the random variable factors X, Y, Z, and S. By using the regularization block, the A-CVAE model has higher robustness against nuisance factors. Therefore, the auxiliary regularization module is used as a so-called pre-shot transfer learning method or domain generalization method to make the AI ​​model robust against unfamiliar nuisance factors. Architecture Exploration

[0031] There are many possible ways to connect the encoder, classifier, decoder, and adversarial network blocks. For example, FIG. 1C shows one exemplary DNN block with another auxiliary model for estimating nuisance factor S from data X. Since the possible number of realizable DNN connectivity rapidly explodes with the size of the AI ​​model, efficient construction of plausible AI models is required. In addition, randomly connected DNN blocks tend to be useless and unjustified. For some embodiments, graph connectivity is explored by using an automated Bayesian graph search method called AutoBayes. The core of AutoBayes is to consider a graphical Bayesian model that captures the probabilistic relationships between random variables representing data features X, task labels Y, nuisance variation labels S, and (latent) latent representations Z. The main objective is to infer task labels Y from measured data features X, which is (partially) hindered by the presence of labeled nuisance variations (e.g., inter-subject / inter-session variations) by S. The latent representations Z (and further Z1, Z2, ..., Z as needed) are then used to estimate the latent representations Z (and further Z1, Z2, ..., Z as needed). L ) is also optionally introduced into these AI models to help capture the underlying relationships between S, X, and Y.

[0032]

number

[0033] The above-mentioned graphical models in Figures 2A, 2B, and 2C do not impose any assumption of potentially inherent independence in the dataset, and are therefore the most comprehensive. However, depending on the underlying independence in the dataset, we may be able to prune some edges in their graphs. For example, if the data has a Markov chain of YX independent of S and Z, it automatically results in Figure 1A. This implies that the most complex inference model with high degrees of freedom does not necessarily perform best across any dataset. It motivates us to consider an extended AutoML framework that automatically explores the best pair of inference factor graph models and corresponding Bayesian graph models that are consistent with the dataset, in addition to other hyperparameter designs.

[0034] AutoBayes starts by exploring any potential Bayesian graph by cutting links in the complete chained graph in Figure 2A and imposing possible independence. Then, we employ the Bayes-Ball algorithm for each hypothetical Bayesian graph to check the conditional independence for different inference strategies, e.g., the complete chained inference graphs in Figures 2B and 2C. Bayes-Ball justifies the reasonable pruning of links in the complete chained inference graphs in Figures 2B and 2C, and also justifies the potential adversarial truncation when Z is independent from S. This process automatically builds the connectivity of the inference block, the generation block, and the adversarial block with sound reasoning, e.g., to build the A-CVAE classifier in Figure 1B from any model in Figure 1C. Exemplary Bayesian Graph Model

[0035] Given sensor measurements such as media, physical, and physiological data, we do not know in advance the true joint probabilities, so we assume one of the possible generative models. Unlike typical AutoML frameworks that search for inference model architectures, AutoBayes aims to explore any such latent graph models to be consistent with the measurement distribution. Even for the four-node case with Y, S, Z, and X, the maximum possible number of graphical models is enormous, so we show several embodiments of such Bayesian graphs in Figures 3A-3K. Each Bayesian graph corresponds to one generative model based on the joint probability factorization.

[0036] Depending on the assumed Bayesian graph, a relevant inference strategy will be determined such that some variables in the inference factor graph are conditionally independent. It allows pruning the links. As shown in Figure 4A-4L, a plausible inference graph model can be automatically generated by the Bayes-Ball algorithm based on each Bayesian graph hypothesis specific to the dataset. For example, the generative model E in Figure 3E can automatically generate the inference factor graph model Ez in Figure 4C. By merging those generative models and inference models, AutoBayes can automatically build a model robust to nuisances based on the A-CVAE in Figure 1B. Bayes-Ball Algorithm

[0037] The system of the present invention relies on the Bayes Ball algorithm to facilitate automatic pruning of links in an inference factor graph through analysis of conditional independence. The Bayes Ball algorithm uses just ten rules to identify conditional independence as shown in Figures 5A-5J. Given a directed Bayes graph, we can determine whether the conditional independence between two disjoint sets of nodes provides conditioning for other nodes by applying the graph separation criterion. Specifically, an undirected path is activated if the Bayes Ball can proceed without encountering a stop arrow symbol in Figures 5A-5J. If there are no active paths between the two sets of nodes when some other conditioning node is shaded, then those sets of random variables are conditionally independent. Using the Bayes Ball algorithm, the present invention generates a list that identifies the independence relations of two disjoint nodes for the AutoBayes algorithm. AutoBayes Algorithm

[0038] FIG. 6 shows the overall procedure of the AutoBayes algorithm described in the pseudocode of Algorithm 1 for a more comprehensive case than only FIGS. 3A-3K and 4A-4L according to some embodiments of the present disclosure. AutoBayes automatically constructs a non-redundant inference factor graph given the hypothesis Bayes graph assumption through the use of the Bayes-Ball algorithm. Depending on the derived conditional independence and the pruned factor graph, the DNN blocks for the encoder, decoder, classifier, nuisance estimator, and adversary are rationally connected. The entire DNN block is trained using adversarial learning in variational Bayes inference. Note that, as an embodiment, the hyperparameters of each DNN block can be further optimized by AutoML in addition to the AutoBayes framework.

[0039] The system of the present invention uses a memory bank to store provisionally calculated values ​​including hyperparameters, trainable variables, intermediate neuron signals, and forward pass signals and backward pass gradients. It reconstructs the DNN block by exploring various Bayesian graphs based on the Bayes-Ball algorithm so that redundant links are pruned to be compact. Depending on the data set, AutoBayes first creates a fully-chained directed Bayesian graph to connect all nodes in a specific permutation order. The system then prunes certain combinations of graph edges in the fully-chained Bayesian graph. Then, the Bayes-Ball algorithm is employed to list conditional independence relations between two disjoint nodes. For each Bayesian graph in the hypothesis, another fully-chained directed factor graph is constructed from the nodes associated with the data signal X, and infers other nodes in a different factorization order. Then, depending on the independence list, pruning of redundant links in the fully-chained factor graph is employed, so that the DNN links can be compact. In another embodiment, redundant links are intentionally maintained and progressively graphed. The pruned Bayes graph and the pruned factor graph are combined to make the generative model and the inference model consistent. Given the combined graphical model, all DNN blocks for the encoder, decoder, classifier, estimator, and adversarial network are associated in the connection to the model. This AutoBayes achieves robust inference against junk and is transferred to new data domains for new test datasets.

[0040] The AutoBayes algorithm can be generalized for more than four node factors. As an example of such an embodiment, the nuisance variation S is calculated by dividing the nuisance variation S1, S2, ..., S2 by the multiple domain side information according to a combination of supervised, semi-supervised, and unsupervised settings. N In another example embodiment, the latent variables are further decomposed into multiple factors of Z1, Z2, ..., Z L1C is one of such embodiments. As an example of an embodiment with decomposed factors, the nuisance fluctuations are grouped into different factors such as subject identification, session number, biological state, environmental state, sensor state, position, orientation, sampling rate, time, and sensitivity.

[0041] When exploring different graphical models, one embodiment uses the output of all the different models explored to improve performance, for example, using a weighted sum to achieve ensemble performance. Yet another embodiment uses an additional DNN block that learns the best weights to combine different graphical models. This embodiment is realized using an attention network to adaptively select the relevant graphical model given the data. Since the original joint probabilities are identical, this embodiment considers consensus equilibrium and voting between different graphical models. For some embodiments, it also recognizes cycle consistency of the encoder / decoder DNN blocks. Abort Mode

[0042] Using AutoBayesian architecture search, we can identify the independence between latent variables Z and nuisance variables S for a given generative model. Auxiliary regularization modules such as adversarial networks and conditional decoders can assist in disentangling the correlation between Z and S for such models. Under the constrained risk minimization framework, there are multiple types of such truncation modes for the auxiliary regularization module to promote the independence between Z and S. In fact, adversarial truncation is not the only way to achieve feature disentanglement. Specifically, we consider some modified learning frameworks, where we implement some notion of independence between the learned representations Z and the nuisance variables S so that the classifier model can achieve similar performance across different domains, for example using the following truncation modes:

[0043]

number

[0044]

number

[0045]

number

[0046] If we have more than two latent representations, the number of truncation modes is naturally increased with the combination of conditional / unconditional and complementary disentanglements.

[0047] The first, marginal truncation mode, captures the simplest notion of a “nuisance-independent representation.” For example, this marginal truncation mode is realized by the adversarial discriminator in the A-CVAE model. This marginal truncation approach does not conflict with the task objective if the distribution of labels does not depend on the nuisance variables, since the nuisance factor S is not useful for downstream tasks to predict Y. However, there may be some correlation between Y and S. Because of this, a representation Z trained to be useful for predicting task label Y may also be informative about S. The second, conditional truncation mode, accounts for this conflict between the task objective and the truncation objective by allowing Z to contain some information about S, but not more than is already implied by the task label Y. For example, the A-CVAE model uses a conditional decoder DNN block to achieve a similar effect of this conditional truncation mode. The third, complementary truncation mode, accounts for this conflict by requiring some parts of the representation Z to be independent of the nuisance variable S while other parts are strongly dependent on the nuisance variables. This truncation mode is illustrated in FIG. 1C.

[0048] These truncation modes lead us to consider a constrained optimization problem that enforces the desired independence. We consider two forms of this restriction: one based on mutual information and one based on the divergence between the two distributions. Specifically, we solve the constrained optimization problem by using Lagrange multipliers as follows:

number

[0049] To estimate independence, we consider several truncation methods for computing mutual information and divergence. For mutual information-based truncation methods, there are several approaches, including but not limited to: · Cross-entropy loss in adversarial nuisance classifiers; · Mutual information neural estimation (MINE); Mutual information gradient estimation (MIGE).

[0050] For some embodiments, in the adversarial nuisance classifier for A-CVAE, the cross-entropy loss is used to estimate the conditional entropy H(s|z). This gives us an estimate for the mutual information, since the mutual information can be decomposed as I(z;s)=H(s)-H(s|z), because the marginal entropy H(s) is constant with respect to the model parameters. The MINE method directly estimates the mutual information, not the cross-entropy, by using a DNN model. However, the main purpose of these truncation methods is to disentangle S from Z, and thus, it is not necessary to explicitly estimate the mutual information for training, but it is necessary to estimate the mutual information gradient. The MIGE method uses a score function estimator to calculate the gradient of the mutual information, and several kernel-based score estimators are known, such as the Spectral Stein Gradient Estimator (SSGE), NuMethod, Tikhonov, Stein Gradient Estimator (SGE), Kernel Exponential Family Estimator (KEF), Nystrom KEF, and Sliced ​​Score Matching (SSM). Kernel-based score estimators have their hyperparameters, such as the kernel length, which can be adaptively selected depending on the dataset.

[0051] Divergence-based truncation methods include several approaches, including but not limited to: Minimum mean discrepancy (MMD) using biased / unbiased kernel estimators; · Pairwise MMD with random subset selection; Boundary equilibrium generative adversarial network (BEGAN) discriminator; · HSIC (Hilbert-Schmidt independence criterion); · Optimal transportation for the Wasserstein distance measure.

[0052] The first two methods rely on a kernel-based estimator of the MMD score, which provides a numerical estimate of the distance between two distributions. It is known that the MMD between two distributions is exactly 0 if these distributions are equivalent. By the definition of conditional probability, the independence z⊥s that we enforce requires that distributions q(z) and q(z|s) are equivalent, or alternatively, that distribution q(z|s i ) and q(z|s j ) is equivalent between any nuisance pair. Thus, we can minimize the MMD between one of these pairs of distributions, enforcing independence of the latent representation Z from the nuisance variables. The first MMD truncation method explores the choice such that q(z)=q(z|s).

[0053] The second pairwise MMD censoring is done by q(z│s i )=q(z|s j ) to find the best fit. To compute an overall score using this "pairwise" approach, we require all combinations of the two distinct values ​​of the nuisance variables and compute the average over these individual terms. To reduce this overhead for computational efficiency, we can consider some approximations of this pairwise MMD truncation method by selecting a subset of the averaging pairs. Figures 7A and 7B show pseudocode for two exemplary subset approximation algorithms. The first algorithm in Figure 7A selects s i ,s jWe use a parameter b ∈ [0,1] that controls the Bernoulli distribution to select a random subset of all possible pairs of (i ≠ j). We call it "Bernoulli" subset selection. The second algorithm in Figure 7B uses an integer d ∈ {1,…,M} that controls the number of junk values ​​included, and considers all combinations in this subset. We call it "clique" subset selection.

[0054] For some embodiments, in the third divergence-based truncation method, we use a neural discriminator based on the BEGAN model. In BEGAN, the discriminator is parameterized as an autoencoder network, which provides a quantitative measure of the divergence between the true and generated data distributions by comparing its own average autoencoder loss on real and fake data. This corresponds to an estimate of the Wasserstein-1 distance between the true and fake autoencoder losses, which provides a stable training signal to enable the generator to align its generated data distribution with the true data distribution. To measure the truncation score, we can use this approach to provide a proxy measure of the divergence between q(s) and q(z|s). Similar to MMD, minimizing this distance allows us to reduce the dependence of S and Z. Automated Transfer Learning: Auto Transfer

[0055] The present disclosure is based on the recognition that there are many algorithms and methods for transfer learning frameworks to make AI models robust to domain shifts and nuisance variations. For example, for various pre-shot regularization methods, there are different truncation modes and truncation methods to disentangle nuisance factors from latent variables as described above. The present disclosure is also based on the recognition that due to the no-free-lunch theorem, there is no single transfer learning approach that can achieve the best performance across any arbitrary dataset. Therefore, the core of the present invention is to automatically explore different transfer learning approaches suitable for the target dataset in addition to the architecture search based on the AutoBayes framework. The method and system of the present invention is called Auto-Transfer, which performs an automated search for the best transfer learning approach across a set of algorithms.

[0056] FIG. 8 shows an exemplary schematic diagram of the auto-transfer framework. The AI ​​model has a main pipeline for predicting Y from X via an encoder model f and a classifier model g. The encoder model and the classifier model are specified by some trainable parameters. The latent variables Z are generated by the encoder model in the middle layer of the main pipeline. There is a set of auxiliary regularization modules or blocks for disentangling nuisance variations S from the latent variables Z. For example, data X is a measurement from an electroencephalogram (EEG) sensor of subject ID S for predicting motor imagery class Y for a brain-computer interface system. The key element in the auto-transfer of the present disclosure is to use a set of different regularization modules for exploration, because a certain truncation algorithm may work well in some situations but impair task prediction performance in different situations. The set of auxiliary regularization modules are based on different truncation modes, such as marginal truncation, conditional truncation, and complementary truncation, and are based on different truncation methods, including but not limited to adversarial truncation, MINE truncation, MIGE truncation, BEGAN discriminator truncation, MMD truncation, pairwise MMD truncation, HSIC truncation, and optimal transport truncation. In some embodiments, multiple truncation algorithms are used simultaneously, similar to the A-CVAE model.

[0057] The latent variable Z should be discriminative enough to predict Y, while Z should be invariant among different nuisance variations S. For example, if the distribution of Z is well clustered depending on the task label Y, it generally results in higher task classification performance. However, if the cluster distribution is sensitive to subject differences when changing the brain-computer interface from subject S1 to another subject S2, it may have lower generalizability for a completely new and unfamiliar subject. Although a set of different censoring modules may implement a subject-invariant latent representation Z, some of them may over-censor the nuisance factors, which in turn may degrade the task performance. The present invention enables the auto-transfer framework to automatically find the best truncation module from a set of regularization modules. For example, the best regularization module may be identified by using external optimization methods, including but not limited to reinforcement learning, evolutionary strategies, differential evolution, particle swarms, genetic algorithms, annealing, Bayesian optimization, hyperbands, and multi-objective Lamarckian evolution, to explore different combinations of discrete and continuous hyperparameter values ​​that identify the regularization module. Specifically, a set of best module pairs can be derived automatically by measuring expected task performance in a validation dataset. In some embodiments, the best regularization modules are further combined by ensemble stacking, such as linear regression, multilayer perceptrons, or attention networks, in a cross-validation setting.

[0058] 9A, 9B, and 9C show example pseudocodes describing the adversarial truncation method used as one of the regularization modules in the marginal truncation mode, the conditional truncation mode, and the complementary truncation mode, respectively. These adversarial truncation modules consider minimizing the conditional mutual information between Z and S given Y using an adversarial nuisance classifier model that maps a latent representation Z to a probability distribution over nuisance variables S. Specifically, we train the parameters of the adversarial module to minimize the standard cross-entropy loss for the prediction task. This can be seen as minimizing an upper bound on the conditional entropy H(s|z).

[0059] 10A, 10B, and 10C show example pseudocodes describing the MIGE truncation method used as one of the regularization modules in the marginal truncation mode, the conditional truncation mode, and the complementary truncation mode, respectively. Considering the difficulty of estimating mutual information in high dimensions, MIGE provides an efficient method for directly estimating the gradient of the mutual information. This is sufficient for regularization where an objective function containing the mutual information term will be minimized by gradient descent. Specifically, MIGE calculates the gradient of the mutual information by sampling from an implicit pushforward distribution q(x,y,s) for sampling tuples (x,y,s) from the data distribution and its latent representation Z. The MIGE truncation method uses several score function estimators, including but not limited to SSGE, k-score, the Nu method; Tikhonov, and Stein. One advantage of MIGE truncation includes the fact that MIGE truncation does not require alternating optimization used for adversarial training, which is often unstable or sensitive to adversarial coefficients.

[0060] 11A, 11B, and 11C show example pseudocodes describing the MMD truncation method used as one of the regularization modules in the marginal, conditional, and complementary truncation modes, respectively. The MMD truncation method serves as a desirable measure of divergence between two distributions because it makes no assumptions about the parametric form of the distributions under measurement and it can be efficiently and easily approximated using a kernel estimator from a batch of samples. MMD is an integrated probability metric that describes the divergence between two distributions as the difference between the expected values ​​of a test function under each distribution (from a class of functions for some worst case). The MMD truncation method provides an unbiased empirical estimate of the squared MMD score using a unit sphere in a generalized reproducing kernel (kernel) Hilbert space using a suitable kernel function. Note that this estimate includes the hyperparameters that define the kernel, such as the length scale of the radial basis function (RBF) kernel matrix. This length scale can be adjusted by several methods, such as a median heuristic. Specifically, each time we construct a kernel matrix for a batch of samples, we set the length scale to the median pairwise L2 distance between points in that batch. To compute this conditional truncation penalty for a batch of encoded examples, we compute a term for each class-conditional subset of the batch and average these terms. We weight each term in the average using the inverse class frequency, which corresponds to enforcing a uniform class prior and takes into account possible class imbalance in our batching procedure.

[0061] 12A, 12B, and 12C show exemplary pseudocodes describing the pairwise MMD truncation method used as one of the regularization modules in the marginal truncation mode, the conditional truncation mode, and the complementary truncation mode, respectively. Similar to the MMD truncation method, the pairwise MMD truncation method calculates a penalty to minimize the average divergence between each nuisance conditional distribution using a quantitative surrogate. This pairwise MMD truncation approach enforces conditional independence by calculating like terms for each class conditional subset of the batch. As before, in some embodiments, we may take a weighted average between classes to account for possible class imbalance in our sample batch. Subset selection is achieved, for example, by the Bernoulli and clique approximations in FIG. 7A and FIG. 7B.

[0062] 13A, 13B, and 13C show example pseudocode describing the pairwise MMD truncation method used as one of the regularization modules in the marginal, conditional, and complementary truncation modes, respectively. BEGAN uses an adversarial training scheme to learn a generative model. The generator network tries to approximately map samples from a Gaussian distribution in its latent space to samples from a target data distribution, while the discriminator network tries to distinguish between true and fake data samples. The key element of this model is that it uses an autoencoder as a discriminator, with a training objective designed for the discriminator to compute a lower bound on the Wasserstein-1 distance between the distributions of its autoencoder losses for the true and generated data. In other words, the discriminator distinguishes between the two data distributions by trying to learn an autoencoder map that works well only for the "true" data distribution, while the generator tries to generate data that is consistent with the "true" data distribution and is therefore well preserved by this autoencoder map. They further stabilize the training of their discriminator model by introducing a trade-off parameter to adaptively scale the magnitude of the discriminator's loss term on the true and generated data. This allows for successful training without the need for common GAN training tricks such as custom scheduling or pre-training of one of the models. The role of the discriminator is to provide a surrogate objective so that the generator can bring two distributions from different domains closer together. This can be easily adapted to provide a signal that allows the encoder model to minimize divergence. We use the alternating optimization algorithm from BEGAN, but replace the distribution of the "true" data with q(z) and the distribution of the "generated" data with q(z|s). We calculate the loss term in this manner for each possible value of the nuisance variables and take the average of these values. Note that in alternating optimization, the discriminator and the encoder are optimized in separate steps using the two loss terms.The BEGAN optimization algorithm includes additional inputs that control the relative magnitudes of these loss terms to maintain a balance between these two models. Automated pre-processing / post-processing

[0063] The above description of the auto-transfer framework in the present invention is particularly suitable for pre-shot transfer learning, also known as domain generalization, in the case where there is no test dataset available in the new target domain. Nevertheless, auto-transfer can also improve post-shot transfer learning, also known as online domain adaptation, for high resilience to domain shifts. Post-shot learning includes zero-shot learning, where unlabeled data in the target domain is available, and few-shot learning, where some labeled data in the target domain is available to fine-tune the pre-trained AI model. For some embodiments, post-shot fine-tuning is performed online during the progress of the testing phase, when new data is available, with or without task labels. In the post-shot adaptation phase, the pre-trained AI model optimized by auto-transfer is further updated by a set of calibration datasets in a new user or target domain. The updates are achieved through domain adaptation techniques including, but not limited to, pseudo-labeling, soft labeling, confusion minimization, entropy minimization, feature normalization, weighted z-scoring, continuous learning with elastic weight integration, FixMatch, MixUp, label propagation, adaptive layer freezing, hypernetwork adaptation, latent space clustering, quantization, sparsification, zero-shot semi-supervised updates, and few-shot supervised fine-tuning.

[0064] According to some embodiments of the present invention, in a similar manner to exploring different truncation methods, auto-transfer can search for the best post-shot adaptation method among different available approaches. Figures 14A, 14B, 14C, and 14D show an exemplary set of post-processing modules to select at the post-shot stage. The selection is achieved by various optimization methods, including Bayesian optimization for a new validation dataset in the source or target domain.

[0065] FIG. 14A shows an exemplary schematic diagram of the FixMatch method. Weakly augmented data is fed into an AI model to obtain predictions. If the predicted score is above a threshold, the prediction is converted to a one-hot pseudo-label. We then calculate the model's predictions for the strong augmentation of the same data. The model is trained to align its predictions for the strongly augmented version with the pseudo-labels via cross-entropy loss minimization.

[0066] Figure 14B shows another exemplary post-processing method based on semi-supervised learning with compact latent space clustering, which dynamically builds a graph in the latent space at each training iteration, propagates the labels to capture the structure of the manifold, and regularizes it to form a single compact cluster per class, facilitating separation.

[0067] Fig. 14C shows another example of a post-processing method based on continuous learning with elastic weight integration. It ensures that task A is memorized while training on task B. The training trajectory is plotted in parameter space, where the parameter region leads to good performance on task A and task B. If we perform gradient steps according to task B only, we may minimize the loss on task B but lose what we have learned on task A. On the other hand, if we constrain each weight with the same coefficient, the imposed restrictions are too strict and we cannot memorize task A without incurring significant loss on task A, at the expense of not learning task B. Thus, elastic weight integration, on the contrary, finds a solution for task B without incurring significant loss on task A by explicitly calculating how important the weights are for task A.

[0068] In the FixMatch illustration, a weakly augmented image (top) is fed into the model to obtain a prediction. If the model assigns a probability to any class above a threshold (dotted line), the prediction is converted to a one-hot pseudo label. The model's prediction for a strongly augmented version of the same image (bottom) is then computed. The model is trained to align its prediction for the strongly augmented version with the pseudo label via cross-entropy loss.

[0069] Figure 14D shows yet another example of a post-processing method based on label propagation for semi-supervised learning. Triangles indicate labeled training data and circles indicate unlabeled training data. Ground truth labels are propagated to generate pseudo labels inferred by diffusion, which are used to train an AI model according to the confidence of the pseudo label predictions on the manifold.

[0070] According to some embodiments of the invention, the method dynamically builds a graph in the latent space of the network at each training iteration, propagating the labels to capture the structure of the manifold and regularizing it to form a single compact cluster per class, facilitating separation.

[0071] Elastic weight consolidation (EWC) ensures that task A is memorized while training on task B. The training trajectory is illustrated in the schematic parameter space, where the parameter region leads to good performance on tasks A and B. After learning the first task, the parameters are θ * A If we take the gradient step according to task B only (arrow (C)), the method minimizes the loss for task B. If we constrain each weight with the same coefficient (arrow (B)), the restrictions imposed are too strict and we cannot memorize task A without paying the cost of not learning task B. EWC does the opposite: by explicitly calculating how important the weights are for task A, we find a solution for task B without incurring significant loss for task A (arrow (A)).

[0072] In label propagation for a very simple toy example of a manifold, triangles indicate labeled training data and circles indicate unlabeled training data. The top figure shows the color-coded ground truth for labeled points and gray for unlabeled points. The bottom figure shows the color-coded pseudo-labels inferred by diffusion that are used to train the CNN. In this case, the size reflects the certainty of the pseudo-label prediction.

[0073] In addition to post-processing, autotransfer can explore different pre-processing approaches before feeding the raw data to the AI ​​model. Pre-processing methods include, but are not limited to, data normalization, data augmentation, auto-augmentation, a universal adversarial example (UAE), spatial filtering such as common spatial pattern filtering, principal component analysis, independent component analysis, short-time Fourier transform, filter bank, vector autoregressive filter, auto-attention mapping, robust z-scoring, spatio-temporal filtering, and wavelet transform. For example, a probabilistic UAE that adversarially disrupts task classification is used as a data augmentation to tackle more challenging artifacts in the dataset. There are many associated hyperparameters to specify pre-processing. For example, the continuous wavelet transform can have a choice of filter bank resolution and mother wavelet kernel (such as the Mexican hat wavelet shown in FIG. 15D, the Morlet wavelet shown in FIG. 15E, and the Gauss8 wavelet shown in FIG. 15F). Since there are many preprocessing approaches with many associated hyperparameters, an automated search is needed to find the best approach without intensive human effort. Figure 15 shows an exemplary set of preprocessing modules used for automatic selection. In some embodiments, auto-transfer also automatically explores various such preprocessing methods so that the AI ​​model can achieve high accuracy in task prediction while achieving robustness against domain shifts. In some embodiments, the selection is achieved by Bayesian optimization.

[0074] FIG. 15A shows an example of a preprocessing method based on auto-augmentation. It uses a search method (e.g., reinforcement learning) to find a better data augmentation policy. An auxiliary controller model (e.g., a recurrent neural network (RNN)) predicts the augmentation policy from the search space. A child network with a fixed architecture is trained toward convergence to achieve accuracy. The accuracy score will be used as a reward value with a policy gradient method to update the controller model so that it can generate better policies over time. The augmentation policies include, but are not limited to, noise injection, space-time shifting / masking, interpolation / extrapolation, and quantization.

[0075] FIG. 15B shows another example of a pre-processing module based on MixUp. It augments the training data by overlapping two separate data with randomly sampled mixing coefficients. The key idea of ​​MixUp is to mix the task label Y in addition to the data X. For some embodiments, MixUp uses multiple data instances, and three or more are combined with more mixing parameters. Yet another embodiment uses an auxiliary DNN model to mix multiple samples from the training data and generate a nonlinear mapping to be augmented.

[0076] Figure 15C shows another example of augmentation based on the UAE framework. The auxiliary DNN model is trained as an offline generator to generate a generic adversary given a training dataset. In the online stage, the DNN model can generate the adversary without backpropagation or model-specific gradients while it tries to perturb the task accuracy as much as possible as a worst-case artifact under a constrained confusion limit. The trained DNN model is then used to augment the training data when the main AI model is trained based on the augmented data. Due to the adversarial attack by the UAE framework, the main AI model can be more generalized to adversarial domain shifts.

[0077] The figure shows an overview of our framework, which uses a search method (e.g., reinforcement learning) to search for better data augmentation policies. A controller RNN predicts an augmentation policy that is trained towards convergence to achieve accuracy R. The reward R will be used with a policy gradient method to update the controller so that it can generate better policies over time. Model realization example

[0078] Each of the DNN blocks is configured with hyperparameters to specify a set of layers with trainable variables and interconnected neuron nodes to pass signals from layer to layer sequentially. The trainable variables are numerically optimized using gradient methods such as stochastic gradient descent, adaptive momentum, adaptive gradient, adaptive boundary, Nesterov accelerated gradient, and root mean square propagation. The gradient methods update the trainable parameters of the DNN blocks by using the training data such that the output of the DNN blocks provides smaller loss values ​​such as mean squared error, cross entropy, structural similarity, negative log likelihood, absolute error, cross covariance, clustering loss, divergence, hinge loss, Huber loss, negative sampling, Wasserstein distance, and triplet loss. The multiple loss functions are further weighted with several regularization coefficients according to a training schedule policy.

[0079] In some embodiments, the DNN block is reconfigurable according to hyperparameters such that the DNN block is configured with a set of fully connected layers, convolutional layers, graph convolutional layers, recurrent layers, loopy connections, skip connections, and initiation layers with a set of nonlinear activations including rectified linear transformation, hyperbolic tangent, sigmoid, gated linear, softmax, and threshold. The DNN block is further regularized using a set of dropout, swapout, zoneout, blockout, dropconnect, noise injection, jitter, and batch normalization. In yet another embodiment, the layer parameters are further quantized to reduce the size of the memory as specified by the adjustable hyperparameters. For another embodiment of link concatenation, the system uses multidimensional tensor projection with dimension-wise trainable linear filters to convert lower dimensional signals to higher dimensional signals for dimension-mismatched links.

[0080] Another embodiment integrates AutoML with AutoBayes and AutoTransfer for hyperparameter search and learning scheduling of each DNN block. Note that AutoTransfer and AutoBayes can be easily integrated with AutoML to optimize any hyperparameter of an individual DNN block. More specifically, the system modifies hyperparameters by using reinforcement learning, evolution strategies, differential evolution, particle swarms, genetic algorithms, annealing, Bayesian optimization, hyperbands, and multi-objective Lamarckian evolution to explore different combinations of discrete and continuous hyperparameter values.

[0081] The system of the present invention also provides a further test step to adapt the trained DNN block as a refinement post-training step by unpacking some trainable variables so that the DNN block can be robust to new data sets with new nuisance variations such as new subjects. This embodiment can reduce the calibration time requirements for new users of the HMI system. Yet another embodiment uses the exploration of different pre-processing methods. Exemplary System

[0082] FIG. 16 is a block diagram illustrating an example of a system 500 for automated construction of artificial neural network architectures, according to some embodiments of the present disclosure. The system 500 includes a set of interfaces and data links 105 configured to receive and transmit signals, at least one processor 120, a memory (or a set of memory banks) 130, and a storage 140. The processor 120 is connected to the memory 130 to execute computer executable programs and algorithms stored in the storage 140. The set of interfaces and data links 105 may include a human machine interface (HMI) 110 and a network interface controller 150. The computer executable programs and algorithms stored in the storage 140 may be a reconfigurable deep neural network (DNN) 141, hyperparameters 142, a scheduling criterion 143, forward / backward data 144, a temporary cache 145, an abort module 146, an autotransition algorithm 147, and a pre-processing / post-processing module 148.

[0083] The system 500 can receive signals through a set of interfaces and data links. The signals may be a data set of training data, validation data, and test data, and the signals include a set of random factors in a multidimensional signal X, some of which are associated with task labels Y for identifying and nuisance variations S from different domains.

[0084] In some cases, each of the reconfigurable DNN blocks (DNN) 141 is configured to encode a multidimensional signal X into a latent variable Z, decode the latent variable Z to reconstruct the multidimensional signal X, classify a task label Y, estimate nuisance variance S, regularize and estimate nuisance variance S, or select a graphical model. In this case, the memory bank further includes provisional calculations including hyperparameters, trainable variables, intermediate neuron signals, and forward pass signals and backward pass gradients.

[0085] At least one processor 120 is connected to the interface and memory bank 130 and configured to present signals and datasets to the reconfigurable DNN block 141. The at least one processor 120 also performs a Bayesian graph search using a Bayes-Ball algorithm to reconfigure the DNN block to make it compact by modifying hyper-parameters 142 in the memory bank 130, where redundant links are pruned. The auto-transfer explores different auxiliary regularization and pre-processing / post-processing modules to improve robustness against nuisance variations.

[0086] The system 500 may be applied to the design of a human-machine interface (HMI) through the analysis of the user's physiological data. The system 500 may receive physiological data 195B as the user's physiological data via the network 190 and a set of interfaces and data links 105. In some embodiments, the system 500 may receive electroencephalogram (EEG) and electromyogram (EMG) as the user's physiological data from a set of sensors 111.

[0087] The above-described embodiments of the present invention may be implemented in any of many ways. For example, the embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code may be executed on any suitable processor or collection of processors, whether the processors are provided in a single computer or distributed among multiple computers. Such a processor may be implemented as an integrated circuit having one or more processors within an integrated circuit component. However, the processor may be implemented using circuits in any suitable format.

[0088] Also, embodiments of the invention may be embodied as a method, examples of which are provided. The operations performed as part of the method may be ordered in any suitable manner. Thus, while an example embodiment may show operations as sequential, embodiments may be constructed in which operations are performed in an order different from that illustrated, including even including performing some operations simultaneously.

[0089] The use of ordinal terms such as "first," "second," etc. in the claims to modify claim elements does not, by itself, imply any priority, precedence, or ordering of one claim element over another claim element, or the temporal order in which method operations are performed, but is merely used as a label to distinguish one claim element having a certain name from another element having the same name (except for the use of the ordinal terminology).

[0090] Although the invention has been described by way of examples of preferred embodiments, it is to be understood that various other adaptations and modifications can be made within the spirit and scope of the invention.

[0091] Therefore, it is the object of the appended claims to cover all such variations and modifications as come within the true spirit and scope of the invention.

Claims

1. 1. A system for automated construction of artificial neural network architectures, comprising: The system further comprises a set of interfaces and data links configured to transmit and receive signals, the signals including a data set of training data, validation data, and test data, the signals including a set of random variable factors in a multi-dimensional signal X, some of the random variable factors being associated with task labels Y for identifying and nuisance variance S, the system further comprising: The system further comprises a set of memory banks for storing a set of reconfigurable deep neural network (DNN) blocks, each of the reconfigurable DNN blocks being configured with a main task pipeline module for identifying the task label Y from the multi-dimensional signal X, and a set of auxiliary regularization modules for adjusting the disentanglement between a plurality of latent variables Z and the nuisance variations S, the memory banks further comprising hyper-parameters, trainable variables, intermediate neuron signals, and preliminary computed values ​​including forward pass signals and backward pass gradients, and the system further comprises: A system comprising at least one processor connected to the interface and the memory bank and configured to present the signal and the dataset to the reconfigurable DNN block, wherein the at least one processor is configured to perform a search across a set of graphical models, a set of pre-shot regularization methods, a set of pre-processing methods, a set of post-processing methods, and a set of post-shot adaptation methods to reconfigure the reconfigurable DNN block so that task prediction is insensitive to the nuisance fluctuations S by modifying the hyper-parameters in the memory bank.

2. The at least one processor further comprises: modifying the hyperparameters to identify the set of graphical models representing a Bayesian graph model and an inference factor graph based on a Bayes-Ball algorithm; Modifying the reconfigurable DNN block by linking graph nodes with graph edges to associate with the random variable factors for the multi-dimensional signal X, the task labels Y, the nuisance variations S, and the latent variables Z according to the Bayesian graph model and the inference factor graph; training the reconfigurable DNN block on the training data using variational sampling and gradient methods; selecting the hyperparameters based on the output of the reconfigurable DNN block for the validation data; 2. The system of claim 1, further comprising: a step of testing the trained reconfigurable DNN block on the ongoing test data and new incoming data to be transferred with robustness against junk.

3. The at least one processor further comprises: performing a step of modifying the hyper-parameters to identify the set of pre-shot regularization methods based on different truncation modes and truncation methods, the truncation modes being based on a marginal truncation mode, a conditional truncation mode, a complementary truncation mode, or a combination thereof, and the truncation methods being based on a divergence truncation method, a mutual information truncation method, and variations thereof; and the at least one processor further comprising: associating the set of auxiliary regularization modules with the reconfigurable DNN block such that at least one of the latent nodes Z is disentangled from at least one of the nuisance variations S according to the set of pre-shot regularization methods; training the reconfigurable DNN block based on the training data using the set of auxiliary regularization modules; and selecting, for the validation data, the hyper-parameters for a set of the truncation modes and a set of the truncation methods based on an output of the reconfigurable DNN block.

4. 4. The system of claim 3, wherein the truncation method comprises an adversarial truncation method, a mutual information neural estimation (MINE) truncation method, a mutual information gradient estimation (MIGE) truncation method, a minimum mean discrepancy (MMD) truncation method, a pairwise maximum mean discrepancy (MMD) truncation method, a boundary equilibrium generative adversarial network (BEGAN) discriminator truncation method, a Hilbert-Schmidt independence criterion (HSIC) truncation method, an optimal transport truncation method, and variations thereof.

5. The at least one processor revising the hyperparameters to identify the set of preprocessing methods based on spatial filtering, spatio-temporal filtering, wavelet transform, vector autoregressive filter, self-attention mapping, robust z-scoring, normalization, data augmentation, generalized adversarial examples, and variations thereof; and modifying the training data, validation data, and test data according to the set of pre-processing methods and providing them in the reconfigurable DNN block.

6. The system of claim 1 , wherein the set of post-processing methods includes cross-validation voting, ensemble stacking, score averaging, and variations thereof.

7. 2. The system of claim 1, wherein the set of post-shot adaptation methods includes pseudo-labeling, soft labeling, confusion minimization, entropy minimization, feature normalization, weighted z-scoring, elastic weight integration, label propagation, adaptive layer freezing, hypernetwork adaptation, latent space clustering, quantization, sparsification, and variations thereof, and the reconfigurable DNN block is refined by unfreezing the combinations of trainable variables so that the reconfigurable DNN block adapts to a new domain dataset.

8. The system of claim 2, wherein the variational sampling is employed for the latent variables with independent distributions specified by exponential or non-exponential distribution families as their prior distributions for a reparameterization trick, and for categorical variables of unknown nuisance variations and task labels using a Gumbel softmax trick to generate near-one-hot vectors based on a random number generator and a softmax temperature.

9. The system of claim 2 , wherein the linking further comprises a step of multi-dimensional tensor projection using a plurality of trainable linear or bilinear filters to transform the lower dimensional signals for dimensionally mismatched links.

10. 2. The system of claim 1, wherein the reconfigurable DNN block is configured with a combination of fully connected layers, convolutional layers, graph convolutional layers, recurrent layers, loopy connections, skip connections, and initiation layers with a set of nonlinear activations including rectified linear transformations, hyperbolic tangent, sigmoid, gated linear, softmax, and thresholding, and is regularized using a combination of dropout, swap out, zone out, block out, drop connect, noise injection, jitter, and batch normalization.

11. 3. The system of claim 2, wherein the training step includes updating trainable parameters of the reconfigurable DNN block by using the training data such that an output of the reconfigurable DNN block provides a smaller loss value in a combination of objective functions, the objective functions further including a combination of mean squared error, cross entropy, structural similarity, negative log-likelihood, absolute error, cross-covariance, clustering loss, divergence, hinge loss, Huber loss, negative sampling, Wasserstein distance, and triplet loss, and the loss functions are weighted with a plurality of regularization coefficients adjusted according to a specified training schedule.

12. The system of claim 2 , wherein the gradient method employs a combination of stochastic gradient descent, adaptive momentum, adaptive gradient, adaptive boundary, Nesterov accelerated gradient method, and root-mean-square propagation to optimize the trainable parameters of the reconfigurable DNN block.

13. The data set is Media data such as images, photographs, movies, text, characters, voice, music, audio, sounds, and variations thereof; Physical data such as radio waves, optical signals, electrical pulses, temperature, pressure, acceleration, velocity, vibration, force, and deformations thereof; Physiological data such as heart rate, blood pressure, mass, hydration, electroencephalogram, electromyogram, electrocardiogram, mechanomyogram, electrooculogram, galvanic skin response, magnetoencephalogram, electrocorticogram, and variations thereof; The system of claim 1 further comprising a combination of sensor measurements comprising:

14. The system of claim 1 , wherein the nuisance variations include a set of subject identification information, a session number, a biological state, an environmental state, a sensor state, a position, an orientation, a sampling rate, a time, and a sensitivity.

15. 2. The system of claim 1, wherein each of the reconfigurable DNN blocks further includes hyperparameters specifying a set of layers having a set of artificial neuron nodes, where pairs of neuron nodes from adjacent layers are interconnected with a plurality of trainable variables and activation functions to sequentially pass signals from the previous layer to the next layer.

16. The nuisance fluctuation S is a fluctuation S as a plurality of domain side information according to a combination of a supervised setting, a semi-supervised setting, and an unsupervised setting. 1 , S 2 , …, S N The latent variables are further decomposed into multiple factors of Z 1 , Z 2 , …, Z L The system of claim 1 , wherein the factor is further decomposed into a plurality of factors of:

17. 3. The system of claim 2, wherein the step of modifying the hyper-parameters employs a combination of reinforcement learning, evolutionary strategies, differential evolution, particle swarms, genetic algorithms, annealing, Bayesian optimization, hyperbands, and multi-objective Lamarckian evolution to explore different combinations of discrete and continuous hyper-parameter values.

18. 2. The system of claim 1 , wherein the set of hyperparameters includes a set of training schedules including adaptive control of learning rate, regularization weights, factorization permutations, and a policy for pruning low priority links by measuring discrepancy between the training data and the validation data using belief propagation.

19. 1. A computer-implemented method for the automated construction of artificial neural network architectures, comprising: The method includes providing a dataset of training data, validation data, and test data, the dataset including a set of random variable factors in a multidimensional signal X, some of the random variable factors being associated with a task label Y for identifying, and nuisance variation S, the method further comprising: configuring a set of reconfigurable deep neural network (DNN) blocks for identifying the task label Y from the multi-dimensional signal X, the set of reconfigurable DNN blocks including a set of auxiliary regularization modules for adjusting disentanglement between a plurality of latent variables Z and the nuisance variation S, the method further comprising: training the set of reconfigurable DNN blocks via stochastic gradient optimization on the training data for accurate task prediction; and searching the set of auxiliary regularization modules for the validation data to search for best hyper-parameters such that the task prediction is insensitive to the nuisance variance S.

20. The set of auxiliary regularization modules is based on different truncation modes and truncation methods, the truncation modes being: including a marginal truncation mode, a conditional truncation mode, a complementary truncation mode, or a combination thereof; The truncation method includes: It is based on divergence truncation methods, mutual information truncation methods, and their variants, The truncation method further comprises:

20. The method of claim 19, including an adversarial truncation method, a mutual information neural estimator truncation method, a mutual information gradient estimator truncation method, a maximum mean discrepancy truncation method, a pairwise maximum mean discrepancy truncation method, a bounded balanced generative adversarial network discriminator truncation method, a Hilbert-Schmidt independence criterion truncation method, an optimal transport truncation method, and variations thereof.

Citation Information

Patent Citations

  • Transition learning system

    JP2021174180A