Hyper-dimensional and hyper-spherical graphical models

Hyperdimensional PGMs address scalability issues in probabilistic graphical models by optimizing variational parameters within the model, enabling efficient posterior inference and flexible conditioning, thus enhancing the quality of generated samples and inferred distributions in high-dimensional tasks.

WO2026161370A1PCT designated stage Publication Date: 2026-07-30RICHARDS WILLIAM D
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
RICHARDS WILLIAM D
Filing Date
2026-01-20
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing probabilistic graphical models face scalability challenges in high-dimensional settings, particularly in tasks like high-resolution image generation and large-scale language modeling, due to computationally intractable posterior inference over latent variables, limiting their application in resource-constrained or latency-sensitive environments.

Method used

Hyperdimensional probabilistic graphical models (PGMs) enable efficient posterior inference in complex, high-dimensional latent spaces by optimizing variational parameters directly within the model, rather than relying on neural-network-based recognition models, allowing scalable inference and flexible conditioning on arbitrary subsets of input variables.

Benefits of technology

This approach allows for deeper and more expressive latent-variable models, improving the quality of generated samples and inferred distributions, and expanding applicability to high-dimensional generative tasks without requiring retraining, while maintaining computational efficiency and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2026011847_30072026_PF_FP_ABST
    Figure US2026011847_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Partial input data corresponding to a subset of variables of a data sample are embedded into a variable space of a probabilistic graphical model containing a plurality of transformation invariant components. Variational parameters associated with one or more latent variables of the probabilistic graphical model are determined by optimizing a probabilistic objective reflecting likelihood of the partial input data by adjusting variational parameters governing the probability distribution over latent variables in the probabilistic graphical model. Latent variable values are sampled from a posterior distribution defined by the variational parameters. Output data is generated by converting the sampled latent variables into output data corresponding to an application domain.
Need to check novelty before this filing date? Find Prior Art

Description

HYPER-DIMENSIONAL AND HYPER- SPHERICAL GRAPHICAL MODELS CROSS REFERENCE TO OTHER APPLICATIONS

[0001] This application claims priority to U.S. Patent Application No. 19 / 444,735 entitled HYPER-DIMENSIONAL AND HYPER-SPHERICAL GRAPHICAL MODELS filed January 09, 2026. which claims priority to U.S. Provisional Patent Application No. 63 / 749,184 entitled HYPER-DIMENSIONAL AND HYPER-SPHERICAL GRAPHICAL MODELS filed January 24, 2025, each of which is incorporated herein by reference for all purposes.BACKGROUND OF THE INVENTION

[0002] Recent advances in machine learning have led to the development of large-scale neural network architectures for a wide range of generative tasks, such as text generation, image synthesis, and video generation. Such architectures typically rely on deep, feed-forward or attention-based networks that implicitly encode both a probabilistic model of observed data and an associated inference process over internal hidden representations. As model capacity and task complexity increase, these approaches require substantial computational resources, memory, and training data, which can limit their practicality in resource-constrained or latency-sensitive environments.

[0003] Probabilistic graphical models, by contrast, explicitly represent latent variables and their conditional dependencies, enabling structured reasoning and interpretability.However, in high-dimensional settings, accurate posterior inference over latent variables often becomes computationally intractable, particularly as the number of variables and interactions increases. As a result, existing graphical-model based generative techniques have encountered significant scalability challenges and have been difficult to apply effectively to large, complex data domains such as high-resolution image generation and large-scale language modeling.BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Various embodiments of the invention are disclosed in the following detailed description and the accompanying drawings.

[0005] FIG. 1 illustrates an example probabilistic graphical model.

[0006] FIG. 2 shows a flow diagram of an example process for generating a sample from a hyper-dimensional neural network in accordance with some embodiments.

[0007] FIG. 3 shows a flow diagram schematic for an example training procedure for determining the weights of a hyperdimensional graphical model in accordance with some embodiments.

[0008] FIG. 4 is a block diagram illustrating a system to generate output samples conditioned on partial inputs using a probabilistic graphical model in accordance with some embodiments.

[0009] FIG. 5 is a flow diagram illustrating a process to generate output samples conditioned on partially specified input data using a probabilistic graphical model in accordance with some embodiments.

[0010] FIG. 6 illustrates an example data flow within a graphical model system configured to generate output data samples conditioned on partially specified input data, in accordance with some embodiments.DETAILED DESCRIPTION

[0011] The invention can be implemented in numerous ways, including as a process; an apparatus; a system; a composition of matter; a computer program product embodied on a computer readable storage medium; and / or a processor, such as a processor configured to execute instructions stored on and / or provided by a memory coupled to the processor. In this specification, these implementations, or any other form that the invention may take, may be referred to as techniques. In general, the order of the steps of disclosed processes may be altered within the scope of the invention. Unless stated otherwise, a component such as a processor or a memory described as being configured to perform a task may be implemented as a general component that is temporarily configured to perform the task at a given time or a specific component that is manufactured to perform the task. As used herein, the term "processor’ refers to one or more devices, circuits, and / or processing cores configured to process data, such as computer program instructions.

[0012] The techniques described herein may be implemented using one or morecomputing devices that are communicatively coupled over a network, such that different portions of the techniques are performed by different devices. For example, operations may be performed by client devices, servers, cloud computing platforms, edge devices, gateways, or any combination thereof, and data and / or control signals may be exchanged between such devices via wired or wireless communication links. Accordingly, the techniques may be implemented in distributed, client-server, cloud-based, edge-based, or hybrid computing environments, and references to a '"system,” ""apparatus,” or ‘"processor” encompass collections of network-connected devices that cooperatively perform the disclosed operations.

[0013] A detailed description of one or more embodiments of the invention is provided below along with accompanying figures that illustrate the principles of the invention. The invention is described in connection with such embodiments, but the invention is not limited to any embodiment. The scope of the invention is limited only by the claims and the invention encompasses numerous alternatives, modifications and equivalents. Numerous specific details are set forth in the following description in order to provide a thorough understanding of the invention. These details are provided for the purpose of example and the invention may be practiced according to the claims without some or all of these specific details. For the purpose of clarity, technical material that is known in the technical fields related to the invention has not been described in detail so that the invention is not unnecessarily obscured.

[0014] Aspects of this disclosure relate generally to probabilistic graphical models (PGMs), generative neural networks, and hybrid generative modeling techniques. Such models are commonly trained to leam a probability distribution over observed data by maximizing the likelihood of a training dataset or an similar objective, in some cases taking into account auxiliary objectives, including for example being more preferable to human annotators.

[0015] In many implementations, training is performed by minimizing a divergence between a distribution of generated samples and an underlying data distribution. For latent-variable models, this often involves estimating a posterior distribution over latent variables and optimizing the evidence lower bound (ELBO) of the data log likelihood.

[0016] For larger models, optimization is often expressed as maximization of the ELBO, which may be written as: / pe(x,z)\■^'0,yX.x') ^z-q.yfzlx)°% / z|xJ

[0017] Evaluation of this quantity requires knowledge of the posterior distribution q^ / zlx) of the latent variables z given observed data x. If the posterior qv.z\x) is exact, the ELBO Ld iV / (x) corresponds to the true log likelihood of the data. However, for large or highdimensional models, exact computation for this posterior is generally intractable.

[0018] As a result, practical implementations typically employ approximate posterior distributions, often using variational approximations with restricted functional forms. In such cases, the ELBO serves as a lower bound on the true likelihood. Objective functions may also be augmented with additional terms to encourage specific properties of the trained model, such as sparsity, smoothness, or stability, or may be replaced entirely by alternative surrogate objectives.

[0019] Training of PGMs, as well as inference for tasks, such as marginalization or classification, requires calculation of or sampling from the posterior distribution q(z|x) given a data point x. Model parameters 6 may then be optimized to maximize the likelihood of training data using gradient descent or other methods, such as contrastive divergence.

[0020] Sigmoid belief networks are one class of PGMs that have received significant attention in attempts to scale graphical models to greater size and complexity. These models consist of layers of stochastic binary variables connected by weighted interactions. While sampling from such models is relatively straightforward, exact computation of the posterior distribution q(z|x) becomes intractable as model size increases, with computational complexity scaling exponentially with the number of variables.

[0021] Even when variational inference methods are employed, sigmoid belief networks and related models have proven difficult to optimize at scale, severely limiting their depth and resolution. As a result, such models have not been effectively applied to highdimensional tasks, such as high-resolution image generation or large-scale language modeling.

[0022] Variational Autoencoders represent an approach that bridges PGMs and neural networks by parameterizing the generative process as a neural network transformation of samples from a simpler latent probability distribution. A recognition model c / ^(z|x). alsoparameterized by a neural network with weights φ, is trained to approximate the posterior distribution over latent variables.

[0023] VAEs are trained by optimizing the ELBO through back-propagation of gradients through both the generative network and the recognition network. While this approach enables scalable training parameterization of the recognition model by a neural network limits flexibility. In particular, adapting the model to variations, such as partially missing data, alternative conditioning variables, or inference tasks such as classification or inpainting may require retraining or architectural modification.

[0024] Neural networks more generally have been scaled to increasingly large parameter counts and applied successfully across a diverse range of high-dimensional generative tasks. In such feed-forward models, the network parameters implicitly encode both a probabilistic model of the data and the inference procedure used to map inputs to outputs. As a consequence, a significant portion of the representative capacity of large neural networks may be devoted to learning inference behavior itself, in addition to modeling the data distribution. This contributes to increased computational cost, memory requirements, and training data demands as models size and task complexity grow.

[0025] In contrast, PGMs explicitly represent latent variables and their interactions, enabling structured reasoning and flexible conditioning. However, the computational cost of calculating or approximating posterior distributions in conventional PGMs has limited their scalability relative to neural -network-based approaches. Accordingly, existing PGM-based generative techniques have encountered significant challenges when applied to large, structured, or high-dimensional datasets, in part due to the potentially exponential scaling behavior of posterior inference with respect to the number of latent variables.

[0026] The systems and methods described herein provide PGMs that address these limitations by enabling efficient posterior inference in complex, high-dimensional latent spaces. In some embodiments, structured high-dimensional latent variables and interaction functions that permit posterior distributions or variational approximations thereof to be computed efficiently are employed, without reliance on a separate neural-network-based recognition model.

[0027] The disclosed systems and methods generate output data conditioned on partially specified input data using a PGM. Rather than requiring complete input samples, thesystem operates on arbitrary subsets of observed variables and infers missing or unobserved components in a probabilistically consistent manner.

[0028] Observed input data is incorporated directly into the variable space of the PGM, thereby conditioning the model on known values while leaving other variables unspecified. Variational parameters associated with latent variables are then determined through an optimization process that reflects the likelihood of the observed data under the learned model. Based on these parameters, a posterior distribution over latent variables is formed, capturing uncertainty associated with unobserved portions of the data.

[0029] The system generates output samples by sampling latent and output variables from the posterior distribution and converting the sampled representations into applicationspecific output forms. This approach enables flexible generation, prediction, completion, or classification of data across a wide range of domains, including images, text, audio, timeseries data, and multimodal inputs, without requiring retraining when different subsets of input variables are provided. The disclosed techniques support scalable inference in highdimensional latent spaces while maintaining consistency with partially specified inputs.

[0030] In some embodiments, posterior inference is performed using optimization procedures that scale polynomially rather than exponentially with model size, thereby enabling substantially larger graphical models than previously practical. This improvement in posterior inference scalability enables the training and deployment of deeper and more expressive latent-variable models, improving the quality of generated samples and inferred distributions, and expanding applicability to high-dimensional generative tasks.

[0031] Moreover, because posterior parameters may be determined through direct optimization rather than through a fixed feed-forward recognition network, embodiments described herein provide increased flexibility after training, including the ability to condition on arbitrary subsets of variables or to perform marginalization without retraining model parameters. Such flexibility enables efficient reuse of a trained PGM for multiple inference tasks, including conditional generation, classification, inpainting, and anomaly detection, using the learned parameters. Accordingly, the disclosed systems and methods bridge advantages of PGMs and large-scale neural networks by providing scalable generative modeling with efficient inference, flexible conditioning, and applicability to complex, high-dimensional data domains.

[0032] FIG. 1 shows a schematic of an example probabilistic graphical model with interactions between a set of observable variables 101 and latent variables 102. The interactions between variables may be directed (as in belief networks), undirected (as in Markov random fields), factors (as in factor graphs) or a mixture of various interaction types. Conditioning variables may be selected from either variables 101 or 102. Variables 103 are observed, conditioning variables. The subset 103 of variables over which the model is conditioned on is arbitrary, and may vary between particular data points during training and during conditional sampling.

[0033] FIG. 2 shows a flow diagram of an example process 200 for generating a sample from a hyper-dimensional neural network in accordance with some embodiments. The system receives a specified set of weights (which may be the result of a training procedure 300), as well as optionally a set of conditioning values for a subset of variables 103 (step 201). For each conditioning variable supplied, the system sets the value of its corresponding variable from 101 to the supplied value (step 202). It then generates a sample comprising values for each of the remaining variables in 101 and 102 (step 203). The generation process may be any process for generating samples from PGMs, such as Markov chain Monte Carlo, Gibbs sampling, hierarchical sampling, etc., provided that the probability function is consistent with the description in this specification. The system may also generate a distribution over the remaining variables corresponding to the input values.

[0034] For example, the sy stem may be a machine translation system, whereby the variables 101 may be a plurality of sequences of words of a plurality of languages. In such a system, the output may comprise a sequence of words in a target language, conditioned on a sequence of words in an original language (variables 103).

[0035] As another example, the system may be a speech recognition or generation system, where a subset of variables 101 represent audio data and another subset of variables 101 represent graphemes, words, or other characteristics corresponding to the audio data. To be configured as a speech recognition system, it may generate sample graphemes or word sequences conditioned on the audio data, in which case variables 103 represent audio data. To be configured as a speech generation system, it may generate audio data conditioned on the graphemes or word sequences, in which case variables 103 represent graphemes or word sequences.

[0036] As another example, the system may be an image recognition or generation system, where a subset of variables 101 represent the pixels in an image, and another subset of variables 101 represents the contents of the image, for example a written description of the image or the location and classification of various objects within the image. The system may be configured to generate images conditioned on a description of the image (variables 103), or to generate descriptions of the contents of an image, conditioned on the image itself (variables 103). The system may also be configured to receive a partial image (variables 103), and output samples of the full image (variables 101), and may or may not be conditioned on an image description (variables 103).

[0037] As another example, the system may be a time series modeling system that is robust to missing input data, for example it may be used to predict missing or future data or stock prices in a financial model. In such an example, variables 101 may represent financial data or stock prices, and process 200 produces distributions over any missing data in the conditioning input data (variables 103, step 201).

[0038] As another example, the system may be a language modeling system which may be configured to complete sentences, interact with a user of the system via chat, or output larger bodies of text or media by generating output sequences conditioned on a text input sequence. In such a system, variables 101 may represent sentences, sequences of interactions with a user, or pairs of input prompts and output text or media respectively.

[0039] As another example, the system may be a recommendation engine, with the system trained on user data and preferences, and producing as output samples from or distributions over a user’s expected rating for a website or other piece of content. In such a system, variables 101 may represent triplets of user data, content, and ratings, with variables 103 representing the user data and content.

[0040] As another example, the system may be an outlier or anomaly detection system, with the system trained on typical examples from a data distribution, and outputting the ELBO or probability of a given input sample.

[0041] As another example, the system may be an unsupervised embedding system, where variables 101 may represent high dimensional data, and the variables 102 a lowerdimensional representation.

[0042] The variables 101 may represent any encoding of the data, as well as the data itself. For example, they may be formed from a principal component analysis dimensionality reduction of the data, a binary encoding of the data, or the outputs of a neural network encoder.

[0043] In particular, hyperdimensional PGMs allow scaling of latent-variable models to high-dimensional structured latent spaces using a specific form of pe(x,z) so that the posterior distribution q,px(z\x) or a variational approximation to it may be determined more efficiently by optimization of the marginalized likelihood Pe(x) or surrogate objective. This optimization may be performed by gradient descent, Gibbs sampling, or any other method. This optimization does not need to find the global minimum or converge exactly to a local minimum — approximate solutions are sufficient in many embodiments.

[0044] This specification uses the notation q^x) for the posterior to distinguish that the variational parameters ip are per-example, and are not parameterizing a (global) neural network, as in the case variational auto encoders qpXx^z).

[0045] Thus, hyperdimensional PGMs may compute variational parameters directly by optimization of:ix= arg max Ez_ lnn (PdCx'ZA (1)

[0046] This optimization can be made computationally feasible by constructing the objective function such that it contains a plurality of invariant components. An invariant component can be broadly defined to be an intermediate value obtained during the computation of the objective that is invariant to a group of transformations of its inputs. In many embodiments, this invariance is clearly apparent from the form of the conditional probability distributions employed by the model as well as the factorization of the variational posterior approximation, but in others it is more nuanced.

[0047] The calculation of the objective function can be broken down into a number of intermediate component functions fc(ip). which are computed and combined, for example, in some embodiments by summation of components of the ELBOm F r i V£(W = 2 feW) W) = Ez~^(z)c=QrVi / AZi).

[0048] A component / c(’ / is considered to be invariant if there exists a continuous group of mappings T E G that operates onwithout changing the resulting value, i.e. fc^ = fc(T^m, T ^ G

[0049] In some embodiments, a plurality of components, particularly those between only latent variables, are invariant the same group of mappings G.

[0050] In some embodiments, Hyperdimensional PGMs construct p9(x, z) and ^,X(Z| ) such that the interactions between latent variables (102) are (at least approximately) invariant to a transformation of each latent variable by any transformation T within a continuous group of such transformations Cj. When considering graphical model structures with the data as leaf nodes with directed edges leading to them, p9(x, z) may be factorized as the product of a conditional probability function ge(x\z). individual latent variable bias terms p (z(). and an interaction function f9(z), i.e. that p9(x, z) = g9(x\z)f9(z) f[i Pt (.zi), with / 0(z) « / 0(T(z))vz, T G g

[0051] Note that in some embodiments, the distribution over latent variables may be conditional on x, and so will have different overall factorization, but similar transformationinvariance the latent-variable interaction term f9z). In some embodiments, the bias terms P (Zi) may be uniformly distributed and thus not necessarily explicitly included in the formulation.

[0052] In many embodiments, these transformations are elements of the special orthogonal group SO(N), and the probability function p is constructed relatively simply by being functions of dot-products between variables, such as p(z z) = f(Zikzjk) (using Einstein summation notation).

[0053] This is of course not the only way to construct a component with the necessary invariance. Consider the simple model with 2 latent variables, parameterized by an arbitrary matrix Ap(z z2) oc exp(— z^Az2)

[0054] This is generally not invariant to transformation of both z^ and z2by the same element of SO(N), but p is invariant to a transformation that maps z' =BA^. andz2= BZ2, for any B that is an element of SO(N), leaving N-l degrees of freedom in specifying z^ and z2.

[0055] The described transformation invariance ensures that degrees of freedom exist in the optimization of (1), allowing separate optimization of relationships between interacting variables and between entire neighborhoods of variables. These additional degrees of freedom reduce the propensity for the optimization over the posterior parameters ip to get stuck in a local minimum, improving the quality of the posterior estimate from (1) over prior art and allowing use of much larger graphical models with many more parameters. The combination of improved posterior estimates and larger model size ultimately results in improvements to the quality of generated samples and posterior distribution estimates.

[0056] In some embodiments, the ELBO or surrogate optimization objective, rather than the probability function, obeys this invariance. In some embodiments, each term in the factorization of fez) obeys such an invariance.

[0057] Despite the clear improvements to the performance of gradient descent optimization of the posterior in Hyperdimensional PGMs. they are not restricted to use of gradient descent methods; sampling and other methods also benefit from improved performance via the reduction in number and severity of local minima. For the case of sampling and Monte Carlo methods this results in achieving equilibrium or ergodicity with fewer sampling steps.

[0058] Hyperdimensional PGMs may be homogenous, in that the entire network obeys such an invariance, or inhomogeneous, with different factors in the probability distribution obeying different forms of transformation invariance, or even differing dimensionality of variables.

[0059] As with other neural networks, hyperdimensional PGMs also require that the probability functions within hyperdimensional neural networks have nonlinear components. Note that this nonlinearity may be achieved through nonlinearities in the interactions between variables, or as a consequence of restrictions in variable domain - for example the nonlinearity of a sigmoid function arises from linear energy interactions but a restricted domain of {0,1} or {-1,1} depending on the parameterization. This distinguishes them from for example, factorized gaussian linear models, and allows them to capture more complex data distributions.

[0060] The requirement that T is an element of a non-degenerate continuous space of transformations excludes sigmoid belief networks from this definition; though the transformation f(z) = — z preserves probabilities, the space of these transformations is nil-dimensional and discontinuous with the identity transformation.

[0061] In some embodiments, the posterior q^,x(z|x) may be calculated as the product of factors of each of these high-dimensional variables, i.e. q^(z|x) ocipx)

[0062] The posterior may be parameterized in such a way that it does not exactly match the true posterior distribution, i.e. that it is not conjugate to pg(x, z) - in such a case the parameters of the recognition model may be optimized variationally to minimize a divergence to or from the true posterior distribution.

[0063] In some embodiments, pg(x, ) is constructed by lifting each of the latent variables ztinto a higher-dimensional space, such that ztG IRW, and constructing the probability function p (x, z) by interactions between them. This PGM may of course equivalently be described as one with additional constraints or interactions between them to enforce the smoothness required to optimize eqn. 2 (e g. by enforcing that particular sets of single-dimensional variables have a fixed 12-norm); as it is one of the simplest methods to smooth the functions, for ease of discussion in the following we will consider each z Eto be a single hyper-dimensional latent variable, rather than a collection of lower-dimensional variables with such a constraint between them.

[0064] There are clear tradeoffs in selection of the dimensionality of the latent variables. Larger values for N typically improve the smoothness of the optimization, but come at increased computational cost, hence slower training and inference. In many embodiments, selecting the dimensionality such that N « width, where width is the maximum number of interactions involving a single latent variable, provides a good balance between the tradeoffs.

[0065] The function pg(x, z) may be defined such that it obeys rotational symmetry about the origin, I.e. that interactions between variables are invariant under transformations T in SO(N'. In some embodiments, pg(x, z) may be constructed by functions where the individual variables interact only through their relative directions and magnitudes.

[0066] In some embodiments, the latent variables are constrained to have unit length.II Zj ||= 1, though any such restriction on magnitude is equivalent up to rescaling of the parameters 6. Additionally or alternately, the probability distribution of each variable may be parameterized by a set of weights WLJand (optionally) biases bij such thatp9(z0 oc exp l (Zt, bi) + X^ij< Zt, Zj} ) (3)\ J / where (%;, Xj} denotes an inner (dot) product between vectors xtand Xj, and we use 6 to refer to all trainable parameters of the model, i.e.b. Within such models, the latent variables take von Mises-Fisher (VMF) distributions, with parameters determined by their parents. We will in the following refer to this subset of embodiments - models with unit constrained, and linear (in energy) pairwise interactions as hyperspherical neural networks, and provide more detailed description of their construction and methods that allows their training without sampling from the posterior distribution. In many instances, calculation of optimization gradients without sampling is preferable as it reduces the variance of gradient estimates, though in cases where approximation is necessary there is a tradeoff with optimization accuracy.

[0067] Note that eqn. 3 uses z;and refers specifically to the latent vectors, but the output nodes may have a similar formulation, though typically in a lower dimensional space N=1 or 2.

[0068] In some embodiments, the latent variables take on a conditional probability distribution according to a scaled gaussian, where

[0069] p(zf|z:i) = N(p / a, I / a),p = Wz.L,a = N / 2 + / ( / V / 2)2+ || p ||2

[0070] This has similar moments and symmetry to the von-Mises Fisher distribution above, but has the benefit of being non-zero over ns', and so is compatible with multivariate gaussian variational posterior distributions. In this embodiment, the ELBO of the latent parameters is again invariant to rotations of the space.

[0071] In some embodiments, the interactions are restricted such that the system forms a directed acyclic graph (e.g. that WLJ= 0: t > j in eq. 3). This restriction allows the direct sampling from the probability distribution, without requiring computationally more intensive Monte Carlo sampling methods. With such a restriction, we define the parents of anode i to be the nodes influencing its distribution, so for the case of Wtj = 0: 1 > j. we can define the parents of node i as pa(i) = {0,1,..., i — 1}.

[0072] In some embodiments, interactions are specified such that variables are arranged into sequential layers by disallowing interactions between variables in the same layer.

[0073] In some embodiments, parameters from 0 are shared between multiple factors, e.g. to enforce translation equi variance in parts of the model similar to that imposed in convolutional neural networks.

[0074] In some embodiments, particularly those involving norm-constrained variables, the posterior may be specified such that it factorizes about each variable according to a von Mises-Fisher distribution, i.e. as ( (z x) oc exp(cipizi')

[0075] In some embodiments, the variational posterior is conditioned on the value of parent variables, i.e. ^(zjz.px). In many preferred embodiments, the form of the conditional variational posterior is similar to that of the probability function p. For example, in the case of a scaled gaussian probability function, the variational posterior may be

[0076] ^(zjzj = = + A / / 2 + J(A / / 2)2+ || 111

[0077] In some embodiments, the variational parameters also modify the gaussian covariance matrix. In some embodiments, the form of the covariance may be structured so as to reduce computational, communication, and / or storage costs.

[0078] The model may be trained to maximize the log-likelihood of the data, in some embodiments through mini batch or batch gradient descent. FIG. 3 shows a flow diagram schematic for an example training procedure for determining the weights of a hyperdimensional graphical model in accordance with some embodiments. It shows the steps for training based on a single data point at a time, but it may also be trained batch-wise or through other methods. The system receives a training data point (step 301), which specifies the value over all or a subset of variables 101. It then finds variational parameters specifying q(z|x) to optimize the ELBO or surrogate objective (step 302). The system then updates the global parameters 0 to maximize the ELBO or surrogate objective (step 303).

[0079] While this process of ELBO optimization appears similar to that used in Variational Autoencoders, the described specification differs significantly in that instead of using a second neural network to determine q,p(z\x) the procedure described here uses an optimization routine to find q(z\x) for each data point x.

[0080] The update in step 303 may be carried out via a variety of methods, including using gradient descent, stochastic gradient descent, or one of the many well-known neural-network optimization routines (e.g. Adam, RMSprop, Lion, etc.), or through methods such as variational expectation maximization, or techniques employing noisy estimates of the gradients with respect to 9 and < >.

[0081] In other embodiments, the model may be trained or pre-trained lay er- wise, with techniques such as contrastive divergence.

[0082] Some embodiments include an augmenting neural-network based recognition model qcp z\x) to seed starting values for the optimization routine or as estimates of q^x(z|x). This has benefits over traditional structured variational auto encoders in that the training signal for learning q<p(z\x) is more informative, as it is directly finding the optimum qx / ,xz\x') rather than indirectly by back-propagation through the ELBO or other objective.

[0083] Given a generative model architecture and method for computing c / (z|x). it is possible to extend the capability of the system to other machine learning tasks relying on conditioning or marginalization of the graphical model, including classification or inpainting. For example, by considering one or a set of the latent variables as denoting class membership, the probability distribution over this set, conditional on the data x can be used to classify the data; similarly, the distribution over latent variables conditional on a partially observed data point implies a probability distribution over the missing data. By providing a direct estimate of p(x), the system is also trivially able to detect anomalies as data points with low likelihood under the model.

[0084] Because hyperspherical and hyperdimensional networks can determine the posterior parametersxvia an optimization process rather than a feedforward network, they are more easily able to handle missing data. Reconfiguring the system to perform posterior predictions conditioned on a different subset of data components, or marginalization of predictions over a subset of components, may be performed with the same learned modelparameters 9. This is in contrast to variational autoencoders, where the recognition model is trained specifically for such tasks. Hence, in some embodiments, training data points may be missing some components of x, and in other embodiments, the trained model parameters 0 may be used to determine probability distributions over missing data components, such as for inpainting or classification tasks.

[0085] FIG. 4 is a block diagram illustrating a system to generate output samples conditioned on partial inputs using a probabilistic graphical model in accordance with some embodiments. In the example shown, system 400 includes a client device 402 configured to transmit input data and a graphical model system 412 configured to generate one or more output samples conditioned on the received input data.

[0086] In some embodiments, system 400 is configured for machine translation, in which the input data provided by client device 402 comprises a sequence of tokens in a source language, and graphical model system 412 generates a corresponding sequence of tokens in a target language conditioned on the input sequence.

[0087] In some embodiments, graphical model system 412 is configured for speech recognition or speech generation. In such embodiments, the input provided by client device 402 includes audio features representing spoken content, textual tokens representing linguistic content, or a combination thereof. When configured for speech recognition, graphical model system 412 generates a sequence of textual tokens conditioned on input audio features. When configured for speech generation, graphical model system 412 generates audio features or synthesized audio conditioned on an input sequence of textual tokens.

[0088] In some embodiments, graphical model system 412 is configured for image recognition or image generation. In such embodiments, input received from client device 402 includes pixel data representing an image, textual or symbolic descriptions of the contents of an image, labels indicating object locations or classifications, or partial versions of any such data. When configured for image generation, graphical model system 412 generates a complete image conditioned on a textual description, a detection map, or other semantic information. When configured for image recognition or captioning, graphical model system 412 generates textual or semantic information describing the contents of an input image. In some embodiments, system 400 receives a partially specified image and generates one ormore samples completing the missing regions of the image.

[0089] In some embodiments, graphical model system 412 is configured for timeseries modeling, including applications involving missing, noisy, or irregularly sampled input data. For example, system 400 may estimate or predict financial indicators, sensor readings, medical measurements, environmental data, or other temporal sequences. In such embodiments, system 400 generates predicted or imputed values for absent or future timeseries elements by producing samples or distributions conditioned on the observed portions of the input sequence, using a sample generation process such as process 200 described herein.

[0090] In some embodiments, graphical model system 412 is configured for language modeling or conversational interaction, including generation of text continuations, dialogue responses, or extended sequences of texts or media conditioned on one or more input tokens or prompts provided by client device 402. In some embodiments, system 400 generates predicted next tokens, complete message responses, or multi-modal outputs without requiring retraining of the underlying probabilistic graphical model when different subsets of input variables are specified.

[0091] In some embodiments, graphical model system 412 is configured as a recommendation engine, trained using historical user-interaction information, preference data, or other contextual signals. In such embodiments, system 400 generates affinity scores, rankings, or probability' distributions representing expected user responses to candidate items, conditioned on available input information.

[0092] In some embodiments, graphical model system 412 is configured for outlier or anomaly detection. For example, system 400 may be trained using representative samples of normal or typical data, and may generate an anomaly score, likelihood estimate, or other measure indicating whether an input sample is consistent with the learned data distribution. In some embodiments, such output may be derived from a conditional or joint probability with the input data, enabling detection of rare or aty pical patterns without requiring retraining when different subsets of input features are observed.

[0093] In some embodiments, system 400 is configured as an unsupervised embedding system, in which high-dimensional input variables are mapped to loyver-dimensional latent representations, and output samples are generated conditioned on partial observations in either space.

[0094] Graphical model system 412 includes optimization subsystem 414. The optimization subsystem determines values of variational parameters ip by function optimization. In some embodiments, the optimization is of the marginalized probability over the unobserved variables in set D, i.e. the indices of variables not specified in x. i.e. maximum likelihood variational inference.ip = maxf^PCx^dXi

[0095] In particular, hyperdimensional PGMs allow scaling of latent-variable models to high-dimensional structured latent spaces using a specific form of pe(x,z) so that the posterior distribution q,px(z\x') or a variational approximation to it may be determined more efficiently by optimization of the marginalized likelihood Pg( ) or surrogate objective. This optimization may be performed by gradient descent, Gibbs sampling, or any other method. This optimization does not need to find the global minimum or converge exactly to a local minimum — approximate solutions are sufficient in many embodiments.

[0096] Graphical model system 412 includes sampling subsystem 416 configured to determine one or more output samples using variational parameters determined by optimization subsystem 414.

[0097] In some embodiments, graphical model system 412 forms a directed acyclic graph, a neural network system of sampling subsystem 416 iteratively constructs output data sample by determining a set of variational parameters from the optimization subsystem 414, and sampling a value for each position xt. In some embodiments, the optimization subsystem 414 is used to determine a new set of variational (sampling) parameters at intervals between such sampling steps.

[0098] In some embodiments, sampling subsystem 416 implements a Monte Carlo sampling algorithm to produce a sample from the learned (directed or undirected) graphical model, and then convert it from the embedding space back to the application space.

[0099] FIG. 5 is a flow diagram illustrating a process to generate output samples conditioned on partially specified input data using a probabilistic graphical model in accordance with some embodiments.

[0100] At 502, partial input data is received. In some embodiments, the partial input data corresponds to an image in which a subset of pixel values is specified. In some embodiments, the partial input data corresponds to time-series values such as financial indicators or sensor measurements. In some embodiments, the partial input data corresponds to a sequence of word or sentence tokens. In some embodiments, the partial input data corresponds to conversational input associated with a dialogue system.

[0101] The partial input data may correspond to a subset of observable variables corresponding to known components of a data sample. For example, the input may include a portion of an image, a sequence of text tokens, audio features, or other observed data elements while other components remain unspecified. The set of variables provided as conditioning input is arbitrary and can vary between data points during both training and inference. In some embodiments, these observed variables correspond to the set of conditioning variables 103, which may be selected from either observable variables 101 or latent variables 102, depending on the application.

[0102] Optionally, output generation may be further conditioned on contextual input, prior outputs, or intermediate predictions.

[0103] At 504, the input is embedded into the model variables such that the observed values become part of the PGM’s variable space. For example, a hyperdimensional generative model may be configured to randomly sample images from a learned distribution, to infer unobserved pixel values for image inpainting, to generate media conditioned on a prompt, or to determine a classification by producing a probability distribution over candidate labels.

[0104] The embedding may employ any suitable representation, including pixel intensities, text tokens, audio features, or intermediate representations such as principalcomponent embeddings or neural-network-derived embeddings. By incorporating observed information directly into the variable structure of the probabilistic graphical model, the system conditions the distribution of remaining variables, including latent variables, on the supplied values. In some embodiments, the embedding step designates the embedded values as conditioning variables, and the particular subset used for conditioning may differ for each data instance.

[0105] At 506, variational parameters associated with one or more latent variables aredetermined by optimizing a probabilistic objective reflecting the likelihood of the observed data under the learned model. In some embodiments, the variational parameters are obtained by maximizing a marginalized likelihood or evidence lower bound (ELBO) over unobserved components of the data. Unlike approaches that rely on a separate recognition network to approximate posterior distributions, the disclosed embodiments determine variational parameters directly through optimization over latent-variable distributions within the graphical model itself. In some embodiments, the optimization exploits structural or symmetry properties of the latent interaction space, enabling scalable approximation of posterior distributions. The variational distribution may, for example, approximate a von Mises-Fisher distribution or another factorized distribution suitable for efficient inference and sampling.

[0106] In various embodiments, the quality of estimates of posterior distribution over latent variables is improved, thus ultimately improving the quality of generated samples as well as predictions of classification or missing data.

[0107] At 508, posterior distribution is determined based on the variational parameters obtained from the optimization process. The posterior distribution represents a conditional probability distribution of the latent variables given the observed input data, the selected conditioning variables, and the structure of the probabilistic graphical model. This posterior distribution characterizes uncertainty over unobserved variables and provides a probabilistic basis for generating samples that are consistent with the partially specified input.

[0108] At 510, the latent variables are sampled according to the posterior distribution. In some embodiments, sampling is performed sequentially or iteratively in an order consistent with dependencies defined by the probabilistic graphical model, such as a directed acyclic graph structure. The sampled latent variable values are selected such that they are statistically consistent with the observed input data and reflect uncertainty associated with unobserved portions of the data.

[0109] At 512, the sampled values are converted to output form corresponding to an application domain. This conversion may include decoding latent representations into pixel values, text tokens, audio signals, class labels, or other domain-specific outputs. The resulting output sample is consistent with the observed input data and reflects the probabilistic structure learned by the graphical model.

[0110] FIG. 6 illustrates an example data flow within a graphical model system configured to generate output data samples conditioned on partially specified input data, in accordance with some embodiments.

[0111] As shown, input data is received in a partially specified form, such that values are provided for only a subset of variables associated with a data sample. The input data is embedded into a variable space of the graphical model to produce embedded input data, which represents the observed components of the data sample within the probabilistic graphical model.

[0112] The embedded input data is provided to a sampling subsystem of the graphical model system. The sampling subsystem includes a probabilistic graphical model comprising a plurality of observable variables and latent variables arranged according to a dependency structure, such as a directed acyclic graph. In the illustrated example, a subset of variables (e.g.. conditioning variables 103) corresponds to observed or fixed values derived from the embedded input data, while remaining variables (e g., latent variables 102 and observable variables 101) are unobserved and subject to inference and sampling.

[0113] Using variational parameters determined by an optimization subsystem (not shown), the sampling subsystem samples values for latent variables and, optionally, additional observable variables in an order consistent with the dependency structure of the graphical model. Sampling is performed such that the generated values are statistically consistent with the observed embedded input data and the learned joint distribution encoded by the graphical model.

[0114] The sampled variable values are combined with the embedded input data to form embedded output data, representing a completed or inferred version of the data sample within the embedding space of the model. The embedded output data is subsequently converted into output (sampled) data in an application-specific form, such as an image, sequence of tokens, audio signal, time-series values, or other structured output. The resulting output sample is consistent with the partially specified input data and reflects one realization drawn from the conditional distribution defined by the graphical model.HYPERSPHERICAL NETWORK COMPUTATIONS

[0115] Many of the quantities described above required for training and samplingfrom hyperspherical networks do not have obvious, computationally tractable procedures for their calculations, so we describe some here:Sampling

[0116] For the particular case of exponential interactions with unit-normal, independent, latent variables, the form of the probability distribution of each variable is equivalent to a von Mises-Fisher distribution with mean direction ptand concentration parameter κi≥ 0.

[0117] p(zi) = Cd(κi)exp(κiμiTzi) (2)

[0118] Where

[0120] and where Ivis the modified Bessel function of the first kind of order v.

[0121] The von Mises-Fisher parameters KLpLmay be obtained from each node’s parents:

[0122] K il = bt + X ZjWijjepa( )

[0123] For directed acyclic graphical models, samples may be obtained by a hierarchical sampling procedure, starting at the root node and sampling each variable. For models containing loops, sampling is repeated until the system reaches equilibrium. This may be, for example, by Gibbs sampling.

[0124] Sampling from eqn. 2 may be done, for example, according to the procedure outlined in “Fast Python sampler for the von Mises Fisher Distribution’’ by Pinzon et al. Approximation of the ELBO

[0125] The ELBO may be broken down into a Kullback-Leibler divergence term of the variational distribution q from the one imposed by its parents, and the data log likelihood conditional on the variational distribution. These terms may be estimated separately.(pg(x, z)\ PeW JEz~qv,x(Z|x) log1®z~<7i / ix(z|x) [l°g(. P0, WZIX) = ~DKL(iq^xII pe) + ^z~q^x(z\x)[log(. Pe O|z))]KL Divergence

[0126] DKL(q,px|| p0) = fZq4,x(z)logq^x(z)dz - fzqlPx(z)logp0(z)dz

[0127] The first term is relatively simple to compute in the mean-field setting, as it is simply the von Mises-Fisher entropy

[0128] Jz<7^(z)Zo^^x(z)dz = {logfd(x; p, K))x= logCd(jc) + KAd(K)Id / zW AdW^d / 2-l(.K)

[0129] The second term is somewhat more involved, with calculation of pg(z requiring marginalization of the node's parents:1r( \Pe(.zi) = ^ Jz. Yl[q^j)]exp 2 HZi7(Z(,z7) dz,>i1J \i*j /

[0130] In some embodiments, the variational distribution is a von Mises-Fisher distribution, soP(Zi) = jJ- YllCd^exp^Kjgj'Zj + VKozJzi)]dzo...dz,-

[0131] Grouping terms containing z0into a separate integral yields1J II [Q (K^expfKj^Zj + Wijzjz^dz^... dZj [J Q(K0)exp(K0 / S0z0+ VFi0zT0Zi)dz0]

[0132] And by similarity to the von Mises-Fisher distribution noting that the last term integrates to a concentration parameter1 f n [Q (jcj)exp[Kjp}zj+ Wijzjz^dz^.. dzj [Q(K0) / Q(II KOPO + Wiozt||)]JE1.. J= ^n[cd(Kj) / cd(ii Kjfij + Wtjzt ID]

[0133] We can further approximate the log of the normalization constant by a quadratic, i.e. 1 / Cd(|| K7q7+ W / i7Zj ||) = expty-z^ + zjA^zJ, obtaining1~ ^n[Q(*7 W(yJ Zi + z'iAjZi^dZi

[0134] So finally (with the CdKj') terms dropped as they do not depend on z)logpe^Zt) ~ yTZj + z}Azt- logZtWhere y = £y7and A = AjJ J

[0135] Noting the axial symmetry about each p^ Yj and Aj may be obtained from a 3-point quadratic fit to Cd(|| KjPj + W^Zi ||) evaluated at (for example) zt— ppzt1 ppzt— -Pj.

[0136] The normalization constant Ztmay be estimated numerically with Monte Carlo simulation or by integration of the quadratic approximation.Zi = S^expCy^Zi + z]AZi)dZi

[0137] The integral is equivalent to the normalization constant of a Fisher-Bingham distribution, and can be calculated by a number of means, including numerical approximation or approximation by the saddle point method such as the ones described in " Saddlepoint Approximations for the Bingham and Fisher-Bingham Normalising Constants” by Kume et al. / d1 / 1 d y2Ci(2,y) = 21 / / 27z4d-1) / 2[Kg N j f] Gi—t)-1 / 2[exp [ — t + — £ —l- —v / u=i ) y4 i=iAi- tc2(4,y) = c1(A,y)(l + T), c3(A,y) = c1(,y)exp(T)

[0138] Where A are the eigenvalues of A and y is replaced by QTy (transforming the previously obtained y into the space of A’s eigenvectors) and the j th derivatives of the cumulant generating function for independent non central Xi is given byd (0 - 1)! 1 y! Yi ]^(t) z1=1 [ 2 (A;- ty+4 (2i - t)Hij1 59T =BP< - 24Pj= ^en / (K^y / 2

[0139] Where t is the solution in (— co, A^) to the equation'. This may be determined by any number of numerical methods.

[0140] The final integration over z;may be estimated using, for example, a sampling method.

[0141] Alternately, the second term in the KL divergence may be estimated for VMF-distributed posteriors byJzq z) / oc?p0(z)dz(I 2J HM + ss = Sy(l - Ad(Kj))W? and gp= Hj{Ad(<Kj^jWji')

[0142] This can be derived by estimating the expected value of the concentration parameter by calculating the concentration parameter for the expected value of K, and tends to have reasonable performance while being much faster than the previous method.Data log-likelihood

[0143] Typically, the training data for a neural network will not consist of a set of unit-length variables in IKW, but will usually be of lower dimension (often either binary as in the case of black-and-white images or single dimensional, as in the case of grayscale images). In this case, we reduce the dimensionality of the output variables.

[0144] As before, the data log-likelihood term Ez^c / rzjx)[logp(x^z')] may be approximated by Monte Carlo estimation, but it is computationally cheaper to obtain gradients for the internal optimization routine with a closed-form solution. One such approach is to use a gaussian approximation for the parents.JEz~<j(Z|x)[^(p(Xi|z))] = fzq(z)a [ xtXWijZj \ dz\ J /

[0145] This integral is over all parents z, and is generally computationally intractable. Approximating each parent by a Gaussian distribution with mean and standard deviation calculated from its underlying von Mises-Fisher distribution allows replacing the integral over all parents with one over a single normal distribution (computed via a sum of the parent normal distributions) which may, for the case of binarized outputs, be further reduced to a single dimension and estimated by a probit function.= SKNN ( S / Tj, Sv7- j er I x^Wu I dz= / 2^Sum°(z)dz « < P ( / z / ^8 / 7T + V2)Where Nsumis a gaussian distribution obtained from the weighted sum of the parent distributions z approximated by a gaussian distribution and projected onto xt, g and v the mean and covariance of that projection, and < P the probit function.

[0146] This probit function estimate ignores correlations between output elements. While it often produces adequate results for systems with independent variational posterior distributions, better results can be obtained using importance sampling estimates, or other algorithms for calculating multivariate normal orthant probabilities.

[0147] Although the foregoing embodiments have been described in some detail for purposes of clarity of understanding, the invention is not limited to the details provided. There are many alternative ways of implementing the invention. The disclosed embodiments are illustrative and not restrictive.

Claims

1. CLAIMS1. A method, comprising:embedding partial input data corresponding to a subset of variables of a data sample into a variable space of a probabilistic graphical model containing a plurality of transformation invariant components;determining variational parameters associated with one or more latent variables of the probabilistic graphical model by optimizing a probabilistic objective reflecting likelihood of the partial input data by adjusting variational parameters governing the probability distribution over latent variables in the probabilistic graphical model;sampling latent variable values from a posterior distribution defined by the variational parameters; andgenerating output data by converting the sampled latent variables into output data corresponding to an application domain.

2. The method of claim 1, further comprising receiving the partial input data.

3. The method of claim 1, wherein remaining variables of the data sample are unspecified and the subset of variables is designated as conditioning variables.

4. The method of claim 1, wherein the sampling is performed in accordance with dependencies of the probabilistic graphical model.

5. The method of claim 1, wherein the output data is statistically consistent with the partial input data without retraining model parameters of the probabilistic graphical model.

6. The method of claim 1, wherein the partial input data corresponds to an image in which a subset of pixel values is specified, time-series data, a sequence of word or sentence tokens or conversational input associated with a dialogue system.

7. The method of claim 1, wherein the variational parameters are obtained by maximizing a marginalized likelihood or evidence lower bound over unobserved components of the partial input data.

8. The method of claim 1, further comprising determining the posterior distribution.

9. The method of claim 8, wherein the posterior distribution represents a conditional probability distribution of the one or more latent variables given observed portions of the partial input data, selected conditioning variables, and a structure of the probabilisticgraphical model.

10. The method of claim 1, wherein the sampling is performed sequentially or iteratively in an order consistent with dependencies defined by the probabilistic graphical model.

11. The method of claim 1, wherein the sampled latent variable values are selected to be statistically consistent with the observed portions of the partial input data and reflect uncertainty associated with unobserved portions of the partial input data12. The method of claim 1, wherein generating the output data comprises generating one or more of: image data, text data, audio data, time-series data, classification outputs, or recommendation outputs.

13. The method of claim 1, wherein sampling the latent variable values comprises performing Monte Carlo sampling from the posterior distribution.

14. The method of claim 1, wherein embedding the partial input data comprises mapping observed input values to corresponding nodes of the probabilistic graphical model while leaving other nodes unassigned.

15. A system, comprising:a processor configured to:embed partial input data corresponding to a subset of variables of a data sample into a variable space of a probabilistic graphical model containing a plurality of transformation invariant components;determine variational parameters associated with one or more latent variables of the probabilistic graphical model by optimizing a probabilistic objective reflecting likelihood of the partial input data by adjusting variational parameters governing the probability distribution over latent variables in the probabilistic graphical model; sample latent variable values from a posterior distribution defined by the variational parameters; andgenerate output data by converting the sampled latent variables into output data corresponding to an application domain; anda memory coupled to the processor and configured to provide the processor with instructions.

16. The system of claim 15, wherein remaining variables of the data sample are unspecified and the subset of variables is designated as conditioning variables.

17. The system of claim 15, wherein the partial input data corresponds to an image in which a subset of pixel values is specified, time-series data, a sequence of word or sentence tokens or conversational input associated with a dialogue system.

18. The system of claim 15, wherein the processor is configured to perform the sampling sequentially or iteratively in an order consistent with dependencies defined by the probabilistic graphical model.

19. The system of claim 15, wherein to generate the output data, the processor is further configured to generate one or more of: image data, text data, audio data, time-series data, classification outputs, or recommendation outputs.

20. A computer program product embodied in a non-transitory computer readable medium and comprising computer instructions for:embedding partial input data corresponding to a subset of variables of a data sample into a variable space of a probabilistic graphical model containing a plurality of transformation invariant components;determining variational parameters associated with one or more latent variables of the probabilistic graphical model by optimizing a probabilistic objective reflecting likelihood of the partial input data by adjusting variational parameters governing the probability distribution over latent variables in the probabilistic graphical model;sampling latent variable values from a posterior distribution defined by the variational parameters; andgenerating output data by converting the sampled latent variables into output data corresponding to an application domain.