Environment-agnostic causal representation machine learning models in interactive systems

By training a causal representation model to generate embedding representations in a common space and segmenting observations, the model becomes environment-agnostic, addressing the lack of generalizability in existing models and improving adaptability and efficiency across varied environments.

WO2026059894A1PCT designated stage Publication Date: 2026-03-19QUALCOMM TECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Causal representation machine learning models are typically trained for a specific environment and lack generalizability across various environments, failing to accurately generate observations due to dynamic changes in object numbers, types, and positions.

Method used

Training a causal representation model to generate embedding representations of images into a common embedding space using an encoder neural network, allowing prediction of future environmental states through a predictive neural network, and segmenting observations into segments associated with unique causal variables.

Benefits of technology

Enables accurate, actionable observations across diverse environments, reducing computational resources and maintaining a single model for multiple environments, thus enhancing adaptability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025045519_19032026_PF_FP_ABST
    Figure US2025045519_19032026_PF_FP_ABST
Patent Text Reader

Abstract

Certain aspects of the present disclosure provide techniques and apparatuses for observing an environment and predicting a future state of the environment using environment-agnostic causal representation machine learning models. An example method generally includes receiving, at a computing system, an image of an environment and an indication of an action to be performed in the environment. Using an encoder neural network trained to generate embedding representations of images from a plurality of environments into a common embedding space, an embedding representation of the image is generated. Using a predictive neural network, causal variables for a future state of the environment are predicted based on the action and the embedding representation of the image. Based on the predicted causal variables, the future state of the environment after execution of the action is predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Client Ref. No.: 2407433WO 1 ENVIRONMENT-AGNOSTIC CAUSAL REPRESENTATION MACHINE LEARNING MODELS IN INTERACTIVE SYSTEMS CROSS-REFERENCE TO RELATED APPLICATION(S)

[0001] This application claims priority to Greece Application No.20240100622, filed September 10, 2024, which is hereby incorporated by reference herein. INTRODUCTION

[0002] Aspects of the present disclosure relate to neural networks.

[0003] Machine learning models may be used to observe an environment and relationships between different objects in the environment. For example, causal representation learning attempts to identify causal variables and the relationships between these causal variables from an unstructured observation, such as an image or video stream including content captured from an environment in which an observation is to be made and actions are to be performed based on the observation. These actions may include controlling a robotic system to perform a task within an environment, performing various self-driving or other autonomous vehicle-related tasks, or the like.

[0004] Many causal representation machine learning models are trained to learn a representation of a specific environment in which observations are to be made and in which actions are to be performed based on such observations. Because these causal representation machine learning models are trained to make observations within the context of a specific environment, these machine learning models may not be generalizable across a variety of environments. For example, because environments in which observations are made may be dynamic, with varying numbers of objects, varying types of objects, varying positions of certain objects in time, and the like, these causal representation models may not accurately generate observations in these environments because these models may be trained on an environment different from that from which images for observation are obtained. BRIEF SUMMARY

[0005] Certain aspects of the present disclosure provide a processor-implemented method for observing environments using machine learning models. The method generally includes receiving, at a computing system, an image of an environment and an indication of an action to be performed in the environment. Using an encoder neural P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 2 network trained to generate embedding representations of images from a plurality of environments into a common embedding space, an embedding representation of the image is generated. Using a predictive neural network, causal variables for a future state of the environment are predicted based on the action and the embedding representation of the image. Based on the predicted causal variables, the future state of the environment after execution of the action is predicted.

[0006] Certain aspects of the present disclosure provide a processing system. The processing system generally includes one or more memories comprising processor- executable instructions and one or more processors coupled to the one or more memories. The one or more processors are configured to execute the processor-executable instructions and cause the processing system to: receive an image of an environment and an indication of an action to be performed in the environment; generate, using an encoder neural network trained to generate embedding representations of images from a plurality of environments into a common embedding space, an embedding representation of the image; predict, using a predictive neural network, causal variables for a future state of the environment based on the action and the embedding representation of the image; and predict, based on the predicted causal variables, the future state of the environment after execution of the action.

[0007] Certain aspects of the present disclosure provide a processing system. The processing system generally includes: means for receiving, at a computing system, an image of an environment and an indication of an action to be performed in the environment; means for generating, using an encoder neural network trained to generate embedding representations of images from a plurality of environments into a common embedding space, an embedding representation of the image; means for predicting, using a predictive neural network, causal variables for a future state of the environment based on the action and the embedding representation of the image; and means for predicting, based on the predicted causal variables, the future state of the environment after execution of the action.

[0008] Other aspects provide processing systems configured to perform the aforementioned methods as well as those described herein; non-transitory, computer- readable media comprising instructions that, when executed by one or more processors of a processing system, cause the processing system to perform the aforementioned methods as well as those described herein; a computer program product embodied on a P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 3 computer-readable storage medium comprising code for performing the aforementioned methods as well as those further described herein; and a processing system comprising means for performing the aforementioned methods as well as those further described herein.

[0009] The following description and the related drawings set forth in detail certain illustrative features of one or more aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The appended figures depict example features of certain aspects of the present disclosure and are therefore not to be considered limiting of the scope of this disclosure.

[0011] FIG. 1 illustrates an environment-agnostic causal representation machine learning model trained to generate an observation of an environment, according to certain aspects of the present disclosure.

[0012] FIG. 2 illustrates an environment-agnostic causal representation machine learning model trained to generate an observation of an environment based on segmentation of an input image into a plurality of segments associated with different objects in the environment, according to certain aspects of the present disclosure.

[0013] FIG.3 illustrates an example of causal variables predicted for an object in an environment using an environment-agnostic causal representation machine learning model, according to certain aspects of the present disclosure.

[0014] FIG. 4 illustrates example operations for generating an observation of an environment using an environment-agnostic causal representation machine learning model, according to certain aspects of the present disclosure.

[0015] FIG.5 depicts an example processing system configured to perform various aspects of the present disclosure.

[0016] To facilitate understanding, identical reference numerals have been used, where possible, to designate identical elements that are common to the drawings. It is contemplated that elements and features of one aspect may be beneficially incorporated in other aspects without further recitation. P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 4 DETAILED DESCRIPTION

[0017] Aspects of the present disclosure provide apparatuses, methods, processing systems, and computer-readable mediums for using machine learning models to generate observations of a variety of environments in which a computing system, such as a robotic system, an autonomous vehicle, or the like, operates.

[0018] As discussed, causal representation machine learning models are typically trained to generate observations with respect to a specific environment. These observations may be generated within a structured latent space, also referred to as causal representations. Generally, samples in the structured latent space represent the various actions which can be performed by or with respect to a specific object in an environment. Because these models are typically trained to generate observations with respect to a specific environment, these models may generate accurate observations from that specific environment, but may not be generalizable to other environments. For example, a causal representation machine learning model trained based on a training data set of observations of a specific environment obtained from a specific perspective may not generate accurate observations of different environments or of the specific environment captured from a different perspective (e.g., a different camera / imaging angle).

[0019] Because causal representation machine learning models trained on a single environment may not be generalizable to other environments, these causal representation machine learning models may not be able to adapt to new, unseen environments. To allow these models to adapt to different environments, a causal representation machine learning model may be trained on multiple environments. However, a multi-environment causal representation machine learning model may still not generalize across a wide variety of environments. For example, a training data set used to train a multi-environment causal representation model may share some causal structures, such as objects, interactions between objects when such objects collide, and the like.

[0020] To allow for causal representation machine learning models to generate observations across a variety of environments with varying causal structures from a predicted causal representation of a future state of the environment, certain aspects of the present disclosure provide techniques for predicting a future state of an environment (e.g., defined by causal variables for the future state of the environment) based on embedding representations of an environment generated by a model that is environment-agnostic. By doing so, the embedding representations of the environment may be located in a common P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 5 embedding space, which allows for a causal representation machine learning model to generalize across a variety of environments. In some aspects, these causal representation machine learning models may generate observations for an environment based on segmentation of a representation of the environment into a plurality of segments, with each segment being associated with a unique set of causal variables.

[0021] By training a causal representation machine learning model to generate observations of an environment in which a computing system operates in an environment- agnostic manner, certain aspects of the present disclosure may provide for the generation of accurate, actionable observations of an environment using a single machine learning model. These observations may be used, for example, by various other computing systems to perform various actions, such as using a robot or other autonomous system to perform a task within the environment, controlling an autonomous vehicle, or the like. Because the causal representation machine learning model is environment-agnostic, new machine learning models need not be trained to account for new environments in which the causal representation machine learning model is being deployed. Thus, certain aspects of the present disclosure may allow for a single model to be trained to perform tasks across different environments instead of training, deploying, and maintaining multiple environment-specific models, which may reduce computational resource utilization involved in training and using causal representation machine learning models to generate observations across different environments. For example, by maintaining a single causal representation machine learning model that is environment-agnostic, certain aspects of the present disclosure may reduce the amount of processor cycles, memory, and power expended in training multiple environment-specific causal representation machine learning models and may reduce the amount of memory used in storing multiple environment-specific causal representation machine learning models on a computing device on which these models are deployed. Example Environment-Agnostic Causal Representation Machine Learning Models for Observing an Environment

[0022] FIG. 1 illustrates an environment-agnostic causal representation machine learning model 100 trained to generate an observation of an environment, according to certain aspects of the present disclosure. P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 6

[0023] Generally, a causal representation model 100 may be defined as a model ℳ ൌ^ℎ, ^^^ which includes a set of causal variables ^^ ൌ ^^^^, ^^ଶ,andan injective observation function ℎ:^^ → ^^. ^^ ⊆ ℝ^ that generally represents the spaceof possible observations. To allow the causal representation model 100 to generate observations for a variety of environments, an encoder 114 may be trained to encodeobservations of a variety of environments into a latent space ^^^ which is an unconstrainedspace (e.g., implemented via an arbitrarily defined neural network). The encoder 114 may be trained, for example, using a variety of techniques, such as contrastive learning techniques in which the encoder 114 is trained based on relationships between positive and negative pairs of instances (e.g., a door of a microwave being opened or closed, a device being on or off, etc.), maximum-likelihood estimation with reconstruction, or the like.

[0024] To allow for the causal representation model 100 to be environment-agnostic,a finite universe of causal models ^^ ൌ ^ℳ^, … ,ℳ^^^ may be defined, and observationsmay be sampled from this universe. The model may be deemed to successfully observethe universe ^^ if, for each model ℳ ∈ ^^, the encoder 114 maps observations ^^to a latent spacesuch that an invertible transformation ^^ℳexists between thelearned latent space of the encoder 114 and a true latent space ^^ (e.g., that ^^ℳ^^̂^^ ൌ ^^).In other words, the encoder 114 may be shared across different environments so that the causal representation model 100 is environment-agnostic and generalizable across different environments, including environments not included in a training data set used to train the causal representation model 100. Generally, in training the encoder 114 toencode an input representing an environment into the latent space ^^^ , it may be assumedthat any two models ℳ in ^^ have non-overlapping observation spaces in the latent space^^^ .

[0025] In some aspects, the causal representation model 100 may be trained using causal representation learning techniques in which an encoder 114 is trained to minimize the distance to a target in a latent space. Generally, an objective for training the causalrepresentation model 100 may be defined as a method that identifies each model ℳ ∈ ^^by solving for the objective function: ^^^ℳ ൌ arg m^^in ^^^^,^^^∼^ℳ^^,^^^^dist൫^^^^^^^,^^^^^^, ^^^^൯൧P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 7 where dist represents a distance measure (e.g., mean squared error) and ^^ represents an arbitrary target function shared across the models ℳ in the universe of models ^^. Theuniverse of models ^^ may be identifiable from input samples ^^^,^^^^ ∼ ^^^^^^^,^^^^. Basedon a factorization of a sample distribution to ^^^^^^^, ^^^^ ൌ^^^^^ℳ^^^ℳ^^^, ^^^^, it maybe seen that a monotonic constraint allows for the objective to be represented according to the expression:

[0026] Thus, if an encoder 114, represented as ^^^, that minimizes each individual model exists, any encoder ^^^ minimizing the overall objective function minimizes each individual model. Based on this observation, a variety of causal representation learning techniques can allow for a model to be environment-agnostic and thus automatically generalize across environments, including environments that are not included in the training data set used to train the encoder 114. These techniques may include, for example, contrastive learning, for which ^^^ is an augmentation of ^^ and models the target^^^^^^, ^^^^ ൌ ^^^^^^^^ , sparse perturbation, for which ^^^ is a perturbed image ^^^^ withperturbation ^^^ఋ and models the target ^^^^^^,^^^^ ൌ ^^^^^^^^^ െ ^^^ఋ, and the like.

[0027] However, for some models, such as those that use a decoder 134 and a latent distribution to generate a causal representation of an environment, it may be seen thatdecoding, unlike encoding, may not be unique, as multiple models ℳ can sharesubspaces in the latent space ^^. To account for the sharing of subspaces by different models, a paired observation technique may be used to identify the model from which anobservation is derived. In such a case, the universe of models ^^ may be observedaccording to the expression:

[0028] A function ^^ that maps an observation to the index ^^ of a modelthatgenerated the observation may allow for the learning of a mixture of|^^|representationmodels, one for each ℳ ∈ ^^ , and the appropriate model may be selected for anobservation ^^ based on ^^^^^^, with the universe of models ^^ becoming identifiable fromobservations ^^ ∈ ⋃ℳ∈^^. In some aspects, ^^ may be inferred from mutualinformation inferred from different observations. To allow for such an inference to be P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 8 made, the encoder 114 may be trained based on solving a constraint optimization problem for a representation function:

[0029] The representation function ^^^may be subject to the limitation^^^^భ,^మ^∈^^^^^^^^^^^^^^ ൌ ^^^^^^ଶ^^^ ൌ 1. The representations of two observations from thesame model may thus be forced to be equal, and thus, a representation of two observations may be a function of the model index. By integrating ^^^into the encoder 114, thus, theuniverse may be identified by learning a mixture of |^^^^^^^^|^^ ∈ ⋃ ^^ℳℳ∈^^ ^| ൌ |^^|representation models.

[0030] Based on the observations above, the causal representation model 100 may be trained to generate a prediction of an environment after the performance of an action 122.The prediction may be generated based on predicted interaction variables ^^௧ ൌ ^^^^^^^^^௧^generated by the action multi-layer perceptron (MLP) 124 and a transition prior 126 in the latent space of the causal representation model 100. To do so, at time stepan observation 112 may be input into the encoder 114 to generate: (i) a latent space representation 116, including a feature vector ^^ℳ௧for the observation 112, representing a causal representation of the environment depicted in the observation 112 and (ii) an encoder output ^̂^௧representing an encoding of the observation 112. To allow for the MLP 124 and transition prior 126 to predict a future state of the environment (e.g., at time step ^^௧ା^130), a randomly sampled observation 120 ^^ఛmay be encoded by the encoder 114 into a latent space representation 118 ^^ℳఛ. An alignment loss between the feature vector ^^ℳ௧of the latent space representation 116 and the latent space representation 118 ^^ℳఛof the randomly sampled observation 120 may be used by the MLP 124 and the transition prior 126 to generate the prediction of the future state of the environment at time step ^^௧ା^130.

[0031] Generally, the MLP 124 may transform the identified action 122 by modelinga probability distributionthrough the prediction of the interactionvariables ^^௧ ൌ ^^^^^^^^^௧^ and the transition prior 126 in the latent space, represented bythe expression ^^൫^^^௧ା^ห^^^௧ , ^^௧൯. The output ^̂^௧ା^ 132 of the transition prior 126 may bepassed to the decoder 134 for reconstruction into the predicted state of the environment at time step ^^௧ା^130. The output 132 of the transition prior, ^̂^௧ା^, generally represents P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 9 an encoding in the latent space ^^ for the predicted state of the environment at time step ^^௧ା^130. To decode the predicted state 136 of the environment at time step ^^௧ା^130, the decoder 134 may use the alignment loss between ^^ℳ௧and ^^ℳఛdiscussed above and the output ^̂^௧ା^132 of the transition prior 126 to decode the latent representation of the environment at time step ^^௧ା^130, based on which a computing system can take one or more actions (e.g., to perform another action in the environment).

[0032] FIG. 2 illustrates an environment-agnostic causal representation machine learning model 200 trained to generate an observation of an environment based on segmentation of an input image into a plurality of segments associated with different objects in the environment, according to certain aspects of the present disclosure.

[0033] Generally, to allow for the causal representation machine learning model 200 to be generalizable across different environments, generalizability may be enforced within an encoder used to encode an observation of an environment into a latent space ^^. To enforce generalizability in an encoder, the relationships between objects and causal variables may be leveraged by segmenting an observation 210 of an environment into a plurality of segments 2201through 220N(collectively referred to as “segments 220”). This segmentation may be performed, for example, using various object detection and / or semantic segmentation models which can detect instances of various objects from an observation of an environment, such as an image or other visual data depicting an environment. Generally, each segment 220 may correspond to a specific object in the environment with which an action can be performed. For example, the observation 210 of the environment may include a cabinet, a microwave, and a slice of bread (amongst others, not illustrated in FIG. 2), each of which may have unique causal attributes and interactions between different objects represented by causal relationships. A causal representation of each object in the environment may, for example, include state variables associated with different objects in the environment, state variables associated with the environment itself (e.g., temperature, humidity, other global parameters defining the environment, etc.).

[0034] To allow for the causal representation machine learning model 200 to generate an observation of an environment, object segmentation techniques can be used to decompose the observation 210 into the plurality of segments 220. By decomposing the observation 210 of the environment into the plurality of segments 220, the causalvariables ^^ defining the environment for which the observation 210 is made may beP+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 10decomposed into a plurality of blocks ^^^^|^^ ∈ ℬ^, where ℬ is a partition of ^1, … , ^^௭^and an observation function ℎ: ^^ → ^^ implemented by the encoder 230 is a mixing ofindividual observation functions across the different blocks ^^^. Thus, the observation function may be represented by the equation:represents a binary masking function with

[0035] In segmenting the observation 210 of the environment into the plurality of segments 220, it may be assumed that any portion of the observation 210 belongs to a specific segment from the plurality of segments 220. This assumption may be enforced by the binary masking function ^^^, which may restrict some interactions between objects, such as overlapping shadows in the environment or the like. By segmenting the observation 210 of the environment into the plurality of segments 220, the causal representation machine learning model 200 may allow for objects to be moved within the environment and for other changes to the appearance of objects in the environment.

[0036] Within the universe ^^ of models ℳ , objects may be summarized acrossmodels as the set ^^^^, where an object ^^ ൌ 〈^^ை,ℎை,^^ை〉 ∈ ^^^^ is defined by a maskedobservation function and a latent space. The observation space of any object ^^ைmayrepresent the space ^^இ^^^^ ⊙ ℎை^^^^ for all ^^ ∈ ^^ை . Thus, a model ℳ ∈ ^^ may bedefined as a composition of a set of objects in ^^^^, and each model ℳ may share objects in the universe of objects ^^^^. Based on the definition of objects across the environmentswhich are covered by the universe ^^ of models ℳ , a segmentation function used tosegment the observation 210 into the segments 220 may be represented by the function^^ௌ^where ^^denotes a power set. The segmentation function ^^ௌ^generally causally segments the observations ^^ of a model ℳ if, for each observation ^^, the set of predicted masks is equal to the masks of the observation function.

[0037] Generally, the alignment of different sets of masks between a first time step ^^and a second time step ^^ ^ 1 may be represented by ^^^^^ୗ^^^^^^,^^ୗ^^^^ଶ^^. ^^ generallyreturns an order for which ∀^^ ∈ 1..^^^^^^^^^ଶ^^ ^ . Thus, for a causal representation machine learning model 200 that canP+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 11identify each model ℳ ∈ ^^ individually and a segmentation function ^^ௌ^, the universe^^ may be identifiable from a dataset of pairs of observations that share the same model ℳ and have aligned object masks.

[0038] Because each portion of the observation 210 may belong to only one object (represented by a segment 220), an encoding of the observation 210 of an environment may be obtained via a concatenation of individual encodings of segments 220 generated by a segmentation function applied to the observation 210. Thus, the encoders 230, represented as ^^^^^^, map a segment 220 of an observation 210 to the causal variables 2401 through 240N (collectively referred to as “causal variables 240”) associated with the respective segment 220 (that is, an object represented with a specific segment 220 of the observation 210). By encoding each segment 220 separately (and thus, encoding each object in the observation 210 of the environment separately), different objects may be composed such that a composition of the causal representations 240 of the objects in the observation 210 of the environment results in a representation that identifies the causal variables of the environment as a whole.

[0039] Generally, in the causal representation machine learning model 200, atransformation ^^ைfor each object ^^ exists such that ^^ை ൌ ^^ை^^̂^ை^. When concatenatingthe segments 220 (that is, concatenating the objects in the environment depicted by the observation 210), a representation of the environment depicted by the observation 210may be represented by the expression: ^̂^ ൌ ^^^^^^^^^^^^ ^ௌ ^^^^ ^^^^,^^^ ^^^^^ ^ ^^^ௌ^^^^^ଶ^^^^, … ^.Because of the identifiability of ^^ை, there exists a transformation and a permutation suchthata permutation oflatents in the latent space into which segments 220 are encoded by the encoders 230 is within a class of functions that are identifiable, the representation ^̂^ of the composition of objects in the environment depicted by the observation 210 may identify the model ℳ relevant to the environment depicted by the observation 210.

[0040] Thus, to generate a prediction of the state of an environment at time step ^^ ^1 based on the observation 210 of the environment at time step ^^, each object in the universe of objects may follow the same observation function, and thus, a single unified encoder 230 may be applied to each object identified in the observation 210 of the environment. For matching techniques, such as sparse perturbation or contrastivelearning, latent space representations of objects between a base observation ^^ and aP+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 12perturbed observation ^^^ may be aligned. To do so, object representations betweendifferent images at different time steps may be aligned. However, because observation pairs between a base observation and a perturbed observation may not be available, blocks of latents may be matched based on Hungarian matching techniques in which a maximum-weight matching in a graph may be identified.

[0041] For temporal techniques, latents may be matched between different time steps, taking into account interactions between objects that may occur over time and different environments with differing numbers of objects. To do so, a graph neural network 250 may be used in the transition prior (e.g., the transition prior 126 discussed above with respect to FIG.1) to generate the causal representations 2601through 260N(collectively referred to as “causal representations 260”) for a subsequent time step. Generally, the graph neural network 250 may be implemented in self-attention layers for cross-object communication and in multi-layer perceptron layers for per-object processing, which allows for interactions to be modeled between different objects over time. The predictionsor observations at time ^^ may be matched with the latent variables at time step ^^ ^ 1based, for example, on Hungarian matching techniques and a minimization of a negative log likelihood or on object representations if paired observations are available.

[0042] FIG.3 illustrates an example 300 of causal variables predicted for an object in an environment using an environment-agnostic causal representation machine learning model, according to certain aspects of the present disclosure.

[0043] As illustrated, in the example 300, an observation 310 may be segmented into a plurality of segments, including a segment 320 corresponding to a specific object in the environment depicted by the observation 310. For example, the segment 320, as illustrated, may correspond to a microwave oven in an observation 310 of a kitchen, including cabinets, a stove, and other objects in the environment. The encoder 230, as discussed above, may be a single encoder that encodes a representation of any object in a universe of objects into a causal representation 330. The causal representation 330 generated by the encoder 230 for the segment 320 may be a plurality of variables representing, for example, the state of the object in the segment 320. For example, the state of the object in the segment 320, and thus the causal variables in the causal representation 330, may represent changeable state information associated with the object. For the microwave depicted in the segment 320, the causal variables may define whether the door of the microwave is open or closed, whether the microwave is on or off, P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 13 and the like. It should be recognized that the example of the microwave depicted in FIG. 3 is but an example, and the encoder 230 may generate causal representations of other objects that may be included in a training data set of depictions of environments which are used to train the encoder 230. Example Operations for Generating Observations of an Environment Using an Environment-Agnostic Causal Representation Machine Learning Model

[0044] FIG.4 illustrates example operations 400 for generating an observation of an environment using an environment-agnostic causal representation machine learning model (e.g., causal representation machine learning model 100 or 200), according to certain aspects of the present disclosure. The operations 400 may be performed on a computing device on which a machine learning model may be trained, such as a server computer, a cluster of physical computing instances, one or more cloud computing instances, or the like.

[0045] As illustrated, the operations 400 begin at block 410, with receiving, at a computing system, an image of an environment and an indication of an action (e.g., action 122) to be performed in the environment.

[0046] At block 420, the operations 400 proceed with generating, using an encoder neural network (e.g., encoder 114 or 230) trained to generate embedding representations of images from a plurality of environments into a common embedding space, an embedding representation of the image.

[0047] At block 430, the operations 400 proceed with predicting, using a predictive neural network (e.g., the transition prior 126), causal variables for a future state of the environment based on the action and the embedding representation of the image. In some aspects, predicting the causal variables for the future state of the environment at block 430 is further based on one or more global environmental variables associated with the environment.

[0048] At block 440, the operations 400 proceed with predicting, based on the predicted causal variables, the future state of the environment after execution of the action.

[0049] In some aspects, the operations 400 further include segmenting, using an image segmentation machine learning model, the image of the environment into a P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 14 plurality of segments (e.g., segments 220). In this case, generating the embedding representation of the image at block 420 may involve generating a plurality of embedding representations of the image based on embedding each segment of the plurality of segments into a latent space. Segmenting the image of the environment into the plurality of segments may, in some aspects, include segmenting the image based on objects detected in the environment.

[0050] In some aspects, the predictive neural network may be a graph neural network (e.g., graph neural network 250) trained to predict the causal variables for the future state of the environment for each segment of the plurality of segments based on interactions between different objects in the environment.

[0051] In some aspects, the causal variables comprise one or more variables describing a respective object associated with a respective segment of the plurality of segments. The variables may include, for example, information describing a state of various properties of the object. These various properties may include, for example, changeable properties of the object, whether the object is active (on) or inactive (off), or the like. In some aspects, the causal variables may include information about whether an object is movable or fixed within the environment and position information associated with movable objects.

[0052] In some aspects, the predicted causal variables for each segment of the plurality of segments comprise a permutation and a transformation for each respective segment.

[0053] In some aspects, the causal variables comprise a plurality of variables describing the environment as a whole. For example, the variables describing the environment as a whole may include information such as a temperature of the environment, a humidity of the environment, an amount of light in the environment, or other variables which may affect the observations and the behavior of objects in the environment.

[0054] In some aspects, the future state of the environment after execution of the action comprises a perturbed version of the received image representing the future state of the environment.

[0055] In some aspects, the future state of the environment after execution of the action comprises a predicted image representing the environment at a time subsequent to P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 15 a time associated with the received image. The predicted image may, for example, be represented by an embedding in a latent space, such as the latent space into which the encoder encodes a received image into the embedding representations. The predicted image may be decoded from the embedding of the predicted image in the latent space.

[0056] In some aspects, the image of the environment depicts an environment different from environments used to train the encoder neural network. The encoder neural network may thus allow for generalizable, cross-environment observations and predictions of future observations of an environment using a machine learning model.

[0057] In some aspects, the image of the environment depicts an environment included in a training data set used to train the encoder neural network from a perspective different from a perspective of the environment included in the training data set.

[0058] In some aspects, the action to be performed involves controlling a robot (or robotic system) to manipulate an object from a first state to a second state (e.g., move an object from a first position to a second position). In this case, the image of the environment may illustrate the object in the first state, and the predicted future state of the environment may represent the object in the second state. Example Processing Systems for Generating Observations of an Environment Using an Environment-Agnostic Causal Representation Machine Learning Model

[0059] FIG. 5 depicts an example processing system 500 configured to perform various aspects of the present disclosure, including, for example, the techniques and methods described with respect to FIGS.1-4. In some aspects, the processing system 500 may train, implement, or provide a set of machine learning models, such as the environment-agnostic causal representation machine learning model 100 of FIG.1, the environment-agnostic causal representation machine learning model 200 of FIG.2, or the like. Although depicted as a single system for conceptual clarity, in at least some aspects, as discussed above, the operations described below with respect to the processing system 500 may be distributed across any number of devices.

[0060] The processing system 500 includes a central processing unit (CPU) 502, which in some examples may be a multi-core CPU. Instructions executed at the CPU 502 may be loaded, for example, from a program memory associated with the CPU 502 or may be loaded from a partition of memory 524. P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 16

[0061] The processing system 500 also includes additional processing components tailored to specific functions, such as a graphics processing unit (GPU) 504, a digital signal processor (DSP) 506, a neural processing unit (NPU) 508, a multimedia processing unit 510, and a wireless connectivity component 512.

[0062] An NPU, such as NPU 508, is generally a specialized circuit configured for implementing control and arithmetic logic for executing machine learning algorithms, such as algorithms for processing artificial neural networks (ANNs), deep neural networks (DNNs), random forests (RFs), and the like. An NPU may sometimes alternatively be referred to as a neural signal processor (NSP), tensor processing unit (TPU), neural network processor (NNP), intelligence processing unit (IPU), vision processing unit (VPU), or graph processing unit.

[0063] NPUs, such as the NPU 508, are configured to accelerate the performance of common machine learning tasks, such as image classification, machine translation, object detection, and various other predictive models. In some examples, a plurality of NPUs may be instantiated on a single chip, such as a system-on-a-chip (SoC), while in other examples the NPUs may be part of a dedicated neural-network accelerator.

[0064] NPUs may be optimized for training or inference, or in some cases configured to balance performance between both. For NPUs that are capable of performing both training and inference, the two tasks may still generally be performed independently.

[0065] NPUs designed to accelerate training are generally configured to accelerate the optimization of new models, which is a highly compute-intensive operation that involves inputting an existing dataset (often labeled or tagged), iterating over the dataset, and then adjusting model parameters, such as weights and biases, in order to improve model performance. Generally, optimizing based on a wrong prediction involves propagating back through the layers of the model and determining gradients to reduce the prediction error.

[0066] NPUs designed to accelerate inference are generally configured to operate on complete models. Such NPUs may thus be configured to input a new piece of data and rapidly process this new data through an already trained model to generate a model output (e.g., an inference).

[0067] In some implementations, the NPU 508 is a part of one or more of the CPU 502, the GPU 504, and / or the DSP 506. P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 17

[0068] In some examples, the wireless connectivity component 512 may include subcomponents, for example, for third generation (3G) connectivity, fourth generation (4G) connectivity (e.g., Long-Term Evolution (LTE)), fifth generation connectivity (e.g., 5G or New Radio (NR)), Wi-Fi connectivity, Bluetooth connectivity, and other wireless transmission standards. The wireless connectivity component 512 is further coupled to one or more antennas 514.

[0069] The processing system 500 may also include one or more sensors and / or sensor processing units 516 associated with any manner of sensor, one or more image signal processors (ISPs) 518 associated with any manner of image sensor, and / or a navigation component 520, which may include satellite-based positioning system components (e.g., GPS or GLONASS), as well as inertial positioning system components.

[0070] The processing system 500 may also include one or more input and / or output devices 522, such as screens, touch-sensitive surfaces (including touch-sensitive displays), physical buttons, speakers, microphones, and the like.

[0071] In some examples, one or more of the processors of the processing system 500 may be based on an ARM or RISC-V instruction set.

[0072] The processing system 500 also includes the memory 524, which is representative of one or more static and / or dynamic memories, such as a dynamic random access memory, a flash-based static memory, and the like. In this example, the memory 524 includes computer-executable components, which may be executed by one or more of the aforementioned processors of the processing system 500.

[0073] In particular, in this example, the memory 524 includes an image receiving component 524A, an embedding generating component 524B, a causal variable predicting component 524C, a future state predicting component 524D, and machine learning models 524E. Though depicted as discrete components for conceptual clarity in FIG. 5, the illustrated components (and others not depicted) may be collectively or individually implemented in various aspects.

[0074] Generally, the processing system 500 and / or components thereof may be configured to perform the methods described herein.

[0075] Notably, in other aspects, components of the processing system 500 may be omitted, such as where the processing system 500 is a server computer or the like. For example, the multimedia processing unit 510, the wireless connectivity component 512, P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 18 the sensor processing units 516, the ISPs 518, and / or the navigation component 520 may be omitted in other aspects. Further, components of the processing system 500 may be distributed between multiple devices. Example Clauses

[0076] Implementation details of various aspects of the present disclosure are described in the following numbered clauses:

[0077] Clause 1: A processor-implemented method for machine learning, comprising: receiving, at a computing system, an image of an environment and an indication of an action to be performed in the environment; generating, using an encoder neural network trained to generate embedding representations of images from a plurality of environments into a common embedding space, an embedding representation of the image; predicting, using a predictive neural network, causal variables for a future state of the environment based on the action and the embedding representation of the image; and based on the predicted causal variables, predicting the future state of the environment after execution of the action.

[0078] Clause 2: The method of Clause 1, further comprising: segmenting, using an image segmentation machine learning model, the image of the environment into a plurality of segments, wherein generating the embedding representation of the image comprises generating a plurality of embedding representations of the image based on embedding each segment of the plurality of segments into a latent space.

[0079] Clause 3: The method of Clause 2, wherein segmenting the image of the environment into the plurality of segments comprises segmenting the image based on objects detected in the environment.

[0080] Clause 4: The method of Clause 2 or 3, wherein the predictive neural network comprises a graph neural network trained to predict the causal variables for the future state of the environment for each segment of the plurality of segments based on interactions between different objects in the environment.

[0081] Clause 5: The method of any of Clauses 2 through 4, wherein the causal variables comprise one or more variables describing a respective object associated with a respective segment of the plurality of segments. P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 19

[0082] Clause 6: The method of any of Clauses 2 through 5, wherein the predicted causal variables for each segment of the plurality of segments comprise a permutation and a transformation for each respective segment.

[0083] Clause 7: The method of any of Clauses 1 through 6, wherein the causal variables comprise a plurality of variables describing the environment as a whole.

[0084] Clause 8: The method of any of Clauses 1 through 7, wherein the future state of the environment after execution of the action comprises a perturbed version of the received image representing the future state of the environment.

[0085] Clause 9: The method of any of Clauses 1 through 8, wherein the future state of the environment after execution of the action comprises a predicted image representing the environment at a time subsequent to a time associated with the received image.

[0086] Clause 10: The method of any of Clauses 1 through 9, wherein the environment depicted in the image is different from environments used to train the encoder neural network.

[0087] Clause 11: The method of any of Clauses 1 through 10, wherein the environment depicted in the image is included in a training data set used to train the encoder neural network, but is from a different perspective than the environment included in the training data set.

[0088] Clause 12: The method of any of Clauses 1 through 11, wherein predicting the causal variables for the future state of the environment is further based on one or more global environmental variables associated with the environment.

[0089] Clause 13: The method of any of Clauses 1 through 12, wherein the action to be performed comprises controlling a robot to manipulate an object from a first state to a second state, wherein the image of the environment illustrates the object in the first state, and wherein the predicted future state of the environment represents the object in the second state.

[0090] Clause 14: A processing system comprising: at least one memory comprising computer-executable instructions; and one or more processors coupled to the at least one memory and configured to execute the computer-executable instructions and cause the processing system to perform a method in accordance with any of Clauses 1 through 13. P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 20

[0091] Clause 15: A processing system comprising means for performing a method in accordance with any of Clauses 1 through 13.

[0092] Clause 16: A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by one or more processors of a processing system, cause the processing system to perform a method in accordance with any of Clauses 1 through 13.

[0093] Clause 17: A computer program product embodied on a computer-readable storage medium comprising code for performing a method in accordance with any of Clauses 1 through 13. Additional Considerations

[0094] The preceding description is provided to enable any person skilled in the art to practice the various aspects described herein. The examples discussed herein are not limiting of the scope, applicability, or aspects set forth in the claims. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. For example, changes may be made in the function and arrangement of elements discussed without departing from the scope of the disclosure. Various examples may omit, substitute, or add various procedures or components as appropriate. For instance, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. Also, features described with respect to some examples may be combined in some other examples. For example, an apparatus may be implemented, or a method may be practiced, using any number of the aspects set forth herein. In addition, the scope of the disclosure is intended to cover such an apparatus or method that is practiced using other structure, functionality, or structure and functionality in addition to, or other than, the various aspects of the disclosure set forth herein. It should be understood that any aspect of the disclosure disclosed herein may be embodied by one or more elements of a claim.

[0095] As used herein, the word “exemplary” means “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects.

[0096] As used herein, a phrase referring to “at least one of” a list of items refers to any combination of those items, including single members. As an example, “at least one P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 21 of: a, b, or c” is intended to cover a, b, c, a-b, a-c, b-c, and a-b-c, as well as any combination with multiples of the same element (e.g., a-a, a-a-a, a-a-b, a-a-c, a-b-b, a-c-c, b-b, b-b-b, b-b-c, c-c, and c-c-c or any other ordering of a, b, and c).

[0097] As used herein, the term “determining” encompasses a wide variety of actions. For example, “determining” may include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, a database, or another data structure), ascertaining, and the like. Also, “determining” may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), and the like. Also, “determining” may include resolving, selecting, choosing, establishing, and the like.

[0098] The methods disclosed herein comprise one or more steps or actions for achieving the methods. The method steps and / or actions may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims. Further, the various operations of methods described above may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software component(s) and / or module(s), including, but not limited to a circuit, an application specific integrated circuit (ASIC), or processor. Generally, where there are operations illustrated in figures, those operations may have corresponding counterpart means-plus-function components with similar numbering.

[0099] The following claims are not intended to be limited to the aspects shown herein, but are to be accorded the full scope consistent with the language of the claims. Within a claim, reference to an element in the singular is not intended to mean “one and only one” unless specifically so stated, but rather “one or more.” Unless specifically stated otherwise, the term “some” refers to one or more. No claim element is to be construed under the provisions of 35 U.S.C. § 112(f) unless the element is expressly recited using the phrase “means for” or, in the case of a method claim, the element is recited using the phrase “step for.” All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims. P+S Ref. No.: QUAL / 2407433PC

Claims

Client Ref. No.: 2407433WO 22 WHAT IS CLAIMED IS:

1. A processor-implemented method for machine learning, comprising: receiving, at a computing system, an image of an environment and an indication of an action to be performed in the environment; generating, using an encoder neural network trained to generate embedding representations of images from a plurality of environments into a common embedding space, an embedding representation of the image; predicting, using a predictive neural network, causal variables for a future state of the environment based on the action and the embedding representation of the image; and based on the predicted causal variables, predicting the future state of the environment after execution of the action.

2. The method of Claim 1, further comprising segmenting, using an image segmentation machine learning model, the image of the environment into a plurality of segments, wherein generating the embedding representation of the image comprises generating a plurality of embedding representations of the image based on embedding each segment of the plurality of segments into a latent space.

3. The method of Claim 2, wherein segmenting the image of the environment into the plurality of segments comprises segmenting the image based on objects detected in the environment.

4. The method of Claim 2, wherein the predictive neural network comprises a graph neural network trained to predict the causal variables for the future state of the environment for each segment of the plurality of segments based on interactions between different objects in the environment.

5. The method of Claim 2, wherein the causal variables comprise one or more variables describing a respective object associated with a respective segment of the plurality of segments.

6. The method of Claim 2, wherein the predicted causal variables for each segment of the plurality of segments comprise a permutation and a transformation for each respective segment. P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 23 7. The method of Claim 1, wherein the causal variables comprise a plurality of variables describing the environment as a whole.

8. The method of Claim 1, wherein the future state of the environment after execution of the action comprises a perturbed version of the received image representing the future state of the environment.

9. The method of Claim 1, wherein the future state of the environment after execution of the action comprises a predicted image representing the environment at a time subsequent to a time associated with the received image.

10. The method of Claim 1, wherein: the environment depicted in the image is different from environments used to train the encoder neural network; or the environment depicted in the image is included in a training data set used to train the encoder neural network, but is from a different perspective than the environment included in the training data set.

11. The method of Claim 1, wherein predicting the causal variables for the future state of the environment is further based on one or more global environmental variables associated with the environment.

12. A processing system comprising: one or more memories comprising processor-executable instructions; and one or more processors coupled to the one or more memories and configured to execute the processor-executable instructions and cause the processing system to: receive an image of an environment and an indication of an action to be performed in the environment; generate, using an encoder neural network trained to generate embedding representations of images from a plurality of environments into a common embedding space, an embedding representation of the image; predict, using a predictive neural network, causal variables for a future state of the environment based on the action and the embedding representation of the image; and predict, based on the predicted causal variables, the future state of the environment after execution of the action. P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 24 13. The processing system of Claim 12, wherein: the one or more processors are further configured to execute the processor- executable instructions and cause the processing system to segment, using an image segmentation machine learning model, the image of the environment into a plurality of segments; and to generate the embedding representation of the image, the one or more processors are configured to execute the processor-executable instructions and cause the processing system to generate a plurality of embedding representations of the image based on embedding each segment of the plurality of segments into a latent space.

14. The processing system of Claim 13, wherein to segment the image of the environment into the plurality of segments, the one or more processors are configured to execute the processor-executable instructions and cause the processing system to segment the image based on objects detected in the environment.

15. The processing system of Claim 13, wherein the predictive neural network comprises a graph neural network trained to predict the causal variables for the future state of the environment for each segment of the plurality of segments based on interactions between different objects in the environment.

16. The processing system of Claim 13, wherein the causal variables comprise one or more variables describing a respective object associated with a respective segment of the plurality of segments.

17. The processing system of Claim 13, wherein the predicted causal variables for each segment of the plurality of segments comprise a permutation and a transformation for each respective segment.

18. The processing system of Claim 12, wherein the future state of the environment after execution of the action comprises: a perturbed version of the received image representing the future state of the environment; or a predicted image representing the environment at a time subsequent to a time associated with the received image. P+S Ref. No.: QUAL / 2407433PCClient Ref. No.: 2407433WO 25 19. The processing system of Claim 12, wherein the environment depicted in the image is: different from environments used to train the encoder neural network; or included in a training data set used to train the encoder neural network, but is from a different perspective than the environment included in the training data set.

20. The processing system of Claim 12, wherein the action to be performed comprises controlling a robot to manipulate an object from a first state to a second state, wherein the image of the environment illustrates the object in the first state, and wherein the predicted future state of the environment represents the object in the second state. P+S Ref. No.: QUAL / 2407433PC