Isovariant trajectory optimization with diffusion model

Through the SE(3) group equivariant diffusion model, the environmental symmetry of the robot system is used for trajectory planning, which solves the problems of low sample efficiency and poor generalization ability in the existing technology and achieves efficient trajectory planning and task adaptability.

CN120641910APending Publication Date: 2025-09-12QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380093446.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-27
Filing Date
2023-12-26
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

When using learned models for path planning, existing technologies have difficulty effectively utilizing the geometric symmetry of the environment, resulting in low sample efficiency, slow training speed and poor generalization ability, especially in robotic systems where it is difficult to generalize to new tasks.

Method used

An equivariant diffusion model based on the SE(3) group is adopted to separate state-action pairs into geometric data types through a training preparation engine, and an equivariant denoising network is used to generate output data, and trajectory planning is performed in combination with the symmetry of the environment.

Benefits of technology

It improves sample efficiency, training speed and model generalization ability, can better adapt to different task environments, and achieve efficient trajectory planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120641910A_ABST
    Figure CN120641910A_ABST
Patent Text Reader

Abstract

Systems and techniques for modeling tasks using geometries are described herein. An example method includes receiving, via a training preparation engine, a training dataset including state action pairs; separating, via the training preparation engine, the pairs of state actions from the training data set into geometric data types; converting the geometric data types into internal representations via the training preparation engine; processing the internal representations via an equivariant denoising network to generate output data; and transforming the output data into a data representation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to solutions to modeling and planning problems. For example, aspects of the present disclosure relate to an equivariant diffuser model that can solve modeling and planning problems by accounting for the symmetric geometry of any given problem, such as motion, navigation, and object manipulation of a system or device (e.g., a vehicle, robotic system or device, etc.). Background Art

[0002] Path modeling and planning can be used by various systems or devices, such as vehicles, robotic systems, aircraft or drones, and / or other systems or devices. For example, a vehicle or robotic system can determine a path for navigation purposes. In some cases, a learned model (e.g., a machine learning model, such as a neural network model) can be used to perform path planning. In some cases, the learned model can be input into a classical trajectory optimization routine. Summary of the Invention

[0003] Systems and techniques are described for providing diffusion models that account for the geometry of an environment, such as for robotics applications. For example, these systems and techniques can exploit symmetries within the structure. These systems and techniques can generate diffusion models for planning that include equivariance constraints based on the SE(3) group (a special Euclidean group with three elements), which results in group-invariant density on trajectories. Such diffusion models can result in improved sample efficiency, training speed, and / or model generalization, among other benefits.

[0004] Various systems (e.g., robots) operate in a structured world and often solve tasks with spatial, temporal, and permutation symmetries. Most reinforcement learning algorithms do not take this structure into account. Despite significant success in idealized scenarios, learning algorithms typically require extensive training and have poor generalization capabilities. To improve sample efficiency, robustness, and generalization capabilities, various aspects provide an algorithm for model-based reinforcement learning and planning that is robust to spatial symmetry groups SE(3) (a special Euclidean group with 3 elements), discrete time translation groups Z, and permutation groups S. n In some aspects, these systems and techniques can be based on the diffuser paradigm, which treats the learning of the dynamics model and policy as a single generative modeling problem and trains the diffusion model to solve problems that do not take into account geometric structure when generalizing. An equivariant architecture is employed to invariantly enforce both the dynamics model and the policy. Conditioning and classifier-based guidance allow the method to softly violate equivariance for specific tasks as needed. The equivariant robotic diffuser algorithm is demonstrated on navigation and object manipulation tasks. Compared to an unstructured diffuser baseline, the new model improves final task performance, is more sample-efficient, and generalizes better across symmetry groups.

[0005] In some examples, a processor-implemented method for modeling a task using geometry includes: receiving a training dataset comprising state-action pairs via a training preparation engine; separating the state-action pairs from the training dataset into geometric data types via the training preparation engine; converting the geometric data types into internal representations via the training preparation engine; processing the internal representations via an equivariant denoising network to generate output data; and transforming the output data into a data representation.

[0006] In some examples, an apparatus for modeling a task using symmetries in geometric structures may include at least one memory (e.g., configured in a circuit) and at least one processor coupled to the at least one memory and configured to: receive a training dataset comprising state-action pairs via a training preparation engine; separate the state-action pairs from the training dataset into geometric data types via the training preparation engine; convert the geometric data types into internal representations via the training preparation engine; process the internal representations via an equivariant denoising network to generate output data; and transform the output data into a data representation.

[0007] In some examples, a non-transitory computer-readable medium is provided having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: receive a training dataset comprising state-action pairs via a training preparation engine; separate the state-action pairs from the training dataset into geometric data types via the training preparation engine; convert the geometric data types into internal representations via the training preparation engine; process the internal representations via an equivariant denoising network to generate output data; and transform the output data into a data representation.

[0008] In some examples, an apparatus for performing object detection is provided. The apparatus includes: means for receiving a training dataset comprising state-action pairs via a training preparation engine; means for separating the state-action pairs from the training dataset into geometric data types via the training preparation engine; means for converting the geometric data types into internal representations via the training preparation engine; means for processing the internal representations via an equivariant denoising network to generate output data; and means for transforming the output data into a data representation.

[0009] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. This subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.

[0010] The foregoing and other features and aspects will become more apparent upon reference to the following description, claims and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Illustrative examples of the present application are described in detail below with reference to the following drawings:

[0012] Figure 1 is a diagram illustrating an example of a diffuser;

[0013] Figure 2 is a diagram illustrating an example of how a diffuser according to aspects of the present disclosure may sample a plan by iteratively denoising a two-dimensional array;

[0014] Figure 3 is a diagram illustrating an example of learning a long-term planned diffuser operation according to aspects of the present disclosure;

[0015] Figure 4 is a diagram illustrating how a diffusion model according to aspects of the present disclosure will perform a forward diffusion process and a reverse denoising process;

[0016] Figure 5 is a diagram illustrating an example of planning using a diffusion model according to aspects of the present disclosure, showing the coupling between modeling and planning;

[0017] Figure 6 is a diagram illustrating an example of how symmetry in geometric structures may be exploited according to aspects of the present disclosure;

[0018] Figure 7 is a diagram illustrating an example of how to generate a trained equivariant diffuser according to aspects of the present disclosure;

[0019] Figure 8A is a diagram illustrating an example method associated with utilizing symmetry of geometric structures according to aspects of the present disclosure;

[0020] Figure 8B is a diagram illustrating another example method associated with using symmetry of geometric structures according to aspects of the present disclosure;

[0021] Figure 9A is a diagram illustrating an example of a geometric U-Net that can view state-action trajectories as images according to various aspects of the present disclosure;

[0022] Figure 9B is a diagram illustrating an example of a geometric U-Net that can view state-action trajectories as images according to various aspects of the present disclosure;

[0023] Figure 10 is a diagram illustrating an example of state decomposition according to aspects of the present disclosure;

[0024] Figure 11 is a diagram illustrating an example of equivariant trajectory generation according to aspects of the present disclosure;

[0025] Figure 12A is a diagram illustrating example usage related to a robot according to aspects of the present disclosure;

[0026] Figure 12B is a diagram illustrating an example use in connection with robotics according to aspects of the present disclosure, showing trajectories generated by an equivariant diffuser model; and

[0027] Figure 13 is a diagram illustrating an example of a computing system according to aspects of the present disclosure. DETAILED DESCRIPTION

[0028] The following provides certain aspects of the present disclosure. Some of these aspects can be applied independently, and some of them can be applied in combination, which will be apparent to those skilled in the art. In the following description, specific details are set forth for explanation purposes to provide a thorough understanding of various aspects of the application. However, it is apparent that various aspects can be practiced without these specific details. Each drawing and description is not intended to be restrictive.

[0029] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of the present disclosure. Instead, the following description of exemplary aspects will provide those skilled in the art with a description that can be used to implement the exemplary aspects. It should be understood that various changes may be made to the function and arrangement of elements without departing from the spirit and scope of the present application as set forth in the appended claims.

[0030] Utilizing learned models for path or route planning can be a conceptually simple framework for reinforcement learning and data-driven decision making. The appeal of using learned models for route planning can be that learning techniques are employed only in areas where they are most mature and effective, such as for approximating unknown environmental dynamics, a process essentially equivalent to a supervised learning problem. The learned model can then be input into classic trajectory optimization routines, which are also well understood in their original context. However, this combination may not achieve the desired results. For example, because powerful trajectory optimizers can utilize learned models, plans generated using such procedures may look more like adversarial examples than optimal trajectories. Consequently, contemporary model-based reinforcement learning algorithms often inherit more from model-free methods (such as value functions and policy gradients) than from trajectory optimization toolboxes. Techniques that do rely on online planning often use simple gradient-free trajectory optimization routines, such as random shooting or cross-entropy methods, to avoid the aforementioned issues.

[0031] One approach to data-driven trajectory optimization is to train models that are directly applicable to trajectory optimization, in the sense that sampling from the model and using it for planning become almost identical. Achieving this goal may require changes in how the model is designed. Because learned dynamics models are often used as proxies for the dynamics of the environment, improvements are often achieved by constructing the model in terms of the underlying causal processes. Instead, some have considered how to design models that are tailored to the planning problem in which the model will be used. For example, because the model will ultimately be used for planning, the action distribution may be as important as the state dynamics, and long-term accuracy may be more important than single-step error. On the other hand, it may be beneficial for the model to remain agnostic to the reward function, allowing it to be used on multiple tasks, including those not seen during training. Furthermore, it may be beneficial to design the model so that its plans, not just its predictions, improve with experience and are resistant to the myopic failure modes of standard shot-based planning algorithms.

[0032] In some cases, trajectory-level diffusion probability models and diffusers can be used for planning. While standard model-based planning techniques autoregressively predict forward in time, "diffusers" or "diffusion models" predict all time steps of the plan simultaneously. The iterative sampling process of the diffuser model enables flexible conditioning, allowing the auxiliary guide to modify the sampling process to recover trajectories with high rewards or that satisfy a set of constraints. The above conception of data-driven trajectory optimization has some attractive properties, such as long-term scalability, task composability, temporal composability, and efficient non-greedy planning. Janner et al. disclosed some methods in the paper "Planning with Diffusion for Flexible Behavior Synthesis" held at the 39th International Conference on Machine Learning held in Baltimore, Maryland in 2022, which is incorporated herein by reference.

[0033] Model-based reinforcement learning methods in the era of deep learning are often based on the availability of approximate dynamics models that, among other important objectives, can also perform planning. Standard methods often require large amounts of training data to be useful. Some diffusion models demonstrate the advantages of tightly coupling the modeling and planning problems by training powerful diffusion models based on offline datasets. However, in many real-world applications (such as robotics), the environment includes geometric structures with symmetric forms. The symmetries of the environment are not explicitly exploited in current diffusion models. When standard models are used to approximate the dynamics of robotic locomotion or other tasks, large amounts of data are required for training before the model is ultimately useful. The systems and techniques described in this paper extend the use of diffusion models for planning under equivariant constraints based on the symmetries of the SE(3) group (a special Euclidean group with 3 elements), which leads to group-invariant density on trajectories, thereby improving sample efficiency, training speed, and model generalization.

[0034] In general, equivariance refers to taking into account the symmetries of the world (such as rotations, reflections and / or translations of state space) and exploiting these symmetries. One reason for exploiting symmetries is that the world is full of symmetries. The laws of physics are the same everywhere in space and time. For example, it is known that gravity causes objects to move up or down in a particular direction. In many cases, the laws of physics are symmetric under translations and rotations of spatial coordinates and under time shifts relative to an object (such as a robot or an object that can be manipulated by a robot). These symmetries also exist in many dynamic environments. For example, a robotic gripper can generally move an object from left to right in a similar way to how the gripper moves an object from top to bottom. Similarly, the navigation pattern of a quadrupedal device has nothing to do with whether it is moving east or north.

[0035] This paper describes systems and techniques for introducing planning algorithms that incorporate the symmetry structure of the environment. This inductive bias based on geometric symmetries improves sample efficiency and generalization performance. For example, by exploiting symmetries in the robot's movement or known physical properties of the environment, it is not necessary to approximate all the dynamics of the robot's movement, requiring less training data to generate the model. The benefit of incorporating known symmetries of the world is that it can improve the training model by requiring less data. The training process can also leverage prior knowledge of various symmetries of the environment to improve the training process.

[0036] In some aspects, the disclosed systems and techniques are based on the diffuser method as described herein. For example, the diffuser method can unify or solve the problem of learning a world model and the problem of using traditional model-based reinforcement learning (RL) to plan individual steps in the world model. The diffuser model method can be based on treating model planning as a generative modeling problem. Other methods train a diffusion model of state-action trajectories based on an offline dataset. By conditioning this diffusion model on an initial state and a final state, these systems and techniques can generate behaviors that cause an agent (e.g., a system or device, such as a vehicle, a robotic system or device, a drone, etc.) to enter another state from one state. In addition, these systems and techniques can generate samples that maximize any reward function by sampling from the model under the guidance of a classifier. For example, the classifier guidance can provide rewards at one or more states in a specific trajectory, and these rewards can be maximized to determine the optimal final trajectory for completing a task. A trajectory can refer to a series of states and actions that a system or device (e.g., a vehicle, a robotic device, a drone, etc.) may encounter.

[0037] The systems and techniques described herein can provide strong performance on long time horizon problems and high flexibility at test time. In some cases, the original diffuser method may still require a large amount of training data, possibly due to the lack of inductive bias about the symmetric structure of certain problems. In some aspects, these systems and techniques provide an equivariant diffuser that is trajectory-based. Planning algorithm for the invariant diffusion model. SE(3) can express the symmetry of spatial translation and rotation SE(3). represents the discrete time translation symmetry, and S n It can represent a permutation group on n objects. The invariance constraint can be enforced by an invariant basis density and an equivariant denoising network.

[0038] The solutions provided by the systems and techniques disclosed herein can be applied to different scenarios, such as offline reinforcement learning, model-based reinforcement learning, and trajectory generation. In some aspects, these systems and techniques can combine the principles of equivariant deep learning, diffusion models, and planning with deep learning. In some cases, these systems and techniques can pre-plan the entire trajectory (e.g., including a series or sequence of states and actions) of a system or device (such as a robotic system for picking up and moving objects). Various types of systems or devices can implement these systems and techniques, such as robotic arms or hands, quadrupeds or bipedal devices, robotic vacuum cleaners, vehicles (e.g., autonomous or semi-autonomous vehicles), drones, etc. These systems and techniques can be used to implement various tasks, such as object manipulation (e.g., picking up and placing objects, surgical procedures, etc.), movement or navigation of systems or devices, etc. As described above, a fundamental property of geometric structures is exploited, such as SE(3) symmetry, because the world (or robot or other environment) behaves similarly at different locations in space. In some aspects, there can be permutation symmetry, where a system or device (e.g., a robotic system) should behave the same when interacting with various similar objects.

[0039] Additional aspects of the disclosure are described in more detail below with reference to the accompanying drawings.

[0040] Figure 1 An example of an environment 100 in which a diffuser can be applied is shown. The diffuser can plan a route by iteratively refining a trajectory. As mentioned above, planning using a learned model is a fundamental framework for reinforcement learning and data-driven decision making. This approach employs learning techniques to approximate the unknown dynamics of the environment, which is equivalent to a supervised learning problem. The learned model can then be input into a classic trajectory optimization routine. However, because powerful trajectory optimizers utilize learned models, routes planned using such optimizers can be problematic, as the resulting routes may look more like adversarial examples than optimal trajectories. Standard model-based reinforcement learning algorithms can inherit more from model-free methods (such as value functions and policy gradients) than from trajectory optimization data. Reliance on online planning often uses gradient-free trajectory optimization routines, such as random shooting or cross-entropy methods, to avoid intractable problems. Standard methods for planning using diffusion for flexible behavior synthesis have been developed to address these issues, but these methods still require large amounts of data for training and do not generalize easily to new tasks or trajectories.

[0041] exist Figure 1 In FIG. 1 , environment 100 includes a structure 102 (eg, a robotic arm or other structure) having a receiving component 106 that can receive an object 104 (eg, a box), such as by grasping the object 104 . Figure 1The environment 100 illustrated in FIG. 1 provides an example of a structure 102 performing a task (e.g., picking up an object 104) that may follow one or more trajectories. For example, track points 110 are shown, collectively representing an original trajectory of structure 102. Each of track points 110 represents a state based on the action (e.g., the position of structure 102 based on a particular movement of structure 102). The original trajectory may have many different states based on a diffuse or noisy environment, in which structure 102 may move to many different possible points along the trajectory in order to receive object 104 (e.g., grasp object 104). Based on performing a denoising 108 process, the model may iteratively refine the trajectory such that, in a new state represented by track point 112, the trajectory becomes more focused or refined. In further iterations, the trajectory may be further refined, as shown by track point 114. The reverse process may result in diffusion 116, in which a more specific trajectory (e.g., the trajectory represented by track point 114 or the trajectory represented by track point 112) may diffuse and become random.

[0042] The diffuser method for data-driven trajectory optimization includes the core idea of ​​training a model that is directly applicable to trajectory optimization, in the sense that sampling from the model and planning with the model become almost identical. The learned dynamics model is typically used as a proxy for the dynamics of the environment. For diffusion models, the method considers how to design a model that is compatible with the planning problem for which the model will be used. As described above, the present disclosure introduces the idea of ​​solving problems associated with traditional diffusion models by using equivariant diffusion models. An example apparatus for using a diffusion model that exploits symmetries in a geometric structure may include at least one memory (e.g., a memory configured in a circuit, such as described below) Figure 13 ) and at least one processor (e.g., one or more of system memory 1315, memories 1320, 1325, and / or cache 1311). Figure 13 The processor 1312 is coupled to the at least one memory and is configured to: via the training preparation engine 704 (eg, described below Figure 7 The engine) receives a training dataset consisting of state-action pairs (e.g. Figure 7 702 of the training data set); separating the state-action pairs from the training data set into geometric data types via the training preparation engine; converting the geometric data types into internal representations via the training preparation engine; and performing the denoising via an equivariant denoising network (e.g., Figure 7 The equivariant diffuser 714 of the embodiment processes these internal representations to generate output data; and transforms the output data into a data representation. Further details about this apparatus and associated methods are described below.

[0043] Figure 2Blocks and components 200 are illustrated, showing various inputs 206 and outputs 210 of a diffuser model or diffuser 208. Trajectory-level diffusion probability models may be referred to as diffusers or diffuser models. While standard model-based planning techniques autoregressively predict forward in time, the diffuser 208 predicts all time steps of the plan simultaneously, as shown by output 210. The iterative sampling process of the diffuser 208 enables flexible conditioning, allowing the auxiliary guide to modify the sampling process to recover trajectories with high rewards or that satisfy a set of constraints. Denoising 202 Figure 2 204. The plan horizon 204 may progress from left to right in time. The diffuser 208 samples the plan by iteratively denoising 202 a two-dimensional array of a variable number of state-action pairs. The small receptive field 212 constrains the model to enforce local consistency only during a single denoising step. By combining many denoising steps together, local consistency can promote global coherence of samples of the plan horizon 204. The optional guidance function Can be used to bias plans toward those that optimize a test-time objective or satisfy a set of constraints.

[0044] The concept of data-driven trajectory optimization described above has various attractive properties. For example, to achieve long-term scalability, the diffuser 208 can be trained for the accuracy of the trajectories it generates rather than its single-step error. Therefore, the diffuser 208 is not affected by the compounding rollout error of the single-step dynamics model and can scale better with long planning horizons. To achieve task composability, the reward function provides auxiliary gradients to be used when sampling the plan, thereby achieving an intuitive planning approach by simultaneously combining multiple rewards by adding their gradients together. When the diffuser 208 generates globally coherent trajectories by iteratively improving local consistency, temporal composability can be achieved, allowing the diffuser 208 to generalize to new trajectories by piecing together subsequences within the distribution. Finally, efficient non-greedy planning is achieved by blurring the boundary between the model and the planner, and improving the training process of model predictions also has the effect of improving the planning ability of the diffuser 208. This design produces a learned planner that can solve long-term, sparse reward problems that have proven difficult for many conventional planning methods.

[0045] Figure 3A graphical representation 300 of the learned long-range planning characteristics of the diffusion planner or diffuser 208 is illustrated. As shown, three different states of a trajectory (such as a first state 304, a second state 316, and a third state 318) are associated with moving a robot or other device from a first point 306 to a destination 312. In the first state 304, there are many random states or positions (e.g., position 308) to which the device can move. These random states or positions can be generated by a training engine (e.g., Figure 7 314 ). A wall or obstacle 310 may require navigation to reach a destination 312 . Through the denoising process 302 , the system may iteratively move to a next state corresponding to a trajectory 314 shown in a more defined path, as illustrated in the second state 316 . The trajectory 314 is tighter and becomes less random. The third state 318 illustrates a more defined path or trajectory 317 relative to the trajectory 314 shown in the second state 316 . The system may include a neural network that performs the denoising process, wherein at each iteration of the neural network, the trajectory (e.g., trajectory 314 or 317) becomes less noisy until a clear trajectory 317 of a defined path is achieved, as shown in the third state 318 . The process may include calling the neural network and calling a reward function to obtain a series of states and actions that achieve the task while appropriately maximizing the reward.

[0046] The single-step model is typically used as a proxy for the true environmental dynamics f and is therefore not particularly tied to any planning algorithm. In contrast, the planning routine in the diffusion planning algorithm can be closely tied to the specific functional characteristics of the diffuser 208. Because the planning method is almost identical to sampling (the only difference being that it is guided by a perturbation or reward function r(τ)), the effectiveness of the diffuser as a long-term predictor directly translates into effective long-term planning. Figure 3 In the example shown, learned planning is beneficial in the goal-attainment scenario, demonstrating that the diffuser 208 is able to generate feasible trajectories in the type of sparse reward scenarios that are known to be difficult for shooting-based methods.

[0047] The denoising diffusion model, or diffuser 208, is designed for trajectory data and an associated probabilistic framework for behavior synthesis. While unconventional compared to the types of models typically used in deep model-based reinforcement learning, diffuser 208 has many useful properties and is particularly effective in offline control scenarios requiring long-term reasoning and flexibility when testing. This paper describes systems and techniques for improving diffuser 208 to further exploit symmetries associated with geometric structures, which are particularly applicable to robotics, although other environments are also contemplated as being within the scope of this application.

[0048] Figure 4Two sets of images 400 are provided, which show the forward diffusion process (which is fixed) and the reverse diffusion process (which is learned) of the diffusion model. Figure 4 As shown in the forward diffusion process of FIG, noise 403 is gradually added to the first set of images 402 at different time steps in a total of T time steps (e.g., forming a Markov chain), thereby generating a series of noisy samples X1 to X T .

[0049] From a training perspective, the diffusion model will take an image and will slowly add noise to the image to destroy the information in the image. In some aspects, the noise 403 is Gaussian noise. Each time step may correspond to Figure 4 Each successive image of the first set of images 402 is shown. Figure 4 The initial image X0 is an image of a cat. Noise 403 is added to each image (corresponding to the noisy samples X1 to X T ) results in a gradual diffusion of pixels in each image until the final image (corresponding to sample X T ) until the noise distribution is basically matched. For example, by adding noise, as the time step becomes larger, each data sample X1 to X T Gradually loses its distinguishable features, eventually leading to the final sample X T Equivalent to the target noise distribution, such as unit variance zero-centered Gaussian

[0050] The second set of images 404 shows the reverse diffusion process, where X T is the starting point of a noisy image (e.g., an image with Gaussian noise). The diffusion model can be trained to reverse the diffusion process (e.g., by training the model p θ -(x t-1 |x t )) to generate new data. In some aspects, the diffusion model can be trained by finding the inverse Markov transition that maximizes the likelihood of the training data. By traversing backward along the time step chain, the diffusion model can generate new data. For example, Figure 4 As shown, the inverse diffusion process is continued to generate X0 as an image of a cat. In other cases, the input data and output data may vary based on the task for which the diffusion model is trained.

[0051] As described above, the diffusion model is trained to denoise or restore the original image X0 in a progressive process, as shown in the second set of images 404. In some aspects, the neural network of the diffusion model can be trained to t-1 Restore X in case t , such as shown in the following example equation:

[0052]

[0053] The diffusion kernel can be defined as follows:

[0054] definition

[0055] Sampling can be defined as follows:

[0056]

[0057] In some cases, β t The value schedule (also called noise schedule) is designed so that and

[0058] The diffusion model is run in an iterative manner to progressively generate the input image X0. In some examples, the model may have twenty steps. However, in other examples, the number of steps may vary.

[0059] Figure 5 A diagram 500 illustrating planning using a diffusion model is illustrated. The Y or vertical axis may represent, from top to bottom, a denoising process 502, which begins with a noisy first set of images 506. Diagram 500 may represent a set of state-action pairs to which random noise may be added. As an example, first set of images 506 may also represent a series of physical objects (such as a robotic arm) and potential movements or joint angles of joints 507 within the robotic arm 509, represented by different points in first set of images 506.

[0060] The robot state can be specified by a set of joint angles. For example, one of the joint angles may represent a rotation of the base about the vertical z-axis. The angles may be represented as a ρ1 vector in the xy plane. Additionally, the direction of gravity (the z-axis itself) may be added as another ρ1 vector, which is also the normal direction of the surface (e.g., a table) on which one or more objects (e.g., n objects) lie. These vectors combined define the pose of the base of the robot arm. Rotating the direction of gravity and the robot and object poses according to SO(3) can be interpreted as a passive coordinate transformation, or as an active rotation of the entire scene (including gravity). SO(3) refers to a "special orthogonal" group with three dimensions. This solution provides effective symmetry because the laws of physics are invariant to the transformation.

[0061] n objects can be translated and rotated. Therefore, the poses of these objects can be determined by the translation relative to the reference pose and the rotation r∈SO(3). The translation is given as a vector by the global rotation g∈SO(3) represented by ρ1. The rotation pose is obtained by left-multiplying to transform. The SO(3) pose is not a Euclidean space, but a non-trivial manifold. Even if diffusion on the manifold is possible, the technique described in this paper simplifies the problem by embedding the pose in Euclidean space. The embedding is done by taking the first two columns of the pose rotation matrix r∈SO(3). These columns can each be transformed as a vector under the representation ρ1, forming an equivariant embedding ι: The resulting image is two orthogonal 3-vectors with unit norm. Through the Gram-Schmidt process, the equivariant mapping π: can be defined such that the equivariant map is the left inverse of the embedding: By combining with the translation, the rotational and translational pose of each object is thus embedded as three ρ1 vectors.

[0062] Noise can be added to the trajectory and then the denoising process is enabled. For example, the noise can be Gaussian noise. The X or horizontal axis from left to right can represent the planning horizon 504. The planning horizon 504 can be a specific amount of time. The second set of images 508 shows a more refined and less noisy positioning of the physical objects and the differences between them. The (final) third set of images 510 can show a denoised representation of the physical objects and the movement or action of the joints of the physical objects. Diagram 500 can illustrate a tighter coupling between modeling and planning, where the function f(τ) can represent a diffusion model and another function (such as r(τ)) can represent some auxiliary information, such as the inevitable result of the reward function described below, and the final function is the product of these two other functions, such as p(τ)=f(τ)r(τ). Here, τ is the trajectory or a series of states and actions. Figure 5 It illustrates how models can guide processes to produce valuable outputs or plans based on noisy inputs.

[0063] Incorporating symmetry structures into robot planning can have several aspects. First, the architecture can support various possible ways in which the system's properties behave under symmetry. Robot and object properties can generally be partitioned into scalar, vector, and quaternion representations of SE(3) (a special Euclidean group with 3 elements). The latter is used, for example, to parameterize the orientation of three-dimensional objects. The techniques described herein include the use of equivariant networks for scalar and vector representations, and introduce novel equivariant processing of quaternion data. In some aspects, training data can be labeled based on geometric data types (e.g., scalar, vector, quaternion), and these labels can be used to cause a neural network used to process the data to exploit the symmetries of the environment, and based on these labels, the neural network can maintain equivariance as described herein.

[0064] Figure 6is a diagram 600 illustrating an example of how symmetries in geometric structures can be exploited. The diffusion model used here is a specific type of neural network. The approach is to prescribe or pre-specify a mathematical group that describes how an object will transform under certain actions within its environment. For example, a ball in a three-dimensional environment can move on its axis or flip, or the ball can translate to a new position, and these actions are actions that can have some known valid symmetric consequences that can be exploited in the modeling process. The model can provide a "density" that provides the probability of states in space, and samples can be taken from this density, and the system can determine possible trajectories. If a given trajectory exists, a probability can be associated with that trajectory. A group-invariant density can mean that no matter what happens to the object, the likelihood or probability that the object actually appears should be the same.

[0065] The system can "bake" group invariant densities and associated probabilities into the model. Diagram 600 shows a robotic arm 604 and a table 602 that exhibit SE(3) symmetry. In a first system, the table 602 and the robotic arm 604 are shown in a coordinate system (such as an XYZ Cartesian coordinate system including an x-axis 610, a y-axis 608, and a z-axis 606). In the first system, the robotic arm 604 is shown along the Z-axis 606. The first system can be rotated so that the robotic arm 604 is configured or aligned along the X-axis 610 of the coordinate system. For all g in G, the group invariant density satisfies p(g*τ)=p(τ). A group can be all possible rotations or all possible reflections associated with the environment. In some examples, a group can also refer to a symmetric action in which something is expected to remain unchanged. A group is a mathematical concept in which there is a set of different rotations, such as a rotation that rotates a little, a rotation that rotates a lot, or any combination of rotations. Various combinations of groups can be in the SE(3) environment. In some cases, a sufficient condition is to construct a group invariant diffusion model f(τ) using a group equivariant neural network. The methods disclosed herein do not require the task to adhere to symmetry. The system can sample from non-invariant trajectory rewards r(τ). When this information is available, the system can indicate to the robot that trajectories that are reflections of each other are actually all the same when symmetry is applied, which can improve the efficiency of the system.

[0066] In some aspects, a robot can be trained based on data of a robotic arm moving a glass from the left side of a table to the right side of the table. The robot will typically not be able to perform new tasks, such as moving a glass from the right side of a table to the left side, or picking up a glass from the ground and lifting it onto the table. The methods herein take into account that there are symmetries between these movements (right to left or bottom to top) that correspond to trained movements that move a glass from left to right on the table. A neural network can exploit the symmetry of the world to know that it can move a glass from right to left or from the ground to the table based on training on a task that moves a glass from left to right but that has some symmetry to an additional task for which the robot was not specifically trained.

[0067] The second aspect is to properly support the splitting or breaking of symmetries during testing. Even when the environment exhibits symmetries, specific tasks often split or break these symmetries, such as when a robot must move an object to a specific point in space. The disclosed method allows for the splitting or breaking of symmetry groups during testing through task specification.

[0068] Next, we will further discuss the diffusion model. The diffusion model can include two processes. The first process (called the diffusion process) starts with a clean data sample x~q(x0) and gradually injects noise (for example, Gaussian noise) at each time step i∈[T] until the final step T (where the resulting sample consists of pure noise). For an example of the forward diffusion process, see Figure 4 The inverse generative process takes samples from a noise distribution and denoises them by gradually adding back structure until the image (e.g., or a robotics task) is restored to resemble samples from the empirical data distribution p(x).

[0069] Trajectory optimization using diffusion is another concept disclosed in this paper. Given the state s at time step t t and the action a taken at that time step t , can be affected by state s t+1 =f(s t ,a t ). Functions can be trained to perform different tasks in different states to perform the next action at the next time based on the observed conditions. The system uses the geometry of the environment to achieve a specific reward. The goal of trajectory optimization is to find the optimal solution that maximizes the objective (e.g., reward). A series of actions * 0:T , the goal is decomposed into a reward r(s) per time step t ,a t ). This technique corresponds to the following optimization problem shown in equation (1):

[0070]

[0071] Here, T is the planning horizon (e.g., the number of steps, such as one hundred steps), and the system uses the abbreviation τ = (s0, a0, ..., a sT ,a T ) to represent the trajectory. The "τ" value is the trajectory, and T is how long the trajectory is. The "a" value is then the action that the robot takes at a specific time or in a specific state, such as moving a certain amount or picking up an item. So, in state "0", the robot takes action "0". Based on this action, going through a new state "1", the robot may observe a new wall or a different environment because the robot has moved. Then, based on state "1", action "1" is taken. Chain graphs can be used as a modeling choice for modeling trajectories, but other methods can also be used. The "r" function in equation (1) is the reward. This reward can occur at various points along the robot's path. Equation (1) attempts to maximize a set of rewards provided along the path at each state and each action. At each state, a separate reward can be provided, so a * 0:T The value can be the maximum value among various paths along the trajectory that provide the highest reward. If the robot reaches the destination or successfully picks up the object, the system can provide a +1 reward. Figure 3 In the example of FIG, the goal is to move the robot from a first point 306 around obstacles 310 to a destination 312, which may be a final destination. The reward is successfully navigating to the destination 312.

[0072] A practical approach to solving trajectory optimization problems is to incorporate the planning process into the diffusion model. Specifically, the systems and techniques described in this paper decouple the learning of the approximate dynamics f and use offline data to train a powerful diffusion model. Planning is then treated as a conditional sampling problem when auxiliary information in the form of rewards is given. For example, these systems and techniques can learn a diffusion model p for the trajectory θ (τ), the diffusion model can be reused for different tasks in the following way as shown in Equation (2):

[0073]

[0074] Here, h(τ) can represent any prior evidence, a desired outcome (e.g., goal conditioning), or a general reward (which can be related to the reward function r(τ) discussed in this paper).

[0075] Trajectory representations can play the following role: Diffusion models provide information about the design space of models that can be exploited to solve trajectory optimization problems. In previous uses of diffusion models, each trajectory can be considered as an image with a single channel, where the width is the planning horizon T and each column in height corresponds to the concatenation of the state and action at a specific time step (s t ,a t ), the concatenation is flattened into a vector. Specifically, the input and output of the diffuser architecture are given by the following two-dimensional arrays:

[0076]

[0077] The benefit of considering this trajectory representation is that it makes it easy to apply diffusion models for images to RL scenarios. Furthermore, with a variable planning horizon, variable length trajectories can be generated by simply sampling noise of the same dimension.

[0078] The equivariant diffuser algorithm introduced in this paper includes the state-action trajectory τ=(s0,a0,...,s T ,a T ). The invariant diffusion model includes an invariant basis density and an equivariant denoising network. The network can be trained following the standard training algorithm for diffusion models: by adding noise to trajectories, feeding these trajectories into the denoising network, and training the network to predict the original trajectories (or equivalently, the added noise).

[0079] After training, the model can be used to sample trajectories unconditionally, conditionally on initial and goal states, and guided by a classifier to solve a task specified at test time.

[0080] Now refer to Figure 7 We discuss symmetry groups and representations in more detail, equivariant architectures of denoising networks, and unconditional and conditional sampling during network training and testing. Figure 7 700 is a flowchart illustrating how to use symmetry to train and exploit a network. Flowchart 700 generally relates to using a trajectory representation as a chain graph over time. However, as described elsewhere herein, using a chain graph is a modeling choice, and therefore other approaches may also be considered.

[0081] The training data set 702 may be prepared by separating the state or each state into a geometric data type and converting the quaternion into a rotation vector. The state may include the state of the robot arm angle, the state of the object, the state of the position of the object, etc. The training preparation engine 704 may perform these operations on the training data set 702. The symmetry group may be considered It is the product of three different groups: (1) the spatial translation and rotation symmetries SE(3), (2) the discrete time translation symmetry Z, and the permutation group S over n objects n The state may include, for example, the joint angles of a robot, or how an object appears, such as its position or orientation in the environment (represented by a quaternion). In some aspects, an object may have a color. The system may include data such as scalar values ​​that do not change with rotation or translation. In other words, color does not change with the movement or rotation of the object.

[0082] Symmetry groups can be segmented (e.g., soft-broken) in the environment. For example, a symmetry group can be segmented into at least one smaller group based on a certain condition. A segmentation of a symmetry group can also be characterized as a disruption of the symmetry group. For example, the direction of gravity typically segments or disrupts a spatial symmetry SE(3) into smaller groups SE(2), and object-destroying permutation groups can be distinguished. The fundamental principles of modeling invariance can be applied with respect to larger groups and include any symmetry-breaking or symmetry-segmenting effects as part of the data.

[0083] Any observable object that can be rotated under a symmetry can be characterized by how it transforms under the symmetry group. For example, adjustments can be made to the data representation to prepare it for further processing by the network. Spatial positions can be expressed relative to some key object (e.g., the position of the robot's base or center of mass). This structure guarantees equivariance with respect to spatial translations: to achieve SE(3) equivariance, the system only needs to design an SO(3) equivariant architecture. As mentioned earlier, SO(3) refers to a "special orthogonal" group with three dimensions.

[0084] Transformations under rotation (e.g., using SO(3) as a subgroup of SE(3) and including a three-dimensional rotation group) can be related to the disclosed architectural constructs. In some aspects, SO(3) is a composite operation around a three-dimensional Euclidean space. The method supports features in the following SO(3) notations: (1) Scalar: features s that are invariant under rotation R, s→ρ 平 Where (R)s=s (examples include the angle between two robot joints); (2) vector: features in the standard representation of SO(3), v→ρ1(R)v=Rv (examples include position or velocity vectors); and (3) quaternion: features transformed in the quaternion representation, where q R is the quaternion representation of the rotation matrix R, and is the Hamilton product of quaternions. Examples include object orientations. It can be assumed that all trajectories are transformed under the canonical representation of the time translation group.

[0085] Under the permutation group, the object properties are permuted while the robot properties or global properties of the state remain unchanged. Therefore, each property is either S n The trivial representation of , or its standard representation.

[0086] The training preparation engine 704 may further convert each state-action pair (s t ,a t ) separated or labeled as SO(3) scalar SO(3) vector and SO(3) quaternions These terms can be determined or characterized as labels that can be used by the equivariant network 706 for messaging between nodes. Different labels (e.g., scalars, vectors, quaternions) have different transformations, and the method is to transform the input data in an equivariant manner. For example, if the system applies a transformation to the input data type, the input data should be transformed in the same way if the system did not initially apply a transformation but applied a function and then applied a group transformation. Different transformation groups apply different laws, so each transformation group is processed differently by a neural network such as the equivariant network 706. Therefore, knowing the labels can tell the equivariant network 706 which part of the data to process based on the separated geometric data type. Here, the indices t, o, c indicate the trajectory time step, object, and channel, respectively. The value It can be used to represent global features that are not associated with any object (invariants under the permutation group). The channel index c distinguishes multiple features with the same representation in the dataset.

[0087] This transformation generates an internal representation of the data for use by the network. In order to construct equivariant networks that support these representations, it is useful to define new representations that are used internally by the network. Such features (1) Transformation in canonical representation under time shift; (2) Transformation in standard representation under permutation, i.e., w toc →w to′c =∑ o P o′o w toc ; and (3) direct and down-conversion of scalar and vector representations in SO(3):

[0088]

[0089] The denoising model f will have noisy input:

[0090]

[0091] and diffusion time step i is mapped to an estimate of the noise vector that produces:

[0092]

[0093] The denoising model f can perform this technique in at least three steps. First, the training preparation engine 704 transforms the data under its various representations into the internal representation introduced above. Next, these representations are processed using the equivariant network 706. The equivariant network 706 may include a geometric equivariant graph neural network (geometric EGNN). Other neural networks (such as those common to computer vision applications) can use additional geometric tags that are "baked" into the network. Models other than graph neural networks can also be used. In some examples, the trajectory can be represented as a chain graph over time. Other structures can also be used to represent the trajectory. The equivariant network 706 can pass equivariant messages on scalars, and can also pass equivariant messages on position and rotation vectors. Finally, the output data can be transformed from the internal representation to the data representation.

[0094] In some aspects, the training preparation engine 704 can add noise to the state-action pairs in the training dataset 702. The noise can be complete noise, Gaussian noise, or different levels of noise. The training preparation engine 704 can add a small amount of noise, a large amount of noise, different types of noise, and / or noise can be added to all items in the dataset. For example, in Figure 4 The second half of the generation process can in principle be applied to robot training. The system can start from state X T The system starts with a noisy robot behavior in X, and each time the neural network processes the input, the system can change from state X to state X. T Transitioning to a slightly less noisy version of the X4 to X0 output input. A neural network can perform this processing based on the known symmetries of the environment disclosed herein.

[0095] The equivariant network 706 can better exploit the symmetries of the geometry of objects or robots in the environment. Equivariant message passing can be used to pass messages between nodes of the neural network. By passing such messages in the context of a diffusion model (which can be a graph neural network), the neural network can have parts of the network that utilize message passing concepts to maintain constraints to keep the equivariant properties consistent so that the output is related to the input in some geometric way. In some cases, the unconditional generative model can generate unconditional trajectories 708 randomly in some aspects. At test time, the system can use an equivariant classifier 710 (which can be a function of r(τ) or can provide a reward) that indicates the possible reward to the system, which can then obtain a conditional trajectory. As Figure 7As further shown, an equivariant diffuser 714 can generate a conditional trajectory 716 based on output from the equivariant classifier 710 and / or the test-time reward engine 712, which is configured to execute a reward function (e.g., the r(τ) function described above) to improve the generation of a desired trajectory. The equivariant classifier 710 can provide conditional information for trajectory generation. The equivariant diffuser 714 can be related to the "f" function discussed above.

[0096] A geometrically equivariant neural network can include a trajectory represented as a chain graph over time, where the geometric graph includes nodes with bidirectional edges at each time step, where the nodes and features represent a concatenation of states and actions. The number of nodes in the graph can correspond to the planning horizon T. One benefit of this approach is that it enables the system to use information in a natural way. However, using a chain graph is a modeling choice regarding how to represent the trajectory. Different models can be implemented, and the geometrically equivariant GNN is a design choice, as other models are also available for neural networks.

[0097] Figure 8A 8 is a flow chart illustrating an example of a process 800 for operating an equivariant diffuser or for modeling a task using geometry. At block 802, process 800 includes receiving a training dataset comprising state-action pairs, for example, via training preparation engine 704. In some aspects, the spatial positions associated with these state-action pairs can be expressed relative to a key object. For example, the key object can include the location of the robot's base or the robot's center of mass.

[0098] At block 804, process 800 includes separating the state-action pairs from the training data set into geometric data types via the training preparation engine. In some aspects, the geometric data types may include one or more of scalars, vectors, and quaternions. When the geometric data types include quaternions, process 800 may include transforming a quaternion of the quaternions into at least two rotation vectors. In some aspects, transforming the quaternion of the quaternions into at least two rotation vectors may include mapping the quaternion to corresponding elements in a matrix representation and selecting two column vectors of the matrix representation as the two rotation vectors. When the data types include scalars, the scalars may be associated with objects corresponding to the state-action pairs.

[0099] In some aspects, separating the state-action pairs into geometric data types can include or can also be characterized as labeling the training data set with labels identifying the geometric data types. The neural network can then use these labels to process the data in an equivariant manner based on the different labels. Different parts of the neural network can be used to process the data based on the corresponding labels associated with the data, which can ensure that the data is processed by the neural network in an equivariant manner. This technique enables the system to determine trajectories for new tasks that the system may not have been trained for, but that are related to the symmetries of the geometric environment associated with the training data.

[0100] At block 806, the process includes converting the geometric data types into internal representations via the training preparation engine. In some aspects, the internal representations are associated with a symmetry group. In another aspect, the symmetry group can be a product of multiple different groups (e.g., three different groups). For example, the three different groups can include a spatial translation and rotational symmetry group, a discrete time translation symmetry group, and a permutation group on n objects. In some aspects, the symmetry group can be segmented (e.g., soft-broken) in the environment and segmented into at least one smaller symmetry group based on a condition. As examples, the condition can be the direction of gravity or the presence of a distinguishable object.

[0101] In some aspects, this gravity direction may break the spatial symmetry group SE(3) to a smaller group SE(2), and the objects may be distinguished from each other, breaking permutation invariance. This concept can follow from the fundamental principle of modeling invariance with respect to the larger group and including any symmetry breaking effects as input to the network.

[0102] In some aspects, the system may require that spatial positions are always expressed relative to some key object (e.g., the position of the robot's base or center of mass). This requirement guarantees equivariance with respect to spatial translations: to achieve SE(3 equivariance, the system only needs to be implemented as an SO(3) equivariant architecture.

[0103] In some aspects, the spatial translation and rotational symmetry group involves representations that include one or more of scalars, vectors, and quaternions. In some examples, a scalar may be invariant under a rotation associated with an angle between two objects. A vector may use a standard representation associated with position or velocity. A quaternion may represent a transformation as a quaternion associated with an orientation.

[0104] In some aspects, converting the geometric data types to the internal representations via the training preparation engine may include at least one of transforming a canonical representation under a time shift, transforming a canonical representation under a permutation, or transforming using scalar and vector representations.

[0105] At block 808, process 800 includes processing the internal representations via an equivariant denoising network to generate output data. In some aspects, the equivariant denoising network may include alternating types of layers. For example, the alternating types of layers may include one or more of temporal layers, permutation layers, and geometric layers. When the layers include temporal layers, the temporal layers may be one-dimensional convolutions along the trajectory step dimension. When the layers are permutation layers, the permutation layers may allow features belonging to different objects to interact. When the layers are geometric layers, the geometric layers may enable blending between scalar and vector quantities combined under the internal representation.

[0106] At block 810, process 800 includes transforming the output data into a data representation. In some aspects, transforming the output data into the data representation can be performed using a linear mapping. In other aspects, transforming the output data into the data representation using these linear mappings can include outputting a scalar for each input scalar, outputting a vector for each input vector, and outputting a scalar and a vector for each input quaternion.

[0107] In some cases, process 800 may further include generating an equivariant diffusion model by combining the invariant basis density and the equivariant denoising network. In some aspects, the equivariant diffusion model may be trained by adding noise to the state-action pairs to generate noisy trajectories, feeding the noisy trajectories into the equivariant denoising network, and using the equivariant diffusion model to output one or more predicted original trajectories for the state-action pairs.

[0108] Process 800 may also include using the equivariant diffusion model to unconditionally sample trajectories. Alternatively, process 800 may include using the equivariant diffusion model to conditionally sample trajectories based on an initial goal and state. As another alternative, process 800 may include using the equivariant diffusion model to sample trajectories under guidance from a classifier to solve a task. Sampling trajectories under guidance from the classifier to solve the task may be accomplished using test-time rewards and goal conditioning. In another aspect, sampling trajectories under guidance from the classifier to solve the task may include using rewards to specify a new task.

[0109] An apparatus for using a diffusion model that exploits symmetries in a geometric structure may include at least one memory (e.g., a memory configured in a circuit, such as Figure 13 ) and at least one processor (e.g., one or more of system memory 1315, memories 1320, 1325, and / or cache 1311). Figure 13 1312 of the processor), the at least one processor coupled to the at least one memory and configured to: via a training preparation engine (e.g., Figure 7 A training preparation engine 704) receives a training dataset (e.g., training dataset 702) comprising state-action pairs; separates these state-action pairs from the training dataset into geometric data types via the training preparation engine; converts these geometric data types into internal representations via the training preparation engine; processes these internal representations via an equivariant denoising network (e.g., equivariant diffuser 714) to generate output data; and transforms the output data into a data representation.

[0110] Figure 8B Another example process 820 for operating an equivariant diffuser or for modeling a task using geometry is illustrated. At block 822, the process 800 includes transforming an input vector into an intermediate vector using a one-dimensional convolution in the time dimension.

[0111] At block 824, process 820 includes, for each step, constructing all scalars (or inner products) from the intermediate vectors to generate derived scalars. At block 826, process 820 includes concatenating the input scalars with the derived scalars to generate concatenated scalars. At block 828, process 820 includes linearly transforming the concatenated scalars into a set of scalars. At block 830, process 820 includes applying a U-Net architecture (e.g., see

[15] ) to the diffuser architecture. Figure 9A U-Net architecture 900 or Figure 9B The alternative U-Net architecture 931 of FIG. 820 feeds the set of scalars to generate U-Net output scalars. At block 832, process 820 includes setting the U-Net output scalars to the first num_scalrs components of the U-Net output scalars. At block 834, process 820 includes processing the remaining components of the U-Net output scalars into two linear matrices acting on the input vectors.

[0112] In some examples, the processes described herein (e.g., processes 800 / 820 and / or any other processes described herein) can be performed by a computing device, apparatus, or system. In some examples, processes 800 / 820 can be performed by Figure 13800 / 820 and / or any other process described herein). In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to convey and / or receive data, any combination thereof and / or other components. The network interface may be configured to convey and / or receive data based on an Internet Protocol (IP) or other types of data.

[0113] A component of a computing device can be implemented in circuitry. For example, a component may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.

[0114] Process 800 / 820 is illustrated as a logical flow diagram, the operations of which represent a series of operations that can be implemented by hardware, computer instructions, or a combination thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally speaking, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the process.

[0115] Additionally, process 800 / 820 and / or any other process described herein may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed on one or more processors, through hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions that can be executed by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0116] Next, we discuss further details on transforming data into an internal representation. First, the input trajectory is transformed into an internal representation by the training preparation engine 704 in the following manner. Each input quaternion is transformed into two SO(3) vectors by mapping them to the corresponding SO(3) elements in the matrix representation and keeping the first two column vectors. Then, for each object o∈{1,...,n}, for each trajectory step t∈{1,...,T} and each channel c={1,...,n c}, the input is defined in internal representation as as follows:

[0117]

[0118] Here, v toc′ Contains vectors derived from quaternions. Matrix W 1,2,3,4 is learnable and is n×n c ×n s dimensional, n×n c ×n s dimensional, n×n c ×n s Dimensional and n×n c ×n s Here, n s is the number of scalar quantities associated with each object in the trajectory, n v is the number of vector quantities associated with each object, is the number of global scalar quantities, and is the number of global vector quantities. The number of channels n c is a hyperparameter. The number of channels should be chosen as Otherwise the network will generally not be able to model arbitrary denoising functions.

[0119] The second step involves isovariantly processing the internal representation in a geometric EGNN or equivariant network 706. The system then utilizes The equivariant denoising network processes the data. The equivariant denoising network consists of three alternating types of layers. Each type acts on the representation dimension of one of the three symmetry groups, while leaving the other two unchanged. One set of layers can include a temporal layer. The temporal layer consists of a one-dimensional convolution along the trajectory step dimension. These temporal layers can be organized as follows: Figure 9A In the U-Net architecture 900 shown, there is no mixing between features associated with different objects, nor between the four geometric features represented by the internal SO(3).

[0120] The temporal layer can be a time-shifted equivariant convolution along the temporal direction (e.g., along the trajectory steps). Figure 9A The U-Net architecture 900 is illustrated as having various layers 902. The U-Net architecture 900 is illustrated as having repeated blocks 904 (e.g., six repeatable blocks or residual blocks). Each block includes two temporal convolutions, each followed by group normalization (GN) and a final mish nonlinearity. The time step embedding is generated by a single fully connected layer 906 and is added to the activation value of the first time confusion within each block. In some cases, a diffuser (e.g., an equivariant diffuser 714) may have a U-Net architecture 900 having a residual block including temporal convolutions, a block for group normalization, and a block for mish nonlinearity. Mish is a self-regularizing non-monotonic activation function that can play a role in performance and training dynamics as well as neural networks. The U-Net architecture 900 can be geometric and the state-action trajectory can be viewed as an image. In some aspects, the network uses a linear combination of scalar quantities to maintain equivariance. As shown, a repetition block 904 may include a one-dimensional convolutional layer 908 receiving "x" data and a single fully connected layer 906 receiving "t" data, the output of which is combined using a combiner 910 to generate group normalization (GN) and mish data. The GN and mish data are processed by a one-dimensional convolutional layer 912 to generate additional GN and mish data. The additional GN and mish data are added to the "x" data using a combiner 914 to generate the output. The GN and mish data are not mixed between different objects, nor between the four geometric features of each internal representation.

[0121] Figure 9B An alternative U-Net architecture 931 is illustrated.

[0122] Another set of layers includes permutation layers. Permutation equivariant self-attention layers are used on the object dimension. These layers do not blend between different time steps, nor do they blend between the four geometric features of each internal representation. The permutation layers allow features belonging to different objects to interact. There is no blending between features associated with different time steps, nor between the four geometric features of the internal SO(3) representation. The permutation layers can be implemented using a self-attention mechanism.

[0123] Given input w toc , the system can use the learnable weight matrix W K,V,Q calculate:

[0124]

[0125] The computation according to Equation (6) is SO(3) equivariant since the scalar product in the attention weights computes the invariant SO(3) norm.

[0126] Another set of layers includes geometric layers, which handle SO(3) equivariant interactions between different geometric features within each internal representation. These geometric layers do not blend between different time steps, nor between different objects. The geometric layers enable blending between scalar and vector quantities combined under the internal representation, but not between objects or trajectory steps. As performed by the training preparation engine 704, the system first separates the input into SO(3) scalar and vector components, w toc =(s toc ,v toc ) T The engine then constructs all scalars that can be constructed for each object and time step:

[0127] S to ={s toc} c ∪{v toc ·v toc′} c,c′ (7)

[0128] These scalar and vector components are then used as two multi-layer perceptrons (MLPs) and ψ as inputs, and ultimately produces output scalars and vectors:

[0129] s′ toc =φ(S to ) c , (8)

[0130]

[0131] The next step involves mapping to an output representation. The equivariant network outputs the internal representation w toc The system then transforms these internal representations back into data representations using a linear map, similar to equation (5) above. In this way, the system outputs a scalar ∈^is for each input scalar t , output a vector ∈^ for each input vector i v t In addition, for each input quaternion, the system outputs a scalar h′i t and a vector h i t .

[0132] Finally, the quaternion output is calculated as in represents the Hamilton product between quaternions, and (h′ i t ,h i t ) is the real part h′ i t and the plural part h i t The quaternion of .

[0133] Figure 7 The trained model of the equivariant diffuser 714 is shown relative to The trained model can be executed to provide outputs of the model, and even if the model is only trained based on, for example, a certain movement of the robot, the model can exploit physical symmetries to execute or provide trajectories for other possible related movements of the robot. Figure 9A As shown in the single fully connected layer 906 in

[15] , equivariance with respect to temporal translation is induced by selecting a one-dimensional convolution along the temporal direction. Equivariance with respect to permutation is induced by the permutation equivariance of the self-attention block. SO(3) equivariance is similarly "baked" into the message passing operation. It is easy to see that constructing quaternions from scalars and vectors is also an SO(3) equivariant operation.

[0134] Furthermore, the network is equivariant under any sign flip of the input quaternion. This property is important when parameterizing object orientations with quaternions, since each object orientation R can be represented by two quaternions q R and -q R To describe.

[0135] The denoising model is trained based on the simplified loss as follows:

[0136]

[0137] Here, τ is the trajectory from the training data, i~Uniform(0,N) is the diffusion time step, and ∈ is Gaussian noise whose variance depends on i following the noise schedule.

[0138] like Figure 7 As shown, unconditional sampling or processing of unconditional trajectories 708 may occur. To sample from the unconditional trajectory density p(τ), trajectories τN~N(0,1) are first sampled from uncorrelated Gaussian noise. The data is then iteratively denoised as:

[0139]

[0140] where α i , α ˉ i , β i and σ i Depends on the noise schedule.

[0141] like Figure 7 As shown, the output from the equivariant classifier 710 can be provided to a test-time reward engine 712, which can implement target conditioning in the form of equivariant image pairing and provide specific new tasks via rewards. To maximize the reward for any given task, the system follows a number of procedures. For each task, the system trains a regression model J(τ) to estimate the cumulative reward achieved by the trajectory τ. The system then samples from the equivariant diffuser model under the guidance of the classifier from J(τ): Instead of using equation (12), the system can use the following average for denoising:

[0142]

[0143] The system can also handle conditional trajectories 716. The equivariant diffuser 714 allows the system to sample conditional constraints on states or actions. By conditioning on the initial state as described above and maximizing the expected cumulative reward achieved by the test-time reward engine 712, the system can solve reinforcement learning problems. Alternatively, by conditioning on the initial and final states, the system can solve goal-conditioned RL problems even without training a reward predictor.

[0144] Conditional sampling can be performed similarly to unconditional sampling, except that for conditional sampling, the system can fix the desired state and action after each denoising step to use them as conditional settings.

[0145] Figure 9B An example of a more detailed framework for the equivariant diffusion model 930 is illustrated. The architecture of the SE(3)×Z×Sn equivariant denoising network is shown in FIG. Figure 9BAs shown. The input trajectory 932 including features under different representations of the symmetry group can be transformed into a single internal representation via a representation mixer 934. The data is then processed using equivariant blocks 938, 944, 948, 952, 954, 958, 962, 966. In some aspects, the equivariant blocks 938, 944, 948, 952, 954, 958, 962, 966 may include convolutional layers along the time dimension, attention on objects, normalization layers, and geometry layers that mix the scalar and vector components of the internal representation. The equivariant blocks can be combined into an alternative U-Net architecture 931. Conditioning information and diffusion time 941 can be fed into the pipeline from the context block 940 via the context block 940 at each level of the alternative U-Net architecture 931. In some cases, the context block also includes context blocks 950, 956, and 964. The second internal representation of the output 968 can be separated into the original data representation. For simplicity, Figure 9B Some details are not included (including residual connections, downsampling and upsampling layers).

[0146] The equivariant diffusion model 930 can be characterized as Invariant diffusion model. The equivariant diffusion model 930 can be a base density that is invariant with respect to the symmetry group and a denoising model that is equivariant with respect to the symmetry group. This diffusion model then has Invariant probability distribution.

[0147] As a basis density, in some aspects, the method can use a multidimensional standard normal distribution, which has the desired invariance properties. The novel equivariant architecture for the denoising model f can be implemented as a neural network and maps the noisy input trajectory τ and diffusion time steps i to an estimate of the noise vector that generates the input In some aspects, equivariant diffusion model 930 can achieve this result in at least three steps. For example, in a first step, input trajectory 932, including various representations, can be transformed into an internal representation of a symmetry group. In a second step, the data can be processed using an equivariant network within this representation. In a third step, the output can be transformed from the internal representation to the original representation present in the trajectory. Figure 9B The architecture of the equivariant diffusion model 930 is illustrated.

[0148] In some cases, according to a first step, representation mixer 934 may receive input trace 932. The noisy trace input may include features under different representations of the symmetry group. While these input representations can be mirrored for the hidden states of the neural network, the design of the equivariant architecture is greatly simplified if all inputs and outputs are transformed under a single representation. Thus, the methods described herein can decouple the data representation from the representation used internally for computation.

[0149] A single internal representation of the data is introduced936. For each trajectory time step t∈{1,...,H}, for each object i∈{1,...,n}, for each channel c∈{1,...,n c},, the single internal representation 936 may include an SO(3) scalar s toc and a SO(3) vector v toc In some aspects, the dimensions of the internal representation 936 can be channels, objects, and time, such as Figure 9B In some cases, a scalar may be paired with a vector. In other cases, such as for systems where scalar or vector quantities play a larger role, the approach may be to use multiple copies of either representation. The system may be written Under spatial rotation g∈SO(3), these features can thus be transformed into the direct sum of scalar and vector representations

[0150]

[0151] These internal features can be transformed in the canonical representation under time shift and in the standard representation under permutation to Therefore, there are no global (non-object specific) features underlying the internal representation.An example of embedding robot features into the representation described above is described below.

[0152] An example of transforming an input representation into an internal representation will now be described. For example, a first layer in a neural network may transform an input trace 932 (e.g., comprised of The system can pair SO(3) scalars and SO(3) vectors as Features. The system can also remove global features that are not assigned to one of the n objects in the scene by including these global features in the representation of each of the n objects.

[0153] In some examples, for each object o∈{1,…,n}, each trajectory step t∈{1,…,T} and each channel c={1,…,n n}, the system can define the input in internal representation as follows

[0154]

[0155] Matrix W 1,2,3,4 is learnable and its dimensions are or here, is the number of SO(3) scalar quantities associated with each object in the trajectory, is the number of SO(3) vector quantities associated with each object, is the number of scalar quantities associated with the robotic or global properties of the system, and is the number of vectors of this nature. The number of input channels n c Is a hyperparameter. The system can initialize the matrix W i , so that Equation (15) corresponds to the concatenation of all object-specific features and global features along the channel axis at the beginning of training, but making these matrices learnable. Note that Equation (15) is similar to Equation (5) above.

[0156] In some cases, according to the second step, the system may apply Equivariant U-Net. Then, the system can use The equivariant denoising network or equivariant block 938 processes the data. The components of the denoising network include three alternating types of layers. Each type of layer can act on the representation dimension of one of the three symmetry groups while leaving the other two symmetry groups unchanged. In some aspects, the equivariant block 938 includes a temporal layer 974. The temporal layer 974 can include time-shifted equivariant convolutions along the time direction (e.g., along the trajectory steps) and is organized in the alternative U-Net architecture 931. The temporal layer 974 may not mix between different objects or between the four geometric features of each internal representation. The equivariant block 938 may also include object layers. For example, the permutation layer 976 can be referred to as a permuted equivariant self-attention layer on the object dimension. These layers may not mix between different time steps or between the four geometric features of each internal representation. The equivariant block 938 may also include a normalization layer 978. The equivariant block 938 may also include a geometric layer 980, such as an SO(3) equivariant interaction between scalar and vector features within each internal representation. The geometry layer 980 may not be blended between different time steps, nor between different objects.

[0157] In some aspects, the system may use residual / skip connections and a new type of normalization layer 978 that does not violate equivariance. The system may also include a context block 940 that processes conditioning information and embeds it into the internal representation. To perform these functions, the context block 940 may include a mish module 942, a linear module 943, and an embedding module 945. The mish module 942 involves a self-regularized non-monotonic activation function that can play a role in the performance and training dynamics of neural networks.

[0158] The layers described above can be combined into equivariant blocks comprising one instance of each layer, and these equivariant blocks are arranged in an alternative U-Net architecture 931, such as Figure 9BBetween the layers of the alternative U-Net architecture 931, the system can downsample (or upsample) along the trajectory time dimension (e.g., by a factor of 2), increasing (or decreasing) the number of channels accordingly.

[0159] The temporal layer 974 includes one-dimensional convolutions along the temporal dimension of the trajectory. To maintain SO(3) equivariance, these convolutions may not add any bias. In some cases, there is no mixing between features associated with different objects, nor between the four geometric features represented by the internal SO(3) representation.

[0160] The permutation layer 976 allows features belonging to different objects to interact via the equivariant self-attention layer. In some cases, there is no blending between features associated with different time steps, nor between the four geometric features represented by the internal SO(3).

[0161] Given input w toc , the permutation layer uses a learnable weight matrix W K,V,Q calculate:

[0162]

[0163] The output of Equation (16) is SO(3) equivariant because the scalar product in the attention weights computes the invariant SO(3) norm.

[0164] The geometric layers 980 are the third type of layer. These layers enable mixing between scalar and vector quantities that are combined in the internal representation, but not between different objects or across the time dimension. The system constructs an equivariant mapping between scalar and vector inputs and outputs: the system first separates the input into SO(3) scalar and vector components w toc =(s toc ,v toc ) T , and then construct a complete set of SO(3) invariants S by combining scalars and pairwise inner products between vectors to =(s toc ) c ∪{v toc ·v toc′} c,c′ These scalar and vector components are then used as two The system can then generate output scalars and vectors as follows:

[0165]

[0166] The system can approximate any equivariant mapping between SO(3) scalars and vectors under mild assumptions. However, in its original form, the system can become prohibitively expensive because the SO(3) invariant Sto The number of s scales quadratically with the number of channels. Therefore, the system may first linearly transform the input vector into a smaller number of vectors, apply the transformation, and increase the number of channels again using another linear transformation.

[0167] The isovariant block 938 may also be included as isovariant blocks 944, 948, 952, 954, 958, 962, and 966. For example, the output of isovariant block 944 is a second internal representation of the data 946. The output of isovariant block 958 is a third internal representation of the data 960, which is input to another isovariant block 962. The output of isovariant block 966 is a second internal representation of the output 968.

[0168] The next step may include representing a second internal representation of the demixer 970 processed output 968, which may be the internal representation w toc The system can then use a linear mapping to transform these internal representations w toc Transform back to data representation. Global properties, such as robot degrees of freedom, are aggregated from the object-specific internal representation by taking the average, minimum, and maximum values ​​across the object. These three aggregates are then concatenated along the channel dimension. Output trajectory 972 is shown as the output of equivariant diffusion model 930. It may be beneficial to apply additional geometric layers to these aggregated global features before separating them into their original representation.

[0169] In some respects, Figure 9B The diffusion model of can be trained based on offline trajectory data {τ}, which does not need to include reward information. The system can use a simplified diffusion loss Here, τ is the trajectory from the training data, i~Uniform(0,N) is the diffusion time step, and ∈ is Gaussian noise whose variance depends on i following the noise schedule.

[0170] The system uses a diffusion model trained on offline trajectory data to jointly learn the world model and policy. The system can use this diffusion model to solve planning problems (e.g., choosing a sequence of actions to maximize expected reward).

[0171] In some aspects, the system may utilize equivariant diffusion (such as using various features of the diffusion model) to perform planning. For example, the system may utilize the ability to sample from the diffusion model by extracting noisy data from the underlying distribution and iteratively denoising the diffusion model using a learning network, thereby providing the system with trajectories similar to those in the training set. In order for such sampled trajectories to be useful for planning, the sampled trajectories may start at the current state of the environment. For example, the sampling process may be conditioned so that the initial state of the generated trajectory matches the current state. The system may guide the sampling process to solve a specific task specified at test time, which may include training a regression model to map trajectories to rewards under a given task. The sampling iterations may then be biased towards trajectories with high rewards. Combining these elements, the system may use conditional sampling guided by a reward model in a closed sampling loop.

[0172] The system can also exploit the concept of symmetry breaking. By construction, the equivariant diffusion model learns the Invariant density. Unconditional samples will reflect the symmetry property: the likelihood of sampling a trajectory and its rotated or permuted counterpart will be equal.

[0173] The table below illustrates the performance of the Kuka robot on a navigation task and a block stacking problem.

[0174]

[0175] Specific tasks can violate the invariants discussed above. For example, they can be violated by requiring the robot or object to be brought to a specific location. The equivariant diffuser method allows the system to gracefully break symmetries for specific tasks during testing. This soft symmetry breaking can occur through conditionalization, such as by specifying the initial or final state of the sampled trajectories, or by using a non-invariant reward model for guidance during sampling.

[0176] Each of the geometric quantities can be processed individually to construct an equivariant mapping under a prescribed group G (eg, G=SE(3)).

[0177] Figure 10 Various data 1000 are illustrated, which may include scalar quantities that may be associated with angles and torques 1002. Torque, for the purposes of this document, may correspond to the angular force applied between joint coordinates and may be expressed in radians. Thus, as a geometric quantity, torque is invariant to any action on g∈SE(3) and may be considered a scalar quantity.

[0178] Other data 1004 can be represented as vectors and can be associated with positions. In the observation space of the example robot stacking environment, the task may include four cubes whose centers are given as positions in three-dimensional space. The position vectors are transformed in the standard representation under the action of g∈SE(3).

[0179] Additional data 1006 may include quaternions. Quaternions themselves are difficult to reason about using geometric types. A method can be used to convert a quaternion into a 3×3 rotation matrix. The matrix itself can be viewed as a flattened 9-dimensional vector that carries three copies of the standard representation. The other dimension of each cube can correspond to a binary variable corresponding to whether the cube is attached to the robot's arm. From a geometric perspective, a scalar quantity can be used that is invariant to the action of SE(3) and can be related to the attachment state.

[0180] Figure 11 1 is a diagram illustrating an example of equivariant trajectory generation according to aspects of the present disclosure. A first set of noisy data 1102 in a Cartesian coordinate system can be denoised using τ ≈ p(τ) as shown, and the resulting denoised trajectory 1104 in the same Cartesian coordinate system can be shown. The first set of noisy data 1102 can also be shown as having a certain symmetry transformation via p1(g) to transformed data 1106 with a new symmetric position. The transformed data 1106 can then also be denoised via τ ≈ p(τ) to generate an output 1108 in the same Cartesian coordinate system.

[0181] Figure 12A 12 is a diagram illustrating an example use in conjunction with robotics according to various aspects of the present disclosure. For example, a robot may be performing an extraction operation, with the robot having a first state 1202 and a second state 1204 (wherein the robot is performing the action). Robotics learning practitioners may use the concepts disclosed herein to construct data-efficient plans, where symmetry may be helpful. In some aspects, textual temporal task specifications may be used to partition symmetry groups or break equivariance, which also provides some flexibility to the modeler. The architecture provides a general solution for many geometric data types, enabling further use of these geometric data types in different environments, which may include different robots or other types of environments.

[0182] Figure 12B FIG1210 is a diagram illustrating an example use related to a robot according to aspects of the present disclosure, showing a trajectory generated by an equivariant diffuser model. For example, as part of a trajectory or plan, the robot can perform different tasks or have different states 1212, 1214, 1216, 1218, 1220, 1222, 1224, 1226 to move an object from one place to another or perform any type of task.

[0183] Figure 13 is a diagram illustrating an example of a system for implementing certain aspects of the present technology. Specifically, Figure 13 An example of a computing system 1300 is illustrated, which can be any computing device, for example, constituting an internal computing system, a remote computing system, a camera, or any component thereof, wherein the components of the system communicate with each other using a connection 1305. Connection 1305 can be a physical connection using a bus, or a direct connection to processor 1310, such as in a chipset architecture. Connection 1305 can also be a virtual connection, a networked connection, or a logical connection.

[0184] In some aspects, computing system 1300 is a distributed system, wherein the functionality described in this disclosure may be distributed within a data center, multiple data centers, a peer-to-peer network, etc. In some aspects, one or more of the described system components represent a plurality of such components, each of which performs some or all of the functionality of the described component. In some aspects, each component may be a physical or virtual device.

[0185] The example computing system 1300 includes at least one processing unit (CPU or processor), which may be characterized as a processor 1310, and connections 1305 coupling various system components including system memory 1315, such as read-only memory (ROM) memory 1320 and random access memory (RAM) memory 1325, to the processor 1310. The computing system 1300 may include a cache 1311 of high-speed memory directly connected to, proximate to, or integrated as part of the processor 1310.

[0186] Processor 1310 may include any general-purpose processor and hardware or software services (such as services 1332, 1334, and 1336 configured to control processor 1310, stored in storage device 1330), as well as a dedicated processor in which software instructions are incorporated into the actual processor design. Processor 1310 may essentially be a completely independent computing system containing multiple cores or processors, a bus, a memory controller, a cache, etc. Multi-core processors may be symmetric or asymmetric.

[0187] To enable user interaction, computing system 1300 includes input device 1345, which can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. Computing system 1300 may also include output device 1335, which can be one or more of a plurality of output mechanisms. In some cases, a multimodal system can enable a user to provide multiple types of input / output to communicate with computing system 1300. Computing system 1300 may include a communication interface 1340, which generally governs and manages user input and system output.

[0188] The communication interface can perform or facilitate receiving and / or sending wired or wireless communications using wired and / or wireless transceivers, including using audio jacks / plugs, microphone jacks / plugs, universal serial bus (USB) ports / plugs, Ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, dedicated wired ports / plugs, Wireless signal transmission, Low energy (BLE) wireless signal transmission, Wireless signal transmission, radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, WLAN signal transmission, visible light communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), infrared (IR) communication wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / long term evolution (LTE) cellular data network wireless signal transmission, self-organizing network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof.

[0189] The communication interface 1340 may also include one or more GNSS receivers or transceivers for determining the location of the computing system 1300 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States' Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There is no limitation to operating on any particular hardware arrangement, and thus the underlying features herein may be readily substituted for improved hardware or firmware arrangements as they are developed.

[0190] The storage device 1330 may be a non-volatile and / or non-transitory and / or computer-readable memory device and may be a hard disk or other type of computer-readable medium that can store data accessible by a computer, such as a magnetic cassette, a flash memory card, a solid-state memory device, a digital versatile disk, a cassette, a floppy disk, a floppy disk, a hard disk, a magnetic tape, a magnetic stripe / magnetic stripe, any other magnetic storage medium, a flash memory, a memristor memory, any other solid-state memory, a compact disc read-only memory (CD-ROM) disc, a rewritable compact disc (CD) disc, a digital video disc (DVD) disc, a Blu-ray disc (BDD) disc, a holographic disc, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a memory stick card, a smart card chip, a Europay, MasterCard, and Visa (EMV) chip, a Subscriber Identity Module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, RAM, static RAM (SRAM), dynamic RAM (DRAM), ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASH EPROM), cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase change memory (PCM), spin transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or a combination thereof.

[0191] Storage devices 1330 may include software services, servers, services, etc., which, when the code defining such software is executed by processor 1310, causes the system to perform functions. In some aspects, hardware services that perform specific functions may include software components for performing functions stored in a computer-readable medium connected to the necessary hardware components (such as processor 1310, connection 1305, output device 1335, etc.). The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data may be stored and does not include carrier waves and / or transient electronic signals propagated wirelessly or over a wired connection.

[0192] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transient media in which data may be stored and does not include carrier waves and / or transient electronic signals that are propagated wirelessly or over a wired connection. Examples of non-transient media may include, but are not limited to, disks or tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, memory, or storage devices. Computer-readable media may store thereon code and / or machine-executable instructions that may represent any combination of a procedure, function, subroutine, program, routine, subroutine, module, engine, software package, class, or instruction, data structure, or program statement. A code segment may be coupled to another code segment or hardware circuit by passing and / or receiving information, data, independent variables, parameters, or memory contents. Information, independent variables, parameters, data, etc. may be passed, forwarded, or sent via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.

[0193] In some aspects, computer-readable storage devices, media, and memories may include wired or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media specifically excludes media such as power consumption, carrier signals, electromagnetic waves, and signals themselves.

[0194] Specific details are provided in the above description to provide a detailed understanding of the aspects and examples provided herein. However, it will be understood by those skilled in the art that these aspects can be practiced without these specific details. For clarity of explanation, in some cases, the present technology can be presented as including a separate functional block, which includes a device, device assembly, step or routine in a method embodied in a combination of software or hardware and software. Additional components other than those components shown in the accompanying drawings and / or described herein can be used. For example, circuits, systems, networks, processes and other components can be shown as components in block diagram form to avoid confusing these aspects in unnecessary details. In other instances, known circuits, processes, algorithms, structures and techniques can be shown without unnecessary details to avoid confusing various aspects.

[0195] Various aspects may be described above as processes or methods, which may be depicted as flow charts, flowcharts, data flow diagrams, structure diagrams, or block diagrams. Although a flow chart may describe operations as a sequential process, many of the operations may be performed in parallel or concurrently. Furthermore, the order of the operations may be rearranged. A process is terminated when its operations are completed, but a process may have additional steps not included in the accompanying figures. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, termination of the process may correspond to the function returning to the calling function or main function.

[0196] The processes and methods according to the examples described above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtained from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions. Portions of the computer resources used may be accessible over a network. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, source code. Examples of computer-readable media that can be used to store instructions, information used, and / or information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, and the like.

[0197] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein may also be embodied in peripheral devices or add-in cards. By way of further example, such functionality may also be implemented on circuit boards among different chips or different processes executed in a single device.

[0198] Instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.

[0199] In the foregoing description, various aspects of the present application are described with reference to specific aspects thereof, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although the exemplary aspects of the present application have been described in detail herein, it is to be understood that each inventive concept can be implemented and adopted in various other ways, and the appended claims are not intended to be interpreted as including these variations, unless limited by the prior art. The various features and aspects of the application described above can be used individually or in combination. In addition, the various aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader essence and scope of this specification. Therefore, the description and the accompanying drawings should be considered as illustrative rather than restrictive. For illustrative purposes, each method is described in a specific order. It should be appreciated that, in alternative aspects, each method can be performed in a different order than described.

[0200] One of ordinary skill in the art will appreciate that the less than ("<") and greater than (">") symbols or terms used herein may be replaced by less than or equal to ("≤") and greater than or equal to ("≥") symbols, respectively, without departing from the scope of the description.

[0201] Where a component is described as being “configured to” perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., a microprocessor or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0202] The phrase “coupled to” refers to any component being physically connected directly or indirectly to another component, and / or any component being in communication directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0203] Claim language or other language in this disclosure that recites “at least one of” a set and / or “one or more of” a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, claim language reciting “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, or A, B, and C. The language “at least one of” a set and / or “one or more of” a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” or “at least one of A or B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0204] Claim language that recites "at least one processor configured to," "at least one processor configured to," "the processor is configured to," or other language indicates that one processor or multiple processors (in any combination) can perform the associated operations. For example, claim language that recites "at least one processor configured to: X, Y, and Z" means that a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each tasked with a particular subset of operations X, Y, and Z such that the multiple processors together perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language that recites "at least one processor configured to: X, Y, and Z" can mean that any single processor can perform only at least a subset of operations X, Y, and Z.

[0205] The various illustrative logic blocks, modules, engines, circuits, and algorithmic steps described in conjunction with the illustrative examples disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. In order to clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been generally described in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the entire system. Technicians may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be interpreted as departing from the scope of this application.

[0206] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices, or integrated circuit devices with multiple uses, including applications in wireless communication devices and other devices. Any features described as modules, engines, or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be implemented at least in part by a computer-readable data storage medium comprising program code, which includes instructions that, when executed, perform one or more of the above methods, algorithms, and / or operations. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as a random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic or optical data storage medium, or the like. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.

[0207] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Thus, the term "processor," as used herein, may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein.

[0208] Illustrative aspects of the present disclosure include:

[0209] Aspect 1. A processor-implemented method for modeling a task using a geometric structure, the processor-implemented method comprising: receiving a training dataset comprising state-action pairs via a training preparation engine; separating the state-action pairs from the training dataset into geometric data types via the training preparation engine; converting the geometric data types into an internal representation via the training preparation engine; processing the internal representation via an equivariant denoising network to generate output data; and transforming the output data into a data representation.

[0210] Aspect 2. The processor-implemented method of aspect 1, wherein the equivariant denoising network comprises layers of alternating types.

[0211] Aspect 3. The processor-implemented method of aspect 2, wherein the alternating types of layers include temporal layers, permutation layers, and geometric layers.

[0212] Aspect 4. The processor-implemented method of aspect 3, wherein the temporal layer comprises a one-dimensional convolution along a trajectory stride dimension.

[0213] Clause 5. The processor-implemented method of any one of clauses 3 or 4, wherein the permutation layer allows features belonging to different objects to interact.

[0214] Clause 6. The processor-implemented method of any one of clauses 3 to 5, wherein the geometry layer enables blending between scalar and vector quantities combined under the internal representation.

[0215] Aspect 7. The processor-implemented method according to any one of aspects 1 to 6, wherein the geometric data types include scalars, vectors, and quaternions.

[0216] Aspect 8. The processor-implemented method according to aspect 7, further comprising: transforming the quaternion into at least two rotation vectors.

[0217] Aspect 9. A processor-implemented method according to aspect 8, wherein transforming the quaternion into at least two rotation vectors comprises: mapping the quaternion to corresponding elements under a matrix representation; and selecting two column vectors of the matrix representation as the two rotation vectors.

[0218] Clause 10. The processor-implemented method of any one of clauses 8 or 9, wherein the scalar is associated with an object corresponding to the state-action pair.

[0219] Aspect 11. The processor-implemented method of any one of aspects 1 to 10, wherein transforming the output data into the data representation is performed using a linear mapping.

[0220] Aspect 12. A processor-implemented method according to Aspect 11, wherein using the linear mapping to transform the output data into the data representation includes: outputting a scalar for each input scalar, outputting a vector for each input vector, and outputting a scalar and a vector for each input quaternion.

[0221] Aspect 13. The processor-implemented method according to any one of aspects 1 to 12, further comprising: generating an equivariant diffusion model by combining an invariant basis density and the equivariant denoising network.

[0222] Aspect 14. The processor-implemented method of aspect 13, wherein the equivariant diffusion model is trained by adding noise to the state-action pairs to generate noisy trajectories, feeding the noisy trajectories into the equivariant denoising network, and using the equivariant diffusion model to output one or more predicted original trajectories of the state-action pairs.

[0223] Aspect 15. The processor-implemented method of aspect 14, further comprising: unconditionally sampling trajectories using the equivariant diffusion model.

[0224] Aspect 16. The processor-implemented method of aspect 14, further comprising: using the equivariant diffusion model to conditionally sample trajectories based on an initial target and a state.

[0225] Aspect 17. The processor-implemented method according to any one of aspects 14 to 16, further comprising: using the equivariant diffusion model to sample trajectories under guidance from a classifier to solve a task.

[0226] Clause 18. The processor-implemented method of clause 17, wherein sampling trajectories to solve the task under guidance from the classifier comprises using test-time rewards and goal conditioning.

[0227] Clause 19. The processor-implemented method of clause 17, wherein sampling trajectories to solve the task under guidance from the classifier comprises: using a reward to specify a new task.

[0228] Aspect 20. The processor-implemented method of any one of aspects 1 to 19, wherein the internal representation is associated with a symmetry group.

[0229] Aspect 21. The processor-implemented method of aspect 20, wherein the symmetric group is a product of three different groups.

[0230] Aspect 22. The processor-implemented method of aspect 21, wherein the three different groups include a group of spatial translation and rotational symmetries, a group of discrete time translation symmetries, and a group of permutation groups over n objects.

[0231] Aspect 23. The processor-implemented method of aspect 22, wherein the permutation group over n objects is associated with object properties that are permuted while keeping the robot properties or global properties of the state unchanged.

[0232] Clause 24. The processor-implemented method of any one of clauses 22 or 23, wherein the spatial translation and rotational symmetry groups involve representations including scalars, vectors, and quaternions.

[0233] Aspect 25. A processor-implemented method according to Aspect 24, wherein the scalar remains invariant under a rotation associated with an angle between two objects, wherein the vector employs a standard representation associated with position or velocity, and wherein the quaternion is transformed under a quaternion representation associated with orientation.

[0234] Aspect 26. The processor-implemented method of aspect 20, wherein the symmetry group is partitioned into at least one smaller symmetry group based on a condition.

[0235] Clause 27. The processor-implemented method of clause 26, wherein the condition comprises at least one of a direction of gravity or the presence of a distinguishable object.

[0236] Clause 28. The processor-implemented method of any one of clauses 1 to 27, wherein the spatial position associated with the state-action pair is expressed relative to a key object.

[0237] Clause 29. The processor-implemented method of clause 28, wherein the key object comprises the position of a base of the robot or a center of mass of the robot.

[0238] Aspect 30. A processor-implemented method according to any one of Aspects 1 to 29, wherein converting the geometric data type to the internal representation via the training preparation engine includes at least one of the following operations: transforming the canonical representation under time shift, transforming the canonical representation under permutation, or transforming using scalar and vector representations.

[0239] Aspect 31. A device for using a diffusion model that exploits symmetries in geometric structures, the device comprising: at least one memory; and at least one processor, the at least one processor being coupled to the at least one memory and configured to: receive a training data set comprising state-action pairs via a training preparation engine; separate the state-action pairs from the training data set into geometric data types via the training preparation engine; convert the geometric data types into an internal representation via the training preparation engine; process the internal representation via an equivariant denoising network to generate output data; and transform the output data into a data representation.

[0240] Clause 32. The apparatus according to clause 31, wherein the equivariant denoising network comprises layers of alternating types.

[0241] Aspect 33. The apparatus of aspect 32, wherein the alternating types of layers include temporal layers, permutation layers, and geometric layers.

[0242] Aspect 34. The apparatus of aspect 33, wherein the temporal layer comprises a one-dimensional convolution along a trajectory stride dimension.

[0243] Clause 35. The apparatus according to any one of clauses 33 or 34, wherein the permutation layer allows features belonging to different objects to interact.

[0244] Aspect 36. The apparatus according to any one of aspects 33 to 35, wherein the geometry layer enables blending between scalar quantities and vector quantities combined under the internal representation.

[0245] Aspect 37. The apparatus according to any one of aspects 31 to 36, wherein the geometric data types include scalars, vectors, and quaternions.

[0246] Aspect 38. The apparatus according to aspect 37, wherein the at least one processor is configured to: transform a quaternion in the quaternion into at least two rotation vectors.

[0247] Aspect 39. An apparatus according to Aspect 38, wherein, in order to transform the quaternion into at least two rotation vectors, the at least one processor is configured to: map the quaternion to corresponding elements under a matrix representation; and select two column vectors of the matrix representation as the two rotation vectors.

[0248] Aspect 40. An apparatus according to any one of aspects 37 or 28, wherein the scalar is associated with an object associated with the state-action pair.

[0249] Aspect 41. The apparatus according to any one of aspects 31 to 40, wherein the at least one processor is configured to transform the output data into the data representation using a linear mapping.

[0250] Aspect 42. An apparatus according to Aspect 41, wherein, based on using the linear mapping to transform the output data into the data representation, the at least one processor is configured to: output a scalar for each input scalar, output a vector for each input vector, and output a scalar and a vector for each input quaternion.

[0251] Aspect 43. The apparatus according to any one of aspects 31 to 42, wherein the at least one processor is configured to generate an equivariant diffusion model by combining an invariant basis density and the equivariant denoising network.

[0252] Aspect 44. An apparatus according to aspect 43, wherein the at least one processor is configured to train the equivariant diffusion model by adding noise to the state-action pairs to generate noisy trajectories, feeding the noisy trajectories into the equivariant denoising network, and using the equivariant diffusion model to output one or more predicted original trajectories of the state-action pairs.

[0253] Aspect 45. The apparatus of aspect 44, wherein the at least one processor is configured to unconditionally sample trajectories using the equivariant diffusion model.

[0254] Aspect 46. The apparatus of aspect 44, wherein the at least one processor is configured to conditionally sample trajectories based on an initial target and a state using the equivariant diffusion model.

[0255] Aspect 47. The apparatus according to any one of aspects 44 to 46, wherein the at least one processor is configured to use the equivariant diffusion model to sample trajectories under guidance from a classifier to solve a task.

[0256] Aspect 48. The apparatus of aspect 47, wherein to sample trajectories to solve the task under guidance from the classifier, the at least one processor is configured to use test-time reward and target conditioning.

[0257] Aspect 49. The apparatus of aspect 47, wherein to sample trajectories to solve the task under guidance from the classifier, the at least one processor is configured to: specify a new task using a reward.

[0258] Aspect 50. The apparatus according to any one of aspects 31 to 49, wherein the internal representation is associated with a symmetry group.

[0259] Aspect 51. The apparatus of aspect 50, wherein the symmetry group is a product of three different groups.

[0260] Aspect 52. The apparatus according to aspect 51, wherein the three different groups include a group of spatial translation and rotational symmetries, a group of discrete time translation symmetries, and a group of permutation groups on n objects.

[0261] Aspect 53. The apparatus of aspect 52, wherein the permutation group on n objects is associated with object properties that are permuted while keeping the robot properties or global properties of the state unchanged.

[0262] Aspect 54. An apparatus according to any one of aspects 52 or 53, wherein the spatial translation and rotation symmetry groups involve representations including scalars, vectors, and quaternions.

[0263] Aspect 55. An apparatus according to Aspect 54, wherein the scalar remains invariant under a rotation associated with an angle between two objects, wherein the vector uses a standard representation associated with position or velocity, and wherein the quaternion is transformed under a quaternion representation associated with orientation.

[0264] Aspect 56. The apparatus according to aspect 50, wherein the symmetry group is partitioned into at least one smaller symmetry group based on a condition.

[0265] Aspect 57. The apparatus of aspect 56, wherein the condition comprises at least one of the direction of gravity or the presence of a distinguishable object.

[0266] Aspect 58. An apparatus according to any one of aspects 31 to 57, wherein the spatial position associated with the state-action pair is expressed relative to a key object.

[0267] Aspect 59. The apparatus of aspect 58, wherein the key object comprises the position of the base of the robot or the center of mass of the robot.

[0268] Aspect 60. An apparatus according to any one of Aspects 31 to 59, wherein, in order to convert the geometric data type into the internal representation, the at least one processor is configured to perform at least one of the following operations: transforming the canonical representation under time shift, transforming the canonical representation under permutation, or transforming using scalar and vector representations.

[0269] Aspect 61. A non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by one or more processors, causing the one or more processors to perform the operations according to any one of aspects 1 to 30.

[0270] Aspect 62. An apparatus for processing data during a constant-variable diffuser, the apparatus comprising one or more components for performing the operations of any one of Aspects 1 to 30.

[0271] Aspect 61. A non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by one or more processors, causing the one or more processors to perform the operations according to any one of aspects 1 to 30.

[0272] Aspect 62. An apparatus for processing data during a constant-variable diffuser, the apparatus comprising one or more components for performing the operations of any one of Aspects 1 to 30.

Claims

1. A processor-implemented method for modeling a task using geometry, the processor-implemented method comprising: receiving, via a training preparation engine, a training dataset comprising state-action pairs; separating, via the training preparation engine, the state-action pairs from the training dataset into geometric data types; converting the geometric data type into an internal representation via the training preparation engine; processing the internal representation via an equivariant denoising network to generate output data; and The output data is transformed into a data representation.

2. The processor-implemented method of claim 1 , wherein the equivariant denoising network comprises layers of alternating types.

3. The processor-implemented method of claim 2, wherein the alternating types of layers include temporal layers, permutation layers, and geometric layers.

4. The processor-implemented method of claim 3 , wherein the temporal layer comprises a one-dimensional convolution along a trajectory stride dimension, wherein the permutation layer allows features belonging to different objects to interact, and wherein the geometric layer enables blending between scalar and vector quantities combined under the internal representation.

5. The processor-implemented method of claim 1 , wherein the geometric data type comprises a quaternion, and wherein the processor-implemented method further comprises transforming the quaternion into at least two rotation vectors.

6. The processor-implemented method of claim 5 , wherein transforming the quaternion into at least two rotation vectors comprises: Mapping the quaternion to a corresponding element in a matrix representation; as well as Two column vectors represented by the matrix are selected as the at least two rotation vectors.

7. The processor-implemented method of claim 1 , wherein transforming the output data into the data representation is performed using a linear mapping.

8. The processor-implemented method of claim 7 , wherein using the linear mapping to transform the output data into the data representation comprises: Outputs a scalar for each input scalar, a vector for each input vector, and a scalar and a vector for each input quaternion.

9. The processor-implemented method of claim 1 , further comprising: An equivariant diffusion model is generated by combining an invariant basis density and the equivariant denoising network.

10. The processor-implemented method of claim 9, wherein the equivariant diffusion model is trained by adding noise to the state-action pairs to generate noisy trajectories, feeding the noisy trajectories into the equivariant denoising network, and outputting one or more predicted original trajectories of the state-action pairs using the equivariant diffusion model.

11. The processor-implemented method of claim 10 , further comprising: The equivariant diffusion model is used to unconditionally sample trajectories.

12. The processor-implemented method of claim 10 , further comprising: Trajectories are conditionally sampled based on the initial target and state using the equivariant diffusion model.

13. The processor-implemented method of claim 10 , further comprising: The equivariant diffusion model is used to sample trajectories to solve the task under guidance from a classifier.

14. The processor-implemented method of claim 13 , wherein sampling trajectories to solve the task under guidance from the classifier comprises: Use reward and goal condition setting when testing.

15. The processor-implemented method of claim 13, wherein sampling trajectories to solve the task under guidance from the classifier comprises: Use rewards to assign new tasks.

16. The processor-implemented method of claim 1, wherein the internal representation is associated with a symmetry group.

17. The processor-implemented method of claim 16, wherein the symmetric group is a product of three different groups.

18. The processor-implemented method of claim 17, wherein the three different groups include a group of spatial translation and rotational symmetries, a group of discrete time translation symmetries, and a group of permutation groups over n objects.

19. The processor-implemented method of claim 18, wherein the permutation group over n objects is associated with object properties that are permuted while a robot property or a global property of a state remains unchanged.

20. The processor-implemented method of claim 18, wherein the spatial translation and rotational symmetry groups involve representations including scalars, vectors, and quaternions.

21. The processor-implemented method of claim 20, wherein the scalar is invariant under a rotation associated with an angle between two objects, wherein the vector employs a standard representation associated with position or velocity, and wherein the quaternion is transformed under a quaternion representation associated with orientation.

22. The processor-implemented method of claim 16, wherein the symmetry group is partitioned into at least one smaller symmetry group based on a condition.

23. The processor-implemented method of claim 22, wherein the condition comprises at least one of the direction of gravity or the presence of a distinguishable object.

24. The processor-implemented method of claim 1, wherein the spatial positions associated with the state-action pairs are expressed relative to a key object.

25. The processor-implemented method of claim 24, wherein the key object comprises a location of a base of a robot or a center of mass of a robot.

26. A processor-implemented method according to claim 1, wherein converting the geometric data type to the internal representation via the training preparation engine includes at least one of the following operations: transforming the canonical representation under time shift, transforming the canonical representation under permutation, or transforming using scalar and vector representations.

27. An apparatus for using a diffusion model that exploits symmetries in a geometric structure, the apparatus comprising: at least one memory; and at least one processor coupled to at least one memory and configured to: receiving, via a training preparation engine, a training dataset comprising state-action pairs; separating, via the training preparation engine, the state-action pairs from the training dataset into geometric data types; converting the geometric data type into an internal representation via the training preparation engine; processing the internal representation via an equivariant denoising network to generate output data; and The output data is transformed into a data representation.

28. The apparatus of claim 27, wherein the at least one processor is configured to: An equivariant diffusion model is generated by combining an invariant basis density and the equivariant denoising network.

29. A non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by one or more processors, causing the one or more processors to be configured to: receiving, via a training preparation engine, a training dataset comprising state-action pairs; separating, via the training preparation engine, the state-action pairs from the training dataset into geometric data types; converting the geometric data type into an internal representation via the training preparation engine; processing the internal representation via an equivariant denoising network to generate output data; and The output data is transformed into a data representation.

30. An apparatus for processing data during a constant-variable diffuser, the apparatus comprising one or more of: means for receiving, via a training preparation engine, a training data set comprising state-action pairs; means for separating, via the training preparation engine, the state-action pairs from the training dataset into geometric data types; means for converting said geometric data type to an internal representation via said training preparation engine; means for processing the internal representation via an equivariant denoising network to generate output data; and Means for transforming said output data into a data representation.