METHOD FOR SEMANTICALLY ALIGNED GENERATIVE EXTENSION FOR TRAINING STRATEGY MODELS
By generating augmented images that preserve depth and semantic information, the method addresses the 'sim-to-real' and 'real-to-real' gaps in visuomotor strategy training, enhancing robot control models' adaptability and task performance.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2026-03-12
AI Technical Summary
Conventional methods for training visuomotor strategy models face challenges in adapting to real-world environments due to the 'sim-to-real' and 'real-to-real' gaps, leading to inaccurate robot control due to differences in image characteristics such as colors, textures, and lighting, and loss of depth information.
A method involving a trained image-generating model that processes input images to produce augmented images, preserving depth and semantic information, which are used to train machine learning models to enhance their ability to control robots in diverse scenarios.
The augmented images provide diverse datasets for training, enabling strategy models to successfully control robots in various real-world and virtual environments by maintaining depth and semantic information, thus improving task performance.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND Area of the different designs
[0001] The various implementations generally relate to computer science, artificial intelligence (AI) and machine learning and robot control, and in particular to methods for semantically oriented generative extension for training strategy models. Description of the state of the art
[0002] In machine learning, training a visuomotor strategy or policy involves training a machine learning model, also known as a "strategy" or "policy" model, to generate motor actions for controlling a robot that is given image data as input. Once trained, the strategy model can be used to control a robot to perform a task, such as manipulating an object or navigating an environment.
[0003] A conventional approach to teaching a visuomotor strategy involves training a strategy model in a real-world environment using images captured by cameras and demonstrations of robot actions that the strategy model learns to imitate. In some cases, the strategy model can also convert the real-world images into canonical images, which are simplified versions of the real-world images. Because training a strategy model in a real-world environment can be time-consuming and could damage a robot, an alternative approach to teaching a visuomotor strategy is to train the strategy model using training data generated through simulations of the robot in a virtual environment.
[0004] However, a disadvantage of the above approaches is that the trained strategy model may not be able to correctly control the physical robot to perform a task in a real-world environment if the captured images of the real environment differ from the images used during training. For example, the captured images and the training images may differ in terms of the colors or textures of objects, the lighting conditions, or similar aspects. These differences are referred to as the "sim-to-real gap" when the training data used to train the strategy model is generated via simulations, and as the "real-to-real gap" when the training data is generated in a real-world environment.Due to the sim-to-real or real-to-real gap, the trained strategy may not be able to adapt to real-world scenarios that differ from the training data, and therefore correct control of a robot in these different scenarios may not be possible.
[0005] Furthermore, in cases where the strategy model converts captured real-world images into canonical images, the canonical images often differ significantly from the captured images. For example, the canonical images may show objects at different depths than the captured images. Accurate depth information is crucial for a robot to avoid collisions and grasp objects. Consequently, a trained strategy model that converts captured real-world images into canonical images may not be able to correctly control a robot in various scenarios.
[0006] As the above illustrates, more effective methods for training strategy models to control robots to perform tasks are needed in the state of the art. SUMMARY
[0007] One embodiment of the present disclosure presents a computer-implemented method for training machine learning models. The method involves processing one or more input images using a trained image-generating model to produce one or more augmented images. The trained image-generating model generates each augmented image contained within the one or more augmented images, depending on an input image contained within the one or more input images, depth information associated with the input image, semantic information associated with the input image, and text describing an augmentation to be performed on the input image.The procedure further involves performing one or more operations, based on the one or more enhanced images, to train an untrained machine learning model in order to generate a trained machine learning model.
[0008] At least one technical advantage of the disclosed methods compared to the prior art is that the disclosed methods generate augmented images that can provide diverse datasets for training machine learning models, such as strategy models for controlling robots. Using augmented images generated according to the disclosed methods, a strategy model can be trained to control a robot and perform a task more successfully than with strategy models trained using conventional approaches. In particular, the augmented images preserve depth and semantic information from the input images, which is useful for training the strategy model to perform tasks correctly. These technical advantages represent one or more technological improvements over prior art approaches. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] To better understand the features of the various embodiments mentioned above, a more detailed description of the inventive concepts summarized above can be provided with reference to various embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings only illustrate typical embodiments of the inventive concepts and are therefore in no way intended to limit the scope of the invention, and that other equally effective embodiments exist. Fig. Figure 1 illustrates a block diagram of a computer-based system configured to implement one or more aspects of at least one embodiment; Fig. Figure 2 is a more detailed illustration of the machine learning server. Fig. 1 according to various embodiments; Fig. Figure 3 is a more detailed illustration of the calculating device. Fig. 1 according to various embodiments; Fig. Figure 4 illustrates how a strategy model can be trained to control a robot, according to different embodiments; Fig. Figure 5 is a more detailed illustration of the image-generating model from Fig. 1 according to various embodiments; Fig. Figure 6 illustrates exemplary input and output images of the image-generating model. Fig. 1 according to various embodiments; Fig. Figure 7 is a flowchart of process steps for training an image-generating model according to different embodiments; and Fig. Figure 8 is a flowchart of process steps for controlling a robot according to different embodiments. DETAILED DESCRIPTION
[0010] The following description sets out numerous specific details to provide a more thorough understanding of the various embodiments. However, it is evident to a person skilled in the art that the inventive concepts can be implemented without one or more of these specific details. General overview
[0011] Embodiments of the present disclosure provide methods for generating augmented image data and training machine learning models using the augmented image data. In some embodiments, a trained image-generating model takes images and text describing augmentations as input, and the image-generating model generates augmented images based on the input images, depth and semantic features extracted from the input images, and the text describing the augmentations. The image-generating model includes three diffusion modules. For a given image input and text, a first diffusion module is used to generate a feature map based on the input image and text. A second diffusion module is used to generate a second feature map based on the image, depth features extracted from the image, and the text.A third diffusion module is used to generate a third feature map based on the image, semantic features extracted from the image, and the text. A decoder processes the first, second, and third feature maps to produce an augmented image. Any number of augmented images can be generated according to the preceding steps for inclusion in a training dataset. A machine learning model, such as a strategy or policy model for controlling a robot, can then be trained using the training dataset. Once trained, the machine learning model can be used to perform one or more tasks. For example, a trained strategy model can be used to control a robot within a real-world or virtual environment.
[0012] The methods for generating augmented image data and training machine learning models disclosed herein have many real-world applications. For example, these methods can be used to generate augmented image data and train strategy models to control robots in real-world environments or to control robot simulations in virtual environments. As another example, these methods can be used to generate augmented image data and train any technically feasible machine learning models that can benefit from training with the augmented image data.
[0013] The examples above are not intended to be restrictive in any way. As those skilled in the art will recognize, the methods described here for generating augmented image data and training machine learning models can generally be used in any application where trained machine learning models are required or useful. System Overview
[0014] Fig. Figure 1 illustrates a block diagram of a computer-based system 100 designed to implement one or more aspects of at least one embodiment. As shown, the system 100 includes, without limitation, a machine learning server 110, a data storage device 120, and a computing device 140 communicating over a network 130, which may include a wide area network (WAN) such as the Internet, a local area network (LAN), a mobile network, and / or any other suitable network or networks.
[0015] As shown, a model trainer 116 and an image-generating model 119 are executed on one or more processors 112 of the machine learning server 110 and stored in a system memory 114 of the machine learning server 110. The processor(s) 112 receive user input from input devices, such as a keyboard or mouse. During operation, the one or more processors 112 may include one or more primary processors of the machine learning server 110, which control and coordinate the operations of other system components. In particular, the processor(s) 112 may issue instructions that control the operation of one or more graphics processing units (GPUs) (not shown) and / or other parallel processing circuitry (e.g., parallel processing units, deep learning accelerators, etc.).The GPU(s) control circuitry that includes components optimized for graphics and video processing, such as video output circuitry. The GPU(s) can deliver pixels to a display device, which can be any conventional cathode ray tube, liquid crystal display, LED display, and / or the like.
[0016] The system memory 114 of the machine learning server 110 stores content, such as software applications and data, for use by the processor(s) 112 and the GPU(s) and / or other processing units. The system memory 114 can be any type of memory capable of storing data and software applications, such as random-access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash ROM), or any suitable combination thereof. In some embodiments, a memory (not shown) can supplement or replace the system memory 114. The memory can include any number and type of external memory accessible to the processor(s) 112 and / or the GPU.For example, and without limitation, the storage medium may include a Secure Digital Card, an external flash memory, a portable Compact Disc read-only storage medium (CD-ROM), an optical storage device, a magnetic storage device and / or any suitable combination of the foregoing.
[0017] The machine learning server 110 shown here is for illustrative purposes only, and variations and modifications are possible without deviating from the scope of this disclosure. For example, the number of processors 112, the number of GPUs and / or other processing unit types, the number of system memories 114, and / or the number of applications contained in the system memory 114 can be modified as required. Furthermore, the connection topology between the various units in Fig. 1 can be modified as required. In some embodiments, any combination of the processor(s) 112, the system memory 114 and / or the GPU(s) in any type of virtual computing system, distributed computing system and / or cloud computing environment, such as a public, private or hybrid cloud system, can be included and / or replaced.
[0018] In some embodiments, the model trainer 116 is configured to train one or more machine learning models, including an image-generating model 119, which is trained to produce enhanced images for training a strategy model 150, which is trained to control a robot to perform a task. The image-generating model 119 and the strategy model 150 can be trained by the model trainer 116 or by different model trainers in any technically feasible manner. Details of the image-generating model 119 and the strategy model 150, as well as methods for training them, are described below in conjunction with the Fig. Sections 5 and 7-8 are discussed in more detail. Training data and / or trained machine learning models, including the image-generating model 119 and the strategy model 150, can be stored in the data storage 120 or elsewhere. In some embodiments, the data storage 120 can include any storage device or devices, such as hard disk drive(s), flash drive(s), optical storage, network-attached storage (NAS), and / or a storage area network (SAN). Although shown to be accessible via the network 130, the machine learning server 110 can include the data storage 120 in at least one embodiment.
[0019] As shown, the data generator 118, which uses the image-generating model 119, is stored in the system memory 114 and executed on the processor(s) 112 of the server 110 for machine learning. Once trained, the image-generating model 119 can be used in any suitable way, such as in the data generator 118, for use in generating enhanced images.
[0020] As shown, a robot control application 146, which uses the trained strategy model 150, is stored in a system memory 144 and is executed on the processor(s) 142 of the computing device 140. Once trained, the strategy model 150 can be used in any suitable way, such as in the robot control application 146. For example, given sensor data acquired by one or more sensors 180, such as images captured by one or more cameras, the strategy model 150 can be used to control a physical robot 160 to perform a task for which the strategy model 150 has been trained in a real-world environment.
[0021] As shown, the robot 160 includes several connections 161, 163, and 165, which are rigid elements, and joints 162, 164, and 166, which are movable components that can be actuated to effect relative movement between adjacent connections. Additionally, the robot 160 includes several fingers 168i (hereinafter collectively referred to as fingers 168 and individually as fingers 168) that can be controlled to grasp an object. Although an exemplary robot 160 is shown for illustrative purposes, in some embodiments, methods disclosed herein can be applied to control any suitable robot.
[0022] Fig. Figure 2 is a more detailed illustration of Server 110 for machine learning. Fig. 1 according to various embodiments. The Machine Learning Server 110 can include any type of computing system, including, without limitation, a server machine, a server platform, a desktop machine, a laptop machine, a handheld / mobile device, a digital kiosk, an in-vehicle infotainment system, and / or a portable device. In some embodiments, the Machine Learning Server 110 is a server machine located in a data center or cloud computing environment that provides scalable computing resources as a service over a network.
[0023] In various embodiments, the machine learning server 110 includes, without limitation, the processor(s) 112 and the system memory 114, which are coupled to a parallel processing subsystem 212 via a memory bridge 205 and a communication path 213. The memory bridge 205 is further coupled to an I / O (input / output) bridge 207 via a communication path 206, and the I / O bridge 207 is in turn coupled to a switch 216.
[0024] In some embodiments, the I / O bridge 207 is configured to receive user input information from optional input devices 208, such as a keyboard, mouse, touchscreen, sensor data analysis (e.g., evaluating gestures, speech, or other information about one or more applications in a field of view or sensory field of one or more sensors), and / or the like, and to forward the input information to the processor(s) 112 for processing. In some embodiments, the machine learning server 110 may be a server machine in a cloud computing environment. In such embodiments, the machine learning server 110 may not include input devices 208 but may receive equivalent input information by receiving commands (e.g.,in response to one or more inputs from a remote computing device) in the form of messages that are transmitted over a network and received via the network adapter 218. In some embodiments, the switch 216 is configured to provide connections between the I / O bridge 207 and other machine learning components of the server 110, such as a network adapter 218 and various add-on cards 220 and 221.
[0025] In some embodiments, the I / O bridge 207 is coupled to a system disk 214, which can be configured to store content, applications, and data for use by the processor(s) 112 and the parallel processing subsystem 212. In some embodiments, the system disk 214 provides non-volatile memory for applications and data and may include hard disk drives, removable disk drives, flash memory devices, and CD-ROM (Compact Disc Read-Only Memory), DVD-ROM (Digital Versatile Disc-ROM), Blu-ray, HD-DVD (High-Definition DVD), or other magnetic, optical, or solid-state storage devices. In various embodiments, other components, such as Universal Serial Bus or other port connections, Compact Disc drives, Digital Versatile Disc drives, movie recording devices, and the like, may also be connected to the I / O bridge 207.
[0026] In various embodiments, the memory bridge 205 can be a northbridge chip, and the I / O bridge 207 can be a southbridge chip. Additionally, the communication paths 206 and 213, as well as other communication paths within the server 110 for machine learning, can be implemented using any technically suitable protocols, including, without limitation, AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol known in the art.
[0027] In some embodiments, the parallel processing subsystem 212 includes a graphics subsystem that delivers pixels to an optional display device 210, which may be any conventional cathode ray tube, liquid crystal display, light-emitting diode display, and / or the like. In such embodiments, the parallel processing subsystem 212 may include circuits optimized for graphics and video processing, including, for example, video output circuits. Such circuits may be integrated via one or more parallel processing units (PPUs), also referred to herein as parallel processors, included in the parallel processing subsystem 212.
[0028] In some embodiments, the parallel processing subsystem 212 includes circuits optimized for general-purpose and / or computational processing (e.g., those undergoing optimization). Such circuits may be integrated via one or more power processing units (PPUs) included in the parallel processing subsystem 212, configured to perform such general-purpose and / or computational operations. In still other embodiments, the one or more PPUs included in the parallel processing subsystem 212 may be configured to perform graphics processing, general-purpose processing, and / or computational operations. The system memory 114 includes at least one device driver configured to manage the processing operations of the one or more PPUs in the parallel processing subsystem 212. Additionally, the system memory 114 includes the model trainer 116 and the data generator 118.Although described herein mainly in relation to the model trainer 116 and the data generator 118, the methods disclosed herein may also be implemented either completely or partially in other software and / or hardware, such as in the parallel processing subsystem 212.
[0029] In various embodiments, the parallel processing subsystem 212 can be combined with one or more of the other elements from Fig. 2 be integrated to form a single system. For example, the parallel processing subsystem 212 can be integrated with the processor(s) 112 and other interconnect circuitry on a single chip to form a system-on-a-chip (SoC).
[0030] In some embodiments, the processor(s) 112 includes the primary machine learning processor of the server 110, which controls and coordinates operations of other system components. In some embodiments, the processor(s) 112 issues instructions that control the operation of the PPUs. In some embodiments, the communication path 213 is a PCI Express connection in which dedicated lanes are assigned to each PPU. Other communication paths may also be used. The PPU advantageously implements a highly parallel processing architecture, and the PPU can be provided with any amount of local parallel processing memory (PP memory).
[0031] It is understood that the system shown here is for illustrative purposes only and that variations and modifications are possible. The connection topology, including the number and arrangement of bridges, the number of processor(s) 112, and the number of parallel processing subsystems 212, can be modified as required. For example, in some embodiments, the system memory 114 can be connected directly to the processor(s) 112 instead of via the memory bridge 205, and other devices can communicate with the system memory 114 via the memory bridge 205 and the processor(s) 112. In other embodiments, the parallel processing subsystem 212 can be connected to the I / O bridge 207 or directly to the processor(s) 112 instead of the memory bridge 205.In other embodiments, the I / O bridge 207 and the memory bridge 205 can be integrated into a single chip instead of being present as one or more discrete devices. In certain embodiments, one or more can be integrated into a single chip. Fig. The two components shown may not be present. For example, the switch 216 can be omitted, and the network adapter 218 and the add-in cards 220 and 221 can connect directly to the I / O bridge 207. Finally, in certain embodiments, one or more of the following may be omitted: Fig. The components shown in Figure 2 can be implemented as virtualized resources in a virtual computing environment, such as a cloud computing environment. For example, in at least one embodiment, the parallel processing subsystem 212 can be implemented as a virtualized parallel processing subsystem. As a specific example, the parallel processing subsystem 212 can be implemented as virtual graphics processing unit(s) (vGPU(s)) that render graphics on a virtual machine(s) (VM(s)) running on server machine(s), whose GPU(s) and other physical resources are shared across one or more VMs.
[0032] Fig. Figure 3 is a more detailed illustration of the calculating device 140. Fig. 1 according to various embodiments. The computing device 140 can include any type of computing system, including, without limitation, a server machine, a server platform, a desktop machine, a laptop machine, a handheld / mobile device, a digital kiosk, an in-vehicle infotainment system, and / or a portable device. In some embodiments, the computing device 140 is a server machine operated in a data center or cloud computing environment that provides scalable computing resources as a service over a network.
[0033] In various embodiments, the computing device 140 includes, without limitation, the processor(s) 142 and the system memory 144, which are coupled to a parallel processing subsystem 312 via a memory bridge 305 and a communication path 313. The memory bridge 305 is further coupled to an I / O bridge 307 via a communication path 306, and the I / O bridge 307 is in turn coupled to a switch 316.
[0034] In some embodiments, the I / O bridge 307 is configured to receive user input information from optional input devices 308, such as a keyboard, mouse, touchscreen, sensor data analysis (e.g., evaluating gestures, speech, or other information about one or more applications in a field of view or sensory field of one or more sensors), and / or the like, and to forward the input information to the processor(s) 142 for processing. In some embodiments, the computing device 140 may be a server machine in a cloud computing environment. In such embodiments, the computing device 140 may not include input devices 308 but may receive equivalent input information by receiving commands (e.g.,in response to one or more inputs from a remote computing device) in the form of messages that are transmitted over a network and received via the network adapter 318. In some embodiments, the switch 316 is configured to provide connections between the I / O bridge 307 and other components of the computing device 140, such as a network adapter 318 and various add-in cards 320 and 321.
[0035] In some embodiments, the I / O bridge 307 is coupled to a system disk 314, which can be configured to store content, applications, and data for use by the processor(s) 142 and the parallel processing subsystem 312. In some embodiments, the system disk 314 provides non-volatile memory for applications and data and may include hard disk drives, removable disk drives, flash memory devices, and CD-ROM (Compact Disc Read-Only Memory), DVD-ROM (Digital Versatile Disc-ROM), Blu-ray, HD-DVD (High-Definition DVD), or other magnetic, optical, or solid-state storage devices. In various embodiments, other components, such as Universal Serial Bus or other port connections, Compact Disc drives, Digital Versatile Disc drives, movie recording devices, and the like, may also be connected to the I / O bridge 307.
[0036] In various embodiments, the memory bridge 305 can be a northbridge chip, and the I / O bridge 307 can be a southbridge chip. Additionally, the communication paths 306 and 313, as well as other communication paths within the computing device 140, can be implemented using any technically suitable protocols, including, without limitation, AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol known in the art.
[0037] In some embodiments, the parallel processing subsystem 312 includes a graphics subsystem that delivers pixels to an optional display device 310, which may be any conventional cathode ray tube, liquid crystal display, light-emitting diode display, and / or the like. In such embodiments, the parallel processing subsystem 312 may include circuits optimized for graphics and video processing, including, for example, video output circuits. Such circuits may be integrated via one or more power processing units (PPUs), also referred to herein as parallel processors, included in the parallel processing subsystem 312.
[0038] In some embodiments, the parallel processing subsystem 312 includes circuits optimized for general-purpose and / or computational processing (e.g., those undergoing optimization). Such circuits may be integrated via one or more power processing units (PPUs) included in the parallel processing subsystem 312, configured to perform such general-purpose and / or computational operations. In still other embodiments, the one or more PPUs included in the parallel processing subsystem 312 may be configured to perform graphics processing, general-purpose processing, and / or computational operations. The system memory 144 includes at least one device driver configured to manage the processing operations of the one or more PPUs in the parallel processing subsystem 312. Additionally, the system memory 144 includes the robot control application 146.Although described herein mainly in relation to the robot control application 146, the methods disclosed herein may also be implemented either completely or partially in other software and / or hardware, such as in the parallel processing subsystem 312.
[0039] In various embodiments, the parallel processing subsystem 312 can be combined with one or more of the other elements from Fig. 3 can be integrated to form a single system. For example, the parallel processing subsystem 312 can be integrated with the processor(s) 142 and other interconnect circuitry on a single chip to form a SoC.
[0040] In some embodiments, the processor(s) 142 includes the primary processor of the computing device 140, which controls and coordinates operations of other system components. In some embodiments, the processor(s) 142 issues instructions that control the operation of PPUs. In some embodiments, the communication path 313 is a PCI Express connection in which dedicated lanes are assigned to each PPU. Other communication paths may also be used. The PPU advantageously implements a highly parallel processing architecture, and the PPU can be provided with any amount of local parallel processing memory (PP memory).
[0041] It is understood that the system shown here is for illustrative purposes only and that variations and modifications are possible. The connection topology, including the number and arrangement of bridges, the number of processor(s) 142, and the number of parallel processing subsystems 312, can be modified as required. For example, in some embodiments, the system memory 144 can be connected directly to the processor(s) 142 instead of via the memory bridge 305, and other devices can communicate with the system memory 144 via the memory bridge 305 and the processor(s) 142. In other embodiments, the parallel processing subsystem 312 can be connected to the I / O bridge 307 or directly to the processor(s) 142 instead of the memory bridge 305.In other embodiments, the I / O bridge 307 and the memory bridge 305 can be integrated into a single chip instead of existing as one or more discrete devices. In certain embodiments, one or more can be integrated into a single chip. Fig. The three components shown may not be present. For example, the switch 316 can be omitted, and the network adapter 318 and the add-in cards 320 and 321 can connect directly to the I / O bridge 307. Finally, in certain embodiments, one or more of the following may be omitted: Fig. The components shown in Figure 3 can be implemented as virtualized resources in a virtual computing environment, such as a cloud computing environment. For example, in at least one embodiment, the parallel processing subsystem 312 can be implemented as a virtualized parallel processing subsystem. As a specific example, the parallel processing subsystem 312 can be implemented as virtual graphics processing unit(s) (vGPU(s)) that render graphics on a virtual machine (VM(s)) running on a server machine (or machines), whose GPU(s) and other physical resources are shared across one or more VMs. Robot control models trained using semantically targeted generative extension
[0042] Fig. Figure 4 illustrates how a strategy model can be trained to control a robot, according to various embodiments. As shown, the data generator 118 includes, without limitation, the image-generating model 119. The image-generating model 119 is a trained machine learning model designed to take as input an image and text describing a robot task associated with the image, as well as an augmentation to be applied to the image, and to output an augmented image. Details of the image-generating model 119, as well as procedures for training the image-generating model 119, are described below in conjunction with the Fig. 5 and Fig. 7 discussed in more detail.
[0043] During operation, the data generator 118 can receive a set of images, represented as a set 402 of extended images, which includes images assigned to one or more robot tasks and captured by one or more cameras. The camera(s) can include one or more physical cameras, such as cameras included in the sensors 180 assigned to the robot 160, which capture images of real-world environments, and / or one or more virtual cameras that capture images within simulated environments, which may be virtual environments simulating real-world environments. In some embodiments, the image set 402 can include sets of images from different domains, such as real-world images and images from simulations in different simulation environments.
[0044] The data generator 118 processes the image set 402 using the image-generating model 119 to produce augmented images. The augmented images contain the same objects as images from the image set 402, but they have different colors, lighting conditions, and / or textures and can reside in different domains, such as real-world images or simulation images, depending on the text used to generate them. For example, in some embodiments, the data generator 118 can repeatedly input an image from the image set 402 into the image-generating model 119, along with text describing the associated robot task and an augmentation to be applied to the image. In such cases, the text can describe the augmentation in any suitable way, including with any level of detail.For example, the extension can be described generally as transforming the image into another image from a simulated or real-world environment. Alternatively, the extension can be described as transforming the image into another image from a specific domain, such as a specific simulated environment. Likewise, the text can describe the robot task in any suitable way, including with any level of detail. For example, the task can be described generally for a robot in a kitchen. Alternatively, the task can be described specifically for a robot picking up an object in a kitchen. In some embodiments, the text describing the robot task and any extensions to be applied can be generated using one or more templates or in any other technically feasible way.
[0045] As discussed in more detail below, the image-generating model 119 is designed to extract render-invariant features, including depth information and semantic information about different identities of objects (e.g., whether an object is an apple, a bottle, etc.), from images input to the image-generating model 119, and the image-generating model 119 produces the augmented images based on the render-invariant features. Although described here mainly with respect to depth and semantic information as reference examples of invariant features, in some embodiments all technically feasible invariant features can be extracted using computer vision or machine vision techniques, such as surface normals, segmentations, etc.The augmented images contain the same render-invariant features, such as the same object identities and the same depths, as the input images. Once generated, the augmented images, along with image set 402, are contained in a set of 406 augmented images, which is output by data generator 118. The set of 406 augmented images can be stored in data memory 120 or elsewhere.
[0046] The model trainer 116 trains a strategy model, shown as strategy model 150, to control a robot, shown as robot 160, using set 406 of augmented images. Strategy model 150 is a machine learning model trained using set 406 augmented images to generate actions for controlling a robot to perform at least part of a task. Although described here primarily in terms of a strategy model as a reference example, in some embodiments any technically feasible machine learning model can be trained using a set of augmented images. Strategy model 150 can have any suitable architecture and be trained in any technically feasible way.For example, in some embodiments, the image set 402 may include images from expert demonstrations of tasks that the strategy model 150 should learn in order to control the robot 160 to perform these tasks. In such cases, the data generator 118 can add the images from the expert demonstrations to the generated set 406 of extended images, and the model trainer 116 can train the strategy model 150 using an imitative learning procedure to mimic expert actions from the expert demonstrations that correspond to images from the set 406 of extended images input into the strategy model 150.In such cases, the strategy model 150 can be trained using supervised learning to predict actions that mimic expert actions based on observed states containing the images from the set of 406 augmented images. This training can minimize the loss due to imitative learning, which is the difference between the predicted actions and the expert actions. Because the set of 406 augmented images contains relatively diverse images with varying colors, textures, lighting conditions, and so on, the trained strategy model 150 can be more generalized to correctly control a robot in different scenarios.
[0047] Once trained, the strategy model 150 can be used to control a robot in a physical or virtual environment. For illustrative purposes, the strategy model 150 was used in the robot control application 146 to control the robot 160 based on sensor data 407 received from the sensors 180. The sensor data 407 can include images captured by one or more cameras mounted on the robot 160 and / or within its environment. Given sensor data 407 as input, the strategy model 150 generates an action 408, which is a command to control the robot 160 to perform at least part of a task.In some embodiments, the robot control application 146 can transfer the action 408 to a lower-level controller, such as a proportional integral differential controller (PID controller) or a proportional differential controller (PD controller), which controls the actuators of the robot 160 according to the action 408.
[0048] Fig. Figure 5 is a more detailed illustration of the image-generating model 119 from Fig. 1 according to various embodiments. As shown, the image-generating model 119 includes, without limitation, an activation function 506, a diffusion module 508, a depth feature extractor 510, downsample and zero-convolution layers 512, a semantic feature extractor 518, upsample and zero-convolution layers 520, diffusion module copies 514 and 522, zero-convolution layers 516 and 524, an activation function 526, and a decoder 528.
[0049] The image-generating model 119 is a machine learning model, such as an artificial neural network. In operation, the image-generating model 119 accepts as input an image 502 and a text 504 that describes a robot task associated with image 502 and an extension to be applied to the image. Similar to the above description in conjunction with Fig. 4. Text 504 may describe the robot task and the extension in any suitable manner, including with any level of detail, in some embodiments.
[0050] Using image 502 and text 504, the image-generating model 119 produces an image 530 to which the enhancement specified by text 504 is applied. For illustrative purposes, the following processing is performed in parallel: (1) image 502 and text 504 are processed using the activation function 506 to generate features, and the diffusion module 508 performs a denoising diffusion procedure depending on the generated features to produce an initial feature map; (2) The image 502 is processed using the depth feature extractor 510, which is a computer vision module that extracts features indicating the depths of objects in the image 502. The extracted features are further processed using the downsample and zero-convolution layers 512 to generate additional features that are linked to text features generated by the activation function 506.and the diffusion module copy 514 performs a denoising diffusion procedure depending on the linked features to generate an intermediate feature map, which is further processed by the zero-convolution layer 516 to generate a second feature map; and (3) the image 502 is processed using the semantic feature extractor 518, which is a computer vision module that extracts features that provide semantic information about the identities of objects in the image 502; the extracted features are further processed using the upsample and zero-convolution layers 520 to generate additional features linked to text features generated by the activation function 506; and the diffusion module copy 522 performs a denoising diffusion procedure depending on the linked features to generate an intermediate feature map.which is further processed using the zero-convolution layer 524 to generate a third feature map. The first, second, and third feature maps are then linked and processed using the activation function 526 to generate additional features, which the decoder 528 decodes to produce the image 530.
[0051] In some embodiments, the Diffusion Module 508 can include a pretrained text-to-image diffusion model, such as the Stable Diffusion XL model. The underlying mechanism of such a model is based on denoising diffusion probabilistic models (DDPMs), which define a forward diffusion process that gradually adds Gaussian noise to images and a reverse process that learns to remove random noise from images. In particular, the forward process can be defined as: q(xt|xt−1)=N(xt;1−βtxt−1,βtI), where β tthe noise plan is and x t The image is represented at time step t. The inverse process learns to predict the noise ε. θ and can be optimized using: L=Ex0,ε,t[‖ε−εθ(xt,t)‖2], which results in a reverse diffusion method corresponding to a Gaussian distribution: pθ(xt−1|xt,c)=N(xt−1;μθ(xt,t,c),∑θ(xt,t)).
[0052] To enable text-guided generation, the diffusion module 508 can integrate text conditioning through classifier-free control. During inference, noise prediction is guided by: ε^θ=εθ(xt,c)+w(εθ(xt,c)−εθ(xt,∅)), where c is the text conditioning or text condition, Ø represents an unconditional generation, and w is the control scale that controls the alignment strength between the generated image and the text instruction.
[0053] While DDPM models excel at generating diverse images from text instructions, precise spatial control over the generated content remains a challenge. In some embodiments, a ControlNet can be used to overcome this limitation, enabling fine-grained spatial control while retaining the generative capabilities of the base diffusion model. The ControlNet extends conventional diffusion models by introducing additional conditioning paths for control signals. Specifically, the ControlNet can be used to enable conditioning based on depth features, generated using the Depth Feature Extractor 510, and semantic features, generated using the Semantic Feature Extractor 518.Although described here mainly in relation to ControlNet as a reference example, in some embodiments any technically feasible mechanism that enables noise reduction (denoising) by diffusion depending on depth and semantic features can be used.
[0054] The depth feature extractor 510 provides spatial control to ensure that the generated image 530 contains objects with the same geometry and at the same depths as objects in the image 502. In some embodiments, the depth feature extractor 510 can be implemented using any technically feasible machine learning model capable of extracting depth information from an input image, such as a Depth-Anything-v2 model, which serves as a base model for extracting precise depth information from input images. In some embodiments, the backbone of both the original diffusion model and the ControlNet in the diffusion module copy 514 can be a UNet architecture that handles features at multiple resolutions via encoder-decoder paths with skip connections.In such cases, deep conditioning can be integrated into the diffusion process using a modified UNet architecture. εθ(xt,t,c,h)=UNet(xt,t,c)+ZeroConv(Control(ZeroConv(h))), where h represents the depth conditioning extracted by the depth feature extractor 510 (e.g., Depth-Anything-v2). The control module in the diffusion module copy 514 mirrors the UNet architecture but processes only the depth information. The zero-convolution layers in the downsample and zero-convolution layers 512 and the zero-convolution layer 516 are initialized with zeros and serve two purposes: The zero-convolution layers allow for gradual learning of the control signal during training and prevent the depth conditioning from overburdening the original generation process. Consequently, the depth-conditioned generation process can be formulated as: pθ(xt−1|xt,c,h)=N(xt−1;μθ(xt,t,c,h),∑θ(xt,t)), where µ θ The mean value of the denoised image is calculated using depth-aware noise prediction. In some embodiments, the control modules and zero-convolution layers can be trained while the original UNet weights remain frozen, thus preserving the generative capabilities of the base model while adding spatial control.
[0055] The semantic feature extractor 518 provides geometric control to ensure that the generated image 530 contains the same object identities as image 502. In some embodiments, the semantic feature extractor 518 can be implemented using any technically feasible machine learning model, such as the SigLIP (Sigmoid Loss Image Pretraining) model, which is capable of associating text labels with an input image. Similar to the depth conditioning path, the semantic feature extractor 518 can process the semantic features through zero-convolution layers. εθ(xt,t,c,h,s)=Unet(xt,t,c)+ZeroConv(Controldepth(ZeroConv(h)))+ZeroConv(Controlsem(ZeroConv(s))), where s represents the semantic features extracted by the semantic feature extractor 518 (e.g., SigLIP). The semantic control branch transforms the token-based features into spatial representations consistent with the image generation process. Specifically, when SigLIP is used as the semantic feature extractor 518 to accommodate the semantic conditioning mechanism, which differs from the geometry branch due to the token-based nature of SigLIP's representations, the control architecture can be modified to use upsample modules, as demonstrated by the upsample and zero-convolution layers 520.Experience has shown that semantic extractors, such as SigLIP, offer superior semantic alignment when a language-contrastive learning approach is used to train the semantic extractors, that they better capture semantic relationships between text and visual features, i.e., the language-vision alignment inherent in the training can help maintain semantic consistency in the image-generating model 119.
[0056] As described, a feature map is generated by the diffusion module 508, the diffusion module copy 514, which depends on depth features extracted by the depth feature extractor 510, and the diffusion module copy 522, which depends on semantic features extracted by the semantic feature extractor 518. The generated feature maps are then processed using the activation function 526 to generate additional features, which the decoder 528 decodes to produce the image 530. The decoder 528 can be implemented in any technically feasible way, such as with one or more layers of neural networks (e.g., the neural network layers of the decoder from a stable diffusion model).
[0057] In some embodiments, the image-generating model 119 can be trained using images from different image datasets, such as image datasets associated with physical and / or simulated environments, and / or datasets associated with different domains. Any technically feasible training method, such as gradient descent backpropagation or a variation thereof, can be used in some embodiments to train the image-generating model 119. In some embodiments, training the image-generating model 119 can minimize reconstruction loss, which is a difference between an input image and an image generated by the image-generating model 119.Reconstruction loss can be used when the training data does not contain paired images that provide examples of output images with different extensions, so the goal of the training is instead to reconstruct the input images. In some embodiments, premature termination of the training can be used to introduce variance (i.e., randomness) into the outputs of the image-generating model 119. In some embodiments, certain parameters of the image-generating model 119 can remain unchanged during training, while other parameters of the image-generating model 119 are updated. Returning to the example where diffusion module copy 514 and diffusion module copy 522 each contain a diffusion model with the ControlNet, the training can involve updating parameters of the ControlNet while the diffusion model remains unchanged.Furthermore, in some embodiments, parameters of a diffusion module in the diffusion module 508 can remain unchanged during training. Additionally, parameters of the depth feature extractor 510 and the semantic feature extractor 518 can remain unchanged during training in some embodiments.
[0058] Fig. Figure 6 illustrates exemplary input and output images of the image-generating model 119 from Fig. 1 according to various embodiments. As shown, the input images 602 and 604 are derived from two different real datasets, and the images 606, 608, and 610 are derived from three different simulated datasets. If the input is an image 602, 604, 606, 608, or 610 and text specifying the conversion of the input image into a domain of the real dataset associated with input image 602, the image-generating model 119 can generate the images 620. If the input is an image 602, 604, 606, 608, or 610 and text specifying the conversion of the input image into a domain of the real dataset associated with input image 604, the image-generating model 119 can generate the images 622.If the input is an image 602, 604, 606, 608, or 610 and text specifying how to convert the input image into a domain of the simulated dataset associated with input image 606, the image-generating model 119 can generate images 624. If the input is an image 602, 604, 606, 608, or 610 and text specifying how to convert the input image into a domain of the simulated dataset associated with input image 608, the image-generating model 119 can generate images 626. If the input is an image 602, 604, 606, 608 or 610 and a text specifying the conversion of the input image into a domain of the simulated data set associated with input image 610, the image-generating model 119 can generate the images 628.
[0059] Fig. Figure 7 is a flowchart of process steps for training the image-generating model 119 according to various embodiments. Although the process steps are related to the systems of Fig. As described in Figures 1-5, the person skilled in the art will understand that any system designed to perform the process steps in any order falls within the scope of the present embodiments.
[0060] As shown, a procedure 700 begins at step 702, where the model trainer 116 receives one or more sets of images. As described, in some embodiments the image-generating model 119 can be trained using images from different image datasets, such as image datasets associated with physical and / or simulated environments, and / or datasets associated with different domains.
[0061] In step 704, the model trainer 116 selects an image from the set (sets) of images. Then, in step 706, the model trainer 116 processes the selected image using an untrained version of the image-generating model 119 to produce an output image. The image-generating model 119 is described above in conjunction with Fig. 5 described.
[0062] In step 708, the model trainer 116 calculates a reconstruction loss based on the output image and the selected image. The reconstruction loss is a difference (e.g., a pixel-wise difference) between the selected image and the output image produced by the image-generating model 119. As described, a reconstruction loss can be used in some embodiments when the training data does not contain paired images that include examples of output images with different extensions, so that the goal of the training is instead to reconstruct the input images.
[0063] In step 710, the model trainer 116 updates parameters of the image-generating model 119 based on the reconstruction loss. As described, in some embodiments, the parameters of the image-generating model 119 can be updated iteratively in any technically feasible way, such as backpropagation with gradient descent or a variation thereof. In some embodiments, certain parameters of the image-generating model 119 can remain unchanged during training, while other parameters of the image-generating model 119 are updated. For example, if the diffusion modulus copy 514 and the diffusion modulus copy 522 each contain a diffusion model with the ControlNet, parameters of the ControlNet can be updated during training, while parameters of the diffusion model remain unchanged.Furthermore, in some embodiments, parameters of a diffusion module in the diffusion module 508 can remain unchanged during training. Additionally, parameters of the depth feature extractor 510 and the semantic feature extractor 518 can remain unchanged during training in some embodiments.
[0064] If the model trainer 116 determines at step 712 to continue training, then the procedure 700 returns to step 704, with the model trainer 116 selecting a different image from the set (sets) of images. In some embodiments, the model trainer 116 can iteratively update parameters of the vision encoder based on the reconstruction loss until a stop condition is met, such as the training having been completed for a predefined number of iterations, the loss stabilizing, or the like. In some embodiments, premature termination of training can be used to introduce variance (i.e., randomness) into the outputs of the image-generating model 119.
[0065] If, on the other hand, the model trainer 116 decides to stop the training at step 712, then procedure 700 ends.
[0066] Fig. Figure 8 is a flowchart of process steps for controlling a robot according to various embodiments. Although the process steps are related to the systems of Fig. As described in Figures 1-6, the person skilled in the art will understand that any system designed to carry out the process steps in any order falls within the scope of the present embodiments.
[0067] As shown, a method 800 begins at step 802, wherein the data generator 118 receives a set of images. In some embodiments, the set of images may include one or more image data sets that include or are associated with a robot, such as images taken by cameras mounted on the robot or elsewhere within various physical and / or simulated environments.
[0068] In step 804, the data generator 118, using the trained image-generating model 119, creates a set of enhanced images based on the received set of images. As above in conjunction with Fig. As described in section 4, in some embodiments, the data generator 118 can repeatedly input an image from the set of images, along with text describing an associated robot task and an extension to be applied to the image, into the image-generating model 119. The text can describe the extension and the robot task in any suitable way, including with any level of detail. In some embodiments, the text can be generated using one or more templates or in any other technically feasible way. Given the image and the text, the image-generating model 119 produces an extended image, which may be included in the set of extended images.
[0069] In step 806, the model trainer 116 (or another model training application) trains the policy model or strategy model 150 using the set of extended images. The strategy model 150 can be trained in any technically feasible way in some embodiments, such as using an imitative learning method, as described above in conjunction with Fig. 4 described.
[0070] In step 808, the robot control application 146 controls a robot (e.g., the robot 160) using the trained strategy model. As described, in some embodiments, the robot control application 146 can control the robot 160 based on sensor data received from the sensors 180. The sensor data may include images captured by one or more cameras mounted on the robot 160 and / or within its environment. Given sensor data as input, the strategy model 150 generates an action, which is a command to control the robot 160 to perform at least part of a task. In some embodiments, the robot control application 146 can delegate the action to a lower-level controller, such as a PID controller or a PD controller, which controls the robot 160's actuators according to the action.
[0071] In summary, methods for generating augmented image data and training machine learning models using this augmented image data are disclosed. In some embodiments, a trained image-generating model takes images and text describing augmentations as input, and the image-generating model generates augmented images based on the input images, depth and semantic features extracted from the input images, and the text describing the augmentations. The image-generating model includes three diffusion modules. For a given input image and text, a first diffusion module is used to generate a feature map based on the input image and text. A second diffusion module is used to generate a second feature map based on the image, depth features extracted from the image, and the text.A third diffusion module is used to generate a third feature map based on the image, semantic features extracted from the image, and the text. A decoder processes the first, second, and third feature maps to produce an augmented image. Any number of augmented images can be generated according to the preceding steps for inclusion in a training dataset. A machine learning model, such as a strategy or policy model for controlling a robot, can then be trained using the training dataset. Once trained, the machine learning model can be used to perform one or more tasks. For example, a trained strategy model can be used to control a robot within a real-world or virtual environment.
[0072] At least one technical advantage of the disclosed methods compared to the prior art is that the disclosed methods generate augmented images that can provide diverse datasets for training machine learning models, such as strategy models for controlling robots. Using augmented images generated according to the disclosed methods, a strategy model can be trained to control a robot and perform a task more successfully than with strategy models trained using conventional approaches. In particular, the augmented images preserve depth and semantic information from the input images, which is useful for training the strategy model to perform tasks correctly. These technical advantages represent one or more technological improvements over prior art approaches. 1. In some embodiments, a computer-implemented method for training machine learning models comprises processing one or more input images using a trained image-generating model to produce one or more augmented images, wherein the trained image-generating model generates each augmented image contained in the one or more augmented images, depending on an input image contained in the one or more input images, depth information associated with the input image, semantic information associated with the input image, and text describing an augmentation to be performed on the input image; and performing, based on the one or more augmented images, one or more operations to train an untrained machine learning model to produce a trained machine learning model. 2. Computer-implemented method according to sentence 1, wherein the trained image-generating model comprises a first trained machine learning model that extracts the depth information from the input image; and a second trained machine learning model that extracts the semantic information from the input image. 3. A computer-implemented method according to sentence 1 or 2, wherein the generation of each augmented image contained in the one or more augmented images, depending on the input image contained in the one or more input images, comprises generating a first feature map using a first trained diffusion model depending on the input image and the text, generating a second feature map using a second trained diffusion model depending on the input image, the depth information associated with the input image, and the text, generating a third feature map using a third trained diffusion model depending on the input image, the semantic information associated with the input image, and the text; and generating the augmented image based on the first feature map, the second feature map, and the third feature map. 4. Computer-implemented method according to one of sentences 1-3, wherein generating the extended image comprises processing the first feature map, the second feature map and the third feature map using at least one decoder to generate the extended image. 5. Computer-implemented method according to one of sentences 1-4, wherein the one or more input images comprise a plurality of sets of images from one or more real environments and / or from one or more simulated environments. 6. Computer-implemented method according to one of sentences 1-5, wherein the text describes a robot task and / or a physical environment and / or a virtual environment and / or a domain. 7. Computer-implemented method according to one of sentences 1-6, wherein the one or more operations for training the untrained machine learning model include training the untrained machine learning model using the one or more input images. 8. Computer-implemented method according to one of theorems 1-7, wherein the trained model is trained to machine learning to generate actions for controlling a robot to perform at least one task. 9. Computer-implemented method according to any of theorems 1-8, wherein the trained machine learning model is trained to process one or more additional images to generate one or more actions that cause a robot to move. 10. Computer-implemented method according to any of sentences 1-9, further comprising performing, based on one or more additional images and a reconstruction loss, one or more training operations to train an image-generating model in order to generate the trained image-generating model. 11. In some embodiments, one or more non-volatile, computer-readable media store instructions which, when executed by at least one processor, cause the at least one processor to perform the steps of processing one or more input images using a trained image-generating model to produce one or more augmented images, wherein the trained image-generating model generates each augmented image contained in the one or more augmented images, depending on an input image contained in the one or more input images, depth information associated with the input image, semantic information associated with the input image, and text describing an augmentation to be performed on the input image, and performing, based on the one or more augmented images, one or more operations.to train an untrained machine learning model, in order to create a trained machine learning model. 12. The one or more non-volatile computer-readable media according to sentence 11, wherein the trained image-generating model comprises a first trained machine learning model that extracts the depth information from the input image and a second trained machine learning model that extracts the semantic information from the input image. 13. The one or more non-volatile, computer-readable media according to sentence 11 or 12, wherein the generation of each augmented image contained in the one or more augmented images, depending on the input image contained in the one or more input images, comprises generating a first feature map using a first trained diffusion model depending on the input image and the text, generating a second feature map using a second trained diffusion model depending on the input image, the depth information associated with the input image, and the text, generating a third feature map using a third trained diffusion model depending on the input image, the semantic information associated with the input image, and the text, and generating the augmented image based on the first feature map.the second feature card and the third feature card. 14. The one or more non-volatile computer-readable media according to any one of sentences 11-13, wherein the first trained diffusion model, the second trained diffusion model and / or the third trained diffusion model includes a ControlNet model. 15. The one or more non-volatile, computer-readable media according to any one of sentences 11-14, wherein the text describes a robot task and / or a physical environment and / or a virtual environment and / or a domain. 16. The one or more non-volatile, computer-readable media according to any one of sentences 11-15, wherein the trained machine learning model is trained to generate actions for controlling a robot to perform at least one task. 17. The one or more non-volatile, computer-readable media according to any of sentences 11-16, wherein the trained machine learning model is trained to process one or more additional images to generate one or more actions that cause a robot to move. 18. The one or more non-volatile computer-readable media according to any one of sentences 11-17, wherein the one or more operations for training the untrained machine learning model include training the untrained machine learning model using a loss with respect to imitative learning. 19. The one or more non-volatile computer-readable media according to any of sentences 11-18, wherein the semantic information identifies at least one object contained in the input image. 20. In some embodiments, a system includes a memory that stores instructions;and one or more processors which, when executing the instructions, are configured to perform the steps of: processing one or more input images using a trained image-generating model to produce one or more augmented images, wherein the trained image-generating model generates each augmented image contained in the one or more augmented images, depending on an input image contained in the one or more input images, depth information associated with the input image, semantic information associated with the input image, and text describing an augmentation to be performed on the input image; and performing, based on the one or more augmented images, one or more operations to train an untrained machine learning model to produce a trained machine learning model.
[0073] All combinations of claim elements listed in the claims and / or of elements described in this application fall in any way within the intended scope of this disclosure and the protection.
[0074] The descriptions of the various embodiments serve for illustrative purposes but do not claim to be exhaustive or limited to the disclosed embodiments. Many modifications and variations are obvious to the person skilled in the art without departing from the scope and spirit of the described embodiments.
[0075] Aspects of the present embodiments may be designed as a system, method, or computer program product. Accordingly, aspects of the present disclosure may take the form of a purely hardware implementation, a purely software implementation (including firmware, resident software, microcode, etc.), or an embodiment that combines software and hardware aspects, which may be generally referred to herein as a "module," "system," or "computer." Furthermore, all hardware and / or software techniques, methods, functions, components, engines or machines, modules, or systems described in the present disclosure may be implemented as a circuit or circuit group. In addition, aspects of the present disclosure may take the form of a computer program product implemented in one or more computer-readable media containing computer-readable program code.
[0076] Any combination of one or more computer-readable media can be used. The computer-readable medium can be a computer-readable signaling medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not exclusively, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.More specific examples (a non-exhaustive list) of computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, random-access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only storage device (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. For the purposes of this document, computer-readable storage medium can be any tangible medium capable of containing or storing a program for use by or in conjunction with a command-execution system, device, or apparatus.
[0077] Aspects of the present disclosure are described above with reference to flowchart diagrams and / or block diagrams of processes, devices (systems), and computer program products according to embodiments of the disclosure. It is understood that each block of the flowchart diagrams and / or block diagrams, and combinations of blocks in the flowchart diagrams and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be fed to a processor of a general-purpose computer, a specialized computer, or other programmable data processing device to create a machine. When executed by the processor of the computer or other programmable data processing device, the instructions enable the implementation of the functions / actions specified in the flowchart and / or block diagram.Such processors can be, without restriction, general-purpose processors, special-purpose processors, application-specific processors, or field-programmable gate arrays.
[0078] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, procedures, and computer program products according to various embodiments of the present disclosure. In this respect, each block in the flowchart or block diagrams can represent a module, segment, or section of code comprising one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions specified in the blocks may occur out of the order shown in the figures. For example, two blocks shown consecutively may in reality be executed essentially simultaneously, or the blocks may sometimes be executed in reverse order, depending on the functionality involved.It is also noted that each block of the block diagrams and / or flowchart representation and combinations of blocks in the block diagrams and / or flowchart representation may be implemented by special hardware-based systems that perform the specified functions or actions, or by combinations of special hardware and computer instructions.
[0079] While the foregoing relates to embodiments of the present disclosure, other and further embodiments of the disclosure may be developed without departing from the basic scope of the disclosure, and the scope of the disclosure is determined by the following claims.
Claims
[1] Computer-implemented method for training machine learning models, wherein the method comprises: Processing one or more input images using a trained image-generating model to produce one or more augmented images, wherein the trained image-generating model produces each augmented image contained in the one or more augmented images, depending on an input image contained in the one or more input images, depth information associated with the input image, semantic information associated with the input image, and text describing an augmentation to be performed on the input image; and Performing one or more operations, based on the one or more enhanced images, to train an untrained machine learning model in order to generate a trained machine learning model. [2] Computer-implemented method according to claim 1, wherein the trained image-generating model comprises: a first trained machine learning model that extracts the depth information from the input image; and a second trained machine learning model that extracts the semantic information from the input image. [3] Computer-implemented method according to claim 1 or 2, wherein generating each augmented image contained in the one or more augmented images, depending on the input image contained in the one or more input images, comprises: Generating an initial feature map using an initial trained diffusion model, depending on the input image and text; Generating a second feature map using a second trained diffusion model depending on the input image, the depth information associated with the input image, and the text; Generating a third feature map using a third trained diffusion model, depending on the input image, the semantic information associated with the input image, and the text; and Generating the extended image based on the first feature map, the second feature map, and the third feature map. [4] Computer-implemented method according to claim 3, wherein generating the augmented image comprises processing the first feature map, the second feature map and the third feature map using at least one decoder to generate the augmented image. [5] Computer-implemented method according to one of the preceding claims, wherein the one or more input images comprise a plurality of sets of images from one or more real environments and / or from one or more simulated environments. [6] Computer-implemented method according to any of the preceding claims, wherein the text describes a robot task and / or a physical environment and / or a virtual environment and / or a domain. [7] Computer-implemented method according to any of the preceding claims, wherein the one or more operations for training the untrained machine learning model include training the untrained machine learning model using the one or more input images. [8] Computer-implemented method according to any of the preceding claims, wherein the trained model is trained to machine learning to generate actions for controlling a robot to perform at least one task. [9] Computer-implemented method according to any of the preceding claims, wherein the trained machine learning model is trained to process one or more additional images to generate one or more actions that cause a robot to move. [10] Computer-implemented method according to one of the preceding claims, further comprising performing, based on one or more additional images and a reconstruction loss, one or more training operations to train an image-generating model in order to generate the trained image-generating model. [11] One or more non-volatile, computer-readable media that store instructions which, when executed by at least one processor, cause that at least one processor to perform the steps of: Processing one or more input images using a trained image-generating model to produce one or more augmented images, wherein the trained image-generating model produces each augmented image contained in the one or more augmented images, depending on an input image contained in the one or more input images, depth information associated with the input image, semantic information associated with the input image, and text describing an augmentation to be performed on the input image; and Performing one or more operations, based on the one or more enhanced images, to train an untrained machine learning model in order to generate a trained machine learning model. [12] The one or more non-volatile computer-readable media according to claim 11, wherein the trained image-generating model comprises: a first trained machine learning model that extracts the depth information from the input image; and a second trained machine learning model that extracts the semantic information from the input image. [13] The one or more non-volatile computer-readable media according to claim 11 or 12, wherein generating each augmented image contained in the one or more augmented images, depending on the input image contained in the one or more input images, comprises: Generating an initial feature map using an initial trained diffusion model, depending on the input image and text; Generating a second feature map using a second trained diffusion model depending on the input image, the depth information associated with the input image, and the text; Generating a third feature map using a third trained diffusion model, depending on the input image, the semantic information associated with the input image, and the text; and Generating the extended image based on the first feature map, the second feature map, and the third feature map. [14] One or more non-volatile computer-readable media according to claim 13, wherein the first trained diffusion model, the second trained diffusion model and / or the third trained diffusion model comprises a ControlNet model. [15] One or more non-volatile computer-readable media according to any one of claims 11 to 14, wherein the text describes a robot task and / or a physical environment and / or a virtual environment and / or a domain. [16] One or more non-volatile computer-readable media according to any one of claims 11 to 15, wherein the trained machine learning model is trained to generate actions for controlling a robot to perform at least one task. [17] One or more non-volatile computer-readable media according to any one of claims 11 to 16, wherein the trained machine learning model is trained to process one or more additional images to generate one or more actions that cause a robot to move. [18] One or more non-volatile computer-readable media according to any one of claims 11 to 17, wherein one or more operations for training the untrained machine learning model include training the untrained machine learning model using a loss with respect to imitative learning. [19] One or more non-volatile computer-readable media according to any one of claims 11 to 18, wherein the semantic information identifies at least one object contained in the input image. [20] System encompassing: a memory that stores instructions; and one or more processors which, when executing the instructions, are configured to perform the steps of: Processing one or more input images using a trained image-generating model to produce one or more augmented images, wherein the trained image-generating model generates each augmented image contained in the one or more augmented images, depending on an input image contained in the one or more input images, depth information associated with the input image, semantic information associated with the input image, and text describing an augmentation to be performed on the input image, and Performing one or more operations, based on the one or more enhanced images, to train an untrained machine learning model in order to generate a trained machine learning model.