Using an intermediary machine learning model to steer pretrained machine learning model output
By using an intermediary machine learning model to steer and condition the activations of pretrained models, the challenges of training multimodal models with scarce data and high computational costs are addressed, allowing for efficient performance of new tasks without sacrificing native task performance.
Patent Information
- Application Number
- PCT/GB2024/053182
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-26
AI Technical Summary
The development of multimodal machine learning models is challenging due to the scarcity of high-quality multimodal data and the computational expense of training large models, making it impractical to train new models from scratch for every new task.
An intermediary machine learning model is trained and used to facilitate communication and conditioning between different pretrained machine learning models by interleaving its layers with those of the pretrained models, allowing raw activations to be accessed and steered without requiring additional training of the pretrained models.
This approach enables the performance of new tasks using pretrained models without costly retraining, while maintaining the performance of native tasks, and reduces the computational resources needed for training.
Smart Images

Figure GB2024053182_26062025_PF_FP_ABST
Abstract
Description
USING AN INTERMEDIARY MACHINE LEARNING MODEL TO STEER PRETRAINED MACHINE LEARNING MODEL OUTPUTATTORNEY REFERENCE: DEEP-0004-WO-01 Background
[0001] Information used by a machine such as a computer or a mechanical agent such as a robot to perform a task may be encoded in a variety of modalities such as text, images, video, and audio, to name a few. Understanding and interacting with an environment may benefit from the ability to process and reason across these modalities. While recent advances in deep learning have led to significant progress in unimodal tasks, the development of multimodal machine learning model models remains a challenge. One obstacle is the scarcity of high-quality multimodal data. While there is an abundance of unpaired unimodal data, such as text-only or video-only data, collecting and labeling multimodal data is often prohibitively expensive and time-consuming. This is particularly true in robotics, where real-world environments are complex and dynamic, and collecting multi-sensory data can be challenging. In addition, training and / or fine-tuning large multimodal machine learning models can be computationally expensive and require significant resources. This can make it difficult, if not impractical, to train new multimodal machine learning models from scratch for every new task.Summary
[0002] A variety of available pretrained machine learning models (e.g., foundation models) have already been trained on large amounts of unimodal or multimodal data to perform various “native” tasks, and more are being trained every day. It would be beneficial to leverage these pretrained machine learning models to perform new tasks without requiring costly multimodal (or more generally, inter-model) training or sacrificing individual model performance of native tasks. However, different pretrained machine learning models are often developed and trained to perform their native tasks by different institutions, research groups, or even individuals. This can pose a challenge, as no specific assumptions about the pretrained machine learning models’ architectures or their training procedures can be made.
[0003] When tokens are applied as initial inputs across pretrained machine learning models, activations are generated midstream, e.g., between layers of the pretrained machine learning models. In various implementations, “tokens” and “activations” may take various forms, such as encodings and / or embeddings that represent, for instance, units of text such as words,sentences or paragraphs, individual images or portions thereof, image frames from video, audio frames from audio sensor data, sensor data sampled at particular time points, snippets of source code, and so forth.
[0004] Implementations are described herein for what is referred to as “GATS” (Gather, Attend, Scatter). This includes training and / or using an “intermediary” machine learning model to enable communication and / or conditioning between different pretrained machine learning models to perform task(s) that are not native to those pretrained machine learning models. More particularly, but not exclusively, implementations are described herein for interleaving layer(s) of the trainable intermediary machine learning model into layer(s) of pretrained machine learning models so that raw activations between layers of the pretrained machine learning models can be accessed by — and in many cases influenced or “steered” by — the intermediary machine learning model. Tapping into these raw activations (referred to herein as “gathering”) enables the intermediary machine learning model to be trained to facilitate communication and / or conditioning between the pretrained machine learning models, without requiring costly additional training and / or fine-tuning of the pretrained machine learning models themselves.Brief Description of the Drawings
[0005] Fig. 1 schematically depicts an example environment in which disclosed techniques may be employed, in accordance with various implementations.
[0006] Fig. 2 depicts an example robot, in accordance with various implementations.
[0007] Fig. 3 A, Fig. 3B, and Fig. 3C schematically depict an example of how activations may be selected for processing using an intermediary machine learning model configured with selected aspects of the present disclosure.
[0008] Fig. 4 depicts an example of how an intermediary machine learning model configured with selected aspects of the present disclosure may be interleaved with layers of one or more pretrained machine learning models.
[0009] Fig. 5 depicts another example of how an intermediary machine learning model configured with selected aspects of the present disclosure may be interleaved with layers of one or more pretrained machine learning models.
[0010] Fig. 6 depicts an example of how an intermediary machine learning model configured with selected aspects of the present disclosure may be interleaved with layers of pretrained machine learning models that are used to control a robot.
[0011] Fig. 7 depicts an example method for practicing selected aspects of the present disclosure.
[0012] Fig. 8 depicts another example method for practicing selected aspects of the present disclosure.
[0013] Fig. 9 schematically depicts an example architecture of a computer system.Detailed Description
[0014] Implementations are described herein for training and / or using an “intermediary” machine learning model to enable communication and / or conditioning between different pretrained machine learning models to perform task(s) that are not native to those pretrained machine learning models. More particularly, but not exclusively, implementations are described herein for interleaving layer(s) of the trainable intermediary machine learning model into layer(s) of pretrained machine learning models so that raw activations between layers of the pretrained machine learning models can be accessed by — and in many cases influenced or “steered” by — the intermediary machine learning model. Tapping into these raw activations enables the intermediary machine learning model to be trained to facilitate communication and / or conditioning between the pretrained machine learning models, without requiring costly additional training and / or fine-tuning of the pretrained machine learning models themselves.
[0015] Interleaving intermediary machine learning models with pretrained machine learning models as described herein allows raw activations from initial layer(s) of multiple different pretrained machine learning models to be intercepted and processed together using the intermediary machine learning model. In many cases the intermediary machine learning model attends across these raw activations, which means activations from one pretrained machine learning model are conditioned by activations from another pretrained machine learning model, and vice versa. In many instances, the resulting conditioned or “steered” activations may be applied as inputs to subsequent layer(s) of the pretrained machine learning models.Consequently, downstream processing by each pretrained machine learning model is conditioned or “steered” at least in part by other pretrained machine learning model(s).
[0016] In various implementations, the intermediary machine learning model may have significantly fewer (e.g., by order(s) of magnitude) parameters than one or more of the pretrained machine learning models, which can have tens, if not hundreds, of billions of parameters, and several or even dozens of layers. Additionally, the intermediary machine learning model may have significantly fewer layers than the pretrained machine learning model(s). Accordingly, during training, weights of the pretrained machine learning models may be held constant or frozen, while weights of the intermediary machine learning model may be trained. Due to its smaller size (in layers and / or overall number of parameters), the training intermediary machine learning model may require significantly less time and / or resources than fine-tuning thepretrained machine learning model(s). This also avoids degrading the pretrained machine learning models’ performance of their “native” tasks.
[0017] Techniques described herein may be applied to various different tasks. One non-limiting example is robot control in which a robot (or more generally, a “mechanical agent”) is controlled based on a high-level command issued in the form of a natural language request. In some cases, the robot may be controlled using as many as three or more different pretrained models: a language model such as a large language model (LLM) to process the natural language input; one or more perception models to process sensor data captured by one or more sensors on the robot or in its environment, such as a vision model to process image(s) and another perception model to process Light Detection and Ranging (LIDAR) data; and an action model that is trained to process current proprioception values of the robot to predict the robot’s future proprioception values (and thereby control the robot). Raw activations from the three or more pretrained machine learning models may be processed together (e.g., attended across each other) using the intermediary machine learning model to generate, as output of the intermediary machine learning model, steered activations.
[0018] In some cases, tokens applied across, and / or activations generated by, the different pretrained machine learning models may be embeddings that have different dimensionalities. In order for all these different-dimensionality embeddings to be processed using the intermediary machine learning model, the embeddings from at least some of the pretrained machine learning models may be projected into a single common dimensionality. For example, an activation / embedding x, of a sequence of activations / embeddings (n, 2, ... , xt) generated using a pretrained machine learning model m may be projected using a projection pminto a common dimensionality d, such that a size of >m(x;) = d.
[0019] Once the projected activations are processed using the intermediary machine learning model, at least some of the resulting steered activations may be applied as inputs to subsequent layer(s) of at least some of the pretrained machine learning models. For instance, in the working robotics example, steered activations from the perception machine learning model(s) and action machine learning model may be applied as inputs to subsequent layer(s) of these models. Consequently, subsequent layers of the perception machine learning model(s) (e.g., to process images / video and LIDAR) and action machine learning model are influenced by cross-model conditioning. In some implementations, because only a single natural language input is providedat the outset of the episode, the activations of the language model that are steered by the intermediary machine learning model may be discarded, and subsequent layers of the language model may be applied to raw activations of the initial layers of the language model (i.e., the language model activations may remain in their raw form). (This may be true more generally for any model or modality for which steered activations are not needed or desired.) Consequently, the processing performed using the language model may be unaffected by the intermediary machine learning model. In some implementations, the activations of the language model — or any model that is elected to be unsteered — can be cached for one or more instructions, enabling the use of large numbers of parameters without sacrificing inference speed.
[0020] Several examples described herein relate to using techniques described herein to control a machine or apparatus such as a robot. However, this is not meant to be limiting. In various implementations, techniques described herein may be applied in other contexts, including but not limited to: digital audio, image or video processing (e.g., enhancement, analysis); separation of sources in speech signals, or speech recognition; encoding data for reliable and / or efficient transmission or storage (and corresponding decoding); encrypting / decrypting or signing electronic communications; generating keys; optimizing load distribution in a computer network; processing data obtained from physiological sensors, e.g. for medical diagnosis; and providing a genotype estimate based on an analysis of DNA samples.
[0021] For example, the first set of one or more tokens may comprise audio frames, pixels from an image, or image frames, and the one or more operations may comprise digital audio, image or video processing, for example for the purposes of enhancement or analysis. Alternatively for example, the first set of one or more tokens may comprise audio frames, and the one or more operations may comprise separation of sources of speech, or speech recognition, from the audio. The first set of one or more tokens may comprise data packets, and the one or more operations may comprise encoding or decoding for transmission or storage, or encryption or decryption. The first set of one or more tokens may comprise data obtained from physiological sensors and the one or more operations may comprise determining a medical diagnosis.
[0022] The intermediary machine learning model described herein is also not limited to use with robotics, nor to the modalities depicted in Figs. 1-6. In various implementations, an intermediary machine learning model configured with selected aspects of the present disclosure may be symmetrical in a sense that no modality needs to be treated in any special way. Each unimodalmodel processes its own inputs and resulting activations may be used to aid all the other models. The intermediary machine learning model is exposed to these activations taken from different layers, and is trained on how to process them (as opposed to hard-coding which layers are used, e.g., middle ones versus ultimate ones). That being said, through a modest set of hyperparameters, the intermediary machine learning model can be set up to act as highly asymmetrical specialist architectures such as vision-to-text cross-attention models.
[0023] Intermediary machine learning models configured with selected aspects of the present disclosure may also exhibit lightweight inference overhead. A given input embedding may be processed by only one corresponding unimodal model, and hence the inference speed for a considered modality does not depend on sizes of other unimodal models. The processing is, however, conditioned on the other modalities thanks to steering by the intermediary machine learning model. But because only already processed activations from other modalities are gathered by the intermediary model, the only added overhead is caused by applying layers of the intermediary machine learning model. This is relatively lightweight, as the size of the intermediary machine learning model projected embeddings d are small and intermediary local context lengths , , ... , NM are short.
[0024] One advantage of intermediary machine learning models described herein is that they can have significantly fewer layers than the pretrained models they are interleaved with. That means that modalities processed with smaller models would still be relatively fast to process, while still being exposed via intermediary machine learning model layers described herein to rich representations from larger models, such as large scale pretrained LLMs.
[0025] Another advantage of intermediary machine learning models described herein is efficient training. The entire multimodal model including all unimodal transformers and the intermediary machine learning model can be trained very efficiently, since the intermediary machine learning model can be parallelized the same way vanilla transformers are. Additionally, thanks to steering, the intermediary machine learning model may access information stored in weights of frozen pretrained models without the need to spend device memory on their updates.
[0026] In various implementations, layer(s) of the intermediary machine learning model(s) may be defined by hyperparameters such as: number of pretrained (e.g., unimodal) networks connected A / ; local context lengths Ni, Ni, ... , NM,' projected embedding size d; a nonempty subset S of {1, 2, ... , M} specifying which modalities are steered; and / or hyperparametersdescribing the transformer used to process embeddings in the local context G. In some implementations, an intermediary machine learning model configured with selected aspects of the present disclosure may include K gather-attention-scatter layers. In some cases all layers share the same hyperparameters, but in other cases, all or some of them can be different. Linear transformations (2A / of them) may be used for all projection (pm, described below) and rm(described below) functions. Each of the gating functions gmmay also be a linear transformation additionally followed by a layer norm. As such all pm, rm, and gmfunctions depend only on the projected embedding size d (and input embeddings sizes in many instances) and hence do not require additional hyperparameters.
[0027] Intermediary machine learning models described herein may be trained in various environments based on various datasets. In some implementations, various amounts (e.g., 20,000) of episodes from the video game PONG may be sampled as training data. Additionally or alternatively, in some implementations, a simulated tabletop manipulation scenario may be used for training. Such a scenario may include, for instance, a robot arm with a cylindrical endeffector constrained to move in a 2D plane, and some number of objects (e.g., blocks) with different shapes, sizes, and / or colors. The goal in such a scenario may be presented as a language instruction (e.g., “push the blue triangle to the top left corner”).
[0028] Fig. 1 is a schematic diagram components that can cooperate to carry out selected aspects of the present disclosure, in accordance with various implementations. The various components depicted in Fig. 1, particularly those components forming an intermediary system 110, a natural language processing (NLP) system 120, a vision system 130, and a proprioception system 140, may be implemented using any combination of hardware and software. The components of Fig.1 are depicted as being communicatively coupled with each other via one or more networks 199, which may include one or more personal area networks, local area networks, and / or wide area networks (e.g., the Internet). However, this is not meant to be limiting. Various aspects of the present disclosure that are described as being performed by and / or stored on systems 110, 120, 130, and / or 140 can alternatively be performed by and / or stored on a single system, such as intermediary system 110, or on any combinations of systems 110, 120, 130, and / or 140.
[0029] In some implementation, techniques described herein may be used to control various types of machines or apparatus. For example, in some implementations, a robot 100 may be in communication with systems 110, 120, 130, and / or 140. In various implementations, and / or allor parts of systems 110, 120, 130, and / or 140 may be implemented onboard robot 100. Other types of machines or apparatus that are not depicted in Fig. 1 may also be controlled using selected aspects of the present disclosure, such as autonomous vehicles, industrial equipment, climate control systems, medical systems and / or devices, video games, and so forth.
[0030] Robot 100 may take various forms, including but not limited to a telepresence robot (e.g., which may be as simple as a wheeled vehicle equipped with a display and a camera), a robot arm, a multi-pedal robot such as a “robot dog,” an aquatic robot, a wheeled device, a submersible vehicle, an unmanned aerial vehicle (“UAV”), and so forth. One non-limiting example of a mobile robot arm is depicted in Fig. 2. In various implementations, robot 100 may include logic 102. Logic 102 may take various forms, such as a real time controller, one or more processors, one or more field-programmable gate arrays (“FPGA”), one or more application-specific integrated circuits (“ASIC”), and so forth. In some implementations, logic 102 may be operably coupled with memory 103. Memory 103 may take various forms, such as random-access memory (“RAM”), dynamic RAM (“DRAM”), read-only memory (“ROM”), Magnetoresistive RAM (“MRAM”), resistive RAM (“RRAM”), NAND flash memory, and so forth. In some implementations, a robot controller may include, for instance, logic 102 and memory 103 of robot 100.
[0031] In some implementations, logic 102 may be operably coupled with one or more joints 104-1 to 104-N, one or more end effectors 106, and / or one or more sensors 108-1 to 108- M, e.g., via one or more buses 109. As used herein, “joint” 104 of a robot may broadly refer to actuators, motors (e.g., servo motors), shafts, gear trains, pumps (e.g., air or liquid), pistons, drives, propellers, flaps, rotors, or other components that may create and / or undergo propulsion, rotation, and / or motion. Some joints 104 may be independently controllable, although this is not required. In some instances, the more joints robot 100 has, the more degrees of freedom of movement it may have.
[0032] As used herein, “end effector” 106 may refer to a variety of tools that may be operated by robot 100 in order to accomplish various tasks. For example, some robots may be equipped with an end effector 106 that takes the form of a claw with two opposing “fingers” or “digits.” Such a claw is one type of “gripper” known as an “impactive” gripper. Other types of grippers may include but are not limited to “ingressive” (e.g., physically penetrating an object using pins, needles, etc.), “astrictive” (e.g., using suction or vacuum to pick up an object), or “contigutive”(e.g., using surface tension, freezing or adhesive to pick up object). More generally, other types of end effectors may include but are not limited to drills, brushes, force-torque sensors, cutting tools, deburring tools, welding torches, containers, trays, and so forth. In some implementations, end effector 106 may be removable, and various types of modular end effectors may be installed onto robot 100, depending on the circumstances. Some robots, such as some telepresence robots, may not be equipped with end effectors. Instead, some telepresence robots may include displays to render visual representations of the users controlling the telepresence robots, as well as speakers and / or microphones that facilitate the telepresence robot “acting” like the user.
[0033] Sensors 108-1 to 108-M may take various forms, including but not limited to 3D laser scanners (e.g., light detection and ranging, or “LIDAR”) or other 3D vision sensors (e.g., stereographic cameras used to perform stereo visual odometry) configured to provide depth measurements, two-dimensional cameras (e.g., RGB, infrared), light sensors (e.g., passive infrared), force sensors, pressure sensors, pressure wave sensors (e.g., microphones), proximity sensors (also referred to as “distance sensors”), depth sensors, torque sensors, barcode readers, radio frequency identification (“RFID”) readers, radars, range finders, accelerometers, gyroscopes, compasses, position coordinate sensors (e.g., global positioning system, or “GPS”), speedometers, edge detectors, Geiger counters, and so forth. While sensors 108-1 to 108-M are depicted as being integral with robot 100, this is not meant to be limiting.
[0034] In some implementations, intermediary system 110, NLP system 120, vision system 130, and / or proprioception system 140 may include one or more computing devices cooperating to perform selected aspects of the present disclosure. An example of such a computing device is depicted schematically in Fig. 9. In some implementations, one or more of systems 110, 120, 130, and / or 140 may include one or more servers forming part of what is often referred to as a “cloud” infrastructure, or simply “the cloud.” Alternatively, one or more components of systems 110, 120, 130, and / or 140 may be operated by logic 102 of robot 100.
[0035] Machine learning model(s) described herein may take various forms, including, but not limited to, generative language model(s) (sometimes referred to as “large language models,” or “LLMs”) such as PaLM, PALM-E (described in “PaLM-E: An Embodied Multimodal Language Model”, arXiv:2303.0337), BERT, LaMDA, Meena, Gemini, Flamingo (described in “Flamingo: a Visual Language Model for Few-Shot Learning”, arXiv:2204.14198), and / or any other generative language model, such as any other generative model that is encoder-only based,decoder-only based, sequence-to-sequence based and that optionally includes an attention mechanism or other memory. In generative model form, machine learning model(s) may have hundreds of millions, or even hundreds of billions of parameters. In some implementations, machine learning model(s) may include a multi-modal model such as a vision language model (VLM, e.g., Flamingo), visual question answering (VQA) model, and / or a diffusion model (e.g., stable diffusion), which can have any of the aforementioned architectures, and which can be used to process multiple modalities of data, particularly images and text, and / or images and audio for example, to generate one or more modalities of output. An example of a machine learning model 146 that may be used by proprioception prediction process 142 is described in “RT-1: Robotics Transformer for Real-World Control at Scale” (arXiv:2212.06817). Another example is described in “RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control” (arXiv:2307.15818).
[0036] Various combinations of pretrained models may be interleaved with intermediary model layers as described herein. In some implementations, the LLM 126 may be the Chinchilla language model described in “Training Compute-Optimal Large Language Models” (arXiv:2203.15556). In some implementations, the “perception” or “vision” model 636 may be the Phenaki video model described in “Phenaki: Variable Length Video Generation from Open Domain Textual Description” (arXiv:2210.02399). In some implementations, the image tokenizer from the Parti model, described in “Scaling Autoregressive Models for Content-Rich Text-to-Image Generation” (arXiv:2206.10789), may be used as the perception / vision model 136. In some implementations, a vision transformer such as that described in “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale” (arXiv:2010.11929) may be used as the vision model 136. For example, in Fig. 6, the Chinchilla language model can be used as LLM 626, the Phenaki model may be used as the perception / vision model 636, and the RT-1 model may be used as the action model 646. In some implementations, Gemini may be used as one of the pretrained models.
[0037] Intermediary system 110 may include an intermediary engine 112, a lead coordination engine 114, one or more intermediary machine learning models 116, and a parameter engine 118. Intermediary engine 112 may be configured to process various information using one or more intermediary machine learning models 116 to generate steered activations. The information processed by intermediary engine 112 may include, for instance, raw activations generated byinitial layer(s) of pretrained machine learning models. Based on this processing, intermediary engine 112 may generate steered activations that may be applied across subsequent layer(s) of at least some of the same pretrained machine learning models to generate steered machine learning model output that is conditioned based on activations from other pretrained machine learning models.
[0038] In some (but not all) implementations, intermediary engine 112 may also be configured to train intermediary machine learning model(s) 116 based on various signals. In other implementations, other components, such as any of those depicted in Fig. 1 or otherwise, may perform operations attributed herein to intermediary engine 112. In some implementations where steered outputs generated by a combination of pretrained machine learning models and intermediary machine learning model(s) 116 are used to operate a machine or robot, an outcome of the machine or robot operation may be used to determine whether or not the steered output (e.g., sequence of steered output activations) should be used as a training example (e.g., paired with a sequence of initial input tokens) to train intermediary machine learning model(s) 116), or whether the input / output pair should be used as a positive or negative training example. For instance, if a robot is controlled using the architecture depicted in Fig. 6 and successfully performs the requested operation, the sequence of input tokens 660 and a corresponding sequence of output tokens (not depicted in Fig. 6) may be flagged as a positive training example pair, and may be used by intermediary engine 112 to train one or more intermediary machine learning model(s), such as the intermediary machine learning model 616 depicted in Fig. 6.
[0039] In various implementations, layer(s) of intermediary machine learning model(s) 116 may be defined by hyperparameters such as: number of pretrained (e.g., unimodal) networks connected A / ; local context lengths M, , ... , Mr; projected embedding size d, a nonempty subset S of {1, 2, ... , M} specifying which modalities are steered; and / or hyperparameters describing the transformer used to process embeddings in the local context G. In some implementations, intermediary machine learning model 116 may include K gather-attention- scatter layers. In some cases all layers share the same hyperparameters, but in other cases, all or some of them can be different. Linear transformations (2M of them) may be used for all projection (pm, described below) and rm(described below) functions. Each of the gating functions gm may also be a linear transformation additionally followed by a layer norm. As such all pm, rm,and gmfunctions depend only on the projected embedding size d (and input embeddings sizes in many instances) and hence do not require additional hyperparameters.
[0040] Referring back to Fig. 1, lead coordination engine 114 may be configured to coordinate with other coordination engines (e.g., 124, 134, 144) to facilitate the exchange of information amongst systems 110, 120, 130, and / or 140. In particular, the various coordination engines 114, 124, 134, 144 may cooperate to exchange the raw activations generated based on initial layer(s) of various pretrained machine learning models (some of which are depicted in Fig. 1). Those raw activations may be processed (e.g., attended across) by intermediary engine 112 using intermediary machine learning model 116 as described previously to generate steered activations. The various coordination engines 114, 124, 134, 144 may additionally facilitate the exchange of these steered activations, e.g., so that the steered activations can be processed using subsequent layer(s) of at least some of the same pretrained machine learning models to generate steered output(s). In implementations where the pretrained machine learning models and intermediary machine learning model(s) 116 are hosted by the same computer system as intermediary engine 112, one or more of coordination engines 114, 124, 134, 144 may be omitted and / or combined with intermediary engine 112.
[0041] Parameter engine 118 may be operable, e.g., by a user 150 operating a client device 152, by an administrator of intermediary system 110, etc., to modify various parameters (e.g., hyper parameters) used by intermediary engine 112 when using intermediary machine learning model(s) 116 to process raw activations generated by other pretrained machine learning models. For instance, and as will be demonstrated in various figures, it may not always be necessary or desirable to steer the ultimate output of every single pretrained model providing data to intermediary system 110. Accordingly, user 150 may operate client device 152 to set a flag or other parameter that causes raw activations generated by initial layer(s) of pretrained machine learning model(s) — which may still be processed by intermediary engine 112 using intermediary machine learning model(s) 116 — to be applied across subsequent layer(s) of the corresponding pretrained machine learning models, rather than steered outputs. In other implementations, a machine learning model may be trained to predict which modalities should be steered in various local contexts.
[0042] Parameter engine 118 may be operable to modify other hyper parameters as well. One example is a number of pretrained machine learning models into which intermediary machinelearning model(s) 116 are to be interleaved. This may be selected, for instance, at the outset of training intermediary machine learning model(s) 116 to be used with particular combinations of pretrained machine learning models.
[0043] Another example of a parameter that can be adjusted using parameter engine 118 is a local context length of each intermediary machine learning model 116. In some implementations, parameter engine 118 may split the context length of a given intermediary machine learning model 116, evenly or otherwise, over all the different pretrained machine learning models with which the given intermediary machine learning model is interleaved. For instance, if an intermediary machine learning model 116 is interleaved among layers of three different pretrained machine learning models, then a context length of the intermediary machine learning model 116 may be split between at least some activations of each of the three pretrained machine learning models. In some cases this may cause the intermediary machine learning model to attend to embeddings / activations from all models (e.g., modalities) processed so far, even if recent inputs come from a single modal / modality. The memory allocated to each modal / modality can be relatively small, assuming that more long term processing is delegated to the larger pretrained machine learning models.
[0044] Parameter engine 118 may additionally or alternatively be operated to adjust parameters such as a projected embedding (e.g., activation) size d, and / or hyperparameters (e.g., transformer hyperparameters) of intermediary machine learning model 116 itself. These transformer hyperparameters may include, for instance, a number of layers, where those layers are interleaved within pretrained machine learning models, the temperature of the transformer, etc.
[0045] NLP system 120 may include an LLM selection engine 122, a respective coordination engine 124 (which was described previously), one or more pretrained LLMs 126 or other generative language models, and an LLM response engine 128. LLM selection engine 122 may be configured to select which LLM 126 should be used in any given circumstance, e.g., based on a context of user 150, a natural language request issued by user 150, and so forth. LLM response engine 128 may be configured to process natural language using LLM(s) 126 to generate LLM output. LLM output often (but not always) takes the form of a sequence of output tokens, such as embeddings that can be decoded into words.
[0046] Vision system 130 may include a perception process 132, a respective coordination engine 134 (described previously), and one or more perception machine learning models 136. Invarious implementations, perception process 132 may be configured to obtain a sequence of input tokens (e.g., embeddings) that are generated from visual data, such as a single images, sequences of images (e.g., video), LIDAR, etc. Perception process 132 may apply one or more perception machine learning models 136 to this vision data in order to generate various types of output, such as one or more output tokens (e.g., semantically rich embedding(s)) that represent visual features extracted from the input data and / or a larger semantic meaning of an entire image. These outputs may be used for various purposes, such as visual question answering (VQA), generating a control signal for controlling a machine and / or robot 100, etc. While vision system 130 in Fig. 1 is unimodal in that perception process 132 only processes vision data, this is not meant to be limiting; in various implementations, one or more perception machine learning models 136 may take the form a multimodal model such as a VLM or VQA model that processes both visual data and text (e.g., asking a question about the visual data) to provide a textual response.
[0047] Proprioception system 140 may be present in implementations where robot 100 is being controlled using techniques described herein. Proprioception system 140 may be omitted in other circumstances. Proprioception system 140 may include a proprioception prediction process 142, a respective coordination engine 144 (described previously), and one or more proprioception machine learning models 146. In various implementations, proprioception prediction process 142 may process input tokens indicative of a current (or past) proprioception values of robot 100, e.g., along with other data such as data indicative of a task to be performed, state data of the robot’s environment (e.g., generated by perception process 132) etc., to predict future proprioception values (e.g., actions) of robot 100. These future proprioception values may be used as commands for robot 100, e.g., so that robot 100 transitions from its current state into a new state in which robot 100 has the proprioception values predicted by proprioception prediction process 142.
[0048] The example pretrained machine learning model systems 120, 130, 140 depicted in Fig. 1 are for illustrative purposes only and are not meant to be limiting. Other types of pretrained machine learning models, unimodal and / or multimodal, may have their layer(s) interleaved with layers of intermediary machine learning model(s) 116. For instance, digital audio, image or video enhancement or analysis may be performed using respective pretrained machine learning models interleaved with intermediary machine learning model(s) 116 configured with selectedaspects of the present disclosure. Intermediary machine learning model(s) 116 may alternatively be interleaved with pretrained machine learning models that are configured to process audio data (e.g., recordings, audio waveforms, Fast Fourier Transforms (FFTs) of audio data, etc. . This may facilitate improved performance of tasks such as separation of sources in speech signals and / or speech recognition. In some implementations, both a vision system for processing images (e.g., still images, video) and another perception system to process LIDAR (and / or other robot sensor data) may be provided.
[0049] Pretrained machine learning models that are used to encode, decode, encrypt, and / or decrypt data and / or electronic transmissions for reliable and / or efficient transmission or storage may also be interleaved with intermediary machine learning models configured with selected aspects of the present disclosure. In instances where cryptographic keys are generated using multiple different modalities of data, respective pretrained machine learning models in each of those modalities may be combined with intermediary machine learning model(s) 116 configured with selected aspects of the present disclosure. Another application for techniques described herein is optimizing load distribution in a computer network. Yet another application is processing data obtained from physiological sensors, e.g. for medical diagnosis. In implementations where multiple pretrained machine learning models are usable to analyze DNA samples, these models may be interleaved as described herein, e.g., to provide a genotype estimate.
[0050] Fig. 2 depicts a non-limiting example of a robot 200 in the form of a robot arm. An end effector 206 in the form of a gripper claw is removably attached to a sixth joint 204-6 of robot 200. In this example, six joints 204-1 to 204-6 are indicated. However, this is not meant to be limiting, and robots may have any number of joints. In some implementations, robot 200 may be mobile, e.g., by virtue of a wheeled base 255 or other locomotive mechanism. Robot 200 is depicted in Fig. 2 in a particular selected configuration or “pose.”
[0051] As noted previously, intermediary machine learning model(s) 116 configured with selected aspects of the present disclosure may be significantly smaller than the pretrained machine learning models (e.g., 126, 136, 146). For example, an LLM 126 may have billions of parameters, whereas intermediary machine learning model may have orders of magnitude fewer parameters. This may reduce the computational cost of training and / or applying intermediarymachine learning model(s) 116. Another way the computation cost is limited is by limiting how much data is processed using intermediary machine learning model(s) 116.
[0052] Figs. 3-C schematically depicts an example of how a sequence 360 of activations in three different modalities (see the key) can be, as part of the GATS framework described herein, be “gathered,” e.g., by intermediary engine 112, for processing using intermediary machine learning model(s) 116. Time runs from left to right as indicated by the arrow. Fig. 3 A depicts the entire sequence 360 as it might be retrieved, with different sizes and fill patterns representing different modalities of data (e.g., text, vision, audio, proprioception values, etc. .
[0053] Fig. 3B depicts which activations are considered (e.g., the “attend” aspect of GATs) when intermediary engine 112 processes the first modality activation of sequence 360 that is bolded. As shown, the two third modality activations and the two second modality activations that immediately precede the bolded activation are considered, along with one other first modality activation. The activations rendered in dashed lines are not considered during this iteration. The reason these particular activations are considered in this example may be because the context length of the intermediary machine learning model 116 being used here is selected to be six activations, and the intent is to evenly distribute that context length between numbers of activations in each modality. Consequently, the two most recent activations of each modality are selected for processing at each iteration of the intermediary machine learning model 116. The same is true at the later iteration of Fig. 3C, where the bolded activation is from the second modality, and so the other most recent second modality activation, along with the two most recent activations in each of the first and third modalities, are selected.
[0054] In some implementations, activations may first be gathered (in the “gather” aspect of GATs) in accordance with the following. For the input sequence of embeddings (n, 2, ...xt} (which may be activations from the pretrained models), the largest subsequence G may be gathered such that:Only the selected subsequence G may be processed using the intermediary machine learning model 116. Other activations outside of G may be ignored.
[0055] During the “attend” phase of GATS, the next step is to attend over the gathered subsequence G. In some implementations, different sizes of the embeddings in G may be handled as follows. For each modality in there is a projection pmsuch that the size / dimensionality of pm(xi) is d. notably, d may be relatively small and the same for all projections. Each embedding from G may be projected with a corresponding projection, e.g., as follows:which results in obtaining new sequence where all elements have the same dimensionality d. Once the subsequence G is projected, it may be further processed, e.g., by a transformer layer (e.g., self-attention Attention and single feed-forward layer FFW), e.g., for each xL6 G, the following may apply:
[0056] During the “scatter” phase of GATS, the processed embeddings are then “projected back”, e.g., for each modality in there is a function rmsuch that the sizes of the initially input xLand the final rm, G))) are the same. This may ensure that the sizes of the gathered subsequence G and the resulting output sequence match. Before the projected embeddings are output, in some implementations, a gated residual connection may be added. A value of a gate may be, for instance, a scalar [0,1] computed by a gating function gmthat has the same inputs as rm.
[0057] To summarize all GATS layer transformations, each selected xtE G may be processed in the following manner:QmZi) ■ pn(z^) where z(= FFW(Attention pm^Xi)(xi'), G)) is the output of the attention block. Unselected input elements xLg G may remain unaltered.
[0058] Fig. 4 illustrates a simple example of how an intermediary machine learning model layer 416 may be interleaved with layers of other, pretrained machine learning models. In Fig. 4, time once again runs from left to right, and a sequence of activations 460A is provided and includes three modalities. Other possibilities, including different numbers of modalities, are contemplated. In Fig. 4, the sequence 460A of embeddings / activations in the second modality are depicted at front and have been processed using a layer i of a pretrained machine learning model (not depicted). As shown by the solid arrows, the first two activations of the secondmodality are processed during the current iteration of intermediary machine learning model layer 416.
[0059] More generally, if a given layer of the intermediary machine learning model layer 416 lies between the zth and z+1 th layers of a unimodal transformer, the output of the zth layer, instead of being fed immediately to the z+1 th layer, is first process and updated (or “steered”) by the intermediary machine learning model layer 416. Because the same intermediary machine learning model layer 416 is also interleaved with other pretraining machine learning models, currently processed embeddings interact with past embeddings from other modalities, enabling information flow between the models.
[0060] Referring back to Fig. 4, the first and third modalities of activations (which are depicted behind the second modality activations) may be processed similarly using layers z of other respective pretrained machine learning model(s). Note that it is not required that the same layer number of each pretrained machine learning model be applied at each iteration. For example, an intermediary machine learning models (e.g., 116, 416) may be interleaved between first and second layers of one pretrained machine learning model, between seventh and eight layers of another pretrained machine learning model, and so forth. Which layers are selected to be interleaved with intermediary machine learning model(s) may be determined, for instance, by parameter engine 118. In Fig. 4, as shown by the dashed arrows, the two most recent activations of the first modality are processed during the current iteration of intermediary machine learning model layer 416. Likewise, as shown by the dash-dot-dashed arrows, the two most recent activations of the third modality are processed during the current iteration of intermediary machine learning model 416.
[0061] Fig. 5 schematically depicts another example architecture configured with selected aspects of the present disclosure, in which a pretrained language model (LLM) 126 is conditioned on visual features from a vision model 136 to caption images. Time once again passes from left to right. An intermediary machine learning model layer 516(z), which is referred to in Fig. 5 as “GATS” (Gather, Attend, Scatter) is interleaved with layers of a pretrained vision (also referred to herein as “perception”) model 136 and an LLM 126. For a given image, the vision model 136 is applied first to obtain V (integer greater than zero) vision features. That may be followed by the frozen LLM 126 interleaved with intermediary machine learning model layer 516(z). The context length of intermediary machine learning model layer 516(z) may be set tocover all V vision features and the last text embedding output from the first layer of LLM 126. A projected embedding size d may be set to match the LLM 126 activation size, and hence, the projection (pi and back projection (n) transformations can both be identity functions.
[0062] At each iteration, intermediary machine learning model layer 516(z) is used, e.g., by intermediary engine 112, to process one language activation and two vision activations. In the example of Fig. 5, the same two vision activations are processed at each iteration as the desired effect is to steer the language model output based on the vision data, but not to steer the vision model output based on the language data. Accordingly, output of intermediary machine learning model layer 516(z) is applied across subsequent layers of LLM 126, but not perception / vision model 136. Put another way, as each language activation is processed by intermediary machine learning model layer 516(z), the same two vision activations are also processed. In Fig. 5, for instance, when the most recent (and bolded) language activation at bottom is processed, the same two vision activations are processed as when the first language activation (shown at left) was processed. The same is true for subsequent layers (e.g., z+1) of intermediary machine learning model 516.
[0063] Fig. 6 depicts another example architecture that incorporates selected aspects of the present disclosure. In this example, three pretrained models are deployed to aid in the control of a robot: an LLM 626, which may share characteristics with model(s) 126 in Fig. 1; a perception / vision model 636, which may share characteristics with model(s) 136 in Fig. 1; and an action model 646, which may or may not share characteristics with model(s) 146 in Fig. 1. Also depicted is an intermediary machine learning model 616 that is interleaved between layers of the pretrained models 626, 626, 646. As before, time runs from left to right.
[0064] At bottom, true inputs are depicted as a sequence of iterations. First, a natural language input instructing a robot (not depicted, e.g., 100) to pick up a lemon is received. Next comes a sequence of vision frames 670- A, 670-B, ... , 670-N interspersed with a sequence of robot proprioception values 672A, 672B, ... , 672N. These inputs may be processed using various types of encoders (not depicted) into a sequence 660 of tokens. These tokens may take the form of, for instance, continuous or discrete embeddings. Particular numbers of tokens in each modality are depicted for illustrative purposes only, and are not meant to be limiting.
[0065] The four language tokens that encode the natural language input are first processed using layers 1-3 of the LLM 626. Each set of three vision tokens encoding each vision frame 670 arefirst processed using layers 1-2 of perception / vision model 636. Each set of two proprioception tokens encoding proprioception values 672 are first processed using layer 1 of action model 646. In some implementations, the perception / vision model 636 and / or the action model 646 may be equipped with long contexts that enable them to condition their outputs on tokens from both the current time step and on tokens from previous time steps.
[0066] The raw activations of each of these models are intercepted and processed, e.g., by intermediary engine 112, using the interleaved layer 1 of intermediary machine learning model 616. This in turn generates steered activations of interleaved layer 1 of intermediary machine learning model 616. As shown by the arrows above layer 1 of intermediary machine learning model 616, the three steered vision activations are processed based on layers 3-4 of perception / vision model 636. Likewise, the two steered proprioception activations are processed using layer 2 of action model 646.
[0067] However, and similar to Fig. 5, the language activations generated based on layer 1 of intermediary machine learning model 616 are not applied to subsequent layers 4-6 of LLM. This may be because the single natural language input is meant to steer the other modalities, but not vice versa (i.e., there is no value to be gained or reason to steer LLM based on the other modalities). Additionally, the activations of the LLM 626 may be cached for frequent instructions, enabling the use of a larger LLM without sacrificing inference speed. This process may repeat as demonstrated in Fig. 6.
[0068] While not shown in Fig. 6, other modalities may be processed in a similar fashion. For instance, modalities such as audio may be added by, for instance, training parameters of the intermediary machine learning model 616 relating to the new modality from scratch. Parameters of the intermediary machine learning model 616 relating to other modalities for which it was trained previously may not require additional training. Even an abstract modality, uncoupled from real world data, such as a “scratch pad” for extra computations, could be added in a similar fashion. In some implementations, a single modality can be processed by multiple models, like a fast low-resolution video model for dynamics alongside a larger image model for manipulation details that runs periodically. Replacing pretrained models can be done in a similar fashion.
[0069] As noted elsewhere herein, the numbers of layers of unimodal transformers may vary. Accordingly, in some implementations, proportional interleaving may be used to interleave these unimodal transformer layers with intermediary machine learning models 116 as described herein.For example, if the intermediary machine learning model 116 has K layers, and the pretrained networks to be interleaved haveL2, ... , LMlayers, inputs to th layer of the intermediary machine learning model 116 may be embeddings obtained from lk l, lk 2, lk Mlayers (respectively for each modality) where the following applies: lkii= min(max(The kth layer of the intermediary machine learning model 116 may lie between lk mand lk m+ 1th layers of a unimodal transformer m (e.g., 126, 136, 146). The min and max may be used to ensure that the value of each lk iis positive and not larger than LL— 1.
[0070] Referring now to Fig. 7, an example method 700 of practicing selected aspects of the present disclosure is described. For convenience, the operations of the flowchart are described with reference to a system that performs the operations. This system may include various components of various computer systems, including any of systems 110, 120, 130, 140. Moreover, while operations of method 700 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted or added.
[0071] At block 702, the system, e.g., by way of intermediary engine 112 or by another inference engine (e.g., 128, 132, 142, etc.), may apply a first set of one or more tokens as inputs across one or more initial layers of a first pretrained machine learning model to generate a first set of raw activations. One example of this in Fig. 6 is when three vision tokens are applied as inputs across layers 1-2 of perception model 636. At block 704, the system, e.g, by way of intermediary engine 112 or by another inference engine (e.g., 128, 132, 142, etc.), may apply a second set of one or more tokens as inputs across one or more initial layers of a second pretrained machine learning model to generate a second set of raw activations. One example of this in Fig. 6 is when two proprioception tokens are applied as inputs across layer 1 of action model 646. At optional block 706, the system, e.g, by way of intermediary engine 112 or by another inference engine (e.g, 128, 132, 142, etc.), may apply a third set of one or more tokens as inputs across one or more initial layers of a third pretrained machine learning model to generate a third set of raw activations. One example of this in Fig. 6 is when four language tokens are applied as inputs across layers 1-3 of LLM 626.
[0072] At block 708, the system, e.g., by way of intermediary engine 112, may process the first and second sets, and the third set if present (and any additional sets, where applicable) of raw 1activations using an intermediary machine learning model (e.g., 116, 416, 516, 616) to generate first and second sets of steered activations.
[0073] At block 710, the system, e.g., by way of intermediary engine 112 or by another inference engine (e.g., 128, 132, 142, etc.), may apply the first set of steered activations as inputs across one or more subsequent layers of the first pretrained machine learning model to generate first pretrained machine learning model steered output. An example of this was depicted in Fig. 6, where steered vision activations were applied as inputs across layers 3-4 of perception model 636.
[0074] At block 712, the system, e.g., by way of intermediary engine 112 or by another inference engine (e.g., 128, 132, 142, etc.), may apply the second set of steered activations as inputs across one or more subsequent layers of the second pretrained machine learning model to generate second pretrained machine learning model steered output. An example of this was depicted in Fig. 6 where steered proprioception activations were applied as inputs across layer 2 of action model 646.
[0075] Unlike blocks 710 and 712, at block 714, the system , e.g., by way of intermediary engine 112 or by another inference engine (e.g., 128, 132, 142, etc.), may apply the third set of raw (not steered) activations as inputs across subsequent layer(s) of the third pretrained machine learning model to generate third pretrained machine learning output that is unsteered. An example of this was depicted in Fig. 6 where raw language activations were processed using layers 4-6 of LLM 626.
[0076] At block 716, the system may cause one or more processors to carry out one or more operations based on one or more of the first or second pretrained machine learning model steered output. In some cases, the one or more operations may include generating a control sign al for controlling a robot (e.g., 100, 200) to perform a task, and / or actually controlling the robot based on the control signal. In other implementations, the one or more operations may include, for instance: digital audio, image or video enhancement or analysis; separation of sources in speech signals, or speech recognition; encoding data for reliable and / or efficient transmission or storage (and corresponding decoding); encrypting / decrypting or signing electronic communications; generating keys; optimizing load distribution in a computer network; processing data obtained from physiological sensors, e.g. for medical diagnosis; and providing a genotype estimate based on an analysis of DNA samples.
[0077] The operations 702-716 may be performed during inference, e.g., once the intermediary machine learning model has been trained to a suitable accuracy, or during training. For instance, at optional block 718, the system, e.g., by way of intermediary engine 112, may train the intermediary machine learning model based on one or more outcomes of the one or more operations. In some implementations, this training may include reinforcement learning from human feedback (RLHF). As noted previously, if a robot operation failed, that may signal that the input sequence(s) that were applied and output sequence(s) that were generated during the episode should not be used as a training example. By contrast, if the robot operation succeeded, that may signal that the input sequence(s) that were applied and output sequence(s) that were generated during the episode are suitable for use as a training example. In other implementations, rather than training the intermediary machine learning model based on outcomes of downstream operations, the intermediary machine learning model may be trained based on the steered output(s). For example, if ground truth labels are available for one of the steered outputs, then the intermediary machine learning model 116 may be trained based on a comparison of the ground truth labels with the one steered output.
[0078] Referring now to Fig. 8, an example method 800 of practicing selected aspects of the present disclosure depicted in Fig. 5 is described. For convenience, the operations of the flowchart are described with reference to a system that performs the operations. This system may include various components of various computer systems, including video conference system 120. Moreover, while operations of method 800 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted or added.
[0079] At block 802, the system, e.g., by way of intermediary engine 112 or by another inference engine (e.g., 128, 132, 142, etc.), may apply a first set of one or more tokens representing one or more images as inputs across one or more initial layers of a vision machine learning model (e.g., vision layer i of vision model 136 in Fig. 5) to generate a first set of raw activations. At block 804, the system, e.g., by way of intermediary engine 112 or by another inference engine (e.g, 128, 132, 142, etc.), may apply a second set of one or more tokens representing a natural language snippet as inputs across one or more initial layers of a generative language model (e.g., LLM 126 of Fig. 5) to generate a second set of raw activations.
[0080] At block 806, the system, e.g., by way of intermediary engine 112, may process the first and second sets of raw activations using an intermediary machine learning model (e.g., 516 inFig. 5) to generate, respectively, first and second sets of steered activations. At block 808, the system, e.g., by way of intermediary engine 112 or by another inference engine (e.g., 128, 132, 142, etc.), may apply the second set of steered activations as inputs across one or more subsequent layers of the generative language model (126 in Fig. 5) to generate generative steered model output that represents a natural language output. At block 810, the system, e.g., by way of intermediary engine 112 or by another inference engine (e.g., 128, 132, 142, etc.), may apply the first set of raw activations as inputs across one or more subsequent layers of the vision machine learning model (136 in Fig. 5) to generate vision machine learning model unsteered output, causing the natural language output to be rendered at one or more output devices. Blocks 812 and 814 may proceed similarly to blocks 716-718 of Fig. 7.
[0081] Fig. 9 is a block diagram of an example computer system 910. Computer system 910 typically includes at least one processor 914 which communicates with a number of peripheral devices via bus subsystem 912. At least one processor 914 make take various forms, such as a central processing unit (CPU), graphics processing unit (GPU), tensor processing unit (TPU), or a neural processing unit (NPU), to name a few.
[0082] These peripheral devices may include a storage subsystem 924, including, for example, a memory subsystem 925 and a file storage subsystem 926, user interface output devices 920, user interface input devices 922, and a network interface subsystem 916. The input and output devices allow user interaction with computer system 910. Network interface subsystem 916 provides an interface to outside networks and is coupled to corresponding interface devices in other computer systems.
[0083] User interface input devices 922 may include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touch screen incorporated into the display, audio input devices such as voice recognition systems, microphones, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways to input information into computer system 910 or onto a communication network.
[0084] User interface output devices 920 may include a display subsystem, a printer, a fax machine, or non-visual displays such as audio output devices. The display subsystem may include a cathode ray tube (CRT), a flat-panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystemmay also provide non-visual display such as via audio output devices. In general, use of the term "output device" is intended to include all possible types of devices and ways to output information from computer system 910 to the user or to another machine or computer system.
[0085] Storage subsystem 924 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 924 may include the logic to perform selected aspects of method 700 and / or 800, and / or to implement one or more aspects of robot 100 or systems 110, 120, 130, and / or 140. Memory 925 used in the storage subsystem 924 can include a number of memories including a main randomaccess memory (RAM) 930 for storage of instructions and data during program execution and a read only memory (ROM) 932 in which fixed instructions are stored. A file storage subsystem 926 can provide persistent storage for program and data files, and may include a hard disk drive, a CD-ROM drive, an optical drive, or removable media cartridges. Modules implementing the functionality of certain implementations may be stored by file storage subsystem 926 in the storage subsystem 924, or in other machines accessible by the processor(s) 914.
[0086] Bus subsystem 912 provides a mechanism for letting the various components and subsystems of computer system 910 communicate with each other as intended. Although bus subsystem 912 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0087] Computer system 910 can be of varying types including a workstation, server, computing cluster, blade server, server farm, smart phone, smart watch, smart glasses, set top box, tablet computer, laptop, or any other data processing system or computing device. Due to the everchanging nature of computers and networks, the description of computer system 910 depicted in Fig. 9 is intended only as a specific example for purposes of illustrating some implementations. Many other configurations of computer system 910 are possible having more or fewer components than the computer system depicted in Fig. 9.
[0088] In some implementations, a computer implemented method may be provided that includes: applying a first set of one or more tokens as inputs across one or more initial layers of a first pretrained machine learning model to generate a first set of raw activations; applying a second set of one or more tokens as inputs across one or more initial layers of a second pretrained machine learning model to generate a second set of raw activations; processing the first and second sets of raw activations using an intermediary machine learning model togenerate first and second sets of steered activations; applying the first set of steered activations as inputs across one or more subsequent layers of the first pretrained machine learning model to generate first pretrained machine learning model steered output; applying the second set of steered activations as inputs across one or more subsequent layers of the second pretrained machine learning model to generate second pretrained machine learning model steered output; and causing one or more of the processors to carry out one or more operations based on one or more of the first or second pretrained machine learning model steered output.
[0089] In various implementations, the intermediary machine learning model may attend across the first and second sets of raw activations. In various implementations, the one or more operations may include controlling a robot to perform a task. In various implementations, the first set of one or more tokens may convey past or present proprioception values of the robot, and the first pretrained machine learning model may include an action model that is trained to predict future proprioception values of the robot. In various implementations, the second set of one or more tokens may include sensor data captured by one or more sensors of the robot, and the second pretrained machine learning model may be a perception model.
[0090] In various implementations, the method may further include: applying data indicative of a natural language request as input across one or more layers of a language model, wherein the natural language request conveys the task to be performed by the robot; and processing a third set of raw activations generated from the one or more layers of the language model using the intermediary machine learning model to attend across the first, second, and third sets of raw activations.
[0091] In various implementations, the intermediary model may include a transformer model with self-attention. In various implementations, the first set of one or more tokens may be from a first modality and the second set of one or more tokens may be from a second modality that is different than the first modality. In various implementations, the method may include apportioning a context length of the intermediary machine learning model between at least the first and second modalities. In various implementations, the apportioning may include selecting the first and second sets of tokens from a larger superset that includes tokens of both the first and second modalities. In various implementations, the selecting may include selecting the most recent tokens of the first and second modalities from the larger superset as the first and second sets of tokens. 1
[0092] In various implementations, processing the first and second sets of raw activations may include projecting one or both of the first set of raw activations and the second set of raw activations to a shared dimensionality, wherein the intermediary machine learning model is applied to the projected activations. In various implementations, the raw activations of the first and second sets may include embeddings, and the projecting may include projecting at least some of the embeddings to the shared dimensionality. In various implementations, the method may further include projecting at least some raw activations of the intermediary machine learning model to an original dimensionality of the first or second set of tokens to generate the first or second set of steered activations.
[0093] In various implementations, the method may further include: applying a third set of one or more tokens as inputs across one or more layers of a third pretrained machine learning model to generate a third set of raw activations; and processing the third set of raw activations along with the first and second sets of activations using the intermediary machine learning model to generate the first and second sets of steered activations. In various implementations, the method may further include applying the third set of raw activations as inputs across one or more subsequent layers of the third pretrained machine learning model to generate third pretrained machine learning model output.
[0094] In another aspect, a method may be implemented using one or more processors and may include: applying a first set of one or more tokens as inputs across one or more initial layers of a first pretrained machine learning model to generate a first set of raw activations; applying a second set of one or more tokens as inputs across one or more initial layers of a second pretrained machine learning model to generate a second set of raw activations; processing the first and second sets of raw activations using an intermediary machine learning model to generate first and second sets of steered activations; applying the first set of steered activations as inputs across one or more subsequent layers of the first pretrained machine learning model to generate first pretrained machine learning model steered output; applying the second set of steered activations as inputs across one or more subsequent layers of the second pretrained machine learning model to generate second pretrained machine learning model steered output; causing one or more of the processors to carry out one or more operations based on one or more of the first or second pretrained machine learning model steered output; and training theintermediary machine learning model based on one or more outcomes of the one or more operations.
[0095] In various implementations, the one or more operations may include one or more robot operations, and the one or more outcomes comprise one or more results of the one or more robot operations. In various implementations, the one or more operations may include predicting natural language output. In various implementations, the intermediary machine learning model may attend across the first and second sets of raw activations. In various implementations, the intermediary machine learning model may attend across the first and second sets of raw activations in parallel.
[0096] In yet another aspect, a method may be implemented using one or more processors and may include: applying a first set of one or more tokens as inputs across one or more initial layers of a first pretrained machine learning model to generate a first set of raw activations; applying a second set of one or more tokens as inputs across one or more initial layers of a second pretrained machine learning model to generate a second set of raw activations; processing the first and second sets of raw activations using an intermediary machine learning model to generate first and second sets of steered activations; applying the first set of steered activations as inputs across one or more subsequent layers of the first pretrained machine learning model to generate first pretrained machine learning model steered output; applying the second set of steered activations as inputs across one or more subsequent layers of the second pretrained machine learning model to generate second pretrained machine learning model steered output; and training the intermediary machine learning model based on one or more of the first or second pretrained machine learning model steered output.
[0097] In various implementations, the one of the first and second pretrained machine learning model steered outputs includes natural language. In various implementations, the intermediary machine learning model attends across the first and second sets of raw activations. In various implementations, the intermediary machine learning model attends across the first and second sets of raw activations in parallel.
[0098] In another aspect, a method may be implemented using one or more processors and may include: applying a first set of one or more tokens representing one or more images as inputs across one or more initial layers of a vision machine learning model to generate a first set of raw activations; applying a second set of one or more tokens representing a natural language snippetas inputs across one or more initial layers of a generative language model to generate a second set of raw activations; processing the first and second sets of raw activations using an intermediary machine learning model to generate, respectively, first and second sets of steered activations; applying at least the second set of steered activations as inputs across one or more subsequent layers of the generative language model to generate generative steered model output that represents a natural language output; and causing the natural language output to be rendered at one or more output devices.
[0099] Other implementations may include a transitory or non-transitory computer readable storage medium storing instructions executable by a processor to perform a method such as one or more of the methods described above. Yet another implementation may include a control system including memory and one or more processors operable to execute instructions, stored in the memory, to implement one or more modules or engines that, alone or collectively, perform a method such as one or more of the methods described above.
[0100] It should be appreciated that all combinations of the foregoing concepts and additional concepts described in greater detail herein are contemplated as being part of the subject matter disclosed herein. For example, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein.
[0101] While several implementations have been described and illustrated herein, a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein may be utilized, and each of such variations and / or modifications is deemed to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the teachings is / are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. It is, therefore, to be understood that the foregoing implementations are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems,articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the scope of the present disclosure.
Claims
CLAIMSWhat is claimed is:
1. A method implemented using one or more processors and comprising: applying a first set of one or more tokens as inputs across one or more initial layers of a first pretrained machine learning model to generate a first set of raw activations; applying a second set of one or more tokens as inputs across one or more initial layers of a second pretrained machine learning model to generate a second set of raw activations; processing the first and second sets of raw activations using an intermediary machine learning model to generate first and second sets of steered activations; applying the first set of steered activations as inputs across one or more subsequent layers of the first pretrained machine learning model to generate first pretrained machine learning model steered output; applying the second set of steered activations as inputs across one or more subsequent layers of the second pretrained machine learning model to generate second pretrained machine learning model steered output; and causing one or more of the processors to carry out one or more operations based on one or more of the first or second pretrained machine learning model steered output.
2. The method of claim 1 , wherein the intermediary machine learning model attends across the first and second sets of raw activations.
3. The method of claim 1, wherein the one or more operations comprise generating a control signal for controlling a robot to perform a task.
4. The method of claim 3, wherein the first set of one or more tokens convey past or present proprioception values of the robot, and the first pretrained machine learning model comprises an action model that is trained to predict future proprioception values of the robot.
5. The method of claim 4, wherein the second set of one or more tokens comprise sensor data captured by one or more sensors of the robot, and the second pretrained machine learning model comprises a perception model.
6. The method of claim 5, further comprising: applying data indicative of a natural language request as input across one or more layers of a language model, wherein the natural language request conveys the task to be performed by the robot; andprocessing a third set of raw activations generated from the one or more layers of the language model using the intermediary machine learning model to attend across the first, second, and third sets of raw activations.
7. The method of claim 5, wherein the sensor data captured by one or more sensors of the robot comprises LIDAR data captured by a LIDAR sensor.
8. The method of claim 5, wherein the sensor data captured by one or more sensors of the robot comprises vision data captured by a vision sensor.
9. The method of claim 1 , wherein the intermediary model comprises a transformer model with self-attention.
10. The method of claim 1, wherein the first set of one or more tokens are from a first modality and the second set of one or more tokens are from a second modality that is different than the first modality.
11. The method of claim 10, further comprising apportioning a context length of the intermediary machine learning model between at least the first and second modalities.
12. The method of claim 11, wherein the apportioning comprises selecting the first and second sets of tokens from a larger superset that includes tokens of both the first and second modalities.
13. The method of claim 12, wherein the selecting comprises selecting the most recent tokens of the first and second modalities from the larger superset as the first and second sets of tokens.
14. The method of claim 1, wherein processing the first and second sets of raw activations comprises projecting one or both of the first set of raw activations and the second set of raw activations to a shared dimensionality, wherein the intermediary machine learning model is applied to the projected activations.
15. The method of claim 14, wherein the raw activations of the first and second sets comprise embeddings, and the projecting comprises projecting at least some of the embeddings to the shared dimensionality.
16. The method of claim 14, further comprising projecting at least some raw activations of the intermediary machine learning model to an original dimensionality of the first or second set of tokens to generate the first or second set of steered activations.
17. The method of claim 1, further comprising:applying a third set of one or more tokens as inputs across one or more layers of a third pretrained machine learning model to generate a third set of raw activations; and processing the third set of raw activations along with the first and second sets of activations using the intermediary machine learning model to generate the first and second sets of steered activations.
18. The method of claim 17, further comprising applying the third set of raw activations as inputs across one or more subsequent layers of the third pretrained machine learning model to generate third pretrained machine learning model output.
19. A method implemented using one or more processors and comprising: applying a first set of one or more tokens as inputs across one or more initial layers of a first pretrained machine learning model to generate a first set of raw activations; applying a second set of one or more tokens as inputs across one or more initial layers of a second pretrained machine learning model to generate a second set of raw activations; processing the first and second sets of raw activations using an intermediary machine learning model to generate first and second sets of steered activations; applying the first set of steered activations as inputs across one or more subsequent layers of the first pretrained machine learning model to generate first pretrained machine learning model steered output; applying the second set of steered activations as inputs across one or more subsequent layers of the second pretrained machine learning model to generate second pretrained machine learning model steered output; causing one or more of the processors to carry out one or more operations based on one or more of the first or second pretrained machine learning model steered output; and training the intermediary machine learning model based on one or more outcomes of the one or more operations.
20. The method of claim 19, wherein the one or more operations comprise one or more robot operations, and the one or more outcomes comprise one or more results of the one or more robot operations.
21. The method of claim 19, wherein the one or more operations comprise predicting natural language output.
22. The method of claim 19, wherein the intermediary machine learning model attends across the first and second sets of raw activations.
23. The method of claim 22, wherein the intermediary machine learning model attends across the first and second sets of raw activations in parallel.
24. A method implemented using one or more processors and comprising: applying a first set of one or more tokens as inputs across one or more initial layers of a first pretrained machine learning model to generate a first set of raw activations; applying a second set of one or more tokens as inputs across one or more initial layers of a second pretrained machine learning model to generate a second set of raw activations; processing the first and second sets of raw activations using an intermediary machine learning model to generate first and second sets of steered activations; applying the first set of steered activations as inputs across one or more subsequent layers of the first pretrained machine learning model to generate first pretrained machine learning model steered output; applying the second set of steered activations as inputs across one or more subsequent layers of the second pretrained machine learning model to generate second pretrained machine learning model steered output; and training the intermediary machine learning model based on one or more of the first or second pretrained machine learning model steered output.
25. The method of claim 24, wherein the one of the first and second pretrained machine learning model steered outputs comprises natural language.
26. The method of claim 24, wherein the intermediary machine learning model attends across the first and second sets of raw activations.
27. The method of claim 26, wherein the intermediary machine learning model attends across the first and second sets of raw activations in parallel.
28. A method implemented using one or more processors and comprising: applying a first set of one or more tokens representing one or more images as inputs across one or more initial layers of a vision machine learning model to generate a first set of raw activations;applying a second set of one or more tokens representing a natural language snippet as inputs across one or more initial layers of a generative language model to generate a second set of raw activations; processing the first and second sets of raw activations using an intermediary machine learning model to generate, respectively, first and second sets of steered activations; applying at least the second set of steered activations as inputs across one or more subsequent layers of the generative language model to generate generative steered model output that represents a natural language output; and causing the natural language output to be rendered at one or more output devices.
29. The method of claim 28, further comprising applying the first set of raw activations as inputs across one or more subsequent layers of the vision machine learning model to generate vision machine learning model unsteered output.
30. The method of claim 28, further comprising applying the first set of steered activations as inputs across one or more subsequent layers of the vision machine learning model to generate vision machine learning model steered output.
31. A method of controlling a mechanical agent in a real-world environment, comprising: performing the method of any one of claims 1 to 23, wherein the one or more operations comprise generating a control signal for controlling the mechanical agent to perform a task, and controlling the mechanical agent using the control signal.
32. A system comprising one or more processors and memory storing instructions that, in response to execution by the one or more processors, cause the one or more processors to perform any of the methods of claims 1-31.
33. At least one non-transitory computer-readable medium comprising instructions that, in response to execution by one or more processors, cause the one or more processors to perform any of the methods of claims 1-31.
Citation Information
Cited By
Mechanical arm path planning method based on Transform and diffusion model
CN120921369A