Delivery and exchange multi-modal model and use
By using the LoReTTa model for pass-through and exchange machine learning, the problem of missing modalities in multimodal datasets is solved, improving the model's performance in cases of missing modalities. This approach is applicable to fields such as healthcare, infrastructure, and transportation.
Patent Information
- Application Number
- CN202480028125.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-28
- Filing Date
- 2024-04-24
- Publication Date
- 2025-11-25
AI Technical Summary
Existing technologies struggle to effectively handle missing modalities in multimodal datasets, especially when modalities are missing in the training data, leading to a decline in the performance of multimodal models in practical applications.
The LoReTTa model is used to process the input dataset by passing and exchanging machine learning models. It leverages the generation of pre-trained and exchange properties to link modalities and learn the relationship between different data distributions, making it suitable for training multimodal models.
It enables better handling of multimodal combinations in the case of missing modes, improving the downstream performance of the model, especially in safety-critical fields such as healthcare, infrastructure and transportation.
Smart Images

Figure CN121014045A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a computer-implemented method for analyzing input data using a multimodal machine learning model trained with exchange modeling and transitive modeling, and a computer-implemented method for training a multimodal machine learning model using exchange modeling and transitive modeling. Related systems and computer program products are also described. Background Technology
[0002] Integrating multiple modalities (such as images, text, or speech) promises to provide a more comprehensive and holistic understanding of what would otherwise be a challenging problem. By leveraging the unique and complementary strengths of each data type, this approach yields more synergistic representations of underlying phenomena—enabling deep learning systems to excel and adapt to real-world scenarios. Consequently, impressive progress has been made in the field of multimodal learning in recent years. Much of this work focuses on multimodal datasets where all data points have all modalities (see Figure 3a). Adapting methods to data with missing modalities still heavily rely on the existence of at least a subset of samples containing all modal combinations (see Figure 3b). In practice, obtaining such datasets is not easy when considering more than two modalities (e.g., three modalities). In the worst case, we might end up with a dataset containing only modalities A and B, and another disjoint dataset containing modalities B and C. Thus, combinations A and C are unconditionally missing (see Figure 3c).
[0003] This is particularly important in safety-critical sectors such as healthcare, infrastructure, and transportation. For example, in medicine, it's already difficult to obtain large datasets with a single modality to train robust foundational models—such as those used in general computer vision and natural language processing. Finding datasets with two or more modalities is even more challenging. However, combining them is crucial because it's common practice to examine different types of data, such as radiographic images, gene sequences, and microscope slides, to determine the optimal treatment. Similarly, infrastructure monitoring models or autonomous driving systems require input from multiple sensors. However, not all modality combinations exist in the same training set.
[0004] Therefore, new methods are needed to obtain multimodal models that are not affected by the shortcomings of existing technologies. Summary of the Invention
[0005] The inventors seek to find a solution to the problem of pre-trained multimodal neural networks that can handle any combination of modalities at inference time when missing modalities exist in the pre-training data. The inventors are particularly interested in developing a method that can handle cases where there are never-before-seen combinations of modalities (e.g., (A,C) and (A,B,C) in the example on Figure 1c). The inventors seek to develop a method that can do this and additionally obtain a model that achieves better downstream performance given more combinations of modalities, including the never-before-seen combinations (A,C) and (A,B,C).
[0006] To the inventors' knowledge, this scenario has never been considered in the literature. The inventors have designed a novel transitive pre-trained machine learning model, called LoReTTa, to handle such situations, illustrating a transformer-based model. LoReTTa (Linking modalities with a transitive and commutative pre-trained model) can utilize generative pre-training (predicting the next lexical term), the commutative property (A,B)=(B,A), and transitive relations. This allows for the linking and transformation between modalities. The inventors further demonstrate, both theoretically and experimentally, that LoReTTa is well-suited for training multimodal models by learning intra- and inter-terminal relationships from diverse data distributions. As described above, this has profound implications for safety-critical fields such as healthcare, infrastructure, and transportation, where it is common for not all modality combinations to exist on the same training set. With LoReTTa, the inventors present a unique approach to addressing this challenging and under-researched problem.
[0007] Therefore, in a first aspect, a computer-implemented method is provided, comprising: obtaining an input dataset comprising data from one or more modalities selected from a set of at least three modalities; and generating an output dataset by processing the input dataset using a machine learning model including a pass-and-exchange machine learning model, wherein the pass-and-exchange machine learning model is a machine learning model trained to transform an input dataset corresponding to any and each of the modalities in the set of at least three modalities into an output dataset of any and each of the other modalities in the set of at least three modalities, wherein the machine learning model has been trained to transform data between two modalities using training data comprising data pairs corresponding to a first modality and another modality in the two modalities and other data pairs corresponding to the other modality and a second modality in the two modalities, and the machine learning model has been trained to perform a corresponding inverse type modal transformation for any forward type modal transformation performed by the model.
[0008] According to this aspect, a computer-implemented method is also described, comprising: accessing an input dataset; determining a specific input modality of the input dataset; receiving an indication of a specific target output modality; and generating an output dataset by processing the input dataset and recognizing the specific target output modality using a transitive and exchangeable machine learning model, wherein the transitive and exchangeable machine learning model: is configured to transform an input dataset corresponding to any and each of at least three modalities into an output dataset of any and each of the other of the at least three modalities; is transitive because it is configured to learn to transform data between two modalities using training data comprising data pairs corresponding to a first modality and another modality of the two modalities and other data pairs corresponding to the other modality and a second modality of the two modalities; and is exchangeable because it is configured such that, for any forward type modality transformation performed by the transitive and exchangeable machine learning model, the transitive and exchangeable machine learning model is also configured to perform a corresponding reverse type modality transformation.
[0009] The method according to aspects of the invention may have any one or more of the following optional features.
[0010] The machine learning model can be a classification model, a regression model, or a generative model. A classification model can include the transitive and exchange machine learning model configured to take the input dataset as input, and a trained classification model trained to take features extracted from the transitive and exchange machine learning model as input and produce a classification for the input dataset as output. A generative model can be trained to take the input dataset as input and produce a generated output dataset corresponding to the input dataset, wherein the output dataset corresponds to a modality from the set of at least three modalities that is not present in the input dataset. A regression model can include the transitive and exchange machine learning model configured to take the input dataset as input, and a trained regression model trained to take features extracted from the transitive and exchange machine learning model as input and produce regression values for the input dataset as output. The regression model can be a Cox proportional hazards model, and the regression values can be survival predictions. Obtaining the input dataset can include accessing or receiving the input dataset. Therefore, the transitive and exchange machine learning model can be described as a machine learning model that: (i) is configured to transform an input dataset corresponding to any and every modality in the set of at least three modalities into an output dataset corresponding to any and every other modality in the set of at least three modalities; (ii) is transitive because it is configured to learn to transform data between two modalities by using training data, which includes data pairs corresponding to a first modality and another modality in the two modalities and other data pairs corresponding to the other modality and a second modality in the two modalities; and (ii) is exchangeive because it is configured such that for any positive type of modality transformation configured to be performed by the transitive and exchange machine learning model, the transitive and exchange machine learning model is also configured to perform a corresponding inverse type of modality transformation.
[0011] The method may include determining a specific input modality of the input dataset and an instruction to receive a specific target output modality; and generating the output dataset is performed by using the pass-and-exchange machine learning model to process the input dataset and identify the specific target output modality.
[0012] The transfer and exchange machine learning model may have been trained using causal generative modeling, masking modeling, and / or variations thereof. A variation may be causal masking modeling. Causal generative modeling, masking modeling, and / or variations thereof may have used benchmark real data for each of a plurality of data pairs in the training data, comprising concatenated input data of a first mode and a second mode of the data pair, wherein the mode used as the first mode of the data pair is randomly sampled from both modes of the data pair. Alternatively, or otherwise, the benchmark real data may have included first and second concatenated input data for each of the plurality of data pairs in the training data, wherein the mode used as the first mode of the data pair is different between the first and second concatenated input data.
[0013] The pass-and-exchange machine learning model can be a sequence model and / or a deep learning model. This pass-and-exchange machine learning model can include sequence models selected from: recurrent neural networks, models combining recursion and attention mechanisms, convolutional neural networks or convolution-based models, optionally models combining element-wise multiplication (gated) and long convolutions or models combining convolution and attention mechanisms, state-space models or models combining state-space models and attention mechanisms, or transformer-based models. The recurrent neural network can be a Long Short-Term Memory (LSTM) model. A model combining recursion and attention mechanisms can be a Receiving Weighted Key-Value (RWKV) model or a Griffin model. A model combining element-wise multiplication (gated) and long convolutions can be a Hyena model. A model combining convolution and attention mechanisms can be a Striped Hyena model. When the attention-based model is an attention-only model, it can also be referred to as a transformer-based model. The state-space model can be a Mamba or Hawk model. A model combining a state-space model and attention mechanisms can be referred to as a hybrid transformer-state-space model. Examples include the Jamba model.
[0014] The transfer and exchange machine learning model may have been trained using causal masking modeling to: identify subsets of lexical units in the training dataset, where the lexical units in the subsets are located at multiple positions within the training dataset; append the subsets of lexical units to the training dataset; and add masks at multiple positions within the training dataset. The method may further include the identification of masks at multiple positions during training.
[0015] Transmitting and exchanging machine learning models can include transformer models that use self-attention.
[0016] The method may further include receiving an initial dataset and generating an input dataset by lexicalizing the initial dataset.
[0017] Generating the input dataset may include lexicalizing an initial dataset and projecting the lexicalized initial dataset into a predetermined feature space using a parameterized embedding function for each modality. Generating the input dataset may include appending modality-specific terms to the lexicalized and optionally projected data corresponding to each individual modality in the input dataset before and / or after it. Generating the input dataset may include learned positional embeddings.
[0018] The method may further include accessing another input dataset associated with the input dataset; and determining another specific input modality of the other input dataset, wherein the processing performed to generate the output dataset also includes using a pass-and-exchange machine learning model to process the other dataset and the identification of the other specific input modality.
[0019] The transitive and exchange machine learning model may have been trained using a method comprising the following steps: accessing a training dataset comprising first data of a specific input modality and corresponding second data of another modality; generating a first predicted output dataset corresponding to the other modality by processing the first data and recognizing the other modality using the transitive and exchange machine learning model; generating a second predicted output dataset corresponding to the specific target output modality by processing the first predicted output data and recognizing the specific target output modality using the transitive and exchange machine learning model; generating a third predicted output dataset corresponding to the specific input modality by processing the second predicted output data and recognizing the specific input modality using the transitive and exchange machine learning model; and calculating a loss by comparing at least a portion of the third predicted output dataset with the first data.
[0020] The transfer and exchange machine learning model may have been trained using training data that lacks data to associate first data of a specific input modality with second data of a specific output modality. The training data may have been trained using training data that associates data of modality pairs in a set of at least three modalities, wherein for each pair of first and second modalities in the set of at least three modalities, the training data includes data that associates the data of the modality pair that together form a path between the first and second modalities. In an embodiment, for each pair of first and second modalities in the set of at least three modalities, the training data includes data that associates the data of the first modality with a third modality and data that associates the data of the third modality with the second modality. Therefore, the training data may not include training data associated with any single pair of modalities, but may include data associated with the modalities in the missing pair using one or more bridging (also referred to herein as “links”) modalities.
[0021] The transitive and exchange-based machine learning model may have been trained using a first step and a second step, where the first step uses exchange-based modeling and a self-supervised learning objective. In the second step, the machine learning model weights are initialized to the weights learned in the first step, and further training is performed using transitive modeling. The first step may have included training the model to take partially masked concatenated input data of a first and second modality of a data pair in a training dataset as input and to produce a prediction of the masked input data as output. The training dataset includes benchmark ground data for multiple data pairs, each data pair including corresponding data for a pair of modalities, wherein the modality used as the first modality of the data pair is randomly sampled from both modalities of the data pair, and / or the benchmark ground data for each of the multiple data pairs in the training data includes first and second concatenated input data, wherein the modality used as the first modality of the data pair is different between the first and second concatenated input data.
[0022] The second step may have already included training the model for each of a plurality of data pairs in the training dataset, each data pair including first data of a first modality and corresponding second data of a second modality in a set of at least three modalities; using a machine learning model with the second data as input to predict third data of a third modality in a set of at least three modalities corresponding to the second data; optionally, using a pass-and-exchange machine learning model with the predicted third data or predicted subsequent data as input to iteratively predict subsequent data of another modality in a set of at least three modalities corresponding to the predicted third data or predicted subsequent data; using a pass-and-exchange machine learning model with the third data or the most recently predicted subsequent data as input to predict final data of the first modality corresponding to the predicted third data or the most recently predicted subsequent data; a loss for comparing the first data and the predicted final data; wherein the training data includes: (i) pairs including data of the second modality and corresponding data of the third modality, and (ii) This includes pairs of data from the third mode and corresponding data from the first mode, or pairs of data from the latest additional mode and corresponding data from the first mode, as well as corresponding pairs of data from each of the two modes used in the step of iteratively predicting subsequent data from the additional mode.
[0023] Training a model using exchange modeling and self-supervised learning objectives may include, or may have already included, training the model to take as input partially masked concatenated input data of a first and second modality of a data pair in a training dataset and to produce a prediction of the masked input data as output. The training dataset includes benchmark ground data for multiple data pairs, each data pair including corresponding data for a pair of modalities, wherein the modality used as the first modality of the data pair is randomly sampled from both modalities of the data pair, and / or the benchmark ground data for each of the multiple data pairs in the training dataset includes first and second concatenated input data, wherein the modality used as the first modality of the data pair is different between the first and second concatenated input data.
[0024] Using transitive modeling to train a model may include, or may have already included, training the model for each of a plurality of data pairs in a training dataset, each data pair including first data of a first modality and corresponding second data of a second modality in a set of at least three modalities; using a machine learning model with the second data as input to predict third data of a third modality in a set of at least three modalities corresponding to the second data; using a transitive and exchange machine learning model with the third data as input to predict final data of the first modality corresponding to the predicted third data; and calculating a loss that compares the first data and the predicted final data; wherein the training data includes: (i) pairs including data of the second modality and corresponding data of the third modality, and (ii) pairs including data of the third modality and corresponding data of the first modality.
[0025] Using transitive modeling to train a model may include, or may have already included, training the model for each of a plurality of data pairs in a training dataset, each data pair comprising: first data for a first modality and corresponding second data for a second modality in a set of at least three modalities; using a machine learning model with the second data as input to predict third data for a third modality in a set of at least three modalities corresponding to the second data; using a transitive and exchange machine learning model with the predicted third data or predicted subsequent data as input to iteratively predict subsequent data for another modality in a set of at least three modalities corresponding to the predicted third data or predicted subsequent data; using a transitive and exchange machine learning model with the most recently predicted subsequent data as input to predict final data for a first modality corresponding to the most recently predicted subsequent data; and calculating a loss that compares the first data and the predicted final data; wherein the training data comprises: (i) pairs comprising data for the second modality and corresponding data for the third modality, and (ii) This includes pairs of data for the latest additional modality and corresponding data for the first modality, as well as corresponding pairs of data for each of the two modalities used in the step of iteratively predicting subsequent data for the additional modality.
[0026] The transfer and exchange machine learning model may have been trained to learn a first conditional distribution (P(C|A), P(A|C)) for a pair of modalities (A, C), for which the training data does not contain a sufficient pair of training data to correlate the data of the pair of modalities to parameterize the conditional distribution. The learning is performed by predicting the data of one or more additional modalities (B) forming the link by using a learned further conditional distribution (P(C|B)) of the modal pair (C, B) that forms the link between the pair of modalities, and for the modal pair, the training data contains a sufficient pair of training data to correlate the data of the pair of modalities to parameterize the further conditional distribution.
[0027] The machine learning model being transferred and exchanged may have been trained using training data that includes tuples (e.g., pairs) or consists of such tuples. These tuples may include corresponding data for multiple modalities out of at least three modalities. Specifically, the number of tuples containing data for a particular pair of modalities out of at least three modalities present in the training data is less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1%. In an embodiment, tuples do not include tuples containing data for that particular pair of modalities out of at least three modalities.
[0028] This method may further include training a machine learning model.
[0029] One or more modalities may be selected from medical or biological sample data associated with the subject. One or more modalities may be selected from: medical image data, electronic medical record data, gene sequence data, and digital pathology image data. Medical image data may be selected from MRI, CT, or PET imaging data. Therefore, a specific input modality or target output modality may include MRI, CT, or PET imaging. A specific input modality or target output modality may include electronic medical record representation. A specific input modality or target output modality may include gene sequence representation. A specific input modality or target output modality may include digital pathology imaging.
[0030] Machine learning models can be configured to take medical or biological sample data associated with a subject as input and produce instructions for recommended treatment as output. These instructions can be categorized. Medical or biological sample data can include one or more of the following: radiographic images, gene sequences, and microscope slides.
[0031] The machine learning model can be configured to take data collected by any one or more types of sensors of a variety of types associated with a vehicle, optionally an automobile, as input, and generate predicted sensor data of another type of sensor data associated with the vehicle as output. The machine learning model may have been trained using synchronized data from pairs of multiple types of sensors. According to a second aspect, a computer-implemented method is provided, comprising: inputting medical or biological sample data associated with a subject using the method of any embodiment of the first aspect; and providing the subject with diagnostic and / or prognostic and / or treatment recommendations using the output of a machine learning model, wherein the output of the machine learning model indicates a diagnostic, prognostic, or treatment recommendation. According to a third aspect, a computer-implemented method for training a multimodal machine learning model is provided, the method comprising: accessing a training dataset comprising data for each of at least three modalities in a set; and training the machine learning model to transform an input dataset corresponding to any and each of the modalities in the set of at least three modalities into an output dataset of any and each of the other modalities in the set of at least three modalities, wherein the machine learning model is trained to transform data between two modalities using training data comprising data pairs corresponding to a first modality and another modality in the two modalities and other data pairs corresponding to the other modality and a second modality in the two modalities, and the machine learning model is trained to perform a corresponding inverse modal transformation for any forward modal transformation performed by the model being trained to perform.
[0032] According to this aspect, a computer-implemented method for obtaining a trained transitive and exchangeable machine learning model is also described, wherein: the trained transitive and exchangeable machine learning model is configured to transform an input dataset corresponding to any and each of at least three modalities into an output dataset of any and each of at least three other modalities; the trained transitive and exchangeable machine learning model is transitive because it is configured to learn to transform data between two modalities by using training data, which includes data pairs corresponding to a first modality and another modality among the two modalities and other data pairs corresponding to the other modality and a second modality among the two modalities; and the trained transitive and exchangeable machine learning model is exchangeable because it is configured such that for any forward type of modality transformation configured to be performed by the transitive and exchangeable machine learning model, the transitive and exchangeable machine learning model is also configured to perform a corresponding reverse type of modality transformation; and wherein the method includes one or both of the following: (i) The machine learning model is trained using causal masking modeling to: identify a subset of lexical units in the training dataset, wherein the lexical units in the subset are located at multiple positions within the training dataset; append the subset of lexical units to the training dataset; and add masks at multiple positions within the training dataset; and (ii) access the training dataset, which includes first data for a specific input modality of at least three modalities and corresponding second data for another modality of at least three modalities; generate a first predicted output dataset corresponding to the other modality by processing the first data and identifying the other modality using a pass-and-exchange machine learning model; generate a second predicted output dataset corresponding to the specific target output modality by processing the first predicted output data and identifying the specific target output modality using a pass-and-exchange machine learning model; generate a third predicted output dataset corresponding to the specific input modality by processing the second predicted output data and identifying the specific input modality using a pass-and-exchange machine learning model; and compute a loss by comparing at least a portion of the third predicted output dataset with the first data.
[0033] The methods according to this invention may have any of the features described with respect to the first aspect. The methods according to this invention may have any one or more of the following optional features.
[0034] Training a machine learning model may include using causal generative modeling, masking modeling, and / or causal masking modeling, wherein: the causal generative modeling, masking modeling, and / or causal masking modeling uses benchmark real data from the training dataset, the benchmark real data comprising concatenated input data of a first mode and a second mode for each of a plurality of data pairs in the training data, wherein the mode used as the first mode of the data pair is randomly sampled from the two modes of the data pair, and / or the benchmark real data comprising first concatenated input data and second concatenated input data for each of a plurality of data pairs in the training data, wherein the mode used as the first mode of the data pair is different between the first concatenated input data and the second concatenated input data.
[0035] Training can use benchmark real data from the training dataset, which for each of a plurality of data pairs in the training data includes cascaded input data of a first mode and a second mode of the data pair, wherein the mode of the data pair used as the first mode is randomly sampled from the two modes of the data pair with a predetermined probability, optionally 50% probability.
[0036] Training the machine learning model includes learning a first conditional distribution (P(C|A), P(A|C)) for a pair of modalities (A, C), where the training data does not contain a sufficient number of training data pairs that correlate the data of the pair of modalities to parameterize the conditional distribution. This learning is performed by using the machine learning model to predict data of one or more additional modalities (B) that form a link between the pair of modalities, the prediction using a learned further conditional distribution (P(C|B)) for the modal pair (C, B) that forms the link, and for that modal pair, the training data contains a sufficient number of training data pairs that correlate the data of the pair of modalities to parameterize the further conditional distribution.
[0037] Training a machine learning model may include, for each pair of training data, which includes first data of a first modality and corresponding second data of a second modality in a set of at least three modalities: using a pass-and-exchange machine learning model with the second data as input to predict third data of a third modality in a set of at least three modalities corresponding to the second data; using a pass-and-exchange machine learning model with the third data or subsequent data predicted in the latest iteration as input to predict final data of the first modality corresponding to the predicted third data; and calculating a loss that compares the first data and the predicted final data; wherein the training data includes: (i) pairs of data of the second modality and corresponding data of the third modality, and (ii) pairs of data of the third modality and corresponding data of the first modality.
[0038] Training a machine learning model may include, for each data pair comprising training data including data pairs comprising first data of a first modality and corresponding second data of a second modality of a set of at least three modalities: using a pass-and-exchange machine learning model with the second data as input to predict third data of a third modality of a set of at least three modalities corresponding to the second data; using a pass-and-exchange machine learning model with the predicted third data or predicted subsequent data as input to iteratively predict subsequent data of another modality of a set of at least three modalities corresponding to the predicted third data or predicted subsequent data; using a pass-and-exchange machine learning model with the third data or the most recently iteratively predicted subsequent data as input to predict final data of a first modality corresponding to the most recently iteratively predicted subsequent data; and calculating a loss for comparing the first data with the predicted final data; wherein the training data comprises: (i) pairs comprising data of the second modality and corresponding data of the third modality, and (ii) pairs comprising data of the most recently iteratively iteratively iteratively iteratively iteratively iteratively iteratively iteratively iteratively iteratively iteratively predicting subsequent data of the other modality.
[0039] Training a machine learning model may include modeling using causal masks to: identify a subset of lexical units in the training dataset, wherein the lexical units in the subset are located at multiple positions within the training dataset; append the subset of lexical units to the training dataset; and add masks at multiple positions within the training dataset.
[0040] Training a machine learning model may include: accessing a training dataset comprising first data of a specific input modality from at least three modalities and corresponding second data of another modality from at least three modalities; generating a first predicted output dataset by processing the first data and identifying the other modality using a pass-and-exchange machine learning model, wherein the first predicted output dataset corresponds to the other modality; generating a second predicted output dataset by processing the first predicted output data and identifying a specific target output modality using a pass-and-exchange machine learning model, wherein the second predicted output dataset corresponds to the specific target output modality; generating a third predicted output dataset by processing the second predicted output data and identifying the specific input modality using a pass-and-exchange machine learning model, wherein the third predicted output dataset corresponds to the specific input modality; and calculating a loss by comparing at least a portion of the third predicted output dataset with the first data.
[0041] The method may further include training an additional machine learning model, including the trained machine learning model, wherein the additional machine learning model is trained for a specific task and is a classification model or a regression generative model. The classification model may include the trained machine learning model and a classification model trained to take features extracted from a transitive and exchanged machine learning model as input and produce a classification as output for an input dataset provided to the trained machine learning model, optionally wherein the classification model is trained using training data including data from one or more of the at least three modalities and associated benchmark true classification labels. The generative model may be trained to take an input dataset as input and produce a generated output dataset corresponding to the input dataset as output, wherein the output dataset corresponds to a modality from the at least three modalities that is not present in the input dataset, optionally wherein the trained machine learning model is used as the generative model or the trained machine learning model is fine-tuned using further training data to obtain the generative model. The generative model may be trained to take an input dataset including missing data as input and produce a generated output dataset corresponding to the missing data in the input dataset as output, optionally wherein the trained machine learning model is used as the generative model or the trained machine learning model is fine-tuned using further training data to obtain the generative model. The regression model may include the trained machine learning model, and a regression model trained to take features extracted from the trained machine learning model of the input dataset as input and produce regression values for the input dataset as output. In an embodiment, the regression model is a Cox proportional hazards model, and the regression values are survival predictions and / or the regression model is trained using training data, which includes data from one or more of the set of at least three modalities and associated baseline true regression values.
[0042] According to any embodiment of any aspect, acquiring or accessing input data (e.g., medical data or biological sample data) may include a processor that receives input data from a database, a data acquisition device (e.g., a medical data acquisition device, a biological sample data acquisition device, such as a microscope, a sequencer, or a computing device associated therewith), a computer-readable medium, a user interface, or a computing device.
[0043] Any approach may include, for example, providing one or more results of the approach to a user via a user interface. Results may include classification predictions, regression predictions, predicted output data associated with one or more modalities (e.g., the output of a generative model), one or more training parameters of a trained machine learning model, training the machine learning model, and / or any information derived therefrom. The data repository may be a public or private database. Information derived from classification predictions may include one or more of the following: prognostic indications derived from classification obtained using a machine learning model, treatment indications derived from classification obtained using a machine learning model, indications of suitability for clinical trials derived from classification obtained using a machine learning model, etc.
[0044] According to a fourth aspect, a system is provided comprising: one or more processors and one or more computer-readable storage devices storing instructions that cause the one or more processors to perform any of the methods described herein, such as methods of any embodiment of any of the foregoing aspects, particularly including any embodiments of the first, second, and / or third aspects. In some embodiments, the system may include, for example, one or more computers, servers, or cloud-based devices.
[0045] According to a fifth aspect, a non-transitory computer-readable storage medium is provided, which includes machine-executable instructions that, when executed on a processor, cause the processor to perform any of the methods described herein, such as the methods of any embodiment of the first, second, and / or third aspects, including any one or any combination thereof, provided they are compatible, with reference to the optional features listed therein.
[0046] According to a sixth aspect, a computer program including executable code is provided, which, when run on a computer, causes the computer to perform any embodiment of any method described herein, such as any embodiment of the first, second, and / or third aspects, including any one or any combination thereof, provided that they are compatible, with reference to the optional features listed therein.
[0047] The present invention includes combinations of the described aspects and preferred features, unless such combinations are clearly not permitted or explicitly avoided. Attached Figure Description
[0048] Embodiments and experiments illustrating the principles of the invention will now be discussed with reference to the accompanying drawings, in which: Figure 1 shows a flowchart of a method for using and / or obtaining a multimodal model according to this disclosure.
[0049] Figure 2 illustrates an embodiment of the system according to this disclosure.
[0050] Figures 3a, 3b, and 3c illustrate Venn diagrams showing the relationships between datasets with different modalities A, B, and C. Overlapping datasets indicate that the datasets contain samples with aligned modalities (i.e., audio, image, and text files belonging to the same concept). While recent work has focused primarily on datasets where at least some samples have all available modal combinations (a, b), embodiments of this disclosure consider the case where some modal combinations (e.g., (A, C) and (A, B, C)) are completely missing (c).
[0051] Figures 4a, 4b, 4c, and 4d schematically illustrate training strategies for multimodal models according to embodiments of the present disclosure (LoReTTa). The figures illustrate two novel self-supervised strategies: exchange and transitive modeling. (a) The exchange modeling shown applies causal modeling in an exchange manner to generate A from modality B and B from modality A—given aligned input samples (A, B). The same technique is applied for aligned but disjoint data points (B', C). (b) To ensure the model learns bidirectional relationships between input terms, the shown embodiment uses a modified variant of the generated pre-trained model (called causal mask modeling). (c, d) Transitive pre-training is used to learn the conditional joint distribution for any missing modalities. This randomly selects samples and uses linked modalities B to predict missing modalities C, which are then used to reconstruct existing modalities A. The final step ensures that all modalities are properly aligned.
[0052] Figure 5 illustrates the properties of the training and test data used in the example (SVL-MNIST dataset) of this disclosure.
[0053] Figure 6 illustrates the properties of the training and testing data used in the example (TCGA-OMICS dataset) of this disclosure.
[0054] Figure 7 shows pseudocode illustrating the pre-training of a model according to an example of this disclosure. The figure outlines the main ideas and algorithmic steps for training a model using exchange and transitive modeling, and how to integrate both into a causal modeling framework. Detailed Implementation
[0055] Aspects and embodiments of the invention will now be discussed with reference to the accompanying drawings. Other aspects and embodiments will be apparent to those skilled in the art. All documents mentioned herein are incorporated by reference.
[0056] Unless the context otherwise requires, the methods described herein are computer-based. In reality, image analysis using deep learning models and the process of training deep learning models are extremely complex, particularly requiring intricate mathematical analysis of large amounts of data, which makes the methods described herein far beyond the capabilities of mental research. Therefore, apart from the described structural components and user interactions, any methods described herein can be implemented in a computer system (particularly in computer hardware or computer software).
[0057] The term "computer system" includes hardware, software, and data storage devices for embodying the system or performing the methods according to the embodiments described above. For example, a computer system may include one or more processing units, such as a central processing unit (CPU) and / or a graphics processing unit (GPU), input components, output components, and data storage. Preferably, the computer system has a monitor to provide a visual output display. The data storage may include RAM, a disk drive, or other computer-readable media. A computer system may include multiple computing devices connected via a network and capable of communicating with each other on the network. It is explicitly envisioned that a computer system may consist of or include cloud computers.
[0058] The methods described herein may be provided as a computer program, a computer program product, or a computer-readable medium carrying a computer program configured to perform the methods described herein when run on a computer. The term "computer-readable medium" includes, but is not limited to, any non-transitory medium or media that can be directly read and accessed by a computer or computer system. Such media may include, but is not limited to, magnetic storage media such as floppy disks, hard disk storage media, and magnetic tape; optical storage media such as optical discs or CD-ROMs; electronic storage media such as memory, including RAM, ROM, and flash memory; and mixtures and combinations thereof, such as magnetic / optical storage media.
[0059] The features disclosed in the following specification, the following claims, or the accompanying drawings, expressed in their particular form or as appropriate with respect to the components used to perform the disclosed functions or the methods or processes used to obtain the disclosed results, may be used alone or in any combination of such features to implement the invention in its various forms.
[0060] While the invention has been described in conjunction with exemplary embodiments described below, many equivalent modifications and variations will be apparent to those skilled in the art upon presentation of this disclosure. Therefore, the exemplary embodiments of the invention described above are to be considered illustrative rather than restrictive. Various changes may be made to the described embodiments without departing from the spirit and scope of the invention.
[0061] To avoid any ambiguity, any theoretical explanations provided herein are intended to improve the reader's understanding. The inventor does not wish to be bound by any of these theoretical explanations. Any section headings used herein are for organizational purposes only and should not be construed as limiting the topics described.
[0062] Throughout this specification, including the following claims, unless the context otherwise requires, the words “comprise” and “include”, as well as variations such as “comprises” and “comprising” and “including”, should be understood to imply inclusion of the stated integers or steps or groups of integers or steps, but do not exclude any other integers or steps or groups of integers or steps. As used herein and in the appended claims, unless the context explicitly states otherwise, the singular forms “a,” “an,” and “the” include plural referents. A range may be expressed herein as from “about” one particular value and / or to “about” another particular value. When such a range is expressed, another embodiment includes from one particular value and / or to another particular value. Similarly, when a value is expressed as an approximation, the particular value will be understood to form another embodiment by using the antecedent “about.” The term “about” in relation to numerical values is optional and means, for example, + / - 10%. The expression “and / or” as used herein will be considered as a specific disclosure of each of two particular features or components, including or excluding the other. For example, “A and / or B” will be regarded as a specific disclosure of each of (i) A, (ii) B and (iii) A and B, as if each were listed separately in this document.
[0063] This disclosure generally relates to multimodal machine learning models trained using exchange modeling and transitive modeling. Such models may be referred to herein as “transitive and exchange models” or simply “machine learning models” or “models”, since transitivity and exchange are inherent properties of the models described herein, arising from the way they are trained.
[0064] Multimodal machine learning models are models trained using multiple data modalities. For example, image and text data can be used to train a multimodal model. The assumption is that the multimodal model can learn a more robust representation of the underlying phenomena represented in the data. For example, a multimodal model can be used to classify images using both image content and associated caption text, or to predict caption text for an input image. Multimodal models are typically trained using data that includes corresponding observations for each of the multiple modalities being modeled. These can also be called tuples or aligned data. For example, the text and image multimodal model described above is trained using paired image and caption data points. This is both inflexible (because it requires knowing at the start of training which data will be available when the model is deployed) and generally impractical when using more than two modalities (because datasets including aligned data for three or more modalities are rare in real-world scenarios). Collecting multimodal datasets with two paired modalities A and B or B and C is difficult in practice. Obtaining a dataset with three aligned modalities A, B, and C can be more difficult, and importantly, less reliable. The number of potential relationships between the dependent and independent variables scales exponentially with the number of independent variables and the intermediate transformations between them. Therefore, the extent to which a dataset of a given modality can be used to predict data from another modality depends on the complexity and relationships of the data in the two modalities. These complexities become even more pronounced when considering data from more than two datasets. For example, Figure 3a illustrates a scenario where data from three modalities is available. Figure 3b illustrates a scenario where some entries from various modalities have corresponding data from one other modality, some other entries have corresponding data from two other modalities, and other entries have no corresponding data from any other modality. Figure 3c illustrates a scenario where some entries in the first modality have corresponding data in the second modality, but no entries in the first modality have corresponding data in the third modality. It is not uncommon, in practice, that data for a given modality of interest may not be usable for training a model, even if other potentially relevant data are available. For example, in some cases, only modality A and modality B are paired; and in others, only modality B and modality C are paired. Combinations A and C may be unconditionally missing, although predicting data from modality C based on data from modality A (or based on data from modalities A and B) may be meaningful. Furthermore, identifying datasets with reliable associations across modalities is particularly difficult when timing may be important. For example, the extent to which results from multiple medical tests can be meaningfully associated with each result can depend on the relative timing of the tests, the precision of the tests, the reliability of the tests, and / or the condition of the subjects.Therefore, it is common for some of these data modal pairs to be largely or completely lacking in the training dataset used to train multimodal models, and existing multimodal learning methods cannot adequately handle these situations.
[0065] The introduction of transformers has brought about a significant transformation in the performance and capabilities of the fields of Natural Language Processing (NLP) and Computer Vision (CV). These models excel at a wide range of complex tasks, including machine translation, text summarization, question answering, sentiment analysis, natural language generation, image classification, image captioning, semantic segmentation, object detection, and face recognition—achieving state-of-the-art performance in each domain. Transformers have also opened up new possibilities for multimodal learning due to their lack of inductive bias. They can use self-attention and cross-attention mechanisms to process and fuse different types of data, such as audio, images, video, and text. Therefore, transformers are considered universal modality-agnostic learners.
[0066] However, their lack of inductive bias comes at a cost. Unlike other neural network architectures, transformers make few assumptions about the structure of the data. This makes them highly expressive, but they also require large amounts of data to learn effectively and generalize well [Hoffmann et al., 2022]. Therefore, self-supervised learning (SSL) on large amounts of unlabeled data has been proposed for transformers. This strategy has been very successful in NLP, where methods such as mask modeling and causal modeling are used to train large language models (LLMs) such as BERT [Devlin et al., 2019] and GPT [Radford et al., 2018] variants. Recently, attempts have been made to unify mask modeling and causal modeling through frameworks such as prefix modeling, permutation modeling, causal mask modeling [Aghajanyan et al., 2022] or unified language learning. When applied to image classification, lexical-based pre-training methods have proven to outperform existing computer vision methods such as SimCLR [Hayes et al., 2022] and DINO [Caron et al., 2021]. Therefore, these sequence modeling techniques have been successfully adapted and extended to a wide range of fields, including programming, biochemistry, and reinforcement learning.
[0067] Adding or combining different modalities during training opens up three possibilities: positive transfer, cross-modal understanding, and cross-modal generation. Positive transfer occurs when learning in one modality helps improve the performance of another. For example, it has been shown that aligning images with text can achieve better classification performance on ImageNet [Radford et al., 2021b]. Cross-modal understanding refers to the interpretation of one modality in relation to another. Models with this capability can use text to describe images [Alayrac et al., 2022] or infer protein structures from sequences [Jumper et al., 2021]. With cross-modal generation, one can even generate another modality from one modality. Examples are text-to-image generation and text-to-music generation. Many modern multimodal foundational models are based on contrastive learning [Girdhar et al., 2023], masking modeling [Wang et al., 2022], or causal modeling [Yu et al., 2022].
[0068] This disclosure proposes a novel pre-training method that combines advancements in self-supervised learning into a unified framework. In its implementation, LoReTTa employs a combination of causal modeling and masking modeling for self-supervised learning of multimodal models. The proposed method extends these two approaches by seamlessly integrating and inferring modalities through exchange and transitive modeling. This enables the model to learn expressive features and relationships within a multimodal dataset. The pre-trained model can be used for various downstream tasks, such as imputation, classification, survival prediction, or cross-modal generation.
[0069] Multimodal models have been proposed previously. For example, the RNN-based model MCTN (Multimodal Recurrent Transformation Network Model, Pham et al., 2019) is a neural model that uses modality transformation to learn joint representations of multimodal data. This model is based on the insight that the transformation from the source modality to the target modality results in an intermediate representation that captures joint information between modalities using only the source modality at test time. This approach differs from current methods in several ways: it can only be used as a classifier, does not model all modality combinations, and, most importantly, requires all modalities to be aligned and present at training time. On the other hand, the method described in this paper does not require such strong assumptions. Models such as GPT, BERT, and their derivatives (e.g., the transformer-based models CM3 and C2M3 used in the examples below) are able to make multimodal predictions using, for example, cross-attention. However, they do not learn joint representations by generating missing modalities at training time, and therefore their performance is generally poor when missing modality pairs are present in the training data. Utilizing a much larger amount of training data can improve performance, but it still cannot compare to the method described in this paper in terms of leveraging the additional information present in incompletely aligned multimodal data. In other words, given any multimodal training dataset that includes incompletely aligned multimodal data, the method described in this paper is expected to outperform models like BERT and GPT because they are specifically designed to leverage the information present in multimodal data even in the presence of incompletely aligned data by generating missing modalities to connect observation pairs bidirectionally at training time (i.e., using passivity and exchange modeling as described in this paper). Furthermore, these models will not be able to learn the relationships between modalities if a pair of modalities is always missing in the training data. Models designed for multimodal learning, such as CLIP (Khosla et al., 2020), can handle training datasets with missing modalities. CLIP jointly trains the image encoder and text encoder to use a contrastive loss to predict which pairs out of all possible pairs in a set of N texts and N images are actual pairs. This contrastive loss maximizes the cosine similarity of the image and text embeddings for the N real-valued pairs in the set, while minimizing the cosine similarity of the incorrectly paired embeddings. This allows for the learning of surface features but does not explicitly model missing modalities, thus its performance rapidly degrades when the training data includes most or consistently missing modal pairs. Additionally, CLIP requires a large amount of data for training, making it computationally expensive. In contrast, by explicitly modeling missing modalities using transitive and exchange modeling as described herein, the method of this invention can provide a model with improved performance compared to all the models described above, as shown in the examples below. As also shown in the examples below, the performance of the comparative model can be improved by increasing the computational power used for training.For example, while performance does improve with a CLIP training length of 3x and a larger batch size (5x), it still doesn't match the performance of the new model described herein. Therefore, this disclosure provides a model capable of achieving performance levels that can only be achieved by comparing models using significantly larger amounts of training data and / or more computationally intensive training schemes, which are essentially equivalent to the currently described model providing a more computationally efficient method for training multimodal models in the presence of missing modalities in the training data.
[0070] The terms “analysis” and “processing” can be used interchangeably in the context of applying the model described herein to input data, since the described model is trained to analyze the input data by processing the data.
[0071] Machine learning models can be deep learning models. Machine learning models can also be sequence models. Sequence models are machine learning models that take a sequence of data as input. A data sequence is any string of input data that has a sequence (order). This contrasts with a set of input data representing a collection of independent data points. Sequence models are typically deep learning models. Deep learning models can be deep neural networks, such as: (i) recurrent neural networks (RNNs) or RNN-based models, such as models combining recursion and attention mechanisms, such as receive-weighted key-value (RWKV) models (Peng et al., 2023) and Griffin (De et al., 2024) or long short-term memory (LSTM) models; (ii) convolutional neural networks (CNNs) or convolution-based models, such as models combining element-wise multiplication (gated) and long convolutions (convolutions with filter sizes the same as the input), such as Hyena (Poli et al., 2023) or models combining convolution and attention mechanisms, such as Striped Hyena (Poli et al., 2023b); (iii) state-space models, such as Mamba (Gu et al., 2023) or Hawk (De et al., 2024), or hybrid transformer state-space models, such as Jamba (Lieber et al., 2024); or (iv) A fully attention-based model is also known as a transformer-based model. In this embodiment, the machine learning model is a transformer-based model. In other words, any machine learning model as described herein can include a transformer model. Machine learning models can include transformer models that use self-attention. Transformer models can solve complex and diverse tasks such as machine translation, text summarization, question answering, sentiment analysis, natural language generation, image classification, object detection, face recognition, and image captioning—achieving state-of-the-art performance in each domain. Due to the lack of inductive bias, transformers are well-suited for multimodal learning. Different types of data, such as audio, images, text, and video, can be processed and fused together using self-attention and cross-attention mechanisms. Therefore, transformers are considered general-purpose, and the learner is modality-agnostic.
[0072] Transformer-based models are a class of machine learning models, particularly deep learning models, that incorporate attention mechanisms (self-attention and / or cross-attention). Transformer-based models rely primarily or entirely on attention mechanisms to learn from lexical units in context, and typically do not rely on recurrent structures to achieve this. The attention mechanism in a transformer uses multiple scaled dot-product attention units. Each attention unit produces an embedding for each lexical unit, which contains information about the lexical unit itself and a weighted combination of other relevant lexical units (weighted by attention weights). Specifically, each attention unit is associated with three learned weight matrices: query weight W... Q Key weight W K Sum weight W V Embed the input of each word i into x. i Multiply by each of the three weight matrices to obtain the query vector q. i Key vector k i Sum vector v i In the original transformer and ViT, the attention weights from word i to word j are computed as the dot product between qi and kj. Then, through... These weights are scaled, where dk is the dimension of the key vector. The scaled attention weights are passed through the softmax function and multiplied by the value vector, i.e., the output of the attention unit for word i is the weighted sum of the value vectors of all words, weighted by the attention weights from word i to each word in the corresponding word. In the context of this disclosure, any machine learning model that includes the attention mechanism or its derivatives as described above can be referred to as a transformer-based model. For example, a derivative of the original transformer described above includes a relative positional bias B for each head. The learned matrix W Q W K and W V Attention heads are formed. Transformer models typically include multiple attention heads in each of multiple layers. In practice, the attention mechanism can be repeated h times in parallel to extract more information. Each attention head focuses on the lexical units associated with each input lexical unit, and the use of multiple attention heads implies that this operation can be repeated for different definitions of the associated content. Then, the final projection matrix W learned for each multi-head attention unit is used. OThe outputs of multiple attention heads in a layer are concatenated and projected. Additional layers may exist. For example, a fully connected layer typically follows an attention layer. This may include one or more further transformations (e.g., two linear transformations in a typical transformer) and / or one or more activation functions (e.g., ReLU activation in the original transformer, although other activation functions such as GELU have been proposed). Around each of the two sub-layers of the fully connected layer, the original transformer is coupled with residual connections. All these components (multi-head attention and subsequent fully connected layers with residual connections) constitute an encoder or decoder block repeated n times to produce a transformer-based encoder or decoder. The encoder output may be further processed to produce predictions for a specific task. For example, the encoder output may be processed by a decoder with a similar architecture for generating tasks. The decoder block may include a self-attention mechanism, an attention mechanism on top of the encoding, and a feedforward neural network. As another example, a classification head such as a multilayer perceptron (MLP) may be used to process the encoder output for classification. For example, class terms may be appended to the embeddings of the input, and the MLP may be used to compute the logits of class terms from the encoder output.
[0073] As explained above, the trained multimodal model according to this disclosure can be further trained or used for any downstream task for which it can utilize a pre-trained multimodal sequence model. In particular, the pre-trained model can be applied to any of positive transfer, cross-modal understanding, and cross-modal generation (and improve upon prior art methods in their context). Positive transfer refers to a situation where training multimodal learning helps improve performance for a downstream task using the model compared to performing the same task with an equivalent model trained using a single modality. A multimodal model can be used with one or more modalities in downstream tasks. For example, a multimodal model can be trained using aligned image and text data and subsequently used only the images as input for image classification. Cross-modal understanding refers to a situation where one modality is interpreted in relation to another modality. For example, a model can be trained to describe an image using text and answer text-based questions about the image. Cross-modal generation refers to the generation of one or more modalities using one or more other modalities. Positive transfer and cross-modal understanding can be utilized in many downstream tasks, including classification, regression, and so on. Furthermore, cross-modal generation can be used to transform data between modalities.
[0074] Specifically, this disclosure relates to a computer-implemented method for training a multimodal machine learning model using exchange modeling and transitive modeling, wherein the resulting model can be further trained (also known as fine-tuning or transfer learning) and / or included as part of a machine learning model comprising additional machine learning modules (e.g., classification models, regression models, or generative layers) trained for any downstream task of interest. Such a model may be referred to as a “base model,” and the derived model trained for the downstream task may be referred to as a “task-specific model.” The training and use of both types of models are explicitly included in this disclosure.
[0075] Therefore, some embodiments of this disclosure include training and / or using machine learning models to transform data between modalities. The machine learning model can be configured and / or trained to operate in an commutative and / or transitive manner. Thus, the machine learning model can be trained to receive inputs from any combination of multiple modalities (e.g., three or more modalities) as input and transform the inputs into outputs from any of the other multiple modalities. The model can utilize generative pre-training, commutativity (A, B) = (B, A), and transitive relations. This involves linking modalities and transforming between them. Machine learning models may include novel LoReTTa (Linking mOdalities with a tRansitive and commutative pre-training sTrAtegy) methods, as described further below. The machine learning model may have been trained using a self-supervised framework that uses commutativity and / or transitivity rules to transform between different modalities and causal generative pre-training. Thus, it can learn the relation modality A → modality C using modality A → modality B → modality C. Given a dataset containing only combinations (A, B) and (B, C), a transformer (or any general neural sequence model) pre-trained according to the techniques disclosed in this paper can dispose of any modality combination at inference time after fine-tuning for downstream tasks, including never-before-seen pairs (A, C) and triplet (A, B, C) simultaneously matching the performance of an oracle model (a model already trained with combinations of all entries (A, B, C)).
[0076] The methods disclosed herein may include using transfer learning to train a model. Transfer learning is a machine learning technique in which a model is pre-trained to perform a first task and then used (alone or as part of a machine learning model that includes additional machine learning modules, such as classification heads, regression heads, etc.) to perform a second task different from the first task. The model may be used as is for the second task, or it may be at least partially retrained (also known as fine-tuning) for the second task. In such embodiments, the parameters of the model learned for the first task are used as a starting point for optimizing the second task, and some or all of these parameters may be further optimized. Further optimization may freeze some of the parameters learned in the first task (i.e., they do not change during optimization of the second task) and / or apply a lower learning rate, such that the optimization explores a more constrained space around the weights optimized for the first task. For example, the machine learning model may be trained using training data including data from multiple modalities and causal language modeling, mask language modeling or combination methods (such as causal mask modeling), exchange modeling as described herein, and transitive modeling as described herein. The resulting base model can then be directly used for any generative task of interest, such as generating data for one or more modalities of interest using input data from one or more other modalities of interest. Generally, any generative task is possible, such as generating complete data for one or more modalities of interest using input data that includes incomplete data from any one or more modalities of interest. The resulting base model can be fine-tuned for any generative task, such as by fine-tuning any weights of the base model using task-specific training data (e.g., training data that includes data associated with one or more groups of subjects of interest). The resulting base model can be used, respectively, as part of a classification or regression model that additionally includes a classification or regression module. A classification module refers to any machine learning model that takes extracted features from a trained base model (trained exchange and transitive machine learning model) of the input dataset as input and provides an output indicating the class to which the input dataset belongs. For example, in the context of binary classification, the output could be a single probability of belonging to one of two classes. In the context of multi-class classification, the output could include multiple probabilities of belonging to the respective multiple classes. The classification module can be referred to as a "classification head" or "classification layer," particularly in the context of deep neural networks. A classification layer is typically a single layer in a deep neural network. A classification head may include multiple layers. A regression module refers to any machine learning model that takes extracted features from a trained base model (trained exchange and pass-through machine learning model) of interest as input and provides continuous outputs of interest. For example, a regression module can predict continuous values such as the probability of survival over a predetermined time period, odds ratio, predicted patient response to treatment, etc. As a specific example, a regression module could be a survival model, such as the Cox proportional hazards model.The regression module can be referred to as a "regression head" or "regression layer," especially in the context of deep neural networks. A regression layer is typically a single layer in a deep neural network. A regression head may include multiple layers.
[0077] Classification refers to classifying input data between two or more categories using a machine learning model. For example, speech and / or image and / or text data can be classified between multiple categories. A classification method or model can be associated with a level of accuracy evaluated using one or more metrics on one or more test datasets. Any classification accuracy metric known in the art can be used, including, for example, top-k accuracy (where k can be, for example, 1, 2, or 3), classification accuracy (% of correct predictions), or F1 score. Top-k accuracy is the number or percentage of correctly classified labels among the top k labels predicted, ranked by probability for each class. The F1 score is the harmonic mean of precision (the number of true positive predictions divided by the total number of positive predictions—including both true and false positives) and recall (the number of true positive predictions divided by the total number of true positive instances, i.e., the sum of true and false negative predictions) for binary classification. In multi-class classification problems, the F1 score is obtained as the average of the F1 scores for each class (optionally weighted proportionally to class to account for class imbalance). Classification accuracy is the ratio of the number of correct predictions made for a dataset to the total number of predictions made. Regression methods or models can be associated with a level of accuracy evaluated using one or more metrics on one or more test datasets. Any regression accuracy metric known in the art can be used, including, for example, the c-index (especially in the case of survival models), mean squared error, root mean square error, and mean absolute error. The c-index (consistency index) is a generalization of the area under the ROC curve (AUC) that takes censored data into account. The c-index of a survival model represents the model's ability to correctly provide a reliable ranking of survival times based on individual hazard scores. It is a measure of the probability that the predicted event times of two observations (subjects) have the same relative order as their actual event times. The c-index can be calculated as: Where α i T is the risk score of the observed object i. i and T j It is the event time of observed objects i and j, 1 Tj<Ti In T j < T i In the case of , it equals 1; otherwise, it equals 0, and 1 αj>αi In α j > α iThe value is 1 if the condition is met, and 0 otherwise. A value of 1 corresponds to the best possible prediction, and a value of 0.5 corresponds to a random prediction. Mean Squared Error (MSE) is the average of the squared differences between the predicted and expected (benchmark true) values in the dataset. Root Mean Squared Error (RMSE) is the square root of the mean squared error. Mean Absolute Error (MAE) is the average of the absolute differences between the predicted and expected (benchmark true) values in the dataset. RMSE, MSE, and MAE associate better-performing models with lower values of these metrics. Generative sequence models can be associated with a level of accuracy evaluated using one or more metrics on one or more test datasets. Any metric known in the art for evaluating sequence models can be used, such as, for example, perplexity. Perplexity is the exponentially averaged negative log-likelihood of the sequence. For a lexicalized sequence X=(x0, x1,…,x…) t The perplexity of X is Higher confusion is associated with higher-performing models.
[0078] The methods of this disclosure are expected to outperform comparative models trained on the same multimodal training data and / or with the same computational load. Performance can be evaluated using any metric known in the art, depending on the task the model is configured to perform, such as, for example, as explained above. The methods of this disclosure are expected to outperform comparative models used for prediction tasks where the input data for the prediction task includes observations of modality pairs that are not represented or poorly represented in the training data of the model. A modality pair is considered poorly represented in the training dataset when the number of training data pairs including data of that modality pair is significantly less than the number of training data pairs including data of any other modality pair. For example, the number of training data pairs including data of that modality pair may be less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the number of training data pairs including data of any other modality pair. The benefits associated with the methods of this disclosure are expected to exist in any such cases and increase as the amount of training data for any modality pair decreases.
[0079] The method of embodiments of this disclosure includes lexicalizing data before it is provided as input data to a machine learning model. Lexicalization is the process of taking a data representation into a predetermined numerical space. A lexicalized representation may also be referred to as an embedding. For example, a transformer model takes a sequence of lexical terms as input, so when using a transformer-based model, the data must first be lexicalized. Several ways to transform data into lexical terms are known in the art, and any such known method can be used. For example, data can be directly transformed into lexical terms using an underlying raw byte stream (as described by Reed et al., 2022). For example, any of text, names, and other discrete data can be encoded using a lookup table (also known as a dictionary). Specific examples include lexicalizing text data using byte-pair encoding (as in, for example, BERT), SentencePiece (as in, for example, GPT), or WordPiece. As another specific example, raw numerical data can be divided into a predetermined number of cells, each cell associated with a value, thereby forming a dictionary of a size corresponding to the predetermined number of cells. As another example, any low-dimensional audio, image, and other continuous data can be encoded using its raw byte values (as in ImageGPT, Perceiver, for example). For instance, each pixel of an image or the frequency of an audio file can be used as a single term. High-dimensional data can be first reduced in dimensionality and then lexicalized. Dimensionality reduction can be performed by projection into a lower-dimensional space and / or by segmenting the data (e.g., chunking). For example, high-dimensional continuous data such as large images or audio can be chunked and lexicalized first via a codebook or dictionary. A codebook is a vector representation in a discrete latent space. The codebook can be obtained from a vector quantization variational autoencoder (as in DALLE and MusicLM, for example). Note that a codebook can also be used without prior dimensionality reduction. A dictionary is a lookup table in which raw values in the data are mapped to a predetermined number of values in the dictionary. In an embodiment, a learned function is used to obtain the codebook, which projects the input data into a latent space with a predetermined dimension d. This can use multiple trainable weights trained simultaneously (or before) the rest of the machine learning model. For example, weights trained simultaneously with the rest of the model can be used to linearly project a flattened image or image patch into a latent space. A patch (also called a "tile") is a sub-segment of a larger data object, such as an image. Patches can also be referred to as "plots". In the case of an image, a patch can have a predetermined size, typically represented in terms of the number of pixels in each of the two dimensions. Patches may be non-overlapping. A flattened image or patch is a data object corresponding to an image, where a two-dimensional matrix containing the values of image pixels is transformed into a one-dimensional vector by concatenating rows or columns.Dimensionality reduction can be performed by projecting onto a lower-dimensional space, which can be achieved by applying principal component analysis to the training dataset to identify multiple principal components and by following a predetermined set of principal components (e.g., the first 10, 20, 100, 200, 1000 principal components, depending on the expected size of the resulting data) for any input data.
[0080] The methods disclosed herein are applicable to many fields, particularly those where it is difficult or impossible to obtain or collect fully aligned multimodal data. Fully aligned multimodal data refers to multimodal data in which each observation is associated with data from each of the multiple modalities under consideration (strongly aligned data) and / or where a large number of aligned data pairs exist for all possible data modal pairs (weakly aligned data). Multimodal data refers to a dataset that includes data associated with different types of information and / or acquired using different sensors / acquisition methods. For example, multimodal data may include data obtained from multiple sensors measuring different physical, physiological, physicochemical, or biological parameters. As another example, multimodal data may include data in the form of images, text, audio, etc. As another example, multimodal data can include data from multiple modalities selected from: medical images (such as, for example, MRI, CT, or PET imaging data, each of which represents a different modality), electronic medical record data (such as electronic medical record representations, textual and / or categorical data representing medical history, including treatment, comorbidities, etc., one or more of which may represent a different modality), and biological sample data (such as, for example, omics data obtained from samples, or digital pathology imaging data, where each type of omics data or digital pathology imaging data may represent a different modality). Omics data can refer to, for example, gene sequence data (e.g., gene sequence representations, which may include, for example, the identification of the presence of one or more gene sequences or variants in a subject or a sample from a subject, the presence or expression level of one or more gene expression products in a sample from a subject, etc.). Omics data can be selected from: mRNA expression data (e.g., mRNA sequencing data, gene expression array data, RT-PCR data, digital PCR data, etc.), genomic data (e.g., genome sequencing data, such as whole genome sequencing data, whole exome sequencing data, targeted panel sequencing data, or data derived therefrom, such as copy number information, variant information, mutation signature information, etc.), epigenetic data (e.g., DNA and / or histone modification data), proteomics data (e.g., data from reversed-phase protein arrays, mass spectrometry, etc.), and metabolomics data (e.g., data from any metabolite detection assay, including but not limited to mass spectrometry, targeted detection assays, such as glucose or other metabolite detection, etc.). The term "medical image" refers to any digital image of a biological subject. A medical image can be an image of a subject or a portion thereof. For example, a medical image can be a radiographic image, a magnetic resonance imaging (MRI) image, a positron emission tomography (PET) image, or an ultrasound image. Any imaging technique used in a clinical context to make decisions based on morphological features visible in the image may benefit from the methods of this disclosure.As used herein, the term digital pathology encompasses the analysis of histopathological images (images of tissues), cytopathological images (including images of individual cells analyzed for their characteristics), hematopathological images (images of blood or blood cell samples), or anatomical pathological images (images of organs or tissues at various resolution levels). Digital pathology images can be images of biological samples that have been fixed and stained with one or more staining agents. Any staining method known in the art can be used, such as, for example, H&E staining (hematoxylin and eosin), immunofluorescence, silver nitrate, Romanovsky-Gymza staining, and trichrome staining. Medical images and digital pathology images are commonly analyzed in the context of cancer diagnosis, prognosis, and treatment; neurodegenerative disease diagnosis, prognosis, and treatment; and cardiovascular disease diagnosis, prognosis, and treatment. The same applies to omics data and electronic health record data. Therefore, omics data and / or electronic health record data and / or medical image data and / or digital pathology image data are often used individually or in combination to obtain diagnostic, prognostic, treatment recommendations, or patient selection (e.g., for participation in clinical trials) for a wide range of conditions, including but not limited to cancer, neurological disorders, and cardiovascular diseases. Machine learning methods for performing these tasks advantageously benefit from reproducibility, reliability, speed, and often higher accuracy compared to analyses performed by trained healthcare professionals. However, no machine learning method has yet been proposed that can utilize multimodal data to the extent currently achieved by manual analysis, because prior to this disclosure, no multimodal machine learning method was designed specifically to handle training datasets including missing modality pairs. This disclosure provides a solution that addresses this, thereby opening new possibilities for multimodal machine learning-based predictions in the medical field. Therefore, this document also describes methods for processing subject-associated input data for the purposes of disease diagnosis, prognosis, monitoring, or therapy identification, comprising processing said input data using models or methods as described herein. Biological samples or specimens can be any cell-containing sample, whether obtained directly from the subject or from cell or tissue cultures. In embodiments, biological samples are samples previously obtained from the subject. Samples can be samples of tissue or bodily fluids (e.g., biopsies or liquid biopsies).
[0081] Figure 1 illustrates a method for analyzing data using a transitive and exchanged machine learning model as described herein, and a flowchart for providing a method for training a transitive and exchanged machine learning model as described herein.
[0082] A method for training, transferring, and exchanging machine learning models will now be described. At step 100, a training dataset is obtained by a processor. The training data may be received from a user interface, a database, a data repository, a computing device, or a data acquisition device. All subsequent steps are performed by the processor or a processor communicatively coupled to the processor that obtained the training data. The training data comprises data from at least three modalities. Therefore, step 100 may include accessing a training dataset comprising data from each of the at least three modalities. The training data may lack data that associates first data of a particular input modality with second data of a particular output modality. The training data may include, or consist of, data that associates or comprises modality pairs of data from the at least three modalities, wherein for each pair of the first and second modalities in the at least three modalities, the training data includes data that associates the data of the modality pair that together form a path between the first and second modalities. In an embodiment, for each pair of the first and second modalities in the at least three modalities, the training data includes data that associates the data of the first modality with a third modality and data that associates the data of the third modality with the second modality. Therefore, the training data may not include training data that associates any pair of modalities, but may include data that uses one or more bridging (also referred to herein as "links") modalities to associate modalities in a missing pair. The training data may include tuples (e.g., pairs) containing corresponding data for multiple modalities from at least three modalities, or composed of such tuples, where the number of tuples containing data for a specific pair of modalities from at least three modalities is less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% present in the training data. In embodiments, tuples may not include tuples containing data for a specific pair of modalities from at least three modalities (i.e., a pair of modalities is completely or substantially missing from the training data).
[0083] Three or more modalities may be selected from medical or biological sample data associated with the subject. For example, one or more modalities may be selected from: medical image data, electronic medical record data, genetic sequence data, and digital pathology image data, optionally wherein the medical image data is selected from MRI, CT, or PET imaging data.
[0084] At step 110, the training data is used to obtain input data for the machine learning model. This may include one or more of the following optional steps: At step 110A, PCA may be used, for example, to reduce the dimensionality of one or more modalities in the training data. At step 110B, a predefined lexicalization scheme may be used to lexically represent the data for each modality in the training data individually. At step 110C, the lexically represented data obtained at step 110B may be projected into a common feature space across modalities. This may be done using a parameterized embedding function. At step 110D, input data pairs may be generated, each pair including first data from a first modality and associated second data from a second modality. Note that the training data may include data associated with more than two different modalities (also referred to as “tuples”) (i.e., data from more than two modalities of the same subject or object). In such cases, at step 110D, multiple pairs may be sampled from the same tuple. At step 110E, a multi-modality-specific lexical can be added to the pair obtained at step 110D to indicate the start and / or end of a modality. At step 110F, a modality-specific absolute positional encoding can be added to the feature vector obtained at steps 110D or 110E. This can be learned simultaneously with the model's weights.
[0085] The model is trained using pairs obtained from the training data, which include first data for a first modality and associated second data for a second modality. Before providing multimodal pairs (e.g., pairs obtained at steps 110D, 110E, or 110F) as input to the model for training, the order of the two modalities in each pair can be randomly sampled between a first and a second order. Alternatively, pairs corresponding to the two orders can be generated. This is referred to as exchange modeling performed at step 120. Therefore, step 120 can include generating training benchmark ground truth data from the training dataset, which includes concatenated input data for each of multiple data pairs in the training data, comprising the first and second modalities of the data pair, wherein the modality used as the first modality in the data pair is randomly sampled from the two modalities of the data pair with a predetermined probability. The predetermined probability can be 50%. This allows the model to infer the first modality from the second modality and the second modality from the first modality at step 130. Therefore, the model can be trained in the first step 120-130, where exchange modeling and self-supervised learning objectives are used to train the model.
[0086] At step 130, the model can be trained using causal language modeling (CLM), masking modeling (MM), and / or any variant thereof, such as causal masking modeling (CMM), prefix language modeling [PLM, Raffel et al., 2019], permutation modeling [Yang et al., 2020], unified language learning [uniLM, Dong et al., 2019; UL2 / hybrid denoiser, Tay et al., 2023], LLM2Vec [BehnamGhader et al., 2024], and GIVT-Causal / GIVT-MaskGIT [Tschannen et al., 2024]. Therefore, step 130 may include training the model to take partially masked concatenated input data of a first mode and a second mode of a data pair in a training dataset as input and produce a prediction of the masked input data as output. The training dataset includes benchmark real data for multiple data pairs, each data pair including corresponding data for a pair of modes, wherein the mode used as the first mode of the data pair is randomly sampled from the two modes of the data pair, and / or the benchmark real data includes first concatenated input data and second concatenated input data for each of the multiple data pairs in the training data, wherein the mode used as the first mode of the data pair is different between the first concatenated input data and the second concatenated input data.
[0087] In embodiments such as those illustrated below, causal masking modeling (CMM) is used to train the model at step 130. At this step, the model is trained to learn the conditional distribution of any modality pairs present in the training data. Causal generative modeling, masking modeling, and / or variations thereof can use benchmark real data, which includes concatenated input data of a first modality and a second modality for each of a plurality of data pairs in the training data. The modality used as the first modality of the data pair can be randomly sampled from both modalities of the data pair. Alternatively, the benchmark real data can include first and second concatenated input data for each of a plurality of data pairs in the training data, wherein the modality used as the first modality of the data pair is different between the first and second concatenated input data. Any of these options ensures commutativity when training the model using self-supervised objectives such as causal language modeling (CLM), masking modeling (MM), and / or variations thereof. Causal masking modeling may include: identifying a subset of lexical units in the training dataset, wherein the lexical units in the subset are located at multiple positions within the training dataset; appending the subset of lexical units to the training dataset; and adding masks at multiple positions within the training dataset. Causal masking modeling may further include using the identification of multiple mask positions during training.
[0088] At step 140, transitive modeling is used to further train the model. The weights of the model at step 140 may be initialized to the weights learned at the first steps 120-130 using exchange modeling, and further training is performed at step 140 using transitive modeling. Alternatively, exchange modeling and transitive modeling may be combined for training. Training using transitive modeling may include training the model against each of a plurality of data pairs in a training dataset, each data pair including first data of a first modality and corresponding second data of a second modality in a set of at least three modalities: using a machine learning model with the second data as input to predict third data of a third modality in a set of at least three modalities corresponding to the second data; using transitive and exchange machine learning models with the third data as input to predict final data of the first modality corresponding to the predicted third data; and calculating a loss that compares the first data with the predicted final data. In such embodiments, the training data may include: (i) pairs including data of the second modality and corresponding data of the third modality, and (ii) pairs including data of the third modality and corresponding data of the first modality. Training a model using transitive modeling may include training the model against each of a plurality of data pairs in a training dataset, each data pair including first data of a first modality and corresponding second data of a second modality in a set of at least three modalities; using a machine learning model with the second data as input to predict third data of a third modality in a set of at least three modalities corresponding to the second data; using a transitive and exchange machine learning model with the predicted third data or predicted subsequent data as input to iteratively predict subsequent data of another modality in a set of at least three modalities corresponding to the predicted third data or predicted subsequent data; using a transitive and exchange machine learning model with the most recently predicted subsequent data as input to predict final data of the first modality corresponding to the most recently predicted subsequent data; and calculating a loss that compares the first data with the predicted final data. In such embodiments, the training data may include: (i) pairs including data of the second modality and corresponding data of the third modality, and (ii) pairs including data of the most recently predicted other modality and corresponding data of the first modality, and corresponding pairs including data of each of the two modalities used in the step of iteratively predicting subsequent data of the other modality.In other words, the method may include: at step 100 accessing a training dataset comprising first data of a specific input modality and corresponding second data of another modality; and at step 140 generating a first predicted output dataset by processing the first data and identifying the other modality using a pass-and-exchange machine learning model, wherein the first predicted output dataset corresponds to the other modality; generating a second predicted output dataset by processing the first predicted output data and identifying a specific target output modality using a pass-and-exchange machine learning model, wherein the second predicted output dataset corresponds to the specific target output modality; generating a third predicted output dataset by processing the second predicted output data and identifying the specific input modality using a pass-and-exchange machine learning model, wherein the third predicted output dataset corresponds to the specific input modality; and calculating a loss by comparing at least a portion of the third predicted output dataset with the first data.
[0089] The trained model generated by step 140 is a transfer and exchange machine learning model because it is a machine learning model that has been trained to transform an input dataset corresponding to any and every modality in the set of at least three modalities into an output dataset corresponding to any and every other modality in the set of at least three modalities, wherein the machine learning model has been trained to transform data between two modalities by using training data, which includes data pairs corresponding to a first modality and another modality in the two modalities and other data pairs corresponding to another modality and a second modality in the two modalities, and the machine learning model has been trained to perform a corresponding inverse type modal transformation for any forward type modal transformation that the model has been trained to perform. Steps 130 and 140 together can also be described as training a machine learning model to learn a first conditional distribution (P(C|A), P(A|C)) for a pair of modalities (A, C), for which the training data does not contain a sufficient pair of training data to correlate the data of the pair of modalities to parameterize the conditional distribution. The learning is performed by predicting the data of one or more additional modalities (B) forming the link by using a learned further conditional distribution (P(C|B)) of the modal pair (C, B) that forms the link between the pair of modalities, and for the modal pair, the training data contains a sufficient pair of training data to correlate the data of the pair of modalities to parameterize the further conditional distribution.
[0090] The trained model generated in step 140 can be considered the final trained model used on the input dataset, such as for generation tasks like modality padding or transformation. Alternatively, the model can be further trained at step 150 for different generation tasks, and / or included as part of a classification or regression model trained at step 150. The classification model can be a model that includes the model obtained at step 140 (the transitive and exchange machine learning model), and a classification model trained at step 150 to take features extracted from the transitive and exchange machine learning model as input and produce a classification as output for the input dataset provided to the transitive and exchange machine learning model. The classification model can be trained at step 150 using training data that includes data for one or more of the at least three modalities in the set and associated benchmark true classification labels. The generative model may include the trained machine learning model obtained in step 140, or a model composed of the trained machine learning model, which is further trained in step 150 to take the input dataset as input and produce a generated output dataset corresponding to the input dataset as output, wherein the output dataset corresponds to a modality among the at least three modalities that are not present in the input dataset. In other words, the generative model may be trained to take the input dataset as input and produce a generated output dataset corresponding to the input dataset as output, wherein the output dataset corresponds to a modality among the at least three modalities that are not present in the input dataset. Alternatively, or otherwise, the generative model may be trained to take an input dataset including missing data as input and produce a generated output dataset corresponding to the missing data in the input dataset as output. The trained machine learning model obtained in step 140 may be used as the generative model for this purpose, or the trained machine learning model obtained in step 140 may be fine-tuned in step 150 using further training data to obtain the generative model. The regression model can be a model that includes the model obtained at step 140 (the pass-and-exchange machine learning model), and a regression model trained at step 150 to extract features from the pass-and-exchange machine learning model as input and produce regression values for the input dataset as output. The regression model can be a Cox proportional hazards model, and the regression values can be survival predictions. Alternatively, the regression values can be any continuous value associated with the input data, such as probabilities (e.g., the probability of having or developing a disease, the probability of responding to treatment, etc.). The regression model can be trained at step 150 using training data comprising data from one or more of the at least three modalities in the set and associated baseline true regression values.
[0091] For example, at step 150, a machine learning model may be trained to take medical or biological sample data associated with the subject as input and produce an instruction for recommended treatment as output. This instruction may be categorized. Medical or biological sample data may include one or more of the following: radiographic images, gene sequences, and microscope slides. As another example, the machine learning model may be trained at steps 130-140 and optionally further trained at step 150 to take data collected by any one or more types of sensors associated with the vehicle, optionally an automobile, as input and generate predicted sensor data as output for another type of sensor data associated with the vehicle, optionally wherein the machine learning model has been trained using synchronized data from pairs of multiple types of sensors.
[0092] At optional step 160, one or more machine learning models or any parameters thereof obtained as a result of step 140 or 150' are provided to the user or computing device.
[0093] Interchangeability modeling is applied to the training of machine learning models. We ensure commutativity by randomly mixing the concatenation order when appending lexical units from both modalities to the input sequence as described above. This allows the model to infer modality A from B and modality B from A. Random mixing is performed by randomly reversing the default order of the modalities provided as inputs to the machine learning model using a predetermined probability. In a fully interchangeable model, this probability can be 50%, meaning that a pair of inputs is equally likely to be represented as (A, B) or (B, A). This probability can be set to a value different from 50% (e.g., below 50%) to prioritize learning in a predetermined direction. This can be useful in some cases, such as depending on the task the model is intended to be trained for and the combination of modalities intended to be used when testing and / or deploying the model. In embodiments, interchangeability modeling is used to pre-train the model using a self-supervised learning objective. The self-supervised learning objective can be a causal generative modeling objective. Therefore, interchangeability modeling can be used to train the model in the context of a self-supervised learning task, such as causal generative modeling. Causal generative modeling can include one or more of the following: causal language modeling, masked language modeling, and their derivatives.
[0094] Transitive modeling is applied to the training of machine learning models. Consider an example where three modalities A, B, and C are modeled, with only combinations (A, B) and (B, C) present in the training data. Training a machine learning model can be performed by randomly sampling available pairs ((a, b), (b, a), (b, c), or (c, b)) from the training data during training and applying transitive modeling. Transitive modeling involves predicting the missing modality of a multimodal dataset by inferring any missing data points needed to transition between two modalities using learned relationships between the available modalities for their data pairs (also referred to herein as bridging or linking modalities; in the example above, B is the linking modality between modalities A and C). For example, consider the pair (a, b). Since the model is trained to predict modality C from B, pseudo-data point c can be inferred from b (i.e., the value of predicted c is given even though no baseline real data is available to correspond to the value c of data a). Since the model can also predict mode A from C, a new sample c (pseudo-data point) can be used to predict point a again, thus ensuring all modes are aligned. The same principle can be applied to (b, a), where a pseudo-data point c can be inferred from a, which can then be used to predict point b, where the loss between the predicted point b and the baseline true point b is minimized, the loss being associated with the prediction of pseudo-data point c from point a and the prediction of b from pseudo-data point c.
[0095] In one embodiment, obtaining or generating input data at step 110 may include generating the input dataset by lexicalizing the initial dataset (step 110B). In another embodiment, obtaining or generating input data at step 110 may include labeling the initial dataset (step 110B) and projecting the lexicalized initial dataset into a predetermined feature space using a parameterized embedding function for each modality (step 110C). A parameterized embedding function is a function with learned parameters. This may be implemented as one or more additional layers of a deep learning model, the parameters of which are learned during training of the deep learning model. In another embodiment, obtaining or generating input data at step 110 may include generating the input dataset at step 110E by appending modality-specific lexics before and / or after the lexicalized and optionally projected data corresponding to each individual modality in the input dataset. In yet another embodiment, obtaining or generating input data at step 110 may include including learned positional embeddings at step 110F. Learned positional embeddings are positional embeddings learned during the training of a machine learning model.
[0096] In one embodiment, all input data are lexicalized using a selected predetermined scheme. After lexicalization, a parameterized embedding function is used to project all lexical units into a common feature space (e.g., projecting the lexicalized representations of samples a, b, c of modalities A, B, and C into a shared vector space: a -> [a_1, …, a_l], b -> [b_1, …, b_m], and c -> [c_1, …, c_n]). In this embodiment, the embedding function is implemented as a linear layer with parameters learned together with other parameters of the model (i.e., the weights of all layers, including the embedding layer, are learned together). This is advantageously simple. Alternative implementations of the parameterized embedding function include convolutional layers or multiple layers (i.e., additional neural networks for embedding), such as, for example, transformer blocks. Such implementations typically involve more parameters to be trained, but may be beneficial in some cases depending on the complexity of the input data. Furthermore, modality-specific lexical units are added to indicate the start and end of a modality. In this embodiment, the learned modality-specific absolute position encoding is added to the feature vector. The embeddings for each modality are then concatenated upon alignment (e.g., if a and b are aligned, the model input would be s = [a_1,…, a_l; b_1,…, b_m] or s = [b_1,…, b_m; a_1,…, a_l], where the use of both orders ensures commutativity). For example, for audio-image pairs or image-text pairs describing the same concept, their lexical sequences can be concatenated and fed into a model (e.g., a transformer model). However, if only one modality is available, only the lexical sequence of that modality is used as the model input.
[0097] In this embodiment, causal generative modeling is used to train the model. Causal generative modeling refers to the task of predicting the next word in a sequence given a sequence of past words. This can be represented as follows: given a sequence of words s_1:L and parameters θ, it can be expressed as modeling the data using a probabilistic chain rule: log pθ(s1, ..., sL) = sum log pθ(sl |s1, ..., sl−1). As explained above, commutativity in this framework can be ensured by randomly mixing the concatenation order when appending words from both modalities to form the input sequence. This allows the model to infer modality A from B and from A to B. This is referred to as commutative modeling in this paper.
[0098] In this embodiment, causal generative modeling is used to train the model to use causal masking modeling. Causal masking modeling randomly masks a segment (or multiple segments) of lexical units and moves them to the end. This allows the model to still use the information after the mask to predict the masked position, which is impossible in classical generative modeling. This is not necessary for the performance of the model described in this paper, but it is advantageous. In fact, models such as those that generate pre-trained transformers only consider information from the left and ignore information from the right; they lack bidirectional context, which is important for many downstream tasks. The use of causal masking modeling enables the training of general models that can be fine-tuned for many different downstream tasks. Other methods that unify masking modeling and causal modeling can also be used, such as, for example, prefix language modeling (PLM), permutation modeling, unified language learning, LLM2Vec, and GIVT-Causal / GIVT-MaskGIT. The latter are associated with a specific architecture (generative infinite vocabulary transformer, GIVT), while other methods are generally applicable to decoder-only language models.
[0099] In embodiments, the machine learning model may be trained to prioritize each of the exchange and / or transfer processes. Such training may include configuring the inputs to the machine learning model to identify specific target output modalities (e.g., since any given input can be transformed into the output of multiple output modalities). Training may further or alternatively include defining training techniques and / or loss functions (e.g., cross-entropy loss function) to bidirectionally and / or across three or more data types to correlate and / or prioritize data transformations.
[0100] In an embodiment, a machine learning model can be pre-trained using exchange modeling of the self-supervised learning objective (e.g., causal language modeling, masked language modeling, and / or derivatives thereof, as further explained herein), but without transitive modeling. This can utilize available pairs of linked data modalities. The resulting pre-trained model can then be used as a starting point for further training using transitive modeling. In other words, the machine learning model can be pre-trained using a first step and a second step, where the first step uses exchange modeling and a self-supervised learning objective, and in the second step, the machine learning model weights are initialized to the weights learned in the first step, and further training is performed using transitive modeling.
[0101] A method for analyzing data will now be described. At step 100', a processor obtains data to be predicted. The obtained dataset includes data from one or more modalities selected from a set of at least three modalities. This set of at least three modalities corresponds to a set of at least three modalities that the machine learning model used in step 170 below has been trained to process. Step 100' may include receiving data associated with one or more modalities. Step 100' may include obtaining or accessing input data (such as, for example, medical data or biological sample data) by receiving input data from a database, data acquisition device (e.g., a medical data acquisition device, a biological sample data acquisition device, such as a microscope, sequencer, or associated computing device) via a processor from a database, a data acquisition device (e.g., a medical data acquisition device, a biological sample data acquisition device, such as a microscope, a sequencer, or an associated computing device), a computer-readable medium, a user interface, or a computing device. This may include receiving data from a user interface, a database, a data repository, a computing device, or a data acquisition device. All subsequent steps are performed by a processor or a processor communicatively coupled to the processor that obtains the data. At optional step 110', the data is processed to obtain input data for a trained machine learning model as described herein. Step 110' may include one or more of the following: dimensionality reduction at step 110A', lexicalization at step 110B', projection into the common feature space at step 110C', addition of modality-specific lexical units at step 110E', and addition of positional encoding at step 110F'. Any of these steps may be performed as explained above with respect to steps 110A, 110B, 110C, 110E, and 110F.
[0102] At step 170, the input data is processed using a trained machine learning model as described herein, for example, as obtained at step 140 or step 150. Therefore, step 170 may include generating an output dataset by processing the input dataset using a machine learning model that includes a pass-and-exchange machine learning model, wherein the pass-and-exchange machine learning model is a machine learning model trained to transform an input dataset corresponding to any and every modality in the set of at least three modalities into an output dataset for any and every other modality in the set of at least three modalities, wherein the machine learning model has been trained to transform data between two modalities using training data comprising data pairs corresponding to a first modality and another modality of the two modalities and other data pairs corresponding to the other modality and a second modality of the two modalities, and the machine learning model has been trained to perform a corresponding inverse type modality transformation for any forward type modality transformation performed by the model. At optional step 180, derived data, such as prognostic indicators, treatment options, diagnostic indicators, selection of participation in clinical trials, etc., may be obtained from the output of the machine learning model. This step is optional, and the output of the machine learning model obtained at step 170 may already be directly interpretable as one of these types of information. At step 190, the results of any one or more of the preceding steps are provided to the user, for example, via a user interface. For example, the results may include classification predictions, regression predictions, predicted output data associated with one or more modalities (e.g., the output of a generative model), one or more training parameters of the trained machine learning model, training the machine learning model, and / or any information derived therefrom. The data repository may be a public or private database. Information derived from classification predictions may include one or more of the following: prognostic indications derived from classification obtained using the machine learning model, treatment indications derived from classification obtained using the machine learning model, indications of suitability for participation in clinical trials derived from classification obtained using the machine learning model, etc.
[0103] A trained machine learning model can be used to generate an output dataset for a specific target output modality using an input dataset of one or more other modalities. Therefore, step 100' may include determining a specific input modality of the input dataset and an instruction to receive a specific target output modality. Step 170 may include generating the output dataset by processing the input dataset and recognizing the specific target output modality using a passing and exchanging machine learning model.
[0104] In embodiments, as previously described, a machine learning model may be trained and / or used to transform data between modalities. Thus, the model can be used to identify the output of a specific target output modality using an input dataset of (other) specific input modalities. Once the output of the specific target output modality is identified, the output may be output (e.g., transmitted to another device or presented via an interface), the output may trigger or be used to identify actions (e.g., movement of a prosthesis or driving control in a car), and / or the output may be used to generate another result (e.g., it may be output, trigger another action, or be used to identify another action). As an example, a specific target output modality may include a predicted MRI scan, which may then be post-processed to identify a potential diagnosis. Results may be generated to include a potential diagnosis, a predicted MRI scan, and one or more potential treatments (e.g., determined based on the potential diagnosis). The results may be transmitted to a care provider's device.
[0105] Collecting multimodal datasets with two paired modalities A and B or B and C is difficult in practice. Obtaining datasets with three aligned modalities A, B, and C is even more challenging. For example, in healthcare, we sometimes have only gene sequences and microscopic images of one patient, and only gene sequences and radiographic images of another patient. This makes it difficult to integrate and combine all modalities into a large pre-trained model. This disclosure provides a solution to this under-explored problem. Embodiments of this disclosure utilize a self-supervised framework that combines causal generative pre-training with commutativity and transitivity rules to transition between different modalities. Thus, the proposed solution can learn the relation A → C using A → B → C, where, in contact, the training data contains only aligned data for pairs (A, B) and (B, C), but not (or very little) aligned data for pair (A, C). In particular, the proposed method can use generative pre-training, commutativity (A, B) = (B, A), and transitivity relations. This method links and transforms modalities. Given a dataset containing only combinations (A, B) and (B, C), the inventors demonstrate that a transformer pre-trained using the proposed method can handle any modal combination at inference time after fine-tuning for downstream tasks, including never-before-seen pairs (A, C) and triplet states (A, B, C), while matching the performance of oracle models. By learning intra- and lexical relationships from different data distributions, the inventors demonstrate in the examples below that the method can be used as a general and multimodal feature extractor. This has profound implications for safety-critical fields such as healthcare, infrastructure, and transportation. For example, in medicine, it is already difficult to obtain large datasets with a single modality to train powerful base models—such as those used in computer vision and natural language processing. Finding datasets with two or more modalities is even more challenging. However, combining them is important because it is common practice to look at different types of data, such as radiographic images, gene sequences, and surgical samples, to determine the best treatment.
[0106] Adding or combining different modalities during pre-training opens up three possibilities: positive transfer, cross-modal understanding, and cross-modal generation. Positive transfer occurs when learning in one modality helps improve the performance of another. For example, it has been shown that aligning images with text can produce better classification performance on ImageNet (see, for example, CLIP, SimVLM). Cross-modal understanding refers to using one modality to interpret another. For example, one can use text to describe an image (see, for example, Flamingo) or understand a protein from a sequence (see, for example, AlphaFold). And with cross-modal generation, one modality can be used to generate another. For example, text-to-image generation (e.g., DALLE-2) and text-to-music generation (e.g., MusicLM). All these models are based on contrastive learning (e.g., ALIGN, CLIP, FLORENCE), masking modeling (e.g., VisualBERT, ViLBERT, Pixel-BERT, ImageBERT, VL-BERT, VD-BERT, LXMERT, UNITER, VinVL, BeiT-3), or causal modeling (e.g., SimVLM, Flamingo, Gato, DaVinci, CM3, DALLE-1, Parti, Pali, GPT-4). This disclosure provides a novel approach to training models capable of performing each of these types of tasks. This approach unifies existing methods in masking and causal modeling into a general self-supervised framework. The model trained as described herein can then handle different modality combinations in various downstream tasks, such as classification. Compared to previous approaches that only considered the contributions of randomly missing modalities, this approach is designed to handle cases where the entire data pair is always missing. This creates a powerful model that can be configured and trained to predict many types of data or perform any other type of generating task of interest, or handle any different types of modal combinations for various downstream tasks, including but not limited to classification and regression or similar tasks such as survival prediction.
[0107] To illustrate the types of modality transformations that can be supported using a pass-and-change machine learning model (configured and trained according to the various embodiments disclosed herein), the following exemplary use cases are provided. In an embodiment, the machine learning model as described herein (i.e., the pass-and-change machine learning model) is configured to receive a flattened and lexicalized version of any of the three types of MRI scans, PET scans, or CT scans, and generate any other type of scan among the three types. The machine learning model can be trained using MRI / CT scan pairs and CT / PET scan pairs. Once trained, one potential use of the model is to transform an input MRI scan into a predicted PET scan or vice versa, even if concurrent scans of these types are not available. In an embodiment, the machine learning model as described herein (i.e., the pass-and-change machine learning model) is configured to receive embedded data collected by any of the various types of sensors (e.g., cameras, sonar sensors, lidar sensors, etc.) in or on a vehicle. Data pairs are generated by synchronizing sensor readings. Once trained, one potential use of the model is to receive input corresponding to one type of sensor data and then generate data for each of the other types; another potential use is to receive input corresponding to two types of sensor data (e.g., cascaded together) and then generate data for another type. These uses can produce three or more synchronized sensor datasets. Importantly, traditional signal synchronization is error-prone, and the more synchronizations performed, the higher the accumulated error becomes. Therefore, reducing the need for synchronization by using a pass-and-switch machine learning model instead can actually produce more accurate and richer datasets (and reduce the number of sensors required in the vehicle). In an embodiment, the machine learning model as described herein (i.e., the pass-and-switch machine learning model) is configured to receive embedded data collected by any of the various types of sensors (e.g., cameras, accelerometers that detect motion of a part of a user's body, EMG electrodes that detect muscle contractions, etc.). Data pairs are generated by synchronizing sensor readings. Once trained, one potential use of the model is to receive input corresponding to one type of sensor data and then generate data for each of the other types; another potential use is to receive input corresponding to two types of sensor data (e.g., cascaded together) and then generate data for another type. These applications can generate three or more synchronized sensor datasets. As mentioned in the examples above, traditional signal synchronization is prone to error, and the more synchronizations performed, the higher the accumulated error becomes. Therefore, reducing the need for synchronization by instead using pass-and-exchange machine learning models can actually produce more accurate and richer datasets (and reduce the number of sensors required to generate information for the limbs).
[0108] Figure 2 illustrates an embodiment of a system for analyzing data and / or providing a trained machine learning model according to the present disclosure. The system includes a computing device 1 comprising a processor 101 and a computer-readable storage device 102. In the illustrated embodiment, the computing device 1 also includes a user interface 103, exemplified as a screen, but may include any other means such as conveying information to a user via sound or visual signals. The computing device 1 is communicatively connected, for example, via a network to a data acquisition device 3 (such as, for example, a medical image data acquisition device and / or an omics data acquisition device and / or a medical data acquisition device), such as a microscope associated with a histopathology station, a medical imaging device such as an MRI or PET scanner, or an associated computing device, and / or connected to one or more databases or data repositories 2 storing data such as, for example, medical data. The one or more databases 2 may further store one or more of the following: one or more deep learning algorithms, training data, parameters (such as, for example, parameters of a deep learning model used for analyzing data), clinical and / or sample-related information, etc. The data acquisition device 3 may be configured to acquire medical image data, omics data, or any other type of medical data from a patient or sample. The computing device may be a smartphone, tablet, server, personal computer, or other computing device. The computing device is configured to implement methods for analyzing data and / or for training machine learning models, as described herein. In an alternative embodiment, computing device 1 is configured to communicate with a remote computing device (not shown), which itself is configured to implement the methods as described herein. In such cases, the remote computing device may also be configured to send the results of the methods to the computing device. Communication between computing device 1 and the remote computing device may be via a wired or wireless connection and may occur via a local or public network 6 (such as, for example, via the public internet). Data acquisition device 3 and / or data storage device 2 may be wired to computing device 1 or may be able to communicate via a wireless connection (such as, for example, via WiFi and / or via the public internet), as illustrated. The connection between computing device 1 and data acquisition device 3 and / or data storage device 2 may be direct or indirect (such as, for example, via a remote computer).
[0109] The examples presented below are intended to be illustrative and should not be construed as limiting the scope of the claims.
[0110] Example Training multimodal base models is challenging due to the limited availability of multimodal datasets. While many public datasets pair images with text, few combine images with audio or text with audio. Even rarer are datasets that simultaneously align all three modalities. Key domains such as healthcare, infrastructure, or transportation are particularly affected by missing modalities. This makes it difficult to integrate all modalities into a large pre-trained neural network that can be used out of the box or finely tuned for different downstream tasks. This example introduces LoReTTa (Linking modalities with a tRansitive and commutative pre-training sTrAtegy) to address this research shortcoming. The described self-supervised framework unifies causal modeling and masking modeling with commutativity and transitivity rules. This allows the method to transition within and between modalities. Therefore, the pre-trained model is better suited to exploring the true underlying joint probability distribution. Given a dataset containing only disjoint combinations (A,B) and (B,C), LoReTTa can model the relation A↔C using A↔B↔C. Specifically, this example demonstrates that a transformer pre-trained with LoReTTa can handle any modal mixture at inference time, including never-before-seen pairs (A,C) and triplet states (A,B,C). This example extensively evaluates the novel approach on synthetic, medical, and reinforcement learning datasets. Across different domains, the new general-purpose multimodal transformer consistently outperforms strong baselines such as GPT, BERT, and CLIP on tasks involving missing modal tuples.
[0111] method Model Architecture. The LoReTTa implementation used in this example employs the autoregressive transformer decoder architecture described in Radford et al.
[2019] , with some modifications as described below. The model includes modified initialization, pre-normalization, and GELU activation. The inventors chose to use RMSNorm [Zhang et al., 2019] instead of LayerNorm [Ba et al., 2016]. They also replaced StandardAttention [Vaswani et al., 2017] with FlashAttention [Dao et al., 2022] to reduce memory usage and runtime. Following Chowdhery et al. [Chowdhery et al., 2022], the model is trained without dropout and without bias in any dense kernel or layer norm. While any general sequence model can be used for next-word prediction, the inventors chose the transformer [Vaswani et al., 2017], especially because of its simplicity and scalability. However, in theory, this method is also applicable to other backbones, such as the recently proposed RWKV (Peng et al., 2023) and Hyena (Poli et al., 2023).
[0112] Data Representation. Since the transformer takes a sequence of tokens as input, the input data needs to be lexicalized. This example uses common methods such as raw byte encoding [Reed et al., 2022; Jaegle et al., 2021; Chen et al., 2020]; learned lookup tables [Kudo and Richardson, 2018; Sennrich et al., 2016; Schuster and Nakajima, 2012] and vector quantization [Van Den Oord et al., 2017; Razavi et al., 2019; Ramesh et al., 2021; Agostinelli et al., 2023]. After lexicalization, a parameterized embedding function (implemented here as a simple linear layer with learned parameters) is used to project all tokens into a common feature space. Furthermore, modality-specific tokens are added to indicate the start and end of a modality. Modality-specific absolute positional encodings are added to the feature vectors and learned. The representations of each modality are then concatenated, but only if they belong to the same object / subject. For example, if we have image-audio pairs or image-text pairs describing the same concept, we concatenate their lexical sequences and feed them into the transformer. However, if only one modality is available, we only use that modality. In summary, the preprocessing steps used in this example are described below: 1. Lexicalize all samples and project them a, b, c into a common vector space to obtain the lexical embeddings [a1, ..., a2]. l ],[b1,…,b m ] and [c1,…,c n ]. 2. If a and b are aligned (i.e., they are related to the same object / subject), then concatenate them as s = [a1, ..., a l ;b1,…,b m ] or s = [b1,…,b m ;a1,…,a l This ensures commutativity.
[0113] Pre-training strategy. Causal language modeling (CLM) is used to train the transformer model. Given a sequence of lexical units s... 1:L Given the parameter θ, the data is modeled using the probability chain rule given in equation (1): Equation (1) Therefore, the model's task is to predict the next lexical given a previous sequence of lexical terms. As shown in GATO [Reed et al., 2022] and PALM-E [Driess et al., 2023], this goal is modality-independent. The system even learns a rich internal world model [Lin et al., 2023, Li et al., 2023]. Compared to conventional CLM methods, the inventors ensure commutativity by randomly mixing the concatenation order when appending lexical terms from both modalities to the input sequence (step 2 of the preprocessing method described in the "Data Representation" section), as illustrated in Figure 4a. This allows the model to infer B from A as well as modality B from A. The inventors refer to this technique as commutative modeling. Interestingly, this simple modification has not been used in recent multimodal base models.
[0114] Because models that generate pre-trained transformers only consider information from the left and ignore information from the right, they lack bidirectional context [Wang et al., 2022], which is critical for many downstream tasks. Therefore, the inventors combine mask modeling (GPT-style modeling) and causal modeling (BERT-style modeling) into a single framework. This recently developed approach is known in the literature under various names: Intermediate Imputation (FIM) [Bavarian et al., 2022], Causal Mask Modeling (CM3) [Aghajanyan et al., 2022, Fried et al., 2023], Hybrid Unidirectional Modeling (HybUni) [Artetxe et al., 2022], and Span Mask + Language Modeling (SCLM) [Tay et al., 2022]. All these approaches are conceptually identical, differing only in minor implementation details. Therefore, throughout this disclosure, references to causal mask modeling include any implementations of the above-described embodiments and any other implementations of the same concept. Conceptually, the idea is to randomly mask a span (or multiple spans) of lexical units, replace them with placeholder lexical units, and move the original lexical units to the end (illustrated in Figure 4b). This change does not alter the training process based on equation (1); however, it now allows the model to use information from the right side of the mask to predict masked locations, which is not possible in traditional unidirectional causal modeling. Hybrid views not only yield better performance in downstream tasks after fine-tuning but also retain the practicality of language modeling. Causal masking modeling has been extended to large-scale vision-language models capable of generating and filling both text and images with CM3Leon [Yu et al., 2023]. CM3Leon is trained with more data than CM3 and utilizes a larger transformer but uses the same algorithm as CM3.
[0115] Transitive modeling. The inventors then extend several language learning paradigms (masking, prefixing, causality, causal masking) using a new technique they call transitive modeling. The problem of learning in contexts with unconditionally missing combinations, as described above, requires combining modalities A and C—given only disjoint combinations (A,B) and (B',C). The inventors propose a method to randomly sample data points (a,b), (b,a), (b',c), or (c,b') during training and apply transitive modeling. For simplicity, this explanation considers the pair (a,b). Since the model is trained to predict modality C from B, it can infer pseudo-data points from b. Now, we will train the transformer to use the new inference samples. To predict points The approach involves minimizing the loss relative to the original sample a, thereby ensuring that all our modalities are aligned (Figure 4c). In this way, we model the conditional distribution of the missing pair (C,A)—which, due to commutativity, is equal to (A,C). This idea shares some conceptual similarities with cycle consistency. However, this method is much more general because it can be applied to a variety of modal combinations, as long as linked modes exist, as illustrated in Figure 4d.
[0116] In fact, upon closer examination, transitive modeling differs significantly from cycle consistency. CycleGAN incorporates the most popular version of cycle consistency [Zhu et al., 2017]. Given an input a in domain / modality A, one aims to generate an aligned output b in domain / modality B. The model then computes a discriminative loss D on b, and compares the original a with the predicted a. The reconstruction loss L is calculated. MCTN [Pham et al., 2019] uses a different version of cycle consistency: given aligned inputs (a, b) from modes A and B, a is used to predict... And using predicted To predict Then the reconstruction loss L is applied to compare a and And b and On the other hand, the transitive modeling described in this paper begins with aligned modes (a,b) and (b',c), and uses b to predict... And use To predict Then rebuild the loss by combining a with The comparisons are then performed. Because this method uses exchange modeling, it also learns another direction starting with different samples (b', c). The three methods discussed above are summarized below: CycleGAN: Given a, for Model and calculate . MCTN: Given (a, b), for Model and calculate . LoReTTa (the method of this disclosure): given (a,b), for Model and calculate .
[0117] Therefore, the transitive modeling described in this paper is not simply achieved by... Expand to This adds another step to cycle consistency. Conversely, the method... Modeling is performed. Intuitively, this ensures that the model does not "cheat" by reconstructing a from its memory input a. The method described in this paper also does not use recurrent loss because a is not used as input or output. Therefore, the loss is relative to the model's input. It is not cyclic. However, it still ensures that the predicted modalities are generally consistent with the data. This is similar to cyclic consistency, but as seen above, it is not. In summary, LoReTTa uses masking modeling to transition within modalities and causal modeling to transition between modalities and vice versa, thanks to exchange modeling. Furthermore, not all modality combinations need to be available at training time, as the method uses transitive modeling to predict missing modalities and concatenate them to the original data. In the experiments described below, the inventors demonstrate that this method allows for the learning of very rich multimodal feature representations.
[0118] Experimental Evaluation Scheme. The inventors conducted empirical analyses of LoReTTa on the constructed dataset (SVL-MNIST), the real-world medical dataset (TCGA-OMICS), and an offline reinforcement learning dataset with three modalities (offline game dataset; MUGEN-GAME). For downstream tasks, they selected object classification, survival prediction, and cross-modal transformation. These three tasks together cover both discrimination and generation problems. The inventors evaluated the former using linear probing, and the latter in a zero-shot scenario. This allows for direct evaluation of the quality of the learned features. The experimental sequence is as follows: First, the inventors conducted an ablation study on the constructed dataset to analyze the impact of transitive modeling on different modal combinations. Second, they compared LoReTTa with other methods such as masking (BERT), causal (GPT), and contrastive (CLIP) modeling on the real-world medical dataset. Third, they compared the new model with a state-of-the-art autoregressive cross-modal transformer on the offline game dataset. In the appendix, we list all optimized hyperparameters and model configurations. Figure 7 shows the pseudocode for the exchange and transitive modeling described and used in these evaluations. The pseudocode below provides a deeper understanding of how LoReTTa pre-training is implemented. It outlines the main ideas and algorithmic steps for training the model using exchange and transitive modeling. Most importantly, the code demonstrates how to integrate both into a causal modeling framework.
[0119] SVL-MNIST. The new method was tested on a custom dataset derived from real-world data, comprising speech (A), vision (I), and language (T) modalities representing 10 different categories. The speech dataset contains approximately 40,000 spectrograms from AudioMNIST [Khacef et al., 2019], the vision dataset includes 70,000 images from MNIST [LeCun et al., 2010], and the language dataset consists of 130,000 documents from WineReviews [Thoutt, 2018]. We link the datasets using their labels 0-9. This method produces weakly aligned modalities. For clarity, we refer to this dataset as Speech-Vision-Language-MNIST (SVL-MNIST). The datasets are randomly split such that they have identical relationships, as shown in Figure 3c. Specifically, the data was split to obtain bimodal datasets (A, I), (A, T), and (I, T), each with 12,000 non-overlapping samples. All remaining samples are part of the unimodal datasets A, I, and T. Further details of the data structure are shown in Figure 5. The SVL-MNIST dataset was divided into training, validation, and test sets. The validation set was merged with the training set after hyperparameter search. To obtain a multimodal dataset with combinations of completely missing modes, the inventors considered three datasets (I, T), (T, A), and (A, I). The first dataset (I, T) consists of 12,000 paired samples from MNIST and WineReviews. The second dataset (T, A) was similarly constructed and contains 12,000 paired samples from WineReviews and AudioMNIST. An additional 12,000 samples were collected from AudioMNIST and combined with the 12,000 samples from MNIST to obtain (A, I). All remaining samples are part of the unimodal datasets I, T, and A. There is no dataset with all three modalities (I, T, A)—except for the test datasets. Note that there is no sample overlap in the datasets. That is, all datasets have the same relationships as in Figure 3c. For fair comparison, the inventors only use a fully aligned subset of the test set to report the final results, as the other test sets contain different samples and have different sizes. However, they keep the data split from the individual unaligned modalities (right side of Figure 5) to make the datasets more flexible for future experiments.
[0120] TCGA-OMICS. The medical dataset contains omics (e.g., genomics, proteomics, and transcriptomics) sequencing values from the Cancer Genome Atlas Project (TCGA) [Weinstein et al., 2013]. The inventors selected three maximum subsets from up to ~11,000 patients, comprising messenger RNA (mRNA), microRNA (miRNA), and reversed protein arrays (RPPA). They aligned the dataset at the patient level and obtained approximately 7,000 patients with all three modalities. To simulate a setting lacking modality combinations, they split the dataset into two subsets (mRNA, miRNA) and (mRNA, RPPA), each with 3,500 non-overlapping samples. They chose mRNA as the linked modality because it is one of the most widely available gene modalities. TCGA is quite unique as it is the only publicly available medical dataset with multiple aligned modalities. Therefore, it is useful for testing the methods described herein. Further details regarding the data structure are shown in Figure 6. The TCGA-OMICS dataset contains 11,069 mRNAs, 10,824 miRNAs, and 7,790 RPPA samples. The inventors aligned the dataset at the patient level, obtaining 7,030 data points across three modalities. 1,030 of these data points were used for testing. The remaining samples formed part of the training and validation sets (again, the validation set was merged with the training set after hyperparameter tuning). Specifically, the training set consisted of (mRNA, miRNA) and (mRNA, RPPA) datasets, each consisting of 3,000 paired samples. As explained above, the inventors chose mRNA as the linking modality because it is typically the richest and most common modality available in medical datasets. As mentioned above, the inventors only used a fully aligned subset of the test set to report the final results, as other test sets contained different samples and had different sizes. However, they kept the data split from the individual unaligned modalities (right side of Figure 6) to make the dataset more flexible for future experiments.
[0121] MUGEN-GAME. As a large-scale and multimodal dataset, the inventors chose MUGEN-GAME [Hayes et al., 2022], which has 375,000 naturally aligned video (V), audio (A), and text (T) samples collected by reinforcement learning agents in the closed-world platform game CoinRun. The protagonist, Mugen, walks, jumps, collects coins, kills monsters, climbs ladders, and dies—triggering various sound effects superimposed with background music. All their adventures are recorded in three-second video clips and audio tracks. A group of human narrators provides rich textual descriptions for each scene. The inventors used the most important modality in video games—video—as the linking modality and considered disjoint datasets (V, A) and (V, T) in the final experiments. MUGEN-GAME consists of 375,368 fully aligned samples. It was divided into training, validation, and test sets of 349, 666, 12,851, and 12,851, respectively. The inventors randomly paired video and audio files, as well as video and text files. This resulted in two disjoint datasets, each with a size of 174,833, used for training.
[0122] Specific implementation details. For the first two experiments (SVL-MNIST and TCGA-OMICS data), the inventors used a transformer decoder with 8 layers and 8 attention heads, with an embedding dimension of 512 and query, key, and value dimensions of 64 each. This compact architecture has proven effective in various tasks and has been widely used in the literature (see, for example, Li et al., 2023). The inventors applied a single parameterized embedder to encode the tokens and added modality-specific learned positional embeddings (as described in Ramesh et al., 2021, Wang and Cho, 2019, and Aghajanyan et al., 2023). For optimization, the AdamW algorithm was used with a learning rate of 6e-4, a weight decay factor of 0.1, and a gradient clipping of 1. During the first few hundred steps, cosine annealing and linear warm-up were used, and the learning rate underwent a 10-fold decay. For the third experiment (MUGEN-GAME), the inventors used the exact same training hyperparameters and model configuration as the baseline model described in [Hayes et al., 2022]. Specifically, the transformer had 12 layers, 8 attention heads, and an embedding size of 768. The query, key, and value dimensions were each 96.
[0123] Models pre-trained using language learning paradigms share a similar set of hyperparameters. This aligns with current best practices for state-of-the-art (multimodal) base models based on transformers. Specifically, the inventors used the AdamW optimizer with a β of (0.9, 0.95), an ε of 1e-8, a weight decay of 0.1, and a gradient shear of 1.0. The learning rate started at 1e-7, increased linearly to 6e-4, and gradually decayed to 6e-5 according to a cosine schedule. To ensure the model saw batches containing different modalities and modality combinations during training, the inventors accumulated batches across different data loaders. For CLIP, they used a β of (0.9, 0.98), an ε of 1e-6, and a weight decay of 0.2. The maximum learning rate started at 5e-4 and ended at 5e-5. A list of batch sizes, warm-up steps, and total training steps for each experiment is available in Tables 5 and 6. The above settings apply to the SVL-MNIST and TCGA-OMICS experiments. The MUGEN-GAME experiment follows the exact hyperparameters from Hayes et al. 2022—the difference being that training was performed for 142,000 steps.
[0124] Lexicalization Scheme. Due to the low dimensionality of the SVL-MNIST dataset, the inventors directly encoded the raw byte stream by binning continuous values into dictionaries of 256 each. TCGA-OMICS data can be very high-dimensional and is best handled differently. For example, mRNA expression arrays typically contain more than 20,500 entries. Therefore, the inventors first used PCA to reduce the size by a factor of 10, and then binned the values into dictionaries of 1,000. The final input sizes were 2,048, 64, and 16, respectively—preserving their original relative lengths. High-dimensional video and audio samples from MUGEN-GAME were lexically coded using their pre-trained VQ-VAE encoder, exactly the same as Hayes et al., 2022.
[0125] Linear Probe. In the SVL-MNIST experiments, we freeze the backbone and train a linear classifier on top. SGD is used as the optimizer, with Nesterov momentum set to 0.9. The initial learning rate is 0.1, but it is gradually reduced to zero during training via cosine annealing. We do not use weight decay. The batch size is 16, and the number of training epochs (based on the validation set) are listed in Table 7. For the TCGA experiments, we fit a Cox proportional hazards model with elastic net penalty to the extracted features. The weights of the regularization terms L1 and L2 are both set to 0.5. Training stops when one of the following criteria is met: tol = 1e-7 or iter = 100000. The underlying optimization algorithm is based on coordinate descent, which continuously minimizes the objective function one step at a time in one direction. It is particularly effective for problems with many features.
[0126] Theoretical Analysis If there exists a joint probability distribution P(X1,...,X...) N A subset S of samples (Y) is considered to have modalities X1,...,X N The datasets are aligned. Here, Y represents the label space. This form of alignment, where only labels are shared across modalities, is called weak alignment [Castrejon et al., 2016]. Conversely, we consider alignment strong when functional relationships exist at the physical object level [Castrejon et al., 2016]. The definitions of strong and weak alignment are also extended to datasets where we do not simultaneously have all modal combinations. For example, during data collection, we might only recover paired samples (x1,x2) and (x2',x3'), where x2 and x2' do not belong to the same object / subject. Thus, we have datasets of the form (X1,X2) and (X2,X3) instead of (X1,X2,X3). This differs from the case considered in previous work where the dataset still contains some samples of the form (x1,x2,x3). To model unseen relationships (X1, X3) in a self-supervised manner, we cannot simply use techniques such as masking or causal modeling (because there are no aligned observations for modalities X1 and X3). The neural network will have to learn implicitly. Equation (2) The method is to derive it from the probability chain rules. and Based on experience (see [Asai and Hajishirzi, 2020, Li and Tou Ng, 2022, Yanaka et al., 2021]), even within a single modality, neither unidirectional nor bidirectional methods can accomplish this task because it is difficult to model p(x3|x2,x1) due to the lack of a set (X1,X2,X3). This is where the proposed transitive modeling paradigm comes in.
[0127] Suppose we have the form (X1,X2), (X2,X3), ..., (X... N-1 ,X N The dataset X, where X ≠ l when k+1 ≠ l. k and X l Misalignment. By applying causal modeling to all pairs, we can learn the transformation X1→X2→…→X N-1 →X N By adding our proposed exchange modeling, we obtain a stronger bidirectional relationship: X1↔X2↔…↔X N-1 ↔X N Now, let's interpret each modality as a node on a high-dimensional graph. This gives us a connected graph (a minimum spanning tree) where we can traverse from one modality to any other. Therefore, given a real data point x... i In this case, we can use transitive modeling to obtain data from x. i Transform to x j To the pseudo data point x j Sampling is performed. After sampling many pairs (x...) i ,x j After that, we obtain a new dataset (X). i ,X j This can be reused during the pre-training process to learn new transformations X. i ↔X j We apply this concept to all i and j. Finally, we learn the conditional distribution among all samples and obtain a fully connected graph. This concept is applicable to more complex mode combinations, as long as bridging modes exist, i.e., we have a spanning tree (see Figure 4d). We can also apply this concept to all missing tuples (x... k ,...,x l Sampling is performed. Therefore, LoReTTa allows us to approximate any missing joint probability distribution P(X) by sampling and training using combinations of missing modes. k ,...,X l ).
[0128] Ideally, there should be no generalization error after pre-training. However, in practice, if available, we will not be able to accurately recover the baseline truth of the missing modalities with perfect accuracy. For datasets (X1,X2) and (X2,X3), the learned generative model f will produce the following error: Equation (3) Among all e i,j These are all vector-valued error terms. Without losing generality, we do not track the sign of the error, because we can always write... Assuming f is distinguishable, and using a multivariate Taylor expansion, we obtain Equation (4) J f It is the Jacobian matrix of f. By iteratively applying the Taylor rule, a more general error propagation formula for longer chains can be obtained. Intuitively, the error from the first transformation is amplified more and more. However, since we use the predicted x3 to reconstruct x1, we can actually define the error term. To recap, we have a baseline true value x1, for which the following holds true: Equation (5) Because we are given during backpropagation In the case of minimizing the reconstruction error of x1, the error term e 1,2 and e 2,3 They must be bounded. And because they constitute the error term e. 1,3 Therefore, it must also be bounded, provided that f and J f Yes, that's fine.
[0129] Experimental evaluation results SVL-MNIST In the first experiment, the inventors used causal masking modeling (CM3) to train a model for each modality (A, I, and T). They then trained four additional models using exchange causal masking modeling (C2M3) on the bimodal datasets (A, I), (A, T), and (I, T), as well as the trimodal dataset (A, I, T). These seven models served as the unimodal baseline, multimodal baseline, and "upper bound" baseline, respectively. They then trained models simultaneously on (I, T) and (A, T) using T as the linked modality in C2M3, repeating this process for all modality combinations. Next, they initialized LoReTTa with the C2M3 model weights to explicitly merge the bimodal datasets via the linked modalities. This pre-training method significantly accelerated training compared to training the entire LoReTTa from scratch. After pre-training, all models were frozen and evaluated in a low-data scenario with only 1,000 labeled samples (100 per class). Three rounds were performed, each with a new randomized subset, and classification accuracy and perplexity on the test set were reported. It's important to note that we trained the models only based on the modalities they saw during pre-training. This also means that during probing, the linear classifier will see some samples of unseen modality pairs. However, the pre-trained backbone itself does not see these modality pairs. In this way, we can analyze how the model combines previously unseen mixtures of input modalities (modality combinations).
[0130] When pre-trained only on pairs (T, I) and (T, A), LoReTTa achieves a significantly low perplexity of 5.04 on the unseen pair (I, A), as shown in Table 1, thus significantly outperforming the baseline non-transitive model trained using C2M3 (which has a perplexity of 104.27). This trend extends to other combinations as well. For example, a model trained on (A, I) and (A, T) and evaluated on (I, T) has a perplexity of 3.13 (LoReTTa) compared to 9.70 (C2M3). Similarly, training on (I, T) and (I, A) and evaluating on (T, A) yields a perplexity of 11.21 (LoReTTa), equivalent to 30.88 (C2M3). Importantly, LoReTTa consistently achieves similar perplexity for combinations of unseen and seen modalities, demonstrating its ability to effectively learn the missing distribution. A similar trend emerged in downstream tasks (Table 2). In almost all cases, LoReTTa consistently outperformed C2M3 on the unimodal datasets A, I, and T, as well as the bimodal datasets (A, I), (A, T), or (I, T). In short, if any modality pair is missing, it is recommended to train the model using LoReTTa, as it uniquely integrates these modalities and improves accuracy compared to any unimodal model. Indeed, providing a small number of samples with all three modalities during linear probing further increases the classification score of models pre-trained using LoReTTa. Interestingly, despite good perplexity scores, the "upper bound" baseline trained on all aligned modalities generally lags behind LoReTTa. This strongly suggests that LoReTTa is capable of overcoming some form of negative transfer or modality competition.
[0131] Table 1. Perplexity on the SVL-MNIST test set. Subscripts describe the pre-trained dataset with images (I), text (T), or speech (A), while blue highlights unseen modality pairs.
[0132] Table 2. After pre-training, the classifier was trained on the frozen model. Classification accuracy was analyzed in a few-shot scenario with 100 samples per class in three trials of SVL-MNIST. Subscripts indicate the pre-training dataset with image (I), text (T), or speech (A) modalities, while bold values highlight results on modality tuples not seen during pre-training.
[0133] TCGA-OMICS The transformer is pre-trained using C2M3, and LoReTTa is initialized using the weights of the C2M3 model, as in the SVL-MNIST experiments. For GPT, the inventors simply used the C2M3 algorithm but disabled causal masking and exchange switching. BERT uses the same transformer architecture; however, the inventors simply changed the unidirectional attention mask to a bidirectional attention mask. To directly demonstrate the benefits of transitive modeling, the inventors extended both GPT and BERT using the transitive modeling techniques described herein. The resulting models are named T-GPT and T-BERT. These models use transitive modeling but not exchange modeling. The CLIP-based models consist of 3 encoders (each using our same transformer architecture) and a contrastive loss [Khosla et al., 2020] to align features. After pre-training, for each model, the inventors extracted embeddings and fitted a Cox proportional hazards model with a resilient net penalty for each modality combination using all available labels. In 6 out of 7 test sets, the LoReTTa pre-trained transformer achieved the highest c-index (Table 3). Importantly, this includes unseen modality pairs (miRNA, RPPA) and triplet states (mRNA, miRNA, RPPA). The only framework comparable to LoReTTa is T-BERT and CLIP on the unseen dataset (miRNA, RPPA). On the unseen trimodal dataset, T-GPT and CLIP yield good results but lag far behind LoReTTa. Overall, transfer modeling primarily improves the performance of all models. Notably, LoReTTa and CLIP are the only strategies that consistently allow positive or neutral transfer, unlike other frameworks. A major drawback of contrastive learning is the need for a separate encoder for each additional modality. Furthermore, CLIP-based methods only excel with large batch sizes. Here, the inventors compare models with the same computational constraints. To investigate how CLIP can be compared to LoReTTa with its additional computational power, they spent their limited resources on additional computational power, allowing them to scale up CLIP experiments by increasing the batch size by 5x and the number of training steps by 3x. This model, known as Large CLIP (L-CLIP), continuously improves the c-exponential of CLIP but requires 3x more VRAM than LoReTTa. As shown in Table 3, no model achieves the same performance as LoReTTa when performing trimodal predictions or two predictions in three bimodal predictions (the third type is only associated with T-BERT and CLIP at the same accuracy).When using unseen pairs as part of a bimodal input (M, R) or a trimodal input (M, I, R), all models fall far short of LoReTTa. This indicates that the LoReTTa method achieves higher performance compared to the comparative models when using the same amount of computational resources, and while increasing the allocation of computational resources to the comparative models does improve performance, it does not reach the same level as LoReTTa. In other words, the data demonstrate that the method described in this paper achieves multimodal prediction with higher accuracy compared to comparative methods that have the same computational efficiency when fully aligned training data is unavailable. In other words, the proposed method provides a more computationally efficient approach for learning from multimodal data in the presence of one or more missing tuples.
[0134] Compared to the SVL-MNIST experiment, the 3-modal “upper bound” model is less affected by negative transfer or modal competition and achieves optimal results in most modality combinations. However, as explained above, the availability of such fully aligned data is extremely rare in the medical data domain. Therefore, models such as the upper bound model are practically unattainable in most cases.
[0135] Table 3. Cox proportional hazards model trained on features extracted from TCGA-OMICS after pre-training. We used a mixture of mRNA (M), miRNA (I), and RPPA (R) as input. Bold highlights the c-index on unseen tuples, and underlined highlights the second-highest value.
[0136] MUGEN-GAME Recently, there has been a growing interest in cross-modal generation. While much research has focused on text-to-video or video-to-text conversion, other cross-modal tasks, such as video-to-audio, audio-to-video, text-to-audio, or audio-to-text, have not been thoroughly investigated. In this large-scale experiment, the inventors addressed the latter task. They used video as the linking modality and considered disjoint datasets (video, audio) and (video, text). Equally splitting the training set yielded approximately 187,500 pairs in each set. They trained their version of (MM)GPT and LoReTTa using the same optimized hyperparameters and model architecture as the state-of-the-art upper bound baseline MMGPT [Hayes et al., 2022]. As seen in the TCGA-OMICS experiment above, BERT unsurprisingly failed to generate long-range coherent samples (see results for mRNA). Therefore, the inventors did not use BERT in this experiment. They also did not train the CLIP model, as this would require training a text diffusion model (or equivalent generator) on top of a pre-trained encoder, which is neither end-to-end nor straightforward. For GPT, they trained the model to generate video from audio and text from video. In this way, they forced the model to implicitly learn how to subtitle audio tracks. In the case of LoReTTa, the inventors modeled the transition between audio and text directly through transitive modeling. Notably, GPT trained on two bimodal datasets outperformed the audio-to-text baseline in ROUGE, demonstrating strong positive transfer (Table 4). LoReTTa further improved GPT and even rivaled the video-to-text baseline in METEOR and ROUGE. Both models had low BLEU4 values. Since BLEU involves exact n-gram matching, these results are expected given the strict data constraints.
[0137] Table 4. Both GPT and LoReTTa are trained on disjoint (video, audio) and (video, text) pairs on MUGEN-GAME. The inventors then used BLEU4, METEOR, and ROUGE to evaluate the model for an unseen task of audio captioning. MMGPT stands for Upper Limit Model. BLEU4, METEOR, and ROUGE are widely used machine translation metrics for comparing the similarity between two texts at the word level (more precisely, n-gram). Higher values are better. While BLEU4 focuses on exact n-gram matching, METEOR considers synonyms, and ROUGE also considers sentence structure.
[0138] in conclusion Through LoReTTa, the inventors introduce an innovative self-supervised learning framework for multimodal ensembles that learns any missing joint conditional distribution given linked modalities. They demonstrate this theoretically and practically. Specifically, the new method combines commutativity and transitivity rules to model multimodal datasets (A,C) and (A,B,C) given two non-overlapping datasets (A,B) and (B,C). The inventors show that the features extracted by the transformer pre-trained using LoReTTa are highly expressive for all modality combinations, including unseen ones. While they evaluated the new algorithm only on datasets with three modalities, it is more general as it can be used with any number of modalities and any mixture of modalities. LoReTTa is able to learn transforms (X) as long as aligned modal chains exist in the dataset. i →...→X j ),...,(X i →...→X k ) for the missing distribution P(X) j ,...,X k (As shown in Figure 4d) is modeled. To the best of the inventors’ knowledge, no set of contrasting models has been trained in the literature without being based on the CLIP concept—most notably Wav2CLIP [Wu et al., 2022] and its extension ImageBind [Girdhar et al., 2023], as shown above, which are much more computationally intensive.
[0139] These examples demonstrate that LoReTTa can serve as a foundation for training very powerful multimodal generation and discrimination models that will have applications in numerous research areas. For instance, LoReTTa is designed to improve the quality of neural networks that contribute to society, for example, by training more accurate multimodal systems that can help doctors and other practitioners perform their daily tasks. LoReTTa is best suited for use in environments where some modality combinations are missing but at least one connectivity modality is present. The method can still be applied to datasets where some samples have all modality combinations. While the inventors have not analyzed this scenario in this example, LoReTTa likely offers benefits by amplifying those data points that still lack the remaining modalities.
[0140] Table 5. Batch size, number of warm-up steps, and total number of training steps for each model in the SVL-MNIST experiment.
[0141] Table 6. Batch size, number of warm-up steps, and total number of training steps for each model in the TCGA-OMICS experiment.
[0142] Table 7. For each dataset containing different modalities and combinations of modalities, find the number of training rounds for the linear classifier on the SVL-MNIST dataset.
[0143] References Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layernormalization. arXiv preprint arXiv:1607.06450, 2016.4 Peng Bo. Blinkdl / rwkv-lm: 0.01, August 2021. URL https: / / doi.org / 10.5281 / zenodo.5196577 Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022 Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. Flashattention: Fast and memory-efficient exact attention with io-awareness. In Advances in Neural Information Processing Systems (NeurIPS), Volume 35, pp. 16344–16359, 2022. Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré. Hyenahierarchy: Towards larger convolutional language models. arXiv preprint arXiv:2302.10866, 2023. [[ID=⑨]]Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners, 2019. Note: There seems to be a formatting issue with the tag [[ID=⑨]] in the original text which was likely a typo. I've translated it as [[ID=⑨]] for consistency with the original structure. If this was meant to be something else, please correct the original text.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In 31st Conference on Neural Information Processing Systems (NIPS), pp. 6000–6010. Curran Associates, Inc., 2017 Biao Zhang and Rico Sennrich. Root mean square layer normalization. In Advances in Neural Information Processing Systems (NeurIPS), vol. 32, 2019 Armen Aghajanyan, Bernie Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, et al. Cm3: A causal masked multimodal model of the internet. arXiv preprint arXiv:2201.07520, 2022 Armen Aghajanyan, Lili Yu, Alexis Conneau, Wei-Ning Hsu, Karen Hambardzumyan, Susan Zhang, Stephen Roller, Naman Goyal, Omer Levy, and Luke Zettlemoyer. Scaling laws for generative mixed-modal language models. arXiv preprint arXiv:2301.03728, 2023 Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In 2021 IEEE / CVF International Conference on Computer Vision (ICCV), pages 9650–9660. IEEE, 2021. doi: 10.1109 / ICCV48922.2021.00951. Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. In Advances in Neural Information Processing Systems (NeurIPS), volume 35, pages 23716–23736, 2022 Andrea Agostinelli, Timo I Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, et al. Musiclm: Generating music from text. arXiv preprint arXiv:2301.11325, 2023 Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pretraining from pixels. In International Conference on Machine Learning (ICML), pages 1691–1703. PMLR, 2020 Lluis Castrejon, Yusuf Aytar, Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba. Learning aligned cross-modal representations from weakly aligned data. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2940–2949, 2016 Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. An empirical analysis of compute-optimal large language model training. In Advances in Neural Information Processing Systems (NeurIPS), volume 35, pages 30016–30030, 2022 Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186. Association for Computational Linguistics, 2019 Thomas Hayes, Songyang Zhang, Xi Yin, Guan Pang, Sasha Sheng, Harry Yang, Songwei Ge, Qiyuan Hu, and Devi Parikh. Mugen: A playground for video-audio-text multimodal understanding and generation. In 17th European Conference on Computer Vision (ECCV), pages 431–449. Springer, 2022. Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training, 2018 Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), pp. 8748–8763. PMLR, 2021b Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. Imagebind: One embedding space to bind them all. arXiv preprint arXiv:2305.05665, 2023 John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021. Kudo and John Richardson. SentencePiece: A simple and language - independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (EMNLP), pages 66–71. Association for Computational Linguistics, 2018. Lyes Khacef, Laurent Rodriguez, and Benoit Miramond. Written and spoken digits database for multimodal learning. https: / / zenodo.org / record / 3515935, 2019 Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning. In 33th Advances in Neural Information Processing Systems (NeurIPS), pages 18661–18673. Curran Associates, Inc., 2020 Yann LeCun, Corinna Cortes, and Christopher J. C. Burges. Mnist handwritten digit database. http: / / yann.lecun.com / exdb / mnist / , 2010 Ruixi Lin and Hwee Tou Ng. Does bert know that the is-a relation is transitive In the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 94–99. Association for Computational Linguistics, 2022 Jessy Lin, Yuqing Du, Olivia Watkins, Danijar Hafner, Pieter Abbeel, Dan Klein, and Anca Dragan. Learning to model the world with language. arXiv preprint arXiv:2308.01399, 2023 Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Emergent world representations: Exploring a sequence model trained on a synthetic task. In the 11th International Conference on Learning Representations (ICLR), 2023 Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira. Perceiver: General perception with iterative attention. In the International Conference on Machine Learning (ICML), pp. 4651–4664. PMLR, 2021 Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, QiangLiu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. arXiv preprint arXiv:2208.10442, 2022 Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid,Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. Scaling autoregressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789, 2022. Hai Pham, Paul Pu Liang, Thomas Manzini, Louis-Philippe Morency, and Barnabás Póczos. Found in translation: Learning robust joint representations by cyclic translations between modalities. In the 33th AAAI Conference on Artificial Intelligence, pp. 6892–6899. AAAI, 2019. Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio GómezColmenarejo, Alexander Novikov, Gabriel Barth-maron, Mai Giménez, YurySulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, AliRazavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, OriolVinyals, Mahyar Bordbar, and Nando de Freitas. A generalist agent.Transactions on Machine Learning Research (TMLR), 2022. Mike Schuster and Kaisuke Nakajima. Japanese and korean voice search.In 2012 IEEE international conference on acoustics, speech and signalprocessing (ICASSP), pages 5149–5152. IEEE, 2012. Rico Sennrich, Barry Haddow, and Alexandra" Birch. Neural machinetranslation of rare words with subword units.In 54th Annual Meeting of theAssociation for Computational Linguistics (ACL), pages 1715–1725. Associationfor Computational Linguistics, 2016. Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning (ICML), pages 8821–8831. PMLR, 2021. Ali Razavi, Aaron Van den Oord, and Oriol Vinyals. Generating diverse high-fidelity images with vq-vae-2. In Advances in Neural Information Processing Systems (NeurIPS), volume 32, 2019 Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. In Advances in Neural Information Processing Systems (NeurIPS), volume 30, 2017. Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023. Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao, Hyung Won Chung, Iz Beltagy, Julien Launay, and Colin Raffel. What language model architecture and pretraining objective works best for zero-shot generalization In International Conference on Machine Learning (ICML), pages 22964–22984. PMLR, 2022 Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In IEEE International Conference on Computer Vision (ICCV), pages 2223–2232, 2017 Hai Pham, Paul Pu Liang, Thomas Manzini, Louis-Philippe Morency, and Barnabás Póczos. Found in translation: Learning robust joint representations by cyclic translations between modalities. In 33th AAAI Conference on Artificial Intelligence, pages 6892–6899. AAAI, 2019 Hitomi Yanaka, Koji Mineshima, and Kentaro Inui. Exploring transitivity in neural nli models through veridicality. In 16th Conference of the European Chapter of the Association for Computational Linguistics, pp. 920–934. Association for Computational Linguistics, 2021 Zack Thoutt. Wine reviews: 130k wine reviews with variety, location, winery, price, and description. https: / / www.kaggle.com / datasets / zynicide / wine-reviews, 2018. John N Weinstein, Eric A Collisson, Gordon B Mills, Kenna R Shaw, Brad A Ozenberger, Kyle Ellrott, Ilya Shmulevich, Chris Sander, and Joshua M Stuart. The cancer genome atlas pan-cancer analysis project. Nature Genetics, 45(10):1113–1120, September 2013. Alex Wang and Kyunghyun Cho. Bert has a mouth, and it must speak: Bert as a markov random field language model. arXiv preprint arXiv:1902.04094, 2019. Poli, M. et al. 2023b. StripedHyena: Moving Beyond Transformers withHybrid Signal Processing Models. github.com / togethercomputer / stripedhyena.doi: 10.57967 / hf / 1595. December 2023. Lieber et al. 2024. Jamba: A Hybrid Transformer-Mamba Language Model.arXiv:2403.19887. Thursday, March 28, 2024 De et al. 2024. Griffin: Mixing Gated Linear Recurrences with LocalAttention for Efficient Language Models. arXiv:2402.19427v1 [cs.LG] February 29, 2024 Tschannen et al. 2023. GIVT: Generative Infinite-VocabularyTransformers. arXiv:2312.02116. Monday, December 4, 2023. BehnamGhader et al. 2024. LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders. arXiv:2404.05961v1 [cs.CL] April 9, 2024 Yang et al. 2020. XLNet: Generalized Autoregressive Pretraining for Language Understanding. arXiv:1906.08237v2 [cs.CL] Thursday, January 2, 2020 Dong et al. 2019. Unified Language Model Pre-training for Natural Language Understanding and Generation. arXiv:1905.03197. Tuesday, October 15, 2019 Raffel et al. 2023. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer arXiv:1910.10683v4 [cs.LG] Tuesday, September 19, 2023 Tay et al. 2023.UL2: Unifying Language Learning Paradigms. arXiv:2205.05131v3 [cs.CL] Tuesday, February 28, 2023 Mikel Artetxe, Jingfei Du, Naman Goyal, Luke Zettlemoyer, and VesStoyanov. On the role of bidirectionality in language model pre-training. arXiv preprint arXiv:2205.11726, 2022. Akari Asai and Hannaneh Hajishirzi. Logic-guided data augmentation and regularization for consistent question answering. In 58th Annual Meeting of the Association for Computational Linguistics, pages 5642–5650. Association for Computational Linguistics, 2020. Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, and Mark Chen. Efficient training of language models to fill in the middle, 2022. Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, and Mike Lewis. Incoder: A generative model for code infilling and synthesis. Presented at the 11th International Conference on Learning Representations (ICLR), 2023. Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello. Wav2clip: Learning robust audio representations from clip. In 2022-2022IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP), pages 4563–4567. IEEE, 2022 Lili Yu, …, and Armen Aghajanyan. Scaling autoregressive multi-modalmodels: Pretraining and instruction tuning. Technical report, Meta AI, 2023. All references cited in this article are incorporated herein by reference in their entirety, and for all purposes, are equally indicated as to each individual publication or patent or patent application to be incorporated herein by reference in their entirety.
Claims
1. A computer-implemented method, comprising: Obtain an input dataset, which includes data from one or more modalities selected from a set of at least three modalities; as well as The output dataset is generated by processing the input dataset using a machine learning model that includes passing and exchanging machine learning models. The transfer and exchange machine learning model is a machine learning model that has been trained to transform an input dataset corresponding to any and every modality in the set of at least three modalities into an output dataset for any and every other modality in the set of at least three modalities. The machine learning model has been trained to transform data between two modalities using training data, which includes data pairs corresponding to a first modality and another modality in the two modalities, and other data pairs corresponding to the other modality and a second modality in the two modalities. Furthermore, the machine learning model has been trained to perform a corresponding inverse modality transformation for any forward modality transformation performed by the model.
2. The method according to claim 1, wherein the machine learning model is a classification model, a regression model, or a generative model, wherein: The classification model includes the transitive and exchange machine learning model configured to take the input dataset as input, and a trained classification model trained to take features extracted from the transitive and exchange machine learning model as input and produce a classification for the input dataset as output. The generative model is trained to take the input dataset as input and produce a generated output dataset corresponding to the input dataset as output, wherein the output dataset corresponds to a modality from the set of at least three modalities that is not present in the input dataset; or The regression model includes the transitive and exchange machine learning model configured to take the input dataset as input, and a trained regression model trained to take features extracted from the transitive and exchange machine learning model as input and produce regression values for the input dataset as output, optionally wherein the regression model is a Cox proportional hazards model and the regression values are survival predictions.
3. The method of claim 1 or claim 2, wherein the method includes determining a specific input modality of the input dataset and receiving an indication of a specific target output modality; and generating the output dataset is performed by processing the input dataset and identifying the specific target output modality using the pass-and-exchange machine learning model.
4. The method according to any preceding claim, wherein the transmission and exchange machine learning model has been trained using causal generative modeling, masking modeling, and / or variations thereof, optionally wherein the variation is causal masking modeling, wherein: The causal generation modeling, the masking modeling, and / or their variations use benchmark real data, which for each of multiple data pairs in the training data includes concatenated input data of the first and second modes of the data pair. The mode used as the first mode in the data pair is randomly sampled from the two modes of the data pair, and / or the benchmark real data includes a first concatenated input data and a second concatenated input data for each of the plurality of data pairs in the training data, wherein the mode used as the first mode in the data pair is different between the first concatenated input data and the second concatenated input data.
5. The method according to any preceding claim, wherein the transmission and exchange machine learning model has been trained using causal masking modeling, wherein the causal masking modeling includes: Identify a subset of lexical units in the training dataset, wherein the lexical units in the subset are located at multiple positions within the training dataset; Append the lexical subset to the training dataset; as well as Add masks at the plurality of locations in the training dataset.
6. The method of claim 5, wherein the causal mask modeling further comprises using the identification of the plurality of mask locations during the training.
7. The method according to any preceding claim, wherein the passing and exchanging machine learning model is a sequence model, and / or wherein the passing and exchanging machine learning model is a deep learning model, and / or wherein the passing and exchanging machine learning model comprises a sequence model selected from: recurrent neural networks, models combining recursive mechanisms and attention mechanisms, convolutional neural networks or convolution-based models, optionally models combining element-wise multiplication (gated) and long convolution or models combining convolution and attention mechanisms, state-space models or models combining state-space models and attention mechanisms or transformer-based models.
8. The method according to any of the preceding claims, wherein the transfer and exchange machine learning model includes a transformer model using self-attention.
9. The method according to any preceding claim, wherein obtaining the input dataset comprises: Receive the initial dataset; as well as The input dataset is generated by lexicalizing the initial dataset.
10. The method of claim 9, wherein generating the input dataset comprises one or more of the following: (i) Lexicalizing the initial dataset and projecting the lexicalized initial dataset into a predetermined feature space using a parameterized embedding function for each modality; (ii) appending modality-specific lexicals to the front and / or back of the lexicalized and optionally projected data corresponding to each individual modality in the input dataset; and (iii) including learned positional embeddings.
11. The method according to any preceding claim, wherein the transitive and exchange machine learning model has been trained using a first step and a second step, the first step using exchange modeling and a self-supervised learning objective, and in the second step, the machine learning model weights are initialized to the weights learned in the first step and further trained using transitive modeling.
12. The method of claim 11, wherein the first step comprises training the model to take partially masked concatenated input data of a first modality and a second modality of a data pair in a training dataset as input and producing a prediction of the masked input data as output, the training dataset comprising benchmark ground data of a plurality of data pairs, each data pair comprising corresponding data of a pair of modalities, wherein the modality of the data pair used as the first modality is randomly sampled from the two modalities of the data pair, and / or the benchmark ground data for each of the plurality of data pairs in the training data comprises first concatenated input data and second concatenated input data, wherein the modality of the data pair used as the first modality is different between the first concatenated input data and the second concatenated input data; and / or The second step includes training the model for each of a plurality of data pairs in the training dataset, each data pair including first data of a first modality and corresponding second data of a second modality in the set of at least three modalities: The machine learning model, with the second data as input, is used to predict the third data of the third modality in the set of at least three modalities corresponding to the second data; Optionally, the pass-and-exchange machine learning model, with the predicted third data or the predicted subsequent data as input, is used to iteratively predict subsequent data of another modality among the set of at least three modalities corresponding to the predicted third data or the predicted subsequent data. The passing and exchanging machine learning model, which takes the predicted third data or the subsequent data predicted in the latest iteration as input, is used to predict the final data of the first modality corresponding to the third data or the subsequent data predicted in the latest iteration. as well as Calculate the loss for comparing the first data with the predicted final data; The training data includes: (i) pairs of data for the second modality and corresponding data for the third modality; and (ii) pairs of data for the third modality and corresponding data for the first modality, or pairs of data for the latest additional modality and corresponding data for the first modality, and corresponding pairs of data for each of the two modalities used in the step of iteratively predicting subsequent data for the additional modality.
13. The method according to any preceding claim, wherein the transfer and exchange machine learning model has been trained using a method comprising: Access the training dataset, which includes first data for a specific input modality and corresponding second data for another modality; A first predicted output dataset is generated by processing the first data and identifying the other modality using the transmission and exchange machine learning model, wherein the first predicted output dataset corresponds to the other modality; A second predicted output dataset is generated by processing the first predicted output data and identifying the specific target output modality using the transmission and exchange machine learning model, wherein the second predicted output dataset corresponds to the specific target output modality; A third predicted output dataset is generated by processing the second predicted output data and identifying the specific input modality using the transmission and exchange machine learning model, wherein the third predicted output dataset corresponds to the specific input modality; as well as The loss is calculated by comparing at least a portion of the third predicted output dataset with the first data.
14. The method according to any preceding claim, wherein the transfer and exchange machine learning model has been trained using training data lacking data that associates first data of a particular input modality with second data of a particular output modality, and / or The aforementioned transfer and exchange machine learning model has been trained using training data that includes data that associates data pairs of modalities in the set of at least three modalities. For each pair of the first and second modalities in the set of at least three modalities, the training data includes data that associates data of the modal pair that together form a path between the first and second modalities. Optionally, for each pair of the first and second modalities in the set of at least three modalities, the training data includes data that associates data of the first modality with a third modality and data that associates data of the third modality with the second modality.
15. The method according to any preceding claim, wherein the transfer and exchange machine learning model has been trained to learn a first conditional distribution (P(C|A), P(A|C)) of a pair of modalities (A, C), wherein for the pair of modalities, the training data does not contain a sufficient pair of training data to correlate the data of the pair of modalities to parameterize the conditional distribution, the learning is performed by predicting data of one or more additional modalities (B) forming the link by using a learned further conditional distribution (P(C|B)) of the modal pair (C, B) forming the link between the pair of modalities, and for the modal pair, the training data contains a sufficient pair of training data to correlate the data of the pair of modalities to parameterize the further conditional distribution.
16. The method according to any preceding claim, wherein the transfer and exchange machine learning model has been trained using a method comprising, for each data pair comprising data pairs including first data of a first modality and corresponding second data of a second modality in the set of at least three modalities: The transmission and exchange machine learning model, with the second data as input, is used to predict the third data of the third modality in the set of at least three modalities corresponding to the second data; Optionally, the pass-and-exchange machine learning model, with the predicted third data or the predicted subsequent data as input, is used to iteratively predict subsequent data of another modality among the set of at least three modalities corresponding to the predicted third data or the predicted subsequent data. The passing and exchanging machine learning model, which takes the predicted third data or the subsequent data predicted in the latest iteration as input, is used to predict the final data of the first modality corresponding to the third data or the subsequent data predicted in the latest iteration. as well as Calculate the loss for comparing the first data with the predicted final data; The training data includes: (i) pairs of data for the second modality and corresponding data for the third modality; and (ii) pairs of data for the third modality and corresponding data for the first modality, or pairs of data for the latest additional modality and corresponding data for the first modality, and corresponding pairs of data for each of the two modalities used in the step of iteratively predicting subsequent data for the additional modality.
17. The method according to any preceding claim, wherein the transfer and exchange machine learning model has been trained using training data comprising tuples (e.g., pairs), the tuples comprising corresponding data for multiple modalities of the set of at least three modalities, wherein the number of tuples comprising data for a specific pair of modalities of the set of at least three modalities present in the training data is less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1%, and optionally wherein the tuples do not include tuples comprising data for a specific pair of modalities of the set of at least three modalities.
18. The method according to any of the preceding claims, wherein the method further comprises training the machine learning model.
19. The method according to any preceding claim, wherein the one or more modalities are selected from medical or biological sample data associated with a subject, optionally wherein the one or more modalities are selected from: medical image data, electronic medical record data, gene sequence data, and digital pathology image data, optionally wherein the medical image data are selected from MRI, CT, or PET imaging data.
20. The method according to any preceding claim, wherein the machine learning model is configured to take medical or biological sample data associated with a subject as input and produce an instruction for treatment recommended for the subject as output, optionally wherein the instruction is categorized and / or wherein the medical or biological sample data includes one or more of the following: as radiographic images, gene sequences, and microscope slides.
21. The method of any one of claims 1 to 18, wherein the machine learning model is configured to take data collected by any one or more types of sensors of a plurality of types associated with the vehicle, optionally an automobile, as input, and to generate predicted sensor data as output for sensor data of another type associated with the vehicle, optionally wherein the machine learning model has been trained using synchronized data from pairs of said plurality of sensors.
22. A computer-implemented method, comprising: Use the method according to any one of claims 1 to 20 to analyze input medical or biological sample data associated with the subject; The output of the machine learning model is used to provide the subject with a diagnosis and / or prognosis and / or treatment recommendations, wherein the output of the machine learning model indicates the diagnosis, the prognosis, or the treatment recommendations.
23. A computer-implemented method for training a multimodal machine learning model, the method comprising: Access a training dataset that includes data for each of at least three modalities; as well as The machine learning model is trained to transform an input dataset corresponding to any and every modality in the set of at least three modalities into an output dataset for any and every other modality in the set of at least three modalities, wherein the machine learning model is trained to transform data between two modalities using training data, the training data including data pairs corresponding to a first modality and another modality in the two modalities and other data pairs corresponding to the other modality and a second modality in the two modalities, and the machine learning model is trained to perform a corresponding inverse modality transformation for any positive type modality transformation performed by the model being trained to perform.
24. The method of claim 23, wherein training the machine learning model comprises using causal generation modeling, masking modeling, and / or causal masking modeling, wherein: The causal generation modeling, the masking modeling, and / or the causal masking modeling use benchmark real data from the training dataset, wherein the benchmark real data, for each of a plurality of data pairs in the training dataset, includes concatenated input data of a first mode and a second mode of the data pair. The mode used as the first mode in the data pair is randomly sampled from the two modes of the data pair, and / or the benchmark real data includes a first concatenated input data and a second concatenated input data for each of the plurality of data pairs in the training data, wherein the mode used as the first mode in the data pair is different between the first concatenated input data and the second concatenated input data.
25. The method of claim 24, wherein the training uses benchmark real data from the training dataset, the benchmark real data comprising, for each of a plurality of data pairs in the training data, cascaded input data of a first mode and a second mode of the data pair, wherein the mode of the data pair used as the first mode is randomly sampled from the two modes of the data pair with a predetermined probability, optionally with a 50% probability.
26. The method of any one of claims 23 to 25, wherein training the machine learning model comprises learning a first conditional distribution (P(C|A), P(A|C)) of a pair of modalities (A, C), wherein for the pair of modalities, the training data does not contain a sufficient pair of training data to correlate the data of the pair of modalities to parameterize the conditional distribution, the learning is performed by using the machine learning model to predict data of one or more additional modalities (B) forming a link between the pair of modalities, the prediction using a learned further conditional distribution (P(C|B)) of the modal pair (C, B) forming the link, and for the modal pair, the training data contains a sufficient pair of training data to correlate the data of the pair of modalities to parameterize the further conditional distribution.
27. The method of any one of claims 23 to 26, wherein training the machine learning model comprises, for each data pair comprising data pairs of training data, the data pair comprising first data of a first modality and corresponding second data of a second modality in the set of at least three modalities: The transmission and exchange machine learning model, with the second data as input, is used to predict the third data of the third modality in the set of at least three modalities corresponding to the second data; Optionally, the pass-and-exchange machine learning model, with the predicted third data or the predicted subsequent data as input, is used to iteratively predict subsequent data of another modality among the set of at least three modalities corresponding to the predicted third data or the predicted subsequent data. The passing and exchanging machine learning model, which takes the predicted third data or the subsequent data predicted in the latest iteration as input, is used to predict the final data of the first modality corresponding to the third data or the subsequent data predicted in the latest iteration. as well as Calculate the loss for comparing the first data with the predicted final data; The training data includes: (i) pairs of data for the second modality and corresponding data for the third modality; and (ii) pairs of data for the third modality and corresponding data for the first modality, or pairs of data for the latest additional modality and corresponding data for the first modality, and corresponding pairs of data for each of the two modalities used in the step of iteratively predicting subsequent data for the additional modality.
28. The method of any one of claims 23 to 27, wherein training the machine learning model comprises modeling using a causal mask so that: Identify a subset of lexical units in the training dataset, wherein the lexical units in the subset are located at multiple positions within the training dataset; Append the lexical subset to the training dataset; as well as Add masks at the plurality of locations in the training dataset.
29. The method according to any one of claims 23 to 28, wherein training the machine learning model comprises: Access a training dataset that includes first data for a specific input modality among the at least three modalities and corresponding second data for another modality among the at least three modalities; A first predicted output dataset is generated by processing the first data and identifying the other modality using the transmission and exchange machine learning model, wherein the first predicted output dataset corresponds to the other modality; A second predicted output dataset is generated by processing the first predicted output data and identifying the specific target output modality using the transitive and exchange machine learning model, wherein the second predicted output dataset corresponds to the specific target output modality; a third predicted output dataset is generated by processing the second predicted output data and identifying the specific input modality using the transitive and exchange machine learning model, wherein the third predicted output dataset corresponds to the specific input modality; as well as The loss is calculated by comparing at least a portion of the third predicted output dataset with the first data.
30. The method of any one of claims 23 to 29, further comprising training an additional machine learning model including a trained machine learning model, wherein said additional machine learning model is trained for a specific task and is a classification model, a regression model, or a generative model. Choose one of them: The classification model includes the trained machine learning model, and a classification model trained to take features extracted from the pass-and-exchange machine learning model as input and produce a classification for the input dataset provided to the trained machine learning model as output, optionally wherein the classification model is trained using training data including data of one or more of the set of at least three modalities and associated benchmark true classification labels; The generative model is trained to take an input dataset as input and generate a corresponding output dataset as output, wherein the output dataset corresponds to a modality from the set of at least three modalities that is not present in the input dataset, optionally wherein the trained machine learning model is used as the generative model or the trained machine learning model is fine-tuned using further training data to obtain the generative model; or The generative model is trained to take an input dataset including missing data as input and produce a generated output dataset corresponding to the missing data in the input dataset as output, optionally wherein the trained machine learning model is used as the generative model or the trained machine learning model is fine-tuned using further training data to obtain the generative model; or The regression model includes the trained machine learning model, and a regression model trained to take features extracted from the trained machine learning model for an input dataset as input and produce regression values for the input dataset as output, optionally wherein the regression model is a Cox proportional hazards model and the regression values are survival predictions, and / or wherein the regression model is trained using training data including data from one or more of the set of at least three modalities and associated baseline true regression values.
31. A system comprising: A computer-readable storage medium for one or more processors and storing instructions that cause the processor to perform the method according to any one of claims 1 to 30.
32. A non-transitory computer-readable storage medium comprising machine-executable instructions that, when executed on a processor, cause the processor to perform the method according to any one of claims 1 to 30.