Method and device for pre-training large-scale foundation model integrating multimodal omics data
The sparse MoE and Hyena module combination addresses the challenge of handling DNA, RNA, and protein sequences together, enhancing biological data analysis by optimizing parameter efficiency and long-range contextual processing.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE HONG KONG UNIV OF SCI & TECH
- Filing Date
- 2025-11-03
- Publication Date
- 2026-05-15
AI Technical Summary
Existing large language models primarily handle DNA, RNA, and protein sequences in isolation, overlooking intricate interconnections among small molecules, leading to challenges like catastrophic forgetting during multimodal mixed training.
A computer-implemented method using a sparse Mixture-of-Experts (MoE) technique combined with Hyena modules for pre-training a large-scale foundation model, involving tokenization, masking, reconstruction, and hybrid neural-network architectures to learn cross-modal relationships and optimize parameters for efficient processing.
Enhances the extraction of complex patterns and relationships in biological data, improving analysis performance and scalability while reducing computational costs by selectively activating experts and integrating long-range contextual processing.
Smart Images

Figure CN2025132070_15052026_PF_FP_ABST
Abstract
Description
METHOD AND DEVICE FOR PRE-TRAINING LARGE-SCALE FOUNDATION MODEL INTEGRATING MULTIMODAL OMICS DATACROSS REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims priority to, and the benefit of, Provisional Application No. 63 / 716,710 filed in the U.S. Patent and Trademark Office on November 5, 2024, the entire contents of which are incorporated herein by reference.FIELD OF INVENTION
[0002] The present disclosure relates generally to a large-scale foundation model for multimodal omics, and more particularly to computer-implemented methods and devices for pre-training a large-scale foundation model.BACKGROUND
[0003] Rapid advancement of large-scale models has led to significant breakthroughs in biology, particularly in the realm of small molecules associated with genomics. Existing large language models effectively handle the modalities of deoxyribonucleic acid (DNA) sequences, ribonucleic acid (RNA) sequences, and protein sequences. However, these models typically address these modalities in isolation, overlooking the intricate interconnections among small molecules. To exploit these connections, some studies have employed mixed training strategies to learn more discriminative features through complementary information. While these efforts have advanced foundational models in genomics, challenges such as catastrophic forgetting persist due to multimodal mixed training.
[0004] Accordingly, there is a need for a solution that addresses the above and / or other problems and / or provides a useful choice.SUMMARY
[0005] The present disclosure relates to a computer-implemented method and a computing device for large-scale foundation model pre-training of multimodal omics using sparse Mixture-of-Experts (MoE) technique, or a combination of sparse MoE and Hyena techniques.
[0006] In a first aspect of the present disclosure, a computer-implemented method for pre-training a large-scale foundation model is provided. The method comprises: tokenizing datasets to obtain tokenized datasets, wherein the datasets include different biological modalities originating from DNA sequences, RNA sequences, and protein sequences; randomly masking a subset of tokens in each tokenized dataset to generate masked datasets; reconstructing at least one masked token associated with one of the biological modalities based on contextual embeddings derived from unmasked tokens associated with another one of the biological modalities to obtain a pre-trained large-scale foundation model comprising optimized tokenization and reconstruction parameters; pre-training, by the large-scale foundation model comprising a sparse Mixture-of-Experts (MoE) module, input token representations derived from the tokenized datasets to obtain the pre-trained large-scale foundation model further comprising optimized routing and expert parameters.
[0007] In an example, the pre-training step includes pre-training, by the large-scale foundation model comprising at least one Hyena module and the sparse MoE module, by operatively coupling the at least one Hyena module and the sparse module to perform complementary long-range contextual processing and expert-specialized processing, such that outputs generated by one of the at least one Hyena module or the sparse MoE module, are utilized by an other module to generate globally mixed contextual embeddings to obtain the pre-trained large-scale foundation model further comprising optimized Hyena-operator parameters, wherein the other module is a remaining one of the at least one Hyena module and the sparse MoE module or a non-Hyena non-MoE module.
[0008] In an example, the at least one Hyena module includes a first Hyena module and a second Hyena module such that an output of the first Hyena layer forms an input to the sparse MoE layer, and an output of the sparse MoE layer forms an input to the second Hyena layer, wherein the sparse MoE layer comprises a routing network and at least two expert subnetworks connected thereto, each expert subnetwork having a plurality of experts, wherein the pre-training step includes: performing, by the first Hyena module, long-range convolutional context mixing on the input token representations to produce first Hyena-mixed representations; for each of the first Hyena-mixed representations, determining, by the routing network of the MoE, routing logits corresponding to the experts and activating, based on the routing logits, a subset of the experts; processing, by the activated experts of the MoE, the corresponding input token representations to generate expert outputs; combining the expert outputs to obtain a combined MoE output representation; and performing, by the second Hyena module, long-range convolutional remixing on the combined MoE output representations to generate the globally mixed contextual embeddings, such that the pre-trained large-scale foundation model further comprises the optimized Hyena-operator parameters.
[0009] In an example, the determining step comprises: applying an auxiliary loss term to encourage a substantially uniform distribution of the first Hyena-mixed representations among the activated experts.
[0010] In an example, the at least one hybrid neural-network architecture comprises a plurality of hybrid neural-network architectures being connected in sequence such that an output of a preceding hybrid neural-network architecture forms an input to a subsequent hybrid neural-network architecture.
[0011] In an example, the sparse MoE layer comprises a routing network and at least two expert subnetworks connected thereto, each expert subnetwork having a plurality of experts, wherein the pre-training step comprises: for each input token representation, determining, by the routing network, routing logits corresponding to the experts and activating, based on the routing logits, a subset of the experts; processing, by the activated experts, the corresponding input token representations to generate expert outputs; and combining the expert outputs to obtain a combined MoE output representation.
[0012] In an example, the determining step comprises: applying an auxiliary loss term to encourage a substantially uniform distribution of the input token representations among the activated experts.
[0013] In a second aspect of the present disclosure, a computer-implemented method executed by one or more processors for performing an inference operation is provided. The method comprises: receiving, by the one or more processors, input data instances respectively corresponding to different biological modalities originating from DNA sequences, RNA sequences, and protein sequences; encoding, by a pre-trained large-scale foundation model stored in one or more memories, the input data instances to generate modality-specific embeddings; projecting, by the pre-trained large-scale foundation model, the modality-specific embeddings into a joint multi-modal embedding which characterizes inter-relationships among the different biological modalities; and performing, by an inference module of the pre-trained large-scale foundation model, the inference operation comprising classification, regression, and / or prediction of a target biological feature based on the joint multi-modal embedding, wherein the target biological feature comprises a predicted attribute, state, or outcome associated with at least one of the different biological modalities, and wherein the pre-trained large-scale foundation model has been pre-trained by the method of any one of the foregoing examples.
[0014] In a third aspect of the present disclosure, a computing device comprises one or more memories having executable code; and one or more processors coupled to the one or more memories, and configured to execute the executable code to cause the computing device to perform the method of any one of the foregoing examples.
[0015] In a fourth aspect of the present disclosure, a non-transitory computer-readable medium comprises executable code, which when executed by a processor of a computing device, cause the computing device to perform the method of any one of the foregoing examples.BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying figures, where like reference numerals refer to identical or functionally similar elements throughout the separate views and which together with the detailed description below are incorporated in and form part of the specification, serve to illustrate various aspects and to explain various principles and advantages in accordance with the present disclosure.
[0017] FIG. 1 is a flowchart illustrating a method for training a large-scale foundation model, in accordance with aspects of the present disclosure.
[0018] FIG. 2A is a schematic diagram of an exemplary framework for pre-training a large-scale foundation model, in accordance with aspects of the present disclosure.
[0019] FIG. 2B is a schematic diagram of a non-limiting MoE framework which may be utilized in FIG. 2A.
[0020] FIG. 3 is a flowchart illustrating a method for training a large-scale foundation model, in accordance with aspects of the present disclosure.
[0021] FIG. 4A is a schematic diagram of another exemplary framework for pre-training a large-scale foundation model, in accordance with aspects of the present disclosure.
[0022] FIG. 4B is a schematic diagram of a non-limiting example of MoE-Hyena framework which may be utilized in FIG. 4A.
[0023] FIG. 5 is a flowchart illustrating a method for performing an inference operation, in accordance with aspects of the present disclosure
[0024] FIG. 6A shows an exemplary description of datasets used for downstream validation tasks.
[0025] FIG. 6B shows experimental results of conventional specialist or generalist omics foundation model and a foundation model of the present disclosure in relation to validation performance.
[0026] FIG. 6C shows a comparison of average performances of different model variants on three DNA benchmarks.
[0027] FIG. 7 is a schematic diagram of an exemplary computing device that may be used for implementing the methods 100, 300 of FIGs. 1 and 3 and / or framework of FIGs. 2A, 2B, 4A and 4B.DETAILED DESCRIPTION
[0028] Aspects according to the present disclosure will be described, by way of example only, with reference to the drawings. Like reference numerals and characters in the drawings refer to like elements or equivalents.
[0029] Unless specifically stated otherwise, and as apparent from the following, it will be appreciated that throughout the present specification, discussions utilizing terms such as “tokenizing” , “masking” , “reconstructing” , “pre-training” , “performing” , “determining” , “processing” , “combing” , “generating” , “predicting” , “computing” , “applying” , “encoding” , “projecting” , or the like, refer to the action and processes of a computer system, or similar electronic device, that manipulates and transforms data represented as physical quantities within the computer system into other data similarly represented as physical quantities within the computer system or other information storage, transmission or display devices.
[0030] As used herein, the term “large-scale foundation model” refers to a neural-network model trained on extensive multimodal or domain-specific datasets and configured to learn generalizable representations applicable to a wide range of downstream tasks. The term “module” refers to a functional component of a computing or neural-network system configured to perform one or more operations. A module may be implemented in software, firmware, hardware, or any combination thereof, and may include one or more layers, operators, or submodules that collectively accomplish the specified function. The term “hybrid neural-network architecture” refers to a neural network configuration that integrates two or more heterogeneous computational mechanisms within a single functional module to perform complementary learning and / or inference operations. The term “layer” refers to a structural unit within a neural-network architecture configured to perform one or more mathematical transformations on input data. Non-limiting examples may be convolutional, feed-forward, attention, Hyena, or mixture-of-experts components. A layer may output intermediate representations for use by subsequent layers or modules. The term “neural-network architecture” refers to an arrangement or sequence of interconnected layers or modules defining the flow of data and the learning process within a neural network. The architecture specifies how inputs are transformed through parameterized operations to produce outputs and may include recurrent, convolutional, transformer-based, Hyena-based, or hybrid configurations.
[0031] FIG. 1 is a flowchart illustrating a method 100 for training a large-scale foundation model, in accordance with aspects of the present disclosure.
[0032] In block 102, multimodal omics datasets are provided or received. In particular, the method 100 may comprise receiving or providing a plurality of datasets having different biological modalities originating from DNA sequences, RNA sequences, and protein sequences. The datasets may be pre-training datasets described in He, Y., Fang, P., Shan, Y., et al. (2024) . LucaOne: Generalized Biological Foundation Model with Unified Nucleic Acid and Protein Language. bioRxiv, 2024.05.10.592927 (hereinafter LucaOne) which comprise DNA, mRNA, and protein sequences. The nuclei acid data was collected from the RefSeq (Reference Sequence database which is an open access, annotated and curated collection of publicly available nucleotide sequences (DNA, RNA) and their protein products) . Molecular types of entries of this database include DNA, RNA, DNA sequences, RNA sequences, DNA annotations, and RNA annotations. Compared to nucleic acids, proteins contain more extensive information, including sequences, classifications, sites, tertiary structures, homologous regions, domains, and keywords. The sequences, classifications, and keywords are sourced from UniProt (adatabase of protein sequence and functional information) and ColabFold (aprotein prediction algorithm) , while sites, domains, and homologous regions are obtained from InterPro (adatabase of protein families) . Tertiary structures are derived from RCSB Protein Data Bank (RCSB-PDB) and AlphaFold2-Swiss-Prot.
[0033] In block 104, the datasets of block 102 are converted into tokenized datasets to thereby represent biological information of the corresponding modality in a machine-readable, discrete symbolic form. In particular, the method 100 may comprise tokenizing the datasets of block 102 to obtain a plurality of tokenized datasets. Tokenization may be performed by a tokenization module which receives the datasets of block 102. Non-limiting examples of a tokenizer module include Byte Pair Encoding (BPE) tokenizer, WordPiece tokenizer, etc.
[0034] In block 106, the tokenized datasets across different modalities are randomly masked and the resulting masked datasets are reconstructed based on cross-modal contextual embeddings. In particular, the method 100 may comprise randomly masking a subset of tokens in each tokenized dataset to generate a plurality of masked datasets. Each masked dataset includes masked token (s) and unmasked token (s) . Masking may include replacing selected tokens with a special mask symbol, a random substitute token, or leaving a fraction of tokens unchanged. This may be performed by a masking module of the foundation model. The masking module of the foundation model may be implemented to perform masking by replacing selected tokens with a designated mask symbol, substituting them with random tokens, or retaining a fraction of the original tokens unchanged, depending on the masking strategy adopted. Masking may be performed on datasets of different biological modalities such that the masked tokens subsist across some or all of the biological modalities. The masking module may include a multimodal encoder which receives and processes the masked datasets to generate contextual embeddings for both masked and unmasked tokens across the different biological modalities. Contextual embeddings refer to learned numerical representations of tokens that encode information about surrounding context and cross-modality relationships.
[0035] In block 108, cross-modal reconstruction is performed to enable the large-scale foundation model learn correlations among different biological modalities consistent with biological relationships. In particular, the method 100 may further comprise reconstructing or predicting at least one masked token associated with one of the biological modalities, e.g., DNA token, based on contextual embeddings derived from unmasked tokens of another one (different one or more) of the biological modalities, e.g. RNA token and / or protein token. This reconstruction produces reconstructed tokens and may be performed by a reconstruction module of the foundation model. The reconstruction module of the foundation model may be implemented as a neural decoding component that restores masked tokens based on contextual embeddings generated by the encoder. Non-limiting examples of the reconstruction module include a Transformer decoder, an attention-based prediction head, or a feedforward projection layer mapping hidden states to token probabilities.
[0036] The reconstruction step is iterated and a reconstruction loss is optimised during the iterations. By iterative optimization of a reconstruction loss between reconstructed tokens and corresponding ground-truth tokens, parameters associated with the tokenization embeddings and reconstruction layers are progressively updated and thus optimized. Accordingly, the method 100 results in the pre-trained large-scale foundation model comprising optimized tokenization and reconstruction parameters, which enable the model to capture structural and contextual relationships among biological modalities for subsequent fine-tuning or inference tasks.
[0037] In block 110, the large-scale foundation model comprising a sparse Mixture-of-Experts (MoE) module is pre-trained to obtain a pre-trained large-scale foundation model further comprising optimized routing and expert parameters. The MoE module may comprise a routing network and a plurality of expert subnetworks connected thereto. Each expert subnetwork comprises a plurality of experts. In particular, the method 100 may comprise: receiving, by a MoE module, a plurality of input token representations (alternatively referred to as MoE input representations) which are derived from the tokenized datasets. The method 100 may comprise, for each input token representation, determining or computing, by the routing network, a plurality of routing logits corresponding to the experts. Each routing logit may be a numeric score which indicates a degree of correspondence between an input token representation and a particular expert. The method 100 may comprise activating, based on the routing logits, a subset of the experts or selected experts. Only selected experts receive the corresponding input representations, thereby achieving sparse expert activation while maintaining a large overall parameter capacity. The method 100 may comprise processing, by each activated expert, its assigned input representation through its internal feed-forward layers to produce its expert output. The method 100 may comprise combining the expert outputs from the selected subset of experts to obtain a combined MoE output representation. The combining may be based on routing weights associated with the activated experts and computed by the routing network. The combined MoE output representation may be passed to subsequent neural network layers.
[0038] By repeating forward-propagation and backward-propagation updates based on a prediction loss, the parameters of both the routing network and the expert subnetworks are progressively updated and thus optimized. Accordingly, the method 100 results in the pre-trained large-scale foundation model further comprising optimized routing and expert parameters, which enable efficient token routing and specialized feature transformation during subsequent fine-tuning or inference tasks.
[0039] It is to be appreciated that some or all of the following variations to the method 100 of FIG. 1 may be envisaged in other embodiments. For example, the expert outputs may be normalized prior to combining. For example, determining the routing logits may comprise applying an auxiliary loss term to encourage a substantially uniform distribution of the MoE inputs among the activated experts. This improves training efficiency as auxiliary loss promotes equal importance among all experts to ensure that each expert receives an approximately equal count of training examples or counts of training examples within a predetermined threshold.
[0040] FIG. 2A is a schematic diagram of an exemplary frameworkfor pre-training a large-scale foundation model using a MoE encoder (alternatively referred to as MoE module) .
[0041] FIG. 2B shows a non-limiting example of a MoE module having a routing network and two sets of expert subnetworks. The MoE module may comprise a self-attention layer which may be connected to a first normalization layer which may be connected to the routing network. Non-limiting examples of suitable implementation architectures include a Transformer-style attention block with LayerNorm followed by a gating network for expert selection, or a lightweight attention module combined with normalization and softmax-based routing as implemented in standard MoE frameworks. The expert subnetworks may be connected to a second normalization layer.It is to be appreciated that variations to FIG. 2B may be envisaged. For example, the MoE module may include one feed forward neural network or expert subnetwork, or other counts of feed forward neural network, e.g. three or more. For example, one or both of the normalization layer (s) may be omitted.
[0042] FIG. 3 is a flowchart illustrating a method 300 for training a large-scale foundation model, in accordance with aspects of the present disclosure.
[0043] Blocks 302, 304, 306, and 308 of FIG. 3 respectively correspond to blocks 102, 104, 106, and 108 of FIG. 1 and hence the corresponding description of blocks 102 to 108 applies to block 302 to 308.
[0044] Block 310 of FIG. 3 differs from block 110 of FIG. 1 in that, according to block 310, the large-scale foundation model comprises at least one hybrid or composite neural-network architecture, each hybrid neural-network architecture comprising at least one Hyena module and at least one sparse MoE module. Each Hyena module may comprise one or more Hyena layers, each implementing a Hyena operator or equivalent long-range filtering and gating function.
[0045] In the hybrid neural-network architecture, the at least one Hyena module and the at least one MoE module are operatively coupled to perform complementary long-range contextual processing and expert-specialized processing, respectively, such that outputs generated by one module (one of the Hyena module or MoE module) are utilized by another module (aremaining one of the Hyena module or MoE module, or a non-Hyena or non-MoE module) to produce globally mixed contextual embeddings. For example, outputs by a Hyena module may be utilized by a MoE module, or outputs by a MoE module may be utilized by a Hyena module, and / or outputs from a Hyena module and a MoE module may be combined by another module which is a non-Hyena non-MoE module) . Accordingly, the hybrid neural-network architecture facilitates both efficient global-context mixing and parameter-efficient expert specialization within the large-scale foundation model to obtain a pre-trained large-scale foundation model further comprising optimized Hyena-operator parameters, in addition to the optimized tokenization and reconstruction parameters, and the optimized routing and expert parameters.
[0046] The Hyena module (s) and the MoE module (s) may be independent and modular components arranged in any sequence or configuration depending on model configuration. Non-limiting examples include serial (sequential) , and interleaved configurations. In a serial configuration example (FIG. 4B) , an output of a first Hyena module provides an input to a MoE module and an output of the MoE module provides an input to a second Hyena module. Other configurations and other configuration examples may be envisaged and implemented in the hybrid neural-network architecture.
[0047] Accordingly, in block 310, the method 300 may comprise pre-training, by the large-scale foundation model comprising at least one Hyena module and a sparse MoE module, input token representations derived from the tokenized datasets to obtain the pre-trained large-scale foundation model comprising optimized routing and expert parameters, and optimized Hyena-operator parameters. This pre-training step includes or is performed by operatively coupling the at least one Hyena module and the sparse MoE module to perform complementary long-range contextual processing and expert-specialized processing, such that outputs generated by one of the at least one Hyena module or the sparse MoE module, are utilized by an other module to generate globally mixed contextual embeddings. This other module refers to a remaining one of the at least one Hyena module and the sparse MoE module (e.g. other than the particular module which generated the outputs) or a non-Hyena non-MoE module.
[0048] A non-limiting example in which the hybrid neural-network architecture combines Hyena modules and a MoE module in a serial configuration example (see FIG. 4B) is illustrated as follows.
[0049] The method 300 may, in block 310, comprise receiving, by the first Hyena module, a plurality of input token representations derived from the tokenized datasets. The method 300 may comprise performing, by the first Hyena module, long-range convolutional context mixing on the input token representations to produce a plurality of first Hyena-mixed representations. Specifically, the first Hyena module applies parameterized long-range convolutional filters, which may be implicit or learned kernels defined as functions of positional or frequency information. The first Hyena module may further comprise a gating function that performs element-wise modulation between filtered and unfiltered representations to yield first Hyena-mixed representations that capture both local and long-range contextual dependencies across extended token sequences.
[0050] The MoE module may comprise a routing network and a plurality of expert subnetworks connected thereto. Each expert subnetwork comprises a plurality of experts. The sparse MoE module of block 310 may be identical or different from the sparse MoE module of block 110. The method 300 may comprise: receiving, by a MoE module, the first Hyena-mixed representations (alternatively referred to as MoE input representations) . The method 100 may comprise, for each first Hyena-mixed representation, determining or computing, by the routing network, a plurality of routing logits corresponding to the experts. Each routing logit may be a numeric score which indicates a degree of correspondence between a first Hyena-mixed representation and a particular expert. The method 300 may comprise activating, based on the routing logits, a subset of the experts or selected experts. Only the selected experts receive the corresponding first Hyena-mixed representations, thereby achieving sparse expert activation while maintaining a large overall parameter capacity. Each activated expert processes its assigned first Hyena-mixed representation through its internal feed-forward layers to produce its expert output. The method 300 may comprise combining the expert outputs from the selected subset of experts to obtain a combined MoE output representation. The combining may be based on routing weights associated with the activated experts and computed by the routing network. The combined MoE output representation may be passed to subsequent neural network layers.
[0051] By repeating forward-and backward-propagation updates based on a prediction loss, the parameters of both the routing network and the expert subnetworks are progressively updated and thus optimized. Accordingly, the method 300 results in the pre-trained large-scale foundation model further comprising optimized routing and expert parameters, which enable efficient token routing and specialized feature transformation during subsequent inference.
[0052] The method 300 may comprise receiving, by the second Hyena module, a combined MoE output representation. The method 300 may comprise performing, by the second Hyena module, long-range convolutional context mixing on the MoE output representations to produce second Hyena-mixed representations which are globally mixed contextual embeddings. Specifically, the second Hyena module applies parameterized long-range convolutional filters to re-integrate the MoE output representation into a unified representation. The second Hyena module may include a gating function to modulate the recombined representations, ensuring that both the global context and expert-specific details are retained in the final embedding space. The resulting globally mixed contextual embeddings may be utilised for reconstructing masked tokens or for downstream inference tasks.
[0053] With block 310, the entire Hyena-MoE-Hyena sequence is optimized end-to-end using a reconstruction or prediction loss function that compares predicted token identities with corresponding ground-truth tokens. Through repeated forward-and backward-propagation updates, the parameters of the Hyena operators (including convolutional filters and gating weights) , the routing network (including routing logits and normalization weights) , and the expert subnetworks (including feedforward transformation weights) are progressively updated and thus optimized. Accordingly, the pre-training process of block 310 produces a pre-trained large-scale foundation model comprising optimized Hyena-operator parameters, routing-network parameters, and expert-subnetwork parameters. The optimized Hyena-operator parameters enable efficient long-range contextual mixing and global information integration, thereby improving representational quality, training stability, and inference efficiency compared with transformer-based attention mechanisms.
[0054] FIG. 4A is a schematic diagram of an exemplary frameworkfor pre-training a large-scale foundation model using a MoE-Hyena encoder (alternatively referred to as MoE-Hyena module) .
[0055] FIG. 4B shows a non-limiting example of a large-scale foundation model having three hybrid neural-network architecture. Each hybrid neural-network architecture comprises a first Hyena module, a sparse MoE module, and a second Hyena module in a serial configuration. The hybrid neural-network architectures are connected in sequence such that an output of a preceding hybrid neural-network architecture forms an input to a subsequent hybrid neural-network architecture.
[0056] It is to be appreciated that variations to FIG. 4B may be envisaged in other embodiments. For example, the large-scale foundation model may include one hybrid neural-network architecture, or other counts of hybrid neural-network architectures, e.g. two, four, or more. For example, while FIG. 4 shows a count of MoE modules (transformer modules) is lower than a count of Hyena modules in each hybrid neural-network architecture and in the overall model, other embodiments may include MoE modules and Hyena modules with equal count, or a count of MoE modules (transformer modules) being higher than a count of Hyena modules.
[0057] After pre-training process as described in the foregoing, the resulting pre-trained large-scale foundation model may be deployed to perform inference on new biological data.
[0058] FIG. 5 is a flowchart illustrating a method 500 for performing an inference operation, in accordance with aspects of the present disclosure.
[0059] In block 502, the method 500 may comprise receiving input data instances respectively corresponding to different biological modalities originating from DNA sequences, RNA sequences, and protein sequences. The receiving step may be performed by one or more processors of one or more computing devices.
[0060] In block 504, the method 600 may comprise encoding the input data instances to generate modality-specific embeddings. The encoding step may be performed by a pre-trained large-scale foundation model stored in one or more memories of the one or more computing devices. The encoding step may include tokenizing or converting the input instances into machine-interpretable representations and encoding the representations to produce modality-specific embeddings that capture intra-modality sequence patterns and structural features.
[0061] In block 506, the method 500 may comprise projecting the modality-specific embeddings into a joint multi-modal embedding which characterizes inter-relationships among the different biological modalities. This projecting step may be performed by the pre-trained large-scale foundation model. This projection may be performed through learned projection matrices, cross-modal attention mechanisms, or hybrid modules such as Hyena-MoE components that align representations across modalities.
[0062] In block 508, the method 500 may comprise performing the inference operation comprising classification, regression, and / or prediction of a target biological feature based on the joint multi-modal embedding. The target biological feature may comprise a predicted attribute, state, or outcome associated with at least one of the different biological modalities. The pre-trained large-scale foundation model has been pre-trained by the method 100, 300 as described above. This inference operation step may be performed by an inference module of the pre-trained large-scale foundation model. In some examples, the inference module may comprise one or more fully connected layers, attention heads, or expert subnetworks configured to map the joint embedding to a target output space.
[0063] The present disclosure is particularly advantageous at least in the following ways.
[0064] Utilizing a hybrid neural-network architecture (FIGs. 3 and 4) which combines MoE technique and Hyena technology to implement a self-supervised biological language model with concurrent training of different modalities significantly enhances overall analysis performance in the omics field. By effectively utilizing a small number of Transformer modules alongside a large number of Hyena modules, the Hyena-MoE hybrid neural-network architecture addresses scaling issues related to larger multimodal datasets and pre-training data while improving feature extraction and analysis capabilities for long sequence tasks, all while reducing computational costs. Firstly, the Hyena-MoE hybrid neural-network architecture improves scaling issues by selectively activating only a subset of experts for each input, which reduces the overall computational burden. This approach allows for parallel processing across different nodes, enabling simultaneous handling of multiple data batches and increasing throughput. Secondly, MoE technique dynamically allocates resources based on input complexity, directing more power to challenging instances while requiring less for simpler ones. As the training data size increases, the Hyena-MoE hybrid neural-network architecture can scale by adding more experts without significantly enlarging the overall model size, making it highly efficient for large-scale applications.
[0065] Utilizing an MoE architecture (FIGs. 1 and 2) to replace traditional feedforward neural network layers while implementing a self-supervised biological language model with concurrent training of different modalities facilitates learning of shared information across different modalities while preserving the unique properties of each modality. This approach enhances the extraction of complex patterns and relationships inherent in the processes of gene transcription and protein translation.
[0066] Furthermore, the present disclosure provides a novel integration of Hyena and MoE within a unified pre-training framework for multimodal biological data. Specifically, the present disclosure leverages the Hyena operator’s long-range sequence modeling capability to enhance contextual representation before and after sparse expert routing in the MoE module. This synergistic combination allows efficient handling of heterogeneous biological modalities (e.g., DNA, RNA, and protein) while maintaining scalability and precision.Experimental Data
[0067] FIG. 6A shows a table which provides basic information regarding datasets used for downstream testing. The selection of downstream tasks is diverse and encompasses various combinations of the three sequence modalities, allowing for a comprehensive assessment of the pre-trained model's transferability.
[0068] [Rule 91,05.12.2025]FIG. 6B shows a table which illustrates performance of the method of the present disclosure against other specialized or generalist genomics foundational models. Notably, Zhou, Z., Ji, Y., Li, W., Dutta, P., Davuluri, R.V., &Liu, H. (2023) . DNABERT-2: Efficient foundation model and benchmark for multi-species genome (hereinafter DNABERT-2) is a foundational model trained on genomics data, while Zeming Lin et al. Evolutionary-scale prediction of atomic-level protein structure with a language model Science 379, 1123-1130 (2023) (hereinafter ESM-2) is trained on protein data. Both LucaOne and the present disclosure are foundational models trained jointly on DNA, mRNA, and protein data; while results for the ESM-2 model on DNA data are blank, and similarly, the results for DNABERT-2 on protein data are also blank. The results in FIG. 6B indicate that joint training across the three modalities has a significant impact on enhancing the overall performance of the model. Furthermore, a comparison between method of the present disclosure and the results from LucaOne demonstrates that the introduction of MoE, such as in the method 100 of FIG. 1 and / or framework of FIG. 2, contributes positively to performance improvement. The best performance is indicated by text in bold.
[0069] FIG. 6C shows a table which compares average performance of different model variants (Hyena only, MoE only, and Hybrid) on three DNA benchmarks (Genome Understanding Evaluation (GUE) , Nucleotide Transformer Tasks, and Genomics Benchmark. The table shows that average performance of a model with hybrid neural-network architecture comprising Hyena and MoE modules exceeds a Hyena-only model and a MoE-only model.Sparse Mixture-of-Experts (MoE)
[0070] A sparse MoE employs the concept of conditional computation in which only selected experts are activated. A learned gate network G of a routing network is configured to determine which experts E receive segments of an input:
[0071] A gating function is configured to avoid G being 0:Gσ (x) =Softmax (x·Wg)Hyena Operator
[0072] A discrete convolution is a function of two arguments: an input u signal of length L and a learnable filter h. The linear (aperiodic) convolution of a (possibly infinitely long) measurable filter h with a length L input signal u is defined as :
[0073] Typically, ut∈RD, where D denotes the width of the signal or, in deep learning terminology, the number of channels. Without loss of generality, an analysis may relate to single-input single-output (SISO) layers, i.e., where D = 1. The extension to the multiple-input multiple-output (MIMO) case, which is standard in conventional convolutional layers, follows directly. In this setting, the input signal can be represented as a vector u∈RL , and the convolution operation can be expressed as a matrix-vector product between the input signal and the Toeplitz kernel matrix Sh∈RL×L generated by the filter h:
[0074] Explicit convolution refers to the direct use of a filter kernel to perform sliding window calculations on an input signal or image, and it is widely applied in Convolutional Neural Networks (CNNs) . The filter is an explicitly defined matrix or tensor with known weights, and the convolution operation is achieved through element-wise multiplication and accumulation. For a one-dimensional signal u and a filter h, the output y of the explicit convolution can be expressed as:
[0075] However, the number of parameters of the filters scales linearly with filter size, which can be prohibitively expensive computationally. To decouple the parameter count from the filter size, we can alternatively express the filter h as a parametric function of the time step t. Such a parametric function is called implicit. A commonly chosen approach for implicit parametrization is to select filter h as the response function of the linear SSM. The convenient choice of x0= 0 renders the input-output map to a simple convolution: where δ denotes the Kronecker delta. Then the filter h can be identified as: where A, B, C, D are the learned parameters of the filter.
[0076] Hyena is a class of data-controlled operators consisting of a recurrence of multiplicative gating interactions and long convolutions. Let v, x1, …, xN be projections of the input and let h1, …, hN be a set of learnable filters. The definition of Order-N Hyena Operator is:
[0077] The input-output map can be rewritten as y=xN· (hN* (xN-1 · (hN-1* (…) ) ) ) . Besides, the element-wise product in the time domain corresponds to convolution in the frequency domain; for example, we define and as the DFT of x and u respectively, so we have Therefore, Hyena is alternatively applying convolutions in the time and then the frequency domain.
[0078] Hyena operators build on the H3 mechanism developed by Poli, Michael; Massaroli, Stefano; Nguyen, Eric; Fu, Daniel Y.; Dao, Tri; Baccus, Stephen; Bengio, Yoshua; Ermon, Stefano; and Ré, Christopher. “Hyena Hierarchy: Towards Larger Convolutional Language Models. ” Proceedings of the 40th International Conference on Machine Learning (ICML 2023) , Proceedings of Machine Learning Research, Volume 202, pp. 28043-28078, 2023. In the SISO case, let Dq and Dk may be the L-by-L diagonal matrices whose respective main diagonal entries are the respective entries of q and k. H3 implements a surrogate attention matrix through a data-driven, parametrized decomposition into four components: H3 (q, k, v) =A (q, k) v,where Sψ, are the Toeplitz matrices of learnable causal filters ψ, parametrized via SSMs. The surrogate attention mechanism of Hyena-2: q, k, v |→ y yt=qt (ψ*z) t.
[0079] Hyena represents a generalization of the previous equations for an arbitrary number of projections with implicit free-form long filters for the convolutions. The resulting recurrence can be represented in matrix form y=H (u) v. Let and be the Toeplitz matrix corresponding to filterhn. The resulting Hyena recurrence can be rewritten in matrix form:
[0080] Details on the convolution parametrization can be presented. Representing the filters of each Hyena operator as a map from the time domain t to values ht, and learning it with a shallow feed-forward neural network:
[0081] Such an approach builds on the neural implicit representation literature, which has found application in long convolution layers. Advantages include decoupling of filter length and parameter cost.
[0082] FIG. 7 is a schematic diagram of an exemplary computing device 700 that may be utilized for executing and performing the methods 100 and 300 of FIGs. 1 and 3 and / or implementing the frameworks of FIGs. 2A, 2B, 4A and 4B. The following description of the computing device 700 is provided by way of example only and is not intended to be limiting.
[0083] As depicted in FIG. 7, the example computing device 700 may include a processor 704 for executing software routines / programs, e.g., operations performed by various modules and architecture described in the present disclosure. While only a single processor is shown for brevity, the computing device 700 may also be configured as a multi-processor system, i.e., includes multiple processors. The processor 704 is coupled to a communication infrastructure 706 for communication with other components of the computing device 700. The communication infrastructure 706 may include, for example, a communications bus, a crossbar network, or a network.
[0084] The computing device 700 further includes a main memory 708, such as a random-access memory (RAM) , and a secondary memory 710. The secondary memory 710 may include, for example, a hard disk drive 712 and / or a removable storage drive 714, which may include a floppy disk drive, a magnetic tape drive, an optical disk drive, or the like. The removable storage drive 714 reads from and / or writes to a removable storage unit 718, as known in the art. The removable storage unit 718 may include a floppy disk, magnetic tape, optical disk, universal serial bus (USB) flash disk, or the like, which is read by and / or written to by removable storage drive 714. As may be appreciated by skilled persons in the art, the removable storage unit 718 may further include a computer readable storage medium having stored therein computer executable program code instructions and / or data.
[0085] In other aspects, the secondary memory 710 may additionally or alternatively include other similar means for allowing computer programs or other instructions to be loaded into the computing device 700 for execution. Such means may include, for example, a removable storage unit 722 and an associated interface 720. Examples of a removable storage unit 722 and interface 720 may include a USB flash drive and a USB interface, a removable memory chip, e.g., an EPROM or PROM, and associated socket, and other exemplary removable storage units 722 and interfaces 720, which may enable software programs and / or data to be transferred between the removable storage unit 722 and the computing device 700.
[0086] The computing device 700 also includes at least one communication interface 724. The communication interface 724 allows software programs and data to be transferred between computing device 700 and external devices, via communication path 726. In various aspects, the communication interface 724 permits data to be transferred between the computing device 700 and a data communication network, such as a public data or private data communication network. The communication interface 724 may be used to exchange data between different computing devices 700 that may together form part of an interconnected computer network. Examples of a communication interface 724 may include a modem, a network interface, e.g., an Ethernet card, a communication port, an antenna with associated circuitry or the like. The communication interface 724 may be configured as wired or wireless. Software and data transferred via the communication interface 724 are in the form of signals, which can be electronic, electromagnetic, optical or other signals capable of being received by communication interface 724. These signals are provided to the communication interface via the communication path 726.
[0087] The computing device 700 may include a display interface 702 configured to perform operations for rendering images to an associated display 730, and an audio interface 732 for performing operations for playing audio content via associated speaker (s) 734.
[0088] Computer readable storage media refers to any non-transitory tangible storage medium that provides recorded instructions and / or data to the computing device 700 for execution and / or processing. Examples of such storage media include floppy disks, USB disk, magnetic tape, CD-ROM, DVD, Blu-rayTM Disc, a hard disk drive, a ROM or integrated circuit, USB memory, a magneto-optical disk, or a computer readable card such as a PCMCIA card or the like, whether or not such devices are internal or external of the computing device 700. Examples of transitory or non-tangible computer readable transmission media that may also participate in the provision of software, application programs, instructions and / or data to the computing device 700 include radio or infra-red transmission channels as well as a network connection to another computer or networked device, and the Internet or Intranets including e-mail transmissions and information recorded on websites and the like.
[0089] The computer programs (also termed computer program code / instruction) are stored in the main memory 708 and / or the secondary memory 710. Computer programs may also be received via the communication interface 724. Such computer programs, when executed, enable the computing device 700 to perform one or more aspects of the present disclosure afore discussed. In various aspects of the present disclosure, the computer programs, which when executed, enable the processor 704 to perform aspect (s) of the present disclosure. Accordingly, such computer programs may represent (logic) controllers of the computing device 700.
[0090] Software may be stored in a computer program product and loaded into the computing device 700, using the removable storage drive 714, the hard disk drive 712, or the interface 720. Alternatively, the computer program product may be downloaded directly onto the computing device 700, via the communication path 726. The software, when executed by the processor 704, causes the computing device 700 to perform aspects of the present disclosure.
[0091] It is to be understood that the computing device 700 in FIG. 7 is presented merely by way of example. Hence, in some aspects, one or more features of the computing device 700 may be omitted. Also, in other aspects, one or more features of the computing device 700 may be combined together, or collocated. Additionally, in some aspects, one or more features of the computing device 700 may be divided into one or more component parts.
[0092] It is to be appreciated that the elements illustrated in FIG. 7 may further function to provide means for performing the various functions of the disclosed methods 100 and 300 in FIGs. 1 and 3, as described in accordance with aspects of the present disclosure. Also, the term “computing device” 700 may include or may refer to a mobile device, a wireless device, a remote device, a handheld device, a smartphone, a tablet computer, a laptop computer, a computer server, a computer terminal, a blade server, among other examples. The computing device 700 described herein may be able to communicate with various types of devices, such as other computing devices that may sometimes act as relays, or work together under configuration to function as a computer cluster for performing high-performance computing.
[0093] All of the methods described herein describe possible implementations, and that the operations and the steps may be rearranged or otherwise modified and that other implementations are possible. Further, aspects from two or more of the methods, if applicable, may be combined.
[0094] Information and signals described herein may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.
[0095] The various illustrative blocks and components described in connection with the disclosure herein may be implemented or performed with a general-purpose processor, a DSP, an ASIC, a CPU, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0096] The functions described herein may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored on or transmitted over as one or more instructions or code on a computer-readable medium. Other examples and implementations are within the scope of the disclosure and appended claims. For example, due to the nature of software, functions described herein may be implemented using software executed by a processor, hardware, firmware, hardwiring, or combinations of any of these. Features implementing functions may also be physically located at various positions, including being distributed such that portions of functions are implemented at different physical locations.
[0097] Computer-readable media includes both non-transitory computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. A non-transitory storage medium may be any available medium that may be accessed by a general-purpose or special purpose computer. By way of example, and not limitation, non-transitory computer-readable media may include RAM, ROM, electrically erasable programmable ROM (EEPROM) , flash memory, compact disk (CD) ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that may be used to carry or store desired program code means in the form of instructions or data structures and that may be accessed by a general-purpose or special-purpose computer, or a general-purpose or special-purpose processor. Also, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) , or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of computer-readable medium. Disk and disc, as used herein, include CD, laser disc, optical disc, digital versatile disc (DVD) , floppy disk and Blu-ray disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above are also included within the scope of computer-readable media.
[0098] As used herein, including in the claims, “or” as used in a list of items (for example, a list of items prefaced by a phrase such as “at least one of” or “one or more of” ) indicates an inclusive list such that, for example, a list of at least one of A, B, or C means A or B or C or AB or AC or BC or ABC (such as, A and B and C) . Also, as used herein, the phrase “based on” shall not be construed as a reference to a closed set of conditions. For example, an example step that is described as “based on condition A” may be based on both a condition A and a condition B without departing from the scope of the present disclosure. In other words, as used herein, the phrase “based on” shall be construed in the same manner as the phrase “based at least in part on” . Identifiers such as "first" , "second" , "third" , etc. are used only as labels, and are not intended to impose numerical requirements on their objects, nor in a way that imposes any relative position or timing sequence between limitations.
[0099] The description set forth herein, in connection with the appended drawings, describes example configurations and does not represent all the examples that may be implemented or that are within the scope of the claims. The term “example” used herein means “serving as an example, instance, or illustration, ” and not “preferred” or “advantageous over other examples” . The detailed description includes specific details for the purpose of providing an understanding of the described techniques. These techniques, however, may be practiced without these specific details. In some instances, known structures and devices are shown in block diagram form in order to avoid obscuring the concepts of the described examples.
[0100] The description herein is provided to enable a person having ordinary skill in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to a person having ordinary skill in the art, and the generic principles defined herein may be applied to other variations without departing from the scope of the disclosure. Thus, the disclosure is not limited to the examples and designs described herein, but to be accorded the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1.A computer-implemented method for pre-training a large-scale foundation model, wherein the method comprises:tokenizing datasets to obtain tokenized datasets, wherein the datasets include different biological modalities originating from DNA sequences, RNA sequences, and protein sequences;randomly masking a subset of tokens in each tokenized dataset to generate masked datasets;reconstructing at least one masked token associated with one of the biological modalities based on contextual embeddings derived from unmasked tokens associated with another one of the biological modalities to obtain a pre-trained large-scale foundation model comprising optimized tokenization and reconstruction parameters; andpre-training, by the large-scale foundation model comprising a sparse Mixture-of-Experts (MoE) module, input token representations derived from the tokenized datasets to obtain the pre-trained large-scale foundation model further comprising optimized routing and expert parameters.2.The computer-implemented method of claim 1, wherein the pre-training step comprises:pre-training, by the large-scale foundation model comprising at least one Hyena module and the sparse MoE module, by operatively coupling the at least one Hyena module and the sparse module to perform complementary long-range contextual processing and expert-specialized processing, such that outputs generated by one of the at least one Hyena module or the sparse MoE module, are utilized by an other module to generate globally mixed contextual embeddings to obtain the pre-trained large-scale foundation model further comprising optimized Hyena-operator parameters, wherein the other module is a remaining one of the at least one Hyena module and the sparse MoE module or a non-Hyena non-MoE module.3.The computer-implemented method of claim 2, wherein the at least one Hyena module includes a first Hyena module and a second Hyena module such that an output of the first Hyena layer forms an input to the sparse MoE layer, and an output of the sparse MoE layer forms an input to the second Hyena layer, wherein the sparse MoE layer comprises a routing network and at least two expert subnetworks connected thereto, each expert subnetwork having a plurality of experts, wherein the pre-training step includes:performing, by the first Hyena module, long-range convolutional context mixing on the input token representations to produce first Hyena-mixed representations;for each of the first Hyena-mixed representations, determining, by the routing network of the MoE, routing logits corresponding to the experts and activating, based on the routing logits, a subset of the experts;processing, by the activated experts of the MoE, the corresponding input token representations to generate expert outputs;combining the expert outputs to obtain a combined MoE output representation; andperforming, by the second Hyena module, long-range convolutional remixing on the combined MoE output representations to generate the globally mixed contextual embeddings,such that the pre-trained large-scale foundation model further comprises the optimized Hyena-operator parameters.4.The computer-implemented method of claim 3, wherein the determining step comprises:applying an auxiliary loss term to encourage a substantially uniform distribution of the first Hyena-mixed representations among the activated experts.5.The computer-implemented method of claim 3, wherein the at least one hybrid neural-network architecture comprises a plurality of hybrid neural-network architectures being connected in sequence such that an output of a preceding hybrid neural-network architecture forms an input to a subsequent hybrid neural-network architecture.6.The computer-implemented method of claim 1, wherein the sparse MoE layer comprises a routing network and at least two expert subnetworks connected thereto, each expert subnetwork having a plurality of experts, wherein the pre-training step comprises:for each input token representation, determining, by the routing network, routing logits corresponding to the experts and activating, based on the routing logits, a subset of the experts;processing, by the activated experts, the corresponding input token representations to generate expert outputs; andcombining the expert outputs to obtain a combined MoE output representation.7.The computer-implemented method of claim 6, wherein the determining step comprises:applying an auxiliary loss term to encourage a substantially uniform distribution of the input token representations among the activated experts.8.A computer-implemented method executed by one or more processors for performing an inference operation, the method comprising:receiving, by the one or more processors, input data instances respectively corresponding to different biological modalities originating from DNA sequences, RNA sequences, and protein sequences;encoding, by a pre-trained large-scale foundation model stored in one or more memories, the input data instances to generate modality-specific embeddings;projecting, by the pre-trained large-scale foundation model, the modality-specific embeddings into a joint multi-modal embedding which characterizes inter-relationships among the different biological modalities; andperforming, by an inference module of the pre-trained large-scale foundation model, the inference operation comprising classification, regression, and / or prediction of a target biological feature based on the joint multi-modal embedding,wherein the target biological feature comprises a predicted attribute, state, or outcome associated with at least one of the different biological modalities, andwherein the pre-trained large-scale foundation model has been pre-trained by the method of any one of claim 1 to claim 7.9.A computing device comprising:one or more memories having executable code; andone or more processors coupled to the one or more memories, and configured to execute the executable code to cause the computing device to perform the method of any one of claim 1 to claim 8.10.A non-transitory computer-readable medium comprising executable code, which when executed by a processor of a computing device, cause the computing device to perform the method of any one of claim 1 to claim 8.