De novo prediction gene annotation method, system and equipment based on hybrid expert network and storage medium
Through the hybrid expert network method, the problem of existing gene annotation methods' dependence and computational efficiency on experimental data is solved, high-precision, cross-lineage gene annotation is achieved, and the efficiency and accuracy of genomic research are improved.
Patent Information
- Application Number
- CN202510477678.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-29
AI Technical Summary
The existing gene annotation methods rely on insufficient experimental data, making it difficult to achieve high-precision annotation in species lacking evidence such as RNA-seq, and are ineffective in computing, so they cannot effectively analyze complex gene structures and cross-lineage gene evolution, resulting in bottlenecks in genomic research.
The hybrid expert network is adopted to extract gene sequence characteristics through stacked convolution layers and multi-head self-attention mechanism, combine multiple expert subnets and relationship computing controllers for dynamic weight allocation, perform nucleotide multi-task learning, and use the Viterbi algorithm for decoding to achieve end-to-end gene structure prediction.
It realizes high-precision gene annotation without external evidence, improves cross-lineage adaptability and computing efficiency, significantly improves the recognition ability of gene regulatory elements and complex gene structures, and reduces computational complexity.
Smart Images

Figure CN120388624A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of gene annotation, and relates to a de novo prediction gene annotation method, system, device and storage medium based on a mixture of experts network. Background Art
[0002] Gene annotation is a fundamental task in genomics research, and its core goal is to systematically identify and locate functional elements in genomic sequences. As an important bridge connecting the original sequence with biological significance, gene annotation provides a reliable basis for subsequent functional research and comparative genomic analysis by accurately identifying key information such as the location of genes, exon-intron structures, transcription start and termination sites, etc. In the early stage of genomics research, gene structure annotation mainly relied on experimental methods. Expressed sequence tag (EST) sequencing and cDNA library analysis were the main means to obtain gene structure information at that time. Although these methods were accurate and reliable, they had low throughput and high costs, making it difficult to meet the needs of large-scale genomic annotation. With the development of computational biology, de novo prediction methods based on computational models have gradually become important tools for gene annotation. These methods use algorithms such as hidden Markov models (HMMs) to predict gene structures by analyzing sequence patterns such as codon usage preferences and splice site characteristics, significantly improving the annotation efficiency. However, relying solely on computational prediction also has limitations. Especially in newly sequenced species lacking training data, the prediction accuracy is often difficult to guarantee.
[0003] In recent years, with the popularization of transcriptome sequencing (RNA-Seq) technology, gene structure annotation has entered a new stage of development. RNA-Seq data can provide direct transcript evidence, which can not only help verify computational prediction results but also discover new alternative splicing isoforms. After aligning transcriptome sequencing data (RNA-Seq reads) to genomic sequences, the accuracy of structure annotation can be significantly improved: splice alignments can indicate intron positions, while increases and decreases in coverage on genomic sequences can mark exon-noncoding region boundaries. Therefore, gene prediction tools represented by AUGUSTUS improve accuracy by combining RNA-Seq data.
[0004] However, these past methods have some core flaws: (1) Over-reliance on experimental data: Existing automated annotation processes (such as the AUGUSTUS, MAKER, and BRAKER series) are highly dependent on experimental evidence like RNA-seq, resulting in poor annotation quality for species with scarce experimental data and annotation biases for spatio-temporally specific expressed genes. (2) Single-dimensional evolutionary modeling: Mainstream methods use Hidden Markov Models (HMMs) or Long Short-Term Memory Networks (LSTMs) with fixed parameters, unable to dynamically capture the multi-dimensional characteristics of eukaryotic gene evolution: neglecting both the dynamic processes of gene fusion / loss in vertical evolution and the cross-lineage impacts of horizontal gene transfer, leading to weak capabilities in tracing the origin of cross-species genes. (3) Inefficient parsing of long genes: Ultra-long genes (containing hundreds of thousands of bases) formed by gene fusion or exon extension contain key biological information, but existing algorithms are limited by local sequence analysis patterns: unable to identify the interactions of distal genomic regulatory elements and difficult to parse the pairing rules of complex splicing sites, resulting in the omission of important functional genes. (4) Lack of cross-lineage integration: Current methods are confined to single-lineage modeling and lack a multi-lineage joint analysis mechanism: unable to simultaneously track the conservative changes of genes in the vertical genetic tree and identify cross-lineage gene networks formed by horizontal transfer, severely restricting the construction accuracy of the whole-genome evolution map. (5) Prominent computational efficiency bottleneck: Existing annotation processes have significant computational bottlenecks in the integration and processing of multi-evidence data (such as cross-species homologous alignment, deep RNA-Seq splicing). Integrated tools represented by MAKER2 often require processing cycles of several weeks to several months when performing whole-genome annotation. This inefficiency not only hinders real-time annotation analysis but also makes it difficult to achieve iterative optimization of large-scale genomic projects. These systematic flaws have made gene annotation a serious bottleneck in the post-genomic era, directly resulting in the lack of annotation for nearly half of the species in the NCBI GenBank database and becoming a major technical bottleneck restricting genomics research. Therefore, there is an urgent need to develop a gene annotation tool that starts directly from genomic sequences, can model long-range dependencies in sequences, and integrate the dynamic associations between different evolutionary branches, thereby breaking through the systematic bottlenecks of traditional methods in computational efficiency, complex gene structure parsing, and cross-lineage gene evolution tracking, and providing accurate and efficient annotation support for genomic sequence analysis and large-scale genomic projects. Summary of the Invention
[0005] An object of the present invention is to overcome the above-mentioned drawbacks of the prior art and provide a de novo prediction gene annotation method, system, device, and storage medium based on a mixture of experts network, which does not rely on any evidence and can perform high-precision gene annotation across taxonomically different lineages.
[0006] To achieve the above object, the present invention adopts the following technical solutions: A method for de novo predicting gene annotation based on a mixture of experts network, comprising the following processes: S1, dividing the entire gene sequence into a number of core regions in sequence, and each core region extends a set distance upstream and downstream thereof; S2, using a stacked convolutional layer to extract local features of each core region, and a positional encoding and a multi-head self-attention mechanism to simulate the long-distance interaction relationship between local features, so as to obtain a gene sequence feature representation that has learned distal dependencies; S3, using multiple expert sub-networks to capture the biological evolution relationship of the gene sequence feature representation that has learned distal dependencies, and using a relationship calculation controller to perform dynamic weight allocation on the biological evolution relationship results obtained by each expert sub-network, so as to obtain a gene sequence feature representation that has learned the biological evolution relationship; S4, performing nucleotide multi-task learning on the gene sequence feature representation that has learned the biological evolution relationship, so as to obtain the prediction probability of the task corresponding information of each nucleotide in the gene sequence; S5, dividing the gene sequence into a number of candidate gene regions according to the prediction probability, and decoding the nucleotide prediction probability within each candidate gene region, so as to obtain the gene structure prediction result of the gene sequence.
[0007] Preferably, in S1, error annotation recognition and marking are performed on the entire genome sequence.
[0008] Preferably, the error categories include: missing untranslated regions, empty-record genes, missing start codons or stop codons, abnormal intron lengths, CDS phase errors, exon out-of-bounds, internal exon overlaps within genes, and CDS-exon mismatches.
[0009] Preferably, in S4, the tasks of nucleotide learning include a class prediction task, a phase prediction task, and a state prediction task.
[0010] Preferably, the class prediction task is specifically: classifying each nucleotide to determine its corresponding functional region; the phase prediction task is specifically: determining which open reading frame phase each nucleotide belongs to, that is, the relative position in the CDS, for determining the translation start and end points; the state prediction task is specifically: judging whether the categories of two adjacent nucleotides are the same, and if different, it indicates that it may be located at the functional boundary.
[0011] Preferably, in S5, preset states, including intergenic states, untranslated region states, start codon states, stop codon states, CDS states, and intron states, each preset state obtains the corresponding emission probability or state transition probability according to the nucleotide multi-task learning result, and decodes according to the emission probability or state transition probability of each preset state, so as to obtain the gene structure prediction result.
[0012] Preferably, the Viterbi algorithm is adopted to perform an optimal state path search in each candidate gene region according to the emission probability or state transition probability of each preset state, output the optimal state path, and deduce the gene structure based on the optimal state path.
[0013] A de novo prediction gene annotation system based on a mixture of experts network, comprising: A context expansion module, configured to sequentially divide the entire gene sequence into a plurality of core regions, and each core region extends a set distance upstream and downstream thereof; A distal information module, configured to extract local features of each core region by using a stacked convolutional layer, and a positional encoding and a multi-head self-attention mechanism to simulate the long-distance interaction relationship between local features, so as to obtain a gene sequence feature representation that has learned distal dependencies; A joint evolution module, configured to capture biological evolution relationships by using a plurality of expert sub-networks for the gene sequence feature representation that has learned distal dependencies, and perform dynamic weight allocation on the biological evolution relationship results obtained by each expert sub-network by using a relationship calculation controller, so as to obtain a gene sequence feature representation that has learned biological evolution relationships; A multi-task learning module, configured to perform nucleotide multi-task learning on the gene sequence feature representation that has learned biological evolution relationships, so as to obtain the prediction probability of the task corresponding information of each nucleotide in the gene sequence; A gene decoding module, configured to divide the gene sequence into a plurality of candidate gene regions according to the prediction probability, decode the nucleotide prediction probability in each candidate gene region, and obtain the gene structure predicted for the gene sequence.
[0014] A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the de novo prediction gene annotation method based on the mixture of experts network are implemented.
[0015] A computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the de novo prediction gene annotation method based on the mixture of experts network are implemented.
[0016] Compared with the prior art, the present invention has the following beneficial effects: The present invention is designed to use only genomic sequences as input, learn information about functional elements and transcriptional features entirely from the inherent patterns of the sequences, achieve high-precision gene annotation without the assistance of any external evidence such as RNA-seq, and be able to perform high-precision gene annotation across taxonomically different lineages. By using context expansion, it can capture cross-regional regulatory signals and long-range structural dependencies in subsequent processes, enhancing the ability to identify gene regulatory elements and complex gene structures; by using the distal information strategy, compared with traditional recurrent neural networks or global convolutional networks, it significantly reduces the computational overhead and parameter scale, improving the scalability of the model; by using the joint evolution strategy, it endows the model with the flexibility of cross-species modeling, avoids the generalization errors that may occur when a unified model structure processes species-specific features, and effectively enhances the species adaptation ability of gene annotation and the depth of sequence evolution modeling.
[0017] Furthermore, a mechanism for automatically identifying annotation errors and masking during training is introduced, which can effectively suppress the interference of incorrect labels on model parameter updates while retaining the ability of these regions to provide context support for other normal regions, thereby enhancing the training stability and generalization ability of the model.
[0018] Furthermore, based on the probability distribution of the multi-task output of the neural network, candidate gene regions are divided, and the Viterbi algorithm is used for optimal path search, which significantly reduces the computational complexity while ensuring biological accuracy. This module does not rely on manually set transition matrices or biological prior rules and is completely driven by the output of the deep model, thus realizing an end-to-end, parameter-adaptive, and fine-grained gene structure prediction process. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic structural diagram of the de novo prediction gene annotation system based on the mixture of experts network described in Embodiment 2 of the present invention.
[0020] Figure 2 It is a schematic diagram of the performance comparison between ANNEVO of the present invention and Augustus among different clades in the RefSeq test set.
[0021] Figure 3 It is a schematic diagram of the benchmark test of ANNEVO of the present invention and the evidence-assisted annotation pipeline on the model species.
[0022] Figure 4 It is a schematic diagram of the computational efficiency comparison between ANNEVO and Augustus of the present invention.
[0023] Figure 5 It is a schematic diagram of the proportion of fully predicted genes divided by gene length groups for ANNEVO and Augustus of the present invention.
[0024] Figure 6 Schematic diagram of the BUSCO integrity comparison between ANNEVO of the present invention and the reference annotations of 793 species.
[0025] Figure 7 Schematic diagram of a representative example of ANNEVO of the present invention correcting gene fusion with incorrect reference annotations. Detailed implementation manners
[0026] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.
[0027] This embodiment provides a de novo prediction gene annotation method based on a mixture of experts network, including the following processes: S1, the entire gene sequence is sequentially divided into a number of core regions, each core region extends a set distance upstream and downstream thereof, and the entire genomic sequence is identified and marked for incorrect annotations. The incorrect categories include: missing untranslated regions, empty-record genes, missing start codons or stop codons, abnormal intron lengths, CDS phase errors, exon out-of-bounds, overlapping exons within genes, and mismatches between CDS and exons.
[0028] A context expansion mechanism for enhancing the representation ability of genomic sequences is proposed. This mechanism adopts a fixed-length sliding window strategy, divides the genomic sequence into several core regions, and extends fixed-length flanking sequences upstream and downstream of each core region, thereby constructing a context expansion window containing information within a 40-kb range. By introducing a total of 20 kb of context sequences on both sides, the model can capture cross-region regulatory signals and long-range structural dependencies, and improve the recognition ability of gene regulatory elements and complex gene structures.
[0029] To improve the model's learning ability for real gene annotation patterns, a mechanism for automatic identification of annotation errors and masking during training is introduced. This mechanism automatically detects and marks eight common types of annotation errors based on a set of predefined rules, including UTR deletions, empty-record genes, start / stop codon deletions, abnormal intron lengths, CDS phase errors, exon out-of-bounds, overlapping exons within / outside genes, and mismatches between CDS and exons, etc. During the model training process, all regions marked as containing annotation errors are masked during the loss function calculation stage, that is, although these regions are used to provide context feature inputs, they do not participate in the error backpropagation of the model. This strategy can effectively suppress the interference of incorrect labels on the update of model parameters, while retaining the ability of these regions to provide context support for other normal regions, thereby improving the training stability and generalization ability of the model.
[0030] S2. Use stacked convolutional layers to extract local features of each core region, and position encoding and multi-head self-attention mechanism to simulate the long-range interaction relationship between local features, obtaining a gene sequence feature representation that learns distal dependencies.
[0031] Adopt a Hi-C inspired remote information module based on local feature compression and multi-head self-attention mechanism to jointly capture local functional signals and long-distance regulatory dependencies in genomic sequences. This module extracts local sequence patterns through a stacked convolutional tower structure, constructs high-dimensional local representation units in units of every 32 bases, and retains key features such as transcription factor binding sites and conserved sequences. On this basis, introduce position encoding and multi-head self-attention mechanism in the Transformer structure to simulate the interdependence between sequences across the entire 40kb context window, providing precise semantic modeling support for complex gene structure recognition.
[0032] S3. Use multiple expert sub-networks to capture biological evolution relationships from the gene sequence feature representation that learns distal dependencies, and use a relationship calculation controller to perform dynamic weight allocation on the biological evolution relationship results obtained by each expert sub-network, obtaining a gene sequence feature representation that learns biological evolution relationships.
[0033] Aiming at the gene structure differences between different species and the complexity of multi-source evolution paths, the joint evolution module based on a mixture of experts network consists of multiple exclusive sub-networks, and each sub-network learns a specific type of evolution pattern. The system introduces a dynamic relationship calculation controller to generate expert weight distributions according to the input sequence content and perform weighted fusion on the outputs of multiple experts. The core innovation of this module lies in that its expert weight allocation mechanism has dynamic adaptability, and can autonomously select the optimal modeling path according to the sequence features of different lineages and different regions. This strategy endows the model with flexibility in cross-species modeling, avoids the generalization errors that may occur when a unified model structure processes species-specific features, and effectively improves the species adaptation ability of gene annotation and the depth of sequence evolution modeling.
[0034] S4. Perform nucleotide multi-task learning on the gene sequence feature representation that learns biological evolution relationships to obtain the predicted probabilities of the task corresponding information of each nucleotide in the gene sequence. The tasks of nucleotide learning include class prediction task, phase prediction task, and state prediction task.
[0035] The class prediction task is specifically: classify each nucleotide to determine its corresponding functional region; the phase prediction task is specifically: determine which open reading frame phase each nucleotide belongs to, that is, the relative position in the CDS, used to determine the translation start and end points; the state prediction task is specifically: judge whether the categories of two adjacent nucleotides are the same, and if different, it means it may be located at the functional boundary.
[0036] S5. Divide a gene sequence into multiple candidate gene regions according to the predicted probability. The preset states include intergenic state, untranslated region state, start codon state, stop codon state, CDS state, and intron state. For each preset state, according to the nucleotide multitask learning results, obtain the corresponding emission probability or state transition probability. Use the Viterbi algorithm to perform an optimal state path search in each candidate gene region according to the emission probability or state transition probability of each preset state, output the optimal state path, and deduce the gene structure from the optimal state path to obtain the gene structure predicted for the gene sequence.
[0037] In another embodiment of the present invention, as Figure 1 shown, a de novo prediction gene annotation system based on a mixture of experts network (ANNEVO) is provided. The de novo prediction gene annotation system based on a mixture of experts network can be used to implement the above-mentioned de novo prediction gene annotation method based on a mixture of experts network. Specifically, the de novo prediction gene annotation system based on a mixture of experts network includes three main modules: a context expansion module, a neural network module, and a gene structure decoding module.
[0038] Context expansion module ( Figure 1a) is a key module in the ANNEVO model for enhancing the genomic sequence representation ability. The input of this component is the complete genome. This component first adopts a sliding window strategy, dividing the entire genome into several core regions in sequence. The length of each core region is fixed at 20 kb (20,000 base pairs). To ensure that the neural network model can obtain sufficient context information when predicting the function of each nucleotide in these core regions, each core region extends 10 kb upstream and downstream respectively, thus forming an extended sequence window with a total length of 40 kb, that is, including a core region and its flanking sequences on both sides. This design can capture potential long-distance regulatory elements and their effects on the core region, thereby enhancing the model's ability to identify complex gene structures and regulatory features. When the core region is located at the edge of the chromosome or sequence, resulting in less than 10 kb of flanking sequence, the system will use a zero vector to fill the insufficient part. This filling operation does not provide any biological information, but can keep the size of the input tensor consistent and avoid interfering with the model training process due to inconsistent lengths. To improve the robustness of the model and reduce the interference of noise information on the training results, ANNEVO introduces an error annotation identification and marking mechanism in the input data preprocessing stage to identify and mark errors in the entire genomic sequence. This mechanism can automatically identify and mark 8 common types of annotation errors before model training, specifically including: (1) Missing untranslated regions (UTRs): that is, the 5' or 3' UTRs are not marked in gene annotation; (2) Empty-record genes: referring to gene records in the annotation that have no exons or coding sequences (CDS); (3) Missing start or stop codons: indicating that there is no legitimate start or stop codon in the CDS region; (4) Abnormal intron length: referring to introns that are too short or do not conform to typical splicing rules; (5) CDS phase error: there is a break or error in the open reading frame of the triplet codon in CDS annotation; (6) Exon out-of-bounds: exons extend beyond the range of their corresponding gene annotation; (7) Overlapping exons within a gene: multiple exons within the same gene overlap in coordinates; (8) Mismatch between CDS and exons: the coordinates or structure of CDS is inconsistent with the exon structure. For regions identified as having annotation errors, a loss masking strategy is adopted during model training. Under this mechanism, although the nucleotide information of these error regions is still input into the model to provide context, they are masked in the calculation of the loss function, that is, they do not participate in the backpropagation of the loss, thus effectively avoiding the adverse effects of incorrect annotations on model parameter updates.This design takes into account both the tolerance of local error information and the promotion of overall structure learning, ensuring that the model can use the information in these regions to provide context support for other regions without being misled by incorrect labels. This component finally divides the complete genome into several independent samples with a length of 40 kb, where the core region of each sample is 20 kb and the flanking regions on both sides are 10 kb each. These data will be input into the neural network component for training and testing.
[0039] The neural network component ( Figure 1 b and Figure 1 d) is the core part of the ANNEVO model for fine-grained analysis of genomic sequences, and its goal is to automatically extract, integrate, and discriminate key functional features from raw sequences of tens of thousands of base pairs. This component adopts a multi-level feature fusion mechanism and a dynamic computing structure to adapt to the complex multi-scale and cross-region biological characteristics in genomic sequences. The entire neural network architecture consists of three tightly coupled functional modules, namely: (1) the distal information module; (2) the co-evolution module; (3) the multi-task learning module. Together, they construct a complete inference process from the raw sequence to the high-resolution annotation.
[0040] (1) Distal information module. This module focuses on jointly modeling local features and long-range dependencies in the sequence. Its design concept stems from the emphasis on long-range interactions in biological research, similar to how Hi-C technology reveals regulatory connections between regions by observing the chromatin spatial structure. In ANNEVO, this interaction is not based on experimental measurements but is simulated by using a deep neural network to perform structural modeling on nucleotide sequences. Its core consists of two stages: Local pattern learning. A set of stacked convolutional towers is used to extract local features of each core region. These convolutional layers are connected by residual connections to alleviate the vanishing gradient problem and facilitate the training of deep networks. After each layer of convolution processing, the model can fuse information over a longer range and finally compress the context of every 32 nucleotides into a high-dimensional vector, serving as a "local representation unit" to effectively capture micro-scale signals such as conservation and transcription factor binding sites; Long-range relationship modeling. After obtaining the local representation, the distal information module further uses the positional encoding and multi-head self-attention mechanism in the Transformer architecture to simulate the long-range interaction relationships between these local features. This design enables the neural network model to obtain remote dependency information from the entire 40kb context, thus effectively capturing complex cross-region patterns such as splicing and regulation. Since the computational cost of self-attention is quadratic in the sequence length, and calculating self-attention after compressing local information can reduce the length by 32 times, this strategy significantly reduces the computational overhead and parameter scale compared to traditional recurrent neural networks or global convolutional networks, improving the scalability of the model. The output of this module is the gene sequence feature representation that has learned distal dependencies and will be input into the next module.
[0041] (2)Co-evolution module. Considering the existence of different levels of genetic events (such as species divergence, horizontal gene transfer, lineage-specific changes, etc.) in the process of genome evolution, this module is designed specifically for modeling multi-source evolutionary patterns. It draws on the Mixture-of-Experts (MoE) framework in deep learning and introduces an expert architecture with adjustable weights. The module contains multiple expert sub-networks, each of which focuses on capturing specific types of evolutionary rules. The relationship calculation controller, which is a dynamic weight allocation mechanism, evaluates the weights of the outputs of each expert according to the structure and content of the input gene sequence feature representation. In essence, it is a lightweight neural network that generates a probability distribution based on the gene sequence features for weighted combination of different expert representations. The design of this controller enables the model to automatically select the most appropriate modeling strategy under different lineage backgrounds, with high biological adaptability. Therefore, during the training phase, the neural network model forms expert sub-networks for different evolutionary patterns such as vertical inheritance and horizontal transfer, which are responsible for learning specific evolutionary relationships. During prediction, the model automatically adjusts its feature representation according to the sequence composition. Finally, the output of this module is a weighted fusion representation of multiple expert sub-network representations, which not only comprehensively reflects the evolutionary background of the sequence itself but also retains the specific contribution information about different sub-lineages (such as different species, different branches), providing a more structured input basis for subsequent multi-task discrimination tasks. The output of this module is the gene sequence feature representation that has learned the biological evolutionary relationship and will be input into the next module.
[0042] (3) Multi-task learning module. To achieve a comprehensive prediction of gene structure, ANNEVO finally introduces a multi-task learning framework to concurrently complete multiple sequence discrimination tasks. This module adopts an extended multi-gate mixture-of-experts architecture, constructing independent prediction channels for different tasks respectively. The three core tasks include: Class prediction task: Classify each nucleotide to determine which of the following functional regions it belongs to: intergenic region, untranslated region (UTR), coding region (CDS), or intron; Phase prediction task: Determine which open reading frame phase (0 / 1 / 2) each nucleotide belongs to, that is, the relative position in the CDS, which is used to determine the translation start and end points; State prediction task: Judge whether the classes of two adjacent nucleotides are the same. If different, it may indicate a possible location at the functional boundary (such as exon-intron junction). This task is extremely crucial for boundary detection. Each task is responsible for by an independent expert network to focus on the specific discrimination features of that task. In addition, all tasks also share a common network branch to capture the common structural features among all tasks, such as splice sites, upstream and downstream regulatory patterns, etc. The advantage of multi-task learning is that the complementarity between different tasks can promote each other: for example, class prediction provides segment information, which helps with phase and state judgment; the boundary information in state prediction in turn optimizes the continuity of class labels. Finally, the model realizes an all-round and multi-level modeling of the functional state of nucleotides through the integration of the three types of prediction results, thereby being able to accurately define the structural composition of genes. The output of this module is the predicted probabilities of three kinds of information for each nucleotide, and these probabilities will be used as the emission probabilities and state transition probabilities during decoding.
[0043] Gene decoding module ( Figure 1 c) is the last step in the ANNEVO model to achieve structured gene annotation. Its task is to integrate the multi-dimensional probability matrix output by the neural network into a set of gene structure models that conform to biological rules. Different from traditional decoding methods that rely on explicit state control and a large number of prior parameter tuning, ANNEVO's decoding framework adopts a probability-driven global optimization strategy and performs an optimal path search in the potential gene region by the Viterbi algorithm to automatically deduce a complete and accurate gene structure.
[0044] To improve decoding efficiency, ANNEVO does not directly perform global state prediction on the entire chromosome or sequence. Instead, it first roughly delimits candidate gene regions (potential gene regions), i.e., fragments that may contain the complete gene structure, through the probability distribution learned by the neural network model. These regions are not precise prediction boundaries but serve as a heuristic segmentation marker, enabling the subsequent decoding process to run in parallel within each candidate region, thus significantly improving decoding efficiency. This "segmentation by gene" strategy has significant advantages over traditional Hidden Markov Model (HMM) methods. Traditional HMM is limited by sequence state dependence and can only be processed in parallel at the contig level; while ANNEVO's method can perform parallel decoding at the single-gene granularity, suitable for large-scale whole-genome annotation tasks. Therefore, the gene decoding module pre-divides potential gene regions based on probability to reduce the search space, and the output is the range of these regions. Subsequent decoding can be performed only within these regions.
[0045] ANNEVO adopts a highly optimized set of minimum state sets to balance biological accuracy and computational complexity. The preset states include 1 intergenic state, 2 untranslated region states (5'UTR and 3'UTR), 1 set of start codon states, 1 set of stop codon states, 3 CDS states (corresponding to different phases), and 5 sets of intron states (covering different splicing types). During the decoding process, the advantage of ANNEVO lies in that both its emission probabilities and transition probabilities are derived from the multi-task prediction results directly output by the neural network model, completely avoiding the need to manually specify parameters or empirical distributions. Each preset state has a corresponding emission probability or transition probability. The emission probability mainly comes from the class prediction output of the multi-task module, and the emission probability mainly comes from the class prediction output of the neural network model. The state membership probability of each nucleotide is determined by its predicted distribution in four types of structures (intergenic, UTR, CDS, intron). Among them, the emission probability of the CDS state also integrates the phase prediction result to ensure the consistency of the translation frame. The transition probability is defined based on the state prediction output of the neural network model. Especially at the boundaries of functional elements, the probability of state transition reflects the neural network model's ability to recognize biological boundaries such as splice sites, start / stop codons, and UTR / CDS interfaces. Since both of these probability matrices are generated by the neural network based on context features and evolutionary signals, without external parameter tuning or manual experience intervention, it effectively avoids the introduction of human biases and ensures the biological rationality of the decoding path and the specificity of individual genes.
[0046] It is decoded by the Viterbi algorithm. In each candidate gene region, an optimal state path search is performed according to the emission probability or state transition probability of each preset state, and the optimal state path is output, and then a complete and accurate gene structure is automatically deduced.
[0047] The entire process, from the input of the original genomic sequence to the completion of structured gene annotation, realizes end-to-end automated processing.
[0048] The ANNEVO described in this embodiment is designed to use only genomic sequences as input, learns information about functional elements and transcriptional features entirely from the intrinsic patterns of the sequences, and can achieve high-precision gene annotation without the assistance of any external evidence such as RNA-seq.
[0049] The ANNEVO described in this embodiment can perform high-precision gene annotation across taxonomically different lineages: ANNEVO was evaluated on 566 species from RefSeq, which cover phylogenetically diverse branches: fungi, embryophytes, invertebrates, vertebrates_mammals, and vertebrates_others. For a strict baseline comparison, Augustus was implemented under optimized conditions, where species-specific training was performed on each test organism using evolutionarily close relatives. ANNEVO showed consistent advantages in all five clades. Compared with Augustus, the average nucleotide-level F1 score increased by 7.5 - 22.3%, the gene-level F1 score increased by 9.9 - 38.5%, and the BUSCO score increased by 9 - 34.5% ( Figure 2 ), and each range represents the minimum to maximum average performance gain among the five clades. Notably, although ANNEVO uses a species-agnostic model (applying the same gene model to all species within a clade), these gains still persist, while Augustus' predictions rely on models tailored to individual species. On 12 model species, ANNEVO, without the need for external evidence, outperformed the evidence-assisted gene annotation pipelines Augustus-Evidence and GeneMark-ETP and achieved performance comparable to that of the evidence-assisted gene annotation pipeline BRAKER3 ( Figure 3 ).
[0050] This performance difference highlights ANNEVO's ability to distill cross-species evolutionary signals into a unified prediction framework. ANNEVO integrates joint evolutionary constraints through its dynamic network architecture, addressing the accuracy limitations inherent in traditional homology-driven methods that require species-specific parameterization.
[0051] The ANNEVO described in this embodiment has an efficient computing framework: under the same 48-thread parallel configuration, ANNEVO is 2 times faster than Augustus ( Figure 4 ). The gene annotation pipeline BRAKER3, which has comparable performance to ANNEVO, is 160 times slower than ANNEVO in annotating Arabidopsis thaliana (BRAKER3: 16 hours; ANNEVO: 6 minutes). This benefits from the parallel optimization of its gene structure decoding component. This efficiency brings practical benefits to large-scale genome projects. ANNEVO can complete the genome annotation of Arabidopsis thaliana in only 6 minutes and the human genome annotation in 4.5 hours, making it an ideal solution for fast and comprehensive genome annotation.
[0052] The ANNEVO described in this embodiment can achieve length-robust gene prediction: more than 1 million protein-coding genes of all mammalian test species are divided into four length groups: 0-10 kb, 10-40 kb, 40-100 kb, and >100 kb, and the fully predicted genes in each group are quantified. Quantitative analysis shows that as the gene length increases, the performance of Augustus drops sharply. Although Augustus has a complete prediction rate of 52.4% for short genes (0-10 kb), its performance for ultra-long genes (>100 kb) drops to only 25.5%, indicating a serious length-dependent limitation. This limitation stems from its short-term memory. In sharp contrast, ANNEVO exhibits remarkable length robustness, maintaining or even improving its prediction accuracy for longer genes ( Figure 5 ). The excellent performance in predicting long genes is directly attributed to the Hi-C-inspired long-range dependence modeling of ANNEVO, which can effectively capture the complex interactions within these extended genomic regions.
[0053] The ANNEVO described in this embodiment can improve the integrity of reference annotations: A key advantage of the ANNEVO ab initio method is that it can avoid annotation errors caused by missing or incomplete external evidence data, thus improving the integrity of reference annotations for many species in RefSeq and Ensembl. Among 793 species, ANNEVO obtained a higher BUSCO score than the reference annotation in 252 species (31.8% of the total) ( Figure 6 ), demonstrating its ability to improve the annotations of all five phylogenetic branches.
[0054] ANNEVO described in this embodiment can correct the errors in the reference annotations. Taking the representative example at the genomic coordinates C5:39,559,113-39,563,423 of the Brassica oleracea genome as an example, the Ensembl annotation wrongly fused two different genes, which is a structural error without splicing evidence to support. However, ANNEVO can correctly predict the correct structure of the gene and independently verify it through RNA-seq evidence ( Figure 7 ).
[0055] In another embodiment of the present invention, a terminal device is provided. The terminal device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used for the operation of the de novo prediction gene annotation method based on the mixture of experts network, including: S1, sequentially dividing the entire gene sequence into several core regions, and each core region extends a set distance upstream and downstream thereof; S2, using a stacked convolutional layer to extract the local features of each core region, position encoding and multi-head self-attention mechanism to simulate the long-distance interaction relationship between local features, and obtaining a gene sequence feature representation that has learned distal dependencies; S3, using multiple expert sub-networks to capture the biological evolution relationship of the gene sequence feature representation that has learned distal dependencies, and using a relationship calculation controller to perform dynamic weight allocation on the biological evolution relationship results obtained by each expert sub-network, and obtaining a gene sequence feature representation that has learned the biological evolution relationship; S4, performing nucleotide multi-task learning on the gene sequence feature representation that has learned the biological evolution relationship, and obtaining the prediction probability of the task corresponding information of each nucleotide in the gene sequence; S5, dividing the gene sequence into multiple candidate gene regions according to the prediction probability, and decoding the nucleotide prediction probability in each candidate gene region to obtain the predicted gene structure of the gene sequence.
[0056] In another embodiment, the present invention further provides a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a terminal device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. Moreover, one or more instructions suitable for being loaded and executed by a processor are stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory.
[0057] One or more instructions stored in the computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the method for de novo predicting gene annotation based on a mixture of experts network in the above embodiments; the one or more instructions in the computer-readable storage medium are loaded and executed by the processor to perform the following steps: S1, the entire gene sequence is sequentially divided into several core regions, and each core region extends a set distance upstream and downstream thereof; S2, a stacked convolutional layer is used to extract the local features of each core region, and positional encoding and a multi-head self-attention mechanism are used to simulate the long-distance interaction relationship between the local features, so as to obtain a gene sequence feature representation that has learned distal dependencies; S3, multiple expert sub-networks are used to capture the biological evolution relationship of the gene sequence feature representation that has learned distal dependencies, and a relationship calculation controller is used to perform dynamic weight allocation on the biological evolution relationship results obtained by each expert sub-network, so as to obtain a gene sequence feature representation that has learned the biological evolution relationship; S4, nucleotide multi-task learning is performed on the gene sequence feature representation that has learned the biological evolution relationship to obtain the prediction probability of the task corresponding information of each nucleotide in the gene sequence; S5, multiple candidate gene regions are divided from the gene sequence according to the prediction probability, and the nucleotide prediction probabilities in each candidate gene region are decoded to obtain the gene structure predicted for the gene sequence.
[0058] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0059] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0060] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0061] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0062] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0063] In the above embodiments of the present application, the descriptions of the respective embodiments have their own focuses. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0064] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0065] The unit described as a separating component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0066] The above description is only a preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
[0067] It should be understood that the above description is for illustrative purposes and not for limitation. Many embodiments and many applications other than the examples provided will be apparent to those skilled in the art upon reading the above description. Therefore, the scope of this patent should not be determined by reference to the above description, but should be determined by reference to the full scope of the foregoing claims and the equivalents thereof. For the sake of completeness, all articles and references, including patent applications and publications of announcements, are incorporated herein by reference. The omission of any aspect of the subject matter disclosed herein in the foregoing claims is not intended to abandon such subject matter, nor should it be considered that the applicant has not considered such subject matter as part of the disclosed inventive subject matter.
Claims
1. A method for de novo prediction of gene annotation based on a mixture of experts network, characterized in that It includes the following processes: S1. Divide the entire gene sequence into several core regions in sequence, and each core region extends a set distance upstream and downstream thereof; S2. Use a stacked convolutional layer to extract the local features of each core region, and use positional encoding and multi-head self-attention mechanism to simulate the long-range interaction relationship between local features, so as to obtain a gene sequence feature representation that has learned distal dependencies; S3. Use multiple expert sub-networks to capture the biological evolution relationship of the gene sequence feature representation that has learned distal dependencies, and use a relationship calculation controller to perform dynamic weight allocation on the biological evolution relationship results obtained by each expert sub-network, so as to obtain a gene sequence feature representation that has learned the biological evolution relationship; S4. Perform nucleotide multi-task learning on the gene sequence feature representation that has learned the biological evolution relationship to obtain the prediction probability of the task corresponding information of each nucleotide in the gene sequence; S5. Divide the gene sequence into multiple candidate gene regions according to the prediction probability, and decode the nucleotide prediction probability in each candidate gene region to obtain the gene structure prediction result of the gene sequence.
2. The method for de novo predicting gene annotation based on a mixture of experts network according to claim 1, wherein In S1, error annotation recognition and marking are performed on the entire genome sequence.
3. The method for de novo prediction of gene annotation based on a mixture of experts network according to claim 2, wherein The error categories include: missing untranslated region, empty record gene, missing start codon or stop codon, abnormal intron length, CDS phase error, exon out-of-bounds, internal exon overlap in the gene, and CDS-exon mismatch.
4. The method for de novo predicting gene annotation based on a mixture of experts network according to claim 1, wherein In S4, the tasks of nucleotide learning include class prediction task, phase prediction task, and status prediction task.
5. The method for de novo predicting gene annotation based on a mixture of experts network according to claim 4, wherein The class prediction task is specifically: classify each nucleotide to determine its corresponding functional region; the phase prediction task is specifically: determine which open reading frame phase each nucleotide belongs to, that is, the relative position in the CDS, which is used to determine the translation start and end points; the status prediction task is specifically: judge whether the categories of two adjacent nucleotides are the same, and if they are different, it means that they may be located at the functional boundary.
6. The method for de novo predicting gene annotation based on a mixture of experts network according to claim 1, wherein In S5, preset states are included, including intergenic state, untranslated region state, start codon state, stop codon state, CDS state, and intron state. Each preset state obtains the corresponding emission probability or state transition probability according to the nucleotide multi-task learning result, and decodes according to the emission probability or state transition probability of each preset state to obtain the gene structure prediction result.
7. The method for de novo predicting gene annotation based on a mixture of experts network according to claim 6, wherein Use the Viterbi algorithm to perform an optimal state path search in each candidate gene region according to the emission probability or state transition probability of each preset state, output the optimal state path, and deduce the gene structure according to the optimal state path.
8. A gene annotation system for de novo prediction based on a mixture of experts network, characterized in that, It includes: A context expansion module for dividing the entire gene sequence into several core regions in sequence, and each core region extends a set distance upstream and downstream thereof; A distal information module for using a stacked convolutional layer to extract the local features of each core region, and using positional encoding and multi-head self-attention mechanism to simulate the long-range interaction relationship between local features, so as to obtain a gene sequence feature representation that has learned distal dependencies; The co-evolution module is used to capture the biological evolution relationship by using multiple expert sub-networks for the gene sequence features that have learned distal dependencies, and use a relationship calculation controller to dynamically allocate weights to the biological evolution relationship results obtained by each expert sub-network, so as to obtain the gene sequence feature representation that has learned the biological evolution relationship; The multi-task learning module is used to perform nucleotide multi-task learning on the gene sequence feature representation that has learned the biological evolution relationship, so as to obtain the prediction probability of the task corresponding information of each nucleotide in the gene sequence; The gene decoding module is used to divide the gene sequence into multiple candidate gene regions according to the prediction probability, and decode the nucleotide prediction probability in each candidate gene region to obtain the gene structure predicted for the gene sequence.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the de novo prediction gene annotation method based on the mixture of experts network according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the de novo prediction gene annotation method based on the mixture of experts network according to any one of claims 1 to 7.