Gene feature-based development dynamic expression mode prediction method, model, system and database structure

By constructing a developmental dynamic expression pattern prediction model based on gene characteristics and using multidimensional gene characteristic information for prediction, the problems of long detection cycle and high cost in existing technologies are solved, and fast and effective prediction of gene developmental dynamic expression patterns is achieved.

CN120636531AActive Publication Date: 2025-09-12BEIJING NORMAL UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510729680.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-12
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Existing technologies in the existing technology are unable to effectively, quickly and cost-effectively predict the dynamic expression patterns of gene development, resulting in long detection cycles and high costs.

Method used

By integrating genomics, transcriptomics and machine learning architecture, a developmental dynamic expression pattern prediction model is constructed, and predictions are made using multidimensional gene feature information, including determining the gene set, obtaining multidimensional gene feature information and inputting it into the prediction model, and outputting the developmental dynamic pattern.

Benefits of technology

The scope of species covered by the training data has been expanded, enabling rapid and effective prediction of dynamic expression patterns of gene development, reducing experimental time and scientific research costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636531A_ABST
    Figure CN120636531A_ABST
Patent Text Reader

Abstract

The invention discloses a development dynamic expression mode prediction method, model and system based on gene characteristics and a database structure, and relates to the technical field of biological information and artificial intelligence crossing, and the method comprises the following steps: determining a gene set of a development dynamic expression mode to be detected; acquiring multi-dimensional basis feature information of a to-be-detected gene; and inputting the multi-dimensional basis feature information into a development dynamic expression mode prediction model, and outputting a development dynamic mode corresponding to the target gene. According to the method disclosed by the invention, the range of species related to training data in a gene expression related model is greatly expanded, and the dynamic expression mode of the gene in biological tissue or organ development is quickly and efficiently predicted at low cost by innovatively utilizing the multidimensional gene characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics and artificial intelligence intersection technology, and in particular to a method, model, system and database structure for predicting developmental dynamic expression patterns based on gene characteristics. Background Art

[0002] The entire genome and the DNA sequence of its genes are largely identical across cells, tissues, and organs of an organism throughout its growth and development. However, gene expression varies significantly across tissue types and developmental stages. In fact, the initiation and termination of developmental programs are largely mediated by changes in gene transcription. Appropriate and appropriate gene transcription events are essential for normal biological processes. Numerous experimental results based on the development of animal tissues or organs indicate that genetic variants associated with human diseases can affect phenotypes by altering the temporal expression of genes. In plants, countless genes modulate their temporal expression to shape growth and development, particularly reproductive development, which is closely linked to crop yield. Therefore, genes with remarkably regular expression changes across organ or tissue developmental stages (i.e., developmentally dynamic genes) provide an excellent entry point for exploring the mechanisms of biological development and the underlying challenges. To some extent, genes with developmentally dynamic expression patterns are also more likely to be key functional molecules regulating biological processes. Therefore, many research teams around the world have invested significant resources, both physical and financial, using next-generation sequencing (NGS), third-generation sequencing (TGS), and single-cell sequencing (scRNA-seq) technologies to understand gene expression and developmental dynamics during organ development in plants and animals (Nature, 2019; Nature, 2019; Genome Biology, 2022; Cell, 2025). Our laboratory's research team has also used NGS and its analytical techniques to analyze the expression and identity of coding and non-coding genes during plant seed development, with some success (International Journal of Molecular Sciences, 2023; Journal of Plant Physiology, 2021; Journal of Plant Physiology, 2021). However, is it possible to quickly and easily understand developmental dynamics of gene expression without relying on complex experiments and expensive sequencing?

[0003] In recent years, with the continuous generation of biological big data and the booming development of artificial intelligence (AI), the application of machine learning (ML) and deep learning (DL) has gradually made it possible to understand gene expression in biological genomes. Recently, a research team at the Arc Institute published a multimodal AI model called Evo, which not only generates nearly realistic genome and gene sequences but also predicts chromatin accessibility, which is directly related to gene expression (Science, 2024; bioRxiv, 2025). In mammals such as humans, there are more examples of using ML or DL ​​to predict gene expression levels. Early models predicted gene expression based on more indirect information, such as chromatin immunoprecipitation data of transcription factors and histone modification data (PNAS, 2009). Deep convolutional neural networks, a popular DL approach, have achieved unprecedented results in predicting gene expression in humans and mice, with two representative works published in Cell Reports and Nature Methods. Xpresso can predict gene expression levels based solely on genomic DNA sequences and help infer transcriptional and post-transcriptional regulatory mechanisms (Cell Reports, 2020). The more well-known Enformer is the first DL model capable of processing DNA sequences up to 100 kb and can also predict gene expression intensity (Nature Methods, 2021). By applying machine learning and DL, cis-regulatory elements on genomic DNA can not only predict gene expression responses to temperature in maize, but also predict gene expression changes during tomato ripening (The Plant Cell, 2022). Last year, a study published in Nature Communications used DNA sequences in the flanking regions of genes to predict gene expression levels with an accuracy exceeding 80% in four species: Arabidopsis thaliana, tomato, sorghum, and maize, demonstrating comparable cross-species performance (Nature Communications, 2024). While it's clear that a range of DL models are indeed more powerful, research indicates that tree-based ML models significantly outperform DL models for processing tabular data (Advances in Neural Information Processing Systems, 2022). Therefore, when conducting research, it is important to avoid blindly relying on DL and to select appropriate model categories for different types of data. In short, using machine learning and deep learning architectures to "understand" the expression patterns of genes in the genome not only provides a new perspective for a deeper understanding of human diseases, but also offers a new strategy for achieving intelligent plant breeding.

[0004] Although the study of gene expression using artificial intelligence technology is in full swing, the inventors have noticed that there are at least the following problems in all the published related excellent works (Example Table 1 has been compared and illustrated by giving examples). First, the species covered by the data training set in the currently published models are still very few, particularly plants. Second, the vast majority of models are clustered around the use of genomic sequences in the hope of completing the gene expression-related prediction tasks, which seriously underestimates other available, valuable, and meaningful information. Third, almost all research teams at home and abroad have focused on the development of models for expression levels or expression differences, and the perspective of focusing on the problem is too one-sided and narrow. In the prior art, there has been no report on the predictive research on gene expression patterns in the development process, which has led to a long process cycle for obtaining developmental dynamic genes, a long time-consuming process, and high cost.

[0005] In summary, the species covered by current gene expression prediction models are not broad enough, the training data types used are too single, and they are unable to complete the important task of predicting the dynamics of gene expression during the developmental process, making it difficult to meet the actual needs of relevant scientific researchers. Summary of the Invention

[0006] In order to overcome the limitations of existing technologies, the technical problem to be solved by the present invention is to provide a method, model, system and database structure for predicting developmental dynamic expression patterns based on gene characteristics. By integrating genomics, transcriptomics, and machine learning architecture, rapid, effective and low-cost prediction of gene developmental dynamic expression patterns can be achieved; in order to fill the current technical gap in the prediction of gene developmental dynamic expression and provide a new tool for life science research, disease mechanism analysis, and intelligent crop breeding.

[0007] One aspect disclosed in the present invention provides a method for predicting developmental dynamic expression patterns based on gene features, comprising the following steps:

[0008] S1. Determine the gene set whose developmental dynamic expression pattern is to be detected;

[0009] S2. Obtaining multidimensional gene feature information of the gene set to be detected;

[0010] S3. Input the multidimensional gene feature information into the developmental dynamic expression pattern prediction model, and output the developmental dynamic pattern corresponding to the target gene, wherein the developmental dynamic expression pattern prediction model is trained based on a multidimensional gene feature information data matrix with a gene dynamic expression pattern label, and the gene dynamic expression pattern label is determined based on the gene expression matrix of the tissue or organ development process of the training sample.

[0011] As a preferred technical solution of the present invention, the gene set in step S1 is any one or any combination of the following:

[0012] (1) The complete set of genes in the newly sequenced and assembled genome;

[0013] (2) differentially expressed gene sets obtained based on high-throughput sequencing;

[0014] (3) A collection of genes that are specifically expressed in biological tissues or organs;

[0015] (4) species-specific or cross-species gene sets that exist within or between species;

[0016] (5) A set of candidate genes compiled through literature research;

[0017] In step S2, multi-dimensional gene feature information of the gene to be detected is obtained. Step S2 specifically includes:

[0018] S2.1: Obtain the numerical features of the gene set to be tested based on the gene sequence and annotation;

[0019] S2.2: Obtain the classification characteristics of the gene set to be tested based on gene properties and annotations;

[0020] S2.3: Integrate the numerical features and categorical features of the gene set to be tested into a data matrix containing multidimensional gene feature information.

[0021] In step S3, as a preferred technical solution of the present invention, the multidimensional gene feature information is input into the developmental dynamic expression pattern prediction model, and the developmental dynamic pattern corresponding to the target gene is output. Step S3 specifically includes:

[0022] S3.1: Loading the data matrix containing multidimensional gene feature information into the developmental dynamic expression pattern prediction model, the developmental dynamic expression pattern prediction model performs mapping calibration on the input features to ensure that the order of the input features is consistent with the order pre-trained by the developmental dynamic expression pattern prediction model;

[0023] S3.2: For each gene, the developmental dynamic expression pattern prediction model traverses the trained decision tree set, obtains the cumulative score of the prediction values ​​of all trees, and converts the cumulative score into a probability value;

[0024] S3.3: Manually set thresholds to determine developmental dynamics of gene expression.

[0025] Another aspect disclosed herein provides a series of related operations for constructing a developmental dynamic expression pattern prediction model, wherein the operations involved in constructing the developmental dynamic expression pattern prediction model include:

[0026] E1: Obtaining the gene expression matrix representing the species' organ development process;

[0027] E2: Acquisition of dynamic gene expression pattern signatures during development;

[0028] E3: Acquisition of multidimensional gene feature information of the whole genome;

[0029] E4: Training methods for prediction models of developmental dynamic expression patterns;

[0030] E5: Strategies for constructing prediction models for developmental dynamic expression patterns.

[0031] In operation E1, according to an embodiment of the present disclosure, a portion of the model training data is obtained based on a gene expression matrix representing the developmental process of an organ of a species. Operation E1 includes:

[0032] E1.1: Obtain transcriptome sequencing data representing the developmental process of tissues or organs of a species;

[0033] E1.2: Filter low-quality sequencing data and retain high-quality sequencing data;

[0034] E1.3: Obtain genomic data for representative species;

[0035] E1.4: Based on genomic data, assemble transcriptomes of representative species and complete transcriptome assembly for tissue or organ development;

[0036] E1.5: Quantitatively analyze the expression of genes and their transcripts to generate a gene expression matrix for organ developmental processes.

[0037] In operation E2, according to an embodiment of the present disclosure, a portion of the model training data is based on the acquisition of gene development dynamic expression pattern labels. Operation E2 includes:

[0038] E2.1: Manually organize the gene expression matrix and sample grouping information to ensure that the data format meets the requirements;

[0039] E2.2: Construct a regression design matrix for subsequent differential expression analysis;

[0040] E2.3: Obtain significantly differentially expressed genes using a generalized linear model;

[0041] E2.4: Through stepwise regression, screen genes whose expression changes significantly during the developmental process, i.e., developmentally dynamic genes.

[0042] In operation E3, according to an embodiment of the present disclosure, model training data is obtained from multi-dimensional gene feature information of the whole genome. Operation E3 includes:

[0043] E3.1: Based on gene sequences and annotations, obtain numerical gene features: length, GC content, and number of exons;

[0044] E3.2: Based on gene properties and annotations, obtain gene classification characteristics: whether it is a protein-coding gene; whether it is a gene that undergoes alternative splicing; whether it is a transposable element-related gene; whether it is an evolutionarily conserved gene;

[0045] E3.3: Integrate the numerical and categorical features of the whole genome gene set into a matrix containing multidimensional gene features.

[0046] In operation E4, according to an embodiment of the present disclosure, a training method for a developmental dynamic expression pattern prediction model includes:

[0047] E4.1: Correlate the dynamic expression pattern of each gene with its respective gene signature to form a data matrix of multidimensional gene signature information with gene dynamic expression pattern labels;

[0048] E4.2: Randomly split the data matrix into a training set and a validation set, with the training set accounting for 80% and the validation set accounting for 20%. Use grid search to tune hyperparameters to obtain the optimal hyperparameter configuration under the current conditions.

[0049] E4.3: Map "non-developmental dynamic" genes and "developmental dynamic" genes into two labels based on their expression patterns during development, and train the model under the optimal hyperparameter configuration;

[0050] E4.4: Monitor training and validation errors, plot error loss curves, use early stopping to determine the optimal number of training rounds to avoid overfitting, and save the trained prediction model.

[0051] E4.5: Assess the contribution of each trait to the developmental dynamic expression pattern;

[0052] E4.6: Input the test data set into the developmental dynamic expression pattern prediction model. The model will calibrate the input features to ensure that the order of the input features is consistent with the order in which the model was pre-trained.

[0053] E4.7: For each gene in the test set, the model traverses the set of trained decision trees and, combined with learning rate scaling, obtains the cumulative score of the predictions of all trees.

[0054] E4.8: Convert the cumulative scores into probability values, use the default threshold of 0.5 to determine whether gene expression is developmental dynamic, calculate the confusion matrix and the area under the receiver operating characteristic (ROC) curve (AUC), and plot the ROC curve.

[0055] E4.9: Considering that the default threshold is not the best threshold for demarcating gene expression dynamics, optimize the sensitivity and specificity of the model by plotting sensitivity and specificity curves and using the intersection of the two as the optimal threshold. Recalculate the confusion matrix and the area under the ROC curve, and plot the optimized ROC curve.

[0056] In operation E5, according to an embodiment of the present disclosure, a strategy for constructing a developmental dynamic expression pattern prediction model includes:

[0057] E5.1: Integrate the multidimensional gene feature information and developmental dynamic expression patterns of whole-genome genes of multiple species into multiple matrices, and construct multiple single-species SSM models through model training;

[0058] E5.2: Integrate the multidimensional gene feature information and developmental dynamic expression patterns of whole-genome genes of multiple species into a matrix, and construct a cross-species model CSM through model training.

[0059] Another aspect disclosed in the present invention provides a developmental dynamic expression pattern prediction model, characterized in that the model is constructed by any of the above-mentioned methods related to constructing a developmental dynamic expression pattern prediction model.

[0060] Another aspect disclosed herein provides a system for predicting the dynamic expression pattern of gene development, comprising:

[0061] M1: Gene set determination module, used to determine the gene set whose developmental dynamic expression pattern is to be detected;

[0062] M2: Gene feature acquisition module, used to obtain multidimensional gene feature information of the gene to be detected;

[0063] M3: Developmental dynamics prediction module, used to input the multidimensional gene feature information into the developmental dynamics expression pattern prediction model and output the developmental dynamics pattern corresponding to the target gene; wherein, the developmental dynamics expression pattern prediction model is trained and constructed based on the above-mentioned E1-E5 operation steps.

[0064] Another aspect disclosed herein provides an interactive and extensible gene signature information database structure for developmental dynamic expression pattern prediction. The database is constructed based on gene signature information and is specifically configured with extensible dimensions oriented toward gene signature information, as well as interactivity or potential interactivity between a gene set for a developmental dynamic expression pattern to be detected and a developmental dynamic expression pattern prediction model.

[0065] The main structure of the gene feature information database is set as follows: the database structure revolves around the GENE_MASTER table, and at the same time, it is set to have specific associations with several other tables, including:

[0066] D1: Core table element: GENE_MASTER; built as the core of the entire database structure, used to store basic information about genes, optionally including the basic gene identifier, name, description, or other basic gene data;

[0067] D2: Association tables and relationship elements; including:

[0068] D2.1: NUMERIC_FEATURES table; Relationship type setting: Build a one-to-many relationship between GENE_MASTER and NUMERIC_FEATURES, using "contains" to describe this relationship; Data content setting: Each record in the GENE_MASTER table can correspond to multiple records in the NUMERIC_FEATURES table; NUMERIC_FEATURES table is set to store optional numerical features of genes;

[0069] D2.2: CATEGORICAL_FEATURES table; Relationship type setting: GENE_MASTER and CATEGORICAL_FEATURES are set to a one-to-many "contains" relationship; Data content setting: CATEGORICAL_FEATURES table is used to store the classification features of genes;

[0070] D2.3: EXTENDED_FEATURES table; Relationship type setting: GENE_MASTER and EXTENDED_FEATURES are in a one-to-many "association" relationship; Data content setting: The EXTENDED_FEATURES table optionally stores extended feature information of genes, including information other than numerical and categorical features;

[0071] D2.4: ANNOTATION_SOURCES table; Relationship type setting: GENE_MASTER and ANNOTATION_SOURCES are set to a one-to-many "reference" relationship; Data content setting: ANNOTATION_SOURCES table stores the sources of gene annotation information;

[0072] The extension and interaction structure of D2.1-2.4 follows the following data relationship: Since the annotation information of a gene may come from multiple different sources, a GENE_MASTER record is also set to allow reference to multiple NUMERIC_FEATURE, CATEGORICAL_FEATURE, EXTENDED_FEATURES, and ANNOTATION_SOURCES records respectively.

[0073] Compared with the prior art, after adopting the above technical solution, the beneficial effects of the present invention include: after determining the gene set of the developmental dynamic expression pattern to be detected, the present invention obtains the multidimensional gene feature information of the gene to be detected, and inputs the multidimensional gene feature information into the developmental dynamic expression pattern prediction model to predict the developmental dynamic expression pattern of the target gene. On the one hand, it greatly expands the species range involved in the training data in the gene expression-related model, integrates multi-species information, and has considerable generalization ability and applicability; on the other hand, it innovatively utilizes multidimensional gene features to quickly and efficiently predict the dynamic expression pattern of genes in the development of biological tissues or organs. Compared with the cumbersome process of sampling developmental stage samples and performing high-throughput transcriptome sequencing in traditional methods, the present invention greatly reduces the experimental time, shortens the scientific research cycle, and significantly reduces the scientific research cost. Therefore, the present invention fills a technical gap in the current field of gene developmental dynamic expression prediction, and at least partially overcomes the technical problems in the related art that the detection of gene expression patterns in the development process is time-consuming and costly. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 This is a flow chart of a method for predicting developmental dynamic expression patterns based on gene features provided in the disclosed embodiments of the present invention.

[0075] Figure 2 This is a block diagram of a database structure for predicting the dynamic expression pattern of gene development provided by the disclosed embodiment of the present invention.

[0076] Figure 3 The present invention discloses a dynamic judgment threshold division and ROC curve for predicting the developmental dynamic expression pattern of barley genome genes in grains using a cross-species model CSM.

[0077] Figure 4 This is a system block diagram for predicting the dynamic expression pattern of gene development provided by the disclosed embodiment of the present invention.

[0078] Figure 5 This is a diagram of the relevant operational framework of a method for predicting developmental dynamic expression patterns based on gene features provided in the disclosed embodiment of the present invention.

[0079] Figure 6 It is an error loss curve and importance feature ranking diagram in the process of building a CSM model covering six species provided by the disclosed embodiment of the present invention.

[0080] Figure 7 It is the dynamic judgment threshold division and ROC curve in the process of building a CSM model covering six species provided by the disclosed embodiment of the present invention.

[0081] Figure 8It is the error loss curve and importance feature ranking diagram of the single-species model SSM and the cross-species model CSM provided by the disclosed embodiment of the present invention.

[0082] Figure 9 It is the dynamic judgment threshold division and ROC curve in the construction process of the single-species model SSM and the cross-species model CSM provided by the disclosed embodiment of the present invention.

[0083] Figure 10 It is a bar chart of evaluation indicators for the prediction of dynamic expression patterns of gene development by the single-species model SSM and the cross-species model CSM provided in the disclosed embodiments of the present invention. DETAILED DESCRIPTION

[0084] The present invention is further described in detail below by way of examples in conjunction with the accompanying drawings. In the description of the following embodiments, specific details such as specific system structures and technologies are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0085] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their collections. It should also be understood that the term "and / or" used in the present specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0086] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0087] The terms used herein (including technical and scientific terms) have the basic meanings commonly understood by those skilled in the art, unless otherwise specified. It should be noted that the terms used herein should be interpreted as having meanings consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner. In addition, in the description of this application specification and the appended claims, the terms "first," "second," "third," etc. are only used to distinguish descriptions and should not be understood as indicating or implying relative importance. References to "one embodiment" or "some embodiments" in this application specification mean that one or more embodiments of the application include specific features, structures, or characteristics described in conjunction with the embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in other embodiments," etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.

[0088] Example 1

[0089] In related technologies, samples are usually taken from developmental stages and high-throughput transcriptome sequencing is performed to obtain the dynamic expression pattern of gene development. However, this method requires a long experimental time and is extremely costly, making it unsuitable for scientific researchers.

[0090] Table 1 Comparison of prediction tools for gene expression-related tasks

[0091]

[0092] In view of this, the present invention provides a method for predicting developmental dynamic expression patterns based on gene features, the steps of which include determining a gene set of developmental dynamic expression patterns to be detected; obtaining multidimensional gene feature information of the genes to be detected; inputting the multidimensional gene feature information into a developmental dynamic expression pattern prediction model, and outputting a developmental dynamic pattern corresponding to the target gene, wherein the developmental dynamic expression pattern prediction model is trained based on a multidimensional gene feature information data matrix with gene dynamic expression pattern labels, and the gene dynamic expression pattern labels are determined based on the gene expression matrix of the tissue or organ development process of the training sample.

[0093] Figure 1 The flowchart of a method for predicting developmental dynamic expression patterns based on gene features is shown, including steps S1-S3:

[0094] Step S1: Determine the gene set whose developmental dynamic expression pattern is to be detected

[0095] Due to differences in research objectives, research foundations, and experimental designs among researchers, the sources and screening criteria of gene sets may also vary. The following are some common sources of target gene sets:

[0096] (1) All genes in the newly sequenced and assembled genome:

[0097] For a newly sequenced and assembled species genome, the genome annotation information only provides information about its basic sequence and genomic structure, and often lacks understanding of its transcriptional level. This gene set typically contains all annotated coding and non-coding genes and is suitable for comprehensive analysis of the dynamic patterns of gene expression during development of the newly sequenced species. To explain and demonstrate steps S1-S3 in detail, the 32,915 protein coding genes (PCGs) and 4,766 long intergenic noncoding genes (lincRNA genes) of the organized barley (Hordeum vulgare) genome were used as the target gene set for detection, and relevant experiments and tests were performed.

[0098] (2) Differentially expressed gene sets obtained based on high-throughput sequencing:

[0099] High-throughput sequencing technologies such as RNA-seq can be used to obtain gene expression data in samples treated with different agents, growing in different environments, or with different genotypes. Differential expression analysis tools (such as DESeq2 and edgeR) can be used to identify sets of genes with significant differential expression. Further tracking the patterns of these differentially expressed genes during tissue and organ development can help identify key regulatory genes.

[0100] (3) A collection of genes that are specifically expressed in biological tissues or organs:

[0101] Certain genes have specific expression patterns in specific tissues or organs. For example, genes involved in plant seed development may be highly expressed in the endosperm, while genes involved in root development may be highly expressed in the root meristem. Focusing on the expression dynamics of genes with tissue-specific expression during development is one of the key clues to understanding organ development.

[0102] (4) Species-specific (or cross-species) gene sets that exist within (or between) species:

[0103] Comparative genomics can identify species-specific genes or genes conserved across species. While species-specific genes may be associated with unique developmental traits and speciation mechanisms within a particular species, conserved genes across species may play core regulatory roles in development across multiple species. Exploring the expression variation of such gene sets during organ development can help reveal evolutionary commonalities and specificities in developmental mechanisms.

[0104] (5) Candidate gene set compiled through literature research:

[0105] In many cases, researchers can quickly access gene sets of interest by searching relevant literature on websites (e.g., CNKI, Google Scholar, PubMed, etc.). For example, certain genes may have been reported to play an important role in various neurodegenerative diseases. Focusing on the dynamic changes in the expression patterns of these genes during organ development can be helpful in understanding their molecular and biological functions.

[0106] Step S2: Obtain multidimensional gene feature information of the gene to be detected

[0107] S2.1: Obtain the numerical features of genes based on gene sequences and annotations;

[0108] Based on the gene ID, sequence, and annotation information to be tested, the length, GC content, and number of exons of the gene's representative transcript were counted one by one. For transcript length, the gene coordinates were extracted and calculated using software such as bedtools or TBtools based on the genomic sequence fasta file and the genome annotation gtf / gff3 file. For GC content, the number of G and C bases in the transcript sequence was counted using software such as BioPython, EMBOSS, SeqKit, or TBtools, and the GC content was calculated according to the formula GC% = (G+C) / (A+T+G+C).

[0109] Exon number can be calculated based on information in the genome annotation gtf file using software such as Gffread, bedtools, or TBtools. The theoretical value of the length feature is a natural number greater than 1, the GC content feature has a value interval of (0, 1), and the exon feature has a value greater than or equal to 1.

[0110] S2.2: Obtain gene classification characteristics based on gene properties and annotations;

[0111] (1) Protein coding feature acquisition:

[0112] According to the genome annotation gtf / gff3 file, if there is detailed annotation of the gene in the annotation (such as annotation: protein coding gene, lncRNA, rRNA, snoRNA, etc.), the feature information can be used directly.

[0113] If the genome annotation file lacks detailed annotation of gene classes, or if researchers have cloned or obtained genes that are not included in the genome annotation, they can submit their results to relevant websites to determine the coding capacity of the gene and its transcripts. CPAT software (https: / / cpat.readthedocs.io / en / latest / ) uses a logistic regression model and the exon and intron sequences of genomic genes to predict the coding capacity of target gene sequences. CPC2 software (http: / / cpc2.cbi.pku.edu.cn / ) analyzes the Fickett score, ORF length, ORF completeness, and isoelectric point of a sequence to predict the presence of a reading frame and its ability to encode proteins. LGC software (https: / / ngdc.cncb.ac.cn / lgc / ) calculates GC content, ORF length, and other metrics to determine the protein-coding potential of transcripts. Additionally, BLAST searches can be performed against the Pfam database to identify potential protein domains. In summary, a combination of various mainstream software tools can be used to preliminarily determine whether a target gene is coding or non-coding.

[0114] (2) Alternative splicing (AS) feature acquisition:

[0115] According to the genome annotation gtf / gff3 file, if a gene locus has only one transcript, then there is no alternative splicing. However, the vast majority of genes in the genome are prone to alternative splicing events, which usually manifest as exon skipping, exon mutual exclusion, intron retention, alternative promoters, alternative terminators, alternative acceptor sites, and alternative donor sites. Simply put, if a gene locus has multiple corresponding transcripts, it can be considered to have an alternative splicing event. In addition, if you want to obtain alternative splicing events for the entire genome or a large-scale gene collection, you can use software such as SUPPA, rMATS, AStalavista, or ASprofile to batch count and generate results.

[0116] (3) Transposable element (TE) feature acquisition:

[0117] Based on the genomic FASTA sequence, BLAST analysis can be performed against the transposable element database to detect the presence of transposable elements and related fragments within the gene sequence. Alternatively, the genome annotation gtf / gff3 file can be converted to a bed file and analyzed using bedtools along with the genome-wide transposable element annotation of the target species to identify genes associated with transposable elements. If genome-wide transposable element annotation of the target species is not available, de novo transposable element annotation of the genome can be performed using EDTA software (https: / / github.com / oushujun / EDTA).

[0118] (4) Evolutionary conservation feature acquisition:

[0119] The evolutionary conservation of genes can be understood from two perspectives. First, the conserved gene sequence can be determined by the homology of nucleic acid sequences between species; second, the conserved gene position can be determined by the genomic collinearity between species. To simultaneously examine gene family division and genomic collinearity across species, analysis can be performed using software such as OrthoFinder (https: / / github.com / davidemms / OrthoFinder) or OrthoMCL (https: / / orthomcl.org / orthomcl / app / ), as well as MCScanX (https: / / github.com / wyp1125 / MCScanX).

[0120] Note that these categorical features need to be recorded in one-hot encoding mode, that is, if the feature exists, it is 1, and if the feature does not exist, it is 0.

[0121] S2.3: Integrate the numerical features and categorical features of the target gene set into a matrix containing multidimensional gene features.

[0122] The standardized numerical features and the encoded categorical features are integrated into a multidimensional gene feature matrix by row (gene) and column (feature). This matrix serves as the input data for the developmental dynamic expression pattern prediction model. Each row of the matrix represents a gene, and each column represents a feature. For example: rows are gene A, gene B, and gene C; columns are coding genes, non-coding genes, with alternative splicing, and without alternative splicing. The example structure of the multidimensional gene feature matrix is ​​given in Table 2:

[0123] Table 2 Example structure of multidimensional gene feature matrix

[0124]

[0125] The multidimensional gene feature information of the gene to be detected is used as a hub carrier for data processing. A special data structure can be constructed based on the gene feature information ( Figure 2 ) and specifically designed for expandable dimensions based on gene signature information, as well as interactivity or potential interactivity between gene sets and developmental expression pattern prediction models for the developmental dynamics of their expression patterns. The main architecture of the gene signature information database is centered around the GENE_MASTER table, which is then linked to several other tables. This includes core table elements, associated tables, and relationship elements.

[0126] D1: Core table elements:

[0127] GENE_MASTER; built as the core of the entire database structure, used to store basic information of genes, optionally including the basic identification, name, description or other basic gene data of the gene.

[0128] D2: Association tables and relationship elements; including:

[0129] D2.1: NUMERIC_FEATURES table;

[0130] Relationship type setting: GENE_MASTER and NUMERIC_FEATURES are built as a one-to-many relationship, and this relationship is described by "contains";

[0131] Data content settings: Each record in the GENE_MASTER table can correspond to multiple records in the NUMERIC_FEATURES table; the NUMERIC_FEATURES table is set to optionally store some numerical features of genes, including but not limited to: length, GC content, and number of exons; the expansion and interaction structure follows the following data relationship: a gene may have multiple different aspects of numerical features, so a GENE_MASTER record is allowed to contain multiple NUMERIC_FEATURES records.

[0132] D2.2: CATEGORICAL_FEATURES table;

[0133] Relationship type setting: GENE_MASTER and CATEGORICAL_FEATURES are set to a one-to-many "contains" relationship;

[0134] Data content settings: The CATEGORICAL_FEATURES table is used to store the classification features of genes, including optional gene coding capacity, variable splicing status, transposable element status, and evolutionary conservation. The expansion and interaction structure follows the following data relationship: Because a gene may have multiple different classification attributes, the GENE_MASTER record is also set to contain multiple CATEGORICAL_FEATURES records.

[0135] D2.3: EXTENDED_FEATURES table;

[0136] Relationship type setting: GENE_MASTER and EXTENDED_FEATURES are a one-to-many "association" relationship;

[0137] Data content setting: The EXTENDED_FEATURES table optionally stores some extended feature information of genes, including other information besides numerical and categorical features, including chromosome position, gene start coordinates, gene end coordinates, transcription direction, and extended information of gene-specific functional annotations, gene performance information in specific environments, or other information; collaboratively, the extension and interaction structure follows the following data relationship: a gene may be associated with multiple extended features, so a GENE_MASTER record can be associated with multiple EXTENDED_FEATURES records.

[0138] D2.4: ANNOTATION_SOURCES table;

[0139] Relationship type setting: GENE_MASTER and ANNOTATION_SOURCES are set to a one-to-many "reference" relationship;

[0140] Data content settings: The ANNOTATION_SOURCES table stores the sources of gene annotation information; the expansion and interaction structure follows the following data relationship: Since the annotation information of a gene may have multiple different sources, a GENE_MASTER record is also set to reference multiple ANNOTATION_SOURCES records.

[0141] Step S3: Input the multidimensional gene feature information into the developmental dynamic expression pattern prediction model and output the developmental dynamic pattern corresponding to the target gene.

[0142] S3.1: Load the multidimensional gene feature matrix into the developmental dynamic expression pattern prediction model. The model calibrates the input features to ensure that the order of the input features is consistent with the order in which the model was pre-trained.

[0143] After checking the appropriate model (single-species model (SSM) or cross-species model (CSM)) for the target species, use readRDS and xgb.load to load manual_training_process_info.rds and xgboost_model_manual.model, respectively, for subsequent prediction of the developmental dynamic expression patterns of the genes to be tested. xgboost_model_manual.model is the actual pre-trained XGBoost model, which contains the decision tree structure and related parameters. Use read_excel and as.data.table to read and convert the multidimensional gene feature matrix and extract specific feature columns from it. These columns are the features used for model training. Use as.matrix to convert the data into a matrix format so that XGBoost can use DMatrix as input. The model maps the input features to ensure that the order of the input features matches the order of the pre-trained features of the model.

[0144] S3.2: For each gene, after traversing the set of trained decision trees, the model obtains the cumulative score of the prediction values ​​of all trees and converts the cumulative score into a probability value;

[0145] Execute the predict function to predict the data using the trained developmental dynamic expression prediction model, ultimately obtaining a predicted value (user_predictions) for each gene, with a value range of (0, 1). Simply put, the trained model consists of multiple weak decision trees. Each data point is input into multiple decision trees, and each tree generates a prediction score. Ultimately, these scores are weighted and accumulated to form a final prediction value for each gene under test.

[0146] S3.3: Manually set thresholds to determine whether gene expression is developmentally dynamic or not;

[0147] Use the ifelse statement to manually convert the probability value into a category: if user_predictions>t, the prediction is "developmental dynamics"; otherwise, the prediction is "non-developmental dynamics". Among them, t can be set to 0.5 by default, or the optimal threshold can be manually found through the sensitivity and specificity curve (generally, it will not deviate too much from the default 0.5). Figure 3As shown in a, the optimal threshold for the example data of barley genes is 0.54. Therefore, if user_predictions>0.54, it is "developmental dynamic"; otherwise it is "non-developmental dynamic". Finally, among the 37,681 genes, 16,980 were predicted to have dynamic expression patterns during grain development. Through the above steps S1-S3, the genes in the barley genome were used as an example to conduct experiments and complete the task of predicting their developmental dynamic expression patterns in grains. In order to examine the effect of the model, a CSM that has been pre-experimented and trained was used. It should be pointed out that the specified CSM does not contain any information and data related to the species barley. The CSM was trained based on grain development-related data of six species, namely corn, millet, wheat, oats, rice, and sorghum. Combined with the real results of the dynamic expression patterns of genes in the genome from real RNA-seq data during the development of barley grains, the confusion matrix was calculated and the area under the ROC curve (AUC) reached 0.7477, the sensitivity was 0.6774, and the specificity was 0.6760 (e.g. Figure 3 b) Therefore, the prediction method provided by the present disclosure achieves for the first time the rapid, effective, and low-cost prediction of the dynamic expression pattern of gene development, overcoming the limitations of existing technologies.

[0148] Therefore, based on the beneficial effects and databases in the embodiments disclosed in the present invention, the present invention also provides a system for predicting the dynamic expression pattern of gene development (e.g. Figure 4 ):

[0149] M1: Gene set determination module, used to determine the gene set whose developmental dynamic expression pattern is to be detected;

[0150] M2: Gene feature acquisition module, used to obtain multidimensional gene feature information of the gene to be detected;

[0151] M3: a developmental dynamics prediction module, used to input the multi-dimensional gene feature information into a developmental dynamics expression pattern prediction model, and output a developmental dynamics pattern corresponding to the target gene.

[0152] Example 2

[0153] Figure 5 A relevant operational framework diagram of a method for predicting developmental dynamic expression patterns based on gene features provided by an embodiment of the present invention is shown, including operations E1-E5.

[0154] Since the developmental dynamic expression pattern prediction model is obtained by training based on multidimensional gene feature information training samples with gene dynamic expression pattern labels, the gene dynamic expression pattern labels are determined based on the gene expression matrix of the tissue or organ development process of the training sample. Therefore, it is first necessary to explain operations E1-E3 related to the acquisition of training data. Here, the relevant data preparation for the construction of a cross-species model CSM based on data of six species, namely corn, millet, wheat, oats, rice, and mulberry, is used as an example to help understand operations E1-E3 in the present embodiment.

[0155] Operation E1: Obtaining a gene expression matrix representing the developmental process of a species' organs

[0156] E1.1: Obtain transcriptome sequencing data representing the developmental process of tissues or organs of a species. NCBI (National Center for Biotechnology Information) in the United States and CNCB (China National Center for Bioinformation) in China are currently the most important databases for storing the vast majority of second-generation sequencing and third-generation sequencing data. Combining relevant literature reports and database searches, organize and download RNA sequencing (RNA-seq) fastq data for different developmental stages of species tissues or organs. A tissue or organ of a species should contain samples from at least three developmental stages, and biological samples from each stage should contain at least two biological replicates.

[0157] E1.2: Filter low-quality sequencing reads. Preferably, use quality control software such as fastqc and fastx_toolkit to filter sequencing reads and retain high-quality reads in the sequencing data.

[0158] E1.3: Obtain genome data for representative species. NCBI and Ensembl contain the genome sequences and annotation data for most currently sequenced plants, animals, and microorganisms. Fasta genome sequences, gtf genome annotation information, CDS sequences, protein amino acid sequences, and non-coding RNA (ncRNA) sequences are all available for representative species.

[0159] E1.4: Assembling the transcriptome of the representative species. Preferably, the sequencing reads are assembled to the reference genome using bowtie2, Tophat2, and cufflinks software to complete the transcriptome assembly of tissue or organ development.

[0160] E1.5: Quantification of expression of genes and their transcripts. Preferably, in each species, the transcript annotation files of different tissues or organs and different developmental stages are merged, and the expression of genes and their transcripts is quantified using cuffdiff under cufflinks. The count values ​​of genes and their transcripts are normalized using the FPKM (Fragments Per Kilobase of Transcript Per MillionFragments) method. Finally, the gene expression matrix of the seed development process of the six species is obtained. This includes (PCG+lincRNAgene): 39,953 genes in corn, 317,821 genes in millet, 126,972 genes in wheat, 84,987 genes in oats, 53,801 genes in rice, and 32,521 genes in safflower.

[0161] Operation E2: Acquisition of dynamic gene expression pattern labels

[0162] E2.1: Based on a generalized linear model (GLM), maSigPro software can identify gene expression patterns that exhibit significant temporal variations. Therefore, using maSigPro and related functions, we used tissue or organ development expression data as input to perform regression analysis across six species to identify developmentally dynamic genes. Gene expression profiles and sample grouping information were manually formatted according to the software's specified input data format.

[0163] E2.2: Construct a regression design matrix. Create a regression design matrix for fitting a regression model, defining the relationship between the experimental groups and the regression variables.

[0164] E2.3: Identification of significantly differentially expressed genes. Given the characteristics of NGS data, the negative binomial distribution (NB) was selected as the distribution function for the GLM model. Regression analysis was performed on each gene using the p.vector function. The F statistic and p-value of the regression model were calculated, and the Benjamini-Hochberg (BH) correction was used to control the error rate of multiple hypothesis testing.

[0165] E2.4: Identify genes that change significantly during development. Use the T.fit function to perform stepwise regression, select the most significant influencing factors, and output the p-value used to test whether the entire regression model is significant and the R value to measure the goodness of fit of the model. 2 It should be noted that R 2 The size of the R describes the ability of the independent variables in the model to explain the changes in the dependent variables, that is, by setting R 2The threshold value is used to classify whether the gene expression is dynamic or not during the organ development process. 2 The value range of is [0, 1]. Here, 0.5 is set as the threshold for classifying dynamically expressed genes. The results are recorded, and the developmental dynamic gene set is ultimately output and saved. The developmental dynamic genes include: 12,799 in maize, 9,811 in millet, 33,849 in wheat, 16,550 in oats, 11,683 in rice, and 6,874 in safflower.

[0166] Operation E3, acquisition of multidimensional gene feature information of the whole genome

[0167] The details of operation E3 are basically the same as those of step S2 and are not described again here.

[0168] Since the developmental dynamic expression pattern prediction model is trained based on multidimensional gene feature information training samples with gene dynamic expression pattern labels, operations E4 and E5 related to the model training and construction process are described below. The training and construction of a cross-species CSM model based on data from six species—maize, millet, wheat, oats, rice, and safflower—is used as an example to facilitate understanding of operations E4 and E5 in the disclosed embodiments.

[0169] Operation E4: Training Method for Predicting Developmental Dynamic Expression Patterns

[0170] The model training process can be at least Training was performed using an i7-8550U CPU @ 1.80GHz and 8GB of RAM, using an R-based framework. This paper integrates the XGBoost (eXtreme Gradient Boosting) architecture to construct a developmental dynamic expression prediction model. XGBoost implements a machine learning algorithm within the gradient boosting framework, boasting high efficiency, flexibility, and portability. This algorithm has been widely used in many machine learning tasks and has performed exceptionally well, achieving results that even surpass deep learning architectures when processing tabular data.

[0171] E4.1: Correlate the dynamic expression pattern of each gene with its respective gene signature to form a data matrix of multidimensional gene signature information with gene dynamic expression pattern labels;

[0172] E4.2: The matrix set was randomly split into a training set (80% of the total number of genes, 146,950 genes) and a validation set (20% of the total number of genes, 36,740 genes). Five-fold cross-validation was used, and grid search was used for hyperparameter search: nrounds = c(1000), eta = c(0.001, 0.01, 0.1), max_depth = c(3, 5, 7), gamma = c(0.3, 0.5, 0.7), colsample_bytree = c(0.3, 0.5, 0.7), min_child_weight = c(1, 3, 5), and subsample = c(0.3, 0.5, 0.7). The appropriate hyperparameter configuration for the model was determined based on the AUC. Among 729 hyperparameter combinations, with hyperparameters eta=0.01, max_depth=7, gamma=0.5, colsample_bytree=0.7, min_child_weight=5, and subsample=0.7 as input, the AUC value can reach a maximum of 0.7002.

[0173] E4.3: Based on the gene expression patterns during development, map the "non-developmental" label to 0 and the "developmental" label to 1. After extracting features and labels, convert the data into a DMatrix format. Train the model using the optimal hyperparameter configuration.

[0174] E4.4: Use watchlist to monitor the training error and validation error, draw the error loss curve, and observe whether it is overfitting. Use early stopping to determine the optimal number of training rounds. If the loss value does not decrease after 10 consecutive rounds, stop training early. Figure 6 As shown in a, 529 is the optimal number of training rounds. Use xgb.save and saveRDS to save the trained model. The output file xgboost_model_manual.model contains the trained CSM model for the six species, and the manual_training_process_info.rds file contains information about the training of the CSM model for the six species.

[0175] E4.5: Use xgb.importance to assess the contribution of each feature to the developmental dynamics of expression. The gain obtained represents the average loss reduction contributed by a feature when used as a split point at all nodes in all trees. This can be preliminarily understood as the contribution of the feature to the final model's predictive power. Figure 6 b shows the contribution of 11 features to the model, among which conservation, length, and GC content contribute the most to the prediction of developmental dynamic expression patterns.

[0176] E4.6: Input the test data (36,740 genes) into the developmental dynamic expression pattern prediction model xgboost_model_manual.model. The model calibrates the input features to ensure that the order of the input features is consistent with the order in which the model was pre-trained.

[0177] E4.7: For each gene in the test set, the model traverses the set of trained decision trees and, combined with learning rate scaling, obtains the cumulative score of the predictions of all trees.

[0178] E4.8: Convert the cumulative scores to probability values. Use the default threshold (0.5) to determine whether gene expression is developmentally dynamic. Calculate the confusion matrix and the area under the receiver operating characteristic (ROC) curve (AUC). Graph the ROC curve.

[0179] E4.9: Considering that the default threshold may not be the best threshold for dividing gene expression dynamics, the sensitivity and specificity of the model are optimized by calculating the sensitivity and specificity at different thresholds and taking the intersection of the two curves as the optimal threshold. Figure 7 a shows that the optimal threshold should be adjusted to 0.54. Recalculate the area under the ROC curve and draw the optimized ROC curve ( Figure 7 b). A series of indicators obtained by calculating the confusion matrix of the test set, AUC has reached 0.6986, and other indicators also show that the model has good predictive ability ( Figure 7 c).

[0180] Operation E5, strategy for constructing a prediction model for developmental dynamic expression patterns

[0181] E5.1: Integrate the multidimensional gene feature information and developmental dynamic expression patterns of the whole genome genes of multiple species into multiple matrices, and construct multiple single-species models (SSM) through model training. That is to say, only for the data of one species, train and construct the SSM based on the multidimensional gene feature information of the gene dynamic expression pattern label. According to the embodiment of the present disclosure, the SSM is trained based on the grain development-related gene expression data specific to seven species: corn, millet, wheat, barley, oats, rice, and safflower. By executing steps S1-S3 and operations E1-E4 as described above, 7 SSM models are trained and constructed, and the error loss curve and importance feature ranking diagram are shown in FIG. Figure 8 As shown in the figure, the dynamic judgment threshold division and ROC curve in the SSM construction process are shown in the figure. Figure 9 shown. Figure 10 The experimental results shown in Figure 1 show that the seven SSM models can perform well in predicting the dynamic expression patterns of gene development in their respective species. For example, the AUC in rice can even exceed 0.8.

[0182] E5.2: Integrate the multidimensional gene feature information and developmental dynamic expression patterns of the whole genome genes of multiple species into a matrix, and construct a cross-species model (CSM) through model training. For mixed data of multiple species, train and construct the CSM based on the multidimensional gene feature information of the gene dynamic expression pattern labels. By executing steps S1-S3 and operations E1-E4 as described above, the CSM model is trained and constructed. The error loss curve and importance feature ranking diagram are shown in the figure below. Figure 8 As shown in the figure, the dynamic judgment threshold division and ROC curve in the CSM construction process are shown in the figure. Figure 9 The experimental results show that the CSM constructed based on seven species, namely corn, millet, wheat, barley, oats, rice, and safflower, has a strong generalization ability and can also effectively predict the developmental dynamic expression pattern of genes in different species ( Figure 10 b). As previously mentioned, the CSM trained on gene expression data related to grain development in six species, namely maize, millet, wheat, oats, rice, and sorghum, has effectively demonstrated its good applicability to new and unknown data (barley). Therefore, for plant species not covered by this disclosure, the CSM can be directly used to predict the developmental dynamic expression patterns of the genes to be tested in the grains of the corresponding species. In summary, the present invention significantly expands the range of species involved in the training data of gene expression-related models, and the large amount of experimental data presented in the examples shows that the present invention has considerable generalization ability and applicability.

[0183] Example 3, code and its description

[0184] The following is a method for predicting developmental dynamic expression patterns based on gene features disclosed in the present invention, and its code corresponds to steps S1-S3:

[0185]

[0186]

[0187] The following is a training method for a developmental dynamic expression pattern prediction model based on gene features disclosed in the present invention, and its code corresponds to the E4 operation:

[0188]

[0189]

[0190]

[0191]

[0192]

[0193]

[0194]

[0195]

[0196] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0197] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A method for predicting developmental dynamic expression patterns based on gene characteristics, characterized in that: The following steps are involved: S1. Determine the gene set whose developmental dynamic expression pattern is to be detected; S2. Obtaining multidimensional gene feature information of the gene set to be detected; S3. Input the multidimensional gene feature information into the developmental dynamic expression pattern prediction model, and output the developmental dynamic pattern corresponding to the target gene, wherein the developmental dynamic expression pattern prediction model is trained based on the multidimensional gene feature information data matrix with gene dynamic expression pattern labels, and the gene dynamic expression pattern labels are determined based on the gene expression matrix of the tissue or organ development process of the training sample.

2. A method for predicting developmental dynamic expression patterns based on gene features according to claim 1, characterized in that: The gene set in step S1 is any one or any combination of the following: (1) The complete set of genes in the newly sequenced and assembled genome; (2) Differentially expressed gene sets obtained based on high-throughput sequencing; (3) A collection of genes that are specifically expressed in biological tissues or organs; (4) Species-specific or cross-species gene sets that exist within or between species; (5) A set of candidate genes compiled through literature research; Step S2 specifically includes: S2.1: Obtain the numerical features of the gene set to be tested based on the gene sequence and annotation; S2.2: Obtain the classification characteristics of the gene set to be tested based on gene properties and annotations; S2.3: Integrate the numerical features and categorical features of the gene set to be tested into a data matrix containing multidimensional gene feature information.

3. The method for predicting developmental dynamic expression patterns based on gene features according to claim 1, wherein: Step S3 specifically includes: S3.1: Load the data matrix containing multidimensional gene feature information into the developmental dynamic expression pattern prediction model. The developmental dynamic expression pattern prediction model performs mapping calibration on the input features to ensure that the order of the input features is consistent with the order pre-trained by the developmental dynamic expression pattern prediction model. S3.2: For each gene, the developmental dynamic expression pattern prediction model traverses the trained decision tree set, obtains the cumulative score of the prediction values ​​of all trees, and converts the cumulative score into a probability value; S3.3: Manually set thresholds to determine developmental dynamics of gene expression.

4. The method for predicting developmental dynamic expression patterns based on gene features according to claim 1, wherein: The operations involved in constructing the developmental dynamic expression pattern prediction model include: E1: Obtaining the gene expression matrix representing the species' organ development process; E2: Acquisition of dynamic gene expression pattern signatures during development; E3: Acquisition of multidimensional gene feature information of the whole genome; E4: Training methods for prediction models of developmental dynamic expression patterns; E5: Strategies for constructing prediction models for developmental dynamic expression patterns.

5. A method for predicting developmental dynamic expression patterns based on gene features according to claim 4, characterized in that: Operation E1 includes: E1.1: Obtain transcriptome sequencing data representing the developmental process of tissues or organs of a species; E1.2: Filter low-quality sequencing data and retain high-quality sequencing data; E1.3: Obtain genomic data for representative species; E1.4: Based on genomic data, assemble transcriptomes of representative species and complete transcriptome assembly for tissue or organ development; E1.5: Quantitatively analyze the expression of genes and their transcripts to generate a gene expression matrix for organ developmental processes.

6. A method for predicting developmental dynamic expression patterns based on gene features according to claim 5, characterized in that: Operation E4 includes: E4.1: Correlate the dynamic expression pattern of each gene with its respective gene signature to form a data matrix of multidimensional gene signature information with gene dynamic expression pattern labels; E4.2: Randomly split the data matrix into a training set and a validation set, with the training set accounting for 80% and the validation set accounting for 20%. Use grid search to perform hyperparameter tuning to obtain the optimal hyperparameter configuration under the current conditions. E4.3: Map "non-developmental" and "developmental" genes into two labels based on their expression patterns during development, and train the model under the optimal hyperparameter configuration. E4.4: Monitor training and validation errors, plot error loss curves, use early stopping to determine the optimal number of training rounds to avoid overfitting, and save the trained prediction model. E4.5: Assess the contribution of each trait to the developmental dynamic expression pattern; E4.6: Input the test data set into the developmental dynamic expression pattern prediction model. The model will calibrate the input features to ensure that the order of the input features is consistent with the order in which the model was pre-trained. E4.7: For each gene in the test set, the model traverses the set of trained decision trees and, combined with learning rate scaling, obtains the cumulative score of the predictions of all trees. E4.8: Convert the cumulative scores into probability values, use the default threshold of 0.5 to determine whether gene expression is developmental dynamic, calculate the confusion matrix and the area under the receiver operating characteristic (ROC) curve (AUC), and plot the ROC curve. E4.9: Considering that the default threshold is not the best threshold for demarcating gene expression dynamics, optimize the sensitivity and specificity of the model by plotting sensitivity and specificity curves and using the intersection of the two as the optimal threshold. Recalculate the confusion matrix and the area under the ROC curve, and plot the optimized ROC curve.

7. A method for predicting developmental dynamic expression patterns based on gene features according to claim 6, characterized in that: Operation E5 includes: E5.1: Integrate the multidimensional gene feature information and developmental dynamic expression patterns of whole-genome genes of multiple species into multiple matrices, and construct multiple single-species SSM models through model training; E5.2: Integrate the multidimensional gene feature information and developmental dynamic expression patterns of whole-genome genes of multiple species into a matrix, and construct a cross-species model CSM through model training.

8. A developmental dynamic expression pattern prediction model, characterized by: The model is constructed by the method described in any one of claims 4 to 7.

9. A system for predicting the dynamic expression pattern of gene development, characterized in that: include: M1: Gene set determination module, used to determine the gene set whose developmental dynamic expression pattern is to be detected; M2: Gene feature acquisition module, used to obtain multidimensional gene feature information of the gene to be detected; M3: a developmental dynamics prediction module, configured to input the multidimensional gene feature information into a developmental dynamics expression pattern prediction model and output a developmental dynamics pattern corresponding to the target gene; Wherein, the developmental dynamic expression pattern prediction model is trained and constructed based on the above-mentioned claim 4.

10. An interactive and scalable gene signature information database structure for predicting developmental dynamic expression patterns, characterized by: The database is constructed based on gene feature information and is specifically designed with expandable dimensions for gene feature information, as well as interactivity or potential interactivity between gene sets and developmental dynamic expression pattern prediction models for the developmental dynamic expression patterns to be detected; The main structure of the gene feature information database is set as follows: the database structure revolves around the GENE_MASTER table, and at the same time, it is set to have specific associations with several other tables, including: D1: Core table element: GENE_MASTER; built as the core of the entire database structure, used to store basic information about genes, optionally including the basic gene identifier, name, description, or other basic gene data; D2: Association tables and relationship elements; including: D2.1: NUMERIC_FEATURES table; Relationship type setting: Build a one-to-many relationship between GENE_MASTER and NUMERIC_FEATURES, using "contains" to describe this relationship; Data content setting: Each record in the GENE_MASTER table can correspond to multiple records in the NUMERIC_FEATURES table; NUMERIC_FEATURES table is set to store optional numerical features of genes; D2.2: CATEGORICAL_FEATURES table; Relationship type setting: GENE_MASTER and CATEGORICAL_FEATURES are set to a one-to-many "contains" relationship; Data content setting: The CATEGORICAL_FEATURES table is used to store the classification features of genes; D2.3: EXTENDED_FEATURES table; Relationship type setting: GENE_MASTER and EXTENDED_FEATURES are in a one-to-many "association" relationship; Data content setting: The EXTENDED_FEATURES table optionally stores extended feature information of genes, including information other than numerical and categorical features; D2.4: ANNOTATION_SOURCES table; Relationship type setting: GENE_MASTER and ANNOTATION_SOURCES are set to a one-to-many "reference" relationship; Data content setting: ANNOTATION_SOURCES table stores the sources of gene annotation information; The extension and interaction structure of D2.1-2.4 follows the following data relationship: Since the annotation information of a gene may come from multiple different sources, a GENE_MASTER record is also set to allow reference to multiple NUMERIC_FEATURE, CATEGORICAL_FEATURE, EXTENDED_FEATURES, and ANNOTATION_SOURCES records respectively.

Citation Information

Patent Citations

  • Cancer type prediction system and method based on tissue and organ differentiation hierarchical relationship

    CN110706749A

  • Block texture morphological gene identification method based on figure atlas

    CN118094386A

  • Gene expression level prediction method based on multi-modal data fusion

    CN118447923A

  • Prediction system and electronic equipment for congenital heart disease and delayed intelligence development

    CN119069003A

  • Single cell transcriptome computation and analysis method and system incorporating deep learning model

    WO2022188785A1