Plant flavonoid pathway mining and database construction method and system based on protein language model and multiple omics

By integrating the ESM-2 protein language model and multi-omics data, an extended flavonoid metabolism network map was constructed, which solved the problems of insufficient gene identification and database annotation in existing flavonoid pathway research, and achieved efficient and accurate gene function mining and breeding support.

CN122067604APending Publication Date: 2026-05-19GUANGXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGXI UNIV
Filing Date
2026-02-09
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies in flavonoid pathway research suffer from problems such as time-consuming and laborious gene identification, missed detections in sequence homology alignment, incomplete database annotation, and lack of multi-omics data integration, resulting in insufficient reliability of functional prediction and difficulty in supporting plant metabolic engineering and breeding.

Method used

By employing the ESM-2 protein language model combined with multi-omics data, an extended flavonoid metabolism network map was constructed through sequence homology alignment, deep learning models, and three-dimensional structural modeling. This map was then integrated with a comprehensive database platform to achieve efficient and accurate gene function mining and visualization analysis.

Benefits of technology

It improves the accuracy and coverage of candidate gene identification, enhances the completeness of pathway annotation, simplifies data retrieval and analysis processes, supports plant metabolic engineering and breeding of high-flavonoid crops, and promotes the shift of research towards data-driven approaches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122067604A_ABST
    Figure CN122067604A_ABST
Patent Text Reader

Abstract

The invention discloses a plant flavonoid pathway mining and database construction method and system based on a protein language model and multiple omics. The method comprises the following steps: acquiring genomes, proteomes and annotation data of plants of multiple species, constructing a standardized flavonoid resource library, extracting high-dimensional sequence features by utilizing an ESM-2 protein language model, performing accurate flavonoid enzyme prediction in combination with three-dimensional structure modeling and molecular docking, and constructing an extended flavonoid metabolism network map. And a database platform is integrated based on a front-end and back-end separation architecture. According to the invention, the accuracy and coverage range of flavonoid related gene identification are obviously improved; compared with a traditional annotation method based on sequence similarity comparison, the method has better functional distinguishing capacity under the low homologous background, and generalization annotation can be reduced; complementing modification reaction nodes deleted by the metabolic pathway, and enhancing the integrity of pathway annotation; total factor integration of multi-omics data and a metabolic network is realized, and interactive pathway visualization and function prediction are supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of bioinformatics and plant metabolic engineering, and in particular to a method and system for mining and constructing a database of plant flavonoid pathways based on protein language models and multi-omics. Background Technology

[0002] Flavonoids, as widely distributed secondary metabolites in plants, play a crucial role in physiological processes such as plant growth and development, stress resistance, disease and pest resistance, and pollinator attraction. They also possess significant antioxidant, anti-inflammatory, and cardiovascular protective effects on human health. Therefore, in-depth analysis of the biosynthetic pathways and regulatory mechanisms of plant flavonoids is of great value for crop quality improvement, functional food development, and drug research.

[0003] Currently, research on flavonoid pathways mainly relies on experimental validation and traditional bioinformatics methods. At the experimental level, identifying key enzyme genes through gene knockout and enzyme activity assays is time-consuming and laborious, making it difficult to cover the broad research needs of multiple plant species. In terms of computational methods, existing studies mostly rely on sequence homology alignment (such as BLAST) to screen for flavonoid-related candidate genes across the entire genome. However, this method has significant limitations: on the one hand, homologous enzymes with low sequence similarity (such as functionally conserved but sequence-differentiated modifying enzymes) are easily missed; on the other hand, it lacks the ability to accurately predict enzyme function and cannot effectively distinguish enzyme subtypes that catalyze the same reaction but have different substrate specificities. Furthermore, while mainstream metabolic pathway databases (such as KEGG and MetaCyc) provide a basic framework for flavonoid pathways, their annotations are significantly lacking: for example, the integration of key modification reactions such as flavonoid glycosylation, methylation, and acylation is still insufficient; related information is not fully collected and updated; and multi-omics data and pathway network information are often stored in a scattered manner, lacking unified standards for correlation and evidence.

[0004] In recent years, artificial intelligence technologies, represented by deep learning, have made rapid progress in the field of bioinformatics, becoming a major driving force for gene function identification and annotation. However, existing research largely focuses on using AI technology to classify the functions of single genes or proteins. Although protein language models (such as the ESM series) have shown breakthrough progress in capturing deep sequence features through large-scale pre-training, their application in plant metabolic pathway mining has not yet been systematically developed. Existing research is mostly limited to single sequence feature analysis, without combining three-dimensional structural modeling and molecular docking verification, resulting in insufficient reliability of functional predictions. At the same time, the lack of a comprehensive database platform that integrates multi-omics data, supports interactive pathway visualization, and online functional prediction makes it difficult for researchers to efficiently acquire, analyze, and verify the functions of flavonoid-related genes. These problems severely restrict the comprehensive analysis and application development of flavonoid pathways.

[0005] Therefore, there is an urgent need for an innovative approach that integrates protein language models, multi-omics data, and structural biology analysis to overcome the shortcomings of existing technologies in terms of pathway discovery efficiency, functional annotation accuracy, and database integration, and to provide technical support for plant metabolic engineering and precision breeding. Summary of the Invention

[0006] This invention addresses the shortcomings of existing technologies by providing a method and system for mining and constructing a database of plant flavonoid pathways based on protein language models and multi-omics. It focuses on using artificial intelligence to analyze plant secondary metabolic pathways, integrating genomic, proteomic, and functional annotation data to achieve precise mining and visual analysis of genes related to flavonoid biosynthesis, modification, and regulation. This provides efficient and systematic computational tools and data support for plant metabolic engineering research.

[0007] To achieve the above-mentioned objectives, the technical solution adopted by the present invention is as follows:

[0008] A method for mining and constructing a database of plant flavonoid pathways based on protein language models and multi-omics, the method comprising the following steps:

[0009] S1: Obtain genome, proteome, and annotation data of multiple plant species, perform redundancy removal and format standardization preprocessing, and identify candidate genes related to flavonoid biosynthesis, transport, and regulation across the entire genome based on sequence homology alignment algorithms, and construct a standardized flavonoid enzyme homology sequence library.

[0010] S2: Flavonoid-modifying enzyme protein sequences from the homologous sequence library are input into the pre-downloaded and deployed ESM-2 protein language model to extract high-dimensional representation features of the protein sequences; based on the obtained sequence embedding features, a fully connected neural network classification model is constructed and trained to predict the functional category of the protein sequences; for candidate sequences with positive prediction results, their three-dimensional structural models are further constructed, and their potential binding ability with substrates is verified by molecular docking analysis, thereby forming a comprehensive annotation result that integrates sequence features, structural features, and functional information;

[0011] S3: Based on the KEGG flavonoid pathway, supplement the information on metabolic reaction nodes and corresponding modifying enzymes by integrating literature mining data, construct an extended flavonoid metabolism network map, and map the annotation data to the corresponding nodes of the network map;

[0012] S4: Based on a front-end and back-end separation architecture, a comprehensive database platform is built to store the genomic and proteomic data of the multi-species plants, the annotation data, and the extended flavonoid metabolism network map into the database, and integrate online bioinformatics tools to provide data retrieval, pathway visualization, and functional prediction services.

[0013] Furthermore, in step S1, the multi-species plants include algae, mosses, ferns, gymnosperms, and angiosperms; the preprocessing includes removing redundant sequences, standardizing file formats, and correcting erroneous annotations.

[0014] Further, in step S2, the specific process of inputting the flavonoid-modifying enzyme sequence into the pre-trained ESM protein language model for feature extraction includes: encoding the modifying enzyme protein sequence using the ESM-2 protein language model to extract a high-dimensional embedding representation of the protein sequence; generating a sequence representation vector with a dimension of 1280 by pooling the representation vectors of each amino acid residue in the sequence; constructing a multilayer perceptron classification model containing an input layer, a hidden layer, and an output layer, and inputting the sequence representation vector into the multilayer perceptron classification model for calculation, outputting the prediction result of the target enzyme functional category to which the protein sequence belongs; when the prediction result meets the preset judgment conditions, the corresponding protein sequence is used as a candidate functional enzyme sequence for subsequent three-dimensional structure modeling and molecular docking analysis to assist in its functional annotation.

[0015] Furthermore, in step S3, the construction of the extended flavonoid metabolism network map includes: identifying and supplementing glycosylation, methylation, and acylation modification reactions not included in the existing KEGG pathway based on literature mining data; establishing a dynamic association index between enzyme sequence IDs and pathway nodes, configured to respond to user clicks and display the corresponding enzyme sequence list and functional prediction results.

[0016] Furthermore, in step S4, the integrated database platform specifically includes:

[0017] The backend architecture uses the Django framework to integrate the ESM prediction model interface and the BLAST comparison tool, and uses Celery and Redis to implement asynchronous scheduling of computing tasks.

[0018] The front-end interactive interface uses the Vue.js framework combined with the Mol* plugin and the D3.js library to render the three-dimensional structure of proteins and interactively display metabolic pathways.

[0019] The data storage module uses a MySQL database to store structured data and a file system to store genome sequence files and three-dimensional structure files.

[0020] This invention also discloses a plant flavonoid pathway mining and database system based on protein language models and multi-omics, used to perform the above method, the system comprising:

[0021] The data mining module is used to acquire genome, proteome and annotation file data of multiple plant species, perform redundancy removal and format standardization preprocessing, and identify candidate genes related to flavonoid biosynthesis, transport and regulation in the whole genome based on sequence homology comparison algorithm, and construct a standardized flavonoid enzyme homology sequence library.

[0022] The intelligent annotation module is used to input flavonoid-modifying enzyme protein sequences into the ESM-2 protein language model to extract high-dimensional representation features, construct and train a fully connected neural network classification model to predict functional categories, and construct three-dimensional structural models for candidate sequences that are predicted to be positive. Combined with molecular docking analysis, it verifies their potential binding ability with substrates and forms a comprehensive annotation result that integrates sequence features, structural features and functional information.

[0023] The pathway reconstruction module is used to construct an extended flavonoid metabolism network map based on the KEGG flavonoid pathway framework, integrate literature mining data to supplement glycosylation, methylation and acylation modification reactions, and map the annotation data to the corresponding nodes of the network map.

[0024] The platform service module is used to build a database platform based on a front-end and back-end separation architecture. It stores genomic and proteomic data, annotation data, and extended flavonoid metabolism network maps of multiple plant species into the database, and integrates online bioinformatics tools to provide data retrieval, pathway visualization, and functional prediction services.

[0025] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the above-described method.

[0026] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method.

[0027] Compared with the prior art, the advantages of the present invention are as follows:

[0028] 1. The ESM-2 protein language model is used to extract high-dimensional sequence features, which overcomes the shortcomings of traditional sequence homology alignment methods in missing low-similarity homologous enzymes (such as functionally conserved but sequence-differentiated modified enzymes). This significantly improves the accuracy and coverage of candidate gene identification, and can accurately distinguish enzyme subtypes with different substrate specificities, providing a more reliable basis for functional annotation.

[0029] 2. The system supplements the missing key modification reaction nodes such as glycosylation, methylation and acylation in the KEGG pathway, and constructs an extended flavonoid metabolism network map covering multiple species. This significantly improves the completeness of pathway annotation and cross-species comparative analysis capabilities, and solves the problem of fragmented pathway information in existing databases.

[0030] 3. The constructed comprehensive database platform integrates all elements of genome, proteome, functional annotation and metabolic network, supports interactive pathway visualization and real-time functional prediction, enables users to intuitively associate enzyme sequences with functional information, greatly simplifies data retrieval and analysis processes and improves research efficiency.

[0031] 4. This invention provides precise functional targets for plant metabolic engineering, directly supporting the breeding of high-flavonoid crops and the development of functional foods, promoting the paradigm shift of flavonoid-related research from traditional experiment-driven to data-driven, and accelerating the transformation and application of basic research results to the agricultural biotechnology industry. Attached Figure Description

[0032] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 This is a schematic diagram of the application environment and computer equipment hardware structure of the method of the present invention in Embodiment 1 of the present invention;

[0034] Figure 2 This is a schematic diagram of the overall process of the plant flavonoid pathway mining and database construction method based on protein language model and multi-omics in Embodiment 2 of the present invention;

[0035] Figure 3 This is a flowchart illustrating the intelligent functional annotation and structural verification based on a protein language model in Embodiment 3 of the present invention.

[0036] Figure 4 This is a schematic diagram illustrating the construction logic of the extended flavonoid metabolic network map in Example 4 of the present invention.

[0037] Figure 5 This is a functional module architecture diagram of the integrated database system in Embodiment 5 of the present invention. Detailed Implementation

[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0039] Example 1: Application Environment and Hardware Architecture The method for mining and constructing plant flavonoid pathways based on protein language models and multi-omics provided in this example can be applied to, for example... Figure 1 The computer system environment shown.

[0040] like Figure 1 As shown, the plant flavonoid pathway mining and database construction method based on protein language models and multi-omics provided in this application can be applied in a computer system environment. This application environment mainly includes a server and terminal devices, which communicate via a network.

[0041] The server is used to deploy the comprehensive database system described in this application and is configured to perform backend business logic processing, database management, and AI model inference tasks. In one embodiment, the server can be a standalone physical server, a server cluster consisting of multiple physical servers, or a cloud computing service center. The terminal device is used to provide researchers with an access interface to the web database platform, responding to user operations by submitting sequence analysis requests and displaying visualization results. The terminal device can be a smartphone, tablet, laptop, or desktop computer, but is not limited to these.

[0042] Taking the server as an example, the internal structure diagram of the computer device includes a processor, memory (including non-volatile storage media and internal memory), network interface and database connected through a system bus.

[0043] The processor, which provides computational and control capabilities, is a core component of the computer device. In one embodiment, to accelerate deep learning inference computations such as ESM protein language model feature extraction and Alpha Fold structure prediction, the processor may integrate a high-performance graphics processing unit (GPU) or a tensor processor (TPU).

[0044] The non-volatile storage medium stores an operating system, computer programs, and a database system. When executed by a processor, the computer program implements a method for mining and constructing a database of plant flavonoid pathways based on a protein language model and multi-omics. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium.

[0045] The database is used to store massive amounts of multi-omics data. Specifically, it includes, but is not limited to, genomic data of thousands of plant species, standardized flavonoid metabolic pathway maps, and pre-trained deep learning model parameter files.

[0046] The network interface is used to connect to external terminal devices, and is configured to respond to user HTTP / HTTPS requests and transmit business data, including JSON data format or visual charts.

[0047] Those skilled in the art will understand that Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0048] Example 2: A Method and Flowchart for Plant Flavonoid Pathway Mining and Database Construction Based on Protein Language Model and Multi-omics

[0049] like Figure 2 As shown, a method for mining and constructing a database of plant flavonoid pathways based on protein language models and multi-omics is provided. This method includes the following steps:

[0050] Step S1: Collect sequence data from multiple plant species and construct a standardized homologous sequence library of flavonoid-related enzymes.

[0051] In this step, to construct a high-quality candidate sequence resource library, plant genome, proteome, and annotation file data covering multiple phylogenetic groups, including algae, bryophytes, ferns, gymnosperms, and angiosperms, are first collected from public databases such as NCBI and Phytozome. An automated preprocessing workflow is established to remove redundancy and standardize the format of the raw data, discarding low-quality data files with a data volume below a preset threshold, while also standardizing file naming and storage specifications.

[0052] Subsequently, a seed sequence-based homology sequence screening strategy was adopted to obtain key enzyme protein sequences of known functions in the flavonoid pathway from databases such as UniProt as seed sequences. Addressing the issues of large-scale multi-species protein sequence data and low efficiency of traditional alignment methods, this embodiment preferentially uses the MMseqs2 high-performance sequence alignment tool. A protein sequence index database is constructed using the mmseqscreatedb module, and a high-sensitivity homology sequence search is performed across the entire genome using the mmseqs search module. Combined with the corresponding GFF3 annotation files and CDS sequence information, and integrating gene chromosomal location and coding region information, a standardized homology sequence library of flavonoid-related enzymes is finally constructed.

[0053] Step S2: Sequence feature extraction is performed based on the ESM protein language model, and functional screening is carried out using a deep learning model.

[0054] This step aims to improve the accuracy of predicting the function of flavonoid-modifying enzymes in non-model plants. Specifically, a pre-trained ESM2 protein language model is used as a sequence feature extractor. The flavonoid-related enzyme protein sequences obtained in step S1 are input into the model for encoding to capture the evolutionary information and functionally relevant features hidden in the sequences. A sequence representation vector with a dimension of 1280 is obtained by pooling the amino acid residue representation vector. Based on this sequence representation vector, a multilayer perceptron classification model is further constructed and trained to perform functional screening of candidate protein sequences, obtaining candidate modifying enzyme sequences with target functional features.

[0055] For the candidate modified enzyme sequences obtained from screening, three-dimensional structural modeling and molecular docking analysis can be used to evaluate their potential binding ability with flavonoid substrates, thereby providing structural-level auxiliary evidence for subsequent functional annotation and database construction.

[0056] Subsequently, a multilayer perceptron (MLP) classifier is constructed and trained. The input feature vector is used to output the probability that the target sequence belongs to a specific functional category, such as glycosyltransferase or methyltransferase. For positive sequences with a predicted probability higher than a threshold (e.g., 0.85), the system further utilizes the Alpha Fold prediction engine to generate its 3D structure file. Finally, using the AutoDockVina engine, a preset flavonoid substrate is molecularly docked with the generated protein structure, and the binding energy is calculated. If the binding energy meets preset conditions (e.g., below -7.5 kcal / mol) and the binding site is located within a conserved pocket), the enzyme is determined to have a high-confidence catalytic function.

[0057] Step S3: Integrate literature data to supplement metabolic response nodes, reconstruct the extended flavonoid metabolic network map, and map annotation information.

[0058] In this step, using KEGG map 00941 as the basic framework, we systematically reviewed literature on flavonoid metabolism published in the past decade to extract newly discovered metabolic reaction steps and modifying enzyme information. These new reaction nodes were then added to the KEGG framework to construct an expanded flavonoid metabolism network map incorporating the latest research findings, thereby achieving a link from macroscopic pathways to microscopic genes.

[0059] Step S4: Build a comprehensive database based on a front-end and back-end separation architecture, and integrate multi-omics data to provide retrieval, visualization and online prediction services.

[0060] In this step, a comprehensive web platform was built to make the aforementioned data and models available to researchers. The system adopts a front-end / back-end separation architecture. The back-end, based on the Django framework, handles business logic and algorithm scheduling, while the front-end, based on the Vue.js framework, builds a responsive interactive interface. The system integrates multi-omics data management modules, machine learning prediction modules, and visualization components. It can respond to user requests, providing services such as gene data retrieval and download, real-time rendering of protein 3D structures via the Mol* plugin, interactive viewing of metabolic pathways, and online sequence function prediction services.

[0061] Example 3: Functional annotation and structure-assisted analysis model architecture based on protein language model

[0062] This embodiment focuses on illustrating the implementation principle and model structure of the core algorithm module for functional annotation based on a protein language model in this invention. For example... Figure 3 As shown, the method generally includes a sequence input module, a protein language model feature extraction module, a deep learning classification module, and a structure analysis auxiliary module.

[0063] A sequence feature extraction module was constructed. To extract high-level biological features related to protein function from amino acid sequences, the system uses a pre-trained ESM2 (Evolutionary Scale Modeling) protein language model as the feature extractor. The ESM2 model is built on the Transformer deep learning architecture, and its core is a multi-head self-attention mechanism. This mechanism can dynamically model the interrelationships between different amino acid residues during sequence modeling, thereby capturing potential long-range dependencies and evolutionary information in protein sequences.

[0064] In this embodiment, the computing device first standardizes the input protein amino acid sequence and converts it into a label sequence recognizable by the model. This label sequence is then input into the ESM2 model (version esm2_t33_650M_UR50D) containing 33 Transformer coding layers. During layer-by-layer computation, the model forms a multi-level representation from low-level physicochemical features to high-level semantic features. Finally, the hidden state output of the last Transformer layer is extracted, and the residue representation vectors along the sequence dimension are pooled to map the variable-length protein sequence into a fixed-length 1280-dimensional sequence representation vector. This sequence representation vector comprehensively reflects the evolutionary background and potential functional characteristics of the protein sequence.

[0065] like Figure 3As shown, after obtaining the sequence representation vector, a deep learning classification model based on a multilayer perceptron (MLP) is further constructed for functional screening of protein sequences. The classification model includes an input layer, at least one hidden layer, and an output layer. The input layer receives a 1280-dimensional feature vector from the ESM-2 model; the hidden layer adopts a fully connected structure and introduces a non-linear activation function to enhance the model's ability to express complex feature patterns; a dropout mechanism is introduced in the hidden layer to reduce the risk of overfitting during training and improve its generalization ability on non-model plant data.

[0066] The output layer contains a single neuron used to determine whether the input sequence possesses the functional characteristics of the target flavonoid-modifying enzyme. During model training, a binary cross-entropy loss function is used to constrain the prediction results against the ground truth labels, and a gradient descent-based optimization algorithm is used to iteratively update the model parameters, thereby obtaining a stable functional screening model.

[0067] In this embodiment, for high-confidence candidate sequences obtained through the classification model, further structural evaluation of their potential binding ability to flavonoid substrates can be conducted by combining protein three-dimensional structure modeling and molecular docking analysis. By introducing structural information, supplementary evidence can be provided for subsequent functional annotation and database construction based on the sequence-level prediction results.

[0068] For positive sequences whose classifier prediction scores exceed a preset threshold (e.g., 0.85), the system automatically invokes the AlphaFold inference engine. This engine does not rely on homologous templates but instead uses multiple sequence alignment (MSA) information and paired features of the sequence for iterative inference via the Evoformer module. It predicts the rotation angles and spatial coordinates of amino acid residues in three-dimensional space, ultimately generating a PDB format file containing precise atomic coordinate information. This process achieves a dimensionality upgrade from "one-dimensional sequence information" to "three-dimensional conformational information."

[0069] The system incorporates a pre-processed library of flavonoid substrates. Utilizing the AutoDock Vina engine, a gradient-based local search algorithm is employed to perform conformational searches within the predicted active pocket regions of the protein structure. The system calculates the optimal binding energy between the substrate and enzyme by comprehensively considering a scoring function that integrates physicochemical terms such as Gaussian spatial interactions, van der Waals forces, hydrogen bonds, and hydrophobic interactions. If the binding energy is below a specific threshold (e.g., -7.0 kcal / mol) and the binding site is located in the neighborhood of a conserved catalytic residue, the predicted result is deemed to possess biological activity.

[0070] In this embodiment, the system front end integrates a high-performance visualization component based on WebGL technology to intuitively display the calculation results to the user;

[0071] The system supports rendering protein skeletons in cartoon mode, clearly displaying secondary structural elements such as α-helices and β-sheets; it also supports rendering the van der Waals surface of proteins in surface mode.

[0072] For molecular docking results, the system not only displays the static position, but also calculates and visualizes the intermolecular force network in real time through algorithms. The docked flavonoid substrate molecules are stacked in a rod-shaped model in the protein binding pocket, and key amino acid residues within a 5 Å range around the substrate are automatically identified and their side chains are highlighted.

[0073] The system automatically calculates the non-covalent interactions between the substrate and enzyme residues, uses dashed lines of different colors to draw hydrogen bonds, π-π stacking, and hydrophobic contacts, and marks the distance values ​​between atoms next to the dashed lines, thereby quantifying and analyzing the enzyme's specific recognition mechanism for the substrate.

[0074] In this embodiment, to improve research efficiency, the system implements synchronous interaction between "one-dimensional sequence" and "three-dimensional structure." The bottom of the interface displays the linear amino acid sequence of the enzyme, while the middle displays the three-dimensional structure. When the user hovers the mouse over or clicks on a specific residue on the linear sequence, the corresponding spatial residue in the three-dimensional view is automatically highlighted and centered; conversely, clicking on a residue on the three-dimensional structure will automatically locate the sequence view, achieving an intuitive association from genotype to phenotype.

[0075] Example 4: Construction Logic of Extended Flavonoid Metabolic Network Map

[0076] like Figure 4 As shown, an extended flavonoid metabolism network map is provided. This method aims to solve the problem of outdated pathway information in existing public databases by constructing a dedicated plant flavonoid metabolism network through knowledge reconstruction, topology expansion and dynamic mapping.

[0077] The method first includes basic framework extraction and knowledge base cleaning steps. Specifically, using the flavonoid biosynthesis pathway (Map00941) in the KEGG database as the basic framework, the KGML (KEGG MarkupLanguage) file is extracted using a Python script, and the reaction nodes, compound nodes, and enzyme nodes are parsed to construct an initial topological network.

[0078] The method further includes a topology expansion step based on literature mining. Specifically, to address the issue of lagging updates to the KEGG database, an incremental update mechanism is established. The system retrieves literature published in the field of plant physiology and biochemistry over the past decade, extracting specific modification reactions not included in existing databases. Substrate, product, and catalytic enzyme information are extracted. Based on the extracted information, reaction nodes and connecting edges are added to the initial topology network, thereby forming a broader extended flavonoid metabolism network map. Compared to the basic framework, 50 new reaction nodes and 25 new intermediate metabolites are added.

[0079] Example 5: Functional Module Architecture of a Comprehensive Database System

[0080] like Figure 5 As shown, a functional module architecture for a comprehensive database system is provided. This system is built on a B / S (browser / server) model and a front-end and back-end separation architecture. It mainly includes a multi-omics resource management and retrieval module, an AI intelligent functional prediction and structural analysis module, a flavonoid metabolism network interaction module, and an auxiliary analysis tool integration module.

[0081] The multi-omics resource management and retrieval module addresses the challenges of standardized management and efficient distribution of massive amounts of heterogeneous data. Specifically, this module uses a MySQL relational database to store gene metadata, including gene IDs, species classifications, and functional annotation text, and employs a distributed file system to store large-scale genome FASTA files and GFF3 annotation files. The module incorporates a high-performance search engine configured to respond to user queries, allowing users to initiate searches by species name, gene ID, or functional keywords. Furthermore, the module includes a resource distribution unit, providing batch data download services, supporting resume download, and allowing users to access high-quality omics data processed through standardized procedures.

[0082] The intelligent functional prediction and structure analysis module provides online sequence analysis services. Specifically, considering that both ESM-2 model feature extraction and Alpha Fold structure prediction are computationally intensive tasks, this module integrates a Celery distributed task queue and a Redis message middleware for asynchronous task scheduling. When the backend receives an input sequence, it automatically invokes a deployed deep learning model to perform functional classification prediction. For sequences with positive predictions, the module further triggers structure prediction and molecular docking processes, and through the integration of the Mol* molecular browser component, renders the predicted PDB 3D structure in real time on the front end. The AI ​​intelligent functional prediction and structure analysis module also supports interactive rotation and scaling of the protein model, and displays the predicted substrate-binding pockets and key catalytic residues in a highlighted format.

[0083] The flavonoid metabolism module is used to construct and display a knowledge graph of flavonoid metabolism pathways and a multi-dimensional gene list.

[0084] This module utilizes visualization components to draw metabolic pathway networks including major branches such as chalcones, flavonoids, and anthocyanins. Directed edges represent reaction directions, and new reaction steps supplemented based on literature mining are highlighted. It is configured to automatically display specific information about the enzyme and its metabolites in response to user clicks on specific enzyme nodes, thus enabling interactive navigation from macroscopic pathways to microscopic genes.

[0085] In addition to the pathway maps mentioned above, this module is also equipped with a hierarchical data display unit, which is used to present the structured data of thousands of plant flavonoid regulatory genes obtained through sequence alignment.

[0086] At the species level, this module establishes a complete list of genes indexed by species, and is configured to aggregate and display all flavonoid metabolism-related genes identified in the genome of a target species in response to the user's selection of the target species.

[0087] At the pathway level, the upstream pathway set covers universal precursor synthesis genes for the phenylpropane biosynthesis pathway; the downstream branch set covers specific modification genes for the six major flavonoid synthesis branches. Through this three-level indexing architecture of "pathway-species-gene," users can quickly locate the distribution of gene resources at specific metabolic nodes in a specific species.

[0088] The auxiliary analysis tool integration module is used to integrate general bioinformatics tools to form a closed loop for scientific research. Specifically, this module integrates the NCBIBLAST+ program, providing a graphical configuration interface that allows users to perform homology comparisons between their own sequences and this database. The auxiliary analysis tool integration module also embeds the JBrowse genome browser component, allowing users to view the context of genes at the chromosome level, intuitively displaying the structure of exons and introns, as well as the distribution of neighboring genes. Furthermore, this module integrates the Primer3 core algorithm, which automatically designs primer pairs suitable for PCR amplification or qPCR verification based on the target gene CDS sequence selected by the user, and outputs Tm values ​​and GC content parameters.

[0089] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0090] In another embodiment of the present invention, a terminal device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to achieve a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used in the operation of a method for mining and constructing a plant flavonoid pathway based on protein language models and multi-omics.

[0091] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory). This computer-readable storage medium is a memory device in a terminal device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and extended storage media supported by the terminal device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device.

[0092] One or more instructions stored in a computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the above embodiment regarding a method for mining and constructing a plant flavonoid pathway based on a protein language model and multi-omics; one or more instructions in the computer-readable storage medium are loaded and executed by the processor.

[0093] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0094] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0095] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for mining and constructing a database of plant flavonoid pathways based on protein language models and multi-omics, characterized in that, The method includes the following steps: S1: Obtain genome, proteome, and annotation data of multiple plant species, perform redundancy removal and format standardization preprocessing, and identify candidate genes related to flavonoid biosynthesis, transport, and regulation across the entire genome based on sequence homology alignment algorithms, and construct a standardized flavonoid enzyme homology sequence library. S2: Flavonoid-modifying enzyme protein sequences from the homologous sequence library are input into the pre-downloaded and deployed ESM-2 protein language model to extract high-dimensional representation features of the protein sequences; based on the obtained sequence embedding features, a fully connected neural network classification model is constructed and trained to predict the functional category of the protein sequences; for candidate sequences with positive prediction results, their three-dimensional structural models are further constructed, and their potential binding ability with substrates is verified by molecular docking analysis, thereby forming a comprehensive annotation result that integrates sequence features, structural features, and functional information; S3: Based on the KEGG flavonoid pathway, supplement the information on metabolic reaction nodes and corresponding modifying enzymes by integrating literature mining data, construct an extended flavonoid metabolism network map, and map the annotation data to the corresponding nodes of the network map; S4: Based on a front-end and back-end separation architecture, a comprehensive database platform is built to store the genomic and proteomic data of the multi-species plants, the annotation data, and the extended flavonoid metabolism network map into the database, and integrate online bioinformatics tools to provide data retrieval, pathway visualization, and functional prediction services.

2. The method according to claim 1, characterized in that, In step S1, the multi-species plants include algae, mosses, ferns, gymnosperms, and angiosperms; the preprocessing includes removing redundant sequences, standardizing file formats, and correcting erroneous annotations.

3. The method according to claim 1, characterized in that, In step S2, the specific process of inputting the flavonoid-modifying enzyme sequence into the pre-trained ESM protein language model for feature extraction includes: encoding the modifying enzyme protein sequence using the ESM-2 protein language model to extract a high-dimensional embedding representation of the protein sequence; generating a sequence representation vector with a dimension of 1280 by pooling the representation vectors of each amino acid residue in the sequence; constructing a multilayer perceptron classification model containing an input layer, a hidden layer, and an output layer, and inputting the sequence representation vector into the multilayer perceptron classification model for calculation, outputting the prediction result of the target enzyme functional category to which the protein sequence belongs; when the prediction result meets the preset judgment conditions, the corresponding protein sequence is used as a candidate functional enzyme sequence for subsequent three-dimensional structure modeling and molecular docking analysis to assist in its functional annotation.

4. The method according to claim 1, characterized in that, In step S3, the construction of the extended flavonoid metabolism network map includes: identifying and supplementing glycosylation, methylation, and acylation modification reactions not included in the existing KEGG pathway based on literature mining data; establishing a dynamic association index between enzyme sequence IDs and pathway nodes, and configuring it to display the corresponding enzyme sequence list and functional prediction results in response to user click operations.

5. The method according to claim 1, characterized in that, In step S4, the integrated database platform specifically includes: a backend architecture that uses the Django framework to integrate the ESM prediction model interface and the BLAST alignment tool, and uses Celery and Redis to achieve asynchronous scheduling of computing tasks; a frontend interactive interface that uses the Vue.js framework combined with the Mol* plugin and the D3.js library to render protein three-dimensional structures and interactively display metabolic pathways; and a data storage module that uses a MySQL database to store structured data and a file system to store genome sequence files and three-dimensional structure files.

6. A plant flavonoid pathway mining and database system based on protein language models and multi-omics, characterized in that, The system for performing the method according to any one of claims 1 to 5 comprises: The data mining module is used to acquire genome, proteome and annotation file data of multiple plant species, perform redundancy removal and format standardization preprocessing, and identify candidate genes related to flavonoid biosynthesis, transport and regulation in the whole genome based on sequence homology comparison algorithm, and construct a standardized flavonoid enzyme homology sequence library. The intelligent annotation module is used to input flavonoid-modifying enzyme protein sequences into the ESM-2 protein language model to extract high-dimensional representation features, construct and train a fully connected neural network classification model to predict functional categories, and construct three-dimensional structural models for candidate sequences that are predicted to be positive. Combined with molecular docking analysis, it verifies their potential binding ability with substrates and forms a comprehensive annotation result that integrates sequence features, structural features and functional information. The pathway reconstruction module is used to construct an extended flavonoid metabolism network map based on the KEGG flavonoid pathway framework, integrate literature mining data to supplement glycosylation, methylation and acylation modification reactions, and map the annotation data to the corresponding nodes of the network map. The platform service module is used to build a database platform based on a front-end and back-end separation architecture. It stores genomic and proteomic data, annotation data, and extended flavonoid metabolism network maps of multiple plant species into the database, and integrates online bioinformatics tools to provide data retrieval, pathway visualization, and functional prediction services.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.