Multi-omics data integration plant gene function inference system and method based on large language model

Through a multi-omics data integration system based on a large language model, the problems of insufficient reliability and stability of gene function inference in existing technologies have been solved, efficient and reliable gene function inference has been achieved, and the transparency and accuracy of data integration have been enhanced.

CN120853669APending Publication Date: 2025-10-28HUAZHONG AGRI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510991731.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing technologies struggle to fully integrate multi-omics data for accurate gene function inference, resulting in low reliability and poor stability of gene function inference.

Method used

A multi-omics data integration system based on a large language model is adopted, including data acquisition and processing modules, storage and retrieval modules, analysis modules, and inference and interpretation modules. The reliability and transparency of the data are ensured through standardized processing, MongoDB database, two-layer verification framework and large language model guidance framework.

Benefits of technology

It significantly improves the accuracy and stability of gene function inference, enables the effective integration of multi-omics data, supports interactive iterative optimization, and enhances the transparency and interpretability of the inference process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853669A_ABST
    Figure CN120853669A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-omics data integration plant gene function inference system and method based on a large language model, and the system comprises a data obtaining and processing module which is used for collecting multi-omics data and converting the data into a unified data storage format; the storage retrieval module is constructed based on an enhanced retrieval generation framework and is used for organizing and querying the data through a unified index and presenting biological information of the data in an interpretable format; the analysis module is used for establishing a hierarchical evaluation framework and a double-layer verification framework so as to ensure the reliability of the data and solve the evidence conflict problem in multi-omics data integration; and the inference and explanation module is used for establishing a large language model guide framework and completing plant gene function inference according to a preset priority sequence. According to the method, multiple omics data can be comprehensively integrated for accurate gene function inference, and the reliability and stability of gene function inference are improved under the conditions of different species and limited data availability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene function inference technology, and in particular to a plant gene function inference system and method based on multi-omics data integration using a large-scale language model. Background Art

[0002] With the rapid development of genome sequencing technology, scientists are able to identify an increasing number of genes. However, understanding their functions often lags behind the speed of gene discovery, making gene function inference a bottleneck in modern plant biology research. Meanwhile, biological research continuously generates massive amounts of diverse data, providing a more comprehensive perspective for a deeper understanding of gene function. However, the sheer volume of data and potential information conflicts between datasets significantly increase the time and workload required for gene function analysis. Although several specialized databases exist for storing and managing different types of biological information, these databases often suffer from high operational barriers and weak interactivity.

[0003] In recent years, deep learning technology has developed rapidly, and researchers have begun to try using these methods to integrate different types of data for gene function inference. Deep learning methods, including graph neural networks (GNNs) and convolutional neural networks (CNNs), can automatically learn complex biological relationships from raw data, avoiding tedious manual feature engineering, and are adept at capturing complex relationships between genes by learning from massive amounts of data. These methods can integrate information from multiple data types, providing new perspectives for gene function analysis. However, these methods still face many challenges, especially poor interpretability and dependence on limited data sources. The "black box" nature of deep learning models makes it difficult to analyze and judge the source of their inference results, thus making it difficult to determine the reliability of the results.

[0004] This invention proposes a multi-omics data integration system and method for inferring plant gene function based on a large-scale language model. It can comprehensively integrate multi-omics data to perform accurate gene function inference, and improve the reliability and stability of gene function inference when facing different species and limited data availability. Summary of the Invention

[0005] In view of this, the present invention provides a plant gene function inference system and method based on multi-omics data integration using a large-scale language model, in order to solve the technical problem that existing plant gene function inference methods cannot accurately infer gene functions from fully integrated multi-omics data, resulting in low reliability and poor stability of gene function inference.

[0006] To achieve the above-mentioned technical objectives, the present invention adopts the following technical solution: On the one hand, this invention provides a plant gene function inference system based on multi-omics data integration using a large-scale language model, comprising: The data acquisition and processing module is used to collect multi-omics data and convert the multi-omics data into a unified data storage format through standardization processing, while retaining the biological information in the data; The storage and retrieval module, built on an enhanced retrieval generation framework, is used to store the processed data and organize and query the data through a unified index; it is also used to convert structured data into natural language descriptions through text templates and present the biological information of the data in an interpretable format. The analysis module is used to establish a hierarchical evaluation framework and a two-layer validation framework based on biological knowledge to ensure the reliability of the data and to determine the priority order according to biological relevance in order to resolve the problem of evidence conflict in the integration of multi-omics data. The inference and interpretation module is used to establish a large language model guidance framework, which includes a preset task overview template, a data display template, an inference path template, and an output requirement template. Based on the large language model guidance framework, plant gene function inference is completed in a preset priority order.

[0007] Furthermore, the data acquisition and processing module includes a format conversion unit, which is used to extract and filter attributes for each type of data in the multi-omics data, and to convert the data into a uniform storage format with a consistent structure through standardized operations.

[0008] Furthermore, the storage and retrieval module organizes and indexes biological data using the MongoDB database. Data for each species is categorized into different classes, creating separate collections for different types of data. Cross-references are used between related entities, with gene IDs as a unified index, to enable cross-collection data retrieval.

[0009] Furthermore, the storage and retrieval module also includes a text conversion unit, which is used to set corresponding text templates for each data type, convert different biological data formats into natural language descriptions, retain their respective biological meanings, and present biological information in an interpretable format.

[0010] Furthermore, the multi-omics data includes at least: sequence alignment data, homologous gene data, co-expression data, tissue-specific expression data, TWAS data, and functional annotation data from the GO and KEGG pathways.

[0011] Furthermore, the analysis module includes a priority determination unit and a cross-validation unit; The priority determination unit is used to divide the data into direct inference evidence, credible verification evidence, and auxiliary inference evidence in descending order of priority. The direct inference evidence includes TWAS data and functional annotation data from GO and KEGG pathways. The credible verification evidence includes tissue-specific expression data and homologous gene data. The auxiliary inference evidence includes sequence alignment data and co-expression data. When different types of evidence contradict each other, the evidence with higher priority shall prevail. The cross-validation unit is used to construct a two-layer validation framework including an internal consistency check mechanism and an external validation mechanism. The internal consistency check is used to adjudicate in order of priority when evidence is contradictory and to exclude abnormal data that does not conform to biological rationality. The external validation mechanism is used to compare the inference results with the GO and KEGG databases and call the large language model to perform cross-validation based on prior knowledge to ensure the validity of the data.

[0012] Furthermore, the preset task overview template is used to clearly define the biological function of the target gene and its potential role in specific traits; The data display template is used to organize the multi-omics data of the target gene by category and convert it into a natural language description; The inference path template is used to perform TWAS direct association analysis, homologous gene function inference and expression pattern verification according to priority, and to determine gene function in combination with auxiliary inference. The output requirement template is used to provide a statement framework for the target gene function and a logical framework for explaining the function inference, as well as an output format for confidence assessment based on the consistency of evidence.

[0013] On the other hand, the present invention also provides a method for inferring plant gene function based on multi-omics data integration using a large language model, implemented using any of the above-described multi-omics data integration plant gene function inference systems based on large language models, comprising: Collect multi-omics data, and convert the multi-omics data into a unified data storage format through standardization processing to retain the biological information in the data; The processed data is stored and organized and queried using a unified index; structured data is converted into natural language descriptions using text templates, and the biological information of the data is presented in an interpretable format. Guided by a large language model framework with preset task overview templates, data display templates, inference path templates, and output requirement templates, gene function inference is completed according to preset priority order and a two-layer validation framework, including direct association, homologous gene function inference, and expression pattern validation.

[0014] Furthermore, when there are contradictions among the collected multi-omics data, a priority and multi-level validation process is adopted, including: The multi-omics data were prioritized and weighted, with the highest confidence level assigned to TWAS associations, followed by homologous gene evidence, expression patterns, and sequence similarity data. Cross-validation was performed on the prioritized data, with cross-referencing of GO and KEGG annotations as the benchmark. Biological rationale was then assessed through known pathway mechanisms, and functional inference was supported by at least two independent data sources.

[0015] Furthermore, the method also includes: After each round of inputting the target gene and obtaining the gene function inference result, additional research data is supplemented in the form of free text in subsequent rounds, which is integrated with the evidence used in the previous round of inference and iteratively updated to update the function inference result.

[0016] Compared with existing technologies, the plant gene function inference method based on multi-omics data integration using a large-scale language model proposed in this invention has the following advantages: (1) Significantly improved the accuracy of gene function inference: Through the priority assessment framework, it ensures that the inference is given priority to high reliability data, and the two-layer verification mechanism of internal consistency check and external verification reduces misjudgment and false positive results; (2) Improved stability of gene function inference: Through the data standardization module, the quality of available data is improved by fine processing and filtering of different types of biological data. Even when data availability is limited, the efficient storage and retrieval system based on MongoDB ensures accurate acquisition of all relevant available information of the target gene. Through priority analysis and mind chain prompts, the inference process is made transparent and in line with the professional logic of biology, making it easier for researchers to verify and understand the inference results.

[0017] (3) Effective integration of multi-omics data was achieved: it showed a significant synergistic effect, and the combined use of different data types made the inference performance exceed the sum of the individual uses.

[0018] (4) Support interactive iterative optimization: Through a flexible conversational knowledge update framework, users can supplement additional data in natural language. The system integrates new evidence while retaining the previous inference logic to achieve incremental reasoning. Users can supervise and refine the inference results through literature integration, which improves the application value of this method in actual research.

[0019] This invention significantly improves the accuracy, stability, and practicality of plant gene function inference through innovative methods such as priority assessment, data standardization, multi-omics integration, and interactive optimization. At the same time, it enhances the transparency and interpretability of the inference process, providing an efficient and reliable method for plant genomics research. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the structure of the plant gene function inference system based on a large language model and multi-omics data integration provided by the present invention. Figure 2 A schematic diagram of the complete architecture of the gene function inference framework provided by this invention. Detailed Implementation

[0021] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.

[0022] Example 1 Please see Figure 1 This embodiment provides a plant gene function inference system 100 based on multi-omics data integration using a large-scale language model, comprising: The data acquisition and processing module 101 is used to collect multi-omics data and convert the multi-omics data into a unified data storage format through standardization processing, while retaining the biological information in the data. The storage and retrieval module 102 is built on an enhanced retrieval generation framework. It is used to store the processed data and organize and query the data through a unified index. It is also used to convert structured data into natural language descriptions through text templates and present the biological information of the data in an interpretable format. Analysis module 103 is used to establish a hierarchical evaluation framework and a two-layer validation framework based on biological knowledge to ensure the reliability of the data and to determine the priority order according to biological relevance in order to resolve the evidence conflict problem in the integration of multi-omics data. The inference and interpretation module 104 is used to establish a large language model guidance framework, which includes a preset task overview template, a data display template, an inference path template, and an output requirement template. Based on the large language model guidance framework, the inference of plant gene functions is completed in a preset priority order.

[0023] The system in this embodiment improves the quality of usable data through refined processing and filtering of different types of biological data via a data standardization module. Even under conditions of limited data availability, the efficient storage and retrieval system based on MongoDB ensures accurate acquisition of all relevant and available information about the target gene. Through techniques such as priority analysis and thought chain hints, the inference process is made transparent and conforms to biological professional logic, facilitating researchers' verification and understanding of the inference results. This significantly improves the accuracy, stability, and practicality of plant gene function inference, while enhancing the transparency and interpretability of the inference process, providing an efficient and reliable new method for plant genomics research.

[0024] In a preferred embodiment, the multi-omics data includes at least: sequence alignment data, homologous gene data, co-expression data, tissue-specific expression data, TWAS data, and functional annotation data from the GO and KEGG pathways.

[0025] As a specific example, cotton data were collected: based on the TM-1-HAU (Gossypium hirsutum) reference genome version, including Blast alignment data, co-expression data, tissue-specific gene expression data, TWAS data, and functional annotation data from the GO and KEGG pathways; homologous gene data were obtained from the CuttonMD database.

[0026] Arabidopsis data: TAIR10 reference genome was used; Blast alignment data were obtained from the Ensembl Plants database; co-expression network data were obtained from AraNet; gene expression data were obtained from Arabidopsis developmental transcriptome studies; functional annotation was based on Araport11; GO annotation data were obtained from the Plaza database.

[0027] Rice data: based on the Nipponbare reference genome (MSU7 / IRGSP1.0); Blast alignment data were obtained from the RICERC database constructed by Sichuan Agricultural University; co-expression and expression pattern data were obtained from the RGAP database; functional annotation data were obtained from RAP-DB; GO annotation data were obtained from the Plaza database; TWAS data were obtained from the study by Ming et al. on the effects of regulatory variations on rice panicle structure.

[0028] In a preferred embodiment, the data acquisition and processing module includes a format conversion unit, which is used to extract and filter attributes for each type of data in the multi-omics data, and to convert the data into a uniform storage format with a consistent structure through standardized operations.

[0029] As a specific implementation example, specific attribute extraction and filtering rules are designed for each type of data, including: For sequence alignment data: Extracted attributes include alignment quality (such as alignment score, coverage), alignment location (chromosome, start and end positions), and alignment sequence (such as cDNA sequence); the filtering rule is to remove low-quality alignments (such as alignment scores below a threshold).

[0030] For homologous gene data: Extracted attributes include homologous gene ID, homology relationship type (e.g., orthologous, paralogous), and species origin of homologous genes; the filtering rule is to remove homology relationships with low confidence (e.g., BLAST E-value is higher than the threshold).

[0031] For co-expression data: Extracted attributes include gene pairs and co-expression correlation (such as Pearson correlation coefficient); the filtering rule is to remove gene pairs with low correlation (such as correlation coefficients below a certain threshold).

[0032] Sequence similarity data is filtered based on a significance threshold to retain only high-confidence matches; co-expression data is filtered based on correlation to identify meaningful gene-gene relationships; and expression data is standardized to enable cross-sample comparisons.

[0033] Applying statistical standardization methods (such as Z-score standardization, ComBat algorithm, etc.) ensures the comparability of cross-sample and cross-experiment data. Through this standardization process, the system can effectively integrate and utilize multi-omics data from different sources, providing a unified data foundation for subsequent functional inference.

[0034] Each data type is converted to a consistent JSON format suitable for database storage, preserving the quantitative aspects of the data (such as expression values ​​and statistical indicators) and their biological context. Gene expression profiles are processed into tissue-specific patterns with standardized values ​​and organized into hierarchical structures reflecting biological relationships through functional annotation.

[0035] In a preferred embodiment, the storage and retrieval module organizes and indexes biological data using a MongoDB database. Data for each species is classified into different categories, creating separate sets for different types of data. Cross-references are used between related entities, and gene IDs are used as a unified index to achieve cross-set data retrieval.

[0036] In a preferred embodiment, the storage and retrieval module further includes a text conversion unit, which is used to set corresponding text templates for each data type, convert different biological data formats into natural language descriptions, retain their respective biological meanings, and present biological information in an interpretable format.

[0037] By creating separate collections for different data types and maintaining cross-references between related entities, a database implemented using MongoDB is provided, supporting efficient querying and retrieval. Gene IDs are used as a unified index, and relevant information is retrieved from all collections in species-specific databases. Text template conversion is also implemented to transform structured biological data into natural language descriptions.

[0038] For example, a text template for homologous gene data could be: The homolog of gene [gene name] in [species 1] is [homologous gene name], located at [chromosome position] in [species 2]. Their homology indicates [type of homology relationship], which may indicate similarity in [biological function or evolutionary significance].

[0039] For example, a text template for gene expression data could be: The expression level of gene [gene name] in [sample source] is [expression amount], indicating the [expression status] of this gene in [sample description]. Its expression level [comparison result] compared to other samples may be related to [biological function or disease relevance].

[0040] These templates can be used to present complex biological data in natural language, making the information more intuitive and easier to understand. In practice, they can be adjusted and optimized according to specific data types and application scenarios to better meet the needs of different users.

[0041] In a preferred embodiment, the analysis module includes a priority determination unit and a cross-validation unit; The priority determination unit is used to divide the data into direct inference evidence, credible verification evidence, and auxiliary inference evidence in descending order of priority. The direct inference evidence includes TWAS data and functional annotation data from GO and KEGG pathways. The credible verification evidence includes tissue-specific expression data and homologous gene data. The auxiliary inference evidence includes sequence alignment data and co-expression data. When different types of evidence contradict each other, the evidence with higher priority shall prevail. The cross-validation unit is used to construct a two-layer validation framework including an internal consistency check mechanism and an external validation mechanism. The internal consistency check is used to adjudicate in order of priority when evidence is contradictory and to exclude abnormal data that does not conform to biological rationality. The external validation mechanism is used to compare the inference results with the GO and KEGG databases and call the large language model to perform cross-validation based on prior knowledge to ensure the validity of the data.

[0042] With the above settings, the analysis module achieves the functions of hierarchical feature extraction and chain reasoning.

[0043] In the prioritization analysis mechanism, the overall priority is as follows: direct inference evidence is the highest, followed by credible verification evidence, and then auxiliary inference evidence is the lowest. In judging direct inference evidence, TWAS data (due to its direct phenotypic association) has the highest priority, followed by GO and KEGG annotations. When evidence is contradictory, high-weighted evidence is given priority, and biological plausibility is assessed to ensure that the final inference conforms to known biological mechanisms and principles.

[0044] This hierarchical weighting system ensures that the inference process prioritizes the most directly relevant evidence while making reasonable use of other data sources for verification and supplementation, effectively solving the problem of evidence conflict in the integration of multi-omics data.

[0045] In the cross-validation mechanism, internal consistency checks and external validation systems construct a two-layer validation framework to ensure the reliability of gene function inferences. Internal consistency checks follow a priority order of "TWAS > homologous genes > expression data > BLAST > co-expression," evaluating functional indications from different data sources. When conflicting evidence appears (e.g., TWAS shows "stress-related" but expression data is inconsistent), the system prioritizes higher-priority data and, if necessary, excludes anomalous data that does not conform to biological plausibility.

[0046] External validation compares the inference results with the GO and KEGG databases (this data is also included in the data we provide, although some genes may be missing this type of data), and at the same time allows the large model to verify and judge based on its own prior knowledge, forming a cross-validation mechanism.

[0047] The systematic guidance of the three main inference paths (TWAS direct association, homologous gene function inference, and expression pattern verification) and other secondary auxiliary paths in the structured prompt template enables large models to reason according to biological logic. By utilizing a two-layer verification mechanism of internal consistency checks and external verification, misjudgments and false positive results are reduced.

[0048] In a preferred embodiment, the preset task overview template is used to clearly define the biological function of the target gene and its potential role in a specific trait; The data display template is used to organize the multi-omics data of the target gene by category and convert it into a natural language description; The inference path template is used to perform TWAS direct association analysis, homologous gene function inference and expression pattern verification according to priority, and to determine gene function in combination with auxiliary inference. The output requirement template is used to provide a statement framework for the target gene function and a logical framework for explaining the function inference, as well as an output format for confidence assessment based on the consistency of evidence.

[0049] In some embodiments, this system employs a prompt template architecture, employing a structured large language model guidance framework to precisely control the gene function inference process. This guidance framework comprises four core components: The task overview section explicitly instructs participants to "predict the biological function of the target gene, provide a detailed explanation of its potential role in specific traits"; In the data presentation section, target gene data (expression data, homologous gene data, TWAS data, etc.) are organized by category and converted into natural language descriptions using pre-designed templates; In the inference path section, the guided model executes three main inference paths according to priority: TWAS direct association analysis, homologous gene function inference, and expression pattern verification, followed by auxiliary path inference of other secondary priority data. The final analysis requires a concise and clear functional statement, a structured explanation of the functional inference, and solutions to conflicting evidence. When conflicting evidence is found, the template guides the evaluation in the priority order of "TWAS > homologous genes > expression > BLAST > co-expression" to ensure biological plausibility, and provides a three-level confidence assessment (high, medium, and low) based on the consistency of the evidence.

[0050] The thought chain guidance strategy in the prompt template architecture requires the model to demonstrate the logical process of each reasoning step. The final analysis section provides a quantitative assessment of the support of evidence and a structured output format, making the inference results traceable to specific evidence. This structured prompt architecture effectively solves the illusion problem of large models in specialized applications, enabling them to produce functional inferences supported by reliable evidence.

[0051] The template's task overview section uses explicit instructional language, such as "Your task is to predict the biological function of a target gene by analyzing multiple datasets," setting clear expectations and constraints.

[0052] The data presentation section uses natural language to display all relevant biological data, such as "expression data of the target gene: expression level in leaves is 0.45, and in roots is 0.12", etc., to ensure that the model can fully acquire and correctly understand all the evidence.

[0053] The inference pathway guides the model to perform three main analyses: first, analyzing TWAS data to find direct associations; second, analyzing homologous gene functions to find evolutionarily conserved functions; and finally, verifying the tissue specificity of expression patterns to identify possible developmental roles.

[0054] The output requirements section specifies a standardized output format, including a functional summary, detailed functional description, related biological processes, molecular function, inference basis, and uncertainty analysis, to ensure the interpretability and usability of the results.

[0055] Example 2 This invention also provides a method for inferring plant gene function based on multi-omics data integration using a large language model, implemented using any of the multi-omics data integration plant gene function inference systems based on large language models described in Embodiment 1, including: Step S101: Collect multi-omics data, and convert the multi-omics data into a unified data storage format through standardization processing to retain the biological information in the data; Step S102: Store the processed data and organize and query the data using a unified index; convert the structured data into natural language descriptions using text templates and present the biological information of the data in an interpretable format; Step S103: Based on the large language model guidance framework of the preset task overview template, data display template, inference path template and output requirement template, complete the gene function inference of direct association, homologous gene function inference and expression pattern verification according to the preset priority order and two-layer verification framework.

[0056] As a preferred embodiment, when input multi-omics data are contradictory, a priority and multi-level validation process is adopted, including: The multi-omics data were prioritized and weighted, with the highest confidence level assigned to TWAS associations, followed by homologous gene evidence, expression patterns, and sequence similarity data. Cross-validation was performed on the prioritized data, with cross-referencing of GO and KEGG annotations as the benchmark. Biological rationale was then assessed through known pathway mechanisms, and functional inference was supported by at least two independent data sources.

[0057] In some embodiments, when conflicts cannot be explicitly resolved, the system explicitly records the uncertainty and provides a transparent reasoning process for the selected inference. Users can flexibly adjust the weights according to the specific research context, considering factors such as data type, quality, and quantity.

[0058] In a preferred embodiment, the method further includes: After each round of inputting the target gene and obtaining the gene function inference result, additional research data is supplemented in the form of free text in subsequent rounds, which is integrated with the evidence used in the previous round of inference and iteratively updated to update the function inference result.

[0059] For example, users can directly input information such as "I just discovered that this gene is upregulated under drought conditions" or "The latest research shows that this gene is related to photosynthesis," without needing to follow specific data format requirements. The large model will automatically understand this supplementary information, integrate it with the evidence used in the previous round of inference, and generate updated functional inference results. This conversational iterative optimization mechanism is particularly suitable for situations where the initial prediction results have low confidence or where there are multiple possible functions.

[0060] As researchers acquire new evidence during experiments, they can input these findings into the system at any time, refining the functional inference results progressively. The system retains the previous inference logic and the evidence used, and then uses this information in conjunction with new input to perform incremental reasoning, highlighting the contribution or correction of new evidence to functional understanding.

[0061] This approach not only simplifies the user experience but also fully leverages the advantages of large models in contextual understanding and multi-turn dialogue, enabling the system to assist researchers in gene function discovery in a more natural and intelligent way.

[0062] By comprehensively analyzing multifaceted biological evidence, connecting seemingly unrelated points of evidence, and generating testable hypotheses, the system demonstrates its potential for genuine biological discovery through systematic data integration and analysis.

[0063] like Figure 2 As shown, Figure 2 The framework provides a complete architecture for gene function inference. It mainly consists of two main steps: (1) a preprocessing stage, which processes multi-omics data and stores it in MongoDB databases with different datasets; and (2) an implementation stage, in which the framework retrieves and processes relevant information based on the input gene ID, constructs comprehensive prompts through our prompt generation system, and generates detailed functional inferences using a large language model.

[0064] The present invention discloses a plant gene function inference system and method based on multi-omics data integration using a large-scale language model. Through innovative means such as priority assessment, data standardization, multi-omics integration, and interactive optimization, it significantly improves the accuracy, stability, and practicality of plant gene function inference, while enhancing the transparency and interpretability of the inference process, providing an efficient and reliable new method for plant genomics research.

[0065] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A plant gene function inference system based on multi-omics data integration using a large-scale language model, characterized in that, include: The data acquisition and processing module is used to collect multi-omics data and convert the multi-omics data into a unified data storage format through standardization processing, while retaining the biological information in the data; The storage and retrieval module, built on an enhanced retrieval generation framework, is used to store the processed data and organize and query the data through a unified index; it is also used to convert structured data into natural language descriptions through text templates and present the biological information of the data in an interpretable format. The analysis module is used to establish a hierarchical evaluation framework and a two-layer validation framework based on biological knowledge to ensure the reliability of the data and to determine the priority order according to biological relevance in order to resolve the problem of evidence conflict in the integration of multi-omics data. The inference and interpretation module is used to establish a large language model guidance framework, which includes a preset task overview template, a data display template, an inference path template, and an output requirement template. Based on the large language model guidance framework, plant gene function inference is completed in a preset priority order.

2. The plant gene function inference system based on a large-scale language model and multi-omics data integration according to claim 1, characterized in that, The data acquisition and processing module includes a format conversion unit, which is used to extract and filter attributes for each type of data in the multi-omics data, and convert the data into a uniform storage format with a consistent structure through standardized operations.

3. The plant gene function inference system based on a large-scale language model and multi-omics data integration according to claim 1, characterized in that, The storage and retrieval module organizes and indexes biological data using the MongoDB database. Data for each species is classified into different categories, creating separate collections for different types of data. Cross-references are used between related entities, and gene IDs are used as a unified index to enable cross-collection data retrieval.

4. The plant gene function inference system based on a large-scale language model and multi-omics data integration according to claim 1, characterized in that, The storage and retrieval module also includes a text conversion unit, which is used to set corresponding text templates for each data type, convert different biological data formats into natural language descriptions, retain their respective biological meanings, and present biological information in an interpretable format.

5. The plant gene function inference system based on a large-scale language model and multi-omics data integration according to claim 1, characterized in that, The multi-omics data includes at least: sequence alignment data, homologous gene data, co-expression data, tissue-specific expression data, TWAS data, and functional annotation data from the GO and KEGG pathways.

6. The plant gene function inference system based on a large-scale language model and multi-omics data integration according to claim 5, characterized in that, The analysis module includes a priority determination unit and a cross-validation unit; The priority determination unit is used to divide the data into direct inference evidence, credible verification evidence, and auxiliary inference evidence in descending order of priority. The direct inference evidence includes TWAS data and functional annotation data from GO and KEGG pathways. The credible verification evidence includes tissue-specific expression data and homologous gene data. The auxiliary inference evidence includes sequence alignment data and co-expression data. When different types of evidence contradict each other, the evidence with higher priority shall prevail. The cross-validation unit is used to construct a two-layer validation framework including an internal consistency check mechanism and an external validation mechanism. The internal consistency check is used to adjudicate in order of priority when evidence is contradictory and to exclude abnormal data that does not conform to biological rationality. The external validation mechanism is used to compare the inference results with the GO and KEGG databases and call the large language model to perform cross-validation based on prior knowledge to ensure the validity of the data.

7. The plant gene function inference system based on a large-scale language model and multi-omics data integration according to claim 1, characterized in that, The preset task overview template is used to clearly define the biological function of the target gene and its potential role in specific traits; The data display template is used to organize the multi-omics data of the target gene by category and convert it into a natural language description; The inference path template is used to perform TWAS direct association analysis, homologous gene function inference and expression pattern verification according to priority, and to determine gene function in combination with auxiliary inference. The output requirement template is used to provide a statement framework for the target gene function and a logical framework for explaining the function inference, as well as an output format for confidence assessment based on the consistency of evidence.

8. A method for inferring plant gene function based on multi-omics data integration using a large-scale language model, characterized in that, The system, based on a large-scale language model and integrating multi-omics data for plant gene function inference as described in any one of claims 1-7, includes: Collect multi-omics data, and convert the multi-omics data into a unified data storage format through standardization processing to retain the biological information in the data; The storage and retrieval module, built on an enhanced retrieval generation framework, is used to store the processed data and organize and query the data through a unified index; it is also used to convert structured data into natural language descriptions through text templates and present the biological information of the data in an interpretable format. The analysis module is used to establish a hierarchical evaluation framework and a two-layer validation framework based on biological knowledge to ensure the reliability of the data and to determine the priority order according to biological relevance in order to resolve the problem of evidence conflict in the integration of multi-omics data. The inference and interpretation module is used to establish a large language model guidance framework, which includes a preset task overview template, a data display template, an inference path template, and an output requirement template. Based on the large language model guidance framework, plant gene function inference is completed in a preset priority order.

9. The method for inferring plant gene function based on multi-omics data integration using a large-scale language model according to claim 8, characterized in that, When there are contradictions among the collected multi-omics data, a priority and multi-level validation process is adopted, including: The multi-omics data were prioritized and weighted, with the highest confidence level assigned to TWAS associations, followed by homologous gene evidence, expression patterns, and sequence similarity data. Cross-validation was performed on the prioritized data, with cross-referencing of GO and KEGG annotations as the benchmark. Biological rationale was then assessed through known pathway mechanisms, and functional inference was supported by at least two independent data sources.

10. The method for inferring plant gene function based on multi-omics data integration using a large-scale language model according to claim 8, characterized in that, Also includes: After each round of inputting the target gene and obtaining the gene function inference result, additional research data is supplemented in the form of free text in subsequent rounds, which is integrated with the evidence used in the previous round of inference and iteratively updated to update the function inference result.

Citation Information

Patent Citations

  • Artificial intelligence-based gene and gene cluster function prediction method and device

    CN116030881A

  • Functional gene prediction method based on deep learning

    CN119207549A

  • Big model technology-based biological information analysis system

    CN120260674A

  • Artificial intelligence systems and methods for enabling natural language transcriptomics analysis

    US20250139386A1

  • Identifying sequence functionality

    WO2025129249A1