A data processing method, device and product based on plant gene-trait association
By constructing a multi-dimensional information database and integrating multi-source data to automatically mine gene-trait associations and interpret causal mechanisms, the problems of insufficient timeliness and mechanism interpretation capabilities in existing technologies have been solved, and efficient gene-trait association information processing has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QINGDAO JINONG GENE TECHNOLOGY CO LTD
- Filing Date
- 2026-04-29
- Publication Date
- 2026-07-31
AI Technical Summary
In existing technologies, the analysis of plant gene-trait associations suffers from insufficient timeliness and completeness, making it difficult to leap from statistical associations to causal mechanisms. Furthermore, the ability to interpret the mechanisms of gene-trait associations is weak, and it is unable to effectively distinguish between direct and indirect associations.
By maintaining a multi-dimensional information database and integrating heterogeneous data from multiple sources, we can achieve automated mining of gene-trait associations and interpretation of causal mechanisms, support bidirectional gene-trait retrieval and prediction of unknown gene functions, and provide intelligent recommendations for breeding programs.
It improves the timeliness and accuracy of gene-trait associations, enables information retrieval in multi-dimensional databases, provides comprehensive and efficient information processing, and meets the information processing needs of different users.
Smart Images

Figure CN122493943A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a data processing method, device and product based on plant gene-trait association. Background Technology
[0002] With the rapid development of plant genomics technology, gene-trait association analysis has become a key component of plant functional genomics research and a crucial element in promoting the cultivation of high-quality and high-yield crops and improving agricultural production efficiency.
[0003] In practical plant functional genomics research and crop molecular breeding, the following technical problems still exist: First, the timeliness and completeness of association evidence are severely insufficient. Traditional gene-trait association databases mainly rely on manual compilation and updating, with long update cycles that cannot keep up with the publication speed of the massive amount of research literature in the field of plant breeding. This results in a large amount of reported gene-trait association evidence being missed, failing to provide the latest association information for scientific research and breeding work in a timely manner. Second, the ability to interpret the mechanisms of gene-trait associations is weak, making it difficult to achieve the leap from "statistical association" to "causal mechanism." At present, the direct association evidence of gene regulation of traits mainly comes from experimental verification results in research literature. Most existing analysis platforms can only present the statistical correlation between genes and traits through bioinformatics methods, and cannot effectively distinguish between direct and indirect associations. Summary of the Invention
[0004] This disclosure provides a data processing method, device, and product based on plant gene-trait associations. By maintaining a multi-dimensional information database in the plant field, it provides a foundation for various processing types of genes / traits in the plant field, effectively integrates multi-source heterogeneous data, realizes automated mining of gene-trait associations, has a powerful ability to interpret causal mechanisms, and can also realize bidirectional gene-trait retrieval, prediction of unknown gene functions, and intelligent recommendation of breeding programs.
[0005] According to one aspect of this disclosure, a data processing method based on plant gene-trait associations is provided, comprising: Maintain a multi-dimensional information database, which includes information groups with relationships under multiple dimensions. The information groups in the multi-dimensional information database are obtained by analyzing known literature. The multiple dimensions include at least one of the following: gene-trait dimension, homologous gene dimension, interacting gene dimension, cis-acting element-trait dimension, GO-trait dimension, KEGG-trait dimension, and pathway-trait dimension. A processing request is received, the processing request including a processing type and first information; the processing type includes at least one of information retrieval, unknown gene function prediction and breeding program reasoning, and the first information includes at least one of species, gene and trait; The first information is retrieved from the multi-dimensional information database to determine the second information that is related to the first information. Based on the second information, the processing procedure corresponding to the processing type is executed to obtain the response result corresponding to the processing request.
[0006] According to another aspect of this disclosure, a data processing apparatus based on plant gene-trait associations is provided, comprising: The information database maintenance module is used to maintain a multi-dimensional information database, which includes information groups with relationships under multiple dimensions. The information groups in the multi-dimensional information database are obtained by analyzing known literature. The multiple dimensions include at least one of the following: gene-trait dimension, homologous gene dimension, interacting gene dimension, cis-acting element-trait dimension, GO-trait dimension, KEGG-trait dimension, and pathway-trait dimension. A request receiving module is used to receive a processing request, the processing request including a processing type and first information; the processing type includes at least one of information retrieval, unknown gene function prediction and breeding scheme reasoning, and the first information includes at least one of species, gene and trait; The information retrieval module is used to retrieve information from the first information in the multi-dimensional information database and determine the second information that is related to the first information; The processing module is used to execute the processing procedure corresponding to the processing type based on the second information, and obtain the response result corresponding to the processing request.
[0007] According to another aspect of this disclosure, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data processing method based on plant gene-trait association as described in any embodiment of this disclosure.
[0008] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores computer instructions for causing a processor to execute and implement the data processing method based on plant gene-trait associations as described in any embodiment of this disclosure.
[0009] According to another aspect of this disclosure, a computer program product is provided, which, when executed by a processor, implements the data processing method based on plant gene-trait association as described in any of the embodiments of this disclosure.
[0010] The technical solution provided in this disclosure extracts information from a large amount of known literature in advance to maintain a multi-dimensional information database, providing an information foundation for the analysis of gene-trait associations in the plant field. This simplifies the process of obtaining gene-trait associations in the plant field by eliminating the need for searching and analyzing a large amount of literature. Through at least one of the processing types, such as information retrieval, unknown gene function prediction, and breeding scheme reasoning, multiple processing types can be executed based on the information retrieval in the multi-dimensional information database, meeting the information processing needs of different users for gene-trait associations in the plant field.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this disclosure and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of a data processing method based on plant gene-trait association in an embodiment of this disclosure; Figure 2 This is a schematic diagram of the structure of a data processing device based on plant gene-trait association in an embodiment of this disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation
[0014] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.
[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0016] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0017] Figure 1 This flowchart illustrates a data processing method based on plant gene-trait associations, provided in this embodiment. This embodiment is applicable to situations where a multi-dimensional information database in the plant field is used as a basis to perform one or more data processing procedures, including information retrieval, unknown gene function prediction, and breeding program reasoning. This method can be executed by the plant gene-trait association-based data processing device described in this embodiment. This device can be implemented in software and / or hardware and can be integrated into an electronic device, including but not limited to terminal devices, computer devices, and servers. The terminal devices include, but are not limited to, mobile phones and tablets. Figure 1 As shown, the method specifically includes the following steps: S110, maintain a multi-dimensional information database, which includes information groups with relationships under multiple dimensions. The information groups in the multi-dimensional information database are obtained by analyzing known literature. The multiple dimensions include at least one of the following: gene-trait dimension, homologous gene dimension, interacting gene dimension, cis-acting element-trait dimension, GO-trait dimension, KEGG-trait dimension, and pathway-trait dimension. In this embodiment, a multi-dimensional information database in the field of plants is pre-constructed. This database includes information groups corresponding to multiple dimensions. Here, an information group can be understood as at least two directly related information items extracted from known literature. Optionally, the multi-dimensional database includes multiple sub-databases corresponding to each dimension, and each sub-database includes multiple information groups under one dimension. This multi-dimensional database can be updated in real-time or periodically. By extracting information from incremental literature across multiple dimensions, at least one information group of at least one dimension is obtained from the incremental literature, and the multi-dimensional database is updated to ensure its timeliness and accuracy.
[0018] The multi-dimensional information database includes at least one of the following dimensions: gene-trait dimension, homologous gene dimension, interacting gene dimension, cis-acting element-trait dimension, GO-trait dimension, KEGG-trait dimension, and pathway-trait dimension. Specifically, the sub-database corresponding to the gene-trait dimension includes genes and phenotypic sets with direct relationships; the sub-database corresponding to the homologous gene dimension includes homologous genomes; the sub-database corresponding to the interacting gene dimension includes interacting genomes; the sub-database corresponding to the cis-acting element-trait dimension includes cis-acting elements and phenotypic sets with direct relationships; the sub-database corresponding to the GO-trait dimension includes GO terms and phenotypic sets with direct relationships; the sub-database corresponding to the KEGG-trait dimension includes KEGG pathways and phenotypic sets with direct relationships; and the sub-database corresponding to the pathway-trait dimension includes signal transduction / metabolic pathways and phenotypic sets with direct relationships. It should be noted that a group of information with a direct relationship can be understood as a group of information recorded in a known document that can be directly extracted, rather than a group of information obtained by reasoning from two or more documents.
[0019] Known literature can be understood as published literature in the plant field, which can be obtained from at least one literature publishing platform. Optionally, known documents include multi-species related literature in the plant field, and correspondingly, the multi-dimensional information database includes multi-dimensional information groups across species. In this embodiment, known literature corresponds to different information sources, enabling the interpretation and information extraction of multi-source literature, thereby achieving cross-species information integration, providing a comprehensive information foundation for research in the plant field, and simplifying the process of users collecting and analyzing a large number of documents.
[0020] Optionally, the information groups in the multi-dimensional information database are obtained by analyzing the known documents using an information analysis model; this information analysis model can be a neural network model, such as a large language model. Specifically, the information analysis model is used to analyze the known documents using pre-determined prompt words, extracting information groups corresponding to at least one dimension of the known documents. The prompt words and known documents are input into the information analysis model, which uses the prompt words as auxiliary information to interpret and analyze the known documents, extracting information for each of the multiple dimensions to obtain information groups corresponding to at least one dimension. The multi-dimensional information database is updated using the extracted information groups, that is, the sub-database for each dimension is updated using the information groups corresponding to each dimension.
[0021] In some embodiments, the prompt includes at least one of the following: role positioning description information, task description information, output format description information, basic processing rule description information for information extraction, gene-species association identification rule description information, trait extraction rule description information, and extraction rule description information for information groups under the at least one dimension.
[0022] The role positioning description information can be understood as descriptive content used to define the identity, capabilities, and service convenience of the information analysis model. For example, this role positioning description information could include "You are a professional biological literature information processing system, specifically designed to accurately extract gene function data from scientific research literature and generate structured reports." The task description information can be understood as explanatory content used to clarify the core work objectives and overall processing tasks to be performed. The output format description information can be understood as rules that uniformly constrain the layout, structure, fields, and standardized form of the parsing results. For example, the output format could be JSON. The basic processing rule description information for information extraction can be understood as universally applicable, end-to-end basic parsing and filtering rules. These basic processing rules may include, but are not limited to, the principles of accuracy, completeness, and fidelity to the original text. The gene-species association identification rule description information can be understood as the description information of the specific judgment rules used to identify gene entities and bind them to corresponding plant species. This solves problems such as homologous gene name duplication across different species, inconsistent gene name abbreviations, and cross-species gene confusion. For example, it may include, but is not limited to, identification standards for gene names, gene IDs, and gene aliases. Based on the contextual species background, Latin name, and common crop classification, it automatically binds genes to their corresponding host species. It distinguishes between homologous genes, heterologous genes, and cross-species gene expressions, avoiding incorrect gene-species matching. The trait extraction rule description information can be understood as the description information of the specific extraction rules used to identify plant traits, unifying different expressions of the same trait in different literature and achieving standardization of trait terminology. The information group extraction rule description information under each dimension can be understood as the rule description information used to extract associated information groups under that dimension. This is used to specifically capture the specific associations of each dimension and generate compliant information groups for database entry.
[0023] In some embodiments, the prompt word may be pre-set. When incremental documents are detected, the incremental documents are analyzed based on the information analysis model and the prompt word to obtain incremental information groups in at least one dimension and update the multi-dimensional information database.
[0024] In some embodiments, the prompt word is determined through multiple iterations. The iteration process of the prompt word includes: receiving the prompt word; extracting information from the sample document based on the prompt word using an information analysis model to obtain an information extraction result; obtaining a detection index based on the information extraction result and the information group tags of the sample document, wherein the detection index includes at least one of detection rate and accuracy; updating the prompt word based on the prompt word and the detection index corresponding to each prompt word to obtain an updated prompt word; extracting information from the sample document based on the updated prompt word using an information analysis model to obtain an information extraction result, until the detection index corresponding to the updated prompt word meets a preset information extraction condition.
[0025] The initial prompts can be externally input, and there can be at least one set of prompts. A sample dataset is used to perform indicator detection on at least one set of prompts to determine the detection indicator corresponding to each set of prompts. The sample dataset can include a set number of sample documents and information group labels for these documents. Each information group label can include information group labels corresponding to at least one dimension. For any set of prompts, an information analysis model analyzes the sample documents in the sample dataset based on the prompts to obtain information extraction results. These results can include information groups corresponding to at least one dimension extracted from the sample documents. Detection indicators are calculated based on these extraction results and the information group labels of the sample documents. These indicators include detection rate and accuracy. The detection rate can be understood as the proportion of truly existing information groups in the sample documents that are successfully extracted by the information analysis model. The detection rate is determined by comparing the information extraction results with the information group labels of the sample documents to determine the number of successfully matched information groups. The ratio of the number of successfully matched information groups to the number of information group labels is used to determine the detection rate. Accuracy can be understood as the percentage of correct information groups among the information groups extracted by the information analysis model. The accuracy can be determined by the ratio of the number of successfully matched information groups to the total number of information groups in the information extraction results.
[0026] The detection metrics corresponding to each set of prompt words are compared with the metric thresholds to determine whether the corresponding metric meets the standard. If any metric of a prompt word fails to meet the standard, the prompt word needs to be further updated. Optionally, each set of prompt words and its detection metrics are displayed, and editing operations on the prompt words are received to obtain the updated prompt words. Based on the sample dataset, the updated prompt words are tested to determine the corresponding detection metrics, and this process is repeated until prompt words that meet all detection metrics are obtained.
[0027] S120, receive a processing request, the processing request including a processing type and first information; the processing type includes at least one of information retrieval, unknown gene function prediction and breeding program reasoning, and the first information includes at least one of species, gene and trait.
[0028] The processing request can be user input, including requests for processing needs.
[0029] In some embodiments, receiving a processing request includes: receiving processing requirement description information, converting the processing requirement description information into a processing request, wherein the processing requirement description information may be in at least one of the following forms: text, voice, or image; performing semantic parsing on the processing requirement description information to obtain a processing type and first information; and generating a processing request based on the processing type and the first information.
[0030] In some embodiments, receiving a processing request includes: displaying an interactive interface, the interactive interface including tabs corresponding to at least one processing type; responding to an information input operation on the tab, obtaining input first information, and generating a processing request based on the processing type corresponding to the tab and the first information. Tab switching can be achieved through triggering operations on the tabs in the interactive interface. When one tab is displayed, other tabs are hidden. The displayed tab includes an information input control, through which the user can input first information, which may include at least one of the following: species, gene, and trait in the plant field. Each tab corresponds to a processing type, and a processing request is generated based on the processing type corresponding to the tab and the first information.
[0031] In some embodiments, an intersection interface is displayed, the interactive interface including display areas for each processing type, and each display area for a processing type including an information input control. Users can input first information through the information input control and generate a processing request based on the processing type corresponding to the display area and the first information.
[0032] Optionally, the processing type includes at least one of information retrieval, unknown gene function prediction, and breeding program reasoning. Information retrieval can be understood as a processing method that retrieves information based on first information in a multi-dimensional information database. Unknown gene function prediction can be understood as a processing method that, after retrieving information based on first information in a multi-dimensional information database, predicts the function of genes with unknown functions based on the retrieved second information. Breeding program reasoning can be understood as a processing method that, after retrieving information based on first information in a multi-dimensional information database, generates a breeding program based on the retrieved second information.
[0033] This embodiment provides processing modules / algorithms for different processing types, which can meet the different processing needs of users.
[0034] S130, the first information is retrieved from the multi-dimensional information database to determine the second information that is related to the first information.
[0035] The first information includes genes or traits, and the second information includes genes and / or traits. The second information is related to the first information, and this relationship includes at least one direct relationship and an indirect relationship; that is, the second information includes directly related information and / or indirectly related information of the first information. Directly related information can be understood as information extracted from one set of information, while indirectly related information can be understood as information inferred from two or more sets of information.
[0036] The first information is retrieved from the multi-dimensional information database, including: matching the first information with sub-databases corresponding to multiple dimensions in the multi-dimensional information database to determine the information groups that are successfully matched with the first information; obtaining the direct association information of the first information based on the successfully matched information groups; continuing to match the direct association information of the first information with the multi-dimensional information database; determining the indirect association information of the first information based on the successfully matched information groups; continuing to match the indirect association information of the first information with the multi-dimensional information database; determining the newly added indirect association information of the first information based on the successfully matched information groups; and so on, until no new indirect association information is found; and using the direct association information and the indirect association information of the first information as the second information.
[0037] It is understandable that a multi-dimensional information database includes multi-dimensional information groups across species. By retrieving information from the multi-dimensional information database using the first piece of information, cross-species information retrieval can be achieved. Furthermore, the aforementioned information retrieval includes a multi-round retrieval process, namely, a retrieval process for directly related information and at least one round of retrieval process for indirectly related information. This allows for the retrieval of both directly and indirectly related information from the first piece of information, improving the comprehensiveness of related information retrieval and providing a foundation for information mining in the plant field.
[0038] In some embodiments, the tab page further includes dimension options. Accordingly, retrieving the first information from the multi-dimensional information database includes: in response to a setting operation on the tab page, determining a selected target dimension option, retrieving the first information from the multi-dimensional information database based on the target dimension option, and determining second information that is associated with the first information.
[0039] The above settings operation may include selecting at least one dimension option, where the selected dimension option is the target dimension option. Specifically, the first information is retrieved from the sub-database corresponding to the target dimension option in the multi-dimensional information database to obtain the second information. The retrieval process for the second information will not be elaborated here. By determining the target dimension option and setting the information retrieval scope, interference from invalid information outside the retrieval scope can be avoided.
[0040] S140, based on the second information, execute the processing procedure corresponding to the processing type to obtain the response result corresponding to the processing request.
[0041] Based on the processing type in the processing request, the processing module / algorithm corresponding to that processing type is invoked. The invoked processing module / algorithm processes the second information to obtain the response result corresponding to the processing request.
[0042] In some embodiments, since the information groups in the multi-dimensional information database are obtained through analysis of a large number of documents, the terminology used in different documents is inconsistent and non-standard, resulting in non-standard information in the information groups of the multi-dimensional information database. To address the above technical problem, the second information is standardized to improve its standardization. Optionally, the second information is standardized and / or normalized based on a set database to obtain updated second information. Specifically, gene names in the second information are mapped to the NCBI database, and replacement correction is completed using standardized gene IDs and standard gene names from the NCBI database; trait names in the second information are mapped to the OLS ontology database, and semantic alignment and standardization rewriting of trait descriptions are completed based on the ontology terminology system of the OLS ontology database. Through joint mapping and correction of the two databases, global standardization and semantic normalization of plant gene-trait association data are completed, establishing a unified data standardization system specific to the plant field, solving the problem of messy descriptions of literature source data from the source, and realizing unified management, accurate matching, and efficient reuse of multi-source literature mining data.
[0043] In some embodiments, the method further includes: enhancing the updated second information based on a multi-source functional gene database to obtain enhanced second information; and performing a processing procedure corresponding to the processing type based on the enhanced second information to obtain a response result corresponding to the processing request.
[0044] Multi-source functional gene databases may include, but are not limited to, at least one of the following: NCBI (National Center for Biotechnology Information), PubMed (PublicMEDLINE), UniProt (Universal Protein Resource), BioGRID (Biological General Repository for Interaction Datasets), OrthoDB (Orthologous Database), OMA (Orthologous Matrix), Pfam (Protein families database), IntAct (IntActMolecular Interaction Database), STRING (Search Tool for the Retrieval of Interacting Genes / Proteins), SRA (Sequence Read Archive), and dbSNP (Single Nucleotide Polymorphism). Databases include: Single Nucleotide Polymorphism Database, eggNOG (evolutionary genealogy of genes: Non-supervised Orthologous Groups), Expression Atlas, GO (Gene Ontology), KEGG (Kyoto Encyclopedia of Genes and Genomes), KOBAS (KO-Based Annotation System), and EMBL-EBI (European Molecular Biology Laboratory - European Bioinformatics Institute).The databases include at least one of the following: European Molecular Biology Laboratory - European Institute of Bioinformatics; Xfam (X (cross) Families Database); CDD (Conserved Domain Database); GEO (Gene Expression Omnibus); genebank (Genetic Sequence Data Bank); miRbase (MicroRNA Database); LncRNAdb (Long Non-Coding RNA Database); pept (Peptide Database); Ensembl (Ensembl Genome Browser); MiRanda (MicroRNA Target Prediction Algorithm); Mir2disease (miRNA and Disease Database); and MiRTarBase (miRNA Target Base).
[0045] Based on the aforementioned multi-source functional gene database, and combined with multi-dimensional data such as gene annotation, protein function, trait ontology, metabolic pathways, homologous genes, and non-coding RNA regulation, the updated second information undergoes feature expansion, content completion, and association information mining to enhance the information of the original data, resulting in enhanced second information. This enhanced information may include, but is not limited to, annotation enhancement, association enhancement, feature enhancement, and semantic enhancement. For example, annotation enhancement includes, but is not limited to, supplementing gene function, protein domain, pathway annotation, and ontology terminology descriptions; association enhancement includes, but is not limited to, completing gene-trait, gene-pathway, miRNA-target gene, and homologous gene associations; feature enhancement includes, but is not limited to, supplementing expression profile data, variation information, species conservation, and interaction network characteristics; and semantic enhancement includes, but is not limited to, using authoritative database definitions, completing abbreviations, aliases, and alternative names to eliminate information gaps from literature sources.
[0046] Accordingly, based on the enhanced second information, the processing procedure corresponding to the processing type is executed to obtain the response result corresponding to the processing request, thereby improving the comprehensiveness of information and providing an information basis for the processing procedure corresponding to the processing type.
[0047] Based on the above implementation, in some embodiments, when the processing type is information retrieval, the second information is integrated to obtain the response result, which may include at least some integrated information of the second information. The second information may include information directly related to the first information and / or information indirectly related to the first information.
[0048] In one scenario, the information retrieval processing requirement includes the retrieval of homologous genes and / or interacting genes. Accordingly, the first information includes a first gene, and the second information includes a second gene, where the second gene includes homologous genes and / or interacting genes of the first gene. For example, if the target dimension option includes a homologous gene dimension, the second gene includes homologous genes of the first gene; if the target dimension option includes an interacting gene dimension, the second gene includes interacting genes of the first gene. The number of second genes can be at least one. The first gene is retrieved from the sub-databases corresponding to the homologous gene dimension and / or interacting gene dimension. The retrieved homologous genes and interacting genes are then categorized and integrated to form the response results, which are then displayed in a categorized manner.
[0049] In one scenario, the information retrieval processing requirements include trait detection for genes. Accordingly, the first information includes a first gene, and the second information includes a first trait. The first trait includes at least one of the following: the associated trait of the first gene, the associated trait of the first gene's homologous genes, and the associated trait of the first gene's interacting genes. Specifically, information retrieval of the first gene can be performed at the homologous gene dimension and / or the interacting gene dimension to obtain the homologous genes and / or interacting genes of the first information. Information retrieval of the first gene, the homologous genes and / or interacting genes of the first information can be performed at the gene-trait dimension to obtain the first trait, i.e., the second information. Here, the associated trait of the first gene can be understood as a trait directly related to the first gene, while the associated trait of the first gene's homologous genes and the associated trait of the first gene's interacting genes are traits indirectly related to the first gene.
[0050] The associated traits of the first gene, the associated traits of its homologous genes, and the associated traits of its interacting genes are integrated as the response results. The first traits can be ranked according to the strength of their association with the first gene, and the ranked first traits are then displayed.
[0051] In one scenario, the information retrieval processing requirement includes gene detection for a trait. Accordingly, the first information includes a second trait, and the second information includes a third gene, which includes at least one gene associated with the second trait. Specifically, information retrieval of the second trait can be performed along the gene-trait dimension to obtain genes directly associated with the second trait. Information retrieval of these directly associated genes can then be performed along the homologous gene dimension and / or interacting gene dimension to obtain indirectly associated genes. Further information retrieval of these indirectly associated genes can continue along the homologous gene dimension and / or interacting gene dimension to obtain new indirectly associated genes, until no new indirectly associated genes can be retrieved. The aforementioned directly associated genes and indirectly associated genes are then integrated as the third gene into the response result.
[0052] Optionally, if there are multiple third genes, the third genes can be sorted according to the strength of the association between the genes and the second trait, and the sorted third genes can be displayed.
[0053] In some embodiments, performing a processing procedure corresponding to the processing type based on the second information to obtain a response result corresponding to the processing request includes: when the processing type includes unknown gene function prediction, the first information includes a fourth gene with unknown function, the second information includes a third trait, the third trait includes a trait associated with a fifth gene, and the fifth gene includes homologous genes and / or interacting genes of the fourth gene; and inferring the function prediction information of the fourth gene based on the third trait, as the response result. Here, the second information includes information that is directly related to the first information and / or information that is indirectly related to the first information.
[0054] In one scenario, if an information retrieval is performed on a fourth gene with an unknown function at the homologous gene level, and the second information obtained includes homologous genes of the fourth gene, then the functional information of the homologous genes of the fourth gene is used as the functional prediction information of the fourth gene, i.e., the response result. This functional information of the homologous genes of the fourth gene can be retrieved from the aforementioned multi-source functional gene databases. For example, the functional information of the homologous genes of the fourth gene can be retrieved from at least one of the following databases: OrthoDB, OMA, eggNOG, NCBI, and Ensembl.
[0055] In one scenario, if an unknown fourth gene is retrieved using an interaction gene dimension, and the second information obtained includes at least one interaction gene of the fourth gene, then the at least one interaction gene of the fourth gene is retrieved using cis-acting element-trait dimension, GO-trait dimension, KEGG-trait dimension, and pathway-trait dimension to determine the labeled biological functions, metabolic pathways, signal regulation pathways, and associated trait information of the aforementioned interaction genes. Frequency statistics, functional enrichment, and association weighting are performed on the functional annotation entries of at least one interaction gene to screen out common functional information and common trait regulatory characteristics of at least one interaction gene. Based on the common functional information and common trait regulatory characteristics, the functional prediction information of the unknown fourth gene is inferred.
[0056] In one scenario, functional prediction information based on homologous genes takes precedence over functional prediction information based on interacting genes.
[0057] In one scenario, the functional prediction results obtained based on homologous genes and the functional prediction results obtained based on interacting genes are fused together to obtain the functional prediction information of a fourth gene, which serves as the response result corresponding to the processing request.
[0058] In some embodiments, the processing procedure corresponding to the processing type is executed based on the second information to obtain a response result corresponding to the processing request, including: when the processing type includes breeding scheme reasoning, the first information includes a species and a combination of target traits, the combination of target traits includes at least one fourth trait, and the second information includes a sixth gene that is directly related to each of the fourth traits; the second information is analyzed based on a scheme recommendation model to obtain a breeding scheme that is suitable for the species and the combination of target traits, as the response result.
[0059] For each fourth trait, information is retrieved in the gene-trait dimension to obtain the sixth gene that is directly related to the fourth trait, which serves as the second information.
[0060] In some embodiments, the information set in the gene-trait dimension sub-library includes genes and traits that are directly related, as well as the direction of gene regulation of traits. The direction of regulation may also include promotion, inhibition, and undetermined. The genes, traits, and the direction of gene regulation are contents directly recorded in known literature. For example, the direction of gene regulation can be promotion or inhibition. If the direction of gene regulation is not recorded in known literature, the direction of gene regulation is set to undetermined.
[0061] In some embodiments, each information group in the multi-dimensional knowledge base may carry species information, which can be extracted from known literature to which the information group belongs. Each information group may carry at least one species information. When retrieving information for a specific species, information groups can be first filtered based on species information, and then targeted searches for genes, traits, etc., can be performed on the filtered information groups. When retrieving information across species, a full-scale information retrieval can be performed in the multi-dimensional knowledge base.
[0062] The breeding scheme recommendation model can be understood as a neural network model with scheme generation capabilities, trained based on historical breeding schemes. By inputting the aforementioned second information into the model, at least one breeding scheme is output, such as a top 10 breeding scheme. The ranking of the breeding schemes is based on their confidence scores. A breeding scheme consists of a directional transgenic combination of one or more genes, where each gene exhibits either overexpression or gene knockout. The transgenic direction of each gene can be determined based on the gene's regulatory direction in the second information. For example, in a new variety breeding scenario where the gene's regulation of the target trait is positive, the transgenic direction in the breeding scheme could be overexpression if the gene's regulation of the target trait is positive; conversely, if the gene's regulation of the target trait is negative, the transgenic direction in the breeding scheme could be knockout. For example, in the scenario of negative breeding of a new variety for a target trait, if the gene's regulation of the target trait is to promote it, the transgenic direction in the breeding program can be knockout; conversely, if the gene's regulation of the target trait is to inhibit it, the transgenic direction in the breeding program can be overexpression. Here, positive breeding of the target trait can be understood as breeding the target trait in at least one direction such as higher, larger, or more, while negative breeding of the target trait can be understood as breeding the target trait in at least one direction such as lower, smaller, or fewer.
[0063] In some embodiments, second information and response results are displayed, with the second information used as the basis for analyzing the response results. Optionally, in response to a triggering operation on the second information, the system redirects to a page containing the database to which the second information belongs, where the second information is summarized and displayed.
[0064] In some embodiments, the tabs in the interactive interface also include dimension options and dimension weight settings; the dimension weight settings are used to obtain user-defined dimension weights.
[0065] In some embodiments, the method further includes: in response to a setting operation on the tab, determining a selected target dimension option and a weight of the target dimension option; based on the target dimension option, performing information retrieval on the first information in the multi-dimensional information database to determine second information that is associated with the first information; performing a processing procedure corresponding to the processing type based on the second information to obtain a response result corresponding to the processing request; and setting a confidence score on the response result based on the weight corresponding to the target dimension option.
[0066] For each information item in the second information, the weight of the dimension to which the information item belongs is used as the weight of the information item. The correlation strength between each information item in the second information and the first information is determined. Based on the correlation strength and weight of each information item in the second information, a weighted calculation is performed to obtain the response result and set the confidence score.
[0067] In one scenario, a response result is generated based on local information items in the second set of information. A weighted calculation is then performed based on the association strength and weight corresponding to each local information item to obtain a confidence score for the response result. The confidence score of the response result is displayed along with the response result, allowing users to determine the reliability of the response result based on the confidence score.
[0068] In the above embodiments, corresponding association strength values are set for different degrees of association, such as a first strength value for direct association and a second strength value for indirect association, where the first strength value is greater than the second strength value. Indirect association can include different connection strengths, such as information a-information b-information c-information d, where information a and information b are directly associated, information a and information c are indirectly associated, and information a and information d are indirectly associated. The association strength of information a with information b, information c, and information d decreases sequentially. Different association strengths correspond to different association strength values; the more intermediate connecting information between two information items, the weaker the association strength and the smaller the association strength value.
[0069] The technical solution of this embodiment extracts information from a large amount of known literature in advance to maintain a multi-dimensional information database, providing an information foundation for gene-trait association analysis in the plant field. This eliminates the need for searching and analyzing a large amount of literature, simplifying the process of obtaining gene-trait associations in the plant field. Through at least one processing type, such as information retrieval, unknown gene function prediction, and breeding scheme reasoning, multiple processing types can be executed based on information retrieval from the multi-dimensional information database, meeting the information processing needs of different users regarding gene-trait associations in the plant field.
[0070] Figure 2This is a schematic diagram of a data processing device based on plant gene-trait associations provided in an embodiment of this disclosure. The device specifically includes: a database maintenance module 210, a request receiving module 220, an information retrieval module 230, and a processing module 240.
[0071] The information database maintenance module 210 is used to maintain a multi-dimensional information database, which includes information groups with relationships under multiple dimensions. The information groups in the multi-dimensional information database are obtained by analyzing known literature. The multiple dimensions include at least one of the following: gene-trait dimension, homologous gene dimension, interacting gene dimension, cis-acting element-trait dimension, GO-trait dimension, KEGG-trait dimension, and pathway-trait dimension. The request receiving module 220 is used to receive a processing request, the processing request including a processing type and first information; the processing type includes at least one of information retrieval, unknown gene function prediction and breeding scheme reasoning, and the first information includes at least one of species, gene and trait; Information retrieval module 230 is used to retrieve information from the first information in the multi-dimensional information database and determine second information that is related to the first information; The processing module 240 is used to execute the processing procedure corresponding to the processing type based on the second information to obtain the response result corresponding to the processing request.
[0072] The technical solution of this embodiment extracts information from a large amount of known literature in advance to maintain a multi-dimensional information database, providing an information foundation for gene-trait association analysis in the plant field. This eliminates the need for searching and analyzing a large amount of literature, simplifying the process of obtaining gene-trait associations in the plant field. Through at least one processing type, such as information retrieval, unknown gene function prediction, and breeding scheme reasoning, multiple processing types can be executed based on information retrieval from the multi-dimensional information database, meeting the information processing needs of different users regarding gene-trait associations in the plant field.
[0073] Based on the above embodiments, optionally, the information groups in the multi-dimensional information database are obtained by analyzing the known documents through an information analysis model; The information analysis model is used to analyze the known documents using pre-determined prompt words and extract information groups corresponding to at least one dimension of the known documents; The prompt words include at least one of the following: role positioning description information, task description information, output format description information, basic processing rule description information for information extraction, gene-species association identification rule description information, trait extraction rule description information, and extraction rule description information for information groups under the at least one dimension.
[0074] The device also includes a prompt word iteration module, which receives prompt words, extracts information from sample documents based on the prompt words using an information analysis model, and obtains information extraction results; obtains detection indicators based on the information extraction results and information group tags of the sample documents, the detection indicators including at least one of detection rate and accuracy; updates the prompt words based on the prompt words and the detection indicators corresponding to each prompt word, obtaining updated prompt words; extracts information from the sample documents based on the updated prompt words using an information analysis model, and obtains information extraction results, until the detection indicators corresponding to the updated prompt words meet preset information extraction conditions.
[0075] Based on the above embodiments, optionally, the processing module 240 is further configured to perform standardization and / or normalization processing on the second information based on a set database to obtain updated second information; perform information enhancement on the updated second information based on a multi-source functional gene database to obtain enhanced second information; and execute the processing procedure corresponding to the processing type based on the enhanced second information to obtain the response result corresponding to the processing request.
[0076] Based on the above embodiments, optionally, the request receiving module 220 is specifically used to display an interactive interface, the interactive interface including at least one tab corresponding to each processing type; in response to an information input operation on the tab, it receives the input first information, and generates a processing request based on the processing type corresponding to the tab and the first information; Optionally, the tab may also include dimension options and dimension weight settings; The processing module 240 is further configured to: in response to a setting operation on the tab, determine the selected target dimension option and the weight of the target dimension option; based on the target dimension option, perform information retrieval on the first information in the multi-dimensional information database to determine second information that is associated with the first information; execute the processing procedure corresponding to the processing type based on the second information to obtain a response result corresponding to the processing request; and set a confidence score on the response result based on the weight corresponding to the target dimension option.
[0077] Optionally, the processing module 240 is further configured to: when the processing type is information retrieval, integrate the second information to obtain the response result; Wherein, the first information includes a first gene, the second information includes a second gene, and the second gene includes homologous genes and / or interacting genes of the first gene; Alternatively, the first information includes a first gene, the second information includes a first trait, and the first trait includes at least one of the following: a trait associated with the first gene, a trait associated with a homologous gene of the first gene, and a trait associated with an interacting gene of the first gene. Alternatively, the first information may include a second trait, and the second information may include a third gene, wherein the third gene includes at least one gene associated with the second trait.
[0078] Optionally, the processing module 240 is further configured to: when the processing type includes unknown gene function prediction, the first information includes a fourth gene with unknown function, the second information includes a third trait, the third trait includes a related trait of a fifth gene, and the fifth gene includes homologous genes and / or interacting genes of the fourth gene; and to obtain the function prediction information of the fourth gene based on the inference of the third trait as the response result.
[0079] Optionally, the processing module 240 is further configured to: when the processing type includes breeding scheme reasoning, the first information includes a species and a combination of target traits, the combination of target traits including at least one fourth trait, and the second information includes at least one of a sixth gene, cis-acting element, GO term, KEGG pathway, and signal transduction / metabolism pathway associated with each of the fourth traits; analyze the second information based on the scheme recommendation model to obtain a breeding scheme adapted to the species and the combination of target traits, as the response result.
[0080] The above-described products can perform the methods provided in any embodiment of this disclosure, and have the corresponding functional modules and beneficial effects for performing the methods.
[0081] Figure 3 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0082] like Figure 3As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0083] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as hard disk, magnetic disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0084] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as data processing methods based on plant gene-trait associations.
[0085] In some embodiments, the data processing method based on plant gene-trait associations can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data processing method based on plant gene-trait associations described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the data processing method based on plant gene-trait associations by any other suitable means (e.g., by means of firmware).
[0086] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0087] Computer programs used to implement the methods of this disclosure may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0088] In the context of this disclosure, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0089] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0090] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0091] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0092] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0093] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the data processing method based on plant gene-trait association according to any embodiment of this disclosure.
[0094] In implementing the computer program product, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Python, Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0095] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A data processing method based on plant gene-trait association, characterized in that, include: Maintain a multi-dimensional information database, which includes information groups with relationships under multiple dimensions. The information groups in the multi-dimensional information database are obtained by analyzing known literature. The multiple dimensions include at least one of the following: gene-trait dimension, homologous gene dimension, interacting gene dimension, cis-acting element-trait dimension, GO-trait dimension, KEGG-trait dimension, and pathway-trait dimension. A processing request is received, the processing request including a processing type and first information; the processing type includes at least one of information retrieval, unknown gene function prediction and breeding program reasoning, and the first information includes at least one of species, gene and trait; The first information is retrieved from the multi-dimensional information database to determine the second information that is related to the first information. Based on the second information, the processing procedure corresponding to the processing type is executed to obtain the response result corresponding to the processing request.
2. The method of claim 1, wherein, The information groups in the multi-dimensional information database are obtained by analyzing the known documents through an information analysis model; The information analysis model is used to analyze the known documents using pre-determined prompt words and extract information groups corresponding to at least one dimension of the known documents; The prompt words include at least one of the following: role positioning description information, task description information, output format description information, basic processing rule description information for information extraction, gene-species association identification rule description information, trait extraction rule description information, and extraction rule description information for information groups under the at least one dimension.
3. The method of claim 2, wherein, The prompt words are determined through multiple iterations, and the iteration process of the prompt words includes: Upon receiving prompt words, the information analysis model is used to extract information from the sample documents based on the prompt words, and the information extraction results are obtained. Based on the information extraction results and the information group tags of the sample documents, detection indicators are obtained, and the detection indicators include at least one of detection rate and accuracy. The prompt words are updated based on the prompt words and the corresponding detection indicators to obtain updated prompt words. Information is then extracted from the sample documents using an information analysis model based on the updated prompt words to obtain information extraction results. This process continues until the detection indicators corresponding to the updated prompt words meet the preset information extraction conditions.
4. The method of claim 1, wherein, Based on the second information, the processing procedure corresponding to the processing type is executed to obtain the response result corresponding to the processing request, including: The second information is standardized and / or normalized based on the established database to obtain the updated second information; Information enhancement is performed on the updated second information based on a multi-source functional gene database to obtain enhanced second information. Based on the enhanced second information, the processing procedure corresponding to the processing type is executed to obtain the response result corresponding to the processing request.
5. The method according to claim 1, characterized in that, Receive processing requests, including: The interactive interface includes at least one tab corresponding to each processing type. In response to the information input operation on the tab, the first input information is obtained, and a processing request is generated based on the processing type corresponding to the tab and the first information; In addition, the tab also includes dimension options and dimension weight settings; Based on the second information, the processing procedure corresponding to the processing type is executed to obtain the response result corresponding to the processing request, including: In response to a settings operation on the tab, the selected target dimension option and its weight are determined. Based on the target dimension option, the first information is retrieved from the multi-dimensional information database to determine the second information that is related to the first information; Based on the second information, the processing procedure corresponding to the processing type is executed to obtain the response result corresponding to the processing request, and a confidence score is set on the response result based on the weight corresponding to the target dimension option.
6. The method according to claim 1 or 5, characterized in that, Based on the second information, the processing procedure corresponding to the processing type is executed to obtain the response result corresponding to the processing request, including: When the processing type is information retrieval, the second information is integrated to obtain the response result; Wherein, the first information includes a first gene, the second information includes a second gene, and the second gene includes homologous genes and / or interacting genes of the first gene; Alternatively, the first information includes a first gene, the second information includes a first trait, and the first trait includes at least one of the following: a trait associated with the first gene, a trait associated with a homologous gene of the first gene, and a trait associated with an interacting gene of the first gene. Alternatively, the first information may include a second trait, and the second information may include a third gene, wherein the third gene includes at least one gene associated with the second trait.
7. The method according to claim 1 or 5, characterized in that, Based on the second information, the processing procedure corresponding to the processing type is executed to obtain the response result corresponding to the processing request, including: In cases where the processing type includes prediction of unknown gene function, the first information includes a fourth gene with unknown function, the second information includes a third trait, the third trait includes an associated trait of a fifth gene, and the fifth gene includes homologous genes and / or interacting genes of the fourth gene. The functional prediction information of the fourth gene is obtained based on the reasoning of the third trait, and is used as the response result.
8. The method according to claim 1 or 5, characterized in that, Based on the second information, the processing procedure corresponding to the processing type is executed to obtain the response result corresponding to the processing request, including: When the processing type includes breeding scheme reasoning, the first information includes a species and a combination of target traits, the combination of target traits including at least one fourth trait, and the second information includes a sixth gene for the species that is directly associated with each of the fourth traits. The second information is analyzed based on the scheme recommendation model to obtain a breeding scheme that is suitable for the species and the combination of the target traits, which is the response result.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data processing method based on plant gene-trait association as described in any one of claims 1-8.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the data processing method based on plant gene-trait associations according to any one of claims 1-8.