A Key Gene Network Search Method and System Based on Intelligent Agents
By combining multi-dimensional data scoring and set operations with an intelligent agent approach, a key gene network is constructed, which solves the accuracy problem of single-dimensional gene search in existing technologies and achieves more accurate and objective key gene search.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YAZHOUWAN NATIONAL LABORATORY
- Filing Date
- 2026-03-27
- Publication Date
- 2026-05-26
AI Technical Summary
Existing key gene search methods mostly rely on single-dimensional gene data, which is difficult to meet the needs of actual research and depends on manual assistance, resulting in inaccurate search results.
Using an agent-based approach, we acquire search data from multiple dimensions (omics differential results, target relevance data, and phenotypic theme evidence data), score genes in the whole genome, generate candidate gene sets through set operations, and construct key gene networks based on gene association evidence.
It improves the accuracy and objectivity of key gene search, comprehensively evaluates gene characteristics, reduces human intervention, and the constructed key gene network can systematically visualize core synergistic genes, which meets the needs of actual research.
Smart Images

Figure CN122090956A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of gene search technology, and more specifically, relates to a key gene network search method and system based on intelligent agents. Background Technology
[0002] In crop functional genomics and molecular breeding research, researchers often utilize transcriptome sequencing and multi-omics approaches such as proteomics, metabolomics, and epigenomics to systematically characterize molecular changes under different genotypes or treatment conditions (e.g., stress treatment, developmental stage differences). These omics data typically provide information on differential expression or abundance, pathway enrichment, co-expression relationships, functional annotations, and network associations, which can be used to infer key regulatory factors, synergistic genes, and molecular mechanisms related to target traits.
[0003] In practical applications, a typical task is to identify, based on relevant omics evidence, a set of core partner genes or key genes that are more likely to function synergistically with the target gene or molecule across the entire genome, for subsequent mechanism verification, pathway analysis, and breeding strategy design.
[0004] However, most existing methods for identifying key genes rely on single-dimensional gene data for searching and analysis, and these methods often require manual assistance, making it difficult to find genes that meet actual research needs. Summary of the Invention
[0005] The purpose of this application is to provide a key gene network search method and system based on intelligent agents, so as to improve the accuracy of key gene search.
[0006] A first aspect of this application provides a key gene network search method based on an intelligent agent, comprising: Acquire multiple search datasets for key gene searches of the target gene; the search datasets include: genome-wide omics differential results data, target relevance data, and phenotypic theme evidence data; multiple search datasets correspond to different dimensions; the target gene is genome-wide; For each search data, each gene in the whole genome is scored based on the search data to obtain a score result; based on the score result, a candidate gene set corresponding to the search data is determined from the whole genome. Based on preset set operation rules, set operations are performed on the candidate gene sets corresponding to each search data to obtain the target candidate gene set; For each target candidate gene in the target candidate gene set, the target candidate gene is treated as a network node, and the node attributes of the network node are determined based on the score of the target candidate gene in each candidate gene set. Based on pre-defined gene association evidence and various target candidate genes, edges are generated between network nodes, and key gene networks are obtained based on the node attributes of each network node.
[0007] A second aspect of this application provides a key gene network search system based on an intelligent agent, comprising: The data acquisition module is used to acquire multiple search data sets for key gene searches of the target gene. The search data includes: genome-wide omics differential results data, target relevance data, and phenotypic theme evidence data. The multiple search data sets correspond to different dimensions. The target gene is genome-wide. The first candidate gene search module is used to score each gene in the whole genome based on each search data to obtain a score result; and to determine the candidate gene set corresponding to the search data from the whole genome based on the score result. The second candidate gene search module is used to perform set operations on the candidate gene sets corresponding to each search data based on preset set operation rules to obtain the target candidate gene set. The key gene network determination module is used to identify each target candidate gene in the target candidate gene set as a network node, and to determine the node attributes of the network node based on the score of the target candidate gene in each candidate gene set; to generate edges between each network node based on preset gene association evidence and each target candidate gene, and to obtain the key gene network based on the node attributes of each network node.
[0008] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described agent-based key gene network search method.
[0009] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the agent-based key gene network search method described above.
[0010] The beneficial effects of the agent-based key gene network search method and system provided in this application are as follows: First, this embodiment acquires three types of multi-dimensional search data: omics difference results data, target relevance data, and phenotypic theme evidence data. It comprehensively evaluates genes from three independent perspectives: molecular differences, functional associations, and phenotypic matching, replacing existing single-dimensional analysis, avoiding evaluation bias, and improving the comprehensiveness of gene feature characterization. Second, this embodiment independently completes gene scoring and candidate gene set generation for each dimension of data, and then converges to obtain the target candidate gene set through set operations. This retains the effectiveness of screening at each dimension and obtains the target candidate genes through set operations, reducing manual intervention, improving the objectivity and accuracy of the search, and enhancing the effectiveness and accuracy of key gene search. Finally, this embodiment maps target candidate genes to network nodes, determines node attributes based on multi-dimensional scoring, and generates node edges by combining gene association evidence. The constructed key gene network simultaneously carries the functional association structure between genes and multi-dimensional importance features, systematically visualizing core synergistic genes, and better meeting practical research needs. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A flowchart illustrating a key gene network search method based on an intelligent agent provided in an embodiment of this application; Figure 2 A flowchart illustrating another agent-based key gene network search method provided in an embodiment of this application; Figure 3 A structural block diagram of a key gene network search system based on an intelligent agent provided in an embodiment of this application; Figure 4 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application can also be implemented in embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0014] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.
[0015] Please refer to Figure 1 , Figure 1 The flowchart of a key gene network search method based on an intelligent agent provided in an embodiment of this application is shown. It can be executed by an electronic device. The method can be implemented based on an intelligent agent configured in the electronic device. The method includes: S101-S105.
[0016] S101: Obtain multiple search data for key gene searches used to perform target gene searches.
[0017] In this embodiment, the target gene refers to a specific gene pre-selected from the whole genome set to be analyzed, and it is also the main body of the subsequent key gene search. Key gene search refers to a targeted screening process conducted across the entire genome, with the target gene as the main component, aiming to identify core associated genes that have molecular interactions, functional synergies, and phenotypic co-effects with the target gene. The whole genome refers to the entire set of genes covered in this analysis, including the target gene and other non-target genes to be evaluated.
[0018] In this embodiment, the search data includes: omics difference results data corresponding to the whole genome, target relevance data, and phenotypic theme evidence data; multiple search data correspond to different dimensions; the target gene belongs to the whole genome.
[0019] Among them, the omics differential results data belong to the molecular-level quantitative dimension data. It is a set of results obtained through high-throughput omics (transcriptomics, proteomics, metabolomics, etc.) detection and differential analysis. It is used to characterize the differential expression characteristics and / or differential abundance characteristics of each gene in the whole genome at the omics level. Specifically, it is used to characterize the significance and magnitude of the differences in transcriptional expression level, protein abundance, and metabolite content of each gene in the whole genome under different experimental groups, treatment conditions, or phenotypic backgrounds, reflecting the molecular variation characteristics of genes.
[0020] Target relevance data, belonging to the functional association dimension data, is a qualitative and quantitative dataset obtained based on gene function annotation, co-expression analysis, pathway enrichment, protein interaction verification, and literature text mining. It is used to characterize the functional association characteristics between each gene in the whole genome and the target gene. Specifically, it is used to characterize the tightness and coupling between each gene in the whole genome and the target gene at the molecular function, biological process, and cellular component levels, reflecting the degree of functional association between the gene and the anchor gene.
[0021] Phenotypic theme evidence data, belonging to the phenotypic fit dimension data, is a set of evidence obtained based on phenotypic-gene association annotation, phenotypic enrichment analysis, and empirical literature tracing. It is used to characterize the matching characteristics of the associated phenotypes of each gene in the whole genome and the target gene. Specifically, it is used to characterize the matching degree and association between each gene in the whole genome and the core phenotypic theme (such as resistance, development, yield, etc.) corresponding to the target gene, reflecting the contribution of the gene to the target phenotype.
[0022] S102: For each search data, score each gene in the whole genome based on the search data to obtain the score result; determine the candidate gene set corresponding to the search data from the whole genome based on the score result.
[0023] In this embodiment, for each search data, the quantitative features and qualitative evidence of the current single-dimensional search data are used as the sole evaluation criteria. Values are assigned to all gene individuals within the whole genome coverage, and the strength and fit of the feature performance of a single gene in this dimension are characterized by numerical methods.
[0024] In this embodiment, different scoring rules can be applied to each search data. Taking target relevance data as an example, the scoring rule for target relevance data can be that the stronger the relevance, the higher the score. A mapping relationship can be pre-constructed for determining the score. After obtaining the scoring results, the N genes with the highest scores in the whole genome can be identified as candidate genes corresponding to the search data, and a candidate gene set can be formed.
[0025] S103: Perform set operations on the candidate gene sets corresponding to each search data based on preset set operation rules to obtain the target candidate gene set.
[0026] In this embodiment, the preset set operation rules refer to the standardized operation constraints that are pre-set before the key gene screening process is started, and are used for the convergence screening of multi-dimensional candidate gene sets. In this embodiment, it can refer to intersection operation, union operation, or conditional intersection operation.
[0027] In one embodiment, a set operation is performed on the candidate gene sets corresponding to each search data based on a preset set operation rule to obtain a target candidate gene set, including: taking the intersection or a preset conditional intersection of the candidate gene sets corresponding to the omics difference results data corresponding to the whole genome, the candidate gene sets corresponding to the target correlation data, and the candidate gene sets corresponding to the phenotypic theme evidence data to obtain the target candidate gene set.
[0028] In this embodiment, the default intersection operation can be performed according to a certain order of set operations. For example, firstly, the intersection of the candidate gene set corresponding to the omics difference result data and the candidate gene set corresponding to the target correlation data is calculated to obtain the intermediate set; then, the intersection of the intermediate set and the candidate gene set corresponding to the phenotypic theme evidence data is calculated to obtain the target candidate gene set.
[0029] In this embodiment, the entire genome can also be used as a global reference to perform gene-by-gene comparisons on the three independent candidate gene sets to determine whether a gene individual belongs to all three candidate sets simultaneously; only gene individuals that exist in all three sets are retained to form the basic intersection result, ensuring that the target candidate gene simultaneously possesses significant molecular differences, target gene association, and phenotypic theme fit.
[0030] In this embodiment, if conditional intersection operation is selected, the conditions can be set by the user. For example, within the gene members of the basic intersection, each gene can be further verified to meet the preset additional constraints. Only genes that meet both the set identity requirement and the preset scoring threshold are retained, while gene individuals with low scores and weak evidence in the basic intersection are removed, ultimately resulting in a more accurate target candidate gene set.
[0031] S104: For each target candidate gene in the target candidate gene set, treat the target candidate gene as a network node, and determine the node attributes of the network node based on the score of the target candidate gene in each candidate gene set.
[0032] In this embodiment, each target candidate gene in the target candidate gene set can be understood as a key gene obtained through the search, because each target candidate gene is obtained by screening and searching based on multi-dimensional search data. In order to more clearly reflect the relationship between the target candidate genes and the target genes, and to make it more interpretable, in this embodiment, each target candidate gene is represented in the form of a key gene network. Each target candidate gene is treated as a network node, and the node attributes of the corresponding network node are determined based on the score of each target candidate gene in each candidate gene set.
[0033] S105: Based on the preset gene association evidence and each target candidate gene, generate the edges between each network node, and obtain the key gene network based on the node attributes of each network node.
[0034] In this embodiment, the predefined gene association evidence refers to a set of qualitative and quantitative evidence used to determine the functional associations, interactions, and synergies between genes; including but not limited to pathway co-annotation, knowledge evidence co-occurrence, or network relationship data, which serve as the objective basis for determining whether a connection has been established between nodes. Pathway co-annotation refers to structured functional association evidence formed when multiple genes are uniformly annotated to the same biological pathway, metabolic pathway, or signal regulation pathway. Knowledge evidence co-occurrence refers to co-occurrence-type association evidence formed in knowledge sources such as literature, knowledge bases, and functional texts when different genes are jointly mentioned and jointly associated with a certain biological process / phenotype. Network relationship data originates from existing mature topologies such as gene interaction networks, regulatory networks, and co-expression networks, directly recording gene interactions, regulation, or co-expression relationships.
[0035] In this embodiment, using preset gene association evidence as the judgment criterion, the target candidate genes corresponding to each network node are verified pairwise. For node pairs with valid association evidence, a topological connection is established to generate the edges between nodes. On this basis, each network node carrying node attributes is topologically integrated with the generated edges, so that the key gene network has both the structural expression of the gene association relationship and the quantitative expression of the multi-dimensional screening results of the gene itself, and finally a complete key gene network is formed.
[0036] In this embodiment, each edge in the key gene network is also bound to its corresponding generation basis and evidence pointer, which are determined based on gene association evidence. The generation basis refers to the specific association facts and judgment criteria extracted from preset gene association evidence, used to directly determine the rationality of establishing edges between corresponding network nodes, and used to clarify the gene association relationship represented by the edge. The evidence pointer is a source location identifier constructed based on the original storage structure, entry index, or source identifier of the gene association evidence, which can directly point to the original association evidence entry, data source, or record location corresponding to the edge. After completing the topology generation of the edges between each network node, based on the gene association evidence supporting the generation of the edge, the substantial supporting content of the association relationship and the original evidence source identifier corresponding to the edge are determined respectively; the above two types of information are bound to the corresponding edge, so that each edge in the network has an explanatory basis and evidence source path for the association relationship, improving the interpretability and verifiability of the key gene network.
[0037] As can be concluded from the above, firstly, this embodiment acquires three types of multi-dimensional search data: omics difference results data, target relevance data, and phenotypic theme evidence data. It comprehensively evaluates genes from three independent perspectives: molecular differences, functional associations, and phenotypic matching, replacing existing single-dimensional analysis, avoiding evaluation bias, and improving the comprehensiveness of gene feature characterization. Secondly, this embodiment independently completes gene scoring and candidate gene set generation for each dimension of data, and then converges to obtain the target candidate gene set through set operations. This retains the effectiveness of screening at each dimension and obtains the target candidate genes through set operations, reducing manual intervention, improving the objectivity and accuracy of the search, and enhancing the effect of key gene search. Finally, this embodiment maps target candidate genes to network nodes, determines node attributes based on multi-dimensional scoring, and generates node edges by combining gene association evidence. The constructed key gene network simultaneously carries the functional association structure between genes and multi-dimensional importance features, systematically visualizing core synergistic genes, and better meeting actual research needs.
[0038] In one embodiment of this application, determining the node attributes of a network node based on the score of the target candidate gene in each candidate gene set includes: For each candidate gene set, the ranking position of the target candidate gene in the candidate gene set is determined based on the score of the target candidate gene in the candidate gene set. The standardized score of the target candidate gene is obtained by ranking it in each candidate gene set according to its ranking position. The standardized scores of the target candidate gene in each candidate gene set are scored to obtain the process score and comprehensive score corresponding to the target candidate gene; The standardized score, process score, and comprehensive score of the target candidate gene in each candidate gene set are determined as the node attributes of the network node.
[0039] In this embodiment, for a network node's node attributes, these attributes may not only record the scores of its corresponding target candidate gene in each candidate gene set, but also the process score and overall score during the set operation. The process score refers to an intermediate quantitative indicator obtained by staged fusion operations of the standardized scores of some dimensions, used to characterize the staged comprehensive performance of the target candidate gene at the dimensional combination level. The overall score refers to the final quantitative indicator obtained by complete fusion operations of the standardized scores of all dimensions, used to holistically characterize the comprehensive importance of the target candidate gene in the multi-dimensional screening system.
[0040] In this embodiment, since the scoring rules for different candidate gene sets may differ, direct calculation may lead to data distortion. Therefore, this embodiment determines the ranking position of the target candidate gene in each single dimension based on the internal scoring distribution of each candidate gene set, converting the absolute score into a relative order feature within the dimension. Secondly, a standardization transformation is performed on the ranking position in each dimension to eliminate the difference in evaluation scale between dimensions and form a standardized score with a unified dimension. For example, the target candidate gene ranked first can be set to 10 points, the target candidate gene ranked second can be set to 9 points, and so on. Next, a preset step-by-step scoring operation is performed on the multi-dimensional standardized scores to obtain the process score representing the combined features of the dimensions and the comprehensive score representing the overall features. Finally, the single-dimensional standardized scores, process scores, and comprehensive scores are integrated to jointly constitute the node attributes of the network node, so that the node attributes fully carry the single-dimensional performance, stage fusion performance, and final comprehensive performance of the gene.
[0041] In one embodiment, a scoring operation is performed on the standardized scores of the target candidate gene in each candidate gene set to obtain the process score and comprehensive score corresponding to the target candidate gene, including: The standardized score of the target candidate gene in the candidate gene set corresponding to the omics differential results data and the standardized score of the target candidate gene in the candidate gene set corresponding to the target relevance data are weighted and calculated to obtain the process score corresponding to the target candidate gene. The comprehensive score for the target candidate gene is obtained by weighting the process score and the standardized score of the target candidate gene in the candidate gene set corresponding to the phenotypic theme evidence data.
[0042] In this embodiment, the standardized scores of the target candidate gene in the candidate gene sets corresponding to the search data in three different dimensions can be calculated in the order of set operations performed on the candidate gene sets corresponding to the search data in three different dimensions.
[0043] For example, referring to the aforementioned embodiments, when determining the target candidate gene set, the intersection of the candidate gene set corresponding to the omics difference results data and the candidate gene set corresponding to the target relevance data is first calculated to obtain an intermediate set; then, the intersection of the intermediate set and the candidate gene set corresponding to the phenotypic theme evidence data is calculated to obtain the target candidate gene set. Therefore, in this embodiment, when performing scoring calculations on the standardized scores of the target candidate gene in each candidate gene set, the standardized scores of the target candidate gene in the candidate gene set corresponding to the omics difference results data and the standardized scores of the target candidate gene in the candidate gene set corresponding to the target relevance data should first be weighted to obtain the process score corresponding to the target candidate gene; then, the process score and the standardized scores of the target candidate gene in the candidate gene set corresponding to the phenotypic theme evidence data are weighted to obtain the comprehensive score corresponding to the target candidate gene.
[0044] In this embodiment, the aforementioned sorting process and the calculation of process scores and comprehensive scores can be performed simultaneously with set operations, or they can be performed when determining network node attributes. The weights for the weighted calculation of standardized scores can all be set to 0.5, i.e., evenly distributed.
[0045] In this embodiment, the process score serves as a stage-based intermediate quantitative indicator, which can independently reflect the fusion results of the two dimensions at the molecular level, providing a verifiable intermediate state for the scoring calculation. Compared with the method of directly outputting a single score, the two-level score can clearly define the contribution transmission path of each dimension's score in the fusion process, which aligns with the technical design intention of the embodiment of this application to be fully traceable and verifiable, making it easier for relevant experimental personnel to understand and trace the source.
[0046] As can be seen from the above, the embodiments of this application firstly convert the absolute scores within each candidate gene set into ranking positions within dimensions, and then perform standardized scoring on the ranking positions. This eliminates the differences in evaluation scales between different candidate gene sets, avoids data distortion caused by directly calculating heterogeneous scores, and ensures the comparability and accuracy of cross-dimensional scores. Secondly, a two-level weighted operation is used to generate process scores and comprehensive scores: first, process scores are obtained by fusing scores from the omics difference and target relevance dimensions, and then comprehensive scores are obtained by combining scores from the phenotypic dimension. The process score can independently reflect the results of the dual-dimensional fusion at the molecular level, clarifying the dimensional contribution transmission path. Compared with a single scoring mode, this improves the traceability and verifiability of the scoring operation. Finally, the single-dimensional standardized scores, process scores, and comprehensive scores are integrated into node attributes, fully carrying the single-dimensional performance, stage fusion performance, and final comprehensive performance of genes. This provides more comprehensive quantitative characteristics for key gene networks, improving the interpretability and practical application value of key gene networks.
[0047] In one embodiment of this application, the genome-wide omics differential results data, target relevance data, and phenotypic theme evidence data are data preprocessed in the following manner: Acquire initial omics differential results data, initial target relevance data, and initial phenotypic theme evidence data corresponding to the whole genome; The gene identifiers contained in the initial omics differential results data, initial target relevance data and initial phenotypic theme evidence data are uniformly standardized and mapped, and data that cannot be mapped, has missing fields and / or conflicts are removed during the uniform standardization and mapping process. Data alignment was performed on the initial omics difference results data, initial omics difference results data, and initial phenotypic theme evidence data after removal to obtain genome-wide corresponding omics difference results data, target relevance data, and phenotypic theme evidence data.
[0048] In this embodiment, the initial omics differential results data, initial target relevance data, and initial phenotypic theme evidence data refer to unstandardized, multi-source heterogeneous data directly from original channels such as high-throughput omics detection, functional database retrieval, and phenotypic evidence mining. These data typically suffer from defects such as inconsistent gene identifiers, missing core fields, and / or contradictory information. Therefore, an authoritative public standard gene identifier system can be used as the sole benchmark to establish a one-way unique mapping relationship between multi-source original gene identifiers and standard gene identifiers. This resolves the ambiguity and heterogeneity of gene naming and numbering among heterogeneous data, achieving normalization of gene identity identifiers across the entire dataset. Abnormal entries with mapping failures, missing core fields, and information conflicts are removed, invalid noise data is stripped away, and valid and compliant data is retained. Finally, using the standard gene identifier as a unified index, the cleaned three types of data are precisely matched and aligned, correcting index misalignment issues between data, and ultimately outputting standardized and well-organized screening data suitable for subsequent processes.
[0049] As can be concluded from the above, firstly, this embodiment addresses the core deficiency of inconsistent gene identifiers in the initial data by establishing a one-way unique mapping based on an authoritative public standard identifier system. This resolves naming and numbering ambiguities between heterogeneous data sources, fundamentally eliminating the heterogeneity barrier of multi-source data. Secondly, this embodiment removes abnormal entries with mapping failures, missing fields, and conflicting information, thus stripping away invalid noise data. This avoids interference from incomplete and contradictory information on subsequent gene scoring, purifies the data sample space, and improves the credibility and accuracy of subsequent scoring and screening results. Finally, this embodiment uses standard gene identifiers as a unified index to complete data alignment, correcting the index misalignment problem between different datasets. This constructs structured and regularized screening data, ensuring the accurate execution of subsequent single-dimensional scoring, set convergence, and other processes, while reducing the cost and error of manual data regularization.
[0050] In one embodiment of this application, each gene is scored based on omics differential result data to obtain a scoring result; based on the scoring result, a candidate gene set corresponding to the omics differential result data is determined from the whole genome, including: Based on the differential expression characteristics and / or differential abundance characteristics of each gene in the omics differential results data, the degree of difference of each gene is scored to obtain the score results; based on the preset degree of difference threshold and the score results, the candidate gene set corresponding to the omics differential results data is screened from the whole genome.
[0051] In this embodiment, the differential expression characteristics and differential abundance characteristics of genes in the omics differential results data are used as the quantitative basis. The significance of omics differences of genes is assigned a single-dimensional quantitative calculation according to the preset differential scoring rules. The differential scoring rules and the preset differential degree threshold can be set by the user. The calculated score is positively correlated with the degree of omics difference of genes, providing a numerical basis for gene screening in the omics dimension.
[0052] In this embodiment, if the omics difference results data contains only one of the differential expression features and differential abundance features, it can be calculated directly. If it contains both differential expression features and differential abundance features, the corresponding scores can be calculated separately and the average value can be taken.
[0053] In this embodiment, genes whose scores exceed a preset difference threshold in the scoring results are selected as candidate genes in the candidate gene set corresponding to the omics difference result data.
[0054] In one embodiment of this application, each gene is scored based on target relevance data to obtain a scoring result; based on the scoring result, a candidate gene set corresponding to the target relevance data is determined from the whole genome, including: The correlation between each gene and the target gene in the target correlation data is scored based on the functional association characteristics between each gene and the target gene, and the scoring results are obtained. Based on the preset correlation threshold and the scoring results, the candidate gene set corresponding to the target correlation data is screened from the whole genome.
[0055] In this embodiment, the functional association characteristics between each gene in the whole genome and the target gene are used as the direct evaluation basis, and a single-dimensional quantitative assignment operation is performed according to the preset relevant scoring rules. The relevant scoring rules and preset correlation thresholds can be set by the user. The obtained score is used to characterize the degree of functional association between the gene and the target gene. The score is positively correlated with the association strength, which constitutes the direct quantitative basis for gene screening in this dimension.
[0056] In this embodiment, genes that exceed a preset relevance threshold in the scoring results are selected as candidate genes in the candidate gene set corresponding to the target relevance data.
[0057] In one embodiment of this application, each gene is scored based on phenotypic theme evidence data to obtain a scoring result; based on the scoring result, a candidate gene set corresponding to the phenotypic theme evidence data is determined from the whole genome, including: Based on the matching characteristics of the phenotypes associated with each gene in the phenotypic evidence data and the corresponding phenotypes of the target gene, the phenotypic matching degree of each gene is scored to obtain the scoring results; based on the preset matching degree threshold and the scoring results, the candidate gene set corresponding to the phenotypic evidence data is screened from the whole genome.
[0058] In this embodiment, the matching characteristics of each gene and the target gene's corresponding phenotype are used as the direct evaluation basis. A single-dimensional quantitative assignment operation is carried out according to the preset phenotype scoring rules. The phenotype scoring rules and the preset matching threshold can be set by the user. The score is used to characterize the degree of fit between the gene and the target phenotype. The score is positively correlated with the degree of matching, which constitutes the quantitative basis for gene screening in this dimension.
[0059] In this embodiment, genes that exceed a preset matching threshold in the scoring results are selected as candidate genes in the candidate gene set corresponding to the phenotypic theme evidence data.
[0060] In one embodiment, the key gene network carries data such as sorting tables, standardized scores, process scores, comprehensive scores, evidence pointers, elimination records, standardization methods, and set operation rules for each candidate gene set in the aforementioned embodiments, so as to facilitate relevant experimental personnel to trace the source and conduct research.
[0061] In one embodiment of this application, the agent-based key gene network search method can also be based on, for example, Figure 2 The flowchart shown is implemented.
[0062] Data Input and Preprocessing: Receives initial omics differential results data corresponding to the whole genome, initial target relevance data, and initial phenotypic theme evidence data; performs gene identifier unified standardization mapping, outlier data removal (unable to map, missing fields, data conflicts), and data alignment operations; outputs well-organized three types of screened data.
[0063] Differential processing of omics data: Based on the preprocessed omics differential results data, differential expression features and / or differential abundance features of each gene are extracted, and differential degree scoring is performed; based on the preset differential degree threshold, a candidate gene set corresponding to the omics differential results data is generated, and only genes with significant omics-level differences are retained. Target relevance calculation: Based on the preprocessed target relevance data, extract the functional association features between each gene and the target gene, and perform relevance scoring; based on the preset relevance threshold, generate a candidate gene set corresponding to the target relevance data, and retain only the genes whose functional association with the target gene meets the standard.
[0064] Set convergence: Based on preset set operation rules (intersection or intersection with preset conditions), the intersection operation is performed on the omics differential candidate gene set and the target related candidate gene set to converge to obtain the target candidate gene set, and only high-confidence genes that simultaneously meet the multi-dimensional screening conditions are retained.
[0065] Standardization transformation: For each gene in the target candidate gene set, determine its ranking position in the omics difference candidate gene set, the target relevance candidate gene set (and the phenotypic theme candidate gene set), and convert the ranking position into a standardized score through a preset standardization mapping rule to eliminate the difference in evaluation scale between dimensions and achieve the comparability of cross-dimensional scores.
[0066] First-level weighted fusion: The standardized scores of the target candidate genes in the omics difference dimension and the target relevance dimension are weighted to obtain the process score. This score focuses on the dual-dimensional coupling characteristics of molecular differences and functional associations, and serves as a transitional quantitative indicator for full-dimensional fusion.
[0067] Phenotypic score calculation: Based on the preprocessed phenotypic theme evidence data, the matching features of each gene and the corresponding phenotype associated with the target gene are extracted, the phenotypic matching degree is scored, and the phenotypic dimension quantitative score is obtained, which provides the phenotypic dimension basis for the secondary weighted fusion.
[0068] Secondary weighted fusion and output: The process score obtained from the primary weighted fusion and the standardized score of the phenotypic topic dimension are weighted again to obtain the comprehensive score; the standardized scores of each dimension, the process score, and the comprehensive score are integrated into the node attributes of the network node to complete the quantitative assignment of node features.
[0069] Key gene network search and visualization: Target candidate genes are mapped to network nodes, and edges between nodes are generated based on preset gene association evidence (binding generation basis and evidence pointer). The key gene network is constructed by combining the node attributes of each node, and the network topology and gene importance are visualized, providing intuitive support for the analysis of key gene functions.
[0070] The overall concept of this embodiment is as follows: First, qualified genes under each dimension are screened out by single-dimensional scoring, and then high-confidence candidate genes are purified by set convergence; then, multi-dimensional heterogeneous scores are integrated into interpretable node attributes by standardization and two-level weighted fusion; finally, a key gene network with both structural correlation and quantitative importance is constructed to achieve accurate mining and visualization of core related genes of the target gene.
[0071] Corresponding to the agent-based key gene network search method in the above embodiment, Figure 3This is a structural block diagram of a key gene network search system based on an agent, provided in one embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 3 The key gene network search system 20 based on intelligent agents includes: a data acquisition module 21, a first candidate gene search module 22, a second candidate gene search module 23, and a key gene network determination module 24.
[0072] The data acquisition module 21 is used to acquire multiple search data for key gene searches of the target gene; the search data includes: genome-wide omics difference results data, target relevance data, and phenotypic theme evidence data; multiple search data correspond to different dimensions; the target gene belongs to the whole genome; The first candidate gene search module 22 is used to score each gene in the whole genome based on each search data to obtain a score result; and to determine the candidate gene set corresponding to the search data from the whole genome based on the score result. The second candidate gene search module 23 is used to perform set operations on the candidate gene sets corresponding to each search data based on preset set operation rules to obtain the target candidate gene set. The key gene network determination module 24 is used to identify each target candidate gene in the target candidate gene set as a network node, and to determine the node attributes of the network node based on the score of the target candidate gene in each candidate gene set; to generate edges between each network node based on preset gene association evidence and each target candidate gene, and to obtain the key gene network based on the node attributes of each network node.
[0073] In one embodiment of this application, the second candidate gene search module 23 is specifically used to take the intersection or a preset condition-based intersection of the candidate gene set corresponding to the omics difference results data corresponding to the whole genome, the candidate gene set corresponding to the target relevance data, and the candidate gene set corresponding to the phenotypic theme evidence data to obtain the target candidate gene set.
[0074] In one embodiment of this application, the key gene network determination module 24 is specifically used to determine the ranking position of the target candidate gene in the candidate gene set based on the score of the target candidate gene in the candidate gene set for each candidate gene set. The standardized score of the target candidate gene is obtained by ranking it in each candidate gene set according to its ranking position. The standardized scores of the target candidate gene in each candidate gene set are scored to obtain the process score and comprehensive score corresponding to the target candidate gene; The standardized score, process score, and comprehensive score of the target candidate gene in each candidate gene set are determined as the node attributes of the network node.
[0075] In one embodiment of this application, the key gene network determination module 24 is further used to perform a weighted calculation of the standardized score of the target candidate gene in the candidate gene set corresponding to the omics difference result data and the standardized score of the target candidate gene in the candidate gene set corresponding to the target correlation data to obtain the process score corresponding to the target candidate gene. The comprehensive score for the target candidate gene is obtained by weighting the process score and the standardized score of the target candidate gene in the candidate gene set corresponding to the phenotypic theme evidence data.
[0076] In one embodiment of this application, the agent-based key gene network search system 20 further includes: a data preprocessing module, used to acquire initial omics difference results data, initial target correlation data and initial phenotypic theme evidence data corresponding to the whole genome; The gene identifiers contained in the initial omics differential results data, initial target relevance data and initial phenotypic theme evidence data are uniformly standardized and mapped, and data that cannot be mapped, has missing fields and / or conflicts are removed during the uniform standardization and mapping process. Data alignment was performed on the initial omics difference results, initial target relevance data, and initial phenotypic theme evidence data after removal to obtain omics difference results, target relevance data, and phenotypic theme evidence data corresponding to the whole genome.
[0077] In one embodiment of this application, the omics differential result data is used to characterize the differential expression features and / or differential abundance features of each gene in the whole genome at the omics level; the first candidate gene search module 22 is specifically used to score the degree of difference of each gene based on the differential expression features and / or differential abundance features of each gene in the omics differential result data, and obtain the score result; based on the preset degree of difference threshold and the score result, the candidate gene set corresponding to the omics differential result data is screened from the whole genome.
[0078] In one embodiment of this application, target relevance data is used to characterize the association features between the functions of each gene in the whole genome and the target gene; the first candidate gene search module 22 is specifically used to score the relevance of each gene based on the association features between the functions of each gene in the target relevance data and the target gene, and obtain the scoring results; based on the preset relevance threshold and the scoring results, the candidate gene set corresponding to the target relevance data is screened from the whole genome.
[0079] In one embodiment of this application, phenotypic theme evidence data is used to characterize the matching features of each gene in the whole genome and the associated phenotypes of the target gene; the first candidate gene search module 22 is specifically used to score the phenotypic matching degree of each gene based on the matching features of each gene in the phenotypic theme evidence data and the associated phenotypes of the target gene, and obtain the scoring results; based on the preset matching degree threshold and the scoring results, the candidate gene set corresponding to the phenotypic theme evidence data is screened from the whole genome.
[0080] In one embodiment of this application, each edge in the key gene network is also bound to the corresponding generation basis and evidence pointer, which are determined based on gene association evidence.
[0081] See Figure 4 , Figure 4 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 4 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of each module / unit in the above system embodiments, for example... Figure 3 The functions of the data acquisition module 21, the first candidate gene search module 22, the second candidate gene search module 23, and the key gene network determination module 24 are shown.
[0082] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), but it may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0083] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.
[0084] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store device type information.
[0085] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation method described in the key gene network search method based on intelligent agents provided in the embodiments of this application, or they can execute the implementation method of the electronic device described in the embodiments of this application, which will not be repeated here.
[0086] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0087] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0088] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0089] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0090] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or units, or it may be an electrical, mechanical, or other form of connection.
[0091] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0092] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0093] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0094] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A key gene network search method based on intelligent agents, characterized in that, The method includes: Multiple search data sets are acquired for key gene searches of the target gene; the search data sets include: genome-wide omics differential results data, target relevance data, and phenotypic theme evidence data; the multiple search data sets correspond to different dimensions; the target gene belongs to the genome-wide genome; For each search data, each gene in the whole genome is scored based on the search data to obtain a scoring result; based on the scoring result, a candidate gene set corresponding to the search data is determined from the whole genome. Based on preset set operation rules, set operations are performed on the candidate gene sets corresponding to each search data to obtain the target candidate gene set; For each target candidate gene in the target candidate gene set, the target candidate gene is used as a network node, and the node attributes of the network node are determined based on the score of the target candidate gene in each of the candidate gene sets. Based on the preset gene association evidence and each of the target candidate genes, edges are generated between each of the network nodes, and a key gene network is obtained based on the node attributes of each network node.
2. The agent-based key gene network search method as described in claim 1, characterized in that, The set operation is performed on the candidate gene set corresponding to each search data based on the preset set operation rules to obtain the target candidate gene set, including: The target candidate gene set is obtained by taking the intersection of the candidate gene set corresponding to the omics difference results data of the whole genome, the candidate gene set corresponding to the target relevance data, and the candidate gene set corresponding to the phenotypic theme evidence data, or by taking the intersection of the candidate gene set with preset conditions.
3. The agent-based key gene network search method as described in claim 1, characterized in that, The process of determining the node attributes of the network node based on the score of the target candidate gene in each of the candidate gene sets includes: For each candidate gene set, the ranking position of the target candidate gene in the candidate gene set is determined based on the score of the target candidate gene in the candidate gene set; The standardized score of the target candidate gene in each of the candidate gene sets is obtained by performing a standardized score on the ranking position of the target candidate gene in each of the candidate gene sets. The standardized scores of the target candidate gene in each of the candidate gene sets are scored to obtain the process score and comprehensive score corresponding to the target candidate gene; The standardized score of the target candidate gene in each of the candidate gene sets, the process score, and the comprehensive score are determined as the node attributes of the network node.
4. The agent-based key gene network search method as described in claim 3, characterized in that, The step of performing a scoring operation on the standardized scores of the target candidate gene in each of the candidate gene sets to obtain the process score and comprehensive score corresponding to the target candidate gene includes: The standardized score of the target candidate gene in the candidate gene set corresponding to the omics difference results data and the standardized score of the target candidate gene in the candidate gene set corresponding to the target correlation data are weighted and calculated to obtain the process score corresponding to the target candidate gene. The process score and the standardized score of the target candidate gene in the candidate gene set corresponding to the phenotypic evidence data are weighted and calculated to obtain the comprehensive score corresponding to the target candidate gene.
5. The agent-based key gene network search method as described in claim 1, characterized in that, The genome-wide omics differential results, the target relevance data, and the phenotypic theme evidence data are preprocessed using the following method: Acquire initial omics differential results data, initial target relevance data, and initial phenotypic theme evidence data corresponding to the whole genome; The identifiers of genes contained in the initial omics differential results data, the initial target correlation data, and the initial phenotypic theme evidence data are uniformly standardized and mapped, and data that cannot be mapped, has missing fields, or has data conflicts are removed during the uniform standardization and mapping process. Data alignment is performed on the initial omics difference results data, initial target relevance data, and initial phenotypic theme evidence data after removal to obtain the omics difference results data, target relevance data, and phenotypic theme evidence data corresponding to the whole genome.
6. The agent-based key gene network search method as described in any one of claims 1 to 5, characterized in that, The omics differential results data are used to characterize the differential expression features and / or differential abundance features of each gene in the whole genome at the omics level; Each gene is scored based on the omics difference results data to obtain the scoring results; Based on the scoring results, a candidate gene set corresponding to the omics differential results data is determined from the whole genome, including: The degree of difference of each gene is scored based on the differential expression characteristics and / or differential abundance characteristics of each gene in the omics differential results data, and the scoring results are obtained. Based on a preset threshold for the degree of difference and the scoring results, a candidate gene set corresponding to the omics difference results data is obtained from the whole genome.
7. The agent-based key gene network search method as described in any one of claims 1 to 5, characterized in that, The target relevance data is used to characterize the functional association between each gene in the whole genome and the target gene; Each gene is scored based on the target relevance data to obtain the scoring results; Based on the scoring results, a candidate gene set corresponding to the target relevance data is determined from the whole genome, including: The correlation score is obtained by scoring each gene based on the functional association characteristics between each gene and the target gene in the target correlation data; Based on the preset relevance threshold and the scoring results, a candidate gene set corresponding to the target relevance data is obtained from the whole genome.
8. The agent-based key gene network search method as described in any one of claims 1 to 5, characterized in that, The phenotypic evidence data is used to characterize the matching features of the associated phenotypes of each gene in the whole genome and the target gene; Each gene is scored based on the phenotypic evidence data to obtain a scoring result; Based on the scoring results, a candidate gene set corresponding to the phenotypic theme evidence data is determined from the whole genome, including: Based on the matching features of the phenotypic evidence data of each gene and the corresponding associated phenotype of the target gene, the phenotypic matching degree of each gene is scored to obtain the scoring result; Based on a preset matching threshold and the scoring results, a candidate gene set corresponding to the phenotypic evidence data is obtained from the whole genome.
9. The agent-based key gene network search method as described in claim 1, characterized in that, Each edge in the key gene network is also bound to a corresponding generation basis and evidence pointer, which are determined based on the gene association evidence.
10. A key gene network search system based on intelligent agents, characterized in that, include: The data acquisition module is used to acquire multiple search data for key gene searches of the target gene; The search data includes: omics differential results data corresponding to the whole genome, target relevance data, and phenotypic theme evidence data; the multiple search data correspond to different dimensions; the target gene belongs to the whole genome; The first candidate gene search module is used to score each gene in the whole genome based on each search data to obtain a score result; and to determine the candidate gene set corresponding to the search data from the whole genome based on the score result. The second candidate gene search module is used to perform set operations on the candidate gene sets corresponding to each search data based on preset set operation rules to obtain the target candidate gene set. The key gene network determination module is used to identify each target candidate gene in the target candidate gene set as a network node, and to determine the node attributes of the network node based on the score of the target candidate gene in each candidate gene set; to generate edges between each network node based on preset gene association evidence and each target candidate gene, and to obtain the key gene network based on the node attributes of each network node.