Gene Coding Breeding Prediction via Graph Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current gene prediction methods for crop breeding face challenges in high-dimensional few-sample problems, where traditional statistical analysis struggles to extract effective features from high-dimensional gene data, leading to low accuracy in phenotype prediction and limited improvement in crop yield.
Innovation Solution
A method and device for gene coding breeding prediction based on graph clustering, which constructs an undirected graph from inter-gene correlation strength, performs clustering to extract co-regulated genomes, fuses allele and cluster information, and uses a deep convolutional neural network with weight sharing to enhance feature extraction and prediction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If traditional statistical analysis methods are used for gene prediction, then the method is simple and easy to implement, but the accuracy of phenotype prediction is low due to inability to extract effective features from high-dimensional gene data
Solution Approach 1:
The patent segments the high-dimensional gene data by constructing a gene map that divides genes into functional modules or pathways. This segmentation transforms the overwhelming high-dimensional data into structured, lower-dimensional representations that capture biological meaning, enabling effective feature extraction while maintaining interpretability.
Solution Approach 2:
The patent introduces a new dimension by constructing a gene map that adds structural information to the flat high-dimensional gene data. This gene map dimension organizes genes based on their functional relationships, transforming the data from a high-dimensional vector space into a structured graph representation that reveals underlying patterns and improves prediction accuracy.
2Measurement precision
If deep learning methods are used for gene prediction, then feature extraction capability is enhanced, but the performance is poor due to shortage of samples in crop breeding
Solution Approach 1:
The patent performs preliminary action by pre-constructing the gene map using existing biological knowledge and data before the actual prediction task. This pre-processing step organizes the high-dimensional gene data into a structured format that captures functional relationships, reducing the complexity of the learning task and enabling effective prediction even with limited samples.
Solution Approach 2:
The patent changes the parameters of the prediction approach by transforming the input data representation from raw high-dimensional gene expressions to structured gene map features. This parameter transformation reduces the effective dimensionality and highlights the most biologically relevant features, making the problem tractable with limited samples while maintaining deep learning's feature extraction capabilities.
3Loss of information
If high-dimensional gene data is analyzed directly, then comprehensive gene information is captured, but effective feature extraction becomes difficult and time-consuming
Solution Approach 1:
The patent segments the high-dimensional gene data by organizing genes into functional modules, pathways, or networks based on their biological relationships. This segmentation preserves the comprehensive gene information while reducing the search space for feature extraction, as the structured organization allows algorithms to focus on relevant gene groups rather than individual genes in isolation.
Solution Approach 2:
The patent introduces a structural dimension by constructing a gene map that adds topological relationships to the high-dimensional gene data. This transformation maintains the completeness of gene information while reducing the computational complexity of feature extraction, as the gene map structure provides priors that guide the extraction process and reduce the time required to identify effective features.
Data Source
AI summary
A method and a device for predicting gene coding breeding based on graph clustering. According to the present disclosure, a gene map is constructed based on inter-gene correlation strength; the gene map is subjected to clustering solution to obtain a number of co-regulated genomes and a genome cluster number information of each gene; gene allelic information and genome cluster number information are fused to obtain the gene cluster code of the sample; based on gene cluster code information and biological phenotype information to be predicted, a deep convolutional neural network is constructed to optimize the prediction performance of genetic breeding.


