Gene Coding Breeding Prediction via Graph Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current gene prediction methods for crop breeding face challenges in high-dimensional few-sample problems, where traditional statistical analysis struggles to extract effective features from high-dimensional gene data, leading to low accuracy in phenotype prediction and limited improvement in crop yield.

Innovation Solution

A method and device for gene coding breeding prediction based on graph clustering, which constructs an undirected graph from inter-gene correlation strength, performs clustering to extract co-regulated genomes, fuses allele and cluster information, and uses a deep convolutional neural network with weight sharing to enhance feature extraction and prediction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If traditional statistical analysis methods are used for gene prediction, then the method is simple and easy to implement, but the accuracy of phenotype prediction is low due to inability to extract effective features from high-dimensional gene data

Engineering Contradiction:
Improveease of implementationVSAvoidprediction accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent segments the high-dimensional gene data by constructing a gene map that divides genes into functional modules or pathways. This segmentation transforms the overwhelming high-dimensional data into structured, lower-dimensional representations that capture biological meaning, enabling effective feature extraction while maintaining interpretability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension by constructing a gene map that adds structural information to the flat high-dimensional gene data. This gene map dimension organizes genes based on their functional relationships, transforming the data from a high-dimensional vector space into a structured graph representation that reveals underlying patterns and improves prediction accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If deep learning methods are used for gene prediction, then feature extraction capability is enhanced, but the performance is poor due to shortage of samples in crop breeding

Engineering Contradiction:
Improvefeature extraction capabilityVSAvoidsample quantity
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent performs preliminary action by pre-constructing the gene map using existing biological knowledge and data before the actual prediction task. This pre-processing step organizes the high-dimensional gene data into a structured format that captures functional relationships, reducing the complexity of the learning task and enabling effective prediction even with limited samples.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the parameters of the prediction approach by transforming the input data representation from raw high-dimensional gene expressions to structured gene map features. This parameter transformation reduces the effective dimensionality and highlights the most biologically relevant features, making the problem tractable with limited samples while maintaining deep learning's feature extraction capabilities.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If high-dimensional gene data is analyzed directly, then comprehensive gene information is captured, but effective feature extraction becomes difficult and time-consuming

Engineering Contradiction:
Improvegene information completenessVSAvoidfeature extraction time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent segments the high-dimensional gene data by organizing genes into functional modules, pathways, or networks based on their biological relationships. This segmentation preserves the comprehensive gene information while reducing the search space for feature extraction, as the structured organization allows algorithms to focus on relevant gene groups rather than individual genes in isolation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a structural dimension by constructing a gene map that adds topological relationships to the high-dimensional gene data. This transformation maintains the completeness of gene information while reducing the computational complexity of feature extraction, as the gene map structure provides priors that guide the extraction process and reduce the time required to identify effective features.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20240119314A1Gene coding breeding prediction method and device based on graph clustering
Publication Date: 2024.04.11 ZHEJIANG LAB
  • US20240119314A1 patent drawing
  • US20240119314A1 patent drawing
  • US20240119314A1 patent drawing

AI summary

A method and a device for predicting gene coding breeding based on graph clustering. According to the present disclosure, a gene map is constructed based on inter-gene correlation strength; the gene map is subjected to clustering solution to obtain a number of co-regulated genomes and a genome cluster number information of each gene; gene allelic information and genome cluster number information are fused to obtain the gene cluster code of the sample; based on gene cluster code information and biological phenotype information to be predicted, a deep convolutional neural network is constructed to optimize the prediction performance of genetic breeding.