Phenotype and genotype data correlation analysis method
By using quantified feature subsets and cross-attention mechanisms in the correlation analysis of phenotype and genotype data, combining multiple machine learning algorithms and K-means clustering algorithms, the problems of insufficient linear relationship capture and difficult processing of high-dimensional data in the existing technology are solved, and more efficient phenotypic prediction model performance and lower computing resource consumption are achieved.
Patent Information
- Application Number
- CN202510333204.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has problems such as insufficient linear relationship capture in the correlation analysis of phenotype and genotype data, easy to produce false positive results when processing high-dimensional data, strong dependence on data preprocessing and feature engineering, and high consumption of computing resources, resulting in limited application effect in complex trait research.
The quantized subset of features is used as the input of the cross attention mechanism, and SNP sites are initially screened through multiple machine learning algorithms (ANOVA, random forest, mutual information coefficient, XGBoost), vector quantization is performed in combination with the K-means clustering algorithm, and finally the correlation between SNP sites and phenotypic data is calculated through the cross attention mechanism, and weighted processing is performed to screen out the most relevant features.
Effectively capture the complex relationship between SNP sites and phenotypic data, reduce the risk of overfitting, improve the performance of phenotypic prediction models, and significantly reduce the demand for computing resources.
Smart Images

Figure CN120220819A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of correlation data feature analysis, and specifically relates to a method for analyzing the correlation between genotype and phenotype data. Background Art
[0002] Phenotype data is generally stored in a CSV format file. Each row represents a sample, and each column represents a different phenotypic feature, which contains the numerical values of the target traits of each sample. These data are collected according to the experimental design and are usually obtained through measurement, scoring, or experimental results; genotype data contains the variation information of each sample at multiple gene loci and is usually stored in the VCF format. The VCF format file records the genotype information of each sample at different SNP (single nucleotide polymorphism) loci, and these loci may involve variations related to certain traits; the relationship between phenotype data and genotype data is the core of this study. The genotype data and phenotype data of each sample are paired, which means that each record in the phenotype data corresponds to a specific genotype sample.
[0003] Currently, the mainstream methods for studying the correlation between phenotype and genotype data include genome-wide association study (GWAS), machine learning, and deep learning, etc. GWAS detects the association between SNP loci and phenotypes through statistical methods, and common tools include PLINK and GEMMA. However, GWAS can only capture linear relationships, cannot handle the interaction effects between SNP loci, and is prone to producing false positive results when dealing with high-dimensional data. Important signals may be lost after multiple comparison corrections. Machine learning methods such as random forest and support vector machine can capture non-linear relationships, but are highly dependent on data preprocessing and feature engineering. For example, SNP loci need to be encoded and dimensionally reduced, and are prone to overfitting on small sample datasets. Deep learning methods such as convolutional neural network (CNN) and recurrent neural network (RNN) perform well in processing high-dimensional data, but the training process requires a large amount of computing resources and has high requirements for data quality and sample size, making it difficult to be efficiently applied to large-scale data. These limitations have restricted the application effect of existing technologies in the study of complex traits.
[0004] In view of the above, the present application provides a method for analyzing the correlation between phenotype and genotype data to solve the above problems for studying the relationship between phenotype data and genotype data. Summary of the Invention
[0005] In view of the above situation, to overcome the defects of the prior art, the present invention provides a method for analyzing the correlation between phenotype and genotype data. The present invention uses the quantified feature subset as the input of the cross-attention mechanism to finally screen out SNP loci.
[0006] A method for analyzing the association between phenotype and genotype data, characterized in that it comprises the following steps:
[0007] S1: Collect sample data and preprocess them, remove samples lacking phenotypic data, and then use Z-score standardization, i.e. zero mean standardization, to process the phenotypic data. The standardized data has a mean of 0 and a standard deviation of 1, making the weights of each trait more balanced;
[0008] Preprocess the genotype data in the sample, filter the SNP data, and remove low-quality SNP sites;
[0009] Then the preprocessed phenotypic data and genotypic data samples are aligned, and finally the genotypic data is mapped to integer encoding to convert the genotypic data into a format acceptable to the machine learning model;
[0010] S2: Using four machine learning algorithms, namely ANOVA, Random Forest, MI, and XGBoost, the two data of the samples in S1 were machine learned and analyzed to preliminarily screen out SNP sites with strong correlation with the phenotype, obtain four independent feature subsets, and preliminarily screen out four groups of features with the most predictive ability, thereby improving the performance of the phenotype prediction model and reducing overfitting;
[0011] S3: Use K-means clustering algorithm to construct a codebook, i.e., perform vector quantization processing on the four independent feature subsets initially screened in S2, and map them into four discrete data feature subsets in the form of discrete code words, thereby removing some data noise and unnecessary details, which helps to improve the performance of the phenotype prediction model and reduce the risk of overfitting;
[0012] S4: taking the discrete data after quantization in S3 as input, further screening out important features through the cross-attention mechanism, calculating the correlation between each SNP site and the phenotypic data using the cross-attention mechanism, and weighting the features based on the correlation weights, and finally screening out a set of weighted features, which are the most relevant parts of the phenotypic data and the genotypic data, i.e., the SNP sites finally screened out.
[0013] The above technical solution has the following beneficial effects:
[0014] (1) The present invention utilizes the quantized feature subset as the input of the cross-attention mechanism. Through the cross-attention mechanism, the correlation between each SNP locus and the phenotypic data is calculated, and weights are assigned based on these correlations. The cross-attention mechanism dynamically adjusts the importance of different SNP loci through these weights, enabling SNP loci highly correlated with the phenotypic data to obtain higher weights, while loci with weaker relationships are weakened or ignored. In this way, a set of weighted features is finally screened out, which are the most relevant parts of the phenotypic data and the genotype data, that is, the finally screened SNP loci;
[0015] (2) The present invention conducts preliminary screening by combining multiple machine learning algorithms (ANOVA, random forest, mutual information coefficient, XGBoost), leveraging the advantages of each algorithm to capture the potential associations between SNP loci and phenotypic data. The present invention also uses the K-means clustering algorithm to perform vector quantization processing on the preliminarily screened feature subset, mapping high-dimensional features into discrete codeword forms, effectively removing data noise and unnecessary details while retaining key information, providing cleaner input data for subsequent steps. This multi-stage screening method significantly reduces the overfitting risk and improves the performance of the phenotypic prediction model by gradually optimizing the feature subset.
[0016] (3) The present invention introduces a cross-attention mechanism that can deeply capture the complex relationships between SNP loci and phenotypic data, including non-linear associations and interaction effects between multiple loci. Traditional methods usually have difficulty effectively dealing with these complex relationships, while the cross-attention mechanism dynamically calculates the correlation between SNP loci and phenotypic data and assigns weights, enabling loci highly correlated with the phenotype to receive more attention. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is the overall technical roadmap of the specific implementation manner of the present invention;
[0018] Figure 2 is the example diagram of partial data of genotype and phenotype in the embodiment;
[0019] Figure 3 is the processing flow chart of genotype and phenotype data in the embodiment;
[0020] Figure 4 is the example diagram of the processing results of genotype and phenotype data in the embodiment;
[0021] Figure 5 is the example diagram of the preliminary feature selection results in the embodiment;
[0022] Figure 6 is the example diagram of the vector quantization results in the embodiment;
[0023] Figure 7 Example diagram of the partial result of the feature weights of the cross-attention mechanism in the embodiment. Detailed implementation manners
[0024] The foregoing and other technical contents, features and effects of the present invention will be clearly presented in the following detailed description of the embodiments with reference to the accompanying drawings of the present application. The contents mentioned in the following embodiments are all referenced to the drawings of the specification.
[0025] Embodiment 1, as Figure 2 shown, for the analysis of the characteristics of the existing dataset, phenotypic data is generally stored in a CSV format file. Each row represents a sample, and each column represents a different phenotypic feature. It contains the numerical values of the target traits of each sample. These data are collected according to the experimental design, usually obtained through measurement, scoring or experimental results. Genotype data contains the variation information of each sample at multiple gene loci and is usually stored in the VCF format. The VCF format file records the genotype information of each sample at different SNP (single nucleotide polymorphism) loci, and these loci may be involved in variations related to certain traits. The relationship between phenotypic data and genotype data is the core of this study. The genotype data and phenotypic data of each sample are paired, which means that each record in the phenotypic data corresponds to a specific genotype sample.
[0026] Embodiment 2, as Figure 3 、 Figure 4 shown, on the basis of Embodiment 1, preprocess the sample data. For the preprocessing of phenotypic data, samples may lack key trait records, and these missing phenotypic data cannot be used for model training or analysis. Samples lacking phenotypic data need to be excluded. Then use Z-score normalization (zero-mean normalization) to process the phenotypic data. The formula for Z-score normalization is:
[0027]
[0028] where x is the original value of the sample, μ is the mean of this feature, and σ is the standard deviation of this feature. The normalized data has the characteristics of a mean of 0 and a standard deviation of 1, making the weights of each trait more balanced. For the preprocessing of genotype data, in actual analysis, not all SNP data are reliable and useful. To improve the accuracy of the analysis, first, the SNP data needs to be filtered to exclude low-quality SNP loci. Here, the bcftools v1.21 tool is used for filtering, mainly according to the following criteria:
[0029] ●QUAL≥40: QUAL is the quality score, indicating the variation call quality of the SNP locus. Usually, QUAL
[0030] SNP sites with higher values are considered more reliable. Therefore, SNPs with QUAL ≥ 40 are selected to ensure data quality.
[0031] · Missing rate ≤ 10%: SNP sites with excessive missing values can lead to data instability. Therefore, variant sites with a high missing rate need to be removed. Here, the missing rate threshold is set to 10%, that is, only SNPs with a missing rate lower than or equal to 10% are retained.
[0032] Then sample alignment is performed. Since genotype data and phenotype data come from different sources or experimental treatments, there may be a situation where the sample order is inconsistent. When performing gene-phenotype association analysis, it is necessary to ensure that the samples in the genotype data and phenotype data correspond one by one.
[0033] Finally, the genotypes are mapped to integer encodings. In genotype data, each SNP site can have different genotypes, which are usually represented by such as "0 / 0", "1 / 1", "0 / 1", etc. To convert genotype data into a format acceptable to machine learning models, these genotype values usually need to be mapped to integer encodings. The specific mapping method is as follows:
[0034] ● 0: Represents the homozygous genotype of the reference allele (such as "0 / 0");
[0035] ● 1: Represents the heterozygous genotype, which contains one reference allele and one alternative allele (such as "0 / 1" or "1 / 0");
[0036] ● 2: Represents the homozygous genotype of the alternative allele (such as "1 / 1").
[0037] Example 3, such as Figure 5As shown, based on Example 2, preliminary feature selection is carried out. Feature selection is an important step in machine learning and data analysis, aiming to screen out the most predictive features, improve model performance and reduce overfitting. In this study, we used four different machine learning algorithms - ANOVA, Random Forest, Mutual Information (MI), and XGBoost. Taking gene data and phenotype data as input, we conducted preliminary feature selection. During the feature selection process, we used GridSearchCV for hyperparameter optimization to improve the performance of each model. GridSearchCV traverses the preset parameter grid by exhaustive search, automatically adjusts the hyperparameters of each model, finds the optimal model configuration, ensures the best performance of each algorithm, and thus obtains more accurate feature scores. After feature selection by the above four algorithms, each algorithm will output a set of scores for the features (SNP loci). Based on these scores, we screened out the SNP loci with strong correlation with the phenotype, obtained four independent feature subsets, and ensured that these features can provide sufficient predictive information for subsequent model training.
[0038] ANOVA (Analysis of Variance) is a statistical method used to evaluate whether there are significant differences in the means among multiple groups. During feature selection, ANOVA measures the relevance of a feature to the target variable by comparing the ratio of the variance between different groups and the within-group variance of each feature. Features with larger variances are generally considered to be highly correlated with the target variable and can thus be preferentially selected.
[0039] MI (Mutual Information) is a feature selection method based on information theory used to quantify the dependence relationship between two variables. Mutual information measures the reduction in uncertainty of a feature for the target variable. The larger the mutual information value, the stronger the explanatory power of the feature for the target variable. In feature selection, the MI method can effectively capture non-linear relationships and is thus suitable for the analysis of complex data.
[0040] Random Forest (RM) is a machine learning algorithm based on decision tree ensembles. The feature selection process of Random Forest relies on its built-in feature importance score, that is, by calculating the contribution of each feature to the improvement of prediction accuracy when splitting decision trees, to evaluate the relevance of the feature. Variables with higher feature importance values are considered more critical. Random Forest shows strong robustness in feature selection, especially having significant advantages for high-dimensional data and non-linear relationships.
[0041] XGBoost (eXtreme Gradient Boosting) is an improved algorithm based on Gradient Boosting Decision Trees (GBDT), with high efficiency and powerful feature selection capabilities. In feature selection, XGBoost evaluates feature importance by calculating the contribution of each feature to the reduction of error in model construction. This algorithm performs well in dealing with sparse data and non-linear features, and at the same time supports custom loss functions and regularization techniques, which can avoid overfitting problems. XGBoost is widely used in the analysis of large-scale and high-dimensional data sets in feature selection.
[0042] Example 4, as Figure 6 shown, based on Example 3, vector quantization processing is performed. During the vector quantization process, the K-means clustering algorithm is used to construct a "codebook". The K-means clustering algorithm iteratively divides the data into multiple clusters, and the center point of each cluster serves as a codeword in the codebook, representing the main regions in the data space. The specific process is as follows: First, the K-means algorithm randomly selects initial codewords (cluster centers), and then continuously adjusts the positions of the cluster centers so that each data point is assigned to the cluster center closest to it. The goal of the algorithm is to minimize the sum of the squared errors between all data points within the cluster and their cluster centers, and finally generate a codebook containing several representative codewords. For any data point xi, it will be mapped to the codeword cj closest to it, that is:
[0043]
[0044] where c j is the j-th codeword in the codebook, ||x i - c j || represents the distance from the data point to the codeword, and K is the size of the codebook. In this way, the original data is mapped to discrete codewords.
[0045] Each data point will be replaced by a discrete codeword, and these codewords represent different regions of the input data in the original feature space. Each codeword represents a "region" in the data space, so these discrete identifiers provide a compact representation of the original data. Since the vector quantization process groups similar data points into the same class, the quantized data has lower dimensions and less noise compared to the original data, which not only simplifies the data complexity but also reduces the redundant information in the data. The quantized data retains the main features and structure of the original data, and the data set is significantly smaller in size compared to the original data, with higher computational efficiency. At the same time, the quantized data removes some noise and unnecessary details, which helps to improve the performance of subsequent models, especially reducing the risk of overfitting.
[0046] Example 5, as Figure 7 shown, on the basis of Example 4, a cross-attention mechanism (Cross-Attention) is added. After quantization processing, the obtained discrete data is used as input and enters the cross-attention mechanism to further screen out important features. Specifically, 4 feature subsets F1, F2, F3, and F4 are obtained by quantization processing. Each subset contains SNP sites screened by different feature selection methods. The phenotypic data P is used as the key matrix K and value matrix V of the cross-attention mechanism to guide the weighted processing of the feature subsets. First, for each feature subset Fi (i = 1, 2, 3, 4), the cross-attention is calculated using the phenotypic data P to convert the relationship between the subset and the phenotypic data into weights:
[0047]
[0048] where Q i is the query matrix of the feature subset Fi, K is the key matrix of the phenotypic data P, V is the value matrix of the phenotypic data P, d k is the dimension of the key, Attention i (F i , P) is the new feature of the subset F i after cross-attention weighting, K T is the transpose of the key matrix. The calculation results are the weighted representations Attention1, Attention2, Attention3, and Attention4 of each feature subset, and these representations respectively reflect the importance of each subset under the guidance of the phenotypic data.
[0049] Then, the importance of features that appear in multiple subsets is comprehensively calculated by the method of weighted fusion, and the 4 weighted feature subsets are combined into a complete feature set. The fusion formula is:
[0050]
[0051] where Score i (f) is the cross-attention score of the feature f in the subset F i , is an indicator function indicating whether the feature f appears in the subset Fi (appears as 1, does not appear as 0), and Weight(f) is the final weighted score of the feature f. According to these weights, a final feature set is obtained.
[0052] In summary, in this study, the data processed by vector quantization, that is, the quantized feature subset, serves as the input to the cross-attention mechanism. Through the cross-attention mechanism, the correlation between each SNP locus and the phenotypic data is calculated, and weights are assigned based on these correlations. The model dynamically adjusts the importance of different SNP loci through these weights, such that the SNP loci highly correlated with the phenotypic data obtain higher weights, while the loci with weaker relationships are weakened or ignored. In this way, a set of weighted features are finally screened out, which are the most relevant parts of the phenotypic data and the genotype data, that is, the finally screened SNP loci.
[0053] The above description is only for the purpose of illustrating the present invention. It should be understood that the present invention is not limited to the above embodiments, and various equivalent forms conforming to the idea of the present invention are within the protection scope of the present invention.
Claims
1. A method for analyzing the association between phenotype and genotype data, characterized in that: The following steps are involved: S1: Collect sample data and preprocess them, remove samples lacking phenotypic data, and then use Z-score standardization, i.e. zero mean standardization, to process the phenotypic data. The standardized data has a mean of 0 and a standard deviation of 1, making the weights of each trait more balanced; Preprocess the genotype data in the sample, filter the SNP data, and remove low-quality SNP sites; Then the preprocessed phenotypic data and genotypic data samples are aligned, and finally the genotypic data is mapped to integer encoding to convert the genotypic data into a format acceptable to the machine learning model; S2: Using four machine learning algorithms, namely ANOVA, Random Forest, MI, and XGBoost, the two data of the samples in S1 were machine learned and analyzed to preliminarily screen out SNP sites with strong correlation with the phenotype, obtain four independent feature subsets, and preliminarily screen out four groups of features with the most predictive ability, thereby improving the performance of the phenotype prediction model and reducing overfitting; S3: Use K-means clustering algorithm to construct a codebook, i.e., perform vector quantization processing on the four independent feature subsets initially screened in S2, and map them into four discrete data feature subsets in the form of discrete code words, thereby removing some data noise and unnecessary details, which helps to improve the performance of the phenotype prediction model and reduce the risk of overfitting; S4: taking the discrete data after quantization in S3 as input, further screening out important features through the cross-attention mechanism, calculating the correlation between each SNP site and the phenotypic data using the cross-attention mechanism, and weighting the features based on the correlation weights, and finally screening out a set of weighted features, which are the most relevant parts of the phenotypic data and the genotypic data, i.e., the SNP sites finally screened out.
2. A method for analyzing the association between phenotype and genotype data according to claim 1, characterized in that: The filtering of SNP data in step S1 specifically includes filtering using bcftools v1.21 tool, and the main criteria include QUAL, i.e., quality score ≥ 40, and missing rate ≤ 10%, i.e., only SNPs with missing rate less than or equal to 10% are retained.
3. A method for analyzing the association between phenotype and genotype data according to claim 1, characterized in that: Mapping the genotype data to integer codes in step S1 specifically includes mapping different genotypes of each SNP site in the genotype data to integer codes, and the specific mapping method includes using "0 / 0" to represent the homozygous genotype of the reference allele, using "0 / 1" or "1 / 0" to represent the heterozygous genotype including a reference allele and an alternative allele, and using "1 / 1" to represent the homozygous genotype of the alternative allele.
4. A method for analyzing the association between phenotype and genotype data according to claim 1, characterized in that: The machine learning and data analysis in step S2 specifically include using GridSearchCV, i.e., grid search and cross-validation to perform hyperparameter optimization to improve the performance of the four machine learning algorithms in S2. After feature selection of the four machine learning algorithms in S2, each algorithm will output a set of features, i.e., scores of SNP sites. Based on the scores, SNP sites with a strong correlation with the phenotype are screened out to obtain four independent feature subsets, thereby providing sufficient prediction information for phenotype prediction model training.
5. A method for analyzing the association between phenotype and genotype data according to claim 1, characterized in that: The step S4 further screens the body by the cross-attention mechanism, including recording the four discrete data feature subsets in S3 as F1, F2, F3, and F4, each subset contains SNP sites selected from different feature selection methods, and the phenotypic data is recorded as P as the key matrix K and value matrix V of the cross-attention mechanism, which are used to guide the weighted processing of the four discrete data feature subsets. First, for each feature subset Fi (i=1, 2, 3, 4), the cross-attention is calculated using the phenotypic data P, and the relationship between the subset and the phenotypic data is converted into a weight. The conversion formula is as follows: Among them, Qi is the query matrix of the feature subset Fi, K is the key matrix of the phenotypic data P, V is the value matrix of the phenotypic data P, and d k is the dimension of the key, Attention i (F i ,P) is the subset F after cross attention weighting i The new feature of K T It is the transpose of the key matrix, and the calculation result is the weighted representation of each feature subset, Attention1, Attention2, Attention3, Attention4, which respectively reflects the importance of each subset under the guidance of phenotypic data; Then, the importance of features appearing in multiple subsets is comprehensively calculated by weighted fusion method, and the four weighted feature subsets are merged into a complete feature set. The fusion formula is: Score i (f) is the cross attention score of feature f in subset Fi, An indicator function that indicates whether feature f appears in the subset Fi, if it appears, it is 1, if not, it is 0; Weight(f) is the final weighted score of feature f, and based on these weights, a final feature set is obtained.