Intelligent analysis method based on high-throughput phenotype data
Through the integration of data preprocessing, principal component analysis and intelligent analysis of high-throughput phenotypic data constructed by the evolutionary tree, the problems of complex operations and difficult results integration in the existing technology are solved, efficient and easy-to-use genotype data analysis is achieved, and research efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510781549.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing genotype data analysis methods are complex in operation, lack integration, and are difficult to meet the usage needs of non-professional personnel. The integration of results between different tools is difficult, which affects research efficiency and data interpretation.
It provides an intelligent analysis method based on high-throughput phenotypic data, integrating data preprocessing, principal component analysis and evolution tree construction, and implementing one-stop intelligent analysis through the web interface, supporting data preprocessing, principal component analysis, evolution tree construction and visual display.
降低了分析门槛,提高了数据质量和结果可解释性,显著提高了研究效率和准确性,适用于基因组大数据分析、群体结构研究和分子育种等领域。
Smart Images

Figure CN120299520A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of gene analysis, and particularly relates to an intelligent analysis method based on high-throughput phenotype data. Background Art
[0002] With the rapid development of high-throughput sequencing technology, it has become more convenient and economical to obtain the genomic information of organisms, generating a large amount of genotype data. However, these genotype data often contain missing values, error information, and redundant information, posing challenges to data analysis and population genetics research. Traditional genotype data processing and analysis methods usually require the use of multiple software and tools, with complex operations and inconsistent analysis processes, making it difficult to meet the usage needs of non-professionals.
[0003] Most existing genotype data analysis methods focus on single analysis functions, such as principal component analysis or population structure analysis, lacking an integrated solution; this results in researchers having to frequently switch between different analysis tools, increasing the workload of data conversion and format adjustment, reducing research efficiency. At the same time, it is also difficult to integrate the results between different analysis tools, making it difficult to form a unified data interpretation and visualization display; in addition, most current genotype data analysis tools are command-line tools, requiring a high level of programming ability from users, lacking a friendly graphical interface and interactive operations, restricting the application scope of the tools, making it difficult to support researchers in interpreting and mining complex genomic data, and affecting the research progress in fields such as population genetics and molecular breeding. Summary of the Invention
[0004] The present invention provides an intelligent analysis method based on high-throughput phenotype data, which realizes one-stop intelligent analysis and processing of high-throughput genotype data by integrating functions such as data preprocessing, principal component analysis, and phylogenetic tree construction, reducing the analysis threshold, improving data quality, and enhancing the interpretability of results.
[0005] To achieve the above-mentioned invention purpose, the specific technical solutions are as follows: An intelligent analysis method based on high-throughput phenotype data, the method comprising the following steps: Step S1, preprocess the high-throughput genotype data, calculate the missing rate and heterozygosity of each sample and SNP locus, and filter the genotype data.
[0006] Step S2, perform principal component analysis (PCA) based on the filtered genotype data to reduce the data dimension through orthogonal transformation; construct a cladogram to analyze the evolutionary relationship between samples within the population. Step S3, output the analysis results and visualize the display.
[0007] Further, the preprocessing of the high-throughput genotype data includes: Reading the genotype data file in HMP or VCF format, where the HMP format contains: rs (SNP name), alleles (SNP alleles), chrom (chromosome number), pos (physical position) information; the VCF format contains: CHROM (chromosome), POS (genomic position), ID (variant site ID), REF (reference base), ALT (variant base), FORMAT (genotype format) information.
[0008] Calculating the missing rate (Missing Rate, MR) and heterozygosity (HeterozygosityRate, HR) of each SNP locus; the calculation formula for the locus missing rate is: , where is the number of missing samples at the locus, is the total number of samples; the calculation formula for the locus heterozygosity is: , where is the number of heterozygous genotype samples at the locus; Calculating the missing rate of each sample and heterozygosity ; sample missing rate: , where is the sample the number of missing loci in, is the total number of loci; sample heterozygosity: , where is the sample the number of heterozygous genotype loci in; Calculating the minor allele frequency (Minor Allele Frequency, MAF) of each SNP locus, and its calculation formula is: , where and are the percentages of the major allele and minor allele of this SNP locus in the population, respectively, and .
[0009] Filtering the genotype data according to the preset threshold, setting the sample missing rate threshold and sample heterozygosity threshold to 10%, and removing the samples with a sample missing rate threshold or sample heterozygosity threshold greater than 10%.
[0010] Setting the locus missing rate threshold to 10%, removing the SNP loci with a locus missing rate threshold greater than 10%, setting the minor allele frequency threshold to 5%, and removing the SNP loci with a minor allele frequency less than 5%; generating a filtered genotype data table as the input data for subsequent analysis.
[0011] Further, the principal component analysis (PCA) adopts the following calculation method: Based on the input data, construct a genotype data matrix , which is an n×m matrix, where n is the number of samples, m is the number of SNP loci, and xij represents the genotype value of the j-th SNP locus in the i-th sample; Centrally process the genotype data matrix X to obtain a matrix X̃ , where x̃ij = xij - x̄j, and x̄j is the average genotype value of the j-th SNP locus: x̄j = (1 / n) ∑i=1n xij; Calculate the covariance matrix C : C = (1 / (n - 1)) X̃TX̃ , where C is an m×m matrix, and X̃T is the transpose matrix of X̃; Perform eigenvalue decomposition on the covariance matrix C to solve for the eigenvalues λi and eigenvectors ui , i = 1, 2, …, m , where λi is the eigenvalue and ui is the corresponding eigenvector; sort the eigenvalues λi in descending order and select the first k eigenvectors to form a projection matrix U , where U is an m×k matrix and k is the number of principal components; Calculate the contribution rate corresponding to each eigenvalue. The contribution rate of the k-th principal component is: ηk = λk / ∑i=1m λi , and the cumulative contribution rate of the first k principal components is: ∑i=1k ηi = (∑i=1k λi) / ∑i=1m λi ; calculate the principal component score matrix Y , where Y is an n×k matrix, the j-th column of Y represents the k-th principal component, and the i-th row of Y
[0012] represents the projection coordinates of the i-th sample on the k principal components.Furthermore, the encoding method of genotype values includes: homozygous major allele (such as AA) is encoded as 0, heterozygous genotype (such as Aa) is encoded as 1, and homozygous minor allele (such as aa) is encoded as 2; missing values are encoded as NA and mean imputation is performed.
[0013] Furthermore, constructing the phylogenetic tree and analyzing the evolutionary relationships among samples within the population includes: Based on the input data, calculate the genetic distance matrix between samples , which is a symmetric matrix of dimension , where is the number of samples, denotes the genetic distance between the -th sample and the -th sample, and is calculated using the Shared Allele Distance (SAD): , where is the number of shared non-missing SNP loci, denotes the proportion of shared alleles between the -th sample and the -th sample at the -th SNP locus. When the genotypes of two samples are the same , completely different ; Based on the genetic distance matrix , construct a hierarchical clustering tree, and according to the preset number of groups , , cut the hierarchical clustering tree to divide the samples into groups and generate a phylogenetic tree structure file.
[0014] Generate a phylogenetic tree structure file, which contains the following information: sample names and their corresponding group numbers (from 1 to N); the lengths of each branch, representing the relative magnitude of genetic distance; the bootstrap support rate of each node, indicating the probability of the branch appearing in resampling; and the topological structure of the tree, reflecting the evolutionary relationships among samples.
[0015] Furthermore, applying the hierarchical clustering algorithm (Hierarchical Clustering), based on the genetic distance matrix , construct a tree structure. The specific steps are as follows: Initially, each sample is an independent class; Calculate the distances between classes, find the two classes with the smallest distance, and merge them into a new class; Update the distances between the new class and other classes , , where and are the centroids of classes and respectively; Repeat the above merging process until all samples are grouped into one class.
[0016] Furthermore, when cutting the hierarchical clustering tree according to the preset number of groups , the cutting principle is: maximizing the genetic distance between groups , minimizing the genetic distance within groups , and maximizing the silhouette coefficient (Silhouette Coefficient) of the group division; Among them, is the average distance between sample and other samples in the same group, is the average distance between sample and samples in the nearest neighbor group.
[0017] Furthermore, the visualization display includes: generating an allele frequency distribution map to show the distribution of MAF; generating an SNP density map to show the distribution density of filtered SNPs on each chromosome; generating a PCA score scatter plot with different principal components as coordinate axes to show the genetic relationship between samples; generating a phylogenetic tree diagram to show the evolutionary relationship between samples in a tree structure; generating a population structure bar chart to distinguish different subgroups with different colors and show the subgroup composition ratio of each sample.
[0018] Furthermore, the method realizes interactive operations through a web interface, and users can access and upload data files through a browser, adjust analysis parameters, and download analysis results.
[0019] Furthermore, the method supports outputting analysis result files in multiple formats, including: a filtered genotype data table in HMP format; a PCA score matrix, a phylogenetic tree grouping result, and a population structure ratio matrix in CSV format; and visualization graphics in PNG or PDF format.
[0020] Compared with the prior art, the beneficial effects of the present invention are: The method of the present invention realizes one-stop analysis and processing of genotype data, including the integration of a series of functions such as data preprocessing, principal component analysis, phylogenetic tree construction, and visualization display; the method realizes interactive operations through a web interface, reducing the user's usage threshold; the present invention has important value in the fields of genomic big data analysis, population structure research, and molecular breeding applications, etc., and can significantly improve the efficiency and accuracy of related research. Brief Description of the Drawings
[0021] Figure 1 It is a flowchart of an intelligent analysis method for high-throughput phenotypic data according to the present invention. Detailed Embodiments
[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0023] As Figure 1 shown, it is a flowchart of an intelligent analysis method for high-throughput phenotypic data according to the present invention, and the method includes the following steps: Step S1, preprocess the high-throughput genotype data, calculate the missing rate and heterozygosity of each sample and SNP locus, and filter the genotype data.
[0024] The high-throughput genotype data described in the present invention includes two common formats: HMP (HapMap) and VCF (Variant Call Format). Among them, the HMP format is a common format for storing sequence data, with SNP information stored in rows and sample information stored in columns. The first 11 columns are the attributes of SNPs (rs, alleles, chrom, pos, etc.), and the remaining columns are the genotype information of each sample at each SNP locus. The VCF format contains information such as CHROM (chromosome), POS (genomic position), ID (variant locus ID), REF (reference base), ALT (variant base), FORMAT (genotype format), etc.; it can automatically identify and process input files in these two formats, improving data compatibility and analysis efficiency.
[0025] Step S2, perform principal component analysis based on the filtered genotype data, reduce the data dimension through orthogonal transformation; construct a phylogenetic tree to analyze the evolutionary relationship among samples within the population. Step S3, output the analysis results and visualize them.
[0026] The preprocessing of the high-throughput genotype data includes: Calculating the missing rate MR and the heterozygosity HR of each SNP locus; the calculation formula for the locus missing rate is: , where is the number of missing samples at the locus, is the total number of samples; the calculation formula for the locus heterozygosity is: , where is the number of samples with heterozygous genotypes at the locus; Calculating the missing rate and the heterozygosity of each sample; sample missing rate: , where is the sample the number of missing loci in, is the total number of loci; sample heterozygosity: , where is the sample the number of loci with heterozygous genotypes in.
[0027] Calculating the missing rate and the heterozygosity are important indicators for evaluating data quality. The higher the locus missing rate, the worse the sequencing quality of the locus; the higher the sample missing rate, the more likely there are problems with the DNA quality or sequencing depth of the sample; and the heterozygosity refers to the proportion of heterozygous genotypes in an individual's genome. Too high or too low heterozygosity may indicate problems such as sample contamination or insufficient inbred line purity.
[0028] Calculating the minor allele frequency MAF of each SNP locus, and its calculation formula is: , where and are the percentages of the major allele and the minor allele of the SNP locus in the population respectively, and . The minor allele frequency is an important indicator for evaluating the polymorphism of SNP loci in population genetics. The larger the MAF value, the higher the degree of variation of the locus in the population and the richer the information content; SNP loci with low MAF come from sequencing errors or only exist in very few individuals and provide limited information in population genetics analysis.
[0029] Screening the genotype data according to a preset threshold. The sample missing rate threshold and the sample heterozygosity threshold are set to 10%, and samples with a sample missing rate threshold or a sample heterozygosity threshold greater than 10% are excluded; setting it to 10% is a relatively common empirical value that can exclude samples with poor quality while retaining most of the effective samples.
[0030] The site deletion rate threshold is set to 10%, and SNP sites with a site deletion rate greater than 10% are excluded. The minor allele frequency threshold is set to 5%, and SNP sites with a minor allele frequency less than 5% are excluded; a filtered genotype data table is generated as the input data for subsequent analysis.
[0031] The principal component analysis PCA uses the following calculation method: Based on the input data, a genotype data matrix is constructed , is an n×m matrix, where n is the number of samples, m is the number of SNP sites, and gij represents the genotype value of the ith sample at the jth SNP site; The genotype data matrix is centered to obtain the matrix G', where : g'ij = gij - gj, gj is the average genotype value of the jth SNP site: The covariance matrix C is calculated as: C = G'GT / (n - 1), where C is an m×m matrix, and GT is the transpose matrix of G; The covariance matrix C is eigen-decomposed to solve for the eigenvalues λi and eigenvectors vi, where λi is the eigenvalue and vi is the corresponding eigenvector; the eigenvectors are sorted in descending order of the eigenvalue λi, and the first One main component, The row represents the projection coordinates of the th sample on the main components; by calculating the contribution rate of the main components, the degree of explanation of each main component for the variation of the original data can be evaluated. Usually, the first few main components with a cumulative contribution rate reaching 70%-90% are selected for subsequent analysis and visualization.
[0032] The coding methods of genotype values include: homozygous major allele (such as AA) is coded as 0, heterozygous genotype (such as Aa) is coded as 1, homozygous minor allele (such as aa) is coded as 2; missing values are coded as NA and mean imputation is performed. For missing values, they are coded as NA and mean imputation is used, that is, the mean value of all non-missing samples at this locus is used to replace the missing value, which is a common method for handling missing values in high-throughput genotype data.
[0033] Constructing the phylogenetic tree and analyzing the evolutionary relationships among samples within the population include: Based on the input data, calculate the genetic distance matrix among samples , which is a dimensional symmetric matrix, where is the number of samples, represents the genetic distance between the th sample and the th sample, calculated using the shared allele distance: where is the number of shared non-missing SNP loci, represents the proportion of shared alleles between the th sample and the th sample at the th SNP locus. When the genotypes of two samples are the same , According to the genetic distance matrix construct a hierarchical clustering tree, and according to the preset number of groups , , cut the hierarchical clustering tree to divide the samples into groups and generate a phylogenetic tree structure file.
[0034] Generate a phylogenetic tree structure file, which contains the following information: sample names and their corresponding group numbers (from 1 to N); the lengths of each branch, representing the relative magnitude of genetic distance; the bootstrap support rate of each node, indicating the probability of the appearance of this branch in resampling; the topological structure of the tree, reflecting the evolutionary relationships among samples.
[0035] Apply the hierarchical clustering algorithm based on the genetic distance matrix Construct a dendrogram, and the specific steps are as follows: Initially, each sample is taken as an independent class; Calculate the distances between classes, and find the two classes with the smallest distance and merge them into a new class; Update the distances between the new class and other classes , , where and are the centroids of class and class respectively; Repeat the above merging process until all samples are grouped into one class.
[0036] When cutting the hierarchical clustering tree according to the preset number of groups The cutting principle is: maximizing the genetic distance between populations , minimizing the genetic distance within populations , and maximizing the silhouette coefficient (SilhouetteCoefficient) of the population division; Among them, is the average distance between sample and other samples in the same population, is the average distance between sample and samples in the nearest neighbor population.
[0037] The said visualization display includes: generating an allele frequency distribution map to show the distribution of MAF; generating an SNP density map to show the distribution density of filtered SNPs on each chromosome; generating a PCA score scatter plot with different principal components as coordinate axes to show the genetic relationship between samples; generating a phylogenetic tree map to show the evolutionary relationship between samples in a tree structure; generating a population structure bar chart to distinguish different subgroups with different colors and show the subgroup composition ratio of each sample.
[0038] The allele frequency distribution map usually shows the distribution characteristics of MAF in the form of a histogram or density map, which helps to evaluate the polymorphism level of the data; the SNP density map shows the distribution density of SNPs on each chromosome in the form of a bar chart or heat map, reflecting the spatial distribution characteristics of genomic variations; the PCA score scatter plot uses the first few principal components as the coordinate axes to visually show the distribution of samples in the dimensionality-reduced space, and different colors or shapes can represent samples from different groups or sources; the phylogenetic tree diagram shows the evolutionary relationships between samples in a tree structure, and the branch length represents the genetic distance, and different groups can be marked with different colors; the population structure bar chart shows the proportion of subgroups of each sample in the form of a stacked bar chart, where different colors represent different subgroups and the bar length represents the proportion size.
[0039] The method realizes interactive operations through a web interface. Users can access and upload data files through a browser, adjust analysis parameters, and download analysis results.
[0040] The method supports outputting analysis result files in multiple formats, including: the filtered genotype data table in HMP format; the PCA score matrix, phylogenetic tree grouping results, and population structure proportion matrix in CSV format; and the visualization graphics in PNG or PDF format.
[0041] The method of the present invention can also perform population structure (Structure) analysis based on the filtered genotype data, and preset the number of subgroups ( ) The value can be determined by the following methods: Based on the PCA analysis result, use the Scree Plot to determine the inflection point position; based on the phylogenetic tree analysis result, refer to the number of main branches.
[0042] Use the cross-validation method to calculate the prediction errors corresponding to different values, and select the value with the smallest error Calculate the Bayesian information criterion: or the Akaike information criterion: , select the
[0043] Assume that each subgroup corresponds to a set of allele frequency vectors , where represents the reference allele frequency of subgroup at the th SNP locus, is the number of SNP loci.
[0044] Each sample has a set of subgroup proportion vectors , representing a sample The probability ratios from each subgroup satisfy the constraint conditions: a) , for all and ; b) , for all , for each sample and SNP locus , calculate the probability of observing genotype based on the mixture model: , where is the allele frequency in subgroup under the condition of observing genotype , for a biallelic SNP: a) (Homozygous major allele); b) (Heterozygous genotype); c) (Homozygous minor allele).
[0045] Optimize the model parameters, and the iterative steps are: E step: Fix the allele frequency vector , and update the subgroup proportion vector ; M step: Fix the subgroup proportion vector , and update the allele frequency vector ; Calculate the value of the log-likelihood function: ; Repeat the E step and M step until the change in the log-likelihood function is less than the preset threshold (default ) or reach the maximum number of iterations (default 100 times).
[0046] Generate the population structure proportion matrix , where is a -dimensional matrix, and the element in the -th row and -th column of represents the probability that the -th sample belongs to the -th subgroup, which is used for subsequent visualization and analysis.
[0047] And determine the optimal value by the method ( ).
[0048] This method realizes a friendly human-computer interaction experience through a web interface developed based on Shiny; users do not need to master complex programming skills. They only need to access the specified website through a browser to upload data files, set analysis parameters, obtain analysis results, and visualize graphics.
[0049] The web interface adopts a responsive design, supports real-time parameter adjustment and result preview, greatly improving the analysis efficiency and user experience. In addition, the web interface also provides detailed operation tips and parameter descriptions, reducing the user's learning cost. In terms of data processing, the system adopts an efficient R language backend computing engine, which can process large-scale genotype data and improves the computing speed through optimization means such as parallel computing. For a genotype data file of 1GB in size, the total processing time from data filtering to population structure analysis is about 20 minutes, improving the analysis efficiency compared with traditional analysis methods.
[0050] To further verify the effectiveness of the method of the present invention, a specific application example is given below.
[0051] This example takes rice ( Oryza sativa L.) germplasm resources as the research object, obtains genotype data by using high-throughput sequencing technology, and analyzes it by the method of the present invention; 120 rice varieties from different geographical regions are selected in this experiment, including 65 indica rice ( indica ), 45 japonica rice ( japonica ), and 10 intermediate types. The whole-genome SNP data is obtained by using gene chip technology, and the initial data contains about 85,000 SNP markers and is stored in VCF format.
[0052] The original data in VCF format is uploaded to the web interface for data preprocessing; after calculation, the average missing rate of SNP loci in the original data is 4.8%, the average missing rate of samples is 5.2%, the average heterozygosity of SNP loci is 12.3%, and the average heterozygosity of samples is 11.7%. Data filtering is carried out according to the preset thresholds: the sample missing rate threshold is set to 10%, the sample heterozygosity threshold is set to 10%, the locus missing rate threshold is set to 10%, and the minor allele frequency threshold is set to 5%. After filtering, 3 samples with a missing rate or heterozygosity exceeding the threshold are removed, and 117 samples are retained; 5,423 SNP loci with a missing rate greater than 10% are removed; 28,752 SNP loci with a MAF less than 5% are removed; the final number of effective SNP markers retained is 50,825. The quality of the filtered data is significantly improved. The average missing rate of SNP loci drops to 2.1%, the average missing rate of samples drops to 3.4%, the median of the minor allele frequency is 0.23, and the SNP markers are evenly distributed on 12 chromosomes, with an average of 4,235 markers per chromosome.
[0053] Perform principal component analysis on the filtered genotype data, and calculate the first 10 principal components and their contribution rates. The results show that: the contribution rate of the first principal component (PC1) is 32.6%, mainly reflecting the indica-japonica differentiation; the contribution rate of the second principal component (PC2) is 15.2%, mainly reflecting the geographical regional differentiation; the contribution rate of the third principal component (PC3) is 8.7%, possibly reflecting the difference in cultivation history; the cumulative contribution rate of the first five principal components reaches 68.4%, which can explain most of the genetic variation among samples. The PCA scatter plot clearly divides 117 rice varieties into three main groups: indica rice group, japonica rice group, and intermediate group. This is highly consistent with the traditional morphological classification, verifying the accuracy of the method; in particular, this method can correctly identify varieties with atypical morphological characteristics but clear genetic backgrounds, which is difficult to achieve by traditional classification methods.
[0054] Construct a phylogenetic tree of rice varieties based on the shared allele distance, set the number of groups N = 3, and divide the samples into three groups. The results show that: Group 1 contains 63 samples, which highly coincides with the indica rice group, with a coincidence rate of 96.9%; Group 2 contains 43 samples, which highly coincides with the japonica rice group, with a coincidence rate of 95.6%; Group 3 contains 11 samples, mainly intermediate types and a small number of difficult-to-classify varieties. The topological structure of the phylogenetic tree clearly shows the genetic relationship among rice varieties, and varieties from the same geographical origin tend to cluster on adjacent branches. Bootstrap support rate analysis shows that the support rates of the main branches are all higher than 85%, indicating that the clustering results have high reliability.
[0055] Perform population structure analysis on rice varieties, and determine the optimal number of subpopulations K = 3 through cross-validation and the ∆K method. The analysis results show that: Subpopulation 1 is mainly composed of indica rice varieties, and the Q1 values of pure indica rice varieties are generally higher than 0.85; Subpopulation 2 is mainly composed of japonica rice varieties, and the Q2 values of pure japonica rice varieties are generally higher than 0.80; Subpopulation 3 mainly contains intermediate types and varieties with mixed ancestries. Population structure analysis further refines the genetic relationship among varieties, especially identifying 9 varieties with obvious mixed genetic backgrounds, which show that the proportions of multiple subpopulation components are close in the Q-value matrix. This discovery is of great significance for understanding the domestication history of rice and the variety improvement path.
[0056] This method is not only applicable to rice, but also can be extended and applied to genomic research of various crops such as wheat, corn, and soybeans. It has broad application prospects in breeding processes such as germplasm resource evaluation, parental selection, and molecular marker-assisted selection. Especially in the field of molecular design breeding, this method can effectively integrate genotype and phenotype data, provide technical support for cultivating new varieties with high yield, high quality, and stress resistance, and ultimately promote agricultural sustainable development and food security.
[0057] This embodiment fully demonstrates the effectiveness and superiority of this method in actual breeding applications. It can not only accurately predict the phenotypic traits of advanced inbred lines, but also significantly improve breeding efficiency, save breeding resources, and accelerate the cultivation process of excellent varieties.
[0058] The specific implementation manners described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An intelligent analysis method based on high-throughput phenotypic data, characterized in that, The method comprises the following steps: Step S1, preprocessing the high-throughput genotype data, calculating the deletion rate and heterozygosity of each sample and SNP site, and filtering the genotype data; Step S2, performing principal component analysis based on the filtered genotype data and reducing the data dimension through orthogonal transformation; Construct evolutionary trees to analyze the evolutionary relationships between samples within a population; Step S3: output the analysis results for visual display.
2. The intelligent analysis method based on high-throughput phenotypic data according to claim 1, characterized in that The preprocessing of high-throughput genotype data includes: Calculate the missing rate MR and heterozygosity HR for each SNP locus; , where is the number of missing samples at the locus, is the total number of samples; , where is the number of samples with heterozygous genotypes at the locus; Calculate the missing rate of each sample and heterozygosity ; , where is the number of missing loci in sample , and is the total number of loci; , where is the number of heterozygous genotype loci in sample ; Calculate the minor allele frequency MAF of each SNP locus: , where and are the percentages of the major allele and minor allele of this SNP locus in the population, respectively, and .
3. The intelligent analysis method based on high-throughput phenotypic data according to claim 2, wherein Filter genotype data based on preset thresholds; The sample missing rate threshold and sample heterozygosity threshold were set at 10%, and samples with a sample missing rate threshold or sample heterozygosity threshold greater than 10% were excluded; the site missing rate threshold was set at 10%, and SNP sites with a site missing rate threshold greater than 10% were excluded; the minor allele frequency threshold was set at 5%, and SNP sites with a minor allele frequency less than 5% were excluded; the filtered genotype data table was generated as input data for subsequent analysis.
4. The intelligent analysis method based on high-throughput phenotypic data according to claim 3, wherein Principal component analysis PCA uses the following calculation method: Construct a genotype data matrix based on the input data , is a dimensional matrix, where is the number of samples, is the number of SNP loci, represents the genotype value of the th sample at the th SNP locus; Center the genotype data matrix to obtain matrix , where , is the average genotype value of the th SNP locus: ; Calculate the covariance matrix : , is a matrix of dimension is the transpose matrix of For the covariance matrix perform eigen decomposition to solve for the eigenvalues and eigenvectors , , where is the eigenvalue is the corresponding eigenvector; sort in descending order by the magnitude of the eigenvalue and select the first eigenvectors to form the projection matrix , is a matrix of dimension and the number of principal components, with a value range of 2 - 10; Calculate the contribution rate corresponding to each eigenvalue. The contribution rate of the th principal component is: . The cumulative contribution rate of the first principal components is: ; Calculate the principal component score matrix . is a -dimensional matrix. The th column of represents the th principal component. The th row of represents the projection coordinates of the th sample on the principal components.
5. The intelligent analysis method based on high-throughput phenotypic data according to claim 4, characterized in that The encoding methods of genotype values include: homozygous major allele is encoded as 0, heterozygous genotype is encoded as 1, and homozygous minor allele is encoded as 2; missing values are encoded as NA and mean interpolation is performed.
6. The intelligent analysis method based on high-throughput phenotype data according to claim 5, wherein The construction of the evolutionary tree and the analysis of the evolutionary relationship between samples in the population include: Calculate the genetic distance matrix between samples based on the input data , is a symmetric matrix of dimension , where represents the number of samples, represents the genetic distance between the -th sample and the -th sample. The genetic distance is calculated using the shared allele distance: , where is the number of shared non-missing SNP loci, represents the proportion of shared alleles between the -th sample and the -th sample at the -th SNP locus. When the genotypes of the two samples are the same , completely different ; According to the genetic distance matrix Construct a hierarchical clustering tree, and according to the preset number of groups , , cut the hierarchical clustering tree, and divide the samples into populations to generate an evolutionary tree structure file.
7. The intelligent analysis method based on high-throughput phenotype data according to claim 6, characterized in that According to the genetic distance matrix Construct a tree structure. The specific steps are as follows: Initially, each sample is treated as an independent class; Calculate the distance between each class, find the two classes with the smallest distance and merge them into a new class; Update the distance between the new class and other classes , , where and are the centroids of classes and class respectively; Repeat the above merging process until all samples are classified into one category.
8. The intelligent analysis method based on high-throughput phenotype data according to claim 6, characterized in that According to the preset number of groups When cutting the hierarchical clustering tree, the cutting principle is: maximizing the genetic distance between groups , minimizing the genetic distance within groups , and the silhouette coefficient of group division is maximized; Among them, is the sample and the average distance from other samples in the same population, is the sample and the average distance from the nearest neighbor population sample.
9. The intelligent analysis method based on high-throughput phenotypic data according to claim 7, wherein The visualization display includes: generating an allele frequency distribution map to display the distribution of MAF; generating a SNP density map to display the distribution density of the filtered SNP on each chromosome; generating a PCA score scatter plot to display the genetic relationship between samples with different principal components as coordinate axes; generating an evolutionary tree diagram to display the evolutionary relationship between samples with a tree structure; generating a population structure bar chart to distinguish different subgroups with different colors and display the subgroup composition ratio of each sample.
10. The intelligent analysis method based on high-throughput phenotype data according to claim 9, characterized in that The method supports the output of analysis result files in multiple formats, including: a filtered genotype data table in HMP format; a PCA score matrix, evolutionary tree grouping results, and a population structure ratio matrix in CSV format; and visualization graphics in PNG or PDF format.
Citation Information
Patent Citations
MiRNA and disease incidence relation prediction method based on graph convolutional network
CN114496092A
Hu sheep purebred evaluation method adopting SNP (Single Nucleotide Polymorphism) molecular marker
CN117487930A
Method for analyzing genetic diversity of camellia sinensis based on SNP molecular marker
CN118726557A
Poplar SNP molecular marker combination, whole genome liquid phase chip prepared from poplar SNP molecular marker combination and application of poplar SNP molecular marker combination
CN118895377A
Cloud black goat 40K liquid phase chip and application thereof
CN119040474A