A method and system for analyzing and managing multi-breed genomic data of beef cattle
Through the collection, integration, preprocessing, feature analysis and visualization of genomic data of multiple beef cattle breeds, the problem of genomic data management of multiple beef cattle breeds has been solved, and efficient data support and genetic breeding progress have been achieved.
Patent Information
- Application Number
- CN202510329839.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-03-19
AI Technical Summary
Existing technologies make it difficult to conduct comprehensive, systematic, and efficient analysis and management of genomic data from multiple beef cattle breeds, resulting in difficulty in achieving breakthrough progress in beef cattle genetic breeding research and practice.
A multi-breed genomic data analysis and management method for beef cattle is adopted, including data collection and integration, data preprocessing and standardization, multi-breed genomic feature analysis, and data visualization and report generation. By identifying breeds with rich and poor genetic diversity and mining trait association sites, dynamically updated data management is achieved.
It has achieved comprehensive collection, integration, preprocessing, feature analysis and visual display of genomic data of multiple beef cattle breeds, provided scientific and efficient data support, and provided a powerful tool for beef cattle genetic breeding.
Smart Images

Figure CN120236660B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of beef cattle gene analysis methods, and in particular to a method and system for analyzing and managing genome data of multiple beef cattle breeds. Background Art
[0002] With the rapid development of gene sequencing technology, a vast amount of genomic data has been acquired from beef cattle. Different breeds of beef cattle exhibit significant differences in growth rate, meat quality, and disease resistance, with complex genomic underpinnings underlying these differences. However, there is currently a lack of systematic, efficient, and comprehensive methods for analyzing and managing genomic data from multiple beef cattle breeds. Existing analyses are often limited to simple comparisons of a single breed or a few breeds, making it difficult to deeply explore the commonalities and characteristics across multiple breeds and to fully integrate and utilize the vast amount of genomic data resources. This has resulted in limited breakthroughs in research and practice related to beef cattle genetics and breeding, hindering the efficient development of the beef cattle industry. Summary of the Invention
[0003] The technical problem to be solved by the present invention is how to provide a beef cattle multi-breed genomic data analysis and management method and system that can realize comprehensive, systematic and efficient analysis and management of beef cattle multi-breed genomic data and provide strong support for beef cattle genetic breeding.
[0004] To solve the above technical problems, the technical solution adopted by the present invention is: a method for analyzing and managing genomic data of multiple breeds of beef cattle, comprising the following steps:
[0005] S1, data collection and integration: obtain raw genomic sequence data of beef cattle, establish a genomic data repository, and add unique identification tags to the data of each breed;
[0006] S2, data preprocessing and standardization: detect variant points and perform standardization on the detected variant sites, and encode the genotype data of the variant sites;
[0007] S3, Multi-breed genomic characterization: Identify genetically diverse beef cattle breeds and those with relatively limited genetic diversity, identify potential breed subgroups or breed combinations with similar genetic backgrounds, and identify trait association loci that are common across breeds or specific to a particular breed.
[0008] S4, Data visualization and report generation: Visualize the analysis results.
[0009] The present invention also discloses a beef cattle multi-breed genome data analysis and management system, comprising:
[0010] Data collection and integration module: used to obtain raw genomic sequence data of beef cattle, establish a genomic data repository, and add unique identification tags to the data of each breed;
[0011] Data preprocessing and standardization module: used to detect variant points and perform standardization on the detected variant sites, and at the same time, encode the genotype data of the variant sites;
[0012] Multi-breed genomic signature analysis module: used to identify beef cattle breeds with rich genetic diversity and those with relatively scarce genetic diversity, obtain potential breed subgroups or breed combinations with similar genetic backgrounds, and discover trait association sites that are common in different breeds or specific to specific breeds;
[0013] Data visualization and report generation module: used to visualize the analysis results;
[0014] Data update and maintenance module: used to regularly collect new beef cattle genome data, including new breed sample data or data obtained from more in-depth sequencing analysis of existing breeds.
[0015] The beneficial effects of adopting the above technical solution are: the method and system described in the present invention realize the comprehensive collection, integration, preprocessing, feature analysis, visualization, report generation, and data updating and maintenance of genomic data of multiple beef cattle breeds, and can provide comprehensive, accurate and dynamically updated genomic data support for beef cattle genetic breeding research and practice, which will help promote the scientific and efficient development of the beef cattle industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0017] Figure 1 is an overall flow chart of the method according to an embodiment of the present invention;
[0018] Figure 2 is a flow chart of data preprocessing and standardization in the method according to an embodiment of the present invention;
[0019] Figure 3 This is a flow chart of the multi-variety genome feature analysis method according to an embodiment of the present invention;
[0020] Figure 4 is a flowchart of step S3-1 in the method according to an embodiment of the present invention;
[0021] Figure 5 is a flowchart of step S3-2 in the method according to an embodiment of the present invention;
[0022] Figure 6is a flowchart of step S3-3 in the method according to an embodiment of the present invention;
[0023] Figure 7 It is a principle block diagram of the system described in an embodiment of the present invention. DETAILED DESCRIPTION
[0024] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.
[0025] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0026] like Figure 1 As shown, the embodiment of the present invention discloses a method for analyzing and managing genome data of multiple breeds of beef cattle, comprising the following steps:
[0027] S1, data collection and integration: obtain raw genomic sequence data of beef cattle, establish a genomic data repository, and add unique identification tags to the data of each breed;
[0028] S2, data preprocessing and standardization: detect variant points and perform standardization on the detected variant sites, and encode the genotype data of the variant sites;
[0029] S3, Multi-breed genomic characterization: Identify genetically diverse beef cattle breeds and those with relatively limited genetic diversity, identify potential breed subgroups or breed combinations with similar genetic backgrounds, and identify trait association loci that are common across breeds or specific to a particular breed.
[0030] S4, data visualization and report generation: Visualize the results of the analysis;
[0031] S5, data update and maintenance: Regularly collect new beef cattle genome data, including new breed sample data or data obtained from more in-depth sequencing analysis of existing breeds, and repeat steps S1-S4 to achieve dynamic data updates.
[0032] Furthermore, the S1 data collection and integration includes the following steps: collecting genomic samples from multiple different breeds of beef cattle, using gene sequencing technology to perform whole genome sequencing on the collected samples to obtain raw genome sequence data, performing quality control on the raw data, removing low-quality reads, and at the same time, filtering the raw data to exclude possible contaminated sequences; integrating the quality-controlled and filtered genomic data of different breeds of beef cattle into a unified database, establishing a multi-breed beef cattle genomic data repository, and adding a unique identification tag to the data of each breed.
[0033] In the step S1:
[0034] First, using the MySQL database management system, a master project database was created within the database to store all information related to multi-breed beef cattle genomic data. This database was named the Beef Cattle Genomic Data Repository. For each different beef cattle breed, a separate tablespace was created within the master database. For example, a tablespace named Simmental was created for Simmental cattle, and a tablespace named Charolais was created for Charolais cattle. Each tablespace had the same predefined data structure to ensure data consistency and standardization. These structures included, but were not limited to, a sample number field (used to uniquely identify each individual beef cattle sample), a sequencing data field (for storing quality-controlled and filtered genomic sequencing data), a breed identifier field (to clearly identify the beef cattle breed to which the data belongs), and a collection time field (to record the time when the sample was collected, facilitating data traceability and time series analysis).
[0035] During the data integration process, when the genomic data of a certain breed of beef cattle is imported into the database, the data is stored in the tablespace of the corresponding breed according to the pre-set correspondence between the tablespace and the breed; for the data of each sample, a unique number is given in the sample number field, for example, a combination of numbers and letters, such as SM001 for the first sample of Simmental cattle, CH005 for the fifth sample of Charolais cattle, etc.; in the breed identification field, the name of the breed to which the sample belongs is clearly filled in; at the same time, an index is created based on the breed identification field to locate the corresponding tablespace and extract data when all data of a specific breed need to be queried; and an index is created based on the sample number field to obtain detailed information of a specific sample when processing single sample data or performing association analysis between samples.
[0036] Furthermore, the S2 data preprocessing and standardization includes the following steps:
[0037] Genome assembly is performed on the integrated multi-breed genomic data. For beef cattle breeds with reference genomes, the sequencing data are compared with the reference genome, and the comparison operation is performed using comparison software to identify the variant sites in the sequencing data. For beef cattle breeds without reference genomes, a de novo assembly method is adopted, and the genome draft of the breed is constructed using software, based on which the variant sites are detected. The detected variant sites are standardized, and the naming format and coordinate system of the variant sites are unified. At the same time, the genotype data of the variant sites are encoded.
[0038] Further, such as Figure 2 As shown, the S2 data preprocessing and standardization includes the following steps:
[0039] S2-1, Data processing of beef cattle breeds with reference genomes:
[0040] S2-1-1, Genome assembly and alignment:
[0041] Select a reference genome: Obtain a high-quality beef cattle reference genome sequence from a public database (such as the NCBI genome database), such as the UMD3.1 reference genome version of cattle;
[0042] Data alignment: The sequencing data were first indexed using the professional alignment software BWA (Burrows-WheelerAligner). For short-read data generated by the Illumina sequencing platform, the bwa index command was used to index the reference genome. The quality-controlled and filtered sequencing data were then aligned with the reference genome using the bwamem command.
[0043] Alignment result processing: Samtools software was used to process the SAM (Sequence Alignment / Map) format files generated by the BWA alignment. First, the SAM files were converted to BAM (Binary Alignment / Map format). The BAM files were sorted using the samtools view-Sb command, and the reads were sorted according to chromosome position using the samtools sort command.
[0044] S2-1-2, variant site identification:
[0045] Local realignment: Use GATK's RealignerTargetCreator and IndelRealigner tools to perform local realignment of regions where indels may exist, correct alignment errors caused by indels, and improve the accuracy of variant detection.
[0046] Base quality score recalibration: Use GATK's BaseRecalibrator and PrintReads tools to recalibrate base quality scores. Based on known variant site information and sequencing data characteristics, the quality score of each base is re-evaluated to make subsequent variant detection more reliable.
[0047] S2-1-3, mutation site detection:
[0048] Use the GATK HaplotypeCaller tool to detect variant sites, including single nucleotide polymorphisms (SNPs) and indels. Construct local haplotypes and perform statistical analysis to identify variant sites. Set a minimum coverage during the detection process, for example, -minDP is set to 10, to filter out false positive variant sites that may arise from low coverage areas.
[0049] S2-2, Data processing for beef cattle breeds without reference genomes
[0050] S2-2-1, reassembly:
[0051] Data preprocessing: Trimmomatic software was used to remove adapter sequences and low-quality bases in the sequencing data without a reference genome, such as bases with a quality value lower than 20, and the reads were length-filtered to retain reads longer than 50 bp.
[0052] Assembly software selection and operation: De novo assembly software Velvet was used. First, the k-mer value was estimated based on the characteristics of the sequencing data. The k-mer length was determined to be 51 using the k-mer analysis tool. Then, the Velveth command was used to construct the Velvet hash table, and the velvetg command was used to assemble the genome.
[0053] S2-2-2, mutation site detection:
[0054] Evaluation and optimization of assembly results: Use the evaluation tool QUAST to evaluate the quality of the assembled genome draft, check the completeness and continuity of the assembly, and adjust the assembly parameters based on the evaluation results;
[0055] Variant site identification: Using the reference-free genome variation detection software FreeBayes, the quality-assessed and optimized draft genome data was input into FreeBayes, and the minimum allele frequency was set to detect SNPs and Indels variation sites.
[0056] S2-3, variant site standardization and genotype data encoding
[0057] S2-3-1, variant site standardization:
[0058] Format conversion: Use VCFtools software to convert variant site information from the output formats of various detection tools into a unified standard VCF format to ensure that variant site information from different sources is compatible and comparable in subsequent analyses;
[0059] Coordinate system unification: The coordinates of all variant sites are uniformly corrected and adjusted according to the coordinate system of the reference genome, so that the coordinate information of variant sites in different varieties is consistent under the same coordinate system, even if they are obtained using different assembly methods or reference genome versions;
[0060] S2-3-2, genotype data encoding:
[0061] Determine the coding rules: code the homozygous reference genotype as 0, the homozygous variant genotype as 1, and the heterozygous genotype as 2;
[0062] Coding operation: Write a Python script to read the genotype data of the variant site, convert the genotype of each variant site according to the above coding rules, and store the encoded data in a new file or database field for subsequent data analysis and calculation.
[0063] Further, such as Figure 3 As shown, the S3 multi-variety genome feature analysis includes the following steps:
[0064] S3-1: Based on the standardized variant site data, calculate the genetic diversity index of multiple beef cattle genomes to assess the degree of genetic variation within and between different beef cattle populations. By comparing the genetic diversity indexes of different breeds, identify beef cattle breeds with rich genetic diversity and those with relatively poor genetic diversity.
[0065] S3-2: Perform principal component analysis (PCA) to reduce high-dimensional genomic variation data to a two-dimensional space, displaying the distribution relationships of different beef cattle breeds at the genomic level and determining the closeness of relationships and genetic structure differences between breeds. Combined with cluster analysis methods, beef cattle breeds are clustered to determine the grouping of different breeds at the genomic level, identifying potential breed subgroups or breed combinations with similar genetic backgrounds.
[0066] S3-3: Conduct genome-wide association studies (GWAS) on gene regions or functional elements related to important economic traits of beef cattle, use mixed linear models to detect association signals between variant sites and economic traits, and identify candidate genes and variant sites that are significantly associated with target traits; through comprehensive analysis of multi-breed beef cattle data, discover trait association sites that are common in different beef cattle breeds or unique to specific beef cattle breeds, providing key information for molecular marker-assisted selection and breeding of beef cattle.
[0067] Further, such as Figure 4 As shown, the S3-1 specifically includes the following methods:
[0068] S3-1-1) Calculate the genetic diversity index of multiple breeds of beef cattle genomes:
[0069] 1) Calculate heterozygosity:
[0070] Observed heterozygosity Ho: Count the proportion of heterozygous individuals at each site to the total number of individuals; for each variant site, calculate the number of individuals with heterozygous genotype at that site in the sample, and then divide it by the total number of samples to obtain the observed heterozygosity of the site. Finally, calculate the average of the observed heterozygosities of all sites to obtain the observed heterozygosity of the beef cattle population of this breed; a higher observed heterozygosity means that the genetic differences between individuals in the population are greater and the genetic diversity is rich.
[0071] Expected heterozygosity He: According to the Hardy-Weinberg equilibrium law, in an ideal population, the expected heterozygous frequency at each site is calculated by calculating the allele frequency at each site and then using the formula Get the expected heterozygosity of each site, where p i is the frequency of the i-th allele, n is the number of alleles, and the average is calculated to obtain the expected heterozygosity of the population. The expected heterozygosity reflects the genetic diversity that a population may have under random mating conditions. Compared with the observed heterozygosity, it can be used to assess whether the population conforms to the Hardy-Weinberg equilibrium and whether there are factors that affect genetic diversity, such as selection, migration, and mutation.
[0072] 2) Calculate nucleotide diversity (Pi): This method calculates the average number of nucleotide differences at the same position between any two randomly selected nucleotide sequences in the genome. By comparing the nucleotide diversity of different beef cattle populations, the degree of genetic variation between them at the genomic level is assessed. For each breed, the number of nucleotide differences between all two individuals at all variant sites is counted and then divided by the total number of sites compared and the number of individual pairs to obtain the nucleotide diversity value for that breed. Higher nucleotide diversity indicates greater genetic variation within the population and higher genetic diversity of the breed.
[0073] 3) Calculation of allelic richness (AR): The AR is calculated by averaging the number of alleles observed in a given number of samples within each breed of beef cattle. A random number of individuals are selected, and the number of alleles at each locus is counted. The average number of alleles across all loci is then averaged. Allelic richness reflects the abundance of alleles present in a population and is unaffected by sample size. It can be used to compare genetic diversity between populations of varying sample sizes. A higher AR for a breed indicates greater genetic variation and diversity.
[0074] 4) Inbreeding coefficient F: By analyzing the standardized variant site data, the inbreeding coefficient of each individual is calculated, and then the average inbreeding coefficient of the group is obtained; the higher the inbreeding coefficient, the closer the relationship between individuals in the group, the lower the genetic diversity, and the possible risk of inbreeding depression.
[0075] 5) Linkage disequilibrium (LD): The degree of linkage disequilibrium between genes in a population is assessed by calculating the linkage disequilibrium coefficient between different loci. A higher degree of linkage disequilibrium may mean that the genetic diversity of the population is relatively low because some allele combinations are relatively fixed in the population. A lower degree of linkage disequilibrium indicates that the gene combinations in the population are more diverse and the genetic diversity is richer.
[0076] 6) Population differentiation index Fst: Compares the genetic variation within a population with the total genetic variation to obtain a value between 0 and 1. The closer the value is to 1, the greater the genetic differentiation between populations and the more uneven the distribution of genetic diversity among populations; the closer the value is to 0, the smaller the genetic difference between populations and the more evenly distributed genetic diversity among populations;
[0077] By calculating the above genetic diversity indicators, we can comprehensively evaluate the degree of genetic variation within and between different breeds of beef cattle. By comparing the various genetic diversity indicators of different breeds, if a breed performs relatively high in heterozygosity, nucleotide diversity, allele richness and other indicators, while having a low inbreeding coefficient and linkage disequilibrium, and a large Fst value compared with other breeds, then this breed can be identified as a beef cattle breed with rich genetic diversity; conversely, if a breed performs poorly in various indicators, it can be considered a beef cattle breed with relatively poor genetic diversity.
[0078] S3-1-2) Identify beef cattle breeds with high genetic diversity and those with low genetic diversity:
[0079] A certain weight was set for each indicator, among which the weight of comprehensive heterozygosity was 0.2, the weight of nucleotide diversity was 0.2, the weight of allele richness was 0.2, the weight of inbreeding coefficient was 0.1, the weight of linkage disequilibrium was 0.1, and the weight of population differentiation index was 0.2;
[0080] For each variety, a comprehensive score is calculated according to the set weight based on the comparison of its values on various indicators with other varieties. Varieties with scores higher than the set threshold are identified as varieties with rich genetic diversity, and varieties with scores lower than the set threshold are identified as varieties with relatively poor genetic diversity.
[0081] Further, such as Figure 5 As shown, the S3-2 specifically includes the following methods:
[0082] S3-2-1, principal component analysis method PCA:
[0083] First, based on the standardized single nucleotide polymorphism (SNP) data matrix, the rows in the matrix represent different beef cattle individuals, the columns represent various SNP sites, and each element represents the genotype of the individual at a specific SNP site. The mean of each SNP site is subtracted from the data of the site to make the mean of the data zero, eliminating the influence of the dimensional difference of the data on the results of principal component analysis.
[0084] Calculate the covariance matrix according to the nature of the data. For the covariance matrix, the element C i,j represents the covariance between the i-th and j-th SNP sites, and its calculation formula is:
[0085]
[0086] Where n is the number of samples, x ki is the data of the kth individual at the i-th SNP site, is the mean of the ith SNP site; x kj is the data of the kth individual at the jth SNP site, is the mean of the j-th SNP site;
[0087] S3-2-2, calculate eigenvalues and eigenvectors: perform eigendecomposition on the covariance matrix to obtain eigenvalues and eigenvectors. The eigenvalues represent the variance explained by each principal component, and the eigenvectors are used to define the direction of the principal component. Sort the principal components according to the size of the eigenvalues and select the first three principal components.
[0088] S3-2-3, construct the principal component score matrix: calculate the individual's score on each principal component, construct the principal component score matrix, and obtain the individual's score on a principal component by multiplying the individual's original data vector with the corresponding eigenvector;
[0089] S3-2-4, obtain potential breed combinations: In two-dimensional space, with the first principal component PC1 as the x-axis and the second principal component PC2 as the y-axis, the scores of each beef cattle individual are plotted in a plane coordinate system; beef cattle individuals of different breeds form different distribution areas on this plane due to differences in genomic variation patterns; if individuals of two breeds are distributed relatively close in two-dimensional space, it means that they have a close relationship at the genomic level; if the distribution is far, the relationship is distant, and potential breed subgroups or breed combinations with similar genetic backgrounds are obtained.
[0090] Further, such as Figure 6 As shown, the S3-3 specifically includes the following methods:
[0091] S3-3-1, Phenotypic data collation:
[0092] Determine economic traits: Identify important economic traits of beef cattle, such as growth rate (including daily weight gain, adult weight, etc.), meat quality (such as tenderness, marbling score, fat content, etc.), and reproductive performance (such as conception rate, calving interval, etc.);
[0093] Measuring phenotypic data: Accurately measure and record the phenotypic values of each individual beef cattle for selected economic traits;
[0094] Data quality control: Perform quality control on the collected phenotypic data and remove outliers;
[0095] S3-3-2, Genotype data preparation
[0096] Using the standardized variant site data, the genotype information of each beef cattle individual was organized into a matrix form, with rows representing individuals and columns representing variant sites. The elements in the matrix were the genotype codes of the individual at each variant site.
[0097] S3-3-3, build a mixed linear model
[0098] Determining fixed and random effects: The model includes fixed effects for breed, sex, and age, and considers relatedness between individuals as a random effect. Fixed effects include factors such as breed, sex, and age. Breed can reflect inherent differences in economic traits among different breeds; sex may affect traits such as growth rate and meat quality; and age, which is closely related to growth and development stage, has a systematic impact on economic traits and is therefore included as a fixed effect. Including relatedness between individuals as a random effect allows for more accurate estimation of the association between variant loci and economic traits, avoiding false positive associations due to relatedness.
[0099] Model construction: The mixed linear model is expressed as y = Xβ + Zu + e;
[0100] Where y is the collected economic trait data, X is the fixed effect design matrix (including variables corresponding to factors such as breed, sex, and age), β is the fixed effect vector, Z is the random effect design matrix (related to the kinship between individuals), u is the random effect vector, and e is the residual vector;
[0101] S3-3-4, Correlation Analysis
[0102] Calculate association statistics: Use the constructed mixed linear model to conduct association analysis between each variant site and economic traits, and calculate association statistics to measure the significance of the association between the variant site and the economic traits;
[0103] Determine the significance threshold: The significance threshold is determined by permutation test. When the association statistic exceeds this threshold, it is considered that there is a significant association between the variant site and the economic trait;
[0104] S3-3-5, identification of candidate genes and variant sites
[0105] Locating gene regions: For variant sites significantly associated with economic traits, determine the gene region or nearby functional elements based on the location information of the reference genome, and use bioinformatics databases to query gene location and functional annotation information;
[0106] Screening candidate genes and variant sites: Combine gene functional annotations with known biological knowledge to screen candidate genes associated with target economic traits. At the same time, identify variant sites closely linked to these candidate genes as targets for molecular marker-assisted selection breeding.
[0107] S3-3-6, Comprehensive Analysis of Multiple Varieties
[0108] Analysis of commonalities between breeds: Compare the commonalities in GWAS results of different beef cattle breeds and identify variant loci that are significantly associated with the same economic traits in multiple breeds. These loci represent genetic variants that have a key regulatory effect on economic traits in beef cattle populations.
[0109] Variety-specific analysis: This method simultaneously screens for significant association loci that appear only in specific varieties and are associated with unique economic traits or a variety-specific genetic background.
[0110] Construct breed-specific networks: For breed-specific trait association sites, construct breed-specific gene-gene interaction networks or gene-environment interaction networks to obtain the functional mechanisms of these sites in specific breeds, and then obtain candidate genes and variant sites related to important economic traits of beef cattle.
[0111] Furthermore, the S4 data visualization and report generation includes the following steps:
[0112] The analysis results were visualized, and genetic diversity index distribution maps, PCA maps, and GWAS Manhattan maps were drawn using drawing software to intuitively present the analysis results of multi-breed beef cattle genome data.
[0113] 1) Genetic diversity indicator distribution plotting: The calculated genetic diversity indicators were organized into a data frame format. The data frame lists the genetic diversity indicators and their corresponding values for different beef cattle breeds. The Matplotlib library in the R language was used to plot the genetic diversity indicator distribution map. The distribution maps included bar charts and box plots. The bar chart was used to compare the differences in a genetic diversity indicator between different breeds. The box plot was used to show the distribution of multiple genetic diversity indicators in different breeds.
[0114] 2) Principal component analysis (PCA) plotting: Obtain the results of principal component analysis, including the coordinate values of each individual on the principal component and the breed information to which the individual belongs, organize these data into a suitable data structure, and use the plotting software to plot the PCA plot.
[0115] 3) GWAS Manhattan plot: Organize the results of the GWAS analysis, including the chromosomal location of each variant site, the value of the association statistic, and whether it reaches a significant level. In R language, use the qqman package to draw the Manhattan plot.
[0116] Further, such as Figure 1 As shown, the method further includes:
[0117] S5, Data update and maintenance: Regularly collect new beef cattle genome data, including new breed sample data or data obtained from more in-depth sequencing analysis of existing breeds; repeat steps S1-S4 to dynamically update the data; regularly maintain the data in the database, check the integrity and accuracy of the data, and fix any data errors or loss problems; at the same time, optimize and adjust the data storage structure and analysis process according to the development of data analysis methods and changes in demand.
[0118] Corresponding to the method, Figure 7 As shown, the embodiment of the present invention also discloses a multi-breed beef cattle genome data analysis and management system, including:
[0119] Data collection and integration module 101: used to obtain raw genomic sequence data of beef cattle, establish a genomic data repository, and add a unique identification tag to the data of each breed;
[0120] Data preprocessing and standardization module 102: used to detect variant points and perform standardization on the detected variant sites, and at the same time, perform encoding on the genotype data of the variant sites;
[0121] Multi-breed genomic feature analysis module 103: used to identify beef cattle breeds with rich genetic diversity and those with relatively scarce genetic diversity, obtain potential breed subgroups or breed combinations with similar genetic backgrounds, and discover trait association sites that are common in different breeds or specific to a specific breed;
[0122] Data visualization and report generation module 104: used to visualize the results obtained from the analysis;
[0123] Data update and maintenance module 105: used to regularly collect new beef cattle genome data, including new breed sample data or data obtained from more in-depth sequencing analysis of existing breeds.
[0124] The specific implementation steps of the above modules in the system can refer to the aforementioned analysis and management method, which will not be described in detail here.
[0125] In summary, it is possible to achieve comprehensive, systematic and efficient analysis and management of genomic data of multiple beef cattle breeds, providing strong support for beef cattle genetic breeding.
Claims
1. A method for analyzing and managing genomic data of multiple breeds of beef cattle, characterized in that The steps include: S1, data collection and integration: obtain raw genomic sequence data of beef cattle, establish a genomic data repository, and add unique identification tags to the data of each breed; S2, data preprocessing and standardization: detect variant points and perform standardization on the detected variant sites, and encode the genotype data of the variant sites; S3, Multi-breed genomic characterization: Identify genetically diverse beef cattle breeds and those with relatively limited genetic diversity, identify potential breed subgroups or breed combinations with similar genetic backgrounds, and identify trait association loci that are common across breeds or specific to a particular breed. S4, data visualization and report generation: Visualize the results of the analysis; The S2 data preprocessing and standardization includes the following steps: The integrated multi-breed genomic data are assembled. For beef cattle breeds with reference genomes, the sequencing data are compared with the reference genomes using alignment software to identify variant sites in the sequencing data. For beef cattle breeds without reference genomes, a de novo assembly method is used, and a draft genome of the breed is constructed using software. Based on this, variant sites are detected. The detected variant sites are standardized, with a unified naming format and coordinate system. At the same time, the genotype data of the variant sites are encoded. The S3 multi-variety genome feature analysis includes the following steps: S3-1: Based on the standardized variant site data, calculate the genetic diversity index of multiple beef cattle genomes to assess the degree of genetic variation within and between different beef cattle populations. By comparing the genetic diversity indexes of different breeds, identify beef cattle breeds with rich genetic diversity and those with relatively poor genetic diversity. S3-2: Perform principal component analysis (PCA) to reduce high-dimensional genomic variation data to a two-dimensional space, displaying the distribution relationships of different beef cattle breeds at the genomic level and determining the closeness of relationships and genetic structure differences between breeds. Combined with cluster analysis methods, beef cattle breeds are clustered to determine the grouping of different breeds at the genomic level, identifying potential breed subgroups or breed combinations with similar genetic backgrounds. S3-3: Conduct genome-wide association studies (GWAS) on gene regions or functional elements related to important economic traits of beef cattle, use mixed linear models to detect association signals between variant sites and economic traits, and identify candidate genes and variant sites that are significantly associated with target traits; through comprehensive analysis of multi-breed beef cattle data, discover trait association sites that are common in different beef cattle breeds or unique to specific beef cattle breeds, providing key information for molecular marker-assisted selection and breeding of beef cattle.
2. The method for analyzing and managing multi-breed beef cattle genome data according to claim 1, wherein: The S1 data collection and integration includes the following steps: First, using the MySQL database management system, a general project database was created in the database to store all information related to the multi-breed genomic data of beef cattle, and it was named the Beef Cattle Genomic Data Repository; For each different beef cattle breed, a separate tablespace is created in the overall database. Each tablespace has the same predefined data structure, which includes: a sample number field, used to uniquely identify each individual beef cattle sample; a sequencing data field, used to store genomic sequencing data that has undergone quality control and filtering; a breed identification field, used to identify the beef cattle breed to which the data belongs; and a collection time field, used to record the time information of sample collection. During the data integration process, when the genomic data of a certain breed of beef cattle is imported into the database, the data is stored in the tablespace of the corresponding breed according to the pre-set correspondence between the tablespace and the breed; for the data of each sample, a unique number is assigned to it in the sample number field, and the name of the breed to which the sample belongs is clearly filled in the breed identification field; at the same time, an index is created based on the breed identification field to locate the corresponding tablespace and extract data when all data of a specific breed need to be queried; and an index is created based on the sample number field to obtain detailed information of a specific sample when processing single sample data or performing association analysis between samples.
3. The method for analyzing and managing multi-breed beef cattle genome data according to claim 1, wherein: The S2 data preprocessing and standardization includes the following steps: S2-1, Data processing for beef cattle breeds with reference genomes S2-1-1, Genome assembly and alignment: Select reference genome: obtain high-quality beef cattle reference genome sequences from public databases; Data alignment: Use the professional alignment software BWA to first index the sequencing data. For short-read data generated by the Illumina sequencing platform, use the bwa index command to index the reference genome. Then, use the bwa mem command to align the sequencing data after quality control and filtering with the reference genome. Alignment result processing: Samtools software was used to process the SAM format files generated by BWA alignment. First, the SAM files were converted to BAM format. The samtools view -Sb command was used to sort the BAM files. The samtools sort command was used to sort the reads according to chromosome position. S2-1-2, variant site identification: Local realignment: Use GATK's RealignerTargetCreator and IndelRealigner tools to perform local realignment of regions where indels may exist and correct alignment errors caused by indels; Base quality score recalibration: GATK's BaseRecalibrator and PrintReads tools are used to recalibrate base quality scores. Based on known variant site information and sequencing data characteristics, the quality score of each base is re-evaluated to make subsequent variant detection more reliable. S2-1-3, mutation site detection: Use the GATK HaplotypeCaller tool to detect variant sites, including single nucleotide polymorphisms (SNPs) and indels. Variant sites are identified by constructing local haplotypes and performing statistical analysis. Minimum coverage is set during the detection process to filter out false positive variant sites that may arise from low-coverage regions. S2-2, Data processing for beef cattle breeds without reference genomes S2-2-1, reassembly: Data preprocessing: Trimmomatic software was used to remove adapter sequences and low-quality bases from sequencing data without a reference genome, and the reads were length-filtered. Assembly software selection and operation: De novo assembly software Velvet was used. First, the k-mer value was estimated based on the characteristics of the sequencing data. Then, the Velveth command was used to construct the Velvet hash table. Finally, the velvetg command was used to assemble the genome. S2-2-2, mutation site detection: Evaluation and optimization of assembly results: Use the evaluation tool QUAST to evaluate the quality of the assembled genome draft, check the completeness and continuity of the assembly, and adjust the assembly parameters based on the evaluation results; Variant site identification: Using the reference-free genome variation detection software FreeBayes, the quality-assessed and optimized draft genome data was input into FreeBayes, and the minimum allele frequency was set to detect SNPs and Indels variation sites. S2-3, variant site standardization and genotype data encoding S2-3-1, variant site standardization: Format conversion: Use VCFtools software to convert variant site information from the output formats of various detection tools into a unified standard VCF format; Coordinate system unification: The coordinates of all variant sites are uniformly corrected and adjusted according to the coordinate system of the reference genome, so that the coordinate information of variant sites in different varieties is consistent under the same coordinate system, even if they are obtained using different assembly methods or reference genome versions; S2-3-2, genotype data encoding: Determine the coding rules: code the homozygous reference genotype as 0, the homozygous variant genotype as 1, and the heterozygous genotype as 2; Coding operation: Write a Python script to read the genotype data of the variant site, convert the genotype of each variant site according to the above coding rules, and store the encoded data in a new file or database field for subsequent data analysis and calculation.
4. The method for analyzing and managing multi-breed beef cattle genome data according to claim 1, wherein: The S3-1 specifically includes the following methods: S3-1-1) Calculate genetic diversity indicators for multiple breeds of beef cattle genomes: 1) Calculate heterozygosity: Observed heterozygosity Ho: Count the proportion of heterozygous individuals at each locus to the total number of individuals; for each variant locus, calculate the number of individuals with heterozygous genotype at that locus in the sample, then divide it by the total number of samples to obtain the observed heterozygosity of that locus, and finally average the observed heterozygosity of all loci to obtain the observed heterozygosity of the beef cattle population of that breed; Expected heterozygosity He: According to the Hardy-Weinberg equilibrium law, in an ideal population, the expected heterozygous frequency at each site is calculated by calculating the allele frequency at each site and then using the formula Get the expected heterozygosity of each site, where is the frequency of the i-th allele, n is the number of alleles, and the average is calculated to obtain the expected heterozygosity of the population, which is used to assess whether the population conforms to the Hardy-Weinberg equilibrium and whether there are factors that affect genetic diversity; 2) Calculate nucleotide diversity Pi: Calculate the average number of nucleotide differences at the same position between any two randomly selected nucleotide sequences in the genome. By comparing the nucleotide diversity of different beef cattle populations, the degree of genetic variation between the two at the genomic level is assessed. For each breed, the number of nucleotide differences between all two individuals at all variant sites is counted, and then divided by the total number of sites compared and the number of individual pairs to obtain the nucleotide diversity value of the breed. 3) Calculate allelic richness (AR): Calculate the average number of different alleles observed in a given number of samples in each breed of beef cattle population to obtain the allelic richness (AR); 4) Inbreeding coefficient F: By analyzing the standardized variant site data, the inbreeding coefficient of each individual is calculated, and then the average inbreeding coefficient of the group is obtained; 5) Linkage disequilibrium (LD): The degree of linkage disequilibrium between genes in a population is assessed by calculating the linkage disequilibrium coefficient between different loci. 6) Population differentiation index Fst: This index compares the genetic variation within a population with the total genetic variation, resulting in a value between 0 and 1. Values closer to 1 indicate greater genetic differentiation between populations and a more uneven distribution of genetic diversity among populations. Values closer to 0 indicate smaller genetic differences between populations and a more even distribution of genetic diversity among populations. S3-1-2) Identify genetically diverse beef cattle breeds and those that are relatively less diverse: A certain weight was set for each indicator, among which the weight of comprehensive heterozygosity was 0.2, the weight of nucleotide diversity was 0.2, the weight of allele richness was 0.2, the weight of inbreeding coefficient was 0.1, the weight of linkage disequilibrium was 0.1, and the weight of population differentiation index was 0.2; For each variety, a comprehensive score is calculated according to the set weight based on the comparison of its values on various indicators with other varieties. Varieties with scores higher than the set threshold are identified as varieties with rich genetic diversity, and varieties with scores lower than the set threshold are identified as varieties with relatively poor genetic diversity.
5. The method for analyzing and managing multi-breed beef cattle genome data according to claim 1, wherein: The S3-2 specifically includes the following methods: S3-2-1, principal component analysis method PCA: First, based on the standardized single nucleotide polymorphism (SNP) data matrix, the rows in the matrix represent different beef cattle individuals, the columns represent various SNP sites, and each element represents the genotype of the individual at a specific SNP site; Subtract the mean of each SNP site from the data of the site to make the mean of the data 0; Calculate the covariance matrix according to the nature of the data. For the covariance matrix, the elements Indicates the and The covariance of SNP sites is calculated as follows: , in is the sample size, It is Individuals in the The data of SNP sites, It is The mean of the SNP sites; It is Individuals in the The data of SNP sites, It is The mean of the SNP sites; S3-2-2, calculate eigenvalues and eigenvectors: perform eigendecomposition on the covariance matrix to obtain eigenvalues and eigenvectors. The eigenvalues represent the variance explained by each principal component, and the eigenvectors are used to define the direction of the principal component. Sort the principal components by eigenvalue and select the first three principal components. S3-2-3, construct the principal component score matrix: calculate the individual's score on each principal component, construct the principal component score matrix, and obtain the individual's score on a principal component by multiplying the individual's original data vector with the corresponding eigenvector; S3-2-4, obtain potential breed combinations: In two-dimensional space, with the first principal component PC1 as the x-axis and the second principal component PC2 as the y-axis, the scores of each beef cattle individual are plotted in a plane coordinate system; beef cattle individuals of different breeds form different distribution areas on this plane due to differences in genomic variation patterns; if individuals of two breeds are distributed relatively close in two-dimensional space, it means that they have a close relationship at the genomic level; if the distribution is far, the relationship is distant, and potential breed subgroups or breed combinations with similar genetic backgrounds are obtained.
6. The method for analyzing and managing multi-breed beef cattle genome data according to claim 1, wherein: The S3-3 specifically includes the following methods: S3-3-1, Phenotypic data collation: Determine economic traits: Identify important economic traits of beef cattle, including growth rate, meat quality and reproductive performance; Measuring phenotypic data: Accurately measure and record the phenotypic values of each individual beef cattle for selected economic traits; Data quality control: Perform quality control on the collected phenotypic data and remove outliers; S3-3-2, Genotype data preparation Using the standardized variant site data, the genotype information of each beef cattle individual was organized into a matrix form, with rows representing individuals and columns representing variant sites. The elements in the matrix were the genotype codes of the individual at each variant site. S3-3-3, build a mixed linear model Determine the fixed and random effects: Set the breed, sex, and age of beef cattle as fixed effects and include them in the model, and consider the kinship between individuals as a random effect. Model construction: The mixed linear model is expressed as ; in It is the economic trait data collected. is the fixed-effect design matrix, is the fixed effect vector, is the random effects design matrix, is the random effects vector, is the residual vector; S3-3-4, Correlation Analysis Calculate association statistics: Use the constructed mixed linear model to conduct association analysis between each variant site and economic traits, and calculate association statistics to measure the significance of the association between the variant site and the economic traits; Determine the significance threshold: The significance threshold is determined by permutation test. When the association statistic exceeds this threshold, it is considered that there is a significant association between the variant site and the economic trait; S3-3-5, identification of candidate genes and variant sites Locating gene regions: For variant sites significantly associated with economic traits, determine the gene region or nearby functional elements based on the location information of the reference genome, and use bioinformatics databases to query gene location and functional annotation information; Screening candidate genes and variant sites: Combine gene functional annotations with known biological knowledge to screen candidate genes associated with target economic traits. At the same time, identify variant sites closely linked to these candidate genes as targets for molecular marker-assisted selection breeding. S3-3-6, Comprehensive Analysis of Multiple Varieties Analysis of commonalities between breeds: Compare the commonalities in GWAS results of different beef cattle breeds and identify variant loci that are significantly associated with the same economic traits in multiple breeds. These loci represent genetic variants that have a key regulatory effect on economic traits in beef cattle populations. Variety-specific analysis: This method simultaneously screens for significant association loci that appear only in specific varieties and are associated with unique economic traits or a variety-specific genetic background. Construct breed-specific networks: For breed-specific trait association sites, construct breed-specific gene-gene interaction networks or gene-environment interaction networks to obtain the functional mechanisms of these sites in specific breeds, and then obtain candidate genes and variant sites related to important economic traits of beef cattle.
7. The method for analyzing and managing multi-breed beef cattle genome data according to claim 1, wherein: The S4 data visualization and report generation includes the following steps: 1) Genetic Diversity Indicator Distribution Plotting: Calculated genetic diversity indicators were organized into a data frame format. The data frame contained various genetic diversity indicators and their corresponding values for different beef cattle breeds. The Matplotlib library in the R language was used to plot genetic diversity indicator distribution plots. Distribution plots included bar charts and boxplots. The bar chart was used to compare differences in a specific genetic diversity indicator between breeds, while the boxplot was used to display the distribution of multiple genetic diversity indicators across breeds. 2) Principal component analysis (PCA) plotting: Obtain the results of the principal component analysis, including the coordinate values of each individual on the principal component and the breed information to which the individual belongs. Organize these data into a suitable data structure and use the plotting software to plot the PCA plot. 3) GWAS Manhattan plot: Organize the results of the GWAS analysis, including the chromosomal location of each variant site, the value of the association statistic, and whether it reaches a significant level. In R language, use the qqman package to draw the Manhattan plot.
8. A beef cattle multi-breed genomic data analysis and management system, the analysis and management system using the beef cattle multi-breed genomic data analysis and management method according to any one of claims 1 to 7, characterized in that include: Data collection and integration module: used to obtain raw genomic sequence data of beef cattle, establish a genomic data repository, and add unique identification tags to the data of each breed; Data preprocessing and standardization module: used to detect variant points and perform standardization on the detected variant sites, and at the same time, encode the genotype data of the variant sites; Multi-breed genomic signature analysis module: used to identify beef cattle breeds with rich genetic diversity and those with relatively scarce genetic diversity, obtain potential breed subgroups or breed combinations with similar genetic backgrounds, and discover trait association sites that are common in different breeds or specific to specific breeds; Data visualization and report generation module: used to visualize the analysis results; Data update and maintenance module: used to regularly collect new beef cattle genome data, including new breed sample data or data obtained from more in-depth sequencing analysis of existing breeds.
Citation Information
Patent Citations
Beef cattle fatty acid component candidate marker multi-omics screening method and application thereof
CN115478113A
GBS whole genome association analysis method for buffalo
CN117095746A