A haplotype distance assessment method and device based on second-generation short-read sequences

Through the method based on the second-generation short read long sequence, whole genome resequencing data are imported, mutant sites are detected, data with deviations greater than the threshold are segmented and removed, and distance standardization is carried out, which solves the problem of inaccurate haplotype distance evaluation in the prior art, and achieves high-accuracy multi-dimensional haplotype distance evaluation.

CN117174178BActive Publication Date: 2025-08-15INST OF HORTICULTURE JIANGXI ACAD OF AGRI SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310955789.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-01
Publication Date
2025-08-15
Estimated Expiration
2043-08-01

AI Technical Summary

Technical Problem

The prior art lacks quantitative standards when evaluating diploid genome-wide haplotype distances, resulting in inaccurate results.

Method used

Through a method based on the second-generation short read long sequence, whole genome resequencing data were imported, mutant sites were detected, and the data with deviations greater than the threshold were eliminated, distance standardization was performed, and haplotype distances were calculated at individual and population levels.

Benefits of technology

A multi-dimensional haplotype distance evaluation with high accuracy is achieved, which improves the accuracy and speed of data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117174178B_ABST
    Figure CN117174178B_ABST
Patent Text Reader

Abstract

The present invention provides a haplotype distance assessment method and device based on second-generation short-read sequences. The method mainly comprises: performing mutation site detection on whole-genome resequencing data and converting it into a whole-genome variation map, performing file segmentation to obtain multiple sub-map files, counting the number of mutations to eliminate sub-map files with large data deviations, and analyzing the remaining sub-map files as relatively accurate assessment data, thereby obtaining individual haplotype distances at the local interval and whole-genome levels, and group haplotype distances at the local interval and whole-genome levels. Multi-dimensional haplotype distance data can be obtained with high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the field of gene data processing technology, and specifically to a haplotype distance evaluation method and device based on second-generation short-read length sequences. Background Art

[0002] Typically, diploid organisms possess a pair of homologous chromosomes (denoted as 2n), containing two sets of haplotype sequences. The diploid genome local haplotype distance (i.e., local segment haplotype difference) calculates the degree of differentiation between two haplotypes within a specified local segment (e.g., a 50kb interval); the diploid genome-wide haplotype distance (i.e., genome-wide haplotype difference) calculates the degree of differentiation between two haplotypes across the entire genome. Evaluating the diploid genome-wide haplotype distance is based on evaluating the local haplotype distance of the diploid genome. The principle is to divide the genome into multiple intervals of fixed length and calculate the average degree of haplotype differentiation across all intervals. Existing methods infer the correlation between different haplotype sequences based on the number of base differences between haplotypes, thereby constructing a haplotype network diagram. A connecting line represents the relationship between two haplotypes, and the short lines on the connecting line represent the number of base substitutions required to change from one haplotype to another. Without using a quantitative standard to understand haplotype differentiation, the results obtained by this method are inaccurate. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to address the deficiencies of the existing technology and provide a haplotype distance assessment method and device based on second-generation short-read length sequences.

[0004] The present invention solves the above-mentioned technical problem with the following technical solution: A haplotype distance assessment method based on second-generation short-read sequences comprises the following steps:

[0005] Importing whole-genome resequencing data with second-generation short-read sequence characteristics, the whole-genome resequencing data is M×N, where M is the number of populations and N is the number of samples in each population;

[0006] The M×N whole-genome resequencing data are processed for variant site detection, and the corresponding whole-genome variation maps are obtained according to the detected M×N variant sites;

[0007] Taking M×N whole-genome variation maps as map files to be segmented, and segmenting the map files to be segmented into multiple sub-map files with the same interval length according to the set interval length of the local haplotype distance;

[0008] Count the number of mutations in each sub-atlas file respectively, and remove the sub-atlas files with a number of mutations less than or equal to the mutation threshold, and take the sub-atlas files with a number of mutations greater than the mutation threshold as the sub-atlas files to be processed;

[0009] In the local interval, the sub-map files to be processed are respectively subjected to distance normalization processing to obtain the individual-level genomic local haplotype distance (the individual-level genomic local haplotype distance is the distance value of two haplotypes of the same sample after normalization processing); in the local interval, the individual-level genomic local haplotype distances corresponding to all the sub-map files to be processed belonging to the same group are respectively subjected to distance normalization processing to obtain the group-level genomic local haplotype distance;

[0010] At the whole genome level, the average of the individual-level genomic local haplotype distances corresponding to all sub-map files to be processed is calculated to obtain the individual-level haplotype distance; at the whole genome level, the average of the population-level genomic local haplotype distances corresponding to M populations is calculated to obtain the population-level haplotype distance.

[0011] The beneficial effects of the present invention are as follows: whole-genome resequencing data is converted into a whole-genome variation map after variant site detection, and the file is segmented to obtain multiple sub-map files, the number of variants is counted to eliminate sub-map files with large data deviations, and the remaining sub-map files are analyzed as relatively accurate evaluation data, thereby obtaining individual haplotype distances at the local interval and whole-genome levels, and group haplotype distances at the local interval and whole-genome levels, and multi-dimensional haplotype distance data can be obtained with high accuracy.

[0012] Another technical solution of the present invention to solve the above technical problems is as follows: a haplotype distance assessment device based on second-generation short-read sequences, comprising:

[0013] An import module is used to import whole-genome resequencing data with second-generation short-read sequence characteristics, wherein the whole-genome resequencing data is M×N, where M is the number of populations and N is the number of samples in each population;

[0014] The variant site detection module is used to detect variant sites in M×N whole-genome resequencing data and obtain corresponding whole-genome variation maps based on the detected M×N variant sites;

[0015] A segmentation module is used to use M×N whole-genome variation maps as map files to be segmented, and to segment the map files to be segmented into multiple sub-map files with the same interval length according to the set interval length of the local haplotype distance;

[0016] The elimination module is used to count the number of mutations in each sub-atlas file, and eliminate the sub-atlas files with a number of mutations less than or equal to the mutation threshold, and take the sub-atlas files with a number of mutations greater than the mutation threshold as the sub-atlas files to be processed;

[0017] The distance evaluation processing module is used to perform distance normalization processing on the sub-map files to be processed in the local interval to obtain the local haplotype distance of the individual level genome; in the local interval, the distance normalization processing is performed on the local haplotype distance of the individual level genome corresponding to all the sub-map files to be processed belonging to the same group to obtain the local haplotype distance of the group level genome;

[0018] It is also used to calculate the average of the individual-level local genomic haplotype distances corresponding to all sub-map files to be processed at the whole genome level to obtain the individual-level haplotype distance; at the whole genome level, it is used to calculate the average of the population-level local genomic haplotype distances corresponding to M populations to obtain the population-level haplotype distance.

[0019] Another technical solution of the present invention to solve the above-mentioned technical problems is as follows: a haplotype distance evaluation device based on second-generation short-read sequences, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the haplotype distance evaluation device based on second-generation short-read sequences as described above is implemented.

[0020] Another technical solution of the present invention to solve the above-mentioned technical problem is as follows: a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements a haplotype distance evaluation device based on the second-generation short read sequence as described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A schematic diagram of a flow chart of a haplotype distance evaluation method provided in an embodiment of the present invention;

[0022] Figure 2 A schematic diagram of the functional modules of a haplotype distance evaluation device provided by an embodiment of the present invention;

[0023] Figure 3 Schematic diagram of haplotype distances at the whole genome level for the five populations in the experiment of the present invention;

[0024] Figure 4 Schematic diagram of the whole-genome haplotype distances within 24 individuals in the experiment of the present invention;

[0025] Figure 5 Schematic diagram of the haplotype distances of 24 individuals in the experiment of the present invention relative to population 4. DETAILED DESCRIPTION

[0026] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.

[0027] Figure 1 A schematic diagram of the flow chart of the haplotype distance evaluation method provided in an embodiment of the present invention.

[0028] Example 1:

[0029] like Figure 1 As shown, a haplotype distance assessment method based on second-generation short-read sequences includes the following steps:

[0030] Importing high-depth whole-genome resequencing data with second-generation short-read sequence characteristics, the whole-genome resequencing data is M×N, where M is the number of populations and N is the number of samples in each population;

[0031] The M×N whole-genome resequencing data are processed for variant site detection, and the corresponding whole-genome variation maps are obtained according to the detected M×N variant sites;

[0032] Taking M×N whole-genome variation maps as map files to be segmented, and segmenting the map files to be segmented into multiple sub-map files with the same interval length according to the set interval length of the local haplotype distance;

[0033] Count the number of mutations in each sub-atlas file respectively, and remove the sub-atlas files with a number of mutations less than or equal to the mutation threshold, and take the sub-atlas files with a number of mutations greater than the mutation threshold as the sub-atlas files to be processed;

[0034] In the local interval, the distance normalization processing is performed on the sub-map files to be processed to obtain the individual-level genomic local haplotype distance; in the local interval, the distance normalization processing is performed on the individual-level genomic local haplotype distance corresponding to all the sub-map files to be processed belonging to the same group to obtain the group-level genomic local haplotype distance;

[0035] At the whole genome level, the average of the individual-level genomic local haplotype distances corresponding to all sub-map files to be processed is calculated to obtain the individual-level haplotype distance; at the whole genome level, the average of the population-level genomic local haplotype distances corresponding to M populations is calculated to obtain the population-level haplotype distance.

[0036] In the above embodiment, the whole genome resequencing data is converted into the form of a whole genome variation map after mutation site detection, and the file is segmented to obtain multiple sub-map files. The number of mutations is counted to eliminate sub-map files with large data deviations. The remaining sub-map files are analyzed as evaluation data with relatively high accuracy, thereby obtaining individual haplotype distances at the local interval and whole genome levels, and population haplotype distances at the local interval and whole genome levels, and multi-dimensional haplotype distance data with high accuracy can be obtained.

[0037] The M×N whole-genome resequencing data are subjected to detection and processing of variant sites, and corresponding whole-genome variation maps are obtained according to the detected M×N variant sites, specifically:

[0038] Each of the whole-genome resequencing data is filtered according to the set filtering criteria, and the filtered M×N whole-genome resequencing data are respectively aligned to the reference genome, and the compared M×N whole-genome resequencing data are sorted according to the genome sequence, and the sorted M×N whole-genome resequencing data are respectively detected for mutation sites using the GATK tool, and the whole-genome variation map corresponding to the M×N whole-genome resequencing data is obtained according to the detected M×N mutation sites.

[0039] The specific operation is to import the high-depth whole-genome resequencing data of M×N samples as a FASTQ format file, and filter this FASTQ file using the fastp tool (fastp is a fast tool for NGS data preprocessing) to improve sequence quality, remove low-quality sequences and adapter sequences, and retain high-quality sequences. The purpose is to filter out fragments with low sequencing quality and obtain a clean version of FASTQ data. Subsequently, the FASTQ format sequence is used to use the BWA tool (BWA is a tool for quickly and accurately aligning short sequences to the reference genome). It can efficiently align large-scale sequencing data to the reference genome, generate SAM format data, and obtain the aligned SAM format file. The Samtools tool (Samtools is a feature-rich, efficient and flexible SAM / BAM file processing tool) can realize the conversion, sorting, indexing, statistics, screening, editing and SNP and Indel detection of SAM / BAM format files. After converting the SAM format file to BAM, it is sorted according to the genome sequence to obtain the sorted BAM format file. Only after sorting can the variant sites be detected, thereby constructing a whole-genome variation map. The sorted BAM file and FASTA-formatted genome were input into the compilation and detection GATK tool (GATK, or Genome Analysis Toolkit, is a software toolkit widely used in genomic data analysis. Its main functions include: 1. Variation detection, such as SNPs, Indels, etc.; 2. Variation annotation, such as functional impact and frequency; 3. Genotype quality control and filtering, such as depth, quality, and heterozygosity). A VCF format file of the whole-genome variation map of a single sample was obtained, and the whole-genome variation maps of M×N whole-genome resequencing data (samples) were integrated into a VCF format file, which contains a VCF format file of all variation site information of M×N samples.

[0040] In the above embodiment, the initial whole-genome resequencing data can be converted into a relatively accurate whole-genome variation map that is convenient for subsequent analysis and processing, thereby speeding up the data analysis.

[0041] After obtaining M×N whole-genome variation maps, the M×N whole-genome variation maps are used as the map file to be segmented. The map file to be segmented is segmented into multiple sub-map files with the same interval length according to the set interval length of the local haplotype distance. The specific operation is as follows:

[0042] Input the above VCF file into the haplotype orientation tool BEAGLE to obtain a phasedVCF file. Split the phased VCF file according to the interval size used for local haplotype distance calculation (for example, a 50kb interval) to obtain multiple phased GFF files with the same interval length. Operate on one of these phased VCF files (i.e., the submap file) to subsequently calculate individual-level and population-level local haplotype distances. The word "phased" can be understood as meaning genetic phasing, genotyping, or haplotype typing.

[0043] The following is a detailed introduction on how to calculate:

[0044] The distance normalization process is performed on each of the sub-map files to be processed to obtain the individual-level genomic local haplotype distance, specifically:

[0045] Each haplotype in each submap file to be processed is split into two hypothetical homozygous samples. The two homozygous samples belonging to the same submap file are then input into the VCF2Dis tool in phased VCF format. The VCF2Dis tool then calculates the haplotype distance value and the average haplotype distance value for the submap file to be processed. Each haplotype distance value is then divided by the average haplotype distance value to obtain the individual-level genomic local haplotype distance for the submap file. The term "phased" can be understood as meaning genetic phasing, genotyping, or haplotype typing.

[0046] It should be understood that one sample corresponds to one sub-atlas file.

[0047] The distance normalization process is performed on the individual-level genomic local haplotype distances corresponding to all the sub-map files to be processed belonging to the same group to obtain the group-level genomic local haplotype distances, specifically:

[0048] For M×N whole-genome resequencing data, the two haplotypes in each sub-map file to be processed are split into two hypothetical homozygous samples, and the two homozygous samples belonging to the same sub-map file are respectively input into the VCF2Dis tool in phased VCF format. The VCF2Dis tool is used to calculate the average haplotype distance value of the M×N individuals in the same sub-map file to be processed (i.e., corresponding to the M×N whole-genome resequencing data), and then each haplotype distance value is divided by the average haplotype distance value to obtain the local haplotype distance of each individual, and then the average of the local haplotype distances in the M populations is calculated to obtain the population-level genome local haplotype distances of the M populations.

[0049] It should be understood that the average haplotype distance value in the population-level genomic local haplotype distance is obtained by calculating the average value of M×N individuals, while the average haplotype distance value in the individual-level genomic local haplotype distance is obtained by calculating the average value of N individuals in 1 population.

[0050] The specific operation is to first split the two haplotypes of a sample in the phased VCF format into two hypothetical homozygous samples (for example, sample 24, genotype 0|1, is split into sample 24_hap1, genotype 0|0, and sample 24_hap2, genotype 1|1). After modification, a new phased VCF file is generated. This file is input into the VCF2Dis tool to calculate pairwise haplotype distances (both as ratios between two samples and as comparisons between different samples). Each haplotype distance is normalized by dividing it by the average haplotype distance. The individual-level genomic local haplotype distance is then the standardized distance between two haplotypes in the same sample (for example, the genomic local haplotype distance of sample 24 is the standardized distance between haplotypes 24_hap1 and 24_hap2). The population-level genomic local haplotype distance is the standardized distance between the haplotypes of two populations (for example, for population M1, the haplotypes of all individuals in M1 are compared to obtain the standardized distance V1, which is the local haplotype distance between different populations).

[0051] Individual and population haplotype distances are calculated at the genome-wide level, taking into account the standardized haplotype distances calculated from all phased VCF files. Calculating individual haplotype distances at the genome-wide level involves calculating the standardized individual haplotype distance for each phased VCF file (e.g., for sample 24, the standardized haplotype distance for each phased VCF file) and then averaging the two values. "Phased" can refer to genotyping, haplotype assignment, or genotyping.

[0052] The specific calculation of population-level haplotype distance at the whole genome level is as follows: the standardized population haplotype distance value of each phasedVCF is calculated separately (for example, for population M1, the standardized haplotype distance value of each phased VCF in the M1 population is calculated), and then the average value is taken.

[0053] The method further includes the step of calculating the haplotype distance of an individual (e.g., sample 24) relative to a population (e.g., population M1), specifically:

[0054] The individual-level genomic local haplotype distance corresponding to the sample to be calculated is divided by the population-level genomic local haplotype distance corresponding to the population to be calculated. The result is the haplotype distance of the individual relative to the population.

[0055] In the above embodiment, multi-dimensional haplotype distance data can be obtained.

[0056] The following experiment was carried out according to the method flow of the present invention:

[0057] A total of 72 citrus samples from five populations were subjected to high-depth genome resequencing to obtain the haplotype distances at the whole genome level of the five populations, such as Figure 3 As shown in Figure 2, the haplotype distances of five populations were assessed at the genome-wide level using the method of the present invention. It was found that populations 4 and 5 had significantly decreased haplotype distances relative to populations 1, 2, and 3. This indicates that the haplotype distances between populations 4 and 5 are short, and this method is highly capable of estimating haplotype distances between populations.

[0058] The whole-genome haplotype distances within 24 individuals, e.g. Figure 4 As shown in FIG. 1 , the haplotype distances of 24 samples in population 3 were evaluated at the whole-genome level using the method of the present invention. The haplotype distances of each individual at the whole-genome level were successfully estimated.

[0059] The haplotype distances of the 24 individuals relative to the population, e.g. Figure 5 As shown. The haplotype distances of 24 samples in population 4 were evaluated at the whole genome level by the method of the present invention. The haplotype distances of each individual at the whole genome level were successfully estimated. Figure 4 and Figure 5 correspond Figure 3 , we can see from the results in the figure: Figure 5 The average value of the 24 individuals is close to Figure 3 Middle group 4 reaction; Figure 4 The average value of the 24 individuals is close to Figure 3 Middle group 3 reactions.

[0060] The method of the present invention can obtain multi-dimensional haplotype distance data with high accuracy.

[0061] Example 2:

[0062] like Figure 2 As shown, a haplotype distance assessment device based on second-generation short-read sequences comprises:

[0063] An import module is used to import whole-genome resequencing data with second-generation short-read sequence characteristics, wherein the whole-genome resequencing data is M×N, where M is the number of populations and N is the number of samples in each population;

[0064] The variant site detection module is used to detect variant sites in M×N whole-genome resequencing data and obtain corresponding whole-genome variation maps based on the detected M×N variant sites;

[0065] A segmentation module is used to use M×N whole-genome variation maps as map files to be segmented, and to segment the map files to be segmented into multiple sub-map files with the same interval length according to the set interval length of the local haplotype distance;

[0066] The elimination module is used to count the number of mutations in each sub-atlas file, and eliminate the sub-atlas files with a number of mutations less than or equal to the mutation threshold, and take the sub-atlas files with a number of mutations greater than the mutation threshold as the sub-atlas files to be processed;

[0067] The distance evaluation processing module is used to perform distance normalization processing on the sub-map files to be processed in the local interval to obtain the local haplotype distance of the individual level genome; in the local interval, the distance normalization processing is performed on the local haplotype distance of the individual level genome corresponding to all the sub-map files to be processed belonging to the same group to obtain the local haplotype distance of the group level genome;

[0068] It is also used to calculate the average of the individual-level local genomic haplotype distances corresponding to all sub-map files to be processed at the whole genome level to obtain the individual-level haplotype distance; at the whole genome level, it is used to calculate the average of the population-level local genomic haplotype distances corresponding to M populations to obtain the population-level haplotype distance.

[0069] In the above embodiment, the whole genome resequencing data is converted into the form of a whole genome variation map after mutation site detection, and the file is split to obtain multiple sub-map files. The number of mutations is counted to eliminate sub-map files with large data deviations, and the remaining sub-map files are analyzed as evaluation data with relatively high accuracy, thereby obtaining individual-level genome local haplotype distances, population-level genome local haplotype distances, individual-level haplotype distances and population-level haplotype distances at the local interval and whole genome levels, and multi-dimensional haplotype distance data can be obtained with high accuracy.

[0070] In the variant site detection module, the M×N whole genome resequencing data are processed for variant site detection, and corresponding whole genome variation maps are obtained according to the detected M×N variant sites, specifically:

[0071] Each of the whole-genome resequencing data is filtered according to the set filtering criteria, and the filtered M×N whole-genome resequencing data are respectively aligned to the reference genome, and the compared M×N whole-genome resequencing data are sorted according to the genome sequence, and the sorted M×N whole-genome resequencing data are respectively detected for mutation sites using the GATK tool, and the whole-genome variation map corresponding to the M×N whole-genome resequencing data is obtained according to the detected M×N mutation sites.

[0072] In the distance evaluation processing module, distance normalization processing is performed on the sub-map files to be processed to obtain the individual-level genomic local haplotype distance, specifically:

[0073] The two haplotypes in each sub-map file to be processed are respectively split into two hypothetical homozygous samples, and the two homozygous samples belonging to the same sub-map file are respectively input into the VCF2Dis tool in phased VCF format. The VCF2Dis tool is used to calculate each haplotype distance value and the average haplotype distance value corresponding to the same sub-map file to be processed, and then each haplotype distance value is divided by the average haplotype distance value to obtain the individual-level genomic local haplotype distance corresponding to the same sub-map file.

[0074] The distance evaluation processing module is also used to calculate the haplotype distance of an individual relative to the population, specifically:

[0075] The individual-level genomic local haplotype distance corresponding to the sample to be calculated is divided by the population-level genomic local haplotype distance corresponding to the population to be calculated. The result is the haplotype distance of the individual relative to the population.

[0076] In the distance evaluation processing module, the individual-level genomic local haplotype distances corresponding to all the sub-map files to be processed belonging to the same group are respectively subjected to distance standardization processing to obtain the group-level genomic local haplotype distances, specifically:

[0077] For M×N whole-genome resequencing data, the two haplotypes in each sub-map file to be processed are split into two hypothetical homozygous samples, and the two homozygous samples belonging to the same sub-map file are respectively input into the VCF2Dis tool in phased VCF format. The VCF2Dis tool is used to calculate the average haplotype distance value of the M×N individuals in the same sub-map file to be processed (i.e., corresponding to the M×N whole-genome resequencing data), and then each haplotype distance value is divided by the average haplotype distance value to obtain the local haplotype distance of each individual, and then the average of the local haplotype distances in the M populations is calculated to obtain the population-level genome local haplotype distances of the M populations.

[0078] Example 3:

[0079] A haplotype distance assessment device based on second-generation short-read sequences comprises a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the haplotype distance assessment device based on second-generation short-read sequences as described above is implemented.

[0080] Example 4:

[0081] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements a haplotype distance assessment device based on second-generation short-read sequences as described above.

[0082] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0083] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A haplotype distance assessment method based on second-generation short-read sequences, characterized in that: The steps include: Importing whole-genome resequencing data with second-generation short-read sequence characteristics, the whole-genome resequencing data is M×N, where M is the number of populations and N is the number of samples in each population; The M×N whole-genome resequencing data are processed for variant site detection, and the corresponding whole-genome variation maps are obtained according to the detected M×N variant sites; Taking M×N whole-genome variation maps as map files to be segmented, and segmenting the map files to be segmented into multiple sub-map files with the same interval length according to the set interval length of the local haplotype distance; Count the number of mutations in each sub-atlas file respectively, and remove the sub-atlas files with a number of mutations less than or equal to the mutation threshold, and take the sub-atlas files with a number of mutations greater than the mutation threshold as the sub-atlas files to be processed; In the local interval, the distance normalization processing is performed on the sub-map files to be processed to obtain the individual-level genomic local haplotype distance; in the local interval, the distance normalization processing is performed on the individual-level genomic local haplotype distance corresponding to all the sub-map files to be processed belonging to the same group to obtain the group-level genomic local haplotype distance; At the whole genome level, the average of the individual-level genomic local haplotype distances corresponding to all sub-map files to be processed is calculated to obtain the individual-level haplotype distance; at the whole genome level, the average of the population-level genomic local haplotype distances corresponding to M populations is calculated to obtain the population-level haplotype distance.

2. The haplotype distance evaluation method according to claim 1, wherein The M×N whole-genome resequencing data are subjected to detection and processing of variant sites, and corresponding whole-genome variation maps are obtained according to the detected M×N variant sites, specifically: Each of the whole-genome resequencing data is filtered according to the set filtering criteria, and the filtered M×N whole-genome resequencing data are respectively aligned to the reference genome, and the compared M×N whole-genome resequencing data are sorted according to the genome sequence, and the sorted M×N whole-genome resequencing data are respectively detected for mutation sites using the GATK tool, and the whole-genome variation map corresponding to the M×N whole-genome resequencing data is obtained according to the detected M×N mutation sites.

3. The haplotype distance evaluation method according to claim 1, wherein The distance normalization process is performed on each of the sub-map files to be processed to obtain the individual-level genomic local haplotype distance, specifically: The two haplotypes in each sub-map file to be processed are respectively split into two hypothetical homozygous samples, and the two homozygous samples belonging to the same sub-map file are respectively input into the VCF2Dis tool in phased VCF format. The VCF2Dis tool is used to calculate each haplotype distance value and the average haplotype distance value corresponding to the same sub-map file to be processed, and then each haplotype distance value is divided by the average haplotype distance value to obtain the individual-level genomic local haplotype distance corresponding to the same sub-map file.

4. The haplotype distance evaluation method according to claim 1, wherein The distance normalization process is performed on the individual-level genomic local haplotype distances corresponding to all the sub-map files to be processed belonging to the same group to obtain the group-level genomic local haplotype distances, specifically: For M×N whole-genome resequencing data, the two haplotypes in each sub-map file to be processed are split into two hypothetical homozygous samples, and the two homozygous samples belonging to the same sub-map file are respectively input into the VCF2Dis tool in phased VCF format. The average haplotype distance value of the M×N individuals in the same sub-map file to be processed is calculated by the VCF2Dis tool, and then each haplotype distance value is divided by the average haplotype distance value to obtain the local haplotype distance of each individual. Then, the average of the local haplotype distances in the M populations is calculated to obtain the population-level genome local haplotype distances of the M populations.

5. The haplotype distance evaluation method according to any one of claims 1 to 4, characterized in that: The step of calculating the haplotype distance of an individual relative to the population is also included, specifically: The individual-level genomic local haplotype distance corresponding to the sample to be calculated is divided by the population-level genomic local haplotype distance corresponding to the population to be calculated. The result is the haplotype distance of the individual relative to the population.

6. A haplotype distance assessment device based on second-generation short-read sequences, characterized in that: include: An import module is used to import whole-genome resequencing data with second-generation short-read sequence characteristics, wherein the whole-genome resequencing data is M×N, where M is the number of populations and N is the number of samples in each population; The variant site detection module is used to detect variant sites in M×N whole-genome resequencing data and obtain corresponding whole-genome variation maps based on the detected M×N variant sites; A segmentation module is used to use M×N whole-genome variation maps as map files to be segmented, and to segment the map files to be segmented into multiple sub-map files with the same interval length according to the set interval length of the local haplotype distance; The elimination module is used to count the number of mutations in each sub-atlas file, and eliminate the sub-atlas files with a number of mutations less than or equal to the mutation threshold, and take the sub-atlas files with a number of mutations greater than the mutation threshold as the sub-atlas files to be processed; The distance evaluation processing module is used to perform distance normalization processing on the sub-map files to be processed in the local interval to obtain the local haplotype distance of the individual level genome; in the local interval, the distance normalization processing is performed on the local haplotype distance of the individual level genome corresponding to all the sub-map files to be processed belonging to the same group to obtain the local haplotype distance of the group level genome; It is also used to calculate the average of the individual-level local genomic haplotype distances corresponding to all sub-map files to be processed at the whole genome level to obtain the individual-level haplotype distance; at the whole genome level, it is used to calculate the average of the population-level local genomic haplotype distances corresponding to M populations to obtain the population-level haplotype distance.

7. The haplotype distance evaluation device according to claim 6, characterized in that: In the variant site detection module, the M×N whole genome resequencing data are processed for variant site detection, and corresponding whole genome variation maps are obtained according to the detected M×N variant sites, specifically: Each of the whole-genome resequencing data is filtered according to the set filtering criteria, and the filtered M×N whole-genome resequencing data are respectively aligned to the reference genome, and the compared M×N whole-genome resequencing data are sorted according to the genome sequence, and the sorted M×N whole-genome resequencing data are respectively detected for mutation sites using the GATK tool, and the whole-genome variation map corresponding to the M×N whole-genome resequencing data is obtained according to the detected M×N mutation sites.

8. The haplotype distance evaluation device according to claim 6, characterized in that: In the distance evaluation processing module, distance normalization processing is performed on the sub-map files to be processed to obtain the individual-level genomic local haplotype distance, specifically: The two haplotypes in each sub-map file to be processed are respectively split into two hypothetical homozygous samples, and the two homozygous samples belonging to the same sub-map file are respectively input into the VCF2Dis tool in phased VCF format. The VCF2Dis tool is used to calculate each haplotype distance value and the average haplotype distance value corresponding to the same sub-map file to be processed, and then each haplotype distance value is divided by the average haplotype distance value to obtain the individual-level genomic local haplotype distance corresponding to the same sub-map file.

9. The haplotype distance evaluation device according to any one of claims 6 to 8, characterized in that: In the distance evaluation processing module, the individual-level genomic local haplotype distances corresponding to all the sub-map files to be processed belonging to the same group are respectively subjected to distance standardization processing to obtain the group-level genomic local haplotype distances, specifically: For M×N individuals, the two haplotypes in each sub-map file to be processed are split into two hypothetical homozygous samples, and the two homozygous samples belonging to the same sub-map file are respectively input into the VCF2Dis tool in phased VCF format. The VCF2Dis tool is used to calculate the average haplotype distance value of the M×N individuals in the same sub-map file to be processed, and then each haplotype distance value is divided by the average haplotype distance value to obtain the local haplotype distance of each individual. Then, the average of the local haplotype distances in the M populations is calculated to obtain the population-level genomic local haplotype distances of the M populations.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a haplotype distance assessment method based on second-generation short read sequences as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Genome structure variation detection method, computing device and storage medium

    CN112669902A

  • Genome structure variation distribution detection method and detection device

    CN114141308A