Method for establishing population genome structure variation database, electronic equipment and storage medium

By combining the third-generation sequencing technology and the T2T reference genome, the population structure variation database establishment method is used to solve the problem of difficult to detect population genome structural variation in the existing technology, and the establishment of a high-precision and high-sensitivity population structure variation database is achieved, providing effective guidance for disease research.

CN120072070APending Publication Date: 2025-05-30ANNOROAD GENE TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510229835.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2025-02-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to detect structural variations in population genomes with high sensitivity, especially in the field of structural variations of large fragments.

Method used

Using third-generation sequencing technology and the T2T reference genome, by obtaining the single-sample structural variation data of each sample in the population, at least two single-sample data are merged, the population structure variation merge data is obtained, and data cleaning and classification are carried out based on the sample deletion rate and frequency information, and a high-precision population structure variation database is established.

Benefits of technology

It has achieved high sensitivity to detect structural mutations, providing reliable detection results for structural mutation research in patients with disease, effectively narrowing the screening range of pathogenic mutations, and improving the accuracy and sensitivity of the database.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005291155250000011
    Figure HDA0005291155250000011
  • Figure HDA0005291155250000012
    Figure HDA0005291155250000012
  • Figure HDA0005291155250000021
    Figure HDA0005291155250000021
Patent Text Reader

Abstract

The invention discloses a method for establishing a crowd structure variation database, an electronic device and a storage medium. The method comprises the following steps: acquiring single sample structure variation data of each sample in a crowd; performing crowd merging on the at least two single sample structure variation data to obtain crowd structure variation merged data; obtaining a sample missing rate of each structure variation in the crowd structure variation merged data; and retaining the structure variation of which the sample missing rate is not greater than a first threshold value to obtain a crowd structure variation database. According to the database establishment method, the electronic device and the storage medium, the reliable crowd structure variation database can be established, and the database is high in precision and sensitivity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a method for establishing a population genomic structural variation database, an electronic device, and a storage medium. Background Art

[0002] Variations in the human genome, especially genomic structural variations (SVs), are closely related to aspects such as evolution and disease risk. Therefore, constructing a population genomic structural variation database is of great significance in both scientific research and clinical applications. At present, although international organizations have constructed SNP databases for a large number of populations, there is still a lack of such a database in the field of large-fragment (>50bp) structural variations. Summary of the Invention

[0003] Detecting genomic structural variations in a population is of great significance in both scientific research and clinical applications. On the one hand, it can provide reliable control samples for the study of structural variations in patients with genetic diseases. On the other hand, the annotation and prediction of variant functions will effectively narrow down the screening range of pathogenic mutations and provide effective guidance and assistance for research in related fields. However, how to detect structural variations with high sensitivity is an urgent problem to be solved. This application precisely proposes a solution in combination with third-generation sequencing technology and the T2T reference genome for this problem.

[0004] The detection method provided by this application can provide reliable detection results for the study of structural variations in disease patients, effectively narrow down the screening range of pathogenic mutations, and provide effective guidance for disease research.

[0005] Specifically, this application adopts the following technical solutions.

[0006] 1. In one aspect of this application, at least one embodiment provides a method for establishing a population structural variation database, including:

[0007] Obtaining single-sample structural variation data for each sample in the population;

[0008] Performing population merging on at least two of the single-sample structural variation data to obtain population structural variation merged data;

[0009] Obtaining the sample deletion rate for each structural variation in the population structural variation merged data;

[0010] Retaining the structural variations with a sample deletion rate not greater than a first threshold to obtain a population structural variation database.

[0011] 2. According to the method described in item 1, after obtaining the population structural variation merged data, it further includes:

[0012] Obtain the genotype of each structural variation in the population structural variation combined data in each sample.

[0013] 3. The method according to item 2, obtaining the genotype of each structural variation in the population structural variation combined data in each sample, includes:

[0014] Obtain the depth of each structural variation in each sample and the number of reads supporting the variation;

[0015] Obtain the genotype based on the depth and the number of reads; wherein,

[0016] If the depth < 8, the genotype is. / .;

[0017] If the depth ≥ 8 and the number of reads / depth < 0.1, the genotype is 0 / 0;

[0018] If the depth ≥ 8 and the number of reads / depth ∈ [0.1, 0.9), the genotype is 0 / 1;

[0019] If the depth ≥ 8 and the number of reads / depth ≥ 0.9, the genotype is 1 / 1.

[0020] 4. The method according to any one of items 1 - 3, merging at least two of the single - sample structural variation data for the population to obtain population structural variation combined data, includes:

[0021] Obtain the data of the same variation type in all the single - sample structural variation data;

[0022] Obtain the positions of the data of the same variation type and the duplication rate between them;

[0023] Merge two or more data located in the repetitive element region and with a duplication rate not less than the second threshold; or, merge two or more data located in the non - repetitive element region and with a duplication rate not less than the third threshold.

[0024] 5. The method according to item 4, the breakpoint information of the population structural variation combined data includes:

[0025] The median of the breakpoint information of the two or more data; and / or,

[0026] The extreme value of the breakpoint information of the two or more data; and / or,

[0027] The average value of the breakpoint information of the two or more data.

[0028] 6. The method according to any one of items 1 - 5, further includes:

[0029] Obtain the frequency information of each structural variation in the population structure variation database in the population;

[0030] Classify each structural variation in the multi-sample structure database according to the frequency information:

[0031] Singleton: allele count = 1;

[0032] Rare: allele count > 1 and AF ≤ 0.05;

[0033] Low: AF > 0.05 and AF ≤ 0.1;

[0034] Common: AF > 0.1;

[0035] Mark the classification information of each structural variation in the multi-sample structure variation database.

[0036] 7. According to the method of any one of items 1-7, the first threshold ≤ 15%.

[0037] 8. According to the method of items 4-6, wherein the second threshold is greater than the third threshold;

[0038] Preferably, the second threshold ≥ 80%; and / or,

[0039] Preferably, the third threshold ≥ 50%.

[0040] 9. In one aspect of the present application, at least one embodiment provides an electronic device, including:

[0041] A memory storing computer-executable instructions;

[0042] A processor configured to run the computer-executable instructions, wherein when the computer-executable instructions are run by the processor, the method described in any one of items 1-12 is implemented.

[0043] 10. In one aspect of the present application, at least one embodiment provides a storage medium storing computer-executable instructions, and when the computer-executable instructions are executed by a processor, the method described in any one of items 1-12 is implemented.

[0044] One aspect of the present invention relates to a method for processing structural variations, including the following steps,

[0045] Obtain the sequencing data of the sample;

[0046] Use N structural variation detection methods to detect the sequencing data of the sample to obtain N initial detection results of structural variations of the sample;

[0047] Perform a first merge on the N initial structural variation detection results of the sample to obtain the merged structural variation detection result of the sample;

[0048] Select the merged structural variation detection results supported by at least two structural variation detection methods to obtain the single-sample structural variation results, where N is a positive integer greater than 1.

[0049] In a specific embodiment, in the foregoing method, the structural variation detection methods include:

[0050] At least one alignment-based structural variation detection method; and / or,

[0051] At least one assembly-based structural variation detection method; and / or,

[0052] At least one machine learning-based structural variation detection method.

[0053] In a specific embodiment, in the foregoing method, the alignment-based structural variation detection methods include pbSV and cuteSV;

[0054] And / or, the assembly-based structural variation detection methods include Svim and Mumandco;

[0055] And / or, the machine learning-based structural variation detection method includes Svision.

[0056] In a specific embodiment, in the foregoing method, the first merge includes:

[0057] Obtain the detection data of the same variation type in the N initial structural variation detection results;

[0058] Obtain the positions of the detection data of the same variation type and the duplication rate between them;

[0059] Merge two or more detection data located in the repetitive element region and having a duplication rate greater than or equal to the first threshold; or, merge two or more detection data located in the non-repetitive element region and having a duplication rate greater than or equal to the second threshold.

[0060] In a specific embodiment, in the foregoing method, the breakpoint information of the merged structural variation detection result of the sample includes:

[0061] The median of the breakpoint information of the two or more detection data; and / or,

[0062] The extreme values of the breakpoint information of the two or more detection data; and / or,

[0063] The average of the breakpoint information of the two or more detection data.

[0064] In a specific embodiment, in the foregoing method, it further includes:

[0065] Performing a second merging on the M single-sample structural variation results to obtain a multi-sample structural variation result; wherein, the second merging includes:

[0066] Obtaining the detection data of the same variation type in the M single-sample structural variation results;

[0067] Obtaining the positions of the detection data of the same variation type and the duplication rate between them;

[0068] Merging two or more detection data located in the repetitive element region and having a duplication rate greater than or equal to a first threshold; or, merging two or more detection data located in the non-repetitive element region and having a duplication rate greater than or equal to a second threshold.

[0069] In a specific embodiment, in the foregoing method, the breakpoint information of the multi-sample structural variation result includes:

[0070] The median of the breakpoint information of the two or more detection data; and / or,

[0071] The extreme values of the breakpoint information of the two or more detection data;

[0072] The average value of the breakpoint information of the two or more detection data.

[0073] 8. The method according to item 6 or 7, further including:

[0074] Obtaining the sample deletion rate of each structural variation in the multi-sample structural variation result;

[0075] Removing the structural variations with a sample deletion rate greater than or equal to a third threshold in the multi-sample structural variation result to obtain a multi-sample structural variation database;

[0076] Preferably, the third threshold is greater than or equal to 15%.

[0077] In a specific embodiment, in the foregoing method, before obtaining the sample deletion rate of each structural variation in the multi-sample structural variation result, it further includes:

[0078] Obtaining the genotype of each structural variation in the multi-sample structural variation result in the M samples.

[0079] In a specific embodiment, in the foregoing method, obtaining the depth and the number of reads supporting the variation in the sequencing data of each sample for each structural variation in the multi-sample structural variation result;

[0080] If the depth < 8, the genotype is. / .

[0081] If the depth ≥ 8 and alt_num / depth < 0.1, the genotype is 0 / 0;

[0082] If the depth ≥ 8 and alt_num / depth ∈ [0.1, 0.9), the genotype is 0 / 1;

[0083] If the depth ≥ 8 and alt_num / depth ≥ 0.9, the genotype is 1 / 1.

[0084] In a specific embodiment, in the foregoing method, it further includes:

[0085] Obtaining the frequency information of each structural variation in the multi-sample structural variation database in the M samples;

[0086] Classifying each structural variation in the multi-sample structural database according to the frequency information; wherein, the classification includes:

[0087] Singleton: allele count = 1;

[0088] Rare: allele count > 1 and AF ≤ 0.05;

[0089] Low: AF > 0.05 and AF ≤ 0.1;

[0090] Common: AF > 0.1;

[0091] Marking the classification information of each structural variation in the multi-sample structural variation database.

[0092] In a specific embodiment, in the foregoing method, the first threshold is greater than the second threshold;

[0093] Preferably, the first threshold ≥ 80%; and / or,

[0094] Preferably, the second threshold ≥ 50%

[0095] The inventors of the present application found in the research that there are many problems when the structural variation sets of different individuals in the population are merged, and the merged database often has various problems. Existing merging methods include Survivor or SV-merge. The former will lose the genotype information of the vast majority of individuals, and the latter's merging effect has not been clearly evaluated.

[0096] To at least solve the above problems, the database establishment method, electronic device, and storage medium provided in this application can establish a reliable population structure variation database, and the accuracy and sensitivity of the database are high. On the one hand, the database can provide reliable control samples for the study of structural variations in patients with genetic diseases. On the other hand, the annotation and prediction of variant functions will effectively narrow down the screening range of pathogenic mutations, providing effective guidance and help for research in related fields. Description of the Drawings

[0097] The drawings are used to better understand this application and do not constitute an improper limitation to this application. Among them:

[0098] Figure 1 It is a distribution diagram of sample missing rates obtained by the database establishment method provided in the embodiment of this application;

[0099] Figure 2 It is a frequency distribution diagram of population structure variations obtained by the database establishment method provided in the embodiment of this application;

[0100] Figure 3 It is a classification diagram of population structure variations provided in the embodiment of this application;

[0101] Figure 4a It is a box plot for detecting and counting the structural variations of 64 samples using the cuteSV analysis software;

[0102] Figure 4b It is a box plot for detecting and counting the structural variations of 64 samples using the pbSV analysis software;

[0103] Figure 4c It is a box plot for detecting and counting the structural variations of 64 samples using the Svision analysis software;

[0104] Figure 4d It is a box plot for detecting and counting the structural variations of 64 samples using the Svim analysis software;

[0105] Figure 4e It is a box plot for detecting and counting the structural variations of 64 samples using the Mumsndco analysis software;

[0106] Figure 5a It is a violin plot for detecting and counting the structural variations of 64 samples using the cuteSV analysis software after removing inv and incorporating dup into ins;

[0107] Figure 5b It is a violin plot for detecting and counting the structural variations of 64 samples using the pbSV analysis software after removing inv and incorporating dup into ins;

[0108] Figure 5cAfter removing inv and incorporating dup into ins, a violin plot of the structural variations of 64 samples was detected and statistically analyzed using Svision analysis software;

[0109] Figure 5d After removing inv and incorporating dup into ins, a violin plot of the structural variations of 64 samples was detected and statistically analyzed using Svim analysis software;

[0110] Figure 5e After removing inv and incorporating dup into ins, a violin plot of the structural variations of 64 samples was detected and statistically analyzed using Mumsndco analysis software;

[0111] Figure 6 The result after performing single-sample merging on 64 samples after removing inv and incorporating dup into ins. Detailed implementation manner

[0112] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions involved in this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the specific implementation manners described are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0113] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices. When the quantity of a component is not specifically indicated in the following text of the embodiments of this application, it means that the component can be one or more, or can be understood as at least one. "At least one" means one or more, and "a plurality" means at least two.

[0114] The following describes the method for establishing a population variation database, an electronic device, and a storage medium provided by the present application in combination with specific embodiments. It should be noted that the same components or the same operation processes can adopt the same setting methods. All embodiments of the present application are applicable to the above-mentioned multiple protection subjects, and the same or similar content will not be repeated multiple times in each protection subject. Reference can be made to the descriptions in the corresponding embodiments of other protection subjects.

[0115] Allele count: The number of variant alleles. When it is equal to 1, it means that the structural variation appears only in one sample; when it is greater than 1, it means that the structural variation appears in multiple samples.

[0116] Allele frequency, AF: Gene frequency. In the embodiments of the present application, it can refer to the proportion of different structural variations in the population.

[0117] Number, SV number: The number of structural variations

[0118] Missing rate: Missing rate

[0119] Count: The sum of the number of SVs within the allele frequency range

[0120] At least one embodiment of the present application provides a method for establishing a population structural variation database, including: obtaining the single-sample structural variation data of each sample in the population; merging at least two of the single-sample structural variation data for the population to obtain population structural variation merged data; obtaining the sample missing rate of each structural variation in the population structural variation merged data; and retaining the structural variations with a sample missing rate less than a first threshold to obtain a population structural variation database.

[0121] For example, the above population may include at least two samples. For example, the samples in the population are different from each other. For example, each sample in the population has its own corresponding single-sample structural variation data. For example, the number of single-sample structural variation data obtained is the same as the number of samples in the population and they correspond to each other one by one.

[0122] For example, single-sample structural variation data can be obtained by at least one detection method among next-generation sequencing, PacBio third-generation sequencing, Nanopore third-generation sequencing, chromosomal microarray (CMA), single nucleotide polymorphism array (SNP array), array-based comparative genomic hybridization (aCGH), and multiplex ligation-dependent probe amplification (MLPA).

[0123] For example, structural variations include at least one of deletion (del), duplication (dup), inversion (inv), and insertion (ins). For example, single-sample structural variation data includes more than one structural variation. For example, single-sample structural variation data includes at least one type of structural variation.

[0124] For example, population pooling can be pooling the single-sample structural variation data of all samples within the population set; or, it can also be pooling the single-sample structural variation data of some samples within the population set.

[0125] For example, the population structural variation pooled data obtained after population pooling is a non-redundant structural variation set, and it can be considered that any two or more pieces of population structural variation pooled data cannot be further pooled. For example, among individual single-sample structural variation data, there are often overlapping or partially overlapping structural variations; when constructing a population structural variation database, these overlapping or partially overlapping structural variations need to be pooled, which can improve the accuracy and practicality of the database and avoid interference caused by redundant data.

[0126] For example, the structural variation database after population pooling also needs to be cleaned to further improve the accuracy of the database; for example, some structural variation data with a large sample missing rate needs to be cleaned. For example, structural variation data with a large sample missing rate may be caused by genetic deletions and / or detection deletions. For example, detection deletions may be caused by limitations, errors, insufficient depth, etc. of the detection technology; therefore, when the sample missing rate of a structural variation is greater than the first threshold, it is removed; or rather, structural variations with a sample missing rate not greater than the first threshold are retained, and the set of these retained structural variation data is used as the population structural variation database.

[0127] The method for establishing a population structural variation database provided by at least one embodiment of the present application merges the structural variation data in multiple single samples through the original population merging, removes redundant data while retaining the comprehensiveness of the structural variation data as much as possible, and improves the accuracy of the database; on the other hand, the method provided by at least one embodiment of the present application also cleans the data based on the sample missing rate, further improving the accuracy of the population structural variation database. The database has high precision and sensitivity, can provide reliable control samples for research related to structural variations, and also makes up for the gap in building a population structural variation database.

[0128] Based on the above at least one embodiment, the first threshold ≤ 15%. For example, the first threshold can be 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14% or 15%. For example, retain the structural variations with a sample missing rate not greater than 5%. For example, retain the structural variations with a sample missing rate not greater than 9%. For example, retain the structural variations with a sample missing rate not greater than 10%.

[0129] Based on the above at least one embodiment, after obtaining the population structural variation merged data, it further includes: obtaining the genotype of each structural variation in the population structural variation merged data in each sample. For example, the genotype of each structural variation in each sample can be obtained based on the sequencing data of the sample. For example, the genotype can be. / ., 0 / 0, 0 / 1 or 1 / 1.

[0130] For example, the missing rate can be obtained based on the genotype of each structural variation in each sample. For example, the sample missing rate of each structural variation can be defined as: the number of samples with the genotype. / . of this structural variation / the total number of samples. For example, for a population of 64 samples, the population structural variation merged data contains structural variation A, where the genotype of structural variation A in 30 samples is 0 / 0, the genotype of structural variation A in 30 samples is 0 / 1, and the genotype of structural variation A in 4 samples is. / .; then the sample missing rate of structural variation A in this population is 4 / 64 = 6.25%.

[0131] Based on the above at least one embodiment, obtaining the genotype of each structural variation in the population structural variation merged data in each sample includes: obtaining the depth of each structural variation in each sample and the number of reads supporting the variation; obtaining the genotype based on the depth and the number of reads.

[0132] For example, the above depth can be the sequencing depth; for example, the number of reads can be the number of sequencing reads.

[0133] A smaller sequencing depth may indicate lower credibility of this structural variation. After research by the inventors of the present application, it is found that:

[0134] If the depth < 8, the genotype is. / .;

[0135] If the depth ≥ 8 and the reads count / depth < 0.1, the genotype is 0 / 0;

[0136] If the depth ≥ 8 and the reads count / depth ∈ [0.1, 0.9), the genotype is 0 / 1;

[0137] If the depth ≥ 8 and the reads count / depth ≥ 0.9, the genotype is 1 / 1.

[0138] Based on at least one of the above embodiments, at least two of the single-sample structural variation data are combined for the population to obtain population structural variation combined data, including: obtaining data of the same variation type among all the single-sample structural variation data; obtaining the positions of the data of the same variation type and the duplication rate between them; combining two or more data located in the repetitive element region and having a duplication rate not less than a second threshold; or combining two or more data located in the non-repetitive element region and having a duplication rate not less than a third threshold.

[0139] For example, combining all dels in the single samples includes: obtaining the positions of all dels and the duplication rate between them; for example, dividing dels into those located in the repetitive element region and those located in the non-repetitive element region. For dels located in the repetitive element region, if the duplication rate (such as overlap) between two or more dels exceeds 80%, then the above two or more dels are combined; for dels located in the non-repetitive element region, if the duplication rate (such as overlap) between two or more dels exceeds 50%, then the above two or more dels are combined. Other types of structural variations such as Ins, dup, and inv can also be combined according to the above method.

[0140] The method provided by the embodiments of the present application combines the structural variations that are repeated or partially repeated among the individual single samples, removes the redundant part in the database, can improve the accuracy of the database, and the results are accurate and reliable.

[0141] Based on at least one of the above embodiments, the second threshold is greater than the third threshold. For example, the second threshold ≥ 80%; for example, the third threshold ≥ 50%. For example, the second threshold is 80% and the third threshold is 50%. For example, the second threshold is 90% and the third threshold is 60%. The embodiments of the present application study the respective suitable duplication rate thresholds for the repetitive element region and the non-repetitive element region, and combine the structural variations that meet the corresponding thresholds, further improving the accuracy and precision of the database.

[0142] Based on the above at least one embodiment, the breakpoint information of the population structure variant merged data includes: the median of the breakpoint information of two or more data; and / or, the extreme values of the breakpoint information of two or more data; and / or, the average of the breakpoint information of two or more data.

[0143] For example, for the structural variants after population merging, their breakpoint information is obtained from the breakpoint information of two or more structural variants before merging; for example, the merged breakpoint can be the median of the breakpoints of the above two or more structural variants. For example, the merged breakpoint can be the extreme values of the breakpoints of the above two or more structural variants; for example, it can be the maximum value and / or the minimum value. For example, the merged breakpoint can be the average of the breakpoints of the above two or more structural variants. By analyzing and processing the breakpoint information, the embodiments of the present application determine the breakpoint information of the merged structural variants, and the results are accurate and reliable, which can improve the accuracy and precision of the database.

[0144] Based on the above at least one embodiment, it further includes: obtaining the frequency information of each structural variant in the population structure variant database in the population; classifying each structural variant in the multi-sample structure database according to the frequency information:

[0145] Singleton: allele count = 1;

[0146] Rare: allele count > 1 and AF ≤ 0.05;

[0147] Low: AF > 0.05 and AF ≤ 0.1;

[0148] Common: AF > 0.1;

[0149] Mark the classification information of each structural variant in the multi-sample structure variant database.

[0150] For example, annotate each structural variant in the multi-sample structure variant database. For example, the annotation information includes at least one of the variant rating, classification information, related disease information, penetrance, occurrence frequency, fragment size, locus / region annotation, and mutation-caused function annotation of the structural variant. For example, the annotation and prediction of the variant function will effectively narrow the screening range of pathogenic mutations and provide effective guidance and help for research in related fields.

[0151] For example, in the embodiments of the present application, structural variations are classified according to frequency information. For example, the frequency information includes at least one of allele frequency and allele count. For example, the classification of structural variations includes Singleton, Rare, Low, and Common; for example, Singleton represents that the allele count of the structural variation is 1, that is, this structural variation only appears in one sample; for example, Rare represents that the allele count of the structural variation is greater than 1 and the allele frequency ≤ 0.05, that is, this structural variation appears in more than 1 person and the frequency in the population is small, representing rare; for example, Low represents that the allele frequency of the structural variation is between 0.05 and 0.1, representing low incidence; for example, Common represents that the allele frequency of the structural variation is greater than 0.1, representing common. By using the method provided in the embodiments of the present application to build a population structural variation database, the structural variations in the database can be annotated and classified, which can effectively narrow the screening range and provide effective guidance and help for research in related fields, such as research on disease-related structural variations.

[0152] In one aspect of the present application, at least one embodiment provides an electronic device, including:

[0153] A memory storing computer-executable instructions;

[0154] A processor configured to run the computer-executable instructions, wherein when the computer-executable instructions are run by the processor, the method for establishing the aforementioned population structural variation database is implemented.

[0155] In one aspect of the present application, at least one embodiment provides a storage medium storing computer-executable instructions, and when the computer-executable instructions are executed by a processor, the method for establishing the aforementioned population structural variation database is implemented.

[0156] For example, the processor can control other components in the electronic device to perform desired functions. The processor can be a central processing unit (CPU), a network processor (NP), a tensor processing unit (TPU), or a graphics processing unit (GPU) and other devices with data processing capabilities and / or program execution capabilities; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. For example, the central processing unit (CPU) can be based on architectures such as X86, RISC-V, or ARM.

[0157] For example, the memory may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer-readable instructions may be stored on the computer-readable storage media, and the processor may run the computer-readable instructions to implement various functions of the electronic device. Various application programs and various data may also be stored in the storage media.

[0158] For example, for a detailed description of the process of the electronic device executing the data acquisition method, reference may be made to the relevant description in the embodiments of the data acquisition method, and repeated parts will not be elaborated.

[0159] Constructing a population genomic structural variation dataset is of great significance in both scientific research and clinical applications. On the one hand, the database can provide reliable control samples for the study of structural variations in genetic disease patients. On the other hand, the annotation and prediction of variant functions will effectively narrow down the screening range of pathogenic mutations and provide effective guidance and assistance for research in related fields. However, how to construct a high-precision structural variation database is an urgent problem to be solved. This application is exactly a solution proposed in combination with third-generation sequencing technology and the T2T reference genome for this problem.

[0160] The T2T (Telomere-to-Telomere) genome refers to a 0-gap genome assembled from telomere to telomere level of one or more chromosomes through the combination of various sequencing technologies such as PacBio HiFi, ONT Ultra-long, and Hi-C. The assembly of the T2T genome enables the exploration of unknown areas such as telomeres and centromeres of the genome, and also provides a more in-depth research direction for research.

[0161] The attached drawings or the embodiments part explains the meaning

[0162] The present application will be described below with reference to specific embodiments. These embodiments are merely illustrative and should not be construed as limiting the present invention.

[0163] Embodiment 1

[0164] Step 1

[0165] Obtain the structural variation data of each sample in a population consisting of 64 individuals; the number of inv detections in the single-sample structural variation data is very small, and the dup and ins types are similar. For the convenience of subsequent analysis, merge dup into ins and only retain the two main types of structural variations, del and ins, for subsequent analysis.

[0166] Statistics on the structural variation data of 64 samples are obtained, as shown in Table 1.

[0167] Table 1 Statistical table of structural variations in 64 samples

[0168] All ins del Mean 22323.66 10534.41 11755.13 SD 492.45 244.66 261.88

[0169] Step 2

[0170] Merge the single-sample structural variation data of 64 samples to obtain population structural variation merged data, which is a non-redundant structural variation collection, and the number of structural variations is 146,071.

[0171] Step 3

[0172] Using the genotypes of the non-redundant SVs in each sample after merging, obtain the sample missing rate of each structural variation and conduct statistics. The results are as Figure 1 and Table 2 show; it shows the distribution relationship between the proportion of SVs and the sample missing rate. For example, for 65.18% of the SVs, the sample missing rate is 0.

[0173] Table 2 Statistical table of missing situations after population merging

[0174] Number of deletions Deletion rate Ratio 0 0 65.18% 6 9.38% 84.18% 9 14.06% 87.43% 13 20.31% 90.52%

[0175] Step 4

[0176] Set the first threshold to 10%, and retain the structural variations with a sample missing rate not greater than 10% to obtain a population structural variation database. At this time, the database contains 122,917 structural variations.

[0177] Comprehensively Figure 1 and Table 2, it can be seen that if a structural variation is missing in more than 6 samples, remove this structural variation, that is, the sample missing rate is required to be not greater than 10%.

[0178] Step 5

[0179] For the structural variations that meet the sample missing rate, calculate their frequency distribution to obtain the frequency distribution diagram of the population structural variation database, as Figure 2 shown; Figure 2 The abscissa in is the gene frequency, and the ordinate is the sum of the number of structural variations in the range under each gene frequency.

[0180] For exampleFigure 2 Among them, as the gene frequency increases, the sum of the number of structural variations in each gene frequency interval decreases in turn, indicating that there are many SVs with low gene frequencies in this population collection.

[0181] Step 6

[0182] Set the frequency classification criteria, and divide the population structural variations into four categories. The results are as Figure 3 shown. Among them:

[0183] Singleton: allele count = 1;

[0184] Rare: allele count > 1 and AF ≤ 0.05;

[0185] Low: AF > 0.05 and AF ≤ 0.1;

[0186] Common: AF > 0.1;

[0187] It can be seen that:

[0188] 1) The average number of structural variations in a single sample in Step 1 is 22,323. For a population of 64 samples, the number of structural variations it includes is 64 * 22,323 = 1,473,318; obviously, it is difficult to perform subsequent analysis on such a large number of structural variations and it is difficult to construct a high-precision database. After Step 2, that is, after performing population merging, the number of merged data of population structural variations is only 146,071; the method provided in the embodiments of the present application merges redundant variations and greatly improves the accuracy of the database.

[0189] 2) In Steps 3 and 4, the deletion rates of each structural variation are obtained, and the structural variations with high deletion rates are excluded. It may be due to factors such as insufficient sequencing depth and sequencing preference that these structural variations are missing in some samples. After research, the inventor sets the first threshold at 10%; by removing the structural variations with a deletion rate greater than 10%, the number of structural variations is reduced from 146,071 to 122,917, further improving the accuracy and precision of the database.

[0190] 3) In Step 5, the frequency distribution is calculated for the structural variations that meet the sample deletion rate; through the frequency distribution diagram, the distribution characteristics of the structural variations in this population, the distribution of various structural variations in the population, etc. can be seen, such as the distribution of a certain pathogenic structure in the population.

[0191] 4) In step 6, set the frequency classification criteria to divide the population structural variations into four categories (including Singleton, Rare, Low, and Common). Adding the classified database is more convenient for use. For example, it can provide a reliable control sample library for the study of structural variations in disease patients. By screening low-frequency structural variations, the screening range of pathogenic mutations can be effectively narrowed, providing effective guidance for disease research.

[0192] Example 2

[0193] Based on the above Example 1, before step 1, it further includes:

[0194] Step A

[0195] Perform third-generation sequencing on a population consisting of 64 individuals using Pacbio third-generation sequencing, with an average depth of 20.6×.

[0196] Step B

[0197] Analyze the number of structural variations in each sample using 5 testing methods respectively (for example, 5 structural variation detection and analysis software), including:

[0198] Based on the alignment of cuteSV, detect and count the structural variation situations of 64 samples, as Figure 4a shown.

[0199] Based on the alignment of pbSV, detect and count the structural variation situations of 64 samples, as Figure 4b shown.

[0200] Based on the machine learning of Svision, detect and count the structural variation situations of 64 samples, as Figure 4c shown.

[0201] Based on the assembly of Svim, detect and count the structural variation situations of 64 samples, as Figure 4d shown.

[0202] Based on the assembly of Mumandco, detect and count the structural variation situations of 64 samples, as Figure 4e shown.

[0203] Since the number of inv detections is very small and the dup and ins types are similar, for the convenience of subsequent analysis, merge dup into ins and only retain the two most main types of variations, del and ins, for subsequent analysis. The result after simplifying the structural variation types is as Figure 5a - Figure 5e shown.

[0204] Step C

[0205] The detection results of the five detection methods are merged for a single sample, and the structural variations supported by at least two methods are selected as the results of the single sample; after the 64 single samples are respectively merged for the single sample, the statistical results are as Figure 6 shown in Table 1.

[0206] It can be seen that:

[0207] 1) The structural variations detected by Svim, cuteSV, pbSV and Svision in step B are about 20,000 - 30,000; the structural variations detected by Mumandc are about 10,000 - 12,000. That is to say, for each single sample, if you want to achieve comprehensive detection based on detection methods with different strategies, for example, using 5 software for detection, the detected structural variations are about 90,000 - 132,000; this is only the structural variation of a single sample. For a population set of multiple samples, the number of structural variations is even larger and the analysis difficulty is extremely high.

[0208] 2) After performing the single - sample merging in step C, the number of single - sample SVs can be reduced from 90,000 - 132,000 to about 22,323. It not only comprehensively detects structural variations using different strategies, but also greatly improves the accuracy of structural variation detection, providing strong support for subsequent population analysis and database construction, etc.

[0209] Although the above combination has described the implementation scheme of the present application, the present application is not limited to the above - mentioned specific implementation schemes and application fields. The above - mentioned specific implementation schemes are only illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present application, and these all belong to the scope of protection of the present application.

Claims

1. A method for establishing a population structure variation database, comprising: Obtain single-sample structural variation data for each sample in the population; Merging at least two of the single sample structural variation data into populations to obtain population structural variation merged data; Obtaining the sample missing rate of each structural variation in the combined data of structural variation of the population; The structural variations whose sample missing rate is not greater than a first threshold are retained to obtain a population structural variation database.

2. The method according to claim 1, characterized in that: After obtaining the combined data of population structure variation, the method further includes: The genotype of each structural variation in the combined data of population structural variation in each sample is obtained.

3. The method according to claim 2, characterized in that Obtaining the genotype of each structural variation in the combined data of population structural variation in each sample includes: Obtain the depth of each structural variation in each sample and the number of reads supporting the variation; The genotype is obtained based on the depth and the number of reads; wherein, If the depth is < 8, the genotype is . / .; If the depth is ≥ 8 and the number of reads / depth is < 0.1, the genotype is 0 / 0; If depth ≥ 8 and reads / depth ∈ [0.1, 0.9), the genotype is 0 / 1; If the depth is ≥ 8 and the number of reads / depth is ≥ 0.9, the genotype is 1 / 1.

4. The method according to any one of claims 1 to 3, characterized in that: Merging at least two of the single sample structural variation data into populations to obtain population structural variation merged data, including: Acquire data of the same variation type in all the single-sample structural variation data; Obtaining the locations of the data of the same variation type and the repetition rates therebetween; Two or more data located in the repeated element region and having a repetition rate not less than a second threshold are merged; or, two or more data located in the non-repetitive element region and having a repetition rate not less than a third threshold are merged.

5. The method according to claim 4, characterized in that The breakpoint information of the combined data of population structure variation includes: The median of the breakpoint information of the two or more data; and / or, The extreme values ​​of the breakpoint information of the two or more data; and / or, The average of the breakpoint information of the two or more data.

6. The method according to any one of claims 1 to 5, further comprising: Obtain frequency information of each structural variation in the population in the population structural variation database; Classify each structural variation in the multi-sample structural database according to the frequency information: Singleton: allele count=1; Rare: allele count>1and AF≤0.05; Low: AF>0.05and AF≤0.1; Common: AF>0.1; Classification information of each structural variation is marked in the multi-sample structural variation database.

7. The method according to any one of claims 1 to 7, characterized in that: The first threshold is ≤15%.

8. The method according to claims 4-6, wherein: The second threshold is greater than the third threshold; Preferably, the second threshold value is ≥ 80%; and / or, Preferably, the third threshold is ≥50%.

9. An electronic device, comprising: a memory storing computer executable instructions; A processor is configured to execute the computer executable instructions, wherein the computer executable instructions, when executed by the processor, implement the method according to any one of claims 1 to 12.

10. A storage medium storing computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 12.