A method and device for detecting alpha-globin gene copy number variation

By constructing a combined model of high-throughput sequencing data and using a machine learning algorithm to analyze the sequencing depth information of the homologous recombination region of the α-globin gene cluster, the problem of inaccurate detection of α-globin copy number variation types in existing technologies was solved, and rapid and accurate α-globin copy number variation typing was achieved.

CN114512187BActive Publication Date: 2025-10-10TIANJIN MEDICAL LAB BGI +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210161840.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-22
Publication Date
2025-10-10
Estimated Expiration
2042-02-22

AI Technical Summary

Technical Problem

Existing methods for detecting α-globin gene copy number variations have low throughput and limited detection types. High-throughput sequencing platforms lack effective analysis methods, making it difficult to accurately type α-globin copy number variations, including triplets.

Method used

A combined model based on high-throughput sequencing data was constructed, and the sequencing depth information of the homologous recombination region of the α-globin gene cluster was analyzed through machine learning algorithms. A three-level detection module was constructed to achieve accurate type identification of α-globin gene copy number variations.

Benefits of technology

It achieves rapid and accurate typing of α-globin copy number variations, including triplets, and improves the accuracy and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114512187B_ABST
    Figure CN114512187B_ABST
Patent Text Reader

Abstract

A method and device for detecting alpha-globin gene copy number variation, the method comprising: a primary detection step, comprising predicting whether the sample to be detected exists alpha-globin gene deletion according to the overall feature matrix of the sequencing data of the sample to be detected; a secondary detection step, comprising analyzing the alpha-globin gene copy number of the sample to be detected according to the prediction result of the primary detection step; and a tertiary detection step, comprising analyzing the accurate type of the alpha-globin gene copy number variation of the sample according to the prediction result of the secondary detection step. The present application constructs a combined model for identifying alpha-globin copy number variation, which can quickly and accurately identify the type of alpha-globin copy number variation including triplets using high-throughput sequencing data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics, and in particular to a method and device for detecting α-globin gene copy number variation. Background Art

[0002] Thalassemia is one of the most common human genetic diseases in the world. It is caused by copy number variation or single nucleotide mutation in the α-globin or β-globin gene, resulting in reduced or absent synthesis of α-globin chains or β-globin chains, thus causing hemolytic anemia. It is estimated that about 7% of the world's population carries the pathogenic gene for this disease. Thalassemia is most common in the Mediterranean region, eastern South Asia, the Indian subcontinent and southern China. The most common of these is α-thalassemia (abbreviated as α-thalassemia), which is mainly caused by the deletion of one or two α-globin genes (HBA1 and HBA2) located in the telomeric region of chromosome 16 (16p13.3). Normal people have 4 copies of the α-globin gene (usually written as αα / αα). Individuals who have lost one copy of the α-globin gene (remaining 3 copies) are silent α-thalassemia carriers (including αα / -α 3.7 and αα / -α 4.2 Two types), those who are missing two copies of the α-globin gene are carriers of mild α-thalassemia (including αα / -- SEA 、-α 3.7 / -α 3.7 、-α 4.2 / -α 4.2 、-α 3 . 7 / -α 4.2 Types), patients with 3 copies missing are α-thalassemia intermediate patients (or Hb-H disease) (including -α 3.7 / -- SEA and -α 4.2 / -- SEA Two types), and the deletion of 4 copies is a major α-thalassemia patient (or Hb-Barts fetal hydrops syndrome). Hb-Barts fetal hydrops syndrome can cause fetal hydrops and death. In addition, α-globin triad (5 copies) (including αα / ααα anti3.7 and αα / ααα anti4.2 The combination of α-globin gene mutations and α-globin gene mutations can aggravate the clinical symptoms of patients with β-thalassemia. Therefore, accurate diagnosis of α-globin copy number variation is of great clinical significance.

[0003] Currently, common methods for detecting α-thalassemia include routine blood analysis (HbA2 content, HbF content, mean corpuscular volume (MCV), mean corpuscular hemoglobin (MCH), and abnormal hemoglobin electrophoresis), gap-polymerase chain reaction (gap-PCR), reverse dot blot hybridization (RDB), quantitative PCR (qPCR), and Sanger sequencing. Gap-PCR is the primary method for detecting α-globin copy number variation.

[0004] With the development of high-throughput sequencing technology, more and more products are using high-throughput sequencing platforms to detect α-globin copy number variation. However, there is currently no effective analytical method to accurately type complex α-globin copy number variations, including triplets, in high-throughput sequencing data.

[0005] Existing detection methods include:

[0006] 1) gap-PCR or q-PCR detection;

[0007] 2) High-throughput sequencing platform based on regional sequencing depth analysis;

[0008] 3) High-throughput sequencing platform based on analysis of reads spanning breakpoints.

[0009] The defects of the above-mentioned prior art include:

[0010] Gap-PCR and qPCR have low throughput and limited detection types for α-globin copy number variation detection.

[0011] Existing high-throughput sequencing platform detection products do not have effective analytical means to accurately type complex types of α-globin copy number variations, including triplets. Summary of the Invention

[0012] According to the first aspect, in one embodiment, a method for detecting α-globin gene copy number variation is provided, comprising:

[0013] The primary detection step includes predicting whether the sample to be tested has an α-globin gene deletion based on the overall feature matrix of the sequencing data of the sample to be tested;

[0014] The secondary detection step includes analyzing and obtaining the α-globin gene copy number of the sample to be tested based on the prediction result of the primary prediction step;

[0015] The third-level detection step includes analyzing and obtaining the precise type of α-globin gene copy number variation of the sample to be tested based on the prediction result of the second-level detection step.

[0016] According to the second aspect, in one embodiment, a device for detecting α-globin gene copy number variation is provided, comprising:

[0017] The primary detection module is used to predict whether the sample to be tested has α-globin gene deletion based on the overall feature matrix of the sequencing data of the sample to be tested;

[0018] The secondary detection module is used to analyze and obtain the α-globin gene copy number of the sample to be tested based on the prediction result of the primary prediction step;

[0019] The third-level detection module includes analyzing and obtaining the precise type of the α-globin gene copy number variation of the sample to be tested based on the prediction result of the second-level detection module.

[0020] According to the third aspect, in one embodiment, there is provided an apparatus comprising:

[0021] Memory, used to store programs;

[0022] A processor is configured to implement the method according to the first aspect by executing the program stored in the memory.

[0023] According to the fourth aspect, in one embodiment, a computer-readable storage medium is provided, on which a program is stored. The program can be executed by a processor to implement the method according to the first aspect.

[0024] Based on the method and apparatus for detecting α-globin gene copy number variation described in the above embodiment, the present invention constructs a combined model for identifying α-globin copy number variation. Using high-throughput sequencing data, the model can quickly and accurately identify the precise type of α-globin copy number variation, including triplets. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a detection flow chart of Example 1. DETAILED DESCRIPTION

[0026] The present invention will be further described in detail below by means of specific embodiments in conjunction with the accompanying drawings. Similar elements in different embodiments are numbered with associated similar elements. In the following embodiments, many detailed descriptions are provided to enable the present application to be better understood. However, those skilled in the art will readily appreciate that some of the features may be omitted in different circumstances, or may be replaced by other elements, materials, or methods. In some cases, some operations related to the present application are not shown or described in the specification. This is to avoid the core portion of the present application being overwhelmed by excessive descriptions, and for those skilled in the art, it is not necessary to describe these related operations in detail. They will fully understand the related operations based on the description in the specification and the general technical knowledge in the art.

[0027] In addition, the features, operations, or characteristics described in the specification may be combined in any appropriate manner to form various embodiments. Furthermore, the steps or actions in the method description may be reordered or adjusted in a manner readily apparent to those skilled in the art. Therefore, the various sequences in the specification and drawings are provided solely for the purpose of clearly describing a particular embodiment and are not intended to be mandatory, unless otherwise specified.

[0028] The serial numbers assigned to the components in this document, such as "first", "second", etc., are only used to distinguish the objects described and do not have any order or technical meaning.

[0029] According to the first aspect, in one embodiment, a method for detecting α-globin gene copy number variation is provided, comprising:

[0030] The primary detection step includes predicting whether the sample to be tested has an α-globin gene deletion based on the overall feature matrix of the sequencing data of the sample to be tested;

[0031] The secondary detection step includes analyzing and obtaining the α-globin gene copy number of the sample to be tested based on the prediction result of the primary prediction step;

[0032] The third-level detection step includes analyzing and obtaining the precise type of α-globin gene copy number variation of the sample to be tested based on the predicted results of the second-level detection step.

[0033] In one embodiment, in the secondary detection step, if the prediction result of the primary prediction step indicates that the sample to be tested has an α-globin gene deletion, the first secondary model is used to predict that the number of α-globin gene copies in the sample to be tested is 1, 2 or 3; if the sample to be tested does not have an α-globin gene deletion, the second secondary model is used to predict that the number of α-globin gene copies in the sample to be tested is 4 or 5.

[0034] In one embodiment, in the three-level detection step, if the predicted value of the α-globin gene copy number in the test sample is 1, 2, or 3, the first, second, and third tertiary models are used, respectively, to predict the precise type of the α-globin gene copy number variation in the test sample. If the predicted value of the α-globin gene copy number in the test sample is 5, the fourth tertiary model is used to predict the precise type of the α-globin gene copy number variation in the test sample.

[0035] In one embodiment, in the secondary detection step, if the judgment result of the second secondary model is that the α-globin gene copy number of the test sample is 4, then the α-globin gene copy number variation type of the test sample is predicted to be αα / αα.

[0036] In one embodiment, the precise types of α-globin gene copy number variation predicted by the first three-level model include: 3.7 / -- SEA and -α 4.2 / -- SEA The precise types of α-globin gene copy number variations predicted by the second and third level models include: -α 3.7 / -α 3.7 、-α 4 . 2 / -α 4.2 、-α 3.7 / -α 4.2 and αα / -- SEA The precise types of α-globin gene copy number variation predicted by the third-level model include: αα / -α 3.7 and αα / -α 4.2 The precise types of α-globin gene copy number variations predicted by the fourth and third level models include: αα / ααα anti3.7 and αα / ααα anti4.2 .

[0037] In one embodiment, the sequencing data of the sample to be tested is calibrated data.

[0038] In one embodiment, the calibration method includes:

[0039] The standardization processing step includes standardizing all the data of the sample to be tested to obtain the labeled data;

[0040] The normalization step includes normalizing the standardized data so that each sample vector of the data is transformed into a vector with unit norm and the features are comparable.

[0041] In one embodiment, all data of the sample to be tested include the depth value of each region in the sample to be tested or the number of reads off the machine.

[0042] In one embodiment, the respective regions include an α-gene cluster homologous recombination region.

[0043] In one embodiment, the α-gene cluster homologous recombination region includes but is not limited to at least one of the following regions: X2 box, Y2 box, Z2 box, X1 box, Y1 box, and Z1 box.

[0044] For information on each region, refer to the following documents:

[0045] Higgs,DR,Hill,AV,Bowden,DK,Weatherall,DJ,&Clegg,JB(1984).Independent recombination events between the duplicated human alpha globingenes;implicati ons for their concerted evolution.Nucleic Acids Res,12(18),6965-6977.doi:10.1093 / nar / 12.18.6965.

[0046] Lam,KW,&Jeffreys,AJ(2007).Processes of de novo duplication ofhuman alph a-globin genes.Proc Natl Acad Sci USA,104(26),10950-10955.doi:10.1073 / pnas.0703856104.

[0047] Ou-Yang,H.,Hua,L.,Mo,QH,&Xu,XM(2004).Rapid,accurate genotypingofthe common-alpha(4.2)thalassaemia deletion based on the use of denaturingHPLC.J Cli n Pathol,57(2),159-163.doi:10.1136 / jcp.2003.011130.

[0048] In one embodiment, the genomic coordinates of the X2 box region are Chr16:170062-170720.

[0049] In one embodiment, the genomic coordinates of the Y2 box region are Chr16:171543-172216.

[0050] In one embodiment, the genomic coordinates of the Z2 box region are Chr16:173347-174029.

[0051] In one embodiment, the genomic coordinates of the X1 box region are Chr16:174318-174975.

[0052] In one embodiment, the genomic coordinates of the Y1 box region are Chr16:175100-175770.

[0053] In one embodiment, the genomic coordinates of the Z1 box region are Chr16:177158-177841.

[0054] In one embodiment, said genome comprises a reference genome.

[0055] In one embodiment, the reference genome includes but is not limited to hg38. If it is another reference genome, the above genome coordinates will change, but the region name will remain unchanged.

[0056] In one embodiment, in the standardization step, the standardization includes converting data of different magnitudes into the same magnitude and measuring them uniformly with the calculated Z-Score value, so that data of different magnitudes are comparable.

[0057] In one embodiment, in the normalization step, the Z-Score value is calculated according to the following formula:

[0058] Where x is the individual observation value, μ is the mean of the population data, and δ is the standard deviation of the population data.

[0059] In one embodiment, the normalization step includes dividing each component value in the vector by an L2 regularization factor.

[0060] In one embodiment, in the normalization step, each component value in the vector is divided by an L2 regularization factor according to the following formula:

[0061] Where x is each eigenvalue and n is the length of the vector, that is, the number of features.

[0062] In one embodiment, the methods for constructing the models used in the primary detection step, the secondary detection step, and the tertiary detection step include but are not limited to machine learning algorithms.

[0063] In one embodiment, the machine learning algorithm includes but is not limited to at least one of a gradient boosted decision tree (GBDT), a random forest (RF), and a support vector machine (SVM), preferably a gradient boosted decision tree.

[0064] In one embodiment, in the first-level detection step, the second-level detection step, and the third-level detection step, confidence thresholds for the sub-model results at each level are set, and samples that do not pass the threshold are judged as gray areas, and no results are output.

[0065] In one embodiment, in the primary detection step, the secondary detection step, and the tertiary detection step, the confidence interval of each model is ≥0.99.

[0066] In one embodiment, the sequencing data includes but is not limited to at least one of whole genome sequencing data, whole exome sequencing data, targeted capture sequencing data, and multiplex PCR sequencing data.

[0067] In one embodiment, the sequencing data includes but is not limited to any one of second-generation sequencing data and third-generation sequencing data.

[0068] According to the second aspect, in one embodiment, a device for detecting α-globin gene copy number variation is provided, comprising:

[0069] The primary detection module is used to predict whether the sample to be tested has α-globin gene deletion based on the overall feature matrix of the sequencing data of the sample to be tested;

[0070] The secondary detection module is used to analyze and obtain the α-globin gene copy number of the sample to be tested based on the prediction result of the primary detection module;

[0071] The third-level detection module includes analyzing the predicted results of the second-level detection module to obtain the precise type of α-globin gene copy number variation of the sample to be tested.

[0072] According to the third aspect, in one embodiment, there is provided an apparatus comprising:

[0073] Memory, used to store programs;

[0074] A processor is configured to implement the method according to the first aspect by executing the program stored in the memory.

[0075] According to the fourth aspect, in one embodiment, a computer-readable storage medium is provided, on which a program is stored. The program can be executed by a processor to implement the method according to the first aspect.

[0076] In one embodiment, the present invention provides a method for identifying α-globin copy number variation based on high-throughput sequencing data.

[0077] In one embodiment, the present invention selects a specific segment of the homologous recombination region of the α-gene cluster to obtain sequencing depth information; based on the relative quantitative theoretical hypothesis, a combined classification model is constructed to achieve accurate typing of the α-globin copy number variation of the sample to be tested.

[0078] In one embodiment, the present invention constructs a combined model for identifying α-globin copy number variation, which can quickly and accurately identify α-globin copy number variation types including triplets using high-throughput sequencing data.

[0079] In one embodiment, the present invention is integrated into a high-throughput sequencing detection product for thalassemia, which can be used for screening of α-thalassemia and auxiliary diagnosis of suspected patients.

[0080] Example 1

[0081] Figure 1 The following is a flow chart of the detection of this embodiment. The method of this embodiment mainly includes the following steps:

[0082] 1. Determination of characteristic regions of homologous recombination regions in the α-gene cluster

[0083] The specific regions of the homologous recombination region of the α-gene cluster used in this example were obtained from the NCBI database (https: / / www.ncbi.nlm.nih.gov / ). Appropriate regions were selected after testing, and their nomenclature and genomic coordinates are shown in Table 1.

[0084] Table 1: Specific regions of the homologous recombination region of the α-gene cluster

[0085] Segment Name Genome coordinates (hg38) X2 box Chr16:170062-170720 Y2 box Chr16:171543-172216 Z2 box Chr16:173347-174029 X1 box Chr16:174318-174975 Y1 box Chr16:175100-175770 Z1 box Chr16:177158-177841

[0086] The area used in this embodiment is included in the above-mentioned area.

[0087] 2. Acquisition of regional depth information

[0088] Using whole-genome sequencing data, whole-exome sequencing data with probes covering the region in Table 1, other chip capture sequencing data with probes covering the region in Table 1, or multiplex PCR library sequencing data with amplification regions including Table 1, sequence alignment and other steps were performed to obtain sample depth information for the target region. Each region depth value was used as a feature to input the following bioinformatics model, totaling six features, as detailed in Table 1.

[0089] 3. Data calibration

[0090] To eliminate data differences, all data are first standardized, as shown in Formula 1. Data of different magnitudes are uniformly converted to the same magnitude and measured uniformly using the calculated Z-Score value to ensure comparability between data.

[0091] Formula 1: Where x is the individual observation value, μ is the mean of the population data, and δ is the standard deviation of the population data;

[0092] Then normalize the data so that each sample (vector) of the data is transformed into a vector with unit norm and the features are comparable. Specifically, each component value in the vector is divided by the L2 regularization factor, as shown in Formula 2:

[0093] Formula 2: Where x is each eigenvalue and n is the length of the vector, that is, the number of features.

[0094] 4. Determine the model structure and judge the type of sample to be tested

[0095] A hierarchical judgment model was constructed based on the number of remaining α-globin genes in different α-globin copy number variation categories, as described in detail as follows (see Table 2):

[0096] The first-level model (Package_model) identifies whether the sample is missing α-globin and categorizes the sample into two groups: copy number 1, 2, 3, and copy number 4 or 5. Copy number 1, 2, 3, 4, and 5 refer to the number of α-globin gene copies in the sample, respectively.

[0097] The second-level models (model_123 and model_45) identify the α-globin copy number in the sample. Model_123 categorizes the sample into three groups: 1, 2, and 3 copies; model_45 categorizes the sample into two groups: 4 and 5 copies. Model_45 identifies samples with 4 copies as αα / αα.

[0098] The third level model (model_1, model_2, model_3, model_5) identifies the precise type of α-globin copy number variation in the sample to be tested. Model_1 identifies -α3.7 / -- SEA 、-α4.2 / -- SEA Two types; model_2 recognizes -α3.7 / -α3.7, -α4.2 / -α4.2, -α3.7 / -α4.2, αα / -- SEA Four types; model_3 recognizes two types: αα / -α3.7 and αα / -α4.2; model_5 recognizes two types: αα / αααanti3.7 and αα / αααanti4.2.

[0099] Table 2 Model structure and input and output

[0100]

[0101]

[0102] Each model was constructed using the gradient boosted decision tree (GBDT) method. A large number of samples with known α-globin copy number variation were trained to establish the model and set confidence thresholds for the results of each sub-model.

[0103] When determining the type of sample to be tested, samples that do not pass the threshold are judged as gray areas and no results are output.

[0104] Outlier samples within a particular category may be classified as gray zone and no results are output. The specific reason for this is related to the algorithm and the modeling data. Some samples that share common features but fall outside the 11 output categories are classified as gray zone. This judgment step is intended to improve accuracy.

[0105] This example uses whole-exome sequencing data from 496 individuals with known α-globin copy number variation types. The phenotype information is shown in Table 3. The average depth information for the relevant regions in Table 1 was extracted and calibrated to construct an overall feature matrix. This overall feature matrix was input into the combined model, with a confidence interval of ≥0.99 set for each model. The resulting number of accurate variants and gray areas for each type is shown in Table 4.

[0106] Table 3 Composition of the training set

[0107] Type quantity αα / αα 153 <![CDATA[αα / -α 3.7 ]]> 65 <![CDATA[αα / -α 4.2 ]]> 54 <![CDATA[αα / -- SEA ]]> 32 <![CDATA[-α 3.7 / -a 3.7 ]]> 28 <![CDATA[-α 4.2 / -a 4.2 ]]> 25 <![CDATA[-α 3.7 / -a 4.2 ]]> 24 -α 3.7 / -- SEA ]]> 34 <![CDATA[-α 4.2 / -- SEA ]]> 18 <![CDATA[αα / ααα anti3.7 ]]> 39 <![CDATA[αα / ααα anti4.2 ]]> 24 total 496

[0108] Table 4 Number of accurate judgments and gray areas of each type of model

[0109]

[0110]

[0111] Gray area samples indicate detection failures, no results are output, and they are not included in the accuracy statistics.

[0112] As can be seen from Table 4, the prediction accuracy of this embodiment for each type is as high as 95.24% to 100%, and a total of 99.58%.

[0113] In one embodiment, the present invention accurately determines the α-globin copy number by constructing a bioinformatics model based on deep relative quantitative analysis of specific regions of the homologous recombination region of the α-gene cluster.

[0114] In one embodiment, the present invention constructs a bioinformatics model by performing relative quantitative analysis of the depth of specific regions of the homologous recombination region of the α-globin gene cluster in high-throughput sequencing data, thereby achieving accurate typing of α-globin copy number variations including triplets.

[0115] In one embodiment, the present invention can be widely used in high-throughput sequencing data bioinformatics analysis processes to accurately detect α-globin copy number variations, including triplets, and can be used for screening of α-thalassemia and auxiliary diagnosis of suspected patients.

[0116] In one embodiment, other characteristic segments of the homologous recombination region of the α-gene cluster can be selected.

[0117] In one embodiment, a large amount of high-throughput sequencing data of known α-thalassemia copy number variation types can be used to expand the reference set, or more critical features can be introduced to construct a more complex model.

[0118] In one embodiment, the present invention accurately determines the α-globin copy number by constructing a bioinformatics model based on deep relative quantitative analysis of specific regions of the homologous recombination region of the α-gene cluster.

[0119] In one embodiment, the present invention constructs a bioinformatics model by performing relative quantitative analysis of the depth of specific regions of the homologous recombination region of the α-globin gene cluster in high-throughput sequencing data, thereby achieving accurate typing of α-globin copy number variations including triplets.

[0120] In one embodiment, the present invention can be widely used in high-throughput sequencing data bioinformatics analysis processes to accurately detect α-globin copy number variations, including triplets, and can be used for screening of α-thalassemia and auxiliary diagnosis of suspected patients.

[0121] In one embodiment, other characteristic segments of the homologous recombination region of the α-gene cluster can be selected.

[0122] In one embodiment, a large amount of high-throughput sequencing data of known α-thalassemia copy number variation types can be used to expand the reference set, or more critical features can be introduced to construct a more complex model.

[0123] Those skilled in the art can understand that all or part of the functions of various methods in the above embodiments can be realized by hardware or by a computer program. When all or part of the functions in the above embodiments are realized by a computer program, the program can be stored in a computer readable storage medium, which can include a read-only memory, a random access memory, a magnetic disk, an optical disk, a hard disk, and the like. The above functions are realized by executing the program by a computer. For example, the program is stored in a memory of a device, and the above functions are realized by executing the program in the memory by a processor. In addition, when all or part of the functions in the above embodiments are realized by a computer program, the program can also be stored in a storage medium such as a server, another computer, a disk, an optical disk, a flash disk, or a mobile hard disk, and is saved in a memory of a local device by downloading or copying, or the system of the local device is updated, and the above functions are realized by executing the program in the memory by a processor.

[0124] The above application of specific examples to the present application is described, which is only used to help understand the present application and does not limit the present application. For those skilled in the art, according to the idea of the present application, a number of simple deductions, deformations or substitutions can be made.

Claims

1. A method for detecting α-globin gene copy number variation, characterized in that: include: The primary detection step includes predicting whether the sample to be tested has an α-globin gene deletion based on the overall feature matrix of the sequencing data of the sample to be tested; The secondary detection step includes analyzing and obtaining the α-globin gene copy number of the sample to be tested based on the predicted result of the primary detection step; The third-level detection step includes analyzing and obtaining the precise type of α-globin gene copy number variation of the sample to be tested based on the prediction results of the second-level detection step; In the secondary detection step, if the prediction result of the primary detection step indicates that the sample to be tested has an α-globin gene deletion, the first secondary model is used to predict that the number of α-globin gene copies in the sample to be tested is 1, 2, or 3; if the sample to be tested does not have an α-globin gene deletion, the second secondary model is used to predict that the number of α-globin gene copies in the sample to be tested is 4 or 5; In the three-level detection step, if the predicted value of the α-globin gene copy number in the sample to be tested is 1, 2 or 3, the first, second and third level models are used respectively to predict the precise type of the α-globin gene copy number variation in the sample to be tested; if the predicted value of the α-globin gene copy number in the sample to be tested is 5, the fourth level model is used to predict the precise type of the α-globin gene copy number variation in the sample to be tested; In the secondary detection step, if the judgment result of the second secondary model is that the copy number of the α-globin gene of the sample to be tested is 4, then the copy number variation type of the α-globin gene of the sample to be tested is predicted to be αα / αα; The precise types of α-globin gene copy number variations predicted by the first and third level models include: 3.7 / -- SEA and -α 4.2 / -- SEA The precise types of α-globin gene copy number variations predicted by the second and third level models include: -α 3.7 / -α 3.7 、-α 4.2 / -α 4.2 、-α 3.7 / -α 4.2 and αα / -- SEA The precise types of α-globin gene copy number variation predicted by the third-level model include: αα / -α 3.7 and αα / -α 4.2 The precise types of α-globin gene copy number variations predicted by the fourth and third level models include: αα / ααα anti3.7 and αα / ααα anti4.2 .

2. The method according to claim 1, wherein The sequencing data of the sample to be tested is calibrated data.

3. The method according to claim 2, wherein Calibration methods include: The standardization processing step includes standardizing all the data of the sample to be tested to obtain the labeled data; The normalization step includes normalizing the standardized data so that each sample vector of the data is transformed into a vector of unit norm and each feature is comparable.

4. The method according to claim 3, wherein All data of the sample to be tested include the depth value of each area in the sample to be tested or the number of reads off the machine; The various regions include the α-gene cluster homologous recombination region.

5. The method according to claim 4, wherein The α-gene cluster homologous recombination region includes at least one of the following regions: X2 box, Y2 box, Z2 box, X1 box, Y1 box, and Z1 box.

6. The method according to claim 5, wherein The genomic coordinates of the X2 box region are Chr16:170062-170720; The genomic coordinates of the Y2 box region are Chr16:171543-172216; The genomic coordinates of the Z2 box region are Chr16:173347-174029; The genomic coordinates of the X1 box region are Chr16:174318-174975; The genomic coordinates of the Y1 box region are Chr16:175100-175770; The genomic coordinates of the Z1 box region are Chr16:177158-177841.

7. The method according to claim 6, wherein The genome comprises a reference genome.

8. The method according to claim 7, wherein The reference genome includes hg38.

9. The method according to claim 3, wherein In the standardization process, the standardization process includes converting data of different magnitudes into the same magnitude and measuring them uniformly with the calculated Z-Score value.

10. The method according to claim 9, wherein In the standardization step, the Z-Score value is calculated according to the following formula: , where x is the individual observation value, μ is the mean of the population data, and δ is the standard deviation of the population data.

11. The method according to claim 3, wherein The normalization step involves dividing each component value in the vector by the L2 regularization factor.

12. The method according to claim 11, wherein In the normalization step, each component value in the vector is divided by the L2 regularization factor according to the following formula: , where x is each eigenvalue and n is the length of the vector, that is, the number of features.

13. The method according to any one of claims 1 to 12, wherein: Methods for constructing models used in the primary detection step, the secondary detection step, and the tertiary detection step include machine learning algorithms.

14. The method according to claim 13, wherein The machine learning algorithm includes at least one of a gradient boosting decision tree, a random forest, and a support vector machine.

15. The method according to claim 13, wherein In the first-level detection step, the second-level detection step, and the third-level detection step, the confidence threshold of the sub-model results at each level is set, and the samples that do not pass the threshold are judged as gray areas and no results are output.

16. The method according to claim 15, wherein In the first-level detection step, the second-level detection step, and the third-level detection step, the confidence interval of each model is ≥0.

99.

17. The method according to claim 1, wherein The sequencing data includes at least one of whole genome sequencing data, whole exome sequencing data, targeted capture sequencing data and multiplex PCR sequencing data.

18. The method according to claim 17, wherein The sequencing data includes any one of second-generation sequencing data and third-generation sequencing data.

19. A device for detecting α-globin gene copy number variation, characterized in that: include: The primary detection module is used to predict whether the sample to be tested has α-globin gene deletion based on the overall feature matrix of the sequencing data of the sample to be tested; The secondary detection module is used to analyze and obtain the α-globin gene copy number of the sample to be tested based on the prediction result of the primary detection module; The third-level detection module is used to analyze and obtain the precise type of α-globin gene copy number variation of the sample to be tested based on the prediction result of the second-level detection module; In the secondary detection module, if the prediction result of the primary detection module indicates that the sample to be tested has an α-globin gene deletion, the first secondary model is used to predict that the α-globin gene copy number in the sample to be tested is 1, 2, or 3; if the sample to be tested does not have an α-globin gene deletion, the second secondary model is used to predict that the α-globin gene copy number in the sample to be tested is 4 or 5; In the three-level detection module, if the predicted value of the α-globin gene copy number in the sample to be tested is 1, 2 or 3, the first, second and third level models are used respectively to predict the precise type of the α-globin gene copy number variation in the sample to be tested; if the predicted value of the α-globin gene copy number in the sample to be tested is 5, the fourth level model is used to predict the precise type of the α-globin gene copy number variation in the sample to be tested; In the secondary detection module, if the judgment result of the second secondary model is that the copy number of the α-globin gene of the sample to be tested is 4, then the copy number variation type of the α-globin gene of the sample to be tested is predicted to be αα / αα; The precise types of α-globin gene copy number variations predicted by the first and third level models include: 3.7 / -- SEA and -α 4.2 / -- SEA The precise types of α-globin gene copy number variations predicted by the second and third level models include: -α 3.7 / -α 3.7 、-α 4.2 / -α 4.2 、-α 3.7 / -α 4.2 and αα / -- SEA The precise types of α-globin gene copy number variation predicted by the third-level model include: αα / -α 3.7 and αα / -α 4.2 The precise types of α-globin gene copy number variations predicted by the fourth and third level models include: αα / ααα anti3.7 and αα / ααα anti4.2 .

20. A device for detecting α-globin gene copy number variation, characterized in that: include: Memory, used to store programs; A processor, configured to implement the method according to any one of claims 1 to 18 by executing the program stored in the memory.

21. A computer-readable storage medium, wherein a program is stored on the medium, and the program can be executed by a processor to implement the method according to any one of claims 1 to 18.

Citation Information

Patent Citations

  • Detection kit for capturing gene copy number related to alpha thalassemia

    CN109584957A

  • Method and device for detecting thalassemia gene variation

    CN111326211A

  • Normalization methods for measuring gene copy number and expression

    US20180046754A1