Method and system for detecting CYP2D6 gene polymorphic typing based on third-generation sequencing

By using PacBio HiFi sequencing technology and third-generation sequencing method, the CYP2D6 gene was detected and typified, which solved the problem of inaccurate clinical phenotype prediction caused by haplotype allocation errors in the prior art, and achieved higher detection accuracy and more reliable support for pharmacogenomics research.

CN120015113APending Publication Date: 2025-05-16WUHAN FRASERGEN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411994086.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-16

Smart Images

  • Figure CN120015113A_ABST
    Figure CN120015113A_ABST
Patent Text Reader

Abstract

The invention provides a method and system for detecting CYP2D6 gene polymorphic typing based on third-generation sequencing, and the method comprises the following steps: carrying out data preprocessing on offline data to obtain HiFi reads; the HiFi reads is compared to the GRCh38 reference genome, and a bam file is obtained; each sequence in the compared bam file is filtered, reads and bam format files which stretch across the upstream of the CYP2D6 gene and the downstream of the CYP2D7 gene and meet a preset coverage depth threshold value are reserved, and a haplotype conclusion is obtained; dividing the alleles into haplotype alleles and distribution star alleles, and performing CYP2D6 gene diallele typing to obtain diplotype gene information; and performing activity scoring to obtain phenotype prediction and inference. According to the method, the information of the CYP2D6 gene can be more accurately obtained, the problem of inaccurate clinical phenotype prediction caused by haplotype typing errors is effectively solved, a reliable basis is provided for pharmacogenomics research, clinical diagnosis and individualized medical treatment, and the development of precise medical treatment is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biological detection technology, and in particular to a method and system for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing. Background Art

[0002] The CYP2D6 gene has been extensively studied in the field of pharmacogenomics (PGx) because it directly affects the metabolism of approximately 20% of prescription drugs, covering a variety of important drug classes such as antidepressants, cardiovascular drugs, and anticancer drugs. However, the study of CYP2D6 faces great challenges, which is mainly attributed to its complex genetic structure. The gene has complex polymorphisms and structural variations (SVs), including complete gene deletions and duplications, tandem alleles, and heterozygous (fusion) events of the CYP2D7 gene.

[0003] In addition, since this region is highly homologous to the upstream adjacent genes CYP2D7 and CYP2D8, and its exon sequence is 97% and 92% similar to CYP2D6, respectively, this high sequence similarity makes it difficult for traditional short-read sequencing and conventional genotyping methods to detect this region, making it difficult to accurately capture the true information of the CYP2D6 gene, greatly limiting its in-depth research and accurate interpretation.

[0004] In the application scenario of pharmacogenomics, the haplotype of drug genetic loci such as CYP2D6 is the core basis for the assignment of star (*) alleles, which is then used to predict the individual's metabolic status of drugs. However, errors and ambiguities often occur in the current haplotype assignment process, which is mainly due to inaccurate or missing variant calls and deviations in the phasing process. These problems inevitably lead to inaccurate predictions of clinical phenotypes. Such inaccurate predictions can seriously interfere with the selection of the most appropriate drugs or doses for patients by healthcare providers based on prescription guidelines or clinical decision support systems, thereby affecting the treatment effect and may cause potential medical risks. Although nearly 200 haplotypes or star (*) alleles of CYP2D6 have been defined in PharmVar, existing detection platforms have limited capabilities in identifying new or rare haplotypes and other structural variations (SVs), making it difficult to meet the evolving needs of precision medicine.

[0005] Based on this, the present invention proposes an improved CYP2D6 gene detection method. Summary of the invention

[0006] Based on the above statements, the present invention provides a method for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing to solve the problem of inaccurate clinical phenotype prediction caused by haplotype assignment errors in the prior art.

[0007] The technical solution of the present invention to solve the above technical problems is as follows:

[0008] In a first aspect, the present invention provides a method for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing, comprising the following steps:

[0009] S1: preprocess the offline data to obtain HiFi reads;

[0010] S2: Align the HiFi reads to the GRCh38 reference genome to obtain a bam file;

[0011] S3: Filter each sequence in the aligned bam file, retain the reads and bam format files that span the upstream of the CYP2D6 gene and the downstream of the CYP2D7 gene and meet the preset coverage depth threshold, and draw a conclusion on the haplotype of the sample;

[0012] S4: Divide alleles into haplotypes and assign star alleles to perform CYP2D6 gene di-allelic typing to obtain diploid gene information;

[0013] S5: Performing activity scoring on the diploid gene information to obtain phenotypic prediction and inference.

[0014] Based on the above technical solution, the present invention can also be improved as follows.

[0015] Further, the off-board data is preprocessed, specifically including:

[0016] The raw sequencing data produced by the PacBio SMRT platform were processed to obtain HiFi reads.

[0017] Furthermore, in step S2, the analysis tool for data comparison is minimap.

[0018] Furthermore, step S3 specifically includes:

[0019] S301, dividing the bam file into moving windows of fixed size, calculating the number of reads in each window, and using the value obtained after Z-score standardization as the copy number as a supplementary basis for detecting structural variation;

[0020] S302, performing structural variation signal detection on reads, recording read ID, read soft shear direction, read start coordinate, read end coordinate, alignment chain information, and whether there is structural variation information;

[0021] S303, classifying the reads of the mutated SV samples and structural variations into different types, and adopting different analysis strategies to perform analysis accordingly.

[0022] Further, step S303 specifically includes: performing the following processing on the reads with structural variation information:

[0023] The copy number information obtained in step S301 is compared and visualized; if the copy number of the region covering the CYP2D6 gene is less than the copy number of the surrounding region, it indicates that the CYP2D6 gene is deleted; otherwise, it indicates that the CYP2D6 gene is amplified;

[0024] The reads obtained in step S302 are iteratively parsed to obtain read information; according to the cigar value information, if soft-cut bases exist and these soft-cut bases are aligned to the CYP2D7 gene, it indicates that heterozygosity with the CYP2D7 gene occurs, and the breakpoint information is recorded.

[0025] Furthermore, the following processing is performed on reads without structural variation information:

[0026] For the mutants without structural variation in step S301, extract the variation at the key difference position and mark the information source;

[0027] According to the site information of PharmVar database typing, the mutation status of specific sites is recorded;

[0028] Determine the loci of the CYP2D6 genotype, mark the loci, record the loci base information, and determine the same haplotype, otherwise, they are different haplotypes.

[0029] Furthermore, step S5 specifically includes:

[0030] The phenotypes were assigned using an activity scoring system, where each allele was assigned an “activity value” ranging from 0 to 1;

[0031] If the copy number is known, the activity value of the allele is multiplied by the gene copy number; if the copy number is unknown, the sample is assigned by multiplying by N by default, where N is a positive integer greater than or equal to 2; CYP2D6 AS is defined as the sum of the activity values ​​assigned to each allele;

[0032] CYP2D6 AS was converted into phenotype using a classification system.

[0033] Further, in converting CYP2D6 AS to phenotype using a classification system:

[0034] If the AS score is 0, the individual is a poor metabolizer;

[0035] If the AS is 0.5 points, the individual is an intermediate metabolizer;

[0036] If the AS is 1.0-2.0 points, the individual is a normal metabolizer;

[0037] If the AS is above 2.0 points, the individual is an ultra-rapid metabolizer.

[0038] Furthermore, the sequencing tool used in steps S1 to S4 is the PacBio HiFi sequencing platform.

[0039] In a second aspect, the present invention also provides a system for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing, comprising:

[0040] The preprocessing module is used to preprocess the offline data to obtain HiFi reads;

[0041] A data comparison module is used to compare the HiFi reads to the GRCh38 reference genome, perform data comparison, and obtain a bam file;

[0042] The filtering module is used to filter each sequence in the aligned bam file, retain the reads and bam format files that span the upstream of the CYP2D6 gene and the downstream of the CYP2D7 gene and meet the preset coverage depth threshold, and draw a conclusion on the haplotype of the sample;

[0043] Genotyping module, used to classify alleles into haplotypes and assign star alleles to perform CYP2D6 gene di-allelic typing and obtain diploid gene information;

[0044] The activity scoring module is used to perform activity scoring on the diploid gene information to obtain phenotype prediction and inference.

[0045] Compared with the prior art, the technical solution of this application has the following beneficial technical effects:

[0046] The method for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing provided by the present invention has the following beneficial effects compared with the existing methods:

[0047] The present invention adopts PacBio HiFi sequencing technology, which can effectively capture CYP2D6 variations on a single platform using high-precision long reads; compared with the existing second-generation short-read sequencing technology, PacBio HiFi sequencing technology has significant advantages, and its long reads can successfully cross complex gene regions, effectively reducing the alignment errors caused by short reads during the second-generation sequencing process, thereby greatly improving the accuracy of detection. Compared with other third-generation ONT sequencing platforms, the technical characteristics of PacBioHiFi sequencing technology make it possible to directly and completely decompose and phase complex genes when detecting the CYP2D6 gene without the need for complex assembly or inference processes, thereby achieving haplotype allocation that is unrelated to ancestry. This feature not only helps to improve the accurate determination of known haplotypes, but is also more conducive to the discovery of new or rare haplotypes and other structural variations (SVs), providing strong technical support for in-depth research on the function and drug metabolism mechanism of the CYP2D6 gene.

[0048] Compared with the prior art, the method for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing proposed in the present application can process and generate continuous sequence reads from template molecules of multiple kilobases in length without the need for a DNA fragmentation step; starting from structural variations and small mutations, the typing results are made more accurate, while providing a common phenotypic prediction function; that is, it can more accurately obtain information on the CYP2D6 gene, effectively solving the problem of inaccurate clinical phenotype prediction caused by haplotype assignment errors in the prior art, and providing a more reliable basis for pharmacogenomic research, clinical diagnosis, and personalized medicine, and promoting the development of the field of precision medicine. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0050] Figure 1 A schematic diagram of a process for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing provided in an embodiment of the present invention;

[0051] Figure 2 A schematic diagram of the CYP2D6 gene provided in an embodiment of the present invention;

[0052] Figure 3 A visual schematic diagram of the phenotypic prediction process in the example provided in the embodiment of the present invention;

[0053] Figure 4A haploid typing diagram at the read level provided in an embodiment of the present invention;

[0054] Figure 5 A schematic diagram of the copy number of heterozygous CYP2D6 provided in an embodiment of the present invention;

[0055] Figure 6 A schematic diagram of a system for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0057] The specific implementation of the present invention is described in detail below. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the protection scope of the present invention.

[0058] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0059] In the present invention, the conventional reagents and raw materials used in the experiments are all commercially available.

[0060] Example 1

[0061] Combined with Figure 1 The method for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing provided by the present invention comprises the following steps:

[0062] It should be noted that the sequencing tool used in this embodiment is the PacBio HiFi sequencing platform.

[0063] Step S1: preprocess the offline data to obtain HiFi reads.

[0064] Specifically, the original sequencing data subreads produced by the PacBio SMRT platform are processed to obtain HiFireads.

[0065] The files in this embodiment specifically include: HiFi reads obtained by PacBio sequencing, GRCh38 reference genome sequence files, and CYP2D6 gene polymorphism files downloaded from the PharmVar website.

[0066] Currently, more than 175 alleles have been found in CYP2D6; based on their function, these alleles are divided into four categories: normal functional alleles (*1, *2, *35, etc.), reduced functional alleles (*9, *10, *17, *29, *41, etc.), non-functional alleles (*3, *4, *5, *6, *8, *11, etc.), non-functional alleles (*1×2, *2×2) and alleles of unknown function (*22, *28, *30, etc.).

[0067] According to the mutation type, it can be divided into single nucleotide mutation (*2), deletion (*5), multiple copy number variation (*1×2, *2×2), and repeated sequences that achieve many unbalanced exchange events (*13, *36, *61, etc.) through non-allelic homologous recombination between CYP2D6 and CYP2D7. These mutations will lead to stable and heritable amplification, deletion and heterozygosity (fusion) formation. Figure 2 Shown is the genetic information of CYP2D6.

[0068] This application uses PacBio HIFI reads in BAM format to compare with the reference genome to obtain the aligned bam file, and uses the read variation information to subsequently use for star allocation and activity score calculation. The specific process is combined with Figure 2 And steps S2-S5 are further described in detail:

[0069] Step S2: Align the HiFi reads to the GRCh38 reference genome to obtain a bam file.

[0070] In this step, the analysis tool for data comparison is minimap.

[0071] The CYP2D6*1 allele is considered the reference allele and is often inferred when the variant being interrogated by the genotyping test is not detected, thus serving as an adjunct to variant detection.

[0072] Step S3: Based on the long length of PacBio HiFi reads and the characteristics of the CYP2D6 gene, each sequence in the aligned bam file was filtered, and reads and bam format files that spanned the upstream of the CYP2D6 gene and the downstream of the CYP2D7 gene and met the preset coverage depth threshold were retained to draw conclusions about the haplotype of the sample.

[0073] Step S3 specifically includes:

[0074] Step S301, divide the bam file into fixed-size moving windows, calculate the number of reads in each window, and use the value obtained after Z-score normalization as the copy number, which is used as a supplementary basis for detecting structural variations. Samples with SVs show different depth signals, and mutated SV samples are detected.

[0075] Step S302: perform structural variation signal detection on reads, record read ID, read soft clip direction (softclip end), read start coordinate (rstart), read stop coordinate (rstop), alignment chain (strand) information and whether there is structural variation (no SV / SV) information.

[0076] Step S303: Classify the reads of the SV samples with variations and the structural variations into different types, and adopt different analysis strategies to analyze them accordingly.

[0077] Wherein, step S303 specifically includes: performing the following processing on the reads with structural variation information:

[0078] The copy number information obtained in step S301 is compared and visualized; if the copy number of the region covering the CYP2D6 gene is less than the copy number of the surrounding region, it indicates that the CYP2D6 gene is deleted; otherwise, it indicates that the CYP2D6 gene is amplified;

[0079] The reads obtained in step S302 are iteratively parsed to obtain read information; according to the cigar value information, if soft-cut bases exist and these soft-cut bases are aligned to the CYP2D7 gene, it indicates that heterozygosity with the CYP2D7 gene occurs, and the breakpoint information is recorded.

[0080] Specifically, for the deletion / amplification type: compare the copy number information obtained in the comparison and visualize it. If the copy number of the region covering the CYP2D6 gene is less than the surrounding copy number, it indicates that the CYP2D6 gene is deleted, otherwise it indicates that the CYP2D6 gene is amplified, such as Figure 5 shown.

[0081] For heterozygous (fusion) type: for the obtained reads, the pysam module is used to further iteratively parse the read information. According to the cigar value information, if there are soft-cut bases and these soft-cut bases are aligned to the CYP2D7 gene, it indicates that heterozygosity has occurred with the CYP2D7 gene, such as Figure 5 As shown, the corresponding breakpoint information is recorded.

[0082] It should be noted that for reads without structural variation information, the following processing is performed:

[0083] For the mutants without structural variation in step S301, extract the variation at the key difference position and mark the information source;

[0084] According to the site information of PharmVar database typing, the mutation status of specific sites is recorded;

[0085] Determine the loci of the CYP2D6 genotype, mark the loci, record the loci base information, and determine the same haplotype, otherwise, different haplotypes, such as Figure 4 shown.

[0086] Specifically, for mutants without structural variation, the variation at the key difference position is extracted and the information source is marked; according to the site information of the PharmVar database, the mutation of the specific site is recorded; it is assumed that the sites L1, L2, L3...L that determine the CYP2D6 genotype m , a total of m marker sites, each recording the base information B1, B2, B3...B m ; For the nth read, the base at position m is recorded as B n,m If for read j and read k, traverse the sites so that B j,m and B k,m If they are equal, read j and read k are merged to determine the same haplotype, otherwise they are different haplotypes.

[0087] Combining the above steps, the haplotype conclusion of the sample is obtained.

[0088] Step S4: Divide the alleles into haplotypes and assign star alleles to perform CYP2D6 gene di-allelic typing to obtain diploid gene information.

[0089] Step S5: performing activity scoring on the diploid gene information to obtain phenotype prediction and inference.

[0090] The above steps S4 and S5 specifically include:

[0091] The phenotypes were assigned using an activity scoring system, where each allele was assigned an “activity value” ranging from 0 to 1;

[0092] If the copy number is known, the activity value of the allele is multiplied by the gene copy number; if the copy number is unknown, the sample is assigned by multiplying by N by default, where N is a positive integer greater than or equal to 2; CYP2D6 AS is defined as the sum of the activity values ​​assigned to each allele;

[0093] CYP2D6 AS was converted into phenotype using a classification system.

[0094] Specifically, in converting CYP2D6AS into a phenotype using a classification system:

[0095] If the AS score is 0, the individual is a poor metabolizer;

[0096] If the AS is 0.5 points, the individual is an intermediate metabolizer;

[0097] If the AS is 1.0-2.0 points, the individual is a normal metabolizer;

[0098] If the AS is above 2.0 points, the individual is an ultra-rapid metabolizer.

[0099] Specifically, the system used to convert genotype to phenotype relies on star (*) allele nomenclature (defining which variants are present in the allele), as well as assigning function to the star allele (gain, normal, minus, or no function) and inferring the phenotype based on the identified genotype.

[0100] Phenotypes were assigned using an activity score (AS) system, in which each allele was assigned an “activity value” ranging from 0–1 (e.g., 0 for no function, 0.5 for reduced function, and 1.0 for normal function). In addition, given that CYP2D6 alleles can also have variable copy numbers, if the copy number was known, the activity value of the allele was multiplied by the gene copy number (i.e., ×2, ×3, etc.); if the copy number was unknown, samples may default to a ×2 assignment or be displayed as xN.

[0101] Therefore, CYP2D6 AS is the sum of the activity values ​​assigned to each allele. CYP2D6 AS was converted to phenotype using the following classification system: individuals with an AS score of 0 were poor metabolizers (PMs), individuals with an AS score of 0.5 were intermediate metabolizers (IMs), individuals with an AS score of 1.0, 1.5, and 2.0 were normal metabolizers (NMs), and individuals with an AS score greater than 2 were ultra-rapid metabolizers (UMs).

[0102] Let's take a specific example to explain:

[0103] like Figure 3 Shown is a visualization of the process for predicting phenotypes, where genetic variation in the CYP6D2 gene on the maternal (orange) and paternal (green) alleles is assigned as a haplotype, which combines to form a diplotype (yellow). This diplotype is translated into a predicted phenotype, on which subsequent treatment is based.

[0104] Example 2

[0105] Based on Example 1, this example provides a corresponding system for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing, such as Figure 6 As shown, including:

[0106] The preprocessing module is used to preprocess the offline data to obtain HiFi reads;

[0107] A data comparison module is used to compare the HiFi reads to the GRCh38 reference genome, perform data comparison, and obtain a bam file;

[0108] The filtering module is used to filter each sequence in the aligned bam file, retain the reads and bam format files that span the upstream of the CYP2D6 gene and the downstream of the CYP2D7 gene and meet the preset coverage depth threshold, and draw a conclusion on the haplotype of the sample;

[0109] Genotyping module, used to classify alleles into haplotypes and assign star alleles to perform CYP2D6 gene di-allelic typing and obtain diploid gene information;

[0110] The activity scoring module is used to perform activity scoring on the diploid gene information to obtain phenotype prediction and inference.

[0111] In summary, the method and system for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing provided in the embodiments of the present invention, wherein the PacBio HiFi sequencing technology adopted can effectively capture CYP2D6 variation on a single platform using high-precision long reads.

[0112] Compared with the existing second-generation short-read sequencing technology, PacBio HiFi sequencing technology has significant advantages. Its long reads can successfully cross complex gene regions, effectively reducing the alignment errors caused by short reads during the second-generation sequencing process, thereby greatly improving the accuracy of detection. Compared with other third-generation ONT sequencing platforms, the technical characteristics of PacBio HiFi sequencing technology make it possible to directly and completely decompose and type complex genes when testing the CYP2D6 gene without the need for complex assembly or inference processes, thereby achieving haplotype assignment that is not related to ancestry. This feature not only helps to improve the accuracy of known haplotypes, but also helps to discover new or rare haplotypes and other structural variations, providing strong technical support for in-depth research on the function of the CYP2D6 gene and the mechanism of drug metabolism.

[0113] Compared with the prior art, the method for detecting CYP2D6 gene polymorphism typing based on PacBio HiFi sequencing technology proposed in this application can process and generate continuous sequence reads from template molecules of multiple kilobases in length without the need for a DNA fragmentation step; starting from structural variations and small mutations, the typing results are more accurate, while providing a common phenotypic prediction function; that is, it can more accurately obtain information on the CYP2D6 gene, effectively solving the problem of inaccurate clinical phenotype prediction caused by haplotype assignment errors in the prior art, and providing a more reliable basis for pharmacogenomic research, clinical diagnosis and personalized medicine, and promoting the development of the field of precision medicine.

[0114] In the description of this specification, the description with reference to the terms "specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiment of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing, characterized in that: The steps include: S1: preprocess the offline data to obtain HiFi reads; S2: Compare the HiFi reads to the GRCh38 reference genome to obtain a bam file; S3: Filter each sequence in the compared bam file, retain the reads and bam format files that span the upstream of the CYP2D6 gene and the downstream of the CYP2D7 gene and meet the preset coverage depth threshold, and draw a conclusion on the haplotype of the sample; S4: Divide alleles into haplotypes and assign star alleles to perform CYP2D6 gene di-allelic typing to obtain diploid gene information; S5: Performing activity scoring on the diploid gene information to obtain phenotypic prediction and inference.

2. The method for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing according to claim 1, characterized in that: Preprocessing the off-board data specifically includes: The raw sequencing data produced by the PacBio SMRT platform were processed to obtain HiFi reads.

3. The method for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing according to claim 1, characterized in that: In step S2, the analysis tool for data comparison is minimap.

4. The method for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing according to claim 1, characterized in that: Step S3 specifically includes: S301, dividing the bam file into moving windows of fixed size, calculating the number of reads in each window, and using the value obtained after Z-score standardization as the copy number as a supplementary basis for detecting structural variation; S302, performing structural variation signal detection on reads, recording information such as read ID, read soft shear direction, read start coordinate, read end coordinate, alignment chain information, and whether there is structural variation; S303, classifying the reads of the mutated SV samples and structural variations into different types, and adopting different analysis strategies to perform analysis accordingly.

5. The method for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing according to claim 4, characterized in that: Step S303 specifically includes: performing the following processing on the reads with structural variation information: The copy number information obtained in step S301 is compared and visualized; if the copy number of the region covering the CYP2D6 gene is less than the copy number of the surrounding region, it indicates that the CYP2D6 gene is deleted; otherwise, it indicates that the CYP2D6 gene is amplified; The reads obtained in step S302 are iteratively parsed to obtain read information; according to the cigar value information, if soft-cut bases exist and these soft-cut bases are aligned to the CYP2D7 gene, it indicates that heterozygosity with the CYP2D7 gene occurs, and the breakpoint information is recorded.

6. The method for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing according to claim 5, characterized in that: For reads without structural variation information, perform the following processing: For the mutants without structural variation in step S301, extract the variation at the key difference position and mark the information source; According to the site information of PharmVar database typing, the mutation status of specific sites is recorded; Determine the loci of the CYP2D6 genotype, mark the loci, record the loci base information, and determine the same haplotype, otherwise, they are different haplotypes.

7. The method for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing according to claim 2, characterized in that: Step S5 specifically includes: The phenotype is assigned using an activity scoring system, where each allele is assigned an "activity value" ranging from 0 to 1; If the copy number is known, the activity value of the allele is multiplied by the gene copy number; if the copy number is unknown, the sample is assigned by multiplying by N by default, where N is a positive integer greater than or equal to 2; CYP2D6 AS is defined as the sum of the activity values ​​assigned to each allele; CYP2D6 AS was converted into phenotype using a classification system.

8. The method for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing according to claim 7, characterized in that: In the conversion of CYP2D6 AS into phenotype using the classification system: If the AS score is 0, the individual is a poor metabolizer; If the AS is 0.5 points, the individual is an intermediate metabolizer; If the AS is 1.0-2.0 points, the individual is a normal metabolizer; If the AS is above 2.0 points, the individual is an ultra-rapid metabolizer.

9. The method for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing according to claim 1, characterized in that: The sequencing tool used in steps S1 to S4 is the PacBio HiFi sequencing platform.

10. A system for detecting CYP2D6 gene polymorphism typing based on third-generation sequencing, characterized in that: include: The preprocessing module is used to preprocess the offline data to obtain HiFi reads; A data alignment module, used to align the HiFi reads to the GRCh38 reference genome to obtain a bam file; The filtering module is used to filter each sequence in the aligned bam file, retain the reads and bam format files that span the upstream of the CYP2D6 gene and the downstream of the CYP2D7 gene and meet the preset coverage depth threshold, and draw a conclusion on the haplotype of the sample; Genotyping module, used to classify alleles into haplotypes and assign star alleles to perform CYP2D6 gene di-allelic typing and obtain diploid gene information; The activity scoring module is used to perform activity scoring on the diploid gene information to obtain phenotype prediction and inference.