Haplotype recognition method based on single molecule sequencing technology

By constructing an amplification coefficient adjustment model and using single-molecule sequencing technology, the amplification bias in haplotype analysis was corrected, solving the difficulty of haplotype identification caused by differences in amplification efficiency, and achieving high-accuracy haplotype identification in complex structural regions.

CN121237210APending Publication Date: 2025-12-30BGI GENOMICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511239511.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing haplotype analysis techniques struggle to accurately distinguish and identify dominant and weak haplotypes when faced with differences in amplification efficiency, leading to an imbalance in the proportion of heterozygous SNP sites and affecting the accuracy of genetic information.

Method used

By constructing an amplification coefficient adjustment model, based on single-molecule sequencing technology and multiplex PCR, the amplification bias of different haplotypes is corrected. The amplification efficiency of primer pairs is adjusted using statistical models, mathematical models or machine learning models. Combined with methods such as kernel density estimation and wavelet convolution, haplotype peaks are identified and smoothed.

Benefits of technology

Under conditions of extreme imbalance in amplification efficiency, accurate identification and differentiation of dominant and weak haplotypes improve the overall accuracy and reliability of haplotype typing, ensuring the integrity and precision of genetic information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237210A_ABST
    Figure CN121237210A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biological information, in particular to a haplotype recognition method based on a single molecule sequencing technology. According to single molecule sequencing data of amplicons of known haplotype samples in a target area, the amplification efficiency of each primer pair in the multiple PCR primers on known haplotypes is obtained, then an amplification coefficient adjustment model is constructed according to the amplification efficiency, and the amplification coefficient adjustment model outputs amplification coefficients of each primer pair. The method is used for correcting amplification bias of different haplotypes in haplotype recognition. According to the method provided by the invention, by accurately correcting the efficiency difference of the primer pair in the amplification process, dominant and vulnerable haplotypes can be accurately identified and distinguished under the condition of extremely unbalanced coverage caused by amplification efficiency deviation, so that the overall accuracy and reliability of haplotype typing are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics, and particularly relates to a haplotype identification method based on single molecule sequencing technology. BACKGROUND

[0002] Accurate reconstruction of diploid genomes is critical for various studies, including disease association and population genetics, and most genome assembly algorithms fold diploid genomes with variants encoded at each heterozygous site into a single reference frame. To more accurately study genetic variations, the concept of "haplotype" is proposed, which refers to a combination of a series of genetic variation sites coexisting on a single chromosome. Haplotype analysis can effectively find heterozygous SNP variation sites on a single chromosome and plays an important role in genetic research and medical applications.

[0003] For haplotype analysis, the prior art generally analyzes SNPs in individual genomes to infer the phase relationship of adjacent SNPs on the same chromosome, thereby constructing a haplotype block. However, for amplicon data of the same region, due to the occurrence of structural variations, the amplification efficiency of two different haplotypes is greatly different, which poses a challenge to haplotype typing based on point mutations and variation information. In this case, the haplotype with high amplification efficiency may be over-represented in the sequencing data, while the haplotype with low amplification efficiency may be almost undetectable. Specifically, when the signal of one haplotype is much stronger than the other, the existing method tends to favor the typing result of the dominant haplotype, and ignores or fails to detect the weak haplotype, which is particularly true when dealing with genomic regions with complex structural variations. This results in a serious imbalance in the proportion of heterozygous SNP sites, reduces the accuracy of typing methods that rely on SNPs, and makes it difficult for haplotype typing methods based on point mutations and variation information to accurately distinguish between the two haplotypes, and may also lead to the omission of critical genetic information.

[0004] The existing haplotype analysis technology is generally based on the assumption that the coverage of the two haplotypes in the sequencing data is similar. However, in actual applications, this assumption does not hold when facing extreme amplification efficiency bias. The limitations presented by this assumption make it difficult for traditional haplotype analysis methods to accurately distinguish between two haplotypes when there is a significant difference in amplification efficiency.

[0005] Therefore, there is an urgent need to provide a haplotype identification method with high accuracy under amplification efficiency bias to improve the problem of uneven amplification efficiency of different haplotypes. SUMMARY

[0006] The present application aims to at least solve one of the problems of the related art.

[0007] The first aspect of the present application provides a model construction method for generating a model for correcting amplification bias of different haplotypes in haplotype identification based on single molecule sequencing technology, comprising: obtaining single molecule sequencing data of amplicons of known haplotype samples in a target region, wherein the amplicons are obtained based on a multiplex PCR system comprising a plurality of pairs of multiplex PCR primers; obtaining amplification efficiency of each primer pair in the plurality of pairs of multiplex PCR primers for the known haplotypes according to the single molecule sequencing data; and constructing an amplification coefficient adjustment model according to the amplification efficiency of each primer pair, wherein the amplification coefficient adjustment model outputs amplification coefficients of each primer pair, and is used to correct amplification bias of different haplotypes in haplotype identification, and the amplification coefficient adjustment model is constructed based on the types of each primer pair, the amplification length of each primer pair, and the amplification efficiency of each primer pair,

[0008] In some embodiments, the construction of the amplification coefficient adjustment model is performed by a statistical model, a mathematical model, a machine learning model, other artificial intelligence models, or any combination thereof. In some embodiments, the amplification coefficient adjustment model presents a function of amplification coefficients of each primer pair increasing with amplification length. In some embodiments, the amplification coefficient adjustment model is a linear model. In some embodiments, the amplification coefficient adjustment model is an exponential function.

[0009] In some embodiments, the amplification coefficient includes amplification weight and / or amplification probability.

[0010] The present application also provides a haplotype identification method based on single molecule sequencing technology, comprising: generating an amplification coefficient adjustment model based on single molecule sequencing data of amplicons of known haplotype samples in a target region according to the model construction method described in any of the embodiments of the present application, wherein the amplicons are obtained based on a multiplex PCR system comprising a plurality of pairs of multiplex PCR primers; obtaining length distribution of each amplicon of a to-be-identified haplotype sample in the target region according to single molecule sequencing data of amplicons of the to-be-identified haplotype sample; correcting the length distribution according to the theoretical amplification length of each primer pair based on the length distribution of each amplicon of the to-be-identified haplotype sample, to obtain a haplotype peak of the to-be-identified haplotype sample in the length distribution; calculating amplification coefficients of each primer pair corresponding to the haplotype peak based on the amplification coefficient adjustment model according to the corrected amplification length of each primer pair corresponding to the haplotype peak; obtaining a score distribution of the haplotype peak with the corrected amplification length according to the amplification coefficients of each primer pair corresponding to the haplotype peak; and identifying a haplotype of the to-be-identified haplotype sample according to the score distribution of the haplotype peak, wherein the amplicons of the to-be-identified haplotype sample are obtained based on the multiplex PCR system.

[0011] In some embodiments, identifying the haplotype of the haplotype sample to be identified according to the score distribution of the haplotype peak comprises: smoothing the score distribution of the haplotype peak; obtaining a peak point of the score distribution based on the smoothed score distribution; and identifying the haplotype of the haplotype sample to be identified according to the peak point of the score distribution.

[0012] In some embodiments, the score distribution of the haplotype peak is smoothed using any one or any combination of the following: kernel density estimation, wavelet convolution, histogram smoothing, locally weighted regression, spline interpolation, Gaussian process regression, Fourier transform, and adaptive filtering, preferably kernel density estimation and wavelet convolution.

[0013] In some embodiments, the method further comprises: identifying the haplotype of the haplotype sample to be identified based on a signal recognition threshold line according to the peak point of the score distribution.

[0014] In some embodiments, the signal recognition threshold line is a 10%, 15%, 20%, 25%, 30% or higher proportion threshold line, preferably a 10% proportion threshold line.

[0015] In some embodiments, the method further comprises: quality control filtering of the single molecule sequencing data of the amplicons of the haplotype sample to be identified in the target region, and optionally, retaining the single molecule sequencing data with a sequencing length ≥ 1000 nt, preferably ≥ 1400 nt, and a quality score ≥ 6, preferably ≥ 7, for haplotype identification.

[0016] In some embodiments, based on the length distribution of each amplicon of the haplotype sample to be identified showing two or more initial haplotype peaks, the method comprises: correcting the length distribution according to the theoretical amplification length of each primer pair to integrate the two or more initial haplotype peaks generated by each primer pair into a single haplotype peak, thereby obtaining the haplotype peak of the haplotype sample to be identified in the length distribution.

[0017] In some embodiments, the target region comprises single nucleotide variations, insertions and / or deletions, and / or structural variations, preferably structural variations. In some embodiments, the structural variations comprise segmental duplications, deletions, inversions and / or translocations.

[0018] The embodiments of the present application also provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of the embodiments of the present application described above.

[0019] The embodiment of the present application also provides a non-transitory computer-readable storage medium, the storage medium comprising a program, the program being capable of being executed by a processor to implement the method in any of the foregoing embodiments of the present application.

[0020] The embodiment of the present application also provides a computer program product comprising a computer program, the computer program implementing the method in any of the foregoing embodiments of the present application when executed by a processor.

[0021] The embodiment of the present application also provides a computer program comprising computer program code, which, when executed on a computer, causes the computer to perform the method in any of the foregoing embodiments of the present application.

[0022] The technical solution of the present application achieves the following technical effects:

[0023] The haplotype recognition method based on single molecule sequencing provided by the present application can accurately recognize and distinguish dominant and weak haplotypes under the condition of extremely unbalanced coverage caused by amplification efficiency deviation by precisely correcting the efficiency difference of primer pairs in the amplification process, thereby greatly improving the overall accuracy and reliability of haplotype typing. Even in a complex structure region, the method can realize complete and accurate extraction of haplotype information. The method provided by the embodiment of the present application provides a more accurate and reliable tool for the research of genetic information, and has great application prospect. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0025] Figure 1 A method flowchart for generating a model for correcting amplification bias of different haplotypes in haplotype recognition based on single molecule sequencing technology according to the embodiment of the present application is shown.

[0026] Figure 2 An amplification coefficient adjustment model constructed according to the embodiment of the present application is shown.

[0027] Figure 3 Haplotype peak distribution under a multiplex PCR system according to the embodiment of the present application is shown.

[0028] Figure 4 Score distribution of haplotype peaks after amplification coefficient adjustment according to the embodiment of the present application is shown.

[0029] Figure 5Electrophoresis results of amplified fragments of an exemplary sample according to embodiments of the present application are shown;

[0030] Figure 6 Amplified coefficient-adjusted haplotype peak distribution (left) and electrophoresis results of amplified fragments (right) of a verification sample according to embodiments of the present application are shown;

[0031] Figure 7 Amplified coefficient-adjusted haplotype peak distribution (left) and electrophoresis results of amplified fragments (right) of another verification sample according to embodiments of the present application are shown;

[0032] Figure 8 Amplified coefficient-adjusted haplotype peak distribution (left) and electrophoresis results of amplified fragments (right) of yet another verification sample according to embodiments of the present application are shown;

[0033] Figure 9 Amplified coefficient-adjusted haplotype peak distribution (left) and electrophoresis results of amplified fragments (right) of yet another verification sample according to embodiments of the present application are shown;

[0034] Figure 10 Amplified coefficient-adjusted haplotype peak distribution (left) and electrophoresis results of amplified fragments (right) of yet another verification sample according to embodiments of the present application are shown;

[0035] Figure 11 Regions and types of thalassemia detection according to embodiments of the present application are shown. DETAILED DESCRIPTION

[0036] The present application will be further described in details with specific embodiments, which are only for illustrating the present application and not intended to limit the scope of the present application. The embodiments provided below can be used as a guide for further improvement by those skilled in the art, and do not constitute any limitation on the present application.

[0037] In embodiments of the present application, "multiplex PCR" refers to a technique of simultaneously amplifying multiple target fragments in the same reaction system using multiplex PCR primers on a target region. In embodiments of the present application, the multiplex PCR primers are designed to simultaneously detect multiple variant sites in the target region, which not only broadens the coverage of detection, but also enhances the ability to distinguish different haplotypes, especially in the presence of complex structural variations in the genomic region.

[0038] In this embodiment, "single-molecule sequencing technology" refers to a technology capable of directly sequencing a single nucleic acid molecule (e.g., Heliscope single-molecule sequencing, single-molecule real-time (SMRT) sequencing, nanopore DNA sequencing, or any combination thereof), which can provide long read lengths and high-resolution sequencing data. Based on different sequencing platforms, single-molecule sequencing technology can generate reads up to hundreds of kb in length. In this embodiment, single-molecule sequencing technology enables full-length sequencing of the amplified products from multiplex PCR, thereby accurately obtaining high-quality nucleic acid sequence data of all haplotype fragments within the target region. Subsequent haplotype identification will be based on single-molecule sequencing data and primer pair identification, avoiding the complex process of relying on heterozygous site encoding calculations to obtain haplotype information in conventional techniques. In this embodiment, by combining multiplex PCR and single-molecule sequencing, the efficiency and accuracy of haplotype identification are effectively improved.

[0039] In this application, a "haplotype" refers to a group of closely linked genes or genetic markers located on a single chromosome. These genes or markers are typically inherited together and rarely recombine. Haplotypes are very important in genetic research, especially in the study of complex diseases, population genetics, and evolutionary biology.

[0040] In this embodiment, the target region of the haplotype sample to be identified may contain single nucleotide variants (SNPs), insertions and / or deletions (InDel, generally ≤50bp), and / or structural variants (SNVs, generally >50bp, such as fragment duplications, deletions, inversions, and / or translocations). The haplotype identification method proposed in this embodiment performs excellently for target regions containing various variants, especially complex structural variants, and can ensure a high level of genotyping accuracy under various experimental conditions, especially when facing extreme amplification efficiency deviations.

[0041] This application proposes a model for correcting amplification bias of different haplotypes in haplotype identification. This model is constructed through the following steps:

[0042] S101: Obtain single-molecule sequencing data of amplicones from known haplotype samples within the target region, wherein the amplicones are obtained based on a multiplex PCR system containing multiplex PCR primer pairs.

[0043] In this embodiment, a multiplex PCR system containing multiplex PCR primer pairs is used to amplify the target region of a known haplotype sample to obtain amplicons of the known haplotype sample in this system. These amplicons are then sequenced individually to obtain sequencing data of the multiplex PCR amplification products of the known haplotype sample.

[0044] S102: Based on the single-molecule sequencing data of amplicon of known haplotype samples in the target region, obtain the amplification efficiency of each primer pair in the multiplex PCR primers for known haplotypes.

[0045] In this embodiment, the amplification efficiency of each primer pair for the known haplotype is obtained based on the sequencing data of the multiplex PCR amplification products of a known haplotype sample in the multiplex PCR system. In some embodiments, the amplification efficiency can be determined by the read count, data volume, or sequencing depth of each amplicon.

[0046] S103: Based on the amplification efficiency of each primer pair, an amplification coefficient adjustment model is constructed. This model outputs the amplification coefficient of each primer pair to correct the amplification bias of different haplotypes in haplotype identification.

[0047] In this embodiment, the amplification efficiency of each primer pair is used as an amplification bias parameter. Based on this amplification bias parameter, the type of primer pair, and the amplification length of each primer pair, an amplification coefficient adjustment model is constructed. In this embodiment, the amplification coefficient adjustment model is used to output the amplification coefficient of each primer pair in the multiplex PCR system. This coefficient can be used to adjust the amplification status of each amplicon in subsequent haplotype identification of unknown samples, thereby correcting amplification efficiency deviations and accurately identifying and distinguishing dominant and weak haplotypes even when amplification efficiencies are extremely unbalanced, thereby improving the completeness and accuracy of haplotype identification.

[0048] In some embodiments, an amplification coefficient adjustment model can be constructed using statistical models, such as linear models, mathematical models, such as probabilistic models (e.g., Bayesian models), machine learning models, other artificial intelligence models, or any combination thereof. In some embodiments, the amplification coefficient adjustment model presents a function of the amplification coefficient of each primer pair increasing with the amplification length; that is, theoretically, the longer the target band is amplified, the greater the amplification difficulty and the lower the amplification efficiency, and therefore the amplification adjustment coefficient of that primer pair should be assigned a high value. It is understood that the amplification coefficient adjustment model is related to the primer system, primer type, and amplification length.

[0049] In some specific embodiments, an amplification coefficient adjustment model can be constructed based on a linear model. In some embodiments, the amplification coefficient adjustment model can be constructed in the form of an exponential equation, and the amplification coefficient adjustment model can be obtained by fitting the amplification coefficient and amplification length of each primer pair to the exponential equation.

[0050] In other specific embodiments, an amplification coefficient adjustment model can be constructed based on a machine learning model. In some embodiments, the amplification coefficient adjustment model can be trained by using some or all of the data from known haplotype samples as a training set, where the training set includes at least the amplification lengths of each primer pair for the known haplotype and the corresponding amplification bias labels; and the model parameters can be updated by minimizing the loss function using an optimization algorithm (such as gradient descent). In some embodiments, the remaining data from the known haplotype samples can also be used as a validation set to further optimize the generated amplification coefficient adjustment model. This application does not limit the specific construction steps of the amplification coefficient adjustment model.

[0051] In some embodiments, the amplification coefficient adjustment model can be expressed as:

[0052] Amplification efficiency = α * e^(-β * amplification length) + γ * primer coefficient y

[0053] Here, α, β, and γ are model parameters, which can be obtained through training with known sample data.

[0054] In this embodiment of the application, based on model construction under different algorithms, the amplification coefficient output by the amplification coefficient adjustment model may include amplification weight and / or amplification probability.

[0055] This application's embodiments, through precise modeling of amplification efficiency differences, can accurately correct for primer pair efficiency differences during amplification when identifying unknown haplotypes. This allows for accurate identification and differentiation of dominant and weak haplotypes even under conditions of extreme coverage imbalance caused by amplification efficiency deviations, thereby significantly improving the overall accuracy and reliability of haplotype typing. Even in complex structural regions, the adjustment model proposed in this application's embodiments can achieve complete and accurate extraction of haplotype information. The adjustment model proposed in this application's embodiments provides a more precise and reliable tool for genetic information research and has great application potential.

[0056] Based on the above model, this application also proposes a haplotype identification method based on single-molecule sequencing technology, which may include the following steps:

[0057] S201: Generate an amplification coefficient adjustment model based on single-molecule sequencing data of amplicon from known haplotype samples within the target region, wherein the amplicon is obtained based on a multiplex PCR system containing multiplex PCR primer pairs.

[0058] S202: Based on the single-molecule sequencing data of the amplicon of the haplotype sample to be identified within the target region, obtain the length distribution of each amplicon of the haplotype sample to be identified.

[0059] In this embodiment, based on the same multiplex PCR system used to amplify samples with known haplotypes, amplification is performed using the target region of the sample to be identified as a template to obtain amplicones of the sample to be identified. These amplicones are then sequenced individually to obtain sequencing data of the multiplex PCR amplification products of the sample to be identified. The sequencing data is analyzed to obtain the length distribution of all amplicones in each sample to be identified.

[0060] S203: Based on the length distribution of each amplicon in the haplotype sample to be identified, the length distribution is corrected according to the theoretical amplification length of each primer pair to obtain the haplotype peak of the haplotype sample in the length distribution.

[0061] In this embodiment, due to the application of a multiplex PCR system, a haplotype can be amplified by multiple primer pairs to produce fragments with different amplification lengths and efficiencies. That is, the length distribution of each amplicon in the sample to be identified will show two or more initial haplotype peaks. In this embodiment, by correcting for the theoretical amplification length of the primer pairs, all fragments generated by the same haplotype amplified by different primer pairs are merged into a single haplotype peak, and the haplotype peak distribution is obtained.

[0062] S204: Based on the corrected amplification length of each primer pair corresponding to the haplotype peak, and using the amplification coefficient adjustment model, calculate the amplification coefficient of each primer pair corresponding to the haplotype peak.

[0063] In this embodiment, by inputting the corrected amplification length of each primer pair corresponding to each haplotype peak into the amplification coefficient adjustment model, the amplification coefficient (e.g., amplification weight and / or amplification probability) of each primer pair corresponding to each haplotype peak is obtained. This accurately corrects the efficiency difference of primer pairs during amplification, thereby accurately identifying and distinguishing dominant and weak haplotypes under the condition of extremely unbalanced coverage caused by amplification efficiency deviation.

[0064] S205: Based on the amplification coefficients of each primer pair corresponding to the haplotype peak, obtain the score distribution of the haplotype peak with the corrected amplification length.

[0065] In this embodiment, after allocating amplification coefficients to each amplicon, the score distribution is plotted on a two-dimensional plane, where the y-axis represents the score and the x-axis represents the corrected amplification length, thereby obtaining the score distribution of the haplotype peaks. It is understood that the score Y can be a linear function of the amplification coefficient (primer coefficient) y, and this application does not limit the specific calculation method for the score Y.

[0066] S206: Identify the haplotype of the sample to be identified based on the score distribution of the haplotype peak.

[0067] In some embodiments, step S206 may include: smoothing the score distribution of the haplotype peak; obtaining the peak point of the score distribution based on the smoothed score distribution; and identifying the haplotype of the haplotype sample to be identified based on the peak point of the score distribution.

[0068] In some embodiments, the score distribution of the haplotype peaks can be smoothed using any of the following or any combination thereof: kernel density estimation, wavelet convolution, histogram smoothing, local weighted regression, spline interpolation, Gaussian process regression, Fourier transform, and adaptive filtering, preferably kernel density estimation and wavelet convolution. The smoothing of the score distribution in this application embodiment can effectively reduce the impact of noise and improve the accuracy of the genotyping results. Furthermore, by combining this with the adjustment of the amplification coefficients of each primer pair, this method maintains high accuracy even under conditions of numerous amplified abnormalities, abundant interference signals, and extremely unbalanced data coverage.

[0069] In some embodiments, the method of this application further includes data filtering to ensure high-quality haplotype resolution. In some embodiments, data filtering may include quality control of the raw sequencing data (i.e., single-molecule sequencing data) of each amplicon, for example, selectively retaining single-molecule sequencing data with a sequencing length ≥1000 nt, preferably ≥1400 nt, and a quality score ≥6, preferably ≥7, for haplotype identification. In some embodiments, data filtering further includes setting a signal identification threshold for the peak points of the score distribution, and treating signals above this threshold as valid signals. In some embodiments, the signal identification threshold is a 10%, 15%, 20%, 25%, 30%, or higher percentage threshold, preferably a 10% percentage threshold. For example, a 10% percentage threshold means multiplying the maximum signal intensity value by 10% as the threshold, and considering signal intensities below this threshold as noise or invalid signals. The method of this application, by introducing a multi-step quality control process, further improves the accuracy and reliability of the data, thereby enhancing the accuracy of haplotype identification. It is understood that the method in the embodiments of this application can also introduce other quality control processes, such as setting more stringent filtering thresholds to further improve the quality and reliability of the data. This application does not limit other data quality control processes that may be introduced.

[0070] The haplotype identification method based on single-molecule sequencing proposed in this application accurately identifies and distinguishes dominant and weak haplotypes even under conditions of extreme coverage imbalance caused by amplification efficiency deviations by precisely correcting for differences in primer pair efficiency during amplification. This significantly improves the overall accuracy and reliability of haplotype typing. Even in complex structural regions, this method can achieve complete and accurate extraction of haplotype information. The method proposed in this application provides a more precise and reliable tool for genetic information research and has great application potential.

[0071] Those skilled in the art will understand that all or part of the functions of the various methods in the above embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved. In addition, when all or part of the functions in the above embodiments are implemented by computer programs, the program can also be stored in a storage medium such as a server, another computer, disk, optical disk, flash drive, or portable hard drive, and can be downloaded or copied to the memory of a local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be achieved.

[0072] Another embodiment of this application provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method according to any of the above embodiments.

[0073] Another embodiment of this application provides a non-transitory computer-readable storage medium including a program that can be executed by a processor to implement the method according to any of the above embodiments.

[0074] Unless otherwise specified, the experimental methods in the following embodiments are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product manual. For example, unless otherwise specified, the process parameters in each step of the following embodiments are the default parameters of the corresponding software / script.

[0075] Example

[0076] This embodiment targets the thalassemia gene region for haplotype identification. The specific method is as follows:

[0077] 1. Multiplex PCR

[0078] Targeting thalassemia deficiency and repetitive thalassemia areas ( Figure 11 Standard design of multiplex PCR primers for detecting these common regions ( Figure 11 Based on these primers, haplotype long fragments of thalassemia gene regions in known thalassemia deletion and duplication samples were amplified using multiple primers. The reaction system was set to 25 μL, containing 20-100 ng gDNA, 0.4 μmol / L PCR primer mixture, 0.25 U ApexHF HSDNA polymerase CL (Accurate Biotech, Hunan, China), and 12.5 μL 2×ApexHF CL Buffer (Accurate Biotech, Hunan, China). The PCR reaction conditions were as follows: 8 cycles of initial denaturation at 94℃ for 5 min, denaturation at 98℃ for 10 s, and extension at 68℃ for 10 min; followed by 19 cycles of denaturation at 98℃ for 10 s, annealing at 64℃ for 30 s, annealing at 60℃ for 30 s, and extension at 68℃ for 10 min; finally, a final extension at 68℃ for 10 min was performed, and the PCR products were stored at 4℃. For each sample, 15 μL of the PCR product was taken and purified using 0.5 volume of VAHTSDNAClean Beads (Vazyme Biotech, Nanjing, China), and the concentration was diluted to 40 ng / μL.

[0079] 2. Library construction and single-molecule sequencing

[0080] 2.1 Library Construction

[0081] Library construction was performed using the CycloneSEQ kit (Cyclone). First, following the kit instructions, 2 μg of purified DNA product underwent end repair using EA enzyme (CycloneSEQ, Hangzhou, China). The product was purified using 1.0 volume of DNA cleaning beads and eluted with 62 μL of nuclease-free water. The concentrations of the end-repaired and A-tailed products were measured. Then, the end-repaired and A-tailed products were ligated to adapters using the Cyclone ligation module, purified with 1.0 volume of DNA cleaning beads, and eluted with 22 μL of elution buffer. The eluent was the adapter-ligated library to be tested.

[0082] 2.2 Single-molecule sequencing

[0083] The concentration of the library to be tested was determined. Based on the determined concentration, 900 ng of the library was taken for sequencing, and single-molecule sequencing of the library was performed using a Cyclone WT02 sequencer according to its instruction manual. Each read obtained from sequencing corresponds to an amplicon sequence.

[0084] 3. Data Analysis

[0085] 3.1 Data quality control processing

[0086] For the raw sequencing data of Cyclone WT02, fastp (version 0.23.4) was used for filtering, with the criteria being a read length greater than 1400 nt and a Phred quality score ≥7 (Q score ≥7).

[0087] 3.2 Data Demultiplexing

[0088] The filtered sequencing data contains data from 48 independent samples, each with its own characteristic sequence tag (introduced during library construction). First, primer and sample tag identification is required to distinguish data from different samples and achieve demultiplexing. In this embodiment, a self-built Python script is used with the edlib module for sequence identification, and a 20% error tolerance is set to accommodate sequencing errors. This script is used to classify the filtered sequencing data according to sample origin and to label primer pair information in the name of each read.

[0089] 3.3 Training of the Amplification Coefficient Adjustment Model

[0090] For individual sample data, an exponential equation was fitted based on the amplification efficiency of each primer pair for samples with known missing and repetitive information (the longer the theoretical target band, the greater the amplification difficulty and the lower the amplification efficiency). Data from all 48 samples were used for training to establish an amplification coefficient adjustment model y = 0.05e^(0.0003x) related to the amplification length (e.g., ...). Figure 2 As shown in the figure, x is the amplification length, and y is the amplification coefficient (i.e., primer pair weight) of the primer pair for a known haplotype sample in this reaction system. The amplification coefficient adjustment equation is related to the primer system configuration, primer type, and amplification length.

[0091] 3.4 Haplotype Identification

[0092] The sequencing data of the sample is read, and the data is organized according to the length of each read and the type of primer pair. The length distribution of all reads in the sample is calculated, with the vertical axis representing the read count. Taking one sample THAL250-06 out of 48 samples as an example, the haplotype identification proposed in this embodiment is demonstrated. The known haplotype data of this sample THAL250-06 is αα / -α3.7, that is, one haplotype is of normal length, and the other haplotype has a 3.7kb fragment deletion.

[0093] Figure 3 and Figure 4 The images show the raw read length data for THAL250-06 and the haplotype information derived from the adjusted model and score distribution. Figure 3 As shown, theoretically, each primer pair produces a haplotype peak with a theoretical band amplification, which is reflected as a haplotype peak in the sequencing data. However, since this embodiment uses a multiplex PCR system, a haplotype can be amplified by multiple primer pairs to produce band fragments with different amplification lengths and different amplification efficiencies. It is necessary to use primer pair theoretical amplification length correction to merge all bands of the same haplotype amplified by different primer pairs into a single haplotype peak. Figure 3 In sample THAL250-06, three peaks of different heights were observed according to the amplification length. These peaks were obtained by amplifying two haplotypes of the sample using different primer pairs. By combining the primer pair information, the data of the two haplotypes can be distinguished.

[0094] Furthermore, based on the amplification length and primer pair type, and combined with the amplification coefficient adjustment model, the data of the two haplotypes can be distinguished. After weighting all reads, the score distribution is plotted on a two-dimensional plane, with the y-axis representing the score information and the x-axis representing the corrected amplification length. The score distribution of the haplotype peaks is plotted as follows: Figure 4 As shown, the score distribution is smoothed using KDE, and then the weight distribution is further smoothed using wavelet convolution. Peak points are identified using the smoothed weight data distribution, and haplotype information is identified using a 10% signal identification threshold line. Finally, the reads data are classified into the corresponding sample haplotype data according to their haplotype source.

[0095] In this embodiment, according to Figure 3 The peak value was obtained by calibrating the amplification length. Figure 4 The two haplotype score peaks are formed by Figure 4 It can be seen that the position of the first peak represents the amplification result of the -α3.7 haplotype, and the position of the second peak represents the amplification result of the αα haplotype. This is consistent with... Figure 5 The haplotype results shown in the electrophoresis images are consistent, indicating that the haplotype identification method of this embodiment is effective and accurate.

[0096] Figures 6 to 10 The haplotype identification results using known samples other than the 48 samples used in this embodiment are shown below. The information of the known haplotypes for each sample is as follows:

[0097] THAL00042-271αα / αα

[0098] THAL00042-275αα / -α3.7

[0099] THAL00042-330--SEA / --SEA

[0100] THAL00042-360 THAL / αα

[0101] THAL00042-365--SEA / anti3.7

[0102] Figures 6-10 The left side of each figure shows the haplotype score peak distribution after integration according to the method of this embodiment, and the right side shows the electrophoresis diagram of the amplification result corresponding to that sample. The red label on the right indicates the theoretical amplification length of each haplotype associated with thalassemia. Figures 6-10 As can be seen, the score distributions obtained according to the method of this embodiment are consistent with the electrophoresis results. By comparing the score distributions or electrophoresis results with the labels, the haplotypes of each sample can be determined, and the determined haplotype information is consistent with the known results listed above. This proves the effectiveness and accuracy of the method of this embodiment.

[0103] In summary, by correcting for amplification efficiency bias, this invention can accurately identify and distinguish dominant and weaker haplotypes even when amplification efficiencies are extremely unbalanced. This precise identification capability significantly improves the accuracy of haplotype typing and overcomes the shortcomings of traditional methods when faced with extreme differences in amplification efficiency. This feature ensures that haplotype information can be extracted completely and accurately in complex genomic regions.

[0104] Secondly, this invention utilizes multiplex PCR technology and a linear model to weight primer pair information and read length information, and performs smoothing through kernel density estimation (KDE). This process ensures the reliability of haplotype identification, maintaining a high level of accuracy even under conditions of significant amplification, numerous interfering signals, and extremely unbalanced data coverage. This method effectively reduces the impact of noise and improves the accuracy of genotyping results.

[0105] Furthermore, by employing multiplex PCR technology and identifying and merging haplotypes with multiple primer pairs, this invention enables accurate haplotype identification in highly variable amplification regions. The use of multiple primers not only improves the detection coverage but also enhances the ability to distinguish between different haplotypes, especially in genomic regions with complex structural variations.

[0106] The technical solution of this invention also simplifies the data processing workflow. By automating primer pair identification and haplotype identification steps, the cumbersome heterozygous site encoding calculation process in conventional methods is avoided, thereby improving the overall operational efficiency. This automation and workflow optimization not only reduces the need for manual intervention but also accelerates data processing, making large-scale genome research more efficient.

[0107] Finally, this invention is applicable not only to disease association studies but also to population genetics and other research fields requiring precise haplotype typing. Its broad applicability and adaptability to diverse research needs make it an indispensable tool in related research fields, providing strong support for advancing genomics research.

[0108] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0109] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A model construction method characterized by comprising: The method is used for generating a model for correcting amplification bias of different haplotypes in haplotype identification based on single molecule sequencing technology, and the method comprises: obtaining single molecule sequencing data of amplicons of known haplotype samples in a target region, wherein the amplicons are obtained based on a multiplex PCR system comprising a plurality of pairs of multiplex PCR primers; obtaining amplification efficiency of each primer pair in the plurality of pairs of multiplex PCR primers on the known haplotype according to the single molecule sequencing data; and constructing an amplification coefficient adjustment model according to the amplification efficiency of each primer pair, wherein the amplification coefficient adjustment model outputs amplification coefficients of each primer pair, and is used to correct amplification bias of different haplotypes in haplotype identification.

2. The method of claim 1, wherein, The amplification coefficient adjustment model is constructed based on the types of each primer pair, the amplification length of each primer pair and the amplification efficiency of each primer pair, wherein the construction of the amplification coefficient adjustment model is performed by a statistical model, a mathematical model, a machine learning model, other artificial intelligence models or any combination thereof.

3. The method of claim 2, wherein, The amplification coefficient adjustment model presents a function of amplification coefficients of each primer pair increasing with amplification length, optionally, the amplification coefficient adjustment model is a linear model, optionally, the amplification coefficient adjustment model is an exponential function.

4. The method according to any one of claims 1 to 3, characterized in that, The amplification coefficient comprises an amplification weight and / or an amplification probability.

5. A haplotype calling method based on single molecule sequencing technology, characterized by, The method comprises: generating an amplification coefficient adjustment model based on single molecule sequencing data of amplicons of known haplotype samples in a target region according to the model construction method of any one of claims 1 to 4, wherein the amplicons are obtained based on a multiplex PCR system comprising a plurality of pairs of multiplex PCR primers; obtaining length distribution of each amplicon of a to-be-identified haplotype sample according to single molecule sequencing data of amplicons of the to-be-identified haplotype sample in the target region; correcting the length distribution according to the theoretical amplification length of each primer pair based on the length distribution of each amplicon of the to-be-identified haplotype sample, to obtain a haplotype peak of the to-be-identified haplotype sample in the length distribution; calculating amplification coefficients of each primer pair corresponding to the haplotype peak based on the amplification coefficient adjustment model according to the corrected amplification length of each primer pair corresponding to the haplotype peak; obtaining a score distribution of the haplotype peak with the corrected amplification length according to the amplification coefficients of each primer pair corresponding to the haplotype peak; and identifying a haplotype of the to-be-identified haplotype sample according to the score distribution of the haplotype peak, wherein the amplicons of the to-be-identified haplotype sample are obtained based on the multiplex PCR system.

6. The method of claim 5, wherein, The identification of the haplotype of the to-be-identified haplotype sample according to the score distribution of the haplotype peak comprises: performing smoothing processing on the score distribution of the haplotype peak; obtaining a peak point of the score distribution based on the smoothed score distribution; and identifying the haplotype of the to-be-identified haplotype sample according to the peak point of the score distribution, Optionally, the score distribution of the haplotype peak is smoothed using any one or any combination of the following: kernel density estimation, wavelet convolution, histogram smoothing, locally weighted regression, spline interpolation, Gaussian process regression, Fourier transform, and adaptive filtering, preferably kernel density estimation and wavelet convolution.

7. The method of claim 6, wherein, The method further comprises: identifying, according to the peak points of the score distribution, a haplotype of the haplotype sample to be identified based on a signal recognition threshold line, Optionally, the signal recognition threshold line is a 10%, 15%, 20%, 25%, 30% or higher proportion threshold line, preferably a 10% proportion threshold line, Optionally, the method further comprises: performing quality control filtering on the single molecule sequencing data of the amplicons of the haplotype sample to be identified in the target region, and optionally, retaining the single molecule sequencing data with a sequencing length of ≥1000 nt, preferably a sequencing length of ≥1400 nt and a quality score of ≥6, preferably a quality score of ≥7, for haplotype identification.

8. The method according to any one of claims 5 to 7, characterized in that, Based on the length distribution of each amplicon of the haplotype sample to be identified showing two or more initial haplotype peaks, the method comprises: correcting the length distribution according to the theoretical amplification length of each primer pair to integrate the two or more initial haplotype peaks generated by each primer pair into a single haplotype peak, thereby obtaining the haplotype peak of the haplotype sample to be identified in the length distribution, Optionally, the target region comprises single nucleotide variations, insertions and / or deletions, and / or structural variations, preferably structural variations, Optionally, the structural variations comprise fragment duplication, deletion, inversion and / or translocation.

9. An electronic device, comprising: at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

10. A non-transitory computer-readable storage medium, comprising: The storage medium comprises a program executable by the processor to implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Compositions and Method for Measuring and Calibrating Amplification Bias in Multiplexed PCR Reactions

    US20150017652A1

  • Method for DNA mixture analysis

    US6807490B1