Bioinformatics analysis method and system of massarray pharmacogenomic genotyping
Patent Information
- Application Number
- CN202611013165.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-07-08
AI Technical Summary
因此在实际应用中,存在人工复核导致的实验流转时间增加的现象
本发明开发了一种MassArray药物基因组分型的全新分析方法,相较于MassArray现有的软件结果,能够显著提高分型检出比例和准确性。
Smart Images

Figure CN122531474B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics technology, specifically to a pharmacogenomics analysis method and system for the MassARRAY platform used for single nucleotide polymorphism (SNP) typing and gene mutation detection. Background Technology
[0002] Single nucleotide polymorphisms (SNPs), as the most common form of genetic variation in the genome, play a crucial role in disease susceptibility assessment, pharmacogenomics, forensic individual identification, and population genetics research. The MassARRAY system (originally developed by Sequenom, now part of Agena Bioscience) based on matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF MS) has emerged to address this need. The core of this technology lies in directly measuring the molecular weight of DNA fragments to determine genotypes, avoiding biases introduced by fluorescent labeling or hybridization processes. Its core principle is as follows: the target fragment is amplified by multiplex PCR, followed by a single-base extension reaction, extending only one dideoxynucleotide (ddNTP) complementary to the template at the location immediately adjacent to the SNP site. Because the four bases (A, T, C, G) have different molecular weights, different genotypes will produce extension products with slightly different molecular weights. Finally, by accurately measuring the molecular weight (time of flight) of these products using MALDI-TOF MS, the SNP genotype can be directly determined.
[0003] MassArray boasts a high degree of automation from sample processing to data analysis, achieving an excellent balance between sample throughput, detection multiplicity, and cost. It has become a widely adopted targeted genotyping platform in pharmacogenomics research and clinical testing. It is particularly suitable for rapid, high-volume detection of SNPs, Indels, and copy number variations (CNVs) of known drug-metabolizing enzymes and drug target genes (such as CYP2C9, CYP2C19, CYP2D6, VKORC1, etc.) with clear clinical guidance.
[0004] Existing MassArray analyses are generally based on locus genotyping results provided by its accompanying analysis software (such as Typer). A study from the University of Copenhagen found that MassArray had similar genotyping rates across different DNA ratios, but in 1:1, 1:3, and 1:5 DNA mixtures, the 1:5 mixture showed a higher detection rate for "homozygous traits," while the 1:1 and 1:3 mixtures showed lower detection rates and higher error rates (see [link to study]). Figure 1DNA from two samples (S1 and S2) was mixed in five different ratios (1:5, 1:3, 1:1, 3:1, and 5:1) and genotyping was performed on 29 SNP loci (Reference 1: Scott, Stuart A et al. Development and Analytical Validation of a 29 Gene Clinical Pharmacogenetic Genotyping Panel: Multi-Ethnic Allele and Copy Number Variant Detection. Clinical and translational science vol. 14,1 (2021): 204-213. doi:10.1111 / cts.12844). Besides a lower genotype call rate and lower detection capability for heterozygotes compared to homozygotes, MassArray exhibits a higher false detection rate in low-quality DNA (Reference 2: Johansen P, Andersen J, Børsting C et al. Evaluation of the iPLEX® Sample ID Plus Panel designed for the Sequenom MassARRAY® system. A SNP typing assay developed for human identification and sample tracking based on the SNPforID Panel. Forensic Science International: Genetics, 7, 482-487). Therefore, in practical applications, manual verification increases experimental turnaround time. In view of this, the present invention is proposed. Summary of the Invention
[0005] To address the aforementioned issues, this patent constructs a new indicator—the Second-to-Max Ratio (SMR)—based on the signal-to-noise ratio (SNR), and develops a corresponding analysis method and system for MassArray pharmacogenomics typing. This system can significantly improve the detection efficiency and accuracy of personalized pharmacogenomics testing.
[0006] To achieve the above objectives, the present invention proposes the following specific technical solutions.
[0007] This invention first provides a bioinformatics analysis method for MassArray pharmacogenomics genotyping, comprising the following steps: Step 1: SNP site SMR threshold construction; Step 2: Genotyping based on SMR threshold.
[0008] Furthermore, step 1, the construction of the site SMR threshold, specifically includes: 1) Baseline sample selection: Select baseline samples that cover all target loci and are known to be homozygous and heterozygous; preferably, the number of samples is 20-50; the ratio of homozygous to heterozygous samples is 9-1:1-9.
[0009] 2) Calculate the SMR value for each locus: Use the MassArray platform to detect the aforementioned baseline samples and obtain the offline data for each locus, including the SNR value; sort the SNR values in descending order and define the second-to-MaxRatio (SMR); the SMR = Second_SNR (second-largest SNR) / Max_SNR (maximum SNR). Preferably, the SMR range is 0-1: when the SMR is closer to 0, it indicates that the site is more homozygous; when the SMR is closer to 1, it indicates that the site is more heterozygous.
[0010] 3) Construct SMR thresholds for each site For each locus, the baseline samples are divided into homozygous and heterozygous groups based on the known homozygous results. The dataset consisting of SMR in the interquartile range above the homozygous locus is defined as G1, and the dataset consisting of SMR in the interquartile range below the heterozygous locus is defined as G2. The upper limit of homozygosity and the lower limit of heterozygosity are obtained based on the median values of G1 and G2 and the 3 times MAD interval. Preferably, the upper limit of homozygosity = median value of G1 + 3 * MAD of G1; the lower limit of heterozygosity = median value of G2 - 3 * MAD of G2; The interval between the lower limit of heterozygosity and the upper limit of homozygosity is divided into four equal parts. The upper quartile (Q1) is selected as the threshold for the heterozygous genotype of the target locus, and the lower quartile (Q3) is selected as the threshold for the homozygous genotype of the target locus.
[0011] Furthermore, step 2, genotyping based on site SMR thresholds, specifically includes: a) Detection of the sample to be tested: Mass spectrometry detection of the sample to be tested is performed using the MassArray platform to obtain offline data including SNR value; b) Quality control: Perform quality control on the obtained SNR values; when SNR < 3, do not calculate SMR, and retest the sample to be tested; c) Homozygous analysis: Calculate the SMR of each site in the test sample after quality control, and compare it with the threshold constructed in step 1 to determine the homozygous status of the test site; Preferably, the interpretation criteria are: when SMR ≤ homozygous threshold, the site is determined to be homozygous; when SMR ≥ heterozygous threshold, the site is determined to be heterozygous.
[0012] d) Genotype interpretation: After determining homozygous loci, the genotype of the locus is determined by combining the largest SNR and the second largest SNR. Preferably, the interpretation is as follows: if the locus is homozygous, the locus genotype is homozygous for the genotype corresponding to the largest SNR; if the locus is heterozygous, the locus genotype is heterozygous for the genotypes corresponding to the largest SNR and the second largest SNR; in some specific examples, if the locus is homozygous and the genotype corresponding to the largest SNR is T, then the locus genotype is TT; if the locus is heterozygous and the largest SNR and the second largest SNR correspond to A and G respectively, then the locus genotype is AG.
[0013] This invention also provides a bioinformatics analysis system for MassArray pharmacogenomics genotyping, the system comprising the following modules: Module 1, Site SMR Threshold Construction Module; Module 2, Genotyping Module Based on Site SMR Thresholds.
[0014] Furthermore, the site SMR threshold construction module performs the aforementioned site SMR threshold construction step; the genotyping module performs the aforementioned SMR threshold-based genotyping step.
[0015] Specifically, module 1, the site SMR threshold construction module, includes: 1) Baseline sample selection module: used to select baseline samples that cover all target sites and are known to be homozygous and heterozygous; preferably, the number of samples is 20-50; the ratio of homozygous to heterozygous samples is 9-1:1-9.
[0016] 2) Module for calculating SMR value for each locus: This module calculates the SMR value for each locus by detecting the offline data (including SNR value) for each locus obtained from the aforementioned baseline samples using the MassArray platform. The SNR values are sorted in descending order, and the second-to-max ratio (SMR) is defined. The SMR is defined as: SMR = Second_SNR / Max_SNR. Preferably, the SMR range is 0-1: when the SMR is closer to 0, it indicates that the site is more homozygous; when the SMR is closer to 1, it indicates that the site is more heterozygous.
[0017] 3) Construct SMR threshold modules for each locus: For each locus, the baseline samples are divided into homozygous and heterozygous groups based on the known homozygous results. The dataset consisting of SMR in the interquartile range above the homozygous locus is defined as G1, and the dataset consisting of SMR in the interquartile range below the heterozygous locus is defined as G2. The upper limit of homozygosity and the lower limit of heterozygosity are obtained based on the median values of G1 and G2 and the 3 times MAD interval. Preferably, the upper limit of homozygosity = median value of G1 + 3 * MAD of G1; the lower limit of heterozygosity = median value of G2 - 3 * MAD of G2; The interval between the lower limit of heterozygosity and the upper limit of homozygosity is divided into four equal parts. The upper quartile (Q1) is selected as the threshold for the heterozygous genotype of the target locus, and the lower quartile (Q3) is selected as the threshold for the homozygous genotype of the target locus.
[0018] Furthermore, module 2, the SMR threshold-based genotyping module, includes: 1) Off-stage data acquisition module: The MassArray platform is used to perform mass spectrometry detection on the test sample to acquire off-stage data including SNR values; 2) Quality control module: used to perform quality control on the obtained SNR value; when SNR<3, SMR is not calculated and the sample to be tested is retested; 3) Hybrid heterozygosity analysis module: used to calculate the SMR of each site in the test sample after quality control, and compare it with the threshold constructed in step 1 to determine the homozygosity of the test site; Preferably, the interpretation criteria are as follows: if the SMR of the site to be tested is less than or equal to the homozygous threshold, the site is determined to be homozygous; if the SMR of the site to be tested is greater than or equal to the heterozygous threshold, the site is determined to be heterozygous; if the heterozygous threshold is less than the SMR of the site to be tested and less than the homozygous threshold, it cannot be determined.
[0019] 5) Genotype interpretation module: After determining homozygous locus, it is used to interpret the genotype of the locus by combining the largest SNR and the second largest SNR; Preferably, the interpretation is as follows: if the locus is homozygous, the locus genotype is homozygous for the genotype corresponding to the largest SNR; if the locus is heterozygous, the locus genotype is heterozygous for the genotypes corresponding to the largest SNR and the second largest SNR; in some specific examples, if the locus is homozygous and the genotype corresponding to the largest SNR is T, then the locus genotype is TT; if the locus is heterozygous and the largest SNR and the second largest SNR correspond to A and G respectively, then the locus genotype is AG.
[0020] The present invention also provides the application of any of the above-described methods or bioinformatics analysis systems in SNP locus genotyping.
[0021] The present invention also provides an electronic device, comprising: a processor and a memory; the processor and the memory are connected together, wherein the memory is used to store a computer program, and the processor is used to invoke the computer program to execute the method described in any of the preceding claims.
[0022] The present invention also provides a computer storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, perform the method described in any of the preceding claims.
[0023] Beneficial technical effects of the present invention: This invention develops a novel analytical method for MassArray pharmacogenomics typing, which significantly improves the typing detection rate and accuracy compared to existing MassArray software results.
[0024] This invention proposes an SNR-based index and analysis process. Based on the SMR threshold of the baseline sample, the homozygosity of the locus is first determined, and then the genotype is determined according to the SNR value. Attached Figure Description
[0025] Figure 1 The detection rate of DNA mixture samples with different proportions in Reference 1 is shown in the figure; among them, Figure 1 A represents the genotyping result at the rs1979255 locus. Bold letters indicate that the software successfully identified alleles, while unbold letters indicate that alleles were not identified. Figure 1 B represents the detection rate of 29 SNP sites.
[0026] Figure 2 The flowchart of the MassArray drug genomic typing analysis of this invention.
[0027] Figure 3 A schematic diagram and site information diagram of MassArray nucleic acid mass spectrometry, taking the rs2289669 site as an example.
[0028] Figure 4 Schematic diagram of homozygous / heterozygous peaks at rs2289669 site; the upper, middle and lower figures show the peaks of G (Mass value 4840.20) and A (Mass value 4920.10) when rs2289669 is GG homozygous, AA homozygous, and AG heterozygous, respectively.
[0029] Figure 5 A schematic diagram illustrating the influence of neighboring peaks on the target peak, taking the peak emergence of rs4149056-1 as an example.
[0030] Figure 6 Schematic diagram of SMR distribution; where, Figure 6A represents the ideal scenario for rs5219, where the SMR of the heterozygous site approaches 1 and the SMR of the homozygous site approaches 0. Figure 6 B represents the homozygous locus clustering caused by the difference in SMR characteristics between the rs2241766 homozygous locus GG and TT. Figure 6 The C-value represents the heterozygous site of rs1801252, which exhibits high and low peaks, resulting in the heterozygous site SMR distribution ranging from 0.5 to 1.
[0031] Figure 7 A schematic diagram of threshold construction based on SMR.
[0032] Figure 8 The threshold construction results for the seven sites contained in the MassArray pharmacogenomics panel used to detect the CYP2C19 gene.
[0033] Figure 9 The accuracy of the MassArray pharmacogenomics panel used to detect the CYP2C19 gene was verified using the method of this invention.
[0034] Figure 10 The MassArray pharmacogenomics panel used for detecting the CYP2C19 gene was validated using the anti-interference substance of the method of this invention.
[0035] Figure 11 The sensitivity verification results of the MassArray pharmacogenomics panel used to detect the CYP2C19 gene using the method of this invention.
[0036] Figure 12 , 5 The detection rate and positive concordance rate of the Typer method and the standard in a drug genome panel.
[0037] Figure 13 Schematic diagram of result verification for inconsistency at rs7412 site - based on manual verification.
[0038] Figure 14 Schematic diagram of inconsistent results verification and re-spectral analysis; among which, Figure 14 A is the original mass spectrum peak of EDT55329157B1D1 rs4646994 (ACE_INDEL) which indicates failure in this method; Figure 14 B is rs4646994 (ACE_INDEL). This method indicates a failed mass spectrum after re-interpreting the spectrum.
[0039] Figure 15 Inconsistent result verification - diagram of first-generation sequencing; among which, Figure 15A is the original mass spectrum peak of EDT27338880B1D1 with rs2289669 indicating failure; Figure 15 B represents the first-generation sequencing result of rs2289669 for this sample, showing that the result for this locus is AG.
[0040] Figure 16 The statistical results of sites that were not detected by the method of this invention but could be detected by the Typer method were reviewed. Detailed Implementation
[0041] The embodiments of the present invention will be described in detail below with reference to examples. However, those skilled in the art will understand that the following examples are for illustrative purposes only and should not be considered as limiting the scope of the invention. Unless otherwise specified in the examples, conventional conditions or conditions recommended by the manufacturer are followed. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased on the market.
[0042] Definitions of some terms Unless otherwise defined below, all technical and scientific terms used in the specific embodiments of this invention are intended to have the same meaning as commonly understood by those skilled in the art. While it is believed that the following terms will be well understood by those skilled in the art, the following definitions are set forth to better explain the invention.
[0043] As used in this invention, the indefinite or definite articles used when referring to singular nouns, such as “a” or “a kind”, “the”, include the plural form of the noun.
[0044] As used in this invention, the terms “comprising,” “including,” “having,” “containing,” or “involving” are inclusive or open-ended and do not exclude other unlisted elements or method steps. The term “consisting of” is considered a preferred embodiment of the term “comprising.” If a group is defined below as comprising at least a certain number of embodiments, this should also be understood to disclose a group that preferably consists only of those embodiments.
[0045] The term "approximately" in this invention refers to an accuracy range that, as would be understood by those skilled in the art, still guarantees the technical effects of the features in question. This term typically indicates a deviation from the indicated value of ±10%, preferably ±5%.
[0046] Furthermore, the terms first, second, third, (a), (b), (c), and similar terms used in the specification and claims are for distinguishing similar elements and are not necessary for the order of description or chronological sequence. It should be understood that such terms are interchangeable in appropriate contexts, and the embodiments described in this invention can be implemented in a different order than that described or illustrated in this invention.
[0047] The terms and definitions above are provided merely to aid in understanding the invention. These definitions should not be construed as having a scope less than that understood by those skilled in the art.
[0048] It is understood that, based on the core method of the present invention, the present invention may also include models, apparatuses, devices, or storage media prepared according to this method. Therefore, according to some aspects of the present invention, the present invention also provides an electronic device comprising: a processor and a memory; the processor and the memory are connected, wherein the memory is used to store a computer program, and the processor is used to invoke the computer program to perform the method described in any of the preceding claims. According to other aspects of the present invention, the present invention also provides a computer storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, perform the method described in any of the preceding claims.
[0049] The present invention will now be described in conjunction with specific embodiments.
[0050] Experimental Example
[0051] This invention, through research, has established a set of... Figure 2 The MassArray pharmacogenomics typing analysis method shown below includes the following specific steps: Step 1: SNP site SMR threshold construction 1) Baseline sample selection: First, select 20-50 samples that cover all target loci and are known to be homozygous and heterozygous, with a homozygous to heterozygous ratio of 9-1:1-9.
[0052] 2) Calculate the SMR value for each locus. Considering that pharmacogenomics typically detects the genetic information of a candidate at a pharmacogenomics-related locus, i.e., the germline genotype of the detected locus, in this application context, the detected locus only exists in two states: homozygous and heterozygous. Taking rs2289669 as an example, when the locus is homozygous (AA or GG), the single-base extension produces only one product (A or G), and only one genotype (A or G) shows a significant nucleic acid mass spectrometry peak. In this case, the SNR value corresponding to this base is also significantly higher than other genotypes. When the locus is heterozygous, the single-base extension should produce two products, and the corresponding two genotypes show significant nucleic acid mass spectrometry peaks (A and G). In this case, the SNR corresponding to these two genotypes is higher than other genotypes. Based on this, determining the genotype of a locus only requires focusing on the two most prominent mass spectrometry peaks (i.e., the genotypes corresponding to the largest and second-largest SNRs). See details. Figure 3 and Figure 4 ( Figure 3This is a site information diagram for the rs2289669 site. The Mass value is the mass of the extension product. When the Mass is 4593.00, it corresponds to the unextended primer (UEP) of rs2289669. When the Mass is 4840.20, it corresponds to the product of rs2289669 after single-base extension G. When the Mass is 4920.10, it corresponds to the product of rs2289669 after single-base extension A. Figure 4 This is a schematic diagram of the homozygous and heterozygous peaks at the rs2289669 locus. The upper, middle, and lower figures show the peak values of G (Mass value 4840.20) and A (Mass value 4920.10) when rs2289669 is GG homozygous, AA homozygous, and AG heterozygous, respectively.
[0053] In addition, such as Figure 5 As shown in the figure (the Mass values of T and C at the rs4149056-1 site are 7570.00 and 7586.00 respectively, indicating that the masses of the two extended products are similar, and a significant T peak leads to a higher C peak), the elution of the target peak may be affected by the elution of neighboring sites, causing deviations. Significant elution of neighboring sites leads to a higher target peak. Therefore, analyzing only the SNR value of the target peak cannot avoid the influence of neighboring peaks. This invention proposes a new index, SMR (Second-to-Max Ratio), which calculates the SMR of the baseline sample to confirm the reasonable range of the relative amounts of the maximum and second-maximum peaks in the homozygous and heterozygous states under the current system, further defining the threshold for determining homozygous and heterozygous states. Since this index is a threshold constructed under the current system, it already includes the influence of neighboring peaks, avoiding errors caused by directly using a single SNR for interpretation.
[0054] The SMR value for each locus is calculated as follows: The MassArray platform is used to analyze the aforementioned baseline samples, obtaining offline data for each locus, including the SNR value. The SNR values are sorted in descending order, and SMR (Second-to-Max Ratio) is defined as: SMR = Second_SNR (second-largest SNR) / Max_SNR (maximum SNR). The SMR range is 0-1. When the SMR is closer to 0, the locus is more homozygous; when the SMR is closer to 1, the locus is more heterozygous (see details...). Figure 6 ).
[0055] 3) Construct the SMR threshold for each site Ideally, the SMR of a heterozygous site should be equal to 1, meaning the SNRs of the two heterozygous peaks are equal; the SMR of a homozygous site should be equal to 0, meaning only the homozygous peak has an SNR value, and the SNRs of the other peaks are 0. The SMR distribution of these sites should be as follows: Figure 6 The rs5219 site is shown in A. Under non-ideal conditions, SMR may occur, including but not limited to... Figure 6 B and Figure 6 In case C, Figure 6 B represents the homozygous locus clustering caused by the difference in SMR characteristics between GG and TT, which are homozygous loci of rs2241766. Figure 6 The heterozygous locus C, representing rs1801252, exhibits a high-low peak distribution, resulting in a SMR (Signal Minimum Ratio) for heterozygous loci ranging from 0.5 to 1. Since an SMR closer to 0 indicates a locus closer to ideal homozygous characteristics, and an SMR closer to 1 indicates a locus closer to ideal heterozygous characteristics, this method uses a lower limit for heterozygosity and an upper limit for homozygosity, further narrowing the region between these two limits to solve the binary classification problem of distinguishing between homozygous and heterozygous loci.
[0056] The specific SMR threshold for each locus is constructed as follows: For each SNP locus, the baseline samples are divided into homozygous and heterozygous groups based on the known homozygous heterozygous results. The dataset consisting of SMR within the quartile interval of the homozygous locus is defined as G1, and the dataset consisting of SMR within the quartile interval of the heterozygous locus is defined as G2. The upper limit of homozygosity and the lower limit of heterozygosity are obtained based on the median values of G1 and G2 and the 3-times MAD interval (see details). Figure 7 Wherein, the upper limit of homozygosity = median of G1 + 3 * MAD of G1; the lower limit of heterozygosity = median of G2 - 3 * MAD of G2; further, the interval between the lower limit of heterozygosity and the upper limit of homozygosity is divided into 4 equal parts, and the upper quartile (Q1) is selected as the threshold for the heterozygous genotype of the target locus, and the lower quartile (Q3) is selected as the threshold for the homozygous genotype of the target locus (corresponding to...). Figure 7 (Grayscale area).
[0057] Step 2: Genotyping based on site SMR thresholds 1) Sample detection: Mass spectrometry detection of the sample to be tested is performed using the MassArray platform to obtain offline data including SNR values.
[0058] 2) Quality control: Quality control is performed based on the SNR value obtained in step 2; when SNR < 3, SMR is not calculated and the sample to be tested is retested.
[0059] 3) Hybrid heterozygosity analysis: Calculate the SMR of each locus for the quality-controlled reliable data, and determine the homozygosity of each locus based on the SMR threshold constructed in step 1; The interpretation criteria are as follows: when SMR ≤ homozygous threshold, the site is determined as homozygous; when SMR ≥ heterozygous threshold, the site is determined as heterozygous; when heterozygous threshold < SMR < homozygous threshold, the site cannot be determined.
[0060] 6) Genotype interpretation: after the homozygosity / heterozygosity determination of the site, the genotype of the site is interpreted in combination with the genotypes corresponding to the maximum SNR and the second maximum SNR; The interpretation is as follows: if the site is homozygous, the genotyping of the site is homozygous for the genotype corresponding to the maximum SNR; if the site is heterozygous, the genotyping of the site is heterozygous for the genotypes corresponding to the maximum SNR and the second maximum SNR; for example, if the site is homozygous and the genotype corresponding to the maximum SNR is T, the genotype of the site is TT; if the site is heterozygous and the maximum SNR and the second maximum SNR correspond to A and G respectively, the genotype of the site is AG.
[0061] Example 1: Evaluation of detection performance of CYP2C19 genotyping based on SMR index
[0062] In this example, a MassArray pharmacogenomics Panel for detecting CYP2C19 gene is tested, and the Panel contains a total of 7 sites: rs12248560, rs12769205, rs28399504, rs3758581, rs4244285, rs4986893, rs72552267.
[0063] Specifically, the method of the present invention is used to construct a threshold with 20 baseline samples (for the detailed construction method, refer to Step 1 of the Experimental Example), and the SMR threshold results obtained from 5 sites are as shown in Figure 8 . Subsequently, the analysis flow of the present invention (Step 2 of the Experimental Example, including quality control, homozygosity / heterozygosity analysis, and genotype interpretation) was used to evaluate and verify the detection accuracy, specificity and sensitivity of the Panel.
[0064] 1) In terms of accuracy: in this example, 19 clinical samples (gDNA samples extracted from EDTA whole blood) of wild / mutant genotypes at all 7 sites of the Panel, 1 positive reference substance and 1 standard substance NA12878 were used for verification. The detection results of all samples were consistent between the MassArray platform and the control Sanger sequencing, and the negative coincidence rate (wild type), positive coincidence rate (mutant type) and total coincidence rate were all 100%, as shown in Figure 9 .
[0065] 2) In terms of specificity: to evaluate the influence of common interfering substances, this example selected 2 representative clinical samples, which were EDTA-anticoagulated whole blood samples, and added triglyceride (37mM) dissolved in isopropanol for interference test. Each sample was repeatedly detected 3 times, and the genotype detection results refer to Figure 10As can be seen, the results are completely consistent with the original results without any interfering substances, indicating that the method and analytical process of this invention have strong anti-interference capabilities.
[0066] 3) Limit of Detection / Sensitivity: In this embodiment, gDNA samples extracted from one EDTA-treated whole blood sample were analyzed five times at each of three dosage gradients (10 ng, 20 ng, and 30 ng). The lowest gradient at which all five replicates were able to detect the DNA was defined as the limit of detection for this test. See the appendix for results. Figure 11 As can be seen, the gDNA sample (EDT81303689B1D1) extracted from whole blood with EDTA was consistently detected in 5 replicates across 3 gradients (10 ng / reaction, 20 ng / reaction, and 30 ng / reaction) and was consistent with the results of the first-generation validation. The lowest detection limit of this invention can reach 10 ng / reaction.
[0067] The above results demonstrate that the analytical workflow constructed in this invention performs perfectly in terms of accuracy, specificity, and sensitivity (all concordance rates / compliance rates are 100%), fully proving that the workflow of this invention has extremely high reliability and robustness in clinical testing.
[0068] Example 2: Comparison with the Typer analysis method, a software companion to MassARRAY
[0069] This embodiment utilizes five pharmacogenomics panels: Panel_E (for personalized medication detection of mercaptopurine drugs), Panel_F (for personalized medication detection of rheumatic and immune diseases), Panel_J (for personalized medication detection of neurological and psychiatric diseases), Panel_X (for personalized medication detection of cardiovascular and cerebrovascular diseases), and Panel_Y (for folic acid metabolism gene detection). A total of 3247 samples (including gDNA samples extracted from EDTA whole blood, gDNA samples extracted from oral swabs, and human genomic DNA standard NA12878) were analyzed using the MassArray platform. Based on the threshold established from baseline samples, the analytical method of this invention was used to perform genotyping analysis on loci within the panels, and the results were compared with those from the MassArray manufacturer's accompanying software, Typer, to verify the unique advantages of this invention's method in improving locus detection rate and accuracy.
[0070] The 3,247 samples tested contained 126,606 loci, spanning 269 batches across 5 panels. The detection rate results are as follows: Figure 12 As shown, this method failed to detect only 371 sites, achieving a detection rate of 99.71%, which is significantly higher than the 95.31% detection rate of the Typer software (p<0.001).
[0071] Regarding accuracy, the 3247 samples in this embodiment included 267 human genomic DNA standards—NA12878 (NA12878 is a widely used genomics reference standard, a core sample of the international "HapMap Project" and "1000 Genomes Project," and its sequence information has been repeatedly tested and analyzed by multiple projects and technology platforms worldwide). The site detection accuracy of this method in this standard was significantly higher than that of Typer software (99.95% vs 99.83%, p=0.01539<0.05) (see [link to relevant documentation]). Figure 12 ).
[0072] Although this method outperforms the Typer method in both detection rate and accuracy, considering the extended turnaround time caused by manual verification of undetected sites in practical applications, and the potential for different medication guidance from pharmacogenomics experts, the applicant further verified 371 sites that were not detected by this method in this embodiment. Of these, 206 sites were not detected by either this method or Typer, and 165 sites were not detected by this method but for which Typer provided genotype results. These 165 sites were verified using methods including manual interpretation of mass spectrometry peak diagrams, re-spectral analysis, and first-generation sequencing. The verification results for some sites are as follows: See the results of manual interpretation of the mass spectrometry peaks. Figure 13 The analysis software indicated a failure at the rs7412 locus of ORA24092079K1D1, and Typer gave a genotype result of TT. Figure 13 The original peak diagram of this sample shows that both C and T bases are clearly visible in the manually verified peak diagram, with SNR values of 5.57667 and 24.4512, respectively. Referring to the instruction manual of the domestically registered product "Human CYP2C19 Genotyping Detection Kit (Time-of-Flight Mass Spectrometry)", an SNR > 5 is considered a peak, so the genotype of this locus should be interpreted as CT.
[0073] See the results of the re-spectral analysis. Figure 14 The analysis software indicated that the method of this invention failed at the rs4646994 site of EDT55329157B1D1. Figure 14 A and Figure 14 B represents the original mass spectrum peaks from the failed rs4646994 (ACE_INDEL) test and the mass spectrum peaks from the re-amplified mass spectrum. The original INS (Mass value 8635.60) had an SNR of 46.3147, and WT (Mass value 8143.30) had an SNR of 1.12851. Typer identified the genotype as homozygous for INS (INS.INS). After re-amplification, the SNRs for INS and WT were 39.1033 and 13.1476, respectively, and were manually interpreted as heterozygous (INS.WT).
[0074] See the verification results of the first-generation sequencing. Figure 15 The analysis software indicated that the method of this invention failed at the rs2289669 site of EDT27338880B1D1; Figure 15 A is the original mass spectrum peak diagram of the sample rs2289669 indicating failure. At this time, A (Mass value 4920.10) is significantly elliptical, with an SNR value of 16.3622. The genotype given by Typer is AA. Figure 15 B represents the first-generation sequencing result of rs2289669 for this sample, indicating that the result for this locus should be interpreted as AG.
[0075] Summary of review results as follows Figure 16 As shown, among the 99 verifiable sites, 13 sites were inconsistent with the verification results by Typer, with a verification consistency rate of 86.87% (86 / 99). However, 10 of the inconsistent sites were detected as heterozygous by Typer but were misclassified as homozygous by the verification. This is consistent with the results reported in Reference 1 above, which showed that Typer has relatively weak heterozygous detection performance. This reflects that the method of the present invention has a relatively low heterozygous misclassification and the results are more reliable.
[0076] In summary, the method of this invention demonstrates higher detection rates and accuracy in multiple MassARRAY pharmacogenomics assays. In large-scale clinical applications, it significantly reduces sample turnaround time caused by manual verification. Furthermore, building upon the already high accuracy (99.83%) achieved by Typer, this method further improves accuracy by 0.12 percentage points, particularly enhancing reliability in heterozygous detection, thus providing more robust genotyping support for personalized medicine.
[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A bioinformatics analysis method for MassArray pharmacogenomic genotyping, characterized in that, Includes the following steps: Step 1: Construct the SMR threshold for the proportion of the second largest SNP value; Step 2: Genotyping based on SMR thresholds; Step 1 includes: 1) Baseline sample selection: Select baseline samples that cover all target loci and whose homozygosity and heterozygosity are known; 2) Calculate the SMR value for each locus: Use the MassArray platform to detect the baseline sample and obtain the offline data including the SNR value for each target locus; sort the SNR values in descending order and define the SMR, where SMR = Second_SNR / Max_SNR; 3) Construct SMR thresholds for each locus: For each locus, the baseline samples are divided into homozygous and heterozygous groups based on the known homozygous results. The dataset consisting of SMRs in the upper quartile interval of the homozygous locus is defined as G1, and the dataset consisting of SMRs in the lower quartile interval of the heterozygous locus is defined as G2. The upper limit of homozygosity and the lower limit of heterozygosity are obtained based on the median values of G1 and G2 and the 3-fold MAD interval. The interval between the lower limit of heterozygosity and the upper limit of homozygosity is divided into 4 equal parts. The upper quartile Q1 is selected as the heterozygous genotype threshold for the target locus, and the lower quartile Q3 is selected as the homozygous genotype threshold for the target locus. Step 2 includes: a) Detection of the test sample: Mass spectrometry detection of the test sample is performed using the MassArray platform to obtain offline data including SNR value; b) Quality control: Quality control of the obtained SNR values of the test samples; c) Homozygous analysis: Calculate the SMR of each site in the test sample after quality control, and compare it with the threshold constructed in step 1 to determine the homozygous status of the test site; d) Genotype interpretation: After determining homozygous loci, genotype interpretation of loci is performed by combining the largest and second largest SNR.
2. The bioinformatics analysis method according to claim 1, characterized in that, The SMR range is 0-1; the closer the SMR is to 0, the more homozygous the site is; the closer the SMR is to 1, the more heterozygous the site is.
3. The bioinformatics analysis method according to claim 1, characterized in that, The upper limit of homozygosity = median value of G1 + 3 * MAD of G1; the lower limit of heterozygosity = median value of G2 - 3 * MAD of G2.
4. The bioinformatics analysis method according to claim 1, characterized in that, In the quality control process, when the SNR value is <3, the SMR is not calculated, and the sample to be tested is retested.
5. The bioinformatics analysis method according to claim 1, characterized in that, The homozygous interpretation is as follows: if the SMR of the test site is less than or equal to the homozygous genotype threshold, the site is determined to be homozygous; if the SMR of the test site is greater than or equal to the heterozygous genotype threshold, the site is determined to be heterozygous.
6. The bioinformatics analysis method according to claim 1, characterized in that, The genotype interpretation is as follows: if the locus is homozygous, the locus genotype is homozygous for the genotype corresponding to the largest SNR; if the locus is heterozygous, the locus genotype is heterozygous for the genotypes corresponding to the largest SNR and the second largest SNR.
7. An electronic device, characterized in that, include: Processor and memory; The processor is connected to a memory, wherein the memory is used to store a computer program, and the processor is used to invoke the computer program to execute the bioinformatics analysis method as described in any one of claims 1-6.
8. A computer storage medium, characterized in that, The computer storage medium stores a computer program, the computer program including program instructions, which, when executed by a processor, perform the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Gene detection method for rheumatism immune disease medication based on nucleic acid mass spectrum and application of gene detection method
CN112941182A
Integrated genome analysis method based on genotype filling and low-depth sequencing
CN115798580A