A SNV detection data processing method
Through automated and standardized SNV detection data processing methods, the problem of time-consuming and error-prone traditional data processing processes is solved, and efficient and accurate SNV chip data analysis is achieved, which is suitable for standardized analysis of large-scale samples.
Patent Information
- Application Number
- CN202510176139.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The traditional SNV chip detection data processing process relies on manual operation and non-standardized methods, which makes it time-consuming, error-prone and difficult to achieve efficient analysis of large-scale samples.
Provide a SNV detection data processing method, including data preprocessing, internal quality control, typing threshold establishment and service analysis template interpretation, to realize automated and standardized data processing.
Through automated data processing methods, the efficiency and accuracy of SNV chip data analysis are significantly improved, human errors are reduced, large-scale sample standardization analysis is realized, and scientific basis is provided for SNV gene detection.
Smart Images

Figure FDA0005334395390000011 
Figure FDA0005334395390000021
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of SNV detection data processing, and in particular relates to a SNV detection data processing method. Background Art
[0002] A single nucleotide variation (SNV) refers to a change in a base (A, C, G, T) in the DNA sequence, resulting in a genomic variation. This variation usually occurs randomly, but is sometimes caused by environmental factors. Single nucleotide variation is a common form of variation in the genome and has important biological significance. Through the continuous development and improvement of gene sequencing technology, we can better understand the impact of single nucleotide variation on human health and disease, and provide more accurate and personalized guidance for clinical treatment.
[0003] Single nucleotide site variations are usually detected by gene sequencing technology. These technologies include Sanger sequencing and next-generation sequencing (NGS). In Sanger sequencing, DNA fragments are amplified and separated into fragments of different sizes. These fragments are then placed in an electrophoresis tank and separated by electrophoresis. Finally, the DNA sequence is determined by comparing fragments of different lengths. In NGS, DNA samples are broken down into small fragments and sequenced using a high-throughput sequencer. This technology can simultaneously determine a large number of DNA sequences in a short period of time, so it is widely used in genomic research and clinical diagnosis.
[0004] The traditional SNV chip detection data processing process relies on manual operations and non-standardized methods, which is not only time-consuming and error-prone, but also difficult to achieve efficient analysis of large-scale samples. Summary of the invention
[0005] In order to solve the problems in the prior art, the present invention provides a SNV detection data processing method to achieve automated and standardized data processing, thereby improving the efficiency and accuracy of SNV chip analysis.
[0006] The present invention solves the technical problem by adopting the following technical solutions:
[0007] The present invention aims to provide a method for processing SNV detection data, comprising the following steps:
[0008] S1. Data preprocessing: Obtain the fluorescence signal value of each gene detection site and calculate the log2 ratio value of each chip site;
[0009] S2. Internal quality control: The log2 ratio data of the quality control sites on the chip and the "standard line" established by the large sample data are used as "X" and "Y" to establish a linear regression equation;
[0010] S3. Establishment of typing threshold: Use standard products to establish typing thresholds for detection sites;
[0011] S4, service analysis template interpretation: the genetic detection sites of the test sample are corrected by log2 ratio using the linear regression equation established in S2, and genotyping is performed based on the final ratio value after correction of the genetic detection sites of the test sample and the typing threshold.
[0012] Furthermore, in S1, the fluorescence signal value includes the value of each gene detection site in the two fluorescence bands after deducting the background value, which are F532-B and F635-B respectively. The log2 ratio value of each chip site is the logarithm of the value of F532-B divided by F635-B at each point with base 2.
[0013] Furthermore, in S3, the classification threshold is established as follows:
[0014] .
[0015] Furthermore, the log2 ratio correction method in S4 includes: dividing the log2 ratio values of all gene detection sites by the slope value of the linear regression equation to obtain a slope correction ratio, and then subtracting the intercept value of the linear regression equation from the slope correction ratio to obtain a final ratio.
[0016] Furthermore, the typing logic is: if the final ratio value of the test sample > High cut-off, the genotype is T1; if the final ratio value of the test sample < Low cut-off, the genotype is T3; if Low cut-off ≤ the final ratio value of the test sample ≤ High cut-off, the genotype is T2.
[0017] Compared with the prior art, the beneficial technical effects of the present invention are:
[0018] 1. The present invention is based on the raw data of human genome single nucleotide site variation chip detection, pre-processes the data by deducting the background value, and corrects the data through internal quality control points to ensure the accuracy of the results under different experimental conditions. The threshold is established through the core algorithm, and the threshold is used as a standard to convert the value into a genotype, thereby realizing efficient and accurate analysis of SNV chip data;
[0019] 2. The automated data processing method of the present invention, including the steps of data preprocessing, internal quality control, typing threshold establishment, service analysis template interpretation, etc., can significantly improve the efficiency and accuracy of SNV chip data analysis, reduce human errors, realize standardized analysis of large-scale samples, and provide a scientific basis for SNV gene detection.
[0020] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above contents of the present invention and its objectives, features and advantages more obvious and easy to understand, the specific implementation methods of the present invention are listed below. DETAILED DESCRIPTION
[0021] The technical solution of the present invention is further described in detail below in conjunction with specific embodiments. It should be understood that the following embodiments are only exemplary illustrations and explanations of the present invention and should not be construed as limiting the scope of protection of the present invention. All technologies implemented based on the above content of the present invention are included in the scope that the present invention is intended to protect.
[0022] In addition, unless otherwise specified, various raw materials, reagents, instruments and equipment used in the present invention can be purchased from the market or prepared by existing methods.
[0023] A microarray chip scanner (Bio Luxscan 10K / A) with dual lasers (wavelength 532nm, wavelength 635nm) is used to scan the chip after molecular hybridization, obtain the fluorescence signal of each sample point on the chip, and enter the data processing process after obtaining the original value of the fluorescence intensity value of the gene chip array site. In the following text, the detection site is the site where the result is finally detected, and the chip site refers to the point on the chip, which can be a detection site or a chip quality control site, and can also be used as a general term for detection sites and chip quality control sites; the chip quality control site is a gene with a known genotype set on the chip, which is used to correct the error of each experiment.
[0024] 1. Data preprocessing:
[0025] (1) Use Excel VBA macro automation programming to automatically bring the raw data in LSR format into the analysis tab to obtain the fluorescence signal value of each gene detection site, namely: foreground value (F532, F635) and background value (B532, B635). The values of the foreground value of each chip site in the two fluorescence bands minus the background value are F532-B and F635-B, respectively.
[0026] Where F532-B: the fluorescence intensity at 532nm minus the background value at this wavelength, that is, the fluorescence intensity of the sample point at 532nm after deducting the background value;
[0027] F635-B: the fluorescence intensity at 635nm minus the background value at this wavelength, that is, the fluorescence intensity of the sample point at 635nm after deducting the background value;
[0028] (2) Take the logarithm of the value of F532-B divided by F635-B at each point with base 2 (log2ratio for short), that is: . Calculate the log2 ratio value of each chip site.
[0029] 2. Internal quality control:
[0030] (1) Obtain the log2 ratio data of the quality control sites on the chip and the "standard line" established by the large sample data as "X" and "Y" respectively to establish a linear regression equation;
[0031] (2) The above is the procedure for establishing the linear regression equation for internal quality control points.
[0032] The linear regression equation is established by taking the log2 ratio data of the quality control site and the "standard line" data established by the large sample data as "X" and "Y" respectively, and using the slope and intercept of this linear regression equation to correct the log2 ratio value of the SNV detection site of the test sample.
[0033] 3. Establishment of typing threshold:
[0034] (1) Use the Coriell standard to establish the typing threshold (cut-off value) of the detection site, and calculate the accuracy rate of matching by comparing with known types.
[0035] The method for establishing the classification threshold (cut-off value) is shown in Table 1:
[0036] Table 1
[0037] ,
[0038] The corresponding conditions of T1, T2, and T3 in the above Table 1 are as follows: Table 2, Table 3, Table 4-1, and Table 4-2:
[0039] Table 2
[0040] ,
[0041] Table 3
[0042] ,
[0043] Table 4-1
[0044] ,
[0045] Table 4-2
[0046] ,
[0047] (2) Calculate the cut-off threshold using the algorithm in Table 1 above: Classify the corrected log2 ratio values of each detection site of each sample into the specified type T1, T2, and T3 fields through the standard genotype of the coriell standard product, and calculate the mean value and standard deviation. Establish the cut-off threshold.
[0048] 4. Service analysis template interpretation: Perform log2 ratio correction on the LSR file of the test sample, and perform genotype typing based on whether the final ratio value of the test sample after correction is greater than the High cut-off or less than the Low cut-off.
[0049] Automated programming divides the log2 ratio values of all test sample SNV detection sites by the slope value of the linear regression equation to obtain the "slope-corrected ratio", and then subtracts the intercept value of the linear regression equation from the "slope-corrected ratio" to obtain the "final ratio"; correction is performed through the slope and intercept of the linear regression equation to eliminate the deviation caused by different red and green fluorescence intensities due to uneven hybridization or different focal lengths during scanner scanning;
[0050] This "final ratio" is automatically summarized, and the median value is taken for each corrected final ratio value of all test samples. This median value is the "standard line" of the log2 ratio value of the SNV detection site.
[0051] Genotyping logic: If the final ratio value of the test sample > High cut-off, the genotype is T1; if the final ratio value of the test sample < Low cut-off, the genotype is T3; if Low cut-off ≤ the final ratio value of the test sample ≤ High cut-off, the genotype is T2.
[0052] Through the present invention, the image file can be directly converted into a data file for corresponding processing and analysis, and then the genotype detection type can be accurately and quickly determined. After establishing the cut-off threshold through 5 coriell standard products (samples with known standard genotypes), three samples (with known genotypes) were tested, and it was confirmed that the correct rate of the sample genotypes determined by this method reached 100%, as shown in Tables 5, 6, 7, and 8 below.
[0053] Table 5
[0054] ,
[0055] Table 6
[0056] ,
[0057] Table 7
[0058] ,
[0059] Table 8
[0060] .
[0061] Tables 5 and 6 above are Coriell standards (samples with known standard genotypes); Tables 7 and 8 are test samples (known genotypes). In Tables 7 and 8, the test results of the three samples (genotype results of the test samples) are compared with the known genotypes to determine the accuracy of the algorithm of the present invention. The final ratio value after correction of the test sample is the value after the original data of the test sample is calculated and corrected according to the algorithm of the present invention. The genotype result can be obtained by combining the threshold value of Table 1.
[0062] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0063] The embodiments of the present invention are described above, but the present invention is not limited to the above-mentioned specific implementation modes. The above-mentioned specific implementation modes are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which are all within the protection of the present invention.
Claims
1. A method for processing SNV detection data, characterized in that: It includes the following steps: S1. Data preprocessing: Obtain the fluorescence signal values of each gene detection site, and calculate the log2ratio value of each chip site; S2. Internal quality control: Establish a linear regression equation with the log2 ratio data of the quality control sites on the obtained chip and the "standard line" established from the large sample data as "X" and "Y" respectively; S3. Cut-off value establishment: Establish the cut-off value of the detection site with the standard product; S4. Service analysis template interpretation: Correct the log2ratio of the gene detection sites of the test sample with the linear regression equation established in S2, and perform genotype typing based on the final ratio value of the gene detection sites of the test sample after correction and the cut-off value; In S1, the fluorescence signal value includes the value F532 - B obtained by subtracting the background value from the foreground value of each gene detection site in the 532nm fluorescence band, and the value F635 - B obtained by subtracting the background value from the foreground value of each gene detection site in the 635nm fluorescence band. The log2 ratio value of each chip site is the logarithm to the base 2 of the value obtained by dividing the F532 - B of each point by F635 - B; In S3, the cut-off value establishment method is: The log2 ratio correction method in S4 includes: Divide the log2 ratio values of all gene detection sites by the slope value of the linear regression equation to obtain the slope-corrected ratio, and then subtract the intercept value of the linear regression equation from the slope-corrected ratio to obtain the final ratio; Genotyping logic: If the final ratio value of the test sample > High cut-off, the genotype is T1; if the final ratio value of the test sample < Low cut-off, the genotype is T3; if Low cut-off ≤ the final ratio value of the test sample ≤ High cut-off, the genotype is T2.
Citation Information
Patent Citations
Method for genotyping forest populations on basis of gene CNV (copy number variation) sites
CN106480221A
Whole-chromosome genotyping chip for synchronously detecting multiple birth defect genetic diseases and method and application of whole-chromosome genotyping chip
CN114196736A