Sample homology determination model, method for establishing same and application thereof

By establishing a sample homology determination model and utilizing SNV data and SVM classification, the problem of sample homology determination in high-throughput pathogen metagenomic detection was solved, achieving rapid and accurate homology determination at low sequencing depth and reducing costs.

CN114944188BActive Publication Date: 2026-02-06GZ VISION GENE TECH CO LTD +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210543729.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-19
Publication Date
2026-02-06
Estimated Expiration
2042-05-19

AI Technical Summary

Technical Problem

In high-throughput metagenomic detection of pathogens, how can we accurately determine sample homology when the sample sequencing depth is low, and avoid increasing experimental steps and prolonging the report issuance time?

Method used

A homology determination model for samples was established. Using raw sequencing data, a classification model based on mismatch rate was constructed through SNV data collection, model building, and parameter optimization. This included sample collection, SNV data collection, target sequence region selection, SNV filtering conditions, and genotype differences. Support vector machine (SVM) was used for determination.

Benefits of technology

This technology enables accurate determination of sample homology at low sequencing depths, improving the accuracy and efficiency of the determination results, reducing costs, and avoiding the need for additional experimental steps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114944188B_ABST
    Figure CN114944188B_ABST
Patent Text Reader

Abstract

The present application relates to a sample homology determination model and its establishment method and application, belong to gene detection technical field. The method comprises the following steps: sample collection: two samples of the same source are used as positive sample pair to form a positive sample set, and two samples of different sources are used as negative sample pair to form a negative sample set; SNV data collection: the sequencing data is compared to the human reference genome to obtain the single nucleotide variation site SNV condition of each sample, any sample pair is selected, the ratio of the number of typing inconsistent sites and the number of common detection sites is the mismatch rate; Model construction: using sample sequencing data volume, target sequence region, SNV filtering condition and genotype difference as model parameter condition, using mismatch rate as determination index, according to the gradient range of sample common detection site quantity corresponding to matched mismatch rate, the classification model is constructed. The model can be used in the scene with low sample sequencing depth, complete the homology determination of sample, has the advantages of low cost and high efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of gene detection, in particular to a sample homology determination model and a method for establishing the same and application thereof. BACKGROUND

[0002] In recent years, with the popularization of high-throughput sequencing data and various application research and development, pathogen metagenomics detection emerges as a brand-new microbial detection method in major hospitals. However, the whole process of pathogen metagenomics is complex, including nucleic acid extraction, host removal, construction of high-throughput sequencing library, and finally sequencing to bioinformatics analysis. In the whole experimental process, there are corresponding quality control indicators to monitor whether the experimental process is problematic. However, it is inevitable that human operation errors will occur, such as sample label filling errors, confusion of suction solutions, etc. And when a patient sends one or more samples for detection, in order to exclude the case that human operation or other factors cause large differences in detected bacterial spectrum, it is necessary to determine whether two or more samples come from the same patient.

[0003] The problem brought by this is how to distinguish the homology (the same patient) of the sample under the condition of high-throughput pathogen metagenomics ultra-low sequencing depth.

[0004] In the conventional technology, STR is commonly used in the field of judicial expertise to distinguish different samples, and in population genetics research, a combination of high polymorphic single nucleotide sites is commonly used to distinguish different samples. Due to the high timeliness requirement and short sequencing read length of metagenomic sequencing, it is not possible to meet the conventional STR and SNP typing. Capillary electrophoresis is also commonly used in forensic laboratories to perform STR typing on samples, which requires increasing the number of experiments and the amount of samples used. Due to the particularity of pathogen metagenomic samples, it is generally difficult to obtain or obtain a small amount, such as cerebrospinal fluid or alveolar fluid. At the same time, considering the timeliness of pathogen metagenomic samples, the timeliness requirement of the whole process from receiving the sample to issuing the report is very high. Therefore, increasing the analysis steps cannot delay the report issuing time, so the analysis steps for determining homology cannot be increased too much. SUMMARY

[0005] Therefore, it is necessary to provide a method for establishing a sample homology determination model to solve the above problems. The sample homology determination model established by the method can be used in a scenario with low sequencing depth of samples (such as metagenomic detection), without the need for additional experiments, and the original data can be used to determine whether the samples come from the same patient.

[0006] A method for establishing a sample homology determination model, comprising the following steps:

[0007] Sample collection: two samples of the same source are used as a positive sample pair, and two samples of different sources are used as a negative sample pair. A number of positive sample pairs are collected to form a positive sample set, and a number of negative sample pairs are collected to form a negative sample set.

[0008] SNV data collection: the sequencing data obtained by sequencing the above samples based on the same sequencing method is aligned to the human reference genome to obtain the alignment of each sample sequence in the human genome and the single nucleotide variation (SNV) of each sample. Arbitrarily select a sample pair, and calculate the ratio of the number of typing inconsistent sites to the number of common detection sites, which is called the mismatch rate.

[0009] Model construction: using sample sequencing data volume, target sequence region, SNV filtering condition and genotype difference as model parameter conditions, using mismatch rate as a judgment index, and according to the gradient range of the number of common detection sites of the sample, the corresponding mismatch rate is matched to construct a classification model, which is a sample homology determination model.

[0010] The above model building method includes all key parameters from high-throughput sequencing data to SNV generation in the training model, which can adapt to data of different application directions, find key parameters in each set, and obtain low-noise mismatch rate as the input of the determination model, thereby improving the accuracy of the model.

[0011] In one embodiment, the target sequence region is determined by the following method: obtaining a polymorphic site set of the same race as the sample source from the database, and taking the exon site with a polymorphic site percentage of 30%-70% and / or the genomic site with a polymorphic site percentage of 10%-90% as the target sequence region. It can be understood that due to the difference in the distribution of polymorphic sites of different races, the corresponding race data should be used for analysis to improve the performance of the model.

[0012] In one embodiment, when the sample is derived from a critically ill patient, the target sequence region further includes a mitochondrial sequence region. Since the mitochondria of critically ill patients are abnormally active, a large amount of mitochondrial data will appear, so the mitochondrial sequence is also added to the model construction to improve the accuracy of the model determination.

[0013] In one of the embodiments, the SNV filtering condition is that: filtering removes sites with sequencing depth of 3x or less, and filtering removes sites with sequencing quality of 15 or less. For metagenomic detection, the metagenomic data volume is generally 20M sequence numbers, read length 50bp, and the average depth of coverage to the human genome is 0.3X. In whole exome (WES) for analyzing SNV / SNP method is basically 60-100X, and in whole genome (WGS) application is 10-30X data volume. Therefore, the filtering SNV condition is limited according to the above method, which can avoid the genotype typing inaccuracy caused by ultra-low depth, and remove the situation of insufficient data availability caused by excessive data removal.

[0014] In one of the embodiments, the genotype difference includes: heterozygote and homozygote; and when the sequencing depth is less than 3x or less, the site of heterozygote typing is filtered out. Genotype is generally divided into heterozygote and homozygote. Since the genotype selection cannot distinguish germline mutations, somatic mutations and the like, the genotype difference is analyzed when the model is constructed; and when the depth is insufficient, the accuracy of the heterozygote is lower, which affects the calculation of the mismatch rate, so the quality value of the heterozygote typing can be filtered out.

[0015] In one of the embodiments, in the model construction step, the orthogonal test method of the control variable is used to control the change of one parameter under the condition that other parameters are fixed, and the sample mismatch rate under each parameter change is iteratively analyzed.

[0016] In one of the embodiments, in the model construction step, according to the mismatch rate value under the parameter change condition, a support vector machine model SVM binary classification model is used to construct the classification model. It can be understood that after the parameter type is determined, the specific modeling method can be constructed according to the conventional SVM model.

[0017] In one of the embodiments, after the model construction step, the method further includes the following model optimization step: another sample set is collected to form a verification sample set, the sample homology determination model is verified by using the verification sample set, and the model is optimized according to the verification result.

[0018] The application also discloses a homology determination model established by the method for establishing the sample homology determination model.

[0019] The application also discloses an application of the homology determination model in sample homology determination in metagenomic detection.

[0020] It can be understood that the above homology determination model can be used in all types of clinical sample scenes and the like requiring homology determination, and when used in metagenomic detection, the homology determination can be realized without additional detection in the case of low metagenomic sequencing depth and small data volume.

[0021] Compared with the prior art, the present application has the following beneficial effects:

[0022] The sample homology determination model obtained by the sample homology determination model establishment method of the present application can be used in a scene with low sample sequencing depth (such as metagenomic detection), without the need for additional experiments, and the homology determination of whether the sample is from the same patient can be completed using original data, which has the advantages of accurate determination results, low cost and high efficiency and speed. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 The figure is a sample homology determination model establishment flowchart in Example 1.

[0024] Figure 2 The figure is a model PPV performance evaluation schematic diagram in Example 1.

[0025] Figure 3 The figure is a model NPV performance evaluation schematic diagram in Example 1.

[0026] Figure 4 The figure is a specific sample pair analysis determination schematic diagram in Example 2. DETAILED DESCRIPTION

[0027] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the accompanying drawings. The preferred embodiments of the present application are shown in the drawings. However, the present application can be realized in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided so that the disclosure of the present application can be more thoroughly and completely understood.

[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terminology used in the description of the present application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0029] Unless otherwise specified, the reagents used in the following examples are commercially available; unless otherwise specified, the methods used in the following examples can be realized by conventional methods.

[0030] Definition: The sequencing data obtained by sequencing based on the same sequencing method refers to the sequencing technology with the same sequencing means, such as macrogenomic sequencing data or probe capture sequencing data.

[0031] Example 1

[0032] A sample homology determination model is obtained by the following method and is subjected to performance evaluation, and a flowchart is shown as Figure 1

[0033] I. Sample collection

[0034] 1. Construct a training data set, and take two samples from the same source as a positive sample pair, such as a patient who may have multiple different sample types (such as alveolar lavage fluid, cerebrospinal fluid, nasopharyngeal swab, blood, etc.), and take two samples as a positive sample pair, collect several positive sample pairs to form a positive sample set, and in this embodiment, 240 samples of 120 patients are collected, and 120 positive sample pairs are taken as a positive sample set.

[0035] Collect samples from different sources (i.e. different patients), and in this embodiment, the above 240 samples of 120 patients are matched with different types of samples from different patients, and (120 x 119) / 2 pairs of negative sample pairs are formed to form a negative sample set, wherein (120 x 119) / 2 is the number of sample pairs formed by any sample of 120 patients and different samples of other patients.

[0036] 2. Construct a verification data set, and collect 1137 clinical samples as a verification data set for standby.

[0037] II. SNV data collection

[0038] 1. The model construction method is not the same as the machine learning model construction. The model itself is constructed for the macrogenomic data produced by the laboratory, and is used for homology determination analysis. Commonly used methods produce common data, which need to be reconstructed, such as the sequence data obtained by the same capture detection method in the capture product. In this embodiment, the above samples are sequenced based on the macrogenomic detection method to obtain sequencing data.

[0039] 2. Use the open source software fastp to remove sequencing adapter sequences and sequences with low quality values from the macrogenomic raw sequencing data.

[0040] 3. Use the open source software bwa to align the above obtained sequencing data to the human reference genome (version hg19) to obtain the complete alignment of the sample sequence in the human genome.

[0041] ​4. After removing repetitive sequences using the open-source software samtools, the sequences were compared with those processed by bcftools to analyze single nucleotide variants (SNVs) at the alignment sites, thus obtaining the SNV information for each sample.

[0042] 5. Randomly select sample pairs from the above sample sets (positive sample set and negative sample set) to find common detection sites, count the ratio of the number of inconsistent sites to the number of common detection sites, and record it as the mismatch rate.

[0043] III. Model Construction

[0044] Using sample sequencing data volume, target sequence region, SNV filtering conditions, and genotype differences as model parameters, mismatch rate as the criterion, and based on the gradient range of the number of common detection sites in samples, corresponding to the mismatch rate, i.e., gradient, multiple batches of analysis are generated to construct a classification model, which is the sample homology determination model. The specific construction method is as follows.

[0045] 1. Sequencing data volume

[0046] The amount of sequencing data determines how many common detection sites are available in the final sample pairs for subsequent judgment, and also the performance of the model's judgment. Therefore, different mismatch rates are corresponding to different gradient ranges in data volume, and models are built accordingly.

[0047] 2. Determination of the target sequence region (bed)

[0048] The selection of the detection region affects the overall analysis time and the efficiency of model judgment. For example, the target sequence region determines the model analysis time. The larger the target sequence region, the longer the analysis time. In addition, the larger the target sequence region, the greater the background signal noise, which affects the calculation of the mismatch rate.

[0049] Polymorphic loci sets from the same ethnic group (such as East Asians) were obtained from the 1000 Genomes Project and the Genomad database. The target sequence regions were whole-exome (WES) loci with a polymorphic percentage of 30%-70% and whole-genome (WGS) loci with a polymorphic percentage of 10%-90%.

[0050] In this embodiment, due to the characteristics of pathogen metagenomic samples, the mitochondria of critically ill patients are abnormally active, resulting in a large amount of mitochondrial data. Therefore, mitochondrial sequences are also included in the model construction.

[0051] 3. Determining SNV filtering conditions (filtering thresholds)

[0052] The macro genome data amount is generally 20M sequence numbers, read length 50bp, and the average depth of coverage to human genome is 0.3X. The method for analyzing SNV / SNP in WES is basically 60-100X, and the data amount in WGS application is 10-30X.

[0053] Therefore, the SNV filtering conditions in the embodiment are as follows:

[0054] Depth filtering: remove sites below 3x.

[0055] Quality value filtering: remove sites with quality value below 15.

[0056] 4. Genotype difference

[0057] After filtering SNV, heterozygotes (such as AT / AC / AG / TC / TG / CG, etc.) and homozygotes (such as AA / TT / CC / GG) are obtained. Since it is impossible to distinguish germline mutations, somatic mutations, etc. in genotype selection, homozygotes or heterozygotes are included in genotype typing to construct an analysis model.

[0058] However, in the typing method with low confidence (more common high throughput), the accuracy of heterozygotes is lower, which affects the calculation of mismatch rate. Therefore, when the depth is insufficient (for example, less than 3x), the quality value of heterozygote typing can be filtered out.

[0059] 5. Modeling method

[0060] The above different data amounts, target sequence regions (SNV site set), SNV filtering conditions, genotype differences are used as model parameters, and different common detection site quantity gradients and corresponding mismatch rate gradients are used to generate different batches of tasks, and the mismatch rate results in different tasks are obtained.

[0061] After obtaining the results in each different task, a support vector machine (SVM) is used to construct a classification model, and finally the optimal model is obtained.

[0062] The optimal model finally obtained includes the target sequence region (SNV site set, bed), SNV filtering condition, genotype difference, and the SVM determination model established by using different site data amounts corresponding to different mismatch rates to obtain the most suitable determination threshold corresponding to different site data amounts, that is, the sample homology determination model in the gradient range of different site quantities.

[0063] Four, model optimization

[0064] The sample homology determination model obtained in the above steps is verified by using the verification data set in step one.

[0065] 1. Method

[0066] The verification method is to substitute the verification data set into the trained various parameter models, obtain the mismatch rate results of the verification data set in the current model, and determine the identity according to the SVM classification model.

[0067] 2. Results

[0068] The verification results are shown in Figures 2-3 , the PPV determination performance schematic diagram is Figure 2 , and the NPV determination performance schematic diagram is Figure 3 , where the horizontal coordinate is the common detected site interval (i.e., the data size gradient, L1000: sample pairs with common detected sites below 1000; L5000: sample pairs with common detected sites between 1000 and 5000; L8000: sample pairs with common detected sites between 5000 and 8000; L10000: sample pairs with common detected sites between 8000 and 10000; L15000: sample pairs with common detected sites between 10000 and 15000; B15000: sample pairs with common detected sites above 15000), and the vertical coordinate is the AUC value determined by the sample homology determination model (SVM classification model).

[0069] From the results, it can be seen that when the number of common detected sites of the comparison sample pairs is between 5000 and 8000, the performance of the model is:

[0070] The PPV (reliability of assuming the same patient) is 99%.

[0071] The NPV (reliability of assuming the same patient) is 67.5%.

[0072] When the number of common detected sites of the comparison sample pairs is greater than 15000 sites, the performance of the model is:

[0073] The PPV (reliability of assuming the same patient) is 99.8%.

[0074] The NPV (reliability of assuming the same patient) is 94.7%.

[0075] Example 2

[0076] Randomly selected one batch of samples, a total of 22 samples, including 5 patients with 2 different samples, that is, 22 samples, 10 samples can be respectively corresponded to 5 patients. The other 12 are samples of another 12 patients, and the determination model of example 1 is used for classification determination.

[0077] The results are shown in Figure 4As shown, the results show that the model generates a total of (22*21) / 2 sample pairs of results, in which 5 patients are correctly determined as the same person by two different sample models; other sample pairs are respectively determined as non-same person by the model. That is, A1 and A2, G1 and G2, I1 and I2, J1 and J2, M1 and M2 are respectively different samples from the same person, and other sample pairs with other numbers are sample pairs of other patients. The results are consistent with the actual sample sending situation.

[0078] The technical features of the above-described embodiments can be combined in any manner. To make the description concise, all possible combinations of the technical features in the above-described embodiments are not described, but as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present disclosure.

[0079] The above-described embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be pointed out that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for establishing a sample homology determination model, characterized by, The method comprises the following steps: Sample collection: two samples of the same source are used as a positive sample pair, and two samples of different sources are used as a negative sample pair; a plurality of positive sample pairs are collected to form a positive sample set, and a plurality of negative sample pairs are collected to form a negative sample set; SNV data collection: the sequencing data obtained by sequencing the above samples based on the same sequencing method are aligned to the human reference genome to obtain the alignment of each sample sequence in the human genome and the single nucleotide variation (SNV) of each sample; any sample pair is selected, and the ratio of the number of typing inconsistent sites to the number of common detection sites is calculated and recorded as the mismatch rate; Model construction: the sample sequencing data volume, the target sequence region, the SNV filtering condition and the genotype difference are used as model parameter conditions, the mismatch rate is used as a judgment index, and a classification model is constructed according to the gradient range of the number of common detection sites of the sample, the matched mismatch rate, which is a sample homology judgment model; The target sequence region is determined by the following method: a polymorphic site set of the same race as the sample source is obtained from a database, and the exon site with a polymorphic site percentage of 30%-70% and / or the genomic site with a polymorphic site percentage of 10%-90% are used as the target sequence region; when the sample is derived from a critically ill patient, the target sequence region further comprises a mitochondrial sequence region; The SNV filtering condition is: filtering and removing sites with a sequencing depth of less than 3x, and filtering and removing sites with a sequencing quality of less than 15; The genotype difference includes: heterozygote and homozygote; when the sequencing depth is less than 3x, the heterozygote typing site is filtered and removed.

2. The method of claim 1, wherein, In the model construction step, the orthogonal test method of the control variable is used to control the change of one parameter under the condition that the other parameters are fixed, and the sample mismatch rate under the change of each parameter is obtained by iterative analysis.

3. The method of claim 2, wherein, In the model construction step, the support vector machine (SVM) binary classification model is used to construct the classification model according to the mismatch rate value under the parameter change condition.

4. The method of claim 1, wherein, After the model construction step, the following model optimization step is further included: a plurality of samples are collected to form a verification sample set, the sample homology judgment model is verified by using the verification sample set, and the model is optimized according to the verification result.

5. The sample homology judgment model established by the method of any one of claims 1-4.

6. The application of the homology judgment model of claim 5 in the homology judgment of metagenomic detection samples.

Citation Information

Patent Citations

  • Homologous recombination defect judgment method based on DNA sequencing data

    CN111462823A

  • Detection methods of homologous sequences and tandem repeat sequences in homologous sequences, computer readable medium and application

    CN111508561A