Computer equipment, sequencing data processing method and storage medium
Through comprehensive analysis of sequencing data of cfDNA and white blood cell DNA using multiple mutation detection strategies, the false positive and false negative problems of MRD detection in existing technologies were solved, and efficient and accurate detection of minimal residual lesions was achieved.
Patent Information
- Application Number
- CN202510612706.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies have difficulty in effectively identifying and detecting low-frequency and rare mutations in minimal residual disease (MRD), especially when there is background noise and interference from non-ctDNA factors in cfDNA, resulting in false positive or false negative results.
Multiple mutation detection strategies are used to comprehensively detect the sequencing data of cfDNA and white blood cell DNA. By merging the results of each strategy and combining them with a reference sequence set, background noise and interference from white blood cell DNA are eliminated, thereby improving detection accuracy.
It achieves comprehensive and accurate detection of various mutation phenomena in target samples, avoids false positive or false negative results caused by a single strategy, and improves the accuracy and reliability of detection.
Smart Images

Figure CN120748485A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a computer device, a method for processing sequencing data, and a storage medium. Background Art
[0002] Minimal residual disease (MRD) refers to a small amount of cancer cells remaining in a patient's body after cancer treatment. These cancer cells often cause mutations in the intracellular DNA. Currently, MRD detection is primarily based on high-depth sequencing analysis of the patient's circulating free DNA (cfDNA). This method identifies, at the base level, whether the cfDNA carries circulating tumor DNA (ctDNA), thereby assisting in confirming the presence of MRD. However, due to the extremely low content of ctDNA in cfDNA, avoiding interference from background noise and other non-ctDNA factors, while simultaneously identifying low-frequency and rare mutations belonging to ctDNA among all mutations in cfDNA, is a major issue that needs to be addressed. Summary of the Invention
[0003] The embodiments of the present application disclose a computer device, a method for processing sequencing data, and a storage medium. These can improve the accuracy of mutation detection in target samples by integrating the mutation detection results corresponding to multiple mutation detection strategies, thereby avoiding false positive or false negative results that may be caused by a single mutation detection strategy.
[0004] A first aspect of an embodiment of the present application discloses a computer device, including a memory and a processor, wherein the memory is used to store a computer program. When the computer program is executed by the processor, the processor performs the following steps:
[0005] Obtaining first sequencing data corresponding to circulating free deoxyribonucleic acid (cfDNA) and leukocyte DNA of the target sample;
[0006] Detecting the first sequencing data corresponding to the cfDNA and the white blood cell DNA, respectively, according to multiple mutation detection strategies, to obtain mutation detection results corresponding to the cfDNA and each of the mutation detection strategies, and mutation detection results corresponding to the white blood cell DNA and each of the mutation detection strategies;
[0007] Combining the cfDNA with the mutation detection results corresponding to each of the mutation detection strategies to obtain a first detection result corresponding to the cfDNA;
[0008] combining the white blood cell DNA with the mutation detection results corresponding to each of the mutation detection strategies to obtain a second detection result corresponding to the white blood cell DNA;
[0009] Determine a target detection result corresponding to the target sample based on the first detection result, the second detection result, and a preset reference sequence set.
[0010] In some possible embodiments, obtaining first sequencing data corresponding to circulating free deoxyribonucleic acid (cfDNA) and leukocyte DNA of the target sample includes:
[0011] Obtaining second sequencing data corresponding to the cfDNA and leukocyte DNA corresponding to the target sample;
[0012] Extracting the unique molecular marker UMI sequences corresponding to the cfDNA and leukocyte DNA respectively;
[0013] The second sequencing data corresponding to the cfDNA is deduplicated based on the UMI sequence corresponding to the cfDNA to obtain the first sequencing data corresponding to the cfDNA; and the second sequencing data corresponding to the white blood cell DNA is deduplicated based on the UMI sequence corresponding to the white blood cell DNA to obtain the first sequencing data corresponding to the white blood cell DNA.
[0014] In some possible embodiments, before obtaining the second sequencing data corresponding to the cfDNA and the leukocyte DNA corresponding to the target sample, the processor is further configured to perform the following steps:
[0015] performing sequence alignment on the second sequencing data corresponding to the cfDNA and the reference sequencing data to obtain a first alignment result corresponding to the cfDNA, and performing data analysis on the first alignment result corresponding to the cfDNA to obtain a first quality assessment parameter corresponding to the cfDNA;
[0016] performing sequence alignment on the second sequencing data corresponding to the white blood cell DNA and the reference sequencing data to obtain a first alignment result corresponding to the white blood cell DNA, and performing data analysis on the first alignment result corresponding to the white blood cell DNA to obtain a first quality assessment parameter corresponding to the white blood cell DNA;
[0017] After performing deduplication processing on the second sequencing data corresponding to the cfDNA and the white blood cell DNA respectively according to the UMI sequence to obtain the first sequencing data corresponding to the cfDNA and the white blood cell DNA respectively, the processor is further configured to perform the following steps:
[0018] performing sequence alignment on the first sequencing data corresponding to the cfDNA and the reference sequencing data to obtain a second alignment result corresponding to the cfDNA, and performing data analysis on the second alignment result corresponding to the cfDNA to obtain a second quality assessment parameter corresponding to the cfDNA;
[0019] performing sequence alignment on the first sequencing data corresponding to the white blood cell DNA and the reference sequencing data to obtain a second alignment result corresponding to the white blood cell DNA, and performing data analysis on the second alignment result corresponding to the white blood cell DNA to obtain a second quality assessment parameter corresponding to the white blood cell DNA;
[0020] Before detecting the first sequencing data corresponding to the cfDNA and the leukocyte DNA respectively according to the multiple mutation detection strategies, the processor is further configured to perform the following steps:
[0021] If the first quality assessment parameters corresponding to the cfDNA and the white blood cell DNA respectively and / or the second quality assessment parameters corresponding to the cfDNA and the white blood cell DNA respectively do not meet the preset first quality requirements, the second sequencing data corresponding to the cfDNA and the white blood cell DNA corresponding to the target sample are re-obtained.
[0022] In some possible embodiments, the first sequencing data includes multiple sequence fragments, each of which includes multiple sites; the multiple mutation detection strategies include a first mutation detection strategy and a second mutation detection strategy;
[0023] Detecting the first sequencing data corresponding to the cfDNA according to multiple mutation detection strategies to obtain mutation detection results corresponding to the cfDNA and each of the mutation detection strategies, including:
[0024] Based on the first mutation detection strategy, comparatively analyzing the first sequencing data corresponding to the cfDNA and the white blood cell DNA, respectively, to obtain a plurality of candidate mutation sites, screening the plurality of candidate mutation sites using a classification model to obtain a plurality of first mutation sites, and using the plurality of first mutation sites as first mutation detection results corresponding to the cfDNA and the first mutation detection strategy;
[0025] Based on the second mutation detection strategy, the mutation frequency of each site in the first sequencing data corresponding to the cfDNA is calculated, and the second mutation site whose mutation frequency meets the frequency condition is screened out. The target sequence fragments with small insertions or deletions in the first sequencing data corresponding to the cfDNA are detected, and the target sequence fragments are locally realigned to obtain the third mutation site in the target sequence fragment whose mutation frequency meets the frequency condition. The second mutation site and the third mutation site are used as the second mutation detection result corresponding to the cfDNA.
[0026] In some possible embodiments, before combining the cfDNA with the mutation detection results corresponding to each mutation detection strategy to obtain a first detection result corresponding to the cfDNA, the processor is further configured to perform the following steps:
[0027] Filtering the first mutation detection results to remove first mutation sites that do not meet a preset second quality requirement in the first mutation detection results, to obtain a first filtered result;
[0028] Filtering the second mutation detection result according to the first filtering result, removing mutation sites in the second mutation detection result that are consistent with the first filtering result, to obtain a second filtering result;
[0029] The step of combining the cfDNA with the mutation detection results corresponding to each of the mutation detection strategies to obtain a first detection result corresponding to the cfDNA includes:
[0030] Converting the first filtering result and the second filtering result corresponding to the cfDNA into a unified data format;
[0031] The format-converted first filtering result and the second filtering result are combined to obtain a first detection result corresponding to the cfDNA.
[0032] In some possible embodiments, the first sequencing data further includes sequencing depth and base recognition quality corresponding to each site of each sequence fragment, the preset second quality requirement includes a sequencing depth greater than or equal to a depth threshold, and a base recognition quality greater than or equal to a quality threshold; filtering the first mutation detection result to remove first mutation sites that do not meet the preset second quality requirement in the first mutation detection result to obtain a first filtering result includes:
[0033] If the sequencing depth corresponding to the target first mutation site in the first mutation detection result is less than the depth threshold or the corresponding base recognition quality is less than the quality threshold, then the target first mutation site is removed from the first mutation detection result; the target first mutation site is any one of the multiple first mutation sites in the first mutation detection result;
[0034] The remaining first mutation sites in the first mutation detection result are used as the first filtering result.
[0035] In some possible embodiments, filtering the second mutation detection result according to the first filtering result, and removing mutation sites in the second mutation detection result that are consistent with the first filtering result, to obtain a second filtering result, includes:
[0036] Comparing the first filtering result with the second mutation detection result, and removing the same mutation sites in the second mutation detection result as in the first filtering result, to obtain a third mutation detection result;
[0037] The third mutation detection result is compared with a preset reference sequence set to determine a fourth mutation site in the third mutation detection result that matches the preset reference sequence set, and each of the fourth mutation sites is used as a second filtering result.
[0038] In some possible embodiments, determining the target detection result corresponding to the target sample according to the first detection result, the second detection result, and a preset reference sequence set includes:
[0039] Annotating the first test result and the second test result to obtain initial annotation results corresponding to the cfDNA and the leukocyte DNA, respectively, the initial annotation results including multiple mutation sites and gene names corresponding to the sequence fragments where the multiple mutation sites are located;
[0040] Determining identical mutation sites in the initial annotation results corresponding to the cfDNA and the leukocyte DNA, respectively, and removing the identical mutation sites from the initial annotation results corresponding to the cfDNA to obtain a target annotation result corresponding to the cfDNA;
[0041] According to the first sequencing data corresponding to the cfDNA, counting the mutation data corresponding to each mutation site in the target annotation result;
[0042] According to the mutation condition, the mutation sites whose mutation data meet the mutation condition are screened from the target annotation results;
[0043] Comparing each mutation site that meets the mutation condition with a preset reference sequence set to determine a fifth mutation site that matches the reference sequence set;
[0044] Obtain the gene name corresponding to each sequence fragment where the fifth mutation site is located to obtain the target detection result corresponding to the target sample.
[0045] A second aspect of the embodiments of the present application discloses a method for processing sequencing data, which is applied to a computer device, and the method comprises:
[0046] The computer device obtains first sequencing data corresponding to circulating free deoxyribonucleic acid (cfDNA) and leukocyte DNA of the target sample respectively;
[0047] The computer device detects the first sequencing data corresponding to the cfDNA and the white blood cell DNA respectively according to multiple mutation detection strategies, and obtains the mutation detection results corresponding to the cfDNA and each of the mutation detection strategies, and the mutation detection results corresponding to the white blood cell DNA and each of the mutation detection strategies;
[0048] The computer device combines the cfDNA with the mutation detection results corresponding to each of the mutation detection strategies to obtain a first detection result corresponding to the cfDNA;
[0049] The computer device combines the white blood cell DNA with the mutation detection results corresponding to each of the mutation detection strategies to obtain a second detection result corresponding to the white blood cell DNA;
[0050] The computer device determines a target detection result corresponding to the target sample based on the first detection result, the second detection result and a preset reference mutation result.
[0051] A third aspect of the embodiments of the present application discloses a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processor implements the method for processing sequencing data as described in any of the above embodiments.
[0052] The present application provides a computer device, a sequencing data processing method, and a storage medium. The computer device obtains first sequencing data corresponding to circulating free deoxyribonucleic acid (cfDNA) and white blood cell DNA of a target sample, respectively, and detects the first sequencing data corresponding to the cfDNA and white blood cell DNA according to multiple mutation detection strategies, to obtain mutation detection results corresponding to the cfDNA and each mutation detection strategy, and mutation detection results corresponding to the white blood cell DNA and each mutation detection strategy. The computer device can merge the mutation detection results corresponding to the cfDNA and each mutation detection strategy to obtain the first detection result corresponding to the cfDNA, and the computer device can merge the mutation detection results corresponding to the white blood cell DNA and each mutation detection strategy to obtain the first detection result corresponding to the cfDNA. The mutation detection results are merged to obtain a second detection result corresponding to the white blood cell DNA, so that the computer device can determine the target detection result corresponding to the target sample based on the first detection result, the second detection result and the preset reference sequence set. Therefore, in an embodiment of the present application, the computer device can perform mutation detection results on cfDNA and white blood cell DNA respectively through multiple mutation detection strategies, and merge the mutation detection results corresponding to each mutation detection strategy to achieve the advantages of multiple mutation detection strategies, thereby avoiding the false positive results or false negative results that may be brought about by a single mutation detection strategy, and realizing comprehensive and accurate detection of various mutation phenomena included in the target sample, thereby improving the accuracy of mutation detection on the target sample. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0054] Figure 1 A schematic diagram of a DNA sequence provided in an embodiment of the present application;
[0055] Figure 2 A schematic diagram of a mutated sequence fragment provided in an embodiment of the present application;
[0056] Figure 3 A flowchart of a method for processing sequencing data provided in an embodiment of the present application;
[0057] Figure 4 A flowchart of preprocessing the second sequencing data provided in an embodiment of the present application;
[0058] Figure 5 A flowchart for detecting sequencing data according to quality assessment parameters provided in an embodiment of the present application;
[0059] Figure 6 A flowchart of performing mutation detection according to various mutation detection strategies provided in the embodiments of the present application;
[0060] Figure 7 A flowchart of combining the filtering results provided in an embodiment of the present application to obtain the test results corresponding to cfDNA and leukocyte DNA respectively;
[0061] Figure 8 A flowchart for filtering the first mutation detection result provided in an embodiment of the present application;
[0062] Figure 9 A flowchart for filtering the second mutation detection result provided in an embodiment of the present application;
[0063] Figure 10 A flowchart for determining a target detection result corresponding to a target sample provided in an embodiment of the present application;
[0064] Figure 11 A structural block diagram of a sequencing data processing device provided in an embodiment of the present application;
[0065] Figure 12 A structural block diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0066] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0067] It should be noted that the terms "including," "having," and any variations thereof in the embodiments and drawings of this application are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.
[0068] In addition, "at least one" means one or more, and "more" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b and c can mean: a, or b, or c, or a and b, or a and c, or b and c, or a, b and c, where a, b, c can be single or multiple.
[0069] The sequencing data processing method provided in the embodiments of the present application can be applied to computer devices, which may include but are not limited to personal computers, tablet computers, laptop computers, mobile phones, etc.
[0070] Figure 1 This is a schematic diagram of a DNA sequence provided in the examples of this application. Figure 1 As shown, a deoxyribonucleic acid (DNA) sequence 100 includes multiple genomic regions 101, each genomic region 101 is composed of multiple bases, wherein the bases include adenine (A), thymine (T), cytosine (C) and guanine (G). The sequence fragments 103 of the genomic region 101 within the dotted region 102 include sequence fragment D, sequence fragment E, sequence fragment F and sequence fragment G, wherein the base arrangement order of sequence fragment G is ATCGGGTCATGTCA.
[0071] Sequencing may refer to determining the order of the bases in a DNA sequence 100.
[0072] In some embodiments, the DNA sequence 100 may be sequenced using high-throughput sequencing technology, wherein high-throughput sequencing technology may refer to a technology for simultaneously sequencing a large number of sequence fragments in the DNA sequence 100 at one time. High-throughput sequencing technology may include sequencing-by-synthesis technology, semiconductor sequencing technology, and single-molecule real-time sequencing technology.
[0073] The principle of sequencing a DNA sequence 100 using high-throughput sequencing technology is described as follows: After extracting the DNA sequence, the extracted DNA sequence is fragmented by physical or enzymatic methods to obtain sequence fragments of relatively concentrated sizes. After repairing the sequence fragments by adding specific adapter sequences to both ends, the repaired sequence fragments are amplified using polymerase chain reaction (PCR) to complete the construction of a DNA sequence library. The constructed library is then loaded onto sequencing platforms corresponding to different high-throughput sequencing technologies, and sequencing reactions are performed on the multiple sequence fragments contained in the library, thereby completing the sequencing of the DNA sequence. Taking sequencing by synthesis technology as an example, a specific fluorescently labeled deoxyribonucleoside triphosphate (dNTP) is added sequentially in each sequencing cycle, allowing each dNTP to complementarily bind to the sequence fragment through DNA polymerase. Therefore, in each sequencing cycle, by detecting the color and intensity of each added dNTP and following the order of appearance of each fluorescent signal, the fluorescent signal sequence of each sequence fragment can be converted into the corresponding base sequence through a specific algorithm to complete the sequencing of the DNA sequence.
[0074] In some embodiments, sequencing may include single-end sequencing and paired-end sequencing. Single-end sequencing may refer to sequencing from one end of the DNA sequence 100, generating sequence fragments in only one direction. Paired-end sequencing may refer to sequencing from both ends of the DNA sequence 100 simultaneously, generating sequence fragments in both directions. For example, paired-end sequencing may be performed on sequence fragment G to obtain a first sequence fragment ATCGGGTCATGTCA and a second sequence fragment ACTGTACTGGGCTA.
[0075] The sequencing data may include multiple sequence fragments contained in the DNA sequence 100. Each sequence fragment may include multiple sites, and a site may refer to the specific arrangement position of the corresponding base in the sequence fragment. Figure 1 As shown, the sequencing data of the DNA sequence 100 includes the arrangement of each site in the sequence fragment G: ATCGGGTCATGTCA.
[0076] In sequencing data, each sequence fragment is also called a read. The length of a read refers to the number of bases contained in the corresponding sequence fragment. The length of a read can include 50bp, 90bp, 100bp, and 150bp, etc. The length of a read is usually determined by the reading limit of the sequencer or the sequencing technology.
[0077] Alternatively, under the same sequencing technology, the length of each read may be the same or different, but is not limited thereto. In the case where the lengths of the reads are different, the lengths of the reads may be made the same by trimming or padding the reads.
[0078] In the DNA sequence 100, multiple genomic regions 101 may undergo different types of gene mutations due to internal factors such as DNA replication errors, or external factors such as ultraviolet radiation, nitrites, and viruses. The types of gene mutations may include sequence variation and structural variation (SV). Among them, sequence variation may include single nucleotide variation (SNV), small insertion and small deletion (InDel). SV includes long sequence fragment insertion of more than 50bp (bases), copy number variation (CNV), deletion, inversion and translocation, etc.
[0079] Figure 2 Schematic diagram of the sequence fragments that have undergone mutations provided in the examples of this application. Figure 2 As shown, sequence fragment 201 is an SNV mutation, which is manifested as the replacement of a base G in sequence fragment E of sequence fragment 103 with base A; sequence fragment 202 is a small insertion mutation, which is manifested as the additional insertion of a small sequence fragment ACC after the base G ranked at the eleventh position in sequence fragment E of sequence fragment 103; sequence fragment 203 is a small deletion mutation, which is manifested as the deletion of the sequence fragment after the base G ranked at the eleventh position in sequence fragment E of sequence fragment 103; sequence fragment 204 is a deletion mutation, which is manifested as the deletion of sequence fragment E in sequence fragment 103; sequence fragment 205 is a CNV mutation, which is manifested as multiple repetitions of sequence fragment F in sequence fragment 103; sequence fragment 206 is an inversion mutation, which is manifested as the inversion of the order of the bases of sequence fragments F and sequence fragment G of sequence fragment 103; sequence fragment 207 is an ectopic mutation, which is manifested as the position exchange of sequence fragments F and sequence fragment G of sequence fragment 103 with sequence fragments of other DNA or sequence fragments.
[0080] It should be noted that mutations in sequence fragments in genomic regions usually cause normal cells to become diseased, thereby leading to changes in protein function or disorders in cellular physiological processes, and then inducing various diseases. In cases where the content of diseased cells caused by mutations in sequence fragments is high, they can be discovered in a timely manner through traditional methods such as imaging examinations or blood marker tests. However, in the early stages of some diseases or after the disease is cured, the patient's body usually only has extremely low levels of diseased cells. Due to the insufficient sensitivity of traditional methods, it is difficult to detect these extremely low levels of diseased cells. In addition, due to the proliferative nature of diseased cells, diseased cells may gradually accumulate in the patient's body and form new lesions, leading to recurrence or further development of the disease.
[0081] Therefore, mutation detection can be performed on the sequencing data of cfDNA contained in or released by cells to identify whether mutations have occurred in various genomic regions in the cfDNA sequence, and whether the sequence fragments of the mutated genomic regions correspond to sequence fragments of known mutation types in the reference sequence set. In other words, accurate detection of extremely low levels of diseased cells can be achieved by processing and analyzing the sequencing data.
[0082] For example, taking the MRD of a tumor as an example, MRD refers to a small amount of cancer cells remaining in the patient's body after cancer treatment. Since the content of ctDNA released into the blood by MRD is extremely low, traditional methods cannot detect MRD. Therefore, the patient's cfDNA can be sequenced and the obtained sequencing data can be tested for mutations. By analyzing whether the genomic regions where mutations occur in the sequencing data are consistent with the genomic regions with known mutations in ctDNA in the reference sequence set, the processing and analysis of the sequencing data can be completed.
[0083] The embodiments of the present application provide a method for processing sequencing data. A computer device can comprehensively and accurately detect various mutation phenomena included in the target sample by integrating the different advantages of multiple mutation detection strategies, avoiding the false positive or false negative results that may be caused by a single mutation detection strategy, and improving the accuracy of mutation detection in the target sample.
[0084] like Figure 3 As shown, in one embodiment, a method for processing sequencing data is provided, which can be applied to a computer device. The method may include the following steps:
[0085] In step 301 , a computer device obtains first sequencing data corresponding to circulating free deoxyribonucleic acid (cfDNA) and white blood cell DNA of a target sample.
[0086] The target sample may include a peripheral blood sample or a bone marrow fluid sample. cfDNA may refer to small DNA fragments released into the target sample for degradation. Leukocyte DNA may refer to intact DNA carried by normal white blood cells present in the target sample.
[0087] Since white blood cells are a common cell type in the target sample, the information of the genomic regions of white blood cell DNA is usually consistent with that of the genomic regions of normal cells, and can serve as a representative of cfDNA of non-tumor origin. Therefore, the first sequencing data corresponding to white blood cell DNA is used as the control group of the first sequencing data corresponding to cfDNA. Through comparative analysis of the two, mutations existing in the white blood cell DNA can be excluded, the true mutations in cfDNA can be accurately identified, and the accuracy and reliability of mutation detection can be improved.
[0088] In some embodiments, the cfDNA and leukocyte DNA of the target sample can be sequenced separately using exactly the same sequencing technology to obtain first sequencing data corresponding to the cfDNA and leukocyte DNA, respectively.
[0089] Using the same sequencing technology can ensure that the quality of the first sequencing data corresponding to cfDNA and white blood cell DNA is consistent, avoiding deviations in the interpretation of the first sequencing data due to differences in sequencing technology and improving the accuracy of subsequent mutation detection.
[0090] In step 303, the computer device detects the first sequencing data corresponding to the cfDNA and the white blood cell DNA according to multiple mutation detection strategies, and obtains the mutation detection results corresponding to the cfDNA and each mutation detection strategy, as well as the mutation detection results corresponding to the white blood cell DNA and each mutation detection strategy.
[0091] The mutation detection strategy can be used to identify whether mutations occur in each genomic region in the first sequencing data corresponding to cfDNA and white blood cell DNA, respectively.
[0092] The mutation detection results corresponding to each mutation detection strategy may include multiple mutation sites in cfDNA and leukocyte DNA under the corresponding mutation detection strategy. The mutation site may refer to the specific arrangement position of the mutated base in the sequence fragment.
[0093] It is understandable that in the related art, only one mutation detection strategy is usually used for mutation detection. However, although each mutation detection strategy usually has its own unique detection conditions, each mutation detection strategy will have some defects to a greater or lesser extent. Therefore, the accuracy of the mutation detection results is usually improved by optimizing the detection parameters, detection conditions, or algorithm models contained in the mutation detection strategy. In the embodiments of the present application, by adopting multiple mutation detection strategies for mutation detection, not only can the limitations of a single mutation detection strategy such as the inability to cover all mutation types and the low accuracy of the detection results be avoided, but also the unique advantages of multiple mutation detection strategies can be combined to achieve comprehensive and accurate detection of various mutation types, effectively reducing false positives caused by background noise or other non-tumor-derived mutations and false negatives caused by missed detection or erroneous detection, thereby improving the accuracy and reliability of mutation detection.
[0094] In some embodiments, computer recognition can be performed on the first sequencing data corresponding to cfDNA by using multiple mutation detection strategies, and the first sequencing data corresponding to white blood cell DNA can be performed on the first sequencing data corresponding to the white blood cell DNA by using the same multiple mutation detection strategies, thereby ensuring that the mutation detection results identified for cfDNA and white blood cell DNA are comparable under the same mutation detection strategy, thereby avoiding uncertainty caused by differences in mutation detection strategies.
[0095] In some embodiments, the multiple mutation detection strategies may include a first mutation detection strategy and a second mutation detection strategy. The first mutation detection strategy may include a mutation detection strategy that compares cfDNA and leukocyte DNA to identify specific mutations in the cfDNA. Specific mutations may refer to mutations that are present only in cfDNA and not in leukocyte DNA. The second mutation detection strategy may include a mutation detection strategy that performs in-depth mutation analysis on cfDNA and leukocyte DNA to detect low-frequency or rare mutations.
[0096] The computer device can perform mutation detection on the first sequencing data corresponding to cfDNA and white blood cell DNA respectively by adopting a first mutation detection strategy to obtain a first mutation detection result of the cfDNA and white blood cell DNA corresponding to the first mutation detection strategy, and can perform mutation detection on the first sequencing data corresponding to cfDNA and white blood cell DNA respectively by adopting a second mutation detection strategy to obtain a second mutation detection result of the cfDNA and white blood cell DNA corresponding to the second mutation detection strategy.
[0097] By adopting the first mutation detection strategy and the second mutation detection strategy to perform mutation detection on cfDNA and white blood cell DNA, it is possible to detect both specific mutations in the DNA sequence and low-frequency mutations or rare mutations in the DNA sequence, thereby achieving the combination and complementarity of the mutation detection results corresponding to different mutation detection strategies, thereby improving the accuracy of mutation detection.
[0098] In some embodiments, the computer device can implement mutation detection on first sequencing data corresponding to cfDNA and white blood cell DNA, respectively, by calling scripts, software, or software modules corresponding to multiple mutation detection strategies. For example, the computer device can call the TNscope module of Sentieon software to perform mutation detection on the first sequencing data corresponding to cfDNA and white blood cell DNA, respectively, to obtain first mutation detection results corresponding to cfDNA and white blood cell DNA, respectively, and call VarDict software to perform mutation detection on the first sequencing data corresponding to cfDNA and white blood cell DNA, respectively, to obtain second mutation detection results corresponding to cfDNA and white blood cell DNA, respectively.
[0099] In step 305 , the computer device combines the cfDNA with the mutation detection results corresponding to each mutation detection strategy to obtain a first detection result corresponding to the cfDNA.
[0100] The first test result may include all mutation sites contained in cfDNA under multiple mutation detection strategies.
[0101] In some embodiments, the computer device may use the union of the cfDNA and the mutation detection results corresponding to each mutation detection strategy as the first detection result corresponding to the cfDNA. By using the union of the cfDNA and the mutation detection results corresponding to each mutation detection strategy as the first detection result corresponding to the cfDNA, the mutation detection results corresponding to each mutation detection strategy can be deduplicated and then merged, thereby preventing the first detection result from including duplicate mutation detection results.
[0102] In some embodiments, the computer device may filter the mutation detection results corresponding to cfDNA and each mutation detection strategy to obtain a first filtering result corresponding to cfDNA and the first mutation detection strategy, and a second filtering result corresponding to cfDNA and the second mutation detection strategy, and then merge the first filtering result and the second filtering result to obtain the first detection result corresponding to cfDNA.
[0103] It is understandable that due to the limitations of sequencing technology itself and the sensitivity of mutation detection strategies, there are inevitably some false detections and misdetections. Therefore, it is necessary to filter the mutation detection results to exclude mutation sites in the mutation detection results that may be caused by technical errors rather than real mutations, so as to improve the accuracy and reliability of mutation detection.
[0104] In some embodiments, the computer device can merge the cfDNA with the mutation detection results corresponding to each mutation detection strategy by calling merging software. For example, after calling the TNscope module of Sentieon software and VarDict software to perform mutation detection on the first sequencing data corresponding to the cfDNA, the computer device can generate a corresponding VCF (Variant Call Format) result file containing the mutation detection results, and then call Bcftools software to merge the VCF result file corresponding to the TNscope module and the VCF result file corresponding to the VarDict software to obtain a result file corresponding to the first detection result of the cfDNA.
[0105] In step 307 , the computer device combines the white blood cell DNA with the mutation detection results corresponding to each mutation detection strategy to obtain a second detection result corresponding to the white blood cell DNA.
[0106] The second test result may include all mutation sites contained in the leukocyte DNA under multiple mutation detection strategies.
[0107] The step of the computer device merging the white blood cell DNA with the mutation detection results corresponding to each mutation detection strategy is the same as the step of merging the cfDNA with the mutation detection results corresponding to each mutation detection strategy in the above step 305, and will not be repeated here.
[0108] In some embodiments, when the computer device filters the mutation detection results corresponding to cfDNA and each mutation detection strategy, the computer device may filter the mutation detection results corresponding to white blood cell DNA and each mutation detection strategy according to the same filtering conditions to obtain a second detection result corresponding to the white blood cell DNA.
[0109] In step 309 , the computer device determines a target detection result corresponding to the target sample based on the first detection result, the second detection result, and a preset reference sequence set.
[0110] The reference sequence set may refer to a set of mutation sequence fragments corresponding to multiple ctDNAs, which are set according to the target sample. A mutation sequence fragment may refer to a sequence fragment in which a mutation occurs. Alternatively, the reference sequence set may include multiple mutation sequence fragments corresponding to MRDs.
[0111] Target detection results can include MRD-positive and MRD-negative. MRD-positive refers to the presence of MRD in the target sample, while MRD-negative refers to the absence of MRD in the target sample or the presence of cfDNA that is less than the sensitivity of various mutation detection strategies.
[0112] In some embodiments, the computer device may compare and analyze the first test result and the second test result to determine the mutation sites in cfDNA that are different from those in white blood cell DNA, as well as the sequence fragments where the different mutation sites are located.
[0113] Since the first test result includes all mutation sites in cfDNA and the second test result includes all mutation sites in leukocyte DNA, leukocyte DNA can be used as a background control group to remove the same mutation sites in cfDNA and leukocyte DNA from all mutation sites in cfDNA, thereby eliminating the background noise of leukocyte DNA and mutation sites caused by other factors, which can effectively reduce the impact of false positive test results on the detection of target samples.
[0114] In some embodiments, after determining the mutation sites in cfDNA that are different from those in white blood cell DNA, the computer device can match and analyze the sequence fragments where the different mutation sites are located with a preset reference sequence set to determine the target detection results corresponding to the target sample.
[0115] Reference sequence collections can include fixed panels (test kits) and personalized panels. Fixed panels can refer to collections of sequences consisting of multiple fixed sequence regions, which typically cover the most likely mutation regions. Personalized panels can refer to collections of sequences based on the genetic background of the target sample and the known mutation sequences corresponding to ctDNA.
[0116] It is understandable that the reference sequence set contains a variety of known mutation sequence fragments for MRD. Therefore, by comparing whether the sequence fragments corresponding to the mutation sites in cfDNA that are different from those in white blood cell DNA match the multiple mutation sequence fragments contained in the reference sequence set, it is possible to determine whether there are sequence fragments related to MRD in cfDNA, and then determine the target detection results corresponding to the target sample.
[0117] In some embodiments, the computer device can construct a personalized panel based on the genetic background corresponding to the target sample and the mutation sequence corresponding to the known ctDNA during the process of sequencing cfDNA and white blood cell DNA, and determine the target detection result corresponding to the target sample based on the first detection result, the second detection result and the personalized panel.
[0118] By designing a personalized panel based on the genetic background of the target sample and the known mutation sequence corresponding to ctDNA, the mutation detection strategy can be made more targeted, accurately capturing specific mutation sites associated with ctDNA, helping to discover low-frequency or rare mutations, thereby improving the accuracy and efficiency of mutation detection.
[0119] In an embodiment of the present application, a computer device can perform mutation detection results on cfDNA and white blood cell DNA separately through multiple mutation detection strategies, and merge the mutation detection results corresponding to each mutation detection strategy to achieve the combined advantages of multiple mutation detection strategies, thereby avoiding false positive results or false negative results that may be caused by a single mutation detection strategy, and achieving comprehensive and accurate detection of various mutation phenomena included in the target sample, thereby improving the accuracy of mutation detection on the target sample.
[0120] In some embodiments, since the sequencing data obtained by the sequencer usually contains repeated sequence fragments, directly performing mutation detection on the sequencing data obtained by the sequencer will affect the accuracy of mutation detection. Therefore, the computer equipment needs to preprocess the sequencing data obtained by the sequencer and use the preprocessed sequencing data for mutation detection. Figure 4 This is a flowchart of preprocessing the second sequencing data provided in the embodiment of the present application. Figure 4 As shown, the step of obtaining first sequencing data corresponding to circulating free deoxyribonucleic acid cfDNA and leukocyte DNA of the target sample may include the following steps:
[0121] Step 402: Obtain second sequencing data corresponding to the cfDNA and leukocyte DNA corresponding to the target sample.
[0122] The second sequencing data may refer to the original data from the sequencer, that is, the sequencing data without any processing after the sequencer sequences the cfDNA and white blood cell DNA.
[0123] In some embodiments, the computer device may obtain the adapter sequence, base recognition quality, base recognition error rate, etc. corresponding to each sequence fragment in the second sequencing data, and filter each sequence fragment of the second sequencing data.
[0124] Adapter sequences refer to the short sequences used to connect sequence fragments to the sequencing chip during sequencing. The sequencing chip is the physical carrier used to immobilize and detect sequence fragments. Base call quality refers to the probability of correctly calling a base, usually expressed directly as a numerical value. The base call error rate refers to the ratio of incorrectly called bases to the total number of bases in a sequence fragment, for example, 5%, 10%, and 20%.
[0125] The computer device can filter the adapter sequences of each sequence fragment in the second sequencing data and remove sequence fragments with base recognition quality lower than a preset quality threshold and base recognition error rate greater than 10%, thereby completing the filtering of the second sequencing data.
[0126] In step 404, the unique molecular marker UMI sequences corresponding to the cfDNA and the white blood cell DNA are extracted from the second sequencing data corresponding to the cfDNA and the white blood cell DNA, respectively.
[0127] A unique molecular identifier (UMI) sequence refers to a specific short sequence of bases added to each sequence fragment prior to PCR amplification during the construction of cfDNA and leukocyte DNA libraries. UMI sequences can be used to distinguish multiple copies of the same sequence fragment generated through PCR amplification.
[0128] In some embodiments, because UMI sequences are typically designed within specific regions of sequence fragments, a computer device can extract the UMI sequences for each sequence fragment corresponding to cfDNA and leukocyte DNA by setting the positions and lengths of the UMI sequences within the multiple sequence fragments corresponding to cfDNA and leukocyte DNA, respectively. For example, a UMI sequence can be set at the start position of each cfDNA sequence fragment, and the UMI sequence length is 8 bp. The computer device can then extract an 8-bp sequence from the start position of each cfDNA sequence fragment as the UMI sequence.
[0129] In some embodiments, the computer device can extract the UMI sequence of each sequence fragment in the second sequencing data corresponding to cfDNA and white blood cell DNA respectively by calling the UMI extract command of the Sentieon software.
[0130] In step 406, the second sequencing data corresponding to the cfDNA is deduplicated based on the UMI sequence corresponding to the cfDNA to obtain the first sequencing data corresponding to the cfDNA; and the second sequencing data corresponding to the white blood cell DNA is deduplicated based on the UMI sequence corresponding to the white blood cell DNA to obtain the first sequencing data corresponding to the white blood cell DNA.
[0131] The UMI sequence is used to remove repeated sequence fragments introduced by PCR amplification or sequencing, retaining only one representative sequence fragment.
[0132] The computer device can obtain multiple UMI sequences contained in the second sequencing data corresponding to the cfDNA, and sequentially detect whether the multiple sequence fragments contained in the second sequencing data have the same UMI sequence based on the UMI sequence. If the same UMI sequence exists, the multiple sequence fragments are considered to be from the same original sequence fragment (i.e., the sequence fragment before PCR amplification), and only one of the sequence fragments is retained, while the remaining duplicate sequence fragments are removed.
[0133] Optionally, the computer device may further detect whether multiple sequence fragments have similar UMI sequences. Similarity may refer to the same order of multiple bases in the UMI sequence, with only a predetermined number of differences. For example, if two UMI sequences differ by only one base, the UMI sequences corresponding to the two sequence fragments may be considered similar, and one of the sequence fragments may be retained while the other is removed.
[0134] The computer device performs deduplication processing on the second sequencing data corresponding to the white blood cell DNA based on the UMI sequence corresponding to the white blood cell DNA to obtain the first sequencing data corresponding to the white blood cell DNA. The process is the same as the above-mentioned method of performing deduplication processing on the second sequencing data corresponding to the cfDNA based on the UMI sequence corresponding to the cfDNA to obtain the first sequencing data corresponding to the cfDNA, and is not further described here.
[0135] In an embodiment of the present application, a computer device can extract the unique molecular marker UMI sequences corresponding to cfDNA and white blood cell DNA, respectively, and perform deduplication processing on the second sequencing data corresponding to cfDNA and white blood cell DNA, respectively, based on the UMI sequences corresponding to cfDNA and white blood cell DNA, respectively, thereby obtaining the first sequencing data corresponding to cfDNA and white blood cell DNA, respectively. This can effectively remove repeated sequence fragments introduced by PCR amplification and sequencing processes, reduce the redundancy of sequencing data, and improve the efficiency of subsequent mutation detection on the sequencing data.
[0136] In some embodiments, the computer device may perform quality control (referred to as quality control) on the first sequencing data and the second sequencing data corresponding to cfDNA and leukocyte DNA, respectively, to detect whether the quality of the first sequencing data and the second sequencing data meets the quality requirements required for detecting the target sample. Figure 5 This is a flow chart of detecting sequencing data according to quality assessment parameters provided in the embodiment of the present application. Figure 5 As shown, the method may further include the following steps:
[0137] Step 501: Obtain second sequencing data corresponding to cfDNA and leukocyte DNA respectively.
[0138] Step 503: Perform sequence alignment on the second sequencing data corresponding to the cfDNA and the reference sequencing data to obtain a first alignment result corresponding to the cfDNA, and perform data analysis on the first alignment result corresponding to the cfDNA to obtain a first quality assessment parameter corresponding to the cfDNA.
[0139] The reference sequencing data may include sequencing data corresponding to a human reference genome. It is understood that by performing sequence alignment between the second sequencing data corresponding to the cfDNA and the reference sequencing data, various information such as the positions and mutation status of multiple sequence fragments contained in the second sequencing data corresponding to the cfDNA in the human reference genome can be determined, thereby determining the first alignment result based on the above information.
[0140] The alignment results may include header annotation information and alignment data information. The header annotation information may include the version of the alignment software used and information corresponding to the reference sequence being aligned. The alignment data information may include the name of the sequence fragment contained in the second sequencing data corresponding to the cfDNA, the starting position of the sequence fragment in the reference sequencing data, and the alignment quality. A higher alignment quality indicates a higher degree of confidence in the alignment of the read to the current position.
[0141] Quality assessment parameters may include sequencing coverage, capture efficiency, and uniformity. Sequencing coverage may refer to the proportion of a sequence fragment in the second sequencing data corresponding to cfDNA that covers the target region, and the target region may include the long sequence corresponding to cfDNA or the long sequence corresponding to the human reference genome. Capture efficiency may refer to the proportion of multiple sequence fragments in the second sequencing data corresponding to cfDNA that actually fall within the multiple sequence fragments contained in the human reference genome. Uniformity may refer to a measure of the uniformity of the sequencing depth of multiple sequence fragments contained in the second sequencing data corresponding to cfDNA in the target region. The higher the uniformity, the more uniform the sequencing depth of each site in the second sequencing data. Sequencing depth may refer to the number of times a site in the second sequencing data corresponding to cfDNA is sequenced.
[0142] In some embodiments, the computer device may call the BWA module of the Sentieon software to perform sequence alignment on the second sequencing data corresponding to the cfDNA with the reference sequencing data to obtain a SAM (Sequencing Alignment Format) file containing the first alignment result, and after sorting the SAM file, convert it into a BAM (Binary Alignment / Map) file, and then call the Bamdst software to perform data analysis on the BAM file to obtain the first quality assessment parameter corresponding to the cfDNA.
[0143] It is understandable that since the data scale of the reference sequencing data is usually extremely large, and although the content of cfDNA in the target sample is very low, the length of a single sequence fragment in cfDNA is usually between 160bp and 200bp, therefore, when the second sequencing data corresponding to cfDNA contains multiple sequence fragments, the data volume of the first alignment result generated by sequence alignment is very large. In addition, the SAM file containing the first alignment result obtained by the computer device is in a plain text format. Reading the SAM file will occupy a large amount of operating resources of the computer device, and it is not organized and stored in a specific order during the alignment process. The first alignment result in the SAM file is out of order. Therefore, the SAM file cannot be used directly for data analysis, but needs to be sorted according to a specific format and further compressed into a binary format to form a BAM file. Only through the BAM file can data analysis be performed to obtain the first quality assessment parameter corresponding to cfDNA.
[0144] Step 505 , performing sequence alignment on the second sequencing data corresponding to the white blood cell DNA and the reference sequencing data to obtain a first alignment result corresponding to the white blood cell DNA, and performing data analysis on the first alignment result corresponding to the white blood cell DNA to obtain a first quality assessment parameter corresponding to the white blood cell DNA.
[0145] The method for obtaining the first quality assessment parameter corresponding to the white blood cell DNA is the same as the method for obtaining the first quality assessment parameter corresponding to the cfDNA in step 503, and will not be repeated here.
[0146] Step 507 , determining whether the first quality assessment parameters corresponding to the cfDNA and the leukocyte DNA respectively meet the preset first quality requirements. If not, executing step 509 ; if so, executing step 511 .
[0147] The first quality requirement may include sequencing coverage, sequencing depth, capture efficiency and uniformity being greater than the corresponding preset thresholds, for example, sequencing coverage must be greater than 70%, sequencing depth must be greater than 100,000X, etc., where 100,000X may mean that a site has been sequenced 100,000 times.
[0148] If the first quality assessment parameters corresponding to the cfDNA and white blood cell DNA respectively meet the preset first quality requirements, it indicates that the data quality of the second sequencing data corresponding to the cfDNA and white blood cell DNA respectively is good, and the accuracy and reliability of the second sequencing data are high, which can provide a reliable basis for subsequent mutation detection. Therefore, subsequent deduplication processing using UMI sequences can be performed.
[0149] If the first quality assessment parameters corresponding to the cfDNA and the leukocyte DNA do not meet the preset first quality requirements, it means that there are quality problems in the second sequencing data corresponding to the cfDNA and the leukocyte DNA, which may cause false negative problems such as missed detection and wrong detection in the subsequent mutation detection process. Therefore, it is necessary to execute step 509 to re-acquire the second sequencing data corresponding to the cfDNA and the leukocyte DNA corresponding to the target sample.
[0150] In step 509 , the second sequencing data corresponding to the cfDNA and the leukocyte DNA corresponding to the target sample are re-acquired, and step 503 is performed again.
[0151] Re-obtaining the second sequencing data corresponding to the cfDNA and the leukocyte DNA corresponding to the target sample may refer to re-extracting the cfDNA and the leukocyte DNA corresponding to the target sample, reconstructing the libraries corresponding to the cfDNA and the leukocyte DNA, and re-sequencing the cfDNA and the leukocyte DNA to obtain the second sequencing data corresponding to the cfDNA and the leukocyte DNA, respectively.
[0152] Step 511: Obtain first sequencing data corresponding to cfDNA and white blood cell DNA respectively.
[0153] In some embodiments, when the second sequencing data corresponding to the cfDNA and the white blood cell DNA respectively meet a preset first quality requirement, deduplication can be performed using the UMI sequence to obtain the first sequencing data corresponding to the cfDNA and the white blood cell DNA respectively.
[0154] In step 513, the first sequencing data corresponding to the cfDNA is aligned with the reference sequencing data to obtain a second alignment result corresponding to the cfDNA, and data analysis is performed on the second alignment result corresponding to the cfDNA to obtain a second quality assessment parameter corresponding to the cfDNA.
[0155] Step 515 , performing sequence alignment on the first sequencing data corresponding to the white blood cell DNA and the reference sequencing data to obtain a second alignment result corresponding to the white blood cell DNA, and performing data analysis on the second alignment result corresponding to the white blood cell DNA to obtain a second quality assessment parameter corresponding to the white blood cell DNA.
[0156] Step 517 , determining whether the second quality assessment parameters corresponding to the cfDNA and the leukocyte DNA respectively meet the preset first quality requirement. If not, execute step 509 ; if so, execute step 519 .
[0157] Since the method of obtaining the second quality assessment parameters corresponding to cfDNA and white blood cell DNA respectively in steps 513 to 517 is the same as the method of obtaining the first quality assessment parameters corresponding to cfDNA and white blood cell DNA respectively in steps 503 to 509, they will not be repeated here.
[0158] It should be noted that after performing deduplication processing based on the UMI sequence to obtain the first sequencing data corresponding to cfDNA and white blood cell DNA, respectively, and then obtaining the second quality assessment parameters corresponding to cfDNA and white blood cell DNA, respectively, not only can the deduplication effect of the UMI sequence be evaluated, but the second quality assessment parameters can also be used to perform a more comprehensive verification of the deduplication-free first sequencing data, thereby improving the data quality of the first sequencing data.
[0159] Step 519 : Determine whether the first sequencing data corresponding to the cfDNA and the leukocyte DNA respectively meet a preset first quality requirement.
[0160] In some embodiments, when it is determined that the first sequencing data corresponding to cfDNA and white blood cell DNA respectively meet the preset first quality requirement, the computer device can also determine whether the preset third quality requirement is met by comparing the first quality assessment parameter and the second quality assessment parameter corresponding to the cfDNA, and / or the first quality assessment parameter and the second quality assessment parameter corresponding to the white blood cell DNA.
[0161] The quality assessment parameter may also include the number of sequence fragments included in the second sequencing data or the number of sequence fragments included in the first sequencing data. The third quality requirement may include that the ratio of the number of sequence fragments included in the first sequencing data corresponding to cfDNA to the number of sequence fragments included in the second sequencing data is greater than a preset first ratio, and / or the ratio of the number of sequence fragments included in the first sequencing data corresponding to leukocyte DNA to the number of sequence fragments included in the second sequencing data is greater than a preset second ratio.
[0162] It is understandable that in the second sequencing data corresponding to cfDNA and leukocyte DNA, respectively, the number of repeated sequence fragments should be relatively controllable and will not occupy the vast majority of the second sequencing data. For example, the number of repeated sequence fragments may account for 5%, 15%, 25% or 50% of the total number of sequence fragments contained in the second sequencing data.
[0163] Therefore, if the preset third quality requirement is not met, it generally indicates that the computer device has removed a large number of duplicate sequence fragments from the second sequencing data based on the UMI sequence, resulting in a significant difference between the first quality assessment parameter corresponding to the second sequencing data and the second quality assessment parameter corresponding to the first sequencing data. In other words, various problems may have occurred during the construction of the target sample's cfDNA and leukocyte DNA libraries, or during the sequencing of the cfDNA and leukocyte DNA, or during the extraction of the UMI sequence, such as contamination of the target sample with DNA sequences from external sources or improper PCR amplification conditions, resulting in a large number of duplicate sequence fragments. In such cases, step 509 must also be performed to reacquire the second sequencing data corresponding to the target sample's cfDNA and leukocyte DNA, respectively.
[0164] In some embodiments, if the second sequencing data corresponding to cfDNA and white blood cell DNA respectively meet the preset first quality requirements, the computer device can detect the first sequencing data corresponding to cfDNA and white blood cell DNA respectively through multiple mutation detection strategies to ensure the accuracy and reliability of mutation detection.
[0165] In an embodiment of the present application, the computer device can perform quality control on the second sequencing data and the first sequencing data before and after deduplication of the second sequencing data using the UMI sequence to obtain corresponding quality assessment parameters, thereby determining whether the second sequencing data or the first sequencing data corresponding to cfDNA and white blood cell DNA meets a preset first quality requirement. This can effectively evaluate the overall quality of the sequencing data before and after UMI deduplication, ensuring the reliability and effectiveness of subsequent mutation detection.
[0166] In some embodiments, the computer device can perform mutation detection on the first sequencing data corresponding to cfDNA and white blood cell DNA respectively through a first mutation detection strategy and a second mutation detection strategy, wherein the first mutation detection strategy and the second mutation detection strategy can be used to detect whether there is a mutation site of the sequence fragment corresponding to ctDNA in the first sequencing data corresponding to cfDNA. Figure 6 This is a flow chart of performing mutation detection according to various mutation detection strategies provided in the embodiments of this application. Figure 6 As shown, the step of detecting the first sequencing data corresponding to the cfDNA according to multiple mutation detection strategies to obtain the mutation detection results corresponding to the cfDNA and each mutation detection strategy may include the following steps:
[0167] Step 602: Based on the first mutation detection strategy, the first sequencing data corresponding to the cfDNA and the white blood cell DNA are compared and analyzed to obtain multiple candidate mutation sites, and the multiple candidate mutation sites are screened by the classification model to obtain multiple first mutation sites. The multiple first mutation sites are used as the first mutation detection results corresponding to the cfDNA and the first mutation detection strategy.
[0168] A candidate mutation site may refer to a mutation site with a specific mutation in the first sequencing data corresponding to cfDNA, that is, the candidate mutation site only appears in the first sequencing data corresponding to cfDNA, but does not appear in the first sequencing data corresponding to leukocyte DNA.
[0169] The computer device can compare and analyze the first sequencing data corresponding to cfDNA and white blood cell DNA, that is, use the first sequencing data corresponding to white blood cell DNA as the background control group of the first sequencing data corresponding to cfDNA, and remove the mutation sites that are the same in cfDNA and white blood cell DNA from all mutation sites of cfDNA to obtain multiple candidate mutation sites.
[0170] It is understood that cfDNA can include cell-free DNA from normal cells, cell-free DNA from non-tumor cells, and ctDNA. Therefore, the mutation sites in each sequence fragment of the second sequencing data corresponding to cfDNA also include the mutation sites corresponding to the aforementioned DNA. By comparing and analyzing the first sequencing data corresponding to cfDNA and white blood cell DNA, the mutation sites corresponding to cell-free DNA from normal cells and cell-free DNA from non-tumor cells can be removed, while retaining the mutation sites belonging to ctDNA.
[0171] In some embodiments, the classification model may include a model trained by a machine learning algorithm based on an MRD positive data set, and the classification model may be used to distinguish whether the candidate mutation site belongs to a mutation site in a sequence fragment corresponding to the ctDNA. Wherein, the machine learning algorithm may include a random forest or a gradient boosting decision tree, and the MRD positive data set may include MRD mutation site data classified as true positive, and MRD mutation site data classified as false positive. The first mutation site may include a mutation site in a sequence fragment corresponding to the ctDNA.
[0172] It should be noted that although the mutation sites belonging to ctDNA can be used as candidate mutation sites by comparing and analyzing the first sequencing data corresponding to cfDNA and leukocyte DNA respectively, due to various factors such as sequencing errors, target sample contamination or leukocyte DNA background mutations, there may be some false-positive mutation sites in the candidate mutation sites that are not actually derived from tumor cells. Therefore, a classification model is needed to distinguish and remove false-positive mutation sites in the candidate mutation sites.
[0173] In some embodiments, since the computer device screens multiple candidate mutation sites through a classification model to obtain multiple first mutation sites belonging to sequence fragments corresponding to ctDNA, the computer device can use the multiple first mutation sites as the first mutation detection results corresponding to the ctDNA and the first mutation detection strategy.
[0174] Since the first mutation detection strategy is to first screen candidate mutation sites and then screen multiple candidate mutation sites through a classification model to determine the first mutation site, the first mutation detection strategy has higher specificity, but has certain shortcomings in sensitivity. Specificity can refer to the accuracy of mutation site detection, that is, distinguishing whether the mutation site belongs to the sequence fragment corresponding to ctDNA, and sensitivity can refer to the number of mutation sites.
[0175] In some embodiments, since the first mutation detection strategy is to compare and analyze the first sequencing data corresponding to cfDNA and white blood cell DNA respectively, the first mutation detection result corresponding to cfDNA and the first mutation detection strategy is the same as the first mutation detection result corresponding to white blood cell DNA and the first mutation detection strategy.
[0176] In some embodiments, the computer device can call the TNscope module of the Sentieon software to perform comparative analysis on the first sequencing data corresponding to the cfDNA and the white blood cell DNA, respectively, to determine the first mutation detection result corresponding to the cfDNA and the first mutation detection strategy.
[0177] Step 604: Based on the second mutation detection strategy, the mutation frequency of each site in the first sequencing data corresponding to the cfDNA is calculated, and the second mutation site whose mutation frequency meets the frequency condition is screened out. The target sequence fragments with small insertions or deletions in the first sequencing data corresponding to the cfDNA are detected, and the target sequence fragments are locally realigned to obtain the third mutation site whose mutation frequency meets the frequency condition in the target sequence fragment. The second mutation site and the third mutation site are used as the second mutation detection result corresponding to the cfDNA.
[0178] Mutation frequency can refer to the ratio of the number of sequence fragments at a mutation site to the total number of sequence fragments in the first sequencing data corresponding to cfDNA, also known as the variant allele frequency (VAF). The frequency condition can refer to the mutation frequency being greater than a preset frequency threshold.
[0179] The mutation frequency and frequency condition can constitute a heuristic condition, and the computer device can use the heuristic condition to filter out potential mutations that do not meet the preset threshold requirements in each mutation site in the first sequencing data corresponding to cfDNA, so as to obtain a second mutation site that meets the sequence fragment corresponding to ctDNA.
[0180] In some embodiments, the computer device can also calculate the average position and alignment quality of each site in the first sequencing data corresponding to the cfDNA, screen out the second mutation site whose average position meets the position condition requirements and whose alignment quality is greater than a preset alignment quality threshold, and use the second mutation site and the third mutation site as the second mutation detection result corresponding to the cfDNA.
[0181] The average position refers to the average position of each site in the first sequencing data corresponding to the cfDNA in each sequence fragment. The position condition requirement may include that the difference between the average position of a site and the position of the site in each sequence fragment is less than a preset difference. If the position condition requirement is met, it means that multiple sequence fragments support the same mutation site at a similar position, and this mutation site is likely the actual mutation site.
[0182] It can be understood that by setting a variety of heuristic conditions, the reliability of each mutation site can be comprehensively evaluated from multiple dimensions, and it can be ensured that the second mutation site screened out is consistent with the mutation site of the sequence fragment corresponding to ctDNA, thereby realizing mutation detection of ctDNA.
[0183] In some embodiments, the computer device may compare and analyze the first sequencing data corresponding to the cfDNA with the reference sequencing data, and then perform local realignment on the target sequence fragments based on the target sequence fragments containing InDels in the first sequencing data corresponding to the detected cfDNA to obtain a third mutation site in the target sequence fragment whose mutation frequency meets the frequency condition.
[0184] Local realignment may refer to using a local realignment algorithm to intercept a small fragment near the mutation site of the InDel in the target sequence fragment and realign it with the reference sequencing data.
[0185] It is understood that when comparing the first sequencing data corresponding to cfDNA with the reference sequencing data, a global alignment algorithm is typically used. This algorithm aligns the entire sequence fragment with the reference sequencing data, focusing on matching and locating the entire sequence fragment. However, the alignment analysis quality of complex mutation sites, such as InDel mutations, is poor. Therefore, a local re-alignment algorithm is required for further alignment analysis. By reducing the length of the aligned sequence fragments, a detailed and accurate analysis of complex mutation sites, such as InDel mutations, is achieved, thereby improving the accuracy of mutation detection results.
[0186] In some embodiments, the one or more heuristic conditions and local re-alignment algorithms used by the computer device for cfDNA are the same as the one or more heuristic conditions and local re-alignment algorithms used for white blood cell DNA to ensure the accuracy of the mutation detection results.
[0187] During mutation detection using the second mutation detection strategy on the first sequencing data corresponding to cfDNA and leukocyte DNA, if one or more heuristic conditions and local re-alignment algorithms differ, the resulting second and third mutation sites corresponding to cfDNA may differ significantly from those corresponding to leukocyte DNA due to differences in sensitivity and specificity between the heuristic conditions and local re-alignment algorithms. These differences prevent leukocyte DNA from accurately serving as a background control for cfDNA, affecting the accuracy of removing non-tumor-derived mutation sites and leading to biased identification of mutation sites corresponding to ctDNA.
[0188] Since the second mutation detection strategy uses one or more heuristic conditions and local re-alignment algorithms to perform multiple screenings on the first sequencing data corresponding to cfDNA and leukocyte DNA, respectively, to determine the second and third mutation sites corresponding to cfDNA, as well as the second and third mutation sites corresponding to leukocyte DNA, the number of second and third mutation sites is usually greater than the number of first mutation sites detected by the first mutation detection strategy, that is, the second mutation detection strategy has higher sensitivity. However, due to the relatively loose threshold setting of the heuristic conditions and the high sensitivity of the local re-alignment algorithm to local regions of the sequence fragments, the second mutation detection strategy is more likely to introduce false-positive and false-negative mutation sites, thereby reducing the specificity of the second mutation detection strategy.
[0189] In some embodiments, the computer device can call VarDict software to compare and analyze the first sequencing data corresponding to cfDNA and white blood cell DNA with the reference sequencing data, respectively, to determine the second mutation detection result corresponding to the cfDNA and the second mutation detection strategy, as well as the second mutation detection result corresponding to the white blood cell DNA and the second mutation detection strategy.
[0190] In an embodiment of the present application, the computer device can accurately analyze each mutation site in the first sequencing data of cfDNA based on the first mutation detection strategy, accurately determine the mutation site of the specific mutation in the cfDNA, and make up for the defect of insufficient sensitivity in detecting specific mutations in the second mutation detection strategy. The computer device can perform in-depth mutation analysis on cfDNA and white blood cell DNA based on the second mutation detection strategy, accurately determine low-frequency or rare mutation sites in cfDNA and white blood cell DNA, and avoid the deficiency of the first mutation detection strategy that is unable to detect low-frequency or rare mutation sites due to insufficient detection sensitivity for low-frequency or rare mutations due to excessive specificity, thereby achieving complementarity between the first mutation detection strategy and the second mutation detection strategy, and thereby improving the accuracy of mutation detection in target samples.
[0191] In some embodiments, the computer device may filter the mutation detection results corresponding to each mutation detection strategy and merge them according to the corresponding filtering results to ensure the accuracy of the detection of each mutation site in cfDNA and white blood cell DNA. Figure 7 This is a flow chart of combining the filtering results provided in the embodiment of the present application to obtain the detection results corresponding to cfDNA and leukocyte DNA. Figure 7 As shown, the method may include the following steps:
[0192] Step 701 : Filter the first mutation detection result to remove first mutation sites in the first mutation detection result that do not meet a preset second quality requirement, so as to obtain a first filtered result.
[0193] The computer device can obtain the quality score corresponding to each first mutation site included in the first mutation detection result, and determine whether the quality score corresponding to each first mutation site meets the preset second quality requirement. If not, the corresponding first mutation site is removed; if so, the first mutation site is retained until all first mutation sites are detected. The computer device can use the remaining first mutation sites as the first filtering result.
[0194] It should be noted that the first quality requirement is used to ensure that the multiple sequence fragments included in the first sequencing data have a certain data quality, while the second quality requirement is used to ensure the confidence of the multiple first mutation sites included in the first mutation detection result. Therefore, even if the first sequencing data is obtained through the embodiments corresponding to steps 402 to 406, the computer device still needs to execute step 701 to perform a second filtering on the first mutation detection result.
[0195] In some embodiments, the first sequencing data further includes the sequencing depth and base recognition quality corresponding to each site of each sequence fragment, the quality score corresponding to each first mutation site may include the corresponding sequencing depth and base recognition quality, and the preset second quality requirement includes the sequencing depth being greater than or equal to the depth threshold, and the base recognition quality being greater than or equal to the quality threshold. Figure 8 As shown, step 701 may include the following steps:
[0196] Step 802: Obtain the sequencing depth or base recognition quality corresponding to the target first mutation site.
[0197] The target first mutation site may be any one of the multiple first mutation sites in the first mutation detection result.
[0198] Since the sequencer can obtain the sequencing depth and base recognition quality corresponding to each site during the process of sequencing cfDNA and white blood cell DNA, the computer equipment can directly obtain the sequencing depth or base recognition quality corresponding to the target first mutation site.
[0199] Step 804: Determine whether the sequencing depth corresponding to the target first mutation site is less than a depth threshold, or whether the base recognition quality corresponding to the target first mutation site is less than a quality threshold. If so, execute step 806; if not, execute step 808.
[0200] The sequencing depth corresponding to the target first mutation site can reflect the number of times the target first mutation site has been sequenced. If only a few sequence fragments have the target first mutation site, it means that these sequence fragments may have caused the target first mutation site to mutate due to various factors such as sequencing errors, but in fact the site has not mutated. Therefore, by setting the depth threshold, the first mutation sites corresponding to fewer sequence fragments can be filtered out.
[0201] The base recognition quality corresponding to the target first mutation site can reflect the probability that the base of the target first mutation site is correctly recognized. The lower the base recognition quality, the higher the probability that the base corresponding to the target first mutation site is misidentified, and the lower the authenticity of the mutation site. Therefore, by setting the quality threshold, those first mutation sites that may be misidentified due to low base recognition quality can be filtered out.
[0202] Step 806: remove the target first mutation site from the first mutation detection result.
[0203] Exemplarily, the first mutation detection result includes the first mutation site A, the first mutation site B, and the first mutation site C. The computer device may remove the first mutation site A from the first mutation detection result, so that the first mutation detection result after removal includes the first mutation site B and the first mutation site C.
[0204] Step 808 : Using the next first mutation site of the target first mutation site as a new target first mutation site, and re-performing step 802 .
[0205] The computer device may sequentially test each first mutation site in the first mutation detection result until all first mutation sites in the first mutation detection result have been tested. Alternatively, the computer device may simultaneously test each first mutation site in the first mutation detection result, but is not limited thereto.
[0206] Step 810 : When the target first mutation site is the last target first mutation site in the first mutation detection result, the remaining first mutation sites in the first mutation detection result are used as the first filtering result.
[0207] The first filtered results include multiple first mutation sites corresponding to cfDNA that meet the second quality requirement. Since only first mutation sites that do not meet the preset second quality requirement are removed from the first mutation test results, without making any formal changes to the first mutation sites or the first mutation test results, the first filtered results are actually a subset of the first mutation test results.
[0208] Exemplarily, after removing the first mutation site A, the computer device may use the remaining first mutation site B and the first mutation site C as the first filtering result.
[0209] In some embodiments, since the first mutation detection result corresponding to the white blood cell DNA is the same as the first mutation detection result corresponding to the cfDNA, the first filtering result corresponding to the cfDNA is also the first filtering result corresponding to the white blood cell DNA.
[0210] By judging whether the preset second quality requirement is met based on the sequencing depth and base recognition quality corresponding to each first mutation site in the first mutation detection result, and thus determining whether to be filtered, first mutation sites with low authenticity can be further filtered out, which can effectively improve the accuracy and credibility of mutation detection.
[0211] Step 703 : Filter the second mutation detection result according to the first filtering result, and remove the mutation sites in the second mutation detection result that are consistent with the first filtering result, so as to obtain a second filtering result.
[0212] The computer device can filter the second mutation detection result through the first filtering result, remove the mutation sites in the second mutation detection result that are consistent with the first filtering result, avoid repeated filtering of mutation sites with consistent results, and at the same time, perform targeted filtering on mutation sites that are inconsistent with the first filtering result, thereby improving the filtering efficiency and the accuracy of the filtering results.
[0213] In some embodiments, the computer device may match the mutation sites in the second mutation detection result that are inconsistent with the first filtering result with the reference sequence set to further determine whether the mutation sites that are inconsistent with the first filtering result belong to the sequence fragment corresponding to the ctDNA.
[0214] like Figure 9 As shown, step 703 may include the following steps:
[0215] Step 901 : Compare the first filtering result and the second mutation detection result, and remove the mutation sites in the second mutation detection result that are the same as those in the first filtering result to obtain a third mutation detection result.
[0216] The third mutation detection result may include mutation sites with specific mutations in the first sequencing data corresponding to cfDNA. The third mutation detection result is a subset of the second mutation detection result. It should be noted that the mutation sites with specific mutations in the first mutation detection result are determined by comparing and analyzing the first sequencing data corresponding to cfDNA and white blood cell DNA, respectively, while the mutation sites with specific mutations in the third mutation detection result are determined by comparing and analyzing the first sequencing data corresponding to cfDNA and white blood cell DNA, respectively, with the reference sequencing data.
[0217] Step 903 : Compare the third mutation detection result with a preset reference sequence set, determine fourth mutation sites in the third mutation detection result that match the reference sequence set, and use each fourth mutation site as a second filtering result.
[0218] In some embodiments, the preset reference sequence set may include a personalized panel, and the computer device may compare the third mutation detection result with the personalized panel, and use the fourth mutation site in the third mutation detection result that matches the reference sequence set as the second filtering result.
[0219] By comparing the fourth mutation site with specific mutations in the third mutation detection results with the personalized panel, the low-frequency mutations or rare mutations screened by the second mutation detection strategy can be preliminarily verified in advance to determine whether there are mutation sequence fragments corresponding to ctDNA, thereby reducing subsequent unnecessary verification steps and improving the overall detection efficiency of the target sample.
[0220] Step 705: Convert the first filtering result and the second filtering result corresponding to the cfDNA into a unified data format.
[0221] Since a computer device can implement mutation detection for the first sequencing data corresponding to cfDNA and white blood cell DNA respectively by calling scripts, software or software modules corresponding to multiple mutation detection strategies, and the data format corresponding to the mutation detection results corresponding to each mutation detection strategy may be different, filtering the mutation detection results does not change the data format of the mutation detection results, that is, the first filtering result has the same data format as the corresponding first mutation detection result, and the second filtering result has the same data format as the corresponding second mutation detection result. Therefore, when the data format corresponding to the mutation detection results corresponding to each mutation detection strategy is different, it is necessary to convert the data format corresponding to the mutation detection results or the data format corresponding to the filtering results into a unified data format to facilitate processing in subsequent steps.
[0222] In some embodiments, the computer device may convert the first filtering result and the second filtering result corresponding to the cfDNA into a unified VCF file.
[0223] Step 707: Combine the format-converted first filtering result and the second filtering result to obtain a first detection result corresponding to cfDNA.
[0224] In some embodiments, the computer device can call Bcftools software to merge the VCF result file corresponding to the first filtering result and the VCF result file corresponding to the second filtering result to obtain a VCF result file of the first detection result corresponding to cfDNA.
[0225] In an embodiment of the present application, the computer device may adopt different filtering methods for different mutation detection strategies before merging cfDNA with the mutation detection results corresponding to each mutation detection strategy, and combine the filtering results corresponding to each mutation detection strategy to obtain a first detection result corresponding to the cfDNA. This can not only further filter out false-positive mutation sites and false-negative mutation sites, improve the accuracy and credibility of the mutation detection results, but also integrate the advantages of different mutation detection strategies, make up for the shortcomings of a single mutation detection strategy, and make the final detection result more reliable.
[0226] In some embodiments, after merging the mutation detection results corresponding to different mutation detection strategies into corresponding detection results, the computer device may determine the target detection result corresponding to the target sample based on the first detection result, the second detection result, and a preset reference sequence set. Figure 10 This is a flow chart of determining the target detection result corresponding to the target sample provided in the embodiment of the present application. Figure 10 As shown, the step of determining the target detection result corresponding to the target sample based on the first detection result, the second detection result and the preset reference sequence set includes the following steps:
[0227] Step 1002: Annotate the first test result and the second test result to obtain initial annotation results corresponding to cfDNA and leukocyte DNA, respectively.
[0228] The initial annotation results include multiple mutation sites and the gene names corresponding to the sequence fragments where the mutation sites are located.
[0229] Because the first and second test results only include information about the mutation site and its location, and do not include further information such as the mutation site and the sequence fragment where the mutation site is located, or the impact of the mutation on the genomic region where the mutation site is located, when inspectors review the first and second test results, they are usually forced to manually query the test results to determine the biological information corresponding to the mutation site, which increases the complexity and time consumption of testing and analysis of target samples. Therefore, in the process of determining the target test result corresponding to the target sample, the first and second test results can be annotated to automatically obtain and integrate the biological information corresponding to each mutation site, thereby providing more direct and comprehensive information support for subsequent testing and analysis.
[0230] In some embodiments, the computer device can call Annovar software to annotate the VCF result files corresponding to the first test result and the second test result, respectively, to obtain initial annotation results corresponding to cfDNA and leukocyte DNA, respectively.
[0231] Optionally, the initial annotation results may also include refGene and cytoBand, etc., wherein refGene may include annotation results such as gene position and structure, mutation function classification and transcript information related to mutation based on the refGene database, and cytoBand may include annotation results such as cytogenetic bands annotated based on the cytoBand database.
[0232] Step 1004 , determining the same mutation sites in the initial annotation results corresponding to cfDNA and leukocyte DNA, respectively, and removing the same mutation sites from the initial annotation results corresponding to cfDNA to obtain the target annotation results corresponding to cfDNA.
[0233] Step 1006: Based on the first sequencing data corresponding to the cfDNA, the mutation data corresponding to each mutation site in the target annotation result is counted.
[0234] Mutation data may include the number of mutation sequences and mutation frequency, where the number of mutation sequences may refer to the total number of sequence fragments where a mutation site is located, which can reflect the coverage and support level of the mutation site in the target sample.
[0235] If a mutation site is supported by a large number of sequence fragments, that is, the number of mutation sequences is high, it means that there are multiple independent sequence fragments in the target sample that simultaneously support the same mutation site, reducing the possibility of false positive mutation sites due to sequencing errors or alignment errors, then this mutation site is more likely to be true.
[0236] Step 1008 : Based on the mutation condition, the mutation sites whose mutation data meet the mutation condition are screened from the target annotation results.
[0237] The mutation condition may refer to the number of mutated sequences at the mutation site being greater than a preset sequence number threshold, and / or the mutation frequency at the mutation site being greater than a preset frequency threshold.
[0238] In some embodiments, the computer device may determine whether the number of mutant sequences at a mutation site in the target annotation result is greater than a preset sequence number threshold, or determine whether the mutation frequency of a mutation site in the target annotation result is greater than a preset frequency threshold. If so, it is determined that the mutation data of the mutation site in the target annotation result meets the mutation condition, that is, the mutation site is a real mutation. If not, it is determined that the mutation data of the mutation site in the target annotation result does not meet the mutation condition, and the mutation site is removed.
[0239] For example, when the sequencing depth is 200,000X, the computer device can detect each mutation site in the target annotation result based on the mutation conditions with a sequence number threshold of 2 and a frequency threshold of 0.0005%, so as to screen out mutation sites whose mutation data meet the mutation conditions from the target annotation results.
[0240] Step 1010: Compare each mutation site that meets the mutation condition with a preset reference sequence set to determine a fifth mutation site that matches the reference sequence set.
[0241] The fifth mutation site may refer to a mutation site in the sequence fragment corresponding to ctDNA.
[0242] In some embodiments, the computer device may compare each mutation site that meets the mutation condition with the personalized panel to determine a fifth mutation site that matches the personalized panel among each mutation site that meets the mutation condition.
[0243] Since the personalized panel is designed based on the genetic background of the target sample and the known mutation sequence corresponding to the ctDNA, the personalized panel can accurately detect whether there is a mutation site corresponding to the ctDNA in each mutation site that meets the mutation conditions.
[0244] In some embodiments, when a fourth mutation site exists among all mutation sites that meet the mutation conditions, the computer device may directly determine the fourth mutation site as the fifth mutation site without the need for comparison and matching with the personalized panel.
[0245] Step 1012: Obtain the gene name corresponding to the sequence fragment where each fifth mutation site is located to obtain the target detection result corresponding to the target sample.
[0246] Target detection results can include positive and negative results with the name of the mutated gene. A positive result can indicate the presence of a sequence fragment corresponding to ctDNA in the target sample, while a negative result can indicate the absence of a sequence fragment corresponding to ctDNA in the target sample, or the presence of a sequence fragment corresponding to ctDNA that is less than the sensitivity of various mutation detection strategies.
[0247] The computer device can determine that a sequence fragment corresponding to ctDNA exists in the target sample when it is determined that a fifth mutation site exists in cfDNA, and determine that a target detection result corresponding to the target sample is a positive result; the computer device can determine that a sequence fragment corresponding to ctDNA does not exist in the target sample when it is determined that a fifth mutation site does not exist in cfDNA, and determine that a target detection result corresponding to the target sample is a negative result.
[0248] In some embodiments, when it is determined that a fifth mutation site exists in multiple sequence fragments corresponding to the cfDNA of the target sample, the computer device can generate a target detection result that MRD exists in the target sample and that the MRD in the target sample is caused by a mutation in the genomic region of the gene name corresponding to the sequence fragment where each fifth mutation site is located.
[0249] Table 1
[0250]
[0251] For example, Table 1 shows the target detection results corresponding to the target samples based on the first detection results corresponding to the cfDNA of the four target samples, the second detection results corresponding to the leukocyte DNA, and the personalized panel. As shown in Table 1, the first target sample detected a total of three mutated genes corresponding to the mutation sites corresponding to the ctDNA with a mutation frequency greater than 0.0005%, so the target detection result corresponding to the first target sample can be determined to be positive. The second target sample did not detect the mutated genes corresponding to the mutation sites corresponding to the ctDNA with a mutation frequency greater than 0.005%, so the target detection result corresponding to the second target sample can be determined to be negative. Similarly, the third and fourth target samples both have mutated genes corresponding to the mutation sites corresponding to the ctDNA with a mutation frequency greater than 0.005%, so the target detection results corresponding to the third and fourth target samples can be determined to be positive.
[0252] In an embodiment of the present application, by annotating the test results and then performing deduplication, screening, and matching on the annotation results, the fifth mutation site in the cfDNA that matches the reference sequence set is finally determined, and the gene name corresponding to the sequence fragment where each fifth mutation site is located is used to obtain the target detection result corresponding to the target sample. This not only allows for more accurate and comprehensive fifth mutation sites to be obtained by combining the mutation detection results corresponding to multiple mutation detection strategies, but also allows for accurate judgment of the attribution of the fifth mutation site based on the match between the reference sequence set and the fifth mutation site, thereby improving the accuracy of mutation detection for the target sample.
[0253] Based on the method for processing sequencing data provided in the above embodiment, Figure 11 This is a structural block diagram of a sequencing data processing device provided in an embodiment of the present application. Figure 11 As shown, in one embodiment, a sequencing data processing device 1100 is provided. The sequencing data processing device 1100 can be applied to a computer device. The sequencing data processing device includes a data acquisition module 1101, a mutation detection module 1102, a data merging module 1103 and a result acquisition module 1104.
[0254] The data acquisition module 1101 is used to obtain first sequencing data corresponding to circulating free deoxyribonucleic acid (cfDNA) and leukocyte DNA of the target sample.
[0255] The mutation detection module 1102 is used to detect the first sequencing data corresponding to cfDNA and white blood cell DNA according to multiple mutation detection strategies, and obtain the mutation detection results corresponding to cfDNA and each mutation detection strategy, as well as the mutation detection results corresponding to white blood cell DNA and each mutation detection strategy.
[0256] The data merging module 1103 is used to merge the cfDNA with the mutation detection results corresponding to each mutation detection strategy to obtain a first detection result corresponding to the cfDNA.
[0257] The data merging module 1103 is further configured to merge the white blood cell DNA with the mutation detection results corresponding to each mutation detection strategy to obtain a second detection result corresponding to the white blood cell DNA.
[0258] The result acquisition module 1104 is configured to determine a target detection result corresponding to the target sample based on the first detection result, the second detection result, and a preset reference sequence set.
[0259] In some embodiments, the sequencing data processing apparatus 1100 further includes a sequence extraction module and a data deduplication module.
[0260] The data acquisition module 1101 is also used to obtain second sequencing data corresponding to the cfDNA and white blood cell DNA corresponding to the target sample.
[0261] The sequence extraction module is used to extract the unique molecular marker UMI sequences corresponding to the cfDNA and the white blood cell DNA from the second sequencing data corresponding to the cfDNA and the white blood cell DNA respectively.
[0262] The data deduplication module is used to deduplicate the second sequencing data corresponding to the cfDNA based on the UMI sequence corresponding to the cfDNA to obtain the first sequencing data corresponding to the cfDNA; and to deduplicate the second sequencing data corresponding to the white blood cell DNA based on the UMI sequence corresponding to the white blood cell DNA to obtain the first sequencing data corresponding to the white blood cell DNA.
[0263] In some embodiments, the sequencing data processing apparatus 1100 further includes a calculation module and an evaluation module.
[0264] The calculation module is used to perform sequence comparison between the second sequencing data corresponding to the cfDNA and the reference sequencing data to obtain a first comparison result corresponding to the cfDNA, and perform data analysis on the first comparison result corresponding to the cfDNA to obtain a first quality assessment parameter corresponding to the cfDNA.
[0265] The calculation module is also used to perform sequence comparison between the second sequencing data corresponding to the white blood cell DNA and the reference sequencing data to obtain a first comparison result corresponding to the white blood cell DNA, and perform data analysis on the first comparison result corresponding to the white blood cell DNA to obtain a first quality assessment parameter corresponding to the white blood cell DNA.
[0266] The calculation module is also used to perform sequence comparison between the first sequencing data corresponding to the cfDNA and the reference sequencing data to obtain a second comparison result corresponding to the cfDNA, and perform data analysis on the second comparison result corresponding to the cfDNA to obtain a second quality assessment parameter corresponding to the cfDNA.
[0267] The calculation module is also used to perform sequence comparison between the first sequencing data corresponding to the white blood cell DNA and the reference sequencing data to obtain a second comparison result corresponding to the white blood cell DNA, and perform data analysis on the second comparison result corresponding to the white blood cell DNA to obtain a second quality assessment parameter corresponding to the white blood cell DNA.
[0268] The evaluation module is used to re-acquire the second sequencing data corresponding to the cfDNA and the leukocyte DNA of the target sample if the first quality evaluation parameters corresponding to the cfDNA and the leukocyte DNA and / or the second quality evaluation parameters corresponding to the cfDNA and the leukocyte DNA respectively do not meet the preset first quality requirements.
[0269] In some embodiments, the first sequencing data includes multiple sequence fragments, each sequence fragment includes multiple sites; the multiple mutation detection strategies include a first mutation detection strategy and a second mutation detection strategy.
[0270] The mutation detection module 1102 is also used to compare and analyze the first sequencing data corresponding to cfDNA and white blood cell DNA, respectively, based on the first mutation detection strategy, to obtain multiple candidate mutation sites, and to screen the multiple candidate mutation sites through a classification model to obtain multiple first mutation sites, and to use the multiple first mutation sites as the first mutation detection results corresponding to the cfDNA and the first mutation detection strategy.
[0271] The mutation detection module 1102 is also used to calculate the mutation frequency of each site in the first sequencing data corresponding to the cfDNA based on the second mutation detection strategy, screen out the second mutation site whose mutation frequency meets the frequency condition, detect the target sequence fragments with small insertions or deletions in the first sequencing data corresponding to the cfDNA, perform local realignment on the target sequence fragments, obtain the third mutation site whose mutation frequency meets the frequency condition in the target sequence fragments, and use the second mutation site and the third mutation site as the second mutation detection result corresponding to the cfDNA.
[0272] In some embodiments, the sequencing data processing apparatus 1100 further includes a filtering module.
[0273] The filtering module is used to filter the first mutation detection result and remove the first mutation sites in the first mutation detection result that do not meet the preset second quality requirement to obtain a first filtered result.
[0274] The filtering module is further configured to filter the second mutation detection result according to the first filtering result, and remove mutation sites in the second mutation detection result that are consistent with the first filtering result, so as to obtain a second filtering result.
[0275] The data merging module 1103 is further configured to convert the first filtering result and the second filtering result corresponding to the cfDNA into a unified data format.
[0276] The data merging module 1103 is further configured to combine the format-converted first filtering result and the second filtering result to obtain a first detection result corresponding to cfDNA.
[0277] In some embodiments, the first sequencing data also includes the sequencing depth and base recognition quality corresponding to each site of each sequence fragment, and the preset second quality requirements include the sequencing depth being greater than or equal to the depth threshold, and the base recognition quality being greater than or equal to the quality threshold.
[0278] The filtering module is further configured to remove the target first mutation site from the first mutation detection result if the sequencing depth corresponding to the target first mutation site in the first mutation detection result is less than a depth threshold or the corresponding base recognition quality is less than a quality threshold; the target first mutation site is any one of the multiple first mutation sites in the first mutation detection result.
[0279] The filtering module is further configured to use the remaining first mutation sites in the first mutation detection result as the first filtering result.
[0280] The filtering module is further configured to compare the first filtering result with the second mutation detection result, and remove the mutation sites in the second mutation detection result that are the same as those in the first filtering result, to obtain a third mutation detection result.
[0281] The filtering module is further configured to compare the third mutation detection result with the reference sequence set, determine fourth mutation sites in the third mutation detection result that match the reference sequence set, and use each fourth mutation site as a second filtering result.
[0282] In some embodiments, the result acquisition module 1104 includes an annotation submodule, a statistics submodule, and a generation submodule.
[0283] The annotation submodule is used to annotate the first test result and the second test result to obtain the initial annotation results corresponding to cfDNA and leukocyte DNA, respectively. The initial annotation results include the gene names corresponding to multiple mutation sites and the sequence fragments where the multiple mutation sites are located.
[0284] The filtering module is also used to determine the same mutation sites in the initial annotation results corresponding to cfDNA and leukocyte DNA, and remove the same mutation sites from the initial annotation results corresponding to cfDNA to obtain the target annotation results corresponding to cfDNA.
[0285] The statistical submodule is used to collect the mutation data corresponding to each mutation site in the target annotation result based on the first sequencing data corresponding to the cfDNA.
[0286] The filtering module is also used to filter out mutation sites whose mutation data meet the mutation conditions from the target annotation results according to the mutation conditions.
[0287] The filtering module is further configured to compare each mutation site that meets the mutation condition with a preset reference sequence set to determine a fifth mutation site that matches the reference sequence set.
[0288] The generation submodule is used to obtain the gene name corresponding to the sequence fragment where each fifth mutation site is located, so as to obtain the target detection result corresponding to the target sample.
[0289] Figure 12 This is a structural block diagram of a computer device provided in an embodiment of the present application. Figure 12 As shown, the computer device 1200 may include a memory 1202 and a processor 1201. The memory 1202 stores a computer program. When the computer program is executed by the processor 1201, the computer device 1200 implements the sequencing data processing method described in the above embodiments.
[0290] The processor 1201 may include one or more processing cores. The processor 1201 uses various interfaces and lines to connect the various parts of the entire computer device, and performs various functions of the computer device and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory. Optionally, the processor 1201 can be implemented in at least one hardware form of digital signal processing, field programmable gate array, and programmable logic array. The processor 1201 can integrate one or more combinations of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem is used to handle wireless communications. It is understandable that the above-mentioned modem may not be integrated into the processor 1201, but may be implemented separately through a communication chip.
[0291] Memory 1202 may include random access memory (RAM) or read-only memory (ROM). Memory may be used to store instructions, programs, codes, code sets, or instruction sets. Memory may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described above, and the like. The data storage area may also store data created during use of the computer device.
[0292] An embodiment of the present application discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor implements the sequencing data processing method described in the above embodiments.
[0293] An embodiment of the present application discloses a computer program product, which includes a computer program. When the computer program can be executed by a processor, the processor implements the sequencing data processing method described in the above embodiments.
[0294] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a ROM, or the like.
[0295] The above description is only a specific example of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A computer device, characterized in that: The system comprises a memory and a processor, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the processor performs the following steps: Obtaining first sequencing data corresponding to circulating free deoxyribonucleic acid (cfDNA) and leukocyte DNA of the target sample; Detecting the first sequencing data corresponding to the cfDNA and the white blood cell DNA, respectively, according to multiple mutation detection strategies, to obtain mutation detection results corresponding to the cfDNA and each of the mutation detection strategies, and mutation detection results corresponding to the white blood cell DNA and each of the mutation detection strategies; Combining the cfDNA with the mutation detection results corresponding to each of the mutation detection strategies to obtain a first detection result corresponding to the cfDNA; combining the white blood cell DNA with the mutation detection results corresponding to each of the mutation detection strategies to obtain a second detection result corresponding to the white blood cell DNA; Determine a target detection result corresponding to the target sample based on the first detection result, the second detection result, and a preset reference sequence set.
2. The computer device according to claim 1, wherein: The obtaining of first sequencing data corresponding to circulating free deoxyribonucleic acid (cfDNA) and leukocyte DNA of the target sample includes: Obtaining second sequencing data corresponding to the cfDNA and leukocyte DNA corresponding to the target sample; Extracting, from the second sequencing data corresponding to the cfDNA and the leukocyte DNA, the unique molecular marker UMI sequences corresponding to the cfDNA and the leukocyte DNA, respectively; The second sequencing data corresponding to the cfDNA is deduplicated based on the UMI sequence corresponding to the cfDNA to obtain the first sequencing data corresponding to the cfDNA; and the second sequencing data corresponding to the white blood cell DNA is deduplicated based on the UMI sequence corresponding to the white blood cell DNA to obtain the first sequencing data corresponding to the white blood cell DNA.
3. The computer device according to claim 2, wherein: Before obtaining the second sequencing data corresponding to the cfDNA and the leukocyte DNA corresponding to the target sample, the processor is further configured to perform the following steps: performing sequence alignment on the second sequencing data corresponding to the cfDNA and the reference sequencing data to obtain a first alignment result corresponding to the cfDNA, and performing data analysis on the first alignment result corresponding to the cfDNA to obtain a first quality assessment parameter corresponding to the cfDNA; performing sequence alignment on the second sequencing data corresponding to the white blood cell DNA and the reference sequencing data to obtain a first alignment result corresponding to the white blood cell DNA, and performing data analysis on the first alignment result corresponding to the white blood cell DNA to obtain a first quality assessment parameter corresponding to the white blood cell DNA; After performing deduplication processing on the second sequencing data corresponding to the cfDNA and the white blood cell DNA respectively according to the UMI sequence to obtain the first sequencing data corresponding to the cfDNA and the white blood cell DNA respectively, the processor is further configured to perform the following steps: performing sequence alignment on the first sequencing data corresponding to the cfDNA and the reference sequencing data to obtain a second alignment result corresponding to the cfDNA, and performing data analysis on the second alignment result corresponding to the cfDNA to obtain a second quality assessment parameter corresponding to the cfDNA; performing sequence alignment on the first sequencing data corresponding to the white blood cell DNA and the reference sequencing data to obtain a second alignment result corresponding to the white blood cell DNA, and performing data analysis on the second alignment result corresponding to the white blood cell DNA to obtain a second quality assessment parameter corresponding to the white blood cell DNA; Before detecting the first sequencing data corresponding to the cfDNA and the leukocyte DNA respectively according to the multiple mutation detection strategies, the processor is further configured to perform the following steps: If the first quality assessment parameters corresponding to the cfDNA and the white blood cell DNA respectively and / or the second quality assessment parameters corresponding to the cfDNA and the white blood cell DNA respectively do not meet the preset first quality requirements, the second sequencing data corresponding to the cfDNA and the white blood cell DNA corresponding to the target sample are re-obtained.
4. The computer device according to claim 1, wherein: The first sequencing data includes multiple sequence fragments, each of which includes multiple sites; the multiple mutation detection strategies include a first mutation detection strategy and a second mutation detection strategy; Detecting the first sequencing data corresponding to the cfDNA according to multiple mutation detection strategies to obtain mutation detection results corresponding to the cfDNA and each of the mutation detection strategies, including: Based on the first mutation detection strategy, comparatively analyzing the first sequencing data corresponding to the cfDNA and the white blood cell DNA, respectively, to obtain a plurality of candidate mutation sites, screening the plurality of candidate mutation sites using a classification model to obtain a plurality of first mutation sites, and using the plurality of first mutation sites as first mutation detection results corresponding to the cfDNA and the first mutation detection strategy; Based on the second mutation detection strategy, the mutation frequency of each site in the first sequencing data corresponding to the cfDNA is calculated, and the second mutation site whose mutation frequency meets the frequency condition is screened out. The target sequence fragments with small insertions or deletions in the first sequencing data corresponding to the cfDNA are detected, and the target sequence fragments are locally realigned to obtain the third mutation site in the target sequence fragment whose mutation frequency meets the frequency condition. The second mutation site and the third mutation site are used as the second mutation detection result corresponding to the cfDNA.
5. The computer device according to claim 4, wherein: Before combining the cfDNA with the mutation detection results corresponding to each mutation detection strategy to obtain a first detection result corresponding to the cfDNA, the processor is further configured to perform the following steps: Filtering the first mutation detection results to remove first mutation sites that do not meet a preset second quality requirement in the first mutation detection results, to obtain a first filtered result; Filtering the second mutation detection result according to the first filtering result, removing mutation sites in the second mutation detection result that are consistent with the first filtering result, to obtain a second filtering result; The step of combining the cfDNA with the mutation detection results corresponding to each of the mutation detection strategies to obtain a first detection result corresponding to the cfDNA includes: Converting the first filtering result and the second filtering result corresponding to the cfDNA into a unified data format; The format-converted first filtering result and the second filtering result are combined to obtain a first detection result corresponding to the cfDNA.
6. The computer device according to claim 5, wherein: The first sequencing data further includes a sequencing depth and a base recognition quality corresponding to each site of each sequence fragment, the preset second quality requirement includes a sequencing depth greater than or equal to a depth threshold, and a base recognition quality greater than or equal to a quality threshold; and filtering the first mutation detection results to remove first mutation sites that do not meet the preset second quality requirement in the first mutation detection results to obtain a first filtering result, including: If the sequencing depth corresponding to the target first mutation site in the first mutation detection result is less than the depth threshold or the corresponding base recognition quality is less than the quality threshold, then the target first mutation site is removed from the first mutation detection result; the target first mutation site is any one of the multiple first mutation sites in the first mutation detection result; The remaining first mutation sites in the first mutation detection result are used as the first filtering result.
7. The computer device according to claim 5, wherein: The filtering the second mutation detection result according to the first filtering result, removing mutation sites in the second mutation detection result that are consistent with the first filtering result, to obtain a second filtering result, includes: Comparing the first filtering result with the second mutation detection result, and removing the same mutation sites in the second mutation detection result as in the first filtering result, to obtain a third mutation detection result; The third mutation detection result is compared with a preset reference sequence set to determine a fourth mutation site in the third mutation detection result that matches the reference sequence set, and each of the fourth mutation sites is used as a second filtering result.
8. The computer device according to claim 1, wherein: Determining a target detection result corresponding to the target sample based on the first detection result, the second detection result, and a preset reference sequence set includes: Annotating the first test result and the second test result to obtain initial annotation results corresponding to the cfDNA and the leukocyte DNA, respectively, the initial annotation results including multiple mutation sites and gene names corresponding to the sequence fragments where the multiple mutation sites are located; Determining identical mutation sites in the initial annotation results corresponding to the cfDNA and the leukocyte DNA, respectively, and removing the identical mutation sites from the initial annotation results corresponding to the cfDNA to obtain a target annotation result corresponding to the cfDNA; According to the first sequencing data corresponding to the cfDNA, counting the mutation data corresponding to each mutation site in the target annotation result; According to the mutation condition, the mutation sites whose mutation data meet the mutation condition are screened from the target annotation results; Comparing each mutation site that meets the mutation condition with a preset reference sequence set to determine a fifth mutation site that matches the reference sequence set; Obtain the gene name corresponding to each sequence fragment where the fifth mutation site is located to obtain the target detection result corresponding to the target sample.
9. A method for processing sequencing data, characterized in that: Applied to a computer device, the method comprises: The computer device obtains first sequencing data corresponding to circulating free deoxyribonucleic acid (cfDNA) and leukocyte DNA of the target sample respectively; The computer device detects the first sequencing data corresponding to the cfDNA and the white blood cell DNA respectively according to multiple mutation detection strategies, and obtains the mutation detection results corresponding to the cfDNA and each of the mutation detection strategies, and the mutation detection results corresponding to the white blood cell DNA and each of the mutation detection strategies; The computer device combines the cfDNA with the mutation detection results corresponding to each of the mutation detection strategies to obtain a first detection result corresponding to the cfDNA; The computer device combines the white blood cell DNA with the mutation detection results corresponding to each of the mutation detection strategies to obtain a second detection result corresponding to the white blood cell DNA; The computer device determines a target detection result corresponding to the target sample based on the first detection result, the second detection result and a preset reference mutation result.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for processing sequencing data according to claim 9 is implemented.