Method and device for detecting minimal residual lesion by using tumor information

By acquiring sequencing data from tumor tissue, blood, and plasma samples, and combining genetic data filtering and tumor cell ratio calculation, the problem of high background error rate in NGS technology has been solved, enabling accurate detection and early discovery of minimal residual disease.

CN120917524APending Publication Date: 2025-11-07INOCRAS KOREA INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380095919.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-22
Filing Date
2023-11-08
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing next-generation DNA sequencing (NGS) technologies suffer from high background error rates when detecting low-incidence DNA variations, making it difficult to distinguish whether variations in the blood originate from cancer tissue or are falsely detected, thus making it difficult to detect residual tumor cells.

Method used

By acquiring sequencing data from patients' tumor tissue biopsy samples, normal blood samples, and plasma samples, and utilizing whole-genome sequencing technology, combined with genetic data filtering and tumor cell ratio calculation, the background error rate is reduced, enabling accurate detection of minimal residual disease.

Benefits of technology

It improves the accuracy and early detection capability of detecting minimal residual disease from non-invasive liquid biopsy samples, reduces computational resource and time requirements, and enables accurate prediction of tumor cell proportions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120917524A_ABST
    Figure CN120917524A_ABST
Patent Text Reader

Abstract

The invention relates to a minimal residual lesion detection method using tumor information. The minimal residual lesion detection method using tumor information comprises the following steps: acquiring first sequencing data related to a first sample of a patient; acquiring second sequencing data associated with a second sample of the patient; acquiring third sequencing data associated with a third sample of the patient; and performing minimal residual lesion detection on the patient based on the first sequencing data, the second sequencing data and the third sequencing data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a method and apparatus for detecting minimal residual disease using tumor information, and more particularly, to a method and apparatus for detecting minimal residual disease of a patient using variant information of a tumor tissue and a liquid biopsy sample. BACKGROUND

[0002] Genome analysis techniques are widely used in the medical field, for example, by identifying the genome of an organism to determine the characteristics or matrix of the organism. Recently, medical practices for treating various diseases such as tumors are gradually developing from a traditional prescription-centered approach to precision medicine, in other words, gradually developing into a customized treatment approach that takes into account each patient's genetic information and health records.

[0003] In the field of precision medicine, it is particularly important to obtain a large amount of personal genomic information and perform related clinical analysis. Recently, next-generation DNA sequencing (NGS) technology, which rapidly reads a large amount of DNA information in a parallel manner, is widely used in various medical fields such as cancer screening. Although the NGS technology has the advantage of relatively less cost and time required for analyzing a genome and convenience, there is an inherent limitation that the background error rate is about 0.1% to 1%. That is, even if there is no variation in DNA, there is a statistical limitation (false positive mutation) that one base pair out of 100 to 1000 base pairs is erroneously determined to have a variation (i.e., an error rate of 0.1% to 1%) during the sequencing process.

[0004] Due to the inherent limitation of the NGS technology, the NGS technology has limitations in detecting low-occurrence DNA variations. For example, the concentration of cancer-derived DNA (e.g., ctDNA: Circulation tumor DNA) in the blood of a cancer patient is usually less than 1%, however, it is difficult to distinguish whether the variation detected from the blood is derived from the cancer tissue or is an erroneous detection due to the background error rate using the NGS technology. Therefore, there is a need to introduce a new technology that detects low-occurrence DNA variations (e.g., tumor cell residues, etc.) in the blood of a patient by reducing the background error rate. SUMMARY TECHNICAL PROBLEM

[0005] To solve the problems as described above, the present application aims to provide a method for detecting minimal residual disease using tumor information. A non-transitory computer-readable recording medium having instructions recorded thereon and an apparatus (system). Technical Solution

[0006] The present application can be implemented in various ways including methods, systems (apparatuses), and / or non-transitory computer-readable recording media having instructions recorded thereon.

[0007] The minimal residual disease detection method of an embodiment of the present application is executed by at least one processor, and includes the steps of obtaining first sequencing data related to a first sample of a patient, obtaining second sequencing data related to a second sample of the patient, obtaining third sequencing data related to a third sample of the patient, and performing minimal residual disease detection on the patient based on the first, second, and third sequencing data.

[0008] According to an embodiment of the present application, the first, second, and third sequencing data can be obtained through whole genome sequencing (WGS).

[0009] According to an embodiment of the present application, the first sample is a tumor tissue biopsy sample of the patient, the second sample is a normal blood sample of the patient, and the third sample is a plasma sample of the patient, the plasma sample including cell free Deoxyribo Nucleic Acid (cfDNA) and circulating tumor Deoxyribo Nucleic Acid (ctDNA).

[0010] According to an embodiment of the present application, the first and second samples are samples obtained at a first time point, and the third sample is a sample obtained at a second time point after the first time point, at least one of a surgery or a treatment therapy being performed on the patient between the first time point and the second time point.

[0011] According to an embodiment of the present application, the minimal residual disease detection method further includes the steps of detecting tumor tissue variant information of the patient by comparing the first and second sequencing data, and performing background error filtering on the detected tumor tissue variant information using genetic data related to a plurality of sample patients who are different from the patient.

[0012] According to an embodiment of the present application, the genetic data comprises sequencing data of normal blood samples of a plurality of sample patients, and the step of performing filtering comprises the step of removing, in the tumor tissue variant information, variants detected from the sequencing data of the normal blood samples of the plurality of sample patients, among the detected tumor tissue variants.

[0013] According to an embodiment of the present application, the genetic data comprises sequencing data of plasma samples of a plurality of sample patients, and the step of performing filtering comprises the step of removing, in the tumor tissue variant information, variants detected from the sequencing data of the plasma samples of the plurality of sample patients, among the detected tumor tissue variants.

[0014] According to an embodiment of the present application, the step of performing minimal residual disease detection comprises the steps of calculating a limit of detection value of the patient based on the first sequencing data, the second sequencing data and the third sequencing data; calculating a tumor cell fraction of the patient based on the third sequencing data; correcting the tumor cell fraction using genetic data related to a plurality of sample patients different from the patient; and determining whether the tumor of the patient has relapsed based on the corrected tumor cell fraction and the limit of detection value.

[0015] According to an embodiment of the present application, the limit of detection value indicates a minimum tumor cell fraction that can be detected from a plasma sample of the patient.

[0016] According to an embodiment of the present application, the limit of detection value is calculated based on tumor tissue variants of the patient detected by comparing the first sequencing data and the second sequencing data, and an average sequencing depth of the third sequencing data.

[0017] According to an embodiment of the present application, the step of calculating the tumor cell fraction comprises the steps of determining a number of sequencing reads different from a reference sequence in the third sequencing data; determining a number of sequencing reads identical to the reference sequence in the third sequencing data; and calculating the tumor cell fraction based on the number of sequencing reads different from the reference sequence and the number of sequencing reads identical to the reference sequence.

[0018] According to an embodiment of the present application, the step of determining the number of different sequencing reads comprises the steps of detecting tumor tissue variant information of the patient by comparing the first sequencing data and the second sequencing data; determining a number of sequencing reads including variants detected from the tumor tissue in the third sequencing data based on the tumor tissue variant information; and using the number of sequencing reads including variants detected from the tumor tissue in the third sequencing data as the number of sequencing reads different from the reference sequence.

[0019] According to an embodiment of the present application, the step of correcting the tumor cell fraction includes the steps of: determining a number of sequencing fragments different from a reference sequence in plasma sequencing data included in the genetic data associated with the plurality of sample patients other than the patient; determining a number of sequencing fragments identical to the reference sequence in the plasma sequencing data included in the genetic data associated with the plurality of sample patients other than the patient; calculating a random error rate based on the number of sequencing fragments different from the reference sequence and the number of sequencing fragments identical to the reference sequence; and correcting the tumor cell fraction using the random error rate.

[0020] According to an embodiment of the present application, the step of determining whether the tumor of the patient has relapsed includes the steps of: calculating a confidence interval of a predetermined confidence level for the corrected tumor cell fraction; and in response to a determination that a lower limit of the calculated confidence interval is higher than the calculated detection limit value, determining that the tumor of the patient has relapsed.

[0021] According to an embodiment of the present application, the step of performing the minimal residual disease detection further includes the steps of: generating a first chromosomal arm-level copy number profile based on the first sequencing data; generating a second chromosomal arm-level copy number profile based on the second sequencing data; generating a third chromosomal arm-level copy number profile based on the third sequencing data; and calculating the tumor cell fraction based on the first chromosomal arm-level copy number profile to the third chromosomal arm-level copy number profile.

[0022] According to an embodiment of the present application, the step of performing the minimal residual disease detection further includes the steps of: identifying a first group of structural variations detected from the first sequencing data; identifying a second group of structural variations detected from the third sequencing data; and determining whether the tumor of the patient has relapsed based on a comparison result of the first group of structural variations and the second group of structural variations.

[0023] According to an embodiment of the present application, the step of determining whether the tumor of the patient has relapsed includes the steps of: in the first group of structural variations, if a number or a proportion of structural variations included in the second group of structural variations is equal to or greater than a threshold value, determining that the tumor of the patient has relapsed.

[0024] According to an embodiment of the present application, the first group of structural variations includes at least one of an inversion, a translocation, a duplication, a deletion, or an insertion of a partial region of a genome.

[0025] The non-transitory computer-readable recording medium according to an embodiment of the present disclosure records instructions for performing a method of detecting a minimal residual disease using tumor information.

[0026] The system according to an embodiment of the present disclosure includes a communication module, a memory, and at least one processor connected to the memory and configured to execute at least one program stored in the memory, the at least one program including instructions for acquiring first sequencing data related to a first sample of a patient, acquiring second sequencing data related to a second sample of the patient, acquiring third sequencing data related to a third sample of the patient, and performing a minimal residual disease detection for the patient based on the first sequencing data, the second sequencing data, and the third sequencing data. Effects of the Invention

[0027] According to the embodiments of the present disclosure, as the threshold for detecting a minimal residual disease (MRD) from a non-invasive liquid biopsy sample is lowered by reducing a background error rate through data filtering, tumor cells can be further rapidly detected at an early stage.

[0028] According to the embodiments of the present disclosure, as a proportion of tumor cells is calculated based on filtered tumor tissue variants, time and computational resources required for comparing overall plasma sequencing data with a reference sequence can be reduced.

[0029] According to the embodiments of the present disclosure, as tumor tissue variants of a target patient are filtered using genetic data of a sample patient, a proportion of tumor cells in the target patient can be very accurately predicted.

[0030] According to the embodiments of the present disclosure, as a proportion of tumor cells in a target patient is very accurately predicted using a chromosome arm-level copy number profile of the target patient.

[0031] The effects of the present disclosure are not limited to the above-mentioned effects, and other effects not mentioned herein can be clearly understood by those skilled in the art (referred to as "those skilled in the art") from the contents of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0032] Hereinafter, the embodiments of the present disclosure will be described with reference to the accompanying drawings, in which like reference numerals refer to like elements, but are not limited thereto.

[0033] Figure 1 FIG. 1 is a flowchart illustrating a process of performing a minimal residual disease detection for a target patient according to an embodiment of the present disclosure.

[0034] Figure 2FIG. 1 is a diagram showing an outline of a connection structure of an information processing system according to an embodiment of the present application for providing a micro residual lesion detection service using tumor information, and a plurality of user terminals.

[0035] Figure 3 FIG. 2 is a block diagram showing an internal structure of a user terminal and an information processing system according to an embodiment of the present application.

[0036] Figure 4 FIG. 3 is an example graph showing a change in a proportion of tumor cells in a target patient over time according to an embodiment of the present application.

[0037] Figure 5 FIG. 4 is a graph showing a blood sample of a target patient according to an embodiment of the present application.

[0038] Figure 6 FIG. 5 is a graph showing a process of calculating a detection limit value of a target patient according to an embodiment of the present application.

[0039] Figure 7 FIG. 6 is a graph showing a tumor detection process according to an embodiment of the present application.

[0040] Figure 8 FIG. 7 is an example graph showing a relationship between a variance detected from tumor cells of a target patient and a detection limit value of the target patient according to an embodiment of the present application.

[0041] Figure 9 FIG. 8 is an example graph showing generation of an arm-level copy number profile according to an embodiment of the present application.

[0042] Figure 10 FIG. 9 is an example graph showing calculation of a proportion of tumor cells in a target patient using an arm-level copy number profile according to an embodiment of the present application.

[0043] Figure 11 FIG. 10 is a graph showing a performance verification result of a method of detecting a micro residual lesion of a target patient using genetic data of a sample patient according to an embodiment of the present application.

[0044] Figure 12 FIG. 11 is a graph showing a performance verification result of a method of detecting a micro residual lesion of a target patient using an arm-level copy number profile of the target patient according to an embodiment of the present application.

[0045] Figure 13 FIG. 12 is a flowchart showing a micro residual lesion detection method using tumor information according to an embodiment of the present application. DETAILED DESCRIPTION

[0046] Hereinafter, specific contents for implementing the present application will be described in detail with reference to the accompanying drawings. However, in the following description, when there is a possibility that the gist of the present application can be unnecessarily confused, a specific description of well-known functions or structures will be omitted.

[0047] In the drawings, the same or similar structural elements are given the same reference numerals. Also, in describing the following embodiments, repeated description of the same or similar structural elements will be omitted. However, even if the description of the structural elements is omitted, it does not mean that such structural elements are not included in any of the embodiments.

[0048] Advantages, features and methods of implementing the disclosed embodiments will be apparent from the detailed disclosure, in conjunction with the drawings, which are described below. Figure 1 However, the present application is not limited to the disclosed embodiments, and can be implemented in various ways, and the embodiments are provided to ensure the completeness of the present disclosure, and the present disclosure is provided only to enable those of ordinary skill in the art to fully understand the scope of the present application.

[0049] In the present specification, the terms used will be briefly described and the disclosed embodiments will be specifically described. In the present specification, although common terms generally used at present are selected considering the functions of the present application, the terms used herein can be changed according to the intention of those of ordinary skill in the art, or the custom, the appearance of new technology, etc. Also, in a certain case, there are terms arbitrarily selected by the applicant, and in this case, the meanings are described in detail in the corresponding description part of the present application. Therefore, the terms used in the present specification should be defined based on the meanings of the terms and the content of the entire present specification, and not only the names of the terms.

[0050] In the present specification, unless the context clearly indicates otherwise, the expression of the singular includes the expression of the plural. Also, unless the context clearly indicates otherwise, the expression of the plural includes the expression of the singular. In the content of the entire present specification, when it is indicated that a certain structural element includes any structural element, unless there is a specific contrary description, it means that other structural elements can also be included, and does not exclude other structural elements.

[0051] Also, in the present specification, the term "module" or "part" refers to a software structural element or a hardware structural element, and the "module" or "part" performs certain functions. However, the "module" or "part" is not limited to software or hardware. The "module" or "part" can exist in a storage medium accessible to a processor, or can also reproduce one or more processors. Accordingly, as an example, the "module" or "part" can include at least one of a software structural element, an object-oriented software structural element, a class structural element, and a task structural element, a process, a function, an attribute, a program, a subroutine, a program code segment, a driver, firmware, a microcode, a circuit, data, a database, a data structure, a table, an array, or a variable. The structural elements and the "module" or "part" provided internally can be combined by a smaller number of structural elements and "modules" or "parts" or further separated into additional structural elements and "modules" or "parts".

[0052] According to an embodiment of the present application, the "module" or "part" can be implemented by a processor and a memory. The "processor" is broadly interpreted to include a general processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In various environments, the "processor" can also refer to an application specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), etc. For example, the "processor" can also refer to a combination of a digital signal processor and a microprocessor, a combination of a plurality of microprocessors, a combination of one or more microprocessors with a digital signal processing core, or a combination of other any combination of such processing devices. Also, the "memory" is broadly interpreted to include any electronic component that can store electronic information. The "memory" can also refer to various types of processor-readable media such as random access memory (RAM), read only memory (ROM), non-volatile random access memory (NVRAM), programmable read only memory (PROM), erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), flash memory, magnetic or optical data storage devices, buffers, etc. If the processor reads information from the memory or records information to the memory, the memory is in an electronic communication state with the processor. The memory integrated in the processor is in an electronic communication state with the processor.

[0053] In the present specification, “Whole Genome Sequencing (WGS)” or “Whole Genome Resequencing” refers to a technique for determining the entire sequence of DNA of an organism’s genome. Specifically, Whole Genome Sequencing is used to read and identify the order of nucleotide bases (adenine, cytosine, guanine, and thymine) in the entire genetic material set of a human or an organism. In this case, the entire genetic material set can include all of the genome, non-coding regions, and any additional genetic elements present within the genome. In an embodiment, Whole Genome Sequencing can be performed through a plurality of steps. For example, Whole Genome Sequencing can be performed by extracting DNA from a specific cell and breaking the extracted DNA into smaller fragments to generate millions or billions of short DNA sequences called “reads”. The generated reads can be reconstituted into a whole genome sequence by alignment and assembly. In a specific embodiment, a variety of duplex sequencing methods can be used to minimize the background error rate, whereby very small cancer-derived DNA can be detected from blood. In Whole Genome Sequencing, a Concatenating Original Duplex for Error Correction (CODEC) method, which is an example of duplex sequencing, can be used, and the present specification incorporates by reference PCT Publication No. WO 2022 / 125997 A1 published on June 16, 2022.

[0054] In the present specification, “DNA” is not limited to DNA itself, and can include any nucleic acid and nucleic acid sequence such as RiboNucleic Acid (RNA).

[0055] In the present specification, “circulating free DNA (cfDNA)” refers to circulating free DNA called cell free DNA or circulating DNA. cfDNA refers to small fragments of DNA released from blood and other fluids by cells undergoing normal apoptosis or as a result of specific pathological processes. cfDNA includes healthy cells and diseased cells, and can consist of short fragments of DNA originating from various tissues of the body. cfDNA can originate from organs such as the liver, heart, or lung, and can also originate from tumors (e.g., malignancies such as cancer). For example, after centrifugation of a blood sample, the plasma layer, buffy coat layer, and red blood cell layer can be obtained from the upper layer, and cfDNA can be obtained from the plasma layer.

[0056] In the present specification, “ctDNA” refers to circulating tumor DNA. In particular, ctDNA can be a subset of cfDNA originating from tumor cells. When apoptosis occurs or is experienced, tumor cells can release ctDNA into the blood. ctDNA can carry genomic alterations or mutations characteristic of tumors as its origin. ctDNA can provide useful information about the genetic properties and mutation profile of tumors without performing invasive tissue biopsies. With analysis of ctDNA, the treatment response of a patient can be monitored and minimal residual disease (MRD) can be detected.

[0057] In the present specification, “normal blood sample” or “buffy coat” refers to a specific structural element of a blood sample including white blood cells and platelets. For example, a blood sample can be centrifuged to obtain a normal blood sample. Among them, the normal blood sample can be obtained from a thin layer formed between the red blood cell layer present at the bottom and the plasma layer present at the upper end of the centrifuged sample. The normal blood sample includes white blood cells having nuclei including DNA, includes a higher concentration of nucleated cells, and can provide a source of genomic DNA from cells of a specific individual. Accordingly, when WGS is performed, DNA can be extracted from the normal blood sample to obtain genomic DNA of a specific individual, which can be used as reference DNA or control DNA, and can be used to compare and identify cfDNA or ctDNA extracted from the same person, genomic variations or mutations present in tumor tissue DNA extracted from the same person. It can be assumed that the normal blood sample does not contain cancer cells.

[0058] In the present specification, "sequencing depth" or sequencing coverage refers to the average number of times each base of a genome is read during sequencing. For example, sequencing depth can represent the number of times a particular nucleotide (A, T, C, or G) located at a particular position of a genome is sequenced. Sequencing depth is a parameter that affects accuracy and reliability. Generally, sequencing depth is measured by "X", where 1X coverage indicates that each base is sequenced an average of once, and 10X coverage indicates that each base is sequenced an average of 10 times.

[0059] In the present specification, "copy number analysis" refers to a technique of acquiring a copy number profile from WGS data by performing copy number analysis on a specific DNA fragment or region within a genome.

[0060] In an embodiment, copy number analysis can include a plurality of steps. First, a DNA isolation step of isolating DNA from a cell or tissue of interest can be performed. Next, after the extracted DNA is fragmented into smaller fragments, a sequencing step of performing sequencing on the DNA fragments can be performed. Through this step, a large number of sequencing fragments as short sequences of DNA can be generated. According to embodiments, a step of amplifying and duplicating DNA fragments using a technique such as a polymerase chain reaction (PCR) can be performed before the DNA is fragmented into a plurality of fragments and the sequencing step is performed.

[0061] Subsequently, a sequencing fragment alignment step of aligning the generated sequencing fragments to a reference sequence can be performed. This step can include a process of comparing the sequencing fragments to known DNA sequences of the reference sequence in order to determine the original positions of the sequencing fragments. If the sequencing fragments are aligned, a copy number analysis step of analyzing the sequencing coverage or depth of the whole genome can be performed using a software tool. Sequencing coverage can represent the number of sequencing fragments aligned to a specific genomic region. Then, a normalization step can be performed on the coverage data in order to account for variations in sequencing depth and other technical biases. Thereby, accurate comparisons between different genomic regions can be achieved.

[0062] Next, a computational algorithm can be applied to the normalized coverage data to identify a copy number calling step that represents genomic regions of copy number variation. Such an algorithm can compare the observed coverage to a predicted coverage based on a diploid genome. Subsequently, the copy number profile can be visualized using a plot or heatmap that shows genomic regions where there is an increase or loss of copy number, which can be referred to as a visualization step and an interpretation step. Such visualized data can provide insights related to genomic changes such as amplification, deletion, or duplication.

[0063] In the present specification, the "limit of detection value" refers to the lowest detection concentration at which the presence or absence of the analysis target substance can be confirmed in the analysis process of the subject substance. For example, the "limit of detection value" refers to the lowest proportion of tumor cells that can be detected from the plasma sample of a tumor patient.

[0064] Figure 1 An example of a process for performing minimal residual disease detection on a target patient 110 according to an embodiment of the present application. The target patient 110 refers to a patient having a tumor tissue (e.g., a cancer tissue). Also, the target patient 110 refers to a subject on whom minimal residual disease detection is performed after treatment 120, 130 or surgery on the tumor tissue.

[0065] According to an embodiment, at a first time point before treatment 120, 130 or surgery on the tumor tissue, a tumor tissue biopsy sample and a normal blood sample can be collected from the target patient 110. The normal blood sample can be obtained by performing centrifugal separation on blood collected from the target patient. Next, tumor tissue sequencing data 112 and normal blood sequencing data 114 can be generated based on the tumor tissue biopsy sample and the normal blood sample. Each of the sequencing data can be obtained by whole genome sequencing (WGS).

[0066] Subsequently, tumor tissue variation information of the target patient 110 can be detected based on the obtained tumor tissue sequencing data 112 and normal blood sequencing data 114. For the detected tumor tissue variation information, background error filtering can be performed using genetic data related to a plurality of sample patients different from the target patient 110. As the background error rate is reduced by performing background error filtering, it can be confirmed at a time point before the background error is performed whether the tumor of the target patient 110 has relapsed, etc. Later, refer toFigure 6 Detailed description of specific processes of detecting tumor tissue variant information and filtering tumor tissue variant information.

[0067] In an embodiment, after collecting the tumor tissue biopsy sample and the normal blood sample from the target patient 110, at least one of the treatment 120, 130 and / or the surgery 140 can be performed on the target patient 110. As the treatment 120, 130 and / or the surgery 140 is performed on the target patient 110, the proportion of tumor cells in the target patient 110 can be reduced. However, over time, the proportion of tumor cells in the target patient 110 can increase again.

[0068] To quickly respond to the recurrence of the disease caused by the increase of tumor cells in the body, it is particularly important to detect the increase of the proportion of tumor cells at an early stage. To this end, after the treatment 120, 130 and / or the surgery 140 is performed, a blood plasma sample of the target patient 110 can be collected. The blood plasma sample can be obtained by performing centrifugal separation on the blood 122, 132, 142, 146 collected from the target patient 110. In this case, the blood plasma sample can include cell free Deoxyribo Nucleic Acid (cfDNA) and circulating tumor Deoxyribo Nucleic Acid (ctDNA). The blood plasma sample can be collected from the target patient 110 at a specified period or on demand for analysis, and can be used to track whether the tumor has recurred. Specifically, blood plasma sample sequencing data 124, 134, 144, 148 can be generated based on the blood plasma sample included in the blood 122, 132, 142, 146, and whether the target patient 110 needs additional treatment and / or surgery can be determined based on the obtained blood plasma sample sequencing data 124, 134, 144, 148.

[0069] For example, the blood plasma sample sequencing data 124 can be obtained from the blood 122 collected at a second time point after the first treatment 120. Subsequently, a detection limit value of the tumor tissue of the target patient 110 can be calculated based on the tumor tissue variant information and the blood plasma sample sequencing data 124, and minimal residual disease detection of the target patient 110 can be performed. Later, with reference to Figure 6 Detailed description of specific processes of calculating the detection limit value of the tumor tissue of the target patient 110. Also, with reference to Figure 7 Detailed description of specific processes of performing minimal residual disease detection on the target patient 110.

[0070] After the first treatment 120, if a minimal residual disease is detected from the target patient 110 (or, in the case where it is judged that the tumor has relapsed), a second treatment 130 can be performed / undertaken. Then, at a third time point after the second treatment 130 is performed, blood 132 can be collected from the target patient 110 and the minimal residual disease detection can be repeatedly performed.

[0071] In yet another example, based on the plasma sample sequencing data 144 obtained from the blood 142 collected at a fourth time point after the surgery 140 is performed, it can be impossible to detect the minimal residual disease. In this case, it can be judged that the tumor of the target patient 110 has not relapsed, and no additional treatment and / or surgery is needed. Then, the minimal residual disease detection can be performed again based on the plasma sample sequencing data 148 obtained from the blood 146 collected at a fifth time point after the fourth time point.

[0072] That is, blood can be collected from the target patient 110 who has completed the treatment 120, 130 and / or the surgery 140 at a predetermined period or on demand, and plasma sample sequencing data 124, 134, 144, 148 can be generated based on the plasma sample extracted from the collected blood. Subsequently, the minimal residual disease detection can be performed based on the plasma sample sequencing data 124, 134, 144, 148. In the case where the minimal residual disease is detected, additional treatment and / or surgery can be performed / undertaken for the target patient 110.

[0073] Figure 2 A schematic diagram of a connection structure of an information processing system 230 for providing a minimal residual disease detection service using tumor information in communication with a plurality of user terminals 210_1, 210_2, 210_3 according to an embodiment of the present disclosure. The information processing system 230 can include a system (systems) capable of providing a minimal residual disease detection service using tumor information. In an embodiment, the information processing system 230 can include a computer executable program (e.g., a downloadable application) related to a minimal residual disease detection service using tumor information and one or more server devices and / or databases capable of storing, providing, and executing data or one or more distributed computing devices and / or distributed databases based on a cloud computing service. For example, the information processing system 230 can include an additional system (e.g., a server) for a minimal residual disease detection service using tumor information.

[0074] A minimal residual disease detection service using tumor information and the like provided by the information processing system 230 can be provided to users through an application and the like installed in the plurality of user terminals 210_1, 210_2, 210_3, respectively. For example, the plurality of user terminals 210_1, 210_2, 210_3 can be user terminals of medical institution professionals, user terminals of patients, and the like.

[0075] The plurality of user terminals 210_1, 210_2, 210_3 can communicate with the information processing system 230 through a network 220. The network 220 can realize communication between the plurality of user terminals 210_1, 210_2, 210_3 and the information processing system 230. For example, the network 220 can be realized by a wired network such as Ethernet, a power line communication network, a telephone line communication device, and RS-serial communication, a mobile communication network, a wireless LAN (WLAN), a mobile hotspot (Wi-Fi), Bluetooth, ZigBee, or a combination thereof, according to a setting environment. The communication method is not limited, and can include near distance wireless communication between the user terminals 210_1, 210_2, 210_3, in addition to the communication network (for example, a mobile communication network, a wired Internet, a wireless Internet, a broadcast network, a satellite network, or the like) that the network 220 can include.

[0076] In Figure 2 , as examples of the user terminals, although the mobile phone terminal 210_1, the tablet terminal 210_2, and the personal computer (PC) terminal 210_3 are illustrated, the user terminals 210_1, 210_2, 210_3 are not limited thereto, and can be any computing device capable of wired and / or wireless communication and capable of executing or installing an application or the like. For example, the user terminals can include a smartphone, a mobile phone, a navigator, a vehicle black box, a computer, a notebook computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a tablet, a game console, a wearable device, an internet of things (IoT) device, a virtual reality (VR) device, an augmented reality (AR) device, or the like. Also, in Figure 2 , although a case where three user terminals 210_1, 210_2, 210_3 communicate with the information processing system 230 through the network 220 is illustrated, the number of user terminals that communicate with the information processing system 230 through the network 220 is not limited thereto.

[0077] Figure 3A block diagram showing internal structures of a user terminal 210 and an information processing system 230 according to an embodiment of the present application. The user terminal 210 can also be referred to as any computing device capable of execution and capable of wired communication / wireless communication, such as a mobile phone terminal 210_1, a tablet terminal 210_2, a PC terminal 210_3, and the like, which can include Figure 2 a processor 314, a communication module 316, and an input / output interface 318. Similarly, the information processing system 230 can include a memory 332, a processor 334, a communication module 336, and an input / output interface 338. As shown, the user terminal 210 and the information processing system 230 can transmit and / or receive information and / or data through the network 220 using the respective communication modules 316, 336. Also, the input / output device 320 can input information and / or data to the user terminal 210 through the input / output interface 318, or can output information and / or data generated by the user terminal 210. Figure 3

[0078] The memories 312, 332 can include any non-transitory computer-readable recording medium. According to an embodiment, the memories 312, 332 can include a read only memory (ROM), a magnetic disk drive, a solid state drive (SSD), or a flash memory, and the like non-volatile mass storage device. As another example, the read only memory, the solid state drive, the flash memory, the magnetic disk drive, and the like non-volatile mass storage device can be included in the user terminal 210 or the information processing system 230 as an additional permanent storage device different from the memory. Also, the memories 312, 332 can store an operating system and at least one program code (e.g., a code for a minimal residual lesion detection service, and the like).

[0079] Such software structural elements can be loaded from an additional computer-readable recording medium different from the memories 312, 332. Such an additional computer-readable recording medium can include a recording medium directly connectable to the user terminal 210 and the information processing system 230, such as a floppy disk drive, a magnetic disk, a magnetic tape, a DVD / CD-ROM drive, and / or a memory card, and the like computer-readable recording medium. As another example, the software structural elements can also be loaded in the memories 312, 332 through the communication modules 316, 336, rather than a computer-readable recording medium. For example, at least one program can be loaded in the memories 312, 332 based on a computer program installed through a file provided by a developer or a file distribution system that distributes application installation files, such as an application program related to a minimal residual lesion detection service for application tumor information, and the like, through the network 220.​

[0080] The processors 314, 334 can process instructions of a computer program by performing basic arithmetical, logical, and input / output operations. The instructions can be provided to the processors 314, 334 by the memories 312, 332 or the communication modules 316, 336. For example, the processors 314, 334 can execute instructions received based on program codes stored in the memories 312, 332 or the like recording devices.

[0081] The communication modules 316, 336 can provide structures or functions that enable the user terminal 210 and the information processing system 230 to communicate with each other through the network 220, and can provide structures or functions that enable the user terminal 210 and / or the information processing system 230 to communicate with other user terminals or other systems (e.g., additional cloud systems or the like). As an example, a request or data (e.g., a micro residual lesion detection request, a genetic data request, or the like) generated by the processor 314 of the user terminal 210 based on program codes stored in the memory 312 or the like recording device can be transmitted to the information processing system 230 through the network 220 according to the control of the communication module 316. Conversely, the user terminal 210 can receive control information or instructions provided by the processor 334 of the information processing system 230 through the communication module 336 and the network 220 through the communication module 316 of the user terminal 210.

[0082] The input / output interface 318 refers to a device for connecting the input / output device 320. As an example, the input device can include a camera having an audio sensor and / or an image sensor, a keyboard, a microphone, a mouse, or the like device, and the output device can include a display, a speaker, a haptic feedback device, or the like device. As a further example, the input / output interface 318 can be an interface device of a device in which structures or functions for performing input functions and output functions are integrated into one, such as a touch screen or the like. For example, as the processor 314 of the user terminal 210 executes computer program instructions loaded in the memory 312, a service screen or the like constituted based on information and / or data provided by the information processing system 230 or other user terminals can be displayed on the display through the input / output interface 318. In Figure 3 In the embodiment, although the input / output device 320 is not included in the user terminal 210, it is not limited thereto, and can be integrated into one device with the user terminal 210. Also, the input / output interface 338 of the information processing system 230 can be connected to the information processing system 230, or can be an interface device of an input / output device (not shown) that the information processing system 230 can include. In Figure 3In the present embodiment, although the input / output interfaces 318 and 338 are separate structural elements from the processors 314 and 334, respectively, the input / output interfaces 318 and 338 can be included in the processors 314 and 334, respectively.

[0083] In comparison with the related art Figure 3 The user terminal 210 and the information processing system 230 can include more structural elements than those shown in the drawings. However, most of the related art structural elements are not explicitly shown. In one embodiment, the user terminal 210 can include at least a part of the input / output devices 320. Also, the user terminal 210 can further include a radio transceiver, a Global Positioning System (GPS) module, a camera, various sensors, a database, and other structural elements. For example, if the user terminal 210 is a smartphone, the user terminal 210 can include structural elements that are commonly included in smartphones, such as an acceleration sensor, a gyro sensor, a microphone module, a camera module, various physical buttons, buttons using a touch pad, an output / output port, a vibrator for vibration, and the like.

[0084] According to one embodiment, the processor 314 of the user terminal 210 can cause an application program or a web browser that provides a micro residual disease detection service of an application tumor information to operate. In this case, program codes related to the corresponding application program can be loaded in the memory 312 of the user terminal 210. During operation of the application program, the processor 314 of the user terminal 210 can receive information and / or data provided from the input / output devices 320 through the input / output interface 318, or can receive information and / or data from the information processing system 230 through the communication module 316, can process the received information and / or data, and store them in the memory 312. Also, such information and / or data can be provided to the information processing system 230 through the communication module 316. For example, the processor 314 can receive genetic data including reference sequence data, sequencing data of normal blood samples of a plurality of sample patients, and / or sequencing data of plasma samples of a plurality of sample patients, micro residual disease detection results, and the like, from the information processing system 230.

[0085] During operation of the application program, the processor 314 can input or receive selected voice data, text, images, videos, and the like, through an input device such as a touch screen, a keyboard, a camera having an audio sensor and / or an image sensor, a microphone, and the like, connected to the input / output interface 318, can store the received voice data, text, images, and / or videos, and the like, in the memory 312 or provide them to the information processing system 230 through the communication module 316 and the network 220.

[0086] The processor 314 of the user terminal 210 can output by transmitting information and / or data to the input / output device 320 through the input / output interface 318. For example, the processor 314 of the user terminal 210 can output processed information and / or data through an output device 320 such as a displayable output device (e.g., a touch screen, a display, etc.), a voiceable output device (e.g., a speaker), etc.

[0087] The processor 334 of the information processing system 230 can manage, process, and / or store information and / or data received from a plurality of user terminals 210 and / or a plurality of external systems. The information and / or data processed by the processor 334 can be provided to the user terminals 210 through the communication module 336 and the network 220.

[0088] Figure 4 An example graph 400 of the proportion of tumor cells in the body of a target patient of an embodiment of the present disclosure over time. In an embodiment, the graph 400 is merely an example for illustrating the present disclosure and can represent the proportion of tumor cells in the body of a general tumor patient over time.

[0089] Detection can be performed using a medical image (e.g., a CT image) of tumor cells in the body of a patient, or detection can be performed using ctDNA in the blood of a patient. Generally, when detecting tumor cells using a medical image, a tumor is displayed on the medical image only when the size of the tumor is greater than a prescribed size. Therefore, a first detection limit value LOD1 when detecting tumor cells using a medical image can be greater than a second detection limit value LOD2 when detecting tumor cells using ctDNA in the blood of a target patient. That is, compared to the case of using a medical image, in the case of using ctDNA, tumor cells in the body can be detected at an earlier stage (in a case where the proportion of tumor cells is lower).

[0090] In consideration of the fact that whether a tumor exists is generally first confirmed using a medical image, if a tumor is detected before the proportion of tumor cells in the body of a target patient reaches the first detection limit value LOD1 (e.g., between LOD1 and LOD2), it can be considered that the tumor is detected at an early stage. In this case, the proportion of tumor cells in the body can be reduced through surgery (tumor removal surgery, etc.) and / or treatment (drug treatment, etc.).

[0091] After the proportion of tumor cells is reduced through surgery and / or treatment, the proportion of tumor cells can increase again over time. For example, after surgery and / or treatment, if the proportion of tumor cells continues to increase and is greater than the first detection limit value LOD1, the tumor can be confirmed again using a medical image. In this case, it can be considered that the tumor "clinically recurs."

[0092] Therefore, it is necessary to track whether tumor cells exist and the proportion of tumor cells before clinical recurrence of the tumor. However, there is a problem in that it is difficult to detect tumor cells using ctDNA due to a background error rate in the sequencing process before the proportion of tumor cells in the body reaches the second detection limit value. In other words, if the proportion of tumor cells in the body is less than the second detection limit value, it is difficult to easily distinguish whether the variation detected from the blood is derived from tumor cells or is an error detection due to the background error rate.

[0093] To solve this problem, if the background error rate is reduced by data filtering, the second detection limit value can be reduced, and tumor cells can be further quickly found at an early stage. That is, in a non-invasive liquid biopsy sample, as the limit value capable of detecting minimal residual disease (MRD) is reduced, not only can the patient's burden be greatly reduced, but tumor cells can be further quickly found at an early stage.

[0094] Figure 5 A graph of a blood sample 500 of a target patient according to an embodiment of the present application. Figure 5 The illustrated blood sample 500 can be a blood sample (liquid biopsy sample) in which blood collected from a target patient is centrifugally separated to achieve layer separation. The centrifugally separated blood sample 500 can include a plasma layer 510, a normal blood layer 520, and a red blood cell layer 530 from the upper layer.

[0095] The blood sample 500 can be collected from the target patient before and / or after surgery and / or treatment of the target patient. In an embodiment, at least a portion of the normal blood layer 520 can be collected from the blood sample 500 before surgery and / or treatment of the target patient. Normal blood sequencing data can be obtained from the collected normal blood layer 520 sample through whole genome sequencing. Subsequently, the obtained normal blood sequencing data can be compared with tumor tissue sequencing data obtained from a tumor tissue biopsy sample of the target patient to detect tumor tissue variation information of the target patient.

[0096] In an embodiment, at least a portion of the plasma layer 510 can be collected from a blood sample 500 taken after surgery and / or treatment of the target patient. In this case, the plasma layer 510 can include cell free Deoxyribo Nucleic Acid (cfDNA) and circulating tumor Deoxyribo Nucleic Acid (ctDNA). Plasma sequencing data can be obtained from the plasma layer 510 sample by whole genome sequencing. Then, the plasma sequencing data can be utilized to detect the cell free Deoxyribo Nucleic Acid and circulating tumor Deoxyribo Nucleic Acid to determine the proportion of tumor cells in the target patient and whether the tumor has recurred, etc.

[0097] Figure 6 A diagram is shown to illustrate a process of calculating a limit of detection (LOD) 680 of a target patient according to an embodiment of the present application. The limit of detection 680 of a target patient refers to the lowest proportion of tumor cells that can be detected from a plasma sample of the target patient. The limit of detection 680 can vary depending on the tumor tissue variance of the target patient, and the higher the accuracy of sequencing data obtained by whole genome sequencing, the more accurately the limit of detection 680 can be grasped.

[0098] As shown, the limit of detection 680 can be calculated based on tumor tissue sequencing data 610, normal blood sequencing data 620, and plasma sequencing data 660. In this case, the tumor tissue sequencing data 610 can be obtained from a tumor tissue biopsy sample collected from a tumor of the target patient by whole genome sequencing (WGS). Also, the normal blood sequencing data 620 can be obtained from a normal blood sample collected from a normal blood layer (e.g., 520) of a blood sample of the target patient by whole genome sequencing (WGS). Additionally, the plasma sequencing data 660 can be obtained from a blood sample collected from a plasma layer (e.g., 510) of a blood sample of the target patient by whole genome sequencing (WGS). Figure 5 Figure 5

[0099] ​​In an embodiment, the detection limit value 680 of the target patient can be calculated based on the tumor tissue variants 630 (specifically, the filtered tumor tissue variants 650) of the target patient and the average sequencing depth 670 of the plasma sequencing data 660. Specifically, as the basis for calculating the detection limit value 680, the tumor tissue variants 630 of the target patient can be detected by comparing the tumor tissue sequencing data 610 with the normal blood sequencing data 620. That is, with the premise that the normal blood sample does not include tumor tissue, the normal blood sequencing data 620 can be used as reference sequencing data or control sequencing data for identifying genomic variants and mutations in relation to the tumor tissue sequencing data 610. The tumor tissue variant 630 information can include position information of the presence of variants and nucleic acid sequence information of the variants.

[0100] Additionally or alternatively, the variants of the ctDNA in the plasma can be detected by comparing the normal blood sequencing data 620 with the plasma sequencing data 660. That is, the normal blood sequencing data 620 can be used as reference sequencing data or control sequencing data in relation to the plasma sequencing data 660. In this case, variants that cannot be detected from the tumor tissue sequencing data 610 can be additionally removed from the variants detected from the plasma sequencing data 660. In an embodiment, the variants detected from the plasma sequencing data 660 are replaced with or added to the tumor tissue variants 630 after the filtering 642, 644, which can be used to calculate the detection limit value 680 of the target patient. Figure 6

[0101] For the detected tumor tissue variants 630, the filtered tumor tissue variants 650 can be generated by performing background error filtering 642, 644 on the tumor tissue variant 630 information detected using genetic data related to a plurality of sample patients that are different from the target patient to be detected for the minimal residual disease. For example, if the target patient is a lung cancer patient, a plurality of sample patients that are different from the target patient are screened as 33 lung cancer patients that are different from the target patient, and the background error filtering can be performed using genetic data of the 33 lung cancer patients. The genetic data can include normal blood sequencing data and plasma sequencing data of the plurality of sample patients. Additionally, the genetic data can include tumor tissue sequencing data of the plurality of sample patients. The genetic data related to the plurality of sample patients for filtering 642, 644 can be obtained from a genetic database 640.

[0102] ​In an embodiment, the first filtering 642 of the tumor tissue variants 630 can be performed using the sequencing data of the normal blood samples of the plurality of sample patients included in the genetic data of the plurality of sample patients. Specifically, the first filtering 642 can be performed by removing the variants detected in the detected tumor tissue variants 630 from the sequencing data of the normal blood samples of the plurality of sample patients. That is, the variants detected from the sequencing data of the normal blood samples of the plurality of sample patients are considered as false positives of the variants within the actual tumor tissue, and such variant information can be removed when the first filtering 642 is performed.

[0103] Additionally, the second filtering 644 can be performed on the tumor tissue variants of the first filtering 642. The second filtering 644 can be performed using the sequencing data of the plasma samples of the plurality of sample patients. Specifically, the second filtering 644 can be performed by removing the variants detected in the detected tumor tissue variants from the sequencing data of the plasma samples of the plurality of sample patients. That is, the variants identical to the variants detected from the sequencing data of the plasma samples of the plurality of sample patients are considered as false positives on the system, and such variants can be removed from the tumor tissue variants of the first filtering 642.

[0104] In Figure 6 In the above-described embodiment, the second filtering 644 is performed after the first filtering 642, but the present disclosure is not limited thereto. For example, the first filtering 642 can be performed after the second filtering 644 is performed. In another example, one of the first filtering 642 or the second filtering 644 can be omitted.

[0105] On the other hand, the ctDNA within the blood of the target patient or the average sequencing depth 670 of the plasma sequencing data 660 of the target patient can be calculated from the plasma sequencing data 660 of the target patient. Then, the detection limit value 680 can be calculated based on the filtered tumor tissue variant 650 information and the average sequencing depth 670 of the plasma sequencing data 660. For example, the detection limit value 680 can be calculated based on the following mathematical expression 1.

[0106] Mathematical Expression 1

[0107] where N denotes the number of variants detected from the tumor tissue (e.g., the filtered tumor tissue variants), and cov can denote the average sequencing depth of the plasma sequencing data.

[0108] Figure 7A diagram illustrating a tumor detection process 780 according to an embodiment of the present invention is provided. As shown, the tumor cell fraction (TCF) 730 of the target patient can be calculated based on plasma sequencing data 710 and a reference sequence 720. Furthermore, a corrected tumor cell fraction 760 can be calculated using a random error rate 750. Subsequently, tumor detection 780 can be performed based on the corrected tumor cell fraction 760 and the patient's detection limit value (LOD) 770. In this case, the detection limit value 770 can be... Figure 6 The target patient's detection limit is calculated to be 680.

[0109] In one embodiment, the tumor cell ratio 730 can be calculated based on plasma sequencing data 710 and a reference sequence 720. Specifically, the number of sequencing reads different from the reference sequence 720 and the number of sequencing reads identical to the reference sequence 720 can be determined from the plasma sequencing data 710. Then, the tumor cell ratio 730 can be calculated based on the number of sequencing reads different from the reference sequence 720 and the number of sequencing reads identical to the reference sequence 720. For example, the tumor cell ratio 730 can be calculated based on the following mathematical formula 2.

[0110] Mathematical formula 2

[0111] Wherein, Tumor Cell Fraction (raw) represents the proportion of tumor cells 730, ∑ alteernative allele count ∑ represents the number of sequencing fragments that differ from the reference sequence 720. reference allele count This indicates the number of sequencing fragments identical to the reference sequence 720.

[0112] Alternatively, the tumor cell ratio of 730 can be based on Figure 6 The number of tumor tissue variations filtered out was calculated to be 650. Figure 6 In the filtered tumor tissue variant 650 information, there are regions of genomic variation from tumor cells and the same information. Therefore, the number of sequencing fragments including variants detected from tumor tissue can be calculated using the filtered tumor tissue variants.

[0113] Specifically, based on the filtered tumor tissue variation information, sequencing fragment data including variations detected from tumor tissue are identified in the patient's plasma sequencing data. The identified sequencing fragment data can be used as the ∑ in mathematical formula 2. alteernative allele countFor example, if the filtered tumor tissue variant number is 2000, regions corresponding to the 2000 tumor tissue variants can be identified in the plasma sequencing data and it can be determined whether the variants exist. For example, if the ctDNA in the plasma is determined to be 20x, the number of sequencing reads including the variants detected from the tumor tissue can be determined from 20*2000=40000 sequencing reads. Using the filtered tumor tissue variant information 650, the number of sequencing reads including the variants detected from the tumor tissue in the plasma sequencing data can be the same or very close to the actual number of sequencing reads different from the reference sequence 720. Thus, the time and computational resources required to compare the entire plasma sequencing data 710 to the reference sequence 720 can be reduced. Figure 6

[0114] On the other hand, the random error rate 750 can be calculated based on genetic data related to a plurality of sample patients different from the target patient and the reference sequence 720. The genetic data related to the plurality of sample patients can be obtained from a genetic database 740. The genetic database 740 can be the same database as the genetic database 640. Figure 6

[0115] The genetic data related to the plurality of sample patients included in the genetic database 740 can include plasma sequencing data of the plurality of sample patients. In an embodiment, after determining the number of sequencing reads different from the reference sequence 720 and the number of sequencing reads identical to the reference sequence 720 from the plasma sequencing data of the plurality of sample patients, the random error rate 750 can be calculated based on the respective numbers of sequencing reads. For example, the random error rate 750 can be calculated based on the following mathematical expression 3.

[0116] Mathematical Expression 3

[0117] wherein Random Error Rate denotes the random error rate 750, Σ alternative allele count denotes the number of sequencing reads different from the reference sequence 720 in the plasma sequencing data of the plurality of sample patients, and Σ reference allele count denotes the number of sequencing reads identical to the reference sequence 720 in the plasma sequencing data of the plurality of sample patients.

[0118] Next, the tumor cell proportion 730 can be corrected using the random error rate 750. For example, the corrected tumor cell proportion 760 can be a value obtained by subtracting the random error rate 750 from the tumor cell proportion 730.

[0119] ​​Subsequently, tumor detection (e.g., determining whether the tumor has recurred) 780 can be performed on the target patient based on the corrected tumor cell proportion 760 and the detection limit value 770. For example, a confidence interval at a predetermined confidence level can be calculated for the corrected tumor cell proportion 760. For example, if a 95% confidence level is used, the confidence interval at the 95% confidence level can be calculated based on the following mathematical formula 4.

[0120] Mathematical expression 4

[0121] Here, TCF represents the corrected proportion of tumor cells, 760.

[0122] Subsequently, in response to the judgment that the lower limit of the confidence interval for the calculated corrected tumor cell proportion 760 is higher than the detection limit value 770, the patient's tumor can be judged to have recurred (tumor detection 780). If the patient's tumor is judged to have recurred, tumor treatment / surgery can be performed on the target patient.

[0123] and Figure 7 Unlike the methods described, tumor detection 780 can be performed using structural variants. In one embodiment, the recurrence of a target patient's tumor can be determined based on a comparison of a first group of structural variants detected from tumor tissue sequencing data and a second group of structural variants detected from plasma sequencing data. For example, if the number or proportion of structural variants included in the second group of structural variants in the first group is above a predetermined threshold, the patient's tumor can be determined to be recurrent. For example, the predetermined threshold can be 0.9. In this case, the structural variants in the first group can include at least one of inversion, translocation, duplication, deletion, or insertion of a region of the genome.

[0124] Figure 8 Schematic graphs 810 and 820 illustrate the relationship between the number of variants detected in tumor cells of a target patient and the detection limit value of the target patient, according to an embodiment of the present invention. Graphs 810 and 820 show the actual results of determining the detection limit value of each sample patient based on ctDNA samples from 33 sample patients. Specifically, the first graph 810 shows the detection limit values ​​of all 33 sample patients, and the second graph 820 shows the detection limit values ​​within a specific range in the first graph 810.

[0125] In the first plot 810, the detection limit values of all the sample patients are in the range of about 1.48E-06 to 1.49E-03, and the median value is about 1.99E-05. Among the 33 sample patients, 29 sample patients (about 87% of the sample patients) have detection limit values less than 1E-4. For the sample patients whose detected ctDNA matches tumor cells with a somatic variant count greater than 20,000, the detection limit value is reduced to 1E-6. That is, the detection limit value of the sample patients is correlated with the somatic variant count of the sample patients. The greater the somatic variant count, the lower the minimum tumor cell proportion that can be detected from the plasma sample of the tumor patient, that is, the detection limit value has a decreasing tendency, and tumor detection is facilitated.

[0126] Figure 9 An example plot of generating arm-level copy number profile 912, 922, 932 for an embodiment of the present application. The arm-level copy number profile can be generated by mapping the sequencing data 910, 920, 930 to a reference genome, as described in the methods described above. Figure 6 and Figure 7 Additionally or alternatively to the methods described above, the arm-level copy number profile can be used to detect minimal residual disease.

[0127] In an embodiment, a first arm-level copy number profile 912 can be generated from tumor tissue sequencing data 910 of a target patient, a second arm-level copy number profile 922 can be generated from normal blood sequencing data 920, and a third arm-level copy number profile 932 can be generated from plasma sequencing data 930. In this case, the tumor tissue sequencing data 910, the normal blood sequencing data 920, and the plasma sequencing data 930 can be obtained from a tumor tissue biopsy sample, a normal blood sample, and a plasma sample of the target patient, respectively.

[0128] In an embodiment, the arm-level copy number profile 912, 922, 932 can be generated by mapping the sequencing data 910, 920, 930 to a reference genome, respectively. For example, for the sequencing data 910, 920, 930, respectively, the genomic position information of the chromosome arms can be obtained, the sequencing coverage (or depth or the number of sequencing reads) of the sequencing reads mapped to each chromosome arm can be counted, and then the normalized arm-level copy number profile 912, 922, 932 can be obtained by dividing the whole genome sequencing coverage of the WGS data.

[0129] Figure 10An example diagram showing the calculation of the tumor cell fraction of a target patient using chromosome arm-level copy number profiles 1010, 1020, 1030 according to an embodiment of the present application. The third chromosome arm-level copy number profile 1030 generated from the plasma sequencing data of the target patient can be considered as a mixture of the first chromosome arm-level copy number profile 1010 generated from the tumor tissue sequencing data of the target patient and the second chromosome arm-level copy number profile 1020 generated from the normal blood sequencing data of the target patient in a certain ratio. For example, as shown in the diagram, the first chromosome arm-level copy number profile 1010 and the second chromosome arm-level copy number profile 1020 can be considered as a mixture in a ratio of a : (1-a) (where a is a real number greater than or equal to 0 and less than or equal to 1). Here, a is the tumor cell fraction (TCF) of the patient (e.g., a = 0.7). The third chromosome arm-level copy number profile 1030 is a mixture of the first chromosome arm-level copy number profile 1010 and the second chromosome arm-level copy number profile 1020 in a ratio of a : (1-a). The tumor cell fraction (TCF) of the target patient can be calculated based on the following mathematical formula 5. Figure 7 The value of A (i.e., the tumor cell fraction) can be calculated based on the non-negative least squares (NNLS) shown in the following mathematical formula 5.

[0130] Mathematical formula 5

[0131] where ||Ax-y||2 represents the Euclidian norm, A represents a matrix of the first chromosome arm-level copy number profile 1010 and the second chromosome arm-level copy number profile 1020, and y can represent a vector related to the third chromosome arm-level copy number profile 1030.

[0132] Figure 11 A diagram showing the performance verification results of the method for detecting the minimal residual disease of a target patient using the genetic data of a sample patient according to an embodiment of the present application. The first graph 1110 shows the verification results of accurately determining the tumor cell fraction (TCF) by the residual disease detection method using the genetic data of a sample patient (e.g., the method shown in Figure 6 and Figure 7 The second graph 1120 shows the verification results of accurately determining the tumor cell fraction by the residual disease detection method using the genetic data of a sample patient after mixing the actual tumor tissue DNA with the normal DNA (experimental mixture).

[0133] As shown in the first graph 1110 and the second graph 1120, in the case of mixed sequencing data and in the case of mixed DNA, it can be confirmed that there is a correlation of 99.9% or more between the correct answer tumor cell proportion (Answer) and the predicted tumor cell proportion (Predicted) of the present application. That is, according to the result of filtering the tumor tissue variation of the target patient using the genetic data of the sample patient, it can be confirmed that the tumor cell proportion in the target patient can be accurately predicted.

[0134] Figure 12 A graph showing the performance verification result of the method for detecting a minimal residual disease of a target patient using a chromosome arm level copy number profile of the target patient according to an embodiment of the present application. The same as the method for detecting a minimal residual disease of a target patient using a chromosome arm level copy number profile of the target patient (for example, the method shown in Figure 11 the same, the performance of the method for detecting a minimal residual disease of a target patient using a chromosome arm level copy number profile of the target patient (for example, the method shown in Figure 9 and Figure 10 may be verified by in silico mixture (first graph 1210) and experimental mixture (second graph 1220).

[0135] As shown in the first graph 1210 and the second graph 1220, in the case of mixed sequencing data or mixed DNA, it can be confirmed that there is a correlation of 99.9% or more between the correct answer tumor cell proportion (Answer) and the predicted tumor cell proportion (Predicted) of the present application. That is, it can be confirmed that the tumor cell proportion in the target patient can be accurately predicted using the chromosome arm level copy number profile of the target patient.

[0136] Figure 13 A flowchart showing a method for detecting a minimal residual disease using tumor information 1130 according to an embodiment of the present application. The method 1300 can be performed by at least one processor of a system for detecting a minimal residual disease using tumor information. The method 1300 can start with the processor acquiring first sequencing data related to a first sample of a patient (step S1310). In this case, the first sample can be a tumor tissue biopsy sample of the patient, and the first sequencing data can be acquired by whole genome sequencing (WGS).

[0137] Next, the processor can acquire second sequencing data related to a second sample of the patient (step S1320). In this case, the second sample is a normal blood sample of the patient, and the second sequencing data can be acquired by whole genome sequencing.

[0138] Subsequently, the processor can acquire third sequencing data related to a third sample of the patient (step S1330). In this case, the third sample is a plasma sample of the patient, and the plasma sample can include cell free Deoxyribo Nucleic Acid (cfDNA) and circulating tumor Deoxyribo Nucleic Acid (ctDNA). The third sequencing data can be acquired through whole genome sequencing.

[0139] In an embodiment, the first sample and the second sample can be samples acquired at a first time point, and the third sample is a sample acquired at a second time point after the first time point. At least one of a surgery or a treatment therapy can be performed on the patient between the first time point and the second time point.

[0140] The processor detects tumor tissue variant information of the patient by comparing the first sequencing data and the second sequencing data, and can perform background error filtering on the detected tumor tissue variant information using genetic data related to a plurality of sample patients distinguished from the patient. In this case, the genetic data can include sequencing data of normal blood samples of the plurality of sample patients and / or sequencing data of plasma samples of the plurality of sample patients.

[0141] In an embodiment, the processor can perform background error filtering by removing, from the tumor tissue variant information, a variant detected from the sequencing data of the normal blood samples of the plurality of sample patients, among the detected tumor tissue variants. In yet another embodiment, the processor can perform background error filtering by removing, from the tumor tissue variant information, a variant detected from the sequencing data of the plasma samples of the plurality of sample patients, among the detected tumor tissue variants.

[0142] Then, the processor can perform a minimal residual disease detection on the patient based on the first sequencing data, the second sequencing data, and the third sequencing data (step S1340). In an embodiment, the processor can calculate a limit of detection value of the patient based on the first sequencing data, the second sequencing data, and the third sequencing data, calculate a tumor cell fraction of the patient based on the third sequencing data, correct the tumor cell fraction using genetic data related to a plurality of sample patients different from the patient, and determine whether the tumor of the patient has relapsed based on the corrected tumor cell fraction and the limit of detection value, whereby the minimal residual disease detection can be performed on the patient. In this case, the limit of detection value can indicate a minimum tumor cell fraction that can be detected from the plasma sample of the patient. Also, the limit of detection value can be calculated based on a tumor tissue variant of the patient detected by comparing the first sequencing data and the second sequencing data and an average sequencing depth of the third sequencing data.

[0143] In an embodiment, the processor can determine a number of sequencing reads different from a reference sequence in the third sequencing data, determine a number of sequencing reads identical to the reference sequence in the third sequencing data, and calculate the tumor cell fraction based on the number of sequencing reads different from the reference sequence and the number of sequencing reads identical to the reference sequence. In this case, the processor detects tumor tissue variant information of the patient by comparing the first sequencing data and the second sequencing data, determines a number of sequencing reads including variants detected from the tumor tissue in the third sequencing data based on the tumor tissue variant information, and uses the number of sequencing reads including the variants detected from the tumor tissue in the third sequencing data as the number of sequencing reads different from the reference sequence.

[0144] In an embodiment, the processor can determine a number of sequencing reads different from a reference sequence in plasma sequencing data included in the genetic data related to a plurality of sample patients different from the patient, determine a number of sequencing reads identical to the reference sequence in the plasma sequencing data included in the genetic data related to a plurality of sample patients different from the patient, calculate a random error rate based on the number of sequencing reads different from the reference sequence and the number of sequencing reads identical to the reference sequence, and correct the tumor cell fraction using the random error rate.

[0145] In one embodiment, the processor calculates a confidence interval of a predetermined confidence level of the corrected tumor cell proportion, and judges that the patient's tumor has recurred in response to a determination that a lower limit of the calculated confidence interval is higher than the calculated detection limit value.

[0146] In one embodiment, the processor identifies a first group of structural variations detected from the first sequencing data, identifies a second group of structural variations detected from the third sequencing data, and judges whether the patient's tumor has recurred based on a comparison result of the first group of structural variations and the second group of structural variations. For example, in the first group of structural variations, if the number or proportion of structural variations included in the second group of structural variations is equal to or greater than a threshold value, it can be determined that the patient's tumor has recurred. In this case, the first group of structural variations can include at least one of inversion, translocation, duplication, deletion, or insertion of a partial region of a genome.

[0147] Alternatively, the processor generates a first chromosomal arm-level copy number profile based on the first sequencing data, generates a second chromosomal arm-level copy number profile based on the second sequencing data, and generates a third chromosomal arm-level copy number profile based on the third sequencing data, and performs the minimal residual disease detection based on the calculation of the tumor cell proportion based on the first chromosomal arm-level copy number profile to the third chromosomal arm-level copy number profile.

[0148] Figure 13 The illustrated flow and the described description are only one example, and in some embodiments, different implementations can be implemented. For example, one or more steps can be omitted, or the order of the steps can be changed, or one or more steps can be performed simultaneously, or one or more steps can be repeated multiple times.

[0149] To enable the computer to execute, the method can be provided by a computer program stored in a computer-readable recording medium. The medium can also be a temporary storage for storing, executing or downloading a computer executable program. Furthermore, the medium can be a plurality of recording units or storage units combined by a single or a plurality of hardware, and is not limited to a medium directly accessible to any computer system, but can be distributed on a network. As an example, the medium includes magnetic media such as a hard disk, a floppy disk and a magnetic tape, optical recording media such as a CD-ROM and a DVD, a magneto optical medium such as a floppy optical disk, and a structure storing program instructions such as a ROM, a RAM and a flash memory. Furthermore, as another example, the medium can be a recording medium or a storage medium managed by an application program store selling an application program or a website or a server providing and selling a plurality of other software.

[0150] The method, work or technology of the present application can be implemented by a plurality of units. For example, the technology can also be implemented by hardware, firmware, software or a combination thereof. In connection with the disclosure of the present application, it should be understood by those skilled in the art to which the present application belongs that the described plurality of exemplary logic blocks, modules, circuits and algorithm steps can also be implemented by electronic hardware, computer software or a combination thereof. In order to clearly illustrate the interchangeability of hardware and software, the above plurality of exemplary structural elements, blocks, modules, circuits and steps are simply described based on the functional point of view. However, whether the function is implemented by hardware or software depends on the design requirements of the specific application and the overall system. Those skilled in the art can also implement the described functions in various ways for each specific application, and therefore such implementation should not be interpreted as departing from the scope of the present application.

[0151] In a hardware instance, the processing unit for executing the technology can also be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), graphics processors (GPUs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices designed to perform the functions of the present application, computers or combinations thereof.

[0152] Accordingly, the various illustrative logical blocks, modules, and circuits described in connection with the disclosure can be implemented or performed with a general purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor can be a microprocessor, but, in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a digital signal processor, a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a digital signal processor core, or any other such configuration.

[0153] The techniques of this disclosure can be implemented in a variety of embodiments as described above. Although the disclosure has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art. Accordingly, functional elements of the specific constructional arrangements are to be understood as merely exemplary, and other arrangements can be devised by those skilled in the art that will be within the scope of the appended claims. In this regard, the various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein can be implemented or performed with a general purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor can be a microprocessor, but, in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a digital signal processor, a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a digital signal processor core, or any other such configuration.

[0154] In the above described embodiments, although the subject embodiments disclosed herein can be applied to one or more independent computer systems, the present disclosure is not limited thereto, and can be implemented in any computing environment, such as a network or distributed computing environment. Also, the subject embodiments disclosed herein can be implemented by a plurality of processing chips or devices, and storage can be performed by a plurality of devices under similar conditions. Such devices can include personal computers, network servers, and portable devices.

[0155] In the present specification, although the disclosure describes matters related to some embodiments, those skilled in the art can make various modifications and changes without departing from the scope of the present disclosure. Also, such modifications and changes belong to the scope of the patent application attached to the present specification.

Claims

1. A method of microresidual lesion detection, performed by at least one processor, the method comprising: comprising the steps of: obtaining first sequencing data associated with a first sample of a patient; obtaining second sequencing data associated with a second sample of the patient; obtaining third sequencing data associated with a third sample of the patient; and performing minimal residual disease detection on the patient based on the first, second and third sequencing data. The first, second and third sequencing data are obtained by whole genome sequencing.

2. The method of microresidual lesion detection of claim 1, wherein, 3. The method of claim 1, wherein: the first sample is a tumor tissue biopsy sample of the patient, the second sample is a normal blood sample of the patient, the third sample is a plasma sample of the patient, the plasma sample comprises circulating cell-free DNA and circulating tumor DNA.

4. The method of claim 1, wherein: the first sample and the second sample are samples obtained at a first time point, the third sample is a sample obtained at a second time point after the first time point, at least one of a surgical or a therapeutic treatment is performed on the patient between the first time point and the second time point. further comprising the steps of:

5. The method of claim 1, wherein the method is used for detecting a minimal residual disease. detecting tumor tissue variant information of the patient by comparing the first and second sequencing data; and performing background error filtering on the detected tumor tissue variant information using genetic data associated with a plurality of sample patients different from the patient.

6. The method of claim 5, wherein: the genetic data comprises sequencing data of normal blood samples of the plurality of sample patients, the step of performing the filtering comprises the step of removing from the detected tumor tissue variant information variants detected from the sequencing data of normal blood samples of the plurality of sample patients.

7. The method of claim 5, wherein: the genetic data comprises sequencing data of plasma samples of the plurality of sample patients, the step of performing the filtering comprises the step of removing from the detected tumor tissue variant information variants detected from the sequencing data of plasma samples of the plurality of sample patients. the step of performing the minimal residual disease detection comprises the steps of:

8. The method of claim 1, wherein the method is used for detecting a minimal residual disease. calculating a limit of detection value for the patient based on the first, second and third sequencing data; calculating a tumor cell fraction for the patient based on the third sequencing data; correcting the tumor cell fraction using genetic data associated with a plurality of sample patients different from the patient; and determining whether the patient has a tumor recurrence based on the corrected tumor cell fraction and the limit of detection value. The limit of detection value represents a minimum tumor cell fraction that can be detected from a plasma sample of the patient.

9. The method of microresidual lesion detection of claim 8, wherein, The limit of detection value is calculated based on the tumor tissue variant information of the patient detected by comparing the first and second sequencing data and an average sequencing depth of the third sequencing data.

10. The method of claim 8, wherein the method is used for detecting a minimal residual disease. ​ 11. The method of claim 8, wherein the method is used for detecting a minimal residual disease. The step of calculating the tumor cell proportion includes the steps of: determining the number of sequencing fragments different from the reference sequence in the third sequencing data; determining the number of sequencing fragments identical to the reference sequence in the third sequencing data; and calculating the tumor cell proportion based on the number of sequencing fragments different from the reference sequence and the number of sequencing fragments identical to the reference sequence.

12. The method of microresidual lesion detection of claim 11, wherein, The step of determining the number of different sequencing fragments includes the steps of: detecting tumor tissue variation information of the patient by comparing the first sequencing data and the second sequencing data; determining the number of sequencing fragments including variations detected from the tumor tissue in the third sequencing data based on the tumor tissue variation information; and using the number of sequencing fragments including variations detected from the tumor tissue in the third sequencing data as the number of sequencing fragments different from the reference sequence.

13. The method of claim 8, wherein the method is used for detecting a minimal residual disease. The step of correcting the tumor cell proportion includes the steps of: determining the number of sequencing fragments different from the reference sequence in the plasma sequencing data included in the genetic data related to a plurality of sample patients different from the patient; determining the number of sequencing fragments identical to the reference sequence in the plasma sequencing data included in the genetic data related to a plurality of sample patients different from the patient; calculating a random error rate based on the number of sequencing fragments different from the reference sequence and the number of sequencing fragments identical to the reference sequence; and correcting the tumor cell proportion using the random error rate.

14. The method of claim 8, wherein the method is used for detecting a minimal residual disease. The step of determining whether the tumor of the patient has relapsed includes the steps of: calculating a confidence interval of a predetermined confidence level of the corrected tumor cell proportion; and in response to a determination that the lower limit of the calculated confidence interval is higher than the calculated detection limit value, determining that the tumor of the patient has relapsed.

15. The method of claim 1, wherein the method is used for detecting a minimal residual disease. The step of performing the minimal residual disease detection further includes the steps of: generating a first chromosome arm level copy number profile based on the first sequencing data; generating a second chromosome arm level copy number profile based on the second sequencing data; generating a third chromosome arm level copy number profile based on the third sequencing data; and calculating a tumor cell proportion based on the first chromosome arm level copy number profile to the third chromosome arm level copy number profile. The step of performing the minimal residual disease detection further includes the steps of:

16. The method of microresidual lesion detection of claim 3, wherein, identifying a first group of structural variations detected from the first sequencing data; identifying a second group of structural variations detected from the third sequencing data; and determining whether the tumor of the patient has relapsed based on a comparison result of the first group of structural variations and the second group of structural variations. The step of determining whether the tumor of the patient has relapsed includes the step of, among the first group of structural variations, if the number or proportion of structural variations included in the second group of structural variations is above a threshold value, determining that the tumor of the patient has relapsed. The first group of structural variations includes at least one of inversion, translocation, duplication, deletion, or insertion of a partial region of a genome.

17. The method of microresidual lesion detection of claim 16, wherein, An instruction for executing the method according to claim 1 in a computer is recorded.

18. The method of microresidual lesion detection of claim 16, wherein, 20. A system, comprising: 19.A non-transitory computer-readable recording medium, characterized by, a communication module; ​ ​ ​ a memory; and at least one processor connected to the memory, configured to execute at least one program readable by the computer, the at least one program comprising instructions to perform the steps of: obtaining first sequencing data associated with a first sample of a patient; obtaining second sequencing data associated with a second sample of the patient; obtaining third sequencing data associated with a third sample of the patient; and performing minimal residual disease detection on the patient based on the first, second, and third sequencing data.

Citation Information

Patent Citations

  • Method for duplex sequencing

    WO2022125997A1