Method and apparatus for detecting minimal residual disease using tumor information

The method addresses the challenge of detecting low-concentration DNA mutations by using whole genome sequencing and data filtering to enhance the accuracy and sensitivity of minimal residual disease detection.

JP2026514368APending Publication Date: 2026-05-11INOCRAS KOREA INC
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
INOCRAS KOREA INC
Filing Date
2023-11-08
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

Existing NGS techniques struggle to accurately detect DNA mutations at very low concentrations due to a background error rate of 0.1% to 1%, making it difficult to distinguish between cancer-derived DNA and false positives, especially in minimal residual disease detection.

Method used

A method involving whole genome sequencing of tumor tissue, normal blood, and plasma samples, combined with data filtering and background error reduction, to accurately detect tumor cells by calculating tumor cell ratios and reducing the detection threshold.

Benefits of technology

Enables earlier detection of tumor cells with high accuracy by lowering the background error rate and improving the sensitivity of minimal residual disease detection in liquid biopsies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026514368000001_ABST
    Figure 2026514368000001_ABST
Patent Text Reader

Abstract

This disclosure relates to a method for detecting minimal residual disease using tumor information. [Solution] A minimal residual disease detection method utilizing tumor information includes the steps of: acquiring first sequencing data associated with a first sample of the patient; acquiring second sequencing data associated with a second sample of the patient; acquiring third sequencing data associated with a third sample of the patient; and performing minimal residual disease detection for the patient based on the first sequencing data, the second sequencing data, and the third sequencing data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method and apparatus for detecting minimal residual disease by utilizing tumor information, and specifically, to a method and apparatus for performing detection of minimal residual disease on a patient by using mutation information of a tumor tissue and a liquid biopsy sample.

Background Art

[0002] Gene analysis technology is widely used in the medical field, such as to determine what characteristics or qualities a living organism has by understanding the genes it possesses. Recently, medical practices for treating various diseases such as tumors have evolved from a traditional prescription-centered approach to precision medicine, that is, a customized treatment form that takes into account the genetic information and health records of individual patients.

[0003] In the field of precision medicine, it is important to obtain a huge amount of personal genetic information and perform related clinical analyses. Recently, the NGS (Next-Generation Sequencing) technique, which can quickly read a large amount of DNA information in parallel, is widely used in various medical fields such as cancer screening. The NGS technique has the advantages of low cost and time required for genome analysis and convenience, but there is a limit in that an inherent background error rate of about 0.1% to 1% occurs. That is, even for DNA without mutations, there is a limit (False Positive Mutation) in that it is erroneously judged that there is a mutation in one base pair out of 100 to 1,000 base pairs statistically during the sequencing process (that is, an error rate of 0.1% to 1%).

[0004] Due to these inherent limitations of NGS techniques, they have difficulty detecting DNA mutations that exist at very low concentrations. For example, cancer-derived DNA (e.g., ctDNA: Circulation tumor DNA) is typically present in the blood of cancer patients at concentrations of less than 1%, making it difficult to distinguish whether mutations detected in blood using NGS techniques originate from cancer tissue or are false positives due to the background error rate. Therefore, there is a need for new techniques that can detect DNA mutations (such as tumor cell residues) present at very low concentrations in the patient's blood by reducing the background error rate. [Prior art documents] [Patent Documents]

[0005] [Patent Document 1] Korean Published Patent Publication No. 10-2015-0017525 [Overview of the project] [Problems that the invention aims to solve]

[0006] This disclosure provides a method for detecting minimal residual disease using tumor information, a computer-readable non-temporary recording medium for recording commands, and an apparatus (system) to solve the above-mentioned problems. [Means for solving the problem]

[0007] This disclosure can be embodied in a variety of ways, including methods, systems (apparatus), or computer-readable non-temporary recording media on which instructions are recorded.

[0008] A method for detecting minimal residual disease (MPD) performed by at least one processor according to one embodiment of the present disclosure includes the steps of: acquiring first sequencing data associated with a first sample of a patient; acquiring second sequencing data associated with a second sample of a patient; acquiring third sequencing data associated with a third sample of a patient; and performing MPD detection for the patient based on the first sequencing data, the second sequencing data, and the third sequencing data.

[0009] A computer-readable non-temporary recording medium is provided, which records instructions for executing a minimal residual disease detection method utilizing tumor information, according to one embodiment of the present disclosure, on a computer.

[0010] A system according to one embodiment of the present disclosure includes a communication module, a memory, and at least one processor connected to the memory and configured to execute at least one computer-readable program contained in the memory, the at least one program including instructions for obtaining first sequencing data associated with a first sample of a patient, obtaining second sequencing data associated with a second sample of a patient, obtaining third sequencing data associated with a third sample of a patient, and performing minimal residual disease detection for the patient based on the first sequencing data, the second sequencing data and the third sequencing data. [Effects of the Invention]

[0011] According to various embodiments of this disclosure, tumor cells can be detected earlier by reducing the background error rate through data filtering and lowering the threshold at which minimal residual disease (MRD) can be detected in non-invasive liquid biopsy samples.

[0012] According to various embodiments of this disclosure, calculating tumor cell ratios based on filtered tumor tissue mutations can reduce the time and computing resources required to compare the entire plasma sequencing data with a reference sequence.

[0013] According to various embodiments of this disclosure, the proportion of tumor cells in a target patient can be predicted with great accuracy by using the genetic data of a sample patient to filter out tumor tissue mutations in that patient.

[0014] According to various embodiments of this disclosure, the proportion of tumor cells in a patient's body can be predicted with great accuracy by utilizing the copy number profile at the chromosomal arm level of the patient.

[0015] The effects of this disclosure are not limited to those mentioned above, and any other effects not mentioned above would be clearly understood by a person with ordinary skill in the art to which this disclosure pertains ("persons skilled in the art") from the wording of the claims. [Brief explanation of the drawing]

[0016] Embodiments of the present disclosure will be described with reference to the accompanying drawings described below, where similar reference numbers indicate similar elements, but are not limited thereto. [Figure 1] This figure illustrates an example of the process of performing minimal residual disease detection on a target patient according to one embodiment of the present disclosure. [Figure 2] This is a schematic diagram showing a configuration in which an information processing system is connected to multiple user terminals in a communicative manner in order to provide a minimal residual disease detection service utilizing tumor information according to one embodiment of the present disclosure. [Figure 3] This is a block diagram showing the internal configuration of a user terminal and an information processing system according to one embodiment of the present disclosure. [Figure 4] This figure shows an example of a graph illustrating the percentage of tumor cells in a patient over time, according to one embodiment of the present disclosure. [Figure 5]A drawing showing a blood sample of a target patient according to an embodiment of the present disclosure. [Figure 6] A diagram showing the process of calculating the detection limit value for a target patient according to an embodiment of the present disclosure. [Figure 7] A diagram showing the tumor detection process according to an embodiment of the present disclosure. [Figure 8] A diagram showing an example of a graph illustrating the relationship between the number of mutations detected in tumor cells of a target patient and the detection limit value of the target patient according to an embodiment of the present disclosure. [Figure 9] A diagram showing an example of generating an arm unit copy number profile according to an embodiment of the present disclosure. [Figure 10] A diagram showing an example of calculating the ratio of in vivo tumor cells of a target patient using a copy number profile at the chromosomal arm level according to an embodiment of the present disclosure. [Figure 11] A diagram showing the performance verification results of a method for detecting minimal residual disease of a target patient using the genetic data of a sample patient according to an embodiment of the present disclosure. [Figure 12] A diagram showing the performance verification results of a method for detecting minimal residual disease of a target patient using the copy number profile at the chromosomal arm level of the target patient according to an embodiment of the present disclosure. [Figure 13] A flowchart showing a method for detecting minimal residual disease utilizing tumor information according to an embodiment of the present disclosure.

Mode for Carrying Out the Invention

[0017] <Summary of the Invention> In one embodiment of the present disclosure, the first sequencing data, the second sequencing data, and the third sequencing data are obtained via whole genome sequencing (WGS).

[0018] In one embodiment of the present disclosure, the first sample is a tumor tissue biopsy sample from the patient, the second sample is a normal blood sample from the patient, and the third sample is a plasma sample from the patient, the plasma sample containing cfDNA (cellfree deoxyribo-nucleic acid) and ctDNA (circulating tumor deoxyribo-nucleic acid).

[0019] In one embodiment of the present disclosure, the first and second samples are samples taken at a first time point, the third sample is a sample taken at a second time point after the first time point, and at least one of surgery or therapeutic treatment is performed on the patient after the first time point and before the second time point.

[0020] One embodiment of the present disclosure further includes the steps of detecting tumor tissue mutation information of a patient by comparing first sequencing data and second sequencing data, and performing background error filtering on the detected tumor tissue mutation information using genetic data associated with a plurality of sample patients that are distinguishable from the patient.

[0021] In one embodiment of the present disclosure, the genetic data includes sequencing data for normal blood samples from multiple sample patients, and the step of performing the filtering includes removing from the tumor tissue mutation information any mutations detected in the sequencing data for normal blood samples from multiple sample patients.

[0022] In one embodiment of the present disclosure, the genetic data includes sequencing data for plasma samples from multiple sample patients, and the step of performing the filtering includes removing from the tumor tissue mutation information any mutations detected in the sequencing data for plasma samples from multiple sample patients.

[0023] In one embodiment of the present disclosure, the step of performing minimal residual disease detection includes: calculating a limit of detection value for a patient based on first sequencing data, second sequencing data, and third sequencing data; calculating a tumor cell fraction (TCF) for a patient based on the third sequencing data; correcting the tumor cell fraction using genetic data associated with a plurality of sample patients distinct from the patient; and determining whether the patient's tumor is recurrent based on the corrected tumor cell fraction and the limit of detection value.

[0024] In one embodiment of this disclosure, the Limit of Detection value means the lowest tumor cell percentage that can be detected in a patient's plasma sample.

[0025] In one embodiment of the present disclosure, the detection limit is calculated based on the number of tumor tissue mutations in the patient detected by comparing the first sequencing data and the second sequencing data, and the average sequencing depth of the third sequencing data.

[0026] In one embodiment of the present disclosure, the step of calculating the tumor cell ratio includes the steps of determining the number of reads in the third sequencing data that differ from the reference sequence, determining the number of reads in the third sequencing data that match the reference sequence, and calculating the tumor cell ratio based on the number of reads that differ from the reference sequence and the number of reads that match the reference sequence.

[0027] In one embodiment of the present disclosure, the step of determining the number of differing reads includes: comparing first sequencing data and second sequencing data to detect tumor tissue mutation information of a patient; determining the number of reads in the third sequencing data that include mutations detected in tumor tissue based on the tumor tissue mutation information; and using the number of reads in the third sequencing data that include mutations detected in tumor tissue as the number of reads that differ from the reference sequence.

[0028] In one embodiment of the present disclosure, the step of correcting the tumor cell ratio includes the steps of: determining the number of reads that differ from a reference sequence in plasma sequencing data included in genetic data associated with multiple sample patients distinct from the patient; determining the number of reads that match a reference sequence in plasma sequencing data included in genetic data associated with multiple sample patients distinct from the patient; calculating a random error rate based on the number of reads that differ from the reference sequence and the number of reads that match the reference sequence; and correcting the tumor cell ratio using the random error rate.

[0029] In one embodiment of the present disclosure, the step of determining whether a patient's tumor has recurred includes the steps of calculating a confidence interval for a predetermined confidence level of the corrected tumor cell ratio, and determining that the patient's tumor has recurred in response to the determination that the lower limit of the calculated confidence interval is higher than the calculated detection limit.

[0030] In one embodiment of the present disclosure, the step of performing minimal residual disease detection further includes generating a chromosomal arm-level copy number profile based on first sequencing data, generating a second chromosomal arm-level copy number profile based on second sequencing data, generating a third chromosomal arm-level copy number profile based on third sequencing data, and calculating tumor cell ratios based on the first to third chromosomal arm-level copy number profiles.

[0031] In one embodiment of the present disclosure, the step of performing minimal residual disease detection further includes identifying a first set of structural mutations detected from first sequencing data, identifying a second set of structural mutations detected from third sequencing data, and determining whether the patient's tumor is likely to recur based on a comparison of the first set of structural mutations and the second set of structural mutations.

[0032] In one embodiment of the present disclosure, the step of determining whether a patient's tumor has recurred includes determining that the patient has recurred if the number or ratio of structural mutations included in the second set of structural mutations out of the first set of structural mutations is equal to or greater than a threshold.

[0033] In one embodiment of the present disclosure, the first set of structural mutations includes at least one of the following: inversion, translocation, duplication, deletion, or insertion of a region of the genome.

[0034] <Details of the invention> The specific details for implementing this disclosure will be described below with reference to the attached drawings. However, in the following explanation, specific descriptions of widely known functions and configurations will be omitted if there is a risk of unnecessarily obscuring the essence of this disclosure.

[0035] In the attached drawings, identical or corresponding components are assigned the same reference numerals. Furthermore, in the following descriptions of embodiments, the description of identical or corresponding components may be omitted. However, the omission of a description of a component does not mean that such a component is not included in any of the embodiments.

[0036] The advantages and features of the disclosed embodiments, and the methods for achieving them, will become clear with reference to the embodiments described below, along with the accompanying drawings. However, this disclosure is not limited to the embodiments disclosed below and may be embodied in a variety of different forms, and these embodiments are provided merely to complete the disclosure and to fully inform the scope of the invention to a person of ordinary skill.

[0037] This specification will briefly explain the terminology used herein and then provide a more detailed description of the disclosed embodiments. The terminology used herein has been selected to the greatest extent possible from commonly used terms, taking into account the function described herein, although this may change depending on the intent of the articulates in the relevant field, case law, the emergence of new technologies, etc. In some cases, the applicant has arbitrarily selected terms, in which case their meaning will be described in detail in the description of the relevant invention. Therefore, the terminology used in this disclosure should not be merely nominal terms, but should be defined based on the meaning of the term and the overall content of this disclosure.

[0038] In this specification, singular expressions include plural expressions unless the context clearly identifies them as singular. Conversely, plural expressions include singular expressions unless the context clearly identifies them as plural. Throughout the specification, when a part is said to contain a certain component, this means that, unless otherwise stated, it may contain other components rather than excluding them.

[0039] Furthermore, the terms “module” or “part” as used in this specification mean software or hardware components, and that a “module” or “part” performs some role. However, the meaning of “module” or “part” is not limited to software or hardware. A “module” or “part” may be configured to reside on an addressable storage medium, or to regenerate one or more processors. Thus, as an example, a “module” or “part” may include components such as software components, object-oriented software components, class components, and task components, and at least one of processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, or variables. The components and the functions provided within a “module” or “part” may be combined into a smaller number of components and a “module” or “part,” or further separated into additional components and a “module” or “part.”

[0040] According to one embodiment of the present disclosure, “module” or “part” may be embodied in a processor and memory. “Processor” should be broadly interpreted to include general-purpose processors, central processing units (CPUs), microprocessors, digital signal processors (DSPs), controllers, microcontrollers, state machines, and the like. In some environments, “processor” may also refer to application-specific semiconductors (ASICs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), and the like. “Processor” may also refer to a combination of processing devices such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors coupled with a DSP core, or any other combination of such configurations. “Memory” should also be broadly interpreted to include any electronic component capable of storing electronic information. "Memory" can also refer to a variety of processor-readable media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage, and registers. Memory is said to be in electronic communication with the processor if the processor can read information from and / or write information to it. Memory integrated into a processor is in electronic communication with the processor.

[0041] In this disclosure, “Whole Genome Sequencing (WGS)” or “whole genome sequencing” may refer to a technique used to determine the entire DNA sequence of an organism’s genome. Specifically, whole genome sequencing may involve deciphering and identifying the sequence of nucleotide bases (adenine, cytosine, guanine, and thymine) within the entire set of genetic material of a human or organism. The entire set of genetic material may include all genes, non-coding regions, and any additional genetic elements present in the genome. In one embodiment, whole genome sequencing may be performed in several steps. For example, whole genome sequencing may be performed by extracting DNA from a particular cell, dividing the extracted DNA into smaller fragments, and generating millions or billions of short DNA sequences, referred to as “reads.” The generated reads may be aligned and assembled to reconstruct the whole genome sequence. In certain embodiments, various duplex sequencing methods may be used during whole genome sequencing to minimize the background error rate, allowing even very minute cancer-derived DNA to be detected in blood. During whole-genome sequencing, the CODEC (Concatenating Original Duplex for Error Correction) method, an example of duplex sequencing, may be used, and in connection therewith, PCT Publication WO2022 / 125997A1, published on June 16, 2022, is incorporated herein by reference.

[0042] In this disclosure, "DNA" is not limited to DNA itself, but may include any nucleic acid or nucleic acid sequence, such as RNA (Ribo-Nucleic Acid).

[0043] In this disclosure, “cfDNA” may refer to circulating free DNA, also known as cell-free DNA or circulating DNA. cfDNA may mean small fragments of DNA released into the bloodstream and other bodily fluids by cells undergoing normal cell death or as a result of certain pathological processes. cfDNA may consist of short DNA fragments derived from diverse tissues of the body, including healthy and diseased cells. cfDNA may originate from organs such as the liver, heart, or lungs, or from tumors (e.g., malignant tumors such as cancer). For example, a blood sample can be centrifuged to obtain a plasma layer, a buffy coat layer, and a red blood cell layer from the top, and cfDNA can be obtained from the plasma layer.

[0044] In this disclosure, “ctDNA” may refer to circulating tumor DNA. ctDNA may be a subset of cfDNA, particularly derived from tumor cells. Tumor cells may shed ctDNA into the bloodstream while undergoing cell death or cell death. ctDNA may be accompanied by genetic alterations or mutations that are characteristic of the tumor from which they originated. ctDNA can provide useful information about the genetic characteristics and mutation profile of a tumor, even without invasive tissue biopsies. By analyzing ctDNA, it may be possible to monitor a patient's response to treatment and detect minimal residual disease (MRD).

[0045] In this disclosure, “normal blood sample” or “buffy coat” may refer to specific components of a blood sample, including leukocytes and platelets. For example, a normal blood sample can be obtained by centrifugation of a blood sample. Here, the normal blood sample can be obtained from the thin layer formed between the red blood cell layer below and the plasma layer above the centrifuged sample. A normal blood sample contains a higher concentration of nucleated cells, including leukocytes with nuclei containing DNA, and may provide a source of genomic DNA derived from the cells of a particular individual. Thus, when performing WGS, DNA can be extracted from a normal blood sample to obtain the genomic DNA of a particular individual, which can be used as reference or control DNA to compare and identify genetic variations or mutations present in cfDNA or ctDNA extracted from the same person, or in the DNA of tumor tissue extracted from the same person. It can be assumed that a normal blood sample does not contain cancer cells.

[0046] In this disclosure, “sequencing depth” or “sequencing coverage” may refer to the average number of times each base of the genome was read during the sequencing process. For example, sequencing depth may indicate how many times a particular nucleotide (A, T, C, or G) at a particular location in the genome was sequenced. Sequencing depth is a parameter that affects the accuracy and reliability of sequencing data. Sequencing depth is generally measured in “X”, where 1X coverage may mean that each base was sequenced an average of 1 time, and 10X coverage may mean that each base was sequenced an average of 10 times.

[0047] In this disclosure, “copy number profiling” may refer to a technique for obtaining a copy number profile through copy number (replica number) analysis of a specific DNA segment or region within the genome using WGS data.

[0048] In one embodiment, copy number profiling may involve several steps. First, a DNA isolation step may be performed to isolate DNA from the cells or tissue of interest. Then, after the extracted DNA has been divided into smaller fragments (fragmentation), a sequencing step may be performed on the DNA fragments. Through this step, a large number of reads, which are short DNA sequences, may be generated. In some embodiments, after the DNA has been fragmented, a step may be performed to amplify and replicate the DNA fragments using a technique such as PCR (Polymerase Chain Reaction) before performing the sequencing step.

[0049] Subsequently, a read alignment step may be performed in which the generated reads are aligned to or mapped to a reference sequence. This step may involve comparing the reads to a known DNA sequence of the reference sequence to determine the reads' original positions. Once the reads are aligned, a copy number analysis step may be performed using software tools to analyze the sequencing coverage or depth across the entire genome. Sequencing coverage may indicate the number of reads aligned to a particular genomic region. Subsequently, a normalization step may be performed on the coverage data to account for variations in sequencing depth and other technical biases. This allows for accurate comparisons between different genomic regions.

[0050] Subsequently, statistical algorithms may be applied to the normalized coverage data to perform a copy number calling step, which identifies genomic regions exhibiting copy number variation. These algorithms may compare the observed coverage to expected coverage based on a diploid genome. The copy number profile may then be visualized using plots or heatmaps showing genomic regions with increased or lost copies, which may be referred to as the visualization and interpretation step. Such visualization data can provide insights into genomic modifications such as amplification, deletion, or duplication.

[0051] In this disclosure, “Limit of Detection” may refer to the minimum detectable concentration at which the presence or absence of the analyte can be confirmed during the analytical process for that analyte. For example, “Limit of Detection” may refer to the lowest tumor cell percentage that can be detected in a plasma sample from a tumor patient.

[0052] Figure 1 shows an example of the process of performing minimal residual disease detection on a target patient (110) according to one embodiment of the present disclosure. The target patient (110) may be a patient with tumor tissue (e.g., cancerous tissue). Alternatively, the target patient (110) may be a patient on whom minimal residual disease detection is performed after treatment (120, 130) or surgery (140) for the tumor tissue has progressed.

[0053] In one embodiment, a tumor tissue biopsy sample and a normal blood sample may be taken from the patient (110) at a first time point before proceeding with treatment (120, 130) or surgery (140) on the tumor tissue. The normal blood sample may be obtained by centrifugation of the blood taken from the patient (110). Subsequently, tumor tissue sequencing data (112) and normal blood sequencing data (114) may be generated based on the tumor tissue biopsy sample and the normal blood sample. Here, each sequencing data may be obtained via whole genome sequencing (WGS).

[0054] Subsequently, tumor tissue mutation information of the target patient (110) can be detected based on the acquired tumor tissue sequencing data (112) and normal blood sequencing data (114). Background error filtering can be performed on the detected tumor tissue mutation information using genetic data associated with multiple sample patients distinct from the target patient (110). By performing background error filtering, the background error rate is reduced, and the possibility of tumor recurrence in the target patient (110) can be confirmed at an earlier point than before background error filtering. The specific process of detecting and filtering tumor tissue mutation information will be described in detail later with reference to Figure 6.

[0055] In one embodiment, after a tumor tissue biopsy sample and a normal blood sample are taken from the patient (110), at least one of the following may be performed on the patient (110): treatment (120, 130) and / or surgery (140). The treatment (120, 130) and / or surgery (140) performed on the patient (110) may reduce the percentage of tumor cells in the patient's body. However, over time, the percentage of tumor cells in the patient's body may increase again.

[0056] To respond quickly to disease recurrence due to an increase in tumor cells in the body, it is important to detect an increase in the tumor cell ratio early. For this reason, plasma samples may be collected from the patient (110) after treatment (120, 130) and / or surgery (140). Plasma samples may be collected by centrifugation of blood (122, 132, 142, 146) taken from the patient (110). At this time, the plasma sample may contain cfDNA (cell-free deoxyribo-nucleic acid) and ctDNA (circulating tumor deoxyribo-nucleic acid). Plasma samples may be collected and analyzed from the patient (110) at regular intervals or as needed and used for follow-up examinations to check for tumor recurrence. Specifically, plasma sample sequencing data (124, 134, 144, 148) is generated based on plasma samples contained in blood (122, 132, 142, 146), and the necessity of additional treatment and / or surgery for the patient (110) can be determined based on the acquired plasma sample sequencing data (124, 134, 144, 148).

[0057] For example, plasma sample sequencing data (124) may be obtained from blood (122) collected at a second time point after the first treatment (120). Subsequently, based on the tumor tissue mutation information and the plasma sample sequencing data (124), the detection limit of the tumor tissue of the target patient (110) can be calculated, and minimal residual disease detection can be performed on the target patient (110). The specific process for calculating the detection limit of the tumor tissue of the target patient (110) will be described in detail using Figure 6. Furthermore, the specific process for performing minimal residual disease detection on the target patient (110) will be described in detail using Figure 7.

[0058] If minimal residual disease is detected in the patient (110) after the first treatment (120) (or if it is determined that the tumor has recurred), a second treatment (130) may be performed. Thereafter, at a third time point after the second treatment (130) has been performed, blood (132) may be collected from the patient (110), and the above-described detection of minimal residual disease may be repeated.

[0059] In another example, minimal residual disease (MI) may not be detected based on plasma sample sequencing data (144) obtained from blood (142) collected at a fourth time point after surgery (140). In this case, it is determined that the tumor has not recurred in the patient (110), and no additional treatment and / or surgery is required. Subsequently, MIFF detection may be performed again based on plasma sample sequencing data (148) obtained from blood (146) collected at a fifth time point after the fourth time point.

[0060] Specifically, blood from a patient (110) who has completed treatment (120, 130) and / or surgery (140) may be collected at predetermined intervals or as needed, and plasma sample sequencing data (124, 134, 144, 148) may be generated based on plasma samples extracted from the collected blood. Subsequently, minimal residual disease detection may be performed based on the plasma sample sequencing data (124, 134, 144, 148). If minimal residual disease is detected, additional treatment and / or surgery may be performed on the patient (110).

[0061] Figure 2 is a schematic diagram showing a configuration in which an information processing system (230) is connected to communicate with multiple user terminals (210_1, 210_2, 210_3) in order to provide a minimal residual disease detection service utilizing tumor information according to one embodiment of the present disclosure. The information processing system (230) may include a system(s) capable of providing a minimal residual disease detection service utilizing tumor information. In one embodiment, the information processing system (230) may include one or more server devices and / or databases, or one or more distributed computing devices and / or distributed databases of a cloud computing service infrastructure, capable of storing, providing, and executing computer-executable programs (e.g., downloadable applications) and data related to the minimal residual disease detection service utilizing tumor information. For example, the information processing system (230) may include a separate system (e.g., a server) for the minimal residual disease detection service utilizing tumor information.

[0062] Services such as minimal residual disease detection utilizing tumor information provided by the information processing system (230) may be provided to users via applications installed on each of the multiple user terminals (210_1, 210_2, 210_3). For example, the multiple user terminals (210_1, 210_2, 210_3) may be user terminals for medical professionals, user terminals for patients, etc.

[0063] Multiple user terminals (210_1, 210_2, 210_3) can communicate with an information processing system (230) via a network (220). The network (220) can be configured to enable communication between the multiple user terminals (210_1, 210_2, 210_3) and the information processing system (230). Depending on the installation environment, the network (220) may consist of wired networks such as Ethernet, Power Line Communication, telephone line communication equipment and RS-serial communication, mobile communication networks, wireless networks such as WLAN (WirelessLAN), Wi-Fi, Bluetooth and ZigBee, or a combination thereof. The communication method is not limited and may include not only communication methods that utilize communication networks that the network (220) may include (for example, mobile communication networks, wired internet, wireless internet, broadcasting networks, satellite networks, etc.), but also short-range wireless communication between user terminals (210_1, 210_2, 210_3).

[0064] In Figure 2, a mobile phone terminal (210_1), a tablet terminal (210_2), and a PC terminal (210_3) are shown as examples of user terminals, but the user terminal (210_1, 210_2, 210_3) can be any computing device capable of wired and / or wireless communication, on which applications can be installed and executed. For example, user terminals may include smartphones, mobile phones, navigation systems, vehicle black boxes, computers, laptop computers, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablet PCs, game consoles, wearable devices, IoT (Internet of Things) devices, VR (virtual reality) devices, AR (augmented reality) devices, and the like. Furthermore, although Figure 2 illustrates three user terminals (210_1, 210_2, 210_3) communicating with the information processing system (230) via the network (220), the configuration is not limited to this, and a different number of user terminals may be configured to communicate with the information processing system (230) via the network (220).

[0065] Figure 3 is a block diagram showing the internal configuration of a user terminal (210) and an information processing system (230) according to one embodiment of the present disclosure. The user terminal (210) can refer to any computing device capable of executing applications and other applications and capable of wired / wireless communication, and may include, for example, the mobile phone terminal (210_1), tablet terminal (210_2), PC terminal (210_3) shown in Figure 2. As illustrated, the user terminal (210) may include memory (312), a processor (314), a communication module (316), and an input / output interface (318). Similarly, the information processing system (230) may include memory (332), a processor (334), a communication module (336), and an input / output interface (338). As illustrated in Figure 3, the user terminal (210) and the information processing system (230) may be configured to communicate information and / or data over a network (220) using their respective communication modules (316, 336). Furthermore, the input / output device (320) may be configured to input information and / or data to the user terminal (210) via the input / output interface (318), or to output information and / or data generated from the user terminal (210).

[0066] The memory (312, 332) may include any non-temporary computer-readable recording medium. According to one embodiment, the memory (312, 332) may include a permanent mass storage device such as a ROM (read-only memory), disk drive, SSD (solid-state drive), or flash memory. In another example, a permanent mass storage device such as a ROM, SSD, flash memory, or disk drive may be included in a user terminal (210) or information processing system (230) as a separate permanent storage device distinct from the memory. The memory (312, 332) may also store an operating system and at least one program code (for example, code for a minimal residual lesion detection service).

[0067] Such software components may be loaded from a computer-readable recording medium separate from memory (312, 332). Such separate computer-readable recording medium may include recording media that can be directly connected to such user terminals (210) and information processing systems (230), and may include computer-readable recording media such as floppy drives, disks, tapes, DVD / CD-ROM drives, and memory cards. As another example, software components may be loaded into memory (312, 332) via a communication module (316, 336) rather than from a computer-readable recording medium. For example, at least one program may be loaded into memory (312, 332) based on a computer program (e.g., an application associated with a minimal residual disease detection service utilizing tumor information) that is installed by a file provided over a network (220) by a developer or a file distribution system that distributes application installation files (e.g., an application associated with a minimal residual disease detection service utilizing tumor information).

[0068] The processors (314, 334) may be configured to process computer program instructions by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processors (314, 334) by memory (312, 332) or communication modules (316, 336). For example, the processors (314, 334) may be configured to execute instructions received according to program code stored in a recording device such as memory (312, 332).

[0069] The communication modules (316, 336) may provide configurations or functions for a user terminal (210) and an information processing system (230) to communicate with each other via a network (220), and may provide configurations or functions for a user terminal (210) and / or an information processing system (230) to communicate with other user terminals or other systems (for example, a separate cloud system). For example, requests or data (e.g., minimal residual disease detection requests, genetic data requests, etc.) generated by the processor (314) of the user terminal (210) according to program code stored in a recording device such as memory (312) may be transmitted to the information processing system (230) via the network (220) under the control of the communication module (316). Conversely, control signals or commands provided under the control of the processor (334) of the information processing system (230) may be received by the user terminal (210) via the communication module (336) and the network (220) through the communication module (316) of the user terminal (210).

[0070] The input / output interface (318) may be a means for interface with an input / output device (320). For example, an input device may include a camera with an audio sensor and / or image sensor, a keyboard, a microphone, a mouse, etc., and an output device may include a display, a speaker, a haptic feedback device, etc. As another example, the input / output interface (318) may be a means for interface with a device in which the configuration or function for performing input and output is integrated into one, such as a touchscreen. In Figure 3, the input / output device (320) is shown not to be included in the user terminal (210), but is not limited to this, and may be configured as a single device with the user terminal (210). Furthermore, the input / output interface (338) of the information processing system (230) may be connected to the information processing system (230) or a means for interface with an input or output device (not shown) that the information processing system (230) may include. In Figure 3, the input / output interfaces (318, 338) are shown as elements configured separately from the processors (314, 334). However, the diagram is not limited to this, and the input / output interfaces (318, 338) can be configured to be included in the processors (314, 334).

[0071] The user terminal (210) and the information processing system (230) may include more components than those shown in Figure 3. However, it is not necessary to clearly illustrate most of the conventional components. In one embodiment, the user terminal (210) may be embodied to include at least some of the input / output devices (320) described above. The user terminal (210) may also further include other components such as a transceiver, a GPS (Global Positioning system) module, a camera, various sensors, and a database. For example, if the user terminal (210) is a smartphone, it may include components that are generally found in smartphones, and the user terminal (210) may be embodied to further include a variety of components such as an accelerometer, a gyroscope, a microphone module, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration.

[0072] In one embodiment, the processor (314) of a user terminal (210) may be configured to run an application or a web browser application that provides a minimal residual disease detection service utilizing tumor information. In this case, program code associated with the application may be loaded into the memory (312) of the user terminal (210). While the application is running, the processor (314) of the user terminal (210) may receive information and / or data provided by an input / output device (320) via an input / output interface (318) or information and / or data from an information processing system (230) via a communication module (316), process the received information and / or data, and store it in memory (312). Such information and / or data may also be provided to the information processing system (230) via the communication module (316). For example, the processor (314) may receive genetic data from the information processing system (230), including reference sequence data, sequencing data for normal blood samples from multiple sample patients and / or sequencing data for plasma samples from multiple sample patients, and minimal residual disease detection results.

[0073] While the application is running, the processor (314) may receive or receive audio data, text, images, video, etc., input via input devices such as a touchscreen, keyboard, camera including audio sensors and / or image sensors, and microphone, which are connected to the input / output interface (318). The received audio data, text, images, and / or video, etc., may be stored in memory (312) or provided to the information processing system (230) via the communication module (316) and network (220).

[0074] The processor (314) of the user terminal (210) may transmit information and / or data to an input / output device (320) via an input / output interface (318) and output it. For example, the processor (314) of the user terminal (210) may output processed information and / or data via an output device (320) such as a display output device (e.g., touchscreen, display, etc.) or an audio output device (e.g., speaker).

[0075] The processor (334) of the information processing system (230) may be configured to manage, process, and / or store information and / or data received from multiple user terminals (210) and / or multiple external systems. The information and / or data processed by the processor (334) may be provided to the user terminals (210) via a communication module (336) and a network (220).

[0076] Figure 4 shows an example of a graph (400) of the percentage of tumor cells in the body over time in a patient according to one embodiment of the present disclosure. In one embodiment, graph (400) may show the percentage of tumor cells in the body over time in a typical tumor patient, as an example for the purpose of explaining the present disclosure.

[0077] Tumor cells in a patient's body can be detected using medical imaging (e.g., CT images) or by using ctDNA in the patient's blood. Generally, when detecting tumor cells using medical imaging, the tumor will only appear on the medical image if it is above a certain size. Therefore, the first limit of detection (LOD1) when detecting tumor cells using medical imaging may be higher than the second limit of detection (LOD2) when detecting tumor cells using ctDNA in the patient's blood. In other words, using ctDNA allows for earlier detection of tumor cells in the body (at a lower tumor cell ratio) than using medical imaging.

[0078] Considering that the presence or absence of a tumor is generally only confirmed through medical imaging, if a tumor is detected before the percentage of tumor cells in the patient's body reaches the first limit of detection (LOD1) (e.g., between LOD1 and LOD2), it can be considered an early detection. In this case, the percentage of tumor cells in the body may decrease through surgery (such as tumor removal surgery) and / or treatment (such as drug therapy).

[0079] Even after surgery and / or treatment has reduced the tumor cell ratio, the tumor cell ratio may increase again over time. For example, if the tumor cell ratio continues to increase after surgery and / or treatment and exceeds the first limit of detection (LOD1), the tumor may be re-identified through medical imaging. In this case, the tumor can be considered to have "clinically recurred."

[0080] Therefore, it is necessary to track the presence or ratio of tumor cells before clinical recurrence of the tumor. However, before the ratio of tumor cells in the body reaches the second detection limit, there is a problem in that it is difficult to detect tumor cells using ctDNA due to the background error rate during the sequencing process. In other words, when the ratio of tumor cells in the body is below the second detection limit, it is not easy to distinguish whether the mutation detected in the blood originates from tumor cells or is a false detection due to the background error rate.

[0081] To address these issues, reducing the background error rate through data filtering can lower the second detection limit, allowing tumor cells to be detected earlier. In other words, by lowering the threshold at which minimal residual disease (MRD) can be detected in non-invasive fluid biopsy samples, tumor cells can be detected earlier while significantly reducing the burden on patients.

[0082] Figure 5 shows a blood sample (500) from a patient according to one embodiment of the present disclosure. The blood sample (500) shown in Figure 5 may be a blood sample (liquid biopsy sample) obtained by centrifuging and separating layers of blood collected from a patient. The centrifuged blood sample (500) may include, from top to bottom, a plasma layer (510), a normal blood layer (520), and a red blood cell layer (530).

[0083] A blood sample (500) may be collected from the patient before or after surgery and / or treatment. In one embodiment, at least a portion of the normal blood layer (520) may be collected from the blood sample (500) of the patient before surgery and / or treatment. Normal blood sequencing data may be obtained from the collected normal blood layer (520) sample by whole-genome sequencing. Subsequently, the obtained normal blood sequencing data may be compared with tumor tissue sequencing data obtained from a tumor tissue biopsy sample of the patient to detect tumor tissue mutation information of the patient.

[0084] In one embodiment, after surgery and / or treatment of a target patient, at least a portion of the plasma layer (510) can be collected from a blood sample (500). The plasma layer (510) may contain cfDNA (cell-free deoxyribo-nucleic acid) and ctDNA (circulating tumor deoxyribo-nucleic acid). Plasma sequencing data can be obtained from the plasma layer (510) sample by whole-genome sequencing. Subsequently, by detecting cfDNA or ctDNA using the plasma sequencing data, the patient's tumor cell ratio or the likelihood of tumor recurrence can be determined.

[0085] Figure 6 shows the process by which the Limit of Detection (LOD) (680) for a target patient is calculated according to one embodiment of the present disclosure. The Limit of Detection (LOD) (680) for a target patient may represent the lowest tumor cell percentage that can be detected in the target patient's plasma sample. The Limit of Detection (LOD) (680) may vary depending on the number of mutations in the target patient's tumor tissue, and the higher the accuracy of the sequencing data obtained by whole-genome sequencing, the more accurately the Limit of Detection (LOD) (680) can be determined.

[0086] As illustrated, the detection limit (680) can be calculated based on tumor tissue sequencing data (610), normal blood sequencing data (620), and plasma sequencing data (660). In this case, tumor tissue sequencing data (610) can be obtained via whole-genome sequencing (WGS) of tumor tissue biopsy samples taken from the tumor of the patient. Normal blood sequencing data (620) can be obtained via whole-genome sequencing (WGS) of normal blood samples taken from the normal blood layer in the patient's blood sample (e.g., 520 in Figure 5). Additionally, plasma sequencing data (660) can be obtained via whole-genome sequencing (WGS) of plasma samples taken from the plasma layer in the patient's blood sample (e.g., 510 in Figure 5).

[0087] In one embodiment, the detection limit (680) for a target patient may be calculated based on the target patient's tumor tissue mutations (630) (specifically, filtered tumor tissue mutations (650)) and the average sequencing depth (670) of the plasma sequencing data (660). Specifically, the tumor tissue mutations (630) of the target patient that form the basis for calculating the detection limit (680) may be detected by comparing tumor tissue sequencing data (610) with normal blood sequencing data (620). That is, it may be assumed that normal blood samples do not contain tumor tissue, and the normal blood sequencing data (620) may be used as reference or control sequencing data to identify genetic mutations or mutations in relation to the tumor tissue sequencing data (610). The tumor tissue mutation (630) information may include location information of the mutation and information of the mutated nucleic acid sequence.

[0088] Additionally or alternatively, normal blood sequencing data (620) and plasma sequencing data (660) may be compared to detect plasma ctDNA mutations. That is, normal blood sequencing data (620) may be used as reference or control sequencing data in relation to plasma sequencing data (660). In this case, mutations detected in plasma sequencing data (660) that were not detected in tumor tissue sequencing data (610) may be additionally removed. In one embodiment, mutations detected in plasma sequencing data (660) may be used to calculate the detection limit (680) for the target patient by substituting for or adding to tumor tissue mutations (630) in Figure 6, and then filtering (642, 644).

[0089] For the detected tumor tissue mutations (630), background error filtering (642, 644) can be performed on the detected tumor tissue mutations (630) information using genetic data associated with multiple sample patients that are distinguished from the target patient for detecting minimal residual disease, thereby generating filtered tumor tissue mutations (650). For example, if the target patient is a lung cancer patient, 33 lung cancer patients who are distinguishable from the target patient can be selected as multiple sample patients, and background error filtering can be performed using the genetic data of these 33 lung cancer patients. The genetic data may include normal blood sequencing data and plasma sequencing data from multiple sample patients. Additionally, the genetic data may include tumor tissue sequencing data from multiple sample patients. The genetic data associated with multiple sample patients used for filtering (642, 644) can be obtained from a genetic database (640).

[0090] In one embodiment, primary filtering (642) of tumor tissue mutations (630) can be performed using sequencing data from normal blood samples of multiple sample patients included in the genetic data of multiple sample patients. Specifically, primary filtering (642) can be performed by removing mutations detected in the sequencing data from normal blood samples of multiple sample patients from the tumor tissue mutations (630). In other words, mutations detected from normal blood samples of multiple sample patients are assumed to be errors rather than actual mutations in tumor tissue, and such mutation information can be removed during primary filtering (642).

[0091] Additionally, secondary filtering (644) may be performed on tumor tissue mutations that have undergone primary filtering (642). Secondary filtering (644) may be performed using sequencing data from plasma samples of multiple sample patients. Specifically, among the detected tumor tissue mutations, those detected in sequencing data from plasma samples of multiple sample patients may be removed from the tumor tissue mutation information that has undergone primary filtering (642). That is, mutations that match those detected in sequencing data from plasma samples of multiple sample patients are assumed to be system errors and may be removed from the tumor tissue mutations that have undergone primary filtering (642).

[0092] In Figure 6, the primary filtering (642) is performed followed by the secondary filtering (644), but this is not the case. For example, the secondary filtering (644) described above may be performed first, followed by the primary filtering (642). In other examples, either the primary filtering (642) or the secondary filtering (644) may be omitted.

[0093] On the other hand, the average sequencing depth (670) of the blood ctDNA or plasma sequencing data (660) of the target patient can be calculated from the plasma sequencing data (660) of the target patient. Subsequently, the detection limit (680) can be calculated based on the filtered tumor tissue mutation (650) information and the average sequencing depth (670) of the plasma sequencing data (660). For example, the detection limit (680) can be calculated based on the following formula 1.

number

[0094] Here, N represents the number of mutations detected in the tumor tissue (e.g., the number of filtered tumor tissue mutations), and cov may represent the average sequencing depth of the plasma sequencing data.

[0095] Figure 7 shows a tumor detection (780) process according to one embodiment of the present disclosure. As shown, a tumor cell fraction (TCF) (730) for a given patient can be calculated based on plasma sequencing data (710) and a reference sequence (720). A corrected tumor cell fraction (760) can also be calculated using a random error rate (750). Subsequently, tumor detection (780) can be performed based on the corrected tumor cell fraction (760) and the patient's limit of detection (LOD) (770). In this case, the limit of detection (770) may be the limit of detection (680) for the given patient calculated in Figure 6.

[0096] In one embodiment, the tumor cell ratio (730) can be calculated based on plasma sequencing data (710) and a reference sequence (720). Specifically, the number of reads in the plasma sequencing data (710) that differ from the reference sequence (720) and the number of reads that match the reference sequence (720) can be determined. Thereafter, the tumor cell ratio (730) can be calculated based on the number of reads that differ from the reference sequence (720) and the number of reads that match the reference sequence (720). For example, the tumor cell ratio (730) can be calculated by the following formula 2.

number

[0097] Here, Tumor Cell Fraction (raw) is the tumor cell ratio (730), JPEG2026514368000004.jpg5128 has a different number of reads from the reference array (720). JPEG2026514368000005.jpg6128 shows the number of reads that match the reference array (720).

[0098] Alternatively, the tumor cell ratio (730) can be calculated based on the filtered tumor tissue mutations (650) in Figure 6. Since the filtered tumor tissue mutations (650) information in Figure 6 includes information associated with the regions in the tumor cell genome where mutations occurred, the filtered tumor tissue mutations can be used to calculate the number of reads containing mutations detected in tumor tissue.

[0099] Specifically, based on the filtered tumor tissue mutation information, the number of reads containing mutations detected in the tumor tissue in the patient's plasma sequencing data is determined, and the number of determined reads is calculated using Equation 2. It can be used as JPEG2026514368000006.jpg5128. For example, when the number of filtered tumor tissue mutations is 2,000, the plasma sequencing data can be used to check the regions corresponding to those 2,000 tumor tissue mutations to determine whether the mutations are present. For example, if plasma ctDNA is sequenced at 20x, the number of reads containing mutations detected in tumor tissue can be determined from 20 * 2,000 = 40,000 reads. The number of reads containing mutations detected in tumor tissue in the plasma sequencing data, determined using the filtered tumor tissue mutation (650) information in Figure 6, may be substantially the same as, or very similar to, the number of reads that differ from, the reference sequence (720). This can reduce the time and computing resources required to compare the entire plasma sequencing data (710) with the reference sequence (720).

[0100] On the other hand, the random error rate (750) can be calculated based on genetic data and reference sequences (720) associated with multiple sample patients distinct from the target patient. The genetic data associated with multiple sample patients can be obtained from a genetic database (740). Here, the genetic database (740) may be the same database as the genetic database (640) in Figure 6.

[0101] The genetic data associated with multiple sample patients included in the genetic database (740) may include plasma sequencing data from multiple sample patients. In one embodiment, after determining the number of reads that differ from the reference sequence (720) and the number of reads that match the reference sequence (720) in the plasma sequencing data from multiple sample patients, the random error rate (750) can be calculated based on each read count. For example, the random error rate (750) can be calculated based on the following formula 3.

number

[0102] Here, the Random Error Rate is the random error rate (750), JPEG2026514368000008.jpg5128 shows the number of reads that differ from the reference sequence (720) in plasma sequencing data from multiple sample patients. JPEG2026514368000009.jpg6128 shows the number of reads matching the reference sequence (720) in plasma sequencing data from multiple sample patients.

[0103] Subsequently, the tumor cell ratio (730) can be corrected using the random error rate (750). For example, the corrected tumor cell ratio (760) may be the tumor cell ratio (730) minus the random error rate (750).

[0104] Subsequently, tumor detection (780) (e.g., determination of tumor recurrence) in the patient can be performed based on the corrected tumor cell ratio (760) and the detection limit (770). For example, a confidence interval for a predetermined confidence level can be calculated for the corrected tumor cell ratio (760). For example, if a 95% confidence level is used, the confidence interval for the 95% confidence level can be calculated via the following formula 4.

number

[0105] Here, TCF represents the corrected tumor cell ratio (760).

[0106] Subsequently, if the lower limit of the confidence interval for the calculated corrected tumor cell ratio (760) is determined to be higher than the detection limit (770), the patient's tumor may be determined to have recurred (tumor detection (780)). If the patient's tumor is determined to have recurred, tumor treatment / surgery may be performed on the patient.

[0107] Contrary to what is illustrated and explained in Figure 7, tumor detection (780) can be performed using structural variants. In one example, the presence or absence of tumor recurrence in a patient can be determined based on a comparison of a first set of structural variants detected from tumor tissue sequencing data and a second set of structural variants detected from plasma sequencing data. For example, if the number or ratio of structural variants in the second set of structural variants within the first set of structural variants is greater than or equal to a predetermined threshold, the patient may be determined to have experienced tumor recurrence. For example, the predetermined threshold may be 0.9. In this case, the first set of structural variants may include at least one of the following: inversion, translocation, duplication, deletion, or insertion of a region of the genome.

[0108] Figure 8 shows an example of graphs (810, 820) illustrating the relationship between the number of mutations detected in tumor cells of a target patient and the detection limit of that patient according to one embodiment of the present disclosure. Graphs (810, 820) show the actual measurement results of the detection limit for each sample patient based on ctDNA samples from 33 sample patients. Specifically, the first graph (810) shows the detection limit for all 33 sample patients, and the second graph (820) shows the detection limit within a specific range from the first graph (810).

[0109] In Graph 1 (810), the detection limit for all sample patients ranges from approximately 1.48E-06 to 1.49E-03, with a median of approximately 1.99E-05. For 29 out of 33 sample patients (approximately 87%), the detection limit is less than 1E-4. For sample patients in whom ctDNA matching tumor cells with more than 20,000 mutations was detected, the detection limit drops to 1E-6. In other words, the detection limit for a sample patient correlates with the somatic variant count of the tumor cells in that patient. A higher mutation count tends to result in a lower detection limit—the lowest tumor cell percentage detectable in the patient's plasma sample—making tumor detection easier.

[0110] Figure 9 shows an example of how an Arm-level Copy Number Profile (912, 922, 932) is generated by one embodiment of the present disclosure. In addition to or as an alternative to the methods described in Figures 6 and 7, minimal residual disease can be detected using the Arm-level Copy Number Profile.

[0111] In one embodiment, a copy number profile (912) at the chromosome 1 arm level can be generated from tumor tissue sequencing data (910) of the target patient, a copy number profile (922) at the chromosome 2 arm level can be generated from normal blood sequencing data (920), and a copy number profile (932) at the chromosome 3 arm level can be generated from plasma sequencing data (930). In this case, the tumor tissue sequencing data (910), normal blood sequencing data (920), and plasma sequencing data (930) can be obtained from tumor tissue biopsy samples, normal blood samples, and plasma samples, respectively, of the target patient.

[0112] In one embodiment, arm-level copy number profiles (912, 922, 932) can be generated by mapping each of the sequencing data (910, 920, 930) to a reference genome. For example, by obtaining genomic position information for each of the sequencing data (910, 920, 930), counting the sequencing coverage (or depth or number of reads) of the reads mapped to each chromosome arm, and then dividing by the genome-wide sequencing coverage of the WGS data, a normalized chromosome arm-level copy number profile (912, 922, 932) can be obtained.

[0113] Figure 10 shows an example of how the in vivo tumor cell fraction of a patient is calculated using chromosome arm-level copy number profiles (1010, 1020, 1030) according to one embodiment of the present disclosure. The copy number profile (1030) at the third chromosome arm level, generated from the plasma sequencing data of the patient, can be considered as a mixture of the copy number profile (1010) at the first chromosome arm level, generated from the tumor tissue sequencing data of the patient, and the copy number profile (1020) at the second chromosome arm level, generated from the normal blood sequencing data of the patient, in a specific ratio. For example, as shown, the copy number profile (1010) at the first chromosome arm level and the copy number profile (1020) at the second chromosome arm level can be considered as a mixture in the ratio a:(1-a) (where a is a real number between 0 and 1), where a is the tumor cell fraction (TCF) to the patient (e.g., 730 in Figure 7). The a value (i.e., the tumor cell ratio) can be calculated using the non-negative least squares (NNLS) method, expressed by equation 5 below.

number

[0114] Here, JPEG2026514368000012.jpg5128 exhibits the Euclidian norm, where A is a matrix associated with the copy number profiles (10¹⁰) of the 1st chromosome arm unit and the copy number profiles (10²⁰) of the 2nd chromosome arm unit, and y is a vector associated with the copy number profile (10³⁰) of the 3rd chromosome arm unit.

[0115] Figure 11 shows the results of performance verification of a method for detecting minimal residual disease in a target patient using the genetic data of a sample patient according to one embodiment of the present disclosure. The first graph (1110) shows the results of verifying how accurately a residual disease detection method using the genetic data of a sample patient (for example, the method illustrated and described in Figures 6 and 7) can measure the tumor cell fraction (TCF) after mixing tumor tissue sequencing data and normal blood sequencing data with known tumor cell ratios in various ratios (in silico mixture). The second graph (1120) shows the results of verifying how accurately a residual disease detection method using sample patient data can measure the tumor cell ratio after mixing actual tumor tissue DNA with normal DNA (experimental mixture).

[0116] As shown in Graph 1 (1110) and Graph 2 (1120), a correlation of over 99.9% was confirmed between the ground truth tumor cell ratio (Answer) and the predicted tumor cell ratio (Predicted) according to this disclosure, both when sequencing data was mixed and when DNA was mixed. In other words, it can be confirmed that the tumor cell ratio in the target patient can be predicted with great accuracy by filtering tumor tissue mutations in the target patient using the genetic data of the sample patient.

[0117] Figure 12 shows the performance verification results of a method for detecting minimal residual disease in a subject patient using a copy number profile at the chromosome arm level according to one embodiment of the present disclosure. Similar to Figure 11, the performance verification of the method for detecting minimal residual disease in a subject patient using a copy number profile at the chromosome arm level (e.g., the method illustrated and described in Figures 9 and 10) was performed on an in silico mixture (first graph (1210)) and an experimental mixture (second graph (1220)).

[0118] As shown in Graph 1 (1210) and Graph 2 (1220), a correlation of over 99.9% was observed between the ground truth tumor cell ratio (Answer) and the predicted tumor cell ratio (Predicted) according to this disclosure, both when sequencing data was mixed or DNA was mixed. In other words, it can be confirmed that the tumor cell ratio in the patient's body can be predicted with great accuracy using the copy number profile at the chromosomal arm level.

[0119] Figure 13 is a flowchart of a minimal residual disease detection method (1300) utilizing tumor information according to one embodiment of the present disclosure. Method (1300) may be performed by at least one processor of a minimal residual disease detection system utilizing tumor information. Method (1300) may be initiated by the processor acquiring first sequencing data associated with a first sample of a patient (S1310). In this case, the first sample may be a tumor tissue biopsy sample of the patient, and the first sequencing data may be acquired via whole genome sequencing (WGS).

[0120] Subsequently, the processor may acquire second sequencing data associated with the patient's second sample (S1320). In this case, the second sample is a normal blood sample from the patient, and the second sequencing data may be acquired via whole-genome sequencing.

[0121] Subsequently, the processor may acquire third sequencing data associated with the patient's third sample (S1330). In this case, the third sample is the patient's plasma sample, which may contain cfDNA (cell-free deoxyribo-nucleic acid) and ctDNA (circulating tumor deoxyribo-nucleic acid). The third sequencing data may be acquired via whole-genome sequencing.

[0122] In one embodiment, the first and second samples are samples taken at the first time point, the third sample is a sample taken at the second time point after the first time point, and at least one of surgery or therapeutic treatment may be performed on the patient after the first time point and before the second time point.

[0123] The processor may compare the first and second sequencing data to detect tumor tissue mutation information in the patient and perform background error filtering on the detected tumor tissue mutation information using genetic data associated with multiple sample patients that are distinguishable from the patient. In this case, the genetic data may include sequencing data for normal blood samples from multiple sample patients and / or sequencing data for plasma samples from multiple sample patients.

[0124] In one embodiment, the processor may perform background error filtering by removing from the tumor tissue mutation information mutations that were detected in sequencing data of normal blood samples from multiple sample patients. In another embodiment, the processor may perform background error filtering by removing from the tumor tissue mutation information mutations that were detected in sequencing data of plasma samples from multiple sample patients.

[0125] Subsequently, the processor may perform minimal residual disease detection for the patient based on the first sequencing data, the second sequencing data, and the third sequencing data (S1340). In one embodiment, the processor may perform minimal residual disease detection for the patient by calculating a limit of detection value for the patient based on the first sequencing data, the second sequencing data, and the third sequencing data, calculating the tumor cell fraction (TCF) for the patient based on the third sequencing data, correcting the tumor cell fraction using genetic data associated with multiple sample patients distinguished from the patient, and determining whether the patient's tumor is recurrent based on the corrected tumor cell fraction and the limit of detection value. In this case, the limit of detection value may represent the lowest tumor cell fraction that can be detected in the patient's plasma sample. The limit of detection value may also be calculated based on the number of tumor tissue mutations in the patient detected by comparing the first sequencing data and the second sequencing data, and the average sequencing depth of the third sequencing data.

[0126] In one embodiment, the processor may determine the number of reads in the third sequencing data that differ from the reference sequence, determine the number of reads in the third sequencing data that match the reference sequence, and calculate the tumor cell ratio based on the number of reads that differ from the reference sequence and the number of reads that match the reference sequence. At this time, the processor may compare the first sequencing data and the second sequencing data to detect tumor tissue mutation information of the patient, determine the number of reads in the third sequencing data that contain mutations detected in the tumor tissue based on the tumor tissue mutation information, and use the number of reads in the third sequencing data that contain mutations detected in the tumor tissue as the number of reads that differ from the reference sequence.

[0127] In one embodiment, the processor determines the number of reads in plasma sequencing data included in genetic data associated with multiple sample patients, which are distinguishable from the patient, that differ from the reference sequence, determines the number of reads in plasma sequencing data included in genetic data associated with multiple sample patients, which differ from the patient, calculates the random error rate based on the number of reads that differ from the reference sequence and the number of reads that match the reference sequence, and can use the random error rate to correct the tumor cell ratio.

[0128] In one embodiment, the processor may determine that the patient's tumor has recurred in response to determining that the lower limit of the calculated confidence interval is higher than the calculated detection limit, after calculating a confidence interval for a predetermined confidence level of the corrected tumor cell ratio.

[0129] In one embodiment, the processor may identify a first set of structural mutations detected from first sequencing data, identify a second set of structural mutations detected from third sequencing data, and determine whether the patient's tumor has recurred based on a comparison of the first set of structural mutations and the second set of structural mutations. For example, the processor may determine that the patient has recurred a tumor if the number or ratio of structural mutations included in the second set of structural mutations among the first set of structural mutations is greater than or equal to a threshold. In this case, the first set of structural mutations may include at least one of the following: inversion, translocation, duplication, deletion, or insertion of a region of the genome.

[0130] Alternatively, the processor may perform minimal residual disease detection by generating a chromosomal arm-level copy number profile based on first sequencing data, a second chromosomal arm-level copy number profile based on second sequencing data, a third chromosomal arm-level copy number profile based on third sequencing data, and calculating the tumor cell ratio based on the first to third chromosomal arm-level copy number profiles.

[0131] The flowchart shown in Figure 13 and the explanation described above are merely examples and may be implemented differently in some embodiments. For example, one or more steps may be omitted, the order of the steps may be changed, one or more steps may be performed redundantly, or one or more steps may be performed repeatedly.

[0132] The methods described above may be provided as computer programs stored on a computer-readable recording medium for execution on a computer. The medium may continuously store computer-executable programs or temporarily store them for execution or download. The medium may also be a variety of recording or storage means in the form of a combination of one or more hardware components, and may not be limited to a medium directly connected to a computer system, but may be distributed on a network. Examples of mediums may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical mediums such as floptical disks; and devices configured to store program instructions, including ROM, RAM, and flash memory. Other examples of mediums include recording or storage media managed by app stores that distribute applications and other sites and servers that supply or distribute various software.

[0133] The methods, operations, or techniques described herein can be embodied by a variety of means. For example, these techniques can be embodied in hardware, firmware, software, or a combination thereof. A person of ordinary skill will understand that the various exemplary logical blocks, modules, circuits, and algorithmic steps described in conjunction with the disclosure may be embodied in electronic hardware, computer software, or a combination thereof. To clearly illustrate such interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been generally described above in terms of their functional aspects. Whether such functions are embodied as hardware or software depends on the design requirements imposed on the particular application and the overall system. A person of ordinary skill may embodied the functions described in a variety of ways for their respective specific applications, but such embodiments should not be construed as exceeding the scope of this disclosure.

[0134] In hardware implementation, the processing units used to perform the techniques may be embodied in one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described herein, computers, or combinations thereof.

[0135] Accordingly, the diverse exemplary logic blocks, modules, and circuits described in conjunction with this disclosure may be embodied or performed by any combination of general-purpose processors, DSPs, ASICs, FPGAs or other programmable logic devices, discrete gates and transistor logic, discrete hardware components, or any combination of those designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, a processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be embodied as a combination of computing devices, e.g., a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other combination of configurations.

[0136] In firmware and / or software embodiments, the technique may be embodied as instructions stored on a computer-readable medium such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, compact disc (CD), or magnetic or optical data storage devices. The instructions may be executable by one or more processors, causing the processors to perform certain aspects of the functions described herein.

[0137] Although the embodiments described above are described as utilizing aspects of the subject matter currently disclosed in one or more standalone computer systems, the disclosure is not limited to and can be embodied in any computing environment, such as networks or distributed computing environments. Furthermore, aspects of the subject matter in the disclosure can be embodied in multiple processing chips or devices, and storage may be similarly affected across multiple devices. These devices may include PCs, network servers, and portable devices.

[0138] While this disclosure has been described in relation to some embodiments, various modifications and alterations are possible, provided that they do not deviate from the scope of this disclosure as understandable to a person with ordinary skill in the art to which the invention of this disclosure pertains. Such modifications and alterations should be considered to fall within the scope of the claims appended to this specification.

Claims

1. A method for detecting minimal residual disease, which is performed by at least one processor, A step of obtaining first sequencing data associated with the first sample of the patient, The steps include obtaining second sequencing data associated with a second sample of the aforementioned patient, The steps include obtaining third sequencing data associated with the third sample of the aforementioned patient, The steps include: performing minimal residual disease detection for the patient based on the first sequencing data, the second sequencing data, and the third sequencing data; A method for detecting minimal residual lesions, including the above.

2. The method for detecting minimal residual disease according to claim 1, wherein the first sequencing data, the second sequencing data, and the third sequencing data are obtained via whole genome sequencing (WGS).

3. The first sample is a tumor tissue biopsy sample from the patient. The second sample is a normal blood sample from the patient. The third sample is a plasma sample from the patient. The method for detecting minimal residual disease according to claim 1, wherein the plasma sample comprises cfDNA (cell-free Deoxyribo Nucleic Acid) and ctDNA (circulating tumor Deoxyribo Nucleic Acid).

4. The first and second samples are samples acquired at the first time point. The third sample is a sample acquired at a second time point after the first time point. The method for detecting minimal residual disease according to claim 1, wherein at least one of surgery or therapeutic treatment is performed on the patient after the first time point and before the second time point.

5. A step of comparing the first sequencing data and the second sequencing data to detect tumor tissue mutation information of the patient, The steps include: performing background error filtering on the detected tumor tissue mutation information using genetic data associated with multiple sample patients that are distinguished from the aforementioned patient; The method for detecting minute residual lesions according to claim 1, further comprising:

6. The genetic data includes sequencing data for normal blood samples from the plurality of sample patients. The step of performing the aforementioned filtering is: The step of removing from the tumor tissue mutation information mutations that were detected in sequencing data for normal blood samples from the plurality of sample patients among the detected tumor tissue mutations. A method for detecting minute residual lesions according to claim 5, including the method described in claim 5.

7. The genetic data includes sequencing data for plasma samples from the plurality of sample patients. The step of performing the aforementioned filtering is: The step of removing from the tumor tissue mutation information mutations that are detected in the sequencing data of the plasma samples of the multiple sample patients among the detected tumor tissue mutations. A method for detecting minute residual lesions according to claim 5, including the method described in claim 5.

8. The step of performing the detection of minute residual lesions is, A step of calculating a Limit of Detection value for the patient based on the first sequencing data, the second sequencing data, and the third sequencing data, A step of calculating the tumor cell fraction (TCF) for the patient based on the third sequencing data, The steps include correcting the tumor cell ratio using genetic data associated with multiple sample patients distinct from the aforementioned patient, A step of determining whether the patient's tumor is recurrent based on the corrected tumor cell ratio and the detection limit, A method for detecting minute residual lesions according to claim 1, including the method described in claim 1.

9. The method for detecting minimal residual disease according to claim 8, wherein the Limit of Detection value means the lowest tumor cell percentage that can be detected in the patient's plasma sample.

10. The method for detecting minimal residual disease according to claim 8, wherein the detection limit is calculated based on the number of tumor tissue mutations of the patient detected by comparing the first sequencing data and the second sequencing data and the average sequencing depth of the third sequencing data.

11. The step of calculating the tumor cell ratio is: The third sequencing data includes a step of determining the number of reads that differ from the reference sequence, The third step of determining the number of reads in the sequencing data that match the reference sequence, A step of calculating the tumor cell ratio based on the number of reads that differ from the reference sequence and the number of reads that match the reference sequence, A method for detecting minute residual lesions according to claim 8, including the method described in claim 8.

12. The step of determining the number of differing leads is: A step of comparing the first sequencing data and the second sequencing data to detect tumor tissue mutation information of the patient, The steps include determining the number of reads containing mutations detected in tumor tissue in the third sequencing data based on the tumor tissue mutation information, The third step of using the number of reads containing mutations detected in tumor tissue in the third sequencing data as the number of reads that differ from the reference sequence, A method for detecting minute residual lesions according to claim 11, including the method described in claim 11.

13. The step of correcting the tumor cell ratio is: The steps include determining the number of reads that differ from the reference sequence in plasma sequencing data included in the genetic data associated with multiple sample patients distinct from the aforementioned patient, The steps include determining the number of reads that match a reference sequence in plasma sequencing data included in the genetic data associated with multiple sample patients distinct from the aforementioned patient, A step of calculating the random error rate based on the number of reads that differ from the reference sequence and the number of reads that match the reference sequence, A step of correcting the tumor cell ratio using the aforementioned random error rate, A method for detecting minute residual lesions according to claim 8, including the method described in claim 8.

14. The step of determining whether the tumor in the aforementioned patient is likely to recur is: The steps include: calculating the confidence interval of a predetermined confidence level for the corrected tumor cell ratio; In response to determining that the lower limit of the calculated confidence interval is higher than the calculated detection limit, the step of determining that the patient's tumor has recurred, A method for detecting minute residual lesions according to claim 8, including the method described in claim 8.

15. The step of performing the detection of minute residual lesions is, The steps include generating a chromosomal arm-level copy number profile based on the first sequencing data, The steps include generating a copy number profile at the chromosome 2 arm level based on the second sequencing data, The steps include generating a copy number profile at the third chromosome arm level based on the third sequencing data, A step of calculating the tumor cell ratio based on the copy number profiles at the level of the first to third chromosome arms, The method for detecting minute residual lesions according to claim 1, further comprising:

16. The step of performing the detection of minute residual lesions is, The steps include identifying a first set of structural mutations detected from the first sequencing data, The steps include identifying a second set of structural mutations detected from the third sequencing data, A step of determining whether the patient's tumor is likely to recur based on the results of comparing the first set of structural mutations with the second set of structural mutations, The method for detecting minute residual lesions according to claim 3, further comprising:

17. The step of determining whether the tumor in the aforementioned patient is likely to recur is: If the number or ratio of structural mutations included in the second set of structural mutations among the first set of structural mutations is greater than or equal to a threshold, the step of determining that the tumor has recurred in the patient. A method for detecting minute residual lesions according to claim 16, including the method described in claim 16.

18. The method for detecting minimal residual disease according to claim 16, wherein the first set of structural mutations includes at least one of inversion, translocation, duplication, deletion, or insertion of a region of the genome.

19. A computer-readable non-temporary recording medium that records instructions for executing the method according to claim 1 on a computer.

20. It is a system, Communication module and Memory and A processor connected to the memory and configured to execute at least one computer-readable program contained in the memory, Includes, The aforementioned at least one program, First sequencing data associated with the first sample of the patient was obtained. Second sequencing data related to the second sample of the aforementioned patient was obtained. Third sequencing data associated with the third sample of the aforementioned patient is obtained, A system including commands for performing minimal residual disease detection on the patient based on the first sequencing data, the second sequencing data, and the third sequencing data.