Method and apparatus for training machine learning model for detecting true positive variant in cell sample

A machine learning model trained on FFPE-processed tissues corrects noise in mutation detection data by distinguishing true positive mutations, ensuring accurate genomic profiling and enabling applications in personalized medicine and genetic research.

WO2025154893A1PCT designated stage expired Publication Date: 2025-07-24INOCRAS KOREA INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/012126
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-19
Filing Date
2024-08-14
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

The use of Formalin-Fixed, Paraffin-Embedded (FFPE)-processed tissues for genetic analysis in clinical settings leads to DNA damage, causing noise in mutation detection data and inaccurate analysis results due to cross-linking, fragmentation, and mutation of DNA bases, which are not present in Fresh-Frozen (FF)-processed tissues.

Method used

A machine learning model is trained using reference mutation candidate information and annotation information to distinguish between true positive and false positive mutations in FFPE-processed tissues, employing a method that includes obtaining reference mutation candidate information, generating annotation information, and training the model with labeled data to correct noise and errors in mutation detection.

Benefits of technology

This approach allows for accurate genomic profiling on FFPE-processed tissues, providing precise genetic information for medical research and clinical applications without requiring additional equipment, and enabling applications in early disease diagnosis, personalized medicine, and genetic disease research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024012126_24072025_PF_FP_ABST
    Figure KR2024012126_24072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method for training a machine learning model. The method for training a machine learning model comprises the steps of: acquiring reference variant candidate information of a reference sample; generating annotation information associated with a reference variant candidate; generating training data on the basis of the acquired reference variant candidate information and the generated annotation information; and training the machine learning model by using the generated training data.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for training a machine learning model for detecting true positive mutations in cell samples

[0001] The present disclosure relates to a method and apparatus for training a machine learning model, and more particularly, to a method and apparatus for training a machine learning model using reference mutation candidate information and annotation information.

[0002] Genetic information analysis technology is widely used in the medical field, such as determining the characteristics or temperament of an organism by understanding its genetic information. Recently, medical practices aimed at understanding the causes and treating various diseases, such as cancer, are evolving from a traditional prescription-centered approach to precision medicine, a form of personalized treatment that considers an individual's genetic information and health history. The field of precision medicine relies on acquiring massive amounts of individual genetic information and performing related clinical analyses. These key elements are key to accelerating the advancement of precision medicine technology.

[0003] In particular, when performing whole-genome analysis on tissues collected from an individual, the so-called "fresh frozen (FF)" processing method, in which the tissue is frozen immediately after collection, is widely used. FF-processed tissue is known to be the optimal tissue processing method for whole-genome analysis because it is frozen immediately after collection, which results in less DNA damage to cells within the tissue. However, there is the problem that FF-processing or storing collected tissue requires facilities or equipment, such as nitrogen tanks, that are not or are difficult to obtain in clinical settings.

[0004] On the other hand, when resecting tumor tissue or conducting a biopsy to analyze the genetic information of an individual, medical institutions usually process the tissue (such as tumor tissue) collected from the individual into FFPE (Formalin-Fixed, Paraffin-Embedded) for long-term storage and utilize the FFPE-processed tissue for follow-up testing or academic research purposes. When processing tissue collected through the FFPE method, not only does it not require a large cost and effort for processing and storing the collected tissue, but the tissue can be stored for a long period of time with most of the genetic information within the tissue maintained, making it easy to utilize the collected tissue in the future (such as re-examination or re-analysis of the tissue).

[0005] However, during the long-term storage of tissues collected from an individual through FFPE processing, various types of damage can occur to the DNA within the tissue, such as cross-linking, where different parts of DNA chemically become entangled with each other, fragmentation, where DNA is cut into smaller pieces, and mutations in DNA bases due to other non-biological causes.

[0006] Due to DNA damage as described above, when performing whole-genome analysis and mutation analysis on FFPE-processed tissue, noise can be introduced into mutation detection data, leading to inaccurate and distorted analysis results. This noise is typically not detected in whole-genome analysis data from FFPE-processed tissue. Therefore, to obtain undistorted analysis results from FFPE-processed tissue, it is necessary to effectively process or remove noise in mutation detection data.

[0007] The present disclosure provides a method for training a machine learning model to solve the above-described problems, a computer-readable non-transitory recording medium recording commands, and a device (system).

[0008] The present disclosure can be implemented in various ways, including a computer-readable, non-transitory recording medium having recorded thereon a method, system (device), or instructions.

[0009] A method for training a machine learning model, which is executed by at least one processor, comprises the steps of: obtaining reference mutation candidate information of a reference sample; generating annotation information associated with the reference mutation candidate; generating training data based on the obtained reference mutation candidate information and the generated annotation information; and training the machine learning model using the generated training data.

[0010] In one embodiment of the present disclosure, the reference sample includes a reference normal sample and a reference abnormal sample collected from the same individual, and the obtained reference variant candidate information is determined based on first reference sequencing data associated with the reference normal sample and the reference abnormal sample and second reference sequencing data using a variant detection module.

[0011] In one embodiment of the present disclosure, the mutation detection module includes a plurality of detection modules, the obtained reference mutation candidate information is determined by unioning reference mutation sub-candidate information obtained using the plurality of detection modules, and the obtained reference mutation sub-candidate information is determined by applying first reference sequencing data and second reference sequencing data to each of the plurality of detection modules.

[0012] In one embodiment of the present disclosure, the step of generating annotation information includes the step of determining a plurality of reads in which at least some of the mapped positions overlap with positions of reference mutation candidates, and the step of generating first annotation information associated with the determined plurality of reads.

[0013] In one embodiment of the present disclosure, the plurality of reads include a plurality of variant reads that are different from a reference genome, and the first annotation information includes at least one of a minimum value of an insert size of the plurality of variant reads, a maximum value of an insert size of the plurality of variant reads, or a number of paired reads that satisfy a specific condition among the plurality of variant reads, wherein the specific condition includes a condition in which a first read and a second read of the paired reads are aligned in a forward direction and a reverse direction, respectively, and an insert size of the paired read is between a lower threshold value and an upper threshold value.

[0014] In one embodiment of the present disclosure, the step of generating annotation information includes the step of receiving normal tissue genomic data (PON: Panel of Normals) generated from a plurality of sequencing data associated with a plurality of normal samples, and the step of generating second annotation information associated with the normal tissue genomic data.

[0015] In one embodiment of the present disclosure, the step of generating annotation information includes the step of receiving FFPE processed tissue genomic data (POF: Panel of FFPEs) generated from a plurality of sequencing data associated with a plurality of Formalin-Fixed, Paraffin-Embedded (FFPE) processed samples, and the step of generating third annotation information associated with the FFPE processed tissue genomic data.

[0016] In one embodiment of the present disclosure, the third annotation information includes the number of samples among a plurality of FFPE-processed samples in which a Variant Allele Frequency (VAF) for a specific position in the base sequence within the sample is less than a predetermined threshold.

[0017] In one embodiment of the present disclosure, the third annotation information includes the number of samples among a plurality of FFPE-processed samples having a predetermined number of variant reads at a specific location on the base sequence within the sample.

[0018] In one embodiment of the present disclosure, the step of generating annotation information includes the step of generating fourth annotation information including information associated with a mutation type of a reference mutation candidate and sequence context information of the reference mutation candidate.

[0019] In one embodiment of the present disclosure, the step of generating learning data includes the step of labeling classification information for reference variant candidates.

[0020] In one embodiment of the present disclosure, the reference sample is an FFPE-processed sample, and the labeling step includes labeling the reference variant candidate as a true positive variant in response to determining that at least a portion of the reference variant candidate information and at least a portion of information associated with one of the variant candidates in the FF (Fresh-Frozen)-processed sample correspond to each other, and the FF-processed sample is a sample corresponding to the FFPE-processed sample.

[0021] In one embodiment of the present disclosure, the labeling step further includes a step of labeling the reference variant candidate as a false positive variant in response to determining that the reference variant candidate information and the information associated with any one of the variant candidates in the FF-processed sample do not correspond to each other.

[0022] In one embodiment of the present disclosure, the step of generating learning data further includes the step of extracting features of a reference variant candidate based on reference variant candidate information and annotation information, and the step of including a data set including the reference variant candidate information, the extracted features of the reference variant candidate, and labeled classification information in the learning data.

[0023] In one embodiment of the present disclosure, the machine learning model includes a plurality of classifiers, and the step of training the machine learning model includes the steps of inputting reference mutation candidate information and features of the reference mutation candidate into each of the plurality of classifiers, determining a classification result indicating whether the reference mutation candidate is a true positive mutation using an output result from at least one classifier among the plurality of classifiers, and adjusting parameters of the machine learning model based on the classification result and classification information labeled on the reference mutation candidate.

[0024] In one embodiment of the present disclosure, a machine learning model receives target mutation candidate information and features of the target mutation candidate in a target sample, and outputs a classification result indicating whether the target mutation candidate is a true positive mutation.

[0025] In one embodiment of the present disclosure, the target sample includes a target normal sample and a target abnormal sample collected from the same individual, and the target variant candidate information is determined based on first target sequencing data associated with the target normal sample and second target sequencing data associated with the target abnormal sample using a variant detection module.

[0026] In one embodiment of the present disclosure, the target sample is an FFPE processed sample.

[0027] A method for genomic profiling through detection of a true-positive mutation in a cell sample, executed by at least one processor according to one embodiment of the present disclosure, comprises the steps of: obtaining target mutation candidate information of a target sample; determining a classification result indicating whether the target mutation candidate is a true-positive mutation using a machine learning model; and performing genomic profiling on the target sample based on the determined classification result, wherein the machine learning model is trained to determine whether the reference mutation candidate is a true-positive mutation using reference mutation candidate information of a reference sample and annotation information associated with the reference mutation candidate.

[0028] In one embodiment of the present disclosure, the method further includes providing at least one of disease diagnosis information, treatment strategy information, prognosis prediction information, or drug responsiveness prediction information of an individual from whom a target sample was collected based on the results of performing genetic profiling.

[0029] A computer-readable non-transitory recording medium having recorded thereon instructions for executing a method for training a machine learning model and / or a method for genetic profiling through detection of true-positive mutations in a cell sample according to one embodiment of the present disclosure is provided.

[0030] In one embodiment of the present disclosure, a device comprises a memory and at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory, wherein the at least one program comprises instructions for obtaining reference mutation candidate information of a reference sample, generating annotation information associated with the reference mutation candidate information, generating learning data based on the obtained reference mutation candidate information and the generated annotation information, and training a machine learning model using the generated learning data.

[0031] According to various embodiments of the present disclosure, if a specific mutation candidate in an abnormal sample is determined to be a false positive mutation, the specific mutation candidate is deleted / filtered from the mutation candidate list, thereby determining a highly accurate mutation list.

[0032] According to various embodiments of the present disclosure, by correcting noise or errors that may occur when performing whole-genome analysis on FFPE-processed tissue, undistorted analysis results can be derived, as in whole-genome analysis data of FFPE-processed tissue.

[0033] According to various embodiments of the present disclosure, not only can whole genome analysis be performed with high accuracy on a large amount of FFPE tissues secured and accumulated by medical institutions and biobanks, but whole genome analysis on tissue samples from patients, etc. can also be performed without changing the tissue sample processing procedures of medical institutions using only facilities equipped in a typical clinical setting.

[0034] According to various embodiments of the present disclosure, genome profiling can be applied in various fields such as early diagnosis of diseases, personalized medicine, genetic disease research, drug development, etc., and by performing genome profiling based on classification results indicating whether a target mutation candidate is a true mutation, precise and extensive genetic information can be provided for medical research and clinical applications.

[0035] The effects of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs (referred to as “one skilled in the art”) from the description of the claims.

[0036] Embodiments of the present disclosure will be described below with reference to the accompanying drawings, wherein like reference numerals represent similar elements, but are not limited thereto.

[0037] FIG. 1 is a diagram showing an example in which a classification result of a mutation candidate is determined using a machine learning model according to one embodiment of the present disclosure.

[0038] FIG. 2 is a block diagram showing the internal configuration of a computing device for performing a learning process and an inference process of a machine learning model according to one embodiment of the present disclosure.

[0039] FIG. 3 is a diagram illustrating an example of a mutation detection module according to one embodiment of the present disclosure.

[0040] FIG. 4 is a diagram illustrating an example of an annotation module according to one embodiment of the present disclosure.

[0041] FIG. 5 is a diagram illustrating an example of a feature extraction module according to one embodiment of the present disclosure.

[0042] FIG. 6 is a diagram showing an example of outputting a classification result of a mutation candidate using a machine learning model according to one embodiment of the present disclosure.

[0043] FIG. 7 is a diagram showing a detailed configuration of a machine learning model according to one embodiment of the present disclosure.

[0044] FIG. 8 is a diagram illustrating an example of a learning process of a machine learning model according to one embodiment of the present disclosure.

[0045] FIG. 9 is a diagram illustrating an example of an inference process of a machine learning model according to one embodiment of the present disclosure.

[0046] FIG. 10 is a diagram illustrating an example of learning data according to one embodiment of the present disclosure.

[0047] FIG. 11 is a diagram showing the performance evaluation results of a learned machine learning model according to one embodiment of the present disclosure.

[0048] FIG. 12 is an exemplary diagram showing an artificial neural network model according to one embodiment of the present disclosure.

[0049] FIG. 13 is a flowchart illustrating a method for training a machine learning model according to one embodiment of the present disclosure.

[0050] FIG. 14 is a flowchart illustrating a method for genetic profiling through detection of true positive mutations in cell samples according to one embodiment of the present disclosure.

[0051] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the attached drawings. However, in the following description, specific descriptions of widely known functions or configurations will be omitted if they may unnecessarily obscure the gist of the present disclosure.

[0052] In the attached drawings, identical or corresponding components are assigned the same reference numerals. Furthermore, in the description of the embodiments below, duplicate descriptions of identical or corresponding components may be omitted. However, even if a description of a component is omitted, it is not intended that such component is not included in any embodiment.

[0053] The advantages and features of the disclosed embodiments, and methods for achieving them, will become clearer with reference to the embodiments described below, along with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure the completeness of the disclosure and to fully inform those skilled in the art of the scope of the invention.

[0054] The terms used in this specification will be briefly explained, followed by a detailed description of the disclosed embodiments. The terms used in this specification have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of engineers working in the relevant field, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.

[0055] In this specification, singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, plural expressions include singular expressions unless the context clearly indicates otherwise. When a part of the specification is said to include a component, this does not exclude other components, but rather implies that other components may be included, unless otherwise specifically stated.

[0056] Also, the term 'module' or 'part' used in the specification means a software or hardware component, and the 'module' or 'part' performs certain roles. However, the 'module' or 'part' is not limited to software or hardware. The 'module' or 'part' may be configured to reside on an addressable storage medium and may be configured to execute one or more processors. Thus, as an example, the 'module' or 'part' may include at least one of components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, or variables. The functionality provided within the components and 'modules' or 'parts' may be combined into a smaller number of components and 'modules' or 'parts', or further separated into additional components and 'modules' or 'parts'.

[0057] According to one embodiment of the present disclosure, a 'module' or 'unit' may be implemented as a processor and a memory. 'Processor' should be broadly construed to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, and the like. In some circumstances, a 'processor' may also refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), and the like. A 'processor' may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors in conjunction with a DSP core, or any other such combination of configurations. In addition, 'memory' should be broadly construed to include any electronic component capable of storing electronic information. 'Memory' may refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage, registers, etc. Memory is said to be in electronic communication with the processor if the processor can read information from, and / or write information to, the memory. Memory integrated in a processor is in electronic communication with the processor.

[0058] In the present disclosure, the "system" may include, but is not limited to, at least one of a server device and a cloud device. For example, the system may be comprised of one or more server devices. As another example, the system may be comprised of one or more cloud devices. As yet another example, the system may be configured and operated by a combination of a server device and a cloud device.

[0059] In addition, terms such as first, second, A, B, (a), (b), etc. used in the following embodiments are only used to distinguish certain components from other components, and the nature, order, or sequence of the components are not limited by the terms.

[0060] Additionally, in the embodiments below, when it is described that a component is 'connected', 'coupled' or 'connected' to another component, it should be understood that the component may be directly connected or connected to the other component, but another component may also be 'connected', 'coupled' or 'connected' between each component.

[0061] Additionally, the terms 'comprises' and / or 'comprising' used in the following embodiments do not exclude the presence or addition of one or more other components, steps, operations and / or elements.

[0062] In the present disclosure, 'each of the plurality of As' may refer to each of all components included in the plurality of As, or may refer to each of some components included in the plurality of As.

[0063] Before describing various embodiments of the present disclosure, the terminology used will be explained.

[0064] In this disclosure, "Whole Genome Sequencing (WGS)" or "whole genome sequencing" may refer to a technique used to determine the entire DNA sequence of a genome. Specifically, whole genome sequencing may involve deciphering and identifying the order of nucleotide bases (adenine, cytosine, guanine, and thymine) within the entire set of genetic material of a human or organism. The entire set of genetic material may include all genes, noncoding regions, and any additional genetic elements present within the genome. In one embodiment, whole genome sequencing may be performed through several steps. For example, whole genome sequencing may be performed by extracting DNA from a specific cell or the like, fragmenting the extracted DNA into smaller pieces, and generating millions or billions of short DNA sequences, referred to as "reads." The resulting reads may be aligned and assembled to reconstruct the entire genome sequence.

[0065] In the present disclosure, 'sequencing data' may refer to data associated with a DNA (Deoxyribo Nucleic Acid) sequence or RNA (Ribo Nucleic Acid) sequence of a specific individual analyzed through a sequencing process.

[0066] In the present disclosure, 'X sample sequencing data' may refer to sequencing data generated through a sequencing process for 'X sample'.

[0067] In the present disclosure, 'abnormal cell' may refer to an abnormal cell having a size, shape, structure, function, etc. different from a normal cell, and 'abnormal sample' may refer to a sample containing an abnormal cell. Abnormal cells may be caused by various factors such as genetic mutation, infection, and exposure to toxins, and may include various types of abnormal cells such as cancer cells, tumor cells, necrotic cells, senescent cells, aneuploid cells, hyperplastic cells, and hypertrophic cells.

[0068] In the present disclosure, 'variation' may refer to various types of mutations, such as point mutations including single-nucleotide variants (hereinafter referred to as "SNVs"), short insertion-and-deletion variants (hereinafter referred to as "INDELs"), structural variations, and / or copy-number variants (CNVs).

[0069] In the present disclosure, a "mutation candidate" may refer to a DNA or RNA sequence with a probability of being a mutation greater than or equal to a predetermined threshold probability. For example, a mutation candidate may be a sequence with a probability of being a point mutation greater than or equal to a predetermined threshold probability.

[0070] In the present disclosure, 'annotation information' may refer to information for extracting features of mutation candidates, and may be distinguished from 'label' or 'classification information', which is a teacher signal (correct answer) given to data for learning a machine learning model in the present disclosure.

[0071] In the present disclosure, 'sequence context' may refer to a base sequence including one or more neighboring nucleotides surrounding a specific DNA or RNA base sequence.

[0072] In the present disclosure, 'union' may refer to a union operation.

[0073] In the present disclosure, genomic profiling may refer to a process of investigating genetic variations, structural changes, gene expression patterns, and / or other genetic information by analyzing the genome or DNA sequence of an individual.

[0074] In correspondence with the term 'target X' used in the inference process, the term used in the learning process can be defined as 'reference X'. For example, the 'target sample' may be a sample that is the target for inferring whether a target mutation candidate in the sample is a true positive mutation using a trained machine learning model, and the 'reference sample' may be a sample sequenced to generate training data used for training the machine learning model. Meanwhile, the term 'X' used without the expression 'target' or 'reference' and associated with the training process or the inference process of the machine learning model may refer to 'target X' and / or 'reference X' unless otherwise stated, and should be interpreted according to the context in which the term is used.

[0075] In this disclosure, ‘information associated with X’ and ‘X information’ may be used interchangeably with the same meaning.

[0076] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the attached drawings.

[0077] FIG. 1 is a diagram showing an example in which a classification result (132) of a mutation candidate is determined using a machine learning model (130) according to one embodiment of the present disclosure.

[0078] The entity (110) may be an entity having mutated tissue or cells, such as tumor tissue (e.g., cancer tissue). The entity (110) may be an entity providing a reference sample that serves as the basis for generating learning data in the learning process of a machine learning model (130), or may be a target entity for determining true positive mutations in a cell sample in an inference process using a learned machine learning model (130). The entity (110) is not limited to a human and may be any living organism.

[0079] From the subject (110), an abnormal sample, such as a tumor tissue biopsy sample, may be collected. The abnormal sample may be a sample containing abnormal cells that are the target of mutation detection.

[0080] Additionally, a normal sample may be collected from the same individual (110) from which the abnormal sample was collected. It may be assumed that the normal sample does not contain abnormal cells. The normal sample may include a normal blood sample or a normal cell sample, etc. As an example, the normal blood sample may be a thin buffy coat formed between a red blood cell layer at the bottom and a plasma layer at the top of a centrifuged sample collected from the individual (110).

[0081] Abnormal samples and normal samples collected from an individual (110) may be processed with FFPE (Formalin-Fixed, Paraffin-Embedded) or FF (Fresh-Frozen). Subsequently, abnormal sample sequencing data (112) may be generated from the FFPE or FF processed abnormal sample, and normal sample sequencing data (114) may be generated from the FFPE or FF processed normal sample. Here, the sequencing data (112, 114) may be obtained through whole genome sequencing (WGS) and / or target panel sequencing (TPS).

[0082] Sequencing data (112, 114) may be reference sequencing data that serves as the basis for generating learning data in the learning process of a machine learning model (130), or may be target sequencing data that is analyzed to detect / determine true positive mutations in cell samples in the inference process using the learned machine learning model (130).

[0083] For example, the sequencing data (112, 114) may correspond to the FFPE sample sequencing data (812) and the FF sample sequencing data (814) in FIG. 8, which illustrates the learning process of the machine learning model (130). In another example, the sequencing data (112, 114) may correspond to the FFPE sample sequencing data (910) in FIG. 9, which illustrates the inference process of the learned machine learning model (130). This will be described in detail later with reference to FIGS. 8 and 9, respectively.

[0084] The mutation detection module (120) can determine mutation candidate information (122) using abnormal sample sequencing data (112) and normal sample sequencing data (114). For example, the normal sample sequencing data (114) can be used as sequencing data to be compared with the abnormal sample sequencing data (112), so that genetic variations or mutations present in the DNA / RNA in the abnormal sample can be compared and identified. Alternatively, by performing deep sequencing using the abnormal sample, comparing it with a known mutation database, etc., the genetic variations or mutations present in the DNA / RNA in the abnormal sample can be identified using only the abnormal sample sequencing data (112) without using the normal sample sequencing data (114).

[0085] The mutation candidate information (122) determined in the mutation detection module (120) may include location information of the mutation candidate, reference allele information at the location of the mutation candidate, altered allele information corresponding to the mutation candidate, etc. The location information of the mutation candidate may include chromosome information (e.g., chromosome number) on which the mutation candidate is located, location information within the chromosome (e.g., 1-based position), etc. The mutation detection module (120) will be described in detail below using FIG. 3.

[0086] The machine learning model (130) can output a classification result (132) indicating whether the mutation candidate is a true positive mutation based on mutation candidate information (122). If a specific mutation candidate is determined to be a false positive mutation among a mutation candidate list including multiple mutation candidates presumed to be mutations present in an abnormal sample, the specific mutation candidate can be deleted / filtered from the mutation candidate list. Through this, a mutation list including mutations presumed to be actual mutations in the abnormal sample can be determined.

[0087] The detailed configuration and operation of the machine learning model (130) will be described in detail later using FIG. 6, FIG. 7, and FIG. 12.

[0088] FIG. 2 is a block diagram showing the internal configuration of a computing device (200) for performing a learning process and an inference process of a machine learning model according to one embodiment of the present disclosure. The computing device (200) may include a memory (210), a processor (220), a communication module (230), and an input / output interface (240). In one embodiment, the computing device (200) may be composed of a plurality of distributed computing devices, and each of the memory (210), the processor (220), the communication module (230), and the input / output interface (240) illustrated in FIG. 2 may collectively refer to a plurality of memories, a plurality of processors, etc. included in the plurality of distributed computing devices.

[0089] As illustrated in FIG. 2, the computing device (200) may be configured to communicate information and / or data with an external database, etc., via a network using the communication module (230). In one embodiment, the computing device (200) may be configured to communicate information and / or data with an external database, etc., via a network using the communication module (230). For example, the computing device (200) may be connected to a database including normal tissue genome data (PON: Panel of Normals) generated from a plurality of sequencing data associated with a plurality of normal samples (e.g., corresponding to the database (630) of FIG. 6)) and a database including feature extraction information (e.g., corresponding to the database (750) of FIG. 7), thereby transmitting and receiving information and / or data to and from each other.

[0090] The memory (210) may include any non-transitory computer-readable recording medium. According to one embodiment, the memory (210) may include a non-volatile mass storage device such as a random access memory (RAM), a read only memory (ROM), a disk drive, a solid state drive (SSD), a flash memory, etc. As another example, a non-volatile mass storage device such as a ROM, an SSD, a flash memory, a disk drive, etc. may be included in the computing device (200) as a separate permanent storage device distinct from the memory. In addition, an operating system and at least one program code may be stored in the memory (210).

[0091] These software components may be loaded from a computer-readable recording medium separate from the memory (210). This separate computer-readable recording medium may include a recording medium directly connectable to the computing device (200), for example, a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, a memory card, etc. As another example, the software components may be loaded into the memory (210) via a communication module (230) other than a computer-readable recording medium. For example, at least one program may be loaded into the memory (210) based on a computer program that is installed by files provided by developers or a file distribution system that distributes installation files of applications via the communication module (230).

[0092] The processor (220) may be configured to process commands, data, etc. of a computer program by performing basic arithmetic, logic, and input / output operations. The commands may be provided to a user terminal (not shown) or another external system via a memory (210) or a communication module (230). In addition, the processor (220) may be configured to manage, process, and / or store information and / or data received from multiple user terminals and / or multiple external systems.

[0093] The communication module (230) may provide a configuration or function for a user terminal (not shown) and a computing device (200) to communicate with each other through a network, and may provide a configuration or function for the computing device (200) to communicate with an external system.

[0094] In addition, the input / output interface (240) of the computing device (200) may be a means for interfacing with a device (not shown) for input or output that is connected to the computing device (200) or that the computing device (200) may include. In FIG. 2, the input / output interface (240) is illustrated as an element configured separately from the processor (220), but is not limited thereto, and the input / output interface (240) may be configured to be included in the processor (220).

[0095] The computing device (200) may include more components than those shown in FIG. 2. However, it is not necessary to explicitly illustrate most of the conventional components.

[0096] The processor (220) may include various types of modules, such as a mutation detection module, a filter module, annotation module, labeling module, and training module, as illustrated in FIGS. 3 to 9.

[0097] This specification first specifically describes the operation of the mutation detection module, annotation module, feature extraction module, and machine learning model, which can be utilized in the learning and inference processes of a machine learning model, using Figures 3 through 7. Next, the learning process of the machine learning model is described using Figure 8, and the inference process of the machine learning model is described using Figure 9.

[0098] FIG. 3 is a diagram illustrating an example of a variant detection module (310) according to one embodiment of the present disclosure. The variant detection module (310) can determine variant candidate information within a specific sample from sequencing data (300) associated with the specific sample. A "variant candidate" may be a sequence estimated to be a variant contained within the specific sample.

[0099] The mutation detection module (310) can be utilized in the learning process and inference process of a machine learning model. For example, the mutation detection module (310) can determine reference mutation candidate information within a reference sample from reference sequencing data associated with the reference sample during the learning process of the machine learning model. Similarly, the mutation detection module (310) can determine target mutation candidate information within a target sample from target sequencing data associated with the target sample during the inference process of the machine learning model.

[0100] The sequencing data (300) may include sequencing data associated with an abnormal sample collected from a specific individual (e.g., 112 in FIG. 1) and sequencing data associated with a normal sample collected from the same individual (e.g., 114 in FIG. 1). For example, the variant detection module (310) may determine variant candidate information (340) by comparing the sequencing data associated with the abnormal sample with the sequencing data associated with the normal sample, assuming that the normal sample does not contain abnormal cells.

[0101] Abnormal and normal samples may be Formalin-Fixed, Paraffin-Embedded (FFPE) processed samples or Fresh-Frozen (FF) processed samples. When samples are FFPE processed, various types of damage may occur in the DNA / RNA. Therefore, the number of mutation candidates determined using the sequencing data of the FFPE-processed sample may be greater than the number of mutation candidates determined using the sequencing data of the FFPE-processed sample. In other words, the mutation candidates determined using the sequencing data of the FFPE-processed sample may include mutation candidates determined using the sequencing data of the FFPE-processed sample and artifacts generated during the FFPE processing and storage process. Since artifacts generated during the FFPE processing and storage process correspond to noise when determining the mutation list in the abnormal sample, such noise may be removed later by a machine learning model or the like according to the present disclosure.

[0102] The mutation detection module (310) may include a plurality of detection modules (310_1 to 310_n) (n is any natural number). The plurality of detection modules (310_1 to 310_n) may each receive sequencing data (300) and determine mutation sub-candidate information (320_1 to 320_n). That is, the 'mutation sub-candidate' may refer to a sequence estimated to be a mutation, which is determined by each individual detection module. Since each of the plurality of detection modules (310_1 to 310_n) may have a different process for comparing sequencing data, the first to nth mutation sub-candidate information (320_1 to 320_n) may be different from each other. The first to nth mutation sub-candidate information (320_1 to 320_n) may include information associated with the detection module in which each mutation sub-candidate is determined.

[0103] Mutation candidate information (340) can be determined based on mutation sub-candidate information (320_1 to 320_n). For example, the mutation detection module (310) can determine mutation candidate information (340) by integrating (330) mutation sub-candidate information (320_1 to 320_n). For example, if the mutation sub-candidates determined from the first detection module (310_1) include mutation a and mutation b, the mutation sub-candidates determined from the second detection module (310_2) include mutation a and mutation c, and the mutation sub-candidates determined from the n-th detection module (310_n) include mutation b and mutation d, the mutation candidate information (340) may include information related to mutation a, mutation b, mutation c, and mutation d as a union of mutation sub-candidate information (320_1 to 320_n). Through this, all mutation candidate information that is likely to be an actual mutation can be collected.

[0104] In contrast, the mutation detection module (310) may determine mutation sub-candidate information commonly determined by any number (e.g., two) or more of the plurality of detection modules (310_1 to 310_n) as mutation candidate information (340). For example, according to the mutation sub-candidate example above, the mutation candidate information (340) may include information related to mutation a and mutation b. Through this, mutation candidate information with a high probability of being an actual mutation can be collected.

[0105] The mutation candidate information (340) may include location information of the mutation candidate, reference allele information at the location of the reference mutation candidate, altered allele information corresponding to the reference mutation candidate, reliability information of the mutation candidate, quality information (e.g., Phred quality score) of the mutation candidate, genotype information of the mutation candidate, read count information associated with the mutation candidate, and / or information of a detection module from among a plurality of detection modules (310_1 to 310_n) in which the mutation candidate is determined, etc. The location information of the mutation candidate may include chromosome information (e.g., chromosome number) on which the mutation candidate is located and / or location information (e.g., 1-based position) within the chromosome.

[0106] Additionally, the mutation candidate information (340) may include filter information. For example, each of the plurality of detection modules (310_1 to 310_n) may perform filtering to determine whether the mutation sub-candidate satisfies a plurality of predetermined quality indicators, and the filter information may include information associated with quality indicators that the mutation sub-candidate satisfies or does not satisfy in each of the plurality of detection modules (310_1 to 310_n). For example, the filter information may include information associated with whether the mutation sub-candidate passes filters such as whether it is a mutation candidate based on weak evidence, whether it is a mutation candidate that has experienced slippage, whether it occurred adjacent to other mutation candidates (clustered events), whether it is a haplotype, and whether it is a mutation candidate of the germline lineage. In response to the mutation sub-candidate passing all filters, the filter information of the mutation sub-candidate may be displayed as 'PASS'.

[0107] In FIG. 3, the mutation detection module (310) is illustrated as including multiple detection modules (310_1 to 310_n), but this is not limited thereto, and the mutation detection module (310) may also be configured as a single detection module. In this case, the mutation sub-candidate information generated from a single detection module may be identical to the mutation candidate information (340).

[0108] FIG. 4 is a diagram illustrating an example of an annotation module (420) according to one embodiment of the present disclosure. The annotation module (420) can generate annotation information (442, 444, 446, 448, 450) based on mutation candidate information (410) and / or information received from a database (430). The database (430) can include multiple databases containing different types of information (e.g., genetic databases such as Ensembl, RefSeq, etc.).

[0109] For example, the annotation module (420) can generate annotation information (442, 444, 446, 448, 450) based on reference mutation candidate information and / or information received from the database (430) in the learning process of the machine learning model. Similarly, the annotation module (420) can generate annotation information (442, 444, 446, 448, 450) based on target mutation candidate information and / or information received from the database (430) in the inference process of the machine learning model. At least a part of the generated annotation information (442, 444, 446, 448, 450) can be input into the machine learning model as features of the reference mutation candidate or the target mutation candidate, and can be used as data that serves as the basis for learning or inference of the machine learning model.

[0110] The mutation candidate information (410) may correspond to the mutation candidate information (340) of FIG. 3. The mutation candidate information (410) may be determined from FFPE-processed abnormal samples and normal samples, or from FFPE-processed abnormal samples and normal samples, using a mutation detection module.

[0111] The annotation module (420) may include a first annotation module (422) and a second annotation module (424). The first annotation module (422) and the second annotation module (424) are described separately by function for convenience of explanation, but this is to help understanding of the invention and does not necessarily mean that they are physically separated, and is not limited thereto. For example, the first annotation module (422) and the second annotation module (424) may be configured as a single module. In another example, the second annotation module (424) may include multiple modules that output second to fifth annotation information (444, 446, 448, 450), respectively.

[0112] The first annotation module (422) can extract known information related to the mutation candidate from the database (430) based on the mutation candidate information (410) and generate first annotation information (442) including the extracted information. For example, the first annotation information (442) can include information related to the effect of the mutation candidate on a protein (or amino acid sequence), genetic pattern, allelic variation, frequency of occurrence in a specific population, risk level, etc.

[0113] Additionally or alternatively, the first annotation information (442) may include information on the biological consequences of the mutation candidate. For example, the first annotation module (422) may extract sequence context information of the mutation candidate from the database (430) based on the mutation candidate information (410), and then use the extracted sequence context information to generate the first annotation information (442) including the biological consequences of the mutation candidate.

[0114] The database (430) referenced in the first annotation module (422) may include a genetic database such as Ensembl or RefSeq.

[0115] The second annotation module (424) can generate annotation information including information related to a sequence context associated with a mutation candidate and / or a state of the mutation candidate.

[0116] In one embodiment, the second annotation information (444) generated by the second annotation module (424) may include information associated with a plurality of reads whose mapping and alignment positions overlap with the positions of the mutation candidates, the sequencing results of the sample (e.g., a reference sample or a target sample including the mutation candidate). For example, the second annotation information (444) may include information such as position from 5'-end, position from 3'-end, etc., which indicate how far the position of the mutation candidate is from the start position of the sequencing read in which the specific mutation candidate is found.

[0117] In another example, the second annotation information (444) may include information associated with a variant read, which is a read that is different from a reference genome (e.g., includes a mutation) among a plurality of reads included in the sequencing data of a sample including a mutation candidate, and information associated with a non-variant read (hereinafter, also referred to as a “reference read”) having a position that overlaps with the position of the variant read.

[0118] Information associated with reference / variant reads may include information such as the number of reference / variant reads, minimum, median, maximum mapping quality associated with reference / variant reads, base quality statistics (e.g., median, mean), percentage of clipping bases, number of unmatched bases of mapped reads (e.g., minimum, median, maximum), insert size statistics (e.g., first quartile to third quartile), number of properly paired reads, number of chimeric reads, which are reads in which different parts of the reads are properly aligned to different reference genomes. "Insert size" may mean the distance on the reference genome between paired reads. "Insert size" can be understood as the sum of the length of Read 1, the length of Read 2, and the length of the unsequenced portion between Read 1 and Read 2. "Properly paired reads" can mean paired reads in which the paired Read 1 and Read 2 are well aligned in the forward and reverse directions, respectively, and the insert size does not deviate significantly from the expected value (e.g., the insert size is between the lower threshold and the upper threshold). The expected value of the insert size can mean the expected distance between Read 1 and Read 2 according to the average length of the fragments cut into a certain size when cutting DNA into fragments of a certain size (size selection) during the process of producing a DNA sequencing library.

[0119] Information associated with reference / variant leads can be stored and categorized under various headings. For example, the information associated with the reference / variant reads described above is ref_readN, ref_minMQ, ref_medMQ, ref_maxMQ, ref_medBQ, ref_meanBQ, ref_clip_pct, ref_mismatch_min, ref_mismatch_med, ref_mismatch_max, ref_f1_n, ref_f2_n, ref_r1_n, ref_r2_n, ref_isize_lq, ref_isize_uq, ref_isize_min, ref_isize_max, ref_ppair_n, ref_chim_n, var_readN, var_minMQ, var_medMQ, var_maxMQ, var_medBQ, var_meanBQ, var_clip_pct, var_mismatch_min, var_mismatch_med, It can be stored with item names such as var_mismatch_max, var_f1_n, var_f2_n, var_r1_n, var_r2_n, var_isize_lq, var_isize_uq, var_isize_min, var_isize_max, var_ppair_n, var_chim_n, pf5p_med, pf3p_med, etc., and each item name can be defined as in Tables 1 and 2. Table 1 shows an example of information associated with a reference read, and Table 2 shows an example of information associated with a variant read.

[0120] Item Name Description ref_readN Number of reference reads ref_minMQ Minimum mapping quality of reference reads ref_medMQ Median mapping quality of reference reads ref_maxMQ Maximum mapping quality of reference reads ref_minBQ Minimum base quality of reference reads ref_medBQ Median base quality of reference reads ref_meanBQ Average base quality of reference reads ref_maxBQ Maximum base quality of reference reads ref_clip_pct Percentage of clipped bases of reference reads ref_mismatch_min Minimum number of bases not matched to the reference genome in reference reads ref_mismatch_med Median number of bases not matched to the reference genome in reference reads ref_mismatch_max In reference reads The maximum number of bases that do not match the reference genome ref_f1_n The number of reference reads that are the first read among read pairs and are simultaneously aligned to the reference genome in the forward direction ref_f2_n The number of reference reads that are the second read among read pairs and are simultaneously aligned to the reference genome in the forward direction ref_r1_n The number of reference reads that are the first read among read pairs and are simultaneously aligned to the reference genome in the backward direction ref_r2_n The number of reference reads that are the second read among read pairs and are simultaneously aligned to the reference genome in the backward directionNumber of reference reads ref_isize_lq First quartile (25%) of insert size of reference reads ref_isize_uq Third quartile (75%) of insert size of reference reads ref_isize_min Minimum insert size of reference reads ref_isize_max Maximum insert size of reference reads ref_ppair_n Number of properly paired reference reads ref_chim_n Number of reference reads corresponding to chimeric reads,

[0121] Item Name Description var_readN Number of variant reads var_min Minimum mapping quality of variant reads var_med Median mapping quality of variant reads var_max Maximum mapping quality of variant reads var_minB Minimum base quality of variant reads var_med Median base quality of variant reads var_meanB Average base quality of variant reads var_max Maximum base quality of variant reads var_clip_pct Percentage of clipped bases of variant reads var_mismatch_min Minimum number of bases not matched to the reference genome in variant reads var_mismatch_med Median number of bases not matched to the reference genome in variant reads var_mismatch_max Maximum number of bases not matched to the reference genome in variant reads Maximum valuevar_f1_nThe number of variant reads that are the first read among read pairs and are simultaneously aligned to the reference genome in the forward directionvar_f2_nThe number of variant reads that are the second read among read pairs and are simultaneously aligned to the reference genome in the forward directionvar_r1_nThe number of variant reads that are the first read among read pairs and are simultaneously aligned to the reference genome in the backward directionvar_r2_nThe number of variant reads that are the second read among read pairs and are simultaneously aligned to the reference genome in the backward directionNumber of variant readsvar_isize_lq1st quartile (25%) of insert size of variant readsvar_isize_uq3rd quartile (75%) of insert size of variant readsvar_isize_minMinimum insert size of variant readsvar_isize_maxMaximum insert size of variant readsvar_ppair_nNumber of properly paired variant readsvar_chim_nNumber of variant reads corresponding to chimeric readspf5p_medMedian distance of variant position of variant read from 5'-endpf3p_medMedian distance of variant position of variant read from 3'-end

[0122] As described in Table 1 and Table 2, among the detailed names constituting each item name, “ref_” may indicate that it is information related to a reference lead, and “var_” may indicate that it is information related to a variant lead. "readN" represents the number of reads, "MQ" and "BQ" represent the mapping quality and base quality, respectively, "min", "med", "mean", and "max" represent the minimum, median, average, and maximum, respectively, "clip_pct" represents the percentage of clipped bases, "mismatch" represents the number of bases that do not match the reference genome, "fk (where k is a natural number)" represents that a specific read is the kth read in a read pair and is aligned to the reference genome in the forward direction, and similarly, "rl (where l is a natural number)" represents that a specific read is the lth read in a read pair and is aligned to the reference genome in the reverse direction, "isize" represents the insert size, and "lq" and "uq" represent the first quartile (25%) and the third quartile, respectively. "ppair" represents a properly paired read, "chim" represents a chimeric read, and "pf5p" and "pf3p" represent the distance from the 5'-end of the mutant read and the distance from the 3'-end of the mutant read, respectively. In addition, the types of information associated with the reference / mutant read can be defined by combining the detailed names described above. For example, "ref_mismatch_mean" can represent the average number of bases in the reference read that do not match the reference genome.Additionally, information associated with reference / variant leads that is additionally defined by combining the detailed names described above, without being limited to the types of information described above, may be included in the second annotation information (444).

[0123] In one embodiment, the third annotation information (446) generated by the second annotation module (424) may include at least a portion of normal tissue genome data (PON: Panel Of Normals) generated from a plurality of sequencing data associated with a plurality of normal samples. The normal tissue genome data may be whole genome sequencing data obtained from a database (430) and may include information associated with characteristics that are common across the plurality of normal samples, i.e., information reflecting characteristics of a group of normal samples.

[0124] For example, normal tissue genome data is the sum of all read depths for any position in multiple normal samples that are the basis for constructing the normal tissue genome data (referred to as 'PON_dpsum'), the number of samples with a non-zero read depth for the corresponding position among multiple normal samples (referred to as 'PON_dpN'), the number of samples with a read depth of 10 or more for the corresponding position among multiple normal samples (referred to as 'PON_dp10N'), the sum of all variant reads for the corresponding position among multiple normal samples (referred to as 'PON_varsum'), the number of samples with at least one variant read for the corresponding position among multiple normal samples (referred to as 'PON_varN'), the number of samples with a variant allele frequency (VAF) of less than 0.2 for the corresponding position among multiple normal samples (referred to as 'PON_var0.2lN'), and the number of samples with a VAF of 0.2 or more for the corresponding position among multiple normal samples. The number of samples (referred to as 'PON_var0.2hN') and / or the number of samples with two variant reads for that position among a plurality of normal samples (referred to as 'PON_var2N'). The reference values ​​for calculating 'PON_dp10N', 'PON_var0.2lN', 'PON_var0.2hN', and 'PON_var2N' are described as 10, 0.2, 0.2, and 2, respectively, but these can be set arbitrarily. For example, the normal tissue genome data can include the number of samples (referred to as 'PON_var0.25lN') for any position among a plurality of normal samples.

[0125] In one embodiment, the third annotation information (446) may include data associated with the location of a specific mutation candidate among normal tissue genome data (e.g., the sum of all read depths for the location of a specific mutation candidate among a plurality of normal samples, the number of samples in which the read depth for the location of a specific mutation candidate among a plurality of normal samples is not 0, etc.). Different third annotation information (446) corresponding to each mutation candidate at different locations may be generated. The third annotation information (446) corresponding to each mutation candidate may be input into a machine learning model as a feature of each mutation candidate, and may be used as data that serves as the basis for learning or inference of the machine learning model.

[0126] In one embodiment, the fourth annotation information (448) generated by the second annotation module (424) may include at least a portion of FFPE processed tissue genome data (Panel of FFPEs: POF) generated from a plurality of sequencing data associated with a plurality of Formalin-Fixed, Paraffin-Embedded (FFPE) processed samples. The plurality of FFPE processed samples may be normal samples.

[0127] Alternatively, the FFPE-processed multiple samples may be abnormal samples (e.g., tumor cell samples). If the FFPE-processed multiple samples are abnormal samples, information associated with mutations in the abnormal samples may be included in the genomic data. Therefore, by removing information associated with mutations using fresh-frozen (FF)-processed abnormal samples corresponding to the FFPE-processed abnormal samples, data substantially identical to that generated using normal samples for FFPE-processed tissue genomic data can be generated. In this case, the FFPE-processed abnormal samples may be directly converted from FF-processed abnormal samples.

[0128] The FFPE processed tissue genome data may be whole genome sequencing data obtained from a database (430) and may include information associated with characteristics that are common across multiple FFPE samples, i.e., information reflecting characteristics of a population of FFPE processed samples.

[0129] For example, FFPE processed tissue genome data may be obtained by summing all read depths for any position in multiple FFPE processed samples that are the basis of the FFPE processed tissue genome data (referred to as 'POF_dpsum'), the number of samples among multiple FFPE samples whose read depth for the corresponding position is not 0 (referred to as 'POF_dpN'), the number of samples among multiple FFPE samples whose read depth for the corresponding position is 10 or more (referred to as 'POF_dp10N'), the sum of all variant reads for the corresponding position in multiple FFPE samples (referred to as 'POF_varsum'), the number of samples among multiple FFPE samples that have at least one variant read for the corresponding position (referred to as 'POF_varN'), the number of samples among multiple FFPE samples whose VAF for the corresponding position is less than 0.2 (referred to as 'POF_var0.2lN'), and the VAF for the corresponding position in multiple FFPE samples. The number of samples with a VAF of 0.2 or greater (referred to as 'POF_var0.2hN') and / or the number of samples with two variant reads for the position among multiple FFPE samples (referred to as 'POF_var2N'), etc. may be included. The reference values ​​for calculating 'POF_dp10N', 'POF_var0.2lN', 'POF_var0.2hN', and 'POF_var2N' are described as 10, 0.2, 0.2, and 2, respectively, but these may be set arbitrarily. For example, the FFPE processed tissue genome data may include the number of samples with a VAF of less than 0.25 for any position among multiple FFPE processed samples (referred to as 'POF_var0.25lN').

[0130] In one embodiment, the fourth annotation information (448) may include data associated with the location of a specific mutation candidate among FFPE processed tissue genome data (e.g., the sum of all read depths for the location of a specific mutation candidate among a plurality of FFPE processed samples, the number of samples in which the read depth for the location of a specific mutation candidate is not 0 among a plurality of FFPE samples, etc.). Different fourth annotation information (448) may be generated corresponding to each mutation candidate at different locations. The fourth annotation information (448) corresponding to each mutation candidate may be input into a machine learning model as a feature of each mutation candidate and used as data that serves as the basis for learning or inference of the machine learning model. In one embodiment, the fifth annotation information (450) generated by the second annotation module (424) may include information associated with the mutation type of the mutation candidate and sequence context information of the mutation candidate.

[0131] For example, if the mutation candidate is a SNV, information associated with the mutation type may include pattern information of the SNV. For example, the pattern information may include information on the bases before and after the mutation, such as 'A->C', 'C->A', 'G->T', 'T->G', and 'G->U'.

[0132] In contrast, when the mutation candidate is an INDEL, information associated with the mutation type may include information such as whether the mutation candidate corresponds to a deletion mutation or an insertion mutation, the length of the deletion sequence if it is a deletion mutation, and the length of the insertion sequence if it is an insertion mutation.

[0133] If the mutation candidate is an SNV, the sequence context information may include information on the flanking base sequence of a certain length (e.g., 3 bp, 5 bp) that includes the mutation candidate. For example, if 'A (adenine)' at a specific position is the mutation candidate, the sequence context information may be expressed as 'CpApG' (where 'p' represents a phosphodiester bond), etc. (i.e., the bases adjacent to the mutation candidate are C (cytosine) and G (guanine)).

[0134] In contrast, when the mutation candidate is an INDEL, the sequence context information may include information about whether the sequence before or after the mutation candidate position is a repeated sequence or has a microhomology pattern. For example, if the repeated sequence is 'A' and the repeat length is 4, the sequence context information may be expressed as 'AAAA', and if the repeated sequence is 'ACG' and the repeat length is 3, the sequence context information may be expressed as 'ACGACGACG'. In an example with a 2bp ('AG') microhomology pattern, the sequence context information may be expressed as 'AG{deleted sequence}AG', where the before and after of the mutation candidate are the same.

[0135] The annotation information generated by the annotation module (420) is not limited to that shown and described in FIG. 4, and some information may be omitted or added.

[0136] FIG. 5 is a diagram illustrating an example of a feature extraction module (530) according to one embodiment of the present disclosure. The feature extraction module (530) may extract features (540) of a mutation candidate based on annotation information (520) (additionally, mutation candidate information (510)). The mutation candidate information (510) may correspond to the mutation candidate information (340) of FIG. 3 and may be determined from abnormal samples and normal samples processed through FFPE. Additionally, the mutation candidate information (510) may be information filtered by a filter module (e.g., 842, 846 of FIG. 8). The annotation information (520) may be generated by the annotation module (420) of FIG. 4.

[0137] The feature extraction module (530) can extract features of reference variant candidates based on reference variant candidate information and associated annotation information during the learning process of a machine learning model. Similarly, the feature extraction module (530) can extract features of target variant candidates based on target variant candidate information and associated annotation information during the inference process of a machine learning model.

[0138] The features (540) of the mutation candidate can be divided into features expressed as categorical variables and features expressed as numeric variables. In one embodiment, categorical variables can be expressed by mapping them to numeric variables through one-hot encoding, etc.

[0139] The feature extraction module (530) can generate features (540) of the mutation candidate based on all or part of the annotation information (520) corresponding to the mutation candidate information (510). The feature extraction module (530) can generate features (540) of the mutation candidate through a feature extraction process that uses all or part of the annotation information (520) as is or in a processed / modified form. Additionally, if the mutation candidate is a biological mutation known to occur repeatedly in the same location in the same form (e.g., a hotspot mutation), the mutation candidate is likely to correspond to an actual mutation, and therefore the features (540) of the mutation candidate may include information indicating that the mutation candidate corresponds to a true positive mutation.

[0140] The feature extraction module (530) can store feature extraction information (550) associated with the feature extraction process that generates features (540) of mutation candidates using annotation information (520) in a database (560). For example, the feature extraction information (550) can include information associated with the type of annotation information to be extracted, operations to be performed on the annotation information, etc. The feature extraction module (530) can perform the same feature extraction process for newly input mutation candidates and annotation information by using the feature extraction information (550) by referencing the database (560) when extracting features of new mutation candidates. Additionally, the feature extraction module (530) can store features (540) of the generated mutation candidates in the database (560).

[0141] The feature extraction module (530) may perform an additional process (e.g., refinement) related to the features (540) of the mutation candidates based on the information stored in the database (560). As an example, the feature extraction module (530) may perform a data standardization process, such as Z-score standardization, on the features (540) of the mutation candidates. For example, the feature extraction module (530) may use the mean (μ) and standard deviation (σ) of the features (540) of the mutation candidates to replace the value x associated with the features (540) of the mutation candidates as z = (x - μ) / σ. At this time, the mean and standard deviation values ​​for the features (540) of the mutation candidates may be stored in the database (560). Additionally, the feature extraction module (530) may perform a log transformation to normalize the features (540) of the mutation candidates before the data standardization process.

[0142] As another example, the feature extraction module (530) may perform a machine learning technique, such as domain adaptation (DA), on the features (540) of the mutation candidates. For example, a difference greater than a threshold may occur in the distribution of arbitrary feature values ​​between a data domain (Source Domain) containing data used in the learning process of a machine learning model and a data domain (Target Domain) containing new data. In this case, the feature extraction module (530) may adjust the distribution of feature values ​​of the Source Domain and / or the Target Domain in a direction that reduces this difference. Through this, the problem of performance degradation of the machine learning model due to the difference in the feature value distribution can be alleviated.

[0143] Data standardization processes such as the Z-score standardization described above or machine learning techniques such as domain adaptation (DA) may not be applied to some features (e.g., features mapped to numeric variables through one-hot encoding, etc.) among the features (540) of the mutation candidates.

[0144] FIG. 6 is a diagram illustrating an example of outputting a classification result (640) of a mutation candidate using a machine learning model (630) according to one embodiment of the present disclosure. The machine learning model (630) receives mutation candidate information (610) and features (620) of the mutation candidate within a sample, and can output a classification result (640) indicating whether the mutation candidate is a true positive mutation (or whether the mutation candidate is an artifact resulting from FFPE processing).

[0145] For example, during the learning process, the machine learning model (630) may receive reference mutation candidate information and the characteristics of the reference mutation candidate, and output a classification result indicating whether the reference mutation candidate is a true positive mutation. Based on the classification result indicating whether the reference mutation candidate is a true positive mutation and the classification information labeled on the reference mutation candidate, the parameters or hyperparameters of the machine learning model (630) may be adjusted. This will be described in detail later with reference to FIG. 8 .

[0146] Similarly, the machine learning model (630) learned by the learning process described above can receive target mutation candidate information and features of the target mutation candidate in the inference process, and output a classification result indicating whether the target mutation candidate is a true positive mutation. The binary classification output result is generated through a series of decision processes starting from the raw output or score internally in the classifier, and can be considered to include cases where the output result is the raw output or score. Therefore, the following description regarding the raw output or score may also apply to the case of binary classification.

[0147] FIG. 7 is a diagram illustrating a detailed configuration of a machine learning model (700) according to one embodiment of the present disclosure. The machine learning model (700) may correspond to the machine learning model (600) of FIG. 6. The machine learning model (700) may include a plurality of classifiers (710_1 to 710_n) and a meta-classifier (730) connected thereto.

[0148] Each of the plurality of classifiers (710_1 to 710_n) may be input with mutation candidate information (610 in FIG. 6) and features of the mutation candidate (620 in FIG. 6). The plurality of classifiers (710_1 to 710_n) may output a plurality of output results (720_1 to 720_n) indicating whether the mutation candidate is a true positive mutation (or whether the mutation candidate is an artifact due to FFPE processing) based on the mutation candidate information and the features of the mutation candidate. The plurality of output results (720_1 to 720_n) may be classified in a binary class manner such as TRUE / FALSE, or may be raw outputs having values ​​between 0 and 1.

[0149] Thereafter, the meta classifier (730) can determine a classification result (740) (corresponding to 640 in FIG. 6) indicating whether the mutation candidate is a true positive mutation by using the output result from at least one of the plurality of classifiers (710_1 to 710_n).

[0150] The meta classifier (730) may be a classifier trained based on training data that includes at least some of the plurality of output results (720_1 to 720_n) as features. The meta classifier (730) may be trained based on classification results (740) produced using the training data and labeled classification information for variant candidates.

[0151] The classification result (740) may indicate whether the variant candidate is a true positive variant. The classification result (740) may be determined to be a true positive variant in response to a determination that a raw output or score (associated with the probability that the variant candidate is a true positive variant) determined from the meta classifier (730) is greater than a decision threshold for determining the classification result (740). In another embodiment, the variant candidate may be determined to be a true positive variant in response to a determination that the probability that the variant candidate is a true positive variant is higher than a threshold probability or that the probability that the variant candidate is an artifact is lower than a threshold probability. The probability that the variant candidate is a true positive variant or the probability that the variant candidate is an artifact may be determined from the raw output or score determined from the meta classifier (730). The classification result (740) may be displayed as 'TRUE', meaning that the mutation candidate is a true positive mutation, or as 'FALSE', meaning that the mutation candidate is not an artifact.

[0152] The decision threshold for determining the classification result (740) can be determined by comparing the raw score determined by the meta-classifier (730) with the labeled classification information. For example, the decision threshold can be determined as a value between 0 and 1 at which the F1-score or the AUC (Area Under the Curve) of the ROC (Receiver Operating Characteristic) curve is maximized. In this case, if there are multiple values ​​at which the F1-score or the AUC of the ROC curve is maximized, the decision threshold can be determined as the median value, average value, etc.

[0153] Each of the plurality of classifiers (710_1 to 710_n) and the meta classifier (730) may be implemented as a neural network such as a regression analysis model (e.g., Elastic-Net Logistic Regression), a support vector machine (SVM), a random forest, gradient boosting (e.g., XGBoost, LightGBM, etc.), or a multilayer perceptron (MLP). Without being limited thereto, each of the plurality of classifiers (710_1 to 710_n) and the meta classifier (730) may be a general machine learning-based classifier including a deep neural network (DNN). In addition, the meta classifier (730) may include a rule-based algorithm as well as machine learning. Additionally or alternatively, the meta classifier (730) can output a classification result (740) by averaging or voting multiple output results (720_1 to 720_n).

[0154] The training data for performing the learning process of the machine learning model (700) may include reference mutation candidate information, characteristics of the mutation candidates, and classification information labeled for the reference mutation candidates. The training data may be divided into a training set and a validation set. Additionally, the training data may be further divided to create a test set separate from the validation set.

[0155] For example, training data can be divided and used through k-fold cross validation, which divides the data into k parts, uses k-1 parts as a training set, and 1 part as a validation set, and repeats this process k times to obtain k performance indicators. The above process can be repeated n times (n-repeated k-fold cross validation), and the data can be randomly shuffled before each data split. The entire training data can be referred to as an "epoch," and each of the k data sets into which the training data is divided can be referred to as a "batch."

[0156] Each of the plurality of divided batches can be trained with a different type of classifier among the plurality of classifiers (710_1 to 710_n). The types and number of the plurality of classifiers (710_1 to 710_n) can be variable, and the number of the plurality of classifiers (710_1 to 710_n) can be determined based on the number of the plurality of divided batches. For example, when there are N types of available classifiers and k-fold cross validation is repeated n times for the training data (n-repeated k-fold cross validation), the number (N*n*k) of the plurality of classifiers (710_1 to 710_n) can be determined based on the values ​​of N, k, and n. For example, if there are 5 types of classifiers and 10-fold cross-validation is repeated 10 times, the number of multiple classifiers (710_1 to 710_n) can be 5*10*10=500.

[0157] Alternatively, the machine learning model (700) may include a classifier and a meta-classifier (730), in which case the meta-classifier (730) may be an indicator function.

[0158] While the learning process of the machine learning model (700) is in progress using the k-fold cross-validation method, the hyperparameters of multiple classifiers (710_1 to 710_n) may be optimized, and model selection may be performed based on whether the classification result (740) matches the actual classification. For example, when generating the classification result (740), it may be determined whether to use the output results of each of the multiple classifiers (710_1 to 710_n), and if so, with what weights.

[0159] In one embodiment, when the training data includes multiple data sets for multiple samples, considering that the actual number of variants and their distributions are different in each of the multiple samples, the training data may be partitioned using a weighted sampling method or a stratified sampling method. Additionally, to prevent class imbalance problems from occurring for each sample due to different numbers of actual variants included in each sample, some data may be oversampled or undersampled, or weight balancing may be performed when calculating loss. Alternatively, synthetic sampling may be performed to modify the characteristics of candidate variants or generate new data when oversampling.

[0160] Alternatively, before the training data is split into multiple batches, multiple data sets for multiple samples can be concatenated into a single data set and then split. In this case, the training data can be split into multiple batches based on multiple mutation candidates within the data set (e.g., rows in Figure 10).

[0161] In the inference process of the machine learning model, each of the plurality of classifiers (710_1 to 710_n) can receive target mutation candidate information and features of the target mutation candidate in the target sample and output a result indicating whether the target mutation candidate is a true positive mutation (or an artifact due to FFPE processing).

[0162] For each of the multiple classifiers (710_1 to 710_n), a raw output or score may be produced, for example, as a result indicating whether the target mutation candidate is a true positive mutation. The produced raw output or score may be mapped or converted into a probability value indicating that the mutation candidate is a true positive mutation (or an artifact).

[0163] In one embodiment, each of the plurality of classifiers (710_1 to 710_n) may perform binary classification on the produced raw output, score, or probability value using a cut-off value determined by cross validation of the learning process, and output the result of the binary classification as an output result (720_1 to 720_n). The meta classifier (730) may output a classification result (740) by voting on the binary classification result output as the output result (720_1 to 720_n).

[0164] In another embodiment, the raw output, score, or probability value representing the result indicating whether the target variant candidate is a true positive variant from each of the plurality of classifiers (710_1 to 710_n) may be averaged using a meta classifier (730), an optimal cutoff value may be calculated using a portion of the learning data, and a classification result (740) may be output based on the averaged value and the calculated cutoff value.

[0165] In another embodiment, a classification result (740) may be produced using a learned meta-learner model by receiving raw outputs, scores, or probability values ​​representing results indicating whether a target mutation candidate is a true positive mutation from each of a plurality of classifiers (710_1 to 710_n) as input features.

[0166] FIG. 8 is a diagram illustrating an example of a learning process of a machine learning model (870) according to one embodiment of the present disclosure. FFPE sample sequencing data (812) (e.g., sequencing data associated with FFPE-treated normal samples and sequencing data associated with FFPE-treated abnormal samples) serves as reference sample sequencing data, and learning data of the machine learning model (870) can be generated using the FFPE sample sequencing data (812).

[0167] The variant detection module (820) (corresponding to 310 in FIG. 3) can determine candidate variant information (822) within an FFPE sample from FFPE sample sequencing data (812). Similarly, the variant detection module (820) can determine candidate variant information (824) within an FFPE sample from FF sample sequencing data (814) (e.g., sequencing data associated with FF-processed normal samples and sequencing data associated with FF-processed abnormal samples).

[0168] FFPE and FF samples may be corresponding samples. For example, FFPE and FF samples may be collected from the same individual. Additionally, FFPE samples may be directly converted from FF samples or collected with some time difference from FF samples.

[0169] An annotation module (830) (corresponding to 420 in FIG. 4) can receive FFPE sample mutation candidate information (822) as input and output FFPE sample annotation information (832). The FFPE sample annotation information (832) can be used to filter mutation candidate information (822) within the FFPE sample by being transmitted to the first filter module (842).

[0170] Similarly, the annotation module (830) can receive FF sample mutation candidate information (824) as input and output FF sample annotation information (834). The FF sample annotation information (834) can be used to filter mutation candidate information (824) within the FF sample by being transmitted to the second filter module (846).

[0171] The first filter module (842) can receive mutation candidate information (822) within an FFPE sample as input and output filtered mutation candidate information (844) within the FFPE sample. Similarly, the second filter module (846) can receive mutation candidate information (824) within an FF sample and output filtered mutation candidate information (848) within the FF sample. That is, the first filter module (842) and the second filter module (846) can filter some of the mutation candidates within the sample to remove noise (e.g., artifacts generated during FFPE processing), thereby improving the learning accuracy of the machine learning model (870).

[0172] In one embodiment, the first filter module (842) and the second filter module (846) may perform filtering based on filter information generated by the variant detection module (820) among the variant candidate information (822, 824). The filter information may include information associated with quality indicators that the variant candidate satisfies or fails to satisfy. For example, the filter information may include information associated with whether the variant candidate passes the filter, such as whether the variant candidate is a weak evidence-based variant candidate, whether the variant candidate is a slippage-induced variant candidate, whether the variant candidate occurs in proximity to other variant candidates (clustered events), whether the variant candidate is a haplotype, and whether the variant candidate is a germline variant candidate. In response to the variant candidate passing all filters, the filter information may be displayed as 'PASS'.

[0173] For example, the first filter module (842) and the second filter module (846) may filter out mutation candidates that do not meet specific quality indicators by referring to the filter information described above. For example, the first filter module (842) and the second filter module (846) may filter out mutation candidates whose filter information is not marked as 'PASS'.

[0174] In one embodiment, the first filter module (842) and the second filter module (846) can filter out at least some of the variant candidates that are artifacts based on the Panel of Normals (PON) and / or Panel of FFPEs (POF) in the FFPE sample annotation information (832) and the FFPE sample annotation information (834) (e.g., generated by the second annotation module (424) of FIG. 4).

[0175] In one embodiment, the first filter module (842) and the second filter module (846) (alternatively, the mutation detection module (820)) may determine, as filtered mutation candidate information (844, 848), mutation sub-candidate information commonly determined by a predetermined number (e.g., two) or more of the detection modules (corresponding to 310_1 to 310_n of FIG. 3) included in the mutation detection module (820) among the mutation candidate information (822, 824) generated by unioning mutation sub-candidate information determined by the plurality of detection modules (corresponding to 310_1 to 310_n of FIG. 3) included in the mutation detection module (820). Additionally, the filtering performed by the first filter module (842) and the second filter module (846) may be performed based on the types and / or numbers of the plurality of detection modules.

[0176] Additionally, the second filter module (846) may perform filtering based on filtering conditions associated with the FF sample annotation information (834). At this time, the filtering conditions associated with the FF sample annotation information (834) may be set by being corrected and optimized based on various environmental contextual variables such as a sequencing platform, a library preparation method, sequencing depth, the status of the tissue sample, and sample purity. The filtering associated with the FF sample annotation information (834) may be performed using a rule-based algorithm or a machine learning-based model.

[0177] The feature extraction module (850) (corresponding to 530 in FIG. 5) can extract features (852) of a mutation candidate based on the FFPE sample annotation information (832) (additionally, filtered mutation candidate information (844) within the FFPE sample). Alternatively, when the first filter module (842) is omitted, the feature extraction module (850) can extract features (852) of a mutation candidate based on the mutation candidate information (822) within the FFPE sample and the FFPE sample annotation information (832). The extracted features (852) of the mutation candidate, together with the filtered mutation candidate information (844) within the FFPE sample, can be used as part of the training data of the machine learning model (870).

[0178] The labeling module (860) can label classification information (862) for a reference mutation candidate, which is part of the training data of a machine learning model (870). The classification information (862) can indicate whether the reference mutation candidate is a true positive mutation or a false positive mutation.

[0179] In one embodiment, in response to determining that at least a portion of information about a particular variant candidate in an FFPE sample corresponds to at least a portion of information associated with one of the variant candidates in an FFPE-processed sample, the labeling module (860) may label the particular variant candidate as a true positive variant. In other words, the particular variant candidate may be determined to be an actual variant, as not having occurred due to a difference in sample processing methods (FFPE, FF).

[0180] For example, among the specific mutation candidate information in the FFPE sample, the location information of the mutation candidate (e.g., chromosome information and location information within the chromosome where the mutation candidate is located), reference allele information at the location of the mutation candidate, and altered allele information corresponding to the mutation candidate may correspond to or be identical to information associated with any mutation candidate in the FFPE-processed sample. In this case, the labeling module (860) may label the specific mutation candidate as a true-positive mutation.

[0181] If a particular variant candidate is a true positive variant, that particular variant candidate can be labeled as 'TRUE' (i.e., indicating that it is a true positive variant) or 'FALSE' (i.e., indicating that it is not an artifact found in FFPE samples).

[0182] On the other hand, in response to determining that the specific mutation candidate information in the FFPE sample and the information associated with any one of the mutation candidates in the FFPE processed sample do not correspond to each other, the labeling module (860) may label the specific mutation candidate as a false positive mutation.

[0183] For example, any one of the location information of a variant candidate in an FFPE sample, the reference allele information at the location of the variant candidate, and the altered allele information corresponding to the variant candidate may not correspond to any information associated with any arbitrary variant candidate in the FFPE-processed sample. In this case, the labeling module (860) may label the specific variant candidate as a false positive variant. Alternatively, the specific variant candidate may remain unlabeled.

[0184] In one embodiment, the classification information (862) generated by the labeling module (860) may not be modified or altered after generation. For example, if all variant candidates within an FFPE sample are labeled with a confidence level above a threshold, or if the entire learning process is performed once, the generated classification information (862) may not be modified or altered after generation.

[0185] In another embodiment, the classification information (862) generated by the labeling module (860) may be modified / changed after generation. For example, in response to determining that at least some of the classification information (862) is mislabeled or that the probability of mislabeling is greater than a threshold, the classification information (862) may be modified / changed after generation.

[0186] In one example, if the training process of the machine learning model (870) is repeated multiple times, the classification information (862) can be modified / changed after generation. In another example, if a teacher model corresponding to or similar to the machine learning model (870) is built and trained using a subset of the training data including data labeled with a confidence level higher than a threshold, and a process of labeling variant candidates labeled with a confidence level lower than a threshold or unlabeled variant candidates using the trained teacher model is performed (e.g., noisy label training), the classification information (862) can be modified / changed even after initial generation.

[0187] The machine learning model (870) can receive filtered mutation candidate information (844) and features (852) of the mutation candidate within the FFPE sample as input, and determine and output a classification result (872) indicating whether the mutation candidate is a true positive mutation. The training module (880) can train the machine learning model (870) by performing parameter adjustment (882) of the machine learning model (870) based on the classification result (872) of the mutation candidate and the classification information (862) labeled on the mutation candidate.

[0188] Some of the configurations illustrated in FIG. 8 may be omitted. For example, if the first filter module (842) is omitted, the variant candidate information (822) within the FFPE sample may be used in place of the filtered variant candidate information (844) within the FFPE sample described above, and if the second filter module (846) is omitted, the variant candidate information (824) within the FF sample may be used in place of the filtered variant candidate information (848) within the FF sample described above.

[0189] FIG. 9 is a diagram illustrating an example of an inference process of a machine learning model (960) according to one embodiment of the present disclosure. The mutation detection module (920), the annotation module (930), the filter module (940), and the feature extraction module (950) of FIG. 9 correspond to the mutation detection module (820), the annotation module (830), the first filter module (842), and the feature extraction module (850) of FIG. 8, respectively, and the description thereof using FIG. 8 may be omitted.

[0190] The FFPE sample sequencing data (910) may correspond to newly provided sequencing data in the inference process of the learned machine learning model (960), as target sample sequencing data. Using the learned machine learning model (960), a classification result (962) indicating whether a target mutation candidate of an FFPE sample is a true positive mutation (or an artifact due to FFPE processing, etc.) may be output from the FFPE sample sequencing data (910). That is, the learned machine learning model (960) can infer whether a mutation candidate existing in the sample is an actual mutation or an artifact caused by external factors such as the FFPE processing process.

[0191] Specifically, the mutation detection module (920) outputs mutation candidate information (922) within the FFPE sample using the FFPE sample sequencing data (910), and the mutation candidate information (922) within the FFPE sample can be transmitted to the annotation module (930) and the filter module (940).

[0192] The annotation module (930) can generate FFPE sample annotation information (932) based on the mutation candidate information (922) within the FFPE sample. The filter module (940) can receive the mutation candidate information (922) within the FFPE sample and the FFPE sample annotation information (932) as input and filter some of the plurality of mutation candidates, thereby generating filtered mutation candidate information (942) within the FFPE sample.

[0193] The feature extraction module (950) can generate features (952) of variant candidates using FFPE sample annotation information (932) (additionally, filtered variant candidate information (942) within the FFPE sample).

[0194] The machine learning model (960) can receive filtered mutation candidate information (942) and mutation candidate features (952) within the FFPE sample as input, and output a mutation candidate classification result (962). The feature extraction module (950) can extract the mutation candidate features (952) through the same feature extraction process using the feature extraction information used in the learning process of the machine learning model (960).

[0195] A specific mutation candidate may be a biological mutation known to occur repeatedly in the same location and form (e.g., a hotspot mutation). In this case, since the specific mutation candidate is likely to correspond to a true mutation, it can be input into a machine learning model (960) with a label indicating that it corresponds to a true positive mutation.

[0196] Thereafter, based on the classification results (962), some of the multiple mutation candidates may be filtered / removed from the mutation candidate list, thereby determining a list of mutations inferred as actual mutations in the FFPE sample.

[0197] Each module within the system illustrated in FIGS. 3 to 9 is merely an example, and in some embodiments, other modules may be additionally included in addition to the illustrated modules, and some components may be omitted. For example, if some of the above internal components are omitted, the functions of the omitted internal components may be configured to be performed by other modules or processors of other computing devices. In addition, although each module is described by function in FIGS. 8 and 9, this is intended to aid understanding of the invention and does not necessarily imply that they are physically separated, nor is it limited thereto.

[0198] FIG. 10 is a diagram illustrating an example of learning data (1000) according to one embodiment of the present disclosure. As illustrated in FIG. 10, the learning data (1000) may be configured as a data matrix in a table format. In FIG. 10, the rows of the data matrix represent unique numbers (1 to 13633) assigned to each reference mutation candidate, and the columns represent reference mutation candidate information, characteristics of the reference mutation candidates, and classification information, but the present invention is not limited to this format. For example, the learning data may be implemented as a transpose matrix of the above-described data matrix, or may be implemented in various formats such as a multidimensional vector, an array, or a data frame.

[0199] The learning data (1000) may include reference mutation candidate information, features of the mutation candidate, and classification information labeled on the reference mutation candidate.

[0200] For example, the CHROM, POS, REF, and ALT items can represent, in order, among the reference mutation candidate information, the chromosome information where the mutation candidate is located, the location information within the chromosome, the reference allele information at the location of the mutation candidate, and the altered allele information corresponding to the mutation candidate.

[0201] Multiple items starting with 'FILTER' may represent filter information generated by a mutation detection module (e.g., 310 in FIG. 3) among reference mutation candidate information. For example, all mutation candidates shown in the learning data (1000) in FIG. 10 can be considered to have passed all filters of the mutation detection module (FILTER_PASS: 1).

[0202] The 'label' entry can indicate the classification information labeled on the reference variant candidate.

[0203] Items other than those described above may represent the characteristics of mutation candidates extracted from annotation information. The types of information indicated by each item representing the characteristics of mutation candidates are described in detail in Figure 4.

[0204] The training data (1000) may be a concatenation of multiple data sets for multiple samples into a single data set. In this case, the training data (1000) may be divided into multiple batches based on the rows of the data matrix, and then input into each of the multiple classifiers of the machine learning model.

[0205] FIG. 11 is a diagram showing the performance evaluation results of a learned machine learning model according to one embodiment of the present disclosure. Each of the plurality of dots indicated in the graphs (1112, 1114, 1122, 1124, 1132, 1134) represents, as a single dot, the value of the variance (x-axis) measured using FF samples and the value of the variance (y-axis) measured using FFPE samples corresponding to the FF samples before and after filtering of mutation candidates using the machine learning model. That is, the closer each of the plurality of dots is to the y=x line indicated in each graph, the more similar the variance value measured using FF samples and the variance value measured using the corresponding FFPE samples can be determined to be. Each of the plurality of dots can represent the variance measured using FFPE samples and FF samples collected from each of a plurality of different individuals.

[0206] Specifically, each of the plurality of dots in the first graph (1112) and the second graph (1114) represents the number of SNVs measured using FF samples and the number of SNVs measured using FFPE samples corresponding to the FF samples before and after filtering of mutation candidates using the machine learning model. Similarly, each of the plurality of dots in the third graph (1122) and the fourth graph (1124) represents the number of INDELs measured using FF samples and the number of INDELs measured using FFPE samples corresponding to the FF samples before and after filtering of mutation candidates using the machine learning model.

[0207] Referring to the first graph (1112), it can be confirmed that the number of SNVs measured using the corresponding FFPE sample is more likely to be greater than the number of SNVs measured using the FF sample before filtering out mutation candidates using the machine learning model. Similarly, referring to the third graph (1122), it can be confirmed that the number of INDELs measured using the corresponding FFPE sample is more likely to be greater than the number of INDELs measured using the FF sample before filtering out mutation candidates using the machine learning model. This may result from various types of damage, such as cross-linking, fragmentation, and other non-biological base mutations that occur within the sample due to the FFPE processing process.

[0208] On the other hand, referring to the second graph (1114), it can be confirmed that the difference between the number of SNVs measured using FF samples and the number of SNVs measured using the corresponding FFPE samples after filtering out mutation candidates using the machine learning model has decreased (i.e., the data tends to align in the direction of the y=x line). Similarly, referring to the fourth graph (1124), it can be confirmed that the difference between the number of INDELs measured using FF samples and the number of INDELs measured using the corresponding FFPE samples has decreased after filtering out mutation candidates using the machine learning model. That is, it can be confirmed through the second graph (1114) and the fourth graph (1124) that noise (or artifacts) generated due to the FFPE processing process have been removed by filtering out false positive mutations in cell samples using the machine learning model.

[0209] Each of the plurality of dots in the fifth graph (1132) and the sixth graph (1134) represents the HRD (Homologous Recombination Deficiency) score measured using the FF sample and the HRD score measured using the FFPE sample corresponding to the FF sample before and after filtering of the mutation candidate using the machine learning model.

[0210] Referring to the sixth graph (1134), it can be seen that after filtering out mutation candidates using the machine learning model, the difference between the HRD score measured using the FF sample and the HRD score measured using the corresponding FFPE sample is reduced.

[0211] Table 3 below shows performance evaluation metrics associated with single nucleotide variants (SNVs), base insertion variants, and deletion variants (INDELs) according to the performance of variant candidate filtering using the machine learning model of the present disclosure.

[0212] Sensitivity Specificity PPVF1 Single nucleotide variant (SNV) 0.97 0.87 0.91 0.94 Insertion and deletion variant (INDEL) 0.91 0.91 0.92 0.91

[0213] In Table 3, sensitivity is defined as TP / (TP + FN), specificity is defined as TN / (TN + FP), positive predictive value (PPV) is defined as TP / (TP + FP), and F1-score is defined as 2 * sensitivity * positive predictive value / (sensitivity + positive predictive value). At this time, TP (True Positive), TN (True Negative), FP (False Positive), and FN (False Negative) may refer to the number correctly predicted as an actual mutation, the number correctly predicted as not an actual mutation, the number incorrectly predicted as an actual mutation, and the number incorrectly predicted as not an actual mutation, as a result of comparing the classification result using the machine learning model according to the present disclosure with the actual classification information (Ground Truth) for mutation candidates determined using FFPE samples, respectively.

[0214] Table 4 below shows the concordance index calculated before performing mutation candidate filtering using the machine learning model according to the present disclosure, and Table 5 shows the concordance index calculated after performing mutation candidate filtering using the machine learning model according to the present disclosure.

[0215] Agreement 95% Confidence Interval HRD 0.60 [0.47, 0.70] TMBS NV: 0.01 INDEL: 0.00 SNV: [0.00, 0.02] INDEL: [0.00, 0.00]

[0216] Agreement 95% Confidence Interval HRD 0.99 [0.98, 0.99] TMBS NV: 0.96 INDEL: 0.87 SNV: [0.93, 0.98] INDEL: [0.78, 0.92]

[0217] The concordance in Tables 4 and 5 is an indicator of how similar the values ​​and value tendencies of two variables are, and can be defined as the Lin's concordance correlation coefficient value. When the concordance is defined as the Lin's concordance correlation coefficient value, the closer the multiple points in each graph (1112, 1114, 1122, 1124, 1132, 1134) are to the y=x line, the closer the concordance value is to 1, and the closer the multiple points are to the y=x line, the closer the value is to 0. In addition, even if the distribution tendency of multiple points is linear, the concordance value may approach 0 as the distribution of multiple points is farther from the y=x line.

[0218] Referring to Tables 4 and 5, it can be seen that the HRD concordance improved significantly from 0.60 to 0.99 before and after performing mutation candidate filtering using the machine learning model, and the TMB (Tumor Mutation Burden) also improved significantly from 0.01 to 0.96 for SNV and from 0.00 to 0.87 for INDEL.

[0219] FIG. 12 is an exemplary diagram illustrating an artificial neural network model (1200) according to one embodiment of the present disclosure. The artificial neural network model (1200) is an example of a machine learning model, and is a statistical learning algorithm implemented based on machine learning technology and the structure of a biological neural network, or a structure that executes the algorithm.

[0220] According to one embodiment, the above-described machine learning model (or classifier within the machine learning model) may be generated in the form of an artificial neural network model (1200). For example, the artificial neural network model (1200) may receive mutation candidate information and features of the mutation candidate, and output a classification result indicating whether the mutation candidate is a true positive mutation (or whether the mutation candidate is an artifact resulting from FFPE processing).

[0221] According to one embodiment, the artificial neural network model (1200) may represent a machine learning model having problem-solving capabilities by learning that nodes, which are artificial neurons that form a network by combining synapses like in a biological neural network, repeatedly adjust the weights of synapses so that the error between the correct output corresponding to a specific input and the inferred output is reduced. For example, the artificial neural network model (1200) may include any probability model, neural network model, etc. used in artificial intelligence learning methods such as machine learning and deep learning.

[0222] In one embodiment, the artificial neural network model (1200) can be implemented as a multilayer perceptron (MLP) composed of multiple layers of nodes and connections therebetween. The artificial neural network model (1200) according to the present embodiment can be implemented using one of various artificial neural network model structures including MLP, but is not limited thereto. As illustrated in FIG. 4, the artificial neural network model (1200) is composed of an input layer (1220) that receives an input signal or data (1210) from the outside, an output layer (1240) that outputs an output signal or data (1250) corresponding to the input data, and n hidden layers (1230_1 to 1230_n) located between the input layer (1220) and the output layer (1240) that receive signals from the input layer (1220), extract features, and transmit them to the output layer (1240) (where n is a positive integer). Here, the output layer (1240) receives signals from the hidden layers (1230_1 to 1230_n) and outputs them to the outside.

[0223] The learning method of the artificial neural network model (1200) includes a supervised learning method that learns to optimize problem solving through input of a teacher signal (correct answer), and an unsupervised learning method that does not require a teacher signal. According to one embodiment, the artificial neural network model (1200) can be learned based on a learning data set including reference mutation candidate information, characteristics of the mutation candidate, and classification information labeled on the reference mutation candidate, and can be learned by a supervised learning method using the classification information among the learning data set as a teacher signal.

[0224] According to one embodiment, the input variables of the artificial neural network model (1200) may include mutation candidate information and characteristics of the mutation candidate. When the input variables described above are input through the input layer (1220), the output variables output from the output layer (1240) of the artificial neural network model (1200) may be a classification result indicating whether the mutation candidate is a true positive mutation (or whether the mutation candidate is an artifact resulting from FFPE processing).

[0225] In this way, a plurality of input variables and a plurality of corresponding output variables are respectively matched in the input layer (1220) and the output layer (1240) of the artificial neural network model (1200), and the synapse values ​​between the nodes included in the input layer (1220), the hidden layers (1230_1 to 1230_n), and the output layer (1240) are adjusted, so that learning can be performed so that the correct output corresponding to a specific input can be extracted. Through this learning process, the characteristics hidden in the input variables of the artificial neural network model (1200) can be identified, and the synapse values ​​(or weights) between the nodes of the artificial neural network model (1200) can be adjusted so that the error between the output variables calculated based on the input variables and the target output (e.g., labeled classification information) is reduced. In addition, the artificial neural network model (1200) learns an algorithm that receives mutation candidate information and the characteristics of the mutation candidate as input, and can be learned in a manner that minimizes the loss with the classification information.

[0226] Using the artificial neural network model (1200) learned in this way, a classification result indicating whether the mutation candidate is a true positive mutation (or whether the mutation candidate is an artifact resulting from FFPE processing) can be extracted.

[0227] FIG. 13 is a flowchart illustrating a method (1300) for training a machine learning model according to one embodiment of the present disclosure. The method (1300) may be performed by at least one processor. The method (1300) may be initiated by the processor determining reference mutation candidate information within a reference sample (S1310). The reference mutation candidate may be a sequence having a probability of corresponding to a point mutation encompassing SNV and INDEL greater than a predetermined threshold probability. The reference mutation candidate information may include at least one of position information of the reference mutation candidate, reference allele information at the position of the reference mutation candidate, or altered allele information corresponding to the reference mutation candidate.

[0228] In one embodiment, the reference sample includes a reference normal sample and a reference abnormal sample collected from the same individual, the reference sequencing data includes first reference sequencing data associated with the reference normal sample and second reference sequencing data associated with the reference abnormal sample, and the processor can determine reference variant candidate information based on the first reference sequencing data and the second reference sequencing data using a variant detection module.

[0229] In one embodiment, the variant detection module includes a plurality of detection modules, and the processor determines reference variant sub-candidate information by applying first reference sequencing data and second reference sequencing data to each of the plurality of detection modules, and determines reference variant candidate information by unioning the reference variant sub-candidate information. Alternatively, the processor determines reference variant sub-candidate information by applying first reference sequencing data and second reference sequencing data to each of the plurality of detection modules, and determines reference variant sub-candidate information commonly determined by two or more of the plurality of detection modules as the reference variant candidate information.

[0230] The processor can generate annotation information for learning a machine learning model (S1330).

[0231] In one embodiment, the processor can determine a plurality of reads, at least some of which mapped positions overlap with positions of reference variant candidates, and generate first annotation information associated with the determined plurality of reads. For example, the plurality of reads include a plurality of variant reads that are different from a reference genome, and the first annotation information includes at least one of a minimum insert size of the plurality of variant reads, a maximum insert size of the plurality of variant reads, or a number of paired reads among the plurality of variant reads that satisfy a specific condition, wherein the specific condition may include a condition in which the first read and the second read of the paired read are aligned in a forward and reverse direction, respectively, and the insert size of the paired read is between a lower threshold and an upper threshold.

[0232] In one embodiment, the processor may receive normal tissue genomic data (PON: Panel of Normals) generated from a plurality of sequencing data associated with a plurality of normal samples, and generate second annotation information associated with the normal tissue genomic data.

[0233] In one embodiment, a processor may receive FFPE-processed tissue genomic data (Panel of FFPEs: POF) generated from a plurality of sequencing data associated with a plurality of Formalin-Fixed, Paraffin-Embedded (FFPE) processed samples, and generate third annotation information associated with the FFPE-processed tissue genomic data. The third annotation information may include a number of samples, among the plurality of FFPE-processed samples, having a Variant Allele Frequency (VAF) for a specific position in a base sequence within the sample below a predetermined threshold. The third annotation information may include a number of samples, among the plurality of FFPE-processed samples, having a predetermined number of variant reads at a specific position in a base sequence within the sample.

[0234] In one embodiment, the processor may generate fourth annotation information including information associated with a mutation type of the reference mutation candidate and sequence context information of the reference mutation candidate.

[0235] The processor can generate learning data based on the determined reference mutation candidate information and the generated annotation information (S1340).

[0236] In one embodiment, the processor may label classification information for a reference variant candidate. For example, the reference sample may be a FFPE-processed sample, and the processor may label the reference variant candidate as a true positive variant in response to determining that at least a portion of the reference variant candidate information corresponds to at least a portion of information associated with one of the variant candidates in the Fresh-Frozen (FFPE)-processed sample. In this case, the FF-processed sample may correspond to the FFPE-processed sample.

[0237] In one embodiment, the processor may label the reference variant candidate as a false positive variant in response to determining that the reference variant candidate information and the information associated with any one of the variant candidates in the FF-processed sample do not correspond to each other.

[0238] In one embodiment, the processor can generate training data by extracting features of the reference variant candidate based on reference variant candidate information and annotation information, and including a data set including the reference variant candidate information, the extracted features of the reference variant candidate, and labeled classification information in the training data.

[0239] The processor can train a machine learning model using the generated training data (S1350). In one embodiment, the machine learning model includes a plurality of classifiers, and the processor inputs reference mutation candidate information and features of the reference mutation candidate into each of the plurality of classifiers, determines a classification result indicating whether the reference mutation candidate is a true positive mutation using an output result from at least one of the plurality of classifiers, and adjusts parameters or hyperparameters of the machine learning model based on the classification result and classification information labeled on the reference mutation candidate.

[0240] In one embodiment, a machine learning model may receive target mutation candidate information and features of the target mutation candidate within a target sample, and output a classification result indicating whether the target mutation candidate is a true-positive mutation. The target sample may include a target normal sample and a target abnormal sample collected from the same individual, and the target mutation candidate information may be determined based on first target sequencing data associated with the target normal sample and second target sequencing data associated with the target abnormal sample using a mutation detection module. The target sample may be an FFPE-processed sample.

[0241] FIG. 14 is a flowchart illustrating a method (1400) for genome profiling through detection of true-positive mutations in a cell sample according to one embodiment of the present disclosure. The method (1400) may be performed by at least one processor. The method (1400) may begin when the processor acquires target mutation candidate information of a target sample (S1410). The target sample may be an FFPE-processed sample.

[0242] In one embodiment, the target sample includes a target abnormality sample, and target mutation candidate information may be determined based on target sequencing data associated with the target abnormality sample. For example, the target mutation candidate information may be determined based on target sequencing data associated with the target abnormality sample, such as by performing deep sequencing, comparing with a database of known mutations, etc.

[0243] In another embodiment, the target sample includes a target normal sample and a target abnormal sample collected from the same individual, and the target variant candidate information can be determined based on first target sequencing data associated with the target normal sample and second target sequencing data associated with the target abnormal sample using a variant detection module.

[0244] Thereafter, the processor can use a machine learning model to determine a classification result indicating whether the target mutation candidate is a true-positive mutation (S1420). The machine learning model may be a model trained by the method for training a machine learning model (1300) of FIG. 13. For example, the machine learning model may be a model trained to determine whether a reference mutation candidate is a true-positive mutation using reference mutation candidate information of a reference sample and annotation information associated with the reference mutation candidate. Through this, noise or errors that may occur during whole-genome analysis can be corrected, thereby producing undistorted analysis results.

[0245] Thereafter, the processor may perform genomic profiling on the target sample based on the determined classification results (S1430). For example, the processor may perform genomic profiling on the target sample using a mutation list from which some mutation candidates (e.g., mutation candidates determined to be false positive mutations) have been deleted / filtered based on the determined classification results.

[0246] As an example of genomic profiling performed based on the determined classification results, the processor can identify and analyze genetic variations such as SNVs, INDELs, and chromosomal rearrangements in the genome of a target sample. As another example, the processor can investigate and identify how, when, and to what extent genes of an individual are expressed within the individual. As another example, the processor can analyze epigenetic changes such as DNA methylation and histone modifications to determine factors that affect gene expression in the individual. As another example, the processor can identify changes in the genome structure of an individual by investigating large-scale chromosomal rearrangements or copy number changes. As another example, the processor can understand the genetic characteristics of an individual and predict disease risk, drug response, etc.

[0247] The above-described genetic profiling can be applied in various fields such as early diagnosis of diseases, personalized medicine, genetic disease research, and drug development. By performing genetic profiling based on the classification results indicating whether the target mutation candidate is a true mutation, precise and extensive genetic information can be provided for medical research and clinical applications.

[0248] Thereafter, based on the results of performing the genomic profiling, the processor may provide at least one of disease diagnosis information, treatment strategy information, prognosis prediction information, or drug responsiveness prediction information of the individual from whom the target sample was collected (S1440). For example, the processor may confirm the presence of BRCA1 and BRCA2 gene mutations in the target sample based on the results of the genomic profiling performed in step S1430, and may diagnose that the individual from whom the target sample was collected has a high risk of developing breast cancer and ovarian cancer (an example of providing disease diagnosis information). The processor may determine the timing, method, and type of treatment based on the results of performing the genomic profiling of the individual, thereby treating the disease more effectively (an example of providing treatment strategy information). The processor may provide prognosis prediction information by predicting the likely course of the disease. The processor may predict how an individual will react to different types and dosages of drugs based on the individual's genetic makeup, and adjust / determine the prescription of the drug, thereby increasing the efficacy of the drug and reducing side effects (an example of providing drug responsiveness prediction information).

[0249] The flowcharts and descriptions illustrated in Figures 13 and 14 are merely examples, and some embodiments may be implemented differently. For example, one or more steps may be omitted, the order of each step may be changed, one or more steps may be performed in an overlapping manner, or one or more steps may be repeated multiple times.

[0250] The above-described method may be provided as a computer program stored on a computer-readable recording medium for execution on a computer. The medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program instructions, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.

[0251] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, these techniques may be implemented in hardware, firmware, software, or a combination thereof. Those skilled in the art will appreciate that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and the design requirements imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementations should not be construed as departing from the scope of the present disclosure.

[0252] In a hardware implementation, the processing units used to perform the techniques may be implemented within one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described herein, a computer, or a combination thereof.

[0253] Accordingly, the various exemplary logical blocks, modules, and circuits described in connection with the present disclosure may be implemented or performed by any combination of a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or those designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0254] In a firmware and / or software implementation, the techniques may be implemented as instructions stored on a computer-readable medium, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, a compact disc (CD), a magnetic or optical data storage device, etc. The instructions may be executable by one or more processors and may cause the processor(s) to perform certain aspects of the functionality described herein.

[0255] While the embodiments described above have been described as utilizing aspects of the presently disclosed subject matter in one or more standalone computer systems, the present disclosure is not limited thereto and may be implemented in conjunction with any computing environment, such as a network or distributed computing environment. Furthermore, aspects of the present disclosure may be implemented in multiple processing chips or devices, and storage may be similarly affected across multiple devices. Such devices may include personal computers, network servers, and portable devices.

[0256] While the present disclosure has been described in connection with certain embodiments herein, various modifications and variations may be made without departing from the scope of the present disclosure, which would be apparent to those skilled in the art. Furthermore, such modifications and variations are intended to fall within the scope of the claims appended to this specification.

Claims

1. A method for training a machine learning model, which is executed by at least one processor, A step of obtaining reference mutation candidate information of a reference sample; A step of generating annotation information associated with the above reference mutation candidate; A step of generating learning data based on the obtained reference mutation candidate information and the generated annotation information; and Step of training a machine learning model using the above generated learning data A method for training a machine learning model, comprising:

2. In paragraph 1, The above reference sample is, Includes reference normal samples and reference abnormal samples collected from the same individual, The above acquired reference mutation candidate information is, A method for training a machine learning model, the machine learning model being determined based on first reference sequencing data associated with the reference normal sample and second reference sequencing data associated with the reference abnormal sample, using a variant detection module.

3. In paragraph 2, The above mutation detection module comprises a plurality of detection modules, The above acquired reference mutation candidate information is, It is determined by integrating (union) the reference variant sub-candidate information obtained using the above multiple detection modules, The above-mentioned acquired reference mutation sub-candidate information is, A method for training a machine learning model, wherein the machine learning model is determined by applying the first reference sequencing data and the second reference sequencing data to each of the plurality of detection modules.

4. In paragraph 1, The steps for generating the above annotation information are: determining a plurality of reads in which at least some of the mapped positions overlap with the positions of the reference mutation candidates; and A step of generating first annotation information associated with the plurality of leads determined above A method for training a machine learning model, comprising:

5. In paragraph 4, The above plurality of reads include a plurality of variant reads that are different from the reference genome, The above first annotation information is, It includes at least one of the minimum value of the insert size of the plurality of mutant reads, the maximum value of the insert size of the plurality of mutant reads, or the number of paired reads satisfying a specific condition among the plurality of mutant reads, The above specific conditions are, A method for training a machine learning model, comprising the conditions that the first lead and the second lead of the paired lead are aligned in the forward and reverse directions, respectively, and the insert size of the paired lead is between a lower limit threshold and an upper limit threshold.

6. In paragraph 1, The steps for generating the above annotation information are: A step of receiving normal tissue genome data (PON: Panel of Normals) generated from a plurality of sequencing data associated with a plurality of normal samples; and A step of generating second annotation information associated with the above normal tissue genome data. A method for training a machine learning model, comprising:

7. In paragraph 1, The steps for generating the above annotation information are: A step of receiving FFPE processed tissue genome data (POF: Panel of FFPEs) generated from a plurality of sequencing data associated with a plurality of FFPE (Formalin-Fixed, Paraffin-Embedded) processed samples; and A step of generating third annotation information associated with the above FFPE processed tissue genome data. A method for training a machine learning model, comprising:

8. In paragraph 7, The above third annotation information is, A method for training a machine learning model, the method comprising: training a plurality of FFPE-processed samples, wherein the number of samples having a variant allele frequency (VAF) for a specific position in the base sequence within the sample is less than a predetermined threshold.

9. In paragraph 7, The above third annotation information is, A method for training a machine learning model, the machine learning model comprising a number of samples having a predetermined number of mutant reads at a specific location in the base sequence of the sample among the plurality of FFPE processed samples.

10. In paragraph 1, The steps for generating the above annotation information are: A step of generating fourth annotation information including information related to the mutation type of the above reference mutation candidate and sequence context information of the above reference mutation candidate. A method for training a machine learning model, comprising:

11. In paragraph 1, The steps for generating the above learning data are: Step of labeling classification information for the above reference mutation candidates A method for training a machine learning model, comprising:

12. In paragraph 11, The above reference sample is an FFPE processed sample, The above labeling steps are: A step of labeling the reference mutation candidate as a true positive mutation in response to determining that at least a part of the reference mutation candidate information and at least a part of the information associated with one of the mutation candidates in the FF (Fresh-Frozen) processed sample correspond to each other. Including, A method for training a machine learning model, wherein the above FF processed sample is a sample corresponding to the above FFPE processed sample.

13. In paragraph 12, The above labeling steps are: In response to determining that the above reference mutation candidate information and the information associated with one of the mutation candidates in the FF-processed sample do not correspond to each other, a step of labeling the above reference mutation candidate as a false positive mutation A method for training a machine learning model, further comprising:

14. In paragraph 11, The steps for generating the above learning data are: A step of extracting features of the reference mutation candidate based on the above reference mutation candidate information and the annotation information; and A step of including a data set including the above reference mutation candidate information, the features of the extracted reference mutation candidate, and the labeled classification information in the above learning data. A method for training a machine learning model, further comprising:

15. In paragraph 14, The above machine learning model includes multiple classifiers, The step of training the above machine learning model is: A step of inputting the above reference mutation candidate information and the features of the above reference mutation candidate into each of the plurality of classifiers; A step of determining a classification result indicating whether the reference mutation candidate is a true positive mutation by using the output result from at least one classifier among the plurality of classifiers; and A step of adjusting the parameters of the machine learning model based on the classification result and the classification information labeled on the reference mutation candidate. A method for training a machine learning model, comprising:

16. In paragraph 1, The above machine learning model is, A method for training a machine learning model, which receives target mutation candidate information and features of the target mutation candidate in a target sample, and outputs a classification result indicating whether the target mutation candidate is a true positive mutation.

17. In paragraph 16, The above target sample is, Contains target normal samples and target abnormal samples collected from the same individual, The above target mutation candidate information is: A method for training a machine learning model, the machine learning model being determined based on first target sequencing data associated with the target normal sample and second target sequencing data associated with the target abnormal sample, using a mutation detection module.

18. In paragraph 16, A method for training a machine learning model, wherein the above target sample is an FFPE processed sample.

19. A method for genetic profiling through detection of a true positive mutation in a cell sample, executed by at least one processor, A step of obtaining target mutation candidate information of a target sample; A step of using a machine learning model to determine a classification result indicating whether the target mutation candidate is a true positive mutation; and A step of performing genomic profiling on the target sample based on the classification result determined above. Including, The above machine learning model is, A genome profiling method, which learns to determine whether the reference mutation candidate is a true-positive mutation by using reference mutation candidate information of a reference sample and annotation information associated with the reference mutation candidate.

20. In paragraph 19, A step of providing at least one of disease diagnosis information, treatment strategy information, prognosis prediction information, or drug responsiveness prediction information of the individual from which the target sample was collected based on the results of performing the above genetic profiling. A method of genetic profiling, further comprising:

21. A computer-readable, non-transitory recording medium having recorded thereon commands for executing the method according to Article 1 on a computer.

22. In the device, memory; and At least one processor connected to said memory and configured to execute at least one computer-readable program contained in said memory Including, At least one of the above programs, Obtain reference mutation candidate information of the reference sample, Generate annotation information associated with the above reference mutation candidate information, Generate learning data based on the above-obtained reference mutation candidate information and the above-generated annotation information, A device including a command for training a machine learning model using the generated learning data.

Citation Information

Patent Citations

  • Superconducting current fault limiter with direct cooling structure

    KR1020230139118A

  • Method and apparatus for display virtual realiry contents on user terminal based on the determination that the user terminal is located in the pre-determined customized region

    KR1020240030345A

  • Aquaponics system by applying heat pump for water temperature optimization for growth and development

    KR102070359B1

  • Systems and methods for artificial intelligence powered molecular workflow verifying slide and block quality for testing

    US20220293249A1