Method for determining amplifiable region in multiple target nucleic acid sequences of target nucleic acid molecule, method for obtaining prediction model to be used for determining amplifiable region, method for providing information on amplifiable region including priority, and computer device for performing same

A computer-based method determines amplifiable regions in nucleic acid sequences to enhance oligonucleotide design for pathogen detection, addressing inefficiencies in existing technologies and improving detection accuracy and efficiency.

WO2026038895A1PCT designated stage Publication Date: 2026-02-19SEEGENE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/012323
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-16
Filing Date
2025-08-13
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Existing molecular diagnostic technologies face challenges in detecting pathogens with genetic diversity due to the inefficiency and low accuracy in selecting regions for designing oligonucleotides, leading to false-negative results and increased time and cost.

Method used

A method using a computer device to determine an amplifiable region in nucleic acid sequences by analyzing alignment results, obtaining feature data, and utilizing a prediction model to predict amplification efficiency, thereby identifying regions suitable for oligonucleotide design.

Benefits of technology

Improves detection accuracy and efficiency by identifying high-priority amplifiable regions for oligonucleotide design, reducing time and cost by providing immediate access to optimized regions for pathogen detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025012323_19022026_PF_FP_ABST
    Figure KR2025012323_19022026_PF_FP_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a method, performed by a computer device, for determining an amplifiable region available for designing an oligonucleotide in multiple nucleic acid sequences of a target nucleic acid molecule, the method comprising the steps of: determining, on the basis of an alignment result of the multiple nucleic acid sequences, a region of interest to be used for determining an amplifiable region; obtaining feature data predicted to affect an amplification reaction in the region of interest; providing the feature data as input data to a prediction model; obtaining, from the prediction model, output data including an amplification reaction efficiency predicted for the amplification reaction in the region of interest; and determining, on the basis of the amplification reaction efficiency, the amplifiable region for detection of the target nucleic acid molecule in the alignment result.
Need to check novelty before this filing date? Find Prior Art

Description

A method for determining an amplifiable region from a plurality of target nucleic acid sequences of a target nucleic acid molecule, a method for obtaining a prediction model used for determining an amplifiable region, a method for providing information on an amplifiable region including a priority, and a computer device for performing the same.

[0001] The present invention relates to a method for determining an amplifiable region in a plurality of target nucleic acid sequences of a target nucleic acid molecule, a method for obtaining a prediction model used for determining an amplifiable region, a method for providing information on an amplifiable region including a priority, and a computer device for performing the same.

[0002] Various technologies have been developed to detect and identify target nucleic acid molecules of pathogens, collectively known as molecular diagnosis. Most molecular diagnostic technologies utilize target nucleic acid molecule-hybridizing oligonucleotides, such as primers and probes.

[0003] Molecular diagnostic technologies have made significant progress to date. However, there are still technical challenges to be addressed in the diagnosis of pathogens whose genomes exhibit genetic diversity or variability.

[0004] Genetic diversity, or genetic variability, has been reported in various genomes. In particular, genetic diversity occurs most frequently in viral genomes (Bastien N. et al., Journal of Clinical Microbiology, 42:3532(2004); Peret TC. et al., Journal of Infectious Diseases, 185:1660(2002); Ebihara T. et al., Journal of Clinical Microbiology, 42:126(2004); Jenny-Avital ER. et al. Clinical Infectious Diseases, 32:1227(2001); Duffy S. et al., Nat. Rev. Genet. 9(4):267-76(2008); Tong YG et al., Nature. 22:526(2015)).

[0005] When detecting pathogens with genetic diversity, designing and using oligonucleotides based on the nucleic acid sequence of a specific target nucleic acid molecule of that pathogen can result in false-negative results. Therefore, to determine the presence of a specific pathogen in an unknown sample, probes or primers must be designed based on all or as many known genetically diverse nucleic acid sequences of a single target nucleic acid molecule of that specific pathogen. Two major methods have been developed to detect target nucleic acid molecules exhibiting such genetic diversity.

[0006] The first method involves designing degenerate oligonucleotides. Typically, a region containing sequences with sequence similarity is identified in the alignment of all nucleic acid sequences of a specific gene with genetic diversity. A degenerate primer or probe (containing degenerate bases at the mutation site) that hybridizes to this region is used to detect the specific gene with the desired coverage. If the degenerate oligonucleotide fails to detect the specific gene with the desired coverage, the second method below is used.

[0007] The second method detects a target nucleic acid molecule using multiple oligonucleotides that hybridize to multiple nucleic acid sequences of the target nucleic acid molecule that exhibit genetic diversity. For example, when targeting the M gene of the influenza A virus, all known nucleic acid sequences for the M gene are aligned, and probes that can cover all of these nucleic acid sequences are designed. In this case, since a single probe cannot cover all of the diverse sequences of the M gene, multiple probes (probes with different probing positions) are designed. In addition, in this case, degenerate bases are sometimes introduced into the multiple probes to further expand coverage.

[0008] In order to select either of the two methods described above, it is most important to determine a region within the target nucleic acid sequences within which oligonucleotides can be designed to detect diverse nucleic acid sequences of the target nucleic acid molecule using oligonucleotides or a combination thereof. Considering the convenience, efficiency, and cost-effectiveness of analysis, it is preferable to select a region within which oligonucleotides can be designed from a plurality of target nucleic acid sequences having sequence similarity, and design an oligonucleotide to be used in either of the two methods described above from this region.

[0009] In the past, in order to detect the nucleic acid sequences of diverse target nucleic acid molecules, analysts visually inspected the alignment results of multiple target nucleic acid sequences, selected a conservative region excluding positions where mutant bases existed, and then designed oligonucleotides covering the multiple target nucleic acid sequences in the region, and determined the positions and number of degenerate bases to be applied to the designed oligonucleotides to expand target coverage, or determined the optimal combination of oligonucleotides.

[0010] However, in cases where the number of target nucleic acid sequences is large, the conventional method has the problem of taking a long time and having low accuracy in selecting a conservative region for designing an oligonucleotide that covers the maximum sequence. In addition, when the genetic diversity of multiple target nucleic acid sequences is large, there are disadvantages such as not being able to select a region for designing an oligonucleotide, introducing a large number of degenerate bases into a single oligonucleotide, or combining a large number of oligonucleotides, resulting in low accuracy and economy. In addition, there are problems such as not being able to select a region for designing an oligonucleotide, and having to select a target nucleic acid molecule other than the intended target nucleic acid molecule, which consumes time.

[0011] Numerous references and patents are cited and cited throughout this specification. The disclosures of these references and patents are incorporated herein by reference in their entirety to further clarify the state of the art and the scope of the present invention.

[0012] The problem to be solved according to one embodiment is to solve the above-described problems and / or limitations, and includes providing a region in a nucleic acid sequence that can efficiently design oligonucleotides (e.g., primers and probes) used to amplify and detect target nucleic acid molecules, particularly target nucleic acid molecules exhibiting genetic diversity.

[0013] Another problem to be solved, according to another embodiment, involves providing a region in a nucleic acid sequence that can design an oligonucleotide having higher detection accuracy by taking into account various factors predicted to affect amplification efficiency.

[0014] Another problem to be solved, according to another embodiment, involves efficiently providing regions in a nucleic acid sequence from which oligonucleotides can be designed using an artificial intelligence-based learning model.

[0015] Another problem to be solved in accordance with another embodiment includes providing an environment in which regions in a nucleic acid sequence that can be used for designing oligonucleotides for each target nucleic acid molecule are stored and can be immediately provided when a user requests them.

[0016] However, the problems to be solved by the present invention are not limited to those mentioned above, and other problems to be solved that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present invention pertains from the description below.

[0017] According to one embodiment of the present invention, a method for determining an amplifiable region usable for designing an oligonucleotide from a plurality of nucleic acid sequences of a target nucleic acid molecule, which is performed by a computer device, is provided. The method comprises the steps of: determining a region of interest used for determining the amplifiable region based on alignment results of a plurality of nucleic acid sequences; obtaining feature data predicted to affect an amplification reaction in the region of interest; providing the feature data as input data to a prediction model; obtaining output data including an amplification reaction efficiency predicted for an amplification reaction in the region of interest from the prediction model; and determining the amplifiable region for detection of the target nucleic acid molecule from the alignment results based on the amplification reaction efficiency.

[0018] In one embodiment, the region of interest may include (a) a conserved region included in the alignment result, (b) a fragmented region of a predetermined length segmented from the conserved region, or (c) a fragmented region of a predetermined length segmented from the entire region of the plurality of nucleic acid sequences.

[0019] In one embodiment, the step of determining the region of interest may include: obtaining a conservative region including sequences having a predetermined level or higher of sequence similarity from the alignment result; determining a plurality of fragment regions of a predetermined length that are sequentially fragmented or fragmented so that some of the fragment regions overlap each other within the conservative region; and determining each of the plurality of fragment regions as the region of interest.

[0020] In one embodiment, the predetermined length may be greater than or equal to 50 bp and less than or equal to 150 bp, and the mutually overlapping length may be greater than or equal to 10 bp and less than or equal to 50 bp.

[0021] In one embodiment, the step of obtaining the output data may obtain output data including an amplification reaction efficiency predicted for an amplification reaction in each of the plurality of fragment regions, and the step of determining the amplifiable region may determine the amplifiable region within the conservative region based on the amplification reaction efficiency of each of the plurality of fragment regions.

[0022] In one embodiment, the step of obtaining the feature data generates the feature data using information of the region of interest, and the information of the region of interest may include at least one selected from the group consisting of a nucleic acid sequence of the region of interest, a length of the region of interest, a biological category of an organism having the target nucleic acid molecule, a nucleic acid type of the target nucleic acid molecule, a gene type of the target nucleic acid molecule, and reaction conditions for the amplification reaction.

[0023] In one embodiment, the nucleic acid sequence of the region of interest may include at least one selected from the group consisting of (a) a sequence determined from a nucleic acid sequence that satisfies a predetermined length condition among the plurality of nucleic acid sequences, (b) a sequence determined from a unique genome sequence of an organism having the target nucleic acid molecule, (c) a sequence determined based on a grouping result for the plurality of nucleic acid sequences, and (d) a sequence determined from a nucleic acid sequence that contains the fewest non-conserved bases among the plurality of nucleic acid sequences.

[0024] In one embodiment, the feature data may include at least one selected from the group consisting of (a) a nucleic acid sequence of the region of interest; (b) thermodynamic data on the formation of an n-order structure (wherein n is an integer greater than or equal to 2) in the nucleic acid sequence of the region of interest; (c) distance data between a predetermined position in the nucleic acid sequence of the region of interest and a position of the n-order structure; (d) reaction conditions including a reaction medium used in the amplification reaction; the reaction medium including at least one material selected from the group consisting of a pH-related material, an ionic strength-related material, an enzyme, and an enzyme stabilization-related material; (f) a type of oligonucleotide that binds to the nucleic acid sequence of the region of interest; (g) a GC content of the nucleic acid sequence of the region of interest; (h) a variation score for the region of interest; and (i) a conservation score for the region of interest.

[0025] In one embodiment, n is 2, and the n-th structure may include at least one selected from the group consisting of a hairpin loop, an internal loop, a bulge loop, multi-loops, a G-quadruplex, and combinations thereof.

[0026] In one embodiment, the thermodynamic data may be expressed in the form of a change in thermodynamic free energy.

[0027] In one embodiment, the thermodynamic data may include thermodynamic data for an n-order structure present in a nucleic acid sequence of an amplicon region corresponding to the region of interest, and / or thermodynamic data for an n-order structure present in a nucleic acid sequence of an oligo bidding region determined from the amplicon region.

[0028] In one embodiment, the mutation score for the region of interest may be calculated based on at least one selected from the group consisting of (a) a mutation position at which a non-conservative base is located in the region of interest of the plurality of nucleic acid sequences, (b) a mutation type indicating a type of the non-conservative base, (c) the number of the non-conservative base and / or the mutation type at the mutation position, and (d) the number of the mutation positions in the region of interest of the plurality of nucleic acid sequences.

[0029] In one embodiment, the conservation score for the region of interest may be calculated based on an entropy value that measures the degree of uncertainty from a probability distribution of similarities of base sequences in the region of interest among the plurality of nucleic acid sequences.

[0030] In one embodiment, the step of obtaining the feature data may be performed by a preprocessing unit configured to calculate the value of each of the plurality of features included in the feature data using a plurality of previously stored mathematical formulas, or may be performed by a model trained to output the value of each of the plurality of features when information on the region of interest is input.

[0031] In one embodiment, the learned model may be operated integrally with the prediction model as a model included in the prediction model, or may be operated independently from the prediction model as a model distinct from the prediction model.

[0032] In one embodiment, the prediction model is trained using a plurality of training data sets, and each of the plurality of training data sets may include training input data including feature data predicted to affect an amplification reaction for a training nucleic acid sequence and training answer data including an amplification reaction efficiency for an amplification reaction for the training nucleic acid sequence.

[0033] In one embodiment, the feature data included in the training input data may include at least one selected from the group consisting of (a) the training nucleic acid sequence; (b) thermodynamic data on the formation of an n-order structure (wherein n is an integer greater than or equal to 2) in the training nucleic acid sequence; (c) distance data between a predetermined position in the training nucleic acid sequence and a position of the n-order structure; (d) reaction conditions including a reaction medium used for an amplification reaction for the training nucleic acid sequence; the reaction medium including at least one material selected from the group consisting of a pH-related material, an ionic strength-related material, an enzyme, and an enzyme stabilization-related material; (e) a type of oligonucleotide bound to the training nucleic acid sequence, (f) a GC content of the training nucleic acid sequence, (g) a variation score for a training sequence group including the training nucleic acid sequence, and (h) a conservation score for the training sequence group.

[0034] In one embodiment, the amplification reaction efficiency included in the training correct answer data may include numerical data on the efficiency level of the amplification reaction for the training nucleic acid sequence and / or label data on the high and low of the efficiency level.

[0035] In one embodiment, the amplification reaction efficiency included in the training correct answer data can be calculated using (a) the difference in amplification points determined from two or more data sets obtained from two or more amplification reactions for the training nucleic acid sequence, or (b) the difference between a signal pattern determined from the two or more data sets and a preset reference pattern.

[0036] In one embodiment, the region of interest is plural, and the step of determining the amplifiable region may determine the amplifiable region by using one or more regions of interest among the plurality of regions of interest whose amplification reaction efficiency satisfies a preset first criterion.

[0037] In one embodiment, the step of determining the amplifiable region may include (a) determining the merged region as the amplifiable region if the region merged from the one or more regions of interest satisfies a preset second criterion, or (b) determining each of the one or more regions of interest as the amplifiable region.

[0038] In one embodiment, the second criterion may include (a) a criterion for the length of the merged region, and / or (b) a criterion for a statistical operation of the amplification reaction efficiency of each of the one or more regions of interest.

[0039] In one embodiment, the step of determining the amplifiable region may include the step of: if there is no region of interest among the plurality of regions of interest whose amplification reaction efficiency satisfies the first criterion, redetermining the region of interest so that the length of the region of interest is reduced; re-acquiring the amplification reaction efficiency by re-performing the step of obtaining the feature data or the step of obtaining the output data based on the re-determined region of interest; and determining the amplifiable region using one or more regions of interest among the re-determined regions of interest whose amplification reaction efficiency satisfies the first criterion.

[0040] In one embodiment, in the step of obtaining the feature data, first feature data and second feature data each affecting an amplification reaction in each of an amplicon region and an oligo bidding region determined from the region of interest are obtained, and in the step of providing the feature data to the prediction model, each of the first feature data and the second feature data is provided as input data to the prediction model, and in the step of obtaining the output data, each of a first amplification reaction efficiency and a second amplification reaction efficiency predicted for an amplification reaction in each of the amplicon region and the oligo bidding region is obtained from the prediction model, and the amplification reaction efficiency in the region of interest can be determined based on the first amplification reaction efficiency and the second amplification reaction efficiency.

[0041] In one embodiment, the prediction model includes a first prediction model trained to predict an amplification reaction efficiency for an amplification reaction in an amplicon sequence using first training datasets including thermodynamic data for an n-th structure present in the amplicon sequence, and / or a second prediction model trained to predict an amplification reaction efficiency for an amplification reaction in an oligo bidding sequence using second training datasets including thermodynamic data for an n-th structure present in an oligo bidding sequence determined from the amplicon sequence, and in the step of obtaining the output data, the first amplification reaction efficiency can be obtained from the first prediction model, and the second amplification reaction efficiency can be obtained from the second prediction model.

[0042] In one embodiment, the method further comprises a step of obtaining and providing information on the amplifiable region, wherein the information on the amplifiable region may include at least one selected from the group consisting of (a) information on the location, base sequence, score, and bindable oligonucleotide of the region of interest included in the amplifiable region, and (b) information on the location, base sequence, score, and bindable oligonucleotide of the amplifiable region.

[0043] In one embodiment, the score of the region of interest may be calculated using at least one selected from the group consisting of (i) numerical data included in the amplification reaction efficiency of the region of interest, (ii) values ​​of each of a plurality of features included in the feature data, and (iii) contributions of each of the plurality of features in the prediction model to the output of the amplification reaction efficiency, and the score of the amplifiable region may be calculated based on the score of the region of interest included in the amplifiable region.

[0044] In one embodiment, the score of the amplifiable region includes a score for each of a plurality of preset items, and the plurality of items may include at least one item selected from the group consisting of (i) a first amplification reaction efficiency in an amplicon region included in the amplification reaction efficiency, (ii) a second amplification reaction efficiency in an oligo binding region included in the amplification reaction efficiency, (iii) a variation score for the amplifiable region, and (iv) a conservation score for the amplifiable region.

[0045] In one embodiment, the prediction model may include at least one selected from the group consisting of a Random Forest (RF), a Logistic Regressor (LR), a Gradient Boosting Classifier (GBC), a Decision Tree Classifier (DTC), a Gaussian Naive Bayes (GNB), a Support Vector Classifier (SVC), and a Neural Network (NN).

[0046] According to one embodiment of the present invention, a computer device is provided, comprising: a memory storing at least one instruction; and a processor. The at least one instruction, when executed by the processor, causes the processor to perform the following operations, wherein the operations may include: determining a region of interest to be used for determining an amplifiable region based on alignment results of a plurality of nucleic acid sequences of a target nucleic acid molecule; obtaining feature data predicted to affect an amplification reaction in the region of interest; providing the feature data as input data to a prediction model; obtaining output data including an amplification reaction efficiency predicted for an amplification reaction in the region of interest from the prediction model; and determining the amplifiable region for detection of the target nucleic acid molecule from the alignment results based on the amplification reaction efficiency.

[0047] According to one embodiment of the present invention, a computer-readable, non-transitory recording medium storing a computer program is provided. The computer program includes instructions that, when executed by one or more processors, cause the one or more processors to perform a method for determining an amplifiable region usable for designing an oligonucleotide from a plurality of nucleic acid sequences of a target nucleic acid molecule, the method comprising: determining a region of interest used for determining an amplifiable region based on alignment results of the plurality of nucleic acid sequences; obtaining feature data predicted to affect an amplification reaction in the region of interest; providing the feature data as input data to a prediction model; obtaining output data from the prediction model, the output data including an amplification reaction efficiency predicted for an amplification reaction in the region of interest; and determining the amplifiable region for detection of the target nucleic acid molecule from the alignment results based on the amplification reaction efficiency.

[0048] According to one embodiment of the present invention, a method for obtaining a prediction model that provides predicted amplification efficiency for an amplification reaction performed by a computing device is provided. The method comprises the steps of: obtaining a plurality of training data sets; and obtaining a prediction model learned to predict the amplification efficiency for the amplification reaction using the plurality of training data sets, wherein each of the plurality of training data sets includes (a) training input data including feature data predicted to affect the amplification reaction of a training nucleic acid sequence, and (b) training answer data including the amplification efficiency for the amplification reaction in the training nucleic acid sequence, and the feature data included in the training input data may include thermodynamic data for the formation of an n-order structure (wherein n is an integer equal to or greater than 2) in the training nucleic acid sequence.

[0049] In one embodiment, the feature data included in the training input data may further include at least one selected from the group consisting of (a) the training nucleic acid sequence; (b) distance data between a predetermined position in the training nucleic acid sequence and a position of the n-th structure; (c) reaction conditions including a reaction medium used for an amplification reaction for the training nucleic acid sequence; the reaction medium including at least one material selected from the group consisting of a pH-related material, an ionic strength-related material, an enzyme, and an enzyme stabilization-related material; (d) a type of oligonucleotide bound to the training nucleic acid sequence, (e) a GC content of the training nucleic acid sequence, (f) a variation score for a training sequence group including the training nucleic acid sequence, and (g) a conservation score for the training sequence group.

[0050] In one embodiment, the amplification reaction efficiency included in the training correct answer data can be calculated using (a) the difference in amplification points determined from two or more data sets obtained from two or more amplification reactions for the training nucleic acid sequence, or (b) the difference between a signal pattern determined from the two or more data sets and a preset reference pattern.

[0051] In one embodiment, the thermodynamic data includes first thermodynamic data for the n-th structure present in an amplicon sequence determined from the training nucleic acid sequence and second thermodynamic data for the n-th structure present in an oligo bidding sequence determined from the amplicon sequence, and the step of obtaining the prediction model may include: obtaining a first prediction model learned to predict an amplification reaction efficiency for an amplification reaction in the amplicon sequence by using a plurality of first training datasets including the first thermodynamic data among the plurality of training datasets; and obtaining a second prediction model learned to predict an amplification reaction efficiency for an amplification reaction in the oligo bidding sequence by using a plurality of second training datasets including the second thermodynamic data among the plurality of training datasets.

[0052] According to one embodiment of the present invention, a method for providing information on amplifiable regions that can be used in designing an oligonucleotide for detecting a target nucleic acid molecule, which is performed by a computing device, is provided. The method comprises the steps of: determining amplifiable regions based on a plurality of nucleic acid sequences of the target nucleic acid molecule; determining priorities for using the amplifiable regions in designing the oligonucleotide; and storing information on the amplifiable regions including the priorities in a database, wherein the priorities can be determined based on thermodynamic data on the formation of an n-order structure (wherein n is an integer greater than or equal to 2) in each of the amplifiable regions.

[0053] In one embodiment, n is 2, and the n-th structure may include at least one selected from the group consisting of a hairpin loop, an internal loop, a bulge loop, multi-loops, a G-quadruplex, and combinations thereof.

[0054] In one embodiment, the priority may be determined by further using at least one selected from the group consisting of (a) distance data between a given position in the nucleic acid sequence of each amplifiable region and the position of the n-th structure; (b) GC content of the nucleic acid sequence of each amplifiable region; (c) variation score for each amplifiable region; and (d) conservation score for each amplifiable region.

[0055] In one embodiment, the step of determining the amplifiable regions is performed at least in part using a prediction model learned to predict the amplification reaction efficiency using feature data that affects the amplification reaction of the nucleic acid sequence, and the priority can be further determined using the amplification reaction efficiency output from the prediction model.

[0056] In one embodiment, the priority may be determined using at least one selected from the group consisting of (i) numerical data included in the amplification reaction efficiency, (ii) values ​​of each of the plurality of features included in the feature data, and (iii) contributions of each of the plurality of features in the prediction model to the output of the amplification reaction efficiency.

[0057] In one embodiment, the method may further include the steps of: receiving a user input for a target nucleic acid molecule from a user terminal; searching for information on an amplifiable region associated with the user input from the database; and providing information on the amplifiable region according to the priority based on the search result to the user terminal.

[0058] In one embodiment, there may be multiple amplifiable regions according to the above priorities.

[0059] In one embodiment, the priority is determined based on scores of a plurality of preset items calculated for each of the amplifiable areas, and information of the amplifiable area according to the priority can be updated based on a selection input for the plurality of items received from the user terminal.

[0060] In one embodiment, the method may further include the steps of: receiving feedback information about the information of the provided amplifiable area from the user terminal; and updating the priority based on the feedback information.

[0061] In one embodiment, the step of updating the priority may change the weights used to determine the priority, update the priorities of the amplifiable regions, or perform additional learning on a prediction model used to determine the amplifiable region, based on the results of oligonucleotide design using the amplifiable region included in the feedback information.

[0062] In one embodiment, the method may further include a step of performing a design process of the oligonucleotide based on an amplifiable region selected by the user terminal from among the plurality of amplifiable regions according to the priority.

[0063] According to one embodiment of the present invention, amplifiable regions can be determined based on regions of interest predicted to exhibit high amplification efficiency based on feature data predicted to influence the amplification reaction. Accordingly, oligonucleotide design can be achieved using these amplifiable regions, further improving the detection accuracy of target nucleic acid molecules.

[0064] Furthermore, by utilizing a model trained on datasets obtained under specific nucleic acid sequences and reaction conditions, features difficult to analyze by humans can be reflected and utilized to predict amplification reaction efficiency, resulting in effective prediction performance. Furthermore, by applying the thermodynamic energy for secondary structure formation in nucleic acid sequences and the reaction conditions used in amplification reactions to the aforementioned feature data, prediction accuracy can be further improved.

[0065] Furthermore, by providing information on amplifiable regions for each target nucleic acid molecule, users can immediately access information on high-priority amplifiable regions when design requirements arise, thereby enhancing user convenience. Furthermore, since information on amplifiable regions is provided according to priority, subsequent oligonucleotide design can be selectively conducted based on high-priority amplifiable regions, thereby largely eliminating the subsequent steps previously performed for numerous regions, thereby saving time and money.

[0066] The effects of the present invention are not limited to the effects described above, and should be understood to include all effects that can be inferred from the detailed description of the present invention or the composition of the invention described in the claims.

[0067] FIG. 1 schematically illustrates a block diagram of a computer device according to one embodiment.

[0068] FIG. 2 illustrates an exemplary flowchart for determining an amplifiable area by a computer device according to the first embodiment.

[0069] FIG. 3 is a diagram illustrating alignment results of multiple nucleic acid sequences according to one embodiment.

[0070] FIG. 4 illustrates an exemplary manner in which a computer device determines a region of interest based on alignment results according to one embodiment.

[0071] FIG. 5 is a drawing illustrating a secondary structure according to one embodiment.

[0072] FIG. 6 is a diagram illustrating thermodynamic data for the formation of n-th order structures in a nucleic acid sequence of a region of interest according to one embodiment.

[0073] FIG. 7 is a conceptual diagram illustrating a process by which a prediction model according to one embodiment predicts an amplification reaction efficiency for an amplification reaction in a region of interest.

[0074] FIG. 8 illustrates an exemplary manner in which a computer device determines an amplifiable region according to one embodiment.

[0075] FIG. 9 is a diagram illustrating information on multiple amplifiable regions for a target nucleic acid molecule according to one embodiment.

[0076] FIG. 10 is a diagram illustrating information of amplifiable areas including priorities according to one embodiment.

[0077] FIG. 11 illustrates an exemplary flowchart of a computer device according to a second embodiment training a prediction model.

[0078] FIG. 12 is a diagram illustrating an amplification reaction efficiency calculated by using the difference in amplification points according to one embodiment.

[0079] FIG. 13 is a diagram for explaining the amplification reaction efficiency calculated by using the difference between the signal pattern and the reference pattern according to one embodiment.

[0080] Figure 14 illustrates a conceptual diagram of a learning process of a prediction model according to one embodiment.

[0081] Figure 15 is a structural diagram for explaining the operation of a computer device according to one embodiment.

[0082] FIG. 16 illustrates an exemplary flowchart for providing information on amplifiable regions available for oligonucleotide design for detection of target nucleic acid molecules by a computer device according to a third embodiment.

[0083] Figure 17 is a diagram illustrating a contribution according to one embodiment.

[0084]

[0085] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined solely by the scope of the claims.

[0086] When describing embodiments of the present invention, detailed descriptions of known functions or configurations will be omitted if they are deemed to unnecessarily obscure the gist of the invention. Furthermore, the terms described below are defined in light of their functions in the embodiments of the present invention and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the overall content of this specification.

[0087] Before explaining Figure 1, let us look at the terms used herein.

[0088] As used herein, the term "target nucleic acid molecule" refers to a nucleotide molecule within an organism to be detected. Target nucleic acid molecules are generally given a specific name and include the entire genome and all nucleotide molecules that make up the genome (e.g., genes, pseudogenes, non-coding sequence molecules, untranslated regions, and portions of the genome). Target nucleic acid molecules include, for example, nucleic acids from prokaryotes, eukaryotes, viruses, or viroids.

[0089] The term "organism" may encompass any form of life or organism that is to be analyzed, obtained, or detected. For example, an organism may refer to a living organism belonging to a genus, species, subspecies, subtype, genotype, sirotype, strain, isolate, or cultivar.

[0090] Organisms include, for example, prokaryotic cells (e.g., Mycoplasma pneumoniae, Chlamydophila pneumoniae, Legionella pneumophila, Haemophilus influenzae, Streptococcus pneumoniae, Bordetella pertussis, Bordetella parapertussis, Neisseria meningitidis, Listeria monocytogenes, Streptococcus agalactiae, Campylobacter, Clostridium difficile, Clostridium perfringens, Salmonella, Escherichia coli, Shigella, Vibrio, Yersinia enterocolitica, Aeromonas, Chlamydia trachomatis, Neisseria gonorrhoeae, Trichomonas vaginalis, Mycoplasma hominis, Mycoplasma genitalium, Ureaplasma urealyticum, Ureaplasma parvum, Mycobacterium tuberculosis), eukaryotic cells (e.g. Protozoa and parasites, fungi, yeasts, higher plants, lower animals, and higher animals including mammals and humans), viruses or viroids. Examples of parasites among the above eukaryotic cells include Giardia lamblia, Entamoeba histolytica, Cryptosporidium, Blastocystis hominis, Dientamoeba fragilis, and Cyclospora cayetanensis.Examples of the viruses include influenza A virus (Flu A), influenza B virus (Flu B), respiratory syncytial virus A (RSV A), respiratory syncytial virus B (RSV B), parainfluenza virus 1 (PIV 1), parainfluenza virus 2 (PIV 2), parainfluenza virus 3 (PIV 3), parainfluenza virus 4 (PIV 4), metapneumovirus (MPV), human enterovirus (HEV), human bocavirus (HBoV), human rhinovirus (HRV), coronaviruses and adenoviruses that cause respiratory diseases; and norovirus, rotavirus, adenovirus, astrovirus and sapovirus that cause gastrointestinal diseases. Additionally, examples of the viruses include human papillomavirus (HPV), Middle East respiratory syndrome-related coronavirus (MERSCoV), Dengue virus, Herpes simplex virus (HSV), Human herpes virus (HHV), Epstein-Barr virus (EMV), Varicella zoster virus (VZV), Cytomegalovirus (CMV), HIV, hepatitis virus, and poliovirus.

[0091] In one embodiment, the organism may be a GBS serotype, a bacterial colony, or v600e. The organism in the present disclosure may include various analysis targets, such as the aforementioned viruses, bacteria, and humans, and may also be a specific region of a gene cut using CRISPR technology, and is not limited to the examples described above.

[0092] The term "nucleic acid sequence" refers to a target nucleic acid molecule represented by a specific nucleic acid sequence, for example, a sequence of bases, which are components of nucleotides. For example, the term "nucleic acid sequence" in the present disclosure may be used interchangeably with the term "base sequence." Each individual base constituting the nucleic acid sequence may correspond to one of four types of bases, for example, A, G, C, and T.

[0093] A single target nucleic acid molecule, for example, a single target gene, may have one specific target nucleic acid sequence, or, in the case of a target nucleic acid molecule exhibiting genetic diversity or genetic variability, may have multiple diverse target nucleic acid sequences.

[0094] The plurality of nucleic acid sequences of the target nucleic acid molecule in the present disclosure are nucleic acid sequences having sequence similarity. Specifically, the nucleic acid sequences having sequence similarity may be a plurality of nucleic acid sequences of one target nucleic acid molecule or a plurality of nucleic acid sequences of two or more target nucleic acid molecules.

[0095] According to one embodiment of the present invention, the plurality of nucleic acid sequences are nucleic acid sequences having sequence similarity to one target nucleic acid molecule having genetic diversity.

[0096] For example, the plurality of nucleic acid sequences used in the present invention are a plurality of nucleic acid sequences having sequence similarity of target nucleic acid molecules exhibiting genetic diversity, such as the genome sequence of a virus. For example, when detecting an influenza A virus and selecting the M gene as the target nucleic acid molecule, nucleic acid sequences according to the diversity of the M gene of the influenza A virus can be used in the present invention. Not only the full-length nucleic acid sequence of the M gene of the influenza A virus but also a partial sequence can be used. Influenza A viruses include various subtypes and variants, and their genome sequences differ from each other. Therefore, in order to detect influenza A viruses without false negative results, the region in the nucleic acid sequence for designing an oligonucleotide must be determined by considering the various nucleic acid sequences of the target nucleic acid molecules of the influenza A virus according to such genetic diversity.

[0097] More specifically, the plurality of nucleic acid sequences are the entire genome sequence, a partial genome sequence, or the nucleic acid sequences of one gene of a virus or bacterium having genetic diversity.

[0098] According to one embodiment of the present invention, the plurality of nucleic acid sequences are a plurality of nucleic acid sequences corresponding to homologues of a plurality of organisms having the same function, the same structure, or the same gene name. The organism refers to an organism belonging to a genus, species, subspecies, subtype, genotype, serotype, strain, isolate, or cultivar. The homologue includes proteins and nucleic acid molecules. The embodiment uses a plurality of nucleic acid sequences of homologous biomolecules (e.g., proteins or nucleic acids) of a plurality of organisms having the same function (e.g., the biological function of the protein encoded by the nucleic acid sequence), the same structure (e.g., the tertiary structure of the protein encoded by the nucleic acid sequence), or the same gene name in the present invention. For example, a plurality of nucleic acid sequences known for the E5 gene of HPV type 16 can be considered as nucleic acid sequences of isolates of HPV type 16.

[0099] According to one embodiment of the present invention, the plurality of nucleic acid sequences include nucleic acid sequences belonging to a lower class of a biological classification (e.g., genus, species, subtype, genotype, serotype, and subspecies). For example, if the target nucleic acid sequence is HPV type 16, the target nucleic acid sequence may include nucleic acid sequences belonging to a lower class thereof.

[0100] According to one embodiment of the present invention, the plurality of nucleic acid sequences are at least 3, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, or at least 500 nucleic acid sequences. For example, in FIG. 2, the plurality of nucleic acid sequences are sequences 1 to 5.

[0101] Meanwhile, most diagnostic methods using nucleic acids utilize a nucleic acid amplification reaction that amplifies target nucleic acid molecules. A representative example is the polymerase chain reaction (PCR) among nucleic acid amplification reactions, which performs repeated cycles of denaturation of double-stranded DNA, annealing of oligonucleotide primers to a DNA template, and primer extension by DNA polymerase (Mullis et al., U.S. Patent Nos. 4,683,195, 4,683,202, and 4,800,159; Saiki et al., Science 230:1350-1354 (1985)). Other methods for amplifying nucleic acids have been proposed, including Ligase Chain Reaction (LCR), Strand Displacement Amplification (SDA), Nucleic Acid Sequence-Based Amplification (NASBA), Transcription Mediated Amplification (TMA), Recombinase polymerase amplification (RPA), Loop-mediated isothermal amplification (LAMP), and Rolling-Circle Amplification (RCA).

[0102] In some embodiments, the amplification reaction for amplifying a signal indicating the presence of a target nucleic acid molecule may be performed in a manner in which the signal is amplified as the target nucleic acid molecule is amplified (e.g., real-time PCR method). Alternatively, in one embodiment, the amplification reaction may be performed in a manner in which only the signal indicating the presence of the target nucleic acid molecule is amplified without the target nucleic acid molecule being amplified (e.g., CPT method). In this way, the amplification reaction may be accompanied by a signal change, and therefore, the degree of progress of such amplification reaction can be evaluated by measuring the signal change. As such a signal providing means, a signal generating composition containing the label itself or an oligonucleotide to which the label is linked can be used. Various methods for generating a signal indicating the presence of a target analyte using a signal generating composition are known (e.g., TaqMan™ probe method, molecular beacon method, etc.).

[0103] Among PCR-based technologies, real-time PCR is a technology for detecting target nucleic acids in real time. To detect a specific target nucleic acid, a signal generation means is used that emits a detectable fluorescent signal proportional to the amount of target nucleic acid during the PCR reaction. A fluorescent signal proportional to the amount of target nucleic acid is detected at each measurement point (cycle) through real-time PCR, thereby obtaining a dataset containing each measurement point and the signal value at the measurement point. From the dataset, an amplification curve or amplification profile curve indicating the intensity of the fluorescent signal detected versus the measurement point is obtained.

[0104] The term "signal value" means a numerical value of the level of a signal (e.g., signal intensity) actually measured in a cycle of an amplification reaction, or a modified value thereof, according to a certain scale. The modified value may include a mathematically processed signal value of the actually measured signal value (i.e., the signal value of the raw data set), and may include, for example, a logarithmic value or derivatives.

[0105] The term "dataset" refers to a collection of data points obtained from an amplification reaction. A data point represents a single coordinate value including a cycle and a signal value. For example, the dataset may be a collection of data points obtained directly through an amplification reaction performed in the presence of a signal-generating composition, or may be a modified dataset of such a dataset. The dataset may be a portion or all of a plurality of data points obtained by the amplification reaction, or modified data points thereof. The dataset may be plotted, thereby obtaining an amplification curve.

[0106] The term "oligonucleotide" refers to a linear oligomer of natural or modified monomers or linkages, comprising deoxyribonucleotides and ribonucleotides, capable of specifically hybridizing to a target nucleic acid sequence, and which may be naturally occurring or artificially synthesized. For example, a binding oligonucleotide may be used interchangeably with a primer or a probe.

[0107] The term "primer" refers to an oligonucleotide that can act as an initiator of synthesis under conditions that induce the synthesis of a primer extension product complementary to a nucleic acid strand (template), i.e., the presence of nucleotides and a polymerizing agent such as DNA polymerase, and suitable temperature and pH conditions. Furthermore, the term "probe" refers to a single-stranded nucleic acid molecule that contains a portion or portions complementary to a target nucleic acid sequence. Furthermore, the probe may include a label capable of generating a signal for target detection.

[0108] The above oligonucleotide may have a conventional primer and probe structure composed of a sequence that hybridizes to the nucleic acid sequence of a target nucleic acid molecule. Alternatively, the structure of the oligonucleotide may be modified to have a unique structure. For example, the oligonucleotide may have a structure of a scorpion primer, a molecular beacon probe, a sunrise primer, a hybeacon probe, a tagging probe, a DPO primer or probe (WO 2006 / 095981), and a PTO probe (see WO 2012 / 096523).

[0109] The set of oligonucleotides may refer to one or more oligonucleotides, and may be interpreted as referring to a sequence set of oligonucleotides including a forward sequence and a reverse sequence, depending on the embodiment. In one embodiment, the set of oligonucleotides may include a primer set of a forward primer and a reverse primer. For example, the forward primer may be a primer that anneals with an antisense strand, a non-coding strand, or a template strand, and may serve as a starting point for the coding or positive strand of the target analyte. In addition, the reverse primer may be a primer that anneals with the 3' end of the sense strand or the coding strand, and may serve as a starting point for synthesizing a complementary strand of the coding sequence or the non-coding sequence of the target analyte. Here, the aforementioned forward primer and reverse primer may refer to a pair of primers that determine a specific amplification region in the target nucleic acid sequence, and depending on the embodiment, may refer to individual primers that do not function as a pair.

[0110] The term "amplification reaction efficiency" refers to the efficiency level of an amplification reaction. In one embodiment, the efficiency level may refer to the degree to which an amplification reaction is performed efficiently or whether or not the amplification reaction can be performed efficiently. For example, a high amplification reaction efficiency may mean that a factor that inhibits the amplification reaction is relatively small during the amplification reaction process, or a factor that mediates or promotes the amplification reaction is relatively sufficient, so that amplification is performed well and accurate detection of the target nucleic acid molecule is possible. For example, when the efficiency level of the amplification reaction is relatively high, the PCR amplification product may increase according to an ideal growth curve in real-time PCR, a fluorescent signal may be emitted in proportion to the increase in the PCR amplification product, or the point in time (e.g., the amplification point) at which the change in the magnitude of the signal value in the amplification curve significantly increases may be relatively brought forward.

[0111] In one embodiment, the amplification reaction efficiency may refer to a measurable efficiency of the amplification reaction, such as efficiency in terms of input (e.g., concentration of target nucleic acid molecules, concentration of oligonucleotides, etc.) and / or efficiency in terms of output (e.g., increase in fluorescence signal). In some embodiments, the amplification reaction efficiency may be interpreted as encompassing amplification efficiency in the present technical field (e.g., PCR efficiency).

[0112] The term "amplifiable region" refers to a region that can be used to design oligonucleotides (e.g., primers and / or probes) from a plurality of target nucleic acid sequences. In one embodiment, the amplifiable region refers to a series of contiguous sequence regions that are predicted to have good results in terms of amplification efficiency and are therefore predicted to be usable for designing oligonucleotides. In one embodiment, the amplifiable region refers to a region that can be used as a probing site when detecting a target nucleic acid molecule, and that exhibits maximum amplification efficiency with a primer pair and / or probe designed based on the sequence within the region. In one embodiment, the amplifiable region is a region that allows maximum detection accuracy for a target nucleic acid molecule based on amplification efficiency.

[0113] FIG. 1 schematically illustrates a block diagram of a computer device (1000) according to one embodiment.

[0114] Referring to FIG. 1, a computer device (1000) may include a memory (100), a communication unit (200), and a processor (300). The configuration of the computer device (1000) illustrated in FIG. 1 is merely a simplified example. In one embodiment, the computer device (1000) may include other components for performing the computing environment of the computer device (1000), and only some of the disclosed components may constitute the computer device (1000).

[0115] A computer device (1000) may refer to a node that constitutes a system for implementing embodiments of the present disclosure. In one embodiment, the computer device (1000) may include any type of server and / or any type of user terminal. A server may include any type of computing system or computer device, such as a microprocessor, a mainframe computer, a digital processor, a portable device, and a device controller. A user terminal may include any type of terminal capable of interacting with a server or other computer device. User terminals may include, for example, mobile phones, smart phones, laptop computers, personal digital assistants (PDAs), slate PCs, tablet PCs, and ultrabooks.

[0116] The computer device (1000) can perform technical features according to embodiments described below. For example, the computer device (1000) can determine an amplifiable region available for designing an oligonucleotide based on feature data predicted to influence an amplification reaction in the region of interest.

[0117] The memory (100) can store at least one instruction that can be executed by the processor (300). In one embodiment, the memory (100) can store any type of information generated or determined by the processor (300) and any type of information received by the computer device (1000). In one embodiment, the memory (100) can be a storage medium that stores computer software that causes the processor (300) to perform operations according to embodiments of the present disclosure. Accordingly, the memory (100) can mean computer-readable media for storing software codes required to perform embodiments of the present disclosure, data that is the target of execution of the codes, and execution results of the codes.

[0118] In one embodiment, the memory (100) may refer to any type of storage medium. For example, the memory (100) may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., an SD or XD memory, etc.), a RAM (Random Access Memory), a SRAM (Static Random Access Memory), a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), a magnetic memory, a magnetic disk, and an optical disk. The computer device (1000) may also operate in relation to web storage that performs the storage function of the memory (100) on the Internet. The description of the memory described above is merely an example, and the memory (100) in the present disclosure is not limited thereto.

[0119] The communication unit (200) may be configured regardless of the communication mode, such as wired or wireless, and may be configured with various communication networks, such as a personal area network (PAN) and a wide area network (WAN). In addition, the communication unit (200) may operate based on the known World Wide Web, and may also utilize wireless transmission technologies used for short-distance communication, such as infrared (IrDA: Infrared Data Association) or Bluetooth. For example, the communication unit (200) may be responsible for transmitting and receiving data required to perform a technique according to an embodiment of the present disclosure.

[0120] The processor (300) can perform technical features according to embodiments to be described later by executing at least one instruction stored in the memory (100). In one embodiment, the processor (300) may be configured with at least one core and may include a processor for data analysis and / or processing, such as a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), or a tensor processing unit (TPU) of the computer device (1000).

[0121] A processor (300) according to one embodiment may perform operations for learning. In one embodiment, the processor (300) may perform calculations for learning, such as processing independent variables used in a prediction function in machine learning, calculating dependent variables using independent variables, calculating errors, and updating weights. In another embodiment, the processor (300) may perform calculations for learning a neural network, such as processing input data for learning, extracting features from input data, calculating errors, and updating weights of a neural network using backpropagation in deep learning. At least one of the CPU, GPGPU, and TPU of the processor (300) may process operations for learning. For example, the CPU and GPGPU may together process learning of a prediction function or network function, and data classification using the learned function. Furthermore, in one embodiment of the present disclosure, processors of a plurality of computer devices may be used together to process learning of a prediction function or network function, and data classification using the learned function. Additionally, a computer program executed in a computer device according to one embodiment of the present disclosure may be a CPU, GPGPU, or TPU executable program.

[0122] The processor (300) can perform technical features according to embodiments of the present invention by executing at least one instruction stored in the memory (100). In one embodiment, the technical features described throughout the specification can be implemented as one or more computer programs, and instructions and data for executing the same can be stored in the memory (100) and executed by the processor (300), but are not limited thereto. Various technical features performed by each component of the computer device (1000) will be described below with reference to FIGS. 2 to 17.

[0123]

[0124] 1. Technical characteristics for determining the amplifiable area

[0125] A computer device (1000) according to a first embodiment of the present invention can perform technical features for determining an amplifiable region. In one embodiment, the computer device (1000) can determine one or more regions of interest for which an amplification reaction efficiency is to be predicted, predict an amplification reaction efficiency for an amplification reaction in each of the one or more regions of interest using a prediction model, and determine an amplifiable region based on the amplification reaction efficiency of each of the one or more regions of interest.

[0126] FIG. 2 illustrates an exemplary flowchart of a computer device (1000) according to a first embodiment, wherein the steps of FIG. 2 determine an amplifiable area. In one embodiment, the steps of FIG. 2 may be implemented by a single entity, such as in a server. In another embodiment, some of the steps of FIG. 2 may be implemented by multiple entities, such as in a first server and others in a second server, or in a server and others in a user terminal.

[0127] Referring to FIG. 2, in step S210, the computer device (1000) can determine a region of interest from the alignment results of a plurality of nucleic acid sequences. Here, the region of interest represents a candidate region used to determine an amplifiable region, and may correspond to, for example, a region for which it is desired to confirm whether the region is applicable to an amplification reaction by predicting the efficiency of the amplification reaction. In one embodiment, the region of interest may be (a) a conservative region included in the alignment result, (b) a fragment region of a predetermined length divided from a conservative region, or (c) a fragment region of a predetermined length divided from the entire region of the plurality of nucleic acid sequences.

[0128] In step S220, the computer device (1000) may acquire feature data of a region of interest. Here, the feature data represents data predicted to influence an amplification reaction in the region of interest. Here, the term "influence" refers to the extent to which the efficiency level of the amplification reaction is influenced by the feature data of the region of interest or whether the efficiency level is influenced. Throughout the specification, the term "influence" may be broadly interpreted to include changes in the efficiency of the amplification reaction or inhibition of the amplification reaction that occur due to amplification being delayed or interrupted during the amplification reaction process.

[0129] In one embodiment, the feature data of the region of interest may include at least one selected from the group consisting of (a) a nucleic acid sequence of the region of interest; (b) thermodynamic data on the formation of an n-order structure in the nucleic acid sequence of the region of interest; (c) distance data between a given position in the nucleic acid sequence of the region of interest and a position of the n-order structure; (d) reaction conditions including a reaction medium used in an amplification reaction; (e) a type of oligonucleotide that binds to the nucleic acid sequence of the region of interest; (f) a GC content of the nucleic acid sequence of the region of interest; (g) a mutation score for the region of interest; and (h) a conservation score for the region of interest.

[0130] In step S230, the computer device (1000) may provide the acquired feature data of the region of interest as input data to a pre-trained prediction model, and in step S240, may obtain output data including the predicted amplification reaction efficiency for the amplification reaction in the region of interest from the prediction model. According to one embodiment, the prediction model may be trained to predict the amplification reaction efficiency for the amplification reaction in the target region using the feature data of the target region.

[0131] In step S250, the computer device (1000) may determine an amplifiable region for detecting a target nucleic acid molecule based on the amplification reaction efficiency of the region of interest. For example, if multiple regions of interest are determined, the computer device (1000) may determine an amplifiable region based on the region(s) of interest whose amplification reaction efficiency satisfies a predetermined criterion among the multiple regions of interest. The amplifiable region thus determined may be used in the design of an oligonucleotide for detecting the corresponding target nucleic acid molecule.

[0132] Various embodiments of the technical features presented above are described in more detail below.

[0133]

[0134] 1-1. Examples of determining areas of interest

[0135] A computer device (1000) can obtain alignment results of a plurality of nucleic acid sequences of a target nucleic acid molecule and determine one or more regions of interest from a predetermined region in the alignment results.

[0136] FIG. 3 is a diagram illustrating an alignment result (20) of a plurality of nucleic acid sequences (10) according to one embodiment. FIG. 4 illustrates an exemplary method by which a computer device (1000) according to one embodiment determines a region of interest (40) based on the alignment result (20).

[0137] Referring to FIG. 3, a computer device (1000) can obtain alignment results (20) of multiple nucleic acid sequences (10) of a target nucleic acid molecule.

[0138] In one embodiment, the computer device (1000) can collect a plurality of nucleic acid sequences (10) associated with a target nucleic acid molecule from a sequence database and align the collected plurality of nucleic acid sequences (10) to generate the alignment result (20).

[0139] A plurality of nucleic acid sequences (10) can be obtained using various sequence databases. For example, a plurality of target nucleic acid sequences (10) can be collected from publicly accessible databases such as GenBank, the EMBL (European Molecular Biology Laboratory) sequence database, and the DDBJ (DNA DataBank of Japan). For example, based on a search for a target nucleic acid molecule, nucleic acid sequences registered as the sequence of the target nucleic acid molecule can be collected from the database.

[0140] Alignment of nucleic acid sequences can be performed according to various methods (e.g., global alignment and local alignment) and algorithms known in the art. Various methods and algorithms for alignment are described in Smith and Waterman, Adv. Appl. Math. 2:482 (1981); Needleman and Wunsch, J. Mol. Bio. 48:443 (1970); Pearson and Lipman, Methods in Mol. Biol. 24: 307-31 (1988); Higgins and Sharp, Gene 73:237-44 (1988); Higgins and Sharp, CABIOS 5:151-3 (1989); Corpet et al., Nuc. Acids Res. 16:10881-90 (1988); Huang et al., Comp. Appl. BioSci. 8:155-65(1992) and Pearson et al., Meth. Mol. Biol. 24:307-31(1994). The NCBI Basic Local Alignment Search Tool (BLAST) (Altschul et al., J. Mol. Biol. 215:403-10(1990)) is available from NCBI (National Center for Biological Information) and can be used in conjunction with sequence analysis programs such as blastn, blasm, blastx, tblastn, and tblastx on the Internet. BLAST can be accessed at http: / www.ncbi.nlm.nih.gov / BLAST / . Instructions for comparing sequence similarity using this program can be found at http: / www.ncbi.nlm.nih.gov / BLAST / blast_help.html.

[0141] In another embodiment, the computer device (1000) may load the alignment result (20) described above from memory (100) or an external storage device, or may receive it from an external device connected to the computer device (1000) via a network.

[0142] The alignment result (20) may include alignment positions and may include a plurality of nucleic acid sequences (10) arranged according to the alignment positions. Here, the alignment positions represent positions where the corresponding nucleic acid sequences are arranged according to homology between the plurality of nucleic acid sequences (10), and each position may be expressed by a serial number. The alignment positions expressed by the serial numbers can be confirmed at the top of the alignment result (20).

[0143] The computer device (1000) can obtain a conserved region (30) from the alignment result (20). In one embodiment, the computer device (1000) can determine the conserved region (30) based on the conservatism in the alignment result (20). In another embodiment, the conserved region (30) may be included in the alignment result (20).

[0144] In one embodiment, the computer device (1000) can detect regions in the alignment result (20) where the conservation is above a predetermined level or the non-conservation is below a predetermined level, thereby determining one or more conservative regions (30). For example, the conservation (or non-conservation) can be measured for each alignment position in a plurality of aligned nucleic acid sequences (10), and positions where the conservation (or non-conservation) at the alignment positions satisfies a predetermined criterion can be detected, thereby determining that the positions are conservative regions (30).

[0145] The term "conservation" means that the ratio or number of identical specific bases to the total bases at alignment positions of a plurality of nucleic acid sequences (10) is equal to or greater than a predetermined value, or that the ratio or number of different specific bases to the total bases is equal to or less than a predetermined value, or a combination thereof. Specifically, conservation means that the ratio of identical specific bases to the total bases at the alignment position is 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more, or that the number of different specific bases to the total bases is 60 or less, 50 or less, 40 or less, 30 or less, or 20 or less, or a combination thereof. In addition, the term "conservative position" means an alignment position at which nucleotides of a plurality of aligned target nucleic acid sequences exhibit conservation, and the term "conservative base" refers to one base exhibiting conservation.

[0146] The term "non-conservation" refers to a case where conservation is not exhibited at alignment positions of a plurality of nucleic acid sequences (10), and refers to a case where the ratio or number of identical specific bases to the total bases at the alignment positions of the plurality of nucleic acid sequences (10) is less than a predetermined value, or a case where the ratio or number of different specific bases to the total bases is greater than a predetermined value, or a case where it is a combination thereof. Specifically, non-conservation refers to a case where the ratio of identical specific bases to the total bases at the alignment positions is less than 99%, less than 98%, less than 97%, less than 96%, or less than 95%, or a case where the number of different specific bases to the total bases is more than 20, more than 30, more than 40, more than 50, or more than 60, or a case where it is a combination thereof. Additionally, the term “non-conservative position” means an alignment position where nucleotides of a plurality of aligned target nucleic acid sequences exhibit non-conservation, and the term “non-conservative base” means two or more bases exhibiting non-conservation.

[0147] For example, a starting position may be determined in a plurality of aligned nucleic acid sequences (10), and conservation may be measured for each position based on the starting position. In addition, if the conservation at that position is above a predetermined level (e.g., the number or ratio of non-conservative bases is less than a predetermined value), a process of measuring conservation while increasing the position and comparing it with the level may be performed, and the terminal position may be determined by repeating the process until the number of non-conservative positions becomes above a predetermined number (e.g., 5). The position of the conservative region (30) may be determined by the positions from the starting position to the terminal position.

[0148] In another embodiment, the computer device (1000) may perform an entropy-based conservation analysis on the alignment result (20) to determine one or more conservative regions (30). For example, an entropy value indicating the degree of uncertainty may be measured from a probability distribution of the similarity of base sequences by alignment position in a plurality of aligned nucleic acid sequences (10). For example, the Shannon Entropy algorithm may be used in the process of measuring the entropy value, and as the entropy value increases, base similarity may decrease due to greater uncertainty. The computer device (1000) may determine a conservative region (30) based on the alignment positions of nucleic acid sequences having entropy values ​​within a preset value range.

[0149] In addition, the computer device (1000) can group the consecutive alignment positions within the alignment result (20) into multiple groups through cluster analysis on the measured entropy values, and a conservative region (30) defined by the alignment positions corresponding to each group can be determined based on the grouping result. For example, an entropy value is measured for each alignment position in each of 50 nucleic acid sequences aligned according to alignment positions 1 to 7000, and grouping of positional division regions can be performed according to about 4 entropy value ranges through clustering. In addition, one or more conservative regions (30) can be determined by merging the positional division regions of groups corresponding to a predetermined value range among the groups.

[0150] Referring to FIG. 4, the computer device (1000) can determine one or more regions of interest (40) from a conservative region (30). For example, when the length of the conservative region (30) is relatively short, one region of interest (40) having a predetermined length can be determined from the conservative region (30), and when the length of the conservative region (30) is relatively long, multiple regions of interest (40) can be determined.

[0151] In one embodiment, the computer device (1000) can determine a plurality of fragmented regions (see identification numbers 40a to 40d) of a predetermined length within the conservative region (30). In one embodiment, the computer device (1000) can determine a plurality of fragmented regions of a predetermined length (51) within the conservative region (30) that are sequentially fragmented or fragmented so that some of the fragmented regions overlap each other, and determine each of the plurality of fragmented regions as a region of interest (40). For example, the computer device (1000) can divide the fragment regions from the conservative region (30) of alignment positions 1 to 500 into regions having a length of 100 bp and overlapping each other by a length of 30 bp, thereby determining a first region of interest (40a) of alignment positions 1 to 100, a second region of interest (40b) of alignment positions 71 to 170, a third region of interest (40c) of alignment positions 141 to 240, and a fourth region of interest (40d) of alignment positions 211 to 310. As another example, the computer device (1000) can determine a first region of interest (40a) from alignment positions 1 to 100, a second region of interest (40b) from alignment positions 101 to 200, a third region of interest (40c) from alignment positions 201 to 300, and a fourth region of interest (40d) from alignment positions 301 to 400 by dividing the fragment regions of 150 bp in length that are sequentially fragmented from the conservative region (30) from alignment positions 1 to 500.

[0152] In one embodiment, the predetermined length (51) may be 50 bp or more and 150 bp or less. For example, the length of the region of interest (40) may be 80 bp, 90 bp, 100 bp, 110 bp, or 120 bp, which is mainly a level corresponding to the amplicon level. In another embodiment, the predetermined length (51) may be 100 bp or more and 300 bp or less, for example, 100 bp, 200 bp, or 300 bp. Here, the amplicon refers to a product generated by amplification in each cycle during an amplification reaction for a nucleic acid sequence, and, for example, a length corresponding to the amplicon level may be set for each target nucleic acid molecule or for each design.

[0153] In one embodiment, the above-described overlapping length (52) may be greater than or equal to 10 bp and less than or equal to 50 bp. For example, the overlapping length of each of the plurality of regions of interest (40) segmented from the conservative region (30) may be 30 bp, 40 bp, 50 bp, 60 bp, or 70 bp, which is primarily a level corresponding to the primer level. In another embodiment, the overlapping length (52) may be greater than or equal to 10 bp and less than or equal to 100 bp, for example, 10 bp, 25 bp, 50 bp, 75 bp, or 100 bp.

[0154] In one embodiment, the predetermined length (51) and / or the mutually overlapping length (52) described above may be updated based on at least one of the length of the conservative region (30) and the conservativeness in the conservative region (30). For example, when the length of the conservative region (30) is less than a set value, the predetermined length (51) may be updated to be less than a preset initial value (e.g., 100 bp). For another example, when the conservativeness of the conservative region (30) measured above is greater than a preset reference size, the mutually overlapping length (52) may be updated to be less than a preset initial value (e.g., 30 bp).

[0155] Meanwhile, the conservative region (30) illustrated in FIG. 4 may be replaced with other regions described herein. For example, the aforementioned fragment regions may be determined from the entire region included in the alignment result (20), and each fragment region may be segmented to have a starting position randomly selected within the entire region.

[0156]

[0157] 1-2. Examples of Obtaining Feature Data

[0158] The computer device (1000) can acquire feature data predicted to influence an amplification reaction in a region of interest (40). In one embodiment, the feature data represents an independent variable expected or anticipated to influence the result or variable of the amplification reaction. Such expectation or prediction can be broadly interpreted to encompass cases where the amplification reaction may or may not be influenced by differences in environmental factors, such as the presence or absence of a target nucleic acid molecule or reaction conditions, depending on the amplification reaction.

[0159] The computer device (1000) can obtain the feature data using information of the region of interest (40). In one embodiment, when the region of interest (40) is determined, the computer device (1000) can obtain information of the region of interest (40), generate the feature data using some of the information of the region of interest (40), and obtain the feature data from the remainder of the information of the region of interest (40).

[0160] Information of a region of interest (40) according to one embodiment may include at least one selected from the group consisting of a nucleic acid sequence of the region of interest (40), a length of the region of interest (40), a biological category of an organism having a target nucleic acid molecule, a nucleic acid type of the target nucleic acid molecule, a sequence function of the target nucleic acid molecule, and reaction conditions for an amplification reaction.

[0161] The nucleic acid sequence of the region of interest (40) refers to one or more nucleic acid sequences representing the nucleic acid sequences located in the region of interest (40) among a plurality of nucleic acid sequences (10). In one embodiment, the nucleic acid sequence of the region of interest (40) may include at least one selected from the group consisting of the first to fourth sequences below.

[0162] The first sequence may be determined from a nucleic acid sequence that satisfies a predetermined length condition among a plurality of nucleic acid sequences (10). The predetermined length condition may be, for example, a condition that is the longest or longer than a predetermined reference length (e.g., 10,000 bp). For example, the first sequence may include a sequence located in a region of interest (40) among the longest nucleic acid sequences among the plurality of nucleic acid sequences (10).

[0163] The second sequence may be determined from the unique genome sequence of the organism having the target nucleic acid molecule. Here, the unique genome sequence may refer to a sequence initially identified for the organism having the target nucleic acid molecule, or a sequence listed as a representative sequence in a sequence database. This representative sequence may be determined based on sequence accuracy and assembly quality for the species or individual. For example, the second sequence may include a sequence located in the region of interest (40) among the base sequences listed as representative sequences of the species or individual in a public database such as NCBI.

[0164] The third sequence may be determined based on the grouping results for multiple nucleic acid sequences (10). For example, multiple nucleic acid sequences (10) may be grouped based on homology, and a representative sequence may be selected from each group. The third sequence may include the representative sequences from each of these groups.

[0165] The fourth sequence may be determined from a nucleic acid sequence that contains the most conservative bases (or a predetermined ratio or more) or the least non-conservative bases (or a predetermined ratio or less) among the plurality of nucleic acid sequences (10). For example, the number of non-conservative bases at each position within the region of interest (40) is counted for each of the plurality of nucleic acid sequences (10), and the fourth sequence may include a sequence with the smallest sum of the counted non-conservative bases.

[0166] The length of the region of interest (40) can be processed based on the alignment positions corresponding to the region of interest (40) or the number of bases (or base pairs) included in the nucleic acid sequence of the region of interest (40).

[0167] The biological category of an organism having a target nucleic acid molecule represents a biological category to which the organism belongs among a plurality of preset biological categories. In one embodiment, the biological category may be a category located at any one hierarchical level among a plurality of hierarchical levels constituting a biological classification system. For example, the biological classification system may include a classification system expressed as Species, Genus, Family, Order, Class, Phylum, Kingdom, and Domain, and may have a hierarchical structure in which an upper level encompasses a lower level. For example, the biological category of the organism may be at the species level in the biological classification system. In addition, the biological category of the organism may be expressed as a taxonomic name and / or taxonomic identifier (Taxonomic ID) of the organism in the classification system.

[0168] The nucleic acid type of a target nucleic acid molecule indicates the type of nucleic acid contained in the nucleic acid molecule, and can be classified as, for example, DNA or RNA.

[0169] The sequence function of a target nucleic acid molecule indicates the presence or absence of a function of the sequence or the sequence type for the type of function. For example, the sequence function can be classified as a coding gene or a non-coding gene, a gene or a non-gene, or an intron or an exon depending on whether it encodes a relatively significant function. In another example, the sequence function can be classified as a promoter, enhancer, terminator, or transcription factor depending on the type of function, or can be classified as a translated sequence or an untranslated sequence depending on whether it can be translated. These sequence functions are not limited to the examples described above, and various known terms or concepts used in the present technical field to distinguish the presence or absence of a function or type of a sequence can be applied.

[0170] Reaction conditions for an amplification reaction can broadly refer to various pieces of information required during the nucleic acid amplification reaction, such as the reaction environment of the nucleic acid amplification reaction or the conditions for substances added to create the reaction environment. In one embodiment, the reaction conditions can include at least one of a reaction medium, nucleic acid sequence conditions, oligonucleotide conditions, environmental conditions, and other conditions (e.g., temperature, pressure, time) used in the nucleic acid amplification reaction.

[0171] Here, the reaction medium is a material surrounding the reaction environment, and in one embodiment, may include materials that are placed in the reaction vessel to create a reaction environment so that one or more of a plurality of steps (e.g., denaturation step, annealing step, extension step) for the nucleic acid amplification reaction can proceed. In one embodiment, the reaction medium may be one or more materials selected from the group consisting of a pH-related material that affects pH (e.g., tris buffer, EDTA (Ethylene-Diamine-Tetraacetic Acid)), an ionic strength-related material that affects ionic strength (e.g., Mg2+, K+, Na+, NH4+, Cl- as an ionic material), an enzyme used for transfer or linkage of nucleic acids in the nucleic acid amplification reaction (e.g., nuclease, polymerase, ligase, modifying enzyme), and an enzyme stabilization-related material for enzyme stabilization (e.g., sugar). In addition, the nucleic acid sequence conditions may include conditions for at least one of the amount (e.g., concentration) of the nucleic acid sequence in the sample accommodated in the reaction vessel, the type (e.g., species) of the organism (e.g., host) having the nucleic acid sequence, and the type (e.g., species) of the organism (e.g., host) that provided the sample. In addition, the oligonucleotide conditions may include conditions for at least one of the composition of one or more oligonucleotide sets including one or more oligonucleotide sequences (e.g., forward primer sequence, reverse primer sequence), and the amount (e.g., concentration) of the oligonucleotides. In addition, the environmental conditions broadly mean conditions for the surrounding environment such as equipment or experimental space that at least partially affect the nucleic acid amplification reaction, and may include, for example, the type of equipment (e.g., reaction vessel, plate, nucleic acid extraction device, amplification device, etc.) used in the process of the nucleic acid amplification reaction.Additionally, other conditions may be temperature, pressure, and time conditions provided to the reaction vessel for the progress of one or more of the multiple steps for a nucleic acid amplification reaction. The above-described examples are exemplary, and various types of materials, conditions, and information known in the art may be utilized. Furthermore, reaction conditions according to one embodiment may be stored and managed based on an identifier (e.g., the name of the enzyme master mix, an identification number, etc.) for identifying a predetermined combination of the types or values ​​of the above-described components.

[0172] Referring to the above-described embodiments, the computer device (1000) may determine a nucleic acid sequence of a region of interest (40), for example, based on sequences at the position of the region of interest (40) among a plurality of aligned nucleic acid sequences (10), and process the length of the region of interest (40) according to the alignment positions of the determined region of interest (40). In addition, metadata regarding a target nucleic acid molecule is pre-stored in the memory (100), and the computer device (1000) may obtain identifiers of each of a biological category of the organism (e.g., SARS-CoV-2), a nucleic acid type of the target nucleic acid molecule (e.g., RNA), a sequence function (e.g., coding gene), and a reaction condition (e.g., enzyme master mix) from the stored memory (100).

[0173] Using the information of the above-described region of interest (40), the computer device (1000) can obtain the above-described feature data. Hereinafter, embodiments regarding features included in the feature data and methods for obtaining each feature value will be described.

[0174] The above feature data may include thermodynamic data on the formation of an n-order structure in a nucleic acid sequence of a region of interest (40). Here, the n-order structure (n is an integer greater than or equal to 2) may be at least one of a secondary structure and a quaternary structure. In one embodiment, n is 2. In another embodiment, n is 3. In yet another embodiment, n is 4.

[0175] In one embodiment, the n-th structure may be a secondary structure representing an interaction between two or more bases included in a nucleic acid sequence. In one embodiment, the secondary structure may include at least one of a hairpin loop, an internal loop, a bulge loop, multi-loops, a G-quadruplex, and a combination thereof. As examples of such secondary structures, FIG. 5(a) illustrates a hairpin loop, FIG. 5(b) illustrates an internal loop, FIG. 5(c) illustrates a bulge loop, and FIG. 5(d) illustrates a multi-loops. In one embodiment, a G-quadruplex may be formed in a nucleic acid sequence having a relatively large amount of G among the bases, and may include, for example, G-tetrads formed in a helix and formed from one, two, or four strands.

[0176] According to another embodiment, the n-th structure may be a tertiary structure having two or more of the above-described interactions, and may include, for example, a case where the n-th structure is folded by hydrogen bonding in the presence of a secondary structure. In one embodiment, the tertiary structure may include at least one of pseudoknots, a kissing hairpin, and a hairpin-bulge contact. Hereinafter, throughout the specification, a hairpin loop of a secondary structure is mainly exemplified and described as an embodiment of the n-th structure, but the n-th structure of the present disclosure is not limited thereto. According to an embodiment, the n-th structure may apply various types of structures that are known in advance in addition to the above-described embodiments, and may include two or more of a secondary structure to a quaternary structure, or may include a quintuple or more structure, but is not limited to any one of them.

[0177] The thermodynamic data for the formation of the above n-dimensional structure refers to thermodynamic analysis data that directly or indirectly indicates whether or not an n-dimensional structure is formed in the corresponding nucleic acid sequence or the possibility of formation. For example, the thermodynamic data for the n-dimensional structure may include thermodynamic stability when a predetermined secondary structure (e.g., a hairpin loop) is formed by interaction between bases in the corresponding nucleic acid sequence under an environment in which a predetermined heat is applied, entropy or energy of the thermodynamic system when the secondary structure is formed, etc. In one embodiment, the thermodynamic data may be data for a thermodynamic property, which is a thermodynamic parameter used to predict the n-dimensional structure of a nucleic acid sequence based on the thermodynamic principle that the folding phenomenon for forming an n-dimensional structure in a nucleic acid sequence occurs in a low and stable energy state. These thermodynamic properties can be understood as encompassing various properties such as Gibbs free energy, Gibbs free entropy, internal energy, enthalpy, and entropy.

[0178] In one embodiment, the thermodynamic data may be expressed in the form of a change in thermodynamic free energy. For example, thermodynamic free energy may include, but is not limited to, Gibbs free energy, and may also include various other energies such as internal energy, Helmholtz free energy, and specific internal entropy. As an example, the thermodynamic data may include a change in Gibbs free energy (dG) when a hairpin loop is formed in a nucleic acid sequence under certain conditions (e.g., temperature, pressure). In the following, thermodynamic data is mainly exemplified as a change in Gibbs free energy, but is not limited thereto.

[0179] FIG. 6 is a diagram illustrating thermodynamic data for the formation of an n-th order structure in a nucleic acid sequence of a region of interest (40) according to one embodiment.

[0180] Referring to FIG. 6, the thermodynamic data may include thermodynamic data for n-th order structures existing within a range of a nucleic acid sequence. Such thermodynamic data may include, for example, a change in Gibbs free energy (dG) for a secondary structure (see identification number 60) (e.g., a hairpin) formed in the nucleic acid sequence within the entire range of the nucleic acid sequence of the region of interest (40). For example, a smaller change in Gibbs free energy (dG) (e.g., a larger negative value) indicates greater thermodynamic stability of the secondary structure, which may indicate a greater likelihood of formation of the secondary structure during an amplification reaction.

[0181] The computer device (1000) can use the nucleic acid sequence of the region of interest (40) and the corresponding reaction conditions to produce thermodynamic data on the formation of an n-order structure in the nucleic acid sequence of the region of interest (40). For example, the computer device (1000) can apply the nucleic acid sequence of the region of interest (40), the nucleic acid type (e.g., DNA, RNA) of the nucleic acid sequence, and the temperature (e.g., 60 degrees) of the nucleic acid sequence to a pre-stored thermodynamic property production algorithm (e.g., a free energy production algorithm), to produce thermodynamic data (e.g., a change in Gibbs free energy) on the formation of an n-order structure (e.g., a hairpin) (see identification number 60) within the entire range of the nucleic acid sequence at the corresponding temperature.

[0182] In one embodiment, the computer device (1000) can generate thermodynamic data regarding the formation of an arbitrary n-th structure as the thermodynamic data. For example, the computer device (1000) can obtain the change in free energy in the most stable structure of the nucleic acid sequence under the reaction conditions by applying a nucleic acid sequence and the corresponding reaction conditions to a free energy calculation algorithm. For example, the stable structure can be divided into a case where an n-th structure is formed and a case where an n-th structure is not formed, and in the case where an n-th structure is not formed, the change in free energy can be processed as 0, and information regarding the type of the n-th structure can be obtained.

[0183] In another embodiment, the computer device (1000) can generate thermodynamic data regarding the formation of a specific n-th order structure as the thermodynamic data. For example, the computer device (1000) can apply a nucleic acid sequence and corresponding reaction conditions to a free energy calculation algorithm for the formation of a hairpin loop, thereby obtaining the change in free energy when a hairpin loop is formed in the nucleic acid sequence under the corresponding reaction conditions.

[0184] In another embodiment, the computer device (1000) can generate thermodynamic data for the formation of each of a plurality of n-order structures as the thermodynamic data. For example, the computer device (1000) can apply a nucleic acid sequence and corresponding reaction conditions to a free energy calculation algorithm for the formation of each of a preset hairpin loop, an internal loop, a bulge loop, and multi-loops, so as to obtain, under the corresponding reaction conditions, each of a free energy change when a hairpin loop is formed in the nucleic acid sequence, a free energy change when an internal loop is formed, a free energy change when a bulge loop is formed, a free energy change when multi-loops are formed, and a free energy change when a G-quadruplex is formed.

[0185] Various known mathematical formulas can be used in the above-described thermodynamic property calculation algorithm. In one embodiment, the above-described free energy calculation algorithm may be applied with a mathematical formula for calculating the MFE (Minimum Free Energy) representing the thermodynamic energy of the most stable state among all structures that a molecule can have (e.g., the NUPACK algorithm), at least one mathematical formula among the enthalpy calculation mathematical formula and the entropy calculation mathematical formula, and for example, the NUPACK algorithm, the dynamic programming algorithm, and the energy binding data, etc. can be used. For example, by applying data on the base sequence in which the types of bases included in the nucleic acid sequence are listed in order, the temperature, and the reaction conditions for the ionic substance to the MFE calculation algorithm, the change in Gibbs free energy (dG) for the secondary structure formed in the nucleic acid sequence under the conditions of the temperature and the concentration of the ionic substance can be calculated.

[0186] In one embodiment, an amplicon region (41) and / or an oligo biding region (42) may be determined from a region of interest (40). For example, the region of interest (40) may be determined to have a length corresponding to a target amplicon, the amplicon region (41) may correspond to the region of interest (40), and the oligo biding region (42) may include a first oligo biding region (42a) and a second oligo biding region (42b) determined to have a predetermined length range at both ends based on the amplicon region (41). In another example, the region of interest (40) may be determined to have a length greater than a target amplicon, in which case the amplicon region (41) is a part of the region of interest (40).

[0187] In one embodiment, the computer device (1000) can generate thermodynamic data for each of the amplicon region (41) and the oligo bidding region (42). Specifically, the computer device (1000) can generate first thermodynamic data on the formation of an n-order structure in the amplicon region (41) using the nucleic acid sequence of the amplicon region (41) and the corresponding reaction conditions, and can generate second thermodynamic data on the formation of an n-order structure in the oligo bidding region (42) using the nucleic acid sequence of the oligo bidding region (42) and the corresponding reaction conditions. For example, the reaction conditions of the former and the latter may be at least partially different. For example, the reaction temperature used to generate the first thermodynamic data may be set to 60 degrees based on the characteristics of the amplicon produced as a product during the amplification reaction. Additionally, the reaction temperature used to generate the second thermodynamic data may be set to 60 degrees based on the Tm (melting temperature) characteristic of the primer sequence.

[0188] The above feature data may include distance data between a predetermined position in the nucleic acid sequence of the region of interest (40) and the position of the n-th structure. Here, the distance data may be, for example, a number of bases (pairs) counted based on a one-dimensional sequence (e.g., bp) or a value quantifying the linear distance between position coordinates measured based on a multi-dimensional structure. For example, the distance data may be calculated by counting the distance in base units between the position where the base sequence starts based on the 5' direction in the nucleic acid sequence of the region of interest (40) and the position where the hairpin structure starts in the nucleic acid sequence (e.g., the starting position of the stem in a stem-loop structure). In one embodiment, the position of the n-th structure may be obtained in the process of calculating the thermodynamic data, and, for example, the position of the n-th structure may be analyzed together in the process of calculating the Gibbs free energy change for the formation of the n-th structure in the nucleic acid sequence using the MFE calculation mathematical formula.

[0189] The above characteristic data may include the GC content of the nucleic acid sequence of the region of interest (40). Here, the GC content refers to the number or ratio of base pairs containing G or C among the bases of the sequence, and a higher GC content may refer to a higher binding affinity. The computer device (1000) may calculate the GC content by counting the ratio of the number of G and C among A, G, C, and T in the base sequence of the region of interest (40).

[0190] The above characteristic data may include a mutation score for a region of interest (40). Here, the mutation score for a region of interest (40) is a numerical value indicating whether the region of interest (40) can tolerate subsequent mutations that occur in multiple nucleic acid sequences (10), and for example, may indicate whether the region is a region in which primers can be robustly hybridized and an amplification reaction can proceed well even if there is a certain degree of mutation in the sequences.

[0191] In one embodiment, the mutation score for the region of interest (40) can be calculated based on at least one selected from the group consisting of (a) a mutation position, (b) a mutation type, (c) the number of non-conserved bases and / or mutation types at the mutation position, and (d) the number of mutation positions in the region of interest (40) of the plurality of nucleic acid sequences (10). Here, the mutation position indicates the position of a non-conserved base in the region of interest (40) of the plurality of nucleic acid sequences (10), and the mutation type indicates the type of non-conserved base. For example, in FIG. 3, assuming that the positions of the region of interest (40) are 1 to 21, among the alignment positions, the mutation position having a mutation within the sequences is 1 of 12, the mutation type at the mutation position 12 is A, the number of non-conserved bases at the mutation position is 1, and the number of mutation types at the mutation position is 1. The computer device (1000) can calculate a higher mutation score, for example, when the number of mutation positions in the region of interest (40) is smaller, the number of non-conservative bases and the number of mutation types at each mutation position are smaller, and the distance between the mutation position and the start position of the oligo binding region (42) in the 5' direction is greater. In addition, a relatively high weight can be given to whether the mutation position is included in and / or close to the oligo binding region (42), and for example, a high mutation score can be calculated when it is not included.

[0192] The above feature data may include a conservation score for the region of interest (40). Here, the conservation score for the region of interest (40) represents a numerical value of the conservation of the region of interest (40) in a plurality of nucleic acid sequences (10). In one embodiment, the conservation score may be calculated based on the entropy measurement described above. The computer device (1000) may measure an entropy value for the position-specific similarity of base sequences in the region of interest (40) using, for example, the Shannon Entropy algorithm described above, calculate a position-specific score inversely proportional to each entropy value, and perform a statistical operation on the position-specific score to calculate the conservation score. In one embodiment, the Shannon Entropy-based conservation score may also reflect a numerical value indicating whether the region is robust to mutations.

[0193] The above characteristic data may further include at least one of a nucleic acid sequence of a region of interest (40), corresponding reaction conditions, and a type of oligonucleotide that binds to the nucleic acid sequence of the region of interest (40). In one embodiment, the type of oligonucleotide may be classified according to the operating method or composition of the oligonucleotide. For example, when the oligonucleotide is a primer, the type of oligonucleotide may include a dual priming oligonucleotide, an inosine primer, etc. Or, when the oligonucleotide is a probe, the type of oligonucleotide may include a hydrolysis probe (e.g., a TaqMan probe), a hybridization probe (e.g., a molecular beacon), etc. In another embodiment, the type of oligonucleotide may be classified according to the degree to which the oligonucleotide specifically binds to a nucleic acid sequence (e.g., specificity). In another embodiment, the type of oligonucleotide may be classified according to the function performed based on the binding (e.g., primer, probe).

[0194] In one embodiment, the feature data may further include information on the Tm (melting temperature) of the oligonucleotide sequence. Here, Tm is a temperature indicating the binding affinity between a nucleic acid sequence and an oligonucleotide sequence, and refers to the dynamic equilibrium temperature when the double helix structure of DNA (or RNA and a strand complementary to the RNA) in the nucleic acid sequence dissociates into single strands, and the double helix structure and the single strand each exist at 50%. The computer device (1000) may, for example, apply a nucleic acid sequence and an oligonucleotide sequence (e.g., a forward primer sequence, a reverse primer sequence) to a pre-stored Tm prediction algorithm to produce a predicted value for the Tm of the oligonucleotide based on the oligonucleotide sequence and the binding state with the nucleic acid sequence. This Tm prediction algorithm is an algorithm for predicting the Tm actually measured under the corresponding reaction conditions, and may provide a modeled Tm value by applying the corresponding reaction conditions. For example, the Biopython algorithm may be utilized.

[0195] The above characteristic data can be divided into data produced based on calculations by a computer device (1000) and data obtained from information of a region of interest (40) or memory (100). For example, the thermodynamic data, distance data, reaction conditions, mutation scores, and conservation scores may be the former, and the nucleic acid sequence of the region of interest (40), the type of oligonucleotide, and the reaction conditions may be the latter.

[0196] Meanwhile, the technical feature of the computer device (1000) for acquiring the feature data may be implemented in the form of a preprocessing unit programmed to extract the feature data, or may be implemented by a model trained to perform the corresponding function. In one embodiment, the process of acquiring the feature data may be performed by a preprocessing unit that calculates the value of each of the plurality of features included in the feature data using a plurality of previously stored mathematical formulas. In another embodiment, the process of acquiring the feature data may be performed by a model trained to output the value of each of the plurality of features included in the feature data when information on the region of interest (40) is input. Such a model may be implemented, for example, as a model included in a prediction model described below and operated integrally by the prediction model, or, for another example, may be implemented as a model distinct from the prediction model and operated independently of the prediction model.

[0197] Meanwhile, the above-described feature data may be preprocessed into a form suitable for calculation in a predictive model. This preprocessing process may include, for example, a process of converting the representation of the above-described data into a numerical or vector form that can be calculated in a machine learning model or a deep learning model. For example, when a nucleic acid sequence is used as input data in a predictive model, a preprocessing process including vectorization of the nucleic acid sequence (e.g., embedding, encoding, tokenization, etc.) may be performed so that it can be calculated in the predictive model. In this preprocessing process, various conventional preprocessing techniques for preprocessing text or images (e.g., text transformation, label transformation, vector transformation, etc.) that are known in the art may be used together. In some embodiments, the process of acquiring the above-described feature data may be understood as part of a preprocessing process for providing input data to a predictive model.

[0198]

[0199] 1-3. Examples of obtaining amplification reaction efficiency using a prediction model

[0200] A predictive model in this specification may refer to any form of computer program that operates based on at least one of one or more functions, network functions, artificial neural networks, and neural networks. In some embodiments, the terms model, function, network function, neural network, and neural network may be used interchangeably. A function may represent a correlation between one or more independent variables and one or more dependent variables, and may define an operation method between them. A neural network is a network in which one or more nodes are interconnected through one or more links to form an input node and an output node relationship within the network. The characteristics of the neural network may be determined based on the number of nodes and links within the neural network, the correlation between the nodes and links, and the weight value assigned to each link. In one embodiment, the neural network may be any of various known forms of neural networks.

[0201] A predictive model according to one embodiment may include at least one of a regression model and a classification model.

[0202] A regression analysis model according to one embodiment may include a simple regression analysis model and / or a multiple regression analysis model. In one embodiment, the regression analysis model may include at least one of a linear regression analysis model, a non-linear regression analysis model, a decision tree model, and an ensemble model. In one embodiment, the linear regression analysis model may include at least one of Robust, Lasso, and Ridge. In addition, the linear regression analysis model may include at least one of a normal equation, a least squares method, and a gradient descent method, and the gradient descent method may include stochastic gradient descent (SGD) and batch gradient descent (BGD). In addition, the non-linear regression analysis model may include at least one of a convolutional neural network (CNN), a recurrent neural network (RNN), and a deep neural network (DNN). Additionally, the ensemble model may include at least one of bagging to reconstruct data for model diversification, random forest to reconstruct data and variables for model diversification, boosting to learn by weighting data with large errors in previous learning, and stacking to use the output of the model as a new independent variable. For example, the boosting family includes GBM (gradient boosting algorithm), AdaBoost, XGBoost, and LightGBM.

[0203] The classification model according to one embodiment may be classified as a binary classification model or a multi-label classification model. The classification model according to one embodiment may include at least one of a logistic regression model, a linear discriminant analysis (LDA) model, a K-nearest neighbor (KNN) model, a naive Bayes model, a decision tree model, an ensemble model, and a support vector machine (SVM) model. For example, the logistic regression model includes softmax regression.

[0204] In one embodiment, the prediction model may include at least one selected from the group consisting of a Random Forest (RF), a Logistic Regressor (LR), a Gradient Boosting Classifier (GBC), a Decision Tree Classifier (DTC), a Gaussian Naive Bayes (GNB), a Support Vector Classifier (SVC), and a Neural Network (NN). However, the description of the prediction model described above is merely an example, and the prediction model in the present disclosure is not limited thereto, and may be various types of machine learning models or deep learning models known in the art.

[0205] According to one embodiment, the predictive model may be learned through at least one of supervised learning, unsupervised learning, semi-supervised learning, self-learning, and reinforcement learning. The learning may be a process in which a machine learning model or deep learning model applies knowledge required to perform a specific action to a predictive function or neural network.

[0206] According to one embodiment, a predictive model may be trained to predict an amplification response efficiency for an amplification response in a target region using feature data of the target region. The predictive model may be trained using multiple training datasets, and each of the multiple training datasets may include training input data including feature data of the target region for training and training answer data including an amplification response efficiency for the amplification response in the target region for training.

[0207] In one embodiment, a predictive model can be trained to minimize output errors. For example, a series of processes may be repeated, such as inputting each training dataset into a predictive model, calculating the error between the output of the predictive model for each training dataset and the expected output, and updating the operation of the predictive model to reduce the error, or a series of processes may be repeated, such as backpropagating the error from the output layer to the input layer to update the weights of each node. In supervised learning, labeled data may be used for each training dataset, and in unsupervised learning, unlabeled data may be used for each training data. In some embodiments, the amount of change in the updated weights may be determined by a learning rate. The calculation of the predictive model for the input data and the updating of the error may constitute a learning cycle (epoch), and the learning rate may be applied differently depending on the number of repetitions of the learning cycle. Embodiments of the learning process of such a predictive model will be described later.

[0208] The computer device (1000) can access the prediction model and exchange data with the prediction model. For example, the computer device (1000) can access the prediction model in a manner such as (a) a method in which a first server (or user terminal) receives and executes a prediction model learned by a second server (or server) from the server (or database), (b) a method in which the first server (or user terminal) communicates with a prediction model executed by the second server (or server), or (c) a method in which a prediction model previously stored in the memory (100) or another storage medium is loaded and executed. The above-described access is not limited to the above-described embodiments and may be implemented in various modified forms.

[0209] The computer device (1000) provides the above-described feature data as input data to a prediction model, and the computer device (1000) can obtain output data including the predicted amplification reaction efficiency for the amplification reaction in the region of interest (40) from the prediction model. For example, the computer device (1000) inputs the feature data of the region of interest (40) into the prediction model, and the prediction model can perform a process of predicting the amplification reaction efficiency for the amplification reaction in the region of interest (40) using the input feature data.

[0210] FIG. 7 is a conceptual diagram illustrating a process by which a prediction model (710) according to one embodiment predicts an amplification reaction efficiency for an amplification reaction in a region of interest (40).

[0211] Referring to FIG. 7, when input data including feature data (720) predicted to affect an amplification reaction in a region of interest (40) is input, the prediction model (710) can output output data including an amplification reaction efficiency (730) predicted for an amplification reaction in a region of interest (40) as a prediction result. As described above, the prediction model (710) may be pre-trained to output an amplification reaction efficiency (730) corresponding to a dependent variable when feature data (720) corresponding to an independent variable is input, and can output the amplification reaction efficiency (730) of the region of interest (40) from the input data based on the relationship between the independent variables and the dependent variables in the prediction function implemented through learning.

[0212] The amplification reaction efficiency (730) of the above-mentioned region of interest (40) may include numerical data on the efficiency level of the amplification reaction in the region of interest (40) and / or classification data on the high and low of the efficiency level.

[0213] In one embodiment, the above numerical data represents a relative value corresponding to the degree of efficiency of the amplification reaction for the nucleic acid sequence. For example, the prediction model (710) may be a model trained based on a regression analysis model, and may be trained to output a value relatively proportional to the efficiency level as numerical data.

[0214] For example, the relative value may mean a numerical representation of the signal pattern in a data set including the cycle-by-cycle signal value for the amplification reaction, or the degree to which the pattern matches a preset reference pattern. A large value means that the cycle-by-cycle signal value in the amplification reaction is predicted to exhibit a relatively ideal signal pattern. As another example, the relative value may mean a numerical representation of the degree to which the amplification point differs in two or more data sets obtained from two or more amplification reactions for the nucleic acid sequence. A large value means that the difference in the amplification point is relatively large, resulting in a relatively inefficient amplification reaction (this will be explained in detail in the section on the correct answer data for training the prediction model later).

[0215] In another embodiment, the above-described numerical data represents a probability value for the efficiency level of the amplification reaction. For example, the prediction model (710) may be a model trained based on a classification model, and may be trained to output a probability value corresponding to the first class, indicating a high efficiency level, as numerical data.

[0216] In one embodiment, the above classification data represents a class corresponding to the amplification reaction among a plurality of classes for high and low efficiency levels. For example, the prediction model (710) is a model learned based on a classification model, and may be learned to output a class whose probability value satisfies a predetermined criterion (e.g., the highest or equal to a predetermined value) among a plurality of classes (e.g., high amplification efficiency level: 2, medium: 1, low: 0), or to output a probability value for each of the classes.

[0217] In another embodiment, a plurality of value ranges for amplification efficiency levels are mapped to a plurality of classes, and the classification data may be a class mapped to a value range to which the numerical data of the amplification reaction efficiency (730) belongs among the plurality of classes. For example, a plurality of efficiency sections and a class corresponding to each section are preset in order of the efficiency levels of the amplification reaction from highest to lowest, and one of the plurality of classes may be output depending on the section to which the probability value of the amplification reaction efficiency (730) belongs.

[0218] Meanwhile, as described above, the nucleic acid sequence of the region of interest (40) according to one embodiment may be plural. In this case, feature data predicted to affect the amplification reaction in each sequence is generated using each of these plural sequences, and the amplification reaction efficiency predicted for the amplification reaction in each sequence is predicted from the prediction model (710) based on each feature data, and the amplification reaction efficiency (730) of the region of interest (40) may be determined according to the result of statistical calculation of each amplification reaction efficiency. For example, the nucleic acid sequence of the region of interest (40) includes the first to fourth sequences, and the first to fourth feature data for the first to fourth sequences may be generated using the first to fourth sequences, respectively. The computer device (1000) may input the first to fourth feature data into the prediction model (710), respectively, and obtain the first to fourth amplification reaction efficiencies of the first to fourth sequences, respectively, from the prediction model (710). The computer device (1000) determines the amplification reaction efficiency (730) based on a majority voting method when the first to fourth amplification reaction efficiencies are classification values, and determines the amplification reaction efficiency (730) of the region of interest (40) by calculating a representative value (e.g., average value, median value, minimum value, etc.) when the first to fourth amplification reaction efficiencies are probability values.

[0219]

[0220] 1-4. Examples of determining the amplifiable area

[0221] The computer device (1000) can determine an amplifiable region by using one or more regions of interest (40) among one or more regions of interest (40) whose amplification reaction efficiency (730) satisfies a preset first criterion. In one embodiment, the first criterion can include a condition for a value range of numerical data or a class of classification data.

[0222] FIG. 8 illustrates an exemplary manner in which a computer device (1000) determines an amplifiable region (70) according to one embodiment.

[0223] Referring to FIG. 8, it can be assumed that the first region of interest (40a) to the fourth region of interest (40d) are determined from the alignment result (20), and the amplification reaction efficiency (730) of each region of interest (40) is predicted.

[0224] In one embodiment, the computer device (1000) may determine each of one or more regions of interest (40) that satisfy a first criterion as an amplifiable region (70). For example, if the probability values ​​of the amplification reaction efficiencies (730) of the first region of interest (40a) to the fourth region of interest (40d) are 15%, 97%, 90%, and 95%, respectively, the computer device (1000) may determine that each of the second region of interest (40b) to the fourth region of interest (40d) that satisfy a value range corresponding to the first criterion (e.g., 90% or more) is an amplifiable region (730).

[0225] Alternatively, in another embodiment, the computer device (1000) may determine a merged region from one or more regions of interest (40) that meet a first criterion as an amplifiable region (70) if the merged region meets a preset second criterion. The second criterion may include, for example, (a) a criterion for the length of the merged region, and / or (b) a criterion for a statistical operation value of an amplification reaction efficiency (730) of each of the one or more regions of interest (40) that meet the first criterion. For example, if the classification values ​​of the amplification reaction efficiency (730) of each of the first region of interest (40a) to the fourth region of interest (40d) are 0, 1, 1, and 1 in order, the computer device (1000) may merge the second region of interest (40b) to the fourth region of interest (40d) having the classification value (e.g., 1) corresponding to the first criterion, and if the merged region (see identification number 70) satisfies the length range (e.g., 300 bp or more) corresponding to the second criterion, the merged region may be determined to be an amplifiable region (70).

[0226] In one embodiment, the computer device (1000) may obtain an amplification reaction efficiency (730) for an amplification reaction in a merged region from one or more regions of interest (40) that satisfy the first criterion described above using a prediction model (710), and if the amplification reaction efficiency (730) in the merged region satisfies the third criterion, the merged region may be determined as an amplifiable region (70). For example, the computer device (1000) may generate feature data including thermodynamic data for the formation of a secondary structure in the entire region by using a nucleic acid sequence corresponding to the entire region (see identification number 70) obtained by merging the second region of interest (40b) to the fourth region of interest (40d) having a classification value (e.g., 1) corresponding to the first criterion and corresponding reaction conditions. In addition, the computer device (1000) inputs the produced feature data into the prediction model (710), and if the amplification reaction efficiency (730) in the entire merged region output from the prediction model (710) has a classification value (e.g., 1) or a probability value (e.g., 70%) corresponding to the third criterion, the merged region can be determined as an amplifiable region (70). In other words, even if the efficiency condition is satisfied in each of the divided regions, if the efficiency condition is not satisfied in the entire merged region, the region can be determined as an amplifiable region (70).

[0227] In another embodiment, when an amplifiable region (70) is determined, the computer device (1000) may obtain an amplification reaction efficiency (730) for an amplification reaction in the amplifiable region (70) using a prediction model (710), and if the obtained amplification reaction efficiency (730) does not satisfy a third criterion, the determination of the amplifiable region (70) may be canceled.

[0228] Traditionally, the process of determining the region to be used in oligonucleotide design primarily focused on assessing target coverage based on conservation. However, this approach often failed to account for characteristics that impact amplification, particularly factors that inhibit amplification (e.g., secondary structure formation), resulting in poor amplification performance.

[0229] However, according to one embodiment of the present invention, the amplifiable region (70) can be determined by dividing the regions of interest (40) corresponding to the amplicon level from the conservative region (30), checking whether the amplification reaction efficiency (730) is at an appropriate level based on the above-described characteristic data for each, and merging the confirmed regions of interest (40). Accordingly, the amplifiable region (70) can be determined centering on sequences that are likely to actually undergo a good amplification reaction, so that when an oligonucleotide designed based on such amplifiable region (70) is used, there is an advantage in that the detection accuracy of the target nucleic acid molecule can be improved.

[0230] Meanwhile, there may be cases where the amplifiable region (70) is not determined for reasons such as the amplification reaction efficiency (730) of the region of interest (40) not meeting the first criterion. In one embodiment, the computer device (1000) may, if there is no determined amplifiable region (70) or the number is less than a predetermined number, re-determine the region of interest (40) and then perform the above processes of calculating the amplification reaction efficiency (730) based thereon again.

[0231] In one embodiment, if there is no region of interest (40) among the regions of interest (40) whose amplification reaction efficiency (730) satisfies the first criterion, the computer device (1000) may redetermine the regions of interest (40) so that the lengths of the regions of interest (40) are reduced. For example, after reducing a predetermined length (51) (e.g., 150 bp) by a unit length (e.g., 50 bp), the regions of interest (40) may be redetermined by applying the reduced predetermined length (51). In this redetermining process, the lengths of mutually overlapping lengths (52) may also increase. For example, the lengths of each region of interest (40) may be reduced but redetermined so that they overlap more.

[0232] In one embodiment, the computer device (1000) can re-acquire the amplification reaction efficiency (730) of the re-determined region of interest (40) by re-performing the step of acquiring the feature data or the step of acquiring the output data based on the re-determined region of interest (40). The computer device (1000) can determine the amplifiable region (70) by using one or more regions of interest (40) among the re-determined regions of interest (40) whose re-acquired amplification reaction efficiency (730) satisfies the first criterion. According to an embodiment, this process can be repeatedly performed by reducing the length of the region of interest (40) within a predetermined length range until the number of amplifiable regions (70) is equal to or greater than a predetermined reference number (e.g., 1, 3).

[0233] The process of re-determining the amplifiable region according to one embodiment may be implemented by first determining the regions of interest (40) to have a relatively large length, and when the amplifiable region (70) is determined, obtaining the amplifiable region (70) by reducing the length of the regions of interest (40). The process of re-determining the amplifiable region according to another embodiment may be implemented by first determining the regions of interest (40) to be continuously segmented, and when the amplifiable region (70) is not determined, obtaining the amplifiable region (70) by changing the regions of interest (40) to overlap each other.

[0234]

[0235] 1-5. Additional embodiments for determining the amplifiable area

[0236] A computer device (1000) according to one embodiment can determine an amplification reaction efficiency (730) of a region of interest (40) based on the amplification reaction efficiency in each of an amplicon region (41) and / or an oligo bidding region (42). This process can be understood with reference to the above-described embodiments, and redundant descriptions will be omitted.

[0237] Specifically, the computer device (1000) can determine an amplicon region (41) and an oligo bidding region (42) from a region of interest (40), and generate first feature data and second feature data that affect an amplification reaction in each of the amplicon region (41) and the oligo bidding region (42). The computer device (1000) can provide each of the first feature data and the second feature data as input data to a prediction model (710), and obtain a first amplification reaction efficiency and a second amplification reaction efficiency in each of the amplicon region (41) and the oligo bidding region (42) from the prediction model (710). In one embodiment, the second feature data may include 2-1 feature data regarding the first oligo biding region (42a) and 2-2 feature data regarding the second oligo biding region (42b), and the second amplification reaction efficiency may include 2-1 amplification reaction efficiency of the first oligo biding region (42a) and 2-2 amplification reaction efficiency of the second oligo biding region (42b).

[0238] The computer device (1000) can determine the amplification reaction efficiency (730) in the region of interest (40) based on the first amplification reaction efficiency and the second amplification reaction efficiency. For example, the computer device (1000) can calculate the average of the class-specific probability values ​​included in each of the first and second amplification reaction efficiencies and determine the class with the highest value and its probability value as the amplification reaction efficiency (730) of the region of interest (40). The average value may be an arithmetic mean result value or a weighted mean result value, and for example, may be a weighted mean result value in which a relatively higher weight is assigned to the second amplification reaction efficiency. As another example, the computer device (1000) can determine the amplification reaction efficiency (730) of the region of interest (40) to be the first class only when the first class (e.g., high efficiency level) is predicted for both the amplicon region (41) and the oligo bidding region (42), and can determine the second class (e.g., low efficiency level) in other cases.

[0239] A prediction model (710) according to one embodiment may include a first prediction model for predicting an amplification reaction efficiency in an amplicon sequence and a second prediction model for predicting an amplification reaction efficiency in an oligo bidding sequence. The first prediction model may be learned to predict an amplification reaction efficiency for an amplification reaction in an amplicon sequence using first training datasets including thermodynamic data for an n-th structure present in an amplicon sequence. In addition, the second prediction model may be learned to predict an amplification reaction efficiency for an amplification reaction in an oligo bidding sequence using second training datasets including thermodynamic data for an n-th structure present in an oligo bidding sequence determined from an amplicon sequence. For example, each of the first and second prediction models is a model learned using training datasets including at least partially different feature data for optimized prediction of an amplification reaction efficiency when a target nucleic acid sequence corresponds to an amplicon and when a target nucleic acid sequence corresponds to an oligo bidding site, respectively. The learning process of the first and second prediction models will also be explained later.

[0240] In the above-described embodiment, the computer device (1000) may provide the first and second feature data to the first and second prediction models, respectively, obtain the first and second amplification reaction efficiencies corresponding to the first and second feature data from the first and second prediction models, respectively, and determine the amplification reaction efficiency (730) of the region of interest (40) based on the first and second amplification reaction efficiencies.

[0241] In one embodiment, the first feature data provided to the first prediction model may include at least one selected from the group consisting of a nucleic acid sequence of the amplicon region (41), thermodynamic data on the formation of an n-order structure in the nucleic acid sequence of the amplicon region (41), thermodynamic data on the state in which the n-order structure is formed in the nucleic acid sequence of the amplicon region (41), distance data between a predetermined position in the nucleic acid sequence of the amplicon region (41) and a position of the n-order structure, reaction conditions, a type of oligonucleotide bound to the amplicon region (41), a GC content of the nucleic acid sequence of the amplicon region (41), a mutation score for the amplicon region (41), and a conservation score for the amplicon region (41).

[0242] In one embodiment, the second feature data provided to the second prediction model is selected from the group consisting of a nucleic acid sequence of the oligo bidding region (42), thermodynamic data on the possibility of formation of an n-order structure in the nucleic acid sequence of the oligo bidding region (42), thermodynamic data on the state in which the n-order structure is formed in the nucleic acid sequence of the oligo bidding region (42), the length of the oligo bidding region (42) in the state in which the n-order structure is formed in the nucleic acid sequence of the oligo bidding region (42), distance data between a predetermined position in the nucleic acid sequence of the oligo bidding region (42) and the position of the n-order structure, a predicted value for Tm of the nucleic acid sequence of the oligo bidding region (42), reaction conditions, the type of oligonucleotide bound to the oligo bidding region (42), the GC content of the nucleic acid sequence of the oligo bidding region (42), a mutation score for the oligo bidding region (42), and a conservation score for the oligo bidding region (42). At least one of the selected ones may be included. The length of the oligo biding region (42) in a state where an n-th structure is formed in the nucleic acid sequence of the above-mentioned oligo biding region (42) may be calculated as, for example, the remaining length (e.g., 10 bp) excluding the length of the stem and loop (e.g., stem 5 bp, loop 15 bp) from the total length (e.g., 30 bp) of the nucleic acid sequence in a state where a secondary structure of a hairpin is formed within the nucleic acid sequence.

[0243] As described above, factors that have a significant influence on the amplification reaction when the target region is the amplicon region and when the primer binding region may be different. According to one embodiment of the present invention, first and second learning models that are trained to more accurately predict the amplification reaction efficiency when the target region is the amplicon region and when the primer binding region are respectively considered by taking these different factors into account can be prepared, thereby enabling the search for an amplifiable region (70) in which the amplification reaction can be performed more efficiently.

[0244] A computer device (1000) according to one embodiment can provide analysis results for patent infringement on a nucleic acid sequence of a region of interest (40). In one embodiment, a database of the computer device (1000) may store base sequence information related to a patent right, and a search engine capable of searching for base sequence information based on the base sequence may be implemented. The computer device (1000) can retrieve base sequence information that at least partially matches the base sequence of the region of interest (40) from the database, and can analyze base similarity between the retrieved base sequence information and the base sequence of the region of interest (40) to calculate a matching degree. For example, when the number of matching bases at each position is equal to or greater than a predetermined ratio (e.g., 70%) of the total number of bases, the computer device (1000) can output a warning message recommending additional patent infringement analysis for the base sequence of the region of interest (40), or can exclude the region of interest (40) in the process of determining an amplifiable region (70).

[0245] Meanwhile, a method for determining an amplifiable region can be performed using the predictive model described herein, as described above. These technical features can be used independently, without being combined with other embodiments described below, depending on the embodiment.

[0246]

[0247] 1-6. Examples of obtaining information on amplifiable areas

[0248] The computer device (1000) can obtain information on an amplifiable region (70). In one embodiment, as the amplifiable region (70) is determined, the computer device (1000) obtains information on the amplifiable region (70) based on information on the corresponding region of interest (40) and an amplification reaction efficiency (730) of the region of interest (40), and associates the obtained information on the amplifiable region (70) with the corresponding target nucleic acid molecule and stores and manages it in the memory (100). In one embodiment, the stored information on the amplifiable region (70) can be provided to a user at a later stage based on a user request.

[0249] Information of an amplifiable region (70) according to one embodiment may include at least one selected from the group consisting of (a) information on the location, base sequence, score, and bindable oligonucleotide of the region of interest (40) included in the amplifiable region (70), and (b) information on the location, base sequence, score, and bindable oligonucleotide of the amplifiable region (70).

[0250] Here, the positions of the regions of interest (40) included in the amplifiable region (70) represent the positions of each of the regions of interest (40) among the alignment positions in the alignment result (20). For example, the positions of the second region of interest (40b) to the fourth region of interest (40d) included in the amplifiable region (70) may be indicated as alignment positions 71 to 170, 141 to 240, and 211 to 310, respectively.

[0251] The base sequence of the region of interest (40) included in the amplifiable region (70) above represents a base sequence arranged according to the alignment position of the region of interest (40). For example, the base sequence is a nucleic acid sequence of the region of interest (40) included in the information of the region of interest (40) described above, and may include one or more of the first to fourth sequences.

[0252] Information on the bindable oligonucleotide of the region of interest (40) included in the amplifiable region (70) above indicates sequence information of a candidate oligonucleotide (e.g., candidate primer, candidate probe) that can be used for detection of the region of interest (40). This sequence information may be expressed as, for example, a complementary base sequence or base sequence pair to the base sequence of the oligo binding region (42) of the region of interest (40) having a second amplification reaction efficiency of a predetermined level or higher, or may be expressed based on position information of the oligo binding region (42), but is not limited thereto. In one embodiment, this sequence information may include information on a sequence set including a forward candidate sequence and a reverse candidate sequence.

[0253] The position of the amplifiable area (70) indicates the position of the amplifiable area (70) among the alignment positions in the alignment result (20). For example, the position may be indicated as alignment position 71 to 310.

[0254] The base sequence of the amplifiable region (70) represents a base sequence arranged according to the alignment position of the amplifiable region (70). In one embodiment, the base sequence of the amplifiable region (70) may include a base sequence of a region of interest (40) merged within the amplifiable region (70), and may include, for example, one or more of the first to fourth sequences of the merged region of interest (40), but is not limited thereto. In another embodiment, the base sequence of the amplifiable region (70) may include a bundle of base sequences at the alignment position of the amplifiable region (70) among a plurality of nucleic acid sequences (10).

[0255] The information of the bindable oligonucleotide of the amplifiable region (70) above represents the sequence information of the candidate oligonucleotide (e.g., candidate primer, candidate probe) that can be used for the detection of the amplifiable region (70). This sequence information may be expressed as, for example, a complementary base sequence or base sequence pair to the base sequence of the oligo binding region (42) of the regions of interest (40) included in the amplifiable region (70) or the regions of interest (40) among the regions of interest (40) having an amplification reaction efficiency (730) above a predetermined level, but is not limited thereto.

[0256] The score of the region of interest (40) included in the amplifiable region (70) above represents an evaluation score for the region of interest (40). In one embodiment, the score of the region of interest (40) may be calculated using at least one selected from the group consisting of (i) numerical data (e.g., probability value) included in the amplification reaction efficiency (730) of the region of interest (40), (ii) the values ​​of each of the plurality of features included in the feature data, and (iii) the contribution of each of the plurality of features to the output of the amplification reaction efficiency in the prediction model (710).

[0257] Here, the contribution can be calculated using a previously stored contribution calculation method. Specifically, the input data of the prediction model (710) includes multiple features (e.g., the change in Gibbs free energy described above, GC content, etc.) as components of the feature data, and the previously stored contribution calculation method can be applied to the prediction model (710) to calculate the contribution of each of these multiple features to the generation of output data including the amplification reaction efficiency (730). For example, the contribution calculation method may include a method of checking the coefficients of independent variables defining the prediction function implemented in the prediction model (710), a method of checking the classification weight of the tree included in the prediction model (710), a method of checking the features, weights, and major object locations of input data dependent on the prediction model (710) (e.g., LRP), a method of performing cause analysis while looking at the output obtained by adjusting the input without depending on the prediction model (710) (e.g., LIME), or a method of extracting explainable features from the prediction model (710) (e.g., SmoothGrad).

[0258] In one embodiment, the contribution may be obtained by calculating the contribution of each feature to a prediction model (710) prepared through learning. In another embodiment, the contribution may be obtained by calculating the contribution of each feature used in the output of an amplification reaction efficiency (730) from a prediction model (710) in response to the input of feature data.

[0259] The score of the region of interest (40) may be, for example, a probability value included in the amplification reaction efficiency (730) of the region of interest (40) converted into a number. For another example, the score of the region of interest (40) may include a score for each of a plurality of features, and each score may be converted into a number by applying a preset standard to the values ​​of each of the plurality of features. For another example, the score of the region of interest (40) may be calculated as a number by applying the values ​​of each of the plurality of features included in the feature data to a predetermined mathematical formula, and the mathematical formula may be designed such that each of the plurality of features is an independent variable and the contribution of each of the plurality of features is applied as a weight for each independent variable. Depending on the embodiment, the score of the region of interest (40) may also be calculated by combining the above-described examples.

[0260] The score of the amplifiable region (70) above represents an evaluation score for the amplifiable region (70). In one embodiment, the score of the amplifiable region (70) may be calculated based on the score of the region of interest (40) included in the amplifiable region (70). For example, the score of the amplifiable region (70) may include all scores of the regions of interest (40) merged into the amplifiable region (70), or may be calculated based on the result of a statistical operation (e.g., an average value) on the scores of each region of interest (40).

[0261] In another embodiment, the score of the amplifiable region (70) may include a score for each of a plurality of preset items. Here, the plurality of items may include at least one item selected from the group consisting of (a) a first amplification reaction efficiency in the amplicon region (41) of the region of interest (40) included in the amplifiable region (70), (b) a second amplification reaction efficiency in the oligo binding region (42) of the corresponding region of interest (40), (c) a mutation score for the amplifiable region (70), and (d) a conservation score for the amplifiable region (70). For example, the score of the first item may be an evaluation score for the amplicon efficiency, which may be a score calculated by averaging the first amplification reaction efficiencies of the regions of interest (40) merged into the amplifiable region (70). The score of the second item may be an evaluation score for the oligo binding efficiency (e.g., primer efficiency), which may be a score calculated by averaging the second amplification reaction efficiencies of the corresponding merged regions of interest (40). The scores of the third item and the fourth item are an evaluation score for whether the sequence can operate robustly against mutations and an evaluation score for sequence conservation, and may be a mutation score and a conservation score calculated according to the examples described above for the nucleic acid sequences of the amplifiable region (70) among the plurality of nucleic acid sequences (10).

[0262] FIG. 9 is a drawing illustrating information on multiple amplifiable regions (70) for a target nucleic acid molecule according to one embodiment.

[0263] Referring to FIG. 9, an identifier is assigned to each of the amplifiable regions (70) for the target nucleic acid molecule, and information on the amplifiable regions (70) including the location of each amplifiable region (70), base sequence, score of the first item (e.g., amplicon efficiency score), score of the second item (e.g., primer efficiency score), score of the third item (e.g., mutation score), and score of the fourth item (e.g., conservation score) can be stored and managed.

[0264] The computer device (1000) can determine the priorities of amplifiable regions (70) for a target nucleic acid molecule and store information on the amplifiable regions (70) including the priorities. Here, the priorities represent priorities for use in designing oligonucleotides, and can represent, for example, an order in which the amplification reaction efficiency in the corresponding region is predicted to be relatively higher.

[0265] In one embodiment, the computer device (1000) can determine the priority of the amplifiable regions (70) based on thermodynamic data on the formation of n-th structure in each of the amplifiable regions (70). For example, the computer device (1000) can calculate the change in Gibbs free energy for the formation of secondary structure in the base sequence of each amplifiable region (70) using the free energy calculation algorithm described above, and can determine the priority of an amplifiable region (70) with a lower possibility of forming secondary structure based on the calculated change in Gibbs free energy value.

[0266] In another embodiment, the computer device (1000) may determine the priority of the amplifiable regions (70) by further including at least one selected from the group consisting of (a) distance data between a predetermined position in the nucleic acid sequence of each amplifiable region (70) and a position of the n-th structure; (b) GC content of the nucleic acid sequence of each amplifiable region (70); (c) a mutation score for each amplifiable region (70); and (d) a conservation score for each amplifiable region (70). For example, the computer device (1000) may calculate a change value of Gibbs free energy for the formation of a secondary structure in the base sequence of each amplifiable region (70), a GC content value, a mutation score, and a conservation score, and may determine a higher priority of an amplifiable region (70) having a relatively higher weighted average score of the calculated values ​​based on preset weights. For example, the weights may be those with the largest change in Gibbs free energy, followed by the mutation score and the conservation score, in that order.

[0267] In another embodiment, the computer device (1000) may further include the amplification reaction efficiency of each of the plurality of amplifiable regions (70) output from the prediction model (710) to determine the priority of the amplifiable regions (70). For example, the computer device (1000) may determine a higher priority for an amplifiable region (70) with a higher score calculated by considering the amplification reaction efficiency.

[0268] In another embodiment, the computer device (1000) may calculate the priority of the amplifiable regions (70) by using at least one selected from the group consisting of (i) numerical data included in the amplification reaction efficiency (710), (ii) the value of each of the plurality of features included in the feature data, and (iii) the contribution degree that each of the plurality of features contributes to the output of the amplification reaction efficiency (730) in the prediction model (710). For example, the computer device (1000) may determine a priority of an amplifiable region (70) having a relatively higher overall score by considering a score calculated as a numerical value by applying the value of each of the plurality of features included in the corresponding feature data to a mathematical formula to which the contribution degree is applied as a weight and the average value of the amplification reaction efficiency of each of the aforementioned amplifiable regions (70) as well as a weighted value.

[0269] In another embodiment, the computer device (1000) may determine the priority of the amplifiable regions (70) based on the score of each of the amplifiable regions (70). In one embodiment, the computer device (1000) may determine the priority based on the score of each of the amplifiable regions (70), and for example, may determine a higher priority for an amplifiable region (70) having a relatively higher score. For example, the score of each of the amplifiable regions (70) may be the score of any one of the first to fourth items described above, an arithmetic mean of the scores of the first to fourth items, or a weighted mean of the scores of the first to fourth items based on preset weights for each item.

[0270] As described above, if the amplification reaction efficiency (730) is predicted for the region of interest (40), the priority may be evaluated for the amplification reaction region (70). For example, in the learning process of the prediction model (710), there may be elements (e.g., mutation score, conservation score) that are not easy to apply as feature data due to reasons such as low correlation with other features or difficulty in quantifying the target nucleic acid sequence for learning. According to an embodiment, elements that are difficult to consider in the process of obtaining the amplification reaction efficiency (730) of the region of interest (40) may be additionally considered in the process of evaluating the amplifiable region (70). Accordingly, there is an advantage in that the priority of the amplifiable regions (70) can be evaluated more accurately based on a wider variety of elements that may affect the efficiency of the amplification reaction.

[0271] FIG. 10 is a diagram illustrating information of amplifiable areas (70) including priorities according to one embodiment.

[0272] Referring to FIG. 10, the computer device (1000) may apply at least some of the scores of the first item (e.g., amplicon efficiency score), the second item (e.g., primer efficiency score), the third item (e.g., mutation score), and the fourth item (e.g., conservation score) of each of the amplifiable regions (70) to a pre-stored mathematical formula to calculate a comprehensive score for distinguishing priorities, and may determine the priorities of the amplifiable regions (70) based on the calculated comprehensive score. For example, a higher priority may be given to an amplifiable region (70) having a higher average value of the scores of the above-described items. In addition, a higher weight may be given to preset items (e.g., amplicon efficiency score and primer efficiency score).

[0273] In another embodiment, priorities may be determined for each group according to the grouped amplifiable areas (70). For example, first to m-th groups matching the first to m-th score ranges (m is a natural number greater than or equal to 2) may be preset, and each amplifiable area (70) may be grouped into a group matching the score range to which the score of the corresponding amplifiable area (70) belongs among the first to m-th groups. In addition, among the groups, the priority of a group having a relatively large value size of the score range may be set higher, and each amplifiable area (70) may follow the priority of the group to which it belongs.

[0274] Meanwhile, the processes described above can be performed for each of a plurality of target nucleic acid molecules. An identifier can be assigned to each of the plurality of target nucleic acid molecules, and information on amplifiable regions (70) mapped to the identifier for each target nucleic acid molecule can be stored and managed.

[0275] The computer device (1000) can store information on amplifiable regions (70) that include priorities. In one embodiment, the computer device (1000) can store and manage information on amplifiable regions (70) in a database of the computer device (1000) based on an identifier of the target nucleic acid molecule and / or an identifier of an organism having the target nucleic acid molecule.

[0276] In one embodiment, the database of the computer device (1000) may include a search engine for retrieving information on amplifiable regions. For example, the search engine may be implemented to provide search results for amplifiable regions from the database based on a predefined search method (e.g., SQL-based keyword search). To this end, indexing items or metadata items for search may be preset based on data items included in at least one of information on the amplifiable region, information on the target nucleic acid molecule, and information on the organism. In one embodiment, an identifier of the target nucleic acid molecule, a taxonomic name of the organism having the target nucleic acid molecule, the number of amplifiable regions (70) for the target nucleic acid molecule, priorities, scores, etc. may be set as metadata items for search. The computer device (1000) may apply such indexing methods or metadata setting methods to the information on the amplifiable regions (70) described above, process it into a searchable form by the search engine, and store it in the database of the computer device (1000).

[0277] In one embodiment, the computer device (1000) can generate an amplifiable region library including information on the amplifiable regions (70) described above for each target nucleic acid molecule, and store and manage the amplifiable region library in a database of the computer device (1000). In one embodiment, the amplifiable region library can be managed based on the metadata items described above. In one embodiment, the amplifiable region library can be categorized by target nucleic acid molecule, and the target nucleic acid molecules can be categorized by organism having the target nucleic acid molecule. In one embodiment, a database environment can be constructed based on this amplifiable region library to provide information on the amplifiable region of a target nucleic acid molecule corresponding to a user request.

[0278]

[0279] 2. Technical features for learning predictive models

[0280] A computer device (1000) according to a second embodiment of the present invention can perform technical features for learning a predictive model used to determine an amplifiable region. In one embodiment, the computer device (1000) can train a predictive model to adjust an amplification reaction efficiency for an amplification reaction in a given nucleic acid sequence based on a given nucleic acid sequence or feature data predicted to affect an amplification reaction for the given nucleic acid sequence. In another embodiment, the computer device (1000) can train a predictive model to adjust an amplification reaction efficiency for an amplification reaction in a given region based on information about a given region of nucleic acid sequences.

[0281] FIG. 11 illustrates an exemplary flowchart of a computer device (1000) according to a second embodiment training a predictive model. In one embodiment, the steps of FIG. 11 may be implemented by a single entity, such as in a server.

[0282] Referring to FIG. 11, in step S1110, the computer device (1000) may obtain a plurality of training data sets for training a prediction model. In one embodiment, each of the plurality of training data sets may include training input data and training correct data corresponding to the training input data. In one embodiment, the training input data may refer to input data (e.g., values ​​of independent variables) input to a prediction model in supervised learning, and the training correct data may refer to data labeled as correct answers in supervised learning (e.g., expected output of a dependent variable or a label determined from the expected output).

[0283] In one embodiment, the training input data may include feature data predicted to affect the amplification reaction of the training nucleic acid sequence, and the training answer data may include the amplification reaction efficiency for the amplification reaction in the training nucleic acid sequence. Throughout the specification, the training nucleic acid sequence mainly refers to a nucleic acid sequence that is a learning target of a prediction model in the training process of the model, and the nucleic acid sequence may be understood as a term that mainly refers to a nucleic acid sequence that is a prediction target of the amplification reaction efficiency or is used for prediction in the process of obtaining a prediction result using a previously trained prediction model, but is not limited thereto.

[0284] In one embodiment, the feature data included in the training input data includes thermodynamic data on the formation of an n-order structure in the training nucleic acid sequence, and in one embodiment, the feature data may further include at least one selected from the group consisting of (a) the training nucleic acid sequence; (b) distance data between a predetermined position in the training nucleic acid sequence and a position of the n-order structure; (c) reaction conditions including a reaction medium used for an amplification reaction for the training nucleic acid sequence; (d) a type of oligonucleotide bound to the training nucleic acid sequence; (e) GC content of the training nucleic acid sequence; (f) a mutation score for a training sequence group including the training nucleic acid sequence; and (g) a conservation score for the training sequence group.

[0285] In one embodiment, the training answer data may include the amplification reaction efficiency for an amplification reaction in the training nucleic acid sequence. In one embodiment, the amplification reaction efficiency for an amplification reaction in the training nucleic acid sequence may include numerical data regarding the efficiency level of the amplification reaction for the training nucleic acid sequence and / or label data regarding the high and low levels of the efficiency level.

[0286] In step S1120, the computer device (1000) may acquire a learned prediction model to predict the amplification reaction efficiency for the amplification reaction using multiple learning datasets. In one embodiment, the learned prediction model may correspond to the prediction model (710).

[0287] A computer device (1000) according to one embodiment can train a predictive model using a plurality of training data. For example, the computer device (1000) can input training input data into a predictive model for each training data, update the predictive model using the output obtained from the predictive model and training correct data, and repeat this process for each of the plurality of training data, thereby training the predictive model.

[0288] In one embodiment, the prediction model may be implemented as a machine learning model based on a multiple regression analysis method. In this case, each feature included in the feature data of the training input data (e.g., the change in Gibbs free energy for the formation of an n-th structure, etc.) corresponds to each independent variable in a prediction function based on a multiple regression analysis method, and the training correct answer data (e.g., the amplification reaction efficiency) may correspond to the expected output of the dependent variable in the corresponding prediction function. During the supervised learning process for the prediction model, the computer device (1000) may update the operation of the prediction model in a direction of reducing the difference by feeding back the difference between the output of the dependent variable by the independent variables of the prediction function and the expected output.

[0289] In another embodiment, the prediction model may be implemented as a DNN-based deep learning model. In this case, training input data may be preprocessed in a predetermined vector or matrix form and provided to the neural network input layer of the prediction model. The prediction model may output multiple class-specific probability values ​​based on features extracted from the training input data through successive layers of the deep neural network. During the supervised learning process for the prediction model, the computer device (1000) may update the weights of the prediction model, etc., so as to minimize the difference between the output and the training correct data.

[0290] Various embodiments of the technical features presented above are described in more detail below. Furthermore, these technical features can be implemented in a manner similar to the feature data acquisition process and preprocessing steps described above, and thus, any redundant description will be omitted.

[0291]

[0292] 2-1. Examples of obtaining a learning dataset

[0293] The computer device (1000) can acquire a plurality of first data groups for a plurality of training datasets. Each first data group includes a predetermined training nucleic acid sequence and reaction conditions used for an amplification reaction for the training nucleic acid sequence, and may further include a type of oligonucleotide that binds to the training nucleic acid sequence. In one embodiment, the computer device (1000) can load the plurality of first data groups from the memory (100) or a storage device, or receive them from another device via the communication unit (200).

[0294] According to one embodiment, the learning nucleic acid sequence may be data obtained from a public database, or data processed, modified, or separated from the public database. For example, virus sequences of a specific species may be collected from a public database such as NCBI, GISAID (Global Initiative for Sharing All Influenza Data), and / or ATCC (American Type Culture Collection), and one or more nucleic acid sequences for learning may be determined by performing alignments on the collected virus sequences to search for conserved regions, etc.

[0295] The computer device (1000) can generate a plurality of second data groups using a plurality of first data groups. In one embodiment, the second data group can include at least one of thermodynamic data on the formation of an n-order structure in a training nucleic acid sequence, distance data between a predetermined position in the training nucleic acid sequence and a position of the n-order structure, GC content of the training nucleic acid sequence, a mutation score for a training sequence group including the training nucleic acid sequence, and a conservation score for the training sequence group.

[0296] The computer device (1000) can calculate thermodynamic data for the formation of an n-dimensional structure in the learning nucleic acid sequence and distance data between a predetermined position in the learning nucleic acid sequence and a position of the n-dimensional structure, based on the learning nucleic acid sequence and reaction conditions of the first data group. As described above, a thermodynamic property calculation algorithm and a distance calculation algorithm, etc., can be used for such calculation. In addition, the thermodynamic data can include at least one of thermodynamic data for the formation of an arbitrary n-dimensional structure, thermodynamic data for the formation of a specific n-dimensional structure, and thermodynamic data for the formation of each of a plurality of n-dimensional structures.

[0297] The computer device (1000) can determine an amplicon sequence from a learning nucleic acid sequence of a first data group, and can determine an oligo binding sequence from the amplicon sequence. The computer device (1000) can use the amplicon sequence and the corresponding reaction conditions to produce first thermodynamic data on the formation of an n-dimensional structure in the amplicon sequence, and use the oligo binding sequence and the corresponding reaction conditions to produce second thermodynamic data on the formation of an n-dimensional structure in the oligo binding sequence.

[0298] The computer device (1000) can use the learning nucleic acid sequence of the first data group to calculate the GC content and / or length of the learning nucleic acid sequence.

[0299] The computer device (1000) can obtain a learning sequence group including the learning nucleic acid sequence based on the learning nucleic acid sequence of the first data group, and calculate a mutation score and a conservation score for each of the learning sequence group. For example, learning nucleic acid sequences for learning target nucleic acid molecules are previously stored in the memory (100), and the alignment result of the learning nucleic acid sequences can be obtained, and from the alignment result, a learning sequence group including other learning nucleic acid sequences in a region corresponding to the position of the learning nucleic acid sequence can be obtained. As described above, the mutation score for the sequence group can be calculated based on the mutation position or type, or the conservation score for the sequence group can be calculated based on an entropy measurement.

[0300] The second data group may include the amplification reaction efficiency for the amplification reaction in the corresponding learning nucleic acid sequence. The computer device (1000) may calculate the amplification reaction efficiency of the second data group based on the first data group, but is not limited thereto. In other embodiments, the amplification reaction efficiency may be included in the first data group.

[0301] Specifically, an amplification reaction is performed to which a learning nucleic acid sequence and a reaction condition are applied in each first data group, and the computer device (1000) can obtain a dataset from the amplification reaction. In one embodiment, the dataset may be a set of coordinate values ​​including cycles and signal values, and may be displayed as coordinate values ​​on a two-dimensional rectangular coordinate system. In the coordinate values, the X-axis may represent the number of cycles, and the Y-axis may represent a signal value (e.g., RFU (relative fluorescence units)) measured or processed in the corresponding cycle. In one embodiment, the first data group further includes the above dataset, and the computer device (1000) can calculate an amplification reaction efficiency for the amplification reaction in the learning nucleic acid sequence from the dataset of the first data group.

[0302] The computer device (1000) can calculate the amplification reaction efficiency of the second data group using the number of cycles corresponding to an amplification point or an amplification region in a dataset for learning nucleic acid sequences. Here, the amplification point refers to a specific number of cycles corresponding to a reaction time or a reaction number during which amplification of a signal value has progressed to a predetermined level or more. In addition, the amplification region refers to a cycle section corresponding to a reaction time section or a reaction number section during which amplification of a signal value has progressed to a predetermined level or more. In some embodiments, the terms amplification point and amplification region may be used interchangeably.

[0303] The number of cycles corresponding to an amplification point or amplification region may include (i) the number of cycles at which the first or second derivative of a curve connecting the signal values ​​for each cycle in the data set is maximum or minimum and / or (ii) a specific number of cycles at which the signal value in the data set reaches a preset threshold.

[0304] The number of cycles corresponding to the above-mentioned amplification point may include the number of cycles that touch the maximum slope in the curve connecting the signal values ​​for each cycle. For example, the number of cycles may include the number of cycles in which a straight line according to the slope at the inflection point in the curve connecting the signal values ​​for each cycle intersects the straight line according to the X-axis or the threshold value. Alternatively, the number of cycles corresponding to the above-mentioned amplification point may include the number of cycles in which the first derivative result or the second derivative result for the curve connecting the signal values ​​for each cycle is maximum or minimum. For example, in a dataset, the first derivative result for the curve connecting the signal values ​​for each cycle is obtained, and the above-mentioned amplification point may include the number of cycles in which the first derivative result or the second derivative result satisfies a specific condition (e.g., the first derivative value is maximum). Examples of such cycle numbers include FDM (first derivative maximum), SR (slope regression)-FDM, SDM (second derivative maximum), and SR-SDM.

[0305] The above-mentioned specific number of cycles includes the number of cycles in which the signal value on the curve connecting the signal values ​​for each cycle included in the dataset reaches a threshold value, and may include, for example, Ct (cycle threshold). Depending on the embodiment, Ct may be interpreted as encompassing the terms CP (cross point), TOP (take-off point), or CQ (quantification cycle) used in the present technical field.

[0306] The amplification region described above may include a cycle section within a curve connecting signal values ​​for each cycle, where the signal value or the slope of the signal value satisfies a predetermined condition. These amplification points and amplification regions are not limited to the aforementioned embodiments, and various known analysis indicators used in the art to analyze amplification results may be applied.

[0307] The computer device (1000) can calculate the amplification reaction efficiency based on the comparison result between the amplification point and the reference amplification point. For example, the amplification reaction efficiency may be the result of subtracting the reference amplification point from the above-mentioned number of cycles. The reference amplification point may refer to an ideal amplification region in which the amplification reaction is efficiently performed in the amplification reaction for the corresponding nucleic acid sequence. In one embodiment, the reference amplification point may be a pre-stored set value, or a measurement value or statistical value determined from data sets obtained for the corresponding nucleic acid sequence.

[0308] The computer device (1000) can calculate the amplification reaction efficiency by using the difference in the amplification point determined from two or more data sets for the learning nucleic acid sequence. Here, the two or more data sets may be obtained from two or more amplification reactions for the same learning nucleic acid sequence. For example, two or more data sets may be obtained from two or more amplification reactions for the nucleic acid sequence performed under the reaction conditions of each first data group. According to an embodiment, the same reaction conditions may mean that the types, concentrations, or sizes of various element conditions (e.g., reaction medium, temperature, etc.) included in the above-described reaction conditions are the same or fall within a predetermined range and are substantially at the same level.

[0309] For example, when an amplification reaction is performed using a plate containing two or more reaction vessels, two or more amplification reactions for one identical nucleic acid sequence may be performed together to obtain two or more data sets together. The number of these two or more amplification reactions (e.g., m1, m2) may be the same or different depending on the respective reaction conditions. In addition, the two or more data sets described above may be obtained together when the corresponding first data group is obtained, or may be obtained by performing an amplification reaction under the corresponding reaction conditions after the first data group is obtained. In addition, the two or more data sets described above may be processed to exclude the influence of one or more preset exclusion variables. The exclusion variables may include, for example, variables related to at least one of a dimer, a Tm of an oligonucleotide set, the number of oligonucleotides, concentration, temperature, and the number of target nucleic acid molecules. The influence of these variables may be checked in various ways, such as adding a reaction vessel to confirm variable control within the plate, or checking a signal indicating whether the variable has been controlled during the amplification reaction. Datasets that are analyzed to contain influences due to excluded variables can be excluded from the training dataset.

[0310] A computer device (1000) can determine the difference in amplification points from two or more datasets for learning nucleic acid sequences and calculate the amplification reaction efficiency using the difference in amplification points. For example, the two or more datasets may include at least one of a case in which amplification is well performed, a case in which amplification is partially performed, and a case in which amplification is not performed for the nucleic acid sequence. The computer device (1000) can calculate the amplification reaction efficiency as a number by measuring the difference in amplification points of cases appearing in the two or more datasets.

[0311] FIG. 12 is a diagram illustrating an amplification reaction efficiency calculated by using the difference in amplification points according to one embodiment.

[0312] Referring to FIG. 12(a), the computer device (1000) can measure the amplification point described above from a curve connecting each of a plurality of data sets for learning nucleic acid sequences, and calculate the amplification reaction efficiency using the comparison result of the measured amplification point values. In one embodiment, the method of obtaining the comparison result may include a method of obtaining a difference between amplification point values ​​(e.g., Cmax - Cmin), a method of performing an arithmetic operation on at least some of the amplification point values ​​according to a preset mathematical formula (e.g., (Cmax - Cmin)*weight, etc.), and a method of performing a statistical operation on one or more representative values ​​(e.g., mean, median, mode, specific quantiles (e.g., first quartile, third quartile), etc.) or representative ranges (e.g., IQR (interquartile range), etc.) using a pre-stored statistical algorithm (e.g., box plot), etc.

[0313] Referring to FIG. 12(b), the computer device (1000) can measure the amplification region described above from a curve connecting each of a plurality of data sets for learning nucleic acid sequences, and calculate the corresponding amplification reaction efficiency using the comparison result between the amplification regions. In one embodiment, a method for obtaining the comparison result may include a method for obtaining the difference between the minimum and maximum cycles when the amplification regions overlap, and a method for obtaining the difference between the aspects of the amplification change (e.g., the slope of the amplification curve) that appear in each amplification region.

[0314] According to another embodiment, a computer device (1000) can calculate an amplification reaction efficiency by using the difference between a signal pattern determined from a dataset for a learning nucleic acid sequence and a preset reference pattern. As described above, the difference means a numerical representation of the signal pattern in the dataset or the degree to which the pattern matches the preset reference pattern. In one embodiment, the computer device (1000) can determine a signal pattern from two or more datasets for a learning nucleic acid sequence. For example, an average signal pattern or a worst signal pattern of the datasets can be obtained, or a signal pattern of a randomly selected dataset can be obtained.

[0315] FIG. 13 is a diagram for explaining the amplification reaction efficiency calculated by using the difference between the signal pattern and the reference pattern according to one embodiment.

[0316] Referring to FIG. 13, the computer device (1000) can calculate the values ​​of parameters (e.g., a1, a2, a3, a4) of a predetermined growth function (e.g., sigmoid function) by applying a signal analysis method (e.g., sigmoid fitting) previously stored in the corresponding dataset. The computer device (1000) can calculate the amplification response efficiency by comparing the signal pattern of the growth curve (1310) according to the values ​​of the calculated parameters with the signal pattern of the reference curve (1320) according to the reference values ​​of each of the preset parameters. For example, in FIG. 13(a), the degree to which the signal patterns of the growth curve (1310) and the reference curve (1320) match is relatively high, which exemplifies a case where the difference between the values ​​of the parameters and the reference values ​​is relatively small, resulting in a relatively ideal signal pattern. In this case, the amplification response efficiency can be calculated to be relatively high. For another example, in Fig. 13(b), the degree of matching between the signal patterns of the growth curve (1310) and the reference curve (1320) is relatively low, which exemplifies a case where the signal patterns are not relatively ideal. In such a case, the amplification reaction efficiency may be calculated to be relatively low.

[0317] In one embodiment, the amplification reaction efficiency described above may be numerical data regarding the efficiency level of the amplification reaction. For example, the amplification reaction efficiency described above may be a numerical value measured from the difference in amplification point or signal pattern according to the aforementioned embodiments.

[0318] In another embodiment, the amplification reaction efficiency may be label data for the high and low of the efficiency level. The label data may be label data for whether the numerical data satisfies a predetermined criterion. The predetermined criterion may include at least one of a condition for comparing the magnitude between the amplification inhibition and a set value and a condition for determining which of a plurality of value ranges it falls into. For example, the amplification reaction efficiency may be labeled with a first label (e.g., 1) indicating that amplification is inhibited if the numerical data (e.g., the difference in Ct values) is greater than or equal to a set value (e.g., 2), and may be labeled with a second label (e.g., 0) indicating that amplification is not inhibited if it is less than the set value. As another example, a plurality of intervals corresponding to a plurality of amplification efficiency levels (e.g., good efficiency, medium efficiency, low efficiency) are preset, and the amplification reaction efficiency can be labeled with any one of the third to fifth labels (e.g., 3, 2, 1) according to the amplification efficiency level corresponding to the interval to which the numerical data (e.g., probability value for efficient performance of the amplification reaction) belongs.

[0319] In another embodiment, the amplification reaction efficiency may include numerical data for the efficiency level of the amplification reaction and / or label data for the high and low of the efficiency level.

[0320] A computer device (1000) can acquire a plurality of training data sets including at least one of a plurality of first data groups and a plurality of second data groups. The computer device (1000) can generate a training data set including training input data and training correct data labeled with the corresponding training input data as a single data set from at least one of each of the first data groups and each of the second data groups.

[0321] The above has described embodiments in which a computer device (1000) generates multiple training data sets, but is not limited thereto. According to another embodiment, the computer device (1000) may load multiple training data sets from a memory (100) or a storage device, or receive them from another device via a communication unit (200).

[0322]

[0323] 2-2. Examples of training a prediction model using a learning dataset

[0324] A computer device (1000) can train a predictive model using multiple training datasets. This process can be performed in a manner similar to the process of obtaining output results based on input by the predictive model (710) described above, and any redundant details will be omitted.

[0325] Figure 14 illustrates a conceptual diagram of the learning process of a prediction model according to one embodiment. Here, the prediction model (1410) may correspond to the prediction model (710) before learning is completed.

[0326] Referring to FIG. 14, a prediction model (1410) can be trained to output an amplification reaction efficiency (1430) for an amplification reaction in a given learning nucleic acid sequence by using learning input data including feature data (1420) that affect the amplification efficiency of a given learning nucleic acid sequence. For example, during the learning process, the prediction model (1410) can apply the Gibbs free energy change amount, distance data, reaction medium, GC content, etc. of the learning input data to the corresponding independent variables in the prediction model (1410) to output an amplification reaction efficiency (1430) as a dependent variable. The amplification reaction efficiency (1440) labeled as learning correct data in the learning input data can be the expected output of the dependent variable, and the prediction model (1410) can be supervised learning by updating the operation of the prediction model (e.g., inversion of coefficient values ​​included in the prediction function, etc.) so that the error between the output of the dependent variable and the expected output is minimized.

[0327] In one embodiment, the prediction model (1410) is implemented as a regression analysis model (e.g., ridge linear regression, lasso linear regression, random forest regression), and the output of the dependent variable and the expected output in the prediction model (1410) can be processed in the form of the above-described numerical data (e.g., the difference value of the amplification point). The ridge or lasso linear regression model is trained by applying additional constraints (penalties) along with the basic conditions of finding the weights (w) and bias (b) that minimize the error (e.g., MSE), and the constraints can be applied with techniques such as L1-norm regularization or L2-norm regularization. The random forest model includes a plurality of different decision trees and an ensemble unit, and classification is performed based on variables and classification conditions at each node of the decision tree, and the ensemble unit can provide a result of ensembling the decisions output from each tree.

[0328] In another embodiment, the prediction model (1410) is implemented as a classification model (e.g., logistic regression, random forest classification), and the output of the dependent variable and the expected output in the prediction model (1410) can be processed in the form of a probability value for a class indicating a high efficiency level and the label data described above. In another embodiment, the prediction model (1410) is implemented as a DNN with a fully connected neural network structure, and, for example, a method may be applied in which input data for training preprocessed in a vector form is converted into a one-dimensional array and then trained with a fully connected multi-layered neural network. The output and the expected output of the prediction model (1410) are processed in the form of numerical data of probability values ​​for each class or label data, and the values ​​of the parameters (e.g., weights and biases) of the artificial neural network of the prediction model (1410) can be updated by a backpropagation method to reduce the error between the output and the expected output.

[0329] In one embodiment, the prediction model (1410) may include a first prediction model for predicting the amplification reaction efficiency in the amplicon sequence and a second prediction model for predicting the amplification reaction efficiency in the oligo bidding sequence. In addition, the plurality of training datasets may include a plurality of first training datasets for training the first prediction model and a plurality of second training datasets for training the second prediction model.

[0330] The learning input data of each first learning dataset may include at least one selected from the group consisting of first feature data predicted to affect an amplification reaction in an amplicon sequence, first-first thermodynamic data on the possibility of forming an n-th structure in the amplicon sequence, first-second thermodynamic data on the state in which an n-th structure is formed in the amplicon sequence, distance data between a predetermined position in the amplicon sequence and a position of the n-th structure, reaction conditions, a type of oligonucleotide bound to the amplicon sequence, a GC content of the amplicon sequence, a mutation score for a learning amplicon sequence group including the amplicon sequence, and a conservation score for a amplicon sequence group. A computer device (1000) trains a first prediction model to predict an amplification reaction efficiency for an amplification reaction in an amplicon sequence using a plurality of first learning data sets, and as a result of the training, when feature data predicted to affect an amplification reaction in a base sequence assumed as an amplicon is input, the first prediction model trained to predict an amplification reaction efficiency in the corresponding base sequence can be obtained.

[0331] The learning input data of each second learning dataset may include at least one selected from the group consisting of second feature data that may affect the amplification reaction in the oligo bidding sequence, 2-1 thermodynamic data on the possibility of forming an n-order structure in the oligo bidding sequence, 2-2 thermodynamic data on the state in which the n-order structure is formed in the oligo bidding sequence, the length of the oligo bidding sequence in the state in which the n-order structure is formed in the oligo bidding sequence, distance data between a predetermined position in the oligo bidding sequence and the position of the n-order structure, a predicted value for Tm of the oligo bidding sequence, reaction conditions, the type of oligonucleotide that binds to the oligo bidding sequence, the GC content of the oligo bidding sequence, a mutation score for the oligo binding sequence group including the oligo bidding sequence, and a conservation score for the oligo binding sequence group. The computer device (1000) trains a second prediction model to predict an amplification reaction efficiency for an amplification reaction in an oligo bidding sequence using a plurality of second learning datasets, and as a result of the training, when feature data predicted to affect an amplification reaction in a base sequence assumed as an oligo bidding site is input, the second prediction model trained to predict the amplification reaction efficiency in the corresponding base sequence can be obtained. Similarly, the calculation method and training process for each feature data can be understood with reference to the embodiments described above.

[0332] Meanwhile, the prediction model of the present disclosure is not limited to the above-described embodiments, and various types of known artificial intelligence models can be applied. In another implementation example, the prediction model may include a Long Short Term Memory (LSTM) network, a Bidirectional Encoder Representations from Transformers (BERT), or a Generative Pre-trained Transformer (GPT). For example, when the prediction model includes a language model, the types and orders of bases included in the target nucleic acid sequence itself can be utilized as input data for learning, and the prediction model can be trained to match the training correct data based on inputs that have been preprocessed (e.g., vectorized) from the training input data including the nucleic acid sequence. Alternatively, the prediction model (1410) may be trained to predict the amplification reaction result by comprehensively analyzing (i) the output of a machine learning model that receives some of the training input data (e.g., thermodynamic data, distance data, etc.) and calculates the result, and (ii) the output of a deep learning model that receives the remaining training input data (e.g., nucleic acid sequence) and calculates the result.

[0333] According to one embodiment, in order to determine the components of feature data included in the learning input data used as input to the prediction model (1410), a process of selecting some of the candidate features based on correlation may be performed. For example, correlations between the candidate features are calculated, and for candidate features having a correlation value greater than or equal to a predetermined value (e.g., 0.8), one candidate feature selected as a representative feature from the candidate features may be applied as a component of the feature data. In this way, candidate features with similar value change patterns are integrated into one, and candidate features with different change patterns are used individually, thereby enabling more efficient result prediction based on a relatively small number of features.

[0334] In the learning process described above, multiple hyperparameters may be used. Hyperparameters may be variables that can be changed by the user and may vary depending on the type of the prediction model (1410). For example, if the prediction model (1410) is a ridge linear regression analysis model, hyperparameters may include alpha (e.g., regularization of the model) or the type of solver (e.g., gradient descent), and if the prediction model (1410) is a random forest regression model, hyperparameters may include the maximum depth value (e.g., depth of a decision tree), the number of features (e.g., the number of features for classification), the minimum samples_leaf value (e.g., the minimum size of data to be included in a leaf node), and the minimum samples_split value (e.g., the minimum size of data to be included in an internal node). For another example, in the case of a DNN model, hyperparameters may include learning rate, cost function, number of learning cycle iterations, weight initialization (e.g., setting the range of weight values ​​to be initialized), number of hidden units (e.g., number of hidden layers, number of nodes in a hidden layer), etc.

[0335] Meanwhile, various known methods can be used to evaluate model performance. For example, for regression models, the MSE evaluation method (e.g., neg_mean_squared_error) can be used. Furthermore, for classification models, evaluation methods based on receiver operating characteristic (ROC), area under the ROC curve (AUC), ROC-AUC, ACU_SD, Specificity, F1, or negative predictive value (NPV) can be used.

[0336] Table 1 below presents performance evaluation scores of a supervised learning prediction model (1410) based on a regression model according to one embodiment. Table 2 shows ROC-AUC for each case according to an exemplary performance evaluation method. However, as mentioned above, it can be seen that various evaluation methods known in the art can be utilized.

[0337] Prediction Model 1 Prediction Model 2 Case 10.82 0.70 Case 20.74 0.80 Case 30.78 0.72 Case 40.66 0.60 Case 50.75 0.78 Case 60.71 0.77 Case 70.77 0.80

[0338] In Table 1, Cases 1 to 7 each represent cases in which prediction models were implemented based on RF (Random Forest), LR (Logistic Regressor), GBC (Gradient Boosting Classifier), DTC (Decision Tree Classifier), GNB (Gaussian Naive Bayes), SVC (Support Vector Classifier), and NN (Neural Network), respectively. As can be seen in Table 1, high evaluation scores were obtained by applying the above-mentioned feature data as input data for learning. In addition, the highest evaluation score was obtained in the case in which the prediction model (1410) was implemented using the random forest model. In Table 1 above, ROC-AUC is described as an example, but it was confirmed that high evaluation scores were shown on average in evaluation cases that used other evaluation methods as well as ROC-AUC.

[0339] As described above, a learned prediction model can be acquired through the above-described learning process. These technical features may, depending on the embodiment, be used independently, without being combined with the method for determining an amplifiable region using the aforementioned prediction model. The computer device (1000) can store and manage the predicted model acquired in the above manner, and provide the predicted model. For example, the computer device (1000) can be implemented to store and manage the predicted model learned by a second server (or server), and provide the predicted model to a first server (or user terminal) requesting the predicted model.

[0340]

[0341] 3. Technical features that provide information on amplifiable areas based on libraries

[0342] A computer device (1000) according to a third embodiment of the present invention can perform technical features for providing information on amplifiable regions available for designing oligonucleotides for detecting target nucleic acid molecules. Here, providing information can be broadly interpreted to encompass processes such as allowing a target to acquire specific information, assisting a target in acquiring specific information, or directly or indirectly transmitting and receiving information to and from a specific target, and encompassing the performance of related operations required in such processes.

[0343] In one embodiment, the computer device (1000) can store information on amplifiable regions for a target nucleic acid molecule in a database and provide information on the amplifiable regions based on a user request.

[0344] FIG. 15 is a structural diagram for explaining the operation of a computer device (1000) according to one embodiment.

[0345] Referring to FIG. 15, a computer device (1000) may be connected to at least one of a database (2000) and one or more user terminals (3000) via a network. Here, the network may be configured through various communication networks such as wired and wireless, and may be configured as various communication networks such as a local area network (LAN), a metropolitan area network (MAN), and a wide area network (WAN).

[0346] The database (2000) may store information on an amplifiable area. The database (2000) according to one embodiment may be one of the components of the computer device (1000). For example, the computer device (1000) may include the database (2000) as an entity that performs data storage and management, and the database (2000) may be included in the computer device (1000) or may exist under the management of the computer device (1000). The database (2000) according to another embodiment may be implemented in a form that exists outside the computer device (1000) and can communicate with the computer device (1000). In this case, the database (2000) may be managed and controlled by a different external server from the computer device (1000). The database (2000) according to another embodiment may be implemented at least partially in the cloud or may include a separate storage server to provide storage space to the computer device (1000).

[0347] The user terminal (3000) may include any type of terminal capable of interacting with a server or other computing device. The user terminal (3000) may include, for example, a mobile phone, a smart phone, a desktop computer, a laptop computer, a personal digital assistant (PDA), a slate PC, a tablet PC, and an ultrabook.

[0348] In one embodiment, the user terminal (3000) may be a terminal belonging to a user account authorized to access information on an amplifiable region. For example, the terminal may be a terminal belonging to a user account that utilizes an amplifiable region library service or an oligonucleotide design service provided by the computer device (1000). In one embodiment, the user terminal (3000) may communicate with the computer device (1000) via an application programmed to implement the aforementioned service, and may receive information on an amplifiable region via the application.

[0349] Figure 16 illustrates an exemplary flowchart for providing information on amplifiable regions available for oligonucleotide design for detection of target nucleic acid molecules by a computer device (1000) according to a third embodiment. The details thereof can be understood by reference to all embodiments described above, and any redundant descriptions will be omitted.

[0350] Referring to FIG. 16, in step 1610, the computer device (1000) can determine amplifiable regions based on a plurality of nucleic acid sequences of the target nucleic acid molecule.

[0351] In one embodiment, the process of determining amplifiable regions can be performed at least in part using a predictive model learned to predict amplification efficiency using feature data that influences the amplification reaction of a nucleic acid sequence.

[0352] In one embodiment, the computer device (1000) can determine a region of interest (40) based on alignment results (20) of a plurality of nucleic acid sequences (10) of a target nucleic acid molecule. The computer device (1000) can obtain feature data (720) predicted to affect an amplification reaction in the region of interest (40). The computer device (1000) can provide the obtained feature data (720) as input data to a prediction model (710) and obtain output data including an amplification reaction efficiency (730) predicted for an amplification reaction in the region of interest (40) from the prediction model (710). The computer device (1000) can determine an amplifiable region (70) for detection of the target nucleic acid molecule in the alignment results (20) based on the amplification reaction efficiency (730).

[0353] At step 1620, the computer device (1000) can determine the priority for use in designing the oligonucleotides of the amplifiable regions.

[0354] In one embodiment, the priority may be determined based on thermodynamic data for the formation of an n-order structure in each of the amplifiable regions. In one embodiment, the n-order structure may comprise at least one selected from the group consisting of a hairpin loop, an internal loop, a bulge loop, multi-loops, a G-quadruplex, and combinations thereof.

[0355] In one embodiment, the thermodynamic data may include thermodynamic data for an n-order structure present in a nucleic acid sequence of an amplicon region determined from each amplifiable region, and / or thermodynamic data for an n-order structure present in a nucleic acid sequence of an oligo bidding region determined from the amplicon region.

[0356] In one embodiment, the computer device (1000) may further include at least one selected from the group consisting of (a) distance data between a predetermined position in the nucleic acid sequence of each amplifiable region and a position of the n-th structure; (b) GC content of the nucleic acid sequence of each amplifiable region; (c) a mutation score for each amplifiable region; and (d) a conservation score for each amplifiable region to determine the priority.

[0357] In one embodiment, the computer device (1000) may further include the amplification reaction efficiency output from the prediction model to determine priorities. For example, the amplification reaction efficiency of a region of interest may be output from the prediction model, and the computer device (1000) may determine priorities of amplifiable regions based on the amplification reaction efficiency of regions of interest included in each amplifiable region. In another example, the amplification reaction efficiency of amplifiable regions may be output from the prediction model, and the computer device (1000) may determine priorities of amplifiable regions based on the amplification reaction efficiency of each amplifiable region.

[0358] In one embodiment, the computer device (1000) can determine the priority using at least one selected from the group consisting of (i) numerical data included in the amplification reaction efficiency, (ii) values ​​of each of the plurality of features included in the feature data, and (iii) contributions of each of the plurality of features in the prediction model to the output of the amplification reaction efficiency.

[0359] In one embodiment, the computer device (1000) may calculate a score for a region of interest based on numerical data included in the amplification reaction efficiency. The computer device (1000) may calculate a score for each amplifiable region based on the score of the region of interest included in each amplifiable region. The computer device (1000) may determine a priority based on the score of each amplifiable region.

[0360] In step 1630, the computer device (1000) may store information on amplifiable regions with priorities in a database (2000). In one embodiment, amplifiable regions are determined for each of a plurality of target nucleic acid molecules, and an amplifiable region library including information on amplifiable regions for each target nucleic acid molecule may be stored and managed in the database (2000).

[0361] The computer device (1000) can provide information on amplifiable areas to the user terminal (3000) based on the amplifiable area library.

[0362] Specifically, the computer device (1000) can receive a user input for a target nucleic acid molecule from the user terminal (3000). In one embodiment, the computer device (1000) can receive a user input for an identifier of the target nucleic acid molecule and / or an identifier of the organism. For example, a list of organisms and / or a list of target nucleic acid molecules may be displayed on the user terminal (3000) based on an amplifiable region library categorized by target nucleic acid molecule for each organism, and the computer device (1000) can receive a selection input for a specific target nucleic acid molecule in the list through the amplifiable region library. As another example, a search window may be displayed on the user terminal (3000), and the computer device (1000) can receive a keyword input for a name of the target nucleic acid molecule and / or a taxonomic name (e.g., species) of the organism through the search window.

[0363] The computer device (1000) can retrieve information on amplifiable regions associated with the user input from the database (2000). For example, the computer device (1000) can retrieve information on amplifiable regions associated with a target nucleic acid molecule selected from the list from the amplifiable region library. As another example, the computer device (1000) can retrieve information on amplifiable regions having indexing items or metadata items corresponding to the keyword input through a search engine.

[0364] The computer device (1000) may provide information on amplifiable areas in order of priority to the user terminal (3000) based on the search results. In one embodiment, the computer device (1000) may provide information on amplifiable areas selected in order of priority from among information on all searched amplifiable areas.

[0365] In one embodiment, there may be multiple amplifiable areas, each with a priority level. For example, information about multiple amplifiable areas with a priority level higher than a predetermined level may be provided to the user terminal (3000). In another example, information about a specific number of amplifiable areas, ranked from highest to lowest priority, may be provided to the user terminal (3000).

[0366] Referring again to FIG. 10, for example, a result screen including information on amplifiable regions according to priority may be output on the user terminal (3000). In one embodiment, the result screen may include information on the amplifiable regions described above, such as the priority, location, sequence, and score of each amplifiable region, and may further include information on target nucleic acid molecules (e.g., gene name, taxonomic name or identifier of the organism, etc.), features of characteristic data used in the process of determining each amplifiable region, their values, algorithms, etc.

[0367] In one embodiment, the information on the amplifiable regions on the results screen may be sorted by priority. For example, the information may be sorted from highest to lowest priority. In one embodiment, if the information on each amplifiable region includes information on regions of interest, the information on each region of interest may include probability values ​​for each class and may be sorted according to the scores of each region of interest. In one embodiment, the results screen may further include information on the contribution of each of the multiple features used in the prediction model.

[0368] The above result screen may be displayed in the form of a table, graph, or image, depending on the implementation method, and the types and scales of these tables and graphs may be different.

[0369] Figure 17 is a diagram illustrating a contribution according to one embodiment.

[0370] Referring to Figure 17, the computer device (1000) can output a comparison screen containing the contributions of multiple features. These contributions can be output in various ways, for example, in the form of various graphs. The contributions of each feature can be processed into a comparable form and displayed on a single screen. For example, the contribution is calculated to be greater in the negative direction as the degree of contribution increases.

[0371] Fig. 17(a) illustrates the contribution of each feature in the first prediction model. For example, the first feature data described above include the type of reverse primer (e.g., rev_type), the type of forward primer (e.g., fwd_type), the distance data between the start position of the amplicon sequence and the position of the n-th structure (e.g., fwd_score), the Gibbs free energy change for the state in which the n-th structure is formed in the forward amplicon sequence (e.g., fwd_overlap_dG), the Gibbs free energy change for the possibility of forming the n-th structure in the amplicon sequence (e.g., amp_dG), and the corresponding reaction medium (e.g., enzyme) and their respective contribution values. As illustrated, in the first prediction model, it can be seen that the contribution of the above-mentioned Gibbs free energy change and distance data is greater than that of other features.

[0372] Fig. 17(b) illustrates the contribution of each feature in the second prediction model. For example, the second feature data described above include the length in the state where an n-order structure is formed in the oligo bidding sequence (e.g., primer_overlap_len), the GC content of the oligo bidding sequence (e.g., primer_GC), the type of oligonucleotide (e.g., primer_type), the predicted value for Tm of the oligo bidding sequence (e.g., primer_mtm), and the Gibbs free energy change for the possibility of forming an n-order structure in the oligo bidding sequence (e.g., primer_dG), and their respective contribution values. As illustrated, in the second prediction model, the contributions of the length and GC content are greater than those of other features.

[0373] Accordingly, rather than being provided with information on all amplifiable regions, users can selectively obtain information on amplifiable regions that are evaluated to have relatively higher amplification efficiency. Consequently, this increases detection accuracy while facilitating the subsequent steps of oligonucleotide design in a cost- and time-efficient manner.

[0374] In one embodiment, the priority is determined based on scores of a plurality of preset items calculated for each of the amplifiable areas, and information on the amplifiable area(s) according to the priority may be updated based on selection inputs for the plurality of items received from the user terminal (3000). For example, a result screen including information on the amplifiable areas according to the priority calculated based on the scores of the first to fourth items described above may be provided to the user terminal (3000). In addition, a selection input from the user terminal (3000) for one or more of the first to fourth items may be received to update the priority, and for example, when the first and second items are selected, the result screen may be updated to include information on the amplifiable areas according to the priority calculated again based on the first and second items.

[0375] In one embodiment, the computer device (1000) may receive feedback information regarding the provided amplifiable region information from the user terminal (3000) and update the priority based on the received feedback information. For example, a design process for an oligonucleotide for detecting a corresponding target nucleic acid molecule may be performed using the provided amplifiable region information. Furthermore, as this design process is performed, an oligonucleotide design result using the amplifiable region information may be obtained. This design result may include, for example, whether the designed oligonucleotide performs well or a numerical value for the performance.

[0376] In one embodiment, the computer device (1000) may change the weights used to determine priorities, update the priorities of the amplifiable regions, or perform additional learning on the prediction model used to determine the amplifiable regions based on the results of oligonucleotide design using the amplifiable regions included in the feedback information. For example, if the feedback information includes a relatively low feedback score, the computer device (1000) may lower the priority, delete the amplifiable region from the list, or place it at the bottom. In addition, if the feedback score is low even though the score of the amplifiable region is relatively high, the computer device (1000) may perform additional learning on the prediction model based on the data.

[0377] A design process of an oligonucleotide for detection of a target nucleic acid molecule according to one embodiment includes a plurality of steps, and the plurality of steps may include (a) a process of determining an amplifiable region for the target nucleic acid molecule, (b) a process of obtaining information on the amplifiable region, and (c) a design process based on the amplifiable region.

[0378] If information on an amplifiable region associated with a user input for a target nucleic acid molecule is not retrieved from the database (2000), the computer device (1000) may determine a target nucleic acid molecule associated with the user input and determine amplifiable regions based on a plurality of nucleic acid sequences of the determined target nucleic acid molecule. For example, steps (a), (b), and (c) of the above-described plurality of processes may be sequentially performed from the beginning.

[0379] When information on an amplifiable region associated with a user input for a target nucleic acid molecule is retrieved from a database (2000), the computer device (1000) may perform an oligonucleotide design process based on an amplifiable region selected by the user terminal (3000) from among a plurality of amplifiable regions according to priority. For example, among the above-described plurality of processes, processes (a) to (b) may be omitted, and process (c) may be subsequently performed based on information on an amplifiable region selected by the user.

[0380] These technical features may be used independently, depending on the embodiment, without being combined with the method for determining the amplifiable region and / or the method for obtaining the prediction model.

[0381]

[0382] According to one aspect of the present invention, a computer device (1000) including a memory (100) storing at least one command and a processor (300) is provided, and by executing at least one command by the processor (300), a region of interest used for determining an amplifiable region is determined based on alignment results of a plurality of nucleic acid sequences, feature data predicted to affect an amplification reaction in the region of interest is obtained, the feature data is provided as input data to a prediction model, output data including an amplification reaction efficiency predicted for an amplification reaction in the region of interest is obtained from the prediction model, and based on the amplification reaction efficiency, an amplifiable region for detection of a target nucleic acid molecule can be determined from the alignment results.

[0383] Each component described in one aspect of the present invention described above overlaps with the method for determining an amplifiable region described with reference to FIGS. 1 to 10, and thus, description thereof is omitted.

[0384] According to one aspect of the present invention, a computer-readable recording medium storing a computer program is provided, wherein the computer program comprises instructions that, when executed by one or more processors, cause the one or more processors to perform a method for determining an amplifiable region available for designing an oligonucleotide from a plurality of nucleic acid sequences of a target nucleic acid molecule, the method may include: determining a region of interest used for determining an amplifiable region based on alignment results of the plurality of nucleic acid sequences; obtaining feature data predicted to affect an amplification reaction in the region of interest; providing the feature data as input data to a prediction model; obtaining output data including an amplification reaction efficiency predicted for an amplification reaction in the region of interest from the prediction model; and determining an amplifiable region for detection of the target nucleic acid molecule from the alignment results based on the amplification reaction efficiency.

[0385] Each component described in one aspect of the present invention described above overlaps with the method for determining an amplifiable region described with reference to FIGS. 1 to 10, and thus, description thereof is omitted.

[0386] According to one aspect of the present invention, a computer device (1000) including a memory (100) storing at least one command and a processor (300) is provided, wherein the at least one command is executed by the processor (300), thereby obtaining a plurality of training data sets, and obtaining a prediction model learned to predict an amplification reaction efficiency for an amplification reaction using the plurality of training data sets, wherein each of the plurality of training data sets includes (a) training input data including feature data predicted to affect an amplification reaction of a training nucleic acid sequence, and (b) training answer data including an amplification reaction efficiency for an amplification reaction in the training nucleic acid sequence, and the feature data included in the training input data may include thermodynamic data for the formation of an n-th order structure in the training nucleic acid sequence.

[0387] Each component described in one aspect of the present invention described above overlaps with the method for obtaining a prediction model described with reference to FIGS. 11 to 14, and thus, description thereof is omitted.

[0388] According to one aspect of the present invention, a computer-readable recording medium storing a computer program is provided, wherein the computer program includes instructions that, when executed by one or more processors, cause the one or more processors to perform a method for obtaining a prediction model that provides an amplification reaction efficiency predicted for an amplification reaction, the method comprising: obtaining a plurality of training data sets; and obtaining a prediction model learned to predict the amplification reaction efficiency for the amplification reaction using the plurality of training data sets, wherein each of the plurality of training data sets includes (a) training input data including feature data predicted to affect the amplification reaction of a training nucleic acid sequence, and (b) training answer data including the amplification reaction efficiency for the amplification reaction in the training nucleic acid sequence, and the feature data included in the training input data may include thermodynamic data for the formation of an n-order structure (wherein n is an integer equal to or greater than 2) in the training nucleic acid sequence.

[0389] Each component described in one aspect of the present invention described above overlaps with the method for obtaining a prediction model described with reference to FIGS. 11 to 14, and thus, description thereof is omitted.

[0390] According to one aspect of the present invention, a computer device (1000) is provided, which includes a memory (100) storing at least one command and a processor (300), and by executing at least one command by the processor (300), amplifiable regions are determined based on a plurality of nucleic acid sequences of a target nucleic acid molecule, priorities for use in designing oligonucleotides of the amplifiable regions are determined, and information on the amplifiable regions including the priorities is stored in a database, wherein the priorities can be determined based on thermodynamic data on the formation of an n-th order structure in each of the amplifiable regions.

[0391] Each component described in one aspect of the present invention described above overlaps with the method of providing information on an amplifiable region described with reference to FIGS. 15 to 17, and thus, description thereof is omitted.

[0392] According to one aspect of the present invention, a computer-readable recording medium storing a computer program is provided, wherein the computer program comprises instructions that, when executed by one or more processors, cause the one or more processors to perform a method for providing information of amplifiable regions available for design of an oligonucleotide for detection of a target nucleic acid molecule, the method comprising: determining amplifiable regions based on a plurality of nucleic acid sequences of the target nucleic acid molecule; determining priorities for use of the amplifiable regions in the design of the oligonucleotide; and storing information of the amplifiable regions including the priorities in a database, wherein, in the step of determining the priorities, the priorities may be determined based on thermodynamic data for the formation of an n-th order structure in each of the amplifiable regions.

[0393] Each component described in one aspect of the present invention described above overlaps with the method of providing information on an amplifiable region described with reference to FIGS. 15 to 17, and thus, description thereof is omitted.

[0394] Meanwhile, those skilled in the art will appreciate that the various exemplary logical blocks, modules, processors, means, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented by electronic hardware, various forms of programs or design code (referred to herein for convenience as software), or a combination of both. To clearly illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0395] The various embodiments presented herein can be implemented as a method, apparatus, or article of manufacture using standard programming and / or engineering techniques. The term article of manufacture includes a computer program, carrier, or media accessible from any computer-readable storage device. For example, computer-readable storage media include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic strips, etc.), optical disks (e.g., CDs, DVDs, etc.), smart cards, and flash memory devices (e.g., EEPROMs, cards, sticks, key drives, etc.). Furthermore, various storage media presented herein include one or more devices and / or other machine-readable media for storing information.

[0396] It should be understood that the specific order or hierarchy of steps in the presented processes is merely an example of exemplary approaches. It should be understood that the specific order or hierarchy of steps in the processes may be rearranged within the scope of the present disclosure based on design priorities. The appended method claims provide elements of various steps in a sample order, but are not intended to be limited to the specific order or hierarchy presented.

Claims

1. A method for determining an amplifiable region available for designing an oligonucleotide from a plurality of nucleic acid sequences of a target nucleic acid molecule, performed by a computer device, A step of determining a region of interest used for determining an amplifiable region based on the alignment results of multiple nucleic acid sequences; A step of acquiring feature data that is predicted to affect the amplification response in the above region of interest; A step of providing the above feature data as input data to a prediction model; A step of obtaining output data including an amplification reaction efficiency predicted for an amplification reaction in the region of interest from the above prediction model; and A step of determining the amplifiable region for detection of the target nucleic acid molecule from the alignment result based on the amplification reaction efficiency, method.

2. In paragraph 1, The above area of ​​interest is (a) a conserved region included in the alignment result, (b) a fragmented region of a predetermined length divided from the conserved region, or (c) a fragmented region of a predetermined length divided from the entire region of the plurality of nucleic acid sequences, method.

3. In paragraph 1, The steps to determine the above area of ​​interest are A step of obtaining a conservative region including sequences having a sequence similarity of a predetermined level or higher from the alignment result; A step of determining a plurality of fragment regions of a predetermined length that are sequentially fragmented or fragmented so that some of them overlap each other within the above-mentioned conservative region; and comprising a step of determining each of the plurality of fragment regions as the region of interest; method.

4. In paragraph 3, The above specified length is 50 bp or more and 150 bp or less, The above mutually overlapping length is 10 bp or more and 50 bp or less, method.

5. In paragraph 3, The step of obtaining the above output data is: Obtain output data including predicted amplification reaction efficiency for the amplification reaction in each of the above multiple fragment regions, The step of determining the amplifiable area is: Based on the amplification reaction efficiency of each of the plurality of fragment regions, the amplifiable region is determined within the conservative region. method.

6. In paragraph 1, The step of obtaining the above feature data is Generate the feature data using information of the above area of ​​interest, Information on the above areas of interest Including at least one selected from the group consisting of a nucleic acid sequence of the region of interest, a length of the region of interest, a biological category of an organism having the target nucleic acid molecule, a nucleic acid type of the target nucleic acid molecule, a gene type of the target nucleic acid molecule, and reaction conditions for the amplification reaction. method.

7. In paragraph 6, The nucleic acid sequence of the above region of interest is (a) a sequence determined from a nucleic acid sequence satisfying a predetermined length condition among the plurality of nucleic acid sequences, (b) a sequence determined from a unique genome sequence of an organism having the target nucleic acid molecule, (c) a sequence determined based on a grouping result for the plurality of nucleic acid sequences, and (d) a sequence determined from a nucleic acid sequence containing the fewest non-conserved bases among the plurality of nucleic acid sequences, comprising at least one selected from the group consisting of method.

8. In paragraph 1, The above feature data is (a) a nucleic acid sequence of the region of interest; (b) thermodynamic data on the formation of an n-order structure (wherein n is an integer greater than or equal to 2) in the nucleic acid sequence of the region of interest; (c) distance data between a predetermined position in the nucleic acid sequence of the region of interest and a position of the n-order structure; (d) reaction conditions including a reaction medium used in the amplification reaction; the reaction medium including at least one material selected from the group consisting of a pH-related material, an ionic strength-related material, an enzyme, and an enzyme stabilization-related material; (f) a type of oligonucleotide that binds to the nucleic acid sequence of the region of interest; (g) a GC content of the nucleic acid sequence of the region of interest; (h) a variation score for the region of interest; and (i) at least one selected from the group consisting of a conservation score for the region of interest. method.

9. In paragraph 8, The above n is 2, The above n-th structure comprises at least one selected from the group consisting of a hairpin loop, an internal loop, a bulge loop, multi-loops, a G-quadruplex, and a combination thereof. method.

10. In paragraph 8, The above thermodynamic data Expressed in the form of change in thermodynamic free energy, method.

11. In paragraph 8, The above thermodynamic data Including thermodynamic data for the n-th structure existing in the nucleic acid sequence of the amplicon region corresponding to the region of interest, and / or thermodynamic data for the n-th structure existing in the nucleic acid sequence of the oligo bidding region determined from the amplicon region. method.

12. In paragraph 8, The above mutation score for the above region of interest is (a) a mutation position at which a non-conservative base is located in the region of interest of the plurality of nucleic acid sequences, (b) a mutation type indicating the type of the non-conservative base, (c) the number of the non-conservative base and / or the mutation type at the mutation position, and (d) the number of the mutation positions in the region of interest of the plurality of nucleic acid sequences, which is calculated based on at least one selected from the group consisting of method.

13. In paragraph 8, The above conservation score for the above area of ​​interest is Based on the entropy value that measures the degree of uncertainty from the probability distribution of the similarity of base sequences in the region of interest among the plurality of nucleic acid sequences, method.

14. In paragraph 1, The step of obtaining the above feature data is It is performed by a preprocessing unit configured to calculate the value of each of the plurality of features included in the feature data using a plurality of previously stored mathematical formulas, or it is performed by a model trained to output the value of each of the plurality of features when information of the region of interest is input. method.

15. In paragraph 14, The above learned model is A model included in the above prediction model, characterized in that it operates integrally by the prediction model, or a model distinct from the prediction model, characterized in that it operates independently from the prediction model. method.

16. In paragraph 1, The above prediction model was trained using multiple learning datasets. Each of the above multiple learning datasets Including training input data including feature data predicted to affect an amplification reaction for a training nucleic acid sequence and training answer data including an amplification reaction efficiency for an amplification reaction for the training nucleic acid sequence. method.

17. In paragraph 16, The above feature data included in the above learning input data is (a) the learning nucleic acid sequence; (b) thermodynamic data on the formation of an n-order structure (wherein n is an integer greater than or equal to 2) in the learning nucleic acid sequence; (c) distance data between a predetermined position in the learning nucleic acid sequence and a position of the n-order structure; (d) reaction conditions including a reaction medium used for an amplification reaction for the learning nucleic acid sequence; the reaction medium including at least one material selected from the group consisting of a pH-related material, an ionic strength-related material, an enzyme, and an enzyme stabilization-related material; (e) at least one selected from the group consisting of a type of oligonucleotide bound to the learning nucleic acid sequence, (f) a GC content of the learning nucleic acid sequence, (g) a variation score for a learning sequence group including the learning nucleic acid sequence, and (h) a conservation score for the learning sequence group. method.

18. In paragraph 16, The amplification reaction efficiency included in the above learning answer data is Including numerical data on the efficiency level of the amplification reaction for the above learning nucleic acid sequence and / or label data on the high and low of the efficiency level, method.

19. In paragraph 16, The amplification reaction efficiency included in the above learning answer data is (a) the difference in amplification points determined from two or more data sets obtained from two or more amplification reactions for the above learning nucleic acid sequence, or (b) calculated using the difference between the signal pattern determined from the two or more data sets and a preset reference pattern. method.

20. In paragraph 1, The above areas of interest are multiple, The step of determining the amplifiable area is Determining the amplifiable region by using one or more regions of interest among the plurality of regions of interest whose amplification reaction efficiency satisfies a preset first criterion, method.

21. In paragraph 20, The step of determining the amplifiable area is (a) if a region merged from one or more regions of interest satisfies a preset second criterion, the merged region is determined as the amplifiable region, or (b) each of the one or more regions of interest is determined as the amplifiable region. method.

22. In paragraph 21, The above second criterion is (a) a criterion for the length of said merged region, and / or (b) a criterion for a statistical operation value of said amplification reaction efficiency of each of said one or more regions of interest, method.

23. In paragraph 20, The step of determining the amplifiable area is A step of re-determining the region of interest so that the length of the region of interest is reduced, when there is no region of interest among the plurality of regions of interest whose amplification reaction efficiency satisfies the first criterion; A step of re-acquiring the amplification reaction efficiency by re-performing the step of acquiring the feature data or the step of acquiring the output data based on the re-determined region of interest; and A step of determining the amplifiable region using one or more regions of interest among the re-determined regions of interest whose re-acquired amplification reaction efficiency satisfies the first criterion, method.

24. In paragraph 20, In the step of acquiring the above feature data, Obtaining first feature data and second feature data each affecting the amplification reaction in the amplicon region and oligo bidding region determined from the above region of interest, respectively, In the step of providing the above feature data to the above prediction model, Providing each of the first feature data and the second feature data as input data to the prediction model, In the step of obtaining the above output data, From the above prediction model, the first amplification reaction efficiency and the second amplification reaction efficiency predicted for the amplification reaction in each of the amplicon region and the oligo bidding region are obtained, respectively. The amplification reaction efficiency in the region of interest is determined based on the first amplification reaction efficiency and the second amplification reaction efficiency. method.

25. In paragraph 24, The above prediction model includes a first prediction model learned to predict an amplification reaction efficiency for an amplification reaction in an amplicon sequence using first training datasets including thermodynamic data for an n-th structure existing in an amplicon sequence, and / or a second prediction model learned to predict an amplification reaction efficiency for an amplification reaction in an oligo bidding sequence using second training datasets including thermodynamic data for an n-th structure existing in an oligo bidding sequence determined from the amplicon sequence. In the step of obtaining the above output data, Obtaining the first amplification reaction efficiency from the first prediction model, and obtaining the second amplification reaction efficiency from the second prediction model. method.

26. In paragraph 1, Further comprising a step of obtaining and providing information of the amplifiable area, Information on the amplifiable area above (a) information on the location, base sequence, score, and bindable oligonucleotide of the region of interest included in the amplifiable region, (b) at least one selected from the group consisting of information on the location, base sequence, score, and bindable oligonucleotide of the amplifiable region, method.

27. In paragraph 26, The score of the region of interest is calculated using at least one selected from the group consisting of (i) numerical data included in the amplification reaction efficiency of the region of interest, (ii) the value of each of the plurality of features included in the feature data, and (iii) the contribution of each of the plurality of features in the prediction model to the output of the amplification reaction efficiency. The score of the amplifiable area is calculated based on the score of the area of ​​interest included in the amplifiable area. method.

28. In paragraph 26, The score of the above amplifiable area includes a score for each of a plurality of preset items, The above multiple items include at least one item selected from the group consisting of (i) a first amplification reaction efficiency in an amplicon region included in the amplification reaction efficiency, (ii) a second amplification reaction efficiency in an oligo binding region included in the amplification reaction efficiency, (iii) a variation score for the amplifiable region, and (iv) a conservation score for the amplifiable region. method.

29. In paragraph 1, The above prediction model is, Containing at least one selected from the group consisting of RF (Random Forest), LR (Logistic Regressor), GBC (Gradient Boosting Classifier), DTC (Decision Tree Classifier), GNB (Gaussian Naive Bayes), SVC (Support Vector Classifier), and NN (Neural Network). method.

30. Memory for storing at least one instruction; and A processor comprising at least one instruction that, when executed by the processor, causes the processor to perform the following operations, the operations comprising: An operation of determining a region of interest used for determining an amplifiable region based on the alignment results of multiple nucleic acid sequences of a target nucleic acid molecule; An operation of acquiring feature data that is predicted to affect the amplification response in the above region of interest; An action of providing the above feature data as input data to a prediction model, An operation of obtaining output data including an amplification reaction efficiency predicted for an amplification reaction in the region of interest from the above prediction model, and An operation of determining the amplifiable region for detection of the target nucleic acid molecule from the alignment result based on the amplification reaction efficiency, Computer devices.

31. A computer-readable non-transitory recording medium having stored thereon a computer program, wherein the computer program comprises instructions that, when executed by one or more processors, cause the one or more processors to perform a method for determining an amplifiable region available for designing an oligonucleotide from a plurality of nucleic acid sequences of a target nucleic acid molecule, the method comprising: A step of determining a region of interest used for determining an amplifiable region based on the alignment results of multiple nucleic acid sequences; A step of acquiring feature data that is predicted to affect the amplification response in the above region of interest; A step of providing the above feature data as input data to a prediction model; A step of obtaining output data including an amplification reaction efficiency predicted for an amplification reaction in the region of interest from the above prediction model; and A step of determining the amplifiable region for detection of the target nucleic acid molecule from the alignment result based on the amplification reaction efficiency, A computer-readable recording medium storing a computer program.

32. A method for obtaining a prediction model that provides predicted amplification reaction efficiency for an amplification reaction performed by a computing device, A step of obtaining multiple learning datasets; and A step of obtaining a prediction model learned to predict the amplification reaction efficiency for the amplification reaction using the above plurality of learning datasets, Each of the plurality of training data sets includes (a) training input data including feature data predicted to affect an amplification reaction of a training nucleic acid sequence and (b) training answer data including an amplification reaction efficiency for an amplification reaction in the training nucleic acid sequence, and the feature data included in the training input data includes thermodynamic data for the formation of an n-dimensional structure (wherein n is an integer greater than or equal to 2) in the training nucleic acid sequence. method.

33. In paragraph 32, The above feature data included in the above learning input data is (a) the learning nucleic acid sequence; (b) distance data between a predetermined position in the learning nucleic acid sequence and a position of the n-th structure; (c) reaction conditions including a reaction medium used for an amplification reaction for the learning nucleic acid sequence; the reaction medium includes at least one material selected from the group consisting of a pH-related material, an ionic strength-related material, an enzyme, and an enzyme stabilization-related material; (d) further including at least one selected from the group consisting of a type of oligonucleotide bound to the learning nucleic acid sequence, (e) a GC content of the learning nucleic acid sequence, (f) a variation score for a learning sequence group including the learning nucleic acid sequence, and (g) a conservation score for the learning sequence group. method.

34. In paragraph 32, The amplification reaction efficiency included in the above learning answer data is (a) the difference in amplification points determined from two or more data sets obtained from two or more amplification reactions for the above learning nucleic acid sequence, or (b) calculated using the difference between the signal pattern determined from the two or more data sets and a preset reference pattern. method.

35. In paragraph 32, The above thermodynamic data Including first thermodynamic data for the n-th structure existing in the amplicon sequence determined from the learning nucleic acid sequence and second thermodynamic data for the n-th structure existing in the oligo bidding sequence determined from the amplicon sequence, The steps for obtaining the above prediction model are A step of obtaining a first prediction model learned to predict the amplification reaction efficiency for the amplification reaction in the amplicon sequence by using a plurality of first learning datasets including the first thermodynamic data among the plurality of learning datasets; and A step of obtaining a second prediction model learned to predict the amplification reaction efficiency for the amplification reaction in the oligo bidding sequence by using a plurality of second learning datasets including the second thermodynamic data among the plurality of learning datasets, method.

36. A method for providing information of an amplifiable region available for designing an oligonucleotide for detection of a target nucleic acid molecule, performed by a computing device, A step of determining amplifiable regions based on multiple nucleic acid sequences of the target nucleic acid molecule; A step of determining the priority for use in the design of the oligonucleotide of the amplifiable regions; and Including a step of storing information of the amplifiable areas including the above priorities in a database, The above priorities are: Determined based on thermodynamic data for the formation of n-th order structures (where n is an integer greater than or equal to 2) in each of the above amplifiable regions, method.

37. In paragraph 36, The above n is 2, The above n-th structure comprises at least one selected from the group consisting of a hairpin loop, an internal loop, a bulge loop, multi-loops, a G-quadruplex, and a combination thereof. method.

38. In paragraph 36, The above priorities are: (a) distance data between a predetermined position in the nucleic acid sequence of each amplifiable region and the position of the n-th structure; (b) GC content of the nucleic acid sequence of each amplifiable region; (c) variation score for each amplifiable region; and (d) conservation score for each amplifiable region, which is determined by using at least one selected from the group consisting of: method.

39. In paragraph 36, The step of determining the above amplifiable areas is: It is performed at least in part by using a prediction model learned to predict the efficiency of an amplification reaction using feature data that affects the amplification reaction of a nucleic acid sequence, The above priorities are: It is further determined by utilizing the amplification reaction efficiency output from the above prediction model. method.

40. In paragraph 39, The above priorities are: (i) numerical data included in the amplification reaction efficiency, (ii) values ​​of each of the plurality of features included in the feature data, and (iii) a contribution degree of each of the plurality of features in the prediction model to the output of the amplification reaction efficiency, which is determined using at least one selected from the group consisting of method.

41. In paragraph 36, A step of receiving a user input for a target nucleic acid molecule from a user terminal; A step of retrieving information of an amplifiable area associated with the user input from the database; and Based on the search results, further comprising a step of providing information on the amplifiable area according to the priority to the user terminal. method.

42. In paragraph 41, The amplifiable area according to the above priority may be multiple, method.

43. In paragraph 41, The above priority is determined based on the scores of multiple preset items calculated for each of the above amplifiable areas, Information on the amplifiable area according to the above priority updated based on the selection input for the plurality of items received from the user terminal, method.

44. In paragraph 41, A step of receiving feedback information about the information of the provided amplifiable area from the user terminal; and Further comprising a step of updating the priority based on the feedback information. method.

45. In paragraph 44, The steps to update the above priorities are Based on the oligonucleotide design results using the amplifiable region included in the feedback information, the weights used to determine the priorities are changed, the priorities of the amplifiable regions are updated, or additional learning is performed on the prediction model used to determine the amplifiable region. method.

46. ​​In paragraph 42, Further comprising a step of performing a design process of the oligonucleotide based on an amplifiable region selected by the user terminal from among the plurality of amplifiable regions according to the above priorities. method.

Citation Information

Patent Citations

  • Notification service server capable of providing access notification service to harmful sites and operating method thereof

    KR102421572B1

  • Artificial intelligence-based chromosomal abnormality detection method

    WO2021107676A1

  • Methods and devices for predicting dimerization in nucleic acid amplification reaction

    WO2024072164A1