Method and system for determining optimal chemical modifications to the base sequence of an RNA therapeutic agent

The method employs an artificial neural network to determine optimal chemical modifications for RNA therapeutic agents, addressing the challenge of regulating target mRNA activity while minimizing side effects and improving therapeutic efficiency and stability.

JP2025519069AActive Publication Date: 2025-06-24SPIDERCORE INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024568482
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-02-09
Filing Date
2023-05-25
Publication Date
2025-06-24
Estimated Expiration
2043-05-25

AI Technical Summary

Technical Problem

RNA therapeutic agents face challenges in effectively regulating target mRNA activity while minimizing side effects on non-target mRNAs, and in determining optimal chemical modifications for improved therapeutic efficiency and stability.

Method used

A method and system using an artificial neural network to predict optimal chemical modifications for RNA therapeutic agents, based on learning data from chemical modification information databases, to enhance therapeutic efficacy and in vivo stability while reducing side effects.

Benefits of technology

The approach allows for the design of RNA therapeutic agents with improved therapeutic effects, reduced side effects, and enhanced in vivo stability by identifying optimal chemical modifications for specific base sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025519069000001_ABST
    Figure 2025519069000001_ABST
Patent Text Reader

Abstract

In a method for determining an optimal chemical modification for a base sequence of an RNA therapeutic agent, a base sequence modification module acquires, as learning data, biological properties when a plurality of chemical modifications are applied to a plurality of base sequences. Among the learning data, at least two sequences to which different chemical modifications are applied to the same base sequence are randomly selected and sequentially input into an artificial neural network, and the output values output for each of the at least two input sequences are compared. The process of training the artificial neural network is repeatedly performed so that the artificial neural network outputs a larger value as the biological properties of the input sequence are better, to generate an optimal chemical modification prediction model. Using the optimal chemical modification prediction model, an optimal chemical modification having excellent biological properties when applied to the base sequence of the RNA therapeutic agent is determined among the first to w-th chemical modifications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the development of new drugs using RNA (Ribonucleic Acid), and more particularly to a method and system for determining an optimal chemical modification for the base sequence of an RNA therapeutic agent.

Background Art

[0002] DNA (Deoxyribonucleic Acid) containing genetic information is a polymer with a double helix structure composed of nucleotides. The genetic information contained in DNA is expressed as a specific protein through a complex process.

[0003] The base sequence contained in DNA contains information on the amino acids that make up a protein. In order for a specific protein to be expressed, a transcription process that copies a specific base sequence of DNA necessary for the expression of the specific protein into mRNA (Messenger Ribonucleic Acid) proceeds, and the copied mRNA binds to ribosomes necessary for protein synthesis, and a translation process in which amino acids corresponding to the base sequence of the mRNA are synthesized to form a protein proceeds. The protein generated by this process moves to a suitable position inside or outside the cell, and finally the gene expression is completed.

[0004] At this time, if the activity of mRNA copied from DNA is reduced, gene expression may be inhibited. Thus, a drug with a new mechanism using RNA that plays a role in transmitting genetic information in the protein formation process of cells is called an RNA therapeutic agent.

[0005] Such RNA therapeutic agents have emerged as a new solution for intractable diseases for which there were no therapeutic agents before, and recently, research on RNA therapeutic agents that can treat intractable diseases such as genetic diseases, cardiovascular diseases, and cancer has been actively promoted.

[0006] Generally, in order to develop an RNA therapeutic agent for a specific disease, after identifying the mRNA related to the synthesis of the protein that causes the specific disease, it is necessary to find an RNA base sequence that can effectively reduce the activity of the identified mRNA.

[0007] However, RNA therapeutic agents have a problem in that, in addition to their intended function of reducing the activity of the target mRNA, they also inhibit the synthesis of normal proteins by reducing the activity of other non-target mRNAs.

[0008] Also, even for RNA therapeutic agents with the same base sequence, when different chemical modifications are applied, the therapeutic efficiency and in vivo stability are different. Therefore, in order to develop an RNA therapeutic agent, it is also necessary to determine what chemical modifications to apply to the base sequence of the RNA therapeutic agent.

Summary of the Invention

Problems to be Solved by the Invention

[0009] Therefore, the technical problem of the present invention is considered in view of such points. One object of the present invention is to provide a base sequence of an RNA therapeutic agent that can effectively regulate the activity of the target mRNA while preventing side effects of regulating the activity of other non-target mRNAs, and to provide an optimal chemical modification that can be applied to the base sequence of the RNA therapeutic agent to improve the therapeutic efficiency and in vivo stability of the RNA therapeutic agent, thereby providing a design method for the RNA therapeutic agent.

[0010] Another object of the present invention is to provide a base sequence of an RNA therapeutic agent that can effectively regulate the activity of the target mRNA while preventing side effects of regulating the activity of other non-target mRNAs, and to provide an optimal chemical modification that can be applied to the base sequence of the RNA therapeutic agent to improve the therapeutic efficiency and in vivo stability of the RNA therapeutic agent, thereby providing a design system for the RNA therapeutic agent.

Means for Solving the Problems

[0011] In order to achieve one of the above-described objects of the present invention, in a method for determining an optimal chemical modification for a base sequence of an RNA therapeutic agent according to an embodiment of the present invention, a base sequence modification module acquires, as learning data, data stored in a chemical modification information database that stores numerical values indicating biological characteristics when a plurality of chemical modifications are applied to a plurality of base sequences, the base sequence modification module randomly selects at least two sequences in which different chemical modifications are applied to the same base sequence from among the learning data and sequentially inputs them into an artificial neural network, compares output values output for each of the at least two sequences input into the artificial neural network, and repeats a process of training the artificial neural network so that the artificial neural network outputs a larger value as the biological characteristics of the sequence input into the artificial neural network are better, thereby generating an optimal chemical modification prediction model, the base sequence modification module receives a base sequence of an RNA therapeutic agent for regulating the activity of a target mRNA (Messenger Ribonucleic Acid) related to the induction of a specific disease, and the base sequence modification module uses the optimal chemical modification prediction model to determine, as an optimal chemical modification, a chemical modification having excellent biological characteristics when applied to the base sequence of the RNA therapeutic agent among the first to w-th (w is an integer of 2 or more) chemical modifications.

[0012] In order to achieve one of the above-described objects of the present invention, a system for determining an optimal chemical modification for a base sequence of an RNA therapeutic agent according to an embodiment of the present invention acquires, as learning data, data stored in a chemical modification information database that stores numerical values indicating biological properties when a plurality of chemical modifications are applied to a plurality of base sequences, randomly selects at least two sequences in which different chemical modifications are applied to the same base sequence from among the learning data, and sequentially inputs them into an artificial neural network, compares output values output for each of the at least two sequences input into the artificial neural network, and repeats a process of training the artificial neural network so that the artificial neural network outputs a larger value as the biological properties of the sequence input into the artificial neural network are better, to generate an optimal chemical modification prediction model. When the base sequence modification module receives the base sequence of an RNA therapeutic agent for regulating the activity of a target mRNA (Messenger Ribonucleic Acid) related to the induction of a specific disease, it uses the optimal chemical modification prediction model to determine, as an optimal chemical modification, a chemical modification having excellent biological properties when applied to the base sequence of the RNA therapeutic agent, among the first to w-th (w is an integer of 2 or more) chemical modifications. [Effects of the Invention]

[0013] The RNA therapeutic agent design system according to an embodiment of the present invention not only provides a base sequence of an RNA therapeutic agent that can effectively prevent side effects while maximizing the therapeutic effect on a specific disease by reducing the activity of a target mRNA, but also provides an optimal chemical modification that can be applied to the base sequence of the RNA therapeutic agent to further improve the therapeutic efficiency and in vivo stability of the RNA therapeutic agent, so that an RNA therapeutic agent with high therapeutic effect, reduced side effects, and improved in vivo stability can be effectively designed. [Brief Description of the Drawings]

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Embodiments for Carrying Out the Invention

[0015] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the accompanying drawings. The same reference numerals are used for the same components on the drawings, and redundant descriptions for the same components are omitted.

[0016] FIG. 1 is a diagram showing an RNA therapeutic agent design system according to an embodiment of the present invention. The RNA therapeutic agent design system 10 shown in FIG. 1 is an RNA therapeutic agent design system that can provide a nucleotide sequence of an RNA therapeutic agent capable of effectively regulating the activity of target mRNA (Messenger Ribonucleic Acid) and an optimal chemical modification that can be applied to the nucleotide sequence of the RNA therapeutic agent to improve the therapeutic efficiency and in vivo stability of the RNA therapeutic agent.

[0017] Referring to FIG. 1, the RNA therapeutic agent design system 10 includes an off-target analysis module 100, a sequence generation module 200, and a sequence modification module 300.

[0018] The off-target analysis module 100 and the sequence generation module 200 receive, from the outside, a target mRNA (T_MRNA) related to the synthesis of a protein that causes a specific disease.

[0019] When the off-target analysis module 100 receives the target mRNA (T_MRNA), the off-target analysis module 100 determines, as a plurality of off-target mRNAs (OFF_T_MRNA), mRNAs having a gene expression pattern similar to the target mRNA (T_MRNA) among the plurality of mRNAs contained in the human body.

[0020] In one embodiment, the off-target analysis module 100 can cluster the plurality of mRNAs into mRNAs having similar gene expression patterns and classify them into first to k-th clusters by using a gene expression database (GE_DB) 110 that pre-stores the gene expression pattern of each of the plurality of mRNAs contained in the human body. Here, k represents an integer of 2 or more.

[0021] Thereafter, the off-target analysis module 100 can determine, as a plurality of off-target mRNAs (OFF_T_MRNA), the mRNAs contained in the cluster to which the target mRNA (T_MRNA) belongs among the first to k-th clusters. Here, the gene expression database 110 may be any database that stores the gene expression pattern of each of the plurality of mRNAs contained in the human body.

[0022] For example, the gene expression database 110 may be Gene Expression Omnibus (GEO), LINCS-L1000, Connectivity Map (CMAP), Human Protein Atlas (HPA), GTEx and FANTOM5, etc.

[0023] In one embodiment, the gene expression database 110 can store in advance the degree to which the expression levels of each of the plurality of mRNAs are regulated by each of the first to m-th drugs. Here, m represents an integer of 2 or more. Here, the first to m-th drugs may be any drugs known to regulate the expression levels of mRNAs contained in the human body.

[0024] In this case, the off-target analysis module 100 reads from the gene expression database 110 the degree to which the expression levels of each of the plurality of mRNAs are regulated by each of the first to m-th drugs, and for each of the plurality of mRNAs, can generate an m-dimensional expression level vector including the degree to which the expression level is regulated by each of the first to m-th drugs.

[0025] In one embodiment, the off-target analysis module 100 can generate first to n-th expression level vectors (V_1, V_2,..., V_n) including information on whether the expression level of each of the plurality of mRNAs significantly increases, significantly decreases, or does not significantly change by each of the first to m-th drugs.

[0026] Thereafter, the off-target analysis module 100 clusters the plurality of expression level vectors corresponding to the plurality of mRNAs into k clusters by grouping the expression level vectors existing at adjacent positions together, and classifies the mRNAs corresponding to the expression level vectors included in the same cluster into the same cluster, thereby classifying the plurality of mRNAs into the first to k-th clusters.

[0027] In one embodiment, the non-target analysis module 100 can classify the plurality of mRNAs into the first to k-th clusters by applying the K-means algorithm to the plurality of expression level vectors corresponding to the plurality of mRNAs to cluster the plurality of expression level vectors into k clusters.

[0028] In one embodiment, the non-target analysis module 100 can classify the plurality of mRNAs into the first to k-th clusters by applying the Gaussian Mixture Model (GMM) algorithm to the plurality of expression level vectors corresponding to the plurality of mRNAs to cluster the plurality of expression level vectors into k clusters.

[0029] In one embodiment, the non-target analysis module 100 can classify the plurality of mRNAs into the first to k-th clusters by applying the DBSCAN (Density Based Spatial Clustering of Applications with Noise) algorithm to the plurality of expression level vectors corresponding to the plurality of mRNAs to cluster the plurality of expression level vectors into k clusters.

[0030] In one embodiment, the non-target analysis module 100 can classify the plurality of mRNAs into the first to k-th clusters by applying the Mean Shift algorithm to the plurality of expression level vectors corresponding to the plurality of mRNAs to cluster the plurality of expression level vectors into k clusters.

[0031] In one embodiment, the non-target analysis module 100 can classify the plurality of mRNAs into the first to k-th clusters by applying the Agglomerative Hierarchical Clustering algorithm to the plurality of expression level vectors corresponding to the plurality of mRNAs to cluster the plurality of expression level vectors into k clusters.

[0032] According to an embodiment, the off-target analysis module 100 may classify the plurality of mRNAs into the first to k-th clusters by clustering the plurality of expression amount vectors into k clusters using other clustering algorithms known in the art.

[0033] According to an embodiment, the off-target analysis module 100 may classify the plurality of mRNAs into the first to k-th clusters by clustering the plurality of expression amount vectors into k clusters using a clustering model using an artificial neural network known in the art. FIG. 2 is a diagram for explaining an example of the operation of the off-target analysis module included in the RNA therapeutic agent design system of FIG. 1.

[0034] FIG. 2 shows, as an example, the operation of the off-target analysis module 100 when the gene expression database 110 stores the degree to which each of the first to n-th mRNAs (MRNA1, MRNA2,..., MRNAn) contained in the human body is regulated in expression amount by each of the first to m-th drugs (DR1, DR2,..., DRm). Here, n represents an integer of 2 or more.

[0035] Referring to FIG. 2, after the off-target analysis module 100 reads from the gene expression database 110 the degree to which each of the first to n-th mRNAs (MRNA1, MRNA2,..., MRNAn) is regulated in expression amount by each of the first to m-th drugs (DR1, DR2,..., DRm), for each of the first to n-th mRNAs (MRNA1, MRNA2,..., MRNAn), it can generate first to n-th expression amount vectors (V_1, V_2,..., V_n) including the degree to which the expression amount is regulated by each of the first to m-th drugs (DR1, DR2,..., DRm). Therefore, each of the first to n-th expression amount vectors (V_1, V_2,..., V_n) may be an m-dimensional vector.

[0036] In one embodiment, the off-target analysis module 100 can generate first to nth expression vectors (V_1, V_2,..., V_n) for each of the first to nth mRNAs (mRNA1, mRNA2,..., mRNA n), which contain information on whether the expression level is significantly increased, significantly decreased, or not significantly changed by each of the first to m drugs (DR1, DR2,..., DRm).

[0037] At this time, the off-target analysis module 100 assigns a first value when the expression level of a specific mRNA is significantly increased by a specific drug, a second value when the expression level of a specific mRNA is significantly decreased by a specific drug, and a third value when the expression level of a specific mRNA does not significantly change by a specific drug, thereby generating the first to nth expression vectors (V_1, V_2,..., V_n) for each of the first to nth mRNAs (mRNA1, mRNA2,..., mRNA n) in the form of a numeric vector.

[0038] FIG. 2 shows, as an example, the case where the first value is "1", the second value is "-1", and the third value is "0".

[0039] Therefore, in the example shown in FIG. 2, the first expression vector (V_1) generated for the first mRNA may be [1, 0,..., -1], the second expression vector (V_2) generated for the second mRNA may be [0, 1,..., 0], and the nth expression vector (V_n) generated for the nth mRNA may be [0, -1,..., 1].

[0040] However, the present invention is not limited thereto, and the first value, the second value, and the third value may be defined by any other numbers instead of "1", "-1", and "0".

[0041] Thereafter, the off-target analysis module 100 can cluster the first to nth expression level vectors (V_1, V_2,..., V_n) corresponding to the first to nth mRNAs (mRNA1, mRNA2,..., mRNA_n) into k clusters by applying various types of clustering algorithms such as the K-means algorithm, Gaussian Mixture Model (GMM) algorithm, DBSCAN (Density Based Spatial Clustering of Applications with Noise) algorithm, Mean Shift algorithm, and Agglomerative Hierarchical Clustering algorithm. According to an embodiment, the number of clusters clustered by the off-target analysis module 100 may be predefined or determined by an administrator's input.

[0042] Since the K-means algorithm, Gaussian Mixture Model (GMM) algorithm, DBSCAN (Density Based Spatial Clustering of Applications with Noise) algorithm, Mean Shift algorithm, and Agglomerative Hierarchical Clustering algorithm are well-known clustering algorithms, a detailed description of the process in which the off-target analysis module 100 applies various types of clustering algorithms to cluster the first to nth expression level vectors (V_1, V_2,..., V_n) into k clusters is omitted here.

[0043] Thereafter, the off-target analysis module 100 can classify the plurality of mRNAs into the first to k clusters by classifying the mRNAs corresponding to the expression level vectors included in the same cluster into the same cluster.

[0044] After classifying the plurality of mRNAs contained in the human body into the first to k-th clusters, the non-target analysis module 100 can determine mRNAs included in the cluster to which the target mRNA (T_MRNA) belongs among the first to k-th clusters as a plurality of non-target mRNAs (OFF_T_MRNA).

[0045] For example, the non-target analysis module 100 can read from the gene expression database 110 the degree to which the expression level of the target mRNA (T_MRNA) is regulated by each of the first to m-th drugs (DR1, DR2,..., DRm), and generate an m-dimensional target expression level vector including the degree to which the expression level of the target mRNA (T_MRNA) is regulated by each of the first to m-th drugs (DR1, DR2,..., DRm).

[0046] In one embodiment, the non-target analysis module 100 can generate the m-dimensional target expression level vector including information on whether the expression level of the target mRNA (T_MRNA) is significantly increased, significantly decreased, or not significantly changed by each of the first to m-th drugs (DR1, DR2,..., DRm).

[0047] At this time, as described above with reference to FIG. 2, when the expression level of the target mRNA (T_MRNA) is significantly increased by a specific drug, the non-target analysis module 100 assigns the first value, and when the expression level of the target mRNA (T_MRNA) is significantly decreased by a specific drug, the non-target analysis module 100 assigns the second value, and when the expression level of the target mRNA (T_MRNA) is not significantly changed by a specific drug, the non-target analysis module 100 assigns the third value, thereby generating the target expression level vector for the target mRNA (T_MRNA) in the form of a numeric vector.

[0048] Thereafter, the off-target analysis module 100 determines, as the target cluster, the cluster corresponding to the center among the centers of the first to k-th clusters that is closest to the target expression level vector, and can determine the mRNAs included in the target cluster as a plurality of off-target mRNAs (OFF_T_MRNA) for the target mRNA (T_MRNA).

[0049] Therefore, the plurality of off-target mRNAs (OFF_T_MRNA) can have a gene expression pattern similar to that of the target mRNA (T_MRNA).

[0050] Referring to FIG. 1 again, the off-target analysis module 100 can provide the plurality of off-target mRNAs (OFF_T_MRNA) determined for the target mRNA (T_MRNA) to the base sequence generation module 200.

[0051] The greater the degree to which a specific RNA therapeutic agent regulates the activity of the target mRNA (T_MRNA) related to the synthesis of the protein that causes a specific disease, the greater the therapeutic effect on the specific disease. However, since the plurality of off-target mRNAs (OFF_T_MRNA) have a gene expression pattern similar to that of the target mRNA (T_MRNA), the specific RNA therapeutic agent is highly likely to cause side effects of inhibiting the synthesis of normal proteins by also regulating the activity of the plurality of off-target mRNAs (OFF_T_MRNA).

[0052] Therefore, based on the target mRNA (T_MRNA) received from the outside and the plurality of off-target mRNAs (OFF_T_MRNA) received from the off-target analysis module 100, the base sequence generation module 200 determines the base sequence (RNA_DRUG_SEQ) of an RNA therapeutic agent that greatly regulates the activity of the target mRNA (T_MRNA) and slightly regulates the activity of the plurality of off-target mRNAs (OFF_T_MRNA).

[0053] As described above, the off-target analysis module 100 determines, as a plurality of off-target mRNAs (OFF_T_MRNA), mRNAs having a gene expression pattern similar to that of the target mRNA (T_MRNA) among the plurality of mRNAs contained in the human body and provides them to the base sequence generation module 200. However, the present invention is not limited to this. According to an embodiment, the base sequence generation module 200 may set all the mRNAs contained in the human body as a plurality of off-target mRNAs (OFF_T_MRNA) and perform the following operations.

[0054] In one embodiment, the base sequence generation module 200 can perform reinforcement learning to determine the base sequence (RNA_DRUG_SEQ) of an RNA therapeutic agent that greatly regulates the activity of the target mRNA (T_MRNA) and slightly regulates the activities of the plurality of off-target mRNAs (OFF_T_MRNA).

[0055] Specifically, the base sequence generation module 200 determines an arbitrary candidate base sequence, and increases the compensation (reward) as the degree to which the candidate base sequence binds to the target mRNA (T_MRNA) and regulates the activity of the target mRNA (T_MRNA) is greater. As the degree to which the candidate base sequence binds to each of the plurality of off-target mRNAs (OFF_T_MRNA) and regulates the activity of each of the plurality of off-target mRNAs (OFF_T_MRNA) is greater, the compensation is decreased, and the compensation for the candidate base sequence can be determined.

[0056] FIG. 3 is a diagram for explaining an example of the operation of the base sequence generation module included in the RNA therapeutic agent design system of FIG. 1. FIG. 3 shows the base sequence of a specific mRNA (MRNA) and the candidate base sequence (C_SEQ).

[0057] As shown in FIG. 3, when the candidate base sequence (C_SEQ) binds to a specific mRNA (MRNA) by binding complementary base pairs between the base sequence of the specific mRNA (MRNA) and the candidate base sequence (C_SEQ), a certain amount of energy can be released.

[0058] The greater the magnitude of the energy released when the candidate base sequence (C_SEQ) binds to a specific mRNA (MRNA), the more likely the candidate base sequence (C_SEQ) binds to the specific mRNA (MRNA) and becomes a more stable state, which means that the candidate base sequence (C_SEQ) can regulate the activity of the specific mRNA (MRNA) to a greater extent by easily binding to the specific mRNA (MRNA).

[0059] Therefore, the greater the magnitude of the energy released when the candidate base sequence (C_SEQ) binds to the target mRNA (T_MRNA), the greater the extent to which the activity of the target mRNA (T_MRNA) is regulated, and the greater the magnitude of the energy released when the candidate base sequence (C_SEQ) binds to each of the plurality of non-target mRNAs (OFF_T_MRNA), the greater the extent to which the activity of each of the plurality of non-target mRNAs (OFF_T_MRNA) can be regulated.

[0060] Therefore, in one embodiment, the base sequence generation module 200 can determine, as the compensation, a value obtained by subtracting a value obtained by multiplying the energy released when the candidate base sequence (C_SEQ) binds to each of the plurality of non-target mRNAs (OFF_T_MRNA) from the energy released when the candidate base sequence (C_SEQ) binds to the target mRNA (T_MRNA).

[0061] More specifically, the energy released when the candidate base sequence (C_SEQ) binds to a specific mRNA (MRNA) is related to the Gibbs free energy (ΔG) for the reaction in which the candidate base sequence (C_SEQ) binds to the specific mRNA (MRNA).

[0062] The Gibbs free energy (ΔG) indicates the spontaneity of a chemical reaction. The greater the negative value of the Gibbs free energy of a chemical reaction, the greater the spontaneity, and the greater the positive value of the Gibbs free energy of a chemical reaction, the greater the non-spontaneity.

[0063] Therefore, the more negative and larger the Gibbs free energy (ΔG) of the reaction in which the candidate base sequence (C_SEQ) binds to a specific mRNA (MRNA) is, the more frequently the binding between the candidate base sequence (C_SEQ) and the specific mRNA (MRNA) occurs spontaneously, which means that the candidate base sequence (C_SEQ) regulates the activity of the specific mRNA (MRNA) to a greater extent. At this time, the amount of energy released may also be large.

[0064] Therefore, in one embodiment, the base sequence generation module 200 can determine the compensation for the candidate base sequence (C_SEQ) using Equation 1 below.

[0065] [Equation 1]

[0066] Here, RW represents the compensation, ΔGon represents the Gibbs free energy of the reaction in which the candidate base sequence (C_SEQ) binds to the target mRNA (T_MRNA), ΔGoff represents the Gibbs free energy of the reaction in which the candidate base sequence (C_SEQ) binds to the non-target mRNA (OFF_T_MRNA), and α represents the weight.

[0067] Therefore, the larger the positive value of the compensation for the candidate base sequence (C_SEQ) is, the greater the degree to which the candidate base sequence (C_SEQ) binds to the target mRNA (T_MRNA) and regulates the activity of the target mRNA (T_MRNA), and the smaller the degree to which it binds to each of the plurality of non-target mRNAs (OFF_T_MRNA) and regulates the activity of each of the plurality of non-target mRNAs (OFF_T_MRNA).

[0068] On the other hand, the base sequence generation module 200 can pre-store the Gibbs free energy for binding reactions between various types of base sequences.

[0069] Therefore, based on the Gibbs free energy for the binding reaction between the pre-stored base sequences, the base sequence generation module 200 can estimate the Gibbs free energy of the reaction in which the candidate base sequence (C_SEQ) binds to the target mRNA (T_MRNA) and the Gibbs free energy of the reaction in which the candidate base sequence (C_SEQ) binds to each of the plurality of non-target mRNAs (OFF_T_MRNA).

[0070] Referring again to FIG. 1, after the base sequence generation module 200 performs the operation as described above with reference to FIG. 3 to determine the compensation for the candidate base sequence, while variously modifying the candidate base sequence, the compensation is calculated in the same manner as described above with reference to FIG. 3, and learning can be performed as to whether the compensation for the modified candidate base sequence has increased or decreased compared to the compensation for the unmodified candidate base sequence.

[0071] For example, the base sequence generation module 200 can learn when the compensation increases when the candidate base sequence is modified in what way and when the compensation decreases when the candidate base sequence is modified in what way.

[0072] In this way, the base sequence generation module 200 repeatedly performs the process of modifying the candidate base sequence in the direction in which the compensation increases while calculating the compensation while variously modifying the candidate base sequence. When the compensation no longer increases due to the modification of the candidate base sequence, the final candidate base sequence can be determined as the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0073] Therefore, the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent generated from the base sequence generation module 200 can effectively prevent side effects by maximizing the therapeutic effect on the specific disease by regulating the activity of the target mRNA (T_MRNA) and minimizing the degree of regulating the activity of the plurality of non-target mRNAs (OFF_T_MRNA).

[0074] On the one hand, the RNA therapeutic agent corresponding to the nucleotide sequence (RNA_DRUG_SEQ) generated by the nucleotide sequence generation module 200 may be various types of RNA therapeutic agents.

[0075] In one embodiment, the RNA therapeutic agent corresponding to the nucleotide sequence (RNA_DRUG_SEQ) generated by the nucleotide sequence generation module 200 can correspond to small interfering RNA (siRNA).

[0076] In one embodiment, the RNA therapeutic agent corresponding to the nucleotide sequence (RNA_DRUG_SEQ) generated by the nucleotide sequence generation module 200 can correspond to antisense oligonucleotide (ASO).

[0077] In one embodiment, the RNA therapeutic agent corresponding to the nucleotide sequence (RNA_DRUG_SEQ) generated by the nucleotide sequence generation module 200 can correspond to Aptamer.

[0078] In one embodiment, the RNA therapeutic agent corresponding to the nucleotide sequence (RNA_DRUG_SEQ) generated by the nucleotide sequence generation module 200 can correspond to CRISPR. The nucleotide sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent generated by the nucleotide sequence generation module 200 can be provided to the nucleotide sequence modification module 300.

[0079] The nucleotide sequence modification module 300 uses a chemical modification information database (CM_DB) 310 that stores numerical values indicating biological characteristics when a plurality of chemical modifications are applied to a plurality of nucleotide sequences, and among the first to wth (where w is an integer of 2 or more) chemical modifications known in advance, the chemical modification with the highest therapeutic efficiency and in vivo stability when applied to the nucleotide sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent is determined as the optimal chemical modification (O_CHEM_MOD).

[0080] The detailed operation of the base sequence modification module 300 for determining the optimal chemical modification (O_CHEM_MOD) will be described later with reference to FIG. 6.

[0081] FIG. 4 is a diagram showing an RNA therapeutic agent design system according to an embodiment of the present invention. Referring to FIG. 4, the RNA therapeutic agent design system 20 includes an off-target analysis module 100, a sequence generation module 200, a sequence modification module 300, and a secondary structure prediction module 400.

[0082] The RNA therapeutic agent design system 20 shown in FIG. 4 is the same as the RNA therapeutic agent design system 10 shown in FIG. 1, except that it further includes a secondary structure prediction module 400 in the RNA therapeutic agent design system 10 shown in FIG. 1.

[0083] Regarding the configuration and operation of the RNA therapeutic agent design system 10 shown in FIG. 1, it has been described in detail with reference to FIGS. 1 to 3. Therefore, here, duplicate descriptions are omitted, and only matters related to the secondary structure prediction module 400 in the RNA therapeutic agent design system 20 of FIG. 4 will be described in detail.

[0084] When the off-target analysis module 100 receives the target mRNA (T_MRNA), the off-target analysis module 100 determines, as a plurality of off-target mRNAs (OFF_T_MRNA), mRNAs having a gene expression pattern similar to the target mRNA (T_MRNA) among a plurality of mRNAs contained in the human body.

[0085] The off-target analysis module 100 included in the RNA therapeutic agent design system 20 of FIG. 4 can operate in the same manner as the off-target analysis module 100 included in the RNA therapeutic agent design system 10 of FIG. 1.

[0086] Regarding the operation of the off-target analysis module 100 included in the RNA therapeutic agent design system 10 of FIG. 1, it has been described above with reference to FIGS. 1 to 3. Therefore, here, a detailed description of the off-target analysis module 100 will be omitted.

[0087] Generally, mRNA exists in a state folded into a mostly complex structure. Specifically, the mRNA strand folds itself and forms hydrogen bonds autonomously between complementary base pairs, thereby existing in a folded state.

[0088] In one embodiment, the secondary structure prediction module 400 can receive a target mRNA (T_MRNA) from the outside.

[0089] In this case, the secondary structure prediction module 400 can predict the secondary structure in which the target mRNA (T_MRNA) is folded.

[0090] In one embodiment, the secondary structure prediction module 400 can predict the secondary structure of the target mRNA (T_MRNA) using a secondary structure database (SS_DB) 410 that pre-stores the folded secondary structures of a plurality of RNAs contained in the human body. For example, the secondary structure database 410 may be an ArchiveII database, an RNAStrAlign database, or the like.

[0091] In this case, the secondary structure prediction module 400 can generate a secondary structure prediction model that outputs the secondary structure for the base sequence of the input mRNA by performing learning using the data stored in the secondary structure database 410 as learning data.

[0092] Thereafter, the secondary structure prediction module 400 can predict the secondary structure of the target mRNA (T_MRNA) by inputting the base sequence of the target mRNA (T_MRNA) into the secondary structure prediction model.

[0093] The secondary structure prediction module 400 can estimate the autonomous binding relationship between the base sequences of the target mRNA (T_MRNA) by predicting the secondary structure in which the target mRNA (T_MRNA) is folded.

[0094] In one embodiment, after predicting the secondary structure in which the target mRNA (T_MRNA) is folded, the secondary structure prediction module 400 has rows and columns corresponding to the length of the base sequence of the target mRNA (T_MRNA), and the target mRNA (T_MRNA) A square matrix (S_MAT) indicating the autonomous binding relationship between the base sequences of the target mRNA (T_MRNA) formed by the folded secondary structure can be generated.

[0095] FIG. 5 is a diagram for explaining an example of the operation of the secondary structure prediction module included in the RNA therapeutic agent design system of FIG. 4. FIG. 5 shows that the target mRNA (T_MRNA) has a base sequence corresponding to G-G-G-A-A-A-C-G-U-U-C-C-G as an example.

[0096] As shown in FIG. 5, the target mRNA (T_MRNA) can exist in a folded state by autonomously forming hydrogen bonds between complementary base pairs.

[0097] In the case of the target mRNA (T_MRNA) shown in FIG. 5, it is shown that guanine, which is the second base, forms a hydrogen bond with cytosine, which is the 12th base, guanine, which is the third base, forms a hydrogen bond with cytosine, which is the 11th base, and adenine, which is the fourth base, forms a hydrogen bond with uracil, which is the 10th base.

[0098] In this case, the secondary structure prediction module 400 can generate a square matrix (S_MAT) having rows and columns corresponding to the length of the base sequence of the target mRNA (T_MRNA), and indicating the autonomous binding relationship between the base sequences of the target mRNA (T_MRNA) formed by the secondary structure in which the target mRNA (T_MRNA) is folded.

[0099] Specifically, as shown in FIG. 5, in the square matrix (S_MAT) generated by the secondary structure prediction module 400, when the i-th base and the j-th base in the base sequence of the target mRNA are bound to each other by the secondary structure of the target mRNA, the elements corresponding to the i-th row and the j-th column and the elements corresponding to the j-th row and the i-th column have a first value, and when the p-th base and the q-th base are not bound to each other, the elements corresponding to the p-th row and the q-th column and the elements corresponding to the q-th row and the p-th column have a second value different from the first value, corresponding to a symmetric matrix. Here, i, j, p, and q represent natural numbers. In one embodiment, as shown in FIG. 5, the first value may be "1" and the second value may be "0".

[0100] Therefore, the square matrix (S_MAT) generated by the secondary structure prediction module 400 can indicate the folded secondary structure of the target mRNA (T_MRNA).

[0101] Referring again to FIG. 4, the autonomous binding relationship between the base sequences of the target mRNA (T_MRNA) estimated by the secondary structure prediction module 400 can be provided to the base sequence generation module 200.

[0102] Based on the autonomous binding relationship between the base sequences of the target mRNA (T_MRNA), the base sequence generation module 200 can determine the base sequence (RNA_DRUG_SEQ) of an RNA therapeutic agent that greatly regulates the activity of the target mRNA (T_MRNA) and slightly regulates the activities of a plurality of non-target mRNAs (OFF_T_MRNA).

[0103] For example, the secondary structure prediction module 400 provides a square matrix (S_MAT) generated for the target mRNA (T_MRNA) to the base sequence generation module 200. Based on the square matrix (S_MAT) for the target mRNA (T_MRNA), the base sequence generation module 200 can determine the base sequence (RNA_DRUG_SEQ) of an RNA therapeutic agent that greatly regulates the activity of the target mRNA (T_MRNA) and slightly regulates the activities of a plurality of non-target mRNAs (OFF_T_MRNA).

[0104] In one embodiment, the base sequence generation module 200 performs reinforcement learning based on the square matrix (S_MAT) for the target mRNA (T_MRNA) to determine the base sequence (RNA_DRUG_SEQ) of an RNA therapeutic agent that greatly regulates the activity of the target mRNA (T_MRNA) and slightly regulates the activities of a plurality of non-target mRNAs (OFF_T_MRNA). An example of the base sequence generation module 200 performing reinforcement learning to determine the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent will be described below. As described above, generally, mRNA exists in a folded state by autonomously forming hydrogen bonds between complementary base pairs.

[0105] Therefore, a general RNA therapeutic agent is more likely to bind to a nucleotide sequence portion that does not form an autonomous bond in the nucleotide sequence of mRNA than to a nucleotide sequence portion that forms an autonomous bond in the nucleotide sequence of mRNA.

[0106] For this reason, in one embodiment, the nucleotide sequence generation module 200 assumes that the RNA therapeutic agent does not bind to a nucleotide sequence portion that forms an autonomous bond in the nucleotide sequence of the target mRNA (T_MRNA), and binds only to a nucleotide sequence portion that does not form an autonomous bond in the nucleotide sequence of the target mRNA (T_MRNA), and can determine the nucleotide sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0107] In this case, the nucleotide sequence generation module 200 can determine, based on the square matrix (S_MAT) for the target mRNA (T_MRNA), a nucleotide sequence that does not form an autonomous bond in the nucleotide sequence of the target mRNA (T_MRNA) as a bindable nucleotide sequence.

[0108] Thereafter, the nucleotide sequence generation module 200 determines an arbitrary candidate nucleotide sequence, and increases the compensation (reward) as the degree to which the candidate nucleotide sequence binds to the bindable nucleotide sequence among the nucleotide sequences of the target mRNA (T_MRNA) and regulates the activity of the target mRNA (T_MRNA) is greater, and decreases the compensation as the degree to which the candidate nucleotide sequence binds to each of the plurality of non-target mRNAs (OFF_T_MRNA) and regulates the activity of each of the plurality of non-target mRNAs (OFF_T_MRNA) is greater, and can determine the compensation for the candidate nucleotide sequence.

[0109] In one embodiment, the nucleotide sequence generation module 200 can determine, as the compensation, a value obtained by subtracting a value obtained by multiplying the energy released when the candidate nucleotide sequence binds to each of the plurality of non-target mRNAs (OFF_T_MRNA) from the energy released when the candidate nucleotide sequence binds to the bindable nucleotide sequence among the nucleotide sequences of the target mRNA (T_MRNA). For example, the base sequence generation module 200 can determine the compensation for the candidate base sequence using the following Equation 2.

[0110]

Equation

[0111] Here, RW represents the compensation, ΔGon represents the Gibbs free energy of the reaction in which the candidate base sequence binds to the bindable base sequence among the base sequences of the target mRNA (T_MRNA), ΔGoff represents the Gibbs free energy of the reaction in which the candidate base sequence binds to the non-target mRNA (OFF_T_MRNA), and α represents the weight.

[0112] Therefore, the larger the positive value of the compensation for the candidate base sequence, the greater the degree to which the candidate base sequence binds to the target mRNA (T_MRNA) and regulates the activity of the target mRNA (T_MRNA), and the smaller the degree to which it binds to each of the plurality of non-target mRNAs (OFF_T_MRNA) and regulates the activity of each of the plurality of non-target mRNAs (OFF_T_MRNA).

[0113] The method by which the base sequence generation module 200 included in the RNA therapeutic agent design system 20 in FIG. 4 calculates the compensation for the candidate base sequence using Equation 2 is the same as the method by which the base sequence generation module 200 included in the RNA therapeutic agent design system 10 in FIG. 1 calculates the compensation for the candidate base sequence using Equation 1, as described above with reference to FIG. 3. Therefore, a duplicate description regarding the specific method by which the base sequence generation module 200 calculates the compensation for the candidate base sequence is omitted here.

[0114] After the base sequence generation module 200 performs the operations as described above and determines the compensation for the candidate base sequence, while diversely modifying the candidate base sequence, the compensation is calculated in the same manner as described above, and learning can be performed on whether the compensation for the modified candidate base sequence increases or decreases compared to the compensation for the candidate base sequence before modification.

[0115] For example, the base sequence generation module 200 can learn about when the compensation increases when the candidate base sequence is modified in what way and when the compensation decreases when the candidate base sequence is modified in what way.

[0116] In this way, the base sequence generation module 200 calculates the compensation while diversely modifying the candidate base sequence, and repeatedly performs the process of modifying the candidate base sequence in the direction in which the compensation increases. When the compensation no longer increases due to the modification of the candidate base sequence, the final candidate base sequence can be determined as the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0117] Therefore, the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent generated from the base sequence generation module 200 can effectively prevent side effects by maximizing the therapeutic effect on the specific disease by regulating the activity of the target mRNA (T_MRNA) and minimizing the degree of regulating the activities of a plurality of non-target mRNAs (OFF_T_MRNA). Hereinafter, another example in which the base sequence generation module 200 performs reinforcement learning to determine the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent will be described. As described above, generally, mRNA exists in a folded state by autonomously forming hydrogen bonds between complementary base pairs.

[0118] Therefore, a general RNA therapeutic agent is more likely to bind to a nucleotide sequence portion that does not form an autonomous bond in the nucleotide sequence of mRNA than to a nucleotide sequence portion that forms an autonomous bond in the nucleotide sequence of mRNA.

[0119] Thus, in one embodiment, the nucleotide sequence generation module 200 can determine the nucleotide sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent in consideration of the fact that the RNA therapeutic agent is more likely to bind to a nucleotide sequence portion that does not form an autonomous bond in the nucleotide sequence of the target mRNA (T_MRNA) than to a nucleotide sequence portion that forms an autonomous bond in the nucleotide sequence of the target mRNA (T_MRNA).

[0120] In this case, the nucleotide sequence generation module 200 determines an arbitrary candidate nucleotide sequence, and increases the compensation (reward) as the degree to which the candidate nucleotide sequence binds to the nucleotide sequence of the target mRNA (T_MRNA) and regulates the activity of the target mRNA (T_MRNA) is greater, regardless of the secondary structure of the target mRNA (T_MRNA) determined based on the square matrix (S_MAT) for the target mRNA (T_MRNA). The compensation can be calculated by decreasing the compensation as the degree to which the candidate nucleotide sequence binds to each of the plurality of non-target mRNAs (OFF_T_MRNA) and regulates the activity of each of the plurality of non-target mRNAs (OFF_T_MRNA) is greater.

[0121] On the one hand, generally, an RNA therapeutic agent is more likely to bind to a base sequence portion that does not form an autonomous bond in the base sequence of mRNA than to a base sequence portion that forms an autonomous bond in the base sequence of mRNA. Therefore, among the bases of the target mRNA (T_MRNA) that bind to the candidate base sequence, the higher the ratio of the bases that form a bond by the target mRNA (T_MRNA) itself, the relatively more difficult it is for the candidate base sequence to bind to the target mRNA (T_MRNA). Among the bases of the target mRNA (T_MRNA) that bind to the candidate base sequence, the lower the ratio of the bases that form a bond by the target mRNA (T_MRNA) itself, the relatively easier it may be for the candidate base sequence to bind to the target mRNA (T_MRNA).

[0122] Therefore, after determining a secondary structure penalty proportional to the ratio of the bases that form a bond by the target mRNA (T_MRNA) itself among the bases of the target mRNA (T_MRNA) that bind to the candidate base sequence based on the square matrix (S_MAT) for the target mRNA (T_MRNA), the base sequence generation module 200 can subtract the secondary structure penalty from the calculated compensation to determine the compensation for the candidate base sequence.

[0123] In one embodiment, the base sequence generation module 200 can subtract a value obtained by multiplying the energy released when the candidate base sequence binds to each of a plurality of non-target mRNAs (OFF_T_MRNA) from the energy released when the candidate base sequence binds to the base sequence of the target mRNA (T_MRNA), and then subtract the secondary structure penalty from the resulting value to determine the compensation. For example, the base sequence generation module 200 can determine the compensation for the candidate base sequence using Equation 3 below.

[0124]

Equation

[0125] Here, RW represents the compensation, ΔGon represents the Gibbs free energy of the reaction in which the candidate base sequence binds to the base sequence of the target mRNA (T_MRNA), ΔGoff represents the Gibbs free energy of the reaction in which the candidate base sequence binds to the non-target mRNA (OFF_T_MRNA), α represents the weight, and SP represents the secondary structure penalty.

[0126] Therefore, the larger the positive value of the compensation for the candidate base sequence, the greater the degree to which the candidate base sequence binds to the target mRNA (T_MRNA) and regulates the activity of the target mRNA (T_MRNA), and the smaller the degree to which it binds to each of the plurality of non-target mRNAs (OFF_T_MRNA) and regulates the activity of each of the plurality of non-target mRNAs (OFF_T_MRNA).

[0127] The method by which the base sequence generation module 200 included in the RNA therapeutic agent design system 20 in FIG. 4 calculates the compensation for the candidate base sequence using Equation 3 is the same as the method by which the base sequence generation module 200 included in the RNA therapeutic agent design system 10 in FIG. 1 calculates the compensation for the candidate base sequence using Equation 1, except that the secondary structure penalty is subtracted from the compensation. Therefore, here, a duplicate description regarding the specific method by which the base sequence generation module 200 calculates the compensation for the candidate base sequence is omitted.

[0128] After determining the compensation for the candidate base sequence by performing the operations as described above, the base sequence generation module 200 can calculate the compensation in the same manner as described above while variously modifying the candidate base sequence, and learn whether the compensation for the modified candidate base sequence has increased or decreased compared to the compensation for the unmodified candidate base sequence.

[0129] For example, the base sequence generation module 200 can learn about when the compensation increases when the candidate base sequence is deformed in what way, and when the compensation decreases when the candidate base sequence is deformed in what way.

[0130] In this way, the base sequence generation module 200 calculates the compensation while diversely deforming the candidate base sequence, and repeatedly performs the process of deforming the candidate base sequence in the direction in which the compensation increases. When the compensation no longer increases due to the deformation of the candidate base sequence, the final candidate base sequence can be determined as the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0131] Therefore, the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent generated from the base sequence generation module 200 can effectively prevent side effects by maximizing the therapeutic effect on the specific disease while minimizing the degree of regulating the activities of a plurality of non-target mRNAs (OFF_T_MRNA) by regulating the activity of the target mRNA (T_MRNA).

[0132] In another embodiment, as shown in FIG. 4, the secondary structure prediction module 400 can receive the target mRNA (T_MRNA) from the outside and receive a plurality of non-target mRNAs (OFF_T_MRNA) for the target mRNA (T_MRNA) from the non-target analysis module 100.

[0133] In this case, the secondary structure prediction module 400 can predict the secondary structure in which the corresponding mRNA is folded for each of the target mRNA (T_MRNA) and the plurality of non-target mRNAs (OFF_T_MRNA).

[0134] In one embodiment, the secondary structure prediction module 400 can predict the secondary structures of the target mRNA (T_MRNA) and a plurality of non-target mRNAs (OFF_T_MRNA) respectively, using a secondary structure database (SS_DB) 410 that pre-stores the folded secondary structures of a plurality of RNAs contained in the human body. For example, the secondary structure database 410 may be an ArchiveII database, an RNAStrAlign database, or the like.

[0135] In this case, the secondary structure prediction module 400 can generate a secondary structure prediction model that outputs the secondary structure for the base sequence of the input mRNA by performing learning using the data stored in the secondary structure database 410 as learning data.

[0136] Thereafter, the secondary structure prediction module 400 can predict the secondary structures of the target mRNA (T_MRNA) and a plurality of non-target mRNAs (OFF_T_MRNA) respectively, by inputting the base sequences of the target mRNA (T_MRNA) and a plurality of non-target mRNAs (OFF_T_MRNA) into the secondary structure prediction model.

[0137] By predicting the secondary structures in which the target mRNA (T_MRNA) and a plurality of non-target mRNAs (OFF_T_MRNA) are respectively folded, the secondary structure prediction module 400 can estimate the autonomous binding relationships between the base sequences of the target mRNA (T_MRNA) and the autonomous binding relationships between the base sequences of each of the plurality of non-target mRNAs (OFF_T_MRNA).

[0138] In one embodiment, after the secondary structure prediction module 400 predicts the secondary structures in which the target mRNA (T_MRNA) and a plurality of non-target mRNAs (OFF_T_MRNA) are folded respectively, it can generate a square matrix (S_MAT) that has rows and columns corresponding to the length of the base sequence of the corresponding mRNA and shows the autonomous binding relationship between the base sequences of the corresponding mRNA formed by the secondary structure in which the corresponding mRNA is folded.

[0139] Specifically, the secondary structure prediction module 400 predicts the secondary structure in which the target mRNA (T_MRNA) is folded, generates a square matrix (S_MAT) that has rows and columns corresponding to the length of the base sequence of the target mRNA (T_MRNA), and shows the autonomous binding relationship between the base sequences of the target mRNA (T_MRNA) formed by the secondary structure in which the target mRNA (T_MRNA) is folded.

[0140] In addition, for each of the plurality of non-target mRNAs (OFF_T_MRNA), the secondary structure prediction module 400 predicts the secondary structure in which the non-target mRNA (OFF_T_MRNA) is folded, generates a square matrix (S_MAT) that has rows and columns corresponding to the length of the base sequence of the non-target mRNA (OFF_T_MRNA), and shows the autonomous binding relationship between the base sequences of the non-target mRNA (OFF_T_MRNA) formed by the secondary structure in which the non-target mRNA (OFF_T_MRNA) is folded.

[0141] The secondary structure prediction module 400 can generate a square matrix (S_MAT) for the target mRNA (T_MRNA) and a square matrix (S_MAT) for each of the plurality of non-target mRNAs (OFF_T_MRNA) in the same manner as described above with reference to FIG. 5.

[0142] The autonomous binding relationships between the base sequences of the target mRNA (T_MRNA) estimated by the secondary structure prediction module 400 and the autonomous binding relationships between the base sequences of each of the plurality of non-target mRNAs (OFF_T_MRNA) can be provided to the base sequence generation module 200.

[0143] Based on the autonomous binding relationships between the base sequences of the target mRNA (T_MRNA) and the autonomous binding relationships between the base sequences of each of the plurality of non-target mRNAs (OFF_T_MRNA), the base sequence generation module 200 can determine the base sequence (RNA_DRUG_SEQ) of an RNA therapeutic agent that greatly regulates the activity of the target mRNA (T_MRNA) and slightly regulates the activity of the plurality of non-target mRNAs (OFF_T_MRNA).

[0144] For example, the secondary structure prediction module 400 provides the square matrix (S_MAT) generated for the target mRNA (T_MRNA) and the square matrix (S_MAT) generated for each of the plurality of non-target mRNAs (OFF_T_MRNA) to the base sequence generation module 200. Based on the square matrix (S_MAT) for the target mRNA (T_MRNA) and the square matrix (S_MAT) for each of the plurality of non-target mRNAs (OFF_T_MRNA), the base sequence generation module 200 can determine the base sequence (RNA_DRUG_SEQ) of an RNA therapeutic agent that greatly regulates the activity of the target mRNA (T_MRNA) and slightly regulates the activity of the plurality of non-target mRNAs (OFF_T_MRNA).

[0145] In one embodiment, the base sequence generation module 200 performs reinforcement learning based on the square matrix (S_MAT) for the target mRNA (T_MRNA) and each of the plurality of non-target mRNAs (OFF_T_MRNA), and can determine the base sequence (RNA_DRUG_SEQ) of an RNA therapeutic agent that greatly regulates the activity of the target mRNA (T_MRNA) and slightly regulates the activity of the plurality of non-target mRNAs (OFF_T_MRNA). Hereinafter, an example in which the base sequence generation module 200 performs reinforcement learning to determine the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent will be described. As described above, generally, mRNA exists in a folded state by autonomously forming hydrogen bonds between complementary base pairs.

[0146] Therefore, it is easier for a general RNA therapeutic agent to bind to a base sequence portion that does not form an autonomous bond in the base sequence of mRNA than to a base sequence portion that forms an autonomous bond in the base sequence of mRNA.

[0147] For this reason, in one embodiment, the base sequence generation module 200 assumes that the RNA therapeutic agent does not bind to the base sequence portions that form autonomous bonds in the base sequences of the target mRNA (T_MRNA) and the plurality of non-target mRNAs (OFF_T_MRNA), and binds only to the base sequence portions that do not form autonomous bonds in the base sequences of the target mRNA (T_MRNA) and the plurality of non-target mRNAs (OFF_T_MRNA), and can determine the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0148] In this case, the base sequence generation module 200 can determine, based on the square matrix (S_MAT) for each of the target mRNA (T_MRNA) and the plurality of non-target mRNAs (OFF_T_MRNA), the base sequences that do not form autonomous bonds in the base sequences of the target mRNA (T_MRNA) and the plurality of non-target mRNAs (OFF_T_MRNA) as bondable base sequences.

[0149] Thereafter, the base sequence generation module 200 determines an arbitrary candidate base sequence, and increases the compensation (reward) as the degree to which the candidate base sequence binds to the bindable base sequence among the base sequences of the target mRNA (T_MRNA) and regulates the activity of the target mRNA (T_MRNA) is greater. As the degree to which the candidate base sequence binds to the bindable base sequence among the base sequences of each of the plurality of non-target mRNAs (OFF_T_MRNA) and regulates the activity of each of the plurality of non-target mRNAs (OFF_T_MRNA) is greater, the compensation is decreased, and the compensation for the candidate base sequence can be determined.

[0150] In one embodiment, the base sequence generation module 200 can determine the compensation as a value obtained by subtracting a value obtained by multiplying the energy released when the candidate base sequence binds to the bindable base sequence among the base sequences of the plurality of non-target mRNAs (OFF_T_MRNA) from the energy released when the candidate base sequence binds to the bindable base sequence among the base sequences of the target mRNA (T_MRNA). For example, the base sequence generation module 200 can determine the compensation for the candidate base sequence using the following Equation 4. [Equation]

[0151] Here, RW represents the compensation, ΔGon represents the Gibbs free energy of the reaction in which the candidate base sequence binds to the bindable base sequence among the base sequences of the target mRNA (T_MRNA), ΔGoff represents the Gibbs free energy of the reaction in which the candidate base sequence binds to the bindable base sequence among the base sequences of the non-target mRNA (OFF_T_MRNA), and α represents the weight.

[0152] Therefore, the greater the compensation for the candidate base sequence has a large positive value, the greater the degree to which the candidate base sequence binds to the target mRNA (T_MRNA) and regulates the activity of the target mRNA (T_MRNA), and the smaller the degree to which it binds to each of the plurality of non-target mRNAs (OFF_T_MRNA) and regulates the activity of each of the plurality of non-target mRNAs (OFF_T_MRNA).

[0153] The method by which the base sequence generation module 200 included in the RNA therapeutic agent design system 20 in FIG. 4 calculates the compensation for the candidate base sequence using the number 4 is the same as the method by which the base sequence generation module 200 included in the RNA therapeutic agent design system 10 in FIG. 1 calculates the compensation for the candidate base sequence using the number 1, as described above with reference to FIG. 3. Therefore, here, a duplicate explanation regarding the specific method by which the base sequence generation module 200 calculates the compensation for the candidate base sequence is omitted.

[0154] After the base sequence generation module 200 performs the operations as described above to determine the compensation for the candidate base sequence, it calculates the compensation in the same manner as described above while diversely modifying the candidate base sequence, and can learn whether the compensation for the modified candidate base sequence has increased or decreased compared to the compensation for the candidate base sequence before modification.

[0155] For example, the base sequence generation module 200 can learn when the compensation increases when the candidate base sequence is modified in what way, and when the compensation decreases when the candidate base sequence is modified in what way.

[0156] In this way, the base sequence generation module 200 calculates the compensation while diversely modifying the candidate base sequence, repeatedly performs the process of modifying the candidate base sequence in the direction in which the compensation increases, and when the compensation no longer increases due to the modification of the candidate base sequence, the final candidate base sequence can be determined as the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0157] Therefore, the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent generated from the base sequence generation module 200 can effectively prevent side effects by maximizing the therapeutic effect on the specific disease by regulating the activity of the target mRNA (T_MRNA) and minimizing the degree of regulating the activities of multiple non-target mRNAs (OFF_T_MRNA).

[0158] Hereinafter, another example in which the base sequence generation module 200 performs reinforcement learning to determine the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent will be described. As described above, generally, mRNA exists in a folded state by autonomously forming hydrogen bonds between complementary base pairs.

[0159] Therefore, it is easier for a general RNA therapeutic agent to bind to a base sequence portion that does not form an autonomous bond in the base sequence of mRNA than to a base sequence portion that forms an autonomous bond in the base sequence of mRNA.

[0160] For this reason, in one embodiment, the base sequence generation module 200 can determine the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent in consideration of the fact that it is easier for the RNA therapeutic agent to bind to a base sequence portion that does not form an autonomous bond in the base sequence of each of the target mRNA (T_MRNA) and the multiple non-target mRNAs (OFF_T_MRNA) than to a base sequence portion that forms an autonomous bond.

[0161] In this case, the base sequence generation module 200 determines an arbitrary candidate base sequence, and regardless of the secondary structure of the target mRNA (T_MRNA) determined based on the square matrix (S_MAT) for the target mRNA (T_MRNA) and the secondary structures of each of the plurality of non-target mRNAs (OFF_T_MRNA) determined based on the square matrix (S_MAT) for each of the plurality of non-target mRNAs (OFF_T_MRNA), the greater the degree to which the candidate base sequence binds to the base sequence of the target mRNA (T_MRNA) and regulates the activity of the target mRNA (T_MRNA), the more the compensation (reward) is increased, and the greater the degree to which the candidate base sequence binds to the base sequence of each of the plurality of non-target mRNAs (OFF_T_MRNA) and regulates the activity of each of the plurality of non-target mRNAs (OFF_T_MRNA), the more the compensation is decreased, and the compensation can be calculated.

[0162] On the other hand, generally, an RNA therapeutic agent is more likely to bind to a base sequence portion that does not form an autonomous bond in the base sequence of mRNA than to a base sequence portion that forms an autonomous bond in the base sequence of mRNA. Therefore, among the bases of the target mRNA (T_MRNA) that bind to the candidate base sequence, the higher the ratio of the bases that form a bond by the target mRNA (T_MRNA) itself, the relatively more difficult it is for the candidate base sequence to bind to the target mRNA (T_MRNA), and among the bases of the target mRNA (T_MRNA) that bind to the candidate base sequence, the lower the ratio of the bases that form a bond by the target mRNA (T_MRNA) itself, the relatively easier it may be for the candidate base sequence to bind to the target mRNA (T_MRNA).

[0163] Similarly, among the bases of the non-target mRNA (OFF_T_MRNA) that binds to the candidate base sequence, the higher the ratio of the bases that form a bond by the non-target mRNA (OFF_T_MRNA) itself, the relatively more difficult it is for the candidate base sequence to bind to the non-target mRNA (OFF_T_MRNA), and among the bases of the non-target mRNA (OFF_T_MRNA) that binds to the candidate base sequence, the lower the ratio of the bases that form a bond by the non-target mRNA (OFF_T_MRNA) itself, the relatively easier it may be for the candidate base sequence to bind to the non-target mRNA (OFF_T_MRNA).

[0164] Therefore, based on the square matrix (S_MAT) for the target mRNA (T_MRNA), the base sequence generation module 200 can determine a target secondary structure penalty proportional to the ratio of the bases of the target mRNA (T_MRNA) that form a bond by the target mRNA (T_MRNA) itself among the bases of the target mRNA (T_MRNA) that bind to the candidate base sequence.

[0165] Similarly, for each of the plurality of non-target mRNAs (OFF_T_MRNA), based on the square matrix (S_MAT) for the non-target mRNA (OFF_T_MRNA), the base sequence generation module 200 can determine a non-target secondary structure penalty proportional to the ratio of the bases that form a bond by the non-target mRNA (OFF_T_MRNA) itself among the bases of the non-target mRNA (OFF_T_MRNA) that bind to the candidate base sequence.

[0166] Thereafter, the base sequence generation module 200 can subtract the target secondary structure penalty for the target mRNA (T_MRNA) from the calculated compensation, sum up the non-target secondary structure penalties for each of the plurality of non-target mRNAs (OFF_T_MRNA), and determine the compensation for the candidate base sequence.

[0167] In one embodiment, the base sequence generation module 200 can determine the compensation by subtracting a value obtained by multiplying the energy released when the candidate base sequence binds to the base sequence of each of the plurality of non-target mRNAs (OFF_T_MRNA) from the energy released when the candidate base sequence binds to the base sequence of the target mRNA (T_MRNA), additionally subtracting the target secondary structure penalty for the target mRNA (T_MRNA), and summing up the non-target secondary structure penalties for each of the plurality of non-target mRNAs (OFF_T_MRNA). For example, the base sequence generation module 200 can determine the compensation for the candidate base sequence using Equation (5) below.

[0168] [Equation]

[0169] Here, RW represents the compensation, ΔGon represents the Gibbs free energy of the reaction in which the candidate base sequence binds to the base sequence of the target mRNA (T_MRNA), ΔGoff represents the Gibbs free energy of the reaction in which the candidate base sequence binds to the base sequence of the non-target mRNA (OFF_T_MRNA), α represents the weight, TSP represents the target secondary structure penalty for the target mRNA (T_MRNA), and OFFTSP represents the non-target secondary structure penalty for the non-target mRNA (OFF_T_MRNA).

[0170] Therefore, it can be said that the greater the positive value of the compensation for the candidate base sequence, the greater the degree to which the candidate base sequence binds to the target mRNA (T_MRNA) and regulates the activity of the target mRNA (T_MRNA), and the smaller the degree to which it binds to each of the plurality of non-target mRNAs (OFF_T_MRNA) and regulates the activity of each of the plurality of non-target mRNAs (OFF_T_MRNA).

[0171] The method by which the base sequence generation module 200 included in the RNA therapeutic agent design system 20 in FIG. 4 calculates the compensation for the candidate base sequence using the above-mentioned number 5 is the same as the method by which the base sequence generation module 200 included in the RNA therapeutic agent design system 10 in FIG. 1 calculates the compensation for the candidate base sequence using the above-mentioned number 1, except that the target secondary structure penalty for the target mRNA (T_MRNA) is subtracted from the compensation, and the non-target secondary structure penalties for each of the plurality of non-target mRNAs (OFF_T_MRNA) are summed up. Therefore, here, duplicate explanations regarding the specific method by which the base sequence generation module 200 calculates the compensation for the candidate base sequence are omitted.

[0172] After the base sequence generation module 200 performs the operations as described above to determine the compensation for the candidate base sequence, it calculates the compensation in the same manner as described above while diversely modifying the candidate base sequence, and can learn whether the compensation for the modified candidate base sequence increases or decreases compared to the compensation for the unmodified candidate base sequence.

[0173] For example, the base sequence generation module 200 can learn when the compensation increases when the candidate base sequence is modified in what way, and when the compensation decreases when the candidate base sequence is modified in what way.

[0174] In this way, the base sequence generation module 200 calculates the compensation while diversely modifying the candidate base sequence, and repeatedly performs the process of modifying the candidate base sequence in the direction in which the compensation increases. When the compensation no longer increases due to the modification of the candidate base sequence, the final candidate base sequence can be determined as the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0175] Therefore, the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent generated from the base sequence generation module 200 can effectively prevent side effects by minimizing the degree of regulating the activities of multiple non-target mRNAs (OFF_T_MRNA) while maximizing the therapeutic effect on the specific disease by regulating the activity of the target mRNA (T_MRNA).

[0176] On the other hand, the RNA therapeutic agent corresponding to the base sequence (RNA_DRUG_SEQ) generated from the base sequence generation module 200 may be various types of RNA therapeutic agents.

[0177] In one embodiment, the RNA therapeutic agent corresponding to the base sequence (RNA_DRUG_SEQ) generated from the base sequence generation module 200 can correspond to small interfering RNA (siRNA).

[0178] In one embodiment, the RNA therapeutic agent corresponding to the base sequence (RNA_DRUG_SEQ) generated from the base sequence generation module 200 can correspond to antisense oligonucleotide (ASO).

[0179] In one embodiment, the RNA therapeutic agent corresponding to the base sequence (RNA_DRUG_SEQ) generated from the base sequence generation module 200 can correspond to Aptamer.

[0180] In one embodiment, the RNA therapeutic agent corresponding to the base sequence (RNA_DRUG_SEQ) generated from the base sequence generation module 200 can correspond to CRISPR. The base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent generated from the base sequence generation module 200 can be provided to the base sequence modification module 300.

[0181] The base sequence modification module 300 uses a chemical modification information database (CM_DB) 310 that stores numerical values indicating biological characteristics when a plurality of chemical modifications are applied to a plurality of base sequences, and among the first to w-th (where w is an integer of 2 or more) chemical modifications that are known in advance, determines the chemical modification with the highest therapeutic efficiency and in vivo stability when applied to the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent as the optimal chemical modification (O_CHEM_MOD).

[0182] FIG. 6 is a diagram for explaining an example of the operation of the base sequence modification module included in the RNA therapeutic agent design system of FIGS. 1 and 4.

[0183] The base sequence modification module 300 included in the RNA therapeutic agent design system 10 of FIG. 1 and the base sequence modification module 300 included in the RNA therapeutic agent design system 20 of FIG. 4 can operate in the same manner as described below with reference to FIG. 6. As shown in FIG. 6, various types of chemical modifications can be applied to the RNA base sequences contained in the human body.

[0184] FIG. 6 shows, as examples, a molecular structural formula when no chemical modification is applied to the RNA base sequence, a molecular structural formula when a 2-O-Methyl modification is applied to the sugar portion at a specific position in the RNA base sequence, and a molecular structural formula when a 2-O-Methoxyethyl modification is applied to the sugar portion at a specific position in the RNA base sequence.

[0185] However, the present invention is not limited thereto, and the first to w-th chemical modifications considered by the base sequence modification module 300 can include all arbitrary chemical modifications such as 2-O-Methyl, 2-O-Methoxyethyl, 2-Cyanoethyl, 5-Methylcytidine, 2-Ome, 2-MOE, etc. that are applicable to the RNA base sequence.

[0186] On the one hand, the chemical modification information database 310 can store numerical values indicating biological characteristics when a plurality of chemical modifications are applied to a plurality of RNA base sequences contained in the human body.

[0187] Specifically, the chemical modification information database 310 can collect and store numerical values indicating biological characteristics measured experimentally when various types of chemical modifications are applied to each of a plurality of RNA base sequences contained in the widely publicized human body.

[0188] In one embodiment, the numerical values indicating the biological characteristics can include at least one of biological activity, inhibition rate indicating efficacy, ED50 indicating drug potency, LD50, IC50, etc., ALT (Alanine aminotransferase), AST (aspartate aminotransferase), Bilirubin, Creatinine, etc. indicating drug toxicity, and elimination half-life.

[0189] For example, the chemical modification information database 310 can collect and store data stored in databases such as the siRNAmod database and the PubChem database. Generally, even for the same RNA base sequence, different biological characteristics can be exhibited depending on the applied chemical modification.

[0190] Therefore, the chemical modification information database 310 can store numerical values indicating the biological characteristics when each of a plurality of chemical modifications is applied to a plurality of RNA base sequences contained in the human body.

[0191] In one embodiment, the base sequence modification module 300 can pre-store all chemical modifications applicable to mRNA contained in the human body as the first to w-th chemical modifications.

[0192] That is, the nucleotide sequence modification module 300 can pre-store all chemical modifications applicable to the sugar portion, phosphate portion, and base portion of the nucleotide sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent as the first to w-th chemical modifications.

[0193] Therefore, the first to w-th chemical modifications can include at least one chemical modification applied to the sugar portion of the nucleotide sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent, at least one chemical modification applied to the phosphate portion of the nucleotide sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent, and at least one chemical modification applied to the base portion of the nucleotide sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0194] The nucleotide sequence modification module 300 can generate an optimal chemical modification prediction model capable of determining the chemical modification having the best biological properties when applied to the nucleotide sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent among the first to w-th chemical modifications by training an artificial neural network using the data stored in the chemical modification information database 310 as training data.

[0195] Specifically, the nucleotide sequence modification module 300 randomly selects at least two sequences in which different chemical modifications are applied to the same nucleotide sequence from the training data obtained from the chemical modification information database 310 and sequentially inputs them into the artificial neural network, compares the output values output for each of the at least two sequences input into the artificial neural network, and repeats the process of training the artificial neural network so that the artificial neural network outputs a larger value as the biological properties of the sequence input into the artificial neural network are better, thereby generating the optimal chemical modification prediction model.

[0196] Hereinafter, taking the process in which the base sequence modification module 300 randomly selects two sequences to which different chemical modifications are applied to the same base sequence among the learning data and trains the artificial neural network to generate the optimal chemical modification prediction model as an example for explanation.

[0197] The base sequence modification module 300 randomly selects a pair of a first sequence and a second sequence to which different chemical modifications are applied to the same base sequence among the learning data, sequentially inputs the first sequence and the second sequence into the artificial neural network, and obtains a first output value output from the artificial neural network for the first sequence and a second output value output from the artificial neural network for the second sequence.

[0198] Thereafter, when the biological property for the first sequence is better than the biological property for the second sequence, the base sequence modification module 300 trains the artificial neural network so that the value obtained by subtracting the second output value from the first output value increases, and when the biological property for the second sequence is better than the biological property for the first sequence, the artificial neural network can be trained so that the value obtained by subtracting the first output value from the second output value increases.

[0199] The base sequence modification module 300 randomly selects at least two sequences to which different chemical modifications are applied to the same base sequence among the learning data, and repeatedly performs the operation of performing the learning process as described above on the artificial neural network using the at least two selected sequences, thereby generating the optimal chemical modification prediction model.

[0200] Therefore, when a sequence to which a specific chemical modification has been applied to a specific base sequence is input, the optimal chemical modification prediction model can output a larger output value as the biological characteristics of the input sequence are better.

[0201] As described above, the base sequence modification module 300 randomly selects two sequences to which different chemical modifications have been applied to the same base sequence from the learning data to train the artificial neural network to generate the optimal chemical modification prediction model. However, it is also possible to randomly select three or more sequences to which different chemical modifications have been applied to the same base sequence from the learning data, input them into the artificial neural network, compare the output values output for each of the at least three or more sequences, train the artificial neural network, and generate the optimal chemical modification prediction model in the same manner as described above.

[0202] As described above, the optimal chemical modification prediction model outputs a larger output value as the biological characteristics of the input sequence are better. However, the present invention is not limited to this.

[0203] According to an embodiment, the base sequence modification module 300 may train the artificial neural network so that the artificial neural network outputs a smaller output value as the biological characteristics of the sequence input to the artificial neural network are better, and generate the optimal chemical modification prediction model. In this case, the optimal chemical modification prediction model can output a smaller output value as the biological characteristics of the input sequence are better. In one embodiment, the artificial neural network may be suitable for a Siamese neural network.

[0204] However, the present invention is not limited thereto, and the base sequence modification module 300 can generate the optimal chemical modification prediction model using various types of artificial neural networks.

[0205] On the other hand, after generating the optimal chemical modification prediction model, when the base sequence modification module 300 receives the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent from the base sequence generation module 200 as shown in FIGS. 1 and 4, using the optimal chemical modification prediction model, among the first to wth chemical modifications, a chemical modification having excellent biological characteristics when applied to the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent can be determined as the optimal chemical modification (O_CHEM_MOD).

[0206] In addition, the base sequence modification module 300 can provide the optimal chemical modification (O_CHEM_MOD) and the position where the optimal chemical modification (O_CHEM_MOD) is applied to the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0207] Hereinafter, an example of a method for the base sequence modification module 300 to determine, as the optimal chemical modification (O_CHEM_MOD), a chemical modification having excellent biological characteristics when applied to the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent among the first to wth chemical modifications using the optimal chemical modification prediction model will be described.

[0208] The base sequence modification module 300 can apply each of the first to wth chemical modifications to the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent to generate a plurality of sequences.

[0209] For example, the base sequence modification module 300 generates sequences corresponding to the results of applying the first chemical modification to all positions where the first chemical modification is applicable to the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent, generates sequences corresponding to the results of applying the second chemical modification to all positions where the second chemical modification is applicable to the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent, and generates sequences corresponding to the results of applying the w-th chemical modification to all positions where the w-th chemical modification is applicable to the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent. Therefore, the number of the plurality of sequences generated from the base sequence modification module 300 may be much more than w.

[0210] Also, each of the plurality of sequences can show the molecular structure for each case where the first to w-th chemical modifications are applied to a plurality of applicable positions in the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0211] On the other hand, it is general that the biological properties are improved when chemical modifications are applied to positions corresponding to the bases at one end and the bases at the other end in the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent, rather than when chemical modifications are applied to the positions corresponding to the bases in the central part of the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0212] Therefore, according to the embodiment, the base sequence modification module 300 may generate the plurality of sequences by applying each of the first to w-th chemical modifications limited to the positions corresponding to the first a (a is an integer) bases and the last b (b is an integer) bases in the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0213] Thereafter, the nucleotide sequence modification module 300 can sequentially input the plurality of sequences into the optimal chemical modification prediction model to obtain a plurality of output values output from the optimal chemical modification prediction model for the plurality of sequences.

[0214] The nucleotide sequence modification module 300 can determine the chemical modification applied to the sequence corresponding to the maximum value of the plurality of output values among the plurality of sequences as the optimal chemical modification (O_CHEM_MOD).

[0215] In addition, the nucleotide sequence modification module 300 can provide the position where the optimal chemical modification (O_CHEM_MOD) and the optimal chemical modification (O_CHEM_MOD) in the sequence corresponding to the maximum value of the plurality of output values are applied to the nucleotide sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0216] According to an example, the nucleotide sequence modification module 300 may determine all chemical modifications applied to the sequence corresponding to the output value having a threshold value or more among the plurality of output values as the optimal chemical modification (O_CHEM_MOD).

[0217] Hereinafter, another example of a method for the nucleotide sequence modification module 300 to determine, using the optimal chemical modification prediction model, the chemical modification having excellent biological characteristics when applied to the nucleotide sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent among the first to w-th chemical modifications as the optimal chemical modification (O_CHEM_MOD) will be described. The nucleotide sequence modification module 300 can generate combinations of a plurality of chemical modifications by combining one or more of the first to w-th chemical modifications.

[0218] For example, the nucleotide sequence modification module 300 can generate combinations of the plurality of chemical modifications by variously combining two or more chemical modifications among the first to w-th chemical modifications.

[0219] Subsequently, the base sequence modification module 300 can generate a plurality of sequences by applying each of the combinations of the plurality of chemical modifications to the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0220] For example, the base sequence modification module 300 applies the first combination of chemical modifications to all positions where two or more chemical modifications included in the first combination of chemical modifications are applicable to the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent, and generates a sequence corresponding to the result of the application. Similarly, for all positions where two or more chemical modifications included in the second combination of chemical modifications are applicable to the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent, a sequence corresponding to the result of the application of the second combination of chemical modifications is generated, and in this way, the plurality of sequences can be generated. Therefore, the number of the plurality of sequences generated from the base sequence modification module 300 can be much more than w.

[0221] Moreover, each of the plurality of sequences can show the molecular structure for each case where each of the combinations of the plurality of chemical modifications is applied to a plurality of applicable positions in the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0222] On the other hand, generally, when chemical modifications are applied to positions corresponding to bases at one end and the other end of the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent, the biological properties are better than when chemical modifications are applied to positions corresponding to the central base in the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0223] Therefore, according to the embodiment, the base sequence modification module 300 may generate the plurality of sequences by applying each of the combinations of the plurality of chemical modifications only to positions corresponding to the first a (a is an integer) bases and the last b (b is an integer) bases in the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0224] Thereafter, the base sequence modification module 300 can sequentially input the plurality of sequences into the optimal chemical modification prediction model to obtain a plurality of output values output from the optimal chemical modification prediction model for the plurality of sequences.

[0225] The base sequence modification module 300 can determine the combination of chemical modifications applied to the sequence corresponding to the maximum value of the plurality of output values among the plurality of sequences as the optimal chemical modification (O_CHEM_MOD).

[0226] In addition, the base sequence modification module 300 can provide the position where the optimal chemical modification (O_CHEM_MOD) and the optimal chemical modification (O_CHEM_MOD) in the sequence corresponding to the maximum value of the plurality of output values are applied to the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0227] According to an embodiment, the base sequence modification module 300 may determine, as the optimal chemical modification (O_CHEM_MOD), all combinations of chemical modifications applied to the sequences corresponding to the output values having a threshold value or more among the plurality of output values.

[0228] As described above with reference to FIGS. 1 to 6, the RNA therapeutic agent design systems 10 and 20 according to the embodiments of the present invention not only provide the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent that can effectively prevent side effects by maximizing the therapeutic effect on the specific disease while minimizing the degree of regulating the activities of a plurality of non-target mRNAs (OFF_T_MRNA) by regulating the activity of the target mRNA (T_MRNA), but also provide, together with the optimal chemical modification (O_CHEM_MOD) that can further improve the therapeutic efficiency and in vivo stability of the RNA therapeutic agent when applied to the base sequence (RNA_DRUG_SEQ) of the RNA therapeutic agent.

[0229] Therefore, the RNA therapeutic agent design systems 10 and 20 according to the embodiments of the present invention can effectively design an RNA therapeutic agent that can reduce side effects and improve in vivo stability while having a high therapeutic effect.

Industrial Applicability

[0230] The present invention can be usefully utilized to design an RNA therapeutic agent that can prevent side effects of regulating the activity of other non-target mRNAs while effectively regulating the activity of the target mRNA.

[0231] As described above, the present invention has been described with reference to the preferred embodiments. However, those having ordinary knowledge in the art can make various modifications and changes to the present invention without departing from the spirit and scope of the present invention described in the following claims.

Claims

1. A step of obtaining, as learning data, data stored in a chemical modification information database in which a base sequence modification module stores numerical values indicating biological characteristics when a plurality of chemical modifications are applied to a plurality of base sequences; a step of repeatedly performing a process of training the artificial neural network so that the artificial neural network outputs a larger value as the biological characteristics of the sequence input to the artificial neural network are better, by randomly selecting at least two sequences in which different chemical modifications are applied to the same base sequence among the learning data, sequentially inputting the at least two sequences into the artificial neural network, and comparing output values output for each of the at least two sequences input to the artificial neural network, thereby generating an optimal chemical modification prediction model; a step of receiving, by the base sequence modification module, a base sequence of an RNA therapeutic agent for regulating the activity of a target mRNA (Messenger Ribonucleic Acid) related to the induction of a specific disease; a step of determining, by the base sequence modification module, as an optimal chemical modification, a chemical modification having excellent biological characteristics when applied to the base sequence of the RNA therapeutic agent, among the first to w-th (w is an integer of 2 or more) chemical modifications, using the optimal chemical modification prediction model. A method for determining an optimal chemical modification for a base sequence of an RNA therapeutic agent, characterized by the above.

2. The step in which the base sequence modification module generates the optimal chemical modification prediction model includes: a step of randomly selecting a pair of a first sequence and a second sequence in which different chemical modifications are applied to the same base sequence among the learning data; a step of sequentially inputting the first sequence and the second sequence into the artificial neural network, and obtaining a first output value output from the artificial neural network for the first sequence and a second output value output from the artificial neural network for the second sequence; When the biological property for the first sequence is better than the biological property for the second sequence, the artificial neural network is trained such that the value obtained by subtracting the second output value from the first output value increases; when the biological property for the second sequence is better than the biological property for the first sequence, the artificial neural network is trained such that the value obtained by subtracting the first output value from the second output value increases, and generating an optimal chemical modification prediction model. The method for determining an optimal chemical modification for the base sequence of the RNA therapeutic agent according to claim 1.

3. The artificial neural network corresponds to a Siamese neural network. The method for determining an optimal chemical modification for the base sequence of the RNA therapeutic agent according to claim 1.

4. The numerical value indicating the biological property includes at least one of biological activity and inhibition rate indicating efficacy, ED50, LD50, and IC50 indicating drug potency, ALT (Alanine aminotransferase), AST (aspartate aminotransferase), Bilirubin, and Creatinine indicating drug toxicity, and elimination half-life. The method for determining an optimal chemical modification for the base sequence of the RNA therapeutic agent according to claim 1.

5. The step in which the base sequence modification module determines, as the optimal chemical modification, the chemical modification having the best biological property when applied to the base sequence of the RNA therapeutic agent among the first to w-th chemical modifications using the optimal chemical modification prediction model includes: applying each of the first to w-th chemical modifications to the base sequence of the RNA therapeutic agent to generate a plurality of sequences; sequentially inputting the plurality of sequences into the optimal chemical modification prediction model; obtaining a plurality of output values output from the optimal chemical modification prediction model for the plurality of sequences. Determining, as the optimal chemical modification, the chemical modification applied to the sequence corresponding to the maximum value of the plurality of output values among the plurality of sequences. A method for determining an optimal chemical modification for the base sequence of the RNA therapeutic agent according to claim 1.

6. The step in which the base sequence modification module applies each of the first to w-th chemical modifications to the base sequence of the RNA therapeutic agent to generate the plurality of sequences includes: including the step of applying each of the first to w-th chemical modifications to generate the plurality of sequences, limited to positions corresponding to the first a (a is an integer) bases and the last b (b is an integer) bases in the base sequence of the RNA therapeutic agent. A method for determining an optimal chemical modification for the base sequence of the RNA therapeutic agent according to claim 5.

7. The step in which the base sequence modification module uses the optimal chemical modification prediction model to determine, as the optimal chemical modification, the chemical modification having the best biological properties when applied to the base sequence of the RNA therapeutic agent among the first to w-th chemical modifications includes: generating a plurality of combinations of chemical modifications by combining one or more of the first to w-th chemical modifications; applying each of the plurality of combinations of chemical modifications to the base sequence of the RNA therapeutic agent to generate a plurality of sequences; sequentially inputting the plurality of sequences into the optimal chemical modification prediction model; obtaining a plurality of output values output from the optimal chemical modification prediction model for the plurality of sequences; and determining, as the optimal chemical modification, the combination of chemical modifications applied to the sequence corresponding to the maximum value of the plurality of output values among the plurality of sequences. A method for determining an optimal chemical modification for the base sequence of the RNA therapeutic agent according to claim 1.

8. The step in which the base sequence modification module applies each of the plurality of combinations of chemical modifications to the base sequence of the RNA therapeutic agent to generate the plurality of sequences includes: including the step of applying each of the plurality of combinations of chemical modifications to generate the plurality of sequences, limited to positions corresponding to the first a (a is an integer) bases and the last b (b is an integer) bases in the base sequence of the RNA therapeutic agent. Method for determining optimal chemical modification for the base sequence of the RNA therapeutic agent according to claim 7.

9. The first to w-th chemical modifications include at least one chemical modification applied to the sugar portion of the base sequence of the RNA therapeutic agent, at least one chemical modification applied to the phosphate portion of the base sequence of the RNA therapeutic agent, and at least one chemical modification applied to the base portion of the base sequence of the RNA therapeutic agent. Method for determining optimal chemical modification for the base sequence of the RNA therapeutic agent according to claim 1.

10. The base sequence generation module further includes a step of determining the base sequence of the RNA therapeutic agent, wherein the degree of regulating the activity of the target mRNA is large, and the degree of regulating the activities of a plurality of non-target mRNAs other than the target mRNA is small. The step in which the base sequence generation module determines the base sequence of the RNA therapeutic agent is as follows: Determining a candidate base sequence; Increasing the reward as the degree of binding of the candidate base sequence to the target mRNA and regulating the activity of the target mRNA increases, and decreasing the reward as the degree of binding of the candidate base sequence to each of the plurality of non-target mRNAs and regulating the activity of each of the plurality of non-target mRNAs increases, and determining the reward for the candidate base sequence. Repeatedly performing a process of calculating the reward while diversely modifying the candidate base sequence and modifying the candidate base sequence in a direction in which the reward increases. When the compensation no longer increases due to the modification of the candidate base sequence, determining the final candidate base sequence as the base sequence of the RNA therapeutic agent. Method for determining optimal chemical modification for the base sequence of the RNA therapeutic agent according to claim 1.

11. The base sequence generation module includes a step of determining the base sequence of the RNA therapeutic agent, wherein the degree of regulating the activity of the target mRNA is large, and the degree of regulating the activities of a plurality of non-target mRNAs other than the target mRNA is small. The secondary structure prediction module further includes a step of predicting the secondary structure in which the target mRNA is folded and estimating the autonomous binding relationship between the base sequences of the target mRNA. The step in which the base sequence generation module determines the base sequence of the RNA therapeutic agent is as follows: The step in which the base sequence generation module determines the base sequence of the RNA therapeutic agent, based on the autonomous binding relationship between the base sequences of the target mRNA, includes a step of determining the base sequence of the RNA therapeutic agent that greatly regulates the activity of the target mRNA and slightly regulates the activities of the plurality of non-target mRNAs. A method for determining an optimal chemical modification for the base sequence of the RNA therapeutic agent according to claim 1.

12. The step in which the base sequence generation module determines the base sequence of the RNA therapeutic agent, based on the autonomous binding relationship between the base sequences of the target mRNA, is as follows: A step of determining a candidate base sequence; A step of increasing the compensation (reward) as the degree to which the candidate base sequence binds to the base sequence of the target mRNA and regulates the activity of the target mRNA increases, and decreasing the compensation as the degree to which the candidate base sequence binds to each of the plurality of non-target mRNAs and regulates the activity of each of the plurality of non-target mRNAs increases; A step of determining a secondary structure penalty proportional to the ratio of the bases that form a bond by the target mRNA itself among the bases of the target mRNA that bind to the candidate base sequence, based on the autonomous binding relationship between the base sequences of the target mRNA; A step of subtracting the secondary structure penalty from the compensation to determine the compensation for the candidate base sequence; A step of repeatedly performing a process of calculating the compensation while diversely modifying the candidate base sequence and modifying the candidate base sequence in a direction in which the compensation increases; A step of determining the final candidate base sequence as the base sequence of the RNA therapeutic agent when the compensation no longer increases due to the modification of the candidate base sequence. A method for determining an optimal chemical modification for the base sequence of the RNA therapeutic agent according to claim 11.

13. A step in which the base sequence generation module determines the base sequence of the RNA therapeutic agent that greatly regulates the activity of the target mRNA and slightly regulates the activities of a plurality of non-target mRNAs other than the target mRNA. The secondary structure prediction module further includes the step of predicting, for each of the target mRNA and the plurality of non-target mRNAs, the secondary structure (folding) in which the corresponding mRNA is folded, and estimating the autonomous binding relationship between the base sequences of the corresponding mRNA. The step in which the base sequence generation module determines the base sequence of the RNA therapeutic agent is as follows. The step in which the base sequence generation module determines the base sequence of the RNA therapeutic agent based on the autonomous binding relationship between the base sequences of each of the target mRNA and the plurality of non-target mRNAs includes the step of determining the base sequence of the RNA therapeutic agent that greatly regulates the activity of the target mRNA and slightly regulates the activity of the plurality of non-target mRNAs. A method for determining an optimal chemical modification for the base sequence of the RNA therapeutic agent according to claim 1.

14. The step in which the base sequence generation module determines the base sequence of the RNA therapeutic agent based on the autonomous binding relationship between the base sequences of each of the target mRNA and the plurality of non-target mRNAs is as follows. The step of determining a candidate base sequence. The step of increasing the reward as the degree of binding between the candidate base sequence and the base sequence of the target mRNA and regulating the activity of the target mRNA increases, and decreasing the reward as the degree of binding between the candidate base sequence and the base sequence of each of the plurality of non-target mRNAs and regulating the activity of each of the plurality of non-target mRNAs increases. The step of determining a target secondary structure penalty proportional to the ratio of the bases that form a bond by the target mRNA itself among the bases of the target mRNA that bind to the candidate base sequence based on the autonomous binding relationship between the base sequences of the target mRNA. For each of the plurality of non-target mRNAs, the step of determining a non-target secondary structure penalty proportional to the ratio of the bases that form a bond by the non-target mRNA itself among the bases of the non-target mRNA that bind to the candidate base sequence based on the autonomous binding relationship between the base sequences of the non-target mRNA. The step of subtracting the target secondary structure penalty from the reward, summing up the non-target secondary structure penalties for each of the plurality of non-target mRNAs, and determining the reward for the candidate base sequence. Calculating the compensation while diversely modifying the candidate base sequence, and repeating the process of modifying the candidate base sequence in the direction in which the compensation increases; When the compensation no longer increases due to the modification to the candidate base sequence, determining the final candidate base sequence as the base sequence of the RNA therapeutic agent. A method for determining an optimal chemical modification for the base sequence of the RNA therapeutic agent according to claim 13.

15. The non-target analysis module further includes a step of determining, as the plurality of non-target mRNAs, mRNAs having a gene expression pattern similar to the target mRNA among the plurality of mRNAs contained in the human body. A method for determining an optimal chemical modification for the base sequence of the RNA therapeutic agent according to any one of claims 10, 11, and 13.

16. Obtaining, as learning data, data stored in a chemical modification information database that stores numerical values indicating biological characteristics when a plurality of chemical modifications are applied to a plurality of base sequences, randomly selecting at least two sequences in which different chemical modifications are applied to the same base sequence from among the learning data, and sequentially inputting them into an artificial neural network, comparing the output values output for each of the at least two sequences input into the artificial neural network, and repeating the process of training the artificial neural network so that the artificial neural network outputs a larger value as the biological characteristics of the sequence input into the artificial neural network are better, to generate an optimal chemical modification prediction model, including a base sequence modification module; When the base sequence modification module receives the base sequence of an RNA therapeutic agent for regulating the activity of a target mRNA (Messenger Ribonucleic Acid) related to the induction of a specific disease, using the optimal chemical modification prediction model, among the first to wth (w is an integer of 2 or more) chemical modifications, determining a chemical modification having excellent biological characteristics when applied to the base sequence of the RNA therapeutic agent as the optimal chemical modification. A system for determining an optimal chemical modification for the base sequence of an RNA therapeutic agent, characterized by the above.

Citation Information

Patent Citations

  • Predicting Response to Chemotherapy Using Gene Expression Markers

    JP2008520192A