Method and device for manufacturing antiviral agent

WO2026177604A1PCT designated stage Publication Date: 2026-08-27CALICI CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/095072
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-18
Filing Date
2026-02-19
Publication Date
2026-08-27

Smart Images

  • Figure KR2026095072_27082026_PF_FP_ABST
    Figure KR2026095072_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A method and a device for manufacturing an antiviral agent are provided. The method for manufacturing an antiviral agent is performed by a computing device including a processor, a memory and a storage device, and may comprise the steps of: acquiring structural data relating to a plurality of target proteins corresponding to a plurality of virus-derived proteins; receiving structural data of a plurality of compounds from a pre-prepared compound structure library; predicting binding affinity for target protein-compound pairs combined from the plurality of target proteins and the plurality of compounds; calculating a weighted score by applying a plurality of weights determined differently according to expression levels of the plurality of virus-derived proteins to the binding affinity; determining a compound to be used as an antiviral agent from among the plurality of compounds according to the weighted score; and generating manufacturing process data of the antiviral agent so as to include chemical property values relating to the determined compound, and storing the manufacturing process data in the storage device.
Need to check novelty before this filing date? Find Prior Art

Description

Method and apparatus for manufacturing antiviral agents

[0001] Cross-citation with related applications

[0002] This application claims the benefit of priority based on Korean Patent Application No. 10-2025-0020638 dated February 18, 2025, and all contents disclosed in the document of said Korean patent application are incorporated herein as part of this specification.

[0003] The disclosure relates to a method and apparatus for manufacturing an antiviral agent.

[0004] The COVID-19 pandemic has highlighted the importance of developing effective antiviral drugs. Paxlovid and Remdesivir are antiviral drugs that target the non-structural proteins of the coronavirus. Paxlovid interferes with the processing and maturation of viral proteins by selectively inhibiting 3CL protease, a protease essential for SARS-CoV-2 replication. Meanwhile, Remdesivir inhibits replication and proliferation by blocking viral RNA synthesis through the targeting of RNA-dependent RNA polymerase (RdRp). Although these drugs demonstrate high antiviral efficacy based on a high-specificity approach targeting a single protein, they have the disadvantage that their therapeutic effect may decrease if resistance to a specific protein develops. Accordingly, there is a need to develop antiviral drugs that can overcome the limitations of the single-target approach and exert a broader and sustained effect.

[0005] The objective is to provide a method and apparatus for manufacturing an antiviral agent that can reduce the potential for viral resistance development and provide a more potent antiviral effect by designing the antiviral agent using a multi-target approach.

[0006] A method for manufacturing an antiviral agent according to one embodiment is a method for manufacturing an antiviral agent performed by a computing device comprising a processor, a memory, and a storage device, wherein the processor acquires sequence or structural data regarding a plurality of target proteins corresponding to a plurality of virus-derived proteins; the processor receives structural data of a plurality of compounds from a pre-prepared compound structure library; the processor predicts a binding affinity for a target protein-compound pair formed from the plurality of target proteins and the plurality of compounds; the processor calculates a weighted score by applying a plurality of weights, which are determined differently according to the expression amount of the plurality of virus-derived proteins, to the binding affinity; the processor determines a compound to be used as the antiviral agent among the plurality of compounds according to the weighted score; and the processor generates manufacturing process data for the antiviral agent to include chemical characteristic values ​​regarding the determined compound and stores the manufacturing process data in the storage device.

[0007] In some embodiments, the plurality of weights may include values ​​obtained by performing RNA (Ribonucleic Acid) copy number sequencing, wherein the expression levels of the plurality of target proteins are normalized as a ratio to the total expression levels.

[0008] In some embodiments, the plurality of weights may include a value obtained by performing a predetermined operation on a value obtained by normalizing the expression amount of each of the plurality of target proteins at a first time point and a value obtained by normalizing the expression amount of each of the plurality of target proteins at a second time point different from the first time point.

[0009] In some embodiments, the step of obtaining the sequence or structural data may include the step of obtaining the sequence or structural data regarding the plurality of target proteins corresponding to the non-structural proteins among the plurality of virus-derived proteins.

[0010] In some embodiments, the non-structural protein may include at least two of the coronavirus NSP1, NSP2, NSP3, NSP4, NSP5, NSP6, NSP7, NSP8, NSP9, NSP10, NSP12, NSP13, NSP14, NSP15, and NSP16.

[0011] In some embodiments, the step of obtaining the sequence or structural data may include the step of obtaining the sequence or structural data regarding the plurality of target proteins corresponding to the structural protein among the plurality of virus-derived proteins.

[0012] In some embodiments, the step of predicting the binding force may include: the step of the processor performing docking for the target protein-compound pair to predict the binding pose for the target protein-compound pair; the step of the processor loading a pre-trained binding affinity prediction model into the memory; and the step of the processor inputting the docked target protein-compound pair into the binding affinity prediction model and obtaining the binding affinity value output from the binding affinity prediction model as the binding force.

[0013] In some embodiments, the bond affinity prediction model may be trained to take a 4-dimensional tensor containing the 3-dimensional coordinates of an atom and one training feature vector as input and output a predicted value of the bond affinity.

[0014] In some embodiments, the binding affinity prediction model may be trained using experimental values ​​regarding the binding strength of the protein and the compound as training data.

[0015] In some embodiments, the step of calculating the weighted score may include: the processor loading a plurality of binding strength values ​​obtained for each of the plurality of target proteins for each compound into the memory as a plurality of first scores; the processor loading a plurality of weights into the memory; the processor calculating a plurality of second scores by applying the plurality of weights to the plurality of first scores loaded into the memory; the processor calculating a final score determined as a single value for each compound from the plurality of second scores; and the processor calculating the final score as the weighted score.

[0016] An antiviral agent manufacturing apparatus according to one embodiment comprises: a processor; and a memory in which an instruction executed by the processor is loaded. The instruction is executed by the processor to cause the processor to obtain sequence or structural data regarding a plurality of target proteins corresponding to a plurality of virus-derived proteins, receive structural data of a plurality of compounds from a pre-prepared compound structure library, predict binding strength for target protein-compound pairs combined from the plurality of target proteins and the plurality of compounds, calculate a weighted score by applying a plurality of weights, which are determined differently according to the expression amount of the plurality of virus-derived proteins, to the binding strength, determine a compound to be used as the antiviral agent among the plurality of compounds according to the weighted score, generate manufacturing process data for the antiviral agent including chemical characteristic values ​​regarding the determined compound, and store the manufacturing process data in a storage device.

[0017] In some embodiments, the plurality of weights may include values ​​obtained by performing RNA (Ribonucleic Acid) copy number sequencing, wherein the expression levels of the plurality of target proteins are normalized as a ratio to the total expression levels.

[0018] In some embodiments, the plurality of weights may include a value obtained by performing a predetermined operation on a value obtained by normalizing the expression amount of each of the plurality of target proteins at a first time point and a value obtained by normalizing the expression amount of each of the plurality of target proteins at a second time point different from the first time point.

[0019] In some embodiments, obtaining the sequence or structural data may include obtaining sequence or structural data regarding the plurality of target proteins corresponding to non-structural proteins among the plurality of virus-derived proteins.

[0020] In some embodiments, the non-structural protein may include at least two of the coronavirus NSP1, NSP2, NSP3, NSP4, NSP5, NSP6, NSP7, NSP8, NSP9, NSP10, NSP12, NSP13, NSP14, NSP15, and NSP16.

[0021] In some embodiments, obtaining the sequence or structural data may include obtaining the sequence or structural data regarding the plurality of target proteins corresponding to the structural proteins among the plurality of virus-derived proteins.

[0022] In some embodiments, predicting the binding strength may include performing docking on the target protein-compound pair to predict the binding pose for the target protein-compound pair, loading a pre-trained binding affinity prediction model into the memory, inputting the docked target protein-compound pair into the binding affinity prediction model, and obtaining the binding affinity value output from the binding affinity prediction model as the binding strength.

[0023] In some embodiments, the bond affinity prediction model may be trained to take a 4-dimensional tensor containing the 3-dimensional coordinates of an atom and one training feature vector as input and output a predicted value of the bond affinity.

[0024] In some embodiments, calculating the weighted score may include, for each compound, loading a plurality of binding strength values ​​obtained for each of the plurality of target proteins into the memory as a plurality of first scores, loading a plurality of weights into the memory, applying the plurality of weights to the plurality of first scores loaded into the memory to calculate a plurality of second scores, calculating a final score determined as a single value for each compound from the plurality of second scores, and calculating the final score as the weighted score.

[0025] A computer-readable recording medium according to one embodiment is a computer-readable recording medium that records instructions executed by a computing device including a processor, memory, and a storage device, wherein the instructions are executed by the computing device to enable the computing device to acquire sequence or structural data regarding a plurality of target proteins corresponding to a plurality of virus-derived proteins, receive structural data of a plurality of compounds from a pre-prepared compound structure library, predict binding strength for target protein-compound pairs combined from the plurality of target proteins and the plurality of compounds, calculate a weighted score by applying a plurality of weights, which are determined differently according to the expression amount of the plurality of virus-derived proteins, to the binding strength, determine a compound to be used as an antiviral agent among the plurality of compounds according to the weighted score, generate manufacturing process data for the antiviral agent including chemical characteristic values ​​regarding the determined compound, and store the manufacturing process data in the storage device.

[0026] According to the embodiments, antiviral agents prepared using a multi-target approach are not dependent on a single specific protein, and thus can significantly reduce the potential for viral resistance development and provide a more potent antiviral effect. Furthermore, because they are developed to exert optimal action based on the relative importance and expression levels of target proteins, they can deliver a broader and more sustained effect than conventional single-target antiviral agents. Additionally, through their multi-target characteristics, they can provide a significant advancement in inhibiting the entire viral replication mechanism while simultaneously suppressing the development of resistance.

[0027] FIG. 1 is a drawing for explaining an antiviral agent manufacturing apparatus according to one embodiment.

[0028] FIGS. 2 and FIGS. 3 are drawings for explaining the operation of an antiviral drug manufacturing apparatus according to one embodiment.

[0029] FIG. 4 is a flowchart illustrating a method for manufacturing an antiviral agent according to one embodiment.

[0030] FIGS. 5 and 6 are drawings for explaining the operation of an antiviral drug manufacturing apparatus according to one embodiment.

[0031] FIG. 7 is a flowchart illustrating a method for manufacturing an antiviral agent according to one embodiment.

[0032] FIG. 8 is a diagram illustrating an example of implementing a binding affinity prediction model used in the manufacture of an antiviral agent according to one embodiment.

[0033] FIG. 9 is a block diagram illustrating a computing device according to one embodiment.

[0034] Embodiments of the present invention are described below with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present invention may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present invention in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.

[0035] Throughout the specification and claims, when a part is described as "comprising" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another.

[0036] Terms such as "...part," "...device," and "module" as described in the specification may refer to a unit capable of processing at least one function or operation described in this specification, and may be implemented as hardware or circuit, software, or a combination of hardware or circuit and software. Additionally, at least some components or functions of the antiviral agent manufacturing method and apparatus according to the embodiments described below may be implemented as a program or software, and the program or software may be stored on a computer-readable recording medium.

[0037] FIG. 1 is a drawing for explaining an antiviral agent manufacturing apparatus according to one embodiment.

[0038] Referring to FIG. 1, an antiviral agent manufacturing apparatus (10) according to one embodiment may execute program code or instructions loaded into one or more memory devices through one or more processors. For example, the antiviral agent manufacturing apparatus (10) may be implemented as a computing device (50) as described below in relation to FIG. 9. In this case, one or more processors may correspond to a processor (510) of the computing device (50), and one or more memory devices may correspond to a memory (520) of the computing device (50). The program code or instructions may be executed by one or more processors to determine a compound to be used as an antiviral agent and to derive manufacturing process data for the antiviral agent therefrom. In this specification, the term "module" is used to logically distinguish these functions performed by the program code or instructions.

[0039] The antiviral drug manufacturing device (10) can discover a compound suitable for maximizing the overall inhibition effect on viral proliferation and manufacture an antiviral drug using it. That is, manufacturing process data is configured including the compound discovered by the antiviral drug manufacturing device (10), and the manufacturing process data can be used for the actual manufacture of an antiviral drug. To this end, the antiviral drug manufacturing device (10) may include a structural data acquisition module (101), a binding force prediction module (102), a compound determination module for the antiviral drug (103), and a manufacturing process data storage module (104).

[0040] The structural data acquisition module (101) can acquire sequence or structural data (e.g., three-dimensional structural data) regarding a plurality of target proteins corresponding to a plurality of virus-derived proteins. Taking coronavirus as an example, the plurality of virus-derived proteins may include structural proteins and non-structural proteins.

[0041] Structural proteins are proteins that constitute the structure of the virus and can be produced by translation from coding sequences located in the latter part of the viral genome. Examples of structural proteins include the spike protein (S), envelope protein (E), membrane protein (M), and nucleocapsid protein (N). The spike protein (S) plays a crucial role in the virus binding to and penetrating host cells, and as the major surface antigen of the coronavirus, it serves as an important target in the immune response (a target of the coronavirus vaccine). The envelope protein (E) forms the viral envelope and performs important functions during the assembly of the viral particle and its release into host cells. The membrane protein (M) contributes to maintaining the structure of the virus and determining the shape of the particle. The nucleocapsid protein (N) protects the viral genetic material (RNA) and plays an essential role in the process of viral replication.

[0042] Meanwhile, non-structural proteins are proteins essential for viral replication and survival, and can be produced from long polyproteins translated from ORF1a and ORF1b in the anterior part of the viral genome. ORF1a and ORF1b are the largest genetic regions of the virus, encoding non-structural proteins and capable of producing enzymes and regulatory proteins essential for viral replication. ORF1a accounts for approximately 67% of the total genome and can encode non-structural proteins. ORF1b accounts for the remaining approximately 33% and is linked to ORF1a but is translated via ribosomal frameshift. This is a translation mechanism in which ribosomes shift frames at the end of ORF1a to translate ORF1b.

[0043] Specifically, the genome of the coronavirus consists of single-stranded positive RNA approximately 30 kb in size, which can generate various non-structural proteins essential for viral replication and transcription. These processes can be regulated primarily by the translation of two open reading frames, ORF1a and ORF1b, and ribosomal frameshifting. ORF1a and ORF1b, located at the 5' end of the genome, account for about two-thirds of the viral genome and can encode key proteins required for viral replication. Between ORF1a and ORF1b, there exist slippery sequences and RNA structures that induce ribosomal frameshifting. These elements cause ribosomes to undergo a -1 frameshift during translation, enabling the sequential translation of ORF1a and ORF1b. pp1a, generated by the translation of ORF1a, and pp1ab, generated by the sequential translation of ORF1a and ORF1b, can be long polyproteins. This polyprotein can be cleaved by its own built-in proteases (Plpro of nsp3 and 3CLpro of nsp5) to form 16 non-structural proteins. When examining the relationship between the expression levels of Nsp1 to Nsp16 and ORF1a and ORF1b, ORF1a includes Nsp1 to Nsp11 and undergoes relatively high translation. On the other hand, ORF1b includes Nsp12 to Nsp16, but undergoes relatively low translation due to frameshift.

[0044] Examples of non-structural proteins of the coronavirus include NSP1, NSP2, NSP3, NSP4, NSP5, NSP6, NSP7, NSP8, NSP9, NSP10, NSP12, NSP13, NSP14, NSP15, and NSP16.

[0045] NSP1 can inhibit the translation of host mRNA by binding to the 40S ribosomal subunit. Additionally, it can interfere with host protein synthesis by inducing the cleavage of the 5' UTR and degrading the cleaved mRNA. On the other hand, viral mRNA can be protected as it is not cleaved thanks to the 5'-leader sequence. Through these actions, NSP1 can induce immune evasion by suppressing the expression of interferons and cytokines.

[0046] NSP2 can regulate cell survival pathways by interacting with host proteins PHB and PHB2. Additionally, it can regulate mitochondrial function by affecting mitochondrial membrane stability and energy metabolism. By regulating the cellular stress response, NSP2 inhibits host cell apoptosis and creates a cellular environment suitable for viral replication, thereby evading the host immune system.

[0047] In the case of NSP3, it can generate NSP1 through NSP3 by acting as a PLpro (Papain-like Protease) and cleaving polyproteins. Additionally, it can alter the stability of host proteins by removing Lys48 and Lys63 ubiquitin chains and ISG15. NSP3 can form replication vesicles (DMVs) together with NSP4 to generate a replication-transcription complex (RTC). Along with this, IRF3 and NF- K By inhibiting signaling pathway B, interferon (IFN) and cytokine expression can be suppressed, thereby inducing immune evasion.

[0048] In the case of NSP3, it can generate NSP1 through NSP3 by acting as a PLpro (Papain-like Protease) and cleaving polyproteins. Additionally, it can alter the stability of host proteins by removing Lys48 and Lys63 ubiquitin chains and ISG15. NSP3 can form replication vesicles (DMVs) together with NSP4 to generate a replication-transcription complex (RTC). Along with this, IRF3 and NF- K By inhibiting signaling pathway B, interferon (IFN) and cytokine expression can be suppressed, thereby inducing immune evasion.

[0049] In the case of NSP4, it can form a DMV together with NSP3 to provide a site for viral replication. Additionally, it can modify the endoplasmic reticulum (ER) membrane to create a bimembrane structure. By linking the RTC with the DMV, NSP4 can support an environment suitable for RNA replication and transcription. Furthermore, it can evade the host immune system by isolating RNA within the DMV.

[0050] In the case of NSP5, it can isolate NSP4 through NSP16 by cleaving the C-terminal polyprotein. The cleavage site can be determined by recognizing the [ILMVF]-Q-|-[SGACN] sequence. Additionally, it can regulate protein-protein interactions by binding to ADP-ribose-1''-phosphate (ADRP). NSP5 can support RNA synthesis and transcription through the activation of NSP4 through NSP16.

[0051] In the case of NSP6, it can induce the formation of autophagocytos in the endoplasmic reticulum (ER) to provide the space necessary for viral replication. Additionally, it can inhibit the movement of autophagocytos into lysosomes by restricting membrane expansion. NSP6, together with NSP3 and NSP4, forms DMVs to provide suitable space for viral RNA replication and transcription. Furthermore, it can protect viral RNA from the host immune system by altering the autophagy pathway.

[0052] In the case of NSP7, it can support primer synthesis by forming a 16-mer complex with NSP8. Additionally, it can assist in the early stages of RNA synthesis by synthesizing RNA primers (6-10nt). NSP7 can improve the speed and accuracy of RNA synthesis by interacting with NSP12 (RdRp). Furthermore, it can support viral RNA replication and transcription by forming a replication-transcription complex (RTC) with NSP8 and NSP12.

[0053] NSP8 can form a 16-mer complex with NSP7 to perform the role of primase and support RNA primer synthesis. Additionally, it can synthesize not only short RNA primers but also long RNA sequences. NSP8 can interact with NSP12 (RdRp) to improve the speed and accuracy of RNA synthesis. Furthermore, it can form an RTC with NSP7 and NSP12 to support viral RNA replication and transcription.

[0054] NSP9 can maintain RNA stability by non-specifically binding to single-stranded RNA (ssRNA). Additionally, it can support RNA chain elongation and replication by cooperating with RNA polymerase (RdRp). NSP9 can maintain the structural stability of RNA by interacting with the replication-transcription complex (RTC). Furthermore, by binding to ssRNA, it can protect RNA and evade host immune recognition.

[0055] In the case of NSP10, it can correct RNA errors and prevent mutations by stimulating the 3'-5' ExoN activity of NSP14. Additionally, it can promote 2'-O methylation of the mRNA cap structure by stimulating the 2'-O-MTase activity of NSP16. NSP10 can assist in the formation of the 5' cap structure of viral mRNA by supporting the activities of NSP14 and NSP16. Furthermore, it can stabilize the structure and function of the RTC by cooperating with NSP12, NSP14, and NSP16.

[0056] NSP12 is an RNA-dependent RNA polymerase (RdRp) that can serve as a key enzyme responsible for RNA replication and transcription. Through RNA replication, it can replicate the viral genome by converting positive RNA (+) to negative RNA (-) and back to positive RNA (+). Additionally, it can support the transcription of structural proteins (S, E, M, N) by transcribing downstream genomic RNA (sgRNA). NSP12 can cooperate with NSP7, NSP8, NSP9, and NSP10 to help stabilize the replication-transcription RTC.

[0057] NSP13 acts as a helicase and can generate ssRNA by dissociating RNA and DNA double strands in the 5'->3' direction. Additionally, it can support RNA replication and transcription in cooperation with NSP12. NSP13 can stabilize the structure and function of the RTC by cooperating with NSP12, NSP7, and NSP8. It possesses a zinc-binding domain (ZBD) at the N-terminus to assist helicase activity, and Mg 2+ It can provide energy for RNA and DNA dissociation through ATPase-dependent activity.

[0058] NSP14, acting as an exonuclease (ExoN), can correct RNA replication errors and prevent mutations through 3'-5' ExoN activity. Additionally, by acting as a guanine N7-methyltransferase (N7-MTase) and performing 5'-cap methylation, it can support the stability of viral mRNA and immune evasion. NSP14 collaborates with NSP12 (RdRp) to perform RNA replication correction and prevent mutations. Furthermore, it can improve the translation and stability of viral mRNA by forming a 7-methylguanosine (m7G) cap structure. NSP14 can form and stabilize the RTC in cooperation with NSP12, NSP10, NSP16, and others.

[0059] NSP15 specifically recognizes uridine (U) residues to cleave RNA, and in this process, it can generate a 2'-3'-cyclic phosphate and a 5'-OH. Additionally, by cleaving dsRNA, it can inhibit interferon signaling, thereby supporting RNA processing and immune evasion. NSP15 can induce immune evasion by evading PRR recognition and inhibiting the interferon (IFN) pathway. Furthermore, it can cooperate with NSP12 and NSP13 to form and stabilize the RTC.

[0060] NSP16 acts as a 2'-O-methyltransferase (2'-O-MTase) to perform 2'-O methylation of the mRNA cap structure, thereby increasing mRNA translation efficiency and supporting immune evasion. Additionally, it can form the m7GpppNm cap structure together with NSP14 (N7-MTase). NSP16 can induce immune evasion by evading PRR recognition and inhibiting the IFN pathway. Furthermore, it can stabilize the structure and function of the RTC in cooperation with NSP10 and NSP14.

[0061] When targeting structural proteins, the S protein, in particular, can be a crucial target in vaccine development. Inducing an immune response against the S protein can effectively block the process by which the virus binds to and penetrates host cells. In fact, many COVID-19 vaccines utilize an approach that aims to prevent infection by inducing the formation of antibodies against the S protein. However, the S protein is a region where mutations frequently occur. For instance, in several SARS-CoV-2 variants (e.g., Delta and Omicron), mutations in the S protein can lead to a reduction in the efficacy of existing vaccines or antibody therapies.

[0062] In contrast, when targeting non-structural proteins, they can be utilized as stable therapeutic targets because they play an essential role in viral replication and survival and exhibit relatively few mutations. For example, drugs that inhibit RNA-dependent RNA polymerase (RdRp), a viral replication enzyme, or proteases can suppress viral replication, making them highly likely to become effective therapeutic agents less affected by mutations. However, since non-structural proteins are not exposed to the outside of the viral particle, it is difficult to induce an immune response through antibodies. Therefore, non-structural proteins may be more suitable for drug development, such as antivirals, rather than vaccine development.

[0063] In some embodiments, the structural data acquisition module (101) may acquire sequence or structural data regarding a plurality of target proteins corresponding to non-structural proteins among a plurality of virus-derived proteins. In some embodiments, the structural data acquisition module (101) may acquire sequence or structural data regarding a plurality of target proteins corresponding to at least two of NSP1, NSP2, NSP3, NSP4, NSP5, NSP6, NSP7, NSP8, NSP9, NSP10, NSP12, NSP13, NSP14, NSP15, and NSP16 of the coronavirus. It is notable that the scope of the present invention is not limited to a specific non-structural protein selected from NSP1, NSP2, NSP3, NSP4, NSP5, NSP6, NSP7, NSP8, NSP9, NSP10, NSP12, NSP13, NSP14, NSP15, and NSP16. In this specification, when citing examples regarding coronavirus, the case of selecting eight types, NSP1, NSP2, NSP5, NSP9, NSP10, NSP12, NSP13, and NSP14, will be mainly described, but this is solely for the clarity and convenience of explanation and is merely a simple example.

[0064] In some embodiments, target proteins may be selected based on relative expression levels. For example, mRNA expression level (mRNA-seq) analysis showed that the expression level of ORF1a was relatively higher than that of ORF1b, because ORF1a is expressed by normal ribosomal translation, whereas ORF1b is translated by -1 ribosomal frameshifting. Translation level (RPF-seq and QTI-seq) analysis also showed that the translation level of ORF1a was higher than that of ORF1b, because more pp1a generated by ORF1a is translated. Since only some ribosomes perform frameshifting to translate pp1ab, the expression level of ORF1b protein derived from pp1ab is relatively lower. Translation efficiency (TE) analysis showed that the translation efficiency of ORF1a was measured to be relatively higher than that of ORF1b. This is interpreted as meaning that while ORF1a efficiently synthesizes proteins through the general ribosomal translation process, the translation efficiency of ORF1b is lower because it relies on an additional process called ribosomal frameshifting. Through such differences in expression levels, dominant proteins among the coronavirus's non-structural proteins that are important for viral replication and immune evasion can be identified, and target proteins may be selected by considering their relative expression levels.

[0065] In some other embodiments, the structural data acquisition module (101) can acquire sequence or structural data regarding a plurality of target proteins corresponding to structural proteins among a plurality of virus-derived proteins.

[0066] The structure data acquisition module (101) may also receive structure data (e.g., three-dimensional structure data) of a plurality of compounds from a pre-prepared compound structure library. In some embodiments, the compounds may be various natural compounds derived from nature. The compound structure library may store and manage information regarding the structures of these compounds.

[0067] The binding strength prediction module (102) can predict the binding strength for a target protein-compound pair composed of a plurality of target proteins and a plurality of compounds. In some embodiments, the binding strength prediction module (102) can predict the binding affinity for a target protein-compound pair composed of a plurality of target proteins and a plurality of compounds. Here, binding affinity may be an indicator of how strongly a target protein (i.e., the active site of the target protein) and a compound (i.e., a ligand) bind. In this regard, Kd (Dissociation Constant) is a dissociation constant representing the strength of the ligand binding to the target protein, and may represent the concentration at which the ligand binds to about half of the protein. A smaller Kd value may indicate a stronger binding between the ligand and the protein. That is, a low Kd implies a high binding affinity and may indicate that the ligand binds stably to the protein. Ki (Inhibition Constant) is an inhibition constant representing the ability of an antiviral agent to bind to a target protein and inhibit its activity; a lower Ki value may indicate that the antiviral agent binds more strongly to the target protein and inhibits its activity more severely. IC50 (Half Maximal Inhibitory Concentration) may refer to the concentration required for an antiviral agent to inhibit the function of a target protein by 50%. A lower IC50 value indicates that the function of the target protein can be inhibited by more than half at a lower concentration, which may imply that the antiviral agent acts efficiently.

[0068] The binding force prediction module (102) can predict the binding pose of a target protein-compound pair by performing docking on the target protein-compound pair. Here, docking (or molecular docking) may be a computer simulation method that predicts the optimal binding method and pose when a ligand and a target protein bind. Here, the pose may represent the position and orientation in which the ligand binds to the active site (or pocket) of the target protein. The binding force prediction module (102) can define the active site on the target protein where the ligand, i.e., the compound, can bind, and calculate the binding pose and energy by placing the compound at various angles and orientations on the active site of the target protein. The binding force prediction module (102) can select the pose with the lowest binding energy among the binding poses derived from the docking simulation results.

[0069] The binding strength prediction module (102) can load a pre-trained binding affinity prediction model into memory. Here, the binding affinity prediction model may be trained using experimental values ​​regarding the binding strength of proteins and compounds (e.g., Kd, ​​Ki, IC50) as training data, along with protein structures and compound binding structures included in, for example, the PDBbind dataset. The binding affinity prediction model can provide concentration-based kinetic prediction values ​​for a pair in units of uM when a new target protein-compound pair is input. Accordingly, the binding strength prediction module (102) can input a docked target protein-compound pair into the binding affinity prediction model and obtain the output of the binding affinity prediction model as binding affinity. That is, the binding strength prediction module (102) can predict the binding affinity between a target protein and a compound after the active site of the target protein has been selected and the binding of the compound to the active site of the target protein has been satisfied.

[0070] In discovering candidate substances, the binding energy for a specific ligand is calculated as a physical constant in kcal / mol, which is a value that predicts the physical binding strength between a single protein and a single ligand. However, without concentration-based kinetic values ​​that explain how strongly multiple target protein molecules bind to and detach from multiple ligand molecules, experimenters must perform separate experiments at various concentrations ranging from low to high concentrations. In particular, it is difficult to observe the reaction below the effective concentration, and cells may die if the concentration exceeds a certain level, making the process of finding the optimal concentration very complex. Since the binding strength prediction module (102) provides concentration-based kinetic prediction values ​​for target protein-compound pairs in uM units, the user can conveniently estimate how much to dissolve and conduct the experiment based on the molecular weight of the compound.

[0071] The antiviral compound determination module (103) can determine a compound to be used as an antiviral agent among a plurality of compounds based on binding strength (e.g., binding affinity). Specifically, the antiviral compound determination module (103) can determine a compound that shows high binding performance, for example using binding affinity as an indicator, as a compound to be used as an antiviral agent. This binding strength-based compound determination method may be suitable when the number of target proteins is one or a small number.

[0072] When there are multiple target proteins, it is necessary to reflect not only the binding strength but also the interactions with multiple target proteins. Therefore, when the number of target proteins exceeds a preset number, the antiviral compound determination module (103) can determine a compound to be used as an antiviral agent among multiple compounds based on a score calculated by applying weights to the binding strength.

[0073] Specifically, the antiviral compound determination module (103) can calculate a weighted score by applying a plurality of weights, which are determined differently according to the expression amount of a plurality of virus-derived proteins, to the binding strength (e.g., binding affinity). That is, when a plurality of virus-derived proteins include, for example, a first virus-derived protein, a second virus-derived protein, and a third virus-derived protein, a first weight may be determined for the first virus-derived protein, a second weight may be determined for the second virus-derived protein, and a third weight may be determined for the third virus-derived protein. The antiviral compound determination module (103) can calculate a first binding affinity with a first target protein corresponding to the first virus-derived protein for the first compound, and calculate a first weighted score by applying a first weight to the first binding affinity. Additionally, the antiviral compound determination module (103) can calculate a second binding affinity with a second target protein corresponding to a second virus-derived protein for the first compound, and apply a second weight to the second binding affinity to calculate a second weight applied score. Additionally, the antiviral compound determination module (103) can calculate a third binding affinity with a third target protein corresponding to a third virus-derived protein for the first compound, and apply a third weight to the third binding affinity to calculate a third weight applied score. The antiviral compound determination module (103) can calculate a final score for the first compound by combining the first weight applied score, the second weight applied score, and the third weight applied score.

[0074] The antiviral compound determination module (103) can perform the operation performed on the first compound on the second compound, which is another candidate substance. Specifically, the antiviral compound determination module (103) can calculate the first binding affinity with the first target protein corresponding to the first virus-derived protein for the second compound, and calculate the first weighted score by applying the first weight to the first binding affinity. In addition, the antiviral compound determination module (103) can calculate the second binding affinity with the second target protein corresponding to the second virus-derived protein for the second compound, and calculate the second weighted score by applying the second weight to the second binding affinity. In addition, the antiviral compound determination module (103) can calculate the third binding affinity with the third target protein corresponding to the third virus-derived protein for the second compound, and calculate the third weighted score by applying the third weight to the third binding affinity. The antiviral compound determination module (103) can calculate a final score for a second compound by combining a first weighted score, a second weighted score, and a third weighted score. In the same way, the antiviral compound determination module (103) can calculate a final score for a third compound to a Nth compound (N is a natural number).

[0075] The antiviral compound determination module (103) can determine which compound to use as an antiviral agent among a plurality of compounds by considering the final score for the first compound up to the final score for the Nth compound. For example, if the final score among the first to Nth compounds is lowest for the second compound, followed by the third compound, the first compound, the fourth compound, and the fifth compound in order of increasing, the antiviral compound determination module (103) can determine the second compound as the compound to use as an antiviral agent. Sometimes, a plurality of compounds with low final scores may be determined as compounds to use as antiviral agents. For example, the antiviral compound determination module (103) may determine three compounds with low final scores, namely the second compound, the third compound, and the first compound, as compounds to use as antiviral agents.

[0076] That is, the antiviral compound determination module (103) can load multiple binding affinity values ​​obtained for each of the multiple virus-derived proteins for each compound into memory as multiple first scores. Additionally, the antiviral compound determination module (103) can load multiple predetermined weights into memory. The antiviral compound determination module (103) can calculate multiple second scores by applying multiple weights to the multiple first scores loaded in memory, and calculate a final score determined as a single value for each compound from the multiple second scores. For example, the final score may be determined as a value corresponding to the average value of the multiple second scores, or as a value to which additional calculations are processed based on the average value. The antiviral compound determination module (103) can calculate the final score calculated in this manner as a weighted score.

[0077] In some embodiments, a plurality of weights may include values ​​obtained by normalizing the expression levels of a plurality of target proteins, obtained by performing RNA (Ribonucleic Acid) copy number sequencing, into a ratio to the total expression level. RNA copy number sequencing is a technique that evaluates the expression level of a specific protein by quantifying the number of copies of RNA molecules within a sample, and can produce expression level data related to specific genes or proteins on the genome by analyzing RNA with high precision. In particular, the expression level of the coronavirus can be precisely measured through RNA copy number sequencing.

[0078] For example, a plurality of weights may include relative expression levels determined by normalizing based on the copy number of each of at least two non-structural proteins selected from NSP1, NSP2, NSP3, NSP4, NSP5, NSP6, NSP7, NSP8, NSP9, NSP10, NSP12, NSP13, NSP14, NSP15, and NSP16 relative to the copy number of all virus-derived proteins of coronavirus obtained by performing RNA copy number sequencing.

[0079] For example, a plurality of weights may include the relative expression levels of NSP1, NSP2, NSP5, NSP9, NSP10, NSP12, NSP13, and NSP14, which are determined by normalizing based on the copy number of each of NSP1, NSP2, NSP5, NSP9, NSP10, NSP12, NSP13, and NSP14 relative to the copy number of all virus-derived proteins of the coronavirus obtained by performing RNA copy number sequencing.

[0080] Alternatively, for example, a plurality of weights may include the relative expression levels of NSP2, NSP5, NSP9, NSP10, and NSP13, respectively, determined by normalizing based on the copy numbers of NSP2, NSP5, NSP9, NSP10, and NSP13 relative to the copy numbers of all virus-derived proteins of the coronavirus obtained by performing RNA copy number sequencing. In this case, exemplary values ​​obtained through RNA copy number sequencing may be as follows.

[0081] Protein NSP2 NSP5 NSP9 NSP10 NSP13 Expression level before normalization 2.0 3.0 2.0 2.0 3.0 Weight after normalization 0.059 0.088 0.059 0.059 0.088

[0082] After normalizing the expression ratio of the total protein, the weighted score of each protein can be calculated by multiplying the normalized weights of the five candidates by the binding affinity of each protein.

[0083] Protein NSP2 NSP5 NSP9 NSP10 NSP13 Weight after normalization 0.059 0.088 0.059 0.059 0.088 Affinity for binding with compound 0.324 7 0.64 03 0.209 20.20.14 67 Weighted score 0.019 20.056 30.012 30.0118 0.0129 Final score 0.0889

[0084] Meanwhile, in some embodiments, a plurality of weights may include a value obtained by performing a predetermined operation on a value obtained by normalizing the expression amount of a plurality of target proteins at a first time point and a value obtained by normalizing the expression amount of a plurality of target proteins at a second time point different from the first time point.

[0085] For example, a plurality of weights may include a value normalized based on the number of copies of each of at least two non-structural proteins selected from NSP1, NSP2, NSP3, NSP4, NSP5, NSP6, NSP7, NSP8, NSP9, NSP10, NSP12, NSP13, NSP14, NSP15, and NSP16 relative to the number of copies of all virus-derived proteins in 6 hpi, and a value normalized based on the number of copies of each of at least two non-structural proteins selected from NSP1, NSP2, NSP3, NSP4, NSP5, NSP6, NSP7, NSP8, NSP9, NSP10, NSP12, NSP13, NSP14, NSP15, and NSP16 relative to the number of copies of all virus-derived proteins in 24 hpi, on which a predetermined operation (e.g., average, weighted average, etc.) is performed.

[0086] For example, multiple weights may include values ​​normalized based on the number of copies of each of NSP1, NSP2, NSP5, NSP9, NSP10, NSP12, NSP13, and NSP14 relative to the number of copies of all virus-derived proteins in 6 hpi, and values ​​normalized based on the number of copies of each of NSP1, NSP2, NSP5, NSP9, NSP10, NSP12, NSP13, and NSP14 relative to the number of copies of all virus-derived proteins in 24 hpi, on which a predetermined operation (e.g., average, weighted average, etc.) has been performed.

[0087] Alternatively, for example, a plurality of weights may include values ​​obtained by performing a predetermined operation (e.g., average, weighted average, etc.) on a value normalized based on the copy number of each of NSP2, NSP5, NSP9, NSP10, and NSP13 relative to the copy number of all virus-derived proteins in 6 hpi, and a value normalized based on the copy number of each of NSP2, NSP5, NSP9, NSP10, and NSP13 relative to the copy number of all virus-derived proteins in 24 hpi.

[0088] The manufacturing process data storage module (104) can generate manufacturing process data for the antiviral agent (which may additionally include reaction pathways, formulation requirements, mass production-related data, and quality control data in addition to chemical characteristic values, in addition to chemical characteristic values) to include chemical characteristic values ​​(chemical structure, molecular formula, molecular weight, etc.) regarding the determined compound, and can store the manufacturing process data in a storage device. Specifically, the manufacturing process data storage module (104) can generate manufacturing process data that can function as design data for the manufacture of the antiviral agent by using information about the compound determined by the antiviral agent compound determination module (103). The formulation data generated by the manufacturing process data storage module (104) can be transmitted to and utilized by another electronic device or computing device that controls the process of actually producing the antiviral agent.

[0089] In some embodiments, when there is only one type of compound determined by the antiviral compound determination module (103), the manufacturing process data may be configured to include information regarding said compound. In other embodiments, when there are several types of compounds determined by the antiviral compound determination module (103), the manufacturing process data may be configured to include information regarding all of said compounds.

[0090] In some embodiments, the antiviral compound determination module (103) has the following minimization function Min k It can operate based on.

[0091]

[0092] Here, ω i is a weight that is determined differently depending on the expression levels of multiple virus-derived proteins, and D ij can represent binding affinity for target protein-compound pairs. Min k is a minimization function, and a lower value can be understood as indicating a stronger bond. i and n can represent the item number and total number of items in the weight matrix. For example, for a total of 8 virus-derived proteins, i increases from 1 to n (8). ω1 represents a weight based on the relative expression level of the first virus-derived protein, and ω2 represents a weight based on the relative expression level of the second virus-derived protein. j and m can represent the item number and total number of items of candidate ligands. For example, if there are 30,000 candidate ligands, j can increase from 1 to m (30,000).

[0093] FIGS. 2 and FIGS. 3 are drawings for explaining the operation of an antiviral drug manufacturing apparatus according to one embodiment.

[0094] Referring to FIG. 2, the structural data acquisition module (101) can acquire sequence or structural data regarding a plurality of target proteins corresponding to at least two of the coronavirus NSP1, NSP2, NSP3, NSP4, NSP5, NSP6, NSP7, NSP8, NSP9, NSP10, NSP12, NSP13, NSP14, NSP15, and NSP16, for example, as illustrated, a plurality of target proteins corresponding to NSP1, NSP2, NSP5, NSP9, NSP10, NSP12, NSP13, and NSP14.

[0095] Next, referring to FIG. 3, the binding force prediction module (102) can predict the binding energy for the target protein-compound pair by performing docking on the target protein-compound pair and predict the binding affinity between the docked target protein and compound using a pre-trained binding affinity prediction model.

[0096] For example, referring to Table (T1), the binding affinities for virus-derived protein 1, virus-derived protein 2, virus-derived protein 3, ..., and virus-derived protein 8 for compound (CH0001) can be predicted as 'SC00001_1', 'SC00001_2', 'SC00001_3', ..., 'SC00001_8', respectively. Here, virus-derived protein 1, virus-derived protein 2, virus-derived protein 3, ..., and virus-derived protein 8 may correspond to NSP1, NSP2, NSP5, NSP9, NSP10, NSP12, NSP13, and NSP14, but this is merely illustrative.

[0097] And the average binding affinity 'SC00001_A' for 'SC00001_1', 'SC00001_2', 'SC00001_3', ..., 'SC00001_8' can be calculated. Next, for compound (CH0002), the binding affinities for virus-derived protein 1, virus-derived protein 2, virus-derived protein 3, ..., and virus-derived protein 8 can be predicted as 'SC00002_1', 'SC00002_2', 'SC00002_3', ..., 'SC00002_8', respectively. And the average binding affinity 'SC00002_A' for 'SC00002_1', 'SC00002_2', 'SC00002_3', ..., 'SC00002_8' can be calculated. By repeating this process, the binding affinities for virus-derived protein 1, virus-derived protein 2, virus-derived protein 3, ..., and virus-derived protein 8 for the compound (CH37576) are predicted as 'SC37576_1', 'SC37576_2', 'SC37576_3', ..., 'SC37576_8', respectively, and the average binding affinity 'SC37576_A' for 'SC37576_1', 'SC37576_2', 'SC37576_3', ..., 'SC37576_8' can be calculated.

[0098] The antiviral compound determination module (103) can determine which compound to use as an antiviral agent among a plurality of compounds (CH0001 to CH37576) based on the average binding affinity indicated by region (A1).

[0099] FIG. 4 is a flowchart illustrating a method for manufacturing an antiviral agent according to one embodiment.

[0100] Referring to FIG. 4, a method for manufacturing an antiviral agent according to one embodiment may include the steps of: obtaining sequence or structural data regarding a plurality of target proteins corresponding to a plurality of virus-derived proteins (S401); receiving structural data of a plurality of compounds from a pre-prepared compound structure library (S402); predicting a binding force for a target protein-compound pair combined from a plurality of target proteins and a plurality of compounds (S403); determining a compound to be used as an antiviral agent among a plurality of compounds based on the binding force (S404); generating manufacturing process data for the antiviral agent to include chemical characteristic values ​​regarding the determined compound (S405); and storing the manufacturing process data in a storage device (S406).

[0101] For more detailed information regarding the above method, reference may be made to other embodiments described in this specification; therefore, redundant details are omitted here.

[0102] FIGS. 5 and 6 are drawings for explaining the operation of an antiviral drug manufacturing apparatus according to one embodiment.

[0103] Referring to FIG. 5, when there are multiple target proteins, it is necessary to reflect not only binding affinity but also interactions with multiple target proteins. Therefore, when the number of target proteins exceeds a preset number, the antiviral compound determination module (103) can determine a compound to be used as an antiviral agent among multiple compounds based on a score calculated by applying weights to binding affinity.

[0104] Referring to the table (T1) shown in FIG. 3, the binding affinities for virus-derived protein 1, virus-derived protein 2, virus-derived protein 3, ..., and virus-derived protein 8 for the compound (CH0001) can be predicted as 'SC00001_1', 'SC00001_2', 'SC00001_3', ..., 'SC00001_8', respectively. Here, virus-derived protein 1, virus-derived protein 2, virus-derived protein 3, ..., and virus-derived protein 8 may correspond to NSP1, NSP2, NSP5, NSP9, NSP10, NSP12, NSP13, and NSP14, but this is merely illustrative.

[0105] And the average binding affinity 'SC00001_A' for 'SC00001_1', 'SC00001_2', 'SC00001_3', ..., 'SC00001_8' can be calculated. Next, for compound (CH0002), the binding affinities for virus-derived protein 1, virus-derived protein 2, virus-derived protein 3, ..., and virus-derived protein 8 can be predicted as 'SC00002_1', 'SC00002_2', 'SC00002_3', ..., 'SC00002_8', respectively. And the average binding affinity 'SC00002_A' for 'SC00002_1', 'SC00002_2', 'SC00002_3', ..., 'SC00002_8' can be calculated. By repeating this process, the binding affinities for virus-derived protein 1, virus-derived protein 2, virus-derived protein 3, ..., and virus-derived protein 8 for the compound (CH37576) are predicted as 'SC37576_1', 'SC37576_2', 'SC37576_3', ..., 'SC37576_8', respectively, and the average binding affinity 'SC37576_A' for 'SC37576_1', 'SC37576_2', 'SC37576_3', ..., 'SC37576_8' can be calculated.

[0106] Next, referring to the table (T2) shown in FIG. 5, relative expression levels 'EL01_1', 'EL01_2', 'EL01_3', 'EL01_4', and 'EL01_5' can be calculated from five samples for the virus-derived protein (VP01). Then, the average relative expression level 'EL01_A' for 'EL01_1', 'EL01_2', 'EL01_3', 'EL01_4', and 'EL01_5' can be calculated. Next, relative expression levels 'EL02_1', 'EL02_2', 'EL02_3', 'EL02_4', and 'EL02_5' can be calculated from five samples for the virus-derived protein (VP02). And the average relative expression amount 'EL02_A' for 'EL02_1', 'EL02_2', 'EL02_3', 'EL02_4', and 'EL02_5' can be calculated. By repeating this process, relative expression amounts 'EL8_1', 'EL8_2', 'EL8_3', 'EL8_4', and 'EL8_5' are obtained from five samples for the virus-derived protein (VP8), and the average relative expression amount 'EL8_A' for 'EL8_1', 'EL8_2', 'EL8_3', 'EL8_4', and 'EL8_5' can be calculated. That is, the area indicated by region (A2) can represent the average relative expression amount for each virus-derived protein. Here, virus-derived protein (VP01), virus-derived protein (VP02), virus-derived protein (VP03), ..., and virus-derived virus-derived protein (VP08) may correspond to NSP1, NSP2, NSP5, NSP9, NSP10, NSP12, NSP13, and NSP14, but this is merely illustrative.

[0107] Next, referring to the table (T3) shown in FIG. 6, the weighted scores for virus-derived protein 1, virus-derived protein 2, virus-derived protein 3, ..., and virus-derived protein 8 for compound (CH0001) can be predicted as 'SC00001_1' X 'EL01_A', 'SC00001_2' X 'EL02_A', 'SC00001_3' X 'EL03_A', ..., 'SC00001_8' X 'EL8_A', respectively. Here, virus-derived protein 1, virus-derived protein 2, virus-derived protein 3, ..., and virus-derived protein 8 may correspond to NSP1, NSP2, NSP5, NSP9, NSP10, NSP12, NSP13, and NSP14, but this is merely illustrative.

[0108] And the average weighted score 'SC00001_F' for 'SC00001_1' X 'EL01_A', 'SC00001_2' X 'EL02_A', 'SC00001_3' X 'EL03_A', ..., 'SC00001_8' X 'EL8_A' can be calculated. Next, for the compound (CH0002), the weighted scores for virus-derived protein 1, virus-derived protein 2, virus-derived protein 3, ..., and virus-derived protein 8 can be predicted as 'SC00002_1' X 'EL01_A', 'SC00002_2' X 'EL02_A', 'SC00002_3' X 'EL03_A', ..., 'SC00002_8' X 'EL8_A', respectively. And the average weighted score 'SC00002_F' for 'SC00002_1' X 'EL01_A', 'SC00002_2' X 'EL02_A', 'SC00002_3' X 'EL03_A', ..., 'SC00002_8' X 'EL8_A' can be calculated. By repeating this process, the weighted scores for virus-derived protein 1, virus-derived protein 2, virus-derived protein 3, ..., and virus-derived protein 8 for compound (CH37576) are predicted as 'SC37576_1' X 'EL01_A', 'SC37576_2' X 'EL02_A', 'SC37576_3' X 'EL03_A', ..., 'SC37576_8' X 'EL8_A', respectively, and the average weighted score 'SC37576_F' for 'SC37576_1' X 'EL01_A', 'SC37576_2' X 'EL02_A', 'SC37576_3' X 'EL03_A', ..., 'SC37576_8' X 'EL8_A' can be calculated. That is, the area (A3) indicates the average weighted score, i.e., the final score.

[0109] The antiviral compound determination module (103) can determine which compound to use as an antiviral agent among a plurality of compounds (CH0001 to CH37576) based on the final score indicated by the area (A3).

[0110] FIG. 7 is a flowchart illustrating a method for manufacturing an antiviral agent according to one embodiment.

[0111] Referring to FIG. 7, a method for manufacturing an antiviral agent according to one embodiment may include the steps of: obtaining sequence or structural data regarding a plurality of target proteins corresponding to a plurality of virus-derived proteins (S701); receiving structural data of a plurality of compounds from a pre-prepared compound structure library (S702); predicting binding strength for target protein-compound pairs combined from a plurality of target proteins and a plurality of compounds (S703); calculating a weighted score by applying a plurality of weights, which are determined differently according to the expression amount of the plurality of virus-derived proteins, to the binding strength (S704); determining a compound to be used as an antiviral agent among a plurality of compounds according to the weighted score (S705); generating manufacturing process data for the antiviral agent to include chemical characteristic values ​​regarding the determined compound (S706); and storing the manufacturing process data in a storage device (S707).

[0112] For more detailed information regarding the above method, reference may be made to other embodiments described in this specification; therefore, redundant details are omitted here.

[0113] FIG. 8 is a diagram illustrating an example of implementing a binding affinity prediction model used in the manufacture of an antiviral agent according to one embodiment.

[0114] Referring to FIG. 8, the binding affinity prediction model used by the binding strength prediction module (102) to predict binding affinity may be trained to take a 4-dimensional tensor containing the 3-dimensional coordinates of atoms and a single training feature vector as input and output a predicted value of binding affinity. Specifically, the binding affinity prediction model may include a Deep 4D Convolutional Neural Network that predicts the binding affinity between a protein and a ligand using a single output neuron. The neural network may largely include a convolutional layer section and a dense layer section. The convolutional layer section may generate a feature map that highlights the spatial occurrence of a specific pattern within the data by identifying patterns encoded by the filters of each convolutional layer. For example, the model uses three convolutional layers with 64, 128, and 256 filters, respectively, and the output of the last convolutional layer can be passed as input to a dense layer block after undergoing a flattening process. Next, the dense layer can include, for example, three layers composed of 1,000, 500, and 200 neurons, respectively. Dropout was applied to each dense layer, and the dropout probability was set to 0.5. Additionally, L2 weighted regularization with λ = 0.001 is applied to prevent overfitting. Since the ReLU activation function is used in both the convolutional layers and the dense layer, non-linearity can be introduced into the model.

[0115] In some embodiments, the 4-dimensional tensor used in the binding affinity prediction model may be constructed by combining the 3-dimensional coordinates of atoms and a training feature vector. Here, the 3-dimensional coordinates provide positional information for each atom and may be an essential element for understanding spatial interactions between proteins and ligands. The training feature vector contains 19 characteristics for each atom, which can enable the model to understand various chemical interactions between atoms. These 19 characteristics may include information such as atomic type (e.g., C, N, O, etc.), bonding state, hybridization state, partial charge, hydrophilicity / hydrophobicity, polarity, atomic radius, electronegativity, polarizability, atomic mass, ionization energy, oxidation state, covalent radius, intra-molecular position, electron affinity, hybrid bond length, rotatability, steric hindrance, and hydrogen bond donor / acceptor ability. By reflecting the different characteristics of each atom, this vector can accurately predict the strength of the binding between a protein and a ligand. Therefore, the 4-dimensional tensor can provide the complex data necessary to predict protein-ligand binding affinity by combining spatial information (3 dimensions) and chemical property information (1 dimension) of atoms.This method may be designed so that the model can predict binding affinity even in new protein-ligand combinations by learning binding data of various protein-ligand complexes, particularly through the PDBBind dataset.

[0116] FIG. 9 is a block diagram illustrating a computing device according to one embodiment.

[0117] Referring to FIG. 9, the method and apparatus for manufacturing an antiviral agent according to the embodiments may be implemented using a computing device (50). This computing device (50) may be implemented as various types of electronic devices, servers, or similar devices, and its functions may be implemented through a combination of software and hardware.

[0118] The computing device (50) may include at least one of a processor (501) communicating via a bus (509), a memory (502), a storage device (503), a display device (504), a network interface device (505) providing access to a network (40) for communication with other entities, and an input / output interface device (506) providing a user input interface or a user output interface. Of course, the computer device (50) may additionally include any electronic device necessary to implement the technical concept described in this specification, although not shown in FIG. 9.

[0119] The processor (501) can be implemented as various types of computing devices, such as an MCU (Micro Controller Unit), AP (Application Processor), CPU (Central Processing Unit), GPU (Graphic Processing Unit), NPU (Neural Processing Unit), QPU (Quantum Processing Unit), etc. The processor (501) is a semiconductor device that executes instructions stored in memory (502) or storage device (503) and can perform a core role in the system. Program code and data stored in memory (502) or storage device (503) instruct the processor (501) to perform specific tasks, thereby enabling the operation of the entire system. The processor (501) can be configured to implement the functions or methods described above in relation to FIGS. 1 through 8.

[0120] The memory (502) and storage device (503) may include various forms of volatile or non-volatile storage media for storing and accessing data of the system. For example, the memory (502) may include read-only memory (ROM) or random access memory (RAM). In some embodiments, the memory (502) may be embedded inside the processor (501), in which case the data transfer speed between the memory (502) and the processor (501) may be very fast. In some other embodiments, the memory (502) may be located outside the processor (501), in which case the memory (502) may be connected to the processor (501) through various data buses or interfaces. Such connection may be made through various known means, for example, a PCIe (Peripheral Component Interconnect Express) interface for high-speed data transfer or a memory controller. Meanwhile, examples of storage devices (503) include HDD (Hard Disk Drive) or SSD (Solid State Drive), and the scope of the present invention is not limited to the elements listed above for the purpose of explanation.

[0121] In some embodiments, at least some components or functions of the antiviral agent manufacturing method and apparatus according to the embodiments may be implemented as a program or software executed on a computing device (50), and the program or software may be stored on a computer-readable recording medium or storage medium. Specifically, a computer-readable recording medium or storage medium according to one embodiment may have a program recorded on it for executing steps included in the implementation of the antiviral agent manufacturing method and apparatus according to the embodiments on a computer including a processor (501) that executes a program or instructions stored in a memory (502) or a storage device (503).

[0122] In some embodiments, at least some of the components or functions of the antiviral agent manufacturing method and apparatus according to the embodiments may be implemented using hardware or circuits of the computing device (50), or may be implemented using separate hardware or circuits that can be electrically connected to the computing device (50).

[0123] According to the embodiments, antiviral agents prepared using a multi-target approach are not dependent on a single specific protein, and thus can significantly reduce the potential for viral resistance development and provide a more potent antiviral effect. Furthermore, because they are developed to exert optimal action based on the relative importance and expression levels of target proteins, they can deliver a broader and more sustained effect than conventional single-target antiviral agents. Additionally, through their multi-target characteristics, they can provide a significant advancement in inhibiting the entire viral replication mechanism while simultaneously suppressing the development of resistance.

[0124] Although embodiments of the present invention have been described in detail above, the scope of the present invention is not limited thereto, and various modifications and improvements by those skilled in the art to which the present invention belongs, using the basic concept of the present invention as defined in the following claims, also fall within the scope of the present invention.

Claims

1. A method for manufacturing an antiviral agent, performed by a computing device including a processor, memory, and storage device, wherein The above processor acquires sequence or structural data regarding a plurality of target proteins corresponding to a plurality of virus-derived proteins; The above processor receives structural data of a plurality of compounds from a pre-prepared compound structure library; The above processor predicts the binding affinity for a target protein-compound pair combined from the plurality of target proteins and the plurality of compounds; The above processor calculates a weighted score by applying a plurality of weights, which are determined differently according to the expression levels of the plurality of virus-derived proteins, to the binding force; The above processor determines a compound to be used as an antiviral agent among the plurality of compounds according to the weighted application score; and The processor comprises the step of generating manufacturing process data for the antiviral agent to include chemical characteristic values ​​regarding the determined compound, and storing the manufacturing process data in the storage device. The above plurality of weights include values ​​obtained by performing RNA (Ribonucleic Acid) copy number sequencing, wherein the expression levels of each of the plurality of target proteins are normalized as a ratio to the total expression levels. Method for manufacturing antiviral agents.

2. In Paragraph 1, A method for manufacturing an antiviral agent, wherein the plurality of weights include values ​​obtained by performing a predetermined operation including an average or a weighted average on a value obtained by normalizing the expression amount of each of the plurality of target proteins at a first time point and a value obtained by normalizing the expression amount of each of the plurality of target proteins at a second time point different from the first time point.

3. In Paragraph 1, The step of obtaining the above sequence or structural data is, A method for manufacturing an antiviral agent comprising the step of obtaining sequence or structural data regarding a plurality of target proteins corresponding to a non-structural protein among the plurality of virus-derived proteins.

4. In Paragraph 3, The above-mentioned non-structural protein is, A method for manufacturing an antiviral agent comprising at least 2 of NSP1, NSP2, NSP3, NSP4, NSP5, NSP6, NSP7, NSP8, NSP9, NSP10, NSP12, NSP13, NSP14, NSP15, and NSP16 of coronavirus.

5. In Paragraph 1, The step of obtaining the above sequence or structural data is, A method for manufacturing an antiviral agent comprising the step of obtaining sequence or structural data regarding a plurality of target proteins corresponding to a structural protein among the plurality of virus-derived proteins.

6. In Paragraph 1, The step of predicting the bonding force above is, The above processor performs docking on the target protein-compound pair to predict the binding pose for the target protein-compound pair; The processor loads a pre-trained coupling affinity prediction model into the memory; and A method for manufacturing an antiviral agent, comprising the step of the processor inputting the docked target protein-compound pair into the binding affinity prediction model and obtaining the binding affinity value output from the binding affinity prediction model as the binding strength.

7. In Paragraph 6, A method for manufacturing an antiviral drug, wherein the above-described binding affinity prediction model is trained to take a 4-dimensional tensor containing the 3-dimensional coordinates of an atom and one training feature vector as input and output a predicted value of the binding affinity.

8. In Paragraph 6, A method for manufacturing an antiviral agent, wherein the above-mentioned binding affinity prediction model is trained using experimental values ​​regarding the binding strength of a protein and a compound as training data.

9. In Paragraph 1, The step of calculating the above-mentioned weighted score is, The above processor loads the values ​​of the plurality of binding forces obtained for each of the plurality of target proteins for each of the above compounds into the memory as a plurality of first scores; The above processor loads the plurality of weights into the memory; The above processor calculates a plurality of second scores by applying a plurality of weights to a plurality of first scores loaded in the memory; The above processor calculates a final score determined as a single value for each compound from the plurality of second scores; and A method for manufacturing an antiviral agent, comprising the step of the processor calculating the final score as the weighted score.

10. As an antiviral agent manufacturing device, processor; and It includes a memory in which instructions executed by the above processor are loaded, and The above instruction is executed by the processor, causing the processor, Sequence or structural data regarding multiple target proteins corresponding to multiple virus-derived proteins are obtained, and Structural data of multiple compounds is provided from a pre-prepared compound structure library, and Predicting binding strength for target protein-compound pairs combined from the plurality of target proteins and the plurality of compounds, and A weighted score is calculated by applying a plurality of weights, which are determined differently according to the expression levels of the plurality of virus-derived proteins, to the above binding force, and Based on the above weighted score, a compound to be used as an antiviral agent among the plurality of compounds is determined, and Generate manufacturing process data for the antiviral agent to include chemical characteristic values ​​regarding the compound determined above, and The above manufacturing process data is stored in a storage device, and The above plurality of weights include values ​​obtained by performing RNA (Ribonucleic Acid) copy number sequencing, wherein the expression levels of each of the plurality of target proteins are normalized as a ratio to the total expression levels. Antiviral drug manufacturing device.

11. In Paragraph 10, An antiviral drug manufacturing apparatus comprising a plurality of weights, wherein the plurality of weights include a value obtained by performing a predetermined operation including an average or a weighted average on a value obtained by normalizing the expression amount of each of the plurality of target proteins at a first time point and a value obtained by normalizing the expression amount of each of the plurality of target proteins at a second time point different from the first time point.

12. In Paragraph 10, Acquiring the above sequence or structural data is, An antiviral agent manufacturing apparatus comprising obtaining sequence or structural data regarding a plurality of target proteins corresponding to a non-structural protein among the plurality of virus-derived proteins.

13. In Paragraph 12, The above-mentioned non-structural protein is, An antiviral agent manufacturing device comprising at least 2 of NSP1, NSP2, NSP3, NSP4, NSP5, NSP6, NSP7, NSP8, NSP9, NSP10, NSP12, NSP13, NSP14, NSP15, and NSP16 of coronavirus.

14. In Paragraph 10, Acquiring the above sequence or structural data is, An antiviral agent manufacturing apparatus comprising obtaining sequence or structural data regarding a plurality of target proteins corresponding to a structural protein among the plurality of virus-derived proteins.

15. In Paragraph 10, Predicting the above bonding force is, Docking is performed on the above target protein-compound pair to predict the binding pose for the above target protein-compound pair, and Load a pre-trained coupling affinity prediction model into the memory, and An antiviral drug manufacturing apparatus comprising inputting the docked target protein-compound pair into the binding affinity prediction model and obtaining the binding affinity value output from the binding affinity prediction model as the binding force.

16. In Paragraph 15, An antiviral drug manufacturing apparatus, wherein the above-described binding affinity prediction model is trained to take a 4-dimensional tensor containing the 3-dimensional coordinates of atoms and one training feature vector as input and output a predicted value of the binding affinity.

17. In Paragraph 10, Calculating the above weighted score is, For each of the above compounds, the values ​​of the plurality of binding forces obtained for each of the plurality of target proteins are loaded into the memory as a plurality of first scores, and Loading the above plurality of weights into the memory, Calculate a plurality of second scores by applying the plurality of weights to the plurality of first scores loaded in the memory, and Calculate a final score determined as a single value for each compound from the plurality of second scores above, and An antiviral drug manufacturing apparatus comprising calculating the above final score as the above weighted score.

18. A computer-readable recording medium that records instructions executed by a computing device including a processor, memory and storage device, The above instruction is executed by the computing device, causing the computing device, Sequence or structural data regarding multiple target proteins corresponding to multiple virus-derived proteins are obtained, and Structural data of multiple compounds is provided from a pre-prepared compound structure library, and Predicting binding strength for target protein-compound pairs combined from the plurality of target proteins and the plurality of compounds, and A weighted score is calculated by applying a plurality of weights, which are determined differently according to the expression levels of the plurality of virus-derived proteins, to the above binding force, and Based on the above weighted score, a compound to be used as an antiviral agent among the plurality of compounds is determined, and Generate manufacturing process data for the antiviral agent to include chemical characteristic values ​​regarding the compound determined above, and The above manufacturing process data is stored in the storage device, and The above plurality of weights include values ​​obtained by performing RNA (Ribonucleic Acid) copy number sequencing, wherein the expression levels of each of the plurality of target proteins are normalized as a ratio to the total expression levels. Computer-readable recording medium.