Estimation device, estimation system, estimation method, and estimation program

The estimation device and system efficiently estimate nucleic acid sequences with high affinity for target proteins using an objective function and quantum annealing, addressing the challenge of vast sequence space and enhancing drug discovery.

JP2026079589APending Publication Date: 2026-05-15NATIONAL INSTITUTE OF ADVANCED INDUSTRIAL SCIENCE & TECHNOLOGY +2
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
NATIONAL INSTITUTE OF ADVANCED INDUSTRIAL SCIENCE & TECHNOLOGY
Filing Date
2024-10-30
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

The number of candidate sequences of nucleic acids with high affinity for a target protein is enormous, making it difficult to experimentally obtain an optimal solution for nucleic acid sequences that exhibit high affinity.

Method used

An estimation device and system utilizing a storage unit, estimation unit, and output unit, along with an objective function, to efficiently estimate nucleic acid sequences that bind strongly to target proteins, leveraging quantum annealing machines for optimization.

Benefits of technology

Enables efficient estimation of nucleic acid sequences with high affinity for target proteins, shortening the drug discovery process and improving the aptamer development efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026079589000001_ABST
    Figure 2026079589000001_ABST
Patent Text Reader

Abstract

The present invention aims to provide a sequencing device that enables the efficient estimation of nucleic acid sequences exhibiting high affinity for target proteins. [Solution] The estimation device according to the present invention is characterized by comprising: a storage unit that stores a set of nucleic acid sequence information and base sequence information of nucleic acids whose affinity to a target protein has been determined by the SELEX method; an estimation unit that estimates the affinity of nucleic acids consisting of base sequence information to a target protein using an objective function set to output a predetermined value for a partial base sequence common to the nucleic acids included in the set of nucleic acid sequence information; and an output unit that can output affinity information indicating the affinity to the target protein.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an estimation device, an estimation system, an estimation method, and an estimation program.

Background Art

[0002] Regarding functional nucleic acids such as aptamers, methods of experimentally collecting nucleic acid sequence information by methods such as the SELEX (Systematic Evolution of Ligands by Exponential Enrichment) method and searching for sequences with high affinity for a target protein have been studied. For example, Patent Document 1 describes selecting an aptamer for the catalytic site of the MMP-9 protein using the SELEX method.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The number of candidate sequences of nucleic acids with high affinity for a target protein is theoretically enormous, and it is difficult to obtain an optimal solution experimentally. Therefore, it is desired to efficiently estimate the sequence of a nucleic acid that exhibits high affinity for a target protein.

[0005] An object of the present invention is to provide an estimation device, an estimation system, an estimation method, and an estimation program that enable efficient estimation of the sequence of a nucleic acid that exhibits high affinity for a target protein.

Means for Solving the Problems

[0006] The present invention provides the following inventions. [1] The estimation device according to the present invention is characterized by comprising: a storage unit that stores a set of nucleic acid sequence information and base sequence information of nucleic acids whose affinity to a target protein has been determined by the SELEX method; an estimation unit that estimates the affinity of nucleic acids consisting of base sequence information to a target protein using an objective function set to output a predetermined value for a partial base sequence common to the nucleic acids included in the set of nucleic acid sequence information; and an output unit that can output affinity information indicating the affinity to a target protein. [2] Furthermore, in the estimation device according to the present invention, it is preferable that the storage unit stores a plurality of base sequence information, the estimation unit extracts base sequence information with high affinity to the target protein based on affinity information, and the output unit outputs the extracted base sequence information. [3] In the estimation device according to the present invention, it is preferable that the objective function is set to output a predetermined value for a partial base sequence that can commonly bind to a predetermined position in the base sequence included in the nucleic acid sequence information group. [4] Furthermore, in the estimation device according to the present invention, it is preferable that the smaller the value of the objective function corresponding to the solution of the objective function, the higher the degree of binding to the base sequence included in the nucleic acid sequence information group. [5] Furthermore, in the estimation device according to the present invention, the objective function is preferably an Ising model type Hamiltonian, and the procedure for finding the solution of the objective function is preferably performed by a quantum annealing machine. [6] Furthermore, the objective function (c(s)) is defined as follows: T is the number of elements in the set of different base sequences, wk is the weight, L is the nucleic acid sequence and the length of the base sequence, l is the length of the subsequence, δ is a variable that is 1 if the two sequences are identical and 0 if they are not identical, s is the base sequence, and s k As a nucleic acid sequence, [Mathematics 1] JPEG2026079589000002.jpg1639 Here, [Math 2] It is preferable that the file name be JPEG2026079589000003.jpg2194. [7] Furthermore, the objective function (c(s)) is defined as follows: T is the number of elements in the set of different base sequences, wk is the weight, L is the nucleic acid sequence and the length of the base sequence, l is the length of the subsequence, P is a positive integer, δ is a variable that is 1 if the two sequences are identical and 0 if they are not identical, s is the base sequence, and s k As a nucleic acid sequence, [Mathematics 3] JPEG2026079589000004.jpg1739 Here, [Mathematics 4] It is preferable that the file name be JPEG2026079589000005.jpg22115. [8] Furthermore, the objective function (c(s)) is defined as follows: T is the number of elements in the set of distinct base sequences, wk is the weight, L is the length of the nucleic acid sequence and base sequence, l is the length of the subsequence, P is a positive integer, x is a binary variable, and s k As a nucleic acid sequence, [Math 5] JPEG2026079589000006.jpg2255 Here, [Mathematics 6] It is preferable that the file name be JPEG2026079589000007.jpg24111. [9] The estimation system according to the present invention is an estimation system comprising a terminal device and an annealing machine, wherein the terminal device has a terminal storage unit that stores a set of nucleic acid sequence information and base sequence information of nucleic acids whose affinity to a target protein has been recognized by the SELEX method, a terminal communication unit that transmits the set of nucleic acid sequence information and base sequence information and receives affinity information indicating affinity to a target protein, and an output unit that can output affinity information, and the annealing machine has an estimation unit that estimates the affinity of nucleic acids consisting of base sequence information to a target protein using an objective function that is set to output a predetermined value for partial base sequences common to nucleic acids included in the set of nucleic acid sequence information.

[10] The estimation method according to the present invention is characterized by storing a set of nucleic acid sequence information and base sequence information of nucleic acids whose affinity to a target protein has been determined by the SELEX method in a storage unit, and using an objective function set to output a predetermined value for a partial base sequence common to the nucleic acids included in the set of nucleic acid sequence information with respect to the base sequence information, estimating the affinity of the nucleic acid consisting of the base sequence information to a target protein, and outputting affinity information indicating the affinity to the target protein.

[11] The estimation program according to the present invention is characterized in that it stores a set of nucleic acid sequence information and base sequence information of nucleic acids whose affinity to a target protein has been determined by the SELEX method in an estimation device, and uses an objective function set to output a predetermined value for a partial base sequence common to the nucleic acids included in the set of nucleic acid sequence information to estimate the affinity of the nucleic acid consisting of the base sequence information to a target protein, and outputs affinity information indicating the affinity to the target protein. [Effects of the Invention]

[0007] According to the present invention, the estimation device, estimation system, estimation method, and estimation program enable the efficient estimation of nucleic acid sequences that exhibit high affinity for target proteins. [Brief explanation of the drawing]

[0008] [Figure 1] This is a functional block diagram of the estimation system according to Embodiment 1. [Figure 2] This flowchart illustrates the process for selecting aptamer drug candidates. [Figure 3] This is a flowchart showing the flow of the estimation process performed by the estimation device. [Figure 4] This figure shows the concept of the subarray combination comparison process shown in Figure 2. [Figure 5] This figure shows an example of the estimation results. [Figure 6] This is a functional block diagram of the estimation system according to Embodiment 2. [Figure 7]This figure shows the concept of the subarray combination comparison process, as shown in Equation 4. [Figure 8] This figure shows the estimation results in Example 1. [Figure 9] This figure shows the activity measurement results of the base sequence obtained in Example 1. [Modes for carrying out the invention]

[0009] The estimation apparatus, estimation system, estimation method, and estimation program according to the present invention will be described below with reference to the drawings. However, it should be noted that the technical scope of the present invention is not limited to these embodiments, but extends to the invention described in the claims and its equivalents.

[0010] (Background of the present invention) In current drug development, the lengthening and increasing costs of the drug discovery process, along with the resulting decline in the research and development capabilities of pharmaceutical companies, are serious problems. One of the reasons for this decline in research and development capabilities is thought to be the difficulty in drug discovery using small-molecule drugs, which have been the mainstream of conventional medicines. Therefore, biopharmaceuticals (antibody drugs) and medium-molecule drugs are mainly being researched as next-generation drugs to replace small-molecule drugs. Medium-molecule drugs are thought to have properties intermediate between small-molecule drugs and antibody drugs. Here, RNA (Ribonucleic Acid) aptamers and DNA (Deoxyribonucleic Acid) aptamers are examples of medium-molecule drugs, and these are collectively called nucleic acid aptamers. Among nucleic acid aptamers, RNA aptamers are generally single-strand nucleic acids with a length of 20 to 50 bases, which form a unique three-dimensional structure based on their base sequence, and bind to the target substance (for example, a disease-related protein, which is one example of a target protein) to exert a pharmacological effect. RNA aptamers are attracting attention as next-generation drugs due to their many advantages over other pharmaceuticals, including high affinity and specificity, applicability to various targets, easy chemical synthesis, and low immunogenicity. For a long time, the only RNA aptamer on the market was Macugen®, a treatment for age-related macular degeneration, but recently a second aptamer was approved by the U.S. Food and Drug Administration (FDA). However, the commercialization of RNA aptamers lags behind that of other pharmaceuticals.

[0011] The development process for nucleic acid aptamers is divided into several stages. The first stage of the development process, basic / exploratory research on aptamers, is a crucial stage for narrowing down aptamer sequences and is further divided into several steps. In this first step, candidate aptamer sequences are obtained using the SELEX method, also known as in vitro artificial evolution. The SELEX method concentrates highly binding aptamers by repeatedly binding and amplifying nucleic acids that strongly bind to the target protein from a random sequence pool. During this process, a sequence analysis device, also known as a high-speed sequencer, can be used to comprehensively measure the sequence information of the sequence pool concentrated in each round. This experimental method is also called HT (High-Throughput)-SELEX. As a result of HT-SELEX, a large amount of sequence information (a collection of nucleic acid sequence information) is obtained in each round. In current aptamer development, it is important to design superior aptamer sequences by utilizing HT-SELEX data.

[0012] There are two main approaches to aptamer development using the SELEX method. The first is a method that ranks sequences obtained experimentally based on some index (score function). While this approach is the most commonly used, it is extremely difficult to design new sequences that are not included in the experimental data. In contrast, a method has been proposed that can probabilistically generate new aptamer sequences using generative models, such as those represented by AI (Artificial Intelligence). However, even with such a method, it is extremely difficult to select the optimal sequence by exhaustively searching the sequence space. For example, if the RNA sequence length is 40, the number of candidate sequences is 10. 24 This is because there are as many as 100,000 elements, and it is extremely difficult for a regular computer to search through all of them. In recent years, annealing machines have been considered as a new method for performing such searches.

[0013] One annealing machine method uses the Ising model, which is a model that describes the properties of magnetic materials and is used to minimize the energy of spin interactions. The Ising model is considered effective in solving combinatorial optimization problems from the perspective of minimizing energy. That is, since many combinatorial optimization problems (e.g., scheduling, routing, clustering, etc.) can be expressed in a form that minimizes an evaluation function such as cost or risk, the same approach as energy minimization in the Ising model can be adopted.

[0014] Sequences that readily bind to target proteins often share subsequences with those contained in the nucleic acid sequence information obtained by the SELEX method. Therefore, selecting nucleic acid sequences with high affinity for target proteins by exhaustively searching the sequence space involves searching for such combinations of common subsequences from the set of nucleic acid sequence information, and can be viewed as a combinatorial optimization problem.

[0015] As described below, an objective function is introduced when solving combinatorial optimization problems. The objective function is a function that indicates the target to be minimized or maximized in a combinatorial optimization problem. In this invention, the sequence of nucleic acids that efficiently exhibit high affinity for a target protein is estimated using the objective function. Preferably, this invention proposes an objective function that can be optimized by quantum annealing (Quadratic Unconstrained Binary Optimization (QUBO)).

[0016] (Embodiment 1) Figure 1 is a functional block diagram of the estimation system 1 according to Embodiment 1. The estimation system 1 includes a sequence analyzer 10 and a terminal device 20, etc. The sequence analyzer 10 and the terminal device 20 are connected to each other so as to be able to communicate via a network N, which is the Internet or an intranet.

[0017] The sequence analyzer 10 is a device for reading the base sequence of DNA or RNA contained in a sample, and is capable of analyzing the order of bases (nucleotides) contained in the base sequence. In one example, the sequence analyzer 10 sequentially binds fluorescently labeled bases to fragmented DNA or RNA, and analyzes which bases have been bound by capturing the fluorescence signal at the time of binding with an optical sensor in the sequence analyzer 10. RNA may also be reverse transcribed into DNA before its base sequence is read. The sequence analyzer 10 outputs the nucleic acid sequence information set, which is the analysis result, to the terminal device 20 via the network N.

[0018] The sequence analysis device 10 described above may be a conventional Sanger-based sequencer, or a next-generation sequencer may be used.

[0019] Terminal device 20 is an example of an estimation device. In one example, terminal device 20 is a personal computer. Terminal device 20 may also be a smartphone or other mobile phone or a tablet-type information terminal. Terminal device 20 includes a terminal communication unit 21, a terminal storage unit 22, a terminal display unit 23, and a terminal processing unit 24, etc. These components are connected via bus B.

[0020] The terminal communication unit 21 is configured to enable the terminal device 20 to communicate with other devices and has a communication interface circuit conforming to wired / wireless communication standards. The communication interface circuit is, for example, a mobile communication system, wired LAN, wireless LAN, LTE (Long Term Evolution), etc. The terminal communication unit 21 supplies signals received via the network N to the terminal processing unit 24 and transmits signals supplied from the terminal processing unit 24 to other devices.

[0021] The terminal storage unit 22 is an example of a storage unit. The terminal storage unit 22 is configured for storing programs and data, and in one example, it includes semiconductor memory. The terminal storage unit 22 stores operating system programs, driver programs, and application programs used for processing by the terminal processing unit 24 as programs. Programs are installed into the terminal storage unit 22 from a computer-readable, non-temporary, portable storage medium such as a CD-ROM or DVD-ROM. Programs may also be stored on a storage medium owned by an external server and installed via the network N.

[0022] The terminal storage unit 22 stores sequence information and base sequence information of nucleic acids whose affinity for the target protein has been determined by the SELEX method. The terminal storage unit 22 may store multiple base sequence information. The nucleic acid sequence information and base sequence information are information about the order of the bases that make up RNA, and the bases include adenine (A), thymine (T), guanine (G), and cytosine (C).

[0023] The terminal display unit 23 is configured for displaying images and includes a display such as a liquid crystal display or an organic EL display, and an output interface circuit that outputs image data to the display. The terminal display unit 23 may also function as an operation interface as a touch panel. The terminal display unit 23 displays images based on signals supplied from the terminal processing unit 24.

[0024] The terminal processing unit 24 is a device that comprehensively controls the operation of the terminal device 20 and comprises one or more processors and their peripheral circuits. In one example, the terminal processing unit 24 comprises a CPU. The terminal processing unit 24 may also comprise a GPU, DSP, LSI, ASIC, FPGA, etc. The terminal processing unit 24 controls the operation of each component and executes various processes so that the various processes of the terminal device 20 are executed in the appropriate procedure based on the program stored in the terminal storage unit 22.

[0025] The terminal processing unit 24 includes an acquisition unit 241, an estimation unit 242, and an output unit 243 as functional blocks. Each of these units is a functional module implemented based on a program executed by the terminal processing unit 24. Each of these units may be implemented in the terminal device 20 as firmware.

[0026] [Aptamer drug candidate selection process] Figure 2 is a flowchart illustrating the process for selecting aptamer drug candidates.

[0027] First, the SELEX method is used to select a set of nucleic acid sequence information to be used in the estimation process described later (step S101). As mentioned above, the SELEX method is a technique for enriching highly binding aptamers by repeatedly binding and amplifying nucleic acids (RNA or DNA) that strongly bind to the target protein from a random sequence pool. In the SELEX method, nucleic acids with random sequences of appropriate length containing sequences for PCR (Polymerase Chain Reaction) primers at both ends are synthesized and associated with the target protein. After washing away the nucleic acids that did not bind, the bound nucleic acids are recovered. In the case of RNA, it is converted to DNA by reverse transcription PCR (RT (Reverse Transcription) PCR). The obtained DNA can be amplified by PCR and used in the next round. In the case of RNA, the amplified DNA is used as a template for the next RNA to be used. By repeating this for several rounds, nucleic acid aptamers that show affinity for the target protein are enriched, and the sequence information set of nucleic acids is obtained for the enriched nucleic acid aptamers by the sequence analyzer 10. The nucleic acid sequence information is output from the sequence analyzer 10 to another device.

[0028] Next, the terminal device 20 performs estimation processing (step S102). The estimation processing will be described later.

[0029] Next, using the base sequence information with high affinity for the target protein output by the terminal device 20, various analyses, processing, tests, and evaluations are performed (step S103). Examples of analyses include measurement of binding properties, processing includes shortening, further SELEX method implementation, and chemical modification, testing includes cell and animal tests, and evaluation includes toxicity tests. This completes the aptamer drug candidate selection process.

[0030] [Estimation Processing] Figure 3 is a flowchart showing an example of estimation processing performed by the terminal device 20 according to Embodiment 1. The processing shown in Figure 3 is realized by the terminal processing unit 24 cooperating with each component of the terminal device 20 based on a program stored in the terminal storage unit 22.

[0031] First, the acquisition unit 241 acquires a set of sequence information of nucleic acids whose affinity for the target protein has been determined by the SELEX method via the network N (step S201).

[0032] The acquisition unit 241 may acquire nucleic acid sequence information without going through the network N. For example, if the nucleic acid sequence information output from the sequence analyzer 10 is stored in a portable memory such as a USB drive, the acquisition unit 241 may acquire the nucleic acid sequence information from the portable memory connected to the terminal device 20 via an interface circuit. The same applies to Embodiment 2, which will be described later.

[0033] Next, the terminal processing unit 24 stores the nucleic acid sequence information and base sequence information in the terminal storage unit 22 (step S202). The terminal processing unit 24 may also store multiple base sequence information in the terminal storage unit 22.

[0034] Next, the estimation unit 242 estimates the affinity of the nucleic acid, consisting of the base sequence information, to the target protein using an objective function that is set to output a predetermined value for partial base sequences common to the nucleic acids included in the set of nucleic acid sequence information (step S203). Details of the process in step S203 will be described later.

[0035] Next, the output unit 243 outputs affinity information indicating the affinity for the target protein (step S204). In other words, the output unit 243 outputs the estimation result from the estimation unit 242. In one example, the output unit 243 displays the estimation result as an image on the terminal display unit 23. The estimation result will be described later. This completes the estimation process.

[0036] [Details of the process in step S203] As mentioned above, selecting the optimal nucleic acid sequence with high affinity for the target protein can be replaced with a combinatorial optimization problem. A combinatorial optimization problem is a problem in which the optimal combination is selected from a combination of multiple options or elements based on predetermined evaluation criteria. Preferably, the combinatorial optimization problem seeks a combination in which each option satisfies the constraints and minimizes or maximizes the overall evaluation index.

[0037] Combinatorial optimization problems are formulated as mathematical models using binary variables. These mathematical models are represented by binary variables, each taking a value of either 0 or 1, and each combination of these variables corresponds to a candidate solution. To quantitatively evaluate the relationships and interactions between the binary variables, an objective function is introduced. The objective function is a value calculated based on the combination of binary variables and defines the evaluation criteria for the problem. The objective function numerically indicates the effectiveness of the candidate solutions, and the solution with the minimum or maximum value of the objective function is the optimal solution.

[0038] In the above-mentioned SELEX method, in each round, sequence information of bases that may bind to the target protein, or in other words, a group of sequence information of nucleic acids that have been found to have an affinity for the target protein, can be obtained. The group of sequence information of nucleic acids obtained by the SELEX method may become a pharmaceutical candidate. However, as described above, it is difficult to design new sequences that are not included in experimental data only by the SELEX method. Here, instead of focusing on the nucleic acid base sequence information itself, by focusing on partial base sequences included in the nucleic acid base sequence information, it becomes possible to select nucleic acid sequences that cannot be obtained by the SELEX method. Hereinafter, for the base sequence information, an objective function that outputs a predetermined value for the partial base sequence common to the nucleic acids included in the nucleic acid sequence information group is set.

[0039] Regarding the objective function, each notation is defined as follows. ·Length of the sequence: L ·Number of rounds: r (r = 0 round is the initial pool) ·Number of elements in the set of different base sequences in the r-th round: T r ·Set of different base sequences in the r-th round: S r :={s k r |k = 1, …, T (r)} ·n k r :Frequency of the sequence s k in the r-th round ·Length of the partial sequence: l 0] ·Partial sequence of length l: l-mer Also, for any sequence s, let s〔i, j〕be the partial sequence from the i-th character to the j-th character of the sequence s (i and j start from 1). Here, s〔i〕:= s〔i, i〕. Further, for two sequences (strings) s and t, a variable δ s,t is introduced, which is 1 when s and t match as strings, that is, when the sequences match, and 0 when they do not match. Note that hereinafter, when only a specific round is focused on, the subscript r may be omitted.

[0040] For any array s of length L, the objective function c(s) is shown in equation 1.

[0041]

number

[0042] The objective function c(s) is the sequence s obtained from the r-round SELEX experimental data (a set of nucleic acid sequence information). k r Score function c k (s) to n k r Weight w defined using k This is a weighted sum with weights taken from it.

[0043] Next, we need the score function c when calculating the right-hand side. k (s) and weight w k We derive the following. First, we define w as a weight based on frequency information. k =n k Let's assume that in this case, the weight w k This is the sequence s in round r. k This is the frequency. k = log n k This is also acceptable. In this case, the weight w k This represents the logarithmic frequency in round r.

[0044] Score function c based on information from subarrays included in experimental data k As (s), we introduce Math 2.

[0045]

number

[0046] Score function c k (s) is s and s k This shows the number of l-mer subarrays that are common to all of them. Score function c k(s) is set so that the energy of sequence s containing many L-mer subsequences in the experimental data is low. For example, nucleic acid sequence s k When =CTATG and the length of the subsequence l=3, the nucleic acid sequence s k The energy is allocated to the 3-mer contained in as follows: • CTA:-1 TAT:-1 ATG: -1 In this case, c k (CTATG) = -3, c k (TATAT) = -2.

[0047] Figure 4 shows the concept of the subsequence combination comparison process shown in Equation 2. Figure 4 shows the subsequence of base sequence s and nucleic acid sequence s when the length of the subsequence l=3. k This shows the comparison process with a subarray. Also, in Figure 4, the left side shows the case where i=2 and the right side shows the case where i=3, separated by an arrow pointing from left to right.

[0048] In the example shown in Figure 4, the estimation unit 242 determines the nucleic acid sequence s for the partial sequence CAC of the base sequence s. k The partial sequences AGC, ATT, etc. are compared in order. If the partial sequence CAC matches the partial sequence s of the nucleic acid, then the partial sequence s k If present, the variable δ becomes 1. The estimation unit 242 determines the nucleic acid sequence s k For this, we compare the subarrays up to j = L - l + 1 (if the length of the array L = 26, then j = 24).

[0049] Next, the estimation unit 242 determines the nucleic acid sequence s for the partial sequence ACG of the base sequence s. k The partial sequences AGC, ATT, etc. are compared in order. The partial sequence ACG is the sequence obtained by shifting the position by one from the partial sequence CAC of base sequence s in Figure 4, in other words, the sequence when i=2 changes to i=3. In this way, the estimation unit 242 compares the partial sequence of base sequence s with the nucleic acid s k The subarray of the array is compared sequentially with j=L-l+1 and i=L-l+1.

[0050] In the example shown in Figure 4, the process of shifting a sub-sequence of the nucleic acid sequence sk and comparing it with a sub-sequence of the base sequence s is shown first, but the order of processing is not limited to this. That is, the estimation unit 242 may first shift the position of the sub-sequence of the base sequence s one position at a time and compare it with the sub-sequence of the nucleic acid sequence sk.

[0051] Score function c k (s) is a subsequence of the base sequence s and the nucleic acid sequence s k The more subarrays that match, the smaller the value becomes. That is, the score function c k (s) is set such that the smaller the score function value corresponding to the solution of the score function, the higher the degree of binding to the base sequence included in the nucleic acid sequence information set. A subsequence of base sequence s and nucleic acid sequence s k A large number of matching subsequences, or in other words, a small score function value corresponding to the solution of the score function, indicates a high degree of binding to the base sequences contained in the nucleic acid sequence information set.

[0052] The objective function c(s) is the score c k Since (s) is the sum of the weights, the objective function c(s) is also set such that the smaller the value of the objective function corresponding to the solution of the objective function, the higher the degree of binding to the base sequence included in the nucleic acid sequence information group.

[0053] [Example of estimation result] Figure 5 shows an example of the estimation results. In the example shown in Figure 5, the top 10 base sequences with high affinity to the target protein are extracted. That is, the estimation unit 242 extracts base sequence information with high affinity to the target protein based on affinity information, and the output unit 243 outputs the extracted base sequence information. Affinity information indicates the affinity to the target protein.

[0054] In the example shown in Figure 5, for each base sequence, an example of affinity information, consisting of scores X1 to X10 and a score rank, is shown. Scores X1 to X10 are negative values, and in the example shown in Figure 5, X1 is the smallest value, followed by X2, X3, ..., X10 in increasing order.

[0055] In the example shown in Figure 5, the top 10 base sequences with high affinity to the target protein were extracted as estimation results. However, the output unit 243 only needs to output the affinity information estimated by the estimation unit 242, and the estimation unit 242 does not need to extract base sequence information with high affinity to the target protein.

[0056] As described above, the terminal device (estimation device) 20 of Embodiment 1 includes an estimation unit 242 that estimates the affinity of nucleic acids consisting of base sequence information to a target protein using an objective function c(s) which is set to output a predetermined value for a partial base sequence common to nucleic acids included in the nucleic acid sequence information group. This enables the terminal device 20 to efficiently estimate nucleic acid sequences that show high affinity to target proteins. Efficiently estimating nucleic acid sequences that show high affinity to target proteins shortens the estimation process in the aptamer drug candidate selection process and contributes to shortening the drug discovery process.

[0057] Preferably, the estimation unit 242 extracts base sequence information with high affinity to the target protein based on affinity information, and the output unit 243 outputs the extracted base sequence information. This allows the terminal device 20 to output only base sequence information with particularly high affinity to the target protein.

[0058] Preferably, the smaller the value of the objective function c(s) corresponding to the solution of the objective function c(s), the higher the degree of binding to the base sequence included in the nucleic acid sequence information set. This enables the terminal device 20 to more accurately estimate nucleic acid sequences that show high affinity to the target protein.

[0059] (Embodiment 2) Figure 6 is a functional block diagram of the estimation system 2 according to Embodiment 2. The estimation system 2 differs from Embodiment 1 in that it further includes an annealing machine 30, among other things. Components similar to those in Embodiment 1 described above are denoted by the same reference numerals, and their descriptions are omitted as appropriate. Similarly, for each of the subsequent modifications, components similar to those in Embodiment 1 to each of the modifications described above are denoted by the same reference numerals, and their descriptions are omitted as appropriate.

[0060] The estimation system 2 includes a sequence analyzer 10, a terminal device 20, and an annealing machine 30, etc. The sequence analyzer 10, the terminal device 20, and the annealing machine 30 are connected to each other via a network N, which is the Internet or an intranet, enabling them to communicate with one another.

[0061] The annealing machine 30 is a dedicated device for finding the ground state of an Ising model Hamiltonian (a function representing energy states). The annealing machine 30 probabilistically finds the optimal variable values ​​by minimizing an objective function that takes a spin variable (+1 or -1) or a binary variable (0 or 1) defined in QUBO form as an argument.

[0062] Annealing machine 30 is, in one example, a quantum annealing machine. A quantum annealing machine is an annealing machine that utilizes quantum effects such as quantum tunneling and quantum superposition. Spin variables or binary variables are realized by qubits, and solutions are searched using the aforementioned quantum effects.

[0063] The annealing machine 30 may also be a simulated annealing machine that does not use quantum effects. For example, the annealing machine 30 may be a digital annealer (registered trademark), a CMOS annealing machine, or the like.

[0064] The annealing machine 30 has an estimation unit 31. The estimation unit 31 has the same function as the estimation unit 242 of the terminal device 20 in Embodiment 1. Therefore, in this embodiment, the terminal device 20 does not need to have an estimation unit 242.

[0065] The estimation system 2 performs the estimation process shown in Figure 3, similar to the terminal device 20 of Embodiment 1. In detail, the terminal processing unit 24 (terminal communication unit 23) of the terminal device 20 transmits the nucleic acid sequence information group and base sequence information to the annealing machine 30 via the network N after the processing in step S202. The terminal processing unit 24 (terminal communication unit 23) may also transmit multiple base sequence information to the annealing machine 30 via the network N.

[0066] The estimation unit 31 of the annealing machine 30 estimates the affinity of nucleic acids, consisting of base sequence information, to a target protein using an objective function c(s) that is set to output a predetermined value for partial base sequences common to nucleic acids included in the set of nucleic acid sequence information. In other words, the procedure for finding the solution to the objective function is performed by the annealing machine 30. The estimation unit 31 may also extract base sequence information with high affinity to the target protein based on the affinity information.

[0067] The annealing machine 30 transmits the estimation results estimated by the estimation unit 31 to the terminal device 20 via the network N. The terminal communication unit 23 receives the estimation results (affinity information) via the network N.

[0068] As described above, the estimation system 2 of Embodiment 2 further comprises an annealing machine 30 having an estimation unit 31, the estimation unit 31 of the annealing machine 30 estimating the affinity of nucleic acids, which consist of base sequence information, to target proteins. As a result, the estimation system 2 uses an annealing machine 30 specialized for annealing processing to estimate the affinity of nucleic acids to target proteins, making it possible to more efficiently estimate nucleic acid sequences that show high affinity to target proteins.

[0069] Preferably, the annealing machine 30 is a quantum annealing machine, and the procedure for finding the solution to the objective function is performed by the quantum annealing machine. Unlike simulated annealing annealing machines, quantum annealing machines are less likely to fall into so-called local optima (the best solution in a certain range, not the overall optimal solution). Therefore, by having the estimation system 2 perform the procedure for finding the solution to the objective function using a quantum annealing machine, it becomes possible to more efficiently estimate nucleic acid sequences that show high affinity to the target protein.

[0070] (Modification 1 of Embodiment 1) In Embodiment 1, the score function c k (s) was the function shown in Equation 2, but the score function c k (s) may be a function as shown in Equation 4, for example. The notation is the same as in Embodiment 1.

[0071]

number

[0072] The score function c shown in Equation 2 k (s) shows the base sequence s and the nucleic acid sequence s k No particular constraints were placed on the position of the matching subarray l-mer. In contrast, the score function c shown in Equation 4 k (s) considers the subsequence l-mer that starts within the range P (where P is a positive integer) before and after position i of base sequence s. That is, the score function c k The objective function c(s) is set to output a predetermined value for sub-nucleotide sequences that can commonly bind to predetermined positions in the nucleotide sequences included in the nucleic acid sequence information set. This is because the positions where the sub-nucleotide sequence l-mer appears in nucleotide sequences with high affinity for the target protein are often similar to the positions of the sub-sequences in the nucleic acid sequence information. Since P searches for positions similar to the positions of the sub-sequences in the nucleic acid sequence information, a small value is preferable, and in one example, P is a positive integer of 2 or less.

[0073] Figure 7 shows the concept of the subsequence combination comparison process shown in Equation 4. In Figure 7, a subsequence of the base sequence s and the nucleic acid sequence s are shown when length l=3 and P=2. k This shows the comparison process with a subarray. Also, in Figure 7, an example of the process when i=5 is on the left and i=6 is on the right, with an arrow pointing from left to right in between.

[0074] In the example shown in Figure 7, the estimation unit 242 determines the nucleic acid sequence s for the partial sequence GAG ​​of the base sequence s. k The subsequences GGT~ATT are compared in order. Similar to the score function in Math 2, if the subsequence GAG ​​matches the subsequence s of the nucleic acid sequence s, k If it exists in any of the subsequences between GGT and ATT, the variable δ becomes 1. The estimation unit 242 is the nucleic acid sequence s k When comparing the subarrays up to j=min(i+P, L-l+1) (in the example shown in Figure 7, i=5 and P=2 on the left side, so if we assume L=26, the value of i+P is smaller than L-l+1, so j=7), the process for i=6 is executed.

[0075] The estimation unit 242 determines the nucleic acid sequence s for the partial sequence AGA of the base sequence s k The partial sequences GTG~TTC are compared in order. The estimation unit 242, similar to the case of i=5, evaluates the nucleic acid sequence s k The subsequences are compared up to j=min(i+P, L-l+1) (in the example shown in Figure 7, i=6 and P=2 on the right side, so if we assume L=26, the value of i+P is smaller than L-l+1, so j=8), and then the process for i=7 is executed. In this way, the estimation unit 242 sequentially compares the subsequences of the base sequence s and the subsequences of the nucleic acid sk sequence until j=min(i+P, L-l+1) and i=L-l+1 are obtained.

[0076] As explained above, in this modified example, the objective function c(s) (score function c k(s)) is set to output a predetermined value for a partial base sequence that can commonly bind to a predetermined position in the base sequence included in the nucleic acid sequence information set. This enables the estimation device to more efficiently estimate nucleic acid sequences that show high affinity to the target protein.

[0077] (Modification 1 of Embodiment 2) In the modified example 1 of Embodiment 1, the score function c k The objective function c(s) was set to output a predetermined value for sub-nucleotide sequences that can commonly bind to predetermined positions in the nucleotide sequences included in the set of nucleic acid sequence information. k (s) In order to optimize the objective function c(s) using the annealing machine 30 of Embodiment 2, a formulation using binary variables is required.

[0078] In formulating using binary variables, we have 4 × L binary variables x such that the i ∈ {1, 2, ..., L}th base is l when t ∈ {A, T, G, C}. i、t We introduce ∈{0, 1}, (i={1, 2, ..., L}, t={A, T, G, C}. Binary variable x={x i、t Using}, the function shown in Equation 1 can be shown as the objective function in Equation 5.

[0079]

number

[0080] Furthermore, the function H corresponds to the function shown in Equation 4. k (x) is shown by the number 6.

[0081]

number

[0082] Function H shown in Equation 6 k (x) is the function c shown in Equation 4. k (s) is the transformed expression for processing by the annealing machine 30, and function Hk The content of (x) is the score function c shown in Equation 4. k Since it is similar to (s), the function H k The details of the process performed by (x) are omitted.

[0083] As explained above, in this modified version as well, the estimation system enables more efficient estimation of nucleic acid sequences that exhibit high affinity for the target protein.

[0084] (Modifications common to Embodiments 1 and 2) In the embodiments described above, RNA was assumed as an example of nucleic acid. However, the estimation device and estimation system may also estimate the affinity of DNA to a target protein.

[0085] Furthermore, the objective function was set such that the smaller the value of the objective function corresponding to the solution of the objective function, the higher the degree of binding to the base sequences contained in the nucleic acid sequence information group. However, it may also be set so that the larger the value of the objective function corresponding to the solution of the objective function, the higher the degree of binding to the base sequences contained in the nucleic acid sequence information group.

[0086] (Example 1) The estimation process according to the present invention, using the objective functions shown in Equations 5 and 6, was carried out in the experimental environment described below. • Experimental environment OS: Debian GNU / Linux(registered trademark) 11 CPU: Intel® Xeon® Silver 4316 CPU @ 2.30GHz Memory: 500GB Programming language used: Python (registered trademark) Ising machine: Fixstars Amplify AE [1]

[0087] Figure 8 shows the estimation results in Example 1. The estimation results in Figure 8 are for array length L=30, length l=5, P=1, and weight w k =n k This is the estimated result at that time.

[0088] In the estimation results shown in Figure 8, "edit distance" indicates the number of different bases between the sequence output from the sequence analyzer 10 (nucleic acid sequence) and the sequence to be defined (nucleotide sequence). "Aptamer ID" indicates the ID of the sequence output from the sequence analyzer 10 (nucleic acid sequence). For example, if the nucleic acid sequence is "ATCGTACGTAGCTGTGGATCTCGTCGATC" (where "Aptamer ID" is the ID of this sequence), and the nucleotide sequence is "ATCGTACGAAGCTGTGGATCGCGTCGATC" (where the 9th base from the left is changed from "G" to "A", and the 9th base from the right is changed from "T" to "G"), then the edit distance is 2 because there are two different bases between the nucleic acid sequence and the nucleotide sequence.

[0089] In the estimation results shown in Figure 8, the base sequence "GAACGAGATTCTGAGGGTTCTCCTAATACA" had the lowest score, meaning it was the sequence with the highest affinity for the target protein.

[0090] Figure 9 shows the results of activity measurement of the base sequence obtained in Example 1. The graph in Figure 9 shows the results of binding activity measurement using the Biacore® instrument, with the relative binding activity ("Relative binding / capture") relative to the control sequence set to 100 on the vertical axis and the score ranking shown in Figure 8 on the horizontal axis.

[0091] As shown in Figure 9, the nucleotide sequence "TAACGAGATTCTGAGGGTTCTCCTAATAGG" (the sequence that ranked 4th in the estimation results in Figure 8) was found to have high relative binding activity. In addition, other nucleotide sequences also showed high relative binding activity, and the estimation process according to the present invention made it possible to select multiple sequences with activity.

[0092] Those skilled in the art will understand that various changes, substitutions, and modifications can be made without departing from the spirit and scope of the invention. For example, the embodiments and modifications described above may be combined as appropriate within the scope of the invention. [Explanation of Symbols]

[0093] 1, 2 Estimation System 20 Terminal device (estimated device) 22 Terminal storage unit (storage unit) 23 Terminal Communication Unit 24 Terminal Processing Unit 242 Estimation Department 243 Output section 30. Annealing Machine (Quantum Annealing Machine) 31 Estimation part

Claims

1. A memory unit that stores a set of nucleic acid sequence information and base sequence information whose affinity to the target protein has been determined by the SELECT method, An estimation unit estimates the affinity of the nucleic acid consisting of the base sequence information to the target protein, using an objective function set to output a predetermined value for a partial base sequence common to the nucleic acids included in the nucleic acid sequence information group, with respect to the base sequence information. An output unit capable of outputting affinity information indicating affinity to the target protein, An estimation device characterized by having [a certain feature].

2. The memory unit stores a plurality of the base sequence information, The estimation unit extracts the base sequence information that has high affinity for the target protein based on the affinity information, The estimation apparatus according to claim 1, wherein the output unit outputs the extracted base sequence information.

3. The estimation device according to claim 1 or 2, wherein the objective function is set to output a predetermined value for a partial base sequence that can commonly bind to a predetermined position in the base sequence included in the nucleic acid sequence information group.

4. The estimation device according to claim 3, wherein the smaller the value of the objective function corresponding to the solution of the objective function, the higher the degree of binding to the base sequence included in the nucleic acid sequence information group.

5. The estimation apparatus according to claim 4, wherein the objective function is an Ising model type Hamiltonian, and the procedure for finding a solution to the objective function is performed by a quantum annealing machine.

6. The objective function (c(s)) is defined as follows: T is the number of elements in the set of different base sequences, wk is the weight, L is the nucleic acid sequence and the length of the base sequence, l is the length of the subsequence, δ is a variable that is 1 if the two sequences match and 0 if they do not match, s is the base sequence, and s k As a nucleic acid sequence, [Math 1] Here, [Math 2] The estimation device according to claim 1.

7. The objective function (c(s)) is defined as follows: T is the number of elements in the set of different base sequences, wk is the weight, L is the nucleic acid sequence and the length of the base sequence, l is the length of the subsequence, P is a positive integer, δ is a variable that is 1 if the two sequences are identical and 0 if they are not identical, s is the base sequence, and s k As a nucleic acid sequence, [Math 3] Here, [Math 4] The estimation device according to claim 2.

8. The objective function (c(s)) is defined as follows: T is the number of elements in the set of distinct base sequences, wk is the weight, L is the nucleic acid sequence and the length of the base sequence, l is the length of the subsequence, P is a positive integer, x is a binary variable, and s k As a nucleic acid sequence, [Math 5] Here, [Math 6] The estimation device according to claim 4.

9. An estimation system comprising a terminal device and an annealing machine, The aforementioned terminal device is A terminal storage unit that stores a set of nucleic acid sequence information and base sequence information, in which affinity for the target protein has been confirmed by the SELECT method, A terminal communication unit that transmits the nucleic acid sequence information set and the base sequence information, and receives affinity information indicating affinity to the target protein, It has an output unit capable of outputting the affinity information, The aforementioned annealing machine is The system includes an estimation unit that estimates the affinity of the nucleic acid consisting of the base sequence information to the target protein, using an objective function set to output a predetermined value for a partial base sequence common to the nucleic acids included in the nucleic acid sequence information group, based on the base sequence information. An estimation system characterized by the following:

10. The sequence information of nucleic acids whose affinity for the target protein was determined by the SELECT method, and the base sequence information are stored in the memory unit. Using an objective function configured to output a predetermined value for a partial base sequence common to the nucleic acids included in the nucleic acid sequence information group, the affinity of the nucleic acid consisting of the base sequence information to the target protein is estimated. The system outputs affinity information indicating affinity for the target protein. An estimation method characterized by the above.

11. In the estimation device, The sequence information of nucleic acids whose affinity for the target protein was confirmed by the SELECT method, and the base sequence information are stored. Using the aforementioned base sequence information, an objective function is set to output a predetermined value for a partial base sequence common to the nucleic acids included in the nucleic acid sequence information group, thereby estimating the affinity of the nucleic acid consisting of the base sequence information to the target protein. To output affinity information indicating affinity to the aforementioned target protein, An estimation program characterized by the following: