Deep learning model, algorithm, program, and computer
The deep learning model addresses the challenge of searching for compounds with reactivity by incorporating higher-order nucleic acid structures and continuous labels, facilitating efficient compound selection with improved reactivity and reliability.
Patent Information
- Application Number
- PCT/JP2025/013837
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-08
- Filing Date
- 2025-04-04
- Publication Date
- 2025-10-16
AI Technical Summary
Existing deep learning models struggle to efficiently search for compounds that exhibit reactivity to a target sample due to their inability to effectively utilize higher-order structural features of aperiodic polymers like nucleic acids, which are complex and change significantly with subtle differences in primary structure, and are unsuitable for learning reactivity as a label.
A deep learning model that includes higher-order structure of nucleic acids as an explanatory variable and learns reactivity to a target sample as a label, using continuous value labels and experimental measurements, combined with an algorithm and computer program to calculate an evaluation function for selecting compounds based on reliability and variance.
Enables efficient search for compounds with enhanced reactivity to target samples by flexibly extracting structural features, allowing for comprehensive evaluation and selection of candidate compounds with desired properties.
Smart Images

Figure JP2025013837_16102025_PF_FP_ABST
Abstract
Description
Deep learning models, algorithms, programs and computers
[0001] The present invention relates to a deep learning model, algorithm, program, and computer for searching for compounds that exhibit reactivity to a target sample.
[0002] Biomolecules such as proteins, peptides, and nucleic acids are polymers with complex primary structures (sequences) that exhibit aperiodicity, and it is not easy to extract their structural features. For example, Patent Literature 1 describes the construction of a machine learning model that extracts mathematical interactions between substructures in the primary structure of a biomolecule and uses them as explanatory variables to search for biomolecules that satisfy desired properties.
[0003] However, constructing such a machine learning model requires explicit assumption of independent substructural features at the outset, making it difficult to extract complex structural features. Furthermore, the higher-order structures of aperiodic polymers are more complex than the primary structure because they change significantly depending on subtle differences in the primary structure. In particular, nucleic acids are known to transition between multiple higher-order structures in solution in an equilibrium reaction, making the higher-order structural features of nucleic acids particularly complex.
[0004] However, at the same time, the higher-order structure can sometimes explain the properties of aperiodic polymers better than the primary structure. For example, the binding reaction to a target molecule occurs when the three-dimensional surface of the higher-order structure comes into contact with the three-dimensional surface of the target molecule. Therefore, for example, the binding strength of an aperiodic polymer to a target molecule can be more accurately explained by the higher-order structure than by the primary structure.
[0005] Deep learning has made remarkable progress in recent years, making it possible to extract features of the higher-order structure of aperiodic polymers. For example, Non-Patent Document 1 and Patent Document 2 describe how deep learning can be used to flexibly extract higher-order structural features of input nucleic acids without the need for explicit assumptions.
[0006] JP 2016-511884 A Korean Patent Publication No. 20200019294 A
[0007] Michiaki Hamada (corresponding author) published the paper "Deep generative design of RNA family sequences" in "Nature Methods" by Nature Publishing Group on January 18, 2024.
[0008] However, the deep generative models described in Non-Patent Document 1 and Patent Document 2 are unsupervised models. The deep generative model described in Non-Patent Document 1 does not learn reactivity to a target sample as a label, and is limited to extracting structural features that frequently appear in an input nucleic acid group. It is therefore difficult to extract structural features that explain reactivity to a target sample. Therefore, it is inherently unsuitable for searching for compounds that exhibit reactivity to a target sample.
[0009] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a deep learning model, algorithm, program, and computer that can efficiently search for compounds that show reactivity to a target sample.
[0010] To solve the above problems, the present invention has the following features: (1) A deep learning model that includes a higher-order structure of a nucleic acid as an explanatory variable and learns the reactivity of the nucleic acid to a target sample as a label.
[0011] (2) The deep learning model according to claim 1, wherein the reactivity to the target sample is measured by an experiment using immobilized nucleic acids.
[0012] (3) The deep learning model according to claim 3, wherein the length of the nucleic acid is 10 to 100 mer.
[0013] (4) The deep learning model of claim 1, wherein the labels are composed of continuous values.
[0014] (5) An algorithm comprising: a step of calculating an evaluation function of the nucleic acid using the output of a deep learning model that includes a higher-order structure of the nucleic acid as an explanatory variable and that has been trained using the reactivity of the nucleic acid to a target sample as a label; and a step of selecting nucleic acids using the evaluation function.
[0015] (6) The algorithm according to claim 5, wherein the evaluation function is calculated using a value representing the reliability of an output of the deep learning model.
[0016] (7) A program for causing a computer to execute the steps of: using a deep learning model including structural information of nucleic acids as explanatory variables to learn the reactivity of the nucleic acids to a target sample as a label; calculating an evaluation function for the nucleic acids using the output of the deep learning model; and selecting nucleic acids using the evaluation function.
[0017] (8) The program according to claim 7, for executing a procedure of calculating the evaluation function using a value representing reliability of an output of the deep learning model.
[0018] (9) A computer including the program according to claim 7 or claim 8.
[0019] By using the above means, the present invention can provide a deep learning model, algorithm, program, and computer that can efficiently search for compounds that show reactivity to a target sample.
[0020] FIG. 1 is an explanatory diagram showing a deep learning model that includes the higher-order structure of a candidate compound as an explanatory variable and learns using the reactivity of the candidate compound to a target sample as a label. FIG. 2 is a process diagram showing a process of calculating an evaluation function for a candidate compound using the output of the deep learning model of the present invention and selecting a candidate compound using the evaluation function. FIG. 3 is an explanatory diagram showing a deep learning model having an LSTM layer. FIG. 4 is an explanatory diagram showing a deep learning model having a graph convolution layer. FIG. 5 is an explanatory diagram showing a VAE (Variational Auto Encoder) model that uses the primary structure and higher-order structure of a nucleic acid aptamer candidate as explanatory variables. FIG. 6 is a graph showing a comparison between the search results for nucleic acid aptamers using an unsupervised VAE model and a deep regression model. FIG. 7 is a graph showing a comparison between the search for nucleic acid aptamers using different evaluation functions. FIG. 8 is an explanatory diagram showing a deep classification model trained in Example 3. 1 is a graph comparing the fluorescence intensities of 10 candidate compounds with high evaluation functions in the deep regression model shown in FIG. 4 of Example 1 with the fluorescence intensities of 10 candidate compounds selected by the deep classification model shown in FIG. 8 trained in Example 3. FIG. 1 is an explanatory diagram showing an example of the higher-order structure of a nucleic acid molecule. FIG. 2 is an explanatory diagram showing an example of the higher-order structure of a nucleic acid molecule. FIG. 3 is an explanatory diagram showing an example of the higher-order structure of a nucleic acid molecule. FIG. 4 is an explanatory diagram showing a higher-order structure predicted using the mFold web server. FIG. 5 is a graph showing a comparison between the maximum fluorescence intensity value obtained by <Selection of candidate compounds> of Example 6 and the maximum fluorescence intensity signal value obtained by <Selection of candidate compounds by Bayesian optimization> of Example 2. FIG. 6 is an explanatory diagram showing an example of a deep learning model. FIG. 7 is an explanatory diagram showing an example of a deep learning model. FIG. 8 is a graph showing a comparison between the maximum fluorescence intensity signal value obtained by <Selection of candidate compounds> of Example 7 and the maximum fluorescence intensity signal value obtained by <Selection of candidate compounds by Bayesian optimization> of Example 2. FIG. 9 is an explanatory diagram showing an example of a deep learning model. FIG. 10 is an explanatory diagram showing an example of a neural network model. 10 is a graph showing a comparison of the coefficient of determination of a deep neural network model and the coefficient of determination of a neural network model.10 is a graph comparing the fluorescence intensity of binding to IgE for each of the bottom 10 and top 10 predicted values of IgE reactivity, and the fluorescence intensity of binding to IgE for each of the bottom 10 and top 10 predicted values of IgG reactivity.
[0021] Next, an embodiment of the present invention will be described in detail with reference to the accompanying drawings, in which the same reference numerals are used to designate common components in the drawings.
[0022] <Target Sample> The type of target sample is not particularly limited and may be a molecule, a structure composed of several molecules, or a mixture. The target sample may also be a metal, a crystal, an ion, or the like. The use of the target sample is not limited and may be in the medical field or the industrial field. Examples of molecules include biomarker molecules that can be used to diagnose disease or detect pre-disease, biomolecules that can treat disease or alleviate symptoms by inhibiting or enhancing function, and molecules that need to be purified or removed in the manufacture of products. Examples of structures include cells and viruses. Examples of mixtures include biological samples such as whole blood, white blood cells, peripheral blood mononuclear cells, plasma, serum, sputum, breath, urine, semen, saliva, cerebrospinal fluid, amniotic fluid, glandular fluid, lymph, nipple aspirate, bronchial aspirate, synovial fluid, joint aspirate, cells, cell culture medium, cell extract, feces, tissue extract, cerebrospinal fluid, solutions in industrial plant reactors, waste liquids, etc.
[0023] <Compounds Reactive with Target Samples> Compounds reactive with target samples can bind to target samples, detect target samples, or enhance or inhibit the function of the target sample. The compounds reactive with target samples are not limited, but nucleic acids such as nucleic acids, peptides, and antibodies are preferred from the viewpoints of ease of production, storage stability, and the like. The partial structure of a nucleic acid can selectively bind intramolecularly or extramolecularly. This allows for the formation of complex higher-order structures, which can be expected to selectively bind to specific substances, enhance or inhibit their function, and perform various other reactions. Furthermore, nucleic acids are preferred from the viewpoints of ease of production, low cost, storage stability, and the ability to simultaneously evaluate the reactivity of a large number of structures with target samples, since multiple types can be synthesized on a substrate using inkjet technology.
[0024] <Nucleic Acid> The main chain structure and type of base of the nucleic acid are not particularly limited. Examples of the main chain structure of the nucleic acid include DNA, RNA, cross-linked artificial nucleic acid (Locked Nucleic Acid: LNA), peptide nucleic acid (PNA), and analogs thereof. A single nucleic acid aptamer may contain one type of main chain structure or multiple types of main chain structures. Examples of bases include natural bases such as adenine, guanine, cytosine, thymine, and uracil, modified natural bases, and natural bases with modified substituents. That is, the nucleic acid may be an artificial nucleic acid that contains bases with any chemical modification or has an arbitrarily modified main chain structure to improve structural diversity. Some bases may be substituted with any chemical structure, such as a fluorescent molecule. Alternatively, the nucleic acid may not have any bases. Furthermore, the nucleic acid may be modified depending on the purpose, such as preventing decomposition in vivo or immobilizing it on a carrier.
[0025] <<Deep Learning Model>> Figure 1 is an explanatory diagram showing a deep learning model 1 that includes the higher-order structure of candidate compound 2 as an explanatory variable 3 and learns the reactivity of candidate compound 2 to a target sample as a label 4. As shown in Figure 1, deep learning model 1 includes the higher-order structure of candidate compound 2 as an explanatory variable 3 and learns the reactivity of candidate compound 2 to a target sample as a label 4, but the form may be any and is not limited to one. Deep learning model 1 will be described using as an example a case where candidate compound 2 is a natural DNA nucleic acid aptamer.
[0026] <Candidate Compound> A candidate compound (candidate substance) 2 is a compound that is a candidate for a substance that is reactive to a target sample. The structure of the candidate compound 2 is a primary structure, a higher-order structure, or both converted into a numerical value, vector, matrix, or graph structure.
[0027] <Target Molecule> The type of target molecule is not particularly limited, and may be, for example, a target molecule used in the medical field or a target molecule used in the industrial field. A target molecule is a molecule that functions as a biomarker that can diagnose disease or detect pre-disease. Other target molecules include biomolecules that can treat disease or alleviate symptoms by inhibiting or enhancing their function, and molecules that are to be purified or removed in the production of a product.
[0028] The target molecule used in the method for obtaining a compound having molecular recognition ability of the present invention is, for example, a molecule contained in a biological sample. Examples of biological samples include whole blood, leukocytes, peripheral blood mononuclear cells, plasma, serum, sputum, breath, urine, semen, saliva, cerebrospinal fluid, amniotic fluid, glandular fluid, lymph, nipple aspirate, bronchial aspirate, synovial fluid, joint aspirate, cells, cell culture medium, cell extract, feces, tissue extract, cerebrospinal fluid, etc. The target molecule does not have to be a biomolecule such as a protein, but may also be a low molecular weight compound, polymer, etc.
[0029] <Example of Deep Learning Model> Figure 3 is an explanatory diagram showing an example of a deep learning model 1 having an LSTM layer 11. In the example of deep learning model 1 in Figure 3, a primary structure consisting of a sequence of four types of bases in natural DNA: A (adenine), T (thymine), G (guanine), and C (cytosine), and a higher-order structure expressed as a string of characters in dot bracket format, in which bases that form hydrogen bonds (and) and bases that do not form hydrogen bonds (・), are converted into vectors or matrices and used as explanatory variables 3. These are input to deep learning model 1 having an LSTM layer 11 (Long Short Term Memory: LSTM), which excels at extracting features from character strings and time-series information. The output data 5 is passed to the next fully connected layer 12, and a predicted reactivity value can be obtained as output data 5.
[0030] Fig. 4 is an explanatory diagram showing an example of deep learning model 1 having a graph convolution layer 13. In the example of deep learning model 1 in Fig. 4, the higher-order structure of a nucleic acid aptamer is regarded as a graph structure and converted into a vector or matrix, which is input to deep learning model 1 having a graph convolution layer 13. In this case, the primary structure is included in the graph structure as main chain bonds.
[0031] <Explanatory Variables of Deep Learning Model> As described above, the higher-order structure included in the explanatory variables 3 of the deep learning model 1 shown in FIG. 1 may be in any form as long as it contains secondary or higher structural information of the nucleic acid. For example, it may be a character string, such as a dot-bracket notation, a graph structure, atomic coordinates, base coordinates and angles, the strength or probability of bonds formed between bases or between bases and the main chain, or a qualitative variable classifying the higher-order structure into patterns. The method for preparing the higher-order structure may be any, and it may be experimentally measured or identified, simulated, or predicted. Furthermore, as long as higher-order structure information can be extracted using deep learning, any information other than the higher-order structure, such as a primary structure, free energy representing stability, or simulation results, may be included. Furthermore, mathematical transformations or feature extraction may be used. For example, only the information about the higher-order structure may be converted and used, or any form of information converted so as to intermingle information about the higher-order structure with any information other than the higher-order structure, such as a numerical value, vector, matrix, character string, graph, or signal, may be used. Furthermore, multiple higher-order structures may be used for one primary structure. It is known that nucleic acids move between multiple higher-order structures in a solution in an equilibrium reaction. The higher-order structural characteristics of nucleic acids are particularly complex, but by using multiple higher-order structures that a single primary structure can form in solution, it is possible to take into account the group of higher-order structures involved in the equilibrium reaction, thereby enabling more accurate prediction of reactivity. Furthermore, explanatory variable 3 can include any information, such as free energy representing the stability of the higher-order structure or simulation results.
[0032] <Labels of Deep Learning Model> The format of the label 4 used in training the deep learning model 1 may be any format as long as it includes information related to reactivity with the target sample. The format of the label 4 may be, for example, a dissociation constant (KD value) representing affinity to the target sample, the brightness, area, or image of luminescence generated by reaction with the target sample, or the value or waveform of an electrical signal, or a conversion thereof. It is desirable that the label 4 be composed of continuous values rather than discrete values. This is because, for example, if a binary value representing high or low reactivity is used as the label 4 to search for compounds with higher reactivity with the target sample, structural features with high reactivity with the target sample can be extracted. However, it is difficult to extract structural features that can further increase reactivity with the target sample. Furthermore, not all candidate compounds 2 need to have a label 4; some data without a label 4 may be included.
[0033] <Structure of Deep Learning Model> The structure of the deep learning model 1 may be any structure as long as it includes the higher-order structure of the candidate compound 2 as an explanatory variable 3 and learns the reactivity of the candidate compound 2 to a target sample as a label 4. The structure of the deep learning model 1 may include, for example, a recurrent neural network (RNN), a long short-term memory (LSTM) layer 11, a graph convolution layer 13, a convolution layer 14, a deconvolution layer 15, a fully connected layer 12, and a transformer encoder layer, as shown in Figures 3 and 4. Alternatively, the structure of the deep learning model 1 may be a deep regression model 10 or a conditional generative model such as a cVAE (conditional variational autoencoder). The number of layers is not limited, and may include any normalization, activation function, transformation, etc.
[0034] <Output of Deep Learning Model> Figure 2 is a process diagram showing a process of calculating an evaluation function 6 for a candidate compound 2 using output data 5 of a deep learning model 1 of the present invention and selecting a candidate compound 2 using the evaluation function 6. The deep learning model 1 shown in Figure 2 does not need to output any output data 5, or may output any output data 5. The deep learning model 1 may or may not output, for example, a prediction of the reactivity of the candidate compound 2 to a target sample, the structure of a new candidate compound 2 that reacts with the target sample, or a deviation of the prediction, as output data 5.
[0035] <Learning of Deep Learning Model> The deep learning model 1 is generally trained to reduce the loss function. The loss function may be designed arbitrarily, but for the purpose of improving the learning accuracy of the reactivity to the target sample, it is preferable to use, for example, MSE (mean square error) as the loss function or to use a loss function including an MSE term. MSE is expressed by the following equation (1):
[0036] An example of a term in the loss function that reduces other label errors is cross-entropy error. The deep learning model 1 may be trained each time, or a trained model may be reused. Alternatively, only a portion of a trained model may be reused. Multiple deep learning models 1 may also be trained, allowing the deviation representing the variability of predictions to be calculated.
[0037] <<Evaluation Function and Candidate Compound Selection>> The evaluation function 6 shown in FIG. 2 is used to determine the candidate compound 2 to be selected, and its form is not limited. The evaluation function 6 may be, for example, a single numerical value, or may be composed of multiple numerical values, or may be converted into a rank, symbol, or the like. It is desirable that the evaluation function 6 is suitable for searching for compounds that exhibit desired reactivity. FIG. 2 shows examples of candidate compounds 2 and their selection states 7. A "○" mark in the selection state 7 indicates a selected compound, and an "×" mark indicates a non-selected compound. Examples of the evaluation function 6 and the selection of candidate compounds 2 include the following:
[0038]
[0039]
[0040]
[0041] The upper confidence limit function UCB is an acquisition function used when searching for compounds with higher reactivity to a target sample, and candidate compounds 2 with a large value of the upper confidence limit function UCB may be selected.
[0042] In addition to UCB, there are other types of acquisition functions, such as Lower Confidence Bound (hereinafter referred to as "LBC"), Expected Improvement (hereinafter referred to as "EI"), and Probability of Improvement (hereinafter referred to as "PI"). Different acquisition functions may be used depending on the situation. The Lower Confidence Bound function LCB is expressed by the following equation (3):
[0043] The lower limit reliability function LCB is an acquisition function used when searching for candidate compounds 2 (candidate substances) with low reactivity to the target sample, and candidate compounds 2 with a small value of the lower limit reliability function LCB can be selected.
[0044]
[0045] <Selection of Candidate Compounds by Applying D Optimization> It is also possible to select candidate compounds 2 with as much variance as possible. For example, when considering a matrix formed by linking variable vectors, the determinant value is large when each variable is nearly independent. For example, a matrix is created by linking the explanatory variable x for combinations of candidate compounds 2, and the determinant is calculated as the evaluation function 6. By selecting the combination of candidate compounds 2 with the largest determinant, it is possible to select a combination of candidate compounds 2 with a large variance in the explanatory variable x. The vector used to calculate the determinant does not have to be the explanatory variable x as long as it is linked to the candidate compound 2; for example, the vector z(x) from the intermediate or final layer when the explanatory variable x is introduced into the deep regression model 10 may be used. By selecting candidate compounds 2 with as much variance as possible, it is possible to avoid selecting similar candidate compounds 2, thereby enabling a broad search.
[0046] In this way, deep learning model 1 of the present invention includes the higher-order structure of nucleic acid shown in Figure 1 as explanatory variable 3, and learns the reactivity of nucleic acid to a target sample as label 4. With this configuration, deep learning model 1 can flexibly extract structural features important in reactivity to a target sample, thereby efficiently searching for compounds that exhibit reactivity to the target sample.
[0047] In addition, the reactivity of deep learning model 1 to a target sample is measured by an experiment using immobilized nucleic acid.
[0048] With this configuration, the reactivity to the target sample is measured through experiments using immobilized nucleic acids, and since a large number of structures can be synthesized inexpensively on a substrate using inkjet technology, the reactivity to the target sample can be evaluated simultaneously for a large number of structures. This allows for the collection of a sufficient amount of training data for training the deep learning model 1.
[0049] Furthermore, the length of the nucleic acid is preferably 10 to 100 mer. With this configuration, the nucleic acid can take on various higher-order structures and can be synthesized on a substrate with high precision by inkjet printing.
[0050] Furthermore, it is desirable that the label 4 be composed of continuous values. With such a configuration, the label 4 is composed of continuous values, and can take many different values, so it contains a wealth of information. Therefore, the continuous value label 4 can extract structural features that more flexibly explain reactivity to a target sample, as well as structural features that further enhance reactivity to a target sample.
[0051] The algorithm of the present invention also includes the steps of: calculating an evaluation function for nucleic acids using the output of a deep learning model that includes the higher-order structure of nucleic acids as an explanatory variable and that has been trained using the reactivity of nucleic acids to a target sample as a label; and selecting nucleic acids using the evaluation function. According to this configuration, for example, by including a step of calculating an evaluation function 6 using a value representing the reliability of output data 5 from deep learning model 1, the reliability and unknown degree of the search can also be taken into consideration. This allows efficient selection of compound candidates that are likely to have desired properties.
[0052] Furthermore, it is desirable that the evaluation function 6 be calculated using a value representing the reliability of the output data 5 of the deep learning model 1. According to this configuration, the algorithm can determine the selection of a desired candidate compound 2 by calculating the evaluation function 6 using a value representing the reliability of the output data 5 of the deep learning model 1.
[0053] Furthermore, the program of the present invention processes information by having a computer execute the following steps: using a deep learning model 1 including structural information about nucleic acids in explanatory variables x to learn the reactivity of nucleic acids to a target sample as labels 4; calculating an evaluation function 6 for nucleic acids using output data 5 from the deep learning model 1; and selecting nucleic acids using the evaluation function 6. With this configuration, the program of the present invention can also take into account the reliability and unknown degree of the search, for example, by calculating the evaluation function 6 using a value representing the reliability of the output data 5 from the deep learning model 1. This makes it possible to efficiently select compound candidates that are likely to have desired properties.
[0054] Furthermore, it is desirable that the program processes information by executing a procedure for calculating an evaluation function 6 using a value representing the reliability of the output data 5 of the deep learning model 1. According to this configuration, the program can determine a desired candidate compound 2 by calculating the evaluation function 6 using a value representing the reliability of the output data 5 of the deep learning model 1.
[0055] The computer of the present invention also includes the program. With this configuration, the computer, having the program, can select candidate compounds 2 using the evaluation function 6, enabling compounds to be evaluated by integrating multiple indicators (biological activity, physical properties, safety, etc.). This allows for comprehensive evaluation, enabling well-balanced candidate compounds 2 to be selected.
[0056] Figure 6 is a graph showing Example 1 of the present invention, which compares the results of searching for nucleic acid aptamers using an unsupervised VAE model that does not learn reactivity to target samples as labels with the results of searching for nucleic acid aptamers using the deep regression model 10 of the present invention that learns reactivity to target samples as labels.
[0057] In order to search for nucleic acid aptamers with higher binding affinity to immunoglobulin E protein (hereinafter referred to as "IgE protein"), unsupervised deep learning model 1 was compared with a deep learning model 1 of the present invention that includes the higher-order structure of candidate compound 2 as explanatory variable 3 and that learns the reactivity of candidate compound 2 with a target sample as label 4. The binding affinity was measured for each candidate compound 2 selected by the unsupervised model that learns only the primary structure and by a deep regression model 10 that learns the higher-order structure. As a result, as shown in Figure 6, the deep regression model 10 showed a higher measured binding affinity. From this, it can be said that the present invention, which uses deep learning model 1 that includes the higher-order structure of candidate compound 2 as explanatory variable 3 and learns the reactivity of candidate compound 2 with a target sample as label 4, makes it possible to efficiently search for nucleic acids.
[0058] <Preparation of Candidate Compounds> IgE-immobilized magnetic beads were mixed with the DNA library and stirred for 1 hour. The supernatant was then removed using a magnetic stand. The magnetic beads were then washed with phosphate buffer containing Tween 20. Next, 0.15 N aqueous sodium hydroxide solution was added dropwise to dissociate the DNA from the magnetic beads, and the mixture was placed on a magnetic stand to recover the supernatant. 1 N aqueous HCl solution was added dropwise to the recovered solution to neutralize it. This sample was then recovered as DNA that is thought to bind to the target molecule. The recovered DNA was subjected to PCR amplification using biotin-labeled primers to prepare single-stranded fragments, thereby preparing the next round of libraries. After repeating this process three times, the post-screening fractions were processed using the NEBNext® Multiplex Oligos for Illumina kit, and the sequences contained in the fractions were identified using a next-generation sequencer, MiSeq, and designated as candidate compound 2.
[0059] <Preparation of Labels> A microarray with nucleic acid aptamer candidates immobilized on a substrate in an array was obtained from Agilent Technologies. A phosphate buffer solution containing 2% bovine serum albumin (BSA), 0.05% Tween 20, and 100 nM IgE was applied to the microarray and allowed to react for 1 hour at 25°C. The microarray was then washed with phosphate buffer to remove unreacted IgE. Subsequently, 2% BSA, 0.05% Tween 20, and 100 nM fluorescein-labeled anti-IgE antibody were applied to the microarray and allowed to react for 1 hour at 25°C. The microarray was then washed with phosphate buffer to remove unreacted fluorescein-labeled anti-IgE antibody, and the fluorescein fluorescence was photographed using a fluorescence microscope. The photographed images were then subjected to image analysis using machine learning, and the fluorescence intensity of each spot was quantified as a binding reaction signal.
[0060] <Prediction of higher-order structure of nucleic acid aptamer candidate> The higher-order structure of the nucleic acid aptamer candidate was predicted using the RNAfold command of the Vienna RNA program developed by the University of Vienna. Note that the higher-order structure of DNA was predicted by setting parameters for DNA.
[0061] <Selection of candidate compounds using an unsupervised deep learning model> A VAE (Variational Auto Encoder) model was trained, as shown in Figure 5, with the primary structure and higher-order structure of the nucleic acid aptamer candidate as explanatory variable 3. The VAE model was trained so that explanatory variable 3 and output x' were similar. The intermediate layer z was clustered into 10 clusters using a GMM (Gaussian mixture model), and the coordinates of 10 centers were calculated. Each center coordinate was input into the decoder section of the VAE model, and the structure of the nucleic acid aptamer candidate corresponding to each center coordinate was generated and selected as a new candidate.
[0062] <Selection of candidate compounds using the deep learning model of the present invention> The primary structure and higher-order structure of the nucleic acid aptamer candidate were used as explanatory variables 3, and the fluorescence intensity obtained in the above <Preparing labels> step was used as label 4. A deep regression model 10 shown in Figure 4 was trained, and 10 nucleic acid aptamer candidates with high predicted fluorescence intensity values were selected. At this time, the primary structure and higher-order structure of the nucleic acid aptamer candidate were represented in the form of a graph and trained into a deep learning model with a graph convolution layer. In addition, MSE was used as the loss function during training.
[0063] <Comparison of Deep Learning Models> For each of the nucleic acid aptamer candidates selected by the above steps, the above <Preparation of Labels> step was repeated, and the fluorescence was quantified as a binding reaction signal and compared. As a result, as shown in Figure 6, deep learning model 1 of the present invention had a higher maximum binding reactivity. From this, it can be said that deep learning model 1 of the present invention, which includes the higher-order structure of candidate compound 2 as explanatory variable 3 and learns the reactivity of candidate compound 2 to a target sample as label 4, is capable of searching for nucleic acids more efficiently than unsupervised deep learning model 1.
[0064] In searching for nucleic acid aptamers with higher binding affinity to IgE, the selection method for candidate compound 2 was compared. As a result, it was shown that compounds showing reactivity to target samples can be searched for more efficiently by calculating the evaluation function for candidate compounds calculated using the output of the deep learning model of the present invention using a value representing the reliability of the output of the deep learning model.
[0065] <Preparation of Candidate Compound> Candidate compound 2 in Example 2 was prepared in the same manner as in Example 1.
[0066] <Preparation of Label> The label 4 in Example 2 was prepared in the same manner as in Example 1.
[0067] <Creation of new nucleic acid aptamer candidates> One million new nucleic acid aptamer candidates were created by randomly performing base sequence deletions, insertions, mutations, etc. on the primary structure of the nucleic acid aptamer candidates. These new nucleic acid aptamer candidates were not used in training the deep learning model 1 described below, but were used only for prediction.
[0068] <Prediction of higher-order structure of nucleic acid aptamer candidate> Prediction was carried out in the same manner as in Example 1.
[0069]
[0070]
[0071] <Selection of Candidate Compounds by Application of D Optimization> The deep learning model 1 trained in <Selection of Candidate Compounds by Phagocytosis Method> was repurposed as the deep learning model 1. A combination of 1,000 new nucleic acid aptamer candidates was created, and the corresponding final layer vector z(x) of the deep learning model 1 was concatenated to create a matrix and calculate the determinant. This process was repeated 10,000 times. Of the 10,000 determinants obtained, 1,000 new nucleic acid aptamer candidates corresponding to the determinant with the largest value were selected. In the training of the deep learning model in Example 2, the primary structure of the nucleic acid aptamer candidate was expressed as a string of base sequences, and the higher-order structure was expressed as a string of dotted brackets, and these were trained on the deep learning model of Figure 3 with an LSTM layer.
[0072] <Comparison of methods for selecting candidate compounds> Figure 7 is a graph comparing the search for nucleic acid aptamers using different evaluation functions 6. For each candidate compound 2 selected by the above steps, the above <Preparing labels> step was performed again, and the fluorescence intensity was quantified as a binding reaction signal and compared. As a result, as shown in Figure 7, the new candidate compound 2 selected by Bayesian optimization exhibited the highest value of binding reactivity. From this, it can be said that nucleic acids can be searched for more efficiently by using an evaluation function 6 calculated using a value representing the reliability of the output data 5 output from the deep learning model 1 of the present invention.
[0073] In this example, a comparison was made between candidate molecule search using a deep learning model trained using discrete value labels and candidate molecule search using a deep learning model trained using continuous value labels. Non-patent document "Machine learning guided aptamer refinement and discovery" (https: / / www.nature.com / articles / s41467-021-22555-9) reports a case where nucleic acids were searched for by training a deep learning model using discrete value labels related to the binding strength with a target sample. Dissociation constant K D The lower the value, the stronger the binding strength between the candidate compound and the target sample. D A higher value indicates a lower binding affinity between the candidate compound and the target sample. D High binding affinity <128 nM, 128 nM < K D <512 μM medium binding, 512 nM < K D The molecules are classified into three groups, with low binding strengths of <2 μM. The discrete value labels of these groups are trained into a deep classification system to search for nucleic acid molecules with higher binding strengths.
[0074] However, for discrete value labels, e.g., K D A candidate compound with a K D Both candidate compounds, which have a K value of 100 nM and are close to the classification of moderate binding affinity, D<128 nM. Therefore, the deep learning model learns all candidate compounds as having the same binding strength. Therefore, discrete value labels are inherently unsuitable for searching for compounds with higher reactivity to the target sample.
[0075] In Example 1, the deep regression model shown in Figure 4 was trained using continuous value labels of fluorescence intensity to search for nucleic acid aptamers with stronger binding affinity to IgE. In this example, the deep classification model shown in Figure 8 was constructed using discrete value labels that were categorized into three groups (high, middle, and low) in descending order of fluorescence intensity, and candidate molecule searches using the deep regression model and the deep classification model were compared. As a result, it was found that candidate molecules could be searched for more efficiently by learning continuous value labels as in the present invention, rather than by learning discrete value labels.
[0076] <Preparation of Candidate Compound> Candidate compound 2 in Example 3 was prepared in the same manner as in Example 1.
[0077] <Preparation of Labels> In preparing labels in Example 3, continuous value labels prepared in the same manner as in Example 1 were divided into three groups in descending order of continuous value and converted into top, middle, and bottom discrete value labels.
[0078] <Prediction of Higher-Order Structure of Nucleic Acid Aptamer Candidate> Prediction of the higher-order structure of the nucleic acid aptamer candidate in Example 3 was carried out in the same manner as in Example 1.
[0079] <Selection of Candidate Compounds Using a Deep Classification Model> The deep classification model shown in Figure 8 was trained using the primary and secondary structures of the nucleic acid aptamer candidates as explanatory variables 3 and the discrete value labels obtained in the above <Preparing Labels> step as labels 4. The primary and higher-order structures of the nucleic acid aptamer candidates were represented in the form of a graph and trained into a deep learning model with a graph convolutional layer. Furthermore, since this is a multi-class classification model, cross-entropy error was used as the loss function. Ten nucleic acid aptamer candidates whose classification predictions were in the top category among the top, middle, and bottom categories were selected. Since it was not possible to distinguish which nucleic acid aptamer candidates had particularly strong fluorescence intensity within the range of those with high classifications, the ten nucleic acid aptamer candidates were selected randomly.
[0080] <Comparison of methods for selecting candidate compounds> Figure 9 is a graph comparing the fluorescence intensities of 10 candidate compounds with high evaluation function UCBs for the deep regression model shown in Figure 4 of Example 1 with the fluorescence intensities of 10 candidate compounds selected by the deep classification model shown in Figure 8 trained in this example. As a result of the comparison, as shown in Figure 9, candidate compounds predicted to have high intensities by the deep regression model had higher maximum binding reactivity values than candidate compounds predicted to be classified as high (high, middle, or low) by the deep classification model. This suggests that candidate compounds can be more efficiently searched for by training a deep learning model using continuous value labels rather than discrete values.
[0081] Nucleic acids can take on multiple higher-order structures in solution even if they have the same primary structure, indicating that their higher-order structural forms are complex.
[0082] <Calculation of higher-order structure, Gibbs energy, and abundance ratio in solution of nucleic acid molecules> Using the mFold web server, which can predict the higher-order structure of nucleic acids, we predicted the higher-order structures that the following nucleic acid sequence A could take in solution and their Gibbs energy ΔG. Nucleic acid sequence A: AGGTATTGGA GCGGAGCTGG ATGCGCACTA TATATACC As a result, the three higher-order structures shown in Figures 10 to 12 were predicted, and the Gibbs energies were as shown in Table 1.
[0083] The lower the Gibbs energy ΔG, the more stable the higher-order structure is, and the higher the abundance ratio P of that higher-order structure in solution. The abundance ratio P of each higher-order structure i in solution can be calculated from its higher-order Gibbs energy ΔGi using the following formula (4).
[0084] where R is the gas constant, and T is the temperature in Kelvin, and the calculation was performed assuming T = 310.15 K, which corresponds to 37 degrees Celsius. As a result, the abundance ratio Pi was as shown in Table 1 below.
[0085]
[0086] As mentioned in the background section, the binding reaction of a nucleic acid to a target molecule occurs when the three-dimensional surface of the higher-order structure comes into contact with the three-dimensional surface of the target molecule, so the higher-order structure may explain the properties of the nucleic acid more than the primary structure. Therefore, it can be said that the higher-order structure may have a greater explanatory power than the primary structure regarding the binding strength of a nucleic acid to a target molecule.
[0087] However, as shown in Table 1, although nucleic acid sequence A has one type of primary structure, it was found to take on a number of completely different higher-order structures as shown in Figures 10 to 12. Furthermore, from the calculated abundance ratios, it was estimated that, for example, one type of higher-order structure does not account for 90%, but rather exists at proportions of 55.6%, 22.4%, and 22.0%, respectively.
[0088] Although compound higher-order structural information is important for efficiently searching for compounds that are reactive to target samples, nucleic acids, in particular, can take on complex higher-order structural forms in solution, even if the primary structure is the same. For this reason, among deep learning models with excellent flexibility in feature extraction, a deep learning model that can flexibly extract higher-order structural features important for reactivity to target samples, which learns the reactivity of candidate compounds to target samples as labels and uses the higher-order structural information of candidate compounds as explanatory variables, was shown to be suitable.
[0089] This shows that even a small mutation in the primary structure of a nucleic acid can significantly change the higher-order structure. The following nucleic acid sequence B is a nucleic acid sequence A with only one base substitution. Specifically, the 12th C from the 5' end of nucleic acid sequence A was substituted with T, but the substitution rate is low at 2.6%, since this is only one base out of 38 bases.
[0090] Nucleic acid sequence B: AGGTATTGGA GTGGAGCTGG ATGCGCACTA TATATACC. The higher-order structure of this nucleic acid sequence B was predicted using the mFold web server in the same manner as in Example 4. The predicted higher-order structure was shown in Figure 13, which is significantly different from any of Figures 10 to 12. As described in the background and Example 4, the higher-order structure can sometimes be more descriptive of the binding strength of a nucleic acid to a target molecule than the primary structure. However, in nucleic acids, even a single base substitution in the primary structure can significantly change the higher-order structure. This indicates that a deep learning model used to search for compounds, particularly nucleic acids, that are reactive to a target sample should, for example, include the higher-order structure of a candidate compound as an explanatory variable rather than using only the primary structure of the candidate compound as an explanatory variable.
[0091]
[0092] <Preparation of Candidate Compound> Candidate compound 2 in Example 6 was prepared in the same manner as in Example 1.
[0093] <Preparation of Labels> Labels in Example 6 were prepared in the same manner as in Example 1.
[0094] <Prediction of Higher-Order Structure of Nucleic Acid Aptamer Candidate> Prediction of the higher-order structure of the nucleic acid aptamer candidate in Example 6 was carried out in the same manner as in Example 1.
[0095]
[0096] Using this as a loss function, the deep regression model of Figure 3 was trained, as in Example 2, with the primary and secondary structures of the nucleic acid aptamer candidate as explanatory variables 3 and the fluorescence intensity obtained in the above step <Preparing labels> as label 4. In this case, the primary structure of the nucleic acid aptamer candidate was expressed in the form of a character string of base sequences, and the higher-order structure was expressed in the form of a character string in dot bracket notation, and these were trained into the deep learning model of Figure 3 having an LSTM layer.
[0097]
[0098] <Comparison of Candidate Compound Selection Methods> The maximum fluorescence intensity value obtained by the above <Candidate Compound Selection> was compared with the maximum fluorescence intensity signal value obtained by the <Candidate Compound Selection by Bayesian Optimization> in Example 2 in Figure 14. Figure 7 shows that the candidate compound selection method using an evaluation function calculated using a value representing reliability based on the output of multiple deep learning models by Bayesian optimization is efficient. However, Figure 14 shows that the maximum fluorescence intensity value obtained by the candidate compound selection method using an evaluation function calculated using a value representing reliability based on the output of a single deep learning model was comparable. This demonstrates that using a value representing the reliability of the output of a deep learning model in the calculation of the evaluation function enables highly efficient search for candidate compounds, and that the method for preparing the value representing reliability is not limited.
[0099] This paper presents an example of compound discovery using a deep learning model to output new compound structures. Any method can be used as long as it includes the higher-order structure of the candidate compound in the present invention as an explanatory variable and learns the reactivity of the candidate compound to a target sample as a label. Therefore, for example, the structure of a new candidate compound may be output. A deep learning model that learns labels and outputs new data can output corresponding data when a label is input. A non-patent document (Semi-supervised Learning with Deep Generative Models, https: / / arxiv.org / pdf / 1406.5298) reports that when a deep learning model, conditional VAE (cVAE), was trained with handwritten images of the digits 0, 1, 2, 3, . . . 9 and their labels, it was able to output various handwritten images of the number 2, for example, when the number 2 was input.
[0100] Here, in the case of the deep learning models of Figures 3 and 4, it is necessary to prepare a large number of new candidate compounds in advance, predict their higher-order structures and input them into the deep learning model, predict their reactivity to a target sample, and select candidate compounds with high reactivity to the target sample. However, for example, in the case of deep learning models configured as in Figures 15 and 16, which output the structure of a new candidate compound, it is not necessary to handle a large number of candidate compounds, and by specifying a value of reactivity to a desired target sample, it is possible to generate a structure of a new candidate compound that can achieve that reactivity value, which is simple.
[0101] <Preparation of Candidate Compound> Candidate compound 2 in Example 7 was prepared in the same manner as in Example 1.
[0102] <Preparation of Labels> Labels in Example 7 were prepared in the same manner as in Example 1, and were normalized to 0 to 1.
[0103] <Prediction of Higher-Order Structure of Nucleic Acid Aptamer Candidate> Prediction of the higher-order structure of the nucleic acid aptamer candidate in Example 7 was carried out in the same manner as in Example 1.
[0104] <Learning of Deep Learning Model> Each of the deep learning models in FIGS. 15 and 16 was trained as follows.
[0105] <Learning of the deep learning model in Figure 15> The encoder unit in Figure 15 takes as input a matrix obtained by converting the primary structure and higher-order structure of a candidate nucleic acid, and outputs vectors M and S and a predicted reactivity value. Next, the intermediate layer Z is calculated using Z = M + S × ε, where ε is a random number sampled from a standard normal distribution. The decoder unit takes as input both the intermediate layer Z and the predicted reactivity value, and converts them into a generator matrix. This generator matrix may have eight rows, for example, as shown in Equation (6) below. The horizontal direction of the generator matrix corresponds to the position of the bases of the nucleic acid, and the vertical direction corresponds to the probability of the type of base or higher-order structure. The higher-order structure is expressed using dot brackets. Furthermore, only the 5 bases on the 5' side are shown.
[0106]
[0107] The above generator matrix is further converted into primary structures and higher-order structures with high probabilities at each base position, and is expressed in the same format as the input, for example, as shown below, thereby improving interpretability.
[0108] CAGTA,,, .. .. ((,,,
[0109] In learning the deep learning model of FIG. 15, the loss function was obtained by adding an MSE term to the loss function of the VAE model, as shown in the following equation (7).
[0110]
[0111] This deep learning model can use only the encoder or decoder section. The encoder section can predict the reactivity of each candidate compound to a target sample by inputting its primary and higher-order structures. The decoder section can output a new compound structure that can achieve the reactivity value to the target sample by inputting any intermediate layer coordinates and the reactivity value to the target sample.
[0112] In this way, the deep learning model of Figure 15 includes the higher-order structure of the candidate compound in the present invention as an explanatory variable, and by learning the reactivity of the candidate compound to a target sample as a label, it can output a predicted value of reactivity to the target sample and the structure of the candidate compound corresponding to that predicted value.
[0113] <Learning of the deep learning model in Figure 16> The encoder unit in Figure 16 takes as input the primary structure of the candidate nucleic acid, a matrix converted from the higher-order structure, and a label of reactivity to the target sample, and outputs vectors M and S. Next, the intermediate layer Z is calculated using Z = M + S × ε, where ε is a random number sampled from a standard normal distribution. The decoder unit takes as input both the intermediate layer Z and the reactivity to the target sample, and converts them into a generator matrix. This generator matrix, similar to the above <Learning of the deep learning model in Figure 15>, takes an 8-row format and is further converted into primary structures and higher-order structures with high probabilities at each base position. As the loss function, the following equation (8) was used as the loss function of the VAE model.
[0114] This deep learning model can use either the encoder or the decoder alone. The decoder can output a new compound structure that can achieve the reactivity value for the target sample by inputting any intermediate layer coordinates and the reactivity value for the target sample.
[0115] In this way, the deep learning model of Figure 16 includes the higher-order structure of the candidate compound in the present invention as an explanatory variable, and by learning the reactivity of the candidate compound to a target sample as a label, it can output the structure of a candidate compound that can exhibit reactivity to any target sample.
[0116] <Selection of Candidate Compounds> Using the deep learning models of Figures 15 and 16 trained as described above, 10,000 types of candidate compounds were output for each. At this time, labels were normalized to 0 to 1 and used for training, but 1.0 was specified. Random numbers following a standard normal distribution were used as the value of the intermediate layer Z. At this time, the NLL value (negative log-likelihood) is known as a value representing the reliability of the output of the deep learning model, so the NLL value was calculated as an evaluation function. The NLL value is expressed by the following formula (9).
[0117]
[0118]
[0119] <Comparison of Candidate Compound Selection> The maximum fluorescence intensity signal obtained by the above <Candidate Compound Selection> was compared with the maximum fluorescence intensity signal obtained by the <Candidate Compound Selection by Bayesian Optimization> in Example 2 in Figure 17 . Figure 7 shows that when selecting candidate compounds using a deep learning model that does not output new compound structures, selecting candidate compounds using an evaluation function calculated using a value representing the reliability of the output is efficient. Figure 17 also shows that the fluorescence intensity signal was evaluated for candidate compound selection using an evaluation function calculated using a value representing the reliability of the output of a deep learning model that outputs new compound structures. In this case, the maximum value obtained in the case of the <Candidate Compound Selection by Bayesian Optimization> in Example 2 was normalized to 1, but the maximum fluorescence intensity signal values for the deep learning models in Figures 15 and 16 were similar. This indicates that deep learning models that output new compound structures are also useful in compound discovery.
[0120] Furthermore, while the deep learning model in Figure 15 predicts reactivity to the target sample, the deep learning model in Figure 16 does not. However, the maximum values of the fluorescence intensity signals were similar, indicating that the form of the deep learning model is not limited as long as it includes the higher-order structure of the candidate compound as an explanatory variable and learns the reactivity of the candidate compound to the target sample as a label.
[0121] We show that it is better to include higher-order structures as explanatory variables even in deep learning models that output new compound structures.
[0122] <Preparation of Candidate Compound> Candidate compound 2 in Example 8 was prepared in the same manner as in Example 1.
[0123] <Preparation of Labels> Labels in Example 8 were prepared in the same manner as in Example 7.
[0124] <Prediction of Higher-Order Structure of Nucleic Acid Aptamer Candidate> Prediction of the higher-order structure of the nucleic acid aptamer candidate in Example 8 was carried out in the same manner as in Example 1.
[0125] <Learning of the deep learning model in Fig. 18> The deep learning model in Fig. 18 is a modification of the deep learning model in Fig. 16 such that the higher-order structure is not used as an explanatory variable and the higher-order structure is not output. Accordingly, the generator matrix has five rows, as shown in the following formula (10).
[0126] The above generator matrix is further transformed into a probable primary structure at each base position as follows: CAGTA, , ,
[0127] The loss function used in training the deep learning model in Fig. 18 is the same as in <Training the deep learning model> in Fig. 16. However, the generator matrix has five rows.
[0128] Comparison of the Accuracy of the Output of New Compound Structures from the Deep Learning Models of Figures 16 and 18 For the deep learning models of Figures 16 and 18 , a nucleic acid structure not used in learning and a label indicating reactivity to a target sample were input into each deep learning model, and the accuracy of the output compound structure was examined to determine how similar the output compound structure was to the input compound structure. Accuracy was calculated by calculating the percentage of bases in the primary structure of each nucleic acid that matched the input and output, and the average value was used to determine which was higher. As a result, it was found that the deep learning model of the present invention, shown in Figure 16 , including the higher-order structure as an explanatory variable had a 3.64% higher accuracy of the primary structure output than the deep learning model of Figure 18 , which did not include the higher-order structure as an explanatory variable. This indicates that the primary structure alone is insufficient as a structural feature important for reactivity to a target sample, and that including the higher-order structure as an explanatory variable enables more accurate extraction of primary structure features important for reactivity to a target sample. This indicates that it is desirable for the deep learning model of the present invention to include the higher-order structure as an explanatory variable.
[0129] This embodiment of the present invention demonstrates that a deep learning model (deep neural network model) is preferable to a neural network model. Patent document US2021319851A1 lists random forests, logistic regression, linear regression, neural networks, sparsity-driven convex optimization fit, and support vector machines as examples of machine learning models useful for compound discovery. However, for example, neural network models are single-layer models. Unlike deep learning models (deep neural network models), these models do not have a structure with multiple layers stacked on top of each other, and therefore tend to lack expressive power and have low accuracy.
[0130] <Preparation of Label> The label 4 in Example 9 was prepared in the same manner as in Example 1.
[0131] <Prediction of higher-order structure of nucleic acid aptamer candidate> Prediction was carried out in the same manner as in Example 1.
[0132]
[0133]
[0134]
[0135] As a result, as shown in Figure 20, the deep learning (deep neural network) model was 2.83 times more accurate than the neural network model. This indicates that models with multiple layers tend to have higher expressive power and higher accuracy, making them more desirable for compound discovery. In this embodiment, it was found that the deep learning (deep neural network) model is preferable to the neural network model.
[0136] One aspect of the present invention is that it is possible to search for selective candidate compounds. For example, if it is possible to search for a candidate compound that is reactive with a specific target sample A but not with a different specific target sample B, it would be very useful because it would enable the detection of only A in a sample in which A and B are mixed.
[0137] <Preparation of Labels> Label 4 in Example 9 was prepared in the same manner as in Example 1, with the fluorescence intensity indicating binding to IgE being measured. Furthermore, for immunoglobulin G protein (hereinafter referred to as "IgG protein"), IgG was bound to a microarray on which the same nucleic acid aptamer candidate as IgE had been immobilized, and after washing, fluorescein-labeled anti-IgG antibody was applied, followed by washing again, and then photographed under a fluorescence microscope. In this manner, the fluorescence intensity of each spot was quantified as a signal of the binding reaction between IgG and the nucleic acid aptamer candidate.
[0138] <Prediction of higher-order structure of nucleic acid aptamer candidate> Prediction was carried out in the same manner as in Example 1.
[0139]
[0140] <Creation of new nucleic acid aptamer candidates> The same procedure as in Example 1 was carried out.
[0141] <Selection of Selective Compounds> For the candidates created in the above <Creation of New Nucleic Acid Aptamer Candidates>, reactivity to IgE and IgG was predicted using a deep learning model. At this time, the top 100 nucleic acid aptamer candidates with the highest predicted values for reactivity to IgE were selected. Subsequently, 10 nucleic acid aptamer candidates with low predicted values for reactivity to IgG and 10 nucleic acid aptamer candidates with high predicted values for reactivity to IgG were selected. For these nucleic acid aptamer candidates with the lowest and highest predicted values for reactivity to IgG, the above <Preparation of Labels> step was again carried out, targeting IgE and IgG. The fluorescence intensity was then quantified as a binding reaction signal, and the medians were compared. At this time, the top 10 data were normalized to 1.
[0142] As a result, as shown in Figure 21, the median fluorescence intensity representing binding to IgE did not change significantly, but the median fluorescence intensity representing binding to IgG changed by more than 10 times. This demonstrates that this embodiment of the present invention makes it possible to search for compounds that have selectivity for the target sample with which they react.
[0143] Although the embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments and can be modified as appropriate within the scope of the present invention.
[0144] 1 Deep learning model 2 Candidate compound 3 Explanatory variables 4 Label 5 Output data 6 Evaluation function 7 Selection state
Claims
1. A deep learning model that includes the higher-order structure of nucleic acids as an explanatory variable and learns the reactivity of the nucleic acids to target samples as labels.
2. The deep learning model according to claim 1, wherein the reactivity to the target sample is measured by an experiment using immobilized nucleic acid.
3. The deep learning model according to claim 2, wherein the length of the nucleic acid is 10 to 100 mer.
4. The deep learning model of claim 1, wherein the labels are composed of continuous values.
5. An algorithm comprising: a step of calculating an evaluation function of the nucleic acid using the output of a deep learning model that includes the higher-order structure of the nucleic acid as an explanatory variable and that has been trained using the reactivity of the nucleic acid to a target sample as a label; and a step of selecting nucleic acids using the evaluation function.
6. The algorithm according to claim 5, wherein the evaluation function is calculated using a value representing the reliability of the output of the deep learning model.
7. A program for causing a computer to execute the following steps: using a deep learning model that includes structural information of nucleic acids as explanatory variables, to learn the reactivity of the nucleic acids to a target sample as a label; and calculating an evaluation function for the nucleic acids using the output of the deep learning model, and selecting nucleic acids using the evaluation function.
8. The program according to claim 7, for executing the procedure of calculating the evaluation function using a value representing the reliability of the output of the deep learning model.
9. A computer including the program according to claim 7 or claim 8.
Citation Information
Patent Citations
Antibody design system, antibody design method, program and recording medium
JP2004295256A
Method of estimating secondary structure in RNA and program and apparatus therefor
WO2008072713A1
Method for estimating dynamic state of secondary structure of RNA
WO2011062166A1