Methods and apparatus for determining RNA sequences

CN114078569BActive Publication Date: 2026-09-11ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110928748.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-14
Filing Date
2021-08-13
Publication Date
2026-09-11
Estimated Expiration
2041-08-13

AI Technical Summary

Benefits of technology

[0005] In comparison, this invention offers the following advantages: it can explore a much larger search space, meaning it can generate/find a far more diverse range of candidate sequences with practical relevance that were previously unattainable through computational techniques. To date, unbalanced brackets and partial structures have been unmanageable. With this invention, solutions can be defined and found within a 'design task'.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114078569B_ABST
    Figure CN114078569B_ABST
Patent Text Reader

Abstract

Method and device for determining RNA sequences. The invention relates to a method for creating a policy (P) set up for determining the positioning of nucleotides within a primary RNA structure from fragments of a pre-given secondary structure, the method comprising the following steps: initializing the policy (P); providing a task representation (T), wherein the task representation (T) comprises structure constraints (C) of a secondary RNA structure and order constraints (O) of a primary RNA structure; determining a primary candidate RNA sequence (S) from the task representation (T) by means of the policy (P); adapting the policy (P) by means of a reinforcement learning algorithm such that the total loss (L) is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to methods, training devices, computer programs, and machine-readable storage media for determining RNA sequences by means of learned strategies. Background Technology

[0002] Because many functional RNA molecules have been shown to participate in the regulation of transcription, epigenetics, and translation, RNA molecule design has recently attracted interest in medicine, synthetic biology, biotechnology, and bioinformatics. Since the function of RNA depends on its structural properties, the problem in RNA design lies in finding RNA sequences that satisfy given structural constraints.

[0003] The publication, "Learning to Design RNA" by Runge et al. (In International Conference on Learning Representations, 2019), is available online. This paper discloses an algorithm for RNA design problems using "Deep Reinforcement Learning" to train a policy network to sequentially design an entire RNA sequence corresponding to a pre-given target structure. Summary of the Invention

[0004] Advantages of the invention However, all current expressions of RNA design strongly constrain their solution space in such a way that the expression requires structural priority for the whole molecule or at least for the complete form of the desired molecule.

[0005] In comparison, this invention offers the following advantages: it can explore a much larger search space, meaning it can generate / find a far more diverse range of candidate sequences with practical relevance that were previously unattainable through computational techniques. To date, unbalanced brackets and partial structures have been unmanageable. With this invention, solutions can be defined and found within a 'design task'.

[0006] Furthermore, this method can transfer learned knowledge to previous RNA design expression tasks. This allows for more efficient RNA sequence finding. The partial RNA design according to this invention can be understood as a super-problem of reverse RNA design and reverse RNA design with predefined sequences.

[0007] Disclosure of the invention In a first aspect, the present invention relates to a method for creating strategies ( A computer-implemented method, wherein the strategy is set up to determine the location of nucleotides within the primary RNA structure based on a pre-given secondary structure fragment.

[0008] The method includes the following steps: initializing the policy. For example, the policy can be implemented using a neural network. Therefore, the policy can be initialized, for example, by randomly adjusting the weights of the neural network.

[0009] Next is providing a task representation ( ), wherein the task representation ( ) including structural constraints on secondary RNA structures ( ) and the sequence constraints of primary RNA structure ( Next is based on the task representation (). ) by means of the strategy ( Identify primary candidate RNA sequences ( ), where by means of the strategy ( ), using the strategy ( The identified nucleotides gradually occupy the candidate RNA sequence. The site of the primary RNA structure. Next is the sequence constraint ( Determine the candidate RNA sequence ( Sequence loss () ); The (RNA) folding algorithm (Faltungsalgorithmus) F was applied to the candidate RNA sequence. Next, it is necessary to determine the structure of the fold (). ) and pre-given structural constraints ( Structural loss between ) ); According to the sequence loss ( ) and the structural loss ( Determine the total loss ( ).

[0010] Next, the strategy is applied using a reinforcement learning algorithm. Adaptation is performed to optimize the total loss ( ).

[0011] It is proposed that the fragment is related to parameters, which are optimized together when the strategy is optimized.

[0012] In a second aspect, the present invention relates to a strategy for using learned methods according to any one of the preceding claims ( Determine the RNA sequence containing portions of the secondary and primary structures of a given RNA. A method, wherein the strategy is set up to determine the nucleotide localization of RNA based on fragments of secondary structure, the method comprising the steps of: providing the task representation ( And according to the task representation ( The fragments were gradually identified using a strategy to determine candidate RNA sequences. ).

[0013] In other respects, the present invention relates to apparatus and computer program respectively configured to perform the above-described methods, and a machine-readable storage medium having the computer program stored thereon. Attached Figure Description

[0014] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. In the drawings: Figure 1 A schematic diagram illustrating the RNA design problem; Figure 2 An embodiment of the present invention is illustrated schematically; Figure 3 A schematic diagram illustrating the hyperparameter optimization of a reinforcement learning algorithm is shown. Figure 4 A table showing possible hyperparameters is provided; Figure 5 The possible structures of the training device are shown. Detailed Implementation

[0015] In its most basic structural form, RNA consists of a sequence of four nucleotides: adenine (A), guanine (G), cytosine (C), and uracil (U). This nucleotide sequence is called the RNA sequence or primary structure.

[0016] During the construction of RNA sequences as a building block, the functional structure of the RNA molecule is determined by folding, which transforms the RNA sequence into its 3D tertiary structure. The inherent thermodynamic properties of the sequence determine the resulting folding. Hydrogen bridges formed between two corresponding nucleotides represent one of the driving forces in the thermodynamic model and strongly influence the tertiary structure. Structures including these hydrogen bridges are generally referred to as the secondary structure of RNA.

[0017] The problem of finding RNA sequences that fold into the desired secondary structure is known as the RNA design problem or RNA inverse folding.

[0018] Figure 1This diagram illustrates the RNA design problem using the folding algorithm F and punkt-Klammer notation. Given the desired RNA secondary structure represented by punkt-Klammer notation (a), the task is to design an RNA sequence (b) that folds into the desired secondary structure (c).

[0019] Below, "partial RNA design" should be defined and one embodiment of the invention should be described to integrate not only sequence features but also structural features into a simple, common task representation, in particular to support knowledge transfer across different RNA design tasks.

[0020] RNA design considers two search spaces: the sequence space, which includes the nucleotide chain. The structural space consists of a sequence of typical secondary structure features. Composition. It should be noted that the common punctuated notation used here is the one adopted from Ivo Hofacker, Walter Fontana, Peter Stadler, Sebastian Bonhoeffer, Manfred Tacker and Peter Schuster, Fast Folding and Comparison of RNA Secondary Structures (Chemical Monthly) 125:167–188, 02 1994.

[0021] RNA folding algorithm F uses a length of... RNA sequence Mapped to its corresponding secondary structure Transformation occurs between these spaces.

[0022] RNA design focuses on the opposite process: given a sequence with secondary structure characteristics... The goal is to find RNA sequences. This ensures that the RNA sequence satisfies the equation .

[0023] Additional sequence restrictions can be used. To exclude parts of the solution space, this makes RNA design an NP-hard problem, see https: / / www.liebertpub.com / doi / full / 10.1089 / cmb.2019.0420.

[0024] Partial RNA design extends expression by allowing unconstrained domains in the structural space, potentially leading to RNA design tasks involving unbalanced brackets and opening doors for exploration through computer-aided methods. Formally, partial RNA design can be defined as follows: It is an RNA folding algorithm, and It is a length of The structural restriction sequences confine the space of the effective RNA secondary structure to... ,and This refers to nucleotide restriction sequences, which spatially restrict the effective RNA sequence to... Therefore, the goal of some RNA design is to find RNA sequences that satisfy the following equation: .

[0025] Since the goal is to predict plausible RNA sequences for every arbitrary structural and sequence feature (including partially and fully defined structural constraints), a simple but general representation of the RNA design task should be used below to enable knowledge transfer between different RNA design tasks. Therefore, structural and sequence constraints... The two sequences are combined into a common representation. The common representation is also referred to below as the task representation. See Figure 2 .

[0026] Additionally, define the function This function will process each point in the RNA sequence (also referred to below as a stelle). From constraints Mapped to a unique representation (for and Or in all other cases mapped to .

[0027] Additionally, a preprocessing step can be performed using pairing sites, where only one of the interacting nucleotides is known, filled with its complementary pairing partner (according to the Watson-Crick-Basenpaarungsschema). Sites where pairing partners cannot be determined in a trivial manner can be skipped, and the process continues at the next pairing site.

[0028] In reinforcement learning (English: R einforcement LIn learning and action (RL), an agent acts upon a dynamic environment through perception and action. At each step of the interaction, the agent receives an indication of the current state of the environment and selects an action based on this observation. The action changes the state of the environment, and the value of this transition is signaled to the agent as a scalar reward. The agent's ultimate goal is to maximize the long-term measure of the reward signal. Achieving optimal behavior can be a very difficult task because actions can affect state transitions and thus all subsequent rewards. In particular, the agent is not told which action will be in its best interest in the long run, and therefore the agent searches through systematic trials guided by several different algorithms, such as temporal difference learning (TD), Q-learning, or policy gradient methods.

[0029] An RL algorithm for de-RNA folding has been proposed by Runge et al. (see the background section above), which serves as the basis for this invention. In the RL case, the agent's policy (English: Policy, This is similar to an artificial neural network, which, for example, outputs a distribution of actions. The environment can be entirely defined by a decision process that provides a set of available actions, a set of states, a reward function, and a state transition probability matrix. To model a partial RNA design as a reinforcement learning problem, a scheme described by Runge et al. is referenced: the representation of states is based on provided molecular features, and actions correspond to nucleotide locations. Once nucleotides are assigned to all sites, the environment calculates a reward based on Hamming distances, which are then communicated to the agent to update its model. The policy is then adjusted using an RL algorithm to minimize the Hamming distances. The precise representation of the decision process and the architecture of the policy network can be optimized together with other parameters.

[0030] Most deRNA folding algorithms use a structural loss function. To quantify the target structure With RNA sequence The structure obtained by folding The difference between them. The optimal candidate structure (also known as the Minimizer). It has the minimum value of the loss function and corresponds to the value used for the pre-given target structure. The solution to the reverse RNA folding problem.

[0031] A common loss function is the Hamming distance. For partial RNA constructions, the desired structure may be known only partially, and the solution space may additionally be constrained in the sequence space. Therefore, the loss expression previously proposed by Runge et al. is adapted to consider only sites of designed candidate solutions, which are either constrained in the structure space, in the sequence space, or both. A site that is unconstrained is excluded from the calculation of the spacing, and thus from the calculation of the loss function. This can be achieved using an indicator function. Formalized, if the site If constrained, then the indicator function is related to length The sequence return value of constraint C is 1, while if the site If there are no restrictions, return 0.

[0032] Therefore, the loss of partially defined constraints can be obtained by applying the nucleotide constraint sequence. The constrained sites and the designed candidate solutions Between corresponding sites; and within structural constraints Constrained sites and folds of the sequence It is expressed as the sum of Hamming distances between corresponding sites. In length nucleotide boundary conditions In the case of sequences, this leads to sequence loss. : Correspondingly, structural loss It can be expressed as follows: Specific RNA task representation and the specific candidate solutions designed Total loss Therefore, it can be defined as: Minimizer Therefore, it was passed by: Give Sequence constraints can be implemented in different ways. Therefore, dimensions are set for three different schemes for generating candidate solutions in a common configuration space, which are described in the following paragraphs.

[0033] Simple Solution: For the simple solution, the proxy forecast targets the RNA design task. The nucleotides at each site, including the sequence portion.

[0034] Replacement Scheme ('Replacement'): The replacement scheme follows the same strategy as the simple scheme, but once all sites are occupied by nucleotides, the task representation is changed before the designed RNA sequence is rewarded. The sequence portion replaces the corresponding prediction portion of the candidate solution.

[0035] Partial Solution ('Partial'): In the case of the third solution, the task representation The sequence domains are completely ignored, and the agent only forecasts nucleotides for structural parts and unconstrained sites.

[0036] Use RL learning to determine the strategy. Parameters of the neural network (Polciy network) The precise architecture of these networks, along with the representation of the decision process, training hyperparameters, training data distribution, training curriculum, and the algorithm used to formulate sequences, are optimized. A policy gradient method, namely proximal policy optimization (PPO), is used to update a given policy network. parameters Runge et al. have demonstrated to date that meta-learning of RNA design strategies outperforms other learning strategies in terms of speed and accuracy, and this invention now adapts such a strategy to solve some RNA design problems. Specifically, each sampled RL algorithm first learns an RNA design strategy across thousands of local RNA design tasks (alternating sequences and structural motifs). Then, for new, previously unseen design tasks, candidate solutions are sampled from the strategy without any other parameter updates.

[0037] Implementing RL learning methods can be highly sensitive to decisions regarding agent parameters, environment, and training parameters, and expressing RL algorithms for new problems is a difficult and lengthy process because there is no experience with which design decisions yield optimal results. Automating RL expression can drastically reduce this process. To address this problem, an automatic reinforcement learning scheme (autoRL) is proposed, which automatically selects the optimal learning environment for reinforcement to solve a portion of the RNA design problem due to its rich configuration space. In particular, a meta-learning process is defined to jointly optimize the expression of the RL algorithm: in the outer loop, iterative meta-learning samples a configuration that defines the RL algorithm, and this configuration is then used to learn RNA design rules in the inner loop. The rules derived therefrom are evaluated on a validation data set, and the meta-learner observes the validation loss to update its own model accordingly. The meta-learner aims to minimize the validation loss by learning a better configuration for each observation trial, while the learner attempts to maximize its reward for each task on the validation set. More formally, the scheme of the present invention can be expressed as follows.

[0038] A is a set of algorithms for generative RNA sequence design, E is a set of RL learning environments, N is a set of RL learning agents, and D... train It is a series of training data, and C is the set of training courses that define the configuration space: .

[0039] The entire verification group Specific configuration The cost function can then be written as: .

[0040] The goal now is to train a meta-learner L on training data (training data + validation set) for a subset of RNA designs, such that the meta-learner finds its optimal configuration. The optimal configuration minimizes the cost function: .

[0041] Therefore, the search space represents Runge et al.'s extended configuration space and includes five new dimensions.

[0042] The configuration space comprises four components: decisions regarding the agent, the environment, the training data, and the algorithm used for sequence design, which is described below. Figure 4 The table in the document provides an overview.

[0043] Agent Subspace: Each agent in agent subspace A is defined by a specific architecture of the policy network and selected values ​​of a set of training hyperparameters tuned for optimization and regularization. Aside from minor variations, the agent subspace is primarily adapted to the parameters described by Runge et al. The architecture subspace is constructed as follows: (1) The task representation is either encoded, where distinctions are made between paired sites, unpaired sites, and sites with specific nucleotides or wildcards, or processed by an optional embedding layer that transforms the symbol-based representation into a learnable numerical representation for each party. Furthermore, an optional CNN with up to two layers can be selected on the embedding layer, followed by an optional LSTM with up to three layers. Finally, a planar network with one or two layers is added, outputting a distribution of actions. This parameterization covers a wide range of possible neural architectures and keeps the dimensionality of the search space relatively small. Figure 4 The diagram illustrates the search space for the neural architecture used in policy networks. Figure 4 Each path in the diagram corresponds to a specific architecture. The performance of a neural network is strongly dependent on the choice of hyperparameters. Preferably, some of the parameters used to construct the network in the PPO are recorded in a common configuration space: learning rate, batch size, and the strength of entropy regularization.

[0044] Environment subspace: Environment subspace E selects parameterized decision-making processes. The values ​​are defined by the specific values ​​and configuration space of the parameters in the decision-making process. Other parameters are optimized together. In particular, the state representation is optimized by using a state radius that is symmetric about the current site. And the number of sites that constitute the exact composition of each state, which is composed of individual states, is optimized. Furthermore, action semantic parameters can be used to analyze the impact of pair predictions and parameters used to formulate rewards. Finally, transition dynamics are related to and correspondingly defined with respect to multiple parameter decisions in different subspaces.

[0045] Training data subspace: Training data determination: which task distributions to access and which states to survey during training. Preferably, three training groups with different task distributions are recorded in a common configuration space, which training groups are accessible via training data parameters. Since different courses can lead to different performance of a particular ML algorithm, training course parameters can be further introduced to allow selection between random or ordered courses of the training data regarding task length.

[0046] Algorithm selection subspace: Here, we can choose from the solutions described above: simple solution, replacement solution, and partial solution.

[0047] This section describes the behavior more precisely to automatically select the optimal RL algorithm for partial RNA design from a shared configuration space. This is particularly relevant for meta-learning, especially as a basis for… Figure 3 The meta-learner uses the BOHB optimizer. BOHB was chosen because it can handle mixed discrete / continuous search spaces, utilize parallel resources, and further accelerate optimization by beneficially evaluating approximations of the objective function. These so-called low-fidelity approximations can be achieved in various ways, such as by limiting training time, the number of independent repetitions of evaluation, or by using only a small fraction of the available data. The training time for the sampled RL algorithm is preferably limited.

[0048] Dataset: The ultimate goal of the current scheme is to define design RNA candidates for each type of constraint in the sequence and structure space by transferring knowledge between different RNA design tasks. To properly optimize the proposed decisions listed in relation to this goal, training and validation datasets are needed, including tasks containing imbalanced brackets.

[0049] Objective Function: Although RL is known to provide noisy or unreliable results in individual optimization runs, it is preferable to use only a single meta-optimization run and a single validation set. To account for the noisy results of the optimization process, it is preferable to study beforehand three loss expressions for the optimization method: (1) the number of unsolved objectives, (2) the sum of average distances, and (3) the sum of minimum distances. Based on provisional results, variant (3), i.e., the sum of minimum distances, has been shown to be advantageous as the objective for optimization. However, variant (1) can also lead to good results. The number of unsolved objectives during the meta-optimization process is particularly preferably minimized according to variant (1).

[0050] Budget: It was found advantageous to use candidate solutions for 100 previously unseen local RNA design tasks from a validation set of learned RNA design guidelines with specified parameters. To approximate performance with varying degrees of reliability, the clock time used for the training process was limited. Each RL algorithm was thus evaluated within 60 seconds using tasks from the validation set. Finally, the established configuration was selected for evaluating different test sets.

[0051] Parameter Importance: To analyze the importance of each parameter, a functional ANOVA (fANOVA) architecture based on random forest is used. The five most important parameters in the meta-optimization case are the algorithm selection parameter, action semantics parameter, number of LSTM layers, learning rate, and state radius (ordered by their importance).

[0052] Overall, the proposed decision yields a 19-dimensional proposed space, which includes a broad spectrum of neural architectures to represent agents (including elements of recurrent neural networks (RNNs) and convolutional neural networks (CNNs), multiple different environmental representations, three different training data distributions, two training courses, three different algorithms for productive design of RNA sequences, and training hyperparameters. Figure 4 The document provides a complete list of parameters, their types, ranges, and priors.

[0053] Efficient Bayesian optimization methods can be used to optimize RL representations. See, for example, Stefan Falkner, Aaron Klein, and Frank Hutter's BOHB: Robust and efficient hyperparameter optimization at scale (in Jennifer Dy and Andreas Krause, editors, Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pp. 1437–1446, Stockholm Fair, Stockholm, Sweden, July 10-15, 2018. PMLR). It can be found online at: http: / / proceedings.mlr.press / v80 / falkner18a.html.

[0054] In the context of Reinforcement Learning (RL), the agent's policy is approximated by an artificial deep neural network whose output is a distribution of possible actions, representing the current state. The environment can be fully represented by the decision-making process. By definition, the decision-making process comprises a series of states S, a series of available actions A, a reward function R already described above, and a state transition probability matrix P. The following paragraphs describe modeling a partial RNA design as different components of the decision-making process.

[0055] The state space is represented as follows. At each time step t = 0; 1; 2; ::::; T, where T is the final time step of the interaction between the agent and the environment, and the environment provides the state st, which guides the agent when learning the policy. To provide the agent with local information, we can use... Gramm, which is represented by a task Centered on the t-th site, where This is a hyperparameter called the state radius. To enable the construction of this centered n-gram at all locations, padding symbols ("#") can be added at the beginning and end of the task representation.

[0056] The action space consists of four available nucleotides. It is conceivable that, for the mission representative... The pairing sites in the model also use Watson-Crick base pairs (AU, UA, GC, CG).

[0057] The state transition dynamics can be modeled as follows. At each time step t, the state is set to a fixed value. The state is defined by deterministic transitions at various points in the task representation. This is based on the choice of action semantics and the methods used to generate candidate solutions. With the choice of algorithm and state composition, the transition dynamics can be varied and implemented accordingly.

[0058] Figure 5 The diagram schematically illustrates a training device 141 including a provider 71 that provides a training sequence e from a training data set. The training sequence is fed to a monitoring unit 61 to be trained, from which the monitoring unit determines a total loss a. The total loss a and the training sequence e are then fed to an evaluator 74, which determines the parameters of the policy. The parameters are transmitted to the parameter memory P and replaced there. .

[0059] The method executed by the training device 141 can be stored on the machine-readable storage medium 146 as a computer program and can be executed by the processor 145.

Claims

1. A method for creating strategies ( A computer-implemented method, wherein the strategy is set up to determine the location of nucleotides within the primary structure of RNA based on a fragment of a pre-given secondary structure of RNA, the method comprising the following steps: Initialize the strategy; Provide task representation ( ), The task representation mentioned therein ( ) including structural constraints on secondary RNA structures ( ) and the sequence constraints of primary RNA structure ( ); According to the task representation ( ) by means of the strategy ( Identify primary candidate RNA sequences ( ), The strategy mentioned above is used to ( ), using the strategy ( The identified nucleotides gradually occupy the candidate RNA sequence. The site of the primary RNA structure; Regarding the order constraint ( Determine the candidate RNA sequence ( Sequence loss () The sequence loss is determined by summing the Hamming distances between the constrained sites of the sequence constraints and the corresponding sites of the candidate RNA sequence. The folding algorithm F was applied to the candidate RNA sequence. ), Determine the structure of the fold ( ) and pre-given structural constraints ( Structural loss between ) The structural loss is determined by summing the Hamming distances between the constrained sites of the structural constraint and the corresponding sites of the folded structure. According to the sequence loss ( ) and the structural loss ( Determine the total loss ( The total loss is the sum of the sequence loss and the structural loss. The strategy (using reinforcement learning algorithms) Adaptation is performed to optimize the total loss ( ).

2. The method according to claim 1, wherein an indicator function is used to determine the sequence loss ( ) and / or the structural losses ( ).

3. The method according to any one of claims 1 to 2, wherein the sequence loss ( It was determined using Hamming distance.

4. The method according to any one of claims 1 to 2, wherein the total loss ( Divide by the task representation ( The number of constraints.

5. The method according to any one of claims 1 to 2, wherein a meta-learner is used to optimize the hyperparameters of the reinforcement learning algorithm.

6. The method of claim 5, wherein the meta-learner comprises BOHB.

7. A strategy learned by means of the method according to any one of the preceding claims ( To determine the RNA sequence of a given RNA in terms of partial secondary and partial primary structure ( ) The method comprises the following steps: Provide the task representation ( ); According to the task representation ( The fragments were gradually identified using a strategy to determine candidate RNA sequences. ).

8. An apparatus configured to perform the method according to any one of the preceding claims.

9. A computer program product having a computer program configured to perform the method according to any one of claims 1 to 7.

10. A machine-readable storage medium having a computer program stored thereon, the computer program being configured to perform the method according to any one of claims 1 to 7.