Method for predicting bacillus subtilis rbs strength based on machine learning

CN118841085BActive Publication Date: 2026-05-29NINGXIA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NINGXIA UNIVERSITY
Filing Date
2024-06-04
Publication Date
2026-05-29

Smart Images

  • Figure CN118841085B_ABST
    Figure CN118841085B_ABST
Patent Text Reader

Abstract

The application discloses a method for predicting the strength of bacillus subtilis RBS based on machine learning, which comprises the following steps: constructing a bacillus subtilis RBS library; encoding the RBS sequence in the bacillus subtilis RBS library based on an encoder to extract the dinucleotide physicochemical property characteristics, RBS sequence characteristics and RBS sequence secondary structure characteristics of the RBS sequence; and fusing the RBS dinucleotide physicochemical property characteristics, RBS sequence characteristics and RBS sequence secondary structure characteristics to predict the strength of the RBS sequence. RBS is an important element for regulating translation strength, and the prediction of the strength of RBS is of great significance to drug development and biological reactor efficiency prediction. The bacillus subtilis RBS prediction model provided by the application can quickly, accurately and at low cost predict the expression strength of the RBS element.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bioinformatics technology and relates to a method for predicting the RBS intensity of Bacillus subtilis based on machine learning. Background Technology

[0002] Synthetic biology integrates many disciplines such as life science, engineering and information science, and is one of the most promising fields in modern biology. It adopts the idea of ​​engineering, and synthesizes and improves existing systems or systems by constructing standardized components and modules, so as to reveal the laws of life and build a new generation of bioengineering systems, thus creating a new research model for life science [1]. Machine learning plays an important role in the field of synthetic biology, providing biologists and researchers with powerful tools to analyze and design biological systems. Synthetic biology and machine learning have a natural synergistic effect. Machine learning can use large datasets to train models to guide synthetic biology experiments. In recent years, a large number of machine learning applications have been applied to synthetic biology research, including the design of new biological components. The introduction of machine learning has brought faster and more accurate analysis and design tools to synthetic biology research, accelerated the development of the biotechnology field, and promoted the innovation and progress of synthetic biology.

[0003] The design and analysis of regulatory elements play a crucial role in gene circuit design and hold immense potential in drug design and the production of specific products. Regulatory elements typically include promoters, ribosome binding sites (RBS), and enhancers. Different levels of regulatory element strength can influence gene expression to varying degrees, resulting in different metabolic fluxes and thus optimizing product metabolic pathways. Therefore, the rational design and analysis of the sequence structure of regulatory elements to quantitatively regulate gene expression has become a significant challenge.

[0004] Restricted biosynthetic elements (RBSs) are crucial for regulating translation intensity. Artificially designed RBSs can be easily linked to specific genes via primer design and PCR amplification, thereby modulating gene expression intensity at the translational level. Therefore, precise regulation of gene expression largely depends on the design and analysis of RBS regulatory elements. However, compared to traditional model strains like *Bacillus subtilis* and yeast, the construction of standardized biosynthetic elements and related research in *Bacillus subtilis* lags behind. Current tools for predicting the strength of RBS regulatory elements include RBS calculator, RBSDesigner, and UTRdesigner. RBS calculator and UTRdesigner offer online platforms for researchers' convenience. Both use the NUPACK or ViennaRNA toolkit to calculate the free energy of key molecular interactions during translation initiation, then establish statistical thermodynamic models to design and analyze RBS regulatory elements. RBSDesigner uses the UNAFold toolkit to calculate mRNA secondary structures and the free energy of each secondary structure, then calculates the individual probability of each structure formation and the total RBS exposure probability, establishing a steady-state kinetic model. Finally, it uses ordinary differential equations to calculate the probability of mRNA binding to ribosomes (translation efficiency). The tools mentioned above are all based on thermodynamic data to predict the intensity of RBS. The emergence of machine learning and deep learning has provided new methods for predicting RBS intensity and for de novo design of RBS elements. Zhang et al. developed a deep learning framework called TITER using high-throughput sequencing data from HEK293 cells. In this framework, the authors combined deep convolutional and recurrent neural network algorithms to effectively and robustly capture the sequence features of translation initiation, integrating the RBS and its surrounding sequence context into a unified framework. Extensive validation tests show that TITER performs well in detecting translation initiation rates (Spilman correlation coefficient R = 0.234). Furthermore, TITER can successfully identify important sequence motifs for different TIS codons, including the Kozak sequence-like motif for the AUG codon, demonstrating that TITER is a powerful tool for predicting sequence features of translation initiation and identifying potential TIS. Machine learning algorithms also hold great potential in the RBS design process.

[0005] In summary, artificial intelligence technology has demonstrated significant advantages and potential in the prediction and de novo design of key regulatory elements. However, compared with traditional model strains such as Escherichia coli and yeast, research on key regulatory elements of Bacillus subtilis is relatively limited and its development is relatively lagging. Related studies have not established a large number of relationships between RBS sequences and intensities, and no research on RBS regulatory elements of Bacillus subtilis based on artificial intelligence has been reported. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method for predicting the RBS intensity of Bacillus subtilis based on machine learning.

[0007] The technical solution of this invention is summarized as follows:

[0008] A method for predicting the RBS intensity of Bacillus subtilis based on machine learning, comprising:

[0009] Constructing a Bacillus subtilis RBS library;

[0010] The RBS sequence in the Bacillus subtilis RBS library is encoded based on the encoder to extract the dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS sequence.

[0011] The physicochemical properties, RBS sequence characteristics, and RBS sequence secondary structure characteristics of the RBS dinucleotide are fused together to predict the strength of the RBS sequence.

[0012] Optionally, the construction of the Bacillus subtilis RBS library includes:

[0013] The experimental validation data were obtained, and green fluorescent protein was used as the reporter gene. Several Bacillus subtilis RBS sequences with fluorescence intensity values ​​greater than the set fluorescence intensity threshold were selected to form a Bacillus subtilis RBS library.

[0014] Optionally, the encoding of the RBS sequence in the Bacillus subtilis RBS library based on the encoder to extract the dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS sequence includes:

[0015] Based on the encoder, two adjacent bases of the RBS sequence in the Bacillus subtilis RBS library are encoded as a unit to extract the dinucleotide physicochemical properties of the RBS sequence.

[0016] Based on the encoder, three consecutive bases of the RBS sequence in the Bacillus subtilis RBS library are encoded as a unit to extract the RBS sequence features.

[0017] Based on the encoder, secondary structure prediction is performed on the RBS sequences in the Bacillus subtilis RBS library to extract the secondary structure features of the RBS sequences.

[0018] Optionally, the encoding of the RBS sequence in the Bacillus subtilis RBS library based on the encoder to extract the dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS sequence includes:

[0019] Based on the encoder, the dinucleotides of the RBS sequence in the Bacillus subtilis RBS library are encoded to extract the physicochemical characteristics of the dinucleotides of the RBS sequence;

[0020] Based on the encoder, the trinucleotides of the RBS sequence in the Bacillus subtilis RBS library are encoded to extract the RBS sequence features.

[0021] Based on the encoder, secondary structure prediction is performed on the RBS sequences in the Bacillus subtilis RBS library to extract the secondary structure features of the RBS sequences.

[0022] Optionally, the physicochemical properties of the dinucleotide are represented by a first encoding matrix, where each row of the first encoding matrix represents a physicochemical property of the dinucleotide, and each column of the first encoding matrix represents an encoding vector of the base combination that satisfies the physicochemical property of the dinucleotide in each row.

[0023] Optionally, the RBS sequence features are represented by a second coding matrix, where each row of the second coding matrix represents a trinucleotide base combination, and each column of the second coding matrix represents the coding vector of each row of trinucleotide base combinations.

[0024] Optionally, the secondary structure features of the RBS sequence are represented by a folded structure based on nucleotide pairing, wherein “0” is used to represent an unfolded paired base, and “-1” and “1” are used to represent complementary paired bases.

[0025] Optionally, the folded structure is a dotted bracket secondary structure, wherein “.” is used to represent an unfolded paired base, and (” and “)” are used to represent complementary paired bases. By replacing “.” with “0”, (” with “-1”, and “)” with “1” in the dotted bracket secondary structure, “0” is used to represent an unfolded paired base, and “-1” and “1” are used to represent complementary paired bases.

[0026] Optionally, the encoding of the RBS sequence in the Bacillus subtilis RBS library based on the encoder to extract the dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS sequence includes:

[0027] The RBS sequence in the Bacillus subtilis RBS library is encoded using a dimer encoder to extract the dinucleotide physicochemical properties of the RBS sequence.

[0028] Optionally, the encoding of the RBS sequence in the Bacillus subtilis RBS library based on the encoder to extract the dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS sequence includes:

[0029] The RBS sequences in the Bacillus subtilis RBS library are encoded using a triplet encoder to extract the RBS sequence features.

[0030] Optionally, the encoding of the RBS sequence in the Bacillus subtilis RBS library based on the encoder to extract the dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS sequence includes:

[0031] Based on the structure encoder, the RBS sequences in the Bacillus subtilis RBS library are encoded to extract the secondary structure features of the RBS sequences.

[0032] Optionally, the fusion of the physicochemical properties, RBS sequence characteristics, and RBS sequence secondary structure characteristics of the RBS dinucleotide to predict the strength of the RBS sequence includes:

[0033] The fusion feature is obtained by fusing the physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS dinucleotide.

[0034] The fused features are input into the trained sequence intensity prediction model for feature extraction to obtain a feature map, which is then used to predict the intensity of the RBS sequence.

[0035] Optionally, the fusion of the physicochemical properties, RBS sequence characteristics, and RBS sequence secondary structure characteristics of the RBS dinucleotide to predict the strength of the RBS sequence includes:

[0036] The physicochemical properties, RBS sequence characteristics, and RBS sequence secondary structure characteristics of the RBS dinucleotide are spliced ​​together to obtain the fusion characteristics.

[0037] Optionally, the step of inputting the fused features into a trained sequence intensity prediction model for feature extraction to obtain a feature map, and then predicting the intensity of the RBS sequence based on the feature map, includes:

[0038] The fused features are input into the trained sequence intensity prediction model for convolution processing to obtain the convolution output features;

[0039] The convolutional output features are pooled to obtain pooled output features, which are then used to generate an intensity prediction feature map.

[0040] The intensity of the RBS sequence is predicted based on the intensity prediction feature map.

[0041] Optionally, the method further includes: training the target sequence intensity prediction model according to the following steps to obtain the trained sequence intensity prediction model:

[0042] Based on the encoder, the Bacillus subtilis RBS sequence sample is encoded to extract the dinucleotide physicochemical property feature sample, RBS sequence feature sample, and RBS sequence secondary structure feature sample of the RBS sequence;

[0043] The physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS dinucleotide were fused to obtain a fused feature sample.

[0044] The fused feature samples are input into the target sequence intensity prediction model for convolution processing to obtain convolution output feature samples;

[0045] The convolutional output feature samples are pooled to obtain pooled output feature samples, which are then used to generate intensity prediction feature map samples.

[0046] Based on the intensity prediction feature map samples, calculate the predicted intensity of the Bacillus subtilis RBS sequence samples;

[0047] Based on the predicted intensity of the Bacillus subtilis RBS sequence sample and its corresponding intensity label, the target sequence intensity prediction model is optimized until a trained sequence intensity prediction model is obtained.

[0048] Advantages of this invention:

[0049] Resonant stem cells (RBSs) are crucial components regulating translation intensity, and predicting RBS intensity is of great significance for drug development and bioreactor efficiency prediction. The Bacillus subtilis RBS prediction model provided in this invention enables rapid, accurate, and low-cost prediction of RBS expression intensity.

[0050] The model provided by this invention is trained using an artificially constructed Bacillus subtilis RBS library. It comprehensively considers factors such as RBS sequence characteristics, dinucleotide physicochemical properties, and secondary structure to predict the intensity of Bacillus subtilis RBS. The model first encodes the input RBS sequence using different encoding methods, then transforms the RBS sequence into a matrix. Convolutional calculations are then used to extract different features from the input, and more complex features are further extracted to ultimately predict the intensity of the RBS sequence. Attached Figure Description

[0051] Figure 1 This is a schematic diagram of the method for predicting the RBS intensity of Bacillus subtilis based on machine learning, as described in this application.

[0052] Figure 2 This is a schematic diagram illustrating the interaction between the encoder and the sequence strength prediction model used in the embodiments of this application. Detailed Implementation

[0053] The specific embodiments of the present invention will be described in detail below with reference to specific examples and accompanying drawings.

[0054] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0055] Figure 1 This is a schematic diagram of the method for predicting the RBS intensity of Bacillus subtilis based on machine learning, as described in this application. Figure 2 This is a schematic diagram illustrating the interaction between the encoder and the sequence strength prediction model used in the embodiments of this application. Figure 1 , 2 As shown, it includes:

[0056] Constructing a Bacillus subtilis RBS library;

[0057] The RBS sequence in the Bacillus subtilis RBS library is encoded based on the encoder to extract the dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS sequence.

[0058] The physicochemical properties, RBS sequence characteristics, and RBS sequence secondary structure characteristics of the RBS dinucleotide are fused together to predict the strength of the RBS sequence.

[0059] Optionally, the construction of the Bacillus subtilis RBS library includes:

[0060] The experimental validation data were obtained, and green fluorescent protein was used as the reporter gene. Several Bacillus subtilis RBS sequences with fluorescence intensity values ​​greater than the set fluorescence intensity threshold were selected to form a Bacillus subtilis RBS library.

[0061] The aforementioned green fluorescent protein is also referred to as GFP. The number of Bacillus subtilis RBS sequences is determined based on the application scenario; for example, in one scenario, it might be 232 sequences. The fluorescence intensity threshold is also determined based on the application scenario.

[0062] Optionally, the encoding of the RBS sequence in the Bacillus subtilis RBS library based on the encoder to extract the dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS sequence includes:

[0063] Based on the encoder, two adjacent bases of the RBS sequence in the Bacillus subtilis RBS library are encoded as a unit to extract the dinucleotide physicochemical properties of the RBS sequence.

[0064] Based on the encoder, three consecutive bases of the RBS sequence in the Bacillus subtilis RBS library are encoded as a unit to extract the RBS sequence features.

[0065] Based on the encoder, secondary structure prediction is performed on the RBS sequences in the Bacillus subtilis RBS library to extract the secondary structure features of the RBS sequences.

[0066] The encoder described above can encode in various ways, including but not limited to one-hot encoding, binary encoding, and integer encoding. Different encoding methods can reflect different characteristics of the RBS sequence, including but not limited to dinucleotide physicochemical properties, RBS sequence characteristics, and RBS sequence secondary structure characteristics.

[0067] Therefore, when implementing the above encoding, the encoder can be any encoding method that can extract the physicochemical properties of dinucleotides, RBS sequence features, and RBS sequence secondary structure features.

[0068] In one application scenario, considering that two adjacent bases can form a dinucleotide and three consecutive bases can form a trinucleotide, the encoder is used to encode the RBS sequence in the Bacillus subtilis RBS library to extract the dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS sequence, including:

[0069] Based on the encoder, the dinucleotides of the RBS sequence in the Bacillus subtilis RBS library are encoded to extract the physicochemical characteristics of the dinucleotides of the RBS sequence;

[0070] Based on the encoder, the trinucleotides of the RBS sequence in the Bacillus subtilis RBS library are encoded to extract the RBS sequence features of the RBS sequence;

[0071] Based on the encoder, secondary structure prediction is performed on the RBS sequences in the Bacillus subtilis RBS library to extract the secondary structure features of the RBS sequences.

[0072] Optionally, the physicochemical properties of the dinucleotide are represented by a first encoding matrix, where each row of the first encoding matrix represents a physicochemical property of the dinucleotide, and each column of the first encoding matrix represents an encoding vector of the base combination that satisfies the physicochemical property of the dinucleotide in each row.

[0073] Optionally, the RBS sequence features are represented by a second coding matrix, where each row of the second coding matrix represents a trinucleotide base combination, and each column of the second coding matrix represents the coding vector of each row of trinucleotide base combinations.

[0074] Optionally, the secondary structure features of the RBS sequence are represented by a folded structure based on nucleotide pairing, wherein “0” is used to represent an unfolded paired base, and “-1” and “1” are used to represent complementary paired bases.

[0075] Optionally, the folded structure is a dotted bracket secondary structure, wherein “.” is used to represent an unfolded paired base, and (” and “)” are used to represent complementary paired bases. By replacing “.” with “0”, (” with “-1”, and “)” with “1” in the dotted bracket secondary structure, “0” is used to represent an unfolded paired base, and “-1” and “1” are used to represent complementary paired bases.

[0076] In one scenario, a combination of three encoding methods is used to encode the RBS sequence using dimer, triplet, and structure encoding to extract the dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features.

[0077] Therefore, the encoder-based encoding of the RBS sequences in the Bacillus subtilis RBS library is used to extract the dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS sequences, including:

[0078] Based on the dimer encoder, the RBS sequence in the Bacillus subtilis RBS library is dimer encoded to extract the dinucleotide physicochemical properties of the RBS sequence.

[0079] In a specific application scenario, dimer encoding involves encoding two adjacent bases (forming a dinucleotide) in an RBS sequence as a single unit. For example, a 30bp RBS sequence can be dimer encoded. Utilizing the 11 physicochemical properties of dinucleotides, dimer encoding can encode a 30bp RBS sequence into an 11*29 matrix. This matrix, called the dimer matrix (the first encoding matrix mentioned above), represents one of the 11 physicochemical properties of the dinucleotide in the RBS sequence. The following is a program code example for implementing the dimer encoding:

[0080]

[0081]

[0082]

[0083] Optionally, the encoding of the RBS sequence in the Bacillus subtilis RBS library based on the encoder to extract the dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS sequence includes:

[0084] Based on the triplet encoder, the RBS sequences in the Bacillus subtilis RBS library are triplet encoded to extract the RBS sequence features.

[0085] In a specific application scenario, triplet encoding is a method of encoding three consecutive bases (comprising a trinucleotide) in an RBS sequence as a unit. For example, a 30bp RBS sequence can be triplet encoded, resulting in each three bases corresponding to a 64 (4) nucleotide sequence. 3 The vector representation of ) is as follows: 'AAA' is encoded as [1,0,...,0], 'AAT' is encoded as [0,1,...,0], and so on. Through triplet encoding, the RBS sequence vectors of length 30bp can be used to form a 64*28 triplet matrix, which is the second encoding matrix mentioned above. Each row vector in the matrix represents the base combination of the RBS sequence. For example, [1,0,...,0] represents 'AAA', [0,1,...,0] represents 'AAT', and so on, with a total of 64 base combinations.

[0086] The following is an example of program code that implements triplet encoding:

[0087]

[0088] Optionally, the encoding of the RBS sequence in the Bacillus subtilis RBS library based on the encoder to extract the dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS sequence includes:

[0089] Based on the structure encoder, the RBS sequences in the Bacillus subtilis RBS library are structure encoded to extract the secondary structure features of the RBS sequences.

[0090] Structure encoding uses dotted brackets to represent secondary structures in RNA or DNA sequences. RNAfold (http: / / rna.tbi.univie.ac.at / cgi-bin / RNAWebSuite / RNAfold.cgi) predicts the secondary structures of all RBS sequences in a library, directly outputting the dotted bracket secondary structures of all RBS sequences, thus encoding the structure of a 30bp RBS sequence. For example, if the software predicts that the secondary structure of the sequence "AAAATCATTCAATAAAAGGGGAGCTTACCA" is such that nucleotides 18 and 19 pair with nucleotides 28 and 29 to form a folded structure, then its secondary structure can be represented using ".................((........)).".

[0091] The following is an example program code that implements the above structure encoding:

[0092]

[0093] Optionally, the fusion of the physicochemical properties, RBS sequence characteristics, and RBS sequence secondary structure characteristics of the RBS dinucleotide to predict the strength of the RBS sequence includes:

[0094] The fusion feature is obtained by fusing the physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS dinucleotide.

[0095] The fused features are input into the trained sequence intensity prediction model for feature extraction to obtain a feature map, which is then used to predict the intensity of the RBS sequence.

[0096] For example, in a specific scenario, the physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS dinucleotide can be directly spliced ​​together to obtain a fusion feature. This splicing can be done, for example, by directly splicing them horizontally using the torch.cat function. Furthermore, the physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS dinucleotide are represented in a matrix manner. Therefore, the fusion feature is specifically a fusion matrix.

[0097] See above Figure 2 The RBS sequence is triplet encoded. As mentioned earlier, triplet encoding encodes three bases as a group, resulting in a sequence with a length of 64 (4... 3 The vector representation of 'AAA' is used, for example, 'AAA' would be encoded as [1,0,...,0], 'AAT' would be encoded as [0,1,...,0], and so on. Therefore, using triplet encoding, a 30bp RBS sequence can be encoded into a 64*28 matrix. In this matrix, the 64 row vectors represent different base combinations, and the 28 column vectors represent the 30bp sequence, allowing for the reading of 28 adjacent triplet combinations. Therefore, this matrix contains the sequence information of the RBS.

[0098] In this application, considering that in RNA research, specific parts or groups of atoms in the nucleotide sequence may affect the spectral properties of the molecule, unlike dinucleotide sequences which have 11 physicochemical properties such as translation, hydrophilicity, glide, elevation distance, tilt, rolling, torsion, stacking, enthalpy, entropy, and free energy, a "dimer" encoding is performed on the 30bp RBS sequence. Different combinations of dinucleotide bases are represented by vectors of these 11 physicochemical properties. Through dimer encoding, the 30bp RBS sequence can be encoded into an 11*29 matrix, which can be called a dimer matrix. Each row of the matrix represents one of the 11 physicochemical properties of the dinucleotide in the RBS sequence, and the 29 columns represent the 29 adjacent dinucleotide combinations that can be read from the 30bp sequence. Therefore, this matrix contains information on the physicochemical properties of the RBS dinucleotide.

[0099] The RBS sequences are encoded using "structure," with "0" representing unfolded base pairs and "-1" and "1" representing complementary base pairs. RNAfold (http: / / rna.tbi.univie.ac.at / cgi-bin / RNAWebSuite / RNAfold.cgi) is used to predict the secondary structure of all RBS sequences in the library. The software directly outputs the punctuated secondary structure of all RBS sequences in the library, using "." to represent unfolded base pairs and "(" and ")" to represent complementary base pairs. Then, by replacing "." with "0," "(" with "-1," and ")" with "1" in the punctuated secondary structure, the "structure" encoding of the RBS sequences is obtained. This encoding encodes the RBS sequences into a 1*30 row vector, where one row represents one RBS sequence and 30 columns represent whether the bases at each position are folded or unfolded. Therefore, this encoding method contains secondary structure information of the RBS sequences.

[0100] Finally, the three codes are directly concatenated horizontally using the torch.cat function to obtain a 64*87 fusion matrix. The first 28 columns of the fusion matrix are the "triplet" codes of RBS, representing RBS sequence features; the middle 29 columns are the "dimer" codes of RBS, representing the dinucleotide physicochemical properties of the RBS sequence; and the last 30 columns are the "structure" codes of RBS, representing the secondary structure features of the RBS sequence. This concatenated fusion matrix is ​​used as input to the trained sequence intensity prediction model for feature extraction.

[0101] Optionally, the fusion of the physicochemical properties, RBS sequence characteristics, and RBS sequence secondary structure characteristics of the RBS dinucleotide to predict the strength of the RBS sequence includes:

[0102] The physicochemical properties, RBS sequence characteristics, and RBS sequence secondary structure characteristics of the RBS dinucleotide are spliced ​​together to obtain the fusion characteristics.

[0103] Optionally, the step of inputting the fused features into a trained sequence intensity prediction model for feature extraction to obtain a feature map, and then predicting the intensity of the RBS sequence based on the feature map, includes:

[0104] The fused features are input into the trained sequence intensity prediction model for convolution processing to obtain the convolution output features;

[0105] The convolutional output features are pooled to obtain pooled output features, which are then used to generate an intensity prediction feature map.

[0106] The intensity of the RBS sequence is predicted based on the intensity prediction feature map.

[0107] Optionally, the method further includes: training the target sequence intensity prediction model according to the following steps to obtain the trained sequence intensity prediction model:

[0108] Based on the encoder, the Bacillus subtilis RBS sequence sample is encoded to extract the dinucleotide physicochemical property feature sample, RBS sequence feature sample, and RBS sequence secondary structure feature sample of the RBS sequence;

[0109] The physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS dinucleotide were fused to obtain a fused feature sample.

[0110] The fused feature samples are input into the target sequence intensity prediction model for convolution processing to obtain convolution output feature samples;

[0111] The convolutional output feature samples are pooled to obtain pooled output feature samples, which are then used to generate intensity prediction feature map samples.

[0112] Based on the intensity prediction feature map samples, calculate the predicted intensity of the Bacillus subtilis RBS sequence samples;

[0113] Based on the predicted intensity of the Bacillus subtilis RBS sequence sample and its corresponding intensity label, the target sequence intensity prediction model is optimized until a trained sequence intensity prediction model is obtained.

[0114] The training set includes several Bacillus subtilis RBS sequence samples. The Bacillus subtilis RBS sequence samples are encoded (e.g., the three encoding methods mentioned above). The fused feature samples are obtained according to the training steps described above, and then input into the target sequence intensity prediction model for convolution, pooling, and other processing to obtain the predicted intensity of the Bacillus subtilis RBS sequence samples. The target sequence intensity prediction model is further optimized based on the predicted intensity of the Bacillus subtilis RBS sequence samples and their corresponding intensity labels until a trained sequence intensity prediction model is obtained.

[0115] Based on the predicted intensity of the Bacillus subtilis RBS sequence samples and their corresponding intensity labels, the target sequence intensity prediction model is optimized until a fully trained sequence intensity prediction model is obtained. The difference between the model's predicted intensity and its corresponding intensity label is calculated using a loss function (root mean square error). Next, the gradient is calculated using the backpropagation algorithm to update the model parameters and reduce the loss function value. This process is repeated multiple times, with each iteration including the aforementioned forward propagation (including encoding the Bacillus subtilis RBS sequence samples (e.g., the three encoding methods mentioned above), obtaining fused feature samples according to the training steps, and then inputting them into the target sequence intensity prediction model for convolution, pooling, etc., to obtain the predicted intensity of the Bacillus subtilis RBS sequence samples), loss function calculation, backpropagation, and parameter update steps. In one scenario, the model is trained after 70 iterations. Then, the model is evaluated using a test set (the evaluation process is similar to the training process, except that it does not include calculating the loss function and updating model parameters), and the mean absolute error, Pearson correlation index, and variance of the model on the test set are calculated. The prediction model obtained after training achieved a PCCs result of 0.8601 on the test set, with R... 2 The value of 0.57 indicates that the model performs well and there is a strong linear correlation between the predicted intensity and the actual intensity (i.e., the intensity label).

[0116] For the sequence strength prediction model (including but not limited to CNN) trained as described above, we take the target RBS sequence "AGTCTAGTCGATGCTAGCTGCTAGCTAGCT" as an example to predict its strength. To predict the strength of other RBS sequences, simply replace the sequence after "seqs=" in the following code with other predicted RBS sequences.

[0117] After obtaining the trained sequence intensity prediction model (e.g., named RBS-232-MobileNet.pth), run the following code (LoadModel.py) to predict the intensity of the target Bacillus subtilis RBS sequence “AGTCTAGTCGATGCTAGCTGCTAGCTAGCT”:

[0118]

[0119] )#load the parameters from the file

[0120] model.eval()#set the model to evaluation mode

[0121] if model_path.find('triplet')!=-1:

[0122] encoder='triplet'

[0123] elif model_path.find('dimer-1')!=-1:

[0124] encoder='dimer-1'

[0125] elif model_path.find('dimer')!=-1:

[0126] encoder='dimer'

[0127] elif model_path.find('combined')!=-1:

[0128] encoder='combined'

[0129] result=[]

[0130] print(encoder)

[0131] if model_path.find('Ensemble')==-1:

[0132] encoded_sequence=encoding([sequence],encode_type=encoder)#encodethe sequence into a tensor

[0133] intensity=model.forward(encoded_sequence)#pass the tensor to themodel and get the output

[0134] else:

[0135] encoded_sequence1=encoding([sequence],encode_type='dimer')

[0136] encoded_sequence2=encoding([sequence],encode_type='triplet')

[0137] intensity=model.forward(encoded_sequence1,encoded_sequence2)

[0138] result.append(intensity.item())

[0139] print(result)

[0140] The final output prediction result, i.e. the prediction strength, is 1.287, which means that the predicted RBS strength of the next sequence is 1.287.

[0141] The above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and are not intended to limit it. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for predicting the RBS intensity of Bacillus subtilis based on machine learning, characterized in that, include: Constructing a Bacillus subtilis RBS library; The RBS sequence in the Bacillus subtilis RBS library is encoded based on the encoder to extract the dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS sequence. The physicochemical properties, RBS sequence characteristics, and RBS sequence secondary structure characteristics of the RBS dinucleotide are fused together to predict the strength of the RBS sequence. The step of encoding the RBS sequence in the Bacillus subtilis RBS library based on the encoder to extract the dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS sequence includes: Based on the encoder, two adjacent bases of the RBS sequence in the Bacillus subtilis RBS library are encoded as a unit to extract the dinucleotide physicochemical properties of the RBS sequence. Based on the encoder, three consecutive bases of the RBS sequence in the Bacillus subtilis RBS library are encoded as a unit to extract the RBS sequence features. Based on the encoder, secondary structure prediction is performed on the RBS sequences in the Bacillus subtilis RBS library to extract the secondary structure features of the RBS sequences.

2. The method for predicting Bacillus subtilis RBS intensity based on machine learning according to claim 1, characterized in that, The construction of the Bacillus subtilis RBS library includes: The experimental validation data were obtained, and green fluorescent protein was used as the reporter gene. Several Bacillus subtilis RBS sequences with fluorescence intensity values ​​greater than the set fluorescence intensity threshold were selected to form a Bacillus subtilis RBS library.

3. The method for predicting Bacillus subtilis RBS intensity based on machine learning according to claim 1, characterized in that, The encoder-based encoding of the RBS sequences in the Bacillus subtilis RBS library extracts dinucleotide physicochemical properties, RBS sequence features, and RBS sequence secondary structure features, including: Based on the encoder, the dinucleotides of the RBS sequence in the Bacillus subtilis RBS library are encoded to extract the physicochemical characteristics of the dinucleotides of the RBS sequence; Based on the encoder, the trinucleotides of the RBS sequence in the Bacillus subtilis RBS library are encoded to extract the RBS sequence features. Based on the encoder, secondary structure prediction is performed on the RBS sequences in the Bacillus subtilis RBS library to extract the secondary structure features of the RBS sequences.

4. The method for predicting Bacillus subtilis RBS intensity based on machine learning according to claim 3, characterized in that, The physicochemical properties of the dinucleotides are represented by a first coding matrix, where each row of the first coding matrix represents a physicochemical property of a dinucleotide, and each column of the first coding matrix represents a coding vector of the base combination that satisfies the physicochemical property of the dinucleotide in each row. The RBS sequence features are represented by a second coding matrix, where each row of the second coding matrix represents a trinucleotide base combination, and each column of the second coding matrix represents a coding vector of the trinucleotide base combination in each row. The secondary structure features of the RBS sequence are represented by a folded structure based on nucleotide pairing, where "0" represents an unfolded paired base, and "-1" and "1" represent complementary paired bases.

5. The method for predicting Bacillus subtilis RBS intensity based on machine learning according to claim 4, characterized in that, The folded structure is a dotted bracket secondary structure, in which "." is used to represent unfolded paired bases, and "(" and "")" represent complementary paired bases. By replacing "." with "0", "(" with "-1", and "")" with "1" in the dotted bracket secondary structure, "0" is used to represent unfolded paired bases, and "-1" and "1" represent complementary paired bases.

6. The method for predicting Bacillus subtilis RBS intensity based on machine learning according to claim 1, characterized in that, The method of fusing the physicochemical properties, RBS sequence characteristics, and RBS sequence secondary structure characteristics of the RBS dinucleotide to predict the strength of the RBS sequence includes: The fusion feature is obtained by fusing the physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS dinucleotide. The fused features are input into the trained sequence intensity prediction model for feature extraction to obtain a feature map, which is then used to predict the intensity of the RBS sequence.

7. The method for predicting Bacillus subtilis RBS intensity based on machine learning according to claim 6, characterized in that, The method of fusing the physicochemical properties, RBS sequence characteristics, and RBS sequence secondary structure characteristics of the RBS dinucleotide to predict the strength of the RBS sequence includes: The physicochemical properties, RBS sequence characteristics, and RBS sequence secondary structure characteristics of the RBS dinucleotide are spliced ​​together to obtain the fusion characteristics.

8. The method for predicting Bacillus subtilis RBS intensity based on machine learning according to claim 7, characterized in that, The step of inputting the fused features into a trained sequence strength prediction model for feature extraction to obtain a feature map, and then predicting the strength of the RBS sequence based on the feature map, includes: The fused features are input into the trained sequence intensity prediction model for convolution processing to obtain the convolution output features; The convolutional output features are pooled to obtain pooled output features, which are then used to generate an intensity prediction feature map. The intensity of the RBS sequence is predicted based on the intensity prediction feature map.

9. The method for predicting Bacillus subtilis RBS intensity based on machine learning according to claim 8, characterized in that, The method further includes: training the target sequence intensity prediction model according to the following steps to obtain the trained sequence intensity prediction model: Based on the encoder, the Bacillus subtilis RBS sequence sample is encoded to extract the dinucleotide physicochemical property feature sample, RBS sequence feature sample, and RBS sequence secondary structure feature sample of the RBS sequence; The physicochemical properties, RBS sequence features, and RBS sequence secondary structure features of the RBS dinucleotide were fused to obtain a fused feature sample. The fused feature samples are input into the target sequence intensity prediction model for convolution processing to obtain convolution output feature samples; The convolutional output feature samples are pooled to obtain pooled output feature samples, which are then used to generate intensity prediction feature map samples. Based on the intensity prediction feature map samples, calculate the predicted intensity of the Bacillus subtilis RBS sequence samples; Based on the predicted intensity of the Bacillus subtilis RBS sequence sample and its corresponding intensity label, the target sequence intensity prediction model is optimized until a trained sequence intensity prediction model is obtained.