A knowledge-driven and data-driven RNA secondary structure generator and prediction method for the AVR gene of rice blast fungus
By integrating knowledge-driven and data-driven rice blast fungus AVR gene RNA secondary structure generator and prediction methods, and utilizing multiple chips and deep neural network models, the problem of the lack of multi-omics information fusion in existing RNA secondary structure prediction methods was solved, thus achieving precise regulation and efficient prediction of rice blast disease.
Patent Information
- Application Number
- CN202411648509.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-19
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-11-19
AI Technical Summary
At present, the RNA secondary structure prediction methods lack a multi-omics computational model that effectively integrates structural and omics information, resulting in an incomplete analysis of the complex regulatory process of the rice blast fungus AVR gene.
A knowledge-driven and data-driven RNA secondary structure generator for the AVR gene of rice blast fungus was designed. By combining the use of multiple chips and constructing a deep neural network model based on the encoder-decoder architecture, a VAE-GAN hybrid model was introduced for evaluation and optimization, and a multi-source transfer learning strategy was adopted to achieve accurate prediction of RNA secondary structure.
It greatly accelerated the execution speed of the generator, improved the prediction accuracy and computational efficiency, achieved precise control of rice blast disease, and enhanced the practicality and computational efficiency of the model.
Smart Images

Figure CN119626339B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics and artificial intelligence intersection technology, and more specifically, to a knowledge-driven and data-driven integrated generator and prediction method for the rice blast fungus AVR gene RNA secondary structure. Background Art
[0002] Rice blast, a fungal disease caused by the blast fungus, is one of the three major diseases affecting global rice production. Statistics show that it causes annual yield losses of 10% to 30%, with yield reductions reaching 50% in epidemic years. Research has shown that rice and the blast fungus share a gene-for-gene relationship in resistance, with the avirulence gene (AVR gene) being key to determining a plant's specific resistance to the pathogen. As a typical model fungal pathogen, in-depth analysis of the biological characteristics and infection mechanisms of the blast fungus has important scientific and economic implications for the prevention and control of this plant disease and for safeguarding global food security.
[0003] Given the unique evolutionary characteristics of rice blast fungus avirulence genes, it is speculated that their RNA secondary structure may also play an important regulatory role. On the one hand, avirulence genes generally evolve rapidly, have multiple copies, and have functional redundancy, and RNA secondary structure may facilitate their rapid mutation and acquisition of new functions. On the other hand, previous studies have shown that some avirulence genes rely on epigenetic regulation mediated by non-coding RNAs, and RNA secondary structure may be involved in regulating the maturation and function of these non-coding RNAs. Furthermore, as a key factor connecting rice blast infection and rice immunity, variations in the RNA structure of avirulence genes may directly affect host recognition and defense responses. Therefore, RNA secondary structure is likely a key aspect of the evolution and regulation of avirulence genes.
[0004] RNA secondary structure prediction technology has shown significant application prospects in fungal research and crop disease control. In fungal research, by predicting RNA secondary structure, researchers can gain a deeper understanding of its biological characteristics and functions, providing new ideas for the development of antifungal drugs. At the same time, in the field of crop disease control, this technology is widely used in RNAi strategies, which achieve efficient and specific disease control by designing dsRNA targeting specific RNA structures. Furthermore, RNA secondary structure prediction can help monitor and warn of crop disease outbreaks, as well as improve crop resistance to disease through disease-resistant breeding. As the technology continues to improve and develop, RNA secondary structure prediction will play an even more important role in fungal research and crop disease control, providing strong support for agricultural production and food security.
[0005] However, the current mainstream RNA secondary structure prediction methods still have limitations, and there is a lack of multi-omics computational models that effectively integrate structural and omics information. These shortcomings restrict the comprehensive analysis of the complex regulatory processes of pathogens. Summary of the Invention
[0006] The purpose of the present invention is to design and develop a knowledge-driven and data-driven generator for the RNA secondary structure of the rice blast fungus AVR gene, which accelerates the execution speed of the generator by using a combination of multiple chips.
[0007] The present invention also designs and develops a prediction method for the secondary structure of the rice blast fungus AVR gene RNA that integrates knowledge-driven and data-driven methods. The integration of two drivers and multiple learning models improves practicality and computational efficiency, and achieves precise regulation of rice blast disease.
[0008] The technical solution provided by the present invention is:
[0009] A knowledge-driven and data-driven RNA secondary structure generator for the rice blast fungus AVR gene, including:
[0010] A storage module, connected to the PC, for storing RNA raw data and intermediate results; and
[0011] A feature extraction module, connected to the storage module, for extracting sequence features of the RNA raw data;
[0012] a structure generation module, connected to the feature extraction module, for predicting the graph structure of the secondary structure of the rice blast fungus AVR gene RNA;
[0013] The evaluation and optimization module is connected to the structure generation module and is used to optimize the graph structure of the secondary structure of the rice blast fungus AVR gene RNA.
[0014] Preferably, the storage module includes:
[0015] An ARM chip connected to the PC;
[0016] A DRAM chip connected to the ARM chip;
[0017] The feature extraction module includes:
[0018] A first FPGA chip, connected to the ARM chip;
[0019] The structure generation module includes:
[0020] a first AI chip connected to the first FPGA chip;
[0021] a second FPGA chip, connected to the first AI chip;
[0022] The evaluation and optimization module includes:
[0023] a second AI chip connected to the second FPGA chip;
[0024] The third FPGA chip is connected to the second AI chip and is used to output the optimized secondary structure of the rice blast fungus AVR gene RNA.
[0025] Preferably, it also includes:
[0026] A first interface unit, which is provided on the storage module and is used to connect to a PC terminal;
[0027] a second interface unit, which is provided on the feature extraction module and is used to connect with the structure generation module;
[0028] The third interface unit is provided on the structure generation module and is used to connect with the evaluation and optimization module.
[0029] Preferably, the first interface unit includes:
[0030] A USB interface is provided on the ARM chip and is used to connect to a PC;
[0031] A JTAG interface is provided on the ARM chip and is used to debug the ARM chip;
[0032] An Ethernet interface, which is provided on the ARM chip and is used to connect to a PC;
[0033] An RS232 interface is provided on the ARM chip and is used to connect to the DRAM chip;
[0034] The second interface unit includes:
[0035] A first data line interface, provided on the first FPGA chip and configured to connect to the DIN0-D7 pins of the first AI chip;
[0036] A first synchronous clock interface, provided on the first FPGA chip and configured to connect to a clock pin of the first AI chip;
[0037] A first enable signal interface, provided on the first FPGA chip and configured to connect to an enable pin of the first AI chip;
[0038] The third interface unit includes:
[0039] A second data line interface, which is provided on the second FPGA chip and is used to connect to the DIN0-D7 pin of the second AI chip;
[0040] A second synchronous clock interface, which is provided on the second FPGA chip and is used to connect to the clock pin of the second AI chip;
[0041] A second enable signal interface, provided on the second FPGA chip and configured to connect to an enable pin of the second AI chip;
[0042] The second FPGA chip is connected to the first AI chip through an API interface, and the third FPGA chip is connected to the second AI chip through an API interface.
[0043] A method for predicting the secondary structure of the rice blast fungus AVR gene RNA that integrates knowledge-driven and data-driven methods, using the rice blast fungus AVR gene RNA secondary structure generator that integrates knowledge-driven and data-driven methods, comprises the following steps:
[0044] Step 1: Collect RNA secondary structure data, RNA sequence data, thermodynamic characteristic data, kinetic characteristic data, and RNA whole genome sequence data of non-toxic genes and classify them according to data type;
[0045] Step 2: construct a deep neural network model based on an encoder-decoder architecture, and input the classified RNA sequence data, thermodynamic characteristic data, and kinetic characteristic data into the deep neural network model to obtain an RNA secondary structure pairing matrix;
[0046] Step 3: Construct a GAN model and input the RNA secondary structure pairing matrix into the GAN model to obtain an optimized secondary structure.
[0047] Preferably, the step one further includes: cleaning and filtering the classified data.
[0048] Preferably, the construction of the encoder comprises the following steps:
[0049] Step a: constructing a graphical representation of the RNA structure G = {V, E, X} based on the classified RNA sequence data;
[0050] Among them, V is composed of RNA base nodes, E represents the reference edge between base nodes, and X represents the reference edge between each node v i The associated content features, N represents the length of the RNA sequence;
[0051] Step b, pre-processing the thermodynamic characteristic data and the kinetic characteristic data to use them as edges referenced between base nodes;
[0052] Step c: Build a graph convolutional network model;
[0053] The graph convolutional network model includes an input layer, a first graph convolutional layer, a first activation layer, a first autoregressive layer, a second graph convolutional layer, a dropout layer, a second autoregressive layer, two fully connected layers and an output layer, which are connected in sequence.
[0054] Wherein, the input layer satisfies:
[0055]
[0056] Where, degree matrix representing the graphical representation of the RNA structure, represents the adjacency matrix of the graphical representation of the RNA structure, L represents the N × N Laplacian matrix;
[0057] The first autoregressive layer and the second autoregressive layer both satisfy:
[0058]
[0059] Where Z (n+1) is the potential code of the n+1th layer, and when n=0 it is the first autoregressive layer, when n=1 it is the second autoregressive layer, Z (0) =L,W (n) is the weight matrix of the nth layer, represents the adjacency matrix loss function with self-loops, is the adjacency matrix, represents the symmetric normalized degree matrix, φ is the activation function;
[0060] The two fully connected layers satisfy;
[0061] μ i =Z (n+1) W μ ;
[0062] σ i =exp(Z (n+1) W σ );
[0063] Where n = 1, μ i is the mean, σ i is the variance, W μ is the first trainable weight matrix, W σ is the second trainable weight matrix;
[0064] The output layer satisfies:
[0065] Z=μ i +σ i ·∈;
[0066] Where \(Z\) is the latent variable, and \(\in\) represents the noise sampled from the standard normal distribution;
[0067] The loss function of the encoder is:
[0068]
[0069] Where represents the loss function, represents the expected value of the latent variable \(Z\) given the feature matrix \(X\) and the adjacency matrix ; represents the logarithm of the probability of calculating the generated adjacency matrix given the latent variable \(Z\);
[0070] Step d: Input the graphical representation of the RNA structure into the graph convolutional network model to obtain the latent variable of the RNA molecular structure.
[0071] Preferably, the decoder consists of 20 identical structures, and each layer includes a Class A masked convolution, 20 residual modules, a first convolutional layer, a second activation layer, a second convolutional layer, a third activation layer, and an output layer connected in sequence;
[0072] Among them, the size of the Class A masked convolution is \(8\times8\), the residual module is a third convolutional layer, a Class B masked convolution, and a fourth convolutional layer connected in sequence. The convolutional kernels of the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer have a size of \(1\times1\). The size of the Class B masked convolution is \(5\times5\). The second activation layer and the third activation layer are both ReLU functions, and the output layer is a softmax function;
[0073] The loss function of the decoder satisfies:
[0074]
[0075] Where is the true RNA secondary structure diagram, \(x'\) ξ is the sample generated by the encoder, \(0 < \xi < N\), and \(\xi\) represents the position of each node in the RNA secondary structure diagram;
[0076] The input of the decoder is the latent variable of the RNA molecular structure and the true RNA secondary structure data diagram of the non-toxic gene, and the output of the decoder is the graph structure of the RNA secondary structure.
[0077] Preferably, the GAN model includes a generator and a discriminator;
[0078] The generator includes an input layer, a fully connected layer, a reshaping layer, 5 transposed convolutional module layers, and an output layer connected in sequence;
[0079] The discriminator includes an input layer, a fifth convolutional layer, a fifth activation layer, a sixth convolutional layer, a sixth activation layer, a seventh convolutional layer, a seventh activation layer, an eighth convolutional layer, an eighth activation layer, a ninth convolutional layer, a flattening layer, a fully connected layer, and a ninth activation layer, which are connected in sequence;
[0080] The input of the generator is the RNA secondary structure pairing matrix, the output of the generator is a 64x64 RGB image, the input of the discriminator is the output of the generator, and the output of the discriminator is the probability of a real sample;
[0081] The objective function of the generator is:
[0082] f=2JS9P r (c)||P g (c))-2log2;
[0083] Where, JS9P r ||p g ) is the JS divergence, p g To generate the sample distribution, P r is the distribution of the real sample;
[0084] The optimal discriminator satisfies:
[0085]
[0086] Where D * (c) is the optimal discriminator, P r (c) is the probability of the true sample, p g (c) is the probability of generating a sample.
[0087] Preferably, model compression and computational optimization are also included, and the construction thereof specifically includes:
[0088] Step 1: Build a student model:
[0089] Calculate the weights of each layer of the original VAE-GAN model, prune the layers below the threshold, and obtain the pruned student model;
[0090] Step 2: Use the original VAE-GAN model as the teacher model and obtain the similarity measurement function based on the student model:
[0091] Similarity(S1,S2)=α·Match(S1,S2)-β·ΔG;
[0092] Where S1 represents the RNA secondary structure predicted by the teacher model, S2 represents the RNA secondary structure predicted by the student model, α is the first hyperparameter, Match(S1,S2) is the base pair matching degree between the RNA secondary structures predicted by S1 and S2, β is the second hyperparameter, and ΔG represents the difference in free energy.
[0093] Step ③: Obtain the optimal student model based on the similarity measurement function.
[0094] The beneficial effects of the present invention are:
[0095] (1) The present invention designs and develops a knowledge-driven and data-driven RNA secondary structure generator for the AVR gene of rice blast fungus. The generator combines the popular AI chip and FPGA chip, which greatly accelerates the execution speed of the generator. The generator has high practical value, high integration, high computational efficiency, and can accurately predict the RNA secondary structure.
[0096] (2) The present invention designs and develops a prediction method for the secondary structure of the rice blast fungus AVR gene RNA that integrates knowledge-driven and data-driven approaches. It cleverly integrates the two paradigms, designs and uses a deep neural network model based on an encoder-decoder architecture, introduces a VAE-GAN hybrid model to evaluate and optimize the prediction results, and adopts a multi-source transfer learning strategy based on evolutionary relationships and meta-learning to improve the generalization ability of the model. Finally, through model compression and computational optimization, the practicality and computational efficiency of the method are further improved, and the deep integration of structure and multi-omics data is achieved. The relationship between RNA secondary structure and AVR gene evolution and regulation is further explored to achieve precise regulation of rice blast disease. BRIEF DESCRIPTION OF THE DRAWINGS
[0097] Figure 1 This is a schematic diagram of the structure of the rice blast fungus AVR gene RNA secondary structure generator that integrates knowledge-driven and data-driven methods according to the present invention.
[0098] Figure 2 Schematic diagram of the process of the method for predicting the secondary structure of the rice blast fungus AVR gene RNA that integrates knowledge-driven and data-driven methods described in the present invention.
[0099] Figure 3 Schematic diagram of the principle of the knowledge and data dual-driven deep neural network model described in the present invention.
[0100] Figure 4 Schematic diagram of the principle of the VAE-GAN hybrid model described in the present invention.
[0101] Figure 5 Schematic diagram of the principle of the multi-source transfer learning strategy based on evolutionary relationships and meta-learning described in the present invention.
[0102] Figure 6 Schematic diagram of the principles of model compression and computational optimization described in the present invention.
[0103] Figure 7 Schematic diagram of the principle of knowledge distillation described in the present invention. DETAILED DESCRIPTION
[0104] The present invention is described in further detail below so that those skilled in the art can implement the invention with reference to the description.
[0105] like Figure 1 As shown, the present invention provides a rice blast fungus AVR gene RNA secondary structure generator that integrates knowledge-driven and data-driven, which consists of a storage module, a feature extraction module, a structure generation module, an evaluation and optimization module, a first interface unit, a second interface unit and a third interface unit. The storage module is connected to the PC end and the feature extraction module through the first interface unit to realize the input of RNA sequence data and the storage of RNA original data and intermediate results; the feature extraction module is connected to the structure generation module through the second interface unit to realize the extraction of RNA sequence features, data transmission and preliminary prediction of the graph structure of the rice blast fungus AVR gene RNA secondary structure; the structure generation module is connected to the evaluation and optimization module through the third interface unit to receive the graph structure information of the predicted RNA secondary structure and optimize it.
[0106] The storage module includes an ARM chip and a DRAM chip, the feature extraction module is a first FPGA chip, the structure generation module includes a first AI chip and a second FPGA chip, the evaluation and optimization module includes a second AI chip and a third FPGA chip, the first interface unit includes a USB interface, a JTAG interface, an Ethernet interface and an RS232 interface, the second interface unit includes a first data line interface, a first synchronous clock interface and a first enable signal interface, and the third interface unit includes a second data line interface, a second synchronous clock interface and a second enable signal interface.
[0107] The ARM chip is connected to the PC via a USB interface to ensure fast data transmission and effective command interaction; at the same time, the ARM chip is connected to the PC via an Ethernet interface to achieve mutual communication with other devices or networks for data sharing and remote control; the ARM chip is connected to the DRAM chip via an RS232 interface to support serial data transmission. The on-chip memory of the ARM chip and the external DRAM chip are used to store RNA raw data and intermediate results, respectively; the ARM chip is debugged via the JTAG interface, which facilitates developers to perform real-time debugging, troubleshooting and performance optimization, thereby improving the reliability and stability of the system. Through the above series of connections and configurations, the system can achieve efficient data transmission and flexible debugging functions.
[0108] In this embodiment, hexadecimal format is used for data transmission to ensure the accuracy and validity of information. The data transmission adopts the format of 8 data bits and 1 stop bit, which facilitates efficient processing when storing RNA raw data and intermediate results.
[0109] The first FPGA chip (ZYNQ) mainly includes three AXI interfaces, namely AXI_GP, AXI_HP and AXI_ACP. The first FPGA chip uses the AXI_HP interface to connect to the ARM chip (Cortex-M7), and the first FPGA chip and the ARM chip use the PL-PS collaborative protocol to connect, as follows:
[0110] Define the ARM chip as the processing system end (PS), define the first FPGA chip as the programmable logic end (PL), and use a lightweight and fast data transmission method. First, create a block design on the PL end and add the corresponding functional modules. The specific configuration is as follows: In the PS and PL end configurations, open the AXI bus interface; in the input and output configuration, set the first group of input and output voltages to 1.8 volts; in the clock configuration, set the clock frequency to 100 MHz; select the lightweight AXI protocol as the protocol, set the number of block RAM interfaces to 1, select the block RAM controller as the mode, select the dual-port RAM as the memory type, select all outputs for the input and output interfaces, and set the data width to 2. After the configuration is completed, corresponding interfaces are created for the block RAM controller and the general input and output interfaces. Finally, through the automatic connection function, an efficient data transmission system is built to ensure fast and stable communication between the PS and PL ends. A new project is created on the PS end. First, a write enable signal is given to let the PL end write data to the BRAM, and then the PS end reads the data. GPIO is used for control, which actually sends a pulse enable signal to the PL end (GPIO[1] from 0->1->0). GPIO[0] is 0 for write state and 1 for read state. In the Vivado waveform, you can see that the data is being written in an incremental manner. After connecting using the above method, the mapping relationship between the RNA sequence and the knowledge graph is input into the first FPGA chip, and the base pairing and other features are integrated through the lookup table LUT to achieve efficient extraction of RNA sequence information.
[0111] The first FPGA chip is provided with a first data line interface (D0-D7), a first synchronous clock interface (CLK) and a first enable signal interface (CLK). The first data line interface is connected to the DIN0-D7 pin of the first AI chip, the first synchronous clock interface is connected to the clock pin (CLK_IN) of the first AI chip, and the first enable signal interface is connected to the enable pin (ENABLE) of the first AI chip, thereby establishing the first FPGA chip and the first AI chip into an efficient and accurate heterogeneous computing platform for RNA secondary structure prediction.
[0112] The first AI chip is connected to the high-speed serial interface (HSI) of the second FPGA chip through an API interface. Through the API interface, the second FPGA chip can access the bus address of the first AI chip to achieve mutual communication between the two. The second FPGA chip is responsible for processing a large amount of input data and sending it to the first AI chip through the API interface for deep learning inference. After receiving the data, the first AI chip analyzes and predicts the secondary structure of the RNA and returns the result to the second FPGA chip. The second FPGA chip receives the prediction result of the first AI chip and performs subsequent data processing or storage. Through this collaborative working mode, the second FPGA chip and the first AI chip can give full play to their respective advantages and achieve accurate preliminary prediction of the RNA secondary structure.
[0113] The code was modified using the MicroPython format, and the deep neural network model based on the encoder and decoder architecture was input into the first AI chip to receive the RNA sequence features obtained from the feature extraction module. The performance of the first AI chip can reach hundreds of TFLOPS, so through acceleration by the first AI chip, accurate and efficient RNA secondary structure prediction can be achieved.
[0114] The second FPGA chip is provided with a second data line interface, a second synchronization clock interface and a second enable signal interface. The connection between the second FPGA chip and the second AI chip is exactly the same as the connection between the first FPGA chip and the first AI chip. The evaluation and optimization module is connected to the structure generation module through the high-speed serial interface (HSI) in the third interface unit to ensure fast data transmission and real-time communication. The preliminary prediction results of the RNA secondary structure from the structure generation module are received in the evaluation and optimization module, and are input into the VAE-GAN hybrid model for further optimization. The predicted RNA secondary structure is gradually optimized by using a pipeline processing method to improve the stability and functionality of the structure. The second AI chip and the third FPGA chip are connected through an API interface. After the optimization is completed, the final RNA secondary structure data is output through the third FPGA chip. The third FPGA chip can be configured as a serial or parallel output interface to meet different data transmission requirements. Through the above process, the evaluation and optimization module can effectively optimize the RNA secondary structure to ensure that the output data meets the needs of subsequent analysis and application.
[0115] The number of the first data line interfaces and the second data line interfaces are both 8, and the number of the first synchronous clock interfaces, the first enable signal interfaces, the second synchronous clock interfaces and the second enable signal interfaces are both 2.
[0116] The invention designs and develops a knowledge-driven and data-driven RNA secondary structure generator for the AVR gene of rice blast fungus, which has high integration and high computational efficiency and can realize accurate prediction of RNA secondary structure.
[0117] like Figure 2 As shown, the present invention also provides a method for predicting the secondary structure of the rice blast fungus AVR gene RNA by integrating knowledge-driven and data-driven methods, and using the rice blast fungus AVR gene RNA secondary structure generator that integrates knowledge-driven and data-driven methods, the method comprises the following steps:
[0118] Step 1: Collect RNA secondary structure data, RNA sequence data, thermodynamic characteristic data, kinetic characteristic data, and RNA whole genome sequence data of the avirulence gene (AVR) to ensure comprehensive acquisition of various characteristic information related to the avirulence gene and classify them according to data type to support subsequent analysis and research;
[0119] The thermodynamic characteristic data include RNA molecule free energy change, base pairing energy and base stacking energy;
[0120] The kinetic characteristic data is the folding rate of RNA molecules;
[0121] The RNA whole genome sequence data includes: ribosomal RNA (rRNA) sequence, transfer RNA (tRNA) sequence, small nuclear (snRNA) RNA sequence;
[0122] In this embodiment, the RNA structure data, thermodynamic characteristic data, and kinetic characteristic data are collected by an intelligent crawling algorithm, specifically including:
[0123] Step 1: RNA secondary structure data, RNA sequence data, thermodynamic characteristic data, kinetic characteristic data and RNA whole genome sequence data can be searched simultaneously in two ways;
[0124] First, search for relevant literature from academic paper retrieval websites such as PubMed and CNKI using keywords such as gene names and literature titles. Use information extraction and other methods to extract and integrate different types of data content, such as tables and graphics, from the content of existing literature, obtain data packages, and save them locally.
[0125] Second, visit the academic resource website and search for relevant data in NCBI. After receiving the data packet returned by the website, extract and save the location information such as the link address, request header or data transmission method during data transmission to the local computer.
[0126] Saved link addresses and data package details, use pycharm software to write a distributed intelligent crawling algorithm based on the consistent Hash algorithm to submit access requests to the relevant web data source server, specifically including:
[0127] First, in the Python environment, relevant libraries (such as requests, beautifulsoup4, and hashlib) are used for development. A hash ring is created to map the data source server and the crawler node. Each node generates a hash value through a hash function and stores it on the ring. Secondly, the hash value is calculated based on the URL of the web page to be crawled, its position on the hash ring is determined, and the request is assigned to the first node in the clockwise direction. Each crawler node uses the requests library to send a request to the designated server, obtains the returned data, and performs preliminary processing. The obtained data is saved to the local database and annotated using dot-bracket notation. The main program loop is written to continuously monitor and process the URL to be crawled. Through this series of steps, distributed intelligent crawling can be effectively implemented to ensure load balancing.
[0128] Step 2: Classify all data according to data type to obtain RNA secondary structure data, RNA sequence data, RNA thermodynamic characteristic data, kinetic characteristic data and whole genome sequence data, including:
[0129] Perform preliminary classification based on data type (such as RNA secondary structure data, thermodynamic characteristic data, etc.), and clean and filter the data;
[0130] In this embodiment, the data cleaning and filtering includes:
[0131] After categorizing data types, duplicates are removed to ensure that each piece of data is unique. Missing values are then checked and processed, filling or deleting missing records where necessary. Format normalization is then performed to unify date and numeric formats, as well as data type conversion to ensure accurate field types. Outlier detection helps identify and address anomalies in the data, while text processing removes HTML tags and special characters. After categorizing and labeling data according to data type, the cleaned data is stored in a local database and backed up regularly. Processing logs are also recorded for subsequent audits. In the main program loop, data to be processed is continuously monitored, and the above cleaning and filtering steps are repeated to ensure data quality and structural rationality.
[0132] Step 2: Figure 3 As shown in the figure, a knowledge- and data-driven deep neural network model based on the encoder-decoder architecture (VAE) was constructed. The classified RNA sequence data was input into the deep neural network model to obtain preliminary prediction results of the secondary structure of the rice blast fungus AVR gene RNA.
[0133] Among them, the encoder first extracts parameters such as base pairing stability and stacking energy from the thermodynamic characteristic data and kinetic characteristic data of RNA molecules, and standardizes and preprocesses these features. Importantly, thermodynamic and kinetic parameters are assigned to each edge in the obtained graph representation as edge attributes, so that it reflects the stability or energy value of base pairing. In the graph convolution operation, the network will use the attribute information of these edges as weighting coefficients, consider the characteristics of the edges when aggregating the characteristics of neighboring nodes, and guide the network to learn structural representations that conform to physical laws. In order to further optimize the model, thermodynamic and kinetic characteristics are incorporated into the loss function design. By comparing the error between the structure predicted by the model and the true structure and combining the expected value of the features to adjust the model parameters, after iterative training, the model can effectively learn how to integrate these features into the graph convolution network, thereby improving the prediction accuracy of RNA secondary structure.
[0134] The encoder specifically comprises the following steps:
[0135] Step a: Use the ViennaRNA package in SnapGene V6.1 software to import the obtained RNA structure data (RNA sequence data), and obtain the graphical representation of the RNA structure G = {V, E, X} by "converting to single strand", where V is composed of RNA base nodes, E represents the edge referenced between base nodes, and X represents the edge referenced to each node v. i The associated content features, N represents the length of the RNA sequence, that is, each nucleotide is regarded as a node in the graph, and the base pairing relationship and spatial adjacency relationship serve as the edges connecting these nodes, which facilitates the intuitive observation and analysis of the structural characteristics of RNA;
[0136] Step b, pre-processing the thermodynamic characteristic data and the kinetic characteristic data, specifically comprising:
[0137] The thermodynamic and kinetic characteristic data of RNA molecules are standardized to make their mean 0 and standard deviation 1 to eliminate the influence of different dimensions. At the same time, samples with missing values are checked and deleted to ensure the integrity of the data. In the graph representation, each edge is assigned the extracted thermodynamic and kinetic parameters as edge attributes to reflect the stability or energy value of base pairing. At the same time, to ensure that the feature data format meets the input dimension requirements of the graph convolutional network, the principal component analysis (PCA) method is used to scale and reduce the dimension of the data when necessary to maintain numerical stability during training. Finally, the standardized thermodynamic and kinetic characteristics are simply spliced with the original characteristics of the graph nodes, and then principal component analysis (PCA) is used for dimensionality reduction to form a new feature set, which provides high-quality input data for subsequent GCN training, thereby improving the prediction accuracy of RNA secondary structure.
[0138] Step c: construct a graph convolutional network (GCN) model, which includes an input layer, a first graph convolutional layer, a first activation layer, a first autoregressive layer, a second graph convolutional layer, a dropout layer, a second autoregressive layer, two fully connected layers, and an output layer, which are connected in sequence and can effectively perform node feature extraction and classification tasks;
[0139] The input layer is represented by a symmetric normalized Laplace matrix through the graphical representation G of the RNA structure, so that the model can further learn the relationship between each node. The specific conversion formula is:
[0140]
[0141] Where, degree matrix representing the graphical representation of the RNA structure, Represents the adjacency matrix of the graphical representation of the RNA structure, L represents the Laplacian matrix of N × N (number of nodes × feature dimension of each node);
[0142] In this embodiment, the first graph convolution layer has 16 output channels and is used to extract local features of nodes. The activation layer is a ReLU function, which introduces nonlinearity to enhance the model's expressiveness. The second graph convolution layer has 3 output channels and is responsible for further processing node features and outputting the probability distribution of categories. The dropout layer is used to randomly drop some nodes to prevent overfitting and enhance the generalization ability of the model. The autoregressive layer performs latent coding, the first fully connected layer calculates the mean, and the second fully connected layer calculates the variance. The output layer further integrates the outputs of the two fully connected layers to obtain latent variables.
[0143] The autoregressive layer satisfies:
[0144]
[0145] Where Z (n+1) is the potential code of the n+1th layer, and when n=0 it is the first autoregressive layer, when n=1 it is the second autoregressive layer, Z (0) =L,W (n) is the weight matrix of the current nth layer, represents the adjacency matrix loss function with self-loops, is the adjacency matrix, I is The identity matrix, represents the symmetric normalized degree matrix, represents the degree of node i (i.e., the element in the degree matrix), j represents the index used to traverse all nodes connected to node j, and φ is the activation function;
[0146] The two fully connected layers satisfy;
[0147] μ i =Z (n+1) W μ ;
[0148] σ i =exp(Z (n+1) W σ );
[0149] In the formula, since the input of the two fully connected layers is the output of the second autoregressive layer, n = 1, μ i is the mean, σ i is the variance, W μ is the first trainable weight matrix, W σ The second trainable weight matrix is initialized at any time when the model is built. The two weight matrices are continuously updated through the back-propagation algorithm during the training process.
[0150] The output layer satisfies:
[0151] Z=μ i +σ i ·∈;
[0152] Where Z is the latent variable, ∈ represents the noise sampled from the standard normal distribution;
[0153] The loss function of the encoder is:
[0154]
[0155] Where, represents the loss function, Indicates that given the feature matrix X and the adjacency matrix The expected value of the latent variable Z, Indicates that the adjacency matrix is calculated given the latent variable Z The logarithm of the probability of
[0156] Using the encoder's loss function and the Adam optimizer, the training epochs are 200. During the training process, the Laplacian matrix with node features is input through forward propagation. The latent code is obtained through the above convolutional layer. The mean and variance are then calculated based on the latent code to generate latent variables. The functions in pytorch are used to randomly generate, calculate the loss, backpropagate, and update the model parameters. The model outputs the loss value of each training cycle to monitor the training effect. After these features are processed by the graph convolutional network, the information of neighboring nodes is integrated, effectively reflecting the characteristics of RNA secondary structure.
[0157] Through continuous iterative updates, the improved GCN model can effectively capture the topological structure information within RNA molecules. In each round, the GCN model aggregates and transforms node features through latent variables and loss functions, gradually refining the hidden representation of RNA molecules. Eventually, after sufficient learning and aggregation, the model outputs the latent variable Z of the optimized RNA molecular structure. This latent variable not only contains rich structural information and features but also lays a foundation for subsequent bioinformatics analyses such as RNA secondary structure prediction and function annotation.
[0158] The decoder adopts an autoregressive generative model (PixelCNN) to generate the RNA secondary structure bottom-up by gradually predicting the pairing probability of each nucleotide. PixelCNN is a generative neural network that generates one pixel at a time when generating an image and uses previously trained pixels to generate the next pixel. The concept of masked convolution is introduced in PixelCNN to control the pixels that have not been predicted. Two main types of masked convolutions are used: Mask TypeA (Type A masked convolution) and Mask TypeB (Type B masked convolution). Mask TypeA acts on the first layer of convolution, while Mask TypeB acts on all convolution layers except the first layer. Therefore, the decoder consists of 20 identical structures. Each layer includes a Type A masked convolution, 20 residual modules, a first convolutional layer, a second activation layer, a second convolutional layer, a third activation layer, and an output layer connected in sequence.
[0159] Among them, the size of the Type A masked convolution is 8×8. The residual module is composed of a third convolutional layer, a Type B masked convolution, and a fourth convolutional layer connected in sequence. The convolutional kernels of the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer have a size of 1×1. The size of the Type B masked convolution is 5×5. The second activation layer and the third activation layer are both ReLU functions, and the output layer is a softmax function.
[0160] The loss function of the decoder satisfies:
[0161]
[0162] In the formula, is the structure of the true RNA secondary structure diagram, x′ ξ is the sample generated by the encoder, 0 < ξ < N, and ξ represents the positions of each node in the RNA secondary structure diagram structure.
[0163] The input of the decoder is the latent variable Z of the generated RNA molecular structure output by the encoder and the true RNA secondary structure diagram structure x collected previously *, the computational model through each layer will extract and update the understanding of RNA structural features by calculating the cross entropy loss function L * The method measures the difference between the generated sample x′ and the real RNA secondary structure graph x * The difference between them is then calculated based on the loss value obtained by calculating the loss function using the chain rule for back propagation. The model parameters are updated using the Adam optimization algorithm to gradually improve the accuracy of the generation. Specifically, the loss L is calculated starting from the output layer. * The gradient of the output x′ is then passed back to the previous layer, the gradient of each layer is calculated and the parameters of each layer are updated using the Adam optimization algorithm. According to the output structure of each round, the settings of the masked convolution can be dynamically adjusted to finally output the graph structure of the RNA secondary structure.
[0164] Among them, the gradient calculation of back propagation satisfies:
[0165]
[0166] The decoder is trained for 50 rounds and outputs a preliminary predicted graph structure after training is completed.
[0167] The decoder output needs to be converted into a pairing matrix K. The specific steps are as follows: The pairing matrix K has dimensions N*N, where N is the length of the sequence. Based on the obtained RNA secondary structure, we can determine which bases are paired with each other. For example, if base 1 pairs with base 5, the values of row 1, column 5 and row 5, column 1 in the pairing matrix are both set to 1. For unpaired base pairs, the corresponding position in the matrix has a value of 0. In this way, the RNA pairing matrix K can be constructed, in which paired positions are set to 1 and unpaired positions are set to 0.
[0168] Step 3: Evaluate and optimize the predicted structure using a GAN hybrid model. A nearest-neighbor-based RNA energy model is used to accurately assess the thermodynamic stability of the candidate structure. The energy value is used as part of the loss function. The predicted pairing probability matrix is constrained according to the RNA base pairing rules to ensure that the generated structure conforms to physical and chemical principles. Ultimately, the optimized secondary structure of the rice blast fungus AVR gene RNA is obtained. This includes the following steps:
[0169] Step I: Use the BioChen sequence conversion tool to convert the predicted RNA secondary structure graph into RNA secondary structure sequence;
[0170] Step II: Figure 4 As shown, a GAN hybrid model is constructed, which specifically includes:
[0171] Step 1) The output of the VAE, i.e., the PixelCNN generative model in the decoder part, is used to generate the preliminary RNA secondary structure pixel by pixel, denoted as C, and used as the input of the generative adversarial network (GAN);
[0172] Step 2) Build a generative adversarial network (GAN) model for data enhancement and distribution alignment:
[0173] In GAN, a generator (G) and a discriminator (D) are designed. Through continuous adversarial learning between the two, the predicted RNA secondary structure is continuously close to the real sample.
[0174] The main task of the generator is to generate samples that are as realistic as possible. It accepts random noise (usually sampled from a simple distribution, such as a Gaussian distribution) and converts it into samples similar to the training data. The goal is to make the generated samples as close to the real data as possible in the eyes of the discriminator. The generator mainly uses a convolutional neural network (CNN). The generator consists of an input layer, a fully connected layer, a reshape layer, five transposed convolution module layers, and an output layer connected in sequence, and outputs virtual images that are as realistic as possible.
[0175] The input layer is the pairing matrix K obtained by VAE, the transposed convolution module includes a transposed convolution layer, a batch normalization layer and a fourth activation layer connected in sequence, the fourth activation layer is a Relu activation function, and the output layer is a sigmoid activation function;
[0176] That is, the overall process of the generator is to set a fully connected layer after the input layer, map the input K to a vector with 256*16*16 nodes, and further adjust the shape of the mapped vector to (16,16,256) through a reshaping layer, that is, a 16*16 feature map with 256 channels. The reshaped image passes through a batch normalization layer and an activation layer respectively to reduce the overfitting of the model; in the transposed convolution layer part, the input feature map is first converted to a (32,32,128) shape so that it can be upsampled to a larger space. The following batch normalization and ReLU activation function layers are the same as before, continuing to stabilize training and introduce nonlinearity, and then continue the above upsampling process to convert the shape to (64,64,64); the final output layer is used to generate the final image, converting the dimension of the feature map from (64,64,64) to (64,64,3), that is, a 64x64 RGB image.
[0177] The task of the discriminator is to determine whether the input sample is a real sample (from real RNA secondary structure data) or a generated sample (from the generator). It learns the characteristics of real data through training and strives to improve the ability to identify generated data. The goal is to distinguish real samples from generated samples as accurately as possible. The discriminator and the generator also use CNN. The discriminator includes the input layer, the fifth convolutional layer, the fifth activation layer, the sixth convolutional layer, the sixth activation layer, the seventh convolutional layer, the seventh activation layer, the eighth convolutional layer, the eighth activation layer, the ninth convolutional layer, the flattening layer, the fully connected layer and the ninth activation layer connected in sequence;
[0178] The input layer receives a 64*64*3 image. The first layer is a two-dimensional convolutional layer, which uses a 5×5 convolution kernel to output 64 feature channels and downsamples the input feature map (with a stride of 2) to reduce the spatial dimension. Subsequently, the ReLU activation function is used to introduce nonlinearity to the output while avoiding the "neuron death" problem. Subsequently, the model adds a second convolutional layer, and the number of output channels increases to 128. Upsampling is still performed, and the ReLU activation function is applied again to maintain the nonlinear characteristics. Then, the third convolutional layer further increases the number of channels to 256, and downsampling is also performed. After 5 convolutional layers, the feature map is flattened into a one-dimensional vector for subsequent processing. Finally, a fully connected layer is added to output a value indicating the probability that the input image is a true sample (ranging between 0 and 1), which is converted using the sigmoid activation function.
[0179] To perform adversarial training of the generator and discriminator, first initialize the model, instantiate the generator and discriminator, and then optimize the model using the cross entropy loss function and Adam optimizer. The cross entropy loss function can be expressed as follows:
[0180]
[0181] Where, Denotes the expected loss of the discriminator on the real sample, denotes the predicted probability of the discriminator for the real data, where c represents the real RNA secondary structure data collected in step 1. Ideally, the discriminator should be able to maximize the probability D(c) of the real sample c to 1, so this term will have a positive impact on the training of the discriminator; It represents the expected loss of the discriminator on the generated sample, and represents the discriminator's prediction for the generated sample G(k). Here K represents the input RNA secondary structure pairing matrix obtained from the VAE model. Ideally, the discriminator should minimize the probability D(G(k)) of the generated sample to 0. This item encourages the discriminator to improve its ability to distinguish generated samples.
[0182] From the above formula, it can be seen that the optimization directions of the generator and the discriminator are different. The generator hopes to minimize the objective function, while the discriminator hopes to maximize the objective function. In order to find the optimal model, the alternating optimization method of the maximum and minimum game should be considered for optimization. The cost functions J(G) and J(D) for the two optimization objectives can be obtained from the value function:
[0183]
[0184] Where, E k represents the expectation of the pairing matrix K, J(G) and J(D) represent the loss functions of the generator and discriminator respectively, and It is a commonly used factor, mainly used to normalize the loss value. The meanings of other parameters are the same as those of the above cross entropy loss function;
[0185] Through the above two cost functions, we can find the best generator and discriminator:
[0186] First, fix the parameters of the generator G, derive the cost function, and set D(c) = 0. After simplification, we can get the optimal discriminator D * (c) as follows:
[0187]
[0188] Where, P r (c) and p g (c) represents the probability of real samples and generated samples respectively;
[0189] On the basis of the optimal discriminator, the concept of Jensen-Shannon divergence is introduced to find the optimal generator. By introducing the concept of Jensen-Shannon divergence, the optimization problem of the generator can be transformed into minimizing the JS divergence, thereby obtaining the objective function of the generator. The ultimate goal is to make the distribution p of the generated samples g (c) As close as possible to the distribution P of the real sample r (c);
[0190] The formulas for KL divergence and JS divergence are as follows:
[0191]
[0192] in, Represents the distribution of real samples P r The expected value of the preliminary secondary structure C in the KL divergence is used to measure the two probability distributions P r (c) and p g (c) The “distance” between the two, which quantifies the use of pg (c) to approximate P r (c) The effect of JS divergence is a symmetric version of KL divergence, which measures the relationship between two probability distributions P r (c) and p g (c) Similarity between.
[0193] According to the above formula, the objective function of the generator can be obtained as follows:
[0194] 2JS(P r (c)||P g (c))-2log 2;
[0195] Then, according to the above objective function during the training process, the parameters of the generator G are continuously adjusted so that the distribution p of the generated samples is g (c) As close as possible to the distribution P of the real sample r (c);
[0196] Next, multiple rounds of training are carried out. First, the discriminator is trained. A batch of samples are randomly selected from the real RNA secondary structure data obtained in step 1 as real samples. At the same time, the same number of fake samples are generated from the generator. Next, labels are set for real samples and fake samples respectively, the former is set to 1 and the latter is set to 0. Then the discriminator is trained. First, real samples and corresponding real labels are used for training; then, fake samples and corresponding fake labels are used for training. Through the above two steps, the discriminator can learn how to distinguish between real samples and generated fake samples; next, the generator is trained. First, new fake samples are generated through the designed generator network, and labels are created for the target labels of the generator. All labels are 1, indicating that the generated fake samples are expected to be judged as real samples by the discriminator. The goal of continuously training the generator is to make the generated fake samples be considered real by the discriminator; the two are trained against each other to further optimize the quality of samples produced by the generator until the optimal generator is found to improve prediction accuracy;
[0197] Through the above-mentioned GAN model, the generated RNA secondary structure can be further distinguished and optimized from the real RNA secondary structure, and the prediction results can be continuously brought closer to the true value, thereby further improving the prediction accuracy.
[0198] Step 4: Adopt a multi-source transfer learning strategy based on evolutionary relationships and meta-learning to enhance the applicability of the model;
[0199] like Figure 5 As shown in Figure 2, the multi-source transfer learning strategy based on evolutionary relationships and meta-learning is constructed as follows:
[0200] First, the RNA phylogenetic tree is constructed using the NJ neighbor-joining method based on the RNA whole genome sequence data obtained in step 1. The specific idea of the NJ neighbor-joining method is as follows:
[0201] The NJ neighbor joining method is a bottom-up clustering method that treats each individual as a category, and then aggregates them in pairs to finally obtain the entire aggregated developmental tree.
[0202] Assume that there are N * classification units, let D cd Refers to the distance between taxa c and d, L ab Refers to the branch length (distance) between nodes a and b, then the total branch length of the entire tree is:
[0203]
[0204] As clustering progresses, the types of classifications continue to increase, so the total branch length will continue to increase. Suppose there is only one node X at the beginning, and a new node Y is added after the first round of clustering. Then the total branch length S after the new branch length is 12 for:
[0205]
[0206]
[0207] Where, L XY Indicates the length of the newly added branch;
[0208] Calculate all S by the above method cd , and select the smallest pair as the selection for this round. Here, assuming that 12 is the neighbor selected in this round, the two of them and X will form a new classification unit <1,2>, and then calculate the distance between the new classification unit and the remaining classification units:
[0209]
[0210] The above steps are repeated until all taxa are found, and finally visualized to obtain a leaf node of the RNA phylogenetic tree representing a specific RNA sequence, while the internal nodes represent common ancestors or hypothetical ancestral sequences, which reflect the relationship and evolutionary history between different RNA sequences; and the connection between nodes reflects the relationship between different RNA sequences or species. Nodes with closer connections indicate that these sequences are closer in evolution, while nodes with farther connections indicate that their common ancestor may be earlier; the length of the branch usually indicates the time of evolution or the degree of variation. Longer branches indicate that more changes or mutations have occurred during the evolutionary process, while shorter branches indicate fewer changes.
[0211] When using the prediction method of the present invention to predict RNA secondary structure, the RNA phylogenetic tree in step 4 can be used to select the source sequence closest to the target RNA sequence for pre-training, thereby improving the learning efficiency of the target task. 3-5 species with close evolutionary distance and rich data are selected as source domains based on the length of the branches. In each source domain, the VAE-GAN hybrid model constructed in the prediction method is pre-trained with known RNA secondary structure data to learn species-specific RNA structural characteristics and patterns.
[0212] In the domain adaptation phase, the main goal of this module is to optimize the feature alignment between the source domain and the target domain (the species of the RNA sequence to be measured) by introducing phylogenetic information of RNA species. This alignment process helps reduce feature offsets caused by species differences, thereby improving the model's performance in the target domain. The previously designed GAN adversarial generative model is used to perform target alignment between the source and target domains.
[0213] Through the above pre-training, when the model faces a new target domain RNA, the model can start from the parameters obtained by pre-training and quickly adapt to the target domain data through a small number of gradient steps, thereby improving prediction efficiency and accuracy.
[0214] In this embodiment, model compression and computational optimization techniques are used to further improve the practicality and computational efficiency of the model;
[0215] like Figure 6 、 7 As shown, the construction method of model compression and computation optimization technology is as follows:
[0216] ① Apply a progressive model compression strategy, systematically analyze the importance of each model component, and selectively trim and simplify redundant parts;
[0217] Adopt the magnitude-based pruning method, use the size of the weight value to judge its importance, obtain the weight information of the model, and call the parameters() method in PyTorch to calculate the weight of each layer of the model, so as to obtain the importance of each layer. According to the weight, you can set a threshold (here it is set to 0.05, that is, the layers with less than 5% weight will be pruned), or select several layers with the lowest total weight for pruning. For example, you can include layers with weights lower than a set threshold in the pruning range, or select the first few layers with the lowest importance score for processing. Then, for the pruned layers, you can choose to set their weights to zero, or directly remove these layers from the model to simplify the model structure. Finally, after the pruning is completed, the model needs to be retrained and evaluated to ensure that it can maintain good performance while improving efficiency. In this way, evaluating the importance of each layer based on the size of the weight value helps to improve the efficiency and practicality of the model, effectively reduce the complexity of the model and will not cause a great loss in the performance of the model, so as to find a suitable student model.
[0218] In this embodiment, the student model is a pruned VAE-GAN model. Specifically, the last 8 layers of the 20 layers in the VAE decoder PixelCNN model are removed, and the last two of the five convolutional layers in the generator and discriminator of the GAN are removed.
[0219] ②Use knowledge distillation method based on structural similarity:
[0220] Using the original VAE-GAN model as the teacher model, we designed an RNA secondary structure similarity metric function. This function comprehensively considers factors such as base pair matching and free energy difference to evaluate the similarity between the structures generated by the teacher model and the student model. This function guides the student model to generate secondary structures similar to those predicted by the teacher model, thus achieving effective knowledge transfer. The specific implementation method is as follows:
[0221] Hamming distance is used to evaluate the base pair matching between two RNA secondary structures. The calculation formula is as follows:
[0222]
[0223] Where S1 represents the RNA secondary structure predicted by the teacher model, S2 represents the RNA secondary structure predicted by the student model, and N match is the number of base pairs matched, N total is the total number of base pairs;
[0224] RNAfold software was used to calculate the free energy of RNA secondary structures predicted by the teacher model and the student model, and the difference between the two was calculated. The calculation formula is as follows:
[0225] ΔG=|G(S1)-G(S2)|;
[0226] Where ΔG represents the difference in free energy, G(S1) and G(S2) represent the free energy of the structures predicted by the two models, respectively;
[0227] Combining the above two aspects, we can design a comprehensive similarity measurement function, the formula is as follows:
[0228] Similarity(S1,S2)=α·Match(S1,S2)-β·ΔG;
[0229] Where α and β are hyperparameters used to adjust the impact of each part on the final similarity score.
[0230] According to the above similarity measurement function, the student model is continuously optimized until the optimal student model is found, so as to reduce the complexity of the model and the number of model parameters. The specific steps are as follows: Specifically, it is the process of determining the clipping threshold. The threshold can be first set to 0.01. At this time, only two layers can be removed. Then, the student model at this time is used to calculate the similarity score with the original model, that is, the teacher model. If the scores are the same or similar, the threshold is increased to continue to compare the scores. When the threshold is set to 0.06, the similarity scores are quite different, so the threshold is set to 0.05 here to find the optimal student model.
[0231] ③ In terms of hardware optimization, we utilize the parallel processing capabilities of modern computing architectures and adopt technical means such as GPU acceleration and distributed computing to improve model training and inference speed.
[0232] The present invention designs and develops a prediction method for the secondary structure of the AVR gene RNA of rice blast fungus that integrates knowledge-driven and data-driven approaches. It cleverly integrates the two paradigms, designs and uses a deep neural network model based on an encoder-decoder architecture, introduces a VAE-GAN hybrid model to evaluate and optimize the prediction results, and adopts a multi-source transfer learning strategy based on evolutionary relationships and meta-learning to improve the generalization ability of the model. Finally, through model compression and computational optimization, the practicality and computational efficiency of the method are further improved, the deep integration of structure and multi-omics data is achieved, and the relationship between RNA secondary structure and AVR gene evolution and regulation is further explored to achieve precise regulation of rice blast disease.
[0233] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and embodiments shown and described herein.
Claims
1. A method for predicting the secondary structure of the rice blast fungus AVR gene RNA that integrates knowledge-driven and data-driven methods, characterized in that: The steps include: Step 1: Collect RNA secondary structure data, RNA sequence data, thermodynamic characteristic data, kinetic characteristic data, and RNA whole genome sequence data of non-toxic genes and classify them according to data type; Step 2: construct a deep neural network model based on an encoder-decoder architecture, and input the classified RNA sequence data, thermodynamic characteristic data, and kinetic characteristic data into the deep neural network model to obtain an RNA secondary structure pairing matrix; The construction of the encoder includes the following steps: Step a: constructing a graphical representation of the RNA structure G = {V, E, X} based on the classified RNA sequence data; Among them, V is composed of RNA base nodes, E represents the reference edge between base nodes, and X represents the reference edge between each node v i Associated content features, each nucleotide is considered as a node, while base pairing relationships and spatial adjacency relationships serve as edges connecting these nodes; Step b, pre-processing the thermodynamic characteristic data and the kinetic characteristic data to use them as edges referenced between base nodes; Step c: Build a graph convolutional network model; The graph convolutional network model includes an input layer, a first graph convolutional layer, a first activation layer, a first autoregressive layer, a second graph convolutional layer, a dropout layer, a second autoregressive layer, two fully connected layers and an output layer, which are connected in sequence. Wherein, the input layer satisfies: Where, degree matrix representing the graphical representation of the RNA structure, represents the adjacency matrix of the graphical representation of the RNA structure, L represents the N × N Laplacian matrix; The first autoregressive layer and the second autoregressive layer both satisfy: Where Z (n+1) is the potential code of the n+1th layer, and when n=0 it is the first autoregressive layer, when n=1 it is the second autoregressive layer, Z (0) =L,W (n) is the weight matrix of the nth layer, I am The identity matrix, represents the symmetric normalized degree matrix, φ is the activation function; The two fully connected layers satisfy; μ i =Z (n+1) IN μ ; σ i =exp(Z (n+1) IN σ ); Where n = 1, μ i is the mean, σ i is the variance, W μ is the first trainable weight matrix, W σ is the second trainable weight matrix; The output layer satisfies: Z=μ i +s i ·∈; Where Z is the latent variable, ∈ represents the noise sampled from the standard normal distribution; The loss function of the encoder is: Where, represents the loss function, Indicates that given the feature matrix X and the adjacency matrix The expected value of the latent variable Z, Indicates that the adjacency matrix is calculated given the latent variable Z The logarithm of the probability of Step d: Inputting the graphical representation of the RNA structure into the graph convolutional network model to obtain latent variables of the RNA molecular structure; The decoder consists of 20 layers of the same structure, each of which includes a type A mask convolution, 20 residual modules, a first convolutional layer, a second activation layer, a second convolutional layer, a third activation layer and an output layer connected in sequence; Among them, the size of the class A masked convolution is 8×8, the residual module is the third convolution layer, the class B masked convolution and the fourth convolution layer connected in sequence, the convolution kernel size of the first convolution layer, the second convolution layer, the third convolution layer and the fourth convolution layer is 1×1, the size of the class B masked convolution is 5×5, the second activation layer and the third activation layer are both ReLU functions, and the output layer is a softmax function; The loss function of the decoder satisfies: In the formula, is the true RNA secondary structure diagram, and x′ ξ is the sample generated by the encoder. 0 < ξ < N, where ξ represents the positions of each node in the RNA secondary structure diagram; The input of the decoder is the potential variables of RNA molecular structure and the real RNA secondary structure data graph structure of the non-toxic gene, and the output of the decoder is the graph structure of RNA secondary structure; Step 3: Construct a GAN model and input the RNA secondary structure pairing matrix into the GAN model to obtain an optimized secondary structure.
2. The method for predicting the secondary structure of the rice blast fungus AVR gene RNA by integrating knowledge-driven and data-driven methods according to claim 1, wherein: The step one also includes: cleaning and filtering the classified data.
3. The method for predicting the secondary structure of the rice blast fungus AVR gene RNA by integrating knowledge-driven and data-driven methods according to claim 2, wherein: The GAN model includes a generator and a discriminator; The generator includes an input layer, a fully connected layer, a reshape layer, 5 transposed convolution module layers and an output layer connected in sequence; The discriminator includes an input layer, a fifth convolutional layer, a fifth activation layer, a sixth convolutional layer, a sixth activation layer, a seventh convolutional layer, a seventh activation layer, an eighth convolutional layer, an eighth activation layer, a ninth convolutional layer, a flattening layer, a fully connected layer, and a ninth activation layer, which are connected in sequence; The input of the generator is the RNA secondary structure pairing matrix, the output of the generator is a 64x64 RGB image, the input of the discriminator is the output of the generator, and the output of the discriminator is the probability of a real sample; The objective function of the generator is: f=2JS(P r (c)||P g (c))-2log2; Where, JS(P r ||p g ) is the JS divergence, p g To generate the sample distribution, P r is the distribution of the real sample; The optimal discriminator satisfies: Where D * (c) is the optimal discriminator, P r (c) is the probability of the true sample, p g (c) is the probability of generating a sample.
4. The method for predicting the secondary structure of the rice blast fungus AVR gene RNA by integrating knowledge-driven and data-driven methods according to claim 3, wherein: It also includes model compression and computational optimization, and its construction specifically includes: Step 1: Build a student model: Calculate the weights of each layer of the original VAE-GAN model, prune the layers below the threshold, and obtain the pruned student model; Step 2: Use the original VAE-GAN model as the teacher model and obtain the similarity measurement function based on the student model: Similarity(S1,S2)=α·Match(S1,S2)-β·ΔG; Where S1 represents the RNA secondary structure predicted by the teacher model, S2 represents the RNA secondary structure predicted by the student model, α is the first hyperparameter, Match(S1,S2) is the base pair matching degree between the RNA secondary structures predicted by S1 and S2, β is the second hyperparameter, and ΔG represents the difference in free energy. Step ③: Obtain the optimal student model based on the similarity measurement function.
5. A knowledge-driven and data-driven fusion rice blast fungus AVR gene RNA secondary structure generator, using the rice blast fungus AVR gene RNA secondary structure prediction method according to any one of claims 1 to 4, characterized in that: include: A storage module, connected to the PC, for storing RNA raw data and intermediate results; and A feature extraction module, connected to the storage module, for extracting sequence features of the RNA raw data; a structure generation module, connected to the feature extraction module, for predicting the graph structure of the secondary structure of the rice blast fungus AVR gene RNA; The evaluation and optimization module is connected to the structure generation module and is used to optimize the graph structure of the secondary structure of the rice blast fungus AVR gene RNA.
6. The knowledge-driven and data-driven rice blast fungus AVR gene RNA secondary structure generator according to claim 5, characterized in that: The storage module includes: An ARM chip connected to the PC; A DRAM chip connected to the ARM chip; The feature extraction module includes: A first FPGA chip, connected to the ARM chip; The structure generation module includes: a first AI chip connected to the first FPGA chip; a second FPGA chip, connected to the first AI chip; The evaluation and optimization module includes: a second AI chip connected to the second FPGA chip; The third FPGA chip is connected to the second AI chip and is used to output the optimized secondary structure of the rice blast fungus AVR gene RNA.
7. The knowledge-driven and data-driven rice blast fungus AVR gene RNA secondary structure generator according to claim 6, characterized in that: Also includes: A first interface unit, which is provided on the storage module and is used to connect to a PC terminal; a second interface unit, which is provided on the feature extraction module and is used to connect with the structure generation module; The third interface unit is provided on the structure generation module and is used to connect with the evaluation and optimization module.
8. The knowledge-driven and data-driven rice blast fungus AVR gene RNA secondary structure generator according to claim 7, characterized in that: The first interface unit includes: A USB interface is provided on the ARM chip and is used to connect to a PC; A JTAG interface is provided on the ARM chip and is used to debug the ARM chip; An Ethernet interface, which is provided on the ARM chip and is used to connect to a PC; An RS232 interface is provided on the ARM chip and is used to connect to the DRAM chip; The second interface unit includes: A first data line interface, provided on the first FPGA chip and configured to connect to the DIN0-D7 pins of the first AI chip; A first synchronous clock interface, provided on the first FPGA chip and configured to connect to a clock pin of the first AI chip; A first enable signal interface, provided on the first FPGA chip and configured to connect to an enable pin of the first AI chip; The third interface unit includes: A second data line interface, which is provided on the second FPGA chip and is used to connect to the DIN0-D7 pin of the second AI chip; A second synchronous clock interface, provided on the second FPGA chip and configured to connect to a clock pin of the second AI chip; A second enable signal interface, provided on the second FPGA chip and configured to connect to an enable pin of the second AI chip; The second FPGA chip is connected to the first AI chip through an API interface, and the third FPGA chip is connected to the second AI chip through an API interface.
Citation Information
Patent Citations
Hardware acceleration method for predication of RNA second-stage structure with pseudoknot
CN104537278A
Processing method and device for RNA secondary structure prediction
CN115881209A