RNA expression acquisition method, device, electronic device and storage medium

By obtaining and fusing the discrete representation of RNA structure and sequence, the problem of not considering structural information in RNA sequence conversion is solved, and more accurate RNA function prediction and drug development progress are achieved.

CN119360967BActive Publication Date: 2025-10-03BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411303457.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-10-03
Estimated Expiration
2044-09-18

AI Technical Summary

Technical Problem

In the existing technology, RNA sequence conversion does not take into account the structural information of RNA, resulting in poor tertiary structure prediction.

Method used

By obtaining the structure discretization representation and sequence discretization representation of RNA and fusing them, the target representation of RNA is generated. The fusion process can be performed through splicing, weighted summation or using a language model for feature extraction and fusion.

Benefits of technology

It improves the accuracy and reliability of RNA function prediction, provides a theoretical basis for disease diagnosis and treatment, promotes the study of the interaction mechanism between RNA and protein, and thus accelerates the drug development process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360967B_ABST
    Figure CN119360967B_ABST
Patent Text Reader

Abstract

This disclosure provides a method, apparatus, electronic device, and storage medium for obtaining an RNA representation, relating to the fields of artificial intelligence and bioinformatics, particularly biocomputing. The specific implementation scheme comprises: obtaining a discretized representation of the RNA structure based on the RNA structure information; obtaining a discretized representation of the RNA sequence based on the RNA sequence information; and fusing the discretized representation of the RNA structure and sequence to obtain a target representation of the RNA.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of artificial intelligence and bioinformatics, specifically to the field of biocomputing technology, and more particularly to a method, device, electronic device, and storage medium for obtaining RNA representation. Background Art

[0002] Ribonucleic acid (RNA) representation learning uses mathematical and computational methods to convert RNA sequences into feature vectors that can be used for machine learning. This technology is widely used in fields such as RNA sequence classification, structure prediction, and drug design. It helps reveal the functions and regulatory mechanisms of RNA in biological processes and promotes research progress in bioinformatics and drug development. However, existing technologies do not consider RNA structural information during RNA sequence conversion, resulting in poor tertiary structure prediction. Summary of the Invention

[0003] The present disclosure provides a method, device, electronic device, and storage medium for obtaining RNA representation.

[0004] According to one aspect of the present disclosure, a method for obtaining an RNA representation is provided, comprising: obtaining a structural discretization representation of the RNA based on structural information of the RNA; obtaining a sequence discretization representation of the RNA based on sequence information of the RNA; and fusing the structural discretization representation of the RNA and the sequence discretization representation to obtain a target representation of the RNA.

[0005] According to another aspect of the present disclosure, a device for acquiring a representation of RNA is provided, comprising: a first acquisition module for acquiring a structural discretized representation of the RNA based on structural information of the RNA; a second acquisition module for acquiring a sequence discretized representation of the RNA based on sequence information of the RNA; and a fusion module for fusing the structural discretized representation of the RNA and the sequence discretized representation to obtain a target representation of the RNA.

[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the RNA representation acquisition method described in the above-mentioned embodiment.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, on which a computer program / instruction is stored, and the computer instructions are used to enable the computer to execute the RNA representation acquisition method described in the embodiment of the above one aspect.

[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program / instruction, which, when executed by a processor, implements the RNA representation acquisition method described in the embodiment of the first aspect.

[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0011] Figure 1 A schematic diagram of a process for obtaining an RNA representation provided in an embodiment of the present disclosure;

[0012] Figure 2 A schematic diagram of a process for obtaining another RNA representation provided in an embodiment of the present disclosure;

[0013] Figure 3 A schematic diagram of a process flow for training a VQ-VAE model in a method for obtaining RNA representation provided by an embodiment of the present disclosure;

[0014] Figure 4 A schematic diagram of a process for obtaining another RNA representation provided in an embodiment of the present disclosure;

[0015] Figure 5 A schematic diagram of a process for training a target language model in a method for obtaining an RNA representation provided by an embodiment of the present disclosure;

[0016] Figure 6 A schematic diagram of a process for obtaining another RNA representation provided in an embodiment of the present disclosure;

[0017] Figure 7 A schematic diagram of the structure of the target representation of obtaining RNA provided in an embodiment of the present disclosure;

[0018] Figure 8 A schematic diagram of the structure of an RNA representation acquisition device provided in an embodiment of the present disclosure;

[0019] Figure 9 The block diagram is a block diagram of an electronic device used to implement the RNA representation acquisition method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0020] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0021] The following describes the RNA representation acquisition method, device, and electronic device according to the embodiments of the present disclosure with reference to the accompanying drawings.

[0022] Artificial Intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This discipline encompasses both hardware and software technologies. AI hardware technologies generally include computer vision, speech recognition, natural language processing, as well as deep learning / learning, big data processing, and knowledge graphs.

[0023] Biocomputing is a field that leverages the principles and mechanisms of biological systems to solve computational problems. It applies biological properties and processes to computing systems to improve efficiency and performance. The goal of biocomputing is to draw inspiration from biological systems and translate it into new computational methods and technologies to solve complex problems. Biocomputing has a wide range of applications in areas such as optimization, pattern recognition, data analysis, and simulation, and continues to grow and expand.

[0024] Figure 1 A schematic diagram of a process for obtaining an RNA representation provided in an embodiment of the present disclosure.

[0025] like Figure 1 As shown, the method for obtaining the RNA may include:

[0026] S101, obtaining a discretized representation of the RNA structure based on the RNA structural information.

[0027] It should be noted that the execution entity of the RNA representation acquisition method in the embodiments of the present disclosure may be a hardware device with data processing capabilities and / or the necessary software to drive the operation of the hardware device. Optionally, the execution entity may include a server, a user terminal, and other intelligent devices. Optionally, the user terminal includes but is not limited to a mobile phone, a computer, an intelligent voice interaction device, etc. Optionally, the server includes but is not limited to a network server, an application server, a server of a distributed system, or a server integrated with a blockchain, etc. This is not specifically limited in the embodiments of the present disclosure.

[0028] In some implementations, RNA structural information includes secondary structure information and tertiary structure information. The secondary structure information of RNA primarily describes the pairing between bases in the RNA molecule, which is determined by hydrogen bonding interactions between bases within the RNA chain. The tertiary structure information of RNA further describes the shape and conformation of the RNA molecule in three-dimensional space. For example, the tertiary structure information of RNA can be the three-dimensional coordinates of the RNA molecule.

[0029] Optionally, the discretized structural representation of the RNA may be determined based on the secondary structure information and / or tertiary structure information of the RNA.

[0030] In some implementations, a discretized representation of the structure of the RNA can be determined based on a pre-built discretized vocabulary and the structural information of the RNA. The discretized vocabulary contains multiple discretized representations of the structure. The similarity between the RNA and the discretized representations is calculated, and the discretized representation with the highest similarity is used as the discretized representation of the structure of the RNA.

[0031] In some implementations, a pre-trained encoding model can be used to encode the RNA structure information to obtain a discretized representation of the RNA structure. The RNA structure information is input into the pre-trained encoding model, which encodes the RNA. The similarity between the encoded RNA and the discretized representation of the structure in the discretized vocabulary is calculated to determine the discretized representation of the RNA structure.

[0032] S102: Obtain a discretized representation of the RNA sequence based on the RNA sequence information.

[0033] It is understood that the discretization of RNA sequences is a process of converting each nucleic acid in the RNA sequence into a numerical value or symbol. In the discretization, each nucleic acid is given a unique identifier, usually an integer or a specific character.

[0034] Alternatively, the RNA nucleic acids, such as adenine, cytosine, guanine, and uracil, can be determined from the RNA sequence information, and these four nucleic acids can be mapped to a number. For example, if adenine is A, cytosine is C, guanine is G, and uracil is U, then A can be mapped to 1, B to 2, C to 3, and D to 4. In other words, if the RNA sequence is AUG CAU, the corresponding RNA sequence discretization representation is 132142.

[0035] S103, fusing the structure discretization representation and sequence discretization representation of the RNA to obtain the target representation of the RNA.

[0036] Alternatively, the structure discretization representation and sequence discretization representation of the RNA may be concatenated to achieve fusion to obtain the target representation of the RNA. Alternatively, different weights may be assigned to the structure discretization representation and sequence discretization representation of the RNA and a weighted sum may be performed to achieve fusion to obtain the target representation of the RNA.

[0037] Optionally, a language model can be trained to extract structural features and sequence features from the structural discretization representation and the sequence discretization representation, and the features can be fused to generate a target representation of the fused RNA.

[0038] According to the method for obtaining RNA representation provided by the embodiment of the present disclosure, by obtaining the discretized representation of RNA structure and the discretized representation of RNA sequence, and fusing the discretized representation of RNA structure and the discretized representation of RNA sequence, a target representation of RNA can be obtained, thereby achieving the fusion of RNA structural information and sequence information, providing more comprehensive and richer information, helping to more accurately describe the characteristics of RNA molecules, and improving the accuracy and reliability of predicting RNA functions. Furthermore, by fusing RNA structural information and sequence information for representation learning, a theoretical basis can be provided for the diagnosis and treatment of diseases, and it can help to study the interaction mechanism between RNA and protein, thereby accelerating the drug development process.

[0039] Figure 2 A schematic diagram of a process for obtaining an RNA representation provided in an embodiment of the present disclosure.

[0040] like Figure 2 As shown, the method for obtaining the RNA may include:

[0041] S201, generating a structure diagram of the RNA based on the structure information of the RNA.

[0042] In some implementations, RNA-based structural diagrams can intuitively display the spatial conformation of RNA molecules, helping to intuitively understand RNA structure and facilitate prediction of RNA structure. In other words, RNA structural diagrams can be generated based on RNA structural information.

[0043] Optionally, the RNA structural information includes at least one of secondary structure information and tertiary structure information of the RNA. That is, a structural diagram of the corresponding RNA can be generated based on the secondary structure information, a structural diagram of the corresponding RNA can be generated based on the tertiary structure information, or a structural diagram of the corresponding RNA can be generated based on the secondary structure information and the tertiary structure information.

[0044] That is, in response to the structural information including secondary structural information and tertiary structural information, a first structural graph of RNA is constructed according to the secondary structural information, and a second structural graph of RNA is constructed according to the tertiary structural information.

[0045] In some implementations, different representations of the RNA structure can be selected to generate a structure diagram of the RNA. For example, a dotted bracket representation, a computed tomography (CT) file representation, or a graphical representation can be selected to generate a structure diagram of the RNA.

[0046] As you can understand, the dot-bracket notation uses dots and brackets to represent RNA secondary structure, with dots representing unpaired bases and paired brackets representing paired bases. CT file notation is a file format that contains detailed information about RNA structure, including the total number of nucleotides, folding energy, and base pairing. Graphical notation uses specialized bioinformatics software or drawing tools to display RNA structure in graphical form, including two-dimensional plan views and three-dimensional spatial diagrams.

[0047] S202, based on the RNA structure graph and a pre-constructed discretization vocabulary, determining a discretization representation of the RNA structure.

[0048] In some implementations, the discretization vocabulary includes multiple structural discretization representations. The similarity between the structure graph of RNA and the structural discretization representation can be calculated, and the structural discretization representation with the highest similarity can be used as the structural discretization representation of RNA.

[0049] Optionally, by obtaining the quantization vector corresponding to the RNA structure graph and calculating the similarity between the quantization vector and the structure discretization representation to determine the structure discretization representation of the RNA, the complexity of the RNA structure graph can be reduced, and redundant information in the RNA structure graph can be removed, thereby achieving effective compression of the structure graph.

[0050] Optionally, a latent space representation of the RNA is obtained based on the RNA structure graph, and the latent space representation is quantized to obtain a quantized vector, and the structural discretization representation with the highest similarity to the quantized vector is obtained from a pre-constructed discretization vocabulary. Optionally, the similarity between the quantized vector and each structural discretization representation can be calculated, and the similarities can be sorted based on the size of the similarities, and the structural discretization representation with the greatest similarity is used as the structural discretization representation with the highest similarity.

[0051] Furthermore, it can be determined that the structural discretization representation with the highest similarity is the structural discretization representation of RNA.

[0052] In some implementations, an encoding model can be pre-trained and used to encode and quantize the RNA structure graph, obtaining a quantized vector corresponding to the RNA structure graph. The quantized vector is then matched to the discretized structure representation with the highest similarity from the discretized vocabulary, as the discretized structure representation of the RNA. The encoding model can be a Vector Quantized Variational Autoencoder (VQ-VAE) model, which includes an encoder and a quantization layer.

[0053] That is, the RNA structure graph is input into a pre-trained first vector quantized variational autoencoder (VQ-VAE) model, and the encoder in the first VQ-VAE model encodes the structure graph to obtain a latent space representation. Alternatively, a graph neural network (GNN) can be used to encode the structure graph.

[0054] Furthermore, the quantization layer in the first VQ-VAE model quantizes the latent space representation and matches the quantized vector in the discretized vocabulary to obtain the structural discretized representation of RNA, so as to effectively compress RNA data and improve the generation efficiency and quality of the model for the structural discretized representation, which is particularly suitable for the storage and processing of large-scale data.

[0055] S203: Obtain a discretized representation of the RNA sequence according to the RNA sequence information.

[0056] S204: Fusing the structure discretization representation and the sequence discretization representation of the RNA to obtain a target representation of the RNA.

[0057] The relevant contents of steps S203-S204 can be found in the above embodiment and will not be repeated here.

[0058] According to the RNA representation acquisition method provided by the embodiment of the present disclosure, a corresponding structural diagram can be generated based on the structural information of RNA, and the structural diagram can be quantified to determine the discretized structural representation of RNA, and then the discretized structural representation of RNA and the discretized sequence representation of RNA can be fused to obtain the target representation of RNA, thereby realizing the fusion of RNA structural information and sequence information, providing more comprehensive and richer information, helping to more accurately describe the characteristics of RNA molecules, and improving the accuracy and reliability of predicting RNA functions. Furthermore, by fusing RNA structural information and sequence information for representation learning, a theoretical basis can be provided for the diagnosis and treatment of diseases, and it can help to learn the interaction mechanism between RNA and protein, thereby accelerating the drug development process.

[0059] Based on the above embodiments, the present disclosure can explain the training process of the VQ-VAE model, such as Figure 3 As shown, the training process of the VQ-VAE model may include:

[0060] S301: Obtain sample RNA and obtain sample structure information of the sample RNA.

[0061] Alternatively, sample RNA may be extracted from sample cells or sample tissues, and an extraction reagent may be used to extract the sample RNA. For example, RNA may be extracted from animal or plant cells, tissues, or blood as the sample RNA.

[0062] Furthermore, bioinformatics tools can be used to determine the sample structure information of the sample RNA. The secondary structure of the sample RNA can be predicted to obtain the sample structure information. X-ray crystallography, cryo-electron microscopy, or nuclear magnetic resonance can also be used to analyze the tertiary structure of the sample RNA as sample structure information.

[0063] S302: Input the sample structure information into the initial VQ-VAE model, and the encoder in the VQ-VAE model encodes the sample structure information to obtain the sample latent space representation.

[0064] In some implementations, the initial VQ-VAE model structure includes an encoder, a decoder, and a quantization layer. Sample structure information is input into the initial VQ-VAE model, where the encoder encodes the sample structure information to obtain a latent space representation of the sample.

[0065] Optionally, the encoder obtains a latent space representation of the sample by mapping the sample structure information to a continuous vector representation in the latent space.

[0066] S303: The quantization layer in the VQ-VAE model quantizes the sample latent space representation, and matches the sample quantization vector in the reference discretization vocabulary to obtain a sample discretization sequence of the sample RNA.

[0067] In some implementations, the main function of the quantization layer is to map the continuous latent space representation output by the encoder into a discrete latent space, that is, by quantizing the sample latent space representation to obtain a sample quantization vector, and calculating the similarity between the sample quantization vector and the discretization representation in the reference discretization vocabulary, and based on the similarity size, matching the sample quantization vector and the reference discretization vocabulary to obtain a sample discretization sequence of the sample RNA.

[0068] Optionally, the distance between the sample quantized vector and the discretization representation in the reference discretization vocabulary, such as the Euclidean distance, can be calculated, and the discretization representation with the closest distance, that is, the most similar discretization representation, can be determined.

[0069] S304: The decoder in the VQ-VAE model decodes the discretized sequence of the sample to obtain the reduced RNA.

[0070] In some implementations, the decoder in the VQ-VAE model can decode the sample discretization sequence by receiving the sample discretization sequence sent by the quantization layer to reconstruct the sample discretization sequence and obtain the restored RNA.

[0071] Optionally, the reduced RNA may be generated based on a nonlinear transformation, such as by performing a convolution transformation on a discretized sequence of the sample to generate the reduced RNA.

[0072] S305. Adjust the model parameters of the VQ-VAE model according to the sample RNA and the reduced RNA, as well as the sample latent space table and the sample discretization sequence, and continue training until the end condition is met to obtain a second VQ-VAE model. The first VQ-VAE model only includes the encoder and quantization layer in the second VQ-VAE model.

[0073] In some implementations, the loss function of the model can be determined based on the sample RNA and the reduced RNA, as well as the sample latent space table and the sample discretized sequence, and the model parameters can be adjusted based on the loss function to obtain an adjusted VQ-VAE model, and the adjusted VQ-VAE model can be used to continue training until the end condition is met to obtain a second VQ-VAE model. The second VQ-VAE model includes an encoder, a decoder, and a quantization layer. When using the trained second VQ-VAE model, only an encoder and a quantization layer are required, that is, the first VQ-VAE model in the above embodiment only includes an encoder and a quantization layer.

[0074] Optionally, the training end condition may be that the number of training times reaches a threshold, the training end condition may also be that the loss function is less than a set threshold, or the training end condition may also be that the accuracy of the model is greater than an accuracy threshold.

[0075] In some implementations, the loss function of the VQ-VAE model includes a reconstruction loss and a quantization loss. The reconstruction loss can be determined based on the sample RNA and the reduced RNA, and the quantization loss can be determined based on the sample latent space table and the sample discretized sequence. The reconstruction loss is used to measure the reconstruction ability of the decoder, while the quantization loss encourages the encoder to learn to map the continuous latent vector to the reference discretized vocabulary.

[0076] Optionally, mean square error or cross entropy can be used as the reconstruction loss of the VQ-VAE model. Square loss or cross entropy loss can be used as the quantization loss of the VQ-VAE model.

[0077] Furthermore, the VQ-VAE model parameters can be adjusted based on the reconstruction loss and quantization loss. In other words, by co-optimizing the model parameters through the reconstruction loss and quantization loss, the VQ-VAE model can not only learn the discrete representation of RNA, but also significantly reduce storage requirements and improve the quality of generated data while preserving important RNA features.

[0078] According to the RNA representation acquisition method provided by the embodiment of the present disclosure, the initial VQ-VAE model is trained based on the sample RNA and the sample structure information of the sample RNA, and the model parameters are adjusted using the loss function until a trained VQ-VAE model is obtained. This allows the model to automatically process the structural diagram of the received RNA, improve processing efficiency, and obtain a highly accurate discretized structural representation of the RNA.

[0079] Figure 4 A schematic diagram of a process for obtaining an RNA representation provided in an embodiment of the present disclosure.

[0080] like Figure 4 As shown, the method for obtaining the RNA may include:

[0081] S401, obtaining a discretized representation of the RNA structure based on the RNA structure information.

[0082] S402: Obtain a discretized representation of the RNA sequence based on the RNA sequence information.

[0083] The relevant contents of steps S401-S402 can be found in the above embodiment and will not be repeated here.

[0084] S403: Input the structure discretization representation and the sequence discretization representation into the pre-trained target language model, and the target language model fuses the structure discretization representation and the sequence discretization representation to obtain the target representation of RNA.

[0085] In some implementations, a target language model can be pre-trained and used to fuse the structure discretization representation and the sequence discretization representation to obtain the target RNA representation. Alternatively, structural features and sequence features can be extracted from the structure discretization representation and the sequence discretization representation, and the structural features and sequence features can be fused to obtain the target RNA representation.

[0086] Optionally, the features of the structural discretization representation and the sequence discretization representation can be fused through methods such as feature concatenation, feature weighting, and attention mechanism.

[0087] According to the method for obtaining the representation of RNA provided by the embodiment of the present disclosure, by obtaining the discretized representation of the structure of RNA and the discretized representation of the sequence of RNA, and using the pre-trained target language model, the discretized representation of the structure of RNA and the discretized representation of the sequence of RNA are fused to obtain the target representation of RNA, thereby realizing the fusion of the structural information and sequence information of RNA, providing more comprehensive and richer information, helping to more accurately describe the characteristics of RNA molecules, and improving the accuracy and reliability of predicting RNA functions. Furthermore, by fusing the structural information and sequence information of RNA for representation learning, a theoretical basis can be provided for the diagnosis and treatment of diseases, and it can help to learn the interaction mechanism between RNA and protein, thereby accelerating the development of drugs.

[0088] Based on the above embodiments, the present disclosure can explain the training process of the target language model, such as Figure 5 As shown, the training process of the target language model may include:

[0089] S501 , performing word segmentation on the sample discretized representation of the sample RNA to obtain structural word segmentation.

[0090] S502: Segment the discretized representation of the sample to obtain sequence segmentations.

[0091] In some implementations, each structure in the sample discretized representation of the sample RNA can be segmented to obtain structural segmentations of the sample RNA. For example, each hairpin loop, stem, or ring can be considered a structural segmentation. Sequences are extracted from the discretized representation of the sample RNA, and the extracted subsequences are used as sequence segmentations.

[0092] S503: Based on the structural segmentation and the sequence segmentation, the language model is trained until the training is completed to obtain the target language model.

[0093] In some implementations, the structural segmentation and sequence segmentation can be spliced ​​together, and the spliced ​​segmentation sequence can be used to train the language model until the training is completed to obtain the target language model, so as to improve the target language model's ability to understand RNA structure and sequence, and further improve the accuracy of RNA structure prediction using the target language model.

[0094] Alternatively, the structure segmentation words and the sequence segmentation words can be directly concatenated to obtain a first concatenated segmentation sequence, and the language model can be trained based on the first concatenated segmentation sequence. For example, if the structure segmentation words are structure tokens and the sequence segmentation words are sequence tokens, then the first concatenated segmentation sequence is structure tokens + sequence tokens.

[0095] Optionally, the structural segmentation words and the sequence segmentation words can be spliced ​​to obtain a first spliced ​​segmentation sequence, and the structural segmentation words and the sequence segmentation words in the first spliced ​​segmentation sequence can be cross-masked to obtain a second spliced ​​segmentation sequence, and then the language model can be trained based on the second spliced ​​segmentation sequence.

[0096] Optionally, the training can be completed when the number of training times reaches a set number, and the accuracy of the model output result greater than the accuracy threshold can be regarded as the training completion.

[0097] According to the RNA representation acquisition method provided by the embodiment of the present disclosure, by performing word segmentation on the discretized representation of the sample to obtain structural word segmentation and sequence word segmentation, and splicing the structural word segmentation and sequence word segmentation, and using the spliced ​​word segmentation sequence to train the language model, a target language model can be obtained, which can predictably reduce the amount of calculation of the model during the training process, help to accelerate the training process and improve the performance of the model, and further improve the accuracy of the target language model in predicting specific functional regions of RNA sequences and predicting gene expression levels.

[0098] Figure 6 A schematic diagram of a process for obtaining an RNA representation provided in an embodiment of the present disclosure.

[0099] like Figure 6 As shown, the method for obtaining the RNA may include:

[0100] S601, generating a structure diagram of the RNA based on the structure information of the RNA.

[0101] S602: Determine a discretized representation of the RNA structure based on the RNA structure graph and a pre-constructed discretized vocabulary.

[0102] S603: Obtain a discretized representation of the RNA sequence according to the RNA sequence information.

[0103] S604: Input the structure discretization representation and the sequence discretization representation into the pre-trained target language model, and the target language model fuses the structure discretization representation and the sequence discretization representation to obtain the target representation of RNA.

[0104] The relevant contents of steps S601-S604 can be found in the above embodiment and will not be repeated here.

[0105] According to the method for obtaining RNA representation provided by the embodiment of the present disclosure, by obtaining the discretized representation of RNA structure and the discretized representation of RNA sequence, and fusing the discretized representation of RNA structure and the discretized representation of RNA sequence, a target representation of RNA can be obtained, thereby achieving the fusion of RNA structural information and sequence information, providing more comprehensive and richer information, helping to more accurately describe the characteristics of RNA molecules, and improving the accuracy and reliability of predicting RNA functions. Furthermore, by fusing RNA structural information and sequence information for representation learning, a theoretical basis can be provided for the diagnosis and treatment of diseases, and it can help to study the interaction mechanism between RNA and protein, thereby accelerating the drug development process.

[0106] like Figure 7 Schematic diagram of the structure of RNA target representation is shown. Figure 7 The method includes a structure representation module and a target representation module. The structure representation module includes a first VQ-VAE model, and the target representation module includes a target language model. The structure information of RNA is input into the structure representation module, and the first VQ-VAE model in the structure representation module encodes and matches the structure information to determine the structure discretization representation of RNA. The sequence discretization representation of RNA is obtained, and the sequence discretization representation of RNA and the structure discretization representation of RNA are input into the target representation module, and the target language model in the target representation module fuses the sequence discretization representation of RNA and the structure discretization representation of RNA to obtain the target representation of RNA.

[0107] Corresponding to the RNA representation acquisition methods provided in the above-mentioned embodiments, an embodiment of the present disclosure also provides an RNA representation acquisition device. Since the RNA representation acquisition device provided in the embodiment of the present disclosure corresponds to the RNA representation acquisition methods provided in the above-mentioned embodiments, the implementation methods of the above-mentioned RNA representation acquisition methods are also applicable to the RNA representation acquisition device provided in the embodiment of the present disclosure, and will not be described in detail in the following embodiments.

[0108] Figure 8 A schematic diagram of the structure of an RNA representation acquisition device provided in an embodiment of the present disclosure.

[0109] like Figure 8 As shown, the RNA representation acquisition device 800 of the embodiment of the present disclosure includes a first acquisition module 801, a second acquisition module 802 and a fusion module 803.

[0110] A first acquisition module 801 is used to obtain a discretized representation of the structure of the RNA based on the structure information of the RNA;

[0111] A second acquisition module 802 is configured to acquire a discretized representation of the sequence of the RNA according to the sequence information of the RNA;

[0112] The fusion module 803 is used to fuse the structure discretization representation and the sequence discretization representation of the RNA to obtain the target representation of the RNA.

[0113] In one embodiment of the present disclosure, the first acquisition module 801 is further configured to: generate a structural diagram of the RNA according to the structural information of the RNA; and determine a discretized structural representation of the RNA based on the structural diagram of the RNA and a pre-constructed discretized vocabulary.

[0114] In one embodiment of the present disclosure, the first acquisition module 801 is further used to: in response to the structural information including the secondary structure information and the tertiary structure information, construct a first structural diagram of the RNA according to the secondary structure information, and construct a second structural diagram of the RNA according to the tertiary structure information.

[0115] In one embodiment of the present disclosure, the first acquisition module 801 is further used to: obtain a latent space representation of the RNA based on the structural diagram of the RNA; quantize the latent space representation to obtain a quantization vector, and obtain a structural discretization representation with the highest similarity to the quantization vector from a pre-constructed discretization vocabulary; and determine the structural discretization representation with the highest similarity as the structural discretization representation of the RNA.

[0116] In one embodiment of the present disclosure, the first acquisition module 801 is further used to: input the structural graph of the RNA into a pre-trained first vector quantization variational autoencoder VQ-VAE model, and the encoder in the first VQ-VAE model encodes the structural graph to obtain the latent space representation; the quantization layer in the first VQ-VAE model quantizes the latent space representation, and matches the quantized vector in the discretized vocabulary to obtain the structural discretized representation of the RNA.

[0117] In one embodiment of the present disclosure, the first acquisition module 801 is further used to: acquire sample RNA and acquire sample structure information of the sample RNA; input the sample structure information into the initial VQ-VAE model, and the encoder in the VQ-VAE model encodes the sample structure information to obtain a sample latent space representation; the quantization layer in the VQ-VAE model quantizes the sample latent space representation, and matches the sample quantization vector in a reference discretization vocabulary to obtain a sample discretization sequence of the sample RNA; the decoder in the VQ-VAE model decodes the sample discretization sequence to obtain reduced RNA; adjust the model parameters of the VQ-VAE model according to the sample RNA and the reduced RNA, as well as the sample latent space table and the sample discretization sequence, and continue training until the end condition is met to obtain a second VQ-VAE model, wherein the first VQ-VAE model only includes the encoder and quantization layer in the second VQ-VAE model.

[0118] In one embodiment of the present disclosure, the first acquisition module 801 is further used to: determine the reconstruction loss based on the sample RNA and the reduced RNA; determine the quantization loss based on the sample latent space table and the sample discretization sequence; and adjust the model parameters of the VQ-VAE model based on the reconstruction loss and the quantization loss.

[0119] In one embodiment of the present disclosure, the fusion module 803 is further used to: input the structural discretization representation and the sequence discretization representation into a pre-trained target language model, and the target language model fuses the structural discretization representation and the sequence discretization representation to obtain the target representation of the RNA.

[0120] In one embodiment of the present disclosure, the fusion module 803 is further used to: perform word segmentation on the sample discretized representation of the sample RNA to obtain structural word segmentation; perform word segmentation on the sample discretized representation to obtain sequence word segmentation; and train the language model based on the structural word segmentation and sequence word segmentation until the training is completed to obtain the target language model.

[0121] In one embodiment of the present disclosure, the fusion module 803 is further configured to: concatenate the structural segmentation words and the sequence segmentation words to obtain a first concatenated segmentation sequence; and train the language model based on the first concatenated segmentation sequence.

[0122] In one embodiment of the present disclosure, the fusion module 803 is further used to: splice the structural segmentation words and the sequence segmentation words to obtain a first spliced ​​segmentation sequence; cross-mask the structural segmentation words and the sequence segmentation words in the first spliced ​​segmentation sequence to obtain a second spliced ​​segmentation sequence; and train the language model based on the second spliced ​​segmentation sequence.

[0123] According to the RNA representation acquisition device provided by the embodiment of the present disclosure, by obtaining the RNA structure discretization representation and the RNA sequence discretization representation, and fusing the RNA structure discretization representation and the RNA sequence discretization representation, the target representation of the RNA can be obtained, thereby realizing the fusion of RNA structure information and sequence information, providing more comprehensive and richer information, helping to more accurately describe the characteristics of RNA molecules, and improving the accuracy and reliability of predicting RNA functions. Furthermore, by fusing RNA structure information and sequence information for representation learning, a theoretical basis can be provided for the diagnosis and treatment of diseases, and it can help to study the interaction mechanism between RNA and protein, thereby accelerating the drug development process.

[0124] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0125] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0126] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0127] like Figure 9As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to computer programs / instructions stored in a read-only memory (ROM) 902 or computer programs / instructions loaded from a storage unit 906 into a random access memory (RAM) 903. Various programs and data required for the operation of the device 900 can also be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0128] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906 such as a keyboard, a mouse, etc.; an output unit 907 such as various types of displays, speakers, etc.; a storage unit 908 such as a magnetic disk, an optical disk, etc.; and a communication unit 909 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0129] The computing unit 901 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as the RNA representation acquisition method. For example, in some embodiments, the RNA representation acquisition method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 906. In some embodiments, part or all of the computer program / instructions can be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program / instructions are loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the RNA representation acquisition method described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the RNA representation acquisition method by any other appropriate means (e.g., by means of firmware).

[0130] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs / instructions that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0131] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0132] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0133] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0134] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0135] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship arises through computer programs / instructions running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0136] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0137] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for obtaining a representation of ribonucleic acid (RNA), wherein: The method comprises: Generating a structural diagram of the RNA according to the structural information of the RNA; Obtaining a discretized representation of the structure of the RNA based on the structure graph of the RNA and a pre-constructed discretized vocabulary; Obtaining a discretized representation of the RNA sequence according to the RNA sequence information, wherein the discretized representation of the RNA sequence is a process of converting each nucleic acid in the RNA sequence into a numerical value or symbol; fusing the structure discretization representation of the RNA and the sequence discretization representation to obtain a target representation of the RNA; The step of determining the discretized representation of the RNA structure based on the RNA structure graph and a pre-constructed discretized vocabulary includes: Obtaining a latent space representation of the RNA based on the structure graph of the RNA; quantizing the latent space representation to obtain a quantized vector, and obtaining a discretized representation of the structure having the highest similarity to the quantized vector from a pre-constructed discretized vocabulary; The structural discretization representation with the highest similarity is determined as the structural discretization representation of the RNA.

2. The method according to claim 1, wherein The RNA structural information includes at least one of secondary structure information and tertiary structure information of the RNA, and generating a structural diagram of the RNA based on the RNA structural information includes: In response to the structural information including the secondary structure information and the tertiary structure information, a first structural graph of the RNA is constructed according to the secondary structure information, and a second structural graph of the RNA is constructed according to the tertiary structure information.

3. The method according to claim 2, wherein: The method comprises: Inputting the structure graph of the RNA into a pre-trained first VQ-VAE model, and encoding the structure graph by an encoder in the first VQ-VAE model to obtain the latent space representation, wherein the first VQ-VAE model is a vector quantized variational autoencoder; The quantization layer in the first VQ-VAE model quantizes the latent space representation, and matches the quantized vector in the discretized vocabulary to obtain a structural discretized representation of the RNA.

4. The method according to claim 3, wherein: The training process of the first VQ-VAE model includes: Obtaining sample RNA and obtaining sample structure information of the sample RNA; Inputting the sample structure information into an initial VQ-VAE model, and encoding the sample structure information by an encoder in the VQ-VAE model to obtain a sample latent space representation; The sample latent space representation is quantized by a quantization layer in the VQ-VAE model, and the sample quantization vector is matched in a reference discretization vocabulary to obtain a sample discretization sequence of the sample RNA; The decoder in the VQ-VAE model decodes the discretized sequence of the sample to obtain a reduced RNA; According to the sample RNA and the reduced RNA, as well as the sample latent space table and the sample discretization sequence, the model parameters of the VQ-VAE model are adjusted, and training is continued until the end condition is met to obtain a second VQ-VAE model, wherein the first VQ-VAE model only includes the encoder and quantization layer in the second VQ-VAE model.

5. The method according to claim 4, wherein The adjusting of model parameters of the VQ-VAE model according to the sample RNA and the reduced RNA, the sample latent space table and the sample discretization sequence includes: determining a reconstruction loss based on the sample RNA and the reduced RNA; Determining a quantization loss according to the sample latent space table and the sample discretization sequence; The model parameters of the VQ-VAE model are adjusted according to the reconstruction loss and the quantization loss.

6. The method according to claim 1, wherein The step of fusing the structure discretization representation and the sequence discretization representation of the RNA to obtain a target representation of the RNA includes: The structure discretization representation and the sequence discretization representation are input into a pre-trained target language model, and the target language model fuses the structure discretization representation and the sequence discretization representation to obtain a target representation of the RNA.

7. The method according to claim 6, wherein: The training process of the target language model includes: Segment each structure in the sample discretization representation of the sample RNA to obtain structural segmentation; Performing sequence extraction on the discretized representation of the sample to obtain sequence segmentation; Based on the structural word segmentation and sequence word segmentation, the language model is trained until the training is completed to obtain the target language model.

8. The method according to claim 7, wherein: The training of the language model based on the structural segmentation and the sequence segmentation includes: Splicing the structural participles and the sequence participles to obtain a first spliced ​​participle sequence; The language model is trained based on the first concatenated word segmentation sequence.

9. The method according to claim 7, wherein: The training of the language model based on the structural segmentation and the sequence segmentation includes: Splicing the structural participles and the sequence participles to obtain a first spliced ​​participle sequence; Cross-masking the structural segmentation words and the sequence segmentation words in the first concatenated segmentation sequence to obtain a second concatenated segmentation sequence; The language model is trained based on the second concatenated word segmentation sequence.

10. A device for obtaining RNA expression, wherein: The device comprises: A first acquisition module is configured to generate a structure graph of the RNA according to the structure information of the RNA; and obtain a discretized representation of the structure of the RNA based on the structure graph of the RNA and a pre-constructed discretized vocabulary; A second acquisition module is configured to acquire a sequence discretization representation of the RNA according to the sequence information of the RNA, wherein the sequence discretization representation of the RNA is a process of converting each nucleic acid in the RNA sequence into a numerical value or symbol; a fusion module, configured to fuse the structure discretization representation of the RNA and the sequence discretization representation to obtain a target representation of the RNA; The first acquisition module is further configured to: Obtaining a latent space representation of the RNA based on the structure graph of the RNA; quantizing the latent space representation to obtain a quantized vector, and obtaining a discretized representation of the structure having the highest similarity to the quantized vector from a pre-constructed discretized vocabulary; The structural discretization representation with the highest similarity is determined as the structural discretization representation of the RNA.

11. The device according to claim 10, wherein The first acquisition module is further configured to: In response to the structural information including secondary structure information and tertiary structure information, a first structural graph of the RNA is constructed according to the secondary structure information, and a second structural graph of the RNA is constructed according to the tertiary structure information.

12. The device according to claim 11, wherein The first acquisition module is further configured to: Inputting the structure graph of the RNA into a pre-trained first VQ-VAE model, and encoding the structure graph by an encoder in the first VQ-VAE model to obtain the latent space representation, wherein the first VQ-VAE model is a vector quantized variational autoencoder; The quantization layer in the first VQ-VAE model quantizes the latent space representation, and matches the quantized vector in the discretized vocabulary to obtain a structural discretized representation of the RNA.

13. The device according to claim 11, wherein The first acquisition module is further configured to: Obtaining sample RNA and obtaining sample structure information of the sample RNA; Inputting the sample structure information into an initial VQ-VAE model, and encoding the sample structure information by an encoder in the VQ-VAE model to obtain a sample latent space representation; The sample latent space representation is quantized by a quantization layer in the VQ-VAE model, and the sample quantization vector is matched in a reference discretization vocabulary to obtain a sample discretization sequence of the sample RNA; The decoder in the VQ-VAE model decodes the discretized sequence of the sample to obtain a reduced RNA; According to the sample RNA and the reduced RNA, as well as the sample latent space table and the sample discretization sequence, the model parameters of the VQ-VAE model are adjusted, and training is continued until the end condition is met to obtain a second VQ-VAE model, where the first VQ-VAE model only includes the encoder and quantization layer in the second VQ-VAE model.

14. The device according to claim 13, wherein The first acquisition module is further configured to: determining a reconstruction loss based on the sample RNA and the reduced RNA; Determining a quantization loss according to the sample latent space table and the sample discretization sequence; The model parameters of the VQ-VAE model are adjusted according to the reconstruction loss and the quantization loss.

15. The device according to claim 10, wherein The fusion module is further configured to: The structure discretization representation and the sequence discretization representation are input into a pre-trained target language model, and the target language model fuses the structure discretization representation and the sequence discretization representation to obtain a target representation of the RNA.

16. The device according to claim 15, wherein The fusion module is further configured to: Segment each structure in the sample discretization representation of the sample RNA to obtain structural segmentation; Performing sequence extraction on the discretized representation of the sample to obtain sequence segmentation; Based on the structural word segmentation and sequence word segmentation, the language model is trained until the training is completed to obtain the target language model.

17. The device according to claim 16, wherein The fusion module is further configured to: Splicing the structural participles and the sequence participles to obtain a first spliced ​​participle sequence; The language model is trained based on the first concatenated word segmentation sequence.

18. The device according to claim 16, wherein The fusion module is further configured to: Splicing the structural participles and the sequence participles to obtain a first spliced ​​participle sequence; Cross-masking the structural segmentation words and the sequence segmentation words in the first concatenated segmentation sequence to obtain a second concatenated segmentation sequence; The language model is trained based on the second concatenated word segmentation sequence.

19. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.

20. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.

21. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • MicroRNA prediction method based on supervised self-organizing mapping neural network

    CN111477271A

  • Protein expression model pre-training and protein interaction prediction method and device

    CN114333982A