Ribonucleic acid binding site prediction methods and related devices
By combining the sequence and structural information of ribonucleic acid (RNA), the functional correlation of nucleotides in RNA is analyzed, which solves the problem of low accuracy in binding site prediction in existing technologies and achieves higher accuracy in binding site prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, it is difficult to accurately analyze the role of nucleotides in ribonucleic acid based on a single dimension of nucleotide sequence or structural distribution, resulting in low accuracy in predicting binding sites.
By combining the sequence and structural information of nucleotides in ribonucleic acid, the correlation between the two is analyzed through a first binding site prediction model to generate the probability of a nucleotide acting as a binding site, thereby determining the binding site.
It improves the accuracy of nucleotide binding site prediction, provides more accurate technical support, and promotes the development of related technical fields.
Smart Images

Figure CN122117015A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to methods and devices for predicting ribonucleic acid binding sites. Background Technology
[0002] Ribonucleic acid (RNA) is a biological macromolecule present in all living organisms. It serves as a carrier of genetic information, participating in the encoding, decoding, regulation, and expression of genes. Therefore, the occurrence and development of some diseases are closely related to RNA. To treat these RNA-related diseases, researchers have developed various drug molecules that can directly bind to RNA to exert therapeutic effects.
[0003] Therefore, identifying nucleotides that can serve as molecular binding sites within RNA is crucial for RNA-based drug therapy. Related technologies primarily analyze the role of nucleotides based on their sequence distribution within RNA or their structural distribution within the RNA structure, thereby predicting whether a nucleotide can act as a binding site.
[0004] However, the sequence distribution and structural distribution of nucleotides in ribonucleic acid (RNA) usually jointly determine the role of nucleotides in RNA. Predicting based on a single dimension (sequence dimension or structural dimension) makes it difficult to accurately analyze the role of nucleotides, resulting in low accuracy in predicting whether nucleotides can act as binding sites. Summary of the Invention
[0005] To address the aforementioned technical problems, this application provides a method for predicting ribonucleic acid (RNA) binding sites, which can accurately predict binding sites in RNA.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] In a first aspect, embodiments of this application disclose a method for predicting ribonucleic acid binding sites, the method comprising:
[0008] Obtain the sequence information corresponding to the ribonucleic acid, wherein the sequence information is used to identify the arrangement of multiple nucleotides constituting the ribonucleic acid in the ribonucleic acid;
[0009] Obtain the structural information corresponding to the ribonucleic acid, wherein the structural information is used to identify the ribonucleic acid structure composed of the plurality of nucleotides;
[0010] The sequence information and the structural information are input into a first binding site prediction model to generate the binding probabilities corresponding to the plurality of nucleotides respectively. The first binding site prediction model is used to analyze the functional correlation between the nucleotides expressed by the sequence information and the structural information respectively, and to generate the binding probabilities based on the correlation. The binding probabilities are used to characterize the probability of the corresponding nucleotide as a binding site. The binding sites are used to bind with molecules through the formation of interactions.
[0011] The binding sites among the plurality of nucleotides are determined based on the binding probability.
[0012] Secondly, embodiments of this application disclose a method for predicting ribonucleic acid binding sites, the method comprising:
[0013] A first sample ribonucleic acid is obtained, which is composed of multiple sample nucleotides. The first sample ribonucleic acid has corresponding sample binding site information, which is used to identify the binding sites in the multiple sample nucleotides.
[0014] Obtain the sample sequence information and sample structure information corresponding to the first sample ribonucleic acid. The sample sequence information is used to identify the sample arrangement of the plurality of sample nucleotides in the first sample ribonucleic acid. The sample structure information is used to identify the sample ribonucleic acid structure composed of the plurality of sample nucleotides.
[0015] The sample sequence information and the sample structure information are input into the initial prediction model to generate the undetermined binding probabilities corresponding to the plurality of nucleotides. The undetermined binding probabilities are used to characterize the probability that the corresponding sample nucleotide is a binding site. The initial prediction model is used to analyze the correlation between the nucleotide functions expressed by the sample sequence information and the sample structure information, and to generate the undetermined binding probabilities based on the correlation.
[0016] Based on the undetermined binding probability, undetermined binding site information is generated. The undetermined binding site information is used to identify the binding sites predicted by the initial prediction model in the plurality of sample nucleotides.
[0017] Based on the difference between the undetermined binding site information and the sample binding site information, the model parameters corresponding to the initial prediction model are adjusted to obtain a first binding site prediction model, which is used to predict binding sites in ribonucleic acid.
[0018] Thirdly, embodiments of this application disclose a ribonucleic acid binding site prediction device, the device comprising a first acquisition unit, a second acquisition unit, a first generation unit, and a first determination unit:
[0019] The first acquisition unit is used to acquire the sequence information corresponding to the ribonucleic acid, wherein the sequence information is used to identify the arrangement of multiple nucleotides constituting the ribonucleic acid in the ribonucleic acid;
[0020] The second acquisition unit is used to acquire the structural information corresponding to the ribonucleic acid, wherein the structural information is used to identify the ribonucleic acid structure composed of the plurality of nucleotides;
[0021] The first generation unit is used to input the sequence information and the structural information into the first binding site prediction model to generate the binding probabilities corresponding to the plurality of nucleotides respectively. The first binding site prediction model is used to analyze the functional correlation between the nucleotides expressed by the sequence information and the structural information respectively, and generate the binding probability according to the correlation. The binding probability is used to characterize the probability of the corresponding nucleotide as a binding site. The binding site is used to bind with the molecule by forming an interaction.
[0022] The first determining unit is used to determine the binding sites among the plurality of nucleotides based on the binding probability.
[0023] In one possible implementation, the first binding site prediction model includes a feature extraction module and a probability prediction module. The feature extraction module is used to extract sequence features similar to the structural information from the sequence information and to extract structural features similar to the sequence information from the structural information. The probability prediction module is used to predict the binding probability based on the sequence features and the structural features. The sequence features similar to the structural information and the structural features similar to the sequence information are used to characterize the functional association between the nucleotides expressed by the sequence information and the structural information, respectively.
[0024] In one possible implementation, the feature extraction module includes a first extraction module and a second extraction module. The step of extracting sequence features similar to the structural information from the sequence information and extracting structural features similar to the sequence information from the structural information includes:
[0025] The first extraction module performs feature encoding on the sequence information based on the arrangement of the sequence information identifier to extract initial sequence features for characterizing the arrangement, and performs feature encoding on the structural information based on the ribonucleic acid structure of the structural information identifier to extract initial structural features for characterizing the ribonucleic acid structure.
[0026] Feature similarity analysis is performed on the initial sequence features and the initial structural features to obtain attention weights. The attention weights are used to characterize a first degree of similarity between the initial sequence features and the initial structural features. The first degree of similarity is positively correlated with the feature similarity, and the attention weights are positively correlated with the first degree of similarity.
[0027] The second extraction module extracts features from the initial sequence features based on the attention weight to generate the sequence features, and extracts features from the initial structural features based on the attention weight to generate the structural features. The attention weight is used to control the degree of information retention of the initial sequence features when generating the sequence features, and to control the degree of information retention of the initial structural features when generating the structural features. The attention weight is positively correlated with the degree of information retention.
[0028] In one possible implementation, the initial sequence feature is composed of the initial arrangement features corresponding to the plurality of nucleotides respectively, and the initial arrangement features are used to characterize the arrangement characteristics of the corresponding nucleotides in the nucleotide sequence corresponding to the ribonucleic acid;
[0029] The initial structural features are composed of the initial position distribution features corresponding to the plurality of nucleotides, and the initial position distribution features are used to characterize the position distribution characteristics of the corresponding nucleotides in the ribonucleic acid structure;
[0030] The attention weight is used to characterize the similarity between the initial arrangement features corresponding to the plurality of nucleotides and the initial position distribution features corresponding to the plurality of nucleotides, wherein the first nucleotide is any one of the plurality of nucleotides, and the first nucleotide corresponds to the first initial arrangement feature and the first initial position distribution feature. The step of extracting features from the initial sequence features based on the attention weight to generate the sequence features, and extracting features from the initial structural features based on the attention weight to generate the structural features, includes:
[0031] Based on the second similarity between the first initial arrangement feature represented by the attention weight and the initial position distribution features corresponding to the plurality of nucleotides, feature extraction is performed on the first initial arrangement feature to obtain the arrangement feature corresponding to the first nucleotide. The degree of information retention of the arrangement feature corresponding to the first nucleotide on the first initial arrangement feature is positively correlated with the second similarity. The arrangement features corresponding to the plurality of nucleotides are used to form the sequence feature.
[0032] Based on the third similarity between the first initial position distribution feature represented by the attention weight and the initial arrangement features corresponding to the plurality of nucleotides, feature extraction is performed on the first initial position distribution feature to obtain the position distribution feature corresponding to the first nucleotide. The degree to which the position distribution feature corresponding to the first nucleotide retains information about the first initial position distribution feature is positively correlated with the third similarity. The position distribution features corresponding to the plurality of nucleotides are used to form the structural feature.
[0033] In one possible implementation, the arrangement of the multiple nucleotides in the sequence features corresponds to the arrangement of the multiple nucleotides in the ribonucleic acid, and the positional distribution of the multiple nucleotides in the sequence features corresponds to the arrangement of the multiple nucleotides in the ribonucleic acid.
[0034] In one possible implementation, the first nucleotide is any one of the plurality of nucleotides, and the first determining unit is specifically used for:
[0035] Based on the fact that the binding probability corresponding to the first nucleotide is greater than a first probability threshold, the first nucleotide is used as the binding site among the plurality of nucleotides.
[0036] In one possible implementation, the first nucleotide is any one of the plurality of nucleotides, the prediction model is any one of the plurality of prediction models, and the plurality of prediction models use different training samples during model training. The first determining unit is specifically used for:
[0037] Based on the fact that the binding probability of the first nucleotide is greater than the second probability threshold, the result output by the first binding site prediction model is determined to be that the first nucleotide is a binding site among the plurality of nucleotides.
[0038] Based on the fact that more than a preset number of prediction models output the result that the first nucleotide is the binding site among the multiple nucleotides, the first nucleotide is used as the binding site among the multiple nucleotides.
[0039] In one possible implementation, the first acquisition unit is specifically used for:
[0040] Obtain the nucleotide sequence information corresponding to the ribonucleic acid, wherein the nucleotide sequence information includes multiple nucleotide identifiers for identifying the multiple nucleotides, and the arrangement of the multiple nucleotide identifiers in the nucleotide sequence information corresponds to the arrangement of the multiple nucleotides in the ribonucleic acid;
[0041] A start identifier is added at the beginning position of the nucleotide sequence information, and an end identifier is added at the end position of the nucleotide sequence information. The start identifier is used to identify the first nucleotide in the ribonucleic acid, and the end identifier is used to identify the last nucleotide in the ribonucleic acid.
[0042] Based on the mapping relationship between identifiers and digital codes, the identifiers in the nucleotide sequence information are converted into corresponding digital codes to obtain the sequence information.
[0043] In one possible implementation, the structural information includes proximity centrality and degree corresponding to the plurality of nucleotides, wherein the proximity centrality is used to characterize the degree of association between each nucleotide and other nucleotides in the ribonucleic acid structure, and the degree is used to characterize the importance of each nucleotide in the ribonucleic acid structure; the second acquisition unit is specifically used for:
[0044] Generate topological information corresponding to the ribonucleic acid, the topological information including multiple nodes and connection identifiers between nodes, the multiple nodes corresponding one-to-one with the multiple nucleotides, the connection identifiers between nodes corresponding to adjacent nucleotides in the sequence information, the connection identifiers being used to identify the association between nucleotides in the ribonucleic acid structure;
[0045] Based on the fact that there is a non-covalent interaction between any two nucleotides in the ribonucleic acid structure, or that the distance between the two nucleotides is less than a preset distance, the connection identifier is added between the nodes corresponding to the two nucleotides in the topology information;
[0046] Based on the topological information, the proximity centrality and degree corresponding to the plurality of nucleotides are determined respectively. The first nucleotide corresponds to the first node in the topological information. The proximity centrality corresponding to the first nucleotide is inversely correlated with the average number of connection identifiers between the first node and other nodes in the topological information. The degree corresponding to the first nucleotide is positively correlated with the number of connection identifiers connecting the first node in the topological information.
[0047] In one possible implementation, the structural information further includes structural property information corresponding to each of the plurality of nucleotides, wherein the structural property information corresponding to the first nucleotide is used to characterize the nucleotide properties generated based on the structure of the first nucleotide itself, and the nucleotide properties are used to affect the binding ability of the first nucleotide.
[0048] In one possible implementation, the structural property information corresponding to the first nucleotide includes any one or more combinations of the molecular mass, acidity coefficient, accessible surface area, and evolution conservation fraction of the first nucleotide.
[0049] Fourthly, embodiments of this application disclose a ribonucleic acid binding site prediction device, the device comprising a third acquisition unit, a fourth acquisition unit, a second generation unit, a third generation unit, and an adjustment unit:
[0050] The third acquisition unit is used to acquire a first sample ribonucleic acid, which is composed of multiple sample nucleotides. The first sample ribonucleic acid has corresponding sample binding site information, which is used to identify the binding sites in the multiple sample nucleotides.
[0051] The fourth acquisition unit is used to acquire sample sequence information and sample structure information corresponding to the first sample ribonucleic acid. The sample sequence information is used to identify the sample arrangement of the plurality of sample nucleotides in the first sample ribonucleic acid. The sample structure information is used to identify the sample ribonucleic acid structure composed of the plurality of sample nucleotides.
[0052] The second generation unit is used to input the sample sequence information and the sample structure information into the initial prediction model to generate the undetermined binding probabilities corresponding to the plurality of nucleotides respectively. The undetermined binding probabilities are used to characterize the probability that the corresponding sample nucleotide is a binding site. The initial prediction model is used to analyze the correlation between the nucleotide functions expressed by the sample sequence information and the sample structure information respectively, and generate the undetermined binding probabilities based on the correlation.
[0053] The third generation unit is used to generate undetermined binding site information based on the undetermined binding probability. The undetermined binding site information is used to identify the binding sites predicted by the initial prediction model in the plurality of sample nucleotides.
[0054] The adjustment unit is used to adjust the model parameters corresponding to the initial prediction model based on the difference between the information of the undetermined binding site and the information of the sample binding site, so as to obtain a first binding site prediction model. The first binding site prediction model is used to predict binding sites in ribonucleic acid.
[0055] In one possible implementation, the initial prediction model includes an initial feature extraction module and an initial probability prediction module. The initial feature extraction module is used to extract undetermined sequence features similar to the sample structural information from the sample sequence information, and to extract undetermined structural features similar to the sample sequence information from the sample structural information. The undetermined sequence features and the undetermined structural features similar to the sample sequence information are used to characterize the functional association of nucleotides expressed by the sample sequence information and the sample structural information, respectively. The initial probability prediction module is used to make predictions based on the undetermined sequence features and the undetermined structural features to generate the undetermined binding probability. The adjustment unit is specifically used for:
[0056] Based on the difference between the undetermined binding site information and the sample binding site information, the model parameters corresponding to the initial feature extraction module and the initial probability prediction module are adjusted to obtain the feature extraction module and the probability prediction module, which are used to constitute the first binding site prediction model.
[0057] In one possible implementation, the initial feature extraction module includes a first initial extraction module and a second initial extraction module. The step of extracting undetermined sequence features similar to the sample structure information from the sample sequence information, and extracting undetermined structural features similar to the sample sequence information from the sample structure information, includes:
[0058] The first initial extraction module encodes the sequence information based on the sample arrangement method identified by the sample sequence information, extracts the initial sequence features to characterize the sample arrangement method, and encodes the structural information based on the sample ribonucleic acid structure identified by the structural information, extracts the initial structural features to characterize the sample ribonucleic acid structure.
[0059] A feature similarity analysis is performed on the undetermined initial sequence features and the undetermined initial structural features to obtain undetermined attention weights. The undetermined attention weights are used to characterize the third similarity between the undetermined initial sequence features and the undetermined initial structural features. The third similarity is positively correlated with the undetermined feature similarity, and the undetermined attention weights are positively correlated with the first similarity.
[0060] The second initial extraction module extracts features from the undetermined initial sequence features based on the undetermined attention weights to generate the undetermined sequence features, and extracts features from the undetermined initial structural features based on the undetermined attention weights to generate the undetermined structural features. The undetermined attention weights are used to control the degree of information retention of the undetermined initial sequence features when generating the undetermined sequence features, and to control the degree of information retention of the undetermined initial structural features when generating the undetermined structural features. The undetermined attention weights are positively correlated with the degree of information retention.
[0061] The adjustment unit is specifically used for:
[0062] Based on the difference between the undetermined binding site information and the sample binding site information, the model parameters corresponding to the first initial extraction module, the second initial extraction module, and the initial probability prediction module are adjusted to obtain the first extraction module, the second extraction module, and the probability prediction module. The first extraction module and the second extraction module are used to constitute the feature extraction module.
[0063] In one possible implementation, the first initial extraction module includes a multi-layer feature extraction network. The multi-layer feature extraction network performs feature encoding on the sequence information of ribonucleic acid and extracts the corresponding sequence features. The multi-layer feature extraction network consists of a first extraction network and a second extraction network. The first extraction network is used to perform feature encoding based on the sequence information to obtain an intermediate output, and the second extraction network is used to perform feature encoding based on the intermediate output to obtain the sequence features.
[0064] The adjustment unit is specifically used for:
[0065] Based on the difference between the undetermined binding site information and the sample binding site information, the model parameters corresponding to the second extraction network, the second initial extraction module, and the initial probability prediction module are adjusted respectively.
[0066] In one possible implementation, the first binding site prediction model is any one of multiple prediction models, which are used to jointly predict the binding sites in the ribonucleic acid. The device further includes a fifth acquisition unit and a second determination unit.
[0067] The fifth acquisition unit is used to acquire multiple sample ribonucleic acids, each of which has corresponding sample binding site information.
[0068] The second determining unit is used to select from the plurality of sample ribonucleic acids to determine a plurality of sample ribonucleic acid sets, wherein the sample ribonucleic acids included in different sample ribonucleic acid sets are partially different, and the plurality of sample ribonucleic acid sets correspond one-to-one with the plurality of prediction models, wherein the first sample ribonucleic acid is any one of the sample ribonucleic acid sets corresponding to the first binding site prediction model.
[0069] Fifthly, embodiments of this application disclose a computer device, the computer device including a processor and a memory:
[0070] The memory is used to store computer programs and to transfer the computer programs to the processor;
[0071] The processor is configured to execute the ribonucleic acid binding site prediction method according to any one of the first aspects, or to execute the ribonucleic acid binding site prediction method according to any one of the second aspects, according to the instructions in the computer program.
[0072] In a sixth aspect, embodiments of this application disclose a computer-readable storage medium for storing a computer program, the computer program being used to execute the ribonucleic acid binding site prediction method according to any one of the first aspects, or to execute the ribonucleic acid binding site prediction method according to any one of the second aspects;
[0073] In a seventh aspect, embodiments of this application disclose a computer program product including a computer program, which, when run on a computer device, causes the computer device to execute the ribonucleic acid binding site prediction method according to any one of the first aspects, or to execute the ribonucleic acid binding site prediction method according to any one of the second aspects.
[0074] As can be seen from the above technical solution, this application can predict binding sites in ribonucleic acid (RNA) by combining two dimensions: RNA sequence and RNA structure. Information at the RNA sequence level can be characterized by constructing sequence information to identify the arrangement of multiple nucleotides, while information at the RNA structure level can be characterized by structural information to identify the RNA structure. After inputting the sequence and structural information into the first binding site prediction model, the model can analyze the correlation between the functions of the nucleotides expressed by the sequence and structural information, and predict the binding probability of multiple nucleotides as binding sites based on the functional correlation. Finally, the binding probabilities output by the first binding site prediction model are used to determine the binding sites among the multiple nucleotides. Since the role of nucleotides in ribonucleic acid (RNA) is influenced by both the arrangement of nucleotides in the RNA sequence and their positional distribution within the RNA structure, the correlation between the functions of nucleotides expressed by sequence information and structural information can more accurately characterize the role of each nucleotide in RNA. Based on this correlation, it is possible to more accurately analyze whether each nucleotide constituting the RNA can bind to a molecule when the RNA encounters it. This allows for a more accurate prediction of the probability that multiple nucleotides can serve as binding sites, improving the accuracy of binding site prediction. Consequently, this provides a rich technical foundation for binding site-based research and contributes to the technological development of related fields. Attached Figure Description
[0075] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0076] Figure 1 A schematic diagram illustrating a method for predicting ribonucleic acid binding sites in a practical application scenario, as provided in this application embodiment;
[0077] Figure 2 A flowchart illustrating a method for predicting ribonucleic acid binding sites provided in this application embodiment;
[0078] Figure 3 A flowchart illustrating a method for predicting ribonucleic acid binding sites provided in this application embodiment;
[0079] Figure 4 A schematic diagram illustrating a method for predicting ribonucleic acid binding sites provided in an embodiment of this application;
[0080] Figure 5 A schematic diagram illustrating a method for predicting ribonucleic acid binding sites provided in an embodiment of this application;
[0081] Figure 6 A schematic diagram illustrating a method for predicting ribonucleic acid binding sites provided in an embodiment of this application;
[0082] Figure 7 A schematic diagram illustrating a method for predicting ribonucleic acid binding sites provided in an embodiment of this application;
[0083] Figure 8 A schematic diagram illustrating a method for predicting ribonucleic acid binding sites provided in an embodiment of this application;
[0084] Figure 9 A schematic diagram illustrating a method for predicting ribonucleic acid binding sites provided in an embodiment of this application;
[0085] Figure 10 A schematic diagram illustrating a method for predicting ribonucleic acid binding sites provided in an embodiment of this application;
[0086] Figure 11 A schematic diagram illustrating a method for predicting ribonucleic acid binding sites provided in an embodiment of this application;
[0087] Figure 12 A schematic diagram illustrating a method for predicting ribonucleic acid binding sites provided in an embodiment of this application;
[0088] Figure 13 A schematic diagram illustrating a method for predicting ribonucleic acid binding sites provided in an embodiment of this application;
[0089] Figure 14 A flowchart illustrating a method for predicting ribonucleic acid binding sites in a practical application scenario, provided in this application embodiment;
[0090] Figure 15 A schematic diagram illustrating a method for predicting ribonucleic acid binding sites in a practical application scenario, as provided in this application embodiment;
[0091] Figure 16 A schematic diagram illustrating a method for predicting ribonucleic acid binding sites in a practical application scenario, as provided in this application embodiment;
[0092] Figure 17 A schematic diagram illustrating a method for predicting ribonucleic acid binding sites provided in an embodiment of this application;
[0093] Figure 18 A structural block diagram of a ribonucleic acid binding site prediction device provided in this application embodiment;
[0094] Figure 19A structural block diagram of a ribonucleic acid binding site prediction device provided in this application embodiment;
[0095] Figure 20 A structural diagram of a terminal provided in an embodiment of this application;
[0096] Figure 21 This is a structural diagram of a server provided in an embodiment of this application. Detailed Implementation
[0097] The embodiments of this application will now be described with reference to the accompanying drawings.
[0098] Binding sites in ribonucleic acid (RNA) refer to nucleotides that can bind to molecules; nucleotides are the basic building blocks of RNA. Binding sites have crucial applications in multiple technological fields. For example, they can be used for drug testing, utilizing the ability of binding sites to bind to drug molecules to test responses to various targeted drugs for RNA-related diseases and optimize treatment plans. Alternatively, the binding site's ability to bind to drug molecules can be used to develop targeted drugs, leading to the development of specific drugs for various RNA-related diseases. The method for binding site prediction in this application can also be used to construct functional modules in bioinformatics platforms, enabling researchers and developers to effectively predict binding sites, thereby allowing for efficient and high-volume processing of binding site prediction tasks.
[0099] In related technologies, the prediction of binding sites in ribonucleic acid (RNA) mainly relies on knowledge-based prediction methods and experience-based prediction methods. In knowledge-based prediction methods, researchers establish a database of multiple RNAs with identified binding sites. Then, they compare the sequence information of the RNA to be analyzed with the corresponding sequence information of the RNA in the database. The sequence information identifies the arrangement of multiple nucleotides that make up the RNA. Since the arrangement of nucleotides usually determines the evolutionary information carried by the nucleotides, it can determine the role of the nucleotides to a certain extent. Therefore, if the sequence information of two RNAs is similar, the roles of the nucleotides in the two RNAs are usually similar. Thus, the binding sites in the RNA to be analyzed can be determined based on the binding site information in the database.
[0100] In empirical prediction methods, researchers can assess the importance of each nucleotide within the RNA structure based on the correlation between multiple nucleotides within the RNA, thereby predicting binding sites. Based on experience analyzing nucleotide structures, it is generally known that the stronger the correlation between a nucleotide and other nucleotides in the structure, or the greater the number of nucleotides correlated with that nucleotide, the more important that nucleotide is in the RNA structure, and the more crucial its role, thus increasing its probability of serving as a binding site.
[0101] However, in ribonucleic acid (RNA), the role of nucleotides is usually determined by both the arrangement of nucleotides in the RNA sequence and the distribution of nucleotides in the RNA structure. Both dimensions of information have a crucial impact on the role of nucleotides. Therefore, in related technologies, it is difficult to accurately measure the role of nucleotides in RNA by analyzing either sequence information or structural information from a single perspective, and thus it is difficult to accurately predict whether nucleotides can serve as binding sites.
[0102] To address the aforementioned technical issues, this application provides a method for predicting ribonucleic acid (RNA) binding sites. This method inputs structural information characterizing the RNA structure and sequence information characterizing nucleotide arrangement into a first binding site prediction model. The model analyzes the functional correlations between the nucleotides expressed by the sequence and structural information. These functional correlations accurately characterize the roles of nucleotides within the RNA, allowing for a relatively accurate prediction of the probability of each nucleotide acting as a binding site. Furthermore, the model's output probability can be used to predict the binding sites within the RNA, achieving accurate prediction of binding sites and providing high-quality data support for subsequent binding site-based technological applications.
[0103] Understandably, this method can be applied to computer devices capable of predicting binding sites, such as terminal devices or servers. The method can be executed independently by a terminal device or server, or it can be applied to network scenarios where the terminal device and server communicate, executing in cooperation. Terminal devices can be mobile phones, tablets, laptops, desktop computers, etc. Terminal devices can also include various virtual reality devices, such as augmented reality (AR) devices like AR glasses and AR screens, and virtual reality (VR) devices like VR headsets. The server can be understood as an application server or a web server. In actual deployment, the server can be a standalone server, a cluster server, or a cloud server, etc.
[0104] To facilitate understanding of the technical solution provided in this application, the method for predicting ribonucleic acid binding sites provided in this application will be introduced next in conjunction with a practical application scenario.
[0105] See Figure 1 , Figure 1 This is a schematic diagram of a method for predicting ribonucleic acid binding sites in a practical application scenario provided by an embodiment of this application. In this practical application scenario, the computer device can be a server 101 with binding site prediction capabilities.
[0106] The ribonucleic acid (RNA) requiring binding site prediction consists of multiple nucleotides. A, C, G, and U are abbreviations for four nucleotides, which can also be considered nucleotide identifiers, corresponding to adenine (A), cytosine (C), guanine (G), and uracil (U), respectively. Server 101 can first obtain the sequence and structural information of the RNA. The sequence information characterizes the arrangement of the multiple nucleotides in the RNA, and the structural information characterizes the structure of the RNA composed of multiple nucleotides.
[0107] Then, server 101 inputs the sequence information and structural information into the first binding site prediction model to obtain the binding probabilities of each nucleotide. These binding probabilities characterize the probability that a nucleotide will act as a binding site. Specifically, the first binding site prediction model analyzes the arrangement characteristics of each nucleotide based on the nucleotide sequence information and the positional distribution characteristics of each nucleotide within the RNA structure based on the RNA structure represented by the sequence information. Therefore, when analyzing the correlation between sequence and structural information, it is possible to determine the relationship between the nucleotide function influenced by its arrangement characteristics and the nucleotide function influenced by its positional distribution characteristics. Since the arrangement and position of nucleotides within ribonucleic acid (RNA) jointly determine their function, this correlation allows for the analysis of the functional relationships between sequence and structural information. This enables the first binding site prediction model to more accurately analyze the role of nucleotides in RNA, and further, to accurately determine whether each nucleotide has the ability to bind to the target molecule when RNA encounters it. Consequently, it can accurately generate the binding probabilities for multiple nucleotides. These multiple nucleotides include nucleotides 1 through n, where the binding probability of nucleotide 1 represents the probability that nucleotide 1 is a binding site.
[0108] Furthermore, server 101 can determine the binding sites among n nucleotides that can bind to the molecule based on the binding probabilities corresponding to multiple nucleotides determined by the first binding site prediction model. Therefore, this application can accurately analyze the role of nucleotides in ribonucleic acid (RNA) by combining information from both the RNA sequence and structure using the first binding site prediction model, thereby accurately predicting binding sites in RNA. This provides high-quality and comprehensive technical support for various binding site-based technologies, such as aiding in the development of drugs targeting RNA.
[0109] Next, the technical solution provided in this application will be described in detail with reference to the accompanying drawings.
[0110] See Figure 2 , Figure 2 A flowchart of a method for predicting ribonucleic acid binding sites provided in this application embodiment is shown. In this embodiment, the computer device can be any of the above-mentioned computer devices with binding site prediction functions. The method includes:
[0111] S201: Obtain the sequence information corresponding to the ribonucleic acid.
[0112] Sequence information is used to identify the arrangement of multiple nucleotides that make up ribonucleic acid (RNA). Sequence analysis can then be used to analyze the arrangement characteristics of each nucleotide within the RNA, such as identifying which nucleotides are adjacent to it. The RNA used for binding site prediction can be any RNA; no specific limitation is made here.
[0113] In ribonucleic acid (RNA), the arrangement of nucleotides can affect the genetic and evolutionary information they carry. That is, it can affect the information carried by each nucleotide in RNA and the role it plays in the process of heredity and evolution. The same nucleotides usually carry different information in different arrangements, and therefore play different roles.
[0114] S202: Obtain the structural information corresponding to the ribonucleic acid.
[0115] Structural information is used to identify the structure of ribonucleic acid (RNA), which is composed of multiple nucleotides. This structural information allows analysis of the positional distribution of each nucleotide within the RNA structure, such as its importance and its correlation with other nucleotides. In RNA, the structure also influences the evolutionary information it carries and its function; that is, it affects the role of each nucleotide. When the same nucleotide is distributed differently within the RNA structure, its role will also differ. For example, among two identical nucleotides, if one nucleotide has a higher correlation with other nucleotides in the RNA structure, then that nucleotide usually plays a more important role.
[0116] It is important to emphasize that there are multiple methods for obtaining the aforementioned structural and sequence information, which are not limited here but will be described in detail below. For example, ribonucleic acid (RNA) can be sequenced in various ways to detect the nucleotide order and obtain the corresponding sequence information; and the structure of multiple nucleotides within RNA can be observed using tools such as high-powered microscopes, for example, to observe whether non-covalent interactions exist between nucleotides, thereby obtaining the corresponding structural information. In addition, there are many other methods for obtaining sequence and structural information, as described below.
[0117] Therefore, it is evident that the arrangement and positional distribution of each nucleotide in RNA interact with each other, jointly influencing the role of nucleotides within RNA. The same RNA sequence can have different effects depending on the RNA structure. Whether RNA can bind to a molecule depends on the presence of nucleotides capable of binding, i.e., the presence of binding sites. Therefore, accurate and comprehensive analysis of the role of nucleotides is crucial for accurately predicting binding sites in RNA. Currently, both knowledge-based and experience-based methods analyze only a single dimension of sequence or structural information, making accurate prediction of binding sites difficult. Even methods based on two dimensions of information for binding site prediction often simply splice the two dimensions together without considering their interaction. This can easily lead to erroneous analyses based on redundant, unrelated information, ultimately resulting in low accuracy in binding site prediction.
[0118] S203: Input sequence information and structural information into the first binding site prediction model to generate the binding probabilities corresponding to multiple nucleotides.
[0119] In this application, in order to achieve high-accuracy binding site prediction, a computer device can be trained to obtain a prediction model for binding site prediction. Taking the first binding site prediction model under this model as an example, the first binding site prediction model can be any model with the following model functions, and can include any model architecture that can implement the following model functions, without limitation here.
[0120] As can be seen from the above analysis, the interaction between ribonucleic acid (RNA) sequence and ribonucleic acid (RNA) structure is a key factor affecting nucleotide function. Therefore, in this application, the first binding site prediction model can be used to analyze the correlation between sequence information and structural information to analyze the functional correlation of nucleotides expressed by sequence information and structural information respectively, and generate binding probabilities based on the correlation. Since sequence information and structural information can respectively characterize the features of ribonucleic acid in the ribonucleic acid sequence and ribonucleic acid structure dimensions, the function of nucleotides can be analyzed from the ribonucleic acid sequence dimension through sequence information, and the function of nucleotides can be analyzed from the ribonucleic acid structure dimension through structural information. Thus, by combining the correlation between these two parts of information, the correlation between the functions of nucleotides in the two dimensions can be analyzed more accurately, thereby characterizing the role played by the nucleotides constituting ribonucleic acid. Based on this correlation information, computer equipment can more accurately analyze whether multiple nucleotides have the ability to bind to molecules, and finally more accurately determine the binding probabilities corresponding to multiple nucleotides. The binding probability is used to characterize the probability of the corresponding nucleotide acting as a binding site. Taking the first nucleotide as an example, the first nucleotide can be any one of multiple nucleotides, and the binding probability corresponding to the first nucleotide is used to characterize the probability that the first nucleotide is a binding site.
[0121] Conversely, information that is not correlated between the two dimensions will not be the focus of the first binding site prediction model. This uncorrelated information is difficult to characterize the correlation between the functions of nucleotides in the two dimensions. Since this uncorrelated information has low accuracy in characterizing nucleotide functions, it is redundant information in the field of binding site prediction. Therefore, this method can reduce the impact of this redundant information on binding site prediction, thereby avoiding the problem of low accuracy in binding site prediction due to over-analysis of redundant information, and further improving the accuracy of binding site prediction.
[0122] S204: Determine the binding sites in multiple nucleotides based on binding probability.
[0123] Based on the binding probabilities determined by the model, computer equipment can analyze which nucleotides among multiple nucleotides have a higher probability of becoming binding sites, thus enabling relatively accurate analysis of the binding sites among multiple nucleotides. These binding sites are used to bind to molecules. There are various methods for determining binding sites based on binding probabilities, which will be described in detail below and will not be elaborated upon here.
[0124] As can be seen from the above technical solutions, this application can predict binding sites in ribonucleic acid (RNA) by combining both the RNA sequence and the RNA structure. Since the role of nucleotides in RNA is influenced by the interaction between their arrangement in the RNA sequence and their position in the RNA structure, the correlation between the functions of nucleotides expressed by sequence information and structural information can more accurately characterize the function of each nucleotide in RNA. Based on this correlation, it is possible to more accurately analyze whether each nucleotide constituting the RNA can bind to a molecule when the RNA encounters it. This allows for a more accurate prediction of the probability that multiple nucleotides can serve as binding sites, improving the accuracy of binding site prediction. This provides a rich technical foundation for research based on binding sites and helps promote the technological development of related fields.
[0125] After identifying binding sites in multiple nucleotides through the above process, the processing equipment can develop molecular products targeting these binding sites. For example, drug molecules can be developed targeting these binding sites, enabling them to bind directionally to these sites and thus exert their effects on the RNA, thereby achieving gene-level disease treatment. Alternatively, hormone molecules can be developed targeting these binding sites, allowing them to bind directionally to these sites and influence gene function. Furthermore, the processing equipment can also develop molecules that are difficult to drug, allowing them to bypass the pharmaceutical manufacturing process and directly bind to the binding sites.
[0126] The aforementioned first binding site prediction model can be trained using the following method:
[0127] See Figure 3 , Figure 3 The flowchart of a method for predicting ribonucleic acid binding sites provided in this application is shown. In this embodiment, the computer device can be any computer device with model training capabilities for training the desired model. The method includes:
[0128] S301: Obtain the first sample of ribonucleic acid.
[0129] The first sample ribonucleic acid (RNA) can be any type of RNA. It consists of multiple sample nucleotides and has corresponding sample binding site information. This information identifies the binding sites within the multiple sample nucleotides. In other words, the binding sites identified by the sample binding site information are the actual binding sites within the sample RNA.
[0130] S302: Obtain the sample sequence information and sample structure information corresponding to the first sample ribonucleic acid.
[0131] Sample sequence information identifies the arrangement of multiple sample nucleotides within the first sample ribonucleic acid (RNA), while sample structure information identifies the structure of the RNA composed of these nucleotides. Sample sequence information characterizes the RNA's information in the RNA sequence dimension, and sample structure information characterizes its information in the RNA structure dimension. The methods for obtaining sample sequence and structure information can be the same as or different from those for obtaining sequence and structure information. When the methods are the same, the first binding site prediction model can more effectively utilize the model knowledge learned during training for information analysis, thus leading to more accurate prediction results.
[0132] S303: Input the sample sequence information and sample structure information into the initial prediction model to generate the undetermined binding probabilities corresponding to multiple nucleotides.
[0133] The initial prediction model can be any model capable of supporting this training method. Similar to the information analysis process in the model application described above, the initial prediction model is used to analyze the functional correlation between the nucleotides expressed by the sample sequence information and the sample structural information, and to generate the undetermined binding probability based on the correlation. The functional correlations analyzed here are those analyzed by the model itself during training. Before learning sufficient model knowledge, they may not necessarily represent actual functional correlations; therefore, the undetermined binding probability is also only a probability analyzed by the model itself, and not necessarily the actual binding probability corresponding to each sample nucleotide.
[0134] The undetermined binding probability is used to characterize the probability that the corresponding sample nucleotide will be a binding site. Taking the first sample nucleotide as an example, the first sample nucleotide can be any one of multiple sample nucleotides. The undetermined binding probability corresponding to the first sample nucleotide is used to characterize the probability that the first sample nucleotide will be a binding site. The probability here is the probability analyzed by the initial prediction model, and is not necessarily the actual probability.
[0135] S304: Generate information on undetermined binding sites based on undetermined binding probabilities.
[0136] Computer equipment can first determine the sample nucleotides that can serve as binding sites among multiple sample nucleotides based on the probability determined by the initial prediction model, and then obtain the information of the binding sites to be located. The information of the binding sites to be located is used to identify the binding sites predicted by the initial prediction model among multiple sample nucleotides.
[0137] S305: Based on the difference between the information on the undetermined binding site and the information on the sample binding site, adjust the model parameters corresponding to the initial prediction model to obtain the first binding site prediction model.
[0138] Since the binding sites identified by the sample binding site information are the actual binding sites in the sample ribonucleic acid, while the binding sites identified by the undetermined binding site information are the binding sites predicted based on the binding probability analyzed by the initial prediction model, the difference between the binding sites identified by the sample binding site information and the undetermined binding site information can characterize the accuracy of the initial prediction model in analyzing the binding probability. Furthermore, the computer device can adjust the model parameters corresponding to the initial prediction model based on this difference, so that the undetermined binding site information determined by the initial prediction model can gradually approach the sample binding site information. During this parameter adjustment process, the initial prediction model can effectively learn the model knowledge used for accurately analyzing the binding probability, transforming into a first binding site prediction model. This first binding site prediction model is used to predict binding sites in ribonucleic acid, that is, it can be used to predict the binding probability corresponding to each of the multiple nucleotides constituting ribonucleic acid to achieve the prediction of binding sites.
[0139] The technical details involved in this application will be described in detail below.
[0140] First, the probability prediction process involved in this application will be described in detail.
[0141] In one possible implementation, such as Figure 4 As shown, the first binding site prediction model may include a feature extraction module and a probability prediction module. The feature extraction module is used to extract sequence features similar to structural information from sequence information, and structural features similar to sequence information from structural information. These sequence features and structural features can characterize the related information portions in the sequence information and structural information. Through this feature extraction module, unrelated information portions in the sequence information and structural information can be effectively screened out, thereby reducing the impact of redundant information on binding probability analysis. Specifically, the sequence features similar to structural information and the structural features similar to sequence information are used to characterize the correlation between the nucleotide functions expressed by the sequence information and the structural information, respectively. Specifically, the sequence features similar to structural information can express the functional portions of nucleotide functions in the ribonucleic acid sequence dimension that are associated with the nucleotide functions in the ribonucleic acid structural dimension, and the structural features similar to sequence information can express the functional portions of nucleotide functions in the ribonucleic acid structural dimension that are associated with the nucleotide functions in the ribonucleic acid sequence dimension.
[0142] The probability prediction module is used to predict and generate binding probabilities based on sequence and structural features. Thus, through the cooperation of these two model modules, accurate prediction of binding probabilities based on the interrelated information in sequence and structural information can be achieved. This approach, on the one hand, breaks down the entire binding probability prediction process into two steps: feature extraction and feature-based probability prediction. This refines the prediction process, enabling more accurate analysis of structural and sequence information and ensuring the accuracy of binding probabilities. On the other hand, during model training, parameters can be adjusted separately for multiple model modules, refining the parameter tuning process and allowing the model to more effectively learn the model knowledge used for predicting binding probabilities.
[0143] Correspondingly, in one possible implementation, the computer device can train the first binding site prediction model, which includes the feature extraction module and the probability prediction module, in the following manner:
[0144] In this implementation, the initial prediction model includes an initial feature extraction module and an initial probability prediction module. The initial feature extraction module is the same as the feature extraction module before training, and the initial probability prediction module is the same as the probability prediction module before training. The initial feature extraction module is used to extract undetermined sequence features similar to sample structural information from sample sequence information, and to extract undetermined structural features similar to sample sequence information from sample structural information. The undetermined sequence features similar to sample structural information and the undetermined structural features similar to sample sequence information are used to characterize the functional correlation of nucleotides expressed by the sample sequence information and sample structural information, respectively.
[0145] The initial probability prediction module is used to predict based on the undetermined sequence features and undetermined structural features, generating undetermined combination probabilities. When executing step S305, the computer device can execute step S3051 (not shown in the figure), where step S3051 is a possible implementation of step S305, including:
[0146] S3051: Based on the difference between the information of the undetermined binding site and the information of the sample binding site, adjust the model parameters corresponding to the initial feature extraction module and the initial probability prediction module respectively to obtain the feature extraction module and the probability prediction module.
[0147] like Figure 5As shown, the computer device can adjust the model parameters corresponding to the initial feature extraction module and the initial probability prediction module based on the accuracy of the initial prediction model characterized by this difference. On the one hand, the initial feature extraction module can learn how to accurately extract sequence features similar to structural information from sequence information, and how to accurately extract structural features similar to sequence information from structural information. Simultaneously, it can learn how to perform feature extraction in a way that makes the extracted features more helpful for probability prediction. On the other hand, the initial probability prediction module can learn how to accurately predict the binding probability based on sequence features and structural features, thereby achieving more fine-grained parameter adjustment of the model, refining the model's learning process, and ultimately optimizing the model training effect. The feature extraction module and probability prediction module obtained through the training process are used to construct the first binding site prediction model.
[0148] Specifically, in one possible implementation, the first binding site prediction model can be implemented through the following process:
[0149] In this implementation, such as Figure 6 As shown, the feature extraction module can specifically include a first extraction module and a second extraction module. In the processes of extracting sequence features similar to structural information from sequence information and extracting structural features similar to sequence information from structural information, the computer device can use the first extraction module to encode the sequence information based on the arrangement of sequence information identifiers, extracting initial sequence features to characterize the arrangement, and to encode the structural information based on the ribonucleic acid structure of the structural information identifiers, extracting initial structural features to characterize the ribonucleic acid structure. That is, the first extraction module is used to extract initial sequence features to accurately characterize the nucleotide arrangement based on information similar to the ribonucleic acid sequence, and the second extraction module is used to extract initial structural features to accurately characterize the ribonucleic acid structure based on information similar to the ribonucleic acid structure. In this process, redundant information that is poorly characterized for nucleotide arrangement and ribonucleic acid structure can be filtered out first, improving feature simplification and helping to improve the accuracy of subsequent probabilistic analysis.
[0150] Understandably, in the feature dimension, the similarity between two features can usually characterize the degree of association between the two features. The higher the similarity, the higher the degree of association between the two features, and the more related information there is between the two features. The similar feature parts between the two features are the feature parts with a high degree of association.
[0151] Therefore, in this implementation, the computer device can perform feature similarity analysis on the initial sequence features and the initial structural features to obtain attention weights. The attention weights characterize the first degree of similarity between the initial sequence features and the initial structural features. The first degree of similarity is positively correlated with the feature similarity, and the attention weights are also positively correlated with the first degree of similarity. That is, the higher the feature similarity between the initial sequence features and the initial structural features, the greater the first degree of similarity between the two features, and thus the larger the determined attention weight. It should be emphasized that the correlation between the information here is set to fit the following technical solution. When the following technical solution changes, the correlation may also change, all of which fall within the technical scope of this application.
[0152] Through the second extraction module, the computer device can extract features from the initial sequence features based on attention weights to generate sequence features, and extract features from the initial structural features based on attention weights to generate structural features. Attention weights are used to control the degree of information retention of the initial sequence features when generating sequence features, and to control the degree of information retention of the initial structural features when generating structural features. Attention weights are positively correlated with the degree of information retention.
[0153] In other words, a larger attention weight indicates that more features in the initial sequence features are similar to the initial structural features. Therefore, when extracting features from the initial sequence features, the information retained is greater, and more information similar to the initial structural features is preserved in the sequence features. Conversely, a smaller attention weight indicates that more information in the initial sequence features is dissimilar to the initial structural features. When extracting features based on attention weight, more information in the initial sequence features will be filtered out, reducing redundant information unrelated to the initial structural features, thus improving the accuracy of subsequent probability predictions based on the sequence features. Similarly, for structural features, a larger attention weight retains more information similar to the initial structural features and initial sequence features; a smaller attention weight filters out more redundant information unrelated to the initial sequence features, improving the accuracy of probability predictions based on structural features.
[0154] Therefore, the above approach can be divided into two processes: firstly, extracting feature information from the original information for each information dimension to aid in probability prediction; and secondly, extracting features similar to those of the other dimension from the feature information of the two dimensions. These two extraction processes can enhance the filtering of redundant information, enabling the extracted feature information to more accurately and centrally represent the correlation between the two dimensions, thereby improving the accuracy of the probability prediction. Secondly, further subdividing the feature extraction process allows the first binding site prediction model to have more refined model parameters, further improving the accuracy of information processing and enabling it to learn richer model knowledge during training, thus improving the accuracy of the binding probability prediction.
[0155] The training process for the above model structure can be described as follows:
[0156] In one possible implementation, see Figure 7 The aforementioned initial feature extraction module may include a first initial extraction module and a second initial extraction module. The first initial extraction module is used to train the aforementioned first extraction module, and the second initial extraction module is used to train the aforementioned second extraction module. During training, when extracting undetermined sequence features similar to sample structural information from sample sequence information, and extracting undetermined structural features similar to sample sequence information from sample structural information, the computer device can use the first initial extraction module to perform feature encoding on the sequence information based on the sample arrangement method identified by the sample sequence information, extracting undetermined initial sequence features characterizing the sample arrangement method; and to perform feature encoding on the structural information based on the sample ribonucleic acid structure identified by the structural information, extracting undetermined initial structural features characterizing the sample ribonucleic acid structure. The undetermined initial sequence features are the feature information in the sample sequence information analyzed by the first initial extraction module that is helpful for predicting the binding probability. The undetermined initial structural features are the feature information in the sample structural information analyzed by the first initial extraction module that is helpful for the binding probability. The extraction of undetermined initial sequence features and undetermined initial structural features is the first aspect affecting the accuracy of binding probability prediction.
[0157] Then, the computer device can perform a feature similarity analysis on the initial sequence features and the initial structural features to obtain the attention weight. The attention weight is used to characterize the third degree of similarity between the initial sequence features and the initial structural features. The third degree of similarity is positively correlated with the feature similarity, and the attention weight is positively correlated with the first degree of similarity. That is, the higher the feature similarity between the two features analyzed by the first initial extraction module, the larger the determined attention weight, and the higher the characterized third degree of similarity. The accuracy of the attention weight analysis also affects the accuracy of the final combination probability prediction, which is the second aspect affecting the accuracy of the combination probability prediction.
[0158] Finally, the second initial extraction module can extract features from the initial sequence features based on the undetermined attention weights to generate undetermined sequence features, and extract features from the initial structural features based on the undetermined attention weights to generate undetermined structural features. The undetermined attention weights control the degree of information retention of the initial sequence features when generating undetermined sequence features, and the undetermined structural features when generating undetermined structural features. The undetermined attention weights are positively correlated with the degree of information retention; that is, the larger the undetermined attention weights, the greater the degree of information retention of the undetermined sequence features relative to the initial sequence features, thus retaining more information similar to the initial structural features. Feature extraction based on undetermined attention weights is the third aspect affecting the accuracy of the combination probability prediction.
[0159] When executing step S3051, the computer device may execute step S30511 (not shown in the figure). Step S30511 is a possible implementation of step S3051, including:
[0160] S30511: Based on the difference between the information of the undetermined binding site and the information of the sample binding site, adjust the model parameters corresponding to the first initial extraction module, the second initial extraction module and the initial probability prediction module respectively to obtain the first extraction module, the second extraction module and the probability prediction module.
[0161] By adjusting the model parameters of the first initial feature extraction module based on differences, the first initial feature extraction module can learn how to extract sequence features from sequence information and structural features from structural information, extract features that are helpful for the binding probability, and accurately measure the feature similarity between two dimensions of features to determine accurate attention weights, thereby obtaining the first extraction module in the first binding site prediction model.
[0162] By adjusting the model parameters of the second initial feature extraction module based on differences, the module can learn how to accurately extract interrelated features from two dimensions based on attention weights. This allows these interrelated features to more accurately characterize the role of nucleotides in ribonucleic acid, resulting in more accurate binding probabilities in subsequent binding probability analysis based on these interrelated features. The second extraction module described above can be obtained by adjusting the model parameters of the second initial feature extraction module. The first and second extraction modules can be used to construct the aforementioned feature extraction module.
[0163] Therefore, based on this model architecture, computer devices can train the model through four processes: single-dimensional feature extraction, attention weight analysis, similar feature extraction between two dimensions, and feature-based binding probability prediction. This enables the model to learn multiple aspects of model knowledge for accurate binding probability prediction, thereby giving the trained first binding site prediction model a better binding probability prediction capability.
[0164] As mentioned above, the arrangement of nucleotides in the ribonucleic acid sequence dimension and their positional distribution in the ribonucleic acid structure dimension interact and jointly affect the role of nucleotides in ribonucleic acid. Therefore, in one possible implementation, in order to further improve the accuracy of predicting the binding probability of each nucleotide, computer equipment can perform feature extraction and analysis on a nucleotide-by-nucleotide basis.
[0165] In this implementation, the initial sequence feature obtained through feature encoding consists of initial arrangement features corresponding to multiple nucleotides. These initial arrangement features characterize the arrangement of the corresponding nucleotide within the nucleotide sequence corresponding to the ribonucleic acid. Taking the first nucleotide as an example, the first nucleotide corresponds to a first initial arrangement feature, which characterizes the arrangement of the first nucleotide within the nucleotide sequence corresponding to the ribonucleic acid. This allows the role of the first nucleotide to be characterized from the perspective of the ribonucleic acid sequence. For instance, the arrangement features can characterize the contextual information generated by nucleotides adjacent to the first nucleotide in the ribonucleic acid sequence, as well as the genetic information carried by the first nucleotide, all of which affect the role of the first nucleotide.
[0166] The initial structural features obtained through feature encoding consist of initial positional distribution features corresponding to multiple nucleotides. These initial positional distribution features characterize the positional distribution of the corresponding nucleotides within the ribonucleic acid (RNA) structure. The first initial positional distribution feature corresponding to the first nucleotide characterizes the positional distribution of the first nucleotide within the RNA structure, thus enabling the characterization of the role of the first nucleotide from a structural perspective. Various methods can be used to characterize positional distribution features, such as highlighting the importance of the position and its structural relationships with other nucleotides, which will be discussed in detail below.
[0167] In ribonucleic acid (RNA), the function of RNA is determined by multiple nucleotides. Therefore, not only do the arrangement and positional distribution of each nucleotide influence its own function, but the characteristics of multiple nucleotides in both dimensions also affect each other. For example, the positional distribution of nucleotide A may be the same in two ribonucleic acids, but when the arrangement of other nucleotides differs in the two ribonucleic acids, nucleotide A may ultimately play different roles in the two ribonucleic acids. Based on this, in this implementation, the computer device can analyze not only the correlation information between the initial arrangement characteristics and initial positional distribution characteristics corresponding to the same nucleotide, but also the correlation information between the initial arrangement characteristics and initial positional distribution characteristics corresponding to different nucleotides, in order to extract more features that can accurately characterize the function of nucleotides.
[0168] In this implementation, the attention weights determined through feature similarity analysis can be used to characterize the similarity between the initial arrangement features corresponding to multiple nucleotides and the initial positional distribution features corresponding to multiple nucleotides, i.e., when there are n nucleotides, such as Figure 8 As shown, the attention weight can be information of n x n dimensions, including the similarity between the initial arrangement feature corresponding to each nucleotide and the initial position distribution feature corresponding to each of the n nucleotides, as well as the similarity between the initial position distribution feature corresponding to each nucleotide and the initial arrangement feature corresponding to each of the n nucleotides.
[0169] When extracting features from initial sequence features based on attention weights to generate sequence features, and when extracting features from initial structural features based on attention weights to generate structural features, taking the first nucleotide as an example, the first nucleotide corresponds to both the first initial arrangement feature and the first initial position distribution feature. The computer device can extract features from the first initial arrangement feature based on the second similarity between the first initial arrangement feature represented by the attention weights and the initial position distribution features corresponding to multiple nucleotides, thus obtaining the arrangement feature corresponding to the first nucleotide. The degree to which the arrangement feature corresponding to the first nucleotide retains information from the first initial arrangement feature is positively correlated with the second similarity. That is, the greater the second similarity between the first initial arrangement feature and the initial position distribution features of multiple nucleotides, the more information in the first initial arrangement feature can be used to characterize the role of multiple nucleotides in ribonucleic acid. Therefore, when extracting the arrangement feature corresponding to the first nucleotide, more information from the first initial arrangement feature can be retained, enabling the arrangement feature corresponding to the first nucleotide to effectively characterize the arrangement information similar to the ribonucleic acid structure from the ribonucleic acid sequence dimension. This improves the accuracy of binding probability prediction and allows the arrangement feature corresponding to the first nucleotide to express more of the nucleotide functions expressed by the first initial arrangement feature.
[0170] The arrangement features corresponding to multiple nucleotides are used to form a sequence feature, so that the sequence feature can include information parts that are similar to the initial position distribution features in the initial arrangement features corresponding to multiple nucleotides.
[0171] Similarly, computer devices can extract features from the first initial position distribution feature based on the third similarity between the first initial position distribution feature represented by the attention weight and the initial arrangement features corresponding to multiple nucleotides, thus obtaining the position distribution feature corresponding to the first nucleotide. The degree to which the position distribution feature corresponding to the first nucleotide retains information from the first initial position distribution feature is positively correlated with the third similarity. That is, the greater the second similarity between the first initial position distribution feature and the initial arrangement features of multiple nucleotides, the more information in the first initial position distribution feature can be used to characterize the role of multiple nucleotides in ribonucleic acid. Therefore, when extracting the position distribution feature corresponding to the first nucleotide, more information in the first initial position distribution feature can be retained, enabling the position distribution feature corresponding to the first nucleotide to effectively characterize structural information similar to the ribonucleic acid sequence from the ribonucleic acid structural dimension, thereby improving the accuracy of binding probability prediction and enabling the position distribution feature corresponding to the first nucleotide to express more of the nucleotide functions expressed by the first initial position distribution feature.
[0172] The positional distribution features corresponding to multiple nucleotides are used to form structural features, and thus the structural features can include information portions similar to the initial arrangement features in the initial positional distribution features corresponding to multiple nucleotides.
[0173] Through the above methods, computer equipment can further refine the feature extraction process, performing feature extraction for each nucleotide separately. This enables accurate and effective feature extraction of each nucleotide in both the ribonucleic acid sequence and ribonucleic acid structure dimensions. It preserves as much similar information as possible in the two dimensions of features and filters out dissimilar redundant information. This allows the subsequent probability prediction process to more accurately analyze the role of each nucleotide in the ribonucleic acid based on the extracted sequence and structural features, and predict a more accurate binding probability.
[0174] like Figure 9 As shown, in Figure 8 Based on this, the computer device can extract features from the initial arrangement features corresponding to nucleotide 1 based on the similarity between the initial arrangement features corresponding to nucleotide 1 and the initial position distribution features corresponding to n nucleotides. Thus, based on n similarity levels, the information parts in the initial arrangement features that are similar to the initial position distribution features corresponding to n nucleotides can be extracted respectively, and finally integrated to obtain the arrangement features corresponding to nucleotide 1.
[0175] In one possible implementation, to enable the first binding site prediction model to more effectively combine information from both the ribonucleic acid sequence and ribonucleic acid structure dimensions for probabilistic prediction, the computer device can further integrate easily expressible ribonucleic acid sequence information into the features extracted in the above manner, thereby further improving the information expressive power of the features. For example, the arrangement of the aforementioned multiple nucleotides' respective positions in the sequence features can correspond to the arrangement of multiple nucleotides in the ribonucleic acid, and the arrangement of the positional distribution features of the multiple nucleotides' respective positions in the sequence features can correspond to the arrangement of multiple nucleotides in the ribonucleic acid. Figure 10 As shown, in the sequence features and structural features, the distribution order of the arrangement features and positional distribution features corresponding to the n nucleotides is consistent with the arrangement order of the n nucleotides in the ribonucleic acid sequence.
[0176] Therefore, when predicting binding probability based on structural and sequence features, the first binding site prediction model can enhance the understanding of the ribonucleic acid sequence by considering the arrangement of features corresponding to multiple nucleotides, thereby improving the accuracy of probability prediction. Simultaneously, this approach only limits the distribution of multiple features within the structural and sequence features, without increasing the number of features in the structural and sequence features, thus reducing the impact on the speed of binding probability prediction and ensuring the efficiency of binding probability prediction.
[0177] The above content mainly introduced the prediction process of binding probability. Next, we will explain in detail how to determine the binding site in ribonucleic acid based on the binding probability predicted by the first binding site prediction model.
[0178] First, in one possible implementation, the computer device can directly identify nucleotides with a high binding probability as binding sites capable of binding to the molecule. When performing step S204, the computer device can execute step S2041 (not shown in the figure), which is a possible implementation of step S204 and includes:
[0179] S2041: Based on the fact that the binding probability of the first nucleotide is greater than the first probability threshold, the first nucleotide is used as the binding site among multiple nucleotides.
[0180] The computer device can preset a first probability threshold, which is used to determine whether the binding probability of the corresponding nucleotide is relatively large. If the binding probability of the first nucleotide is greater than the first probability threshold, it indicates that the binding probability of the first nucleotide is relatively large. Therefore, the computer device can identify the first nucleotide as a binding site among multiple nucleotides, achieving efficient binding site determination based on binding probability. This determination process can be performed within the first binding site prediction model and integrated into its functionality, or it can be performed outside of the first binding site prediction model; this is not limited here.
[0181] The first probability threshold can be adjusted based on the required accuracy of binding site prediction. For example, when the required accuracy of binding site prediction is high, the first probability threshold can be set to a high probability value, so that only when the binding probability of the nucleotide is high enough will it be identified as a binding site, thus ensuring the accuracy of the output binding sites. When the required accuracy of binding site prediction is low, the first probability threshold can be set to a relatively low probability value, so that as long as the binding probability of the nucleotide is high, it will be identified as a binding site, thus avoiding missing any possible binding sites.
[0182] The above process can determine the binding site using the probability prediction results of a single prediction model. In another possible implementation, to further improve the prediction accuracy of the binding site, the computer device can also combine the probability prediction outputs of multiple prediction models to determine the binding site.
[0183] In this implementation, the first binding site prediction model can be any one of multiple prediction models, all of which are capable of predicting binding probability based on the association information between the two dimensions. Since the multiple prediction models use different training samples during training, the model knowledge learned by each model will differ, leading to variations in the way they analyze the association information between the two dimensions when predicting binding probability—that is, they will analyze the information from different perspectives.
[0184] When executing step S204, the computer device may execute steps S2042-S2043 (not shown in the figure). Steps S2042-S2043 are one possible implementation of step S204, including:
[0185] S2042: Based on the fact that the binding probability of the first nucleotide is greater than the second probability threshold, the result of the prediction model for the first binding site is determined to be that the first nucleotide is a binding site among multiple nucleotides.
[0186] The computer device can set a second probability threshold, which measures whether a nucleotide has a high probability of binding. Similar to the first probability threshold, the second probability threshold can also be adjusted based on the required accuracy of the binding site prediction; this is not limited here. The computer device can first analyze the prediction results of each prediction model for each nucleotide based on the binding probability predicted by each prediction model and the second probability threshold. Taking the first binding site prediction model as an example, if the binding probability of the first nucleotide predicted by the first binding site prediction model is greater than the second probability threshold, then it can be determined that in the output of the first binding site prediction model, the first nucleotide is a binding site among multiple nucleotides.
[0187] S2043: Based on the fact that more than a preset number of prediction models output results that the first nucleotide is the binding site among multiple nucleotides, the first nucleotide is used as the binding site among multiple nucleotides.
[0188] The computer device can be set to a preset number to measure whether a majority of models output the same result. For example, this preset number can be half the number of multiple prediction models. If the output of the preset number of prediction models all indicate that the first nucleotide is a binding site among multiple nucleotides, it can be concluded that most prediction models consider the first nucleotide to be a binding site with a high probability. Furthermore, since different prediction models use different information analysis perspectives when predicting binding probabilities, it can be concluded that under multiple information analysis perspectives, the first nucleotide has a high probability of functioning as a binding site. Therefore, the computer device can identify the first nucleotide as a binding site among multiple nucleotides.
[0189] This method allows computer equipment to combine multiple information analysis perspectives to more accurately predict the binding probability of nucleotides, further improving the prediction accuracy of binding sites and reducing the probability of incorrect binding site predictions due to the inadequacy of a single model's information analysis perspective. Furthermore, this method fully utilizes the powerful information analysis capabilities of multiple prediction models, improving resource utilization. Also, because this method combines the outputs of multiple prediction models to determine binding sites, even if the prediction accuracy of a single model is low, it can be compensated for by the joint analysis of multiple prediction models. This reduces the requirement for the prediction accuracy of a single model, thereby simplifying the model's complexity and enabling its application in a wider range of scenarios. It also helps reduce the need for high model training accuracy, thus lowering the difficulty of model training.
[0190] For example, in one possible implementation, the first binding site prediction model can be any one of multiple prediction models used to jointly predict binding sites in ribonucleic acid. As shown above, each prediction model can individually analyze the binding probability of nucleotides, and then the binding probabilities predicted by the multiple prediction models are combined to predict the binding sites in ribonucleic acid.
[0191] Because analyzing binding sites in RNA is challenging, the number of RNAs with clearly defined binding sites is limited, resulting in a small amount of data suitable for training predictive models. To train a predictive model capable of accurately predicting binding sites even with limited data, computer systems can combine multiple predictive models to jointly predict binding probabilities. This allows multiple models to analyze correlation information from different perspectives, compensating for the insufficient accuracy of a single predictive model.
[0192] The computer device can first acquire multiple sample ribonucleic acids, each with corresponding sample binding site information; that is, whether the nucleotides in the multiple sample ribonucleic acids are binding sites is known. To enable multiple prediction models to analyze information from various perspectives, the computer device needs to train different prediction models using different training samples.
[0193] To improve the utilization rate of training samples and train multiple different prediction models with a limited number of training samples, the computer can select from multiple sample ribonucleic acid (RNA) sets to determine multiple sets of sample RNAs. These sets contain partially different RNAs; that is, some RNAs in different sets are reused, while others remain distinct, thus creating multiple different training samples. Each set of sample RNAs corresponds one-to-one with a prediction model. The first sample RNA is any one of the RNAs in the set corresponding to the first binding site prediction model. As seen in the model training process, each prediction model is trained based on its corresponding set of sample RNAs. Because each set of sample RNAs is unique, different prediction models can learn to analyze the association between two dimensions from different information analysis perspectives and predict binding probabilities. Therefore, by combining the binding probability prediction functions of multiple prediction models, accurate prediction of binding sites can be achieved. Meanwhile, since some sample RNAs are reused in different sample RNA sets, on the one hand, a relatively rich training sample can be constructed based on a small number of sample RNAs, and on the other hand, each sample RNA can be learned multiple times, so that multiple training models can learn the model knowledge brought by the sample in a more comprehensive way.
[0194] The following section will detail the process of obtaining the aforementioned sequence and structural information.
[0195] The process of obtaining sequence information:
[0196] In one possible implementation, when executing step S201, the computer device may execute steps S2011-S2013 (not shown in the figure), where steps S2011-S2013 are a possible implementation of step S201, including:
[0197] S2011: Obtain the nucleotide sequence information corresponding to ribonucleic acid.
[0198] This nucleotide sequence information includes multiple nucleotide identifiers for identifying multiple nucleotides. There is a one-to-one correspondence between nucleotides and nucleotide identifiers. The arrangement of these nucleotide identifiers in the nucleotide sequence information corresponds to the arrangement of the multiple nucleotides in the ribonucleic acid (RNA). For example… Figure 11 As shown, the ribonucleic acid sequence can be characterized by the arrangement of multiple nucleotide markers.
[0199] S2012: Add a start marker at the beginning of the nucleotide sequence information and an end marker at the end of the nucleotide sequence information.
[0200] Since models typically have a strong ability to understand and process numerical codes but a poor ability to understand other types of information, in order to enable models to better understand the sequence information corresponding to ribonucleic acid, computer devices can convert the identifiers in the nucleotide sequence information into corresponding numerical codes to obtain the sequence information.
[0201] Computer devices can add a start marker at the beginning of the nucleotide sequence information and an end marker at the end of the nucleotide sequence information. The start marker is used to identify the first nucleotide in the ribonucleic acid, and the end marker is used to identify the last nucleotide in the ribonucleic acid. In this way, the model can know which nucleotides are nucleotides in the ribonucleic acid, and at the same time, it can more accurately identify the arrangement of multiple nucleotides in the ribonucleic acid sequence, such as the order of multiple nucleotides from beginning to end.
[0202] S2013: Based on the mapping relationship between identifiers and digital codes, the identifiers in the nucleotide sequence information are converted into corresponding digital codes to obtain the sequence information.
[0203] like Figure 11 As shown, finally, the computer device can predetermine the mapping relationship between identifiers and digital codes, which includes the digital codes corresponding to various identifiers. Based on this mapping relationship, the computer device can convert the identifiers in the nucleotide sequence information into their corresponding digital codes to obtain sequence information. This converts identifiers that the model cannot understand into understandable digital codes, helping the model better understand the nucleotide sequence pattern represented by the sequence information, thus enabling the subsequent binding probability prediction process.
[0204] The mapping relationship can be shown in the table below:
[0205]
[0206]
[0207] In this application, the capital letters in the identifiers represent various bases. Since uracil in ribonucleic acid (RNA) and thymine in deoxyribonucleic acid (DNA) have similar functions but different names, computer equipment can replace all identifiers 'T' (thymine) in the original nucleotide sequence information with identifier 'U' (uracil). In this application, there are seven base identifiers (A, U, G, C, X, N, I), where identifier X, representing an unknown base, and identifier I, representing hypopurine, are treated as unknown identifiers. <unk>The beginning is marked as <cls>The end marker is <eos>.
[0208] The process of obtaining structural information:
[0209] Understandably, the role of nucleotides can be characterized primarily through the structure of ribonucleic acid (RNA). One aspect is the importance of the nucleotide's position within the RNA structure; the more important the nucleotide's position, the more crucial its role. The other aspect is the correlation between different nucleotides within the RNA structure; the stronger the correlation between a nucleotide and other nucleotides, the more crucial its role. Based on this, in one possible implementation, a computer device can acquire the structural information of RNA from these two perspectives, thereby enabling the structural information to more effectively characterize the role of nucleotides.
[0210] In this implementation, the structural information can include the proximity centrality and degree corresponding to multiple nucleotides. The proximity centrality is used to characterize the degree of association between each nucleotide and other nucleotides in the ribonucleic acid structure, and the degree is used to characterize the importance of each nucleotide in the ribonucleic acid structure. Thus, the structural information can identify the ribonucleic acid structure through the two dimensions of association degree and structural importance corresponding to multiple nucleotides.
[0211] When performing step S202, the computer device may execute steps S2021-S2023 (not shown in the figure). Steps S2021-S2023 are one possible implementation of step S202, including:
[0212] S2021: Generate the topological information corresponding to the ribonucleic acid.
[0213] The topological information includes multiple nodes and connection identifiers between them, with each node corresponding to a specific nucleotide. Connection identifiers identify the relationships between nucleotides in the ribonucleic acid (RNA) structure; that is, nodes corresponding to two strongly related nucleotides will have a connection identifier. For example, adjacent nucleotides in sequence information usually have strong relationships; therefore, nodes corresponding to adjacent nucleotides in the sequence information will have a connection identifier.
[0214] S2022: Based on the fact that there is a non-covalent interaction between any two nucleotides in the ribonucleic acid structure, or that the distance between two nucleotides is less than a preset distance, a connection identifier is added between the nodes corresponding to the two nucleotides in the topology information.
[0215] It is understandable that some nucleotides in ribonucleic acid (RNA) may have non-covalent interactions, such as hydrogen bonds between two nucleotides. Multiple nucleotides with non-covalent interactions often have a strong correlation. Therefore, based on the existence of non-covalent interactions between two nucleotides, computer devices can add connection markers between the nodes corresponding to the two nucleotides.
[0216] Furthermore, in the structure of ribonucleic acid (RNA), the distance between nucleotides is generally positively correlated with their correlation. Therefore, computer devices can preset a distance to measure whether two nucleotides are close to each other in the RNA structure. If the distance between two nucleotides is less than the preset distance, it indicates that the two nucleotides are close and therefore generally have a strong correlation. The computer device can then add a connection marker between the nodes corresponding to these two nucleotides.
[0217] S2023: Determine the proximity center and degree of multiple nucleotides based on topological information.
[0218] As seen above, the relationships between nucleotides can be identified through the association markers between nodes. It's understandable that the stronger the association between a nucleotide and other nucleotides, the more important its role in the RNA is generally. Therefore, computer devices can characterize the functions of nucleotides through these associations.
[0219] Computer devices can determine the proximity centrality and degree of multiple nucleotides based on the associations represented by connection identifiers in topological information. Taking the first nucleotide as an example, the first nucleotide corresponds to the first node in the topological information. The proximity centrality of the first nucleotide is inversely correlated with the average number of connection identifiers between the first node and other nodes in the topological information. The number of connection identifiers between nodes refers to the number of connection identifiers required to reach other nodes from one node. For example, in... Figure 12 In the diagram, the number of connection markers between the node corresponding to nucleotide 1 and the node corresponding to nucleotide 2 is 3, and the number of connection markers between the node corresponding to nucleotide 1 and the node corresponding to nucleotide 3 is 2. The fewer the number of connection markers between the nodes corresponding to each nucleotide, the more direct the association between the two nucleotides, indicating a stronger association. Therefore, a lower proximity center corresponding to the first nucleotide indicates a greater number of nucleotides with strong associations in the ribonucleic acid structure. Thus, proximity center is inversely correlated with the degree of association it represents. The proximity center can be determined using the following formula:
[0220]
[0221] Where Closeness is the proximity centrality, n is the total number of nodes in the topology, and d(s,v) i ) is from node s to other nodes v i The number of connection identifiers, where ∑ represents summation.
[0222] The degree of the first nucleotide is positively correlated with the number of linkers connecting the first node in the topological information. That is, the more linkers connecting the first node, the more nucleotides directly associated with the first nucleotide in the ribonucleic acid structure, and therefore the stronger the importance of that first nucleotide in the ribonucleic acid structure. In other words, the degree is positively correlated with the structural importance it represents. For example, in ribonucleic acid structures, nucleotides with high degrees often easily form local binding cavities, which may play a very important role in binding to molecules. The degree can be determined by the following formula:
[0223] Degree(v) = k
[0224] Where Degree(v) is the degree of node v, v is a node in the topology, and k is the number of connection identifiers directly connected to node v. Using this method, computer equipment can accurately analyze the proximity center degree and degree of each nucleotide to constitute the structural information of the ribonucleic acid (RNA), thus accurately and effectively characterizing the RNA structure from the two dimensions of correlation and structural importance.
[0225] The aforementioned structural information primarily characterizes the structural features of nucleotides from the perspective of ribonucleic acid (RNA) structure. It is understandable that the structure of the nucleotide itself can influence its binding efficacy to molecules, meaning it can also affect whether a nucleotide can serve as a binding site. For example, the acid-base properties of a nucleotide may affect the charge interaction between the nucleotide and the molecule, thus influencing the binding efficacy. Therefore, in one possible implementation, the computer device can also combine the structural properties of the nucleotide itself to obtain the corresponding structural information of the RNA, so that the first binding site prediction model can consider the influence of the nucleotide's structural properties on the binding probability when predicting the binding probability.
[0226] In this implementation, the structural information may further include structural property information corresponding to multiple nucleotides. The structural property information corresponding to the first nucleotide is used to characterize the nucleotide properties arising from the structure of the first nucleotide itself, and these nucleotide properties influence the binding affinity of the first nucleotide. The structural property information may include various types of information, which will be described in detail below.
[0227] This approach enables the first binding site prediction model to more accurately analyze the binding effect between nucleotides and molecules from two dimensions: the structural characteristics of nucleotides in the ribonucleic acid structure and the structural properties of the nucleotides themselves, when predicting the binding probability using structural information. This allows for a more accurate prediction of the binding probability between nucleotides and molecules, thus improving the accuracy of binding site prediction.
[0228] The structural properties of the first nucleotide can include any combination of one or more of the following: molecular mass, acidity coefficient, accessible surface area, and evolutionary conservation fraction. The molecular mass of the first nucleotide refers to its own mass, which can affect the structural stability of the nucleotide and its interactions with molecules. The molecular masses and acidity coefficients (pKa) of different nucleotides are shown in the table below:
[0229] Nucleotides Molecular mass Acidity coefficient A 507.18 3.5 C 483.156 4.2 G 523.18 10.8 U 484.141 9.2
[0230] The acidity coefficient corresponding to the first nucleotide can characterize the acid-base properties of the first nucleotide; the evolutionary conservation fraction corresponding to the first nucleotide can characterize the stability of the ribonucleic acid sequence composed of the first nucleotide during biological evolution, and thus the stability of the first nucleotide structure; the accessible surface area corresponding to the first nucleotide refers to the surface area of the region on the first nucleotide that can bind to molecules, such as... Figure 17 As shown, the size of the accessible surface area can also affect the ability of nucleotides to bind to molecules.
[0231] Furthermore, in one possible implementation, in order to further improve the efficiency of model training and the accuracy of combined probability prediction, the computer device can fine-tune the model parameters based on the pre-trained model to obtain the prediction model in this application.
[0232] For example, during model training, the first initial extraction module may include a multi-layer feature extraction network. This network is used to encode the sequence information of ribonucleic acid (RNA) and extract corresponding sequence features. Specifically, the multi-layer feature extraction network consists of a first extraction network and a second extraction network. The first extraction network encodes features based on the sequence information to obtain intermediate outputs, and the second extraction network encodes features based on the intermediate outputs to obtain sequence features. That is, before model training, the multi-layer feature extraction network already possesses mature sequence feature extraction capabilities. This network can be obtained through pre-training, for example, by using a model network from an RNA foundation model (RNA-FM). The first and second extraction networks can work together to complete feature extraction from the sequence information; therefore, both networks contain model knowledge applicable to sequence feature extraction.
[0233] When executing step S30511, the computer device can execute step S305111 (not shown in the figure). Step S305111 is one possible implementation of step S30511, including:
[0234] S305111: Based on the difference between the information on the undetermined binding sites and the information on the sample binding sites, adjust the model parameters corresponding to the second extraction network, the second initial extraction module, and the initial probability prediction module, respectively.
[0235] Understandably, since the first extraction network is closer to the sequence information, i.e. closer to the input, the intermediate output extracted is a relatively basic feature. The accuracy of this feature extraction has a significant impact on the accuracy of the sequence features. Therefore, in order to make full use of the effective model knowledge in the first extraction network to ensure the accuracy of sequence feature extraction, the computer device can leave the parameters of this part of the network unadjusted when adjusting the model parameters, so as to retain this more critical model knowledge.
[0236] The second extraction network is closer to the sequence features, i.e., closer to the output. This part of the network is usually used to fine-tune the features, making them more accurately represent specific dimensional feature information. Based on this, to enable the extracted sequence features to provide more assistance in binding probability prediction, the computer device can adjust the second extraction network based on the accuracy of the binding probability prediction, represented by the difference between the information of the undetermined binding site and the information of the sample binding sites. This allows the second extraction network to analyze which features in the intermediate output of the first extraction network are more helpful for the binding probability, thereby strengthening the extraction of these features and thus making the binding probability determined based on sequence features more accurate. Therefore, through this implementation, the computer device can retain the relatively effective model knowledge in the pre-trained model while fine-tuning some model parameters to make the output more suitable for the application scenario of binding probability prediction. Overall, this reduces the amount of model parameter adjustment, thereby improving model training efficiency while ensuring the prediction accuracy of the trained first binding site prediction model.
[0237] For example Figure 13 As shown, the first feature extraction network consists of four layers, and the second feature extraction network consists of two layers. When adjusting the model parameters, only the model parameters corresponding to the last two layers are adjusted, retaining the more basic and key model knowledge in the first four layers, while enabling the last two layers to learn how to make the sequence features more suitable for probability prediction scenarios.
[0238] To facilitate understanding of the technical solution provided in this application, the method for predicting ribonucleic acid binding sites provided in this application will be described in detail below, taking into account a practical application scenario.
[0239] See Figure 14 , Figure 14 This application provides a flowchart of a method for predicting ribonucleic acid binding sites in a practical application scenario. In this scenario, the computer device can be any of the aforementioned computer devices with binding site prediction capabilities, such as a server or terminal device. The method includes:
[0240] S1401: Obtain the ribonucleic acid to be predicted.
[0241] This ribonucleic acid calculation can be performed on any type of ribonucleic acid that requires binding site prediction.
[0242] S1402: Obtain the nucleotide sequence information corresponding to ribonucleic acid.
[0243] The nucleotide sequence information includes multiple nucleotide identifiers, which are used to identify multiple nucleotides that make up the ribonucleic acid. The order of the multiple nucleotide identifiers in the nucleotide sequence information is consistent with the arrangement of the multiple nucleotides in the ribonucleic acid sequence.
[0244] S1403: Obtain the sequence information corresponding to ribonucleic acid based on nucleotide sequence information.
[0245] Computer devices can add start and end identifiers to nucleotide sequence information, and then encode these identifiers digitally to obtain sequence information (Seq) composed of numbers. i This is so that the prediction model can understand it.
[0246] S1404; Generate the topological information corresponding to the ribonucleic acid.
[0247] The topology information includes nodes corresponding to multiple nucleotides, as well as connection identifiers between nodes added based on the association between nucleotides.
[0248] S1405: Determine the proximity center degree and degree of multiple nucleotides based on topological information.
[0249] The method for determining this has been described in detail above and will not be repeated here. By using proximity centrality and degree, the role of nucleotides can be characterized from the perspective of ribonucleic acid structure.
[0250] S1406: Obtain structural property information for multiple nucleotides.
[0251] Structural property information is used to characterize nucleotide information based on the structure of the nucleotide itself, thereby characterizing the role of the nucleotide from the dimension of the nucleotide's own structure.
[0252] S1407: Obtain the structural information of ribonucleic acid based on the structural properties, proximity center, degree, and nucleotide identifier of multiple nucleotides.
[0253] This structural information may include the following:
[0254] 2D proximity centrality and degree;
[0255] Two-dimensional molecular weight and acidity coefficient;
[0256] One-dimensional accessible surface area;
[0257] One-dimensional evolutionary conservation fraction;
[0258] Four-dimensional nucleotide identifiers, such as codes, identify the four nucleotides in a ribonucleic acid sequence, providing a clear characteristic representation for each nucleotide, thus enabling the model to know which nucleotide corresponds to the aforementioned structural information. These elements can then be used to construct a 10-dimensional structural information structure. i .
[0259] S1408: Input sequence information and structural information into multiple prediction models to obtain the combination probabilities output by the multiple prediction models respectively.
[0260] The model architecture of each prediction model can be as follows: Figure 6 As shown, multiple prediction models are trained based on training samples that are not entirely identical. For example, when there are 10 prediction models, the computer can divide the RNA samples from multiple samples into 10 subsets. Each subset is used as a validation set in turn, and the remaining 9 subsets are used for training. This results in 10 different training samples to train 10 prediction models. This method allows for full utilization of the sample RNA, enabling the training of diverse prediction models with fewer training samples, allowing multiple prediction models to analyze binding probabilities from different perspectives.
[0261] After inputting sequence and structural information into the prediction model, initial sequence features and initial structural features can be extracted first. The initial sequence features... The determination method can be shown by the following formula:
[0262]
[0263] Among them, Encoder seq (Seq i This represents the process of feature extraction from the initial sequence features. This is an n x p dimension matrix feature, where n is the number of nucleotides in the ribonucleic acid and p is the feature dimension. Each row of the matrix represents the initial permutation feature corresponding to each nucleotide. In this practical application scenario, the feature dimension can be 640 dimensions, and this feature extraction process can be completed using a pre-trained model with sequence feature extraction capabilities.
[0264] Before extracting the initial structural features, the computer device can first map the structural information to a high-dimensional space through a linear layer, so that the information dimension can match the dimension of the initial sequence features.
[0265] Initial structural features The extraction process can be represented by the following formula:
[0266]
[0267] Among them, Encoder str (W1·Str i +b1) represents the process of extracting initial structural features, where W1 and b1 are the model parameters of the model module used to extract initial structural features in the prediction model. Each row represents the initial positional distribution characteristics of each nucleotide.
[0268] After determining the initial sequence features and initial structural features, the prediction model inputs these features into the attention layer, which is the model module used for attention weight analysis. In the attention layer, an attention score matrix is calculated between two feature matrices using a dot product. This attention score matrix represents the feature similarity between the two features, and the calculation process is shown in the following formula:
[0269]
[0270] in, The attention score matrix representing the initial sequence features and initial structural features corresponding to ribonucleic acid is an n x n feature matrix. The feature information in the i-th column and j-th row of the matrix is used to characterize the feature similarity between the initial arrangement features corresponding to the i-th nucleotide and the initial position distribution features corresponding to the j-th nucleotide.
[0271] Subsequently, the computer device can normalize all elements of the attention score matrix to obtain the attention weight matrix, as shown in the following formula:
[0272]
[0273] in, Let j = 1, 2, ..., n represent row indices and k = 1, 2, ..., n represent column indices. The feature information corresponding to the j-th row and k-th column of the matrix is used to characterize the feature similarity between the initial permutation feature corresponding to the j-th nucleotide and the initial position distribution feature corresponding to the k-th nucleotide, i.e., the degree of similarity between features.
[0274] Subsequently, the computer device can utilize the attention weight matrix A i Features of the initial sequence and initial structural features Perform feature weighting to obtain sequence features and structural features As shown in the formula below:
[0275]
[0276] Sequence features Each row represents the arrangement and structural features of each nucleotide. Each row represents the positional distribution feature corresponding to each nucleotide. In this way, the feature portion of the initial arrangement feature corresponding to each nucleotide that is similar to the initial positional distribution feature corresponding to each of the n nucleotides can be extracted, as well as the feature portion of the initial positional distribution feature corresponding to each nucleotide that is similar to the initial arrangement feature corresponding to each of the n nucleotides. This allows the prediction model to accurately analyze the correlation between nucleotide functions in two dimensions through sequence features and structural features.
[0277] When predicting the combination probability based on sequence features and structural features, computer devices can first fuse the two features and then perform feature optimization, as shown in the following formula:
[0278]
[0279] in, This represents the result of concatenating two feature matrices. E is the process of optimizing the concatenation result of two features. i The optimized feature matrix is used to further extract the parts of the splicing result that are helpful for combining probability prediction.
[0280] Then, the computer device can process the feature matrix E i Normalization is performed to obtain the normalized feature E. norm,i As shown in the formula below:
[0281] E norm,i =LayerNorm(E i )
[0282] Finally, the computer device linearly maps the normalized matrix to a two-dimensional space, and then applies the normalized exponential function (Softmax) to obtain the final assemblage probability, as shown in the following formula:
[0283] P i =Softmax(W·E) norm,i +b)
[0284] Where W is the weight matrix of the linear mapping, and b is the bias term. This is the output probability matrix, representing the probability that each nucleotide belongs to a binding site. Using the above method, the binding probability output by each prediction model can be obtained.
[0285] S1409: Determine the binding site in the ribonucleic acid based on the binding probabilities output by the multiple prediction models.
[0286] Computer equipment can determine the predicted binding site outcome for each model based on the ensemble probability output by each prediction model. Where k is the index of the model. When the number of prediction models is 10, the value of k ranges from 1, 2, ..., 10. This represents the binding site prediction result determined by the k-th prediction model, used to identify whether each nucleotide is a binding site. The computer device can combine the combined binding sites determined by the k models to determine the final binding site, as shown in the following formula:
[0287]
[0288] Where c is the index of all possible category labels, and argmax is... c This indicates that the maximum value has been found. It is an indicator function, when The value is 1 if it equals tag c, and 0 otherwise. In this practical application scenario, tag c indicates that the nucleotide is a binding site. The number of models that identify each nucleotide as a binding site is calculated. Using this formula, the nucleotide most frequently predicted as a binding site can be found. Therefore, it can be concluded that the nucleotide has the highest probability of being a binding site, and the computer device can identify the nucleotide as a binding site.
[0289] To verify the superiority of the prediction method proposed in this application, the computer device can acquire some old sample data as a training set to train multiple prediction models, thereby constructing multiple prediction methods, including prediction method 1 to prediction method 4, where each prediction method is specifically shown below:
[0290] Prediction Method 1: Use existing convolutional neural network models for combined probability prediction;
[0291] Prediction Method 2: Using existing models, predictions are made based solely on sequence features extracted from pre-trained models;
[0292] Prediction Method 3: Based on existing models, the binding probability is predicted using sequence and structural features. However, this method only performs a simple splicing of sequence and structural features without analyzing the correlation between nucleotide functions in the two dimensions.
[0293] Prediction Method 4: The prediction method in this application combines multiple prediction models to perform probability prediction.
[0294] exist Figure 15 neutralization Figure 16 In this study, computer equipment can use the Matthews correlation coefficient to measure the accuracy of each prediction method for the binding site; this prediction accuracy is positively correlated with the Matthews correlation coefficient. Figure 15 The Matthews correlation coefficients for various prediction methods are determined based on the ribonucleic acid in the training set. This coefficients can characterize the generalization ability of various prediction methods within the distribution, and the generalization ability is positively correlated with the Matthews correlation coefficient. Figure 16 This method uses new sample data not present in the old sample data as a validation set. Based on the ribonucleic acid in the validation set, the Matthews correlation coefficients corresponding to various prediction methods are determined. This coefficient characterizes the out-of-distribution generalization ability of the various prediction methods, and this generalization ability is positively correlated with the Matthews correlation coefficient. Therefore, the prediction method provided in this application demonstrates superior performance compared to existing prediction methods, both in terms of in-distribution and out-of-distribution generalization.
[0295] As can be seen from the above, the ribonucleic acid binding site prediction method provided in this application has the following technical advantages:
[0296] 1. This application can predict binding sites in ribonucleic acid (RNA) by combining both the RNA sequence and the RNA structure. Since the role of nucleotides in RNA is influenced by the interaction between the arrangement of nucleotides in the RNA sequence and their position in the RNA structure, the interrelated information in the sequence and structure can more accurately characterize the role of each nucleotide in RNA. Based on this information, it is possible to more accurately analyze whether each nucleotide constituting the RNA can bind to the molecule when the RNA encounters it. This allows for a more accurate prediction of the probability that multiple nucleotides can serve as binding sites, thus improving the accuracy of binding site prediction.
[0297] 2. This application can refine the process of combining probability prediction in various ways, so that during the training process, each model module of the prediction model can learn how to accurately perform combined probability prediction, thereby further improving the prediction accuracy of the prediction model.
[0298] 3. This application can obtain the sequence information and structural information of ribonucleic acid through multiple methods, so that the sequence information and structural information can fully characterize the ribonucleic acid sequence and ribonucleic acid structural information from multiple information dimensions.
[0299] 4. This application can combine multiple prediction models to jointly predict binding sites, which reduces the probability of incorrect binding site prediction due to insufficient accuracy of a single prediction model, and further improves the accuracy of binding site prediction.
[0300] 5. This application can construct multiple different sample sets based on a small number of training samples, so that multiple prediction models can learn to analyze the probability of the binding site corresponding to each nucleotide from different perspectives, and thus train multiple different prediction models based on a small number of training samples to support the above-mentioned multi-model prediction method.
[0301] 6. This application utilizes the strong sequence feature extraction capability of the pre-trained model to extract sequence features, and at the same time adjusts some networks of the pre-trained model network to make the extracted sequence features more suitable for combining probability prediction.
[0302] Based on the above-mentioned application-side ribonucleic acid (RNA) binding site prediction method, this application also provides a ribonucleic acid (RNA) binding site prediction device, see [link to relevant documentation]. Figure 18 , Figure 18 This application provides a structural block diagram of a ribonucleic acid binding site prediction device. The device 1800 includes a first acquisition unit 1801, a second acquisition unit 1802, a first generation unit 1803, and a first determination unit 1804.
[0303] The first acquisition unit 1801 is used to acquire the sequence information corresponding to the ribonucleic acid, wherein the sequence information is used to identify the arrangement of multiple nucleotides constituting the ribonucleic acid in the ribonucleic acid;
[0304] The second acquisition unit 1802 is used to acquire the structural information corresponding to the ribonucleic acid, wherein the structural information is used to identify the ribonucleic acid structure composed of the plurality of nucleotides;
[0305] The first generation unit 1803 is used to input the sequence information and the structural information into the first binding site prediction model to generate the binding probabilities corresponding to the plurality of nucleotides respectively. The first binding site prediction model is used to analyze the functional correlation between the nucleotides expressed by the sequence information and the structural information respectively, and generate the binding probability according to the correlation. The binding probability is used to characterize the probability of the corresponding nucleotide as a binding site. The binding site is used to bind with the molecule by forming an interaction.
[0306] The first determining unit 1804 is used to determine the binding sites among the plurality of nucleotides based on the binding probability.
[0307] In one possible implementation, the first binding site prediction model includes a feature extraction module and a probability prediction module. The feature extraction module is used to extract sequence features similar to the structural information from the sequence information and to extract structural features similar to the sequence information from the structural information. The probability prediction module is used to predict the binding probability based on the sequence features and the structural features. The sequence features similar to the structural information and the structural features similar to the sequence information are used to characterize the functional association between the nucleotides expressed by the sequence information and the structural information, respectively.
[0308] In one possible implementation, the feature extraction module includes a first extraction module and a second extraction module. The step of extracting sequence features similar to the structural information from the sequence information and extracting structural features similar to the sequence information from the structural information includes:
[0309] The first extraction module performs feature encoding on the sequence information based on the arrangement of the sequence information identifier to extract initial sequence features for characterizing the arrangement, and performs feature encoding on the structural information based on the ribonucleic acid structure of the structural information identifier to extract initial structural features for characterizing the ribonucleic acid structure.
[0310] Feature similarity analysis is performed on the initial sequence features and the initial structural features to obtain attention weights. The attention weights are used to characterize a first degree of similarity between the initial sequence features and the initial structural features. The first degree of similarity is positively correlated with the feature similarity, and the attention weights are positively correlated with the first degree of similarity.
[0311] The second extraction module extracts features from the initial sequence features based on the attention weight to generate the sequence features, and extracts features from the initial structural features based on the attention weight to generate the structural features. The attention weight is used to control the degree of information retention of the initial sequence features when generating the sequence features, and to control the degree of information retention of the initial structural features when generating the structural features. The attention weight is positively correlated with the degree of information retention.
[0312] In one possible implementation, the initial sequence feature is composed of the initial arrangement features corresponding to the plurality of nucleotides respectively, and the initial arrangement features are used to characterize the arrangement characteristics of the corresponding nucleotides in the nucleotide sequence corresponding to the ribonucleic acid;
[0313] The initial structural features are composed of the initial position distribution features corresponding to the plurality of nucleotides, and the initial position distribution features are used to characterize the position distribution characteristics of the corresponding nucleotides in the ribonucleic acid structure;
[0314] The attention weight is used to characterize the similarity between the initial arrangement features corresponding to the plurality of nucleotides and the initial position distribution features corresponding to the plurality of nucleotides, wherein the first nucleotide is any one of the plurality of nucleotides, and the first nucleotide corresponds to the first initial arrangement feature and the first initial position distribution feature. The step of extracting features from the initial sequence features based on the attention weight to generate the sequence features, and extracting features from the initial structural features based on the attention weight to generate the structural features, includes:
[0315] Based on the second similarity between the first initial arrangement feature represented by the attention weight and the initial position distribution features corresponding to the plurality of nucleotides, feature extraction is performed on the first initial arrangement feature to obtain the arrangement feature corresponding to the first nucleotide. The degree of information retention of the arrangement feature corresponding to the first nucleotide on the first initial arrangement feature is positively correlated with the second similarity. The arrangement features corresponding to the plurality of nucleotides are used to form the sequence feature.
[0316] Based on the third similarity between the first initial position distribution feature represented by the attention weight and the initial arrangement features corresponding to the plurality of nucleotides, feature extraction is performed on the first initial position distribution feature to obtain the position distribution feature corresponding to the first nucleotide. The degree to which the position distribution feature corresponding to the first nucleotide retains information about the first initial position distribution feature is positively correlated with the third similarity. The position distribution features corresponding to the plurality of nucleotides are used to form the structural feature.
[0317] In one possible implementation, the arrangement of the multiple nucleotides in the sequence features corresponds to the arrangement of the multiple nucleotides in the ribonucleic acid, and the positional distribution of the multiple nucleotides in the sequence features corresponds to the arrangement of the multiple nucleotides in the ribonucleic acid.
[0318] In one possible implementation, the first nucleotide is any one of the plurality of nucleotides, and the first determining unit 1804 is specifically used for:
[0319] Based on the fact that the binding probability corresponding to the first nucleotide is greater than a first probability threshold, the first nucleotide is used as the binding site among the plurality of nucleotides.
[0320] In one possible implementation, the first nucleotide is any one of the plurality of nucleotides, the prediction model is any one of the plurality of prediction models, and the plurality of prediction models use different training samples during model training. The first determining unit 1804 is specifically used for:
[0321] Based on the fact that the binding probability of the first nucleotide is greater than the second probability threshold, the result output by the first binding site prediction model is determined to be that the first nucleotide is a binding site among the plurality of nucleotides.
[0322] Based on the fact that more than a preset number of prediction models output the result that the first nucleotide is the binding site among the multiple nucleotides, the first nucleotide is used as the binding site among the multiple nucleotides.
[0323] In one possible implementation, the first acquisition unit 1801 is specifically used for:
[0324] Obtain the nucleotide sequence information corresponding to the ribonucleic acid, wherein the nucleotide sequence information includes multiple nucleotide identifiers for identifying the multiple nucleotides, and the arrangement of the multiple nucleotide identifiers in the nucleotide sequence information corresponds to the arrangement of the multiple nucleotides in the ribonucleic acid;
[0325] A start identifier is added at the beginning position of the nucleotide sequence information, and an end identifier is added at the end position of the nucleotide sequence information. The start identifier is used to identify the first nucleotide in the ribonucleic acid, and the end identifier is used to identify the last nucleotide in the ribonucleic acid.
[0326] Based on the mapping relationship between identifiers and digital codes, the identifiers in the nucleotide sequence information are converted into corresponding digital codes to obtain the sequence information.
[0327] In one possible implementation, the structural information includes proximity centrality and degree corresponding to the plurality of nucleotides, wherein the proximity centrality is used to characterize the degree of association between each nucleotide and other nucleotides in the ribonucleic acid structure, and the degree is used to characterize the importance of each nucleotide in the ribonucleic acid structure; the second acquisition unit 1802 is specifically used for:
[0328] Generate topological information corresponding to the ribonucleic acid, the topological information including multiple nodes and connection identifiers between nodes, the multiple nodes corresponding one-to-one with the multiple nucleotides, the connection identifiers between nodes corresponding to adjacent nucleotides in the sequence information, the connection identifiers being used to identify the association between nucleotides in the ribonucleic acid structure;
[0329] Based on the fact that there is a non-covalent interaction between any two nucleotides in the ribonucleic acid structure, or that the distance between the two nucleotides is less than a preset distance, the connection identifier is added between the nodes corresponding to the two nucleotides in the topology information;
[0330] Based on the topological information, the proximity centrality and degree corresponding to the plurality of nucleotides are determined respectively. The first nucleotide corresponds to the first node in the topological information. The proximity centrality corresponding to the first nucleotide is inversely correlated with the average number of connection identifiers between the first node and other nodes in the topological information. The degree corresponding to the first nucleotide is positively correlated with the number of connection identifiers connecting the first node in the topological information.
[0331] In one possible implementation, the structural information further includes structural property information corresponding to each of the plurality of nucleotides, wherein the structural property information corresponding to the first nucleotide is used to characterize the nucleotide properties generated based on the structure of the first nucleotide itself, and the nucleotide properties are used to affect the binding ability of the first nucleotide.
[0332] In one possible implementation, the structural property information corresponding to the first nucleotide includes any one or more combinations of the molecular mass, acidity coefficient, accessible surface area, and evolution conservation fraction of the first nucleotide.
[0333] Based on the aforementioned method for predicting ribonucleic acid (RNA) binding sites on the training side, this application also provides a device for predicting RNA binding sites, see [link to relevant documentation]. Figure 19 , Figure 19 This application provides a structural block diagram of a ribonucleic acid binding site prediction device. The device 1900 includes a third acquisition unit 1901, a fourth acquisition unit 1902, a second generation unit 1903, a third generation unit 1904, and an adjustment unit 1905.
[0334] The third acquisition unit 1901 is used to acquire a first sample ribonucleic acid, which is composed of multiple sample nucleotides. The first sample ribonucleic acid has corresponding sample binding site information, which is used to identify the binding sites in the multiple sample nucleotides.
[0335] The fourth acquisition unit 1902 is used to acquire sample sequence information and sample structure information corresponding to the first sample ribonucleic acid. The sample sequence information is used to identify the sample arrangement of the plurality of sample nucleotides in the first sample ribonucleic acid. The sample structure information is used to identify the sample ribonucleic acid structure composed of the plurality of sample nucleotides.
[0336] The second generation unit 1903 is used to input the sample sequence information and the sample structure information into the initial prediction model to generate the undetermined binding probabilities corresponding to the plurality of nucleotides respectively. The undetermined binding probabilities are used to characterize the probability that the corresponding sample nucleotide is a binding site. The initial prediction model is used to analyze the correlation between the nucleotide functions expressed by the sample sequence information and the sample structure information respectively, and generate the undetermined binding probabilities based on the correlation.
[0337] The third generation unit 1904 is used to generate undetermined binding site information based on the undetermined binding probability. The undetermined binding site information is used to identify the binding sites predicted by the initial prediction model in the plurality of sample nucleotides.
[0338] The adjustment unit 1905 is used to adjust the model parameters corresponding to the initial prediction model based on the difference between the information of the undetermined binding site and the information of the sample binding site, so as to obtain a first binding site prediction model. The first binding site prediction model is used to predict binding sites in ribonucleic acid.
[0339] In one possible implementation, the initial prediction model includes an initial feature extraction module and an initial probability prediction module. The initial feature extraction module is used to extract undetermined sequence features similar to the sample structural information from the sample sequence information, and to extract undetermined structural features similar to the sample sequence information from the sample structural information. The undetermined sequence features similar to the sample structural information and the undetermined structural features similar to the sample sequence information are used to characterize the functional association of nucleotides expressed by the sample sequence information and the sample structural information, respectively. The initial probability prediction module is used to make predictions based on the undetermined sequence features and the undetermined structural features to generate the undetermined binding probability. The adjustment unit 1905 is specifically used for:
[0340] Based on the difference between the undetermined binding site information and the sample binding site information, the model parameters corresponding to the initial feature extraction module and the initial probability prediction module are adjusted to obtain the feature extraction module and the probability prediction module, which are used to constitute the first binding site prediction model.
[0341] In one possible implementation, the initial feature extraction module includes a first initial extraction module and a second initial extraction module. The step of extracting undetermined sequence features similar to the sample structure information from the sample sequence information, and extracting undetermined structural features similar to the sample sequence information from the sample structure information, includes:
[0342] The first initial extraction module encodes the sequence information based on the sample arrangement method identified by the sample sequence information, extracts the initial sequence features to characterize the sample arrangement method, and encodes the structural information based on the sample ribonucleic acid structure identified by the structural information, extracts the initial structural features to characterize the sample ribonucleic acid structure.
[0343] A feature similarity analysis is performed on the undetermined initial sequence features and the undetermined initial structural features to obtain undetermined attention weights. The undetermined attention weights are used to characterize the third similarity between the undetermined initial sequence features and the undetermined initial structural features. The third similarity is positively correlated with the undetermined feature similarity, and the undetermined attention weights are positively correlated with the first similarity.
[0344] The second initial extraction module extracts features from the undetermined initial sequence features based on the undetermined attention weights to generate the undetermined sequence features, and extracts features from the undetermined initial structural features based on the undetermined attention weights to generate the undetermined structural features. The undetermined attention weights are used to control the degree of information retention of the undetermined initial sequence features when generating the undetermined sequence features, and to control the degree of information retention of the undetermined initial structural features when generating the undetermined structural features. The undetermined attention weights are positively correlated with the degree of information retention.
[0345] The adjustment unit 1905 is specifically used for:
[0346] Based on the difference between the undetermined binding site information and the sample binding site information, the model parameters corresponding to the first initial extraction module, the second initial extraction module, and the initial probability prediction module are adjusted to obtain the first extraction module, the second extraction module, and the probability prediction module. The first extraction module and the second extraction module are used to constitute the feature extraction module.
[0347] In one possible implementation, the first initial extraction module includes a multi-layer feature extraction network. The multi-layer feature extraction network performs feature encoding on the sequence information of ribonucleic acid and extracts the corresponding sequence features. The multi-layer feature extraction network consists of a first extraction network and a second extraction network. The first extraction network is used to perform feature encoding based on the sequence information to obtain an intermediate output, and the second extraction network is used to perform feature encoding based on the intermediate output to obtain the sequence features.
[0348] The adjustment unit 1905 is specifically used for:
[0349] Based on the difference between the undetermined binding site information and the sample binding site information, the model parameters corresponding to the second extraction network, the second initial extraction module, and the initial probability prediction module are adjusted respectively.
[0350] In one possible implementation, the first binding site prediction model is any one of multiple prediction models, which are used to jointly predict the binding sites in the ribonucleic acid. The device further includes a fifth acquisition unit and a second determination unit.
[0351] The fifth acquisition unit is used to acquire multiple sample ribonucleic acids, each of which has corresponding sample binding site information.
[0352] The second determining unit is used to select from the plurality of sample ribonucleic acids to determine a plurality of sample ribonucleic acid sets, wherein the sample ribonucleic acids included in different sample ribonucleic acid sets are partially different, and the plurality of sample ribonucleic acid sets correspond one-to-one with the plurality of prediction models, wherein the first sample ribonucleic acid is any one of the sample ribonucleic acid sets corresponding to the first binding site prediction model.
[0353] This application also provides a computer device; please refer to [link to relevant documentation]. Figure 20 As shown, the computer device can be a terminal device; for example, a mobile phone can be used as a terminal device.
[0354] Figure 20 This diagram illustrates a partial structural representation of a mobile phone related to the terminal device provided in this embodiment. (Reference) Figure 20 The mobile phone includes components such as a radio frequency (RF) circuit 710, a memory 720, an input unit 730, a display unit 740, a sensor 750, an audio circuit 760, a wireless Fidelity (WiFi) module 770, a processor 780, and a power supply 790. Those skilled in the art will understand that... Figure 20 The mobile phone structure shown does not constitute a limitation on the mobile phone and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0355] The following is combined with Figure 20 A detailed introduction to each component of a mobile phone:
[0356] RF circuit 710 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with processor 780; additionally, it transmits uplink data to the base station. Typically, RF circuit 710 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), and a duplexer. Furthermore, RF circuit 710 can also communicate wirelessly with networks and other devices. The aforementioned wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, and Short Messaging Service (SMS).
[0357] The memory 720 can be used to store software programs and modules. The processor 780 executes various mobile phone functions and data processing by running the software programs and modules stored in the memory 720. The memory 720 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 720 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0358] The input unit 730 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the mobile phone. Specifically, the input unit 730 may include a touch panel 731 and other input devices 732. The touch panel 731, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 731), and drive the corresponding connected devices according to a pre-set program. Optionally, the touch panel 731 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 780, and can also receive and execute commands sent by the processor 780. In addition, the touch panel 731 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 731, the input unit 730 may also include other input devices 732. Specifically, other input devices 732 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0359] The display unit 740 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 740 may include a display panel 741, which may optionally be configured as a Liquid Crystal Display (LCD), Organic Light-Emitting Diode (OLED), or similar display panel. Further, a touch panel 731 may cover the display panel 741. When the touch panel 731 detects a touch operation on or near it, it transmits the information to the processor 780 to determine the type of touch event. Subsequently, the processor 780 provides corresponding visual output on the display panel 741 based on the type of touch event. Although in Figure 20 In this embodiment, the touch panel 731 and the display panel 741 are two separate components to realize the input and output functions of the mobile phone. However, in some embodiments, the touch panel 731 and the display panel 741 can be integrated to realize the input and output functions of the mobile phone.
[0360] The mobile phone may also include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 741 according to the ambient light level, and the proximity sensor can turn off the display panel 741 and / or backlight when the phone is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, taps), etc. Other sensors that may be configured in the mobile phone, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0361] Audio circuit 760, speaker 761, and microphone 762 provide an audio interface between the user and the mobile phone. Audio circuit 760 converts received audio data into electrical signals and transmits them to speaker 761, where speaker 761 converts them into sound signals for output. On the other hand, microphone 762 converts collected sound signals into electrical signals, which are received by audio circuit 760, converted into audio data, and then processed by processor 780 before being transmitted via RF circuit 710 to, for example, another mobile phone, or the audio data can be output to memory 720 for further processing.
[0362] WiFi is a short-range wireless transmission technology. Through the WiFi module 770, mobile phones can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 20 The WiFi module 770 is shown, but it is understood that it is not an essential component of a mobile phone and can be omitted as needed without changing the essence of the invention.
[0363] The processor 780 is the control center of the mobile phone, connecting various parts of the phone through various interfaces and lines. It executes software programs and / or modules stored in the memory 720, and calls data stored in the memory 720 to perform various functions and process data, thereby performing overall detection of the phone. Optionally, the processor 780 may include one or more processing units; preferably, the processor 780 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 780.
[0364] The mobile phone also includes a power supply 790 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 780 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
[0365] Although not shown, mobile phones may also include a camera, Bluetooth module, etc., which will not be described in detail here.
[0366] In this embodiment, the processor 780 included in the terminal device also has the following functions:
[0367] Obtain the sequence information corresponding to the ribonucleic acid, wherein the sequence information is used to identify the arrangement of multiple nucleotides constituting the ribonucleic acid in the ribonucleic acid;
[0368] Obtain the structural information corresponding to the ribonucleic acid, wherein the structural information is used to identify the ribonucleic acid structure composed of the plurality of nucleotides;
[0369] The sequence information and the structural information are input into a first binding site prediction model to generate the binding probabilities corresponding to the plurality of nucleotides respectively. The first binding site prediction model is used to analyze the functional correlation between the nucleotides expressed by the sequence information and the structural information respectively, and to generate the binding probabilities based on the correlation. The binding probabilities are used to characterize the probability of the corresponding nucleotide as a binding site. The binding sites are used to bind with molecules through the formation of interactions.
[0370] The binding sites among the plurality of nucleotides are determined based on the binding probability.
[0371] Alternatively, in this embodiment, the processor 780 included in the terminal device also has the following functions:
[0372] A first sample ribonucleic acid is obtained, which is composed of multiple sample nucleotides. The first sample ribonucleic acid has corresponding sample binding site information, which is used to identify the binding sites in the multiple sample nucleotides.
[0373] Obtain the sample sequence information and sample structure information corresponding to the first sample ribonucleic acid. The sample sequence information is used to identify the sample arrangement of the plurality of sample nucleotides in the first sample ribonucleic acid. The sample structure information is used to identify the sample ribonucleic acid structure composed of the plurality of sample nucleotides.
[0374] The sample sequence information and the sample structure information are input into the initial prediction model to generate the undetermined binding probabilities corresponding to the plurality of nucleotides. The undetermined binding probabilities are used to characterize the probability that the corresponding sample nucleotide is a binding site. The initial prediction model is used to analyze the correlation between the nucleotide functions expressed by the sample sequence information and the sample structure information, and to generate the undetermined binding probabilities based on the correlation.
[0375] Based on the undetermined binding probability, undetermined binding site information is generated. The undetermined binding site information is used to identify the binding sites predicted by the initial prediction model in the plurality of sample nucleotides.
[0376] Based on the difference between the undetermined binding site information and the sample binding site information, the model parameters corresponding to the initial prediction model are adjusted to obtain a first binding site prediction model, which is used to predict binding sites in ribonucleic acid.
[0377] This application also provides a server; please refer to [link / reference]. Figure 21 As shown, Figure 21 This is a structural diagram of a server 800 provided in an embodiment of this application. The server 800 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 822 (e.g., one or more processors) and a memory 832, and one or more storage media 830 (e.g., one or more mass storage devices) for storing application programs 842 or data 844. The memory 832 and storage media 830 can be temporary or persistent storage. The program stored in the storage media 830 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server. Furthermore, the CPU 822 may be configured to communicate with the storage media 830 and execute the series of instruction operations in the storage media 830 on the server 800.
[0378] Server 800 may also include one or more power supplies 826, one or more wired or wireless network interfaces 850, one or more input / output interfaces 858, and / or one or more operating systems 841, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.
[0379] The steps performed by the server in the above embodiments can be based on Figure 21 The server structure shown.
[0380] This application also provides a computer-readable storage medium for storing a computer program that executes any one of the object information display methods described in the foregoing embodiments.
[0381] This application also provides a computer program product including a computer program, which, when run on a computer device, causes the computer device to execute the object information display method described in any of the above embodiments.
[0382] It is understood that in the specific embodiments of this application, data related to user information (such as ribonucleic acid information) is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0383] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium can be at least one of the following media: read-only memory (ROM), RAM, magnetic disk, or optical disk, etc., and other media capable of storing program code.
[0384] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The device and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0385] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.< / eos> < / cls> < / unk>
Claims
1. A method for predicting ribonucleic acid binding sites, characterized in that, The method includes: Obtain the sequence information corresponding to the ribonucleic acid, wherein the sequence information is used to identify the arrangement of multiple nucleotides constituting the ribonucleic acid in the ribonucleic acid; Obtain the structural information corresponding to the ribonucleic acid, wherein the structural information is used to identify the ribonucleic acid structure composed of the plurality of nucleotides; The sequence information and the structural information are input into a first binding site prediction model to generate the binding probabilities corresponding to the plurality of nucleotides respectively. The first binding site prediction model is used to analyze the functional correlation between the nucleotides expressed by the sequence information and the structural information respectively, and to generate the binding probabilities based on the correlation. The binding probabilities are used to characterize the probability of the corresponding nucleotide as a binding site. The binding sites are used to bind with molecules by forming interactions. The binding sites among the plurality of nucleotides are determined based on the binding probability.
2. The method according to claim 1, characterized in that, The first binding site prediction model includes a feature extraction module and a probability prediction module. The feature extraction module is used to extract sequence features similar to the structural information from the sequence information and to extract structural features similar to the sequence information from the structural information. The probability prediction module is used to make predictions based on the sequence features and the structural features to generate the binding probability. The sequence features similar to the structural information and the structural features similar to the sequence information are used to characterize the functional association between the nucleotides expressed by the sequence information and the structural information, respectively.
3. The method according to claim 2, characterized in that, The feature extraction module includes a first extraction module and a second extraction module. The steps of extracting sequence features similar to the structural information from the sequence information and extracting structural features similar to the sequence information from the structural information include: The first extraction module performs feature encoding on the sequence information based on the arrangement of the sequence information identifier to extract initial sequence features for characterizing the arrangement, and performs feature encoding on the structural information based on the ribonucleic acid structure of the structural information identifier to extract initial structural features for characterizing the ribonucleic acid structure. Feature similarity analysis is performed on the initial sequence features and the initial structural features to obtain attention weights. The attention weights are used to characterize a first degree of similarity between the initial sequence features and the initial structural features. The first degree of similarity is positively correlated with the feature similarity, and the attention weights are positively correlated with the first degree of similarity. The second extraction module extracts features from the initial sequence features based on the attention weight to generate the sequence features, and extracts features from the initial structural features based on the attention weight to generate the structural features. The attention weight is used to control the degree of information retention of the initial sequence features when generating the sequence features, and to control the degree of information retention of the initial structural features when generating the structural features. The attention weight is positively correlated with the degree of information retention.
4. The method according to claim 3, characterized in that, The initial sequence feature is composed of the initial arrangement features corresponding to the plurality of nucleotides respectively, and the initial arrangement feature is used to characterize the arrangement characteristics of the corresponding nucleotides in the nucleotide sequence corresponding to the ribonucleic acid; The initial structural features are composed of the initial position distribution features corresponding to the plurality of nucleotides respectively, and the initial position distribution features are used to characterize the position distribution characteristics of the corresponding nucleotides in the ribonucleic acid structure; The attention weight is used to characterize the similarity between the initial arrangement features corresponding to the plurality of nucleotides and the initial position distribution features corresponding to the plurality of nucleotides, wherein the first nucleotide is any one of the plurality of nucleotides, and the first nucleotide corresponds to the first initial arrangement feature and the first initial position distribution feature. The step of extracting features from the initial sequence features based on the attention weight to generate the sequence features, and extracting features from the initial structural features based on the attention weight to generate the structural features, includes: Based on the second similarity between the first initial arrangement feature represented by the attention weight and the initial position distribution features corresponding to the plurality of nucleotides, feature extraction is performed on the first initial arrangement feature to obtain the arrangement feature corresponding to the first nucleotide. The degree of information retention of the arrangement feature corresponding to the first nucleotide on the first initial arrangement feature is positively correlated with the second similarity. The arrangement features corresponding to the plurality of nucleotides are used to form the sequence feature. Based on the third similarity between the first initial position distribution feature represented by the attention weight and the initial arrangement features corresponding to the plurality of nucleotides, feature extraction is performed on the first initial position distribution feature to obtain the position distribution feature corresponding to the first nucleotide. The degree to which the position distribution feature corresponding to the first nucleotide retains information about the first initial position distribution feature is positively correlated with the third similarity. The position distribution features corresponding to the plurality of nucleotides are used to form the structural feature.
5. The method according to claim 4, characterized in that, The arrangement characteristics of the plurality of nucleotides in the sequence characteristics correspond to the arrangement of the plurality of nucleotides in the ribonucleic acid, and the positional distribution characteristics of the plurality of nucleotides in the sequence characteristics correspond to the arrangement of the plurality of nucleotides in the ribonucleic acid.
6. The method according to claim 1, characterized in that, The first nucleotide is any one of the plurality of nucleotides, and the determination of the binding site among the plurality of nucleotides based on the binding probability includes: Based on the fact that the binding probability corresponding to the first nucleotide is greater than a first probability threshold, the first nucleotide is used as the binding site among the plurality of nucleotides.
7. The method according to claim 1, characterized in that, The first nucleotide is any one of the plurality of nucleotides, the prediction model is any one of the plurality of prediction models, and the plurality of prediction models use different training samples during model training. Determining the binding site among the plurality of nucleotides based on the binding probability includes: Based on the fact that the binding probability of the first nucleotide is greater than the second probability threshold, the result output by the first binding site prediction model is determined to be that the first nucleotide is a binding site among the plurality of nucleotides. Based on the fact that more than a preset number of prediction models output the result that the first nucleotide is the binding site among the multiple nucleotides, the first nucleotide is used as the binding site among the multiple nucleotides.
8. The method according to claim 1, characterized in that, The process of obtaining the sequence information corresponding to the ribonucleic acid includes: Obtain the nucleotide sequence information corresponding to the ribonucleic acid, wherein the nucleotide sequence information includes multiple nucleotide identifiers for identifying the multiple nucleotides, and the arrangement of the multiple nucleotide identifiers in the nucleotide sequence information corresponds to the arrangement of the multiple nucleotides in the ribonucleic acid; A start identifier is added at the beginning position of the nucleotide sequence information, and an end identifier is added at the end position of the nucleotide sequence information. The start identifier is used to identify the first nucleotide in the ribonucleic acid, and the end identifier is used to identify the last nucleotide in the ribonucleic acid. Based on the mapping relationship between identifiers and digital codes, the identifiers in the nucleotide sequence information are converted into corresponding digital codes to obtain the sequence information.
9. The method according to claim 1, characterized in that, The structural information includes the proximity centrality and degree corresponding to each of the plurality of nucleotides. The proximity centrality is used to characterize the degree of association between each nucleotide and other nucleotides in the ribonucleic acid structure, and the degree is used to characterize the importance of each nucleotide in the ribonucleic acid structure. Obtaining the structural information corresponding to the ribonucleic acid includes: Generate topological information corresponding to the ribonucleic acid, the topological information including multiple nodes and connection identifiers between nodes, the multiple nodes corresponding one-to-one with the multiple nucleotides, the connection identifiers between nodes corresponding to adjacent nucleotides in the sequence information, the connection identifiers being used to identify the association between nucleotides in the ribonucleic acid structure; Based on the fact that there is a non-covalent interaction between any two nucleotides in the ribonucleic acid structure, or that the distance between the two nucleotides is less than a preset distance, the connection identifier is added between the nodes corresponding to the two nucleotides in the topology information; Based on the topological information, the proximity centrality and degree corresponding to the plurality of nucleotides are determined respectively. The first nucleotide corresponds to the first node in the topological information. The proximity centrality corresponding to the first nucleotide is inversely correlated with the average number of connection identifiers between the first node and other nodes in the topological information. The degree corresponding to the first nucleotide is positively correlated with the number of connection identifiers connecting the first node in the topological information.
10. The method according to claim 1, characterized in that, The structural information also includes structural property information corresponding to the plurality of nucleotides respectively. The structural property information corresponding to the first nucleotide is used to characterize the nucleotide properties generated based on the structure of the first nucleotide itself. The nucleotide properties are used to affect the binding ability of the first nucleotide.
11. The method according to claim 10, characterized in that, The structural property information corresponding to the first nucleotide includes any one or more combinations of the molecular mass, acidity coefficient, accessible surface area, and evolution conservation fraction of the first nucleotide.
12. A method for predicting ribonucleic acid binding sites, characterized in that, The method includes: A first sample ribonucleic acid is obtained, which is composed of multiple sample nucleotides. The first sample ribonucleic acid has corresponding sample binding site information, which is used to identify the binding sites in the multiple sample nucleotides. Obtain the sample sequence information and sample structure information corresponding to the first sample ribonucleic acid. The sample sequence information is used to identify the sample arrangement of the plurality of sample nucleotides in the first sample ribonucleic acid. The sample structure information is used to identify the sample ribonucleic acid structure composed of the plurality of sample nucleotides. The sample sequence information and the sample structure information are input into the initial prediction model to generate the undetermined binding probabilities corresponding to the plurality of nucleotides. The undetermined binding probabilities are used to characterize the probability that the corresponding sample nucleotide is a binding site. The initial prediction model is used to analyze the correlation between the nucleotide functions expressed by the sample sequence information and the sample structure information, and to generate the undetermined binding probabilities based on the correlation. Based on the undetermined binding probability, undetermined binding site information is generated. The undetermined binding site information is used to identify the binding sites predicted by the initial prediction model in the plurality of sample nucleotides. Based on the difference between the undetermined binding site information and the sample binding site information, the model parameters corresponding to the initial prediction model are adjusted to obtain a first binding site prediction model, which is used to predict binding sites in ribonucleic acid.
13. The method according to claim 12, characterized in that, The initial prediction model includes an initial feature extraction module and an initial probability prediction module. The initial feature extraction module is used to extract undetermined sequence features similar to the sample structural information from the sample sequence information, and to extract undetermined structural features similar to the sample sequence information from the sample structural information. The undetermined sequence features similar to the sample structural information and the undetermined structural features similar to the sample sequence information are used to characterize the functional association of nucleotides expressed by the sample sequence information and the sample structural information, respectively. The initial probability prediction module is used to predict based on the undetermined sequence features and the undetermined structural features to generate the undetermined binding probability. Based on the difference between the undetermined binding site information and the sample binding site information, the model parameters corresponding to the initial prediction model are adjusted to obtain a first binding site prediction model, including: Based on the difference between the undetermined binding site information and the sample binding site information, the model parameters corresponding to the initial feature extraction module and the initial probability prediction module are adjusted to obtain the feature extraction module and the probability prediction module, which are used to constitute the first binding site prediction model.
14. The method according to claim 13, characterized in that, The initial feature extraction module includes a first initial extraction module and a second initial extraction module. The steps of extracting undetermined sequence features similar to the sample structure information from the sample sequence information and extracting undetermined structural features similar to the sample sequence information from the sample structure information include: The first initial extraction module encodes the sequence information based on the sample arrangement method identified by the sample sequence information, extracts the initial sequence features to characterize the sample arrangement method, and encodes the structural information based on the sample ribonucleic acid structure identified by the structural information, extracts the initial structural features to characterize the sample ribonucleic acid structure. A feature similarity analysis is performed on the undetermined initial sequence features and the undetermined initial structural features to obtain undetermined attention weights. The undetermined attention weights are used to characterize the third similarity between the undetermined initial sequence features and the undetermined initial structural features. The third similarity is positively correlated with the undetermined feature similarity, and the undetermined attention weights are positively correlated with the first similarity. The second initial extraction module extracts features from the undetermined initial sequence features based on the undetermined attention weights to generate the undetermined sequence features, and extracts features from the undetermined initial structural features based on the undetermined attention weights to generate the undetermined structural features. The undetermined attention weights are used to control the degree of information retention of the undetermined initial sequence features when generating the undetermined sequence features, and to control the degree of information retention of the undetermined initial structural features when generating the undetermined structural features. The undetermined attention weights are positively correlated with the degree of information retention. The step of adjusting the model parameters corresponding to the initial feature extraction module and the initial probability prediction module based on the difference between the undetermined binding site information and the sample binding site information to obtain the feature extraction module and the probability prediction module includes: Based on the difference between the undetermined binding site information and the sample binding site information, the model parameters corresponding to the first initial extraction module, the second initial extraction module, and the initial probability prediction module are adjusted to obtain the first extraction module, the second extraction module, and the probability prediction module. The first extraction module and the second extraction module are used to constitute the feature extraction module.
15. The method according to claim 14, characterized in that, The first initial extraction module includes a multi-layer feature extraction network. The multi-layer feature extraction network performs feature encoding on the sequence information of ribonucleic acid and extracts the corresponding sequence features. The multi-layer feature extraction network consists of a first extraction network and a second extraction network. The first extraction network is used to perform feature encoding based on the sequence information to obtain intermediate output, and the second extraction network is used to perform feature encoding based on the intermediate output to obtain the sequence features. The step of adjusting the model parameters corresponding to the first initial extraction module, the second initial extraction module, and the initial probability prediction module based on the difference between the undetermined binding site information and the sample binding site information to obtain the first extraction module, the second extraction module, and the probability prediction module includes: Based on the difference between the undetermined binding site information and the sample binding site information, the model parameters corresponding to the second extraction network, the second initial extraction module, and the initial probability prediction module are adjusted respectively.
16. The method according to claim 12, characterized in that, The first binding site prediction model is any one of multiple prediction models, which are used to jointly predict the binding sites in the ribonucleic acid. The method further includes: Multiple sample ribonucleic acids are obtained, and each of the multiple sample ribonucleic acids has corresponding sample binding site information; Multiple sample ribonucleic acids are selected from the multiple sample ribonucleic acids to determine multiple sample ribonucleic acid sets. The sample ribonucleic acids included in different sample ribonucleic acid sets are partially different. The multiple sample ribonucleic acid sets correspond one-to-one with the multiple prediction models. The first sample ribonucleic acid is any one of the sample ribonucleic acid sets corresponding to the first binding site prediction model.
17. A ribonucleic acid binding site prediction device, characterized in that, The device includes a first acquisition unit, a second acquisition unit, a first generation unit, and a first determination unit: The first acquisition unit is used to acquire the sequence information corresponding to the ribonucleic acid, wherein the sequence information is used to identify the arrangement of multiple nucleotides constituting the ribonucleic acid in the ribonucleic acid; The second acquisition unit is used to acquire the structural information corresponding to the ribonucleic acid, wherein the structural information is used to identify the ribonucleic acid structure composed of the plurality of nucleotides; The first generation unit is used to input the sequence information and the structural information into the first binding site prediction model to generate the binding probabilities corresponding to the plurality of nucleotides respectively. The first binding site prediction model is used to analyze the functional correlation between the nucleotides expressed by the sequence information and the structural information respectively, and generate the binding probability according to the correlation. The binding probability is used to characterize the probability of the corresponding nucleotide as a binding site. The binding site is used to bind with the molecule by forming an interaction. The first determining unit is used to determine the binding sites among the plurality of nucleotides based on the binding probability.
18. A ribonucleic acid binding site prediction device, characterized in that, The device includes a third acquisition unit, a fourth acquisition unit, a second generation unit, a third generation unit, and an adjustment unit: The third acquisition unit is used to acquire a first sample ribonucleic acid, which is composed of multiple sample nucleotides. The first sample ribonucleic acid has corresponding sample binding site information, which is used to identify the binding sites in the multiple sample nucleotides. The fourth acquisition unit is used to acquire sample sequence information and sample structure information corresponding to the first sample ribonucleic acid. The sample sequence information is used to identify the sample arrangement of the plurality of sample nucleotides in the first sample ribonucleic acid. The sample structure information is used to identify the sample ribonucleic acid structure composed of the plurality of sample nucleotides. The second generation unit is used to input the sample sequence information and the sample structure information into the initial prediction model to generate the undetermined binding probabilities corresponding to the plurality of nucleotides respectively. The undetermined binding probabilities are used to characterize the probability that the corresponding sample nucleotide is a binding site. The initial prediction model is used to analyze the correlation between the nucleotide functions expressed by the sample sequence information and the sample structure information respectively, and generate the undetermined binding probabilities based on the correlation. The third generation unit is used to generate undetermined binding site information based on the undetermined binding probability. The undetermined binding site information is used to identify the binding sites predicted by the initial prediction model in the plurality of sample nucleotides. The adjustment unit is used to adjust the model parameters corresponding to the initial prediction model based on the difference between the information of the undetermined binding site and the information of the sample binding site, so as to obtain a first binding site prediction model. The first binding site prediction model is used to predict binding sites in ribonucleic acid.
19. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program for executing the ribonucleic acid binding site prediction method according to any one of claims 1-11, or for executing the ribonucleic acid binding site prediction method according to any one of claims 12-16.
20. A computer program product comprising a computer program, which, when run on a computer device, causes the computer device to perform the ribonucleic acid binding site prediction method according to any one of claims 1-11, or to perform the ribonucleic acid binding site prediction method according to any one of claims 12-16.