Non-coding RNA and disease association data processing method and database system

By integrating and predicting models, the problem of insufficient coverage and integration of non-coding RNA and disease association databases has been solved, achieving full coverage and visual query of various non-coding RNA data, and providing a comprehensive analysis tool for non-coding RNA and disease association.

CN116612815BActive Publication Date: 2026-02-27XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310301393.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-24
Publication Date
2026-02-27
Estimated Expiration
2043-03-24

AI Technical Summary

Technical Problem

Existing databases of non-coding RNAs and their associations with diseases suffer from low coverage and insufficient data integration. In particular, they lack effective recording and prediction capabilities when dealing with interactions between various non-coding RNAs and their associations with diseases. Furthermore, the inconsistent naming of non-coding RNAs in different databases makes data integration difficult.

Method used

This paper proposes a non-coding RNA and disease association database system. It integrates existing non-coding RNA and disease association databases using multiple name matching and association dictionary methods, utilizes a pre-set association prediction model for data integration and prediction, and combines the Echarts framework for visualization. The system includes a data integration module, a search module, and a relationship prediction module, and supports data querying and association prediction for various non-coding RNAs.

Benefits of technology

It achieves full coverage of various non-coding RNA data, solves the problem of inconsistent non-coding RNA naming in different databases, and provides comprehensive and visualized correlation query and prediction functions to help researchers understand the interaction relationships of non-coding RNAs and their association with diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612815B_ABST
    Figure CN116612815B_ABST
Patent Text Reader

Abstract

The application belongs to the field of data processing, and discloses a non-coding RNA and disease correlation data processing method and a database system. The method comprises the following steps: integrating non-coding RNA data, disease data and non-coding RNA and disease correlation data in existing non-coding RNA and disease correlation databases; searching target non-coding RNA data and target non-coding RNA and disease correlation data from the integrated data according to a non-coding RNA name to be queried; searching target disease data and target non-coding RNA and disease correlation data from the integrated data according to a disease name to be queried; and obtaining a plurality of non-coding RNA names and a plurality of disease names having correlation with a non-coding RNA to be correlated based on a preset correlation prediction model according to a non-coding RNA name to be correlated. The completeness of existing public data sets is arranged, full coverage of existing databases is realized, and the correlation between non-coding RNAs and diseases that cannot be effectively verified can be effectively predicted.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of data processing, and relates to a non-coding RNA and disease association data processing method and database system. BACKGROUND

[0002] As an important regulatory element of life, non-coding RNA plays an extremely important role in the gene expression regulation of almost all physiological or pathological processes such as embryonic development, cell differentiation, metabolism, signal transduction, immune response, cancer and aging by affecting chromatin modification, binding of transcription targets, mRNA splicing and protein translation, etc. The length of non-coding RNA varies from 22 bases (22nt) to thousands of bases, covering microRNAs (miRNAs) with chain structure, long non-coding RNAs (lncRNAs), circular RNAs (circRNAs) with circular structure and the like. Although tens of thousands of human non-coding RNA genes have been discovered, researchers have a relatively complete understanding of only a small number of non-coding RNAs, and a large number of non-coding RNAs have not been completely studied due to their complex molecular mechanisms. The interaction between non-coding RNAs and the relationship between non-coding RNAs and diseases have become a current research hotspot. At present, the most intuitive way to query, record and explore different non-coding RNAs and their relationships with diseases is through a database, in which users can search for different non-coding RNAs of interest, find the association between different non-coding RNAs, compare the relationship between different non-coding RNAs and diseases, and also download relevant data to provide a basis for subsequent integrated analysis.

[0003] In summary, the existing databases recording non-coding RNA have the following types: (1) independent databases recording different non-coding RNAs, such as circBase and miRBase, which record the information of circular RNA (circRNA) or microRNA (miRNA) in non-coding RNA and provide functions such as query and download, but fail to reflect the association between different RNAs and the interaction with diseases. (2) databases recording the relationship between non-coding RNA and diseases, such as Circ2Disease, HMDD, miRCancer and Lnc2Cancer, which record the association between different single non-coding RNA and diseases, but there is a certain lack of RNA information or the relationship between multiple non-coding RNAs. (3) databases recording the relationship between two non-coding RNAs and diseases, such as Lnc2Cancer database. Lnc2Cancer database records the association between lncRNA and circRNA and diseases, but fails to reflect the interaction between multiple non-coding RNAs, and cannot further predict the relationship between non-recorded non-coding RNA and diseases.

[0004] In summary, although the existing non-coding RNA and disease association databases can record and query different non-coding RNAs and their association with diseases, there are still the following shortcomings. First, different non-coding RNAs, especially circular RNA and long non-coding RNA, have different database names between different databases, and the standards are not unified, which brings great barriers to comprehensive data integration analysis, but the existing databases fail to effectively match and align the names. Second, most of the existing databases record the information of single or two non-coding RNAs, and lack of arrangement and relationship analysis of three or more non-coding RNAs and their interaction. Finally, although the relationship between different non-coding RNAs and diseases has been recorded in databases, the association between multiple non-coding RNAs and diseases has not been effectively explored, especially the potential association between unrecorded non-coding RNAs and diseases can provide a guiding direction for future research, but the current databases generally lack this function. Therefore, it is necessary to design and invent a non-coding RNA and disease association database with more comprehensive functions and better data integration. SUMMARY

[0005] The purpose of the present application is to overcome the shortcomings of the existing non-coding RNA and disease association database, which has low coverage and low data integration, and to provide a non-coding RNA and disease association data processing method and database system.

[0006] To achieve the above purpose, the following technical solutions are adopted:

[0007] In a first aspect, the present application provides a non-coding RNA and disease association data processing method, comprising:

[0008] Based on the multi-name matching and association dictionary method, the non-coding RNA data, disease data and non-coding RNA and disease association data in the existing non-coding RNA and disease association database are integrated to obtain integrated data and store them;

[0009] Obtain the non-coding RNA name to be queried or the disease name to be queried, and according to the non-coding RNA name to be queried, retrieve the target non-coding RNA data and the target non-coding RNA and disease association data from the integrated data; or according to the disease name to be queried, retrieve the target disease data and the target non-coding RNA and disease association data from the integrated data;

[0010] Obtain the non-coding RNA name to be associated, and according to the non-coding RNA name to be associated, based on a preset association prediction model, obtain a plurality of non-coding RNA names and a plurality of disease names having an association with the non-coding RNA to be associated.

[0011] Optionally, it further comprises visualizing the target non-coding RNA data and the target non-coding RNA and disease association data, or visualizing the target disease data and the target non-coding RNA and disease association data; and visualizing the plurality of non-coding RNA names and the plurality of disease names having an association with the non-coding RNA to be associated.

[0012] Optionally, when visualizing the target non-coding RNA and disease association data, an association graph network is constructed based on the target non-coding RNA and disease association data using the Echarts framework for visual display.

[0013] Optionally, the visualization of the plurality of non-coding RNA names and the plurality of disease names having an association with the non-coding RNA to be associated comprises: visualizing the plurality of non-coding RNA names and the plurality of disease names having an association with the non-coding RNA to be associated in mode one and / or mode two according to the association possibility size; wherein mode one: visualizing and displaying the plurality of non-coding RNA names and the plurality of disease names having an association with the non-coding RNA to be associated in the form of a table according to the order from large to small of the association possibility size; mode two: constructing a word cloud map of the plurality of disease names having an association with the non-coding RNA to be associated to obtain a first word cloud map, and constructing a word cloud map of the plurality of non-coding RNA names having an association with the non-coding RNA to be associated to obtain a second word cloud map, and visualizing and displaying the first word cloud map and the second word cloud map.

[0014] Optionally, when the target non-coding RNA data is retrieved, a UCSC Genome Browser database search interface and / or an ENBL-EBI database search interface of the target non-coding RNA data is generated; and when the target disease data is retrieved, a MalaCards database search interface of the target disease data is generated.

[0015] Optionally, the NCBI database hyperlinks of the disease names having the association with the non-coding RNA to be associated are further generated.

[0016] Optionally, the association prediction model comprises a data association module, a knowledge graph embedding model based on a transformer structure, a multi-layer graph convolution network and a fully connected layer network.

[0017] The data association module is configured to retrieve, according to the non-coding RNA name to be associated, non-coding RNAs or diseases having an association relationship with the non-coding RNA to be associated from the integrated data, to obtain a plurality of association data, and to combine the non-coding RNA name to be associated and the names of the association data to obtain a plurality of association pairs; the knowledge graph embedding model based on the transformer structure is configured to obtain embedding representations of the association pairs; the multi-layer graph convolution network is configured to optimize the embedding representations of the association pairs by aggregating neighbor node information to obtain optimized embedding representations of the association pairs; and the fully connected layer network is configured to decode the optimized embedding representations of the association pairs to obtain a plurality of non-coding RNA names and a plurality of disease names having the association with the non-coding RNA to be associated.

[0018] In a second aspect, the application provides a non-coding RNA and disease association database system, comprising a cloud server and a front-end server connected in communication; the cloud server is provided with a data integration module, a search module and a relationship prediction module, and the front-end server is provided with an interaction module.

[0019] The data integration module is configured to integrate, based on multi-name matching and association dictionary, non-coding RNA data, disease data and non-coding RNA and disease association data in existing non-coding RNA and disease association databases to obtain integrated data and store the integrated data.

[0020] The interaction module is configured to obtain a non-coding RNA name to be queried or a disease name to be queried and send the non-coding RNA name to be queried or the disease name to be queried to the search module; and obtain a non-coding RNA name to be associated and send the non-coding RNA name to be associated to the relationship prediction module.

[0021] The searching module is configured to retrieve target non-coding RNA data and target non-coding RNA and disease association data from the integrated data according to a non-coding RNA name to be queried and send the data to the interaction module, or retrieve target disease data and target non-coding RNA and disease association data from the integrated data according to a disease name to be queried and send the data to the interaction module.

[0022] The relationship prediction module is configured to obtain a plurality of non-coding RNA names and a plurality of disease names having an association with a non-coding RNA to be associated based on a preset association prediction model according to a non-coding RNA name to be associated and send the names to the interaction module.

[0023] Optionally, the interaction module is further configured to visualize the target non-coding RNA data and the target non-coding RNA and disease association data, or visualize the target disease data and the target non-coding RNA and disease association data, and visualize the plurality of non-coding RNA names and the plurality of disease names having an association with the non-coding RNA to be associated; the target non-coding RNA and disease association data are visualized by constructing an association graph network based on the target non-coding RNA and disease association data using an Echarts framework; the visualization of the plurality of non-coding RNA names and the plurality of disease names having an association with the non-coding RNA to be associated includes: visualizing the plurality of non-coding RNA names and the plurality of disease names having an association with the non-coding RNA to be associated in a manner one and / or manner two according to an association possibility size; the manner one includes: visualizing and displaying the plurality of non-coding RNA names and the plurality of disease names having an association with the non-coding RNA to be associated in a form of a table according to an order from large to small of the association possibility size; the manner two includes: constructing a word cloud diagram of the plurality of disease names having an association with the non-coding RNA to be associated to obtain a first word cloud diagram, constructing a word cloud diagram of the plurality of non-coding RNA names having an association with the non-coding RNA to be associated to obtain a second word cloud diagram, and visualizing and displaying the first word cloud diagram and the second word cloud diagram.

[0024] Optionally, the interaction module is further configured to generate a UCSC Genome Browser database search interface and / or an ENBL-EBI database search interface of the target non-coding RNA data when the target non-coding RNA data is received, generate a MalaCards database search interface of the target disease data when the target disease data is received, and generate an NCBI database hyperlink of each disease name having an association with the non-coding RNA to be associated.

[0025] Compared with the prior art, the present application has the following beneficial effects:

[0026] The non-coding RNA and disease correlation data processing method of the present application is based on multiple name matching and correlation dictionary method, integrates non-coding RNA data, disease data and non-coding RNA and disease correlation data in the existing non-coding RNA and disease correlation database to obtain integrated data, completes the completeness arrangement of the existing public data set, can realize the integration of the three most common non-coding RNA data of miRNA, lncRNA and circRNA, realizes the full coverage of the existing database, and based on the application of multiple name matching, effectively overcomes the renaming of the same non-coding RNA in different databases, greatly reduces the confusion caused by the repeated names, and then the integrated data can be used to realize the comprehensive retrieval of the non-coding RNA and disease correlation data. At the same time, based on the preset correlation prediction model, a plurality of non-coding RNA names and a plurality of disease names having correlation with the non-coding RNA to be associated are obtained, the correlation between non-coding RNAs and diseases that cannot be effectively verified is effectively predicted, and reference is provided for relevant researchers, so as to help researchers better understand the interaction relationship between different non-coding RNAs and the correlation between non-coding RNAs and diseases. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1 The non-coding RNA and disease correlation data processing method of the present application is based on multiple name matching and correlation dictionary method, integrates non-coding RNA data, disease data and non-coding RNA and disease correlation data in the existing non-coding RNA and disease correlation database to obtain integrated data, completes the completeness arrangement of the existing public data set, can realize the integration of the three most common non-coding RNA data of miRNA, lncRNA and circRNA, realizes the full coverage of the existing database, and based on the application of multiple name matching, effectively overcomes the renaming of the same non-coding RNA in different databases, greatly reduces the confusion caused by the repeated names, and then the integrated data can be used to realize the comprehensive retrieval of the non-coding RNA and disease correlation data. At the same time, based on the preset correlation prediction model, a plurality of non-coding RNA names and a plurality of disease names having correlation with the non-coding RNA to be associated are obtained, the correlation between non-coding RNAs and diseases that cannot be effectively verified is effectively predicted, and reference is provided for relevant researchers, so as to help researchers better understand the interaction relationship between different non-coding RNAs and the correlation between non-coding RNAs and diseases that cannot be effectively verified.

[0028] Figure 2 The non-coding RNA and disease correlation data processing method of the present application is based on multiple name matching and correlation dictionary method, integrates non-coding RNA data, disease data and non-coding RNA and disease correlation data in the existing non-coding RNA and disease correlation database to obtain integrated data, completes the completeness arrangement of the existing public data set, can realize the integration of the three most common non-coding RNA data of miRNA, lncRNA and circRNA, realizes the full coverage of the existing database, and based on the application of multiple name matching, effectively overcomes the renaming of the same non-coding RNA in different databases, greatly reduces the confusion caused by the repeated names, and then the integrated data can be used to realize the comprehensive retrieval of the non-coding RNA and disease correlation data. At the same time, based on the preset correlation prediction model, a plurality of non-coding RNA names and a plurality of disease names having correlation with the non-coding RNA to be associated are obtained, the correlation between non-coding RNAs and diseases that cannot be effectively verified is effectively predicted, and reference is provided for relevant researchers, so as to help researchers better understand the interaction relationship between different non-coding RNAs and the correlation between non-coding RNAs and diseases that cannot be effectively verified. DETAILED DESCRIPTION

[0029] In order to make the person skilled in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by the person skilled in the art without creative labor shall belong to the scope of protection of the present application.

[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0031] The present invention will now be described in further detail with reference to the accompanying drawings:

[0032] See Figure 1 In one embodiment of the present invention, a method for processing non-coding RNA and disease association data is provided. Specifically, it is a data processing method that integrates querying, recording, predicting, and displaying the associations between multiple non-coding RNAs and diseases, so as to help researchers better understand the interaction relationships between different non-coding RNAs and their association with diseases.

[0033] Specifically, this method for processing non-coding RNA and disease association data includes the following steps:

[0034] S1: Based on multi-name matching and association dictionary methods, integrate non-coding RNA data, disease data, and non-coding RNA and disease association data from existing non-coding RNA and disease association databases to obtain integrated data and store it.

[0035] S2: Obtain the name of the non-coding RNA to be queried or the name of the disease to be queried, and retrieve the target non-coding RNA data and the association data between the target non-coding RNA and the disease from the integrated data based on the name of the non-coding RNA to be queried; or retrieve the target disease data and the association data between the target non-coding RNA and the disease from the integrated data based on the name of the disease to be queried.

[0036] S3: Obtain the name of the non-coding RNA to be associated, and based on the name of the non-coding RNA to be associated, obtain several non-coding RNA names and several disease names that are associated with the non-coding RNA to be associated, based on the preset association prediction model.

[0037] In summary, the non-coding RNA and disease association data processing method based on multi-name matching and association dictionary method integrates non-coding RNA data, disease data and non-coding RNA and disease association data in the existing non-coding RNA and disease association database to complete the completeness arrangement of the existing public data set, can realize the integration of three most common non-coding RNA data of miRNA, lncRNA and circRNA, realize the full coverage of the existing database, and based on the application of multi-name matching, effectively overcome the renaming of the same non-coding RNA in different databases, greatly reduce the confusion caused by the repeated names, and then based on the integrated data, realize the comprehensive retrieval of the non-coding RNA and disease association data. At the same time, based on the preset association prediction model, a plurality of non-coding RNA names and a plurality of disease names having association with the non-coding RNA to be associated are obtained, the association between non-coding RNAs that cannot be effectively verified and diseases is effectively predicted, and reference basis is provided for relevant researchers, so as to help researchers better understand the interaction relationship between different non-coding RNAs and the association between diseases.

[0038] In a possible implementation, first, from 13 existing non-coding RNA and disease association databases such as circBase, circBank, circad, circR2Cancer, RNADiseasev4.0, starBasev2.0, HMDDv3.2, miRCancer, miR2Diseaes, Lnc2Cancer v3.0 and LncRNADisease, the unique identifier of different non-coding RNA data and the specific name corresponding relationship of the disease are determined through multi-name matching and association dictionary establishment, and then the non-coding RNA data, disease data and non-coding RNA and disease association data in the existing non-coding RNA and disease association database are integrated, and then the non-coding RNA and its interaction map are constructed, and the number of different databases and the data distribution obtained finally are displayed in the form of table image, as shown in Table 1.

[0039] Table 1

[0040] Association pair name Total entries circRNA-disease 1399 miRNA-disease 10154 lncRNA-disease 3280 circRNA-miRNA 1129 miRNA-lncRNA 9506

[0041] In a possible implementation, the non-coding RNA and disease association data processing method further includes visualizing the target non-coding RNA data and the target non-coding RNA and disease association data, or visualizing the target disease data and the target non-coding RNA and disease association data; and visualizing a plurality of non-coding RNA names and a plurality of disease names having association with the non-coding RNA to be associated.

[0042] Optionally, when the target non-coding RNA and disease association data are visualized, the association graph network is constructed based on the target non-coding RNA and disease association data using the Echarts framework to visualize and display. The visualized several non-coding RNA names and several disease names associated with the non-coding RNA to be associated include: visualizing the several non-coding RNA names and several disease names associated with the non-coding RNA to be associated in mode one and / or mode two according to the size of the association possibility; wherein mode one: visualizing and displaying the several non-coding RNA names and several disease names associated with the non-coding RNA to be associated in the form of a table according to the order from large to small of the size of the association possibility; mode two: constructing a word cloud diagram of the several disease names associated with the non-coding RNA to be associated to obtain a first word cloud diagram, and constructing a word cloud diagram of the several non-coding RNA names associated with the non-coding RNA to be associated to obtain a second word cloud diagram, and visualizing and displaying the first word cloud diagram and the second word cloud diagram.

[0043] Specifically, according to the non-coding RNA or disease name input by the user, i.e. the non-coding RNA name to be queried or the disease name to be queried, the integrated data is searched and the search results are displayed in the form of a table. The public relationship between non-coding RNAs and between non-coding RNAs and diseases verified by public experiments is combined to form a visual graph connection structure. In order to more intuitively represent the overall association relationship of the database, the entire integrated data can also be directly retrieved. The summary table and the interactive word cloud diagram formed according to the degree of the entities in the relationship network diagram can be seen, thereby realizing web-based visual display of the integrated data.

[0044] Optionally, different non-coding RNAs and diseases can be searched by item, or different non-coding RNA or disease names can be input, or the non-coding RNA or disease name of interest can be stored in a text file for uploading, and the search results are displayed in the form of a table.

[0045] In one possible implementation, the non-coding RNA and disease association data processing method further includes generating a UCSC Genome Browser database search interface and / or an ENBL-EBI database search interface for the target non-coding RNA data when the target non-coding RNA data is retrieved, and generating a MalaCards database search interface for the target disease data when the target disease data is retrieved. Optionally, a NCBI database hyperlink for each disease name associated with the non-coding RNA to be associated is also generated.

[0046] According to different entries, automatically search for matching UCSC Genome Browser, EMBL-EBI and MalaCards database to visualize the position area, jump to the page display, corresponding gene expression display and corresponding disease information display; at the same time, combined with multi-source data matching and graph network construction, form the non-coding RNA and disease association network, use Echarts framework to build interactive support interaction network for visual display.

[0047] Specifically, the Axios library is used to send Ajax asynchronous request, and the input non-coding RNA name is uploaded to the backend. The backend interface responds to the request and receives the data for searching, sends the results to the front end and performs page rendering to display the results. If the non-coding RNA has complete name and corresponding chromosome position in circBase, circBank and other databases, the search interface provided by UCSC Genome Browser is used, that is, the UCSC Genome Browser database search interface: http: / / genome.mdc-berlin.de / cgi-bin / hgTracks?db=hg19&position=chr4:87967317-87968746 is generated, and the RNA position information is replaced to display the corresponding RNA interval information of interest, which can be jumped to the corresponding page for viewing through the UCSC Genome Browser database search interface.

[0048] Similarly, for different non-coding RNAs, if there are corresponding gene expression records, the ENBL-EBI database is connected, that is, the ENBL-EBI database search interface is generated, and the interface is introduced by replacing the disease keyword: https: / / www.ebi.ac.uk / gxa / experiments / E-MTAB-513 / Results?specific=true&geneQue ry=%255B%257B%2522value%2522%253A%2522XIST%2522%252C%2522catego ry%2522%253A%2522symbol%2522%257D%255D&filterFactors=%257B%257D&cutoff=%257B%2522value%2522%253A0.5%257D&unit=%2522TPM%2522, gene expression screening and display can be realized, and the ENBL-EBI database search interface can be jumped to the corresponding webpage for dynamic expression display page.

[0049] For different diseases, the interface of the MalaCards database is connected, that is, the MalaCards database search interface of the target disease data is generated: https: / / www.malacards.org / card / acne, which displays the specific introduction of the disease, the DiseaseOntology database ID, the MESH database ID and the associated diseases, and the search interface can jump to the corresponding webpage to view the detailed information of the interested disease.

[0050] In a possible implementation, the association prediction model comprises a data association module, a knowledge graph embedding model based on a transformer structure, a multi-layer graph convolution network, and a fully connected layer network.

[0051] The data association module is configured to retrieve non-coding RNAs or diseases associated with the non-coding RNA to be associated from the integrated data according to the non-coding RNA name to be associated, to obtain a plurality of association data; and combine the non-coding RNA name to be associated and the names of the association data to obtain a plurality of association pairs; the knowledge graph embedding model based on the transformer structure is configured to obtain embedding representations of the association pairs; the multi-layer graph convolution network is configured to optimize the embedding representations of the association pairs by aggregating neighbor node information to obtain optimized embedding representations of the association pairs; and the fully connected layer network is configured to decode the optimized embedding representations of the association pairs to obtain a plurality of non-coding RNA names and a plurality of disease names associated with the non-coding RNA to be associated, in addition to the association possibility sizes of the non-coding RNAs associated with the non-coding RNA to be associated and the association possibility sizes of the diseases associated with the non-coding RNA to be associated.

[0052] For example, when the association prediction model inputs a non-coding RNA circACAP2 to be predicted, the association prediction model will first obtain the known association entities of circACAP2 from the database, such as breast cancer (diseases) and hsa-mir-29a-3p (miRNA), and then input each pair of association pairs (such as circACAP2-breast cancer) into the model to obtain the embedding representation of circACAP2. In order to further optimize the embedding representation, the association prediction model introduces a multi-layer graph convolution network-based aggregation structure to further optimize the embedding representation of circACAP2 by aggregating neighbor node information. Finally, the fully connected layer network is connected to decode the optimized embedding representations, thereby predicting the relationship between the non-coding RNA circRNA of interest and a plurality of diseases, and the prediction results are sorted according to the association possibility size (probability value).

[0053] Specifically, the non-coding RNA of interest can be input, or the name of the non-coding RNA to be predicted can be uploaded in the form of a text file, and at the same time, the prediction task is submitted. The data will be transmitted from the front end to the background cloud server through the flask framework, and the task will be uploaded to the cloud server for task prediction and result download through communication between node servers, and then transmitted back to the local platform for subsequent front-end display.

[0054] For the association between non-coding RNAs that are not formally recorded or not experimentally verified and diseases, the association prediction model is used to predict the association between potential unknown relationships and display them in the form of tables and word clouds, forming an interactive word cloud structure.

[0055] The specific prediction process is as follows: first, for the various non-coding RNA relationships of different non-coding RNAs and their connections with diseases, a head-entity-relation-tail entity connection is constructed to form a triple structure, such as (has_circ_0000001, circRNA-disease, acne). Then, a biological knowledge network between non-coding RNAs and diseases is constructed using a knowledge graph structure, and a knowledge graph embedding model based on a transformer structure is used to obtain the embedding representation of each entity. The introduction of the transformer structure allows the knowledge graph entity embedding representation learning to focus more on its relevant low-order or high-order neighbor nodes. Specifically, the transformer structure is composed of a multi-head attention structure, a residual structure, and a feedforward network. The multi-head attention structure first obtains the embedding representation of each entity by selecting the aggregation of neighbor node information through a self-attention mechanism:

[0056]

[0057] where Q is the index Query of the attention mechanism, K is the key value Key, V is the value Value, d k is the embedding dimension.

[0058] At the same time, multiple multi-head attention structures are introduced to obtain more optimal entity embedding representations by aggregating feature information of neighbor nodes in different dimensions:

[0059]

[0060] where Attention i is the weighted output of the i-th sub-attention mechanism, W O is a trainable parameter for dimension conversion.

[0061] The introduction of the residual structure can effectively change the deep model training problem, and the feedforward network is used to change the embedding feature dimension, filter invalid information and aggregate effective information.

[0062]

[0063] wherein, V h is the embedding feature representation of node h, is the aggregation representation of the neighbor node information of node h.

[0064] Through the application of the knowledge graph embedding model and the graph network structure, the final embedding representation of each entity can be obtained. Finally, the full connection layer is connected to decode the embedding representation of each entity, and the relationship between the non-coding RNA of interest and a plurality of diseases is predicted. The prediction result is displayed in two ways according to the possibility of relevance. First, in the form of a standard table, each disease is also connected to the NCBI database through a hyperlink, and the search can be further queried to obtain the non-coding RNA associated with different diseases. In addition, the Echarts framework can be used to construct a word cloud diagram for all associated diseases according to the prediction possibility, and the possibility of different diseases is dynamically displayed for the user to understand at a glance.

[0065] In addition, different files of interest (lists of non-coding RNAs of interest) can be uploaded, and the search results, prediction results and data of the association between different RNAs and diseases can be downloaded.

[0066] Specifically, the file upload supports standard file format upload, and the non-coding RNA required by the user for research is extracted through data parsing and data query. The results of the query can provide association data download. For the non-coding RNA of interest of the user, the RNA-disease association size result list with high accuracy can be obtained by calling the association prediction model, and background download is supported. Finally, the function of downloading and subsequent integration of the existing integrated data of the association between different non-coding RNAs and diseases is also supported.

[0067] The non-coding RNA and disease association data processing method has at least the following advantages:

[0068] The non-coding RNA and disease association data processing method can be used for non-discriminatory query of different non-coding RNAs and diseases, obtaining interaction relationships of different non-coding RNAs and association relationships with existing diseases, and supporting visualization of a relationship diagram; in addition, the UCSC genome browser and the EMBL-EBI database are accessed, so that the positions of different non-coding RNAs on a genome and distribution information of upstream and downstream elements thereof can be viewed, and expression information of different non-coding RNAs in tissues can be queried, visualized and downloaded. In addition, for non-coding RNA and disease relationships not recorded in existing public data sets, a knowledge graph is constructed by using an association prediction model and combining multiple non-coding RNA relationships, potential association relationship representations are obtained by mining multiple group relationships, the association sizes of different non-coding RNAs and diseases are predicted, and dynamic display is supported through an interactive word cloud diagram and a relationship network diagram. In addition, public association data download and problem feedback are supported.

[0069] Referring to Figure 2 In another embodiment of the present application, a non-coding RNA and disease association database system is provided, which can be used to implement the above-mentioned non-coding RNA and disease association data processing method. Specifically, the non-coding RNA and disease association database system includes a cloud server and a front-end server connected in communication; the cloud server is provided with a data integration module, a search module and a relationship prediction module, and the front-end server is provided with an interactive module.

[0070] The data integration module is used to integrate non-coding RNA data, disease data and non-coding RNA and disease association data in existing non-coding RNA and disease association databases based on multi-name matching and association dictionary methods, to obtain integrated data and store the same; the interactive module is used to obtain a non-coding RNA name to be queried or a disease name to be queried, and send the same to the search module; and obtain a non-coding RNA name to be associated and send the same to the relationship prediction module; the search module is used to retrieve target non-coding RNA data and target non-coding RNA and disease association data from the integrated data according to the non-coding RNA name to be queried, and send the same to the interactive module; or retrieve target disease data and target non-coding RNA and disease association data from the integrated data according to the disease name to be queried, and send the same to the interactive module; and the relationship prediction module is used to obtain a plurality of non-coding RNA names and a plurality of disease names having association with the non-coding RNA to be associated based on a preset association prediction model according to the non-coding RNA name to be associated, and send the same to the interactive module.

[0071] In a possible implementation, the interaction module is further configured to visualize the target non-coding RNA data and the target non-coding RNA and disease association data, or visualize the target disease data and the target non-coding RNA and disease association data; and visualize a plurality of non-coding RNA names and a plurality of disease names that have an association with the non-coding RNA to be associated; when the target non-coding RNA and disease association data are visualized, the target non-coding RNA and disease association data are visualized by using an Echarts framework to construct an association graph network; and the plurality of non-coding RNA names and the plurality of disease names that have an association with the non-coding RNA to be associated include: visualizing the plurality of non-coding RNA names and the plurality of disease names that have an association with the non-coding RNA to be associated in a manner one and / or manner two according to the association possibility size; wherein the manner one: visualizing the plurality of non-coding RNA names and the plurality of disease names that have an association with the non-coding RNA to be associated in a form of a table according to the association possibility size from large to small; and the manner two: constructing a word cloud diagram of the plurality of disease names that have an association with the non-coding RNA to be associated to obtain a first word cloud diagram, and constructing a word cloud diagram of the plurality of non-coding RNA names that have an association with the non-coding RNA to be associated to obtain a second word cloud diagram, and visualizing the first word cloud diagram and the second word cloud diagram.

[0072] In a possible implementation, the interaction module is further configured to, when the target non-coding RNA data is received, generate a UCSC Genome Browser database search interface and / or an ENBL-EBI database search interface of the target non-coding RNA data; when the target disease data is received, generate a MalaCards database search interface of the target disease data; and further configured to generate an NCBI database hyperlink of each disease name that has an association with the non-coding RNA to be associated.

[0073] Specifically, the application uses a flask framework and Bootstrap, Layui front-end framework to provide a multi-element non-coding and disease association database system, provides non-coding RNA interaction relationship query visualization, non-coding RNA disease association prediction query visualization and data interaction download functions. The development of the front-end page is completed by using Bootstrap, Layui and Echarts components, in the back-end development work, the flask framework in the Python language is used to complete the back-end data access and association prediction function. Pack Bootstrap and Layui related files and deploy them to the server, send Ajax asynchronous requests through the Axios library, submit the front-end data to the back-end, return the response result after the back-end processing, and complete the data interaction.

[0074] In addition, a help module is also provided, based on which the use instructions of each of the above modules and the content contained in different modules can be viewed, and if there are still some problems and improvement suggestions when the user uses the database system, the user can fill in the user contact information through the information feedback channel module and send it to the server system administrator through the email, so as to promote the further improvement of the database system.

[0075] Regarding the deployment and online of the database system, specifically, the front-end server is formed by Bootstrap, Layui and Echarts frameworks to form a front-end display platform, the back-end, i.e., the cloud server, uses the flask framework to realize data front-end and back-end transmission and processing, in the relationship prediction module, the front-end server is realized to carry out lightweight operation and display through the communication between the cloud server and the intranet node server, the heavy computing work is sent to the high-performance computing node server of the back-end cloud server to carry out, and the result is transmitted back to the front-end server through the low-delay local area network communication to carry out result output and visual display. Finally, the Gunicorn service is used for server port deployment, and the Nginx reverse proxy server is used for node forwarding and hiding to realize the deployment and online of the server.

[0076] In summary, the non-coding RNA and disease correlation data processing method and the database system have the characteristics of comprehensive data coverage, comprehensive functions, strong user friendliness and lightweight.

[0077] Specifically, regarding comprehensive data coverage, three most common non-coding RNA data of miRNA, lncRNA and circRNA are comprehensively integrated, the completeness of the existing public data set is completed through data integration, and full coverage of the existing database is realized. In addition, the renaming of the same non-coding RNA in different databases is also matched and arranged, which greatly reduces the confusion caused by the repeated names. Regarding comprehensive functions, the association construction and query visualization function of the association between the three non-coding RNAs and diseases are realized, and a disease association prediction model under the association relationship of multiple non-coding RNAs is constructed by combining a knowledge graph and a graph convolution model. The model is used to predict the existing, especially the disease-non-coding RNA relationship that cannot be effectively verified and display it to the end user in real time, providing a reference for relevant researchers. Regarding strong user friendliness, a plurality of visualization functions are constructed on the front end, and different types of interactive visualization modules are constructed for different RNA position information, the association relationship between different non-coding RNAs, and the association size between different potential diseases, which greatly improves the user-friendly degree, enhances the user stickiness and reference frequency. Regarding lightness, a lightweight computing, display and interaction page based on a web framework is built by combining the existing lightweight front-end and back-end separation framework and cloud server configuration, which not only supports flexible access of multiple terminals, but also realizes the separation of back-end computing and front-end display, improves the page rendering loading efficiency. In addition, the front-end page of the database system is friendly, the interoperability is strong, the back-end algorithm calculation speed is fast, and the ease of use is strong.

[0078] The embodiments of the foregoing non-coding RNA and disease association data processing method involve all related contents of each step, which can be cited to the function description of the function module corresponding to the non-coding RNA and disease association data processing library system in the embodiments of the present application, and will not be repeated here.

[0079] The division of the modules in the embodiments of the present application is illustrative, and is only a logical function division. In actual implementation, another division mode can be used. In addition, each function module in each embodiment of the present application can be integrated in one processor, or can be a separate physical existence, or two or more modules can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software function module.

[0080] In still another embodiment of the present application, a computer device is provided, which comprises a processor and a memory, the memory being configured to store a computer program, the computer program comprising program instructions, and the processor being configured to execute the program instructions stored in the computer storage medium. The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, etc., which are the computing core and control core of the terminal, and are suitable for implementing one or more instructions, and are specifically suitable for loading and executing one or more instructions in the computer storage medium to implement a corresponding method flow or a corresponding function; the processor in the embodiments of the present application can be used for the operation of the non-coding RNA and disease association data processing method.

[0081] In still another embodiment of the present application, the present application further provides a storage medium, specifically a computer readable storage medium (Memory), which is a memory device in the computer device, and is used for storing programs and data. It can be understood that the computer readable storage medium herein can include the built-in storage medium in the computer device, and of course can also include the expansion storage medium supported by the computer device. The computer readable storage medium provides a storage space, and the storage space stores the operating system of the terminal. Moreover, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer readable storage medium herein can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory. One or more instructions stored in the computer readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the non-coding RNA and disease association data processing method in the above embodiments.

[0082] Those skilled in the art will appreciate that embodiments of the application can be devised for a method, a system, or a computer program product. Accordingly, the present application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.

[0083] The present application is described in reference to the flowchart and / or block diagrams of the method, apparatus (system) and computer program product according to embodiments of the application. It will be understood that each block of the flowchart and / or block diagrams, and combinations of blocks in the flowchart and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing device or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.

[0084] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.

[0085] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.

[0086] Finally, it should be noted that the above-mentioned embodiments are merely intended for describing the technical solutions of the present application, but not for limiting it. Although the present application is described in detail with reference to the above embodiments, those skilled in the field should understand that the specific embodiments of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application, and any modification or equivalent replacement without departing from the spirit and scope of the present application should be covered in the protection scope of the claims of the present application.

Claims

1. A method for processing non-coding RNA and disease association data, characterized by, The method comprises the following steps: Based on multi-name matching and association dictionary method, integrate non-coding RNA data, disease data and non-coding RNA and disease association data in existing non-coding RNA and disease association database to obtain integrated data and store them; Obtain the non-coding RNA name to be queried or the disease name to be queried, retrieve target non-coding RNA data and target non-coding RNA and disease association data from the integrated data according to the non-coding RNA name to be queried, or retrieve target disease data and target non-coding RNA and disease association data from the integrated data according to the disease name to be queried; Obtain the non-coding RNA name to be associated, and based on the preset association prediction model, obtain a plurality of non-coding RNA names and a plurality of disease names having association with the non-coding RNA to be associated according to the non-coding RNA name to be associated; It also includes visualizing target non-coding RNA data and target non-coding RNA and disease association data, or visualizing target disease data and target non-coding RNA and disease association data, and visualizing a plurality of non-coding RNA names and a plurality of disease names having association with the non-coding RNA to be associated; When visualizing the target non-coding RNA and disease association data, an association graph network is constructed based on the target non-coding RNA and disease association data to realize visual display by using Echarts framework; The association prediction model comprises a data association module, a knowledge graph embedding model based on transformer structure, a multi-layer graph convolution network and a fully connected layer network; The data association module is used to retrieve non-coding RNA or disease having association with the non-coding RNA to be associated from the integrated data according to the non-coding RNA name to be associated, to obtain a plurality of association data, and to combine the non-coding RNA name to be associated and each association data name to obtain a plurality of association pairs; the knowledge graph embedding model based on transformer structure is used to obtain embedding representation of each association pair; the multi-layer graph convolution network is used to optimize the embedding representation of each association pair by aggregating neighbor node information to obtain optimized embedding representation of each association pair; and the fully connected layer network is used to decode the optimized embedding representation of each association pair to obtain a plurality of non-coding RNA names and a plurality of disease names having association with the non-coding RNA to be associated.

2. The non-coding RNA and disease association data processing method according to claim 1, characterized in that, The visualization of a plurality of non-coding RNA names and a plurality of disease names having association with the non-coding RNA to be associated comprises: The several non-coding RNA names and the several disease names having the correlation with the non-coding RNA to be correlated are visualized according to the size of the correlation possibility in mode one and / or mode two; wherein, mode one: the several non-coding RNA names and the several disease names having the correlation with the non-coding RNA to be correlated are visualized and displayed in the form of a table according to the size of the correlation possibility from large to small; mode two: a word cloud diagram of the several disease names having the correlation with the non-coding RNA to be correlated is constructed to obtain a first word cloud diagram, and a word cloud diagram of the several non-coding RNA names having the correlation with the non-coding RNA to be correlated is constructed to obtain a second word cloud diagram, and the first word cloud diagram and the second word cloud diagram are visualized and displayed.

3. The non-coding RNA and disease association data processing method of claim 1, wherein, When the target non-coding RNA data is retrieved, a UCSC Genome Browser database search interface and / or an ENBL-EBI database search interface of the target non-coding RNA data is generated; when the target disease data is retrieved, a MalaCards database search interface of the target disease data is generated.

4. The non-coding RNA and disease association data processing method of claim 1, wherein, A NCBI database hyperlink of each disease name having the correlation with the non-coding RNA to be correlated is further generated.

5. A non-coding RNA and disease association database system, comprising: The cloud server and the front-end server are communicatively connected; the cloud server is internally provided with a data integration module, a search module and a relationship prediction module, and the front-end server is internally provided with an interactive module; The data integration module is used for integrating non-coding RNA data, disease data and non-coding RNA and disease correlation data in an existing non-coding RNA and disease correlation database based on multi-name matching and correlation dictionary mode, obtaining integrated data and storing; The interactive module is used for obtaining a non-coding RNA name to be queried or a disease name to be queried, and sending to the search module; and obtaining a non-coding RNA name to be correlated and sending to the relationship prediction module; The search module is used for retrieving target non-coding RNA data and target non-coding RNA and disease correlation data from the integrated data according to the non-coding RNA name to be queried, and sending to the interactive module; or retrieving target disease data and target non-coding RNA and disease correlation data from the integrated data according to the disease name to be queried, and sending to the interactive module; The relationship prediction module is used for obtaining several non-coding RNA names and several disease names having the correlation with the non-coding RNA to be correlated based on a preset correlation prediction model according to the non-coding RNA name to be correlated, and sending to the interactive module; The interactive module is further used for visualizing the target non-coding RNA data and the target non-coding RNA and disease correlation data, or visualizing the target disease data and the target non-coding RNA and disease correlation data; and further used for visualizing the several non-coding RNA names and the several disease names having the correlation with the non-coding RNA to be correlated; When the target non-coding RNA and disease correlation data is visualized, an Echarts framework is used to construct a correlation graph network based on the target non-coding RNA and disease correlation data, and visualized and displayed. The association prediction model comprises a data association module, a knowledge graph embedding model based on a transformer structure, a multi-layer graph convolution network, and a fully connected layer network. The data association module is configured to retrieve, according to a to-be-associated non-coding RNA name, non-coding RNAs or diseases having an association relationship with the to-be-associated non-coding RNA from integrated data to obtain a plurality of association data, and combine the to-be-associated non-coding RNA name and each association data name to obtain a plurality of association pairs; the knowledge graph embedding model based on the transformer structure is configured to obtain embedding representations of the association pairs; the multi-layer graph convolution network is configured to optimize the embedding representations of the association pairs by aggregating neighbor node information to obtain optimized embedding representations of the association pairs; and the fully connected layer network is configured to decode the optimized embedding representations of the association pairs to obtain a plurality of non-coding RNA names and a plurality of disease names having an association with the to-be-associated non-coding RNA.

6. The non-coding RNA and disease association database system of claim 5, wherein, The visualized plurality of non-coding RNA names and plurality of disease names having an association with the to-be-associated non-coding RNA comprise: The plurality of non-coding RNA names and plurality of disease names having an association with the to-be-associated non-coding RNA are visualized according to the association possibility size in a manner one and / or manner two; wherein, manner one: the plurality of non-coding RNA names and plurality of disease names having an association with the to-be-associated non-coding RNA are visualized and displayed in a table form according to the association possibility size from large to small; and manner two: a first word cloud diagram of the plurality of disease names having an association with the to-be-associated non-coding RNA is constructed according to the association possibility size to obtain a first word cloud diagram, and a second word cloud diagram of the plurality of non-coding RNA names having an association with the to-be-associated non-coding RNA is constructed to obtain a second word cloud diagram, and the first word cloud diagram and the second word cloud diagram are visualized and displayed.

7. The non-coding RNA and disease association database system of claim 5, wherein, The interaction module is further configured to generate a UCSC GenomeBrowser database search interface and / or an ENBL-EBI database search interface of target non-coding RNA data when the target non-coding RNA data is received, generate a MalaCards database search interface of target disease data when the target disease data is received, and generate an NCBI database hyperlink of each disease name having an association with the to-be-associated non-coding RNA.

Citation Information

Patent Citations

  • Adult pancreatic derived stromal cells

    CN101188942A

  • Emergency information interaction system based on structured icons

    CN112394847A