Massive parallel processing database for sequence and graph data structures
Through large-scale parallel processing graph database technology, the problem of inefficient data processing during drug reuse is solved, and rapid response and efficient processing of large-scale multimodal data is achieved, which significantly improves the efficiency of drug reuse.
Patent Information
- Application Number
- CN202510201707.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-10
- Filing Date
- 2021-10-26
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art has cumbersomeness in the drug reuse process, making it difficult to quickly process large-scale multimodal data and perform protein sequence matching and analysis, resulting in inefficient drug reuse pipelines.
Using large-scale parallel processing graph database technology, we quickly generate drug hypotheses by scalable graph databases by storing and processing multimodal data represented in the form of knowledge graphs, implementing interactive queries and semantic traversals, and accelerating domain-specific functions, such as protein similarity analysis.
It has achieved accelerating drug reuse pipeline during the epidemic, and can quickly respond and process large-scale data, spanning 4 million proteins, more than 155 billion fact queries, and processing about 30TB of data, significantly improving the efficiency of drug reuse.
Smart Images

Figure CN120126622A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application date of October 26, 2021, the application number of 202111250156.1, and the invention name of "Large-Scale Parallel Processing Database for Sequence and Graph Data Structures". Technical Field
[0002] The present disclosure generally relates to systems and methods for drug repurposing. More specifically, some embodiments relate to large-scale parallel graph databases that process multimodal data represented in the form of knowledge graphs and accelerate domain-specific functions for protein similarity analysis to generate drug hypotheses. Brief Description of the Drawings
[0003] In accordance with one or more respective embodiments, the present disclosure is described in detail with reference to the following drawings. These drawings are provided for illustrative purposes only and depict exemplary or sample embodiments.
[0004] Figure 1 Illustrates an example graph engine in accordance with various embodiments.
[0005] Figure 2 Illustrates an example of graph engine query execution in accordance with various embodiments.
[0006] Figure 3 Illustrates an example implementation of a graph engine (such as CGE) for performing protein sequence analysis in accordance with various embodiments.
[0007] Figure 4 Illustrates an example of a sample query that can be used to perform protein sequence analysis.
[0008] Figure 5 Illustrates an example computing component 500 that can be used to implement protein sequence analysis in accordance with various embodiments.
[0009] Figure 6 Is an example computing component that can be used to implement various features of the embodiments described in the present disclosure.
[0010] The drawings are not exhaustive and do not limit the present disclosure to the exact forms disclosed. Detailed Description
[0011] In the pandemic caused by the novel coronavirus, drug repurposing (investigating existing drugs for new therapeutic purposes) has emerged as the first glimmer of hope in finding medical care. However, drug repurposing is not limited to coronaviruses and can also be used to identify therapeutic agents / drugs for any of a variety of different medical conditions. The drug repurposing pipeline involves understanding the protein structure of the disease-causing organism, interpreting the interaction of the organism's protein structure with the human body, mining the properties of potential drug molecules, connecting the dots in the selected literature to explain the mechanism of action, looking for evidence in assay data, and using data from previous trials to analyze potential safety and efficacy, etc. Traditionally, this process has been done manually and has taken months.
[0012] The laborious nature of this problem is attributed to the time required by life science researchers to: (a) understand the disease-causing organism by matching and comparing protein sequences to previously known or studied disease-causing organisms (over 4 million sequences); (b) handle and process multi-modal big data (protein sequences, proteomic interactions, biochemical pathways, structured data from past clinical trials, etc.); (c) integrate and search for patterns connected across multiple multi-modal multi-terabyte datasets; (d) install, configure, and run a large number of tools (genetics, proteomics, molecular dynamics, data science, etc.) to generate insights; and finally (e) verify and confirm the scientific rigor of pharmacological interpretations. Embodiments of the systems and methods disclosed herein go beyond merely automating the traditional process and provide new technologies and techniques for implementing the repurposing pipeline. Embodiments can use massively parallel processing graph database technology to provide a more rapid response to accelerate the drug repurposing pipeline in a pandemic. This represents a significant improvement over current technology for identifying drug candidates for re-purposing / repurposing for known and novel diseases through protein sequence analysis, which diseases include, for example, novel influenza strains, coronaviruses, genetic rare diseases, etc.
[0013] Embodiments relate to the application of a massively parallel processing graph database for rapid response drug repurposing. Implementations can use a scalable graph database that is configured to host a knowledge graph of medical-related facts integrated from multiple knowledge sources and also act as a computational engine capable of performing in-database protein sequence analysis. Embodiments can be configured to use the graph database for multi-modal drug repurposing based on the processed sequence of a subject's virus or other medical condition, identify other known viruses / medical conditions with similar or matching sequences, and query the properties of compounds and therapeutic agents that interact with those known viruses / medical conditions.
[0014] Embodiments can provide a massively parallel graph database that (a) stores, disposes, hosts, and processes multimodal data represented in the form of a knowledge graph; (b) provides interactive query and semantic traversal capabilities for data-driven discovery; (c) accelerates domain-specific functions such as the Smith-Waterman algorithm for protein similarity analysis, vertex-centric all-graph algorithms for graph theory connectivity and correlation analysis such as PageRank; and (d) runs / executes query workflows across multiple datasets to generate drug hypotheses on the order of seconds rather than months. Embodiments implement an integrated knowledge graph of multiple multimodal life science databases, perform protein sequence matching in parallel, and provide a novel and fast drug repurposing method that can query across more than 4 million proteins and more than 15.5 billion facts while disposing of approximately 30 TB of data.
[0015] Some applications implement a generalizable big data platform for other biomedical discovery problems beyond the COVID-19 pandemic, which allows: (a) a scalable graph database that provides order-of-magnitude computational acceleration and interactivity required for knowledge traversal and discovery; (b) an integrated life science knowledge graph that captures the open science domain of available biomedical facts; (c) hypotheses for potential candidate drugs for the ongoing pandemic; (d) reproducible code and results for future research in the domain of biomedical facts (for viruses, proteins, drugs, biochemical pathways), rather than being limited to the state of practice of disease-specific knowledge graphs.
[0016] Embodiments can be implemented using a Cray graph engine or other similar engines. The Cray Graph Engine (CGE) is an in-memory semantic graph database designed to scale to hundreds of nodes and tens of thousands of processes on a Cray XC supercomputer to support interactive queries on large datasets (about 100 TB). CGE is based on the standardized Resource Description Framework (RDF) format to ingest datasets of N-triples / N-quads and implements queries using the SPARQL query language. RDF data is expressed as a directed graph with "quadruple" labels, where a "quadruple" consists of four fields: subject, predicate, object, and graph. A triple is just a quadruple stored in the "default graph". For example, the following is a simplified version of an example RDF triple from Uniprot COVID-19 data that can be loaded into CGE:
[0017]
[0018] A graph as a data structure can include a network of possible connections. Vertices or nodes generally refer to entities (data, people, enterprises, etc.), and the connections between entities are edges. A graph database can be used to identify entities that are connected to other entities. Typically, local processing can be used to process a small amount of data around the nodes. However, other tasks may involve evaluating the edges / connections on a more comprehensive basis (e.g., in full graph analysis). A semantic graph can include a collection of such triples, where the subject and object represent vertices, and the predicate represents the edge between the vertices. A semantic graph database differs from a relational database in that the underlying data structure is a graph rather than a structured collection of tables. The graph structure makes semantic databases ideal for analyzing loosely connected or schema-less multimodal unstructured and structured data, such as social network interactions or interactions between proteins and genes in a living organism.
[0019] In various embodiments, the CGE can include two main components: a dictionary and a query engine. The dictionary is responsible for building the database, which is the process of extracting the raw N-triple / N-quadruple files from a high-performance Lustre file system and converting them into the internal representation used by the CGE. The dictionary stores the unique RDF strings from the N-triples / N-quadruples and provides a mapping between the unique strings and the integer identifiers used internally by the query engine for quadruples. Most of the dictionary build time may be dominated by Lustre I / O time.
[0020] The CGE query engine processes SPARQL queries and SPARUL update requests, provides several built-in graph algorithms (such as, for example, centrality measurement, page rank, connectivity analysis) that can be applied to query data, and returns results to the user. The core work performed by the query engine can include: matching the basic graph patterns in the SPARQL query and supporting operations on the query results (such as FILTER and ORDER), which allow the user to remove and sort solutions respectively.
[0021] Embodiments can be implemented to interface with the CGE to improve the performance and scalability of supercomputer products with general-purpose processors using high-performance interconnects. Multiple features are added on top of and outside of the traditional CGE to specifically support rapid response drug repurposing.
[0022] Figure 1Illustrates an example graph engine according to various embodiments. The example includes a front end 210 and multiple resources replicated across multiple computer images 212. These resources can include, for example, a deserializer, operators, and a scheduler as shown in the computational image 212. The example graph engine can also include a dictionary 218, intermediate results to be promoted 220, a hash table, and other exciting data structures 222 and a database 224. A storage file system 214 can be used to accommodate the database 224 as well as user space, checkpoints, and other data.
[0023] The front end 210 provides an interface through which a user can interact with the graph engine, such as, for example, by submitting a query and receiving back results from the query. On the back end, the graph engines run on hardware that can be built on top of a partitioned global address space, which can allow the system to treat independent processes and images as its own entities, but you can segment the data and use a communication library to share data across images. The graph engines can be configured to run tens of thousands or hundreds of thousands of graphs in a coordinated manner, where all graphs can run independently on their own subsets and then synchronize using the library when results are needed.
[0024] Figure 2 Illustrates an example of graph engine query execution according to various embodiments. Now, referring to Figure 2 , a user can submit an example query 320. A communication interface 324 can provide communication and control between the front end (e.g., front end 210) compute nodes and the back end compute nodes. In this example, the communication interface can include elements such as a SPARQL Protocol and RDF Query Language (SPARQL) converter, an IP interface with a web browser, a service that displays or forwards SPARQL and command results, a service that generates a low-level query (e.g., RPN), and a service that passes non-SPARQL commands.
[0025] Multiple compute nodes 328 can be provided to perform query operations. In this example, an image is received (one image is shown as image 0 334). The compute nodes can receive, validate, and send the RPN to all images and, when results are obtained, send these results along with pointers to an output file. The compute nodes can include multiple operators 336 that perform operations on the images. The operators 336 included in this example are SCAN, JOIN, MERGE, OPTIONAL, UNION, FILTER, and BIND, although other operations can also be used. These operations can be used to traverse the data in the respective databases 338 in different ways to complete the query. In Figure 2In the illustrated example, the example query 320 is used to identify the person selling DVDs at a store. This query can be transformed (e.g., SPARQL) and submitted to a computing node, which performs various operations (e.g., operator 336) suitable for the query to identify the person selling DVDs.
[0026] In various embodiments, the computing node 328 can also be enhanced to perform protein sequence analysis. This can be implemented as a domain-specific capability into the database to provide sequence analysis for various applications (including, for example, drug identification) and reuse them. Figure 3 An example implementation of a graph engine such as CGE for performing protein sequence analysis according to various embodiments is illustrated. This example also shows example queries 420, which can be submitted via interface 424 to the computing node 428 in CGE. Similar to Figure 3 the example of, in this example, the included operators 436 include SCAN, JOIN, MERGE, OPTIONAL, UNION, FILTER, and BIND, etc. However, different from Figure 3 the example of, in this example, the operator 436 also includes an operator for performing protein sequence analysis. In one example, CGE is modified to define an interface for a function, which can be used as part of an evaluation expression in an operator (such as ORDER or FILTER). This function is called a user-defined function because the user of CGE can write this function to apply domain-specific knowledge to the query results. CGE provides general functions through which the user can rewrite their own functions, and CGE will load them into memory at program startup for use during execution. CGE defines an interface with the function to enable passing parameters to the user function and, for the purpose of evaluating the expression of the operator, allowing the user to return information to CGE. The information returned by the user-defined function to CGE enables the domain-specific function to easily rank or filter the results. In various embodiments, the system can be configured to enable the user to also add user-defined functions to perform custom searches / queries.
[0027] Figure 4 An example of a sample query that can be used to perform protein sequence analysis is illustrated. This can be the one described above in Figure 3An example of query 420 introduced. As seen in this example, at 442, the example query can include the identification of a protein for a given disorder, for which the user wants to identify it as a viable therapeutic agent. In this example, the user identified the SARS2 spike protein, and the query includes a mnemonic, which is 'SPIKE_SARS2' in this example. This part of the query also identifies the protein sequence of the disorder of interest, which is a virus in this case. The core sequence can include the entire protein sequence of the virus or one or more fragments.
[0028] Also as seen in this example, at 454, the query can specify a request to find all other proteins with the same sequence information. This can include a request to find the entire sequence or one or more parts of the sequence to see if there are any other proteins with a matching sequence or matching sequence fragment. At 455, the query requests to compare the disorder of interest (SARS2 spike) with other proteins in the database and score the sequence matches based on their distance to provide a similarity score. At 457, the query requests to return the results and these results are enumerated in descending order based on the similarity score.
[0029] Figure 5 Illustrated is an example computing component 500 that can be used to implement protein sequence analysis according to an embodiment of the disclosed technology. The computing component 500 can be, for example, a server computer, a controller, or any other similar computing component capable of processing data. In Figure 5 the example implementation, the computing component 500 includes a hardware processor 502 and a machine-readable storage medium 504. The hardware processor 502 can be one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices suitable for retrieving and executing instructions stored in the machine-readable storage medium 504. The hardware processor 502 can obtain, decode, and execute instructions (such as instructions 506 to 514) to control the process or operation for combining local parameters to implement protein sequence analysis for rapid response drug repurposing. As an alternative or supplement to retrieving and executing instructions, the hardware processor 502 can include one or more electronic circuits that include electronic components for performing functions of one or more instructions, such as a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other electronic circuits.
[0030] A machine-readable storage medium (such as machine-readable storage medium 504) can be any electronic, magnetic, optical, or other physical storage device that contains or stores executable instructions. Thus, machine-readable storage medium 504 can be, for example, random access memory (RAM), non-volatile RAM (NVRAM), electrically erasable programmable read-only memory (EEPROM), a storage device, an optical disc, etc. In some embodiments, machine-readable storage medium 504 can be a non-transitory storage medium, where the term "non-transitory" does not cover transitory propagated signals. As will be described in detail below, machine-readable storage medium 504 can be encoded with executable instructions, such as instructions 506 to 514.
[0031] Hardware processor 502 can execute instruction 506 to receive a query that includes a sequence of a given virus or other disorder. For example, a protein or gene sequencing device can be used to determine the sequence of a disorder of interest. As described above with respect to Figure 4 the example query in, a query can be constructed that includes an identification of a protein sequence of a given disorder for which the user wants to identify a viable therapeutic agent. The constructed query can also include a mnemonic of the disorder of interest and the protein sequence or a fragment of the protein sequence.
[0032] Hardware processor 502 can execute instruction 508 in response to the received query to search a sequence database to identify other viruses or proteins that have sequences similar to the sequence of the virus (or protein) of interest. As noted above with reference to Figure 4 the system can search the database to identify matches or similarities to the entire sequence or one or more identified fragments of the sequence. With respect to this operation, hardware processor 502 can execute instruction 510, and sequences in the sequence database can be compared with the query sequence (the entire sequence or one or more identified fragments of the sequence) to determine whether any matches can be identified.
[0033] Hardware processor 502 can execute instruction 512 to determine a similarity score based on the comparison. The similarity score can include, for example, a numerical value (e.g., as a number within a defined range or as a percentage, etc.) or other indicator that indicates the degree of similarity of a given sequence to the query sequence. For example, if expressed as a percentage, 100% can represent a perfect match. The system can be configured to return results for proteins with a similarity score greater than a given threshold. For example, the system can identify proteins that match a similarity score greater than 70% (or some other threshold). In various embodiments, the threshold level can also be set by the query.
[0034] The hardware processor 502 may execute instructions 514 to query the graph database to identify therapeutic agents (drugs or molecules) that have an inhibitory effect on proteins having a similarity score greater than the identified threshold with respect to a protein sequence. For example, for each protein having a similarity score greater than the threshold, the system may identify in the graph database therapeutic agents that have an effect on those proteins. The system may return results identifying therapeutic agents having the desired effect and may score the results based on the similarity score. In other words, therapeutic agents known in the database to have the desired effect on proteins in the sequence database may be considered more likely to have the desired effect on the condition of interest.
[0035] Two core CGE improvements may be included in various embodiments to support drug repurposing: (a) support for user-defined functions (UDFs) for protein sequence analysis in the database; (b) the ability to parallelize the execution of such domain-specific UDFs using the SPARQL front-end user interface described above to accelerate and scale horizontally.
[0036] For example, since CGE utilizes Jena's SPARQL query parser interface, the syntax of user-defined functions in CGE may follow Apache Jena guidelines. As part of a query, the SPARQL interface allows the use of custom functions within the query expression to perform domain-specific operations on the data. This is a feature that allows users to define, express, and execute domain-specific mathematical operations to evaluate and rank query results not supported by SPARQL. Such graph operations may be implemented as custom functions defined by URIs in the expression. This capability may be configured to allow users to define their own functions. Calls to these user-defined functions may take the following form:
[0037]
[0038] Two custom user-defined functions (UDFs) (new URIs, arq:user_func) for drug repurposing may be included in CGE to call the invocation of UDFs that exist separately from CGE. A single C interface may define a function named cge_user_eval that CGE may execute as part of an expression. The cge_user_eval function accepts four arguments that provide the total number of arguments, the argument list, the return value, and the return type. This may allow users to pass data from CGE to the UDF, evaluate the arguments, and return the raw values (e.g., boolean, integer, or double precision) that can be used to evaluate the SPARQL expression.
[0039] Because CGE executes in a massively parallel fashion, where tens of thousands of images may be running concurrently, any UDFs retrieved as part of a query can also be executed in parallel over the set of query solutions. The parallel execution of UDFs enables it to scale to datasets that might otherwise be too large to partition data across parallel images. UDFs can also be applied to the entire set of solutions in an embarrassingly parallel manner. Additionally, the distributed execution of computationally intensive algorithms can be achieved via the parallel execution of UDFs by breaking down complex processing tasks into images that require processing a fairly small set of inputs.
[0040] Embodiments can include enhancements to CGE that focus on improving the execution of database operations such as FILTER or GROUP. These operations can enable a user to compare terms found as part of a query match to apply some order or ranking. The original strings of these terms can be stored in a CGE dictionary, which can be implemented as a distributed hash table that distributes the strings across all processes. This distribution of terms creates a significant amount of work to extract the terms from local to the process when the terms are needed as part of an operation such as FILTER or GROUP.
[0041] To improve the overall performance and scalability of these operations, the strings used by the images can be fetched as large chunks in a coordinated manner. Each image can go through the results and create a list of the strings it needs from each of the other images. Then, all images can fetch the required strings from each other as a single chunk rather than issuing a remote fetch for each string individually. This increases the size of each message but significantly reduces the total number of messages required. This communication pattern matches the work done in CGE for core graph operations such as JOIN and MERGE and has been shown in previous studies to significantly improve performance by reducing the number of outstanding messages at once [8]. This CGE improvement is crucial for queries that parallel pairwise compare protein sequences against millions of open science sequences and rank order the result set.
[0042] Embodiments use biomedical or other life science data resources to integrate life science knowledge graphs. For example, an implementation can use available biomedical data resources commonly used in life science and systems biology research for the knowledge graph. The typical workflow for researchers is to perform a search in one database in the database, then construct a query for another database, and iterate. Manually mapping between the ontologies of individual data sources and piecing together the results from multiple query endpoints (or using another database to perform the transformation) is a cumbersome process. In various embodiments, the scalability of CGE enables all relevant databases to be loaded in one environment, enabling seamless cross-database queries. In various embodiments, federated queries can also be used to query across multiple databases. However, due to challenges such as network access from firewall systems, query rate limitations, or simple performance issues with complex federated queries, this method may not be suitable for complex queries. The scalable load time of CGE also supports frequent reloading of datasets, including integrating internal data on top of the background of a common database during the workflow. In addition, the performance and scalability of CGE for the database construction process enable updated data to be quickly pulled in and the database to be fully rebuilt in less than an hour.
[0043] The integrated life science knowledge graph assembled to study potential drug repurposing candidates for COVID-19 is generated from a collection of publicly available databases. A description of the larger databases in the collection and the descriptions specifically mentioned in this document are now described.
[0044] Uniprot: The UniProt database is a collection of functional information about proteins and includes annotations, relationships, and in some cases, the amino acid sequences of the proteins themselves. Proteins are the cornerstone for studying drug-protein structures and interactions. The interactions between proteins are complex and widely connected, so graph representation is particularly useful. Uniprot focuses on human proteins, although other widely studied organisms are also well represented.
[0045] The UniProt Consortium is a collaboration between the European Bioinformatics Institute (EBI), the Swiss Institute of Bioinformatics (SIB), and the Protein Information Resource (PIR). It has been a pioneer in semantic web technologies. Since 2008, Uniprot has been distributed in RDF format. Uniprot grows as more scientific data is added. New Uniprot releases are distributed every four weeks.
[0046] For this study, most of the Uniprot RDF data was from the release on March 19, 2020. There is a new UniProt portal that provides up-to-date information on COVID-19 coronavirus proteins and receptors, which is updated independently of the general UniProt release cycle. For COVID-19 research, this allows us to pull updated COVID data more quickly. The COVID-19 Uniprot data for the knowledge graph was updated on May 22, 2020. The Uniprot database contains approximately 87.6 billion triples. In the form of N-triple (.nt) files, it is roughly 12.7TB on disk. To simplify queries across multiple databases, we merged all named graphs into a single default graph.
[0047] PubChem: PubChem is an open chemical database maintained by the National Institutes of Health (NIH) in the United States. The PubChemRDF project provides RDF-formatted information for the PubChem Compound, Substance, and Bioassay databases. The knowledge graph for this study used the test version V1.6.3 of PubChemRDF downloaded from ftp: / / ftp.ncbi.nlm.nih.gov / pubchem / RDF on March 30, 2020. The PubChemRDF database contains approximately 80 billion RDF triples. In the form of N-triples, this is equivalent to approximately 13TB on disk.
[0048] ChEMBL: ChEMBL is a curated database of bioactive molecules with drug-like properties. It brings together chemical, bioactivity, and genomic data to help translate genomic information into effective new drugs. The data is updated regularly, approximately every 3 to 4 months. The ChEMBL-RDF version 27.0 (May 18, 2020) has been integrated into the knowledge graph for this study. The ChEMBL database contains approximately 539M triples. In the form of N-triples, this is equivalent to approximately 81GB on disk.
[0049] Bio2RDF dataset: Bio2RDF is an open-source project that uses Semantic Web technologies to pull together different datasets from multiple data providers. In addition to providing an online SPARQL endpoint based on Virtuoso for querying across a collection of heterogeneous datasets, Bio2RDF also provides a portal to download transformed RDF data files for the datasets included in the Bio2RDF database. Download the Bio2RDF dataset included in the knowledge graph.
[0050] The complete Bio2RDF collection consists of approximately 11 billion triples across 35 datasets and includes the DrugBank, PubMed, and MESH datasets.
[0051] OrthoDB: OrthoDB (https: / / www.orthodb.org) provides evolutionary and functional annotations of orthologs (i.e., genes that extant species have inherited from their last common ancestor). Since orthologs are the most likely candidates to retain the function of their ancestral genes, OrthoDB aims to narrow down hypotheses about gene function and conduct comparative evolutionary studies.
[0052] The OrthoDB database contains approximately 2.2 billion RDF triples that describe the evolutionary and functional properties of 40 million genes from 15,000 organisms. In the form of N-triples, this amounts to approximately 275 GB on disk.
[0053] Table 1 Characteristics of knowledge graph datasets. Original size before deduplication.
[0054]
[0055]
[0056] BioModels: The BioModels database is a repository of mathematical models representing biological systems. It currently has a collection of models describing processes such as signaling, one or more protein-drug interactions, metabolic pathways, epidemic models, etc. The models hosted by BioModels are typically described in peer-reviewed scientific literature and, in some cases, they are automatically generated from pathway resources (Path2Models). These models are manually curated and semantically enriched through cross-references to external data resources such as databases of publications, compounds, and pathways, ontologies, etc.
[0057] Embodiments provide the ability for CGE to support UDFs, which can be implemented such that queries can be written that combine information from the knowledge graph and apply domain-specific UDFs to the data to better refine the results. For drug repurposing, embodiments can implement a UDF that performs protein sequence similarity to infer connections between proteins. This can be configured to enable the inference of connections between less well-known proteins (such as COVID-19) and proteins that are well-documented in open datasets (such as Uniprot and ChEMBL).
[0058] Embodiments can utilize the Smith-Waterman (SW) protein similarity algorithm to align sequence pairs and calculate the similarity score of the alignment. For two sequences of lengths m and n, the SW algorithm returns the optimal local alignment and similarity score, with a computational time complexity of O(mn). Local alignments are used to provide an alignment that describes the most similar regions within a sequence, as opposed to the end-to-end alignment of a sequence returned by a global alignment. Since SW returns the best local alignment, it is an important component of many aligners. However, the computational complexity limits the extent to which the algorithm can be used to compare large sequence sets.
[0059] The SW algorithm can be desirable in various embodiments because of the preference for using the best local alignment to score similarity and the availability of a highly optimized open-source implementation as a stand-alone C / C++ library that can be loaded by CGE
[25] . Given the highly parallel implementation of CGE, users can query the knowledge graph in a few seconds and perform millions of protein similarity calculations, enabling easy filtering and ranking of solutions by similarity score.
[0060] To normalize the scores, each sequence can be compared to itself. The product of the square roots of these scores is used as the denominator
[26] , as outlined in Listing 1.
[0061]
[0062] Listing 1. Method for normalizing Smith-Waterman scores
[0063] Embodiments of the systems and methods described herein can allow researchers to understand how similar or different the novel coronavirus is from other known viruses. If parts of the protein sequences that make up the novel virus have sequence and functional overlap with other known viruses, the information in the integrated knowledge graph helps us infer searches to identify potential candidate known drugs that inhibit viral pathogenic activity on the known viruses. Now, a simple example query for performing similarity-based inference is provided in the following paragraphs.
[0064] COVID-19 Similarity: The COVID-19 protein sequence consists of several non-structural proteins, envelope proteins, spike proteins, etc. To hypothesize potential drugs that bind to or interact with different parts of the COVID-19 viral proteins, embodiments can first identify open science proteins that have structures similar to the novel COVID-19 mutations. The query in Listing 2 is an example of finding the protein most similar to the COVID-19 spike protein sequence.
[0065]
[0066] SPARQL Query for Ranking Proteins Similar to a Reference Protein
[0067] The example similarity query in Listing 2 first looks up the protein sequence of SPIKE_SARS2 using the Uniprot mnemonic. Next, it retrieves the sequences of all proteins with sequences and names. Finally, each of these sequences is compared to the sequence of SPIKE_SARS2, and the similarity scores are saved in the variable sim. The bind clause saves all the sim values in a temporary table so that they can be used for other operations and returned to the user. In this query, any results with similarity scores less than 0.1 are removed, and the results are returned in descending order by similarity score.
[0068] The protein with the highest similarity score returned is A0A2D1PX97, which is "Bat SARS-like coronavirus", with a similarity score of 0.817. Several of the top-ranked results are bat coronaviruses or coronaviruses of other species, as shown in the top 10 results listed in Table II. The similarity scores drop rapidly from 0.79 to 0.37, which is the point at which Middle East Respiratory Syndrome (MERS), where the similarity score of protein A0A2I6PIX8 (i.e., "Middle East Respiratory Syndrome-related coronavirus") is 0.368, first appears in the results. Several proteins of MERS are followed by several coronaviruses in other species, including cattle, humans, rabbits, and mice, and several non-coronavirus proteins start to appear, such as A0A1B2RX89 of "Infectious bronchitis virus" with a similarity score of 0.322. These scores are consistent with studies indicating that COVID-19 may have originated in bats and is very similar to MERS
[27] .
[0069] Table II
[0070] Top 10 Protein Sequences Most Similar to the COVID-19 Spike
[0071] Protein Scientific name Score A0A2DIPX9T "Ba SARSike coonavins" 0.817 A0A0U2WM2 "SARS-ike corona inis WTV16" 0.817 A0A 2DIPXA9 "Bat SARS-ike coronavins" 0.816 U5WLK5 "Bat SARS-ke coronavins RSHC0147 0.814 A0A2D1PX29 "Bat SARS-ke coronavins" 0.814 U5WHZ7 "Bat SARS-iRe comonavins Rs33677" 0.813 U5W105 "Bat SA RsS-ike coronavius wiVT" 0.813 A0A2D1PXC0 "Bat SARShke corouavins" 0.813 A0A2D1PXDS "Bat SAR-ie coronav ins5" 0.812 A0A4Y6GL47 "Coronavinus BtRs-BetaCoV / YN2018B" 0.812
[0072] COVID-19 Drug Repurposing: After obtaining the similarity analysis results, embodiments can be implemented to utilize the knowledge graph to rank based on similarity scores to find potential drugs that can be repurposed for COVID-19. To this end, using a SPARQL query configured to work backwards rather than looking up all known targets of a given compound, the query starts with an unknown protein and searches for potential compounds that might target it. In this case, the example focuses on compounds that have an inhibitory effect on the protein. The example query in Listing 3 can be used to perform this search.
[0073]
[0074] SPARQL Query in Listing 3 to Find Potential Drugs for Repurposing
[0075] This example query has three main components. First, the inner query at the top searches ChEMBL for information on inhibitory active compounds that have passed a certain development stage. Since the intention is to repurpose existing drugs, these compounds may be limited to those in clinical trial development stages of phase 3 or higher. In the second inner query, all proteins that are known targets of a given compound are compared to the COVID-19 spike protein. The results are put in descending order by similarity to the SPIKE_SARS2 protein, and only the top 150 isotypes of proteins are returned. There are usually multiple sequences for a given protein, and in this case, the top 150 isotypes are associated with approximately the top 50 most similar proteins. The last part of the query again matches the selected proteins with the compounds that target them and the activity information from the first inner query. The final results are returned in descending order by similarity score to highlight compounds that potentially may be repurposed based on their similarity to the COVID-19 spike.
[0076] The reverse query returns compounds that target the proteins, where the similarity scores range from 0.2 down to 0.183. Several of these compounds are for drugs that have been put into clinical trials as they have the potential to be repurposed against COVID-19
[29] . Some of the protein sequences in the highest-scoring proteins against the COVID-19 spike found by the reverse query in clinical trials are also shown in Table III.
[0077] Table III
[0078] Example Drugs in COVID-19 Clinical Trials Present in Current Reverse Query Results
[0079] Protein Compound name Score P52333 BARICITINIB 0.194 P17948 RIBAVIRIN 0.189 P17948 RITONAVIR 0.189 P17948 DEXAMETHASONE 0.189 P17948 AZITHROMYCIN 0.189 P08183 LOPINAVIR 0.187
[0080] One way we validate the results returned by the reverse query is to compare the overlap between the compounds returned and the compounds currently used in clinical trials for COVID-19. Based on the drugs that are currently part of clinical trials in early June, we created a list of 196 unique drugs to compare with our results. For the above query considering the top 150 isoform sequences most similar to the SPIKE_SARS2 protein, the results returned by the knowledge graph included 91 out of 196 compounds (46%). The significant overlap between the compounds found by the reverse query and the clinical trial list also helps to define the score range that can be considered of interest. Since the similarity scores of the proteins found by the reverse query all fall between 0.183 and 0.20, it seems reasonable that compounds targeting other proteins with scores in the same range may also have a beneficial effect on COVID-19.
[0081] New hypothesis (tetanus): One potentially interesting result returned by the reverse query using SPIKE_SARS2 is tetanus toxin, which has the Uniprot identifier P04958 and the mnemonic TETX_CLOTE. The reverse query against the spike returns TETX_CLOTE as the top match, with a similarity score of 0.20. Given the high proportion of asymptomatic COVID-19 positive cases, currently estimated by the Centers for Disease Control and Prevention (CDC) to be 40%
[30] , the TETX_CLOTE result raises an unexpected but interesting hypothesis that the tetanus vaccine may contribute to the asymptomatic rate and reduce the severity of symptoms by enabling the immune system to generate a reasonable response to the virus. According to the CDC, in 2017, approximately 63.4% of adults 19 years and older in the United States had received some form of tetanus vaccine as recommended in the past 10 years, with a significant decline among individuals 65 years and older. Although tetanus is caused by bacteria and COVID-19 is a virus, there are multiple examples of heterologous immunity between bacteria and viruses. This heterologous immunity is at least initially attributed to the amino acid sequence similarity of the T and B cell epitopes of different pathogen antigens.
[0082] Database performance
[0083] To facilitate COVID-19 research, the life science knowledge graph is hosted on a few larger Cray XC-40 development systems. These systems mainly consist of a combination of Intel Broadwell, Skylake, and Cascade Lake processors. The file containing the N-triples used to build the database and the built database is stored on the attached Lustre file system and striped to match the number of available object storage targets (OSTs) in the file system.
[0084] The performance results in this section were run on an internal 370-node (336 compute nodes, 34 service nodes) XC-40 development system with dual-socket 48-core Skylake nodes and a mix of 48-core Skylake nodes and 56-core Skylake nodes in the frequency range of 2.1 GHz to 2.4 GHz. Most of these nodes have 192GB DDR-2666 memory, but 63 Cascade Lake nodes have larger 384GB DDR4-2933 memory. The attached Lustre file system is a Sonexion CS-L300N system with 8 OSTs providing 655TB of storage. Database build and load times are determined by the I / O performance into and out of the Lustre file system, so I / O system performance is an important consideration. The reported query execution times are strict query times and do not include the time required to write the results to the Lustre file system, which is common practice.
[0085] Database Build: As previously mentioned in the CGE Background section, the first step that CGE performs is to build a database from the input N-triple / N-quadruple file set to produce a compiled database in the representation used by the query engine. The original N-triple input files for the life science database are 28.29TB on Lustre. The build process is handled by CGE's dictionary component, and the build process consists of several steps, which are outlined in Table V along with the time (in seconds) for each step.
[0086] Table V
[0087] Time for the Build Steps of the Life Science Database
[0088]
[0089] As the numbers indicate, the build time of the database mainly depends on the time to read the original N-triple files from Lustre, which is expected. The times for the remaining build steps (i.e., ingestion, synchronization, and update) scale well from 128 nodes (2048 images) to 256 nodes (4096 images). The checkpoint of the built database is written to Lustre, so subsequent restarts of CGE with the same database can load the compiled database without having to ingest the original N-triples again. The built database is only about 5.4TB on disk, while the original N-triples are about 28.29TB on disk, and CGE can restart using the built database on 256 nodes (with 16 images per node) in about 568 seconds.
[0090] Spike similarity query: The first query used to test the CGE performance using the life science knowledge graph is the similarity query in Listing 2. This query searches in Uniprot for known proteins that meet specific conditions (such as having a sequence value and a recommended name), and compares them with the sequence value of SPIKE_SARS2. The similarity query found 49,299,877 protein sequences to compare with the SPIKE_SARS2 sequence. Table VI shows the time for the SW calculation for the 49.3 million protein sequences and the total query time when executed on 128 nodes and 256 nodes with 16 images or 32 images per node.
[0091] Looking at the time for the SW calculation, we observe that the time to calculate similarity scales well with both the image count (i.e., cores) and the node count. This is attributed to the fact that the calculation is independent of other calculations, so all images can compute subsets of protein sequences in parallel. The scaling also highlights the advantage of leveraging the SW calculation within the large-scale parallel context of the CGE. If the knowledge graph executed the same SW calculation on a single process, it would likely take approximately 21,709 seconds (i.e., 10.6 × 2048), thus making the query essentially infeasible in a serial context. The strict query time does show reasonable scaling from 128 nodes to 256 nodes with at least 16 images per node, but the query scaling is limited by the performance of the GROUP operator. Since the similarity scores are rounded to three decimal places, a large number of duplicate values need to be removed when storing the values as new variables (i.e., sin query variables) within the CGE. Due to the way the CGE distributes these new variables across images, a large number of duplicates can cause several images to wait for a small number of images to finish processing the values they will store. Additionally, the query scaling is also related to recent performance improvements that enable operations such as GROUP and FILTER to fetch the required strings as blocks rather than individually. Protein sequences can be very long, ranging from hundreds to thousands of amino acids, so fetching these long strings as blocks from remote processes is crucial for preventing communication overhead from dominating query performance.
[0092] Spike Reverse Query: The next query used to test the performance of CGE on the life science knowledge graph is the reverse query from List 3. This query starts with the protein sequence of SPIKE_SARS2 and searches for similar proteins that are targets of acceptable compounds with a desired effect (i.e., inhibitory). Since several joins must be made on large intermediate results to find the unique desired protein or compound, the reverse query is much more complex than the similarity query. Although the complex joins affect the overall query time, the additional filtering imposed by the joins significantly reduces the number of protein sequences that must be compared. CGE has been optimized to filter out solutions early in the query during the scan and join phases by reusing information from previous parts of the query. This optimization can be used to reduce the size of the intermediate results that must be joined, which is useful not only for performance but also for the memory requirements of complex queries like reverse queries. For the case of the SPIKE_SARS2 reverse query, the number of proteins compared is only 1,165,914, which is much less than the 49.3 million proteins compared in the similarity query.
[0093] As shown by the numbers in Table VII, since CGE is able to significantly reduce the number of proteins considered by leveraging information within the knowledge graph, the SW time is a very small fraction of the query time. The strict query time for the reverse query is more than twice that of the similarity query time, which is expected since there are more complex joins in the reverse query. Most of the strict query time is dominated by the joins. For example, on 256 nodes with 16 images per node, the strict query time is 49.0 seconds, of which 34.52 seconds is spent on making joins. Even with complex joins, the performance scales reasonably well from 128 nodes to 256 nodes when using 16 images per node (1.82x speedup), but although the query is faster with 32 images per node, the scaling efficiency is not high. The limited scaling with more cores per node is related to the memory access bottleneck due to images accessing memory on the same node but in different slots.
[0094] Since no other known large semantic graph engines are able to load real-world life science datasets of this magnitude, it is not easy to compare the performance of CGE with other database engines. However, previous benchmarks have clearly shown that CGE is at least an order of magnitude faster than any competitor, especially when performing complex queries [8], using the standard LUBM trillion triple dataset. For the case of LUBM, the typical benchmark query is number 9, which makes multiple complex joins to search the dataset for a certain triangular relationship among entities [9].
[0095] Table VI
[0096] Expansion Results for Spike Similarity Query
[0097]
[0098] Table VII
[0099] Expanded results of reverse lookup
[0100]
[0101] Scalability advantage: The performance and scalability of CGE have been demonstrated in previous studies [8],
[10] and the results discussed above. This large-scale performance is crucial when attempting to search such large datasets interactively. While it is possible to use smaller graph engines to analyze the same data, the time required to execute simple queries would almost certainly be too long to make the system useful. Even when performing the most complex queries across graphs, CGE's ability to scale to hundreds of nodes and tens of thousands of images enables very large datasets to be rapidly ingested and searched within seconds. This scalability is also highly beneficial for UDFs, as it enables domain-specific functions (some of which may be computationally complex) to be easily applied to refine query results. These unparalleled capabilities of CGE [8] enable researchers to search very large real-world datasets to find compounds that can be effectively reused in real time.
[0102] Insights into the problem of drug repurposing
[0103] While this study has focused primarily on the application of knowledge graphs and SW sequence comparisons to COVID-19, these variations and queries can be applied to any number of diseases of interest. For example, if the SPIKE_SARS2 mnemonic in the similarity query in Listing 2 is changed to the Uniprot mnemonic U5TGX1_COWPX for vaccinia virus, the similarity query can be used to find the proteins most similar to a given vaccinia. In this case, running the similarity query returns multiple vaccinia viruses as the top three results, followed by taterapox, and then three camelpox viruses. More interesting results start to appear at the 7th position in the similarity list, namely, the Uniprot mnemonic V5QZD2_9POXV for the protein V5QZD2 for "vaccinia virus WAU86 / 88-1", with a similarity score of 0.892. Due to the similarity of vaccinia virus to variola virus (the pathogen of smallpox), it has been used more in human immunization than any other vaccine.
[0104] Reverse queries can also be used for other diseases to search for potential drugs that can be repurposed. For example, replacing the SPIKE_SARS2 mnemonic with the Uniprot mnemonic KITH_HHV11 for "Human Herpesvirus 1" (HHV1) returns Brivudine as the top compound that inhibits the protein P06479, with a similarity score of 0.984. Brivudine is known to have strong antiviral activity against varicella-zoster virus and herpes simplex virus type 1
[44] . The top 10 query results are actually for various HHV1 proteins with compounds such as Penciclovir and Acyclovir, which are known antiviral drugs targeting herpes simplex virus type 1
[44] ,
[45] .
[0105] These results clearly demonstrate that CGE enables knowledge graphs with SW sequence similarity UDFs to quickly and effectively find potential drugs that can be repurposed to target different diseases. These capabilities of CGE allow researchers to interactively search for potential candidate drugs and apply subject matter expertise to further refine the query results, potentially improving the effectiveness of the repurposed drugs.
[0106] Conclusions and Future Work
[0107] The examples described in this paper demonstrate that CGE, a large-scale parallel semantic graph engine, and other similar engines that can easily ingest and search graph databases of this size enable researchers to perform complex queries across datasets in seconds. Additionally, the performance of CGE allows researchers to interactively query a given database containing 155 billion triples to search for potential hidden connections that may not be found between the nodes of the graph.
[0108] This paper also discusses recent changes made to CGE to enable user-defined functions that allow users to apply domain-specific expertise to operations such as FILTER, GROUP, and ORDER. Our work focuses on leveraging an open-source implementation of the Smith-Waterman sequence similarity algorithm as a UDF within the query to rank the proteins targeted by known compounds based on their similarity to a given reference sequence. Using the SPIKE_SARS2 protein as a reference, we show how to write a query that enables CGE to find potential drugs that can be repurposed for COVID-19 within seconds. Additionally, we demonstrate that these capabilities of CGE are not specific to COVID-19 and that it can easily be used to find potential drugs for repurposing for other known or newly interested diseases.
[0109] Using our two main queries of interest, we demonstrated a powerful extension of the Needleman-Wunsch function and showed good overall scaling of the queries themselves. The scaling tests showed some areas for future attention to improve scaling. First, the GROUP operator has a performance bottleneck due to generating too many duplicates, and as the core count increases, the query does not scale as expected. However, even with these limitations, for complex queries, we were able to show good scaling, further demonstrating the unique capabilities of CGE.
[0110] Finally, we have shown that the combination of the unique capabilities of CGE, including large-scale parallelism, complex query performance, and the scale of data that can be ingested, with protein sequence similarity can enable researchers to quickly and effectively repurpose existing drugs to target new diseases.
[0111] Figure XYZ depicts a block diagram of an example computer system XYZ00 in which various embodiments described herein may be implemented. The computer system XYZ00 includes a bus XYZ02 or other communication mechanism for communicating information, and one or more hardware processors XYZ04 coupled to the bus XYZ02 for processing information. The hardware processor(s) XYZ04 may be, for example, one or more general-purpose microprocessors.
[0112] The computer system XYZ00 also includes a main memory XYZ06, such as random access memory (RAM), cache, and / or other dynamic storage devices, coupled to the bus XYZ02 for storing information and instructions to be executed by the processor XYZ04. The main memory XYZ06 may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by the processor XYZ04. When these instructions are stored in a storage medium accessible to the processor XYZ04, they cause the computer system XYZ00 to be presented as a special-purpose machine customized to perform the operations specified in the instructions.
[0113] The computer system XYZ00 also includes a read-only memory (ROM) XYZ08 or other static storage device coupled to the bus XYZ02 for storing static information and instructions for the processor XYZ04. A storage device XYZ10 is provided, such as a magnetic disk, optical disk, or USB thumb drive (flash memory), etc., and is coupled to the bus XYZ02 for storing information and instructions.
[0114] The computer system XYZ00 can be coupled via a bus XYZ02 to a display XYZ12, such as a liquid crystal display (LCD) (or touch screen), for displaying information to a computer user. An input device XYZ14, including alphanumeric keys and other keys, is coupled to the bus XYZ02 for communicating information and command selections to a processor XYZ04. Another type of user input device is a cursor control XYZ16 (such as a mouse, trackball, or cursor direction keys) for communicating direction information and command selections to the processor XYZ04 and for controlling the movement of a cursor on the display XYZ12. In some embodiments, the same direction information and command selections as those of the cursor control can be implemented by receiving touches on a touch screen without a cursor.
[0115] The computing system XYZ00 can include a user interface module for implementing a GUI, which can be stored as executable software code executed by the computing device(s) in a mass storage device. By way of example, the module and other modules can include components such as software components, object-oriented software components, class components and task components, procedures, functions, attributes, processes, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables.
[0116] In general, as used herein, the words "component", "engine", "system", "database", "data store", etc. can refer to logic embodied in hardware or firmware, or to a collection of software instructions written in a programming language (such as Java, C, or C++) that may have entry and exit points. Software components can be compiled and linked into an executable program, installed in a dynamic link library, or can be written in an interpreted programming language (such as, for example, BASIC, Perl, or Python). It should be understood that software components can call from other components or from themselves, and / or can be called in response to detected events or interrupts. Software components configured to execute on a computing device can be provided on a computer-readable medium (such as a compact disc, digital video disc, flash drive, disk, or any other tangible medium) or as a digital download (and may initially be stored in a compressed or installable format that requires installation, decompression, or decryption before execution). Such software code can be stored, in whole or in part, on the memory device of the executing computing device for execution by the computing device. Software instructions can be embedded in firmware such as an EPROM. It should also be understood that hardware components can be composed of connected logic units such as gates and flip-flops, and / or can be composed of programmable units such as programmable gate arrays or processors.
[0117] The computer system XYZ00 can implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that, when combined with the computer system, causes the computer system XYZ00 to be a special-purpose machine or programs it to be a special-purpose machine. According to one embodiment, the techniques herein are performed by the computer system XYZ00 in response to one or more sequences of one or more instructions included in main memory XYZ06 being executed by the processor(s) XYZ04. Such instructions can be read into main memory XYZ06 from another storage medium, such as storage device XYZ10. Executing the instruction sequence included in main memory XYZ06 causes the processor(s) XYZ04 to perform the processing steps described herein. In an alternative embodiment, hardwired circuitry may be used in place of or in combination with software instructions.
[0118] The term "non-transitory medium" and like terms as used herein refer to any medium that stores data and / or instructions that cause a machine to operate in a particular fashion. Such non-transitory media can include non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device XYZ10. Volatile media includes dynamic memory, such as, for example, main memory XYZ06. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disk drives, solid state drives, magnetic tape, or any other magnetic data storage medium, CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge, and networked versions thereof.
[0119] Non-transitory media is different from but can be used in combination with transmission media. Transmission media participates in the transfer of information between non-transitory media. For example, transmission media includes coaxial cables, copper wire, and optical fibers that include wires that include bus XYZ02. Transmission media can also take the form of acoustic or light waves, such as acoustic or light waves generated during radio wave and infrared data communications.
[0120] Computer system XYZ00 also includes a communication interface XYZ18 coupled to bus XYZ02. The communication interface XYZ18 provides two-way data communication that is communicatively coupled to one or more network links, which are connected to one or more local networks. For example, the communication interface XYZ18 can be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem to provide a data communication connection with a corresponding type of telephone line. As another example, the communication interface XYZ18 can be a Local Area Network (LAN) card to provide a data communication connection with a compatible LAN (or a WAN component communicating with a WAN). A wireless link can also be implemented. In any such implementation, the communication interface XYZ18 transmits and receives electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information.
[0121] Network links typically provide data communication to other data devices through one or more networks. For example, a network link can provide a connection to a host computer or a data device operated by an Internet Service Provider (ISP) through a local network. The ISP in turn provides data communication services through the global packet data communication network now commonly referred to as the "Internet". Both the local network and the Internet use electrical, electromagnetic, or optical signals carrying digital data streams. Signals passing through various networks, network links, and the communication interface XYZ18 carrying digital data into and out of computer system XYZ00 are example forms of transmission media.
[0122] Computer system XYZ00 can send and receive data including program code through one or more networks, network links, and the communication interface XYZ18. In the Internet example, a server can transmit the request code of an application through the Internet, the ISP, the local network, and the communication interface XYZ18.
[0123] The received code can be executed by the processor XYZ04 when it is received, and / or stored in the storage device XYZ10 or other non-volatile memory for later execution.
[0124] Each of the processes, methods, and algorithms described in the foregoing sections can be embodied in code components executed by one or more computer systems or computer processors that include computer hardware, and be fully or partially automated by these code components. One or more computer systems or computer processors can also operate to support the execution of related operations in a "cloud computing" environment or as "software as a service" (SaaS) operations. The processes and algorithms can be implemented partially or fully in dedicated circuitry. The various features and processes described above can be used independently of each other or can be combined in various ways. Different combinations and sub - combinations are intended to fall within the scope of the present disclosure, and in some embodiments, certain method or process blocks can be omitted. The methods and processes described herein are also not limited to any particular order, and the blocks or states associated therewith can be executed in other appropriate orders, or can be executed in parallel, or in some other manner. Blocks or states can be added to or removed from the disclosed example embodiments. The execution of certain operations or processes in an operation or process can be distributed among computer systems or computer processors, not only residing within a single machine but also deployed across multiple machines.
[0125] As used herein, circuitry can be implemented using any form of hardware, software, or a combination thereof. For example, one or more processors, controllers, ASICs, PLAs, PALs, CPLDs, FPGAs, logic components, software routines, or other mechanisms can be implemented as constituting circuitry. In an embodiment, the various circuits described herein can be implemented as discrete circuits, or the described functions and features can be partially or fully shared among one or more circuits. Although various functional features or functional elements may be individually described or claimed as separate circuits, these features and functions can be shared among one or more common circuits, and such description should not be required or imply the need for separate circuits to implement these features or functions. Where the circuitry is implemented fully or partially using software, such software can be implemented to operate with a computing or processing system (such as computer system XYZ00) capable of performing the functions described thereof.
[0126] As used herein, the term "or" can be interpreted to include either the inclusive or the exclusive sense. In addition, the description of a resource, operation, or structure in the singular form should not be construed as excluding the plural form. Unless otherwise specifically stated or understood in the context in which it is used, conditional language such as "can", "could", "might", or "may" generally is intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements, and / or steps.
[0127] Unless otherwise expressly stated, the terms and phrases used in this document and their variants shall be construed as open-ended rather than limiting. Adjectives such as "conventional", "traditional", "normal", "standard", "known", etc., and terms with similar meanings shall not be construed as limiting the items described to a given time period or only to items available since a given time. Instead, they should be understood to cover conventional, traditional, normal, or standard techniques that are available or known at any time, present or future. In some instances, the presence of broadening words and phrases such as "one or more", "at least", "but not limited to", or other similar phrases shall not be construed to mean that a narrower situation was intended or required in instances where such broadening phrases may not be present.
Claims
1. A parallel processing graph database system for protein sequence analysis to determine viable therapeutic agents for a given disorder, the parallel processing graph database system comprises: at least one processor; and a memory including instructions which, when executed, cause the at least one processor to: receive a query from a user, the query including at least one fragment of a protein sequence of the given disorder, wherein the query includes one or more domain-specific functions; construct a sequence database by ingesting data from a file system and transforming the data for sequence mapping to compare a query sequence of the protein sequence of the given disorder with sequences of other known proteins in the sequence database; perform protein similarity analysis using the sequence database by executing a first domain-specific function in the query to determine the similarity between the query sequence and the sequences of the other known proteins in the sequence database, execute a second domain-specific function in the query to: determine a corresponding similarity score based on the similarity between the sequences of the other known proteins and the query sequence, and identify proteins with sequences of the other known proteins having a similarity score higher than a determined threshold, and identify one or more therapeutic agents associated with the identified proteins by querying a parallel processing graph database, the parallel processing graph database including potential therapeutic agents associated with the identified proteins that can have an inhibitory effect on the given disorder, wherein the query of the parallel processing graph database includes the identified proteins, and return a sorted list of at least a subset of the identified therapeutic agents associated with the identified proteins, wherein the subset of the identified therapeutic agents in the sorted list is sorted according to the similarity score of the identified proteins, wherein the one or more domain-specific functions are executed in parallel.
2. The system according to claim 1, wherein identifying the one or more therapeutic agents comprises: identifying drugs in the graph database known to have an inhibitory effect on the identified proteins with similarity scores higher than the determined threshold among the other known proteins.
3. The system according to claim 1, wherein determining viable therapeutic agents for the given disorder comprises: executing a query workflow across multiple datasets to generate drug hypotheses for the identified drugs.
4. The system according to claim 1, wherein querying the graph database comprises: searching the graph database to identify therapeutic agents having a desired effect on other known proteins having the same or similar sequence as the query sequence.
5. The system according to claim 1, wherein transforming the data for sequence mapping includes filtering, grouping, and sorting the sequences of the data executed in parallel with the one or more domain-specific functions.
6. The system according to claim 1, wherein the data from the file system includes N-triple / N-quadruple data.
7. The system according to claim 1, wherein the determined threshold is based on the query sequence.
8. A computing system for protein sequence analysis to determine viable therapeutic agents for a given disorder, the computing system comprises: a hardware processor; and a machine-readable storage medium coupled to the processor and storing a set of instructions, which when executed by the processor, cause the processor to perform operations, the operations including: determining a protein sequence of the given disorder from a query, wherein the query includes one or more domain-specific functions; constructing a sequence database by ingesting data from a file system and transforming the data for sequence mapping, and comparing the query sequence of the protein sequence of the given disorder with sequences of other known proteins in the sequence database; performing protein similarity analysis using the sequence database by executing a first domain-specific function in the query to determine the similarity between the query sequence and the sequences of the other known proteins in the sequence database; executing a second domain-specific function in the query to: determine a corresponding similarity score based on the similarity between the sequences of the other known proteins and the query sequence; and identify proteins of the sequences of the other known proteins having a similarity score higher than a determined threshold, and querying a graph database to identify potential therapeutic agents associated with the identified proteins that can have an inhibitory effect on the given disorder, and reconstructing the sequence database according to updated data in the file system and the result of querying the graph database, wherein the one or more domain-specific functions are executed in parallel.
9. The computing system according to claim 8, identifying the potential therapeutic agent comprises: identifying drugs in the graph database that are known to have an inhibitory effect on the identified proteins having a similarity score higher than the determined threshold among the other known proteins.
10. The computing system according to claim 8, wherein determining viable therapeutic agents for the given disorder comprises: executing a query workflow across multiple data sets to generate drug hypotheses for the identified drugs.
11. The computing system according to claim 8, wherein the operation of querying the graph database comprises: searching the graph database to identify therapeutic agents having a desired effect on other known proteins having the same or similar sequences as the query sequence.
12. The computing system according to claim 8, wherein transforming the data for sequence mapping includes filtering, grouping, and sorting the sequences of the data executed in parallel with the one or more domain-specific functions.
13. The computing system according to claim 8, wherein the data from the file system includes N-triple / N-quadruple data.
14. The computing system according to claim 8, wherein the instructions further cause the processor to perform operations including the following operations: querying a second set of proteins in the reconstructed sequence database, the second set of proteins having a higher likelihood of having an inhibitory effect on the given disorder by having a similarity score higher than a second determined threshold.
15. A non-transitory computer-readable medium stores a set of instructions that, when executed by a computer processing system, cause the computer processing system to perform operations, and the operations include: Determine a protein sequence of a given disease condition from a query, where the query includes one or more domain-specific functions; Construct a sequence database by ingesting data from a file system and converting the data for sequence mapping to compare a query sequence of the protein sequence of the given disease condition with sequences of other known proteins in the sequence database; Use the sequence database to perform protein similarity analysis by executing a first domain-specific function in the query to determine the similarity between the query sequence and the sequences of the other known proteins in the sequence database; Execute a second domain-specific function in the query to: Determine a corresponding similarity score based on the similarity between the sequences of the other known proteins and the query sequence, and Identify proteins of the sequences of the other known proteins having a similarity score higher than a determined threshold, and Query a graph database to identify potential therapeutic agents associated with the identified proteins that can have an inhibitory effect on the given disease condition, and Reconstruct the sequence database according to updated data in the file system and the results of querying the graph database, wherein the one or more domain-specific functions are executed in parallel.
16. The non-transitory computer-readable medium according to claim 15, wherein identifying the potential therapeutic agents that can have an inhibitory effect on the given disease condition includes: Identifying drugs in the graph database that are known to have an inhibitory effect on the identified proteins having a similarity score higher than the determined threshold among the other known proteins.
17. The non-transitory computer-readable medium according to claim 15, wherein determining feasible therapeutic agents for the given disease condition includes: Executing a query workflow across multiple data sets to generate drug hypotheses for the identified drugs.
18. The non-transitory computer-readable medium according to claim 15, wherein the operation of querying the graph database includes: Searching the graph database to identify therapeutic agents having an expected effect on other known proteins having sequences identical or similar to the query sequence.
19. The non-transitory computer-readable medium according to claim 15, wherein converting the data for sequence mapping includes filtering, grouping, and sorting sequences of the data executed in parallel with the one or more domain-specific functions.
20. The non-transitory computer-readable medium according to claim 15, wherein the data from the file system includes N-triple / N-quadruple data.
Citation Information
Patent Citations
System for excavating medicine related with disease gene in computer
CN101989297A
Evolution-based functional proteomics
US20040204861A1
Bioinformatics systems, apparatuses, and methods for performing secondary and / or tertiary processing
US20170270245A1
Method and apparatus for identifying candidate signatures and compounds for drug therapies
WO2020070485A1