Identifying sequence functionality
The system addresses the challenge of predicting gene functionalities by employing a computational model to perform tool queries, thereby enhancing the accuracy and efficiency of genome annotations.
Patent Information
- Application Number
- PCT/AU2024/051374
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-26
AI Technical Summary
Existing methods struggle to accurately predict the functionality of genes lacking ground truth labels, leading to a portion of genes remaining 'uncharacterized' in genome annotations generated by automated methods.
A system and method that utilize a computational model to analyze input queries related to target sequences, perform tool queries using available tools, and generate outputs indicative of the functionality associated with the target sequences.
The method enables efficient identification of gene functionalities by leveraging computational models and tool queries, reducing the reliance on manual expertise and improving the accuracy of genome annotations.
Smart Images

Figure AU2024051374_26062025_PF_FP_ABST
Abstract
Description
IDENTIFYING SEQUENCE FUNCTIONALITYBackground of the Invention
[0001] The present invention relates to a system and method for identifying functionality associated with biological sequences, and in one particular example, to a system and method for annotating nucleotide sequences to identify the function of genes.Description of the Prior Art
[0002] The reference in this specification to any prior publication (or information derived from it), or to any matter which is known, is not, and should not be taken as an acknowledgement or admission or any form of suggestion that the prior publication (or information derived from it) or known matter forms part of the common general knowledge in the field of endeavour to which this specification relates.
[0003] A large number of genes lack ground truth labels, i.e., functions in organisms validated by biological experiments or silico experiments, in public databases. This situation impedes the ability to predict gene functions, particularly when using Al / computational methods. Consequently, a portion of genes remains ‘uncharacterized’ in genome annotations generated by automated methods.
[0004] Whilst some software systems assist with manual genome annotation, they are either focused on providing a platform for users to share / collect information or a tool to find specific types of information. It still requires significant expertise and manual effort from users, who must curate and validate the information provided by these systems.Summary of the Present Invention
[0005] In one broad form, the present invention seeks to provide a method of identifying functionality associated with biological sequences, the method including, in one or more electronic processing devices: receiving an input query relating to a target sequence; using a computational model to: analyse the input query to identify required information; perform a tool query to ascertain the required information by: selecting one of a number of available tools based on the required information; constructing a tool input using the required information;applying the tool input to the tool; and, receiving a tool response from the tool; and, construct a query response, at least in part using the tool response; and, generating an output using the query response, wherein the output is at least partially indicative of functionality relating to the target sequence.
[0006] In one embodiment the method includes, in the one or more electronic processing devices, using the computational model to perform a sequence of tool queries to ascertain the required information.
[0007] In one embodiment the method includes, in the one or more electronic processing devices, analysing a tool response from a previous tool query to generate a tool input for a subsequent tool query.
[0008] In one embodiment the method includes, in the one or more electronic processing devices, using the computational model to: analyse the tool response and determine if the tool response includes the required information; and, one of: construct the query response if the tool response includes the required information; and, perform a further tool query if the tool response does not include the required information.
[0009] In one embodiment the method includes, in the one or more electronic processing devices, performing a further tool query by at least one of: selecting a different one of the tools; and, constructing a different tool input.
[0010] In one embodiment the method includes, in the one or more electronic processing devices: using the computational model to analyse the tool response and determine further required information; and, performing a further tool query to ascertain the further required information.
[0011] In one embodiment the method includes, in the one or more electronic processing devices: applying a pre-prompt to the computational model, the pre-prompt including an indication of available tools and their function; and, applying a prompt to the computational model, the prompt including the input query, wherein the computational model is responsive to the prompt to: analyse the input query; and, perform a tool query using one of the tools specified in the pre-prompt.
[0012] In one embodiment the method includes, in the one or more electronic processing devices, using the computational model to identify required information at least in part by querying a user.
[0013] In one embodiment the method includes, in the one or more electronic processing devices, using the computational model to perform tool queries to at least one of: perform a search to identify homologous sequences, the homologous sequences being homologous with respect to the target sequence; find sequence names relating to one or more homologous sequences; perform a search to identify functions associated with one or more homologous sequences or one or more sequence names; analyse functional relationships between one or more sequences; and, determine a candidate function of the target sequence at least in part based on at least one of identified functions and functional relationships; and, generate a sequence annotation using the candidate function.
[0014] In one embodiment the method includes, in the one or more electronic processing devices, using the computational model to perform a search to validate the candidate function.
[0015] In one embodiment the method includes, in the one or more electronic processing devices, using the computational model to analyse functional relationships between one or more sequences using a knowledge graph.
[0016] In one embodiment the method includes, in the one or more electronic processing devices: receiving an input query including a target biological sequence; and, using the computational model to perform one or more of the following steps: perform a search to identify homologous sequences, the homologous sequences being homologous with respect to the target sequence; identify sequence names for one or more genes associated with the homologous sequences; perform a search to identify functions associated with one or more genes; construct knowledge graphs indicative of functional relationships for the one or more genes; determine a function associated with the one or more genes; perform a search to validate functionality relating to the one or more genes; and, generate a gene annotation.
[0017] In one embodiment the target sequence is at least one of: a gene sequence; a nucleotide sequence; a protein sequence; and, an amino acid sequence.
[0018] In one embodiment the computational model is a large language model.
[0019] In one broad form, the present invention seeks to provide a system for identifying functionality associated with biological sequences, the system including one or more electronic processing devices configured to: receive an input query relating to a target sequence; use a computational model to: analyse the input query to identify required information; perform a tool query to ascertain the required information by: selecting one of a number of available tools based on the required information; constructing a tool input using the required information; applying the tool input to the tool; and, receiving a tool response from the tool; and, construct a query response, at least in part using the tool response; and, generate an output using the query response, wherein the output is at least partially indicative of functionality relating to the target sequence.
[0020] In another broad form, the present invention seeks to provide a non-transitory computer- readable medium comprising computer-executable instructions that when executed perform the method of the preceding paragraphs.
[0021] It will be appreciated that the broad forms of the invention and their respective features can be used in conjunction and / or independently, and reference to separate broad forms is not intended to be limiting. Furthermore, it will be appreciated that features of the method can be performed using the system or apparatus and that features of the system or apparatus can be implemented using the method.Brief Description of the Drawings
[0022] Various examples and embodiments of the present invention will now be described with reference to the accompanying drawings, in which: -
[0023] Figure 1 is a flow chart of an example of a process for identifying functionality associated with biological sequences;
[0024] Figure 2 is a schematic diagram of an example of a system architecture for identifying functionality associated with biological sequences;
[0025] Figure 3 is as schematic diagram of an example of a processing system;
[0026] Figure 4 is a schematic diagram of an example of a client device;
[0027] Figure 5 is a flow chart of an example of a process for performing queries using a computational model;
[0028] Figure 6 is a flow chart of an example of typical stages involved in identifying functionality associated with biological sequences;
[0029] Figure 7 is a schematic diagram of an example of the functionality of a system for identifying functionality associated with biological sequences;
[0030] Figure 8 is an image of an example of a user interface for answering a question by searching scientific papers;
[0031] Figure 9 is a flow chart of an example of the process for answering questions by searching;
[0032] Figure 10 is an image of an example of a user interface illustrating the generation of summary information based on an uploaded paper;
[0033] Figure 11 is an image of an example of a user interface for searching different data sources;
[0034] Figure 12 is a schematic diagram of an example of a knowledge graph;
[0035] Figure 13 is a flowchart of an example of a process of operation for an Al agent; and,
[0036] Figure 14 is a flowchart of an example of a process of human collaboration with an Al agent.Detailed Description of the Preferred Embodiments
[0037] An example of a process for identifying functionality associated with biological sequences will now be described with reference to Figure 1.
[0038] For the purpose of illustration, it is assumed that the process is performed at least in part using one or more electronic processing devices forming part of one or more processingsystems, such as computer systems, or the like. Whilst the system can use multiple processing devices, with processing performed by one or more of the devices, for the purpose of ease of illustration, the following examples will refer to a single device, but it will be appreciated that reference to a singular processing device should be understood to encompass multiple processing devices and vice versa, with processing being distributed between the devices as appropriate.
[0039] In this example, at step 100 the processing device receives an input query relating to a target sequence. The nature of the query and the manner in which this is received will vary depending on the preferred implementation and the example usage. For example, the input query could include all or part of a biological sequence that requires annotation with functional information, such as a nucleotide sequence representing one or more genes and / or an amino acid sequence representing one or more proteins. Additionally and / or alternatively, the query could include the name of a gene or protein, where information is required regarding related genes or proteins, for example, identifying genes that code for certain proteins. The query is typically received at least in part based on user inputs, and may include having a user upload or retrieve all or part of the sequence, for example retrieving the sequence from a data store, a sequencing machine, or the like.
[0040] The following steps 110 to 160 are then performed by the processing device with the assistance of a computational model. The nature of the computational model and the manner in which this is used will vary depending on the preferred implementation. In one example, the computational model includes a large language model, which is configured to use multiple tools, to allow specific tasks to be performed, such as performing searching of databases, performing literature searches and / or reviews, constructing knowledge graphs, or the like. Specific example tasks will be described in more detail below. The computational model can be configured to perform the tasks in any appropriate manner. For example, this could include generating a custom computational model that has been trained using particular training data, and / or could include configuring the computational model using specific prompts, or similar.
[0041] At step 110, the processing device uses the computational model to analyse the input query to identify required information. Specifically, this is performed to allowing thecomputational model to understand the query and identify information that is needed to successfully respond to the query.
[0042] Following this, the processing device uses the computational model to perform a tool query to ascertain the required information. This is achieved by having the computational model select one of a number of available tools based on the required information at step 120. For example, this could involve selecting to search a sequence database if the query includes an unknown sequence, to look for the sequence or homologous sequences. Conversely, this might involve a literature search, for example if an identity of the sequence is known, but the function is not.
[0043] At step 130, the processing device uses the computational model to construct a tool input using the required information. This step will typically involve formulating an input in a format that can be interpreted by the tool, for example creating a search query including relevant search terms, or the like. At step 140, the tool input is applied to the tool, for example by submitting a search query to a database, or the like.
[0044] At step 150, the computational model receives a tool response from the tool, for example, returning search results, or similar. The computational model then uses the tool response to formulate a response to the original query at step 160, for example by extracting relevant information from the tool response. As part of this, the computational model may determine that the query has not been fully responded to, in which case the process might return to step 120, allowing further tool queries to be performed, either by performing additional queries from the same tool and / or by query alternative tools, optionally using information derived from the previous tool response, which in turn allows multiple tool queries to be performed in sequence.
[0045] Finally, at step 170, the processing device generates an output using the query response, wherein the output is at least partially indicative of functionality relating to the target sequence.
[0046] Accordingly, the above described approach can be used to allow functionality associated with biological sequences to be determined using a computational model, such as a large language model. In particular, the computational model is configured to act as an Alagent, facilitating access to a number of tools through a single Al agent interface. In one example, this is achieved by making a number of tools available to the computational model, allowing the model to selectively query these depending on the information submitted by a user in an input query. This in turn allows the computational model to provide enhanced functionality, making the process of identifying functionality associated with biological sequences more straightforward.
[0047] A number of further features will now be described.
[0048] In one example, the method includes having the computational model perform a sequence of tool queries to ascertain the required information. This could be achieved in any suitable manner, but in one particular example, involves having the processing device analyse a tool response from a previous tool query to generate a tool input for a subsequent tool query. This allows the computational model to take the result of one tool query, and use this to perform a subsequent tool query, so that the original input query can be resolved via a series of steps that progressively identify further information, to ultimately allow the input query to be answered.
[0049] Thus, in one example, the processing device can use the computational model to analyse the tool response and determine if the tool response includes the required information. In this instance, if the tool response includes the required information, the query response can be constructed. Otherwise, the processing device can cause the computational model to perform a further tool query if the tool response does not include the required information, for example by selecting a different one of the tools and / or constructing a different tool input.
[0050] In performing further queries, the processing device would typically use the computational model to analyse the tool response, determine further required information and perform a further tool query to ascertain the further required information. However, this is not essential and alternative approaches might be used, such as seeking further input from a user, either by having the user review information provided in a tool response to ascertain if this is sufficient and / or to identify the further required information.
[0051] Performing a sequence of tool queries in this fashion can be important as information is often contained in various different locations and may therefore need to be identified using different approaches. For example, the computational model may need to query a sequence database in order to identify relevant genes, and then search literature or other databases in order to ascertain functions associated with the genes. Additionally and / or alternatively, the form of the query might be important, and so multiple different queries might be required before the required information is provided by a particular tool.
[0052] In one example, the processing device is configured to apply a pre-prompt to the computational model, the pre-prompt including an indication of available tools and their function. This configures the computational model, so that it is then able to make decisions regarding which tool should be used to answer a query, and also how the tools should be used, for example the required format for inputs to the tool. Once the pre-prompt has been applied, the processing device can apply a prompt including the input query to the computational model, with the computational model being responsive to the prompt to analyse the input query and perform a tool query using one of the tools specified in the pre-prompt.
[0053] In one example, the processing device uses the computational model to identify required information at least in part by querying a user. Thus, in this example, the user is identified as a tool that is able to be used in an attempt to resolve queries that other tools are unable to handle. This additional use of user interaction facilitates the overall process, and allows the user to participate in the process in the event that other tools are not able to wholly resolve the input query. This human / model interaction can improve the ability of the system to correctly identify the required information, leading to significantly better outcomes.
[0054] As mentioned above, the computational model can be used to perform a sequence of tool queries. The exact sequence of tool queries performed will vary depending on a range of factors, including the nature of the input query, the functionality of the tools, and the information the tools are able to provide. An example of the sequence of steps performed when annotating a target sequence with information regarding sequence functionality are set out below:1. perform a search to identify homologous sequences, the homologous sequences being homologous with respect to the target sequence;2. find sequence names relating to one or more homologous sequences;3. perform a search to identify functions associated with one or more homologous sequences or one or more sequence names;4. analyse functional relationships between one or more sequences;5. determine a candidate function of the target sequence at least in part based on at least one of identified functions and functional relationships; and,6. generate a sequence annotation using the candidate function.
[0055] However, it will be appreciated that the exact sequence of steps used will vary depending on a range of factors, such as the information sought, the tool responses provided by the tools, or the like.
[0056] Examples of additional steps that can be performed include performing a search to validate the candidate function, for example searching literature or other data sources; or constructing and / or analysing a knowledge graph of relationships between different genes or proteins and their functions, thereby allowing functional relationships to be ascertained, and in turn used in establishing sequence function.
[0057] In one specific example, the method includes having the electronic processing device receive an input query including a target biological sequence, and then use the computational model to perform one or more of: searching to identify homologous sequences, the homologous sequences being homologous with respect to the target sequence; identify sequence names for one or more genes associated with the homologous sequences; perform a search to identify functions associated with one or more genes; construct and / or analyse knowledge graphs indicative of functional relationships for the one or more genes; determine a function associated with the one or more genes; perform a search to validate functionality relating to the one or more genes; and, generate a gene annotation.
[0058] In the above examples, the target sequence is typically a gene sequence, a nucleotide sequence, a protein sequence, an amino acid sequence, or the like, whilst the computational model is typically a large language model, or other similar computational model.
[0059] A specific example of a system for performing the above described functionality will now be described in more detail with reference to Figures 2 to 4.
[0060] In this example, the system 200 includes a number of client devices 220 and processing systems 230, such as one or more servers, in communication with the client devices 220 via one or more communications networks 240.
[0061] It will be appreciated that the configuration of the networks 240 are for the purpose of example only, and in practice the client devices 220 and the processing system 230 can communicate via any appropriate mechanism, such as via wired or wireless connections, including, but not limited to mobile networks, private networks, such as an 802.11 networks, the Internet, LANs, WANs, or the like, as well as via direct or point-to-point connections, such as Bluetooth, or the like.
[0062] Whilst the processing systems 230 are shown single entities, it will be appreciated that in practice each processing system 230 can be distributed over a number of geographically separate locations, for example as part of a cloud based environment. However, the above described arrangement is not essential and other suitable configurations could be used.
[0063] An example of a suitable processing system 230 is shown in Figure 3. In this example, the processing system 230 includes at least one microprocessor 331, a memory 332, an optional input / output device 333, such as a keyboard and / or display, and an external interface 334, interconnected via a bus 335 as shown. In this example the external interface 334 can be utilised for connecting the processing system 230 to peripheral devices, such as the communications networks 240, databases 231, other storage devices, or the like. Although a single external interface 334 is shown, this is for the purpose of example only, and in practice multiple interfaces using various methods (eg. Ethernet, serial, USB, wireless or the like) may be provided.
[0064] In use, the microprocessor 331 executes instructions in the form of applications software stored in the memory 332 to allow the required processes to be performed. The applications software may include one or more software modules, and may be executed in a suitable execution environment, such as an operating system environment, or the like.
[0065] Accordingly, it will be appreciated that the processing system 230 may be formed from any suitable processing system, such as a suitably programmed client device, PC, web server, network server, or the like. In one particular example, the processing system 230 is a standard processing system such as an Intel Architecture based processing system, which executes software applications stored on non-volatile (e.g., hard disk) storage, although this is not essential. However, it will also be understood that the processing system could be any electronic processing device such as a microprocessor, microchip processor, logic gate configuration, firmware optionally associated with implementing logic such as an FPGA (Field Programmable Gate Array), or any other electronic device, system or arrangement.
[0066] As shown in Figure 4, in one example, the client device 220 includes at least one microprocessor 421, a memory 422, an input / output device 423, such as a keyboard and / or display, and an external interface 424, interconnected via a bus 425 as shown. In this example the external interface 424 can be utilised for connecting the client device 220 to peripheral devices, such as the communications networks 240, databases 231, other storage devices, or the like. Although a single external interface 424 is shown, this is for the purpose of example only, and in practice multiple interfaces using various methods (eg. Ethernet, serial, USB, wireless or the like) may be provided.
[0067] In use, the microprocessor 421 executes instructions in the form of applications software stored in the memory 422 to allow for communication with the processing systems 230, as well as to allow user interaction for example through a suitable user interface.
[0068] Accordingly, it will be appreciated that the client devices 220 may be formed from any suitable processing system, such as a suitably programmed PC, Internet terminal, lap-top, or hand-held PC, and in one preferred example is either a tablet, or smart phone, or the like. Thus, in one example, the client device 220 is a standard processing system such as an Intel Architecture based processing system, which executes software applications stored on nonvolatile (e.g., hard disk) storage, although this is not essential. However, it will also be understood that the client devices 220 can be any electronic processing device such as a microprocessor, microchip processor, logic gate configuration, firmware optionally associated with implementing logic such as an FPGA (Field Programmable Gate Array), or any other electronic device, system or arrangement.
[0069] For the purpose of the following examples, it is assumed that one or more processing systems 230 are servers, which communicate with the client devices 220 via a communications network, or the like, depending on the particular network infrastructure available. The servers 230 typically execute applications software for performing required tasks including storing, searching and processing of data, and in one example, can host or act as a gateway to one or more tools, and may also include servers hosting computational models. Actions performed by the servers 230 are performed by the processor 331 in accordance with instructions stored as applications software in the memory 332 and / or input commands received from a user via the VO device 333, or commands received from the client device 220.
[0070] Meanwhile, the user typically interacts with the client device 220 via a GUI (Graphical User Interface), or the like presented on a display of the client device 220, and in one particular example via a browser application that displays webpages, or an App that displays relevant information. Actions performed by the client devices 220 are performed by the processor 421 in accordance with instructions stored as applications software in the memory 422 and / or input commands received from a user via the VO device 423.
[0071] However, it will be appreciated that the above described configuration assumed for the purpose of the following examples is not essential, and numerous other configurations may be used. It will also be appreciated that the partitioning of functionality between the client devices 220, and the servers 230 may vary, depending on the particular implementation.
[0072] An example of the process for using a computational model to access tools to answer an input query will now be described in more detail with reference to Figure 5. In this example, the process is implemented using an Al agent implemented using a computational model, such as a Large Language Model (LLM).
[0073] In this example, at step 500, the Al agent is implemented by applying a pre-prompt to the computational model to configure the computational model to use one or more tools. The form of the pre-prompt will vary depending on the preferred implementation and the nature of the computational model, and a specific example will be discussed in more detail below.
[0074] At step 510 the Al agent receives an input query. In one example, the input query includes a target sequence, such as a protein or gene sequence, which requires annotation with functionality. However, as will also be apparent from the forgoing, the input query could also be of another form, such as a question relating to a gene sequence or protein, or the like.
[0075] At step 520, the input query is configured as a prompt, which is applied to the computational model, causing the computational model to select a tool at step 530, and then generate a tool query at step 540. Thus, for example, this might involve having the computational model generate a search query for a search database based on the target sequence from the input query.
[0076] At step 550, the computational model receives and analyses a response from the tool. Specifically, this is performed to allow the computational model to assess whether the tool response answers the input query or whether additional searching is required at step 530. For example, an initial search of the database might reveal only partially relevant information, in which case further searching of the same and / or different databases might be required.
[0077] If the query is answered, at step 560, the agent can generate an output responding to the original query at step 570. Otherwise, the process can return to step 530 allowing further tool querying to be performed.
[0078] As part of this, the ability to seek user assistance can be embedded within the process, by defining the user as a tool available to the computational model. This allows the model to seek assistance from the user if it is otherwise unable to proceed.
[0079] An example of a sequence tool queries that are used to annotate a target sequence, such as a gene sequence will now be described with reference to Figure 6.
[0080] In this example, at step 600 the computational model performs a search of one or more databases in order to identify sequences that are homologous to the target sequence. Results of this search are used to identify names and synonyms associated with genes in the target sequence and / or homologous genes at step 610. At step 620, the computational model uses the names and synonyms to perform further searching in an attempt to identify functions, with these functions being analysed at step 630 in order to identify functional relationships withother genes. This can be performed by constructing and analysing knowledge graphs, as will be described in more detail below.
[0081] At step 640, the computational model uses results of the analysis to determine a candidate function for one or more genes in the target sequence, with a further search optionally being performed in order to validate the function at step 650. Finally, once this has been completed for genes of interest in the target sequence, an annotation can be generated at step 660.Case Study
[0082] Further example uses will now be described. For the purpose of these examples, the preferred implementation integrates Large Language Model (LLM)-driven automation with domain- specific tools and human intelligence to enhance the accuracy and efficiency of manual curation of gene functions.
[0083] The system can assist users in curating gene functions from public databases, web pages and scientific literature. The curated gene functions would help 1) biologists shape their hypotheses and design biological experiments, and 2) biocurators update current publications based on the latest understanding of the genome. Ultimately, the performance of automated methods is improved with the growth of ground truth data of genome annotations, which speed up the development of live science.
[0084] In this regard, the emergence of large language models (LLMs) has allowed agent development to be improved. Most such works emphasise the autonomy of Al agents, neglecting shortcomings, such as handling complex tasks, a propensity to hallucinate, and not achieving human-like decision-making capabilities.
[0085] In the current arrangement human and Al intelligence can be combined by means of a conversational platform with an Al agent, implemented using the LLM, as illustrated for example in Figure 7. This can alleviate common LLM shortcomings using retrieval-augmented generation and human- augmented generation. The Al agent in the current approach can assist humans in choosing suitable bioinformatics methods for their queries and answering their questions correctly. Ultimately, this can make the arrangement more useful and trustworthy inbiological research tasks, freeing up human professionals to focus on higher-level strategic thinking and creative problem- solving.
[0086] In considering manual curation in genome annotation, there are a number of general subtasks that are typically performed, such as navigating public databases and web pages, searching scientific literature and finding information in scientific papers of interest. To implement this, the current approach can use three components, which will each be described in detail, as well as how these are integrated into a conversation system, allowing users to use all the functions in a flexible and natural way.Searching scientific literature
[0087] In one example, users input their questions and the system will answer the question based on relevant scientific papers (abstract or full paper) identified in a repository such as PubMed. An example of the user interface used in this process is shown in Figure 8.
[0088] To achieve this, an LLM (such as OpenAI API) acts as a mediator to communicate with the PubMed database and answer user’s questions, using the steps shown in Figure 9.
[0089] Specifically, at step 900 the LLM reviews a chat history and new question entered by the user, using this to generate a customised prompt at step 905. This in turn is used to construct a standalone question at step 910, corresponding to a question for a particular tool, generating a corresponding prompt at step 915, which is used to construct a keyword query at step 920, which can be used to search a repository such as PubMed at step 925.
[0090] Results, including related abstracts and papers are received at step 930, with the LLM splitting these into chunks at step 935 and storing them in a data store at step 940. The chunks are reviewed for relevant text at step 945, with results of this being combined with the prior chat history 950, to create a further customised prompt at step 955, which can be used to answer the original query at step 960.Retrieval of information from full scientific papers
[0091] Users can use the system to narrow down papers of interest when searching literature. For those papers, users might like to read in detail. Rather than users reading the full longpapers by themselves, the system uses LLMs to ‘read’ papers for the users and answer any questions they may have. By studying the workflow of manual curation, LLMs can be used to generate the common interest information from the paper when the user uploads it, and an example of an interface showing this can be seen in Figure 10.
[0092] In this example, the information presented includes the citation, summary of the abstract, gene definitions mentioned in the abstract, gene-gene or gene-phenotype relation pairs mentioned in the abstract, evidence and conclusion text mentioned in the abstract. All these meta-data are generated through OpenAI API with specific customized prompts.
[0093] Besides this pre-generated information, the system can answer any questions based on the full paper, as will be described in more detail below.Navigating public databases and web pages
[0094] Similar to the PubMed search described above, the system can use LLMs to communicate with public databases and web pages, as well as generate answers to user’s questions.
[0095] To provide usefulness in genome annotation, this can include performing searches of repositories, such as NCBI databases, GO database, ECO database and Google search. In one example, the LLM can be taught to use Web APIs of NCBI, allowing the system to answer any questions related to NCBI databases, such as gene aliases, gene SNPs, gene-disease relation and nucleotide or amino-acid sequence alignment. This can be achieved using approaches similar to those described in "GeneGPT: Augmenting Large Language Models with Domain Tools for Improved Access to Biomedical Information" by Qiao Jin, Yifan Yang, Qingyu Chen, and Zhiyong Lu arXiv:2304.09667v3.
[0096] To normalize the free text to Gene ontology / evidence and conclusion ontology, the LLM can be taught to call quickGO REST API to find ontology terms for GO or ECO. In one example, this is achieved using the OpenAI API with specific customized prompts.
[0097] To expand search capabilities beyond PubMed and GO databases, the system can leverage additional search engines to retrieve general information most relevant to user queries.One example is using the langchain GoogleSearchAPIWrapper method to integrate general web search engines like Google into the system’s toolkit.
[0098] Rather than search different databases individually, the system provides a very simple way to navigate these databases and web pages together by inputting the question using the interface, as shown for example in Figure 11, and then generating summary information across results from multiple tools. The system uses an LLM -powered Al agent to choose which tools / APIs to use for specific questions.
[0099] The system can also address queries related to knowledge graphs. A biological knowledge graph, with molecular entities and traits as nodes, as shown in Figure 12, effectively maps the intricate relationships and interactions within biological systems. While knowledge graphs sometimes provide a visual representation of complex genetic and phenotypic interconnections, they can be challenging to navigate and interpret in their entirety. The system can extract entities from a user's question, retrieves any relationships related to these entities, and finally answers the question based on this information.
[0100] In one example, this process uses an LLM to extract relationships mentioned in the abstracts of selected papers relevant to the user's queries (interests). These relationships are organized as triplets (head entity, relation, tail entity). An entity could be a molecule, mutant, phenotype, or trait, and the relation indicates a type of interaction between the head and tail entities mentioned in the context, such as "interacts with". Building upon the extracted triplets, the langchain GraphlndexCreator method can be employed to construct a directed graph. The GraphlndexCreator utilizes an LLM model to first extract and index all of the knowledge triplets within a knowledge graph. This enables the system to effectively answer user queries based on the established knowledge graph.Al agent
[0101] As mentioned above, the system uses an LLM-powered Al agent to choose which tools / APIs to use for specific questions. These tools are categorized into: reading the full paper (knowledge base), calling NCBI EUTILS, Google Search, calling quickGO REST API, question-answering against a knowledge graph and asking the user for help, and an example of this is shown in Figure 13.
[0102] In order for this to operate correctly, the name, executable Python function, and description for each tool is defined and included in a pre -prompt. The description demonstrates the functionality of the specific tool, for example, question answering based on full papers. All the names and descriptions of tools will be incorporated into the prompt for the Al agent, which influences the performance of the Al agent in choosing the correct tools. For example, “knowledge base” performs better than “read the full paper” as an instruction to the LLM. For the Al agent, the langchain LLMS ingle Action Agent method is used with the customized prompt as below.Prompt = """You are AgentGPT, a professional research assistant who provides informative answers to users mostly in the biological field . You have access to the following tools :{tools}Use the following format :Question : the input question you must answer Thought : you should always think about what to do Action : the action to take, should be one of [{tool_names}] Action Input : the input to the action Observation : the result of the action . . . (this Thought / Action / Action Input / Observation can repeat N times) Thought : I now know the final answer Final Answer : the final answer to the original input questionBegin ! Remember to give detailed, informative answersPrevious conversation history:{history}New question : {input} {agent_scratchpad}" " " expanded_tools = [ Tool( name = ' Knowledge base ' ,# func=podcast_retriever . run, f unc=kb_chat . run, description="Useful for answering questions about the paper and for details on the paper. "),Tool( name= "Call Entrez Programming Utilities", func=GeneGPT, description="useful for when you need to answer questions about finding gene alias, gene-disease association, gene location, genome DNA alignment, gene name conversion, protein-coding genes or not, gene SNP association and SNP location, which can be done by searching the NCBI database. "),Tool( name = "Google Search",# func=search . run, func = top3_results, description="useful for when you need to answer questions about current events or cannot find the answer from the paper and Entrez Programming Utilities . " ), Tool( name= "Call quickGO REST API", func=quickGO, description="useful for when you need to answer questions about find ontology terms from GO (the Gene Ontology) and ECO (the Evidence & Conclusion Ontology) but you cannot find the answer by Google Search . " ), Tool( name= "Question-answering against a graph", func=KGbot, description="useful for when you need to answer questions about relationship inference from a knowledge graph . " ), Tool( name= "Human", func= input, description= "Provides guidance and assistance to the Al agent or LLM. " ), ]
[0103] The process of the Al agent is shown in Figure 13.
[0104] In this example, at step 1300 a chat history is accessed and a user question provided to the agent at step 1310. A semantic search is performed at step 1320, in conjunction with accessing the tool list from the pre -prompt at step 1330, allowing the agent to decide which tool it will use and generate the corresponding input for this tool at steps 1331-1335.
[0105] Unlike most existing Al agents, humans are included as an explicit tool, which enables the agent to query the human as an option at step 1335, bringing the collaboration between humans and Al for difficult tasks, which often happen in knowledge discovery of complex biology. Importantly, all the tools are independent modules. Therefore, the system can easily incorporate more tools as needed.
[0106] Following this, a response is constructed at step 1340, with this being repeated as needed, until a final answer can be generated at step 1350.
[0107] A further example of the collaborative nature of the interaction process is shown in Figure 14.
[0108] In this example, a query 1400 is supplied by a user, with this being proceed by the agent at 1405 and used to select and generate a tool query at 1410, with this being provided to a tool for execution at 1420. A tool response is generated at 1425 and provided to the agent for review at 1430. In this example, clarification is required, so a request for assistance is provided to the user at 1435, allowing the user to provide information at 1440. The agent reviews the information at 1445, and performs a further tool query at 1450, with the tool executing the query at 1455 and generating a response at 1460. The response is provided to the agent, which reviews the response at 1465, and generates a query response at 1470, providing this to the user at 1475. A user may then provide optional feedback at 1480.
[0109] Throughout this specification and claims which follow, unless the context requires otherwise, the word “comprise”, and variations such as “comprises” or “comprising”, will be understood to imply the inclusion of a stated integer or group of integers or steps but not the exclusion of any other integer or group of integers. As used herein and unless otherwise stated, the term "approximately" means ±20%.
[0110] Persons skilled in the art will appreciate that numerous variations and modifications will become apparent. All such variations and modifications which become apparent to persons skilled in the art, should be considered to fall within the spirit and scope that the invention broadly appearing before described.
Claims
THE CLAIMS DEFINING THE INVENTION ARE AS FOLLOWS:1) A method of identifying functionality associated with biological sequences, the method including, in one or more electronic processing devices: a) receiving an input query relating to a target sequence; b) using a computational model to: i) analyse the input query to identify required information; ii) perform a tool query to ascertain the required information by:(1) selecting one of a number of available tools based on the required information;(2) constructing a tool input using the required information;(3) applying the tool input to the tool; and,(4) receiving a tool response from the tool; and, iii) construct a query response, at least in part using the tool response; and, c) generating an output using the query response, wherein the output is at least partially indicative of functionality relating to the target sequence.2) A method according to claim 1, wherein the method includes, in the one or more electronic processing devices, using the computational model to perform a sequence of tool queries to ascertain the required information.3) A method according to claim 2, wherein the method includes, in the one or more electronic processing devices, analysing a tool response from a previous tool query to generate a tool input for a subsequent tool query.4) A method according to any one of the claims 1 to 3, wherein the method includes, in the one or more electronic processing devices, using the computational model to: a) analyse the tool response and determine if the tool response includes the required information; and, b) one of: i) construct the query response if the tool response includes the required information; and, ii) perform a further tool query if the tool response does not include the required information.5) A method according to claim 4, wherein the method includes, in the one or more electronic processing devices, performing a further tool query by at least one of:a) selecting a different one of the tools; and, b) constructing a different tool input.6) A method according to claim 4 or claim 5, wherein the method includes, in the one or more electronic processing devices: a) using the computational model to analyse the tool response and determine further required information; and, b) performing a further tool query to ascertain the further required information.7) A method according to any one of the claims 1 to 6, wherein the method includes, in the one or more electronic processing devices: a) applying a pre-prompt to the computational model, the pre-prompt including an indication of available tools and their function; and, b) applying a prompt to the computational model, the prompt including the input query, wherein the computational model is responsive to the prompt to: i) analyse the input query; and, ii) perform a tool query using one of the tools specified in the pre-prompt.8) A method according to any one of the claims 1 to 7, wherein the method includes, in the one or more electronic processing devices, using the computational model to identify required information at least in part by querying a user.9) A method according to any one of the claims 1 to 8, wherein the method includes, in the one or more electronic processing devices, using the computational model to perform tool queries to at least one of: a) perform a search to identify homologous sequences, the homologous sequences being homologous with respect to the target sequence; b) find sequence names relating to one or more homologous sequences; c) perform a search to identify functions associated with one or more homologous sequences or one or more sequence names; d) analyse functional relationships between one or more sequences; e) determine a candidate function of the target sequence at least in part based on at least one of identified functions and functional relationships; and, f) generate a sequence annotation using the candidate function.10) A method according to claim 9, wherein the method includes, in the one or more electronic processing devices, using the computational model to perform a search to validate the candidate function.11) A method according to claim 9 or claim 10, wherein the method includes, in the one or more electronic processing devices, using the computational model to analyse functional relationships between one or more sequences using a knowledge graph.12) A method according to any one of the claims 1 to 11, wherein the method includes, in the one or more electronic processing devices: a) receiving an input query including a target biological sequence; and, b) using the computational model to perform one or more of the following steps: i) perform a search to identify homologous sequences, the homologous sequences being homologous with respect to the target sequence; ii) identify sequence names for one or more genes associated with the homologous sequences; iii) perform a search to identify functions associated with one or more genes; iv) construct knowledge graphs indicative of functional relationships for the one or more genes; w v) determine a function associated with the one or more genes; vi) perform a search to validate functionality relating to the one or more genes; and, vii) generate a gene annotation.13) A method according to any one of the claims 1 to 12, wherein the target sequence is at least one of: a) a gene sequence; b) a nucleotide sequence; c) a protein sequence; and, d) an amino acid sequence.14) A method according to any one of the claims 1 to 13, wherein the computational model is a large language model.15)A system for identifying functionality associated with biological sequences, the system including one or more electronic processing devices configured to: a) receive an input query relating to a target sequence;b) use a computational model to: i) analyse the input query to identify required information; ii) perform a tool query to ascertain the required information by:(1) selecting one of a number of available tools based on the required information;(2) constructing a tool input using the required information;(3) applying the tool input to the tool; and,(4) receiving a tool response from the tool; and, iii) construct a query response, at least in part using the tool response; and, c) generate an output using the query response, wherein the output is at least partially indicative of functionality relating to the target sequence.16) A non-transitory computer-readable medium comprising computer-executable instructions that when executed perform the method of any one of claims 1 to 14.
Citation Information
Patent Citations
Anti-galectin-9 antibodies and uses thereof
US20190127472A1
Artificial intelligence platform for protein engineering
US20190259470A1
Protein database search using learned representations
US20220165356A1
RNA-protein interaction prediction method and apparatus, and medium and electronic device
WO2023044931A1
High-throughput prediction of variant effects from conformational dynamics
WO2023064874A1
Cited By
Multi-omics data integration plant gene function inference system and method based on large language model
CN120853669A