Domain task processing method, biological information domain task processing method, computing device, computer readable storage medium and computer program product
By building a task library and using the target generation model to generate call chain information, task units are automatically selected and executed, which solves the problems of insufficient efficiency and accuracy in existing technologies and improves the efficiency and accuracy of domain task processing.
Patent Information
- Application Number
- CN202410362622.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-27
- Publication Date
- 2025-09-30
AI Technical Summary
The efficiency of domain task processing in existing technologies is insufficient and its accuracy is difficult to guarantee. It requires continuous exploration and testing by domain experts to select appropriate task units and calling methods.
By building a task library for the target domain, retrieving domain task units based on task information, and using the target generation model to generate call chain information that conforms to the target task execution logic, the appropriate task unit is automatically selected and the task is executed.
It improves the efficiency and accuracy of domain task processing, reduces the workload of domain experts in manually organizing and updating task libraries, and achieves efficient and highly accurate task execution.
Smart Images

Figure CN120723345A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of data processing technology, and in particular to a field task processing method, a task processing method in the field of bioinformatics, a computing device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the rapid development of data in various fields, massive data has brought new challenges to domain task processing. How to effectively utilize this data has become the key to domain task processing.
[0003] Currently, the development and application of task units are the key to solving this challenge, and tasks are performed by calling task units.
[0004] However, selecting the right task unit and appropriately invoking it to execute tasks requires domain experts to continuously explore, compare, and test different task units and invocation methods, building comprehensive knowledge and experience to effectively execute domain tasks. This approach is inefficient and difficult to guarantee accuracy. Therefore, an efficient and highly accurate domain task processing method is urgently needed. Summary of the Invention
[0005] In light of this, embodiments of this specification provide a domain task processing method. One or more embodiments of this specification also relate to a task processing method in the field of bioinformatics, a domain task processing apparatus, a task processing apparatus in the field of bioinformatics, a computing device, a computer-readable storage medium, and a computer program product to address the technical deficiencies of the prior art, such as insufficient efficiency and difficulty in ensuring accuracy.
[0006] According to a first aspect of an embodiment of this specification, a domain task processing method is provided, including:
[0007] Obtain task information of target tasks in target domain;
[0008] Based on the task information, a domain task unit related to the target task is retrieved from a task library corresponding to the target domain, wherein the task library is constructed based on at least one sample task in the target domain and a sample task unit called by executing the sample task;
[0009] Based on the unit information of the domain task unit, the target generation model is used to generate call chain information that conforms to the target task execution logic;
[0010] Based on the call chain information, the domain task unit is called to execute the target task and obtain the task execution result.
[0011] According to a second aspect of the embodiments of this specification, a task processing method in the field of bioinformatics is provided, comprising:
[0012] Obtain task information of bioinformatics tasks in the field of bioinformatics;
[0013] Based on the task information, searching for domain task tools related to the bioinformatics task from a task library corresponding to the bioinformatics field, wherein the task library is constructed based on at least one sample task in the bioinformatics field and a sample task tool called to execute the sample task;
[0014] Based on the tool information of the domain task tools, the tool call chain information that conforms to the execution logic of the generated bioinformatics task is generated using the target generation model;
[0015] Based on the tool call chain information, the domain task tool is called to execute the bioinformatics task and obtain the task execution result.
[0016] According to a third aspect of the embodiments of this specification, a domain task processing device is provided, including:
[0017] A first acquisition module is configured to acquire task information of a target task in a target domain;
[0018] A first retrieval module is configured to retrieve, based on the task information, domain task units related to the target task from a task library corresponding to the target domain, wherein the task library is constructed based on at least one sample task in the target domain and a sample task unit called by executing the sample task;
[0019] The first generation module is configured to generate call chain information that conforms to the target task execution logic based on the unit information of the domain task unit and using the target generation model;
[0020] The first execution module is configured to call the domain task unit to execute the target task based on the call chain information and obtain the task execution result.
[0021] According to a fourth aspect of the embodiments of this specification, a task processing device in the field of bioinformatics is provided, comprising:
[0022] A second acquisition module is configured to acquire task information of a bioinformatics task in the field of bioinformatics;
[0023] The second retrieval module is configured to retrieve, based on the task information, domain task tools related to the bioinformatics task from a task library corresponding to the bioinformatics field, wherein the task library is constructed based on at least one sample task in the bioinformatics field and a sample task tool called to execute the sample task;
[0024] The second generation module is configured to generate tool call chain information that conforms to the bioinformatics task execution logic based on the tool information of the domain task tool using the target generation model;
[0025] The second execution module is configured to call the domain task tool to execute the bioinformatics task based on the tool call chain information and obtain the task execution result.
[0026] According to a fifth aspect of the embodiments of this specification, there is provided a computing device, including:
[0027] memory and processor;
[0028] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above method are implemented.
[0029] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, and the steps of the above method are implemented when the computer program / instruction is executed by a processor.
[0030] According to a seventh aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.
[0031] In one embodiment of the present specification, a comprehensive task library corresponding to the target domain is automatically constructed based on sample tasks and sample task units. Based on the task information of the target task, domain task units related to the target task are retrieved from the comprehensive task library, and appropriate task units are automatically selected. The target generation model, a deep learning model, is used to generate accurate call chain information that conforms to the execution logic of the target task. Based on the call chain information, the domain task unit is called to execute the target task, thereby realizing the appropriate calling of the task unit to execute the task and improving the efficiency and accuracy of domain task processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 This is a flowchart of a field task processing method provided by one embodiment of this specification;
[0033] Figure 2 This is a schematic diagram of a task library construction in a domain task processing method provided in one embodiment of this specification;
[0034] Figure 3 This is a schematic diagram of constructing a task library in another field task processing method provided by an embodiment of this specification;
[0035] Figure 4This is a flowchart of a task library construction method for a domain task processing method provided by an embodiment of this specification;
[0036] Figure 5 This is a flowchart of a method for processing a domain task provided by an embodiment of this specification;
[0037] Figure 6 This is a flowchart of a task processing method in the field of bioinformatics provided by one embodiment of this specification;
[0038] Figure 7 This is a process flow chart of a method for processing a domain task applied to the field of bioinformatics provided by one embodiment of this specification;
[0039] Figure 8 This is a front-end schematic diagram of a domain task processing method applied to the field of bioinformatics provided by an embodiment of this specification;
[0040] Figure 9 This is a schematic diagram of the structure of a domain task processing device provided by one embodiment of this specification;
[0041] Figure 10 This is a schematic diagram of the structure of a task processing device in the field of bioinformatics provided by one embodiment of this specification;
[0042] Figure 11 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION
[0043] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0044] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0045] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0046] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0047] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model. It is pre-trained on a large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as a large language model (LLM) and a multi-modal pre-training model.
[0048] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0049] First, the terms involved in one or more embodiments of this specification are explained.
[0050] Bioinformatics system: An intelligent system built on a large language model, specifically used in the field of bioinformatics, which can utilize and manage the knowledge graph of task tools in the field of bioinformatics to complete specific bioinformatics tasks.
[0051] Tool knowledge graph in the field of bioinformatics: a structured knowledge system that contains the correspondence between bioinformatics tasks and corresponding tools, aiming to help bioinformatics systems quickly retrieve and generate feasible tool call chains.
[0052] Tool call chain: refers to a series of bioinformatics tools organized in a specific order to perform complex bioinformatics analysis tasks.
[0053] Retrieval-Augmented Generation (RAG): A technique that combines retrieval mechanisms and generative models to improve the model's ability to handle complex tasks by retrieving relevant information to assist the generation process.
[0054] In this specification, a domain task processing method is provided. This specification also involves a task processing method in the field of bioinformatics, a domain task processing device, a task processing device in the field of bioinformatics, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.
[0055] See also Figure 1 , Figure 1 A flowchart of a domain task processing method provided by an embodiment of this specification is shown, including the following specific steps:
[0056] Step 102: Obtain task information of the target task in the target domain.
[0057] The embodiments of this specification are applied to applications, websites or platforms with domain task processing functions, for example, an application or website that deploys a large model, or an application or website that calls the large model through an application programming interface (Application Programming Interface, referred to as API), and executes the target task by calling the domain task unit of the target domain.
[0058] A domain is a specific area of knowledge, such as science, law, finance, e-commerce, etc. Domains can have multiple levels. For example, the bioinformatics domain has a more detailed level than the science domain.
[0059] The target domain is the knowledge domain where the target task to be processed lies. The target domain limits the task library and the domain task units to be retrieved and called to a certain extent. For example, the target domain is the field of bioinformatics.
[0060] The target task is a domain task to be processed in the target domain. The target task needs to be executed through the domain task unit according to a specific process. Multiple subtasks can be divided according to the process. Any subtask can be understood as one or more process steps. For example, the target domain is the bioinformatics field, and the target task is to explore the expression differences of specific gene markers in different cell types (such as neurons and hepatocytes). It includes gene labeling, sampling, threshold determination, gene screening, enrichment analysis, visualization processing and other subtasks.
[0061] The task information of the target task is the content information related to the target task, including but not limited to: task name, task requirements, task data information, task execution process and task execution constraints. For example, if the target task is to explore the expression differences of specific gene markers in different cell types (such as neurons and hepatocytes), the task information includes: task name (exploring the expression differences of specific gene markers in different cell types), cell type list (for example, neurons, hepatocytes, cardiomyocytes, etc.); gene marker list (such as specific transcription factors or disease-related genes); analysis requirements (whether statistical significance testing, visualization results, comparison benchmarks, etc. are required); data source restrictions (database name, time range, literature screening criteria, etc.).
[0062] An optional way to obtain task information of a target task in a target domain is to obtain a task request for the target domain, wherein the task request includes the task information of the target task. The task request for the target domain is an execution instruction request for the target task in the target domain, including the task information of the target task. For example, the target domain is the bioinformatics field, and the task request for the target domain is an execution instruction request for an API call, including the following task information: "Please retrieve and analyze relevant research data on specific gene markers in different cell types (such as neurons, hepatocytes), and compare the expression differences of these markers between different cells." For example, directly receiving a task request for the target domain sent by a front-end user, or generating a task request for the target domain based on the task information of the target task sent by the front-end user, is not limited here.
[0063] For example, a bioinformatics platform uses a large language model through an application programming interface. On the platform's front-end, users select the desired bioinformatics task information, including the cell type, input gene list, and data source. Based on this information, the platform's front-end generates a task request for the target domain and sends it to the platform's server.
[0064] Obtaining the task information of the target task in the target domain provides index information for the subsequent retrieval of domain task units, laying the foundation for the subsequent execution of the target task.
[0065] Step 104: Based on the task information, retrieve domain task units related to the target task from a task library corresponding to the target domain, wherein the task library is constructed based on at least one sample task in the target domain and a sample task unit called to execute the sample task.
[0066] The task library corresponding to the target domain is a systematic storage structure built for the target domain. According to the correspondence between tasks and task units, it records the task information of the sample task and the unit information of the sample task unit called to execute the sample task. The task library corresponding to the target domain includes, but is not limited to, structured data such as text or triples, key-value pairs, knowledge graphs, and lists. For example, the task library is a knowledge graph in the field of bioinformatics, which records a complete gene expression data analysis task, which records the task information of the sample task: comparing gene expression differences between different cell types, as well as the unit information of the sample task unit: "NCBI GEO Data Retrieval API" (to obtain research data), "DESeq2 Analysis Tool" (to perform differential expression analysis), and "Cytoscape Plugin" (to visualize the result network), etc.
[0067] A domain task unit is a component unit capable of executing the target task in the target domain. There must be at least one domain task unit, corresponding to each subtask within the target task. A domain task unit can be a software tool, an algorithm, or a solution, such as a gene tagging tool, a sampling algorithm, a gene screening tool, or an enrichment analysis tool.
[0068] A sample task is a domain subtask in the target domain, represented by one or more process steps. The sample task has been executed through the corresponding sample task unit. For example, the sample task is to analyze the expression changes of key genes in various stages of mouse embryonic development.
[0069] The sample task unit is a component unit used to perform a sample task, which may be a software tool, an algorithm, or a solution. For example, the sample task unit is to utilize RNA-seq data.
[0070] Based on the task information, retrieve the domain task units related to the target task from the task library corresponding to the target domain. The specific method is: based on the task information, retrieve the domain task units related to each subtask in the target task from the correspondence between at least one sample task and sample task unit in the task library.
[0071] For example, based on the name of the bioinformatics task (exploring the expression differences of specific gene markers in different cell types between different cells), the corresponding relationships between multiple sample bioinformatics tasks and sample bioinformatics tools in the knowledge graph of the bioinformatics field are retrieved, and the following field bioinformatics tools are retrieved: NCBI GEO data retrieval API, DESeq2 analysis software package and Cytoscape visualization plug-in.
[0072] A comprehensive task library is automatically constructed based on sample tasks and sample task units. Based on the task information of the target task, domain task units related to the target task are retrieved from the comprehensive task library, and appropriate task units are automatically selected, providing task unit support for the subsequent generation of call chain information.
[0073] Step 106: Based on the unit information of the domain task unit, the target generation model is used to generate call chain information that conforms to the target task execution logic.
[0074] The unit information of a domain task unit is the specific descriptive information or operation parameters related to the task unit, including the unit name, function description, parameter indicators, version information, etc. For example, the unit information of a domain task unit is: "NCBI GEO Data Retrieval API" (to obtain research data), "DESeq2 Analysis Tool" (to perform differential expression analysis), and "Cytoscape Plugin" (to visualize the result network).
[0075] The target generation model is a deep learning model with the function of generating call chain information. The target generation model generates call chain information of a specific format and process based on the input information. The target generation model is pre-trained and can understand the execution process of the target task, and translate the unit information of the domain task unit into a specific format according to the execution process. The target generation model includes but is not limited to: Transformer model, BERT model and large model. The target generation model can be a vertical model of the target domain, which is pre-trained with the training data of the target domain (fine-tuning, transfer learning and reinforcement learning, etc.), or it can be a general model of different fields, which generates call chain information based on the prompts of the task information of the target task, which is not limited here.
[0076] The call chain information is a logical chain of process information that records the processes of each subtask in the target task and the domain task units that execute each subtask. It represents the subtask processes and dependencies of this complex task of the target task. The call chain information has a specific format and process. For example, the call chain information is in JSON format, specifically:
[0077]
[0078]
[0079] Based on the unit information of the domain task unit, the target generation model is used to generate call chain information that conforms to the target task execution logic. One optional method is: input the unit information of the domain task unit into the pre-trained target generation model to generate call chain information that conforms to the target task execution logic. One optional method is: input the unit information of the domain task unit into the target generation model, and generate call chain information that conforms to the target task execution logic under the prompt of the task information of the target task. There is no limitation here.
[0080] For example, the unit information of the domain bioinformatics tools (unit information of the task unit: "NCBI GEO Data Retrieval API" (obtaining research data), "DESeq2 Analysis Tool" (performing differential expression analysis) and "Cytoscape Plug-in" (visualizing the result network), etc.) is input into the universal large language model, and under the prompt of the task information of the bioinformatics task, the call chain information in JSON format that conforms to the execution logic of the bioinformatics task is generated.
[0081] By using the deep learning model of the target generation model, accurate call chain information that conforms to the execution logic of the target task is generated, providing process logic support for the subsequent execution of the target task.
[0082] Step 108: Based on the call chain information, call the domain task unit to execute the target task and obtain the task execution result.
[0083] The task execution result is the result obtained by calling the task units in each domain to execute the target task according to the process logic corresponding to the call chain information. The task execution result is a summary and integration of the results obtained during the execution of the target task, and can be presented in various forms such as data sets, reports, charts, and analysis conclusions. For example, in the bioinformatics field, after executing each subtask according to the call chain information, the task execution result includes: a data analysis report: a detailed document summarizing the expression difference analysis process of a specific gene marker between different cell types, the main findings, and the statistical significance test results; data tables: such as "deseq2_results.csv", which record key data such as the expression level, fold change, P value, and adjusted P value (such as using FDR correction) for each gene across different cell types; visualization charts: such as the network diagram named "expression_network.png", which shows the interactions between genes and the expression differences, intuitively presenting the differences in gene expression patterns in different cell types; raw data and intermediate results: this may also include the research data file "expression_data.json" obtained from the NCBI GEO Data Retrieval API and any other intermediate calculation result files.
[0084] Based on the call chain information, the domain task unit is called to execute the target task and obtain the task execution result. The specific method is: based on the call chain information, the domain task units related to each subtask are called to execute the target task and obtain the task execution result.
[0085] For example, based on the call chain information in JSON format, the "NCBI GEO Data Retrieval API" is called to execute the subtask of obtaining research data, the "DESeq2 Analysis Tool" is called to execute the subtask of performing differential expression analysis, and the "Cytoscape Plug-in" is called to execute the subtask of visualizing the result network to obtain the visualization results of the bioinformatics task.
[0086] In the embodiments of this specification, a comprehensive task library corresponding to the target domain is automatically constructed based on sample tasks and sample task units. Based on the task information of the target task, domain task units related to the target task are retrieved from the comprehensive task library, and appropriate task units are automatically selected. The target generation model, a deep learning model, is used to generate accurate call chain information that conforms to the execution logic of the target task. Based on the call chain information, the domain task unit is called to execute the target task, thereby realizing the appropriate calling of the task unit to execute the task and improving the efficiency and accuracy of domain task processing.
[0087] In an optional embodiment of the present specification, the task library corresponding to the target domain includes a task node of at least one sample task in the target domain and a unit node of a sample task unit called to execute the sample task, and there is a corresponding relationship between the task node and the unit node;
[0088] Correspondingly, step 104 includes the following specific steps:
[0089] Based on the task information, a first task node of a first sample task related to the target task is retrieved from a task library corresponding to the target domain;
[0090] Based on the first task node, the first unit node of the first sample task unit called to execute the first sample task is retrieved from the corresponding relationship, and the first sample task unit is determined to be a domain task unit related to the target task.
[0091] In the embodiment of this specification, the task library is a systematic storage structure that records the task nodes of sample tasks and the unit nodes of sample task units called to execute the sample tasks according to the correspondence between tasks and task units.
[0092] The task node of a sample task is a structured data node for the sample task in the task library. The task node includes the task information of the sample task and serves as an index point for the sample task in the task library to provide retrieval of domain task units. This includes, but is not limited to, metadata such as the task name, task requirements, task data information, task execution process, and task execution constraints. For example, in a knowledge graph, a task node is a node that represents a sample task.
[0093] The unit node of a sample task unit is a structured data node of the sample task unit in the task library. The unit node includes the task unit information of the sample task unit. As the search result of searching for domain task units in the task library, it includes but is not limited to metadata such as the task unit name, task unit parameters, task unit type, and task unit call information. For example, in a knowledge graph, a task node is a node that represents a sample task unit.
[0094] The correspondence between task nodes and unit nodes is a mapping between the sample tasks represented in the task library and the sample task units that call these sample tasks. It shows how a specific sample task completes its subtasks by calling the corresponding sample task units. For example, in the knowledge graph of the bioinformatics field, the task node for the sample task "Performing Differential Expression Analysis" is connected to the unit nodes of sample task units such as "edgeR Analysis Tool," "DESeq2 Analysis Tool," and "limma Analysis Tool" through one or more edges. This indicates that when executing this sample task, the task units represented by these three unit nodes can be used to perform the subtasks of data analysis. The corresponding relationship is represented by the edges connecting the nodes.
[0095] The first sample task is a sample task related to the target task, a subtask of the target task, or a task semantically related to the task information of the target task. The first task node of the first sample task is a structured data node of the first sample task in the task library. The first task node includes the task information of the first sample task and serves as an index point for the first sample task in the task library to provide retrieval of domain task units. For example, "performing differential expression analysis" is the first sample task, and the first task node of the first sample task is node 1.
[0096] The first sample task unit is a component unit used to perform the first sample task. The first unit node of the first sample task unit is a structured data node of the first sample task unit in the task library. The first unit node includes the task unit information of the first sample task unit and serves as a search result for the domain task unit in the task library. For example, "edgeR analysis tool", "DESeq2 analysis tool", and "limma analysis tool" are first sample task units, and the first unit node of the first sample task unit is node 1,1 connected to node 1 via an edge (correspondence relationship).
[0097] Based on the first task node, from the corresponding relationship, retrieve the first unit node of the first sample task unit called to execute the first sample task, and determine that the first sample task unit is a domain task unit related to the target task. The specific method is: from the corresponding relationship of the first task node, retrieve the first unit node of the first sample task unit called to execute the first sample task, and determine that the first sample task unit is a domain task unit related to the target task.
[0098] Exemplarily, based on the name of the bioinformatics task (exploring the expression differences of specific gene markers in different cell types between different cells), the first task node of the first sample bioinformatics task related to the bioinformatics task is retrieved from the knowledge graph of the bioinformatics field: node 1, and from the multiple edges of node 1, the first tool node of the first sample bioinformatics tool called to execute the first sample task is retrieved: node 1.1, and the first sample bioinformatics tool is determined to be the field bioinformatics tool related to the target task: NCBI GEO data retrieval API, DESeq2 analysis software package and Cytoscape visualization plug-in.
[0099] In an embodiment of the present specification, based on the task information of the target task, the first task node of the first sample task related to the target task is retrieved from a comprehensive task library, and then based on the first task node, the first unit node of the first sample task unit called to execute the first sample task is retrieved from the corresponding relationship, and the first sample task unit is determined to be a domain task unit related to the target task, and a more suitable task unit is automatically selected, providing more accurate task unit support for the subsequent generation of call chain information.
[0100] In an optional embodiment of the present specification, before retrieving the first task node of the first sample task related to the target task from the task library corresponding to the target domain based on the task information, the following specific steps are further included:
[0101] Acquire at least one sample task in the target domain and a sample task unit called for executing the sample task;
[0102] Constructing a task node of the sample task based on the sample task, constructing a unit node of the sample task unit based on the sample task unit, and determining that there is a corresponding relationship between the task node and the unit node;
[0103] Based on task nodes, unit nodes and corresponding relationships, a task library corresponding to the target field is constructed.
[0104] In the embodiment of this specification, the task library is a systematic storage structure that records the task nodes of sample tasks and the unit nodes of sample task units called to execute the sample tasks according to the correspondence between tasks and task units.
[0105] Obtain at least one sample task in the target field and the sample task unit called to execute the sample task. One optional method is to directly obtain at least one sample task in the target field and the sample task unit called to execute the sample task. Another optional method is to obtain multiple material files in the target field, extract at least one sample task in the target field and the sample task unit called to execute the sample task from the multiple material files, which is not limited here.
[0106] Based on the task nodes, unit nodes and corresponding relationships, a task library corresponding to the target domain is constructed. The specific method is: based on the corresponding relationships, the task nodes and unit nodes are connected to construct a task library corresponding to the target domain.
[0107] Exemplarily, multiple papers in the field of bioinformatics are obtained, sample bioinformatics tasks in the field of bioinformatics and sample bioinformatics tools called to execute the sample bioinformatics tasks are extracted from the multiple papers, task nodes of the sample bioinformatics tasks and unit nodes of the sample bioinformatics tools are constructed based on the sample bioinformatics tasks and the sample bioinformatics tools, and a corresponding relationship between the task nodes and the unit nodes is determined. Based on the corresponding relationship, the task nodes and the unit nodes are connected to construct a knowledge graph corresponding to the field of bioinformatics.
[0108] In the embodiments of this specification, a comprehensive task library is automatically constructed based on task nodes, unit nodes and corresponding relationships, which reduces the workload of domain experts in manually organizing and updating the task library and improves the efficiency of domain task processing.
[0109] In an optional embodiment of the present specification, before constructing a task library corresponding to the target domain based on the task nodes, unit nodes and corresponding relationships, the following specific steps are also included:
[0110] Acquire an initial task library corresponding to the target domain, wherein the initial task library includes at least one initial task node and at least one initial unit node corresponding to the initial task node;
[0111] Based on task nodes, unit nodes and corresponding relationships, a task library corresponding to the target domain is constructed, including:
[0112] Retrieving a first initial unit node similar to the unit node from the initial task library;
[0113] If not retrieved, the task node and unit node are recorded into the initial task library based on the corresponding relationship to obtain the task library corresponding to the target field.
[0114] The task library is a continuously updated task library, and the embodiments of this specification implement automatic updating of the task library.
[0115] The initial task library for the target domain is a systematic storage structure pre-built for the target domain. Based on the correspondence between tasks and task units, it records the initial task nodes of the initial sample tasks and the initial unit nodes of the initial sample task units invoked to execute the initial sample tasks. The initial task library for the target domain includes, but is not limited to, structured data such as text or triples, key-value pairs, knowledge graphs, and lists.
[0116] The initial task node is a structured data node of the initial sample task in the initial task library. The initial task node includes the initial task information of the initial sample task and serves as an index point for the initial sample task in the initial task library to record unit nodes, including but not limited to metadata information such as the initial task name, initial task requirements, initial task data information, initial task execution process, and constraints on the execution of the initial task.
[0117] The initial unit node is a structured data node for the initial sample task unit in the initial task library. The initial unit node includes the initial task unit information of the initial sample task unit. It serves as an index point for the initial sample task in the initial task library to record task nodes and unit nodes, including but not limited to metadata such as the initial task unit name, initial task unit parameters, initial task unit type, and initial task unit call information. The first initial unit node is an initial unit node in the initial task library that is similar to the unit node.
[0118] It should be noted that the task nodes and unit nodes in the initial task library and the task library are recorded according to a hierarchical structure. For example, the task nodes are at the first level and the unit nodes are at the second level. For another example, the task nodes are at the second level and the unit nodes are at the first level. Furthermore, the task nodes of tasks and subtasks are also at different levels.
[0119] Based on the correspondence, the task nodes and unit nodes are recorded in the initial task library to obtain the task library corresponding to the target field. The specific method is: according to the hierarchical structure of the initial task library, the file node to which the first initial unit node belongs is determined; based on the correspondence, the correspondence between the task nodes and unit nodes and the file node to which the first initial unit node belongs is recorded to obtain the task library corresponding to the target field.
[0120] Figure 2 FIG. 1 shows a schematic diagram of a task library construction in a domain task processing method provided by an embodiment of the present specification. Figure 2 As shown:
[0121] The initial task library includes an initial task node_1, an initial unit node_1, and an initial unit node_2. The initial task node_1 has a corresponding relationship with the initial unit node_1 and the initial unit node_2.
[0122] From the initial task library, the first initial unit node similar to unit node_3 is retrieved, but not found. Based on the correspondence between task node_2 and unit node_3, task node_2 and unit node_3 are recorded in the initial task library to obtain the task library corresponding to the target field.
[0123] Exemplarily, an initial knowledge graph corresponding to the field of bioinformatics is obtained, wherein the initial knowledge graph includes at least one initial task node and at least one initial tool node corresponding to the initial task node. A first initial tool node similar to the tool node is retrieved from the initial knowledge graph. If not retrieved, the file node to which the first initial unit node belongs is determined according to the hierarchical structure of the initial task library, and the task node and the tool node are recorded in the initial knowledge graph to obtain a knowledge graph corresponding to the field of bioinformatics.
[0124] In the embodiments of this specification, task nodes and unit nodes not recorded in the initial task library are automatically recorded, which improves the comprehensiveness of the task library, reduces the workload of domain experts in manually organizing and updating the task library, and improves the efficiency of domain task processing.
[0125] In an optional embodiment of the present specification, after retrieving the first initial unit node similar to the unit node from the initial task library, the following specific steps are further included:
[0126] If retrieved, identifying whether the task node has a similarity with the first initial task node corresponding to the first initial unit node and whether the similarity meets a preset threshold;
[0127] If not, the corresponding relationship between the task node and the first initial unit node is recorded to obtain the task library corresponding to the target domain.
[0128] The task library is a continuously updated task library, and the embodiments of this specification implement automatic updating of the task library.
[0129] The first initial unit node is an initial unit node with high similarity to the unit node in the initial task library. The first initial task node is an initial task node in the initial task library that has a corresponding relationship with the first initial unit node.
[0130] It should be noted that the task nodes and unit nodes in the initial task library and the task library are recorded according to a hierarchical structure. For example, the task nodes are at the first level and the unit nodes are at the second level. For another example, the task nodes are at the second level and the unit nodes are at the first level. Furthermore, the task nodes of tasks and subtasks are also at different levels.
[0131] The correspondence between the task node and the first initial unit node is recorded to obtain the task library corresponding to the target domain. Specifically, the correspondence between the task node and the first initial unit node is recorded according to the hierarchical structure of the initial task library to obtain the task library corresponding to the target domain.
[0132] Figure 3 FIG. 1 shows a schematic diagram of a task library construction in another field task processing method provided by an embodiment of the present specification, such as Figure 3 As shown:
[0133] The initial task library includes an initial task node_1, an initial unit node_1, and an initial unit node_2. The initial task node_1 has a corresponding relationship with the initial unit node_1 and the initial unit node_2.
[0134] From the initial task library, the first initial unit node similar to unit node_3 is retrieved as initial unit node_2. After retrieval, the correspondence between task node_2 and the first initial unit node (initial unit node_2, unit node_3) is recorded to obtain the task library corresponding to the target field.
[0135] Exemplarily, an initial knowledge graph corresponding to the field of bioinformatics is obtained, wherein the initial knowledge graph includes at least one initial task node and at least one initial tool node corresponding to the initial task node. From the initial knowledge graph, a first initial tool node similar to the tool node is retrieved, and according to the hierarchical structure of the initial task library, the correspondence between the task node and the first initial unit node is recorded to obtain a knowledge graph corresponding to the field of bioinformatics.
[0136] In the embodiments of this specification, the correspondence between the task node and the first initial unit node is automatically recorded to obtain a task library corresponding to the target field, thereby improving the comprehensiveness of the task library, reducing the workload of domain experts in manually organizing and updating the task library, and improving the efficiency of domain task processing.
[0137] In an optional embodiment of the present specification, obtaining at least one sample task in a target domain and a sample task unit called for executing the sample task includes the following specific steps:
[0138] Acquire multiple material files in the target area;
[0139] Using the information extraction model, extracting at least one sample task in the target domain and a sample task unit called for executing the sample task from the source file;
[0140] Based on task nodes, unit nodes and corresponding relationships, a task library corresponding to the target domain is constructed, including the following specific steps:
[0141] Integrate the task nodes, unit nodes and corresponding relationships of each material file to obtain the task library corresponding to the target field.
[0142] Source files are various document resources in the target domain that contain information about the actual task execution process, task requirements, data sources, and the tools or algorithms used. Source files are the source of sample tasks and sample task units, and serve as sample sources for building a task library. Source files can be research papers, technical reports, case studies, database entries, or other forms of text records. For example, in the field of bioinformatics, a source file is a scientific paper that describes in detail how to use a sample task unit such as a specific gene expression data analysis tool (such as DESeq2) to study gene expression differences between different cell types, and lists information such as experimental design, data sources, and specific operational steps.
[0143] The task node of a source file is a structured data node representing the sample task mentioned in the source file within a systematic storage structure. It contains key information about the sample task, such as the task name, task requirements, data information, execution process, constraints, and other metadata. For example, for a paper studying gene expression differences between neurons and liver cells, the task node in the constructed task library might include the task name "Comparing gene expression differences between different cell types" and related information such as the experimental design, data screening criteria, and expected analysis results.
[0144] The unit node of the material file is a structured data node in the systematic storage structure that represents the specific tools, algorithms, or solutions called in the material file to execute the sample task. It contains the core attributes of the task unit, such as the task unit name, functional description, parameter indicators, version information, etc. For example, taking the bioinformatics paper as an example, if the paper mentions the use of the "NCBI GEO Data Retrieval API" to obtain research data, the "DESeq2 Analysis Tool" for differential expression analysis, and the "Cytoscape Plug-in" for visualization, then in the constructed knowledge graph, there are three corresponding unit nodes, each of which records the basic information of the corresponding tool or algorithm and its role in the task execution process.
[0145] The corresponding relationship between each material file is the association between the task node and the unit node in the material file represented in the systematic storage structure. It characterizes how a task completes subtasks by calling different task units. The connection between task nodes and unit nodes is usually represented by edges, key-value pairs, or hierarchical relationships. For example, in the above-mentioned bioinformatics example, the task node "Comparing gene expression differences between different cell types" in the paper may be connected to the unit nodes of "NCBI GEO Data Retrieval API", "DESeq2 Analysis Tool" and "Cytoscape Plug-in" through one or more edges, which clearly indicates which tool units are needed to perform this task and their position and role in the task process.
[0146] To obtain multiple material files in the target field, one optional method is to directly manually generate multiple material files in the target field. Another optional method is to use a generation model to generate multiple material files in the target field based on the domain information of the target field. Another optional method is to obtain multiple material files in the target field from an open source database. There is no limitation here.
[0147] An information extraction model is used to extract at least one sample task in the target domain and a sample task unit invoked by executing the sample task from a source file. An optional method is to use the information extraction model to extract at least one sample task in the target domain and a sample task unit invoked by executing the sample task from a source file based on a preset extraction prompt. An extraction prompt is information used to prompt the model to execute a task and a task unit, such as a prompt.
[0148] Integrate the task nodes, unit nodes, and corresponding relationships of each material file to obtain a task library corresponding to the target domain. The task nodes, unit nodes, and corresponding relationships of each material file can be regarded as a node tree. By integrating the task nodes, unit nodes, and corresponding relationships of each material file, a larger node tree is constructed to obtain a comprehensive task library.
[0149] For example, multiple papers in the field of bioinformatics are obtained from open source databases. Based on the preset extraction prompt "from the materials and methods, results or discussion sections of the papers, locate and extract all bioinformatics tools used in the experiment and their version numbers, such as software tools such as Cytoscape", the information extraction model is used to extract multiple sample bioinformatics tasks and sample bioinformatics tools in the field of bioinformatics from multiple papers, integrate the task nodes, tool nodes and corresponding relationships of each paper, and obtain the corresponding knowledge graph in the field of bioinformatics.
[0150] In the embodiments of this specification, a more comprehensive task library is constructed by extracting and integrating the task nodes, unit nodes and corresponding relationships of each material file, providing data support for subsequent retrieval of domain task units related to the target task.
[0151] In an optional embodiment of this specification, the task execution result includes alarm information;
[0152] After step 108, the following specific steps are also included:
[0153] Feedback alarm information to front-end users;
[0154] Receive updated call chain information sent by front-end users based on alarm information;
[0155] Based on the updated call chain information, the domain task unit is called to execute the target task and obtain the updated task execution result.
[0156] Alarm information is a warning notification generated by the system when one or more subtasks fail to complete successfully due to errors, abnormal conditions or other unforeseen problems during the execution of the target task. The alarm information usually includes the error type, error description and possible cause analysis. Alarm information is an important feedback mechanism in the task execution process. It can promptly remind users or backend systems of potential problems and require further investigation or appropriate measures to correct the errors. For example. In the gene expression difference analysis task in the field of bioinformatics, if the "DESeq2 analysis tool" cannot perform effective analysis due to data quality issues, an alarm message is generated, such as: "An error occurred during the data analysis phase: DESeq2 analysis failed."
[0157] Updating the call chain information is to regenerate new call logic and sequence information for executing the target domain task based on the intervention of the front-end user or the automatic correction strategy after receiving the alarm information. When an alarm message appears during the execution of a task guided by the call chain information, the user will replace the domain task unit provided by the alarm message, and then generate new call chain information to execute the target task and continue to execute the task. For example, the call chain information uses a specific version of the RNA-seq data processing tool, but during execution it is found that the tool does not support the current data format enough, resulting in an alarm. The user changes the tool version or selects a more suitable tool on the front end. Based on this change, an updated call chain information will be generated, the original tool will be replaced with the latest version of the compatible tool, and the execution order of subsequent task units will be re-planned.
[0158] The update task execution result is the update task result obtained by calling each update domain task unit to execute the target task according to the process logic corresponding to the update call chain information.
[0159] For example, when executing a gene expression differential analysis task in the bioinformatics field, the "DESeq2 analysis tool" was invoked based on pre-generated call chain information. The analysis failed, generating an alarm message: "Error in data analysis: DESeq2 analysis failed due to incompliance with DESeq2 tool requirements." This alarm message was quickly fed back to the front-end user interface, allowing the user to promptly identify and address the issue. After receiving the alarm message, the front-end user identified the problem and adjusted the call chain information accordingly, such as upgrading the existing DESeq2 version or switching to a different differential expression analysis tool that supports the new data format, such as "edgeR." The user updated the call chain information through the front-end interface. The new call chain information contained the analysis tool and its parameter configuration optimized for the current data format. After receiving the updated call chain information sent by the front-end user based on the alarm message, the execution process was reorganized according to the new call logic and sequence. Following the updated call chain information, the system first ensured that the data had been converted to the format required by the newly selected tool, then invoked the "edgeR analysis tool" to execute the differential expression analysis subtask. Next, the system continued to follow the subsequent steps in the updated call chain to complete the execution of all subtasks.
[0160] In the embodiments of this specification, through front-end interaction, the alarm information is flexibly responded to and the task execution plan is dynamically adjusted according to user intervention, thereby effectively solving the problems encountered during the task execution process and ultimately ensuring the high-quality achievement of the task goals.
[0161] In an optional embodiment of this specification, the task execution result includes alarm information;
[0162] After step 108, the following specific steps are also included:
[0163] Based on the unit information and alarm information of the domain task unit, the target generation model is used to generate the update call chain information that conforms to the target task execution logic;
[0164] Based on the updated call chain information, the domain task unit is called to execute the target task and obtain the updated task execution result.
[0165] Alarm information is a warning notification generated by the system when one or more subtasks fail to complete successfully due to errors, abnormal conditions or other unforeseen problems during the execution of the target task. The alarm information usually includes the error type, error description and possible cause analysis. Alarm information is an important feedback mechanism in the task execution process. It can promptly remind users or backend systems of potential problems and require further investigation or appropriate measures to correct the errors. For example. In the gene expression difference analysis task in the field of bioinformatics, if the "DESeq2 analysis tool" cannot perform effective analysis due to data quality issues, an alarm message is generated, such as: "An error occurred during the data analysis phase: DESeq2 analysis failed."
[0166] The target generation model is a pre-trained deep learning model that can generate corresponding call chain information based on the input unit information and alarm information, ensuring that the task units in each field execute specific target tasks in the correct logical order and configuration. The target generation model can understand and translate the functional descriptions and parameter indicators of task units in different fields, and generate a call chain structure that meets the requirements of the target task execution process, such as models such as Transformer, BERT, and even vertical models for specific fields. In this scenario, the target generation model can generate an updated call chain information containing the correct call order and parameter settings based on the unit information and alarm information "Error in the data analysis phase: DESeq2 analysis failed" of the "NCBI GEO Data Retrieval API", "DESeq2 Analysis Tool" or "edgeR Analysis Tool".
[0167] Updating the call chain information is the new call logic and sequence information that the target generation model regenerates for executing the target domain task after receiving the alarm information. When an alarm message appears during the execution of a task guided by the call chain information, the target generation model will generate updated call chain information to execute the target task and continue to execute the task based on the error type, error description, and possible cause analysis provided by the unit information and the alarm message. For example, the call chain information uses a specific version of an RNA-seq data processing tool, but during execution, it is found that the tool does not support the current data format sufficiently, resulting in an alarm. The target generation model generates an updated call chain information based on the unit information and the alarm message, replaces the original tool with the latest version of the compatible tool, and replans the execution order of subsequent task units.
[0168] The update task execution result is the update task result obtained by calling each update domain task unit to execute the target task according to the process logic corresponding to the update call chain information.
[0169] For example, when performing a gene expression differential analysis task in the bioinformatics field, the "DESeq2 Analysis Tool" is called based on pre-generated call chain information. The analysis fails, and an alarm message is generated: "An error occurred during the data analysis phase: DESeq2 analysis failed because the DESeq2 tool requirements are not met." Based on the unit information and the alarm message, the target generation model generates updated call chain information, such as upgrading the original DESeq2 version or switching to another differential expression analysis tool that supports the new data format, such as "edgeR." The updated call chain information contains the analysis tool and its parameter configuration optimized for the current data format. After receiving the updated call chain information sent by the front-end user based on the alarm message, the execution process is reorganized according to the new call logic and sequence. According to the updated call chain information, the system first ensures that the data has been converted to the format required by the newly selected tool, and then calls the "edgeR Analysis Tool" to execute the differential expression analysis subtask. Next, it continues to follow the subsequent steps in the updated call chain to complete the execution of all subtasks.
[0170] In the embodiments of this specification, by means of model generation, the alarm information is flexibly responded to and the task execution plan is dynamically adjusted according to user intervention, thereby effectively solving the problems encountered during the task execution process and ultimately ensuring the high-quality achievement of the task objectives.
[0171] In an optional embodiment of the present specification, after step 108, the following specific steps are further included:
[0172] Based on the task execution results, a task analysis report is generated using the result analysis model.
[0173] The result analysis model is a pre-trained deep learning model used to automatically analyze, interpret, and summarize task execution results, generating easy-to-understand and easy-to-present task completion reports. During domain task processing, the model can extract key information based on the raw data, intermediate calculation results, and final statistical conclusions generated during the task execution phase, and generate structured report content based on preset templates or dynamically. It can understand task objectives and execution processes, converting complex data into visual charts, conclusion summaries, etc., helping users quickly grasp the overall effect and important findings of task execution. For example, in the bioinformatics gene expression difference analysis task, the result analysis model can automatically generate analysis charts such as heat maps, volcano plots, and statistical test indicators containing gene expression differences between different cell types based on the statistical significance data in the task execution result "DESeq2_results.csv". It can also use these charts to write a detailed text report explaining the analysis process and main conclusions.
[0174] A task analysis report is a comprehensive report document generated by the system based on the results of the task execution after the system completes the target task in a specific field. It is used to comprehensively present the process, methods, results and conclusions of the task execution. A task analysis report usually includes multiple parts such as an introduction to the task background, description of the experimental design and methods, data analysis steps, presentation of key results, statistical tests and interpretations, conclusions and suggestions. It aims to provide a clear and easy-to-understand overview of the results for users to review, review and cite. In addition, the report may also include suggestions for future research directions or practical applications. Taking the gene expression difference analysis task as an example, the task analysis report:
[0175] Introduction: briefly describe the purpose and background of the study;
[0176] Methods: Detailed description of the tools used (e.g., NCBI GEO data retrieval API, DESeq2 analysis tool), data sources, experimental design, and statistical methods;
[0177] Results: Displays the main contents of the data analysis report, such as differential expression data of key genes in tabular form, P values and adjusted P values, and various visual charts (such as expression network diagrams);
[0178] Discussion: Explain the significance of data analysis results and explore the functional relevance and biological significance of differentially expressed genes;
[0179] Conclusion: Summarize the main findings and potential application value of this analysis;
[0180] References: List all data sources and tool references used.
[0181] Based on the task execution results, a task analysis report is generated using the result analysis model. One optional method is: based on the task execution results, under the prompt of report generation and extraction information, a task analysis report is generated using the result analysis model. Another optional method is: based on the task execution results, a task analysis report is generated using a pre-trained result analysis model, wherein the result analysis model is trained based on the sample execution results of the sample task and the labeled task analysis report, which is not limited here.
[0182] For example, in the field of bioinformatics, when a gene expression differential analysis task is completed, the system will collect a series of task execution results, including detailed statistical reports, data tables, and visual charts. For example, the original data analysis report shows the expression changes of specific genes between different cell types and the results of statistical significance tests; the "deseq2_results.csv" file contains the expression levels, fold changes, P values, and corrected P values of all tested genes (such as using the Benjamini-Hochberg method for FDR correction). Based on these rich task execution results, the system will call a pre-trained result analysis model to deeply interpret and integrate the data. The model can automatically identify key findings and generate a structured task analysis report based on a preset template, including: Abstract: A brief overview of the experimental design purpose, methods, and main findings. Materials and methods: A detailed description of the sample source, experimental process, and the specific steps of using the NCBI GEO data retrieval API to obtain the data set and use the DESeq2 tool for differential expression analysis. Results Presentation: Core results will be presented graphically and visually, such as a list of significantly differentially expressed genes, accompanied by heatmaps or volcano plots to illustrate overall expression trends and significance distributions. Gene co-expression network images generated using the Cytoscape plugin will also be provided to visually demonstrate gene interactions. Statistical Analysis and Discussion: The significance of statistical test indicators will be explained, and the functional annotations of significantly differentially expressed genes and the potential biological mechanisms underlying their differential expression across different cell types will be discussed. Conclusion: The key findings of this analysis will be summarized, and their value and significance for scientific research or clinical application will be discussed.
[0183] In the embodiments of this specification, the task analysis report generated by the result analysis model based on the task execution results is both comprehensive and professional, which greatly enhances the richness of domain task execution and improves user experience.
[0184] In an optional embodiment of the present specification, before step 108, the following specific steps are further included:
[0185] Based on the call chain information, generate the call process diagram corresponding to the target task;
[0186] Feedback the call process diagram to the front-end user;
[0187] Receive updated call chain information sent by front-end users for the call process diagram.
[0188] The call process diagram is a visualization of call chain information, used to intuitively present the multiple subtasks required to complete the target task and their execution process, including the domain task unit corresponding to each subtask and its position in the execution process, the preceding and following relationships, and possible conditional branches or loop structures. The call process diagram is constructed through nodes (representing each subtask) and lines (representing the call sequence or dependency relationships), which can clearly reflect the dynamic connection between all subtask units involved in the entire task execution process from start to finish. For example, in the field of bioinformatics, in the analysis of gene expression differences, the call process diagram can show a series of steps: first calling the NCBI GEO data retrieval API to obtain data, then using DESeq2 for differential analysis, and finally using the Cytoscape plug-in to generate a visual network diagram.
[0189] After receiving an alarm, the front-end user regenerates the call chain information based on the call process diagram to create new call logic and sequence information for executing the target domain task. The user will generate new call chain information based on the replacement domain task units provided by the call process diagram to execute the target task and continue task execution.
[0190] For example, the call chain information includes: using the NCBI GEO Data Retrieval API to retrieve publicly available research data containing target genes and corresponding cell types. The DESeq2 analysis tool was then used to process the downloaded data to determine the differential expression of each gene across different cell types. A gene co-expression network diagram was constructed and displayed using the Cytoscape plugin. Based on this call chain information, the system automatically generated a call process diagram, which depicts the entire task flow through a series of nodes (representing subtasks) and lines (indicating the call sequence). Upon receiving this call process diagram, the front-end user can clearly see the domain task units used at each stage and their interrelationships. The front-end user discovered quality issues with a dataset for a certain cell type, resulting in an alert. Based on this feedback, the user decided to switch data sources to the ArrayExpress Database API to obtain relevant research data. The subsequent analysis process was adjusted accordingly, adding a preprocessing step before using DESeq2, using the R package limma for preliminary data cleaning and normalization. The front-end user modified the call process diagram based on the updated logic and sequence and submitted the new call chain information to the server through the user interface. After receiving the updated call chain information, the server reorganized the execution process and, in accordance with the new order and task units specified by the user, sequentially called the "ArrayExpress Database API" to obtain data, "limma" for preprocessing, the "DESeq2 Analysis Tool" to calculate differential expression, and the "Cytoscape Plug-in" for visualization, ultimately completing and optimizing the execution process of the target task.
[0191] In the embodiments of this specification, by generating and sending a call process diagram to the front-end user, the front-end interacts with the user, flexibly responds to alarm information and dynamically adjusts the task execution plan based on user intervention, thereby effectively solving problems encountered during task execution and ultimately ensuring the high-quality achievement of task goals.
[0192] In an optional embodiment of the present specification, after step 106, the following specific steps are further included:
[0193] Based on the call chain information, a call process diagram that conforms to the target task execution logic is generated;
[0194] Feedback the call process diagram to the front-end user;
[0195] Receiving update sample data sent by a front-end user for a call process diagram, wherein the update sample data includes update unit information of an update domain task unit of a sample task and label call chain information;
[0196] The target generation model is trained based on the update unit information and label call chain information.
[0197] Updated sample data is new sample data generated by front-end users adjusting the domain task units or their unit information used during target task execution based on the call process diagram. This updated sample data includes the specific content of the domain task units, as well as their new label call chain information, updated based on actual needs or in response to alerts. For example, in a gene expression differential analysis task in the bioinformatics field, the original call chain information indicated that data should first be acquired through the NCBI GEO Data Retrieval API and then differential analysis should be performed. However, after receiving the call process diagram, the front-end user discovered that the data quality of a certain cell type did not meet their requirements. Therefore, they decided to switch to the ArrayExpress database API and added the R package limma as a preprocessing step. The user updated these changes in the interface and submitted them to the server. This user-adjusted unit information, including the new domain task units "ArrayExpress database API" and "R package limma," and the corresponding new call chain information, constitutes the updated sample data. The server retrains or fine-tunes the target generation model based on this updated sample data to adapt to the new task execution requirements and logic specified by the user.
[0198] Based on the update unit information and label call chain information, the target generation model is trained. The specific method is as follows: based on the update unit information and label call chain information, the target generation model is supervised trained, wherein the supervised training includes forward propagation and backward propagation. In the forward propagation, the predicted call chain information is generated, and the loss value is determined based on the predicted call chain information and the label call chain information. In the backward propagation, based on the loss value, the model parameters of the target generation model are adjusted through the gradient update method to complete the supervised training of the target generation model.
[0199] For example, the user requested a task to perform differential gene expression analysis in the bioinformatics field. Based on the task information, the system generated call chain information, which included the following: "First, obtain data through the NCBI GEO Data Retrieval API, then perform differential analysis using the DESeq2 tool, and finally create a visualization chart using the Cytoscape plugin." Based on this call chain information, a call process diagram corresponding to the differential gene expression analysis was generated, where each node represents a subtask (i.e., a domain task unit), and the lines indicate the call order and dependencies. The call process diagram was displayed to the front-end user, allowing them to view the logical flow of the entire task execution through an intuitive interface. While reviewing the call process diagram, the user discovered that the data quality of the NCBI GEO Data Retrieval API used in the original plan was suboptimal for a specific cell type. Therefore, they decided to replace it with the ArrayExpress database API and add the "R package limma" as a preprocessing step to improve data quality. The user made the corresponding adjustments in the interface and submitted this updated information, which constituted the updated sample data, including the unit information for the new domain task units "ArrayExpress Database API" and "R package limma," as well as the corresponding new label call chain information. Based on the updated sample data provided by the user, the server uses the new domain's task units and their call logic as new training instances to retrain or fine-tune the target generation model. During supervised training, the model performs forward propagation, predicting new call chain information based on the updated unit information; then, it calculates the loss between the predicted call chain information and the actual labeled call chain information. During the backward propagation phase, the model adjusts its parameters based on the loss value using optimization algorithms such as gradient descent, enabling it to better learn and understand the new task execution logic and component selection method proposed by the user.
[0200] In the embodiments of this specification, by calling the process diagram, the front-end user is guided to update the sample data in an interactive manner to train the target generation model, thereby improving its accuracy and adaptability when generating call chain information in the future.
[0201] Figure 4FIG. 1 shows a flow chart of a task library construction process of a domain task processing method provided by an embodiment of this specification, as shown in FIG. Figure 4 As shown:
[0202] Acquire multiple material files in the target domain. Utilize the information extraction model to extract at least one sample task in the target domain and the sample task unit called to execute the sample task from the material file. Acquire the initial task library corresponding to the target domain. From the initial task library, retrieve the first initial unit node that is similar to the unit node. If not retrieved, determine the file node to which the task node belongs from top to bottom, and add the task node and the unit node to the initial task library. If retrieved, retrieve the first initial task node corresponding to the first initial unit node from bottom to top, identify whether the similarity between the task node and the first initial task node corresponding to the first initial unit node meets the preset threshold, and if so, record the correspondence between the task node and the first initial unit node. Then construct the file nodes of each material file.
[0203] Figure 5 FIG. 1 shows a flow chart of a method for processing a domain task provided by an embodiment of the present specification, such as Figure 5 As shown:
[0204] Define the target task of the target domain. Retrieve the task library corresponding to the target domain. Retrieve the domain task units related to the target task from the task library. Organize the unit information of the domain task unit as the model input. Use the target generation model to generate call chain information that conforms to the target task execution logic. Based on the call chain information, call the domain task unit to execute the target task and obtain the task execution result. Is there any alarm information? If not, based on the task execution result, use the result analysis model to generate a task analysis report. If so, feed back the alarm information to the front-end user, receive the updated call chain information sent by the front-end user based on the alarm information, or, based on the unit information and alarm information of the domain task unit, use the target generation model to generate the updated call chain information that conforms to the target task execution logic.
[0205] See also Figure 6 , Figure 6 A flowchart of a domain task processing method provided by an embodiment of this specification is shown, including the following specific steps:
[0206] Step 602: Obtain task information of a bioinformatics task in the bioinformatics field.
[0207] Step 604: Based on the task information, retrieve domain task tools related to the bioinformatics task from a task library corresponding to the bioinformatics field, wherein the task library is constructed based on at least one sample task in the bioinformatics field and a sample task tool called to execute the sample task.
[0208] Step 606: Based on the tool information of the domain task tool, the target generation model is used to generate tool call chain information that conforms to the bioinformatics task execution logic.
[0209] Step 608: Based on the tool call chain information, call the domain task tool to execute the bioinformatics task and obtain the task execution result.
[0210] The embodiments of this specification are applied to applications, websites, or platforms with task processing functions in the field of bioinformatics, for example, an application or website that deploys a large model, or an application or website that calls the large model through an application programming interface (Application Programming Interface, referred to as API) to perform bioinformatics tasks by calling domain task units in the field of bioinformatics.
[0211] The embodiments of this specification are similar to the above Figure 1 The embodiments of the specification are based on the same inventive concept. For the specific contents of steps 602-608, please refer to the above embodiments of the specification and will not be repeated here.
[0212] In the embodiments of this specification, a comprehensive task library corresponding to the bioinformatics field is automatically constructed based on sample tasks and sample task tools. Based on the task information of the bioinformatics task, task tools in the bioinformatics field related to the bioinformatics task are retrieved from the comprehensive task library, and a suitable task tool is automatically selected. Using the target generation model, a deep learning model, accurate call chain information that conforms to the bioinformatics task execution logic is generated. Based on the call chain information, the task tool is called to execute the bioinformatics task, thereby achieving appropriate calling of the task tool to execute the task and improving the efficiency and accuracy of task processing in the bioinformatics field.
[0213] The following combined Figure 7 , taking the application of the domain task processing method provided in this specification in the field of bioinformatics as an example, the domain task processing method is further explained. Figure 7 A flowchart of a process for processing a domain task in the field of bioinformatics provided by one embodiment of this specification is shown, including the following specific steps:
[0214] Step 702: Obtain multiple papers in the field of bioinformatics.
[0215] Step 704: Using the large language model, extract at least one sample task in the field of bioinformatics and a sample tool called to execute the sample task from the paper.
[0216] Step 706: Based on the sample task and the sample tool, construct a task node of the sample task and a tool node of the sample tool, and determine whether there is a corresponding relationship between the task node and the tool node.
[0217] Step 708: Obtain an initial knowledge graph corresponding to the bioinformatics field, wherein the initial knowledge graph includes at least one initial task node and at least one initial tool node corresponding to the initial task node.
[0218] Step 710: Retrieve a first initial tool node similar to the tool node from the initial knowledge graph.
[0219] Step 712: If not retrieved, based on the corresponding relationship, record the task node and tool node into the initial knowledge graph.
[0220] Step 714: If retrieved, identify whether the task node and the first initial task node corresponding to the first initial tool node have a similarity that meets a preset threshold; if not, record the corresponding relationship between the task node and the first initial tool node.
[0221] Step 716: Integrate the task nodes, tool nodes, and corresponding relationships of each paper to obtain the knowledge graph corresponding to the bioinformatics field.
[0222] Step 718: Receive a task request in the bioinformatics field sent by a front-end user, wherein the task request includes task information of the bioinformatics task.
[0223] For example, the task information is "perform functional annotation on the genome sequence of a newly cultivated plant."
[0224] Step 720: Based on the task information, retrieve domain tools related to the bioinformatics task from the knowledge graph corresponding to the bioinformatics field.
[0225] Step 722: Based on the tool information of the domain tool, use the large language model to generate call chain information that conforms to the execution logic of the bioinformatics task.
[0226] Step 724: Based on the call chain information, call the domain tool to execute the bioinformatics task.
[0227] Step 726: If the execution fails, obtain alarm information. Based on the tool information and alarm information of the domain tool, use the large language model to update the call chain information that conforms to the bioinformatics task execution logic. Based on the updated call chain information, call the domain tool to execute the bioinformatics task and obtain the updated task execution result.
[0228] Step 728: Based on the updated task execution results, a task analysis report is generated using the large language model.
[0229] For example, the task analysis report is as follows:
[0230] 1. Project Background and Objectives
[0231] This study functionally annotated the whole-genome sequences of newly cultivated plant varieties, which were obtained and assembled using high-throughput sequencing technology. The primary goal was to reveal their genetic information and potential biological functions to guide breeding improvement, molecular marker development, and subsequent physiological and biochemical function verification.
[0232] 2. Raw Data Processing and Quality Assessment
[0233] Sequence Acquisition and Assembly: High-quality genome sequences were successfully obtained through second- or third-generation sequencing of the samples, and a chromosome-level draft genome was generated using advanced assembly algorithms. Repeat Sequence Identification and Removal: Repeat sequence analysis and filtering were performed on the assembly results using tools such as RepeatMasker to reduce interference factors during the annotation process. Gene Prediction and Structure Analysis: Based on known plant species homology information, gene structure characteristics, and expression data, gene structure prediction was performed using software such as Augustus and GlimmerHMM.
[0234] 3. Key results of functional annotation
[0235] Protein-coding gene annotation
[0236] Gene Location and Structure Description: A total of X protein-coding genes were predicted, distributed across Y chromosomes. The exon and intron distribution and complete CDS regions of each gene are detailed. Protein Functional Domain Prediction: Using tools such as InterProScan, all predicted protein amino acid sequences were searched for functional domains, identifying the biochemical reactions or cellular processes in which they may participate.
[0237] Homology analysis
[0238] By comparing with public databases such as NR and SwissProt, we found that a large number of genes have highly homologous known proteins, based on which we speculate on the possible functions of these new genes in metabolic pathways, signal transduction, etc.
[0239] Transcription factor binding site prediction
[0240] Using JASPAR or other TFBS prediction tools, multiple possible transcription factor binding sites were identified in the gene promoter region, which are important for regulating gene expression patterns.
[0241] 4. Highlights and Discussions
[0242] New Gene Discovery: The report highlights several novel genes not previously reported in existing databases and their potential functional associations. Key Regulatory Networks: Based on transcription factor binding site analysis, a transcriptional regulatory network model was constructed that may influence the development of key traits. Evolutionary Relationship Exploration: Phylogenetic tree construction and gene family analysis revealed evolutionary relationships with closely related species and unique gene clusters.
[0243] In the embodiments of this specification, the following are achieved: 1. Automated construction of domain-specific knowledge graphs: By deeply analyzing bioinformatics literature, especially methodology chapters, it is possible to automatically identify and extract key tasks and analysis tools used in the field of bioinformatics. This method of automatically constructing and updating knowledge graphs can not only reflect the latest technologies and methods in the field of bioinformatics, but also reduce the workload of field experts in manually organizing and updating knowledge bases. 2. Intelligent bioinformatics tool chain recommendation system: Using the constructed knowledge graph, combined with advanced retrieval technology, it is possible to recommend the most appropriate analysis tools and call chains for specific bioinformatics problems. This greatly improves the efficiency of research, especially for new researchers, who can quickly obtain expert-level advice in the field. 3. Application of retrieval-enhanced language models based on bioinformatics domain knowledge: The large language model generated based on retrieval enhancement can combine the specific knowledge and context of the bioinformatics field to generate customized call chains for specific bioinformatics tasks. This method not only improves the accuracy and adaptability of the analysis process, but also makes the automatically generated analysis process more in line with the actual needs of the bioinformatics field. 4. Real-time Error Feedback and Adaptive Optimization: During actual data analysis, tool execution errors are captured in real time and intelligently corrected or suggested by combining domain knowledge graphs and large language models. This real-time error handling and adaptive optimization mechanism significantly improves the robustness of the analysis process and is particularly important for high-throughput data analysis.
[0244] Figure 8 FIG. 1 shows a front-end schematic diagram of a domain task processing method applied to the field of bioinformatics provided by an embodiment of this specification. Figure 8 As shown:
[0245] The front-end interface includes a content display area, an input box, a send control, and an export control. The user enters the target task information ("Functionally annotate a newly cultivated plant genome sequence") in the input box and clicks the send control. This initiates steps 718-728, until the task analysis report is generated and displayed in the display area. The user can then click the export control to export the task analysis report in a file format such as .doc or .pdf.
[0246] Corresponding to the above method embodiment, this specification also provides an embodiment of a domain task processing device, Figure 9FIG1 shows a schematic diagram of the structure of a domain task processing device provided by an embodiment of this specification. Figure 9 As shown, the device includes:
[0247] A first acquisition module 902 is configured to acquire task information of a target task in a target domain;
[0248] A first retrieval module 904 is configured to retrieve domain task units related to the target task from a task library corresponding to the target domain based on the task information, wherein the task library is constructed based on at least one sample task in the target domain and sample task units called by executing the sample task;
[0249] The first generation module 906 is configured to generate call chain information that conforms to the target task execution logic based on the unit information of the domain task unit and using the target generation model;
[0250] The first execution module 908 is configured to call the domain task unit to execute the target task based on the call chain information and obtain the task execution result.
[0251] Optionally, the task library corresponding to the target domain includes a task node of at least one sample task in the target domain and a unit node of a sample task unit called to execute the sample task, and there is a corresponding relationship between the task node and the unit node;
[0252] Correspondingly, the first retrieval module 904 is further configured to:
[0253] Based on the task information, the first task node of the first sample task related to the target task is retrieved from the task library corresponding to the target domain; based on the first task node, the first unit node of the first sample task unit called to execute the first sample task is retrieved from the corresponding relationship, and the first sample task unit is determined to be the domain task unit related to the target task.
[0254] Optionally, the device further comprises:
[0255] The first construction module is configured to obtain at least one sample task in the target domain and a sample task unit called to execute the sample task; construct a task node of the sample task based on the sample task, construct a unit node of the sample task unit based on the sample task unit, and determine that there is a corresponding relationship between the task node and the unit node; based on the task node, the unit node and the corresponding relationship, construct a task library corresponding to the target domain.
[0256] Optionally, the device further comprises:
[0257] An initial task library acquisition module is configured to acquire an initial task library corresponding to a target domain, wherein the initial task library includes at least one initial task node and at least one initial unit node corresponding to the initial task node;
[0258] Correspondingly, the first building block is further configured as follows:
[0259] From the initial task library, retrieve the first initial unit node similar to the unit node; if not retrieved, based on the correspondence, record the task node and the unit node into the initial task library to obtain the task library corresponding to the target field.
[0260] Optionally, the device further comprises:
[0261] The recording module is configured to, if retrieved, identify whether the similarity between the task node and the first initial task node corresponding to the first initial unit node meets a preset threshold; if not, record the correspondence between the task node and the first initial unit node to obtain a task library corresponding to the target field.
[0262] Optionally, the first building block is further configured to:
[0263] Acquire multiple material files in the target domain; use the information extraction model to extract at least one sample task in the target domain and the sample task unit called to execute the sample task from the material files; integrate the task nodes, unit nodes and corresponding relationships of each material file to obtain a task library corresponding to the target domain.
[0264] Optionally, the task execution result includes alarm information;
[0265] Correspondingly, the device further includes:
[0266] The first update execution module is configured to feed back the alarm information to the front-end user; receive the update call chain information sent by the front-end user based on the alarm information; based on the update call chain information, call the domain task unit to execute the target task and obtain the update task execution result.
[0267] Optionally, the task execution result includes alarm information;
[0268] The second update execution module is configured to generate update call chain information that conforms to the target task execution logic based on the unit information and alarm information of the domain task unit using the target generation model; based on the update call chain information, call the domain task unit to execute the target task and obtain the update task execution result.
[0269] Optionally, the device further comprises:
[0270] The report generation module is configured to generate a task analysis report based on the task execution result and using the result analysis model.
[0271] Optionally, the device further comprises:
[0272] The update generation module is configured to generate a call process diagram corresponding to the target task based on the call chain information; feed back the call process diagram to the front-end user; and receive the updated call chain information sent by the front-end user for the call process diagram.
[0273] Optionally, the device further comprises:
[0274] The model training module is configured to generate a call process diagram that conforms to the target task execution logic based on the call chain information; receive updated sample data sent by the front-end user for the call process diagram, wherein the updated sample data includes the updated unit information and label call chain information of the updated domain task unit of the sample task; and train the target generation model based on the updated unit information and the label call chain information.
[0275] In the embodiments of this specification, a comprehensive task library corresponding to the target domain is automatically constructed based on sample tasks and sample task units. Based on the task information of the target task, domain task units related to the target task are retrieved from the comprehensive task library, and appropriate task units are automatically selected. The target generation model, a deep learning model, is used to generate accurate call chain information that conforms to the execution logic of the target task. Based on the call chain information, the domain task unit is called to execute the target task, thereby realizing the appropriate calling of the task unit to execute the task and improving the efficiency and accuracy of domain task processing.
[0276] The above is a schematic diagram of a domain task processing device according to this embodiment. It should be noted that the technical solution of this domain task processing device and the technical solution of the aforementioned domain task processing method are based on the same concept. For details not described in detail in the technical solution of the domain task processing device, please refer to the description of the technical solution of the aforementioned domain task processing method.
[0277] Corresponding to the above method embodiment, this specification also provides an embodiment of a task processing device in the field of bioinformatics, Figure 10 FIG. 1 shows a schematic diagram of a task processing device in the field of bioinformatics provided by an embodiment of this specification. Figure 10 As shown, the device includes:
[0278] The second acquisition module 1002 is configured to acquire task information of a bioinformatics task in the field of bioinformatics;
[0279] The second retrieval module 1004 is configured to retrieve, based on the task information, domain task tools related to the bioinformatics task from a task library corresponding to the bioinformatics field, wherein the task library is constructed based on at least one sample task in the bioinformatics field and a sample task tool called to execute the sample task;
[0280] The second generation module 1006 is configured to generate tool call chain information that conforms to the bioinformatics task execution logic based on the tool information of the domain task tool using the target generation model;
[0281] The second execution module 1008 is configured to call the domain task tool to execute the bioinformatics task based on the tool call chain information and obtain the task execution result.
[0282] In the embodiments of this specification, a comprehensive task library corresponding to the bioinformatics field is automatically constructed based on sample tasks and sample task tools. Based on the task information of the bioinformatics task, task tools in the bioinformatics field related to the bioinformatics task are retrieved from the comprehensive task library, and a suitable task tool is automatically selected. Using the target generation model, a deep learning model, accurate call chain information that conforms to the bioinformatics task execution logic is generated. Based on the call chain information, the task tool is called to execute the bioinformatics task, thereby achieving appropriate calling of the task tool to execute the task and improving the efficiency and accuracy of task processing in the bioinformatics field.
[0283] The above is a schematic diagram of a task processing device in the field of bioinformatics according to this embodiment. It should be noted that the technical solution of this task processing device in the field of bioinformatics and the technical solution of the task processing method in the field of bioinformatics described above share the same concept. For details not described in detail in the technical solution of the task processing device in the field of bioinformatics, please refer to the description of the technical solution of the task processing method in the field of bioinformatics.
[0284] Figure 11 11 is a block diagram of a computing device according to an embodiment of the present disclosure. Components of the computing device 1100 include, but are not limited to, a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 1130, and a database 1150 is used to store data.
[0285] The computing device 1100 also includes an access device 1140 that enables the computing device 1100 to communicate via one or more networks 1160. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1140 may include one or more of any type of network interface (e.g., a network interface controller (NIC)) of wired or wireless type, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC).
[0286] In one embodiment of the present specification, the above components of the computing device 1100 and Figure 11 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 11 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0287] Computing device 1100 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1100 may also be a mobile or stationary server.
[0288] The processor 1120 is configured to execute the following computer program / instruction, which, when executed by the processor, implements the steps of the above-mentioned field task processing method or the task processing method in the field of bioinformatics.
[0289] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device is based on the same concept as the technical solutions of the aforementioned domain task processing method and the task processing method in the field of bioinformatics. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the aforementioned domain task processing method or the task processing method in the field of bioinformatics.
[0290] An embodiment of the present specification further provides a computer-readable storage medium storing a computer program / instruction. When the computer program / instruction is executed by a processor, the steps of the above-mentioned task processing method or the task processing method in the field of bioinformatics are implemented.
[0291] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solutions of the aforementioned domain task processing method and the bioinformatics task processing method. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the aforementioned domain task processing method or the bioinformatics task processing method.
[0292] An embodiment of this specification also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned field task processing method or the task processing method in the field of bioinformatics.
[0293] The above is a schematic diagram of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product shares the same concept as the technical solutions of the aforementioned domain task processing method and the task processing method in the field of bioinformatics. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solutions of the aforementioned domain task processing method or the task processing method in the field of bioinformatics.
[0294] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0295] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0296] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0297] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0298] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A domain task processing method, comprising: Obtain task information of target tasks in target domain; Based on the task information, searching for domain task units related to the target task from a task library corresponding to the target domain, wherein the task library is constructed based on at least one sample task in the target domain and a sample task unit called to execute the sample task; Based on the unit information of the domain task unit, a target generation model is used to generate call chain information that conforms to the target task execution logic; Based on the call chain information, the domain task unit is called to execute the target task and obtain the task execution result.
2. The method according to claim 1, wherein the task library corresponding to the target domain includes a task node of at least one sample task in the target domain and a unit node of a sample task unit called to execute the sample task, and there is a corresponding relationship between the task node and the unit node; The retrieving, based on the task information, domain task units related to the target task from a task library corresponding to the target domain, includes: Retrieving, based on the task information, a first task node of a first sample task related to the target task from a task library corresponding to the target domain; Based on the first task node, the first unit node of the first sample task unit called to execute the first sample task is retrieved from the corresponding relationship, and the first sample task unit is determined to be a domain task unit related to the target task.
3. The method according to claim 2, before retrieving the first task node of the first sample task related to the target task from the task library corresponding to the target domain based on the task information, further comprising: Acquire at least one sample task in the target domain and a sample task unit called for executing the sample task; Constructing a task node of the sample task based on the sample task, constructing a unit node of the sample task unit based on the sample task unit, and determining that there is a corresponding relationship between the task node and the unit node; Based on the task nodes, the unit nodes and the corresponding relationships, a task library corresponding to the target field is constructed.
4. The method according to claim 3, before constructing the task library corresponding to the target domain based on the task nodes, the unit nodes, and the corresponding relationships, further comprising: Acquire an initial task library corresponding to the target domain, wherein the initial task library includes at least one initial task node and at least one initial unit node corresponding to the initial task node; The step of constructing a task library corresponding to the target domain based on the task nodes, the unit nodes, and the corresponding relationships includes: Retrieving a first initial unit node similar to the unit node from the initial task library; If not found, based on the corresponding relationship, the task node and the unit node are recorded in the initial task library to obtain the task library corresponding to the target field.
5. The method according to claim 4, further comprising, after retrieving a first initial unit node similar to the unit node from the initial task library: If retrieved, identifying whether the similarity between the task node and the first initial task node corresponding to the first initial unit node meets a preset threshold; If not, the corresponding relationship between the task node and the first initial unit node is recorded to obtain a task library corresponding to the target domain.
6. The method according to claim 3, wherein the step of obtaining at least one sample task in the target domain and a sample task unit called to execute the sample task comprises: Acquire multiple material files in the target area; Using an information extraction model, extracting at least one sample task in the target domain and a sample task unit called for executing the sample task from the material file; The step of constructing a task library corresponding to the target domain based on the task nodes, the unit nodes, and the corresponding relationships includes: Integrate the task nodes, unit nodes and corresponding relationships of each material file to obtain a task library corresponding to the target field.
7. The method according to claim 1, wherein the task execution result includes alarm information; After calling the domain task unit to execute the target task based on the call chain information and obtaining the task execution result, the method further includes: Feedback the alarm information to the front-end user; Receiving updated call chain information sent by the front-end user based on the alarm information; Based on the updated call chain information, the domain task unit is called to execute the target task to obtain an updated task execution result.
8. The method according to claim 1, wherein the task execution result includes alarm information; After calling the domain task unit to execute the target task based on the call chain information and obtaining the task execution result, the method further includes: Based on the unit information of the domain task unit and the alarm information, using the target generation model to generate update call chain information that conforms to the target task execution logic; Based on the updated call chain information, the domain task unit is called to execute the target task to obtain an updated task execution result.
9. The method according to any one of claims 1 to 8, further comprising: after calling the domain task unit to execute the target task based on the call chain information and obtaining the task execution result; Based on the task execution results, a task analysis report is generated using a result analysis model.
10. The method according to any one of claims 1 to 8, before calling the domain task unit to execute the target task based on the call chain information and obtaining the task execution result, further comprising: Based on the call chain information, generating a call process diagram corresponding to the target task; Feedback the calling process diagram to the front-end user; Receive the updated call chain information sent by the front-end user with respect to the call process diagram.
11. The method according to claim 1, after generating call chain information that conforms to the target task execution logic based on the unit information of the domain task unit using a target generation model, further comprising: Based on the call chain information, a call process diagram is generated that conforms to the target task execution logic; Feedback the calling process diagram to the front-end user; Receiving update sample data sent by the front-end user for the call process diagram, wherein the update sample data includes update unit information and label call chain information of the update domain task unit of the sample task; The target generation model is trained based on the update unit information and the label call chain information.
12. A task processing method in the field of bioinformatics, comprising: Obtain task information of bioinformatics tasks in the field of bioinformatics; Based on the task information, searching for a domain task tool related to the bioinformatics task from a task library corresponding to the bioinformatics field, wherein the task library is constructed based on at least one sample task in the bioinformatics field and a sample task tool called to execute the sample task; Based on the tool information of the domain task tool, a target generation model is used to generate tool call chain information that conforms to the execution logic of the bioinformatics task; Based on the tool call chain information, the domain task tool is called to execute the bioinformatics task and obtain a task execution result.
13. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.
14. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.
15. A computer program product comprising a computer program / instruction, which implements the steps of the method according to any one of claims 1 to 12 when executed by a processor.