Intelligent agent automation system and equipment for transcriptomics regulation and control analysis

By designing the TransAgent intelligent agent automation system, the complexity of tool invocation and data integration in transcriptional regulation analysis is solved, enabling efficient and accurate transcriptional regulation analysis. It integrates multiple tools and data resources, supports dynamic control of complex tasks, and improves the transparency and security of the analysis process.

CN120877882APending Publication Date: 2025-10-31THE FIRST AFFILIATED HOSPITAL HENGYANG MEDICAL SCHOOL UNIV OF SOUTH CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511039158.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing transcriptional regulation analysis methods struggle to handle complex tool call relationships and large-scale epigenetic and expression data integration, and pose challenges for personalized analysis needs and data format conversion. In particular, for researchers lacking professional background, the analysis process is cumbersome and opaque, and lacks security and flexibility.

Method used

TransAgent, an intelligent automated system for transcriptomics regulation analysis, was designed. Through intelligent task management and flexible tool invocation, including functional modules, system prompt modules, and memory modules, it realizes task planning, execution, result integration, and memory storage. It supports modular and atomic operations, integrates a variety of transcription regulation tools and data resources, and provides automatic mode, execution mode, and planning mode. It also combines long and short memory to control the context length.

Benefits of technology

It improves the efficiency and accuracy of transcriptional regulation analysis, supports dynamic control of complex tasks, reduces repetitive work, saves storage space, integrates rich data resources and tools, adapts to different analysis needs, lowers the technical threshold for data processing, and improves the transparency and security of the analysis process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877882A_ABST
    Figure CN120877882A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent medical treatment, in particular to an intelligent agent automation system and equipment for transcriptomics regulation and control analysis. Comprising a function module, a system prompt module and a memory module. The system prompt module obtains transcriptional related data and / or instructions and then executes a planning mode to obtain a task planning path; the system prompt module is converted into a task execution mode, generates a structured task request instruction based on the task planning path, and then transmits the structured task request instruction to the function module; the function module calls one or more function sub-modules in the function module based on the task request instruction to execute a task to obtain a task result; the function module transmits the task result to the system prompt module, and then the system prompt module performs integration to obtain a transcriptomics regulation and control result; and the transcriptomics regulation and control result and the task planning path are stored in the memory module. The application can realize automatic transcriptomics regulation and control analysis, and has good clinical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent healthcare, specifically to an automated intelligent agent system, device, and computer-readable storage medium for transcriptomics regulatory analysis. Background Technology

[0002] Transcriptional regulation is a complex process that determines when and how different genes work to control cell growth; abnormal transcriptional regulation can lead to disease. Analyzing the patterns of transcriptional regulation can not only unravel the mysteries of life but also provide new insights for the development of various fields, such as precision medicine. However, transcriptional regulation research still faces many challenges. For example, the differential regulatory mechanisms among individuals necessitate personalized analysis, and processing large amounts of data requires complex analytical workflows, making it difficult for traditional methods to integrate and cover complex call relationships. Furthermore, because transcriptional regulation involves interactions between multiple omics, researchers need to further process and analyze transcriptome data in addition to epigenomic data such as Chip-seq and ATAC-seq when analyzing regulatory elements such as super enhancers. Downstream analysis also requires the use of numerous specialized transcriptional regulation analysis software programs, such as ROSE, Homer, and DeepTools. The entire process requires professional data preprocessing (such as quality control and standardization), regulatory element annotation, and in-depth analysis of the results. These steps are not only cumbersome but may also require manual adjustments by the user at each stage, making data preprocessing technically challenging. Furthermore, data format conversion between different tools and precise selection among various transcriptional regulation software present challenges, especially for researchers lacking a background in transcriptional regulation analysis.

[0003] In recent years, agents have continued to develop, such as MetaGPT and BioAgents. Recently, large language models (such as DeepSeek and ChatGPT) have made significant progress in multiple fields, driving the rapid development of the life sciences. Among these, various bioinformatics-related intelligent agent tools have emerged, such as CellAgent, BioAgents, and MRAgent. BioAgents uses small language models to handle genomics tasks with near-human expert performance; CellAgent specializes in analyzing single-cell data and can automatically select tools and parameters; and MRAgent can automatically mine causal relationships of diseases from literature. However, they are still limited to analyzing data in specific vertical domains, with relatively simple analysis workflows and straightforward tool-to-tool relationships. They lack flexibility when encountering new data or tasks, and the analysis process is opaque, posing security risks. Most importantly, they cannot handle the complex inter-tool relationships in transcriptional regulation analysis or the integration of large-scale epigenetic and expression data. Summary of the Invention

[0004] To address the aforementioned problems, this invention provides an intelligent agent automation system (TransAgent) for transcriptomics regulation analysis. Through intelligent task management and flexible tool invocation, it effectively solves these problems, enabling researchers to complete complex transcriptional regulation tasks more efficiently. Specifically, it includes: a functional module, a system prompt module, and a memory module.

[0005] After obtaining transcription-related data and / or instructions, the system prompting module executes the planning mode to obtain the task planning path;

[0006] The system prompt module converts to task execution mode and generates a structured task request instruction based on the task planning path, which is then input to the functional module;

[0007] The functional module calls one or more functional sub-modules in the functional module to execute the task and obtain the task result based on the task request instruction;

[0008] The functional module integrates the system prompts module after the task result transmission to obtain transcriptomics regulation results;

[0009] The transcriptomics regulation results and task planning paths are stored in the memory module.

[0010] The task execution mode includes any one of the following: execution mode and automatic mode; the automatic mode is to autonomously complete the task planning and obtain the task result; the execution mode is to complete the task planning and obtain the task result in combination with user instructions.

[0011] Optionally, the user instructions include task adjustment instructions, which include one or more of the following: task modular deletion or addition instructions, task atomic deletion or addition instructions; the transcription regulation task results can be adjusted by dynamically controlling any task process or task stage in the task planning path through user instructions.

[0012] Optionally, the task planning path consists of N task nodes and their relationships, where N is a natural number greater than 1. When a task adjustment instruction is obtained under the execution model, the current task execution node is interrupted, the adjusted task node is updated to form an updated task path, and the task node is judged. If the updated task node is before the current task node, the task is executed at the updated task node, and the task of the current task node is released. If the updated task node is after the current task node or a parallel task node, the current task node continues to execute and the task is executed based on the updated task path.

[0013] Optionally, the task node is determined by obtaining the historical task path from the memory module, comparing the current task node with the adjusted task node position order to obtain the determination result, and storing the updated task path in the memory module after the system prompt module executes based on the determination result.

[0014] Optionally, the task execution order or task module execution order in the task execution mode includes: parallel task execution, serial task execution, and a combination of parallel and serial tasks; the task execution order is formed by dynamic scheduling based on system resources and task completion progress; the task module execution order is formed by dynamic scheduling based on task node relationships.

[0015] Optionally, the task node relationships are obtained based on the task planning path, including: sequential relationships, multi-module fusion relationships, and jump fusion relationships.

[0016] The system prompt module includes an enhanced context management module for managing context messages or instructions, which include one or more of the following: system messages or instructions, assistant messages or instructions, user messages or instructions; and recording of transcriptomic regulatory processes based on context messages or instructions.

[0017] Optionally, the record is as follows: when the system prompt module generates a task planning path in planning mode, the context management module generates a structured task request instruction based on the task planning path. The structured task request instruction includes the thought process and application tools for building each task node, which is recorded as an assistant message or instruction; the task planning path is recorded as a system message or instruction; when the system prompt module is in task execution mode, the output of the execution result is based on the structured output result, and the standardized calling instruction and execution result are recorded as a user message or instruction.

[0018] Optionally, the application tool that requires parameters when used also includes parameter data recorded as assistant messages or instructions;

[0019] Optionally, the system further includes a correction tool, which is connected to the system prompt module and the function module. When it is detected that the output result of the task executed in the task execution mode is not a structured result, the result is converted into a structured result before further processing.

[0020] The system also includes an environmental information module, which is connected to the system prompt module. The environmental information module is used to monitor the system status in real time, including the current language, current working path, and current system time.

[0021] Optionally, the environment information module further includes a custom language module, through which the model's response language type is defined;

[0022] Optionally, when the system prompt module is a task execution module, it also includes environmental information acquisition. The environmental information module obtains detailed information about the current system environment, which is recorded as user messages or instructions. While the structured task request instruction is input to the functional module, the user messages or instructions and assistant messages or instructions are transmitted to the memory module for storage.

[0023] The system prompt module also includes a memory control module, which controls whether system messages or instructions, user messages or instructions, and assistant messages or instructions are stored, and dynamically adjusts the length of the message or instruction context to obtain the specified context content.

[0024] Optionally, the memory control module is connected to the correction tool. When the correction tool detects that the output result of the task executed in the task execution mode is not a structured result, the memory control module cancels the storage of the current output result.

[0025] Optionally, the memory regulation module can be used to combine long and short memories to reduce the context overflow of the model.

[0026] The memory module includes a fuzzy memory module and a precise memory module; the fuzzy memory module is used to store user messages or instructions or system messages or instructions; the precise memory module is used to store a complete list of messages when executing a task;

[0027] Optionally, the fuzzy memory module stores the message or instruction of a task in the form of a memory list, and the precise memory module stores all the messages for executing the task in the form of memory retrieval values;

[0028] Optionally, the memory module further includes a memory retrieval tool. When the system prompts the module to obtain a historical interactive query instruction, the memory retrieval tool in the memory module is used to query the list or search for values.

[0029] Optionally, when the system prompt module obtains transcription-related data or similar instructions of the same type, in the planning mode, it calls the system messages and user messages in the fuzzy memory module to generate a task planning path, and in the task execution mode, it calls the user messages in the fuzzy memory module and enters the automatic mode to execute the task.

[0030] The functional sub-modules include one or more of the following: a parameter adjustment module and a tooltip and definition module; the parameter adjustment module dynamically configures model and tool parameters through task request commands; the tooltip and definition module standardizes call commands or processes through task request commands;

[0031] Optionally, the functional submodule calls the parameter adjustment module and the tooltip definition module based on the task request instruction. The parameter adjustment module configures the transcriptional regulation model and model parameters. The tooltip and definition module then generates standardized call instructions or processes. The standardized call instructions or processes call the transcriptional regulation service to execute the request, obtain the request result, and then send the request result back to the transcriptional regulation model. The model is then used to perform transcriptional regulation analysis to obtain the task result.

[0032] Optionally, the transcriptional regulation service includes predefined transcriptional regulation query tools, including epigenetic annotation query, transcriptional regulation binding region query, and gene expression query.

[0033] Optionally, the transcriptional regulation service is invoked via a message transmission protocol in the cloud.

[0034] Optionally, the transcriptional regulation service also includes a transcriptional regulation software package, which requests the software package to be executed through standardized invocation instructions to obtain the software package analysis results, and then feeds the software package analysis results back into the transcriptional regulation model;

[0035] Optionally, the transcriptional regulation service is virtualized as a whole through a Docker image to obtain L transcriptional regulation services, where L is a natural number greater than 1, and the L transcriptional regulation services independently execute the request instruction based on the calling instruction;

[0036] Optionally, the functional submodule further includes a tool module, including one or more of the following: a remote invocation tool, a real-time display and interaction tool, a file upload tool, a file download tool, and a display tool; which performs remote invocation and / or display and / or upload and / or download of the task results or intermediate process results based on task request instructions.

[0037] The parameter adjustment module also includes a custom parameter module. The custom parameter module is used to configure custom parameters and custom structures for the transcriptional analysis model, and to execute any transcriptional task. When the function module calls the parameter adjustment module, it generates a transcriptional regulation model and model parameters. The custom parameter module is used to make custom adjustments to the transcriptional regulation model and model parameters to obtain a custom transcriptional regulation model.

[0038] The tooltips and definition module includes a tool definition module, through which tools are customized. The customized tools consist of core function definitions and tooltips.

[0039] Optionally, the tool definition module also includes integrated custom tools;

[0040] Optionally, the tooltips and definition module also includes a CLI prompt module, which includes S definitions and invocation methods for individual or integrated transcriptional regulation tools, where S is a natural number greater than or equal to 1; the CLI prompt module generates standardized invocation instructions or procedures, and then the transcriptional regulation service is invoked based on the standardized invocation instructions or procedures;

[0041] Optionally, the CLI prompt module is connected to the tool definition module. After the tool definition module builds a custom tool, it is input to the CLI prompt module, which then generates standardized call instructions or processes.

[0042] The functional submodule also includes a genome regulatory element annotation tool, which includes enhancer annotation data, genetic variation annotation data, chromatin openness annotation data, three-dimensional genome data, and DNA methylation data;

[0043] Optionally, the enhancer annotation data includes one or more of the following: SEdb database, SEA database, dbSuper database, EnhancerAtlas database, HACER database, ENCODE database, FANTOM5 database, DENDB database, ENdb database, and eRNAbase database;

[0044] Optionally, the genetic variation annotation data includes one or more of the following: dbSNP database, GWASCatalog database, GWASdb database, GTEx database, PancanQTL database, seeQTL database, SCAN database, and Oncobase database;

[0045] Optionally, the chromatin openness annotation data includes one or more of the following: ATACdb database, ENCODE database;

[0046] Optionally, the three-dimensional genome data includes one or more of the following: the 4DGenome database, the Oncobase database, and the 3D Genome Browser database;

[0047] Optionally, the DNA methylation data includes one or more of the following: the ENCODE database;

[0048] Optionally, the functional submodule also includes a target gene selection module. When the task node of the execution task path is target gene selection, the target gene selection module is called. The algorithm identifies the association between DNA regulatory elements and target genes to obtain the association relationship. Gene expression profile data is called, and target gene selection is performed based on the association relationship and gene expression profile data.

[0049] The purpose of this invention is to provide a computer device comprising a memory, a processor, and a computer program or instructions stored in the memory, wherein the computer program or instructions are executed by the processor to implement the above-described steps of the intelligent agent automated system module for transcriptomics regulatory analysis.

[0050] The purpose of this invention is to provide a computer-readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions are executed by a processor to implement the above-described intelligent agent automated system module steps for transcriptomics regulatory analysis.

[0051] Advantages of this invention:

[0052] 1. To enable TransAgent to be applied to more complex transcriptional regulation analyses and precise control of the analysis process, this invention provides a precise memory operation function. This invention allows users to dynamically enable or disable the interaction history of a specific step during task execution. This invention offers three options: modular operation, atomized manipulation, and memory deletion operation. Users can choose to disable, enable, or delete an entire execution cycle or an analysis step within a cycle. Specifically, for the fuzzy memory part, each memory atomization operation precisely controls an interaction cycle. For the precise memory part, disabling a specified memory removes the entire assistant's thought process, tool calls, and observation results from the PM. Through precise context task node control, it better meets user needs. For repetitive analyses, deletion reduces context length and saves storage space.

[0053] 2. Transcriptional regulation research is a crucial area of ​​study in life sciences, involving the joint analysis of vast amounts of multi-omics and differentially derived data. For example, constructing transcriptional regulatory networks requires researchers to manually process complex epigenomic and transcriptomic data and utilize various transcriptional regulation tools. However, the complex data types and numerous tool calls necessitate sophisticated data analysis capabilities and a significant, costly, and time-consuming process. More importantly, researchers spend considerable time resolving data input / output format adjustments between different software programs and performing reproducible analyses, further contributing to overall research inefficiency. This invention, through the adoption of the ReAct architecture, enhanced memory modules, and refined memory flow control, enables TransAgent to handle highly complex and long-cycle transcriptional regulation tasks. TransAgent effectively understands human language and automatically develops detailed analysis workflows based on researchers' requirements. Furthermore, through deep interaction, it allows for fine-tuning of tasks in the early stages of analysis. For complex, long-cycle analysis tasks, TransAgent supports atomic-level control of memory, enabling researchers to precisely control the entire analysis process.

[0054] 3. TransAgent integrates large-scale epigenomic annotation data, including enhancers, eRNAs, SNPs, and transcriptional regulators. For each data type, this invention collects data from multiple sources, ultimately resulting in a transcriptional regulation resource library from over dozens of sources. Regarding transcriptome data, this invention collects expression datasets from GTX, TCGA, CCLE, and ENCODE, covering comprehensive gene data types such as human tissues, cancer, and cells. In terms of tools, TransAgent currently integrates dozens of transcriptional regulation tools. Notably, to facilitate rapid tool expansion and the need for customized toolsets, this invention integrates MCP service functionality into TransAgent and packages all tools using Docker into an SSE MCP service called "biotools." Remote access to MCP enables TransAgent to rapidly expand its capabilities and achieve unified cloud and local deployments.

[0055] 4. Regarding process control, TransAgent proposes three execution modes: automatic mode, execution mode, and planning mode. Users can alternate between different modes to complete specific tasks throughout the overall analysis process. Experimental results show that when submitting new requests before or during each user task, using planning mode followed by automatic or execution mode yields better results. During interaction, agents inevitably make erroneous tool calls and execution results. By using memory for precise control, including opening and closing memory, the length of the context can be dynamically adjusted, allowing the agent to focus more on the user-specified context content. This approach is crucial and effective for LLM models with limited context length. Simultaneously, to further retain more interaction history within limited contexts, this invention uses a combination of long and short memory. This significantly reduces model forgetting and context overflow in large language models when analyzing complex transcriptional regulation tasks. Furthermore, the memory retrieval tool provided by this invention can retrieve complete tool call information and results stored in memory at the user's request or the agent's judgment. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0057] Figure 1 This is a schematic diagram of an automated intelligent system for transcriptomics regulatory analysis provided in an embodiment of the present invention;

[0058] Figure 2 This is a schematic diagram of the device provided in an embodiment of the present invention;

[0059] Figure 3 The TransAgent software architecture provided in this embodiment of the invention;

[0060] Figure 4This invention provides TransAgent for analyzing the transcriptional regulatory network driven by super-enhancers in esophageal squamous cell carcinoma. A. A pie chart, from left to right, shows the genomic proportions of super-enhancers identified by ROSE, the control group, and the experimental group. B. A bar chart shows the top 20 target genes of super-enhancers identified by BETA, ranked by regulatory potential score. C. A dot plot at the top shows the super-enhancers identified by ROSE, and the bottom shows the TFs enriched in the super-enhancer regions. D. DeepTools analysis results; the left side of a single image represents the experimental group, and the right side represents the control group. The top of the image shows the overall distribution of various local database region files, including Super Enhancer, Enhancer, eRNA, TFBS, and genes. The bottom of the image shows the coverage of data peaks in different regions. F. A heatmap on the left shows the expression correlation of key master regulators from ESCC RNA-seq sequencing (n=274). A heatmap on the right shows the expression correlation of key master regulators from TCGA ESCA RNA-seq sequencing (n=196).

[0061] Figure 5 This invention provides TransAgent for identifying key transcriptional regulators and their regulatory networks for cardiomyocyte differentiation. A represents uploaded data and simple prompts; TransAgent automatically plans the detailed execution process in planning mode. B represents the genomic proportion of binding sites identified by ChIP-seq for the key transcriptional regulators GATA4, NKX2-5, and TBX5 identified by TRAPT. C represents the average regulatory potential score of the top 40 non-redundant target genes of the key transcriptional regulators identified by BETA software. D shows the overlap of high-potential target genes identified by BETA software in the Venn diagram above. The network diagram below shows the regulatory network of key transcriptional regulators and high-potential target genes. E represents gene expression of key transcriptional regulators in normal human tissues. F shows the expression of high-potential target genes of key transcriptional regulators in normal human tissues. G shows the enrichment of high-potential genes in GO and Pathway. The size of the dots represents the number of genes covered in the set, and the color of the dots represents the enrichment fold. Detailed Implementation

[0062] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0063] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as S101, S102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0064] Figure 1 A schematic diagram of an automated intelligent system for transcriptomics regulatory analysis provided in this embodiment of the invention specifically includes:

[0065] System prompt module: After acquiring transcription-related data and / or instructions, the system prompt module executes the planning mode to obtain the task planning path; the system prompt module converts to the task execution mode and generates a structured task request instruction based on the task planning path, and then inputs it to the functional module;

[0066] In one embodiment, the transcription-related data includes one or more of the following: sequencing data (including RNA-Seq, single-cell RNA-Seq, and third-generation sequencing), microarray data, gene / transcriptome annotation data, sample metadata, differential expression analysis results, functional enrichment analysis results, alternative splicing analysis data, fusion gene detection results, and epitranscriptomics data.

[0067] In one embodiment, when the system prompt module is planning a task path in the planning mode, it acquires external expert knowledge, plans a task path based on the external expert knowledge for transcriptomics-related data, and acquires user messages or instructions. When the user message or instruction is the current task planning path, the system prompt module switches to automatic mode or execution mode in the task execution mode. When the user message or instruction is an adjustment, it acquires the adjustment instruction to adjust the task planning path to obtain the adjusted task planning path, and the system prompt module switches to automatic mode in the task execution model.

[0068] In another embodiment, the system prompting module obtains the user's intent through a question-and-answer interaction in the planning mode, and generates a task planning path based on the user's intent. The user intent is stored as a user message or instruction. The question-and-answer interaction is performed using a large language model.

[0069] In one embodiment, the task execution mode includes any one of the following: execution mode and automatic mode; the automatic mode is to autonomously complete the task planning and obtain the task result throughout the entire process; the execution mode is to complete the task planning and obtain the task result in combination with user instructions.

[0070] In one embodiment, the user instructions include task adjustment instructions, which include one or more of the following: task modular deletion or addition instructions, task atomic deletion or addition instructions; the transcriptional regulation task results are adjusted by dynamically controlling any task process or task stage in the task planning path through user instructions.

[0071] In one embodiment, the task planning path consists of N task nodes and their relationships, where N is a natural number greater than 1. When a task adjustment instruction is obtained under the execution model, the current task execution node is interrupted, the adjusted task node is updated to form an updated task path, and the task node is judged. If the updated task node is before the current task node, the task is executed by jumping to the updated task node and the task of the current task node is released. If the updated task node is after the current task node or a parallel task node, the current task node continues to be executed and the task is executed based on the updated task path.

[0072] In one embodiment, the task node is determined by obtaining the historical task path from the memory module, comparing the current task node with the adjusted task node position order by comparing the historical task path with the updated task path to obtain the determination result, and storing the updated task path in the memory module after the system prompt module executes based on the determination result.

[0073] In one embodiment, the task execution order or task module execution order under the task execution mode includes: parallel task execution, serial task execution, and a combination of parallel and serial tasks; the task execution order is formed by dynamic scheduling based on system resources and task completion progress; the task module execution order is formed by dynamic scheduling based on task node relationships.

[0074] In one embodiment, the task node relationship is obtained based on the task planning path and includes: sequential relationship, multi-module fusion relationship, and jump fusion relationship.

[0075] In another embodiment, when the system prompt module is in planning mode, it generates a discrete task node (subtask) after obtaining the user's intent based on user messages or instructions or through question and answer. When the system prompt module switches to task execution mode, it thinks about the discrete task node and calls the tool to execute to obtain preliminary results. Based on the preliminary results, it conducts further thinking and tool calls. The thinking process and tool calls are carried out synchronously.

[0076] In one embodiment, the system prompt module includes an enhanced context management module for managing context messages or instructions, wherein the context messages or instructions include one or more of the following: system messages or instructions, assistant messages or instructions, user messages or instructions; and recording of transcriptomic regulatory processes based on context messages or instructions.

[0077] In one embodiment, the record is as follows: when the system prompt module generates a task planning path in planning mode, the context management module generates a structured task request instruction based on the task planning path. The structured task request instruction includes the thought process and application tools for building each task node, which are recorded as assistant messages or instructions; the task planning path is recorded as system messages or instructions; when the system prompt module is in task execution mode, the output of the execution result is based on the structured output result, and the standardized call instructions and execution results are recorded as user messages or instructions.

[0078] In one embodiment, the application tool that requires parameters when in use also includes parameter data denoted as assistant messages or instructions.

[0079] In one embodiment, the system further includes a correction tool connected to the system prompt module and the function module. When it is detected that the output result of the task executed in the task execution mode is not a structured result, the correction tool converts the result into a structured result before proceeding to the next step.

[0080] In one embodiment, the system further includes an environmental information module connected to the system prompt module. The environmental information module is used to monitor the system status in real time, including the current language, current working path, and current system time.

[0081] In one embodiment, the environment information module further includes a custom language module, through which the model answer language type is defined.

[0082] In one embodiment, when the system prompt module is a task execution module, it further includes environmental information acquisition. The environmental information module obtains detailed information about the current system environment, which is recorded as user messages or instructions. While the structured task request instruction is input to the functional module, the user messages or instructions and assistant messages or instructions are transmitted to the memory module for storage.

[0083] In one embodiment, the system prompt module further includes a memory control module, which controls whether system messages or instructions, user messages or instructions, and assistant messages or instructions are stored, dynamically adjusts the length of the message or instruction context, and obtains the specified context content.

[0084] In one embodiment, the memory control module is connected to the correction tool. When the correction tool detects that the output result of the task executed in the task execution mode is not a structured result, it cancels the storage of the current output result through the memory control module.

[0085] In one embodiment, the memory regulation module combines long and short memories to reduce context overflow in the model.

[0086] In one specific embodiment, the context messages of the agent (system prompt module) mainly include three role types: system messages, assistant messages, and user messages. To effectively reduce model illusion and address memory loss issues, this invention specifies that the model's output must be in standard JSON format. For each step of the agent's output, it must include three parts: "Thinking," "Tools," and "Parameters." Except for "Parameters," which may be empty depending on the tool, the other two are mandatory. In the "Thinking" part, the model should provide the thought process for the current step and the reason for calling the tool. In the "Tools" part, the model should provide the name of the tool being called. And in the "Parameters" part, the model should provide the parameters required by the called tool. These together constitute the main body of the current assistant message. The execution process and results of the tool are also specified in standard JSON format and saved as user messages. Verification by this invention shows that ensuring the agent's standardized output effectively improves the model's instruction-following ability. Furthermore, if the model does not output in a standardized format, TransAgent will use the standardized format to indicate problems in the model's output and allow the agent to correct them itself. The context related to the incorrect output format is not submitted as memory. System prompts (system messages) and environmental details (user messages) are provided separately in each model call but are not stored in the message list.

[0087] To address the problem of memory loss, this invention defines fuzzy memory (FM) and precise memory (PM) modules. Fuzzy memory retains longer but more vague information, while precise memory retains shorter but more specific details. More specifically, in fuzzy memory, this invention only retains the user's message (the user's task description and the name of the tool invoked), the "thinking" section of the large model assistant, and the index of the corresponding detailed information. These messages explicitly define the task to be performed, the assistant's thought process for invoking the tool, and the name of the tool invoked. Fuzzy memory is included as part of the "memory list" in the system prompts, placed at the end of each large model request.

[0088]

[0089] Precise memory retains the entire message list, but it is truncated according to the user-defined context length when the actual model is requested, retaining only information from the most recent interactions:

[0090]

[0091] Meanwhile, in order to enable the agent to recall details of previous interactions, this invention designs a "memory retrieval" tool. When the agent is asked by the user to recall relevant content or details in the list of past interaction messages, it will call this tool and retrieve details of key messages based on the saved index values.

[0092] To further enhance TransAgent's applicability to more complex transcriptional regulation analyses and enable precise control over the analysis process, this invention provides a precise memory operation function. This invention allows users to dynamically enable or disable the interaction history of a specific step during task execution. The invention offers three options: modular operation, atomic manipulation, and memory deletion operation. Users can choose to disable, enable, or delete an entire execution cycle or an analysis step within a cycle. Specifically, for fuzzy memory, each atomic manipulation precisely controls an interaction cycle, such as removing specified assistant thought content from the FM (Factor Memory). For precise memory, disabling a specified memory removes the entire assistant thought, tool calls, and observation results from the PM (Product Memory). It is worth noting that traditional agents can only handle potential errors or execution failures through global retries or iterative loops, and they cannot dynamically control the length of the global context, leading to the unavailability of complex analyses. With the precise context manipulation of this invention, users can more controllably guide the agent's behavior, and in very long contexts, by disabling intermediate environment steps, the agent can analyze the user's initial needs, such as adding analysis or visualization results. In parallelized processes, such as repetitive analysis, each repetitive analysis process can be shut down as a whole after it ends, thereby greatly reducing the length of the context.

[0093] In one specific embodiment, the transcriptional regulation MCP service:

[0094] This invention provides a general transcriptional regulation MCP service, which defines a rich set of transcriptional regulation data query tools, including epigenetic annotation query (this invention comprehensively collects epigenetic annotation data from over 20 sources, covering enhancers, super enhancers, eRNA, SNPs, risk SNPs, eQTLs, transcription factor binding sites, and CRISPR, etc.), transcription regulator binding region query (this invention manually collects over 17,000 comprehensive transcription regulator binding data), gene expression query (including expression data types from multiple sources such as GTEx, TCGA, CCEL, and ENCODE), and gene coordinate acquisition. In system command calls, this invention also defines a rich set of transcriptional regulation software packages, such as homer, deeptools, ROSE, BETA, and TRAPT developed in this invention, totaling over 30 software transcriptional regulation analysis toolkits. It is worth noting that only the MCP service provided by this invention can call TRAPT and further obtain binding files of key transcription regulators. To further simplify the deployment process and achieve high-performance online server analysis, this invention packages the MCP service into a unified Docker image and uses SSE in MCP for message transmission. It is worth noting that due to the complexity and time-consuming nature of transcriptional regulation analysis, traditional clients cannot obtain real-time execution results from the tools. Compared to similar clients, the TransAgent client innovatively features SSH remote command invocation and real-time interactive display tools, as well as file upload, download, and display tools. Therefore, the TransAgent client is recommended as a dedicated client application for MCP services requiring real-time and long-term transcriptional regulation analysis, because TransAgent provides richer details of tool execution and more convenient file transfer and display functions compared to similar clients.

[0095] In one specific embodiment, transcriptional regulation-related multi-omics internal annotation data: The TransAgent tool constructed a comprehensive annotation system for human genome regulatory elements, which specifically integrated important database resources independently developed by our research group. For enhancer annotation, the system included 2,678,273 super enhancers from the SEdb database independently developed by our research group, while also integrating relevant data from the SEA and dbSuper databases. Ordinary enhancer annotation integrated 14,797,266 enhancers from the EnhancerAtlas, HACER, ENCODE, FANTOM5, DENDB, and ENdb databases, and prominently included 10,399,928 eRNA data from the eRNAbase database independently developed by our research group. The genetic variation annotation includes 37,302,978 common SNPs from the dbSNP database, all of which meet the minimum allele frequency requirement of greater than 0.05. It also integrates 351,728 GWAS risk SNPs from the GWAS Catalog and GWASdb databases, as well as 11,995,221 eQTL sites from the GTEx, PancanQTL, seeQTL, SCAN, and Oncobase databases. For chromatin openness annotation, the system includes over 130,000,000 ATAC-seq open regions from the ATACdb database developed by the research group, and 69,860,705 DNase-sensitive sites from the ENCODE database. The three-dimensional genome data includes 34,342,926 chromatin interaction sites from the 4DGenome and Oncobase databases, and 72,019 TAD structures from the 3D Genome Browser database. The DNA methylation data encompasses 166,855,665 whole-genome methylation sequencing sites and 30,392,523 450K microarray sites from the ENCODE database. Furthermore, TransAgent effectively identifies the associations between DNA regulatory elements and target genes by utilizing the BETA and geneMapper algorithms. By further integrating gene expression profiling data from databases such as GTEx, TCGA, ENCODE, and CCLE, it optimizes the target gene selection process for regulatory elements, providing a complete solution for transcriptional regulation research.

[0096] In one specific embodiment, the system prompt module employs both fixed key system prompts and injectable custom system prompts. The key system prompts include role definitions, core tool definitions, and mode definitions. In the role definition, the invention defines the agent as an all-around AI assistant capable of calling a rich set of tools to complete user tasks, and provides the format and examples for tool calls. In the core tools, the invention defines several tools, including "MCP service call," "ask user questions," "planned mode reply," "wait for user feedback," "memory retrieval," and "task completion" tools. In the mode definition, the invention defines three modes: "automatic mode," "execution mode," and "planned mode." In planned mode, the invention states that the agent's goal is to collect information and obtain context to create a detailed plan to complete the user's task. In this mode, the invention only allows the agent to call the "planned mode reply" tool to answer and create the planned process. The user will review and approve the plan, then switch to execution mode or automatic mode to implement the solution. In execution mode, the invention specifies that the agent can use tools other than the "planned mode reply" tool to complete the user's task. In automatic mode, this invention stipulates that the intelligent agent does not need to ask the user questions and completes the subsequent process autonomously until the mode changes.

[0097] This invention incorporates specific prompts related to transcriptional regulation analysis tasks into its custom system prompts. Key aspects include task requirements and precautions. The task requirements specify the location where the agent's analytical data should be stored and advise the agent to provide analyzable data and analysis options. The precautions stipulate that the agent is prohibited from modifying core data content and remind it to pay attention to the format and differences in MCP and basic tool calls to reduce model illusions. The custom system prompts allow users to quickly fine-tune the agent's behavior based on model and task differences. Furthermore, TransAgent provides a one-click option to load custom system prompts.

[0098] In one specific embodiment, the environmental information module: This invention reserves injection interfaces in both system prompts and environmental detail information. In the system prompts, this invention provides operating system type, platform and architecture information, and a fuzzy memory list, etc. This information is necessary for the model to correctly execute system instructions. For example, this invention needs to remotely execute Linux system instructions on a Windows platform; this information needs to be accurately specified, otherwise it may lead to agent confusion. Environmental detail information is designed to include key environmental information that appears after each user message, including the current system time, the currently used language, and the current working path. This invention reserves an interface for defining the currently used language, thus enabling precise control over the model's response language type.

[0099] Functional Module: The functional module calls one or more functional sub-modules based on the task request instruction to execute the task and obtain the task result; the functional module transmits the task result to the system prompt module, and the system prompt module integrates the results to obtain transcriptomics regulation results;

[0100] In one embodiment, the functional sub-module includes one or more of the following: a parameter adjustment module and a tooltip and definition module; the parameter adjustment module dynamically configures the model and tool parameters through task request commands; the tooltip and definition module standardizes the calling commands or processes through task request commands.

[0101] In one embodiment, the functional submodule calls the parameter adjustment module and the tooltip definition module based on the task request instruction. The parameter adjustment module configures the transcriptional regulation model and model parameters, and the tooltip and definition module generates standardized call instructions or processes. The standardized call instructions or processes call the transcriptional regulation service to execute the request, obtain the request result, and then send the request result back to the transcriptional regulation model to perform transcriptional regulation analysis to obtain the task result.

[0102] In one embodiment, the transcriptional regulation service includes predefined transcriptional regulation query tools, including epigenetic annotation query, transcriptional regulation binding region query, and gene expression query.

[0103] In one embodiment, the transcriptional regulation service is invoked via a message transmission protocol in the cloud.

[0104] In one embodiment, the transcriptional regulation service further includes a transcriptional regulation software package, which requests the software package to be executed through standardized invocation instructions to obtain the software package analysis results, and then feeds the software package analysis results back into the transcriptional regulation model.

[0105] In one embodiment, the transcriptional regulation service is virtualized as a whole through a Docker image to obtain L transcriptional regulation services, where L is a natural number greater than 1, and the L transcriptional regulation services independently execute the request instruction based on the call instruction.

[0106] In one embodiment, the functional submodule further includes a tool module, comprising one or more of the following: a remote invocation tool, a real-time display and interaction tool, a file upload tool, a file download tool, and a display tool; capable of remotely invoking and / or displaying and / or uploading and / or downloading the task results or intermediate process results based on task request instructions.

[0107] In one embodiment, the tool parameters in the parameter adjustment module further include a custom parameter module. The custom parameter module is used to configure custom parameters and custom structures for the transcriptional analysis model, and to execute arbitrary transcriptional tasks using the custom parameters and custom structures. When the function module calls the parameter adjustment module, it generates a transcriptional regulation model and model parameters. The custom parameter module is used to make custom adjustments to the transcriptional regulation model and model parameters to obtain a custom transcriptional regulation model.

[0108] In one embodiment, the tooltips and definition module includes a tool definition module, through which tools are customized. The customized tools consist of core function definitions and tooltips.

[0109] In one embodiment, the tool definition module further includes integrated custom tools.

[0110] In one embodiment, the tooltips and definition module further includes a CLI prompt module, which includes S definitions and invocation methods for individual or integrated transcriptional regulation tools, where S is a natural number greater than or equal to 1; the CLI prompt module generates standardized invocation instructions or procedures, and then the transcriptional regulation service is invoked based on the standardized invocation instructions or procedures.

[0111] In one embodiment, the CLI prompt module is connected to the tool definition module. After the tool definition module builds a custom tool, it is input to the CLI prompt module, which then generates standardized call instructions or processes.

[0112] In one embodiment, the functional submodule further includes a genome regulatory element annotation tool, which includes enhancer annotation data, genetic variation annotation data, chromatin openness annotation data, three-dimensional genome data, and DNA methylation data.

[0113] In one embodiment, the enhancer annotation data includes one or more of the following: SEdb database, SEA database, dbSuper database, EnhancerAtlas database, HACER database, ENCODE database, FANTOM5 database, DENDB database, ENdb database, and eRNAbase database.

[0114] In one embodiment, the genetic variation annotation data includes one or more of the following: dbSNP database, GWAS Catalog database, GWASdb database, GTEX database, PancanQTL database, seeQTL database, SCAN database, and Oncobase database.

[0115] In one embodiment, the chromatin open annotation data includes one or more of the following: the ATACdb database and the ENCODE database.

[0116] In one embodiment, the three-dimensional genome data includes one or more of the following: the 4D Genome database, the Oncobase database, and the 3D Genome Browser database.

[0117] In one embodiment, the DNA methylation data includes one or more of the following: the ENCODE database.

[0118] In one embodiment, the functional submodule further includes a target gene selection module. When the task node of the execution task path is target gene selection, the target gene selection module is called, and the association relationship between DNA regulatory elements and target genes is identified through an algorithm. Gene expression profile data is called, and target gene selection is performed based on the association relationship and gene expression profile data.

[0119] Memory module: The transcriptomics regulatory results and task planning paths are stored in the memory module;

[0120] In one embodiment, the memory module includes a fuzzy memory module and a precise memory module; the fuzzy memory module is used to store user messages or instructions or system messages or instructions; the precise memory module is used to store a complete list of messages when performing a task.

[0121] In one embodiment, the fuzzy memory module stores the message or instruction of a task in the form of a memory list, and the precise memory module stores all messages for executing the task in the form of memory retrieval values;

[0122] In one embodiment, the memory module further includes a memory retrieval tool. When the system prompts the module to obtain a historical interactive query instruction, the memory retrieval tool in the memory module is used to query the list or search for values.

[0123] In one embodiment, when the system prompt module obtains similar transcription-related data or similar instructions, in the planning mode, it calls the system messages and user messages in the fuzzy memory module to generate a task planning path, and in the task execution mode, it calls the user messages in the fuzzy memory module and enters the automatic mode to execute the task.

[0124] In one specific embodiment, the CLI prompt module: To further enhance the adjustability and controllability of the agent, this invention allows users to inject prompts for core system commands to invoke tools. Specifically, this invention provides multiple tools for executing system commands, including "CLI command execution," "file display," "file viewing," "file search," "Python execution," and "network search," among others. "CLI command execution" and "file display" can connect to a remote Docker container via SSH and execute key system commands and invoke tools. To quickly expand the tools, this invention injects various definitions and invocation methods for transcriptional regulation tools into "CLI command execution," greatly reducing the complexity and cumbersomeness of tool definition and providing a unified tool invocation method, thus significantly stabilizing the standardization of large language pattern output. Simultaneously, this invention allows users to customize CLI prompt content and use it through simple configuration in TransAgent. When facing specific complex transcriptional regulation analysis tasks, using this function in conjunction with customized system prompts can effectively reduce the complexity of tool invocation and provide precise control over the analysis process.

[0125] Tool Definition Module: While providing multiple system tools, ransAgent also offers an interface for user-defined tools. Custom tools are defined in the same way as system tools, including core function definitions and tooltips. TransAgent also provides several custom tools, including web search and OCR recognition. This approach allows users to customize their own tools while ensuring they are called at the same level as system tools, further reducing the complexity of tool invocation.

[0126] Model Parameter Module: The transcriptomics regulation model (large language model, default DeepSeek model) includes several important parameters, including context length, temperature, maximum token count, and streaming response, etc. For example, the definition of the temperature parameter should differ under different conditions; lower temperatures imply more stable output, which is crucial for the reproducibility of transcriptional regulation analysis tasks. However, this does not mean that higher temperatures will lead to task failure; the resulting outcomes can vary. Different models may have even more definable parameters. For example, in the DeepSeek-Chat model, the model response type can be defined as fixed in JSON format, which is very important in the current design architecture of this invention. Therefore, this invention reserves interfaces for all model parameters to adapt to different tasks and large language models.

[0127] In one embodiment, the various modules in the system form a complete system, enabling (analysis) interaction of transcription-related data through messages or instructions, functional modules, and tools. For input transcription-related data and multi-task analysis instructions, the system generates a corresponding number of task planning paths based on the instructions. A question-and-answer format is used to determine whether multiple tasks are related. If the messages indicate a relationship, the relevant paths are connected; otherwise, execution is performed in parallel and / or serially based on system resources. To avoid ambiguity, task nodes in the task planning paths are designated as subtasks. The relationships between subtasks are determined by detecting or obtaining messages or instructions. If a relationship exists, the task node relationship (subtask relationship) is marked, and the subtasks are executed serially and / or in parallel.

[0128] In one specific implementation, TransAgent offers three operating modes to adapt to transcriptional regulation analysis scenarios of varying complexity, catering to different research needs: Planning Mode: Accurately captures user needs through deep interaction, such as transcription factor activity prediction, binding prediction, epigenome annotation, and gene expression analysis, and can generate detailed analysis workflows to ensure the scientific rigor and reproducibility of subsequent execution. Execution Mode: Flexibly utilizes various transcriptional regulation tools and provides real-time feedback on issues during execution, ensuring stable progress of the analysis workflow. Automation Mode: Completes the entire analysis process without manual intervention, significantly improving data processing efficiency. This invention applies TransAgent to classic transcriptional regulation analysis tasks such as super-enhancer regulatory network construction and key regulator identification of cardiomyocyte differentiation, achieving reliable results as expected.

[0129] TransAgent possesses the ability to rapidly integrate new tools (including custom system-level tools and MCP services) without modifying the core code, greatly improving scalability. This invention also utilizes Docker virtualization cloud technology to handle computationally intensive tasks, saving local resources while ensuring speed. Regarding model output standardization, TransAgent requires all outputs to be in JSON format, fully recording each step of "thinking-tool invocation-result observation," ensuring the analysis process is traceable and reproducible. In terms of security, the system uses access control and container isolation technology to prevent user error or accidental modification of original data. TransAgent also has significant advantages in dynamic memory management, selectively retaining or hiding previous dialogue memories to prevent inaccurate final analysis results due to memory forgetting or confusion caused by excessively long dialogues or other redundant information. In addition, as an intelligent system designed specifically for the field of transcriptional regulation, it supports zero-code operation, can automatically optimize the analysis process (such as dynamically adjusting peak calling parameters according to the quality of Chip-seq data), flexibly connect tools (such as from Chip-seq differential peak analysis to transcription factor-target gene network construction), and flexibly reuse historical analysis results (such as identified super enhancer regions) through a dynamic memory management system, significantly improving the efficiency of complex tasks. These features make TransAgent an important bridge connecting researchers and computational tools, and provide useful new ways to help researchers solve complex transcriptional regulation problems.

[0130] In one specific embodiment, transcriptional regulation data sources are complex and analytical methods are diverse. Analyzing transcriptional regulation typically requires manually collecting large amounts of specific omics data and performing functional annotation using region annotation methods. However, the complexity of omics and the changing needs of researchers for specific tasks make a unified workflow impossible. To address this challenge, this invention introduces TransAgent, an interactive agent software focused on transcriptional regulation analysis. TransAgent is an agent system software with ReAct as its core architecture. It can quickly invoke various tools defined in this invention and rapidly extend domain-specific tools using the advanced MCP (Model Context Protocol) service to complete researchers' personalized transcriptional regulation analysis tasks. Figure 3 (a). Meanwhile, through virtualization container technology, this invention migrates all computationally intensive tasks to professional-grade servers, and executes cloud-based system-level instructions via the remote invocation tool of this invention's software, achieving a unified approach of rapid local deployment and online analysis. Figure 3(b) Due to the context length limitation of large language models, the agent (system prompting module) exhibits significant "memory loss" in long interactive tasks. Although several large language models have long context windows of 1 million tokens, such as Gemini 2.5 Pro and Qwen2.5-1M, longer contexts are more likely to cause the model to ignore key information. To effectively solve this problem, this invention uses a combination of fuzzy memory and precise memory in TransAgent (b). Figure 3 (c) Precise memory refers to the agent's precise use of messages from the most recent dialogue cycle in a large language model request, while fuzzy memory refers to the agent providing its thought process, tool usage, and an index of precise memory at the end of the system prompt, but not retaining the specific tool usage results. For content the agent needs to recall, this invention provides a memory retrieval tool to retrieve the index of precise memory within fuzzy memory, enabling the agent to further reduce context length while ensuring the retention of key information. More importantly, this invention innovatively proposes a dynamic context memory management function, allowing TransAgent to achieve fine-grained control over the analysis process and the ability to analyze ultra-long contextual interactions by allowing users to manually enable or disable the interaction history of specified steps.

[0131] TransAgent interacts with the backend large language model through natural language dialogue. This invention offers different interaction modes tailored to various research needs, primarily including three modes: automatic mode, execution mode, and planning mode. Figure 3(a) Users can switch between different modes to maximize control over various complex transcriptional regulation analysis tasks. In planning mode, TransAgent interacts deeply with the user based on the current task requirements, repeatedly asking for more detailed information, including the data source, whether to upload their own data, or use a local database. After collecting sufficient information, the agent provides specific sub-tasks and then prompts the user to manually (obtain user instructions) enter execution mode or automatic mode. In execution mode, TransAgent automatically considers and invokes appropriate tools based on the current sub-task; the consideration process and tool invocation are simultaneous. After invoking a tool, the agent observes the execution results and proceeds with the next consideration and tool invocation until all tasks are completed. It is worth noting that the agent may fail during execution; in such cases, the agent will attempt to resolve the problem manually. When encountering unsolvable problems, such as missing input files, it will prompt the user to upload the corresponding files and pause the current task until the user provides feedback, at which point the agent will resume execution. In automatic mode, this invention stipulates that TransAgent is prohibited from interacting with the user; all problems encountered should be manually attempted to be resolved, which is a necessary option when fully automated control is required. Other execution processes remain consistent with the execution mode until the task is completed and the results are output.

[0132] In one specific embodiment, TransAgent analyzes the transcriptional regulatory network driven by super-enhancer in esophageal squamous cell carcinoma:

[0133] Super-enhancer-driven transcriptional regulation plays a crucial role in cancer pathogenesis, and elucidating its molecular mechanisms is essential for discovering oncogenic drivers. To evaluate TransAgent's ability to elucidate disease-related regulatory circuits, this invention analyzes the super-enhancer transcriptional network of esophageal squamous cell carcinoma (ESCC). The study first uploaded raw H3K27ac Chip-seq sequencing data (fastq files) to TransAgent, which then autonomously initiated a full-process analysis, including data quality control (FastQC), sequence alignment (Bowtie2), and peak detection (MACS2). Through a conversational analysis workflow, TransAgent, in conjunction with the above analysis results, can invoke relevant tools, such as the Chipsearer software, based on user prompts to automatically complete code analysis and provide visual visualization. Figure 4 TransAgent also integrates multiple strategies, including gene mapping (genemapper) and BETA, to identify downstream target genes of super enhancers (A). Figure 4(B). Further results from the super-enhancer prediction (ROSE) showed that the identified super-enhancers were highly consistent with previously reported ESCC super-enhancer maps, validating the accuracy of TransAgent in chromatin characterization. Figure 4 (C). To further decode the regulatory circuits, TransAgent was prompted to use the CRCMapper tool for analysis, and the core transcriptional network was successfully reconstructed. By searching for highly expressed genes in ESCC and filtering for potential master regulatory transcription factors identified by CRCMapper, this invention identified 12 key regulatory factors, which is highly similar to previously reported results. Figure 4 (D). By constructing a correlation network using these key regulators and the expression profiles of ESCC and TCGA ESCA (D). Figure 4 The study also identified potential correlations between key master-regulatory transcription factors such as TP63, SOX2, and KLF5—all of which have been confirmed as important oncogenic transcriptional regulators in ESCA. Simultaneously, using DeepTools, the study performed genome binding enrichment analysis on the identified peaks. TransAgent automatically extracted various local annotation data and revealed the binding of these peaks in super-enhancer regions, enhancer regions, eRNA regions, and transcription factor binding sites, identifying significant enrichment in these regions. Figure 4 This further reveals the crucial role of epigenetics in maintaining malignant transcriptional programs. This multilevel analysis approach not only improves predictive accuracy but also effectively reduces the false positive rate common in traditional single-analysis methods. Notably, TransAgent, through seamless integration of a series of computational tools, demonstrates end-to-end automation, analytical robustness, and translational application potential, making it a powerful platform for accelerating the discovery of cancer therapeutic targets.

[0134] In one specific embodiment, TransAgent identifies key transcriptional regulators and their regulatory networks for cardiomyocyte differentiation:

[0135] Understanding the transcriptional regulation of cardiomyocyte differentiation is crucial for cardiovascular research. To demonstrate TransAgent's ability to decipher differentiation-related gene regulatory networks, this invention uses differentially expressed genes (DEGs) during cardiomyocyte differentiation as a case study. The study first inputs the Top 200 differentially expressed genes from RNA-seq data of developing cardiomyocytes under two conditions into TransAgent. In the planning mode, TransAgent automatically generates detailed analysis steps (…). Figure 5(A). After switching execution modes, TransAgent automatically executes and predicts upstream transcriptional regulators using the built-in TRAPT tool, successfully screening for the major cardiac regulators GATA4, NKX2-5, and TBX5—factors known to coordinate heart development through specific transcriptional programs. Following user prompts for genome task distribution analysis, TransAgent plans the analysis workflow and automatically calls the Chipseeker software for analysis, generating visualization code (A). Figure 5 (B) In this invention, significant high expression of key transcriptional regulators and their binding preference to promoters and enhancers were observed in cardiac tissue, demonstrating the importance of key regulators identified by TransAgent analysis in cardiac differentiation. Furthermore, in a further target gene suggestion task, TransAgent constructs a cyclical task flow, calling BETA software multiple times based on the predicted key transcriptional regulators to further infer their downstream target genes using regulatory potential. Figure 5 (C). Meanwhile, the gene regulatory network mapped by TransAgent clearly shows how these transcription factors synergistically regulate cardiomyocyte maturation (C). Figure 5 TransAgent, with user prompts, provides biological interpretations of predicted target genes, further enhancing the credibility of network biology relevance. By prioritizing high-confidence regulatory interactions, it enables researchers to achieve accurate analysis and summarization without manual literature mining or building computational workflows. In addition to upstream transcription factor identification, this invention further prompts TransAgent to perform expression analysis of key target genes. TransAgent can extract target gene expression profiles from the GTEx database and automatically generate visualization charts. Figure 5 The clustering diagram clearly shows that key targets are enriched in heart and muscle tissue, further demonstrating the accuracy of the TransAgent analysis. To verify the functional significance of the prediction network, this invention suggests that TransAgent perform enrichment analysis on the target gene set and find that it is significantly associated with key pathways in heart development, such as myocardial contraction. Figure 5 The system also provides tissue-specific expression profiles of identified regulatory factors (G). Figure 5 The presence of E genes further supports their role in cardiomyocyte differentiation. This further highlights the value of TransAgent in simplifying the regulatory network inference process—a fully interactive, language-based analysis from differential gene analysis to master regulator prediction and functional validation. This platform automates complex bioinformatics processes while maintaining biological interpretability, helping researchers efficiently reveal novel transcriptional regulatory mechanisms in developmental biology and disease.

[0136] In one specific embodiment, TransAgent is an intelligent transcriptional regulation analysis platform driven by a large language model. Its core architecture, through the efficient collaboration of multiple modules, automates the process from data preprocessing to advanced analysis. The platform adopts a modular design, seamlessly integrating functions such as parameter tuning, tool management, environment monitoring, and memory optimization to form a flexible and scalable interactive analysis system. The parameter tuning module dynamically configures model and tool parameters, ensuring the stability and reproducibility of task execution, while collaborating with the system prompt module to guide the model in generating standardized outputs. The tool prompt and definition module standardizes the tool invocation process, and through deep integration with the MCP service, quickly schedules cloud-based or local transcriptional regulation tools (such as ROSE and BETA), adapting to the needs of different computing environments. The environment information module monitors the system status in real time (such as path and language settings), providing contextual support for tool invocation, while the memory module optimizes the model's inference ability for long-cycle tasks through dynamic management of precise and fuzzy memory, avoiding the loss of key information. The system prompt module, as the central hub, integrates the inputs from various modules and generates structured model requests, ensuring the coherence and controllability of the analysis process.

[0137] The core advantage of this architecture lies in its closed-loop interaction mechanism between modules. After a user submits a task via natural language, the system prompts the module to parse the requirements and trigger model planning; tool calls are executed through the MCP service, and the results are stored by the memory module and fed back to subsequent steps; the environment and parameter modules adjust resource allocation in real time to improve efficiency. For example, in super enhancer analysis, TransAgent automatically coordinates Chip-seq data processing, peak annotation, and network construction tools, while retaining intermediate results through memory backtracking, ultimately generating a reproducible regulatory network. This highly collaborative design not only solves the complexity of integrating multi-omics tools but also adapts to diverse analysis scenarios, from basic annotation to high-order network inference, through dynamic parameter optimization and context management, providing an efficient and transparent intelligent analysis paradigm for transcriptional regulation research.

[0138] Figure 2 An embodiment of the present invention provides a schematic diagram of a computer device, specifically including:

[0139] A memory and a processor; the memory is used to store program instructions; the processor is used to invoke the program instructions, when the program instructions are executed any of the above-described steps of the intelligent agent automated system module for transcriptomics regulatory analysis.

[0140] The present invention also discloses a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, comprises any of the above-described steps of the intelligent agent automated system module for transcriptomics regulatory analysis.

[0141] The verification results of this verification embodiment show that assigning inherent weights to indications can improve the performance of this method compared to the default settings. Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, indirect coupling or communication connection of devices or units, and may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated; the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of this embodiment. Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. This program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0142] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0143] The computer device provided by the present invention has been described in detail above. For those skilled in the art, there will be changes in the specific implementation and application scope based on the ideas of the embodiments of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. An automated intelligent system for transcriptomics regulatory analysis, characterized in that, include: Functional modules, system prompt modules, and memory modules; After obtaining transcription-related data and / or instructions, the system prompting module executes the planning mode to obtain the task planning path; The system prompt module converts to task execution mode and generates a structured task request instruction based on the task planning path, which is then input to the functional module; The functional module calls one or more functional sub-modules in the functional module to execute the task and obtain the task result based on the task request instruction; The functional module integrates the system prompts module after the task result transmission to obtain transcriptomics regulation results; The transcriptomics regulation results and task planning paths are stored in the memory module.

2. The automated intelligent agent system for transcriptomics regulatory analysis according to claim 1, characterized in that, The task execution mode includes any one of the following: execution mode and automatic mode; the automatic mode is to autonomously complete the task planning and obtain the task result; the execution mode is to complete the task planning and obtain the task result in combination with user instructions. Optionally, the user instructions include task adjustment instructions, which include one or more of the following: task modular deletion or addition instructions, task atomic deletion or addition instructions; the transcription regulation task results can be adjusted by dynamically controlling any task process or task stage in the task planning path through user instructions. Optionally, the task planning path consists of N task nodes and their relationships, where N is a natural number greater than 1. When a task adjustment instruction is obtained under the execution model, the current task execution node is interrupted, the adjusted task node is updated to form an updated task path, and the task node is judged. If the updated task node is before the current task node, the task is executed at the updated task node, and the task of the current task node is released. If the updated task node is after the current task node or a parallel task node, the current task node continues to execute and the task is executed based on the updated task path. Optionally, the task node is determined by obtaining the historical task path from the memory module, comparing the current task node with the adjusted task node position order to obtain the determination result, and storing the updated task path in the memory module after the system prompt module executes based on the determination result. Optionally, the task execution order or task module execution order in the task execution mode includes: parallel task execution, serial task execution, and a combination of parallel and serial tasks; the task execution order is formed by dynamic scheduling based on system resources and task completion progress; the task module execution order is formed by dynamic scheduling based on task node relationships. Optionally, the task node relationships are obtained based on the task planning path, including: sequential relationships, multi-module fusion relationships, and skip fusion relationships; Optionally, the mode execution process of the system prompt module is replaced by generating a discrete task node after obtaining the user's intent based on user messages or instructions or through question and answer in the planning mode. When the system prompt module switches to the task execution mode, it thinks based on the discrete task node and calls the tool to obtain a preliminary result. Based on the preliminary result, it conducts the next step of thinking and tool calling. The thinking process and tool calling are carried out synchronously.

3. The automated intelligent agent system for transcriptomics regulatory analysis according to claim 1, characterized in that, The system prompt module includes an enhanced context management module for managing context messages or instructions, which include one or more of the following: system messages or instructions, assistant messages or instructions, user messages or instructions; and recording of transcriptomic regulatory processes based on context messages or instructions. Optionally, the record is as follows: when the system prompt module generates a task planning path in planning mode, the context management module generates a structured task request instruction based on the task planning path. The structured task request instruction includes the thought process and application tools for building each task node, which is recorded as an assistant message or instruction; the task planning path is recorded as a system message or instruction; when the system prompt module is in task execution mode, the output of the execution result is based on the structured output result, and the standardized calling instruction and execution result are recorded as a user message or instruction. Optionally, the application tool that requires parameters when used also includes parameter data recorded as assistant messages or instructions; Optionally, the system further includes a correction tool, which is connected to the system prompt module and the function module. When it is detected that the output result of the task executed in the task execution mode is not a structured result, the result is converted into a structured result before further processing. Optionally, the system further includes an environmental information module, which is connected to the system prompt module. The environmental information module is used to monitor the system status in real time, including the current language, current working path, and current system time. Optionally, the environment information module further includes a custom language module, through which the model's response language type is defined; Optionally, when the system prompt module is a task execution module, it also includes environmental information acquisition. The environmental information module obtains detailed information about the current system environment, which is recorded as user messages or instructions. While the structured task request instruction is input to the functional module, the user messages or instructions and assistant messages or instructions are transmitted to the memory module for storage.

4. The automated intelligent agent system for transcriptomics regulatory analysis according to claim 1, characterized in that, The system prompt module also includes a memory control module, which controls whether system messages or instructions, user messages or instructions, and assistant messages or instructions are stored, and dynamically adjusts the length of the message or instruction context to obtain the specified context content. Optionally, the memory control module is connected to the correction tool. When the correction tool detects that the output result of the task executed in the task execution mode is not a structured result, the memory control module cancels the storage of the current output result. Optionally, the memory regulation module can be used to combine long and short memories to reduce the context overflow of the model.

5. The automated intelligent agent system for transcriptomics regulatory analysis according to claim 1, characterized in that, The memory module includes a fuzzy memory module and a precise memory module; the fuzzy memory module is used to store user messages or instructions or system messages or instructions; the precise memory module is used to store a complete list of messages when executing a task; Optionally, the fuzzy memory module stores the message or instruction of a task in the form of a memory list, and the precise memory module stores all the messages for executing the task in the form of memory retrieval values; Optionally, the memory module further includes a memory retrieval tool. When the system prompts the module to obtain a historical interactive query instruction, the memory retrieval tool in the memory module is used to query the list or search for values. Optionally, when the system prompt module obtains transcription-related data or similar instructions of the same type, in the planning mode, it calls the system messages and user messages in the fuzzy memory module to generate a task planning path, and in the task execution mode, it calls the user messages in the fuzzy memory module and enters the automatic mode to execute the task.

6. The automated intelligent agent system for transcriptomics regulatory analysis according to claim 1, characterized in that, The functional sub-modules include one or more of the following: a parameter adjustment module and a tooltip and definition module; the parameter adjustment module dynamically configures model and tool parameters through task request commands; the tooltip and definition module standardizes call commands or processes through task request commands; Optionally, the functional submodule calls the parameter adjustment module and the tooltip definition module based on the task request instruction. The parameter adjustment module configures the transcriptional regulation model and model parameters. The tooltip and definition module then generates standardized call instructions or processes. The standardized call instructions or processes call the transcriptional regulation service to execute the request, obtain the request result, and then send the request result back to the transcriptional regulation model. The model is then used to perform transcriptional regulation analysis to obtain the task result. Optionally, the transcriptional regulation service includes predefined transcriptional regulation query tools, including epigenetic annotation query, transcriptional regulation binding region query, and gene expression query. Optionally, the transcriptional regulation service is invoked via a message transmission protocol in the cloud. Optionally, the transcriptional regulation service also includes a transcriptional regulation software package, which requests the software package to be executed through standardized invocation instructions to obtain the software package analysis results, and then feeds the software package analysis results back into the transcriptional regulation model; Optionally, the transcriptional regulation service is virtualized as a whole through a Docker image to obtain L transcriptional regulation services, where L is a natural number greater than 1, and the L transcriptional regulation services independently execute the request instruction based on the calling instruction; Optionally, the functional submodule further includes a tool module, including one or more of the following: a remote invocation tool, a real-time display and interaction tool, a file upload tool, a file download tool, and a display tool; which performs remote invocation and / or display and / or upload and / or download of the task results or intermediate process results based on task request instructions.

7. The automated intelligent agent system for transcriptomics regulatory analysis according to claim 6, characterized in that, The parameter adjustment module also includes a custom parameter module. The custom parameter module is used to configure custom parameters and custom structures for the transcriptional analysis model, and to execute any transcriptional task. When the function module calls the parameter adjustment module, it generates a transcriptional regulation model and model parameters. The custom parameter module is used to make custom adjustments to the transcriptional regulation model and model parameters to obtain a custom transcriptional regulation model. Optionally, the tooltips and definition module includes a tool definition module, through which tools are customized. The customized tools consist of core function definitions and tooltips. Optionally, the tooltips and definition module also includes a CLI prompt module, which includes S definitions and invocation methods for individual or integrated transcriptional regulation tools, where S is a natural number greater than or equal to 1; the CLI prompt module generates standardized invocation instructions or procedures, and then the transcriptional regulation service is invoked based on the standardized invocation instructions or procedures; Optionally, the CLI prompt module is connected to the tool definition module. After the tool definition module builds a custom tool, it is input to the CLI prompt module, which then generates standardized call instructions or processes.

8. The automated intelligent agent system for transcriptomics regulatory analysis according to claim 1, characterized in that, The functional submodule also includes a genome regulatory element annotation tool, which includes enhancer annotation data, genetic variation annotation data, chromatin openness annotation data, three-dimensional genome data, and DNA methylation data; Optionally, the enhancer annotation data includes one or more of the following: SEdb database, SEA database, dbSuper database, EnhancerAtlas database, HACER database, ENCODE database, FANTOM5 database, DENDB database, ENdb database, and eRNAbase database; Optionally, the genetic variation annotation data includes one or more of the following: dbSNP database, GWAS Catalog database, GWASdb database, GTEX database, PancanQTL database, seeQTL database, SCAN database, and Oncobase database; Optionally, the chromatin openness annotation data includes one or more of the following: ATACdb database, ENCODE database; Optionally, the three-dimensional genome data includes one or more of the following: the 4DGenome database, the Oncobase database, and the 3D Genome Browser database; Optionally, the DNA methylation data includes one or more of the following: the ENCODE database; Optionally, the functional submodule also includes a target gene selection module. When the task node of the execution task path is target gene selection, the target gene selection module is called. The algorithm identifies the association between DNA regulatory elements and target genes to obtain the association relationship. Gene expression profile data is called, and target gene selection is performed based on the association relationship and gene expression profile data.

9. A computer device comprising a memory, a processor, and a computer program or instructions stored in the memory, characterized in that, The computer program or instructions are executed by the processor to implement the module steps of the intelligent agent automated system for transcriptomics regulatory analysis as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, The computer program or instructions are executed by the processor to implement the module steps of the intelligent agent automated system for transcriptomics regulatory analysis as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Analysis method of spatial transcriptome sequencing data

    CN112522371A

  • Data processing method and system for identifying enhancer and super enhancer

    CN115083517A

  • Single cell transcriptome analysis system

    CN118173180A

  • Artificial intelligence systems and methods for enabling natural language transcriptomics analysis

    US20250139386A1