Biological information drawing system based on multiple agents

Through multi-agent system, combined with natural language processing and large models, the problem of low efficiency of biological information drawing tools is solved, and efficient and professional chart generation and data analysis services are achieved.

CN120339451AInactive Publication Date: 2025-07-18BEIJING NOVOGENE TECH CO LTD

Patent Information

Application Number
CN202510799711.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing bioinformatics drawing tools have problems such as low efficiency, complex parameter adjustment, tool independence, and lack of real-time interaction and guidance.

Method used

A biological information drawing system based on multi-agents is adopted, including supervisory agents, drawing agents, customer service agents and data analysis agents. Through natural language processing and large models, user-friendly chart generation and data analysis are achieved.

Benefits of technology

It improves the efficiency and quality of bioinformatics drawing, simplifies the operation process, lowers the threshold for use, provides professional data analysis and information consulting services, and meets the personalized needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339451A_ABST
    Figure CN120339451A_ABST
Patent Text Reader

Abstract

The invention discloses a biological information drawing system based on multiple agents. Relates to the field of artificial intelligence, and comprises a supervisor agent used for receiving a drawing request of a user, analyzing semantics of the drawing request, generating a drawing task and sending the drawing task to a drawing agent; the drawing agent comprises a biological information drawing library and is used for analyzing semantics of the drawing task, determining a biological information drawing template corresponding to the semantics from the biological information drawing library, generating a drawing code based on the semantics and the biological information drawing template and executing the drawing code to obtain a target drawing, the biological information drawing library comprises a plurality of biological information drawing templates. Through the method and the device, the problem of low biological information drawing efficiency in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and more particularly, to a multi-agent based biological information drawing system. Background Art

[0002] In the field of bioinformatics, the results of data analysis need to be presented and interpreted in a graphical way. This not only requires the graph to accurately reflect the characteristics of the data and the analysis results, but also requires the chart to be visually attractive enough for review and academic publication. The biological information drawing methods in the related art, such as directly using R language or Python to write scripts, require users to have certain programming knowledge and bioinformatics professional skills, which results in a relatively high threshold for many non-professional background researchers in drawing.

[0003] A variety of online drawing tools in the related art provide certain convenience for biological information drawing. These tools preset the drawing parameters, and users only need to upload data to generate charts, which simplifies the drawing process. However, these tools have serious limitations. First, the presetting of drawing parameters is too simple to meet the adjustment needs of professional users for the details of complex charts; second, different drawing tools are independent of each other, lacking unified management and data processing capabilities, which makes it necessary for users to switch between different tools when processing multiple charts, and the operation is cumbersome; furthermore, these tools lack real-time interaction and guidance after drawing. When facing problems in the drawing results or the need for further customization and optimization, users have nowhere to start and need to seek the help of professional technical personnel, which affects the efficiency and quality of drawing.

[0004] Aiming at the problem of low efficiency of biological information drawing in the related art, no effective solution has been proposed yet. Summary of the Invention

[0005] The main purpose of this application is to provide a multi-agent based biological information drawing system to solve the problem of low efficiency of biological information drawing in the related art.

[0006] To achieve the above purpose, according to one aspect of this application, a multi-agent based biological information drawing system is provided. The system includes: a supervisor agent, configured to receive a drawing request from a user, analyze the semantics of the drawing request, generate a drawing task, and send the drawing task to a drawing agent; the drawing agent includes a biological information drawing library, configured to analyze the semantics of the drawing task, determine a biological information drawing template corresponding to the semantics from the biological information drawing library, generate drawing code based on the semantics and the biological information drawing template, and execute the drawing code to obtain a target drawing, wherein the biological information drawing library contains a variety of biological information drawing templates.

[0007] Optionally, the drawing agent generates a target drawing in the following manner: analyze the drawing task to obtain drawing data and a drawing type; call the data analysis agent to process the drawing data to obtain a target data analysis result, where the target data analysis result includes at least one of the following: difference analysis result, correlation analysis result, and data statistics result; determine a target bioinformatics drawing template corresponding to the drawing type, and generate drawing code based on the target bioinformatics drawing template, the drawing data, and the target data analysis result, where the drawing code includes at least one of the following: data code, code comment, legend explanation, and parameter setting; run the drawing code through a code execution environment to obtain the target drawing.

[0008] Optionally, the drawing agent is further configured to: receive a code modification instruction from the user, analyze the code modification instruction to obtain modified target code, where the code modification instruction is used to indicate modifying the target drawing; replace the code to be modified in the drawing code with the modified target code to obtain updated drawing code; run the updated drawing code through the code execution environment to obtain an updated target drawing.

[0009] Optionally, the system further includes: a customer service agent, including a bioinformatics knowledge base, configured to receive an information consultation task issued by the supervisor agent, retrieve bioinformatics knowledge associated with the information consultation task from the bioinformatics knowledge base through a retrieval-augmented generation model, and determine the bioinformatics knowledge as the consultation result.

[0010] Optionally, the customer service agent generates the consultation result in the following manner: extract a consultation keyword from the information consultation task, and retrieve a target document associated with the consultation keyword from the bioinformatics knowledge base through the retrieval-augmented generation model; extract associated data corresponding to the consultation requirement in the information consultation task from the target document; generate a consultation result based on the information consultation task and the associated data.

[0011] Optionally, the system further includes: a data analysis agent, including a variety of preset data analysis tools, configured to receive a data analysis task issued by the supervisor agent, call the data analysis tools to perform data processing operations on the data in the data analysis task, and generate a data analysis result.

[0012] Optionally, the data analysis agent generates the data analysis result in the following manner: extract the data to be processed in the data analysis task, preprocess the data to be processed to obtain preprocessed data, where the preprocessing includes at least one of the following: data format and data type recognition and missing value processing; call a target data analysis tool based on the task requirement in the data analysis task to perform a target operation on the preprocessed data to obtain a data analysis result, where the target operation includes at least one of the following: data statistics, difference analysis, correlation analysis, and format conversion.

[0013] Optionally, the biological information knowledge base is obtained in the following manner: obtaining biological information knowledge from multiple data sources, where the biological information knowledge at least includes: bioinformatics knowledge, biological information drawing data, and operation documents of data analysis tools; converting the biological information knowledge into biological information vectors, and constructing a knowledge graph based on the biological information vectors; constructing a document index for each biological information vector, and forming the biological information knowledge base from the knowledge graph and all the document indexes.

[0014] Optionally, the biological information drawing library is obtained in the following manner: obtaining multiple historical biological information drawings, and extracting drawing types from the multiple historical biological information drawings, where the drawing types at least include: volcano plots, heatmaps, and bubble plots; determining a target drawing package, and constructing a basic drawing template for each drawing type through the target drawing package, where the basic drawing template includes at least one of the following: axis setting, color mapping, legend, and text format; configuring parameters for each basic drawing template based on the historical biological information drawings to obtain multiple biological information drawing templates, and constructing the biological information drawing library from the multiple biological information drawing templates, where the parameters include at least one of the following: axis range and scale, color scheme, title font size, and legend position.

[0015] Optionally, the supervisor agent is trained in the following manner: obtaining multiple historical requests and the historical tasks corresponding to each historical request, where the historical tasks include at least one of the following: historical drawing tasks, historical information consultation tasks, and historical data analysis tasks; determining each historical request and the historical task corresponding to the historical request as a set of training samples to obtain multiple sets of training samples; training a large language model through the multiple sets of training samples to obtain the supervisor agent.

[0016] Through the present application, the following system is adopted: a supervisor agent, which is used to receive a user's drawing request, analyze the semantics of the drawing request, generate a drawing task, and send the drawing task to a drawing agent; the drawing agent, which includes a biological information drawing library, is used to analyze the semantics of the drawing task, determine a biological information drawing template corresponding to the semantics from the biological information drawing library, generate drawing code based on the semantics and the biological information drawing template, and execute the drawing code to obtain a target drawing. Among them, the biological information drawing library contains multiple biological information drawing templates, which solves the problem of low efficiency of biological information drawing in the related art. Through the supervisor agent and the drawing agent, it is realized that the user only needs to communicate with the supervisor agent, the supervisor agent generates a drawing task, and the drawing agent can output a target drawing based on the drawing task, thereby achieving the effect of improving the efficiency of biological information drawing. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings constituting a part of this application are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0018] Figure 1 is a schematic diagram of a multi-agent based biological information mapping system provided according to an embodiment of the present application;

[0019] Figure 2 is a schematic diagram of the application process of a multi-agent based biological information mapping system provided according to an embodiment of the present application. Detailed implementation manners

[0020] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0021] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so as to implement the embodiments of the present application described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0023] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.

[0024] It should be noted that the collected information is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application, etc. of the relevant data all comply with the relevant laws, regulations and standards of the relevant regions, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0025] The present invention will be described below in conjunction with preferred implementation steps. Figure 1 It is a schematic diagram of a multi-agent-based biological information drawing system provided according to an embodiment of the present application. As Figure 1 shown, the system includes:

[0026] A supervisor agent, configured to receive a drawing request from a user, analyze the semantics of the drawing request, generate a drawing task, and send the drawing task to a drawing agent.

[0027] In some embodiments, the supervisor agent serves as an entry point for the user to interact with the entire multi-agent-based biological information drawing system. According to the user's conversation content and requirements, it intelligently judges and selects a suitable agent to process the task, coordinates the workflow between agents, and realizes the efficient allocation and collaboration of tasks. The supervisor agent adopts natural language processing technology based on the first large language model to deeply understand the semantics and recognize the intentions of the user's input. Through training on a large amount of text data, it can accurately grasp the key information in the user's language and judge whether the user needs drawing, data analysis, tool usage consultation, or other requirements. For example, when the user sends "Please help me draw a bubble chart to show the gene enrichment analysis results, with the gene names on the horizontal axis, the enrichment degree on the vertical axis, and the color representing the p-value", the supervisor agent not only recognizes the drawing requirement and the specific chart type and data usage, but also understands the various dimensions and parameter requirements of the chart, and then assigns the drawing task to the drawing agent; if the user asks "What types of biological information knowledge analysis can this tool be used for? Can you introduce the application scenarios of each analysis in detail?", the supervisor agent will transfer the task to the customer service agent. During the task allocation process, the supervisor agent also tracks the working status of each agent to ensure the smooth execution and timely feedback of the task.

[0028] The multi-agent-based biological information drawing system of this embodiment adopts a multi-agent architecture design and collaborative working mechanism, combines the natural language processing ability of the large model, and realizes natural and efficient interaction between the user and the supervisor agent, as well as intelligent task allocation and efficient collaboration. The user no longer needs to interact with multiple different tools or modules separately. Just communicate with the supervisor agent in natural language to complete all operations such as drawing, data analysis, and question consultation, greatly simplifying the operation process, reducing the usage threshold, and improving work efficiency.

[0029] A drawing agent, including a biological information drawing library, configured to analyze the semantics of the drawing task, determine a biological information drawing template corresponding to the semantics from the biological information drawing library, generate drawing code based on the semantics and the biological information drawing template, and execute the drawing code to obtain a target drawing, where the biological information drawing library contains multiple biological information drawing templates.

[0030] In some embodiments, the drawing agent can use the R language (a programming language for statistical analysis and graphical representation) for drawing. The drawing agent utilizes the large model's understanding of drawing requirements and code generation capabilities, and has built-in various professional basic codes for bioinformatics drawing, such as enrichment analysis bubble charts, volcano plots, heatmaps, violin plots, etc., which are stored in the bioinformatics drawing library. This agent is bound with a structured output function. In addition to generating charts, it can also output the corresponding R language code (i.e., drawing code), which is convenient for users to view, learn, and modify. At the same time, it provides a legend list to explain in detail the meaning of the legends in the figure, as well as a parameter list covering adjustable drawing parameters such as fonts and titles. Users can further modify and improve the chart according to their needs, and this agent can execute R code to achieve dynamic drawing and chart modification operations, supporting natural language chart modification. The agent generates code directly for rendering after understanding the language. The display of the target drawing can include pictures (preview and download), legend explanations, parameter displays, etc., which are clear and straightforward. The chart modification is directly described in natural language, and the operation is convenient.

[0031] It should be noted that the bioinformatics drawing templates of the drawing agent can be continuously expanded and updated. According to the latest research trends in the bioinformatics field and user feedback, new chart types and drawing styles can be added. At the same time, interfaces can also be provided to allow users to customize drawing templates and integrate them into the bioinformatics drawing library.

[0032] In addition to the bioinformatics field, the multi-agent-based bioinformatics drawing system of this embodiment can also be applied to other fields that require data processing, drawing, and customer service support, such as financial data analysis and visualization, social science data research, medical data analysis, etc. In the financial field, it can help analysts quickly draw various financial data charts, such as stock trend charts, value-at-risk charts, etc., and support natural language chart modification and data analysis consulting; in the social science field, it can be used for drawing demographic charts, visualization of survey results, etc.; in the medical field, it can assist doctors and researchers in drawing patient data charts, disease distribution charts, etc. By adjusting the functions of each agent and the content of the knowledge base, it can quickly adapt to the application requirements of different fields and provide intelligent solutions for data visualization and analysis work in various fields.

[0033] The bioinformatics plotting system based on multi-agent provided by the embodiments of the present application uses a supervisor agent to receive a plotting request from a user, analyze the semantics of the plotting request, generate a plotting task, and send the plotting task to a plotting agent; the plotting agent includes a bioinformatics plotting library, which is used to analyze the semantics of the plotting task, determine a bioinformatics plotting template corresponding to the semantics from the bioinformatics plotting library, generate plotting code based on the semantics and the bioinformatics plotting template, and execute the plotting code to obtain a target plot. Among them, the bioinformatics plotting library contains a variety of bioinformatics plotting templates, which solves the problem of low bioinformatics plotting efficiency in the related art. Through the supervisor agent and the plotting agent, it is realized that the user only needs to communicate with the supervisor agent. The supervisor agent generates a plotting task, and the plotting agent can output a target plot based on the plotting task, thereby achieving the effect of improving the bioinformatics plotting efficiency.

[0034] After the supervisor agent issues a plotting task, the plotting agent generates a target plot based on the plotting task. Optionally, in the bioinformatics plotting system based on multi-agent provided by the embodiments of the present application, the plotting agent generates a target plot in the following manner: analyze the plotting task to obtain plotting data and a plotting type; call a data analysis agent to process the plotting data to obtain a target data analysis result, where the target data analysis result includes at least one of the following: a differential analysis result, a correlation analysis result, and a data statistics result; determine a target bioinformatics plotting template corresponding to the plotting type, and generate plotting code based on the target bioinformatics plotting template, the plotting data, and the target data analysis result, where the plotting code includes at least one of the following: data code, code comments, legend explanations, and parameter settings; run the plotting code through a code execution environment to obtain a target plot.

[0035] In some embodiments, a series of professional plotting templates and functions for the bioinformatics field are developed based on the R language plotting package system, such as ggplot2 (a plotting system), pheatmap (a heatmap plotting tool), etc. Each bioinformatics plotting template is carefully designed and optimized to meet the specification requirements of academic papers for charts, including aspects such as axis settings, label formats, and color combinations. And a large number of carefully optimized prompt words are set in the template to facilitate the large model to generate the code required for plotting according to the user's semantics. During the plotting process, the plotting agent automatically calls the corresponding target bioinformatics plotting template according to the input data and the selected chart type in the plotting task, that is, the plotting data and the plotting type, and combines the target data analysis result provided by the data analysis agent to perform the plotting.

[0036] Meanwhile, using the code generation and annotation functions of R language, it generates plot codes with strong readability and easy modification, and outputs content such as legend explanations and parameter settings in a structured manner, facilitating users to understand and further customize the charts. In addition, this intelligent agent also has a code execution environment where users can directly run the modified code in the software to view the plot effects in real time, improving the efficiency and flexibility of plotting. It also supports modifying plots with natural language. After the plotting intelligent agent understands the natural language instructions of users, it generates corresponding R code for direct rendering, simplifying the plot modification process and improving the convenience and efficiency of plot modification.

[0037] The multiple professional bioinformatics plotting templates built into the plotting intelligent agent meet publication standards. At the same time, it provides plot codes, legends, and parameters with structured output. Users can easily make personalized modifications and customizations, which can meet the diverse needs of different users for chart styles and content, with higher plotting quality and stronger flexibility. It also supports modifying plots with natural language, further facilitating users to adjust and optimize the charts.

[0038] The plotting intelligent agent in this embodiment provides bioinformatics plotting template design, structured output functions (including various R basic reference codes with prompts, legend lists, parameter lists), and dynamic plotting and plot modification capabilities. The underlying reference codes of the bioinformatics plotting templates cover common bioinformatics plotting types (such as enrichment analysis bubble charts, volcano plots, heatmaps, etc.), and common modification requirements are included in the design, which better meets the bioinformatics professional plot modification needs. It has carefully tuned prompts built in, facilitating the large model to easily understand the requirements and convert them into executable codes, and the one-time execution success rate of the codes is extremely high. The reference codes, configuration files, example data, and example pictures are all imported through configuration tables, facilitating the maintenance, update, and optimization of the codes. The close cooperation between the data analysis intelligent agent and the plotting intelligent agent enables the data processing results to be directly used for high-quality plotting, avoiding the cumbersome conversion and manual processing of data between different tools, improving data accuracy, and shortening the cycle from data to charts.

[0039] The plotting intelligent agent can also provide a function for efficiently modifying the target plot. Optionally, in the multi-agent-based bioinformatics plotting system provided in the embodiments of this application, the plotting intelligent agent is further configured to: receive a code modification instruction from a user, analyze the code modification instruction to obtain a modified target code, where the code modification instruction is used to indicate modifying the target plot; replace the code to be modified in the plot code with the modified target code to obtain an updated plot code; and run the updated plot code through the code execution environment to obtain an updated target plot.

[0040] In some embodiments, after a user views the initial target drawing generated by the drawing agent, they may find that some details need to be adjusted, such as changing the size and color of points, the axis range, the title font, etc. The user only needs to describe these modification requirements in natural language as code modification instructions. For example: "Please increase the size of the points of differentially expressed genes in the volcano plot." The drawing agent passes the user's code modification instructions to the second large language model for in-depth analysis. The model converts this natural language instruction into a code modification instruction that can be understood by a computer. For example, for the above instruction "Please increase the size of the points of differentially expressed genes in the volcano plot", the model will parse out the parameters related to the point size in the drawing code that need to be modified, such as the size attribute, and determine that the object to be modified is the points of differentially expressed genes in the volcano plot.

[0041] Based on the results of the analysis by the second large language model, the drawing agent determines the part of the code that needs to be modified and intelligently generates the modified target code. For example, it may generate the following modified code: "Increase the point size of differentially expressed genes in the volcano plot from the original 3 to 5." The drawing agent applies the modification instruction to the original drawing code, replaces the original part with the modified parameters, and generates the updated drawing code. After ensuring the integrity and compatibility of the drawing code, the drawing agent submits it to the code execution environment, such as the drawing environment of the R language. The environment runs the updated drawing code, re-renders the chart according to the new code parameters, and generates the modified target drawing. The user can immediately view the modification effect on the user interface of the drawing agent. If further adjustment is needed, new modification requirements can be issued through natural language instructions.

[0042] The drawing agent in this embodiment, based on the natural language drawing modification function of the large model, generates code incremental modifications and directly renders them by understanding the user's natural language instructions, realizing convenient and efficient chart modification, and meeting the user's requirements for high-quality bioinformatics charts and personalized modification requirements.

[0043] In addition to performing bioinformatics drawing, the multi-agent based bioinformatics drawing system can also provide information consultation services. Optionally, in the multi-agent based bioinformatics drawing system provided in the embodiments of the present application, the system further includes: a customer service agent, including a bioinformatics knowledge base, for receiving information consultation tasks sent by the supervisor agent, retrieving bioinformatics knowledge associated with the information consultation tasks from the bioinformatics knowledge base through a retrieval augmented generation model, and determining the bioinformatics knowledge as the consultation result.

[0044] In some embodiments, by leveraging the semantic understanding and knowledge integration capabilities of large models, a biological information knowledge base is constructed using the RAG (Retrieval-Augmented Generation) technology. It can automatically identify users' questions and retrieve relevant information from the knowledge base to provide answers, offering users a fast and accurate information consultation support. The biological information knowledge base contains a large amount of domain expertise, including but not limited to: the usage methods and tutorials of various bioinformatics tools and software; common bioinformatics data analysis processes and problem-solving solutions; the explanations of biological and medical terms, as well as relevant research progress; detailed guides on chart types, styles, and application scenarios.

[0045] When the customer service agent receives the information consultation task forwarded by the supervisor agent, it first searches in the biological information knowledge base using the query vector to find the document fragments most relevant to the user's question. The RAG model will further analyze the retrieved document fragments, extract key information, and generate a concise and easy-to-understand answer summary to avoid interference from lengthy and irrelevant content. Finally, the model will synthesize the user's question context and the extracted biological information knowledge to generate a natural, fluent, detailed, and accurate consultation result to answer the user's doubts.

[0046] For example, if the user asks "How to display the clustering results of genes in a heatmap?", the RAG model of the customer service agent will retrieve the usage tutorial of the pheatmap package from the knowledge base, especially the part about how to integrate dendrograms and clustering information in the heatmap, and then generate a specific guiding answer such as: "When using the pheatper::pheatmap() function, by setting the cluster_rows=TRUE and cluster_cols=TRUE parameters, you can display the clustering results of rows and columns in the heatmap. In addition, you can control the display of row and column names by adjusting show_rownames and show_colnames, and add additional row and column annotations using annotation_row and annotation_col to help explain the meaning of the clustering."

[0047] The customer service agent in this embodiment provides instant and professional bioinformatics consulting services. By combining the RAG model with the bioinformatics knowledge base, the customer service agent can quickly locate problems, accurately extract relevant knowledge, and generate answers in natural language, significantly improving the efficiency and satisfaction of users in obtaining information. The customer service agent can respond to users' consultations immediately, avoiding the time cost of waiting for manual replies. The bioinformatics knowledge base covers multiple subfields of bioinformatics and can provide users with comprehensive and in-depth professional answers. Using the RAG model, the customer service agent can not only intelligently retrieve relevant knowledge but also automatically generate personalized answers according to specific situations, making the consultation results more in line with users' needs. Through the establishment of the customer service agent, the multi-agent-based bioinformatics drawing system provides a comprehensive and integrated drawing solution for bioinformatics researchers, not only solving the problems of drawing quality and efficiency but also overcoming the challenges in data processing and professional consulting, greatly improving the efficiency and quality of scientific research work.

[0048] After receiving an information consultation task, the customer service agent automatically generates a consultation result based on the information consultation task. Optionally, in the multi-agent-based bioinformatics drawing system provided by the embodiments of the present application, the customer service agent generates a consultation result in the following manner: extracting consultation keywords in the information consultation task, and retrieving target documents associated with the consultation keywords from the bioinformatics knowledge base through a retrieval-augmented generation model; extracting associated data corresponding to the consultation requirements in the information consultation task from the target documents; generating a consultation result based on the information consultation task and the associated data.

[0049] In some embodiments, when the customer service agent receives an information consultation task forwarded by the supervisor agent, it first uses natural language processing techniques to extract the consultation keywords in the task. These keywords may be specific tool names, chart types, data analysis methods, or biological concepts, etc., which constitute the core elements of the user's consultation. For example, from the consultation "I want to know how to use ggplot2 to draw a high-quality volcano plot, especially how to adjust the color and size of the points.", the customer service agent will identify keywords such as "ggplot2", "volcano plot", "color", and "size". After extracting the keywords, the customer service agent will use the retrieval-augmented generation model to search for information such as literature, tutorials, and technical documents related to these keywords in the bioinformatics knowledge base. The RAG model can effectively understand the context and intention of the keywords and screen out the most relevant documents from the vast knowledge base. Taking the adjustment of the volcano plot as an example, the model will find a detailed tutorial on ggplot2 volcano plot drawing and parameter adjustment and use it as the target document.

[0050] From the target document, the customer service agent will further extract data and information directly related to the user's consultation needs. For example, specific drawing code snippets, parameter setting instructions, example datasets, or chart effect examples, etc. In the case of the volcano plot, the customer service agent will find code examples for adjusting the point color and size, as well as the detailed steps on how to implement these adjustments in ggplot2. Finally, the customer service agent inputs the original text of the information consultation task and the associated data extracted from the target document into the first large language model, and requests the model to synthesize a complete consultation result tailored to the user's needs. The first large language model has powerful text generation capabilities and can output the consultation result in a smooth and natural language form based on the input context and data, not only answering the user's specific questions, but also providing additional relevant information and suggestions to help the user better understand and apply.

[0051] The customer service agent utilizes the semantic understanding and knowledge generation capabilities of the large model, adopts the RAG technology, and combines document retrieval. The document retrieval part contains a large number of text materials such as tool usage documents, frequently asked questions, and function descriptions. Through in-depth semantic analysis of these documents, the large model quickly locates the document content related to the user's question and provides answers. For example, when the user asks "How to adjust the size and color of the points in the volcano plot to more intuitively display the significance of gene expression differences", the customer service agent, relying on the understanding ability of the large model, retrieves the drawing parameter part of the volcano plot from the knowledge base through retrieval, combines the specific descriptions in the document, gives detailed adjustment methods and example code snippets, and can further explain the impact of different parameter settings on the visual effect of the chart to help the user better understand and apply.

[0052] The customer service agent in this embodiment enables the multi-agent-based bioinformatics drawing system to not only provide graphic drawing services, but also become a fast query and learning platform for bioinformatics knowledge, solving many knowledge-related problems for users beyond drawing, thus comprehensively improving the efficiency and the quality of the results of bioinformatics research.

[0053] In addition to performing bioinformatics drawing, the multi-agent-based bioinformatics drawing system can also provide data analysis services. Optionally, in the multi-agent-based bioinformatics drawing system provided in the embodiments of the present application, the system further includes: a data analysis agent, including a variety of preset data analysis tools, used to receive data analysis tasks issued by the supervisor agent, call the data analysis tools to perform data processing operations on the data in the data analysis tasks, and generate data analysis results.

[0054] In some embodiments, the data analysis agent is equipped with a variety of pre-set data analysis tools, such as Python (a programming language) tools, which have powerful data analysis capabilities. When the user needs to analyze the raw data, the supervisor agent assigns the data analysis task to the data analysis agent, which can use various data analysis libraries of Python (such as Pandas (a Python library for data manipulation and analysis), NumPy (Numerical Python, a library for numerical computing), etc.) to perform operations such as data preprocessing and statistical analysis on the data. The processed results can be directly transmitted to the graph agent for plotting, ensuring the accuracy and usability of the data. The data analysis agent understands and optimizes the processing strategy for the data analysis task with the assistance of the large model. Between the data analysis agent and the graph agent, through data transmission and format conversion methods, such as the interface design for data interaction and data conversion rules, including the specific steps and methods of how the large model participates in data analysis strategy adjustment and result optimization. Ensure the seamless connection and efficient utilization of data during the analysis and plotting processes, and at the same time, rely on the large model to provide deeper data insights. In addition to using Python tools, the data analysis agent can also integrate other data analysis languages or tools (such as some data analysis packages of the R language) as alternative options to meet the needs of specific users or handle certain complex data analysis scenarios.

[0055] The data analysis agent in this embodiment can automatically parse the task requirements, select appropriate data processing tools, and reduce human errors and time consumption. By integrating data analysis tools, the agent can perform complex statistical modeling and machine learning tasks, providing in-depth data insights for bioinformatics research. The data exchange between agents follows a predefined format, ensuring the consistency and operability of the data during processing at each stage, and accelerating the conversion process from data to graphics. The data analysis agent not only provides pure data results but also can generate preliminary graphs and tables, providing intuitive guidance for the graphing tasks of the graph agent.

[0056] After receiving the data analysis task, the data analysis agent processes the data in the data analysis task. Optionally, in the multi-agent-based bioinformatics graphing system provided in the embodiments of the present application, the data analysis agent generates the data analysis result in the following manner: extracting the data to be processed in the data analysis task, preprocessing the data to be processed to obtain the preprocessed data, where the preprocessing includes at least one of the following: data format and data type recognition and missing value processing; calling a target data analysis tool based on the task requirements in the data analysis task to perform target operations on the preprocessed data to obtain the data analysis result, where the target operations include at least one of the following: data statistics, differential analysis, correlation analysis, and format conversion.

[0057] In some embodiments, the data analysis agent takes advantage of the data analysis ecosystem of Python and integrates Pandas for data cleaning, transformation, and organization, NumPy for numerical calculations, SciPy (a library for scientific computing), and Statsmodels (a library for statistical modeling, data analysis, and the design of econometric applications) for statistical analysis. When receiving a data analysis task assigned by the supervisor agent, it first reads and preliminarily checks the raw data to identify basic information such as the data format, missing value situation, and data type. Then, according to specific analysis requirements, appropriate target data analysis tools are used for processing, such as performing differential analysis and correlation analysis on gene expression data. During the data processing, intermediate results and statistical metrics are generated and organized into a format suitable for plotting, and data transmission is carried out through the interface with the plotting agent to ensure the accurate transfer and efficient utilization of data.

[0058] The data analysis agent in this embodiment adopts the above highly automated and intelligent data analysis process, significantly reducing the user's data processing burden and accelerating the conversion from raw data to meaningful analysis results, providing strong support for bioinformatics research. The close cooperation between the data analysis agent and other agents in the system ensures the coherence of data and the efficiency of analysis work.

[0059] To provide efficient information consulting services, a bioinformatics knowledge base needs to be constructed. Optionally, in the multi-agent-based bioinformatics plotting system provided in the embodiments of the present application, the bioinformatics knowledge base is obtained in the following manner: Obtain bioinformatics knowledge from multiple data sources, where the bioinformatics knowledge at least includes: bioinformatics knowledge, bioinformatics plotting data, and operation documents of data analysis tools; convert the bioinformatics knowledge into bioinformatics vectors, and construct a knowledge graph based on the bioinformatics vectors; construct a document index for each bioinformatics vector, and the bioinformatics knowledge base is composed of the knowledge graph and all document indexes.

[0060] In some embodiments, diverse bioinformatics knowledge sources are collected. These data sources can cover all aspects of bioinformatics, including but not limited to: Bioinformatics knowledge: basic bioinformatics concepts, terms, biological pathways, gene function annotations, etc. Bioinformatics plotting data: typical plotting cases, plotting techniques, parameter setting guides, legend explanations, etc. Operation documents of data analysis tools: user manuals, interface descriptions, script examples, algorithm principles, etc. of commonly used bioinformatics software. The data sources can come from public bioinformatics databases, scientific literature, online tutorials, forum discussions, and contributions from professional communities, ensuring that the knowledge base contains the latest biological knowledge and industry best practices.

[0061] After acquiring biological information knowledge, the unstructured text data is converted into a structured representation that can be understood and manipulated by a computer, namely a biological information vector. Natural language processing techniques such as word embedding, TF-IDF (Term Frequency-Inverse Document Frequency), etc. can be used to transform the words in the text into vectors. Through named entity recognition and dependency analysis, biological entities in the text (such as genes, proteins, disease names) and their relationships are identified, and this information will be used to construct a knowledge graph. The entity vectors, relationship vectors, and context vectors are fused to form a composite biological information vector to comprehensively reflect the semantic content of the data. Based on the biological information vector, a knowledge graph is constructed. A knowledge graph is a structured data representation method that organizes entities, attributes, and relationships in the form of nodes and edges to form a network structure. In a biological information knowledge base, the knowledge graph can: express complex relationships between entities, such as genes and diseases, proteins and proteins, molecular pathways and disease progression, etc.; quickly retrieve knowledge related to specific entities, improving the accuracy and efficiency of the agent's answer to questions; support relationship-based reasoning, and through the graph traversal algorithm, the agent can infer implicit knowledge that is not explicitly stated.

[0062] In addition to the knowledge graph, a document index also needs to be constructed for each biological information vector. A document index refers to marking and indexing the content of the original document to facilitate the rapid location and extraction of relevant document fragments. When constructing a biological information knowledge base, the role of the document index is as follows: accelerating the search speed, even when faced with a large volume of documents, the agent can quickly find the most relevant part of the document for the user's query; improving the accuracy of the search, by analyzing the similarity between the vector and the index, the agent can filter out the most appropriate answer candidates and reduce the interference of irrelevant information; supporting context understanding, through the analysis of the context of the document, the agent can better understand the background and intention of the user's query and provide a more accurate response.

[0063] The final biological information knowledge base consists of a knowledge graph and all document indexes. These two major components complement each other, enabling the agent to quickly respond to user queries: whether the user is asking about the function of a specific gene or seeking a guide to specific drawing skills, the agent can quickly retrieve the most relevant information and document fragments from the knowledge base; providing in-depth interpretation and suggestions: based on the knowledge graph and document indexes, the agent can not only give surface answers, but also deeply analyze the biological meaning behind the question or put forward professional drawing and data analysis suggestions.

[0064] This embodiment constructs a biological information knowledge base, which provides data support for each agent and ensures that the system can provide professional, comprehensive, and personalized bioinformatics services and support.

[0065] To provide efficient plotting services, a bioinformatics plotting library needs to be constructed. Optionally, in the multi-agent based bioinformatics plotting system provided by the embodiments of the present application, the bioinformatics plotting library is obtained in the following manner: Obtain a variety of historical bioinformatics plots, and extract plot types from the variety of historical bioinformatics plots. Among them, the plot types at least include: volcano plots, heatmaps, and bubble plots; Determine a target plotting package, and construct a basic plotting template for each plot type through the target plotting package. Among them, the basic plotting template includes at least one of the following: axis setting, color mapping, legend, and text format; Configure parameters for each basic plotting template based on the historical bioinformatics plots to obtain a variety of bioinformatics plotting templates, and construct a bioinformatics plotting library from the variety of bioinformatics plotting templates. Among them, the parameters include at least one of the following: axis range and scale, color scheme, title font size, and legend position.

[0066] In some embodiments, a large number of historical bioinformatics plots are collected. These historical bioinformatics plots can come from research reports in bioinformatics, academic conferences, professional journals, and open access databases, covering various types of bioinformatics knowledge and their visualization forms. These historical bioinformatics plots can also be reference codes with professional annotations written by senior bioinformatics analysis engineers according to business scenarios and customer requirements. The types of graphics cover common bioinformatics plotting types, such as: volcano plots, heatmaps, bubble plots, etc.; The built-in parameters of the code are comprehensive, including: professional bioinformatics data preprocessing methods and plotting package selection, graphic style settings (fonts, color schemes, axes, etc.). From the historical bioinformatics plots, the system will automatically identify and extract different plot types. Such as volcano plots, heatmaps, and bubble plots, etc. Volcano plots are used to show the significance of gene expression differences. Heatmaps are used for the visualization of expression matrices, which can clearly show the similarities and differences in gene or protein expression among different samples. Bubble plots are used to show high-dimensional data, such as gene enrichment levels, p-values, and gene counts, and the size and color can represent data in different dimensions.

[0067] After determining the target plotting package, the plotting agent will use these packages to construct basic templates for the corresponding plot types. The target plotting package can select professional packages in the R language that are most suitable for bioinformatics plotting, such as ggplot2, pheatmap, etc. The design of the basic plotting template can include: Axis setting: including the range, scale, labels, and direction of the axes to ensure the accuracy and readability of the chart. Color mapping: Assign specific colors to different data levels or categories to enhance the visual effect of the chart and data differentiation. Legend and text format: The layout, style, and content of the legend, as well as the font, size, and alignment of the title, labels, and other text elements, making the chart information clearly and professionally conveyed.

[0068] Based on the actual experience of historical bioinformatics drawing, each basic drawing template is parameterized to meet the specific needs of bioinformatics research. Parameter configuration can include: Axis range and scale: adjust the optimal display range and scale of the axis according to the characteristics of the data to highlight important data changes. Color scheme: select the color scheme that best suits the specific drawing type, such as the red and blue gradient in the volcano map to indicate the upregulation and downregulation of gene expression. Title font size and legend position: optimize the title font size and legend placement to ensure that the chart is both beautiful and practical, easy to interpret and reference.

[0069] A bioinformatics drawing library is constructed by a variety of bioinformatics drawing templates. When a user makes a drawing request, the intelligent body will select an appropriate template based on the drawing type and specific data characteristics, and adjust the parameters in the template through natural language understanding to generate a chart that meets the user's needs.

[0070] The biological information drawing library construction process of this embodiment simplifies the user's drawing operation and ensures the professionalism and scientificity of the chart. Through the experience accumulation of historical biological information drawing, the drawing agent can provide bioinformatics researchers with more intuitive and accurate data visualization tools to promote the progress of scientific research.

[0071] The supervisor agent realizes the collaborative cooperation of various agents by training the first large language model. Optionally, in the multi-agent based bioinformation mapping system provided in the embodiment of the present application, the supervisor agent is trained in the following manner: obtaining multiple historical requests and historical tasks corresponding to each historical request, wherein the historical tasks include at least one of the following: historical mapping tasks, historical information consultation tasks, and historical data analysis tasks; determining each historical request and the historical tasks corresponding to the historical requests as a group of training samples to obtain multiple groups of training samples; training the large language model through the multiple groups of training samples to obtain the supervisor agent.

[0072] In some embodiments, a large number of historical requests and corresponding task instances are collected, which can be extracted from the interaction records between past users and the system, including: Historical drawing tasks: The drawing requirements put forward by users, such as drawing volcano plots, heatmaps or bubble charts, as well as requests for adjusting specific parameters of the charts. Historical information consultation tasks: Questions asked by users, such as drawing skills, bioinformatics concepts, usage methods of analysis tools, etc. Historical data analysis tasks: The processing and analysis operations specified by users on data, such as differential expression analysis, correlation calculation, etc. Each pair of historical requests and their corresponding historical tasks is determined as a set of training samples. In these training samples, the historical requests serve as inputs (i.e., natural language instructions), and the execution details and results of the historical tasks serve as outputs (i.e., the operations that the model should take and the expected outputs). For example: Input: "Please help me draw a volcano plot showing gene expression differences." Output: Call the data analysis agent to calculate gene expression differences, then use the R language drawing agent to generate a volcano plot, and finally return it to the user through the supervisor agent.

[0073] With sufficient training samples, train a large language model: Convert text data into a digital form that can be processed by the model, such as mapping vocabulary to a high-dimensional vector space through word embedding. Use a neural network architecture suitable for large-scale text processing, such as the Transformer model. During the training process, define a loss function to measure the gap between the model prediction and the actual task execution, and use optimization techniques such as backpropagation algorithm and gradient descent to gradually adjust the model parameters to make its prediction of the training data closer and closer to the actual task execution description. Through trial and error or more advanced parameter tuning methods (such as Bayesian optimization), find the optimal hyperparameters such as learning rate and batch size to obtain better training results. During the model training process, use a validation set to monitor the model performance to prevent overfitting; finally, evaluate the generalization ability of the model on the test set to ensure that it can accurately understand and execute various requests from users in the field of bioinformatics.

[0074] In this embodiment, by training the first large language model, the supervisor agent can understand the user's intention, accurately identify the task type, and effectively coordinate the work of each agent, ensuring that the entire system can smoothly and professionally serve the various needs of bioinformatics researchers.

[0075] According to another embodiment of the present application, an application process of the above multi-agent-based bioinformatics drawing system is further provided. Figure 2 It is a schematic diagram of the application process of the multi-agent-based bioinformatics drawing system provided according to the embodiments of the present application, as Figure 2As shown, the user interacts with the supervisor agent through natural language. The supervisor agent recognizes the intent of the user's natural language and routes the task based on the recognized task to other agents, including the data analysis agent, which calls Python-related tools for data analysis. The data analysis agent has pre-set scientific computing libraries such as pandas and numpy. The R language plotting agent is configured with expert-level basic code and R tools, and uses the basic plotting code tuned by experts and the R rendering tool for plotting. The intelligent customer service agent is configured with a local knowledge base, including documents such as plotting methods and bioinformatics for answering users' information consultations.

[0076] In this embodiment, through the supervisor agent, the customer service agent, the data analysis agent, and the R language plotting agent, it is realized that the user only needs to communicate with the supervisor agent, and the supervisor agent intelligently allocates tasks. The customer service agent can quickly answer common questions based on the knowledge base. The data analysis agent uses Python tools to process data and seamlessly connects with the R language plotting agent. The R language plotting agent can output high-quality charts and related codes, legends, and parameter descriptions that meet the publication standards, and supports natural language chart modification. After understanding the language, the agent generates code directly for rendering, which is convenient to generate, thus providing a professional and one-stop bioinformatics plotting solution for users.

[0077] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0078] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0079] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be realized by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for realizing the process Figure 1one or more processes and / or blocks Figure 1 a device for the functions specified in one or more blocks

[0080] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions in the process Figure 1 one or more processes and / or blocks Figure 1 the functions specified in one or more blocks

[0081] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the process Figure 1 one or more processes and / or blocks Figure 1 the functions specified in one or more blocks

[0082] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory

[0083] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media

[0084] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves

[0085] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0086] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0087] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A multi-agent-based biological information mapping system, characterized in that, Including: A supervisor agent, which is used to receive a user's drawing request, analyze the semantics of the drawing request, generate a drawing task, and send the drawing task to a drawing agent; The drawing agent includes a bioinformatics drawing library, which is used to analyze the semantics of the drawing task, determine a bioinformatics drawing template corresponding to the semantics from the bioinformatics drawing library, generate drawing code based on the semantics and the bioinformatics drawing template, and execute the drawing code to obtain a target drawing. Among them, the bioinformatics drawing library contains a variety of bioinformatics drawing templates.

2. The system according to claim 1, wherein The drawing agent generates the target drawing in the following way: Analyze the drawing task to obtain drawing data and a drawing type; Call a data analysis agent to process the drawing data to obtain a target data analysis result, where the target data analysis result includes at least one of the following: a differential analysis result, a correlation analysis result, and a data statistics result; Determine a target bioinformatics drawing template corresponding to the drawing type, and generate the drawing code based on the target bioinformatics drawing template, the drawing data, and the target data analysis result. Among them, the drawing code includes at least one of the following: data code, code comments, legend explanations, and parameter settings; Run the drawing code through a code execution environment to obtain the target drawing.

3. The system according to claim 2, wherein The drawing agent is also used to: Receive the user's code modification instruction, analyze the code modification instruction to obtain a modified target code, where the code modification instruction is used to indicate modifying the target drawing; Replace the code to be modified in the drawing code with the modified target code to obtain an updated drawing code; Run the updated drawing code through a code execution environment to obtain an updated target drawing.

4. The system according to claim 1, wherein The system also includes: A customer service agent, which includes a bioinformatics knowledge base, and is used to receive an information consultation task sent by the supervisor agent, retrieve bioinformatics knowledge associated with the information consultation task from the bioinformatics knowledge base through a retrieval augmented generation model, and determine the bioinformatics knowledge as a consultation result.

5. The system according to claim 4, characterized in that The customer service agent generates the consultation result in the following way: Extract the consultation keywords in the information consultation task, and retrieve a target document associated with the consultation keywords from the bioinformatics knowledge base through the retrieval augmented generation model; Extract the associated data corresponding to the consultation requirements in the information consultation task from the target document; Generate the consultation result based on the information consultation task and the associated data.

6. The system according to claim 1, characterized in that, The system also includes: A data analysis agent, which includes a variety of preset data analysis tools, and is used to receive a data analysis task sent by the supervisor agent, call the data analysis tools to perform data processing operations on the data in the data analysis task, and generate a data analysis result.

7. The system according to claim 6, characterized in that, The data analysis agent generates the data analysis result in the following way: Extract the data to be processed in the data analysis task, and preprocess the data to be processed to obtain the preprocessed data, where the preprocessing includes at least one of the following: data format and data type recognition and missing value processing; Based on the task requirements in the data analysis task, call the target data analysis tool to perform target operations on the preprocessed data to obtain the data analysis result, where the target operations include at least one of the following: data statistics, difference analysis, correlation analysis, and format conversion.

8. The system according to claim 4, wherein The biological information knowledge base is obtained in the following way: Obtain the biological information knowledge of multiple data sources, where the biological information knowledge includes at least: bioinformatics knowledge, biological information drawing data, and data analysis tool operation documents; Convert the biological information knowledge into biological information vectors, and construct a knowledge graph based on the biological information vectors; Construct a document index for each biological information vector, and the knowledge graph and all document indexes constitute the biological information knowledge base.

9. The system according to claim 1, wherein The biological information drawing library is obtained in the following way: Obtain a variety of historical biological information drawings, and extract the drawing types from the variety of historical biological information drawings, where the drawing types include at least: volcano plots, heatmaps, and bubble charts; Determine the target drawing package, and use the target drawing package to construct a basic drawing template for each drawing type, where the basic drawing template includes at least one of the following: axis setting, color mapping, legend, and text format; Configure parameters for each basic drawing template based on the historical biological information drawings to obtain a variety of biological information drawing templates, and construct the biological information drawing library from the variety of biological information drawing templates, where the parameters include at least one of the following: axis range and scale, color scheme, title font size, and legend position.

10. The system according to claim 1, wherein The supervisor intelligent agent is trained in the following way: Obtain multiple historical requests and the historical tasks corresponding to each historical request, where the historical tasks include at least one of the following: historical drawing tasks, historical information consultation tasks, and historical data analysis tasks; Determine each historical request and the historical task corresponding to the historical request as a set of training samples to obtain multiple sets of training samples; Train a large language model through the multiple sets of training samples to obtain the supervisor intelligent agent.

Citation Information

Patent Citations

  • Interactive analysis system and method for transcriptome project with reference genome based on cloud computing platform

    CN109086567A

  • Data visualization chart library system

    CN110866379A

  • Scientific drawing method and system based on multi-agent cooperation and storage medium

    CN119105738A

  • Batch drawing method and device based on multiple intelligent drawing models, equipment and medium

    CN119722430A

  • Multi-agent interactive efficient data analysis system

    CN119988421A

Cited By

  • Intelligent system, method and equipment for assisting multi-step genome data analysis

    CN121306238A

  • Intelligent data analysis method and system based on large language model

    CN122285733A

  • Intelligent data analysis method and system based on large language model

    CN122285733B