A data analysis method, device, storage medium and electronic equipment
By performing semantic analysis and dynamic priority sorting on user questions and intelligently selecting data processing tools, the problem of low tool calling efficiency in the intelligent question-answering system is solved, and an efficient question-answering process and improved user experience are achieved.
Patent Information
- Application Number
- CN202511106017.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-08-07
AI Technical Summary
The tool calling efficiency in the existing intelligent question-answering system is low, resulting in frequent task chain breaks and low work efficiency when multiple tools work together.
By obtaining user questions for semantic analysis, determining the problem domain and feature information, dynamically prioritizing relevant data processing tools, calling the most suitable tool for analysis, and generating answer results.
It improves the accuracy of problem understanding and the intelligence level of tool calling, avoids frequent breaks in the task chain when multiple tools work together, and improves the system's work efficiency and user experience.
Smart Images

Figure CN120611800B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, and particularly relates to a data analysis method and device, a storage medium and an electronic device. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, the current intelligent question answering system shows excellent language processing capability. The intelligent question answering system is an automatic information retrieval and interaction system based on artificial intelligence technology, which can understand the question raised by the user in natural language, and quickly give accurate and relevant answers by analyzing the problem semantics, matching the knowledge base or calling external tools.
[0003] However, in actual application of enterprises, the related intelligent question answering system only supports tool-level integration. With the continuous increase in the number of system integration tools, it becomes more and more difficult to find suitable data processing tools, resulting in low tool calling efficiency, and further leading to frequent task chain breakage and low work efficiency when multiple tools work together. SUMMARY
[0004] The present disclosure provides a data analysis method, device, storage medium and electronic device. The main purpose is to solve the problem of low tool calling efficiency in related technologies, which further leads to frequent task chain breakage and low work efficiency when multiple tools work together.
[0005] In a first aspect, the present application provides a data analysis method, comprising:
[0006] obtaining a target question raised by a user;
[0007] performing semantic analysis on the target question to determine the domain corresponding to the target question and semantic feature information of the target question;
[0008] determining a plurality of data processing tools related to the target question based on the domain corresponding to the target question and the semantic feature information;
[0009] performing dynamic priority sorting on the plurality of data processing tools based on the relevance index of the plurality of data processing tools and the target question and historical use;
[0010] calling a corresponding data processing tool to analyze the target question according to the dynamic priority of the plurality of data processing tools, and generating and displaying the answer result corresponding to the target question.
[0011] In a second aspect, the present application provides a data analysis, comprising: a user interaction layer, a question answering processing layer, a tool calling layer and a service layer.
[0012] The user interaction layer is configured to obtain a target question raised by a user and display a corresponding answer result of the target question to the user.
[0013] The question and answer processing layer is configured to perform semantic analysis on the target question, determine a domain corresponding to the target question and semantic feature information of the target question, determine a plurality of data processing tools related to the target question based on the domain corresponding to the target question and the semantic feature information, and perform dynamic priority sorting on the plurality of data processing tools based on a relevance index of the plurality of data processing tools and the target question and historical use cases.
[0014] The tool calling layer is configured to call corresponding data processing tools to analyze the target question according to the dynamic priority of the plurality of data processing tools, and generate a corresponding answer result of the target question.
[0015] The service layer is configured to provide a plurality of data processing services and the plurality of data processing tools for the system.
[0016] In a third aspect, the present application provides a data analysis device, comprising:
[0017] An obtaining module configured to obtain a target question raised by a user.
[0018] An analysis module configured to perform semantic analysis on the target question, determine a domain corresponding to the target question and semantic feature information of the target question.
[0019] A determination module configured to determine a plurality of data processing tools related to the target question based on the domain corresponding to the target question and the semantic feature information.
[0020] A sorting module configured to perform dynamic priority sorting on the plurality of data processing tools based on a relevance index of the plurality of data processing tools and the target question and historical use cases.
[0021] A calling module configured to call corresponding data processing tools to analyze the target question according to the dynamic priority of the plurality of data processing tools, and generate and display a corresponding answer result of the target question.
[0022] In a fourth aspect, the present application provides a non-volatile computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the method of the first aspect.
[0023] In a fifth aspect, the present application provides an electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the method of the first aspect.
[0024] In a fifth aspect, the present application provides a computer program product, which has stored thereon a computer program, and the computer program is executed by a processor to implement the method of the first aspect.
[0025] The data analysis method, device, storage medium and electronic device provided by the present disclosure, wherein the method comprises: firstly obtaining a target question proposed by a user; performing semantic analysis on the target question to determine the field corresponding to the target question and the semantic feature information of the target question; determining a plurality of data processing tools related to the target question based on the field corresponding to the target question and the semantic feature information; dynamically prioritizing the plurality of data processing tools based on the relevance indicators of the plurality of data processing tools and the target question and the historical usage; calling the corresponding data processing tool to analyze the target question according to the dynamic priority of the plurality of data processing tools, and generating and displaying the answer result corresponding to the target question. Compared with the related art, the present application accurately identifies the field and key information of the target question by combining context semantic analysis, thereby effectively screening the related data processing tools. On this basis, the dynamic priority of the candidate tools is prioritized by comprehensively considering the relevance indicators of the tools and the problems and their historical usage, ensuring that the more suitable and reliable tools are preferentially called to process the problems. Not only the accuracy of problem understanding is improved, but also the intelligent level and execution efficiency of tool calling are enhanced, avoiding frequent task chain breaks in multi-tool collaborative work, improving the work efficiency of the entire system, thereby significantly improving the overall service response capability and user experience of the system.
[0026] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present application, the drawings required in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0028] Figure 1 A flowchart of a data analysis method provided by an embodiment of the present application is shown;
[0029] Figure 2 A flowchart of another data analysis method provided by an embodiment of the present application is shown;
[0030] Figure 3 A flowchart of an example provided by an embodiment of the present application is shown;
[0031] Figure 4 Another example provided by the embodiment of the application is shown in a flowchart.
[0032] Figure 5 An example provided by the embodiment of the application is shown in a schematic diagram.
[0033] Figure 6 Another example provided by the embodiment of the application is shown in a schematic diagram.
[0034] Figure 7 A structural schematic diagram of a data analysis device provided by the embodiment of the application is shown. DETAILED DESCRIPTION
[0035] The technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.
[0036] It should be noted that in the description of the present application, the terms "include", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms "first", "second" and the like in the present application are used to distinguish similar objects, not to describe a specific order or sequence.
[0037] With the continuous evolution of artificial intelligence (AI) technology, current large models have shown excellent language processing capabilities. However, they still have significant limitations in dealing with certain specific scenarios. For example, in dealing with private domain professional knowledge, large models often fall short; there is also a certain lag in obtaining the latest information; when following a fixed process to perform tasks, it is difficult to accurately adapt; and when faced with complex project automatic planning, it is difficult to meet the requirements. To effectively resolve these problems, experts have proposed many tool solutions. However, the calling efficiency of these tools is generally low, and multiple tools often need to be customized for collaborative work. Due to the lack of a unified interface standard, the task chain frequently breaks down. For example, after completing data analysis and generating a table, manual data saving is required to continue subsequent operations, which undoubtedly reduces work efficiency.
[0038] In this context, the Model Context Protocol (MCP) was born. The core goal of this protocol is to build an efficient and stable information transmission channel between large models and external tools. With the MCP protocol, developers no longer need to write complex interface code for each external tool. Different forms of MCP Hosts (which can be applications, proxy programs, web applications, or desktop applications) can easily access a large number of third-party tools. The MCP protocol provides standardized interface specifications, allowing MCP Hosts to interact seamlessly with local or remote resources in a secure and controllable environment.
[0039] However, although the MCP protocol provides a standardized solution for tool invocation, its related implementation still has some shortcomings. Currently, it only supports tool-level integration and fails to properly address key issues such as context semantic association and cross-tool task flow automation. As the number of system integration tools continues to increase, it becomes increasingly difficult to find the most suitable tool. At the same time, the accuracy and execution efficiency of tool parameters generated by large language models (LLMs) are showing a downward trend. In addition, the more MCP services deployed, the more resources consumed, which is undoubtedly a major problem that needs to be solved for enterprise applications with limited resources and high concurrent business demands.
[0040] In related technologies, the MCP protocol follows a client-server architecture, aiming to provide a unified interface and process for intelligent question-answering systems and other applications, supporting tool discovery, invocation execution, bidirectional communication, and context management. Its core consists of MCP hosts (responsible for receiving questions and interacting with models), MCP clients (maintaining a one-to-one connection with servers to invoke specific operations), MCP servers (providing context, tools, and prompt information), and tool sets. MCP decouples the protocol layer from the transport layer through layered design, ensuring communication flexibility and semantic consistency, supporting stdio, HTTP-SSE, and streaming HTTP transmission modes, and adapting to different deployment scenarios from local to cloud. In intelligent question-answering systems, MCP services interact with large language models (LLMs), obtaining tool lists and structured prompt words through the initialization phase. In the query processing phase, LLM generates JSON-formatted tool invocation instructions based on user input questions, requests the server to execute the corresponding tools via the client, and returns the final response to the user based on tool execution results, achieving effective management and coordination of complex multi-tool environments.
[0041] Related technologies fail to address the issues of semantic association and task flow automation. They only support tool-level integration but lack support for contextual semantic associations. They also fail to automate task flows across tools, making tool selection difficult in complex application scenarios and impacting work efficiency. As the number of integrated tools increases, the accuracy and execution efficiency of tool parameters generated by LLM significantly decrease, making it difficult to meet actual business needs. This reduces the practicality and reliability of the system and results in high resource consumption. The more MCP services deployed, the greater the demand for system resources. In resource-constrained and highly concurrent enterprise application scenarios, this not only increases operating costs but also places additional pressure on the company's IT infrastructure.
[0042] Based on the above problems with server configuration, in order to improve the low efficiency of related technical tools, which leads to frequent task chain breaks and low work efficiency when multiple tools work together. This embodiment provides a data analysis method, such as Figure 1 As shown, the method comprises the following steps:
[0043] Step 101: Obtain the target question raised by the user.
[0044] For example, the user interaction layer serves as the system's front-end interface, responsible for direct interaction with users. It receives questions from users through various channels, such as web pages and mobile apps, and provides system-processed answers. This layer ensures that users can easily access the information they need, supports multi-platform access, and enhances system usability and user experience.
[0045] Step 102: Perform semantic analysis on the target question to determine the domain corresponding to the target question and semantic feature information of the target question.
[0046] For example, semantic feature information can be key elements with specific meaning extracted from a user's question or request. These elements can help the system understand the user's intent and provide corresponding answers or solutions. Based on the domain and semantic feature information corresponding to the target question, the system can identify specific content in fields such as healthcare (symptoms, treatment plans), fintech (market trends, investment advice), education (courses, learning resources), law (regulations, contract review), and information technology (programming, troubleshooting). It can then accurately match relevant knowledge and tools with the context to achieve efficient question-answering and intelligent decision-making.
[0047] Step 103: Based on the domain and semantic feature information corresponding to the target problem, determine multiple data processing tools related to the target problem.
[0048] In some examples, the tool information table of the data processing tool can include tool ID, associated service ID, tool name, function description, parameter information, call times, success times, average response time, and answer accuracy rate (based on historical data statistical analysis) and other key indicators, for quantifying the performance and use effect of the tool.
[0049] Step 104, based on the relevance indicators of the plurality of data processing tools to the target problem and the historical use, dynamically prioritizing the plurality of data processing tools.
[0050] In some examples, in order to solve the problem of tool information redundancy, LLM inference efficiency decline and tool selection precision reduction caused by the introduction of a large number of MCP services in the question and answer system, a tool screening mechanism based on domain recognition and dynamic priority sorting can be used. This mechanism first performs context semantic analysis based on historical information and the current problem to determine the business domain to which the problem belongs, and limits subsequent processing only in the tools corresponding to the relevant MCP services in this domain. On this basis, a dynamic priority sorting module based on reinforcement learning can be further introduced to score and sort the candidate tool set in real time according to its relevance and applicability, thereby improving the accuracy and execution efficiency of tool invocation.
[0051] For example, the question and answer processing layer deeply processes the user question received from the front end, including steps such as semantic analysis and intent recognition, to extract key domain information in the question. Based on this information, the scores of each tool are calculated using a pre-set reward function, thereby determining a prioritized tool list. This step is crucial for accurately matching appropriate tools and directly affects the accuracy and efficiency of subsequent processing.
[0052] For example, structured prompts can be constructed to guide the selection and use of appropriate tools by large language models, enhancing the flexibility and accuracy of the question and answer system.
[0053] Step 105, according to the dynamic priority of the plurality of data processing tools, calling the corresponding data processing tool to analyze the target problem and generate and display the answer result corresponding to the target problem.
[0054] For example, the tool invocation layer can be based on the prioritized tool list provided by the question and answer processing layer, and this layer is responsible for the invocation of specific tools. It can communicate with the MCP service layer to request and obtain the resources and data required for task execution. This process may involve the cooperative work of multiple tools to complete complex query or analysis tasks. The design of the tool invocation layer ensures that the system can flexibly utilize various professional tools, improving the system's scalability and adaptability.
[0055] In some examples, the system in this embodiment can also include an MCP service layer at the bottom of the system, which provides a variety of basic services and tools, such as log storage services, data analysis tools, etc. Different services can access the system through standard input and output (stdio), HTTP-based server push events (HTTP-SSE), or HTTP protocols supporting bidirectional streaming. According to the characteristics of each connection method, the system sets differentiated access permissions and resource configurations to optimize resource utilization efficiency and ensure the security and stability of the service.
[0056] Compared with related technologies, the embodiment first acquires a target question proposed by a user; performs semantic analysis on the target question to determine a field corresponding to the target question and semantic feature information of the target question; then determines a plurality of data processing tools related to the target question based on the field corresponding to the target question and the semantic feature information; performs dynamic priority sorting on the plurality of data processing tools based on a relevance index of the plurality of data processing tools and the target question and historical usage; and calls a corresponding data processing tool to analyze the target question according to the dynamic priority of the plurality of data processing tools, and generates and displays a corresponding answer result of the target question. Compared with the related technologies, the embodiment accurately identifies the field and key information of the target question through context semantic analysis, thereby effectively screening the related data processing tools. On this basis, the dynamic priority of the candidate tools is sorted by comprehensively considering the relevance index of the tools and the problem and the historical usage, so as to ensure that the more suitable and reliable tools are preferentially called to process the problem. Not only the accuracy of problem understanding is improved, but also the intelligent level and execution efficiency of tool calling are enhanced, the frequent task chain breakage in multi-tool collaborative work is avoided, and the working efficiency of the entire system is improved, thereby significantly improving the overall service response capability and user experience of the system.
[0057] As a refinement of the embodiment, the dynamic priority of the data processing tool can be determined by, but not limited to, the following manners, such as Figure 2 As shown in Figure 2 A flowchart of a data analysis method provided by the embodiment of the disclosure, comprising:
[0058] Step 201, acquiring a target question proposed by a user.
[0059] For example, the user proposes a question through the user interaction layer, such as "the server performance has been abnormal in the past week, please analyze the reason".
[0060] Step 202, performing semantic analysis on the target question to determine a field corresponding to the target question and semantic feature information of the target question.
[0061] For example, the question and answer processing layer performs semantic analysis and intent recognition on the question, extracts semantic feature information such as "the past week", "server performance", "abnormality", and determines that the user's intent is to analyze the cause of the abnormality.
[0062] Optionally, step 202 can specifically include: using natural language processing to analyze the target question and extract keywords; matching the keywords and understanding the semantics to identify the field corresponding to the target question, and combining context information to determine the intent and requirement corresponding to the target question, and determining the semantic feature information corresponding to the target question.
[0063] For example, the natural language processing technology is used to deeply analyze the target question input by the user, first extract the keywords and key phrases, and then identify the specific field involved in the target question by combining keyword matching and semantic understanding. On this basis, the system further combines context information to analyze and judge the intent of the question and the user's potential requirements, thereby accurately determining the key information corresponding to the target question. This process not only improves the accuracy of question understanding, but also provides a solid foundation for subsequent data processing tool selection, answer generation, and personalized recommendation.
[0064] Step 203, based on the field and semantic feature information corresponding to the target question, determine a plurality of data processing tools related to the target question.
[0065] For example, the question and answer processing layer performs semantic analysis and intent recognition on the question, extracts key information such as "the past week", "server performance", "abnormality", and determines that the user's intent is to analyze the cause of the abnormality; according to the field and key information corresponding to the question, the relevant tools (such as log analysis tools, performance monitoring tools, etc.) are selected from the MCP service layer, and the historical use and relevance are combined to dynamically prioritize.
[0066] Optionally, step 203 can specifically include: based on the field and semantic feature information corresponding to the target question, determining the data processing service required to solve the target question; determining the data processing tool corresponding to the data processing service as the plurality of data processing tools related to the target question.
[0067] In some examples, for the target question raised by the user, the type of data processing service required to solve the question can be determined according to the field corresponding to the question and the extracted semantic feature information (such as time range, theme object, question type, etc.). A question and answer system can introduce thousands of MCP services, each service corresponding to at least one tool, for example, if the question involves "server performance abnormality analysis", the system may need to call log query, index monitoring, trend analysis, etc. Data processing services. Each data processing service usually corresponds to one or more specific tools, which are responsible for performing actual data processing tasks. For example,Figure 3 As shown, after the user inputs the question, the LLM analyzes all the fields involved and obtains the relevant configuration information from the MCP server information table in the database, installs the MCP server, connects it, and creates and connects the MCP client. At the same time, detailed information of the tool list corresponding to the MCP server is added to the database. The tools are sorted according to the priority, and a prompt is generated and passed to the LLM, which determines whether the tools need to be called. If so, all tool names and parameters are extracted, the tools are called through the MCP client, the usage frequency and response time in the tool information table are updated, and all tool results are added to the dialogue and passed to the LLM for processing again.
[0068] Step 204: According to the matching degree of the content of the target question and the function description of the plurality of data processing tools, the relevance index of the plurality of data processing tools to the target question is determined.
[0069] For example, the relevance index comprehensively considers factors such as keyword matching degree, semantic similarity, and functional applicability, and is used to quantitatively evaluate the support capability of different data processing tools for the target question, providing a basis for subsequent tool recommendation and optimization.
[0070] Step 205: The comprehensive score of the plurality of data processing tools is calculated by a preset tool evaluation model in combination with the relevance index and the historical usage situation.
[0071] For example, when evaluating the applicability of data processing tools, the system combines the relevance index and the historical usage situation to calculate the scores of the plurality of candidate tools through a preset tool evaluation model. The relevance index reflects the matching degree of the tool function and the current target question, while the historical usage situation includes dimensions such as tool calling frequency, execution efficiency, and user satisfaction. The evaluation model is based on a weighting algorithm or a machine learning method, dynamically adjusts the weights of each index, and generates a more targeted comprehensive score, so as to ensure that the recommended results meet the current task requirements and take into account the actual application performance of the tools.
[0072] Optionally, the method of the embodiment can further include: in the calling process of the data processing tool, the performance of the tool is monitored in real time and the relevance index and the historical usage situation data are updated; and the tool evaluation model is optimized according to the relevance index and the historical usage situation data.
[0073] In some examples, the system monitors the performance of the tool in real time, dynamically updates the relevance index and the historical usage situation data, and continuously optimizes the tool evaluation model based on these data, thereby improving the accuracy of tool selection and calling efficiency.
[0074] Optionally, the optimization of the tool evaluation model can further include: based on the actual use of the data processing tool, dynamically adjusting the reward function in the tool evaluation model through a reinforcement learning algorithm.
[0075] For example, in order to more comprehensively evaluate the actual performance of the tool, a comprehensive reward function can be designed as the core evaluation index of dynamic ranking. The reward function integrates multiple key factors, including the use frequency, accuracy rate and response time of the tool, which are respectively weighted and combined through adjustable weights α, β and δ to form a unified quantitative evaluation standard:
[0076] (Formula 1)
[0077] In some examples, α, β, δ ∈ [0, 1] and α + β + δ = 1 are satisfied, respectively used to adjust the influence weight of use frequency, accuracy rate and response time in scoring; f represents the use frequency of the tool, reflecting its applicability; acc represents the accuracy rate, reflecting the output quality; t represents the response time, measuring the execution efficiency.
[0078] For example, the initial weights can be set as α = 0.4, β = 0.4 and δ = 0.2 according to historical records, and the parameters can be continuously optimized through a reinforcement learning algorithm, so that the system can adapt to different business scenarios and user needs, and improve the intelligent level and operation efficiency of the question and answer system.
[0079] In some examples, in order to ensure the rationality and adaptability of the weight configuration, multiple determination methods are supported, including but not limited to expert experience method, data analysis method and reinforcement learning adjustment method, etc.
[0080] Optionally, based on the actual use of the data processing tool, the reward function in the tool evaluation model is dynamically adjusted through a reinforcement learning algorithm, which can further include: based on the current performance index of the data processing tool and the cumulative performance improvement of the data processing tool in a preset time period, the weight adjustment action corresponding to the reward function is output through a policy network, and the policy network is updated through a policy gradient method according to the collected trajectory data; based on the weight adjustment action, the state action trajectory of the tool evaluation model is continuously interacted and recorded during the training process, and the parameters and weight configuration of the reward function are iteratively optimized.
[0081] Exemplarily, a reinforcement learning algorithm can be used to optimize the weight coefficients in the tool evaluation model, where the state (State) is composed of a vector [f, acc, t] of current tool performance indicators such as usage frequency f, answer accuracy acc, and response time t, as well as historical weight performance; the action (Action) involves adjusting the weight coefficients α, β, δ, outputting incremental values Δα, Δβ, Δδ by discretization, and setting the constraint conditions α, β, δ ∈ [0, 1] and α + β + δ = 1 to maintain reasonableness; the reward (Reward) includes direct reward The overall improvement of tool performance over a period of time is measured by the long-term cumulative reward for instant feedback; the policy (Policy) uses a neural network as the policy network, which inputs the state and outputs the weight adjustment action, and is trained by the policy gradient method and used in the early stage The policy balances exploration and utilization.
[0082] In some examples, the specific training process can include initializing random weight and policy network parameters, then selecting actions according to the current policy during interaction with the environment and evaluating tool performance under the new weight combination to calculate the reward, recording state, action, and reward data, and finally updating the policy network parameters based on the trajectory data until a stable weight configuration is found, achieving dynamic adjustment of the reward function in the tool evaluation model.
[0083] In some examples, by continuously optimizing the weight configuration of each indicator through the policy network and reinforcement learning algorithm, the system can adaptively adjust the evaluation focus according to different business scenarios, thereby achieving dynamic evaluation and optimal selection of tool performance and improving the intelligent level and operational efficiency of the overall question answering system.
[0084] Step 206, sorting multiple data processing tools according to the comprehensive score and generating a tool priority list.
[0085] For example, the prompt generated according to the tool list sorted by dynamic priority contains three parts: tools_desc describes the information of each available tool in detail, such as name, description, and parameters; tool_examples provides typical tool usage examples to help understand the correct invocation method; tool_rules specifies the rules for tool usage, including priority sorting, invocation format requirements, etc. These parts together form a complete prompt, for example:
[0086] In addition, in order to identify the tool invocation instruction in the LLM output, a regular expression can be used to parse the <tool_use> tag and extract the tool name and parameter information. The regular expression is as follows:
[0087] (Formula Two)
[0088] where [ \s\S ] *? means matching all characters between the characters immediately before and after, <name> ...< / name> between the tool name, <arguments> ...< / arguments> between the tool parameters.
[0089] This way not only ensures the accuracy and efficiency of tool invocation, but also guides the LLM to make intelligent decisions through explicit rules and examples, significantly improving the overall performance and user experience of the question and answer system.
[0090] Step 207, according to the tool priority list, the corresponding data processing tool is called to analyze the target question, and the answer result corresponding to the target question is generated and displayed.
[0091] In some examples, the tool invocation layer sequentially invokes the corresponding tools according to the sorting results, first obtains the server log data in the specified time period from the MCP service layer through the log analysis tool, and analyzes to extract key performance indicators and error information. Then the tool call results are integrated and input into the LLM to generate a structured analysis report, and are displayed to the user through the user interaction layer, so that the user can intuitively understand the reasons for the server performance anomaly, and realize the whole process automation from problem input to intelligent analysis and visual output.
[0092] In some examples, first, MCP Server initialization and configuration, user interaction and tool screening, tool invocation and result integration, and result generation and output. First, in the MCP Server initialization and configuration phase, the system loads the necessary configuration information to complete the installation of the MCP server and establishes a connection, initializes the MCP client and connects it, and then stores the MCP server information and tool information in the database to form the corresponding information table. Then enter the user interaction and tool screening link, after the user raises a question, analyze the historical question and answer records through the LLM, determine the involved field according to the user input, and select the MCP Server and its corresponding tool list in the related field from the database, calculate the score using the reward function and prioritize the tools accordingly, and finally input the sorted tool information into the prompt LLM. In the tool invocation and result integration stage, it is determined whether to call the tool according to the judgment of the LLM, if it needs to be called, the corresponding tool information (such as function name and parameters) is extracted, the MCP client initiates a call request, updates the tool information table in the database and adds the call result to the historical session, this process may be repeated multiple times until there is no need to call the tool. In the result generation and output stage, the system sorts all tool invocation input and output and LLM generated content as the final result returned to the user, completing the entire question and answer process.
[0093] Optionally, the method can further include: determining a service type of the user request according to the data processing service; and determining whether the user has access to the data processing service of the service type based on the access permission of the user.
[0094] For example, the MCP Server information table can include fields such as service ID, service name, service description, field of application (in the form of keywords), installation method, service status, provider, creation time, update time, and available tool list, etc., to comprehensively record the basic attributes and running status of each MCP service, and to display in a visual form in the UI interface, so as to facilitate the user to intuitively view and manage service resources.
[0095] For example, in the MCP service management of the embodiment, differentiated access permissions are set for different connection methods. Taking the server log analysis scenario as an example, the system provides two connection methods: internal network connection and external network connection. The internal network connection has high permission, allowing access to all log data and advanced analysis tools such as professional log mining platforms and performance optimization tools, which is suitable for internal technical personnel to conduct in-depth problem troubleshooting and system tuning. The external network connection has lower permission, allowing only to view public log summary information and use basic analysis tools, which is suitable for external users or partners to obtain limited running status information. Through this permission grading management mechanism, the system ensures data security while improving the rationality of resource use and the applicability of question and answer services, effectively supporting the intelligent question and answer needs of enterprises in various business scenarios.
[0096] In some examples, when the user raises a question through the interactive interface, the system performs semantic analysis and intent recognition to extract the key information in the question and determine the field of application and service type. For example, if the user asks “the server performance has been abnormal in the past week, please analyze the cause”, the system will identify that this is a requirement for server log analysis, and further determine the specific service type required, such as log analysis tools or performance monitoring software. Based on the analysis result, the specific service type that meets the user's demand is determined, such as log analysis, performance monitoring, etc. Then check the user's access permission (internal network connection or external network connection) to determine whether access to the requested service type is allowed. If the user has sufficient permission, access to the corresponding service is allowed and the request is continued to be processed. If the user's permission is insufficient, access is restricted, and an error message or limited service options may be returned.
[0097] Optionally, based on the deployment method corresponding to the data processing service, the user is configured with differentiated setting permissions, wherein the deployment method includes any one of a standard input / output method, a server push event based on hypertext transfer protocol, and a hypertext transfer protocol supporting bidirectional streaming.
[0098] Optionally, the deployment mode corresponding to the data processing service is configured with differentiated setting permissions for the user, and specifically can include: if the deployment mode corresponding to the data processing service is a standard input and output mode, a user with administrator permission is allowed to add, delete, and modify the data processing service; if the deployment mode corresponding to the data processing service is a server push event based on the hypertext transfer protocol or a mode based on the hypertext transfer protocol supporting bidirectional streaming, a user without administrator permission is allowed to add, delete, and modify the data processing service.
[0099] For example, the user permissions can be divided into two levels of administrator (Admin) and ordinary user (User), and MCP services of different levels are constructed respectively.
[0100] For example, at the administrator level, three MCP service deployment modes are supported: standard input and output, server push event (Server-Sent Events, SSE) based on HTTP, and Streamable HTTP supporting bidirectional streaming. Among them, stdio as a local deployment mode is simple to implement and has low latency, but it does not support concurrent use, and if a service instance is initialized independently for each user, it will cause serious waste of resources.
[0101] For example, as shown in Figure 4 , the configuration interface of the Admin account when constructing the MCP service can include basic information such as MCP service name, description, provider, and keywords, as well as selection of installation method (such as npx or uvx command line) and installation package source (source 1, source 2, or source 3), while necessary parameters and environment variables need to be input to ensure correct deployment and operation of the service. A user with administrator permission can add, delete, and modify the data processing service of the standard input and output, and an ordinary user without administrator permission is not allowed to add, delete, and modify the data processing service of the standard input and output.
[0102] For example, at the user level, in order to meet individual needs, an ordinary user can be allowed to add, modify, and delete exclusive services based on the two remote invocation modes of HTTP-SSE or Streamable HTTP, without occupying local resources, and service invocation can be completed through API interface only. At the same time, the User account can view and use the services deployed by the Admin, but does not have the permission to add, delete, or modify them, thereby improving flexibility and expandability while ensuring system stability and security.
[0103] For example, as shown in Figure 5As shown, the configuration options of the common user when creating a dedicated service, including basic information such as MCP service name, description, provider and keywords, and selecting the installation method (SSE or Streamable HTTP) and inputting the service uniform resource locator (URL) and request header and other remote calling parameters.
[0104] Optionally, the embodiment method can further include configuring differentiated resource quotas for services with different resource requirements based on the resource requirements corresponding to the service types.
[0105] In some examples, a dynamic process pool (MCP Server Stdio Process Pool) mechanism can be introduced to uniformly schedule and efficiently manage services deployed by Admin through stdio.
[0106] Optionally, the resource requirements of the plurality of data processing services are evaluated according to the expected load corresponding to the plurality of data processing services, and the dynamic process pool is automatically and elastically scaled based on the resource requirements of the plurality of data processing services and the actual resource usage, the dynamic process pool being used for resource quota management of the plurality of data processing services.
[0107] In some examples, the process pool is initialized with a default number of processes of 1 and a maximum number of processes of N. The number of processes can be flexibly configured according to the service resource occupation, and the process pool state is persistently stored in a database.
[0108] Optionally, the dynamic process pool is automatically and elastically scaled based on the resource requirements of the plurality of services and the actual usage of the plurality of services, and the method can further include: if the actual resource usage is lower than the resource requirement, automatically reducing the size of the dynamic process pool; and if the actual resource usage is higher than the resource requirement, automatically expanding the size of the dynamic process pool.
[0109] For example, in view of the sparsity and burstiness of access to MCP services, the dynamic process pool realizes on-demand elastic scaling, significantly reducing the resource idle rate.
[0110] Optionally, the embodiment method can further include initializing the dynamic process pool and pre-creating a preset number of idle processes in the dynamic process pool for a plurality of services, checking whether there is an idle process available in the dynamic process pool corresponding to the service type requested by the user, assigning a process to the service if there is an idle process available in the dynamic process pool corresponding to the service, and calling the corresponding data processing tool in the process to analyze the target problem.
[0111] For example, in view of the sparsity and burstiness of access to MCP services, the dynamic process pool realizes on-demand elastic scaling, significantly reducing the resource idle rate. Figure 6As shown, when a user makes a service request (such as a file system service), the system first checks whether there is a free process corresponding to the service; if so, the tool is directly called to complete the request, and the process is returned to the process pool as a free process; if there is no free process, it is determined whether the service process has reached the maximum process number limit; if not, a new service process is created and added to the process pool to meet the request; if the limit has been reached, the user needs to queue until a free process appears.
[0112] Optionally, the method of the embodiment can further include: if there is no free process corresponding to the service available in the dynamic process pool, detecting whether the process number of the service has reached the maximum limit, and processing the dynamic process pool according to the detection result.
[0113] Optionally, the processing of the dynamic process pool according to the detection result can further include: if the process number of the service has not reached the maximum limit, creating a new service process and adding the new service process to the process pool as a free process waiting for calling; if the process number of the service has reached the maximum limit, the user task corresponding to the service enters a queuing state to wait for a free process to appear.
[0114] For example, if there is no free process corresponding to the service available in the dynamic process pool, the dynamic expansion and contraction judgment logic is: creating a new process when the upper limit is not reached, and queuing the task when the upper limit is reached.
[0115] In some examples, the embodiment proposes an efficient enterprise intelligent question and answer system and platform based on MCP service. The system has the following innovative features: first, according to different connection modes of MCP service, the system sets differentiated access permissions, and automatically expands and contracts according to device resources through a dynamic process pool. This measure not only effectively solves the waste problem caused by excessive resource occupation, but also ensures that users can efficiently call the required tools. Second, the embodiment introduces a tool dynamic priority sorting mechanism. Through this mechanism, the system can reduce the tool selection range according to the context semantic association, and sort the tools according to the scores to obtain the priority, thereby significantly improving the efficiency and accuracy of tool use. Finally, the embodiment refines the overall process of the question and answer system in the application process.
[0116] Optionally, the method of the embodiment can further include: adjusting the size of the dynamic process pool according to the historical load situation.
[0117] Optionally, the method of the embodiment can further include: periodically checking the dynamic process pool, and removing the abnormal processes in the dynamic process pool according to the checking result.
[0118] In some examples, resource utilization maximization and service availability are ensured by periodically cleaning up zombie processes, dynamically adjusting pool size according to historical load, priority queue scheduling, automatic scaling mechanism, and health check mechanism.
[0119] Optionally, the embodiment method can further include: optimizing the large language model and the tool evaluation model by collecting feedback of the user on the question and answer result of the target problem, the large language model being used for context semantic analysis on the target problem.
[0120] For example, the embodiment has the following features, but is not limited to: first, domain adaptation, by analyzing the domain of the user's question, the relevant MCP service and its corresponding tool list are screened from the database, the tool is accurately matched, and the processing efficiency is improved; second, dynamic decision, the large language model judges whether the tool needs to be called in real time according to the current context, and the optimal tool combination is selected by combining the priority sorting mechanism, and the intelligent response ability of the system is enhanced; third, automatic process, from MCP service initialization, tool screening, calling and execution to result generation, the whole process is highly automated, and the manual intervention is greatly reduced. The process builds a complete closed loop of "user demand→tool scheduling→intelligent response", fully utilizes the flexibility of the MCP tool library, deeply integrates the reasoning and decision-making ability of the LLM, and significantly improves the intelligent level and business adaptability of the question and answer system.
[0121] Compared with the prior art, the embodiment proposes an innovative solution system in view of the significant capability limitations of current large models in private domain knowledge processing, latest information acquisition, fixed process execution, and complex project management. By introducing the MCP service management module in the enterprise question and answer system, a standardized interface specification is constructed to realize seamless access and efficient integration of massive MCP services, thereby significantly improving the system's knowledge processing capability in specific fields, information real-time updating capability, process flexible adaptation capability, and complex task decomposition and execution capability. In view of the problem of excessive resource consumption caused by MCP service deployment, the system adopts a refined resource management strategy, flexibly configures differentiated access permissions and resource quotas according to the connection characteristics and resource requirements of the service, ensures efficient utilization and reasonable allocation of resources, and avoids resource waste and performance bottlenecks. In addition, in the tool calling and collaboration mechanism, the system innovatively introduces a dynamic priority tool routing mechanism, which can intelligently filter and recommend an optimal tool priority list based on deep correlation analysis of context semantics, realize seamless collaboration and automatic scheduling between tools, effectively solve the problems of low efficiency and poor accuracy in tool calling in the prior art, and significantly improve the task execution efficiency and overall system performance.
[0122] The embodiment also provides a data analysis system, comprising: a user interaction layer, a question and answer processing layer, a tool calling layer and a service layer; the user interaction layer is used for obtaining a target question proposed by a user and displaying a corresponding answer result of the target question to the user; the question and answer processing layer is used for performing semantic analysis on the target question, determining a field corresponding to the target question and semantic feature information of the target question; determining a plurality of data processing tools related to the target question based on the field corresponding to the target question and the semantic feature information; performing dynamic priority sorting on the plurality of data processing tools based on a relevance index of the plurality of data processing tools and the target question and historical use cases; the tool calling layer is used for calling corresponding data processing tools to analyze the target question according to the dynamic priority of the plurality of data processing tools, and generating the corresponding answer result of the target question; and the service layer is used for providing a plurality of data processing services and a plurality of data processing tools for the system.
[0123] For example, the user interaction layer serves as a front-end interface of the system and is responsible for direct interaction with the user; the question and answer processing layer performs in-depth processing on the user question received from the front end, including steps such as semantic analysis and intent recognition, to extract key field information in the question, and based on the information, scores of tools are calculated based on a preset reward function, so as to determine a tool list sorted by priority; the tool calling layer can be based on the priority-sorted tool list provided by the question and answer processing layer, and the layer is responsible for calling operations of specific tools; and the layer can communicate with the MCP service layer to request and obtain resources and data required for task execution.
[0124] Optionally, the service layer in the embodiment further comprises: a service management module; the service management module is used for configuring differentiated setting permissions for the user based on a deployment mode corresponding to the data processing service, wherein the deployment mode comprises any one of a standard input and output mode, a mode based on a server push event of a hypertext transfer protocol and a mode based on hypertext transfer protocol support bidirectional streaming.
[0125] Optionally, the service management module is further used for evaluating resource requirements of the plurality of data processing services according to expected loads corresponding to the plurality of data processing services; and performing automatic elastic scaling of a dynamic process pool based on the resource requirements of the plurality of data processing services and actual resource usage, the dynamic process pool being used for resource quota management of the plurality of data processing services.
[0126] Embodiments of the application also provide a data analysis device, which is used as Figure 1 As a specific implementation of the method shown in Figure 7 The device comprises: an acquisition module 31, an analysis module 32, a determination module 33, a sorting module 34 and a calling module 35.
[0127] The acquisition module 31 is configured to acquire a target question proposed by a user.
[0128] The analysis module 32 is configured to perform semantic analysis on the target question, determine a field corresponding to the target question and semantic feature information of the target question.
[0129] The determination module 33 is configured to determine a plurality of data processing tools related to the target question based on the field corresponding to the target question and the semantic feature information.
[0130] The sorting module 34 is configured to dynamically prioritize the plurality of data processing tools based on a relevance indicator of the plurality of data processing tools to the target question and historical usage.
[0131] The calling module 35 is configured to call a corresponding data processing tool to analyze the target question according to the dynamic priority of the plurality of data processing tools, and generate and display a corresponding answer result of the target question.
[0132] In some examples of the present embodiment, the sorting module 34 is specifically configured to determine a relevance indicator of the plurality of data processing tools to the target question according to a matching degree of content of the target question and a function description of the plurality of data processing tools, combine the relevance indicator and historical usage, calculate a comprehensive score of the plurality of data processing tools through a preset tool evaluation model, sort the plurality of data processing tools according to the comprehensive score, and generate a tool priority list. Correspondingly, the calling module 35 is specifically configured to call a corresponding data processing tool to analyze the target question according to the tool priority list, and generate and display a corresponding answer result of the target question.
[0133] In some examples of the present embodiment, the analysis module 32 is specifically configured to parse the target question by using natural language processing, extract keywords, identify a field corresponding to the target question through keyword matching and semantic understanding, and determine an intention and a demand corresponding to the target question in combination with context information, and determine semantic feature information corresponding to the target question.
[0134] In some examples of the present embodiment, the determination module 33 is specifically further configured to determine a data processing service required for solving the target question based on the field corresponding to the target question and the semantic feature information, and determine a data processing tool corresponding to the data processing service as the plurality of data processing tools related to the target question.
[0135] In some examples of the present embodiment, the determining module 33 is specifically further configured to configure differentiated setting permissions for the user based on a deployment mode corresponding to the data processing service, wherein the deployment mode comprises any one of a standard input / output mode, a server push event based on a hypertext transfer protocol, and a hypertext transfer protocol supporting bidirectional streaming mode.
[0136] In some examples of the present embodiment, the determining module 33 is specifically further configured to, if the deployment mode corresponding to the data processing service is the standard input / output mode, allow a user with administrator permission to add, delete, and modify the data processing service; if the deployment mode corresponding to the data processing service is the server push event based on the hypertext transfer protocol or the hypertext transfer protocol supporting bidirectional streaming mode, allow a user without administrator permission to add, delete, and modify the data processing service.
[0137] In some examples of the present embodiment, the determining module 33 is specifically further configured to determine a service type requested by the user according to the data processing service, and judge whether the user has access to the data processing service of the service type based on the access permission of the user.
[0138] In some examples of the present embodiment, the analyzing module 32 is specifically further configured to evaluate resource requirements of a plurality of data processing services according to expected loads corresponding to the plurality of data processing services, and automatically perform elastic scaling of a dynamic process pool for resource quota management of the plurality of data processing services according to the resource requirements and actual resource usage of the plurality of data processing services.
[0139] In some examples of the present embodiment, the analyzing module 32 is specifically further configured to, if the actual resource usage is lower than the resource requirement, automatically reduce the size of the dynamic process pool; if the actual resource usage is higher than the resource requirement, automatically expand the size of the dynamic process pool.
[0140] In some examples of the present embodiment, the calling module 35 is specifically further configured to initialize the dynamic process pool, and pre-create a preset number of idle processes in the dynamic process pool for a plurality of services; based on a service type requested by the user, check whether an idle process corresponding to the service is available in the dynamic process pool; if the idle process corresponding to the service is available in the dynamic process pool, allocate a process for the service, and call a corresponding data processing tool in the process to analyze the target problem; if the idle process corresponding to the service is not available in the dynamic process pool, detect whether the number of processes of the service has reached a maximum limit, and perform processing on the dynamic process pool according to the detection result.
[0141] In some examples of the present embodiment, the calling module 35 is specifically further configured to, if the number of processes of the service does not reach the maximum limit, create a new service process and add the new service process to the process pool as an idle process waiting for calling; if the number of processes of the service has reached the maximum limit, the user task corresponding to the service enters a queuing state waiting for an idle process to appear.
[0142] In some examples of the present embodiment, the calling module 35 is specifically further configured to periodically check the dynamic process pool, and remove abnormal processes in the dynamic process pool according to the checking result.
[0143] In some examples of the present embodiment, the sorting module 34 is specifically further configured to, in the calling process of the data processing tool, monitor the tool performance in real time and update the relevance index and the historical usage data; and optimize the tool evaluation model according to the relevance index and the historical usage data.
[0144] In some examples of the present embodiment, the sorting module 34 is specifically further configured to, based on the current performance index of the data processing tool and the cumulative performance improvement of the data processing tool in a preset time period, output a weight adjustment action corresponding to a reward function in the tool evaluation model through a strategy network, the strategy network is updated according to the collected trajectory data by a strategy gradient method; based on the weight adjustment action, continuously interact with and record the state action trajectory of the tool evaluation model in the training process, and iteratively optimize the parameters and weight configuration of the reward function.
[0145] It should be noted that other corresponding descriptions of each functional unit involved in the data analysis device provided in the present embodiment can be referred to the corresponding description in Figure 1 , which will not be repeated here.
[0146] Based on the above methods as shown in Figure 1 and Figure 2 , accordingly, the present embodiment also provides a non-volatile computer readable storage medium having a computer program stored thereon, which is executed by a processor to implement the above methods as shown in Figure 1 and Figure 2 .
[0147] Based on the above methods as shown in Figure 1 and Figure 2 , accordingly, the present embodiment also provides a computer program product having a computer program stored thereon, which is executed by a processor to implement the above methods as shown in Figure 1 and Figure 2 .
[0148] Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash disk, a mobile hard disk, etc.) and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various implementation scenarios of the present application.
[0149] Based on the methods as shown in Figure 1 and Figure 2 , and the virtual device embodiments as shown in Figure 7 , in order to achieve the above-mentioned purposes, the embodiments of the present application further provide an electronic device such as a personal computer, a server, which includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the methods as shown in Figure 1 and Figure 2 .
[0150] In some embodiments, the above-mentioned entity device can also include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a WI-FI module, etc. The user interface can include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. The optional user interface can also include a USB interface, a card reader interface, etc. The network interface can include a standard wired interface, a wireless interface (such as a WI-FI interface), etc. in some embodiments.
[0151] Those skilled in the art can understand that the above-mentioned entity device structure provided by the embodiments does not constitute a limitation on the entity device, and can include more or fewer components, or combine certain components, or different component arrangements.
[0152] The storage medium can also include an operating system, a network communication module. The operating system is a program that manages the hardware and software resources of the above-mentioned entity device, supports the running of information processing programs and other software and / or programs. The network communication module is used to realize the communication between the components inside the storage medium, and the communication with other hardware and software in the information processing entity device.
[0153] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be realized by means of software with a necessary general hardware platform, or by hardware. By applying the scheme of the present embodiment, compared with the prior art, the present embodiment realizes the rapid response and accurate answer to enterprise-level problems through the efficient enterprise intelligent question and answer system based on the MCP service. The system not only supports the administrator to flexibly manage the service deployment through the Admin account, and ensures the effective use of resources in the multi-user environment through the dynamic process pool technology, avoiding resource waste. Ordinary users can conveniently call services through the User account in the Http server push event or bidirectional streaming mode. The system can automatically identify the field and filter related tools according to the historical information combined with the current problem content, and then evaluate the performance of each tool by using a comprehensive reward function, so as to perform dynamic priority sorting. This process considers multiple dimensions such as frequency of use, accuracy of answer and response time, and allows dynamic adjustment of weight coefficients through reinforcement learning algorithm, further optimizing tool selection. Finally, the system organizes the sorted tool list into a prompt in a structured form and inputs it into the LLM, determines whether to call specific tools and parameters according to the output, updates the database after completing the task processing and returns the result to the user. The whole question and answer process has high automation degree, significantly improves the problem solving efficiency and service quality, and ensures the scalability and flexibility of the system.
[0154] It should be noted that, in this document, the terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.
[0155] The above is only a specific embodiment of the present application, which enables those skilled in the art to understand or implement the present application. Various modifications of these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but will conform to the widest scope consistent with the principles and novel features applied herein.
Claims
1. A data analysis method, characterized in that: include: Get the target questions raised by users; Performing semantic analysis on the target question to determine the domain corresponding to the target question and semantic feature information of the target question; Determining multiple data processing tools related to the target problem based on the domain corresponding to the target problem and the semantic feature information, including: determining a data processing service required to solve the target problem based on the domain corresponding to the target problem and the semantic feature information; determining the data processing tool corresponding to the data processing service as multiple data processing tools related to the target problem; configuring differentiated setting permissions for the user based on the deployment method corresponding to the data processing service, wherein the deployment method includes any one of a standard input and output method, a server push event method based on the Hypertext Transfer Protocol, and a method based on the Hypertext Transfer Protocol supporting two-way streaming transmission; Dynamically prioritizing the plurality of data processing tools based on relevance indicators of the plurality of data processing tools to the target problem and historical usage; According to the dynamic priorities of the multiple data processing tools, the corresponding data processing tools are called to analyze the target question, and the answer results corresponding to the target question are generated and displayed.
2. The method according to claim 1, characterized in that The dynamically prioritizing the plurality of data processing tools based on their relevance to the target problem and their historical usage includes: Determining relevance indicators of the multiple data processing tools to the target problem based on a degree of matching between the content of the target problem and the functional descriptions of the multiple data processing tools; Calculating comprehensive scores of the plurality of data processing tools using a preset tool evaluation model in combination with the relevance index and historical usage; sorting the plurality of data processing tools according to the comprehensive scores and generating a tool priority list; The step of calling corresponding data processing tools to analyze the target question according to the dynamic priorities of the multiple data processing tools, and generating and displaying the answer results corresponding to the target question includes: According to the tool priority list, the corresponding data processing tool is called to analyze the target question, and the answer result corresponding to the target question is generated and displayed.
3. The method according to claim 1, characterized in that The performing semantic analysis on the target question to determine the domain corresponding to the target question and the semantic feature information of the target question includes: Utilize natural language processing to analyze the target question and extract keywords; The domain corresponding to the target question is identified through keyword matching and semantic understanding, and the semantic feature information corresponding to the target question is determined in combination with context information.
4. The method according to claim 1, wherein The configuring of differentiated setting permissions for users based on the deployment mode corresponding to the data processing service includes: If the deployment mode corresponding to the data processing service is the standard input and output mode, users with administrator privileges are allowed to add, delete and modify the data processing service; If the deployment mode corresponding to the data processing service is a server push event mode based on the Hypertext Transfer Protocol or a mode supporting two-way streaming transmission based on the Hypertext Transfer Protocol, users without administrator privileges are allowed to add, delete and modify the data processing service.
5. The method according to claim 1, wherein After determining the data processing service required to solve the target problem based on the domain corresponding to the target problem and the semantic feature information, the method further includes: Determining the type of service requested by the user based on the data processing service; Based on the access rights of the user, it is determined whether the user has the right to access the data processing service of the service type.
6. The method according to claim 1, characterized in that The method further comprises: Evaluating resource requirements of the multiple data processing services based on expected loads corresponding to the multiple data processing services; Automatic elastic expansion and contraction processing is performed on the dynamic process pool according to the resource requirements and actual resource usage of the multiple data processing services. The dynamic process pool is used to manage resource quotas for the multiple data processing services.
7. The method according to claim 6, characterized in that The automatic elastic expansion and contraction of the dynamic process pool according to the resource requirements of the multiple data processing services and the actual usage of the multiple data processing services includes: If the actual resource usage is lower than the resource demand, the size of the dynamic process pool is automatically reduced; If the actual resource usage is higher than the resource demand, the size of the dynamic process pool is automatically expanded.
8. The method according to claim 6, characterized in that Before invoking corresponding data processing tools to analyze the target question according to the dynamic priorities of the multiple data processing tools and generating and displaying an answer result corresponding to the target question, the method further includes: Initializing the dynamic process pool and pre-creating a preset number of idle processes for multiple services in the dynamic process pool; Based on the service type requested by the user, checking whether there is an idle process corresponding to the service available in the dynamic process pool; If there is an idle process corresponding to the service available in the dynamic process pool, a process is allocated to the service, and a corresponding data processing tool in the process is called to analyze the target problem; If no idle process corresponding to the service is available in the dynamic process pool, it is detected whether the number of processes of the service has reached the maximum limit, and the dynamic process pool is processed according to the detection result.
9. The method according to claim 8, characterized in that The processing of the dynamic process pool according to the detection result includes: If the number of processes of the service has not reached the maximum limit, a new service process is created, and the new service process is added to the process pool to become an idle process waiting to be called; If the number of processes of the service has reached the maximum limit, the user task corresponding to the service enters a queue state and waits for an idle process to appear.
10. The method according to claim 6, characterized in that The method further comprises: The dynamic process pool is regularly checked, and abnormal processes in the dynamic process pool are removed according to the check result.
11. The method according to claim 2, characterized in that The method further comprises: During the invocation of the data processing tool, real-time monitoring of the tool performance is performed and the relevance index and historical usage data are updated; The tool evaluation model is optimized based on the relevance index and the historical usage data.
12. The method according to claim 11, characterized in that The optimizing the tool evaluation model includes: Based on the current performance indicators of the data processing tool and the cumulative performance improvement of the data processing tool over a preset time period, a weight adjustment action corresponding to the reward function in the tool evaluation model is output through a policy network, and the policy network is updated according to the collected trajectory data using a policy gradient method; Based on the weight adjustment action, the state action trajectory of the tool evaluation model is continuously interacted and recorded during the training process, and the parameters and weight configuration of the reward function are iteratively optimized.
13. A data analysis system, characterized in that: include: User interaction layer, question and answer processing layer, tool calling layer, and service layer; The user interaction layer is used to obtain the target question raised by the user and display the answer result corresponding to the target question to the user; The question-answering processing layer is used to perform semantic analysis on the target question to determine the domain corresponding to the target question and the semantic feature information of the target question; Determining multiple data processing tools related to the target problem based on the domain corresponding to the target problem and the semantic feature information; dynamically prioritizing the multiple data processing tools based on relevance indicators of the multiple data processing tools to the target problem and historical usage; The tool calling layer is used to call corresponding data processing tools according to the dynamic priorities of the multiple data processing tools to analyze the target question and generate an answer result corresponding to the target question; The service layer is used to provide the system with multiple data processing services and the multiple data processing tools. The service layer includes a service management module. The service management module is used to configure differentiated setting permissions for users based on the deployment method corresponding to the data processing service, wherein the deployment method includes any one of a standard input and output method, a server push event method based on the Hypertext Transfer Protocol, and a method based on the Hypertext Transfer Protocol that supports two-way streaming transmission.
14. The system according to claim 13, wherein: The service management module is also used to evaluate the resource requirements of the multiple data processing services based on the expected loads corresponding to the multiple data processing services; and automatically elastically scale the dynamic process pool based on the resource requirements and actual resource utilization of the multiple data processing services. The dynamic process pool is used to manage resource quotas for the multiple data processing services.
15. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
16. An electronic device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 12 is implemented.
17. A computer program product having a computer program stored thereon, characterized in that: When the computer program product is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Multi-source fusion question and answer method and device, storage medium and electronic device
CN119128098A