A method and apparatus for querying intrinsically secure databases based on large models and the MCP protocol.
By combining a large model with the MCP protocol, an intrinsically secure database query method has been developed, which solves the problem of cross-database queries for non-technical personnel. This enables secure and reliable queries of multi-source heterogeneous data, improving the intelligence and security of the query process.
Patent Information
- Application Number
- CN202511119948.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-08-12
AI Technical Summary
Existing database query methods are insufficient to meet the query needs of non-technical personnel, lack support for cross-database collaborative queries, and pose risks of information leakage and data security.
An intrinsically secure database query method based on a large model and the MCP protocol is adopted. Through the collaborative work of the MCP client and server, synonym rewriting and intent recognition are performed to construct a heterogeneous intelligent agent set for database querying. Finally, the final result is generated through an adjudication strategy, realizing secure and reliable querying of multi-source heterogeneous data.
It enables fully controllable and secure data querying of multi-source heterogeneous data without requiring users to have coding skills, significantly improving the unified query and collaborative processing capabilities across database environments and reducing security risks during the query process.
Smart Images

Figure CN120631919B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of data query and artificial intelligence technology, and in particular to an intrinsically secure database query method and apparatus based on large models and the MCP protocol. Background Technology
[0002] With the continuous development of "Internet Plus," data is increasingly becoming a key production factor. Database systems, as the core carrier for the circulation and application of data, are used to store, manage, and query massive amounts of structured and semi-structured data, and have been widely applied in various industries and scenarios. As the scale of data continues to expand and business needs become increasingly complex, users are placing higher demands on the intelligence level of database queries.
[0003] Currently, existing database query methods mainly fall into two categories: The first is the traditional database query method, which typically relies on Structured Query Language (SQL) and predefined data schemas. The query process often requires pre-creating indexes, designing query statements, optimizing execution plans, and relying on the query engine of a Database Management System (DBMS) for parsing and execution. While this method is relatively mature in terms of efficiency and execution control, it demands a high level of user expertise, making it difficult to meet the query needs of non-technical users. The second is the Text2SQL method based on models such as Seq2Seq. This type of method transforms users' natural language queries into corresponding SQL statements, greatly lowering the barrier to entry. However, most existing Text2SQL models are only built for single databases, making it difficult to adapt to heterogeneous structures and multi-source data environments. They lack support for cross-database collaborative queries and may pose risks such as semantic misjudgment, field out-of-bounds errors, and unauthorized access when generating SQL query statements, potentially leading to information leakage and data security issues. Summary of the Invention
[0004] To address the limitations of existing database query methods in meeting the query needs of non-technical users or lacking support for cross-database collaborative queries, as well as their inability to adapt to heterogeneous structures and multi-source data environments, which can lead to potential information leakage and data security issues, this invention proposes an intrinsically secure database query method and apparatus based on a large model and the MCP protocol. This method enables fully controllable and secure data queries on multi-source heterogeneous data without requiring users to have coding skills.
[0005] In a first aspect, the present invention provides an intrinsically secure database query method based on a large model and the MCP protocol, comprising:
[0006] Step 1: Obtain the user's input question information; wherein, the question information includes text information and voice information;
[0007] Step 2: The MCP client receives the question information, performs synonym rewriting and intent recognition on the question information to obtain a request information set, and sends the request information set together with a preset MCP connection credential set to the MCP server; wherein, the request information set contains N request information messages that are synonymous with the question information;
[0008] Step 3: The MCP server receives the request information set and parses the MCP connection credential set. Based on the request information set, it constructs a heterogeneous intelligent agent set. Each heterogeneous intelligent agent in the heterogeneous intelligent agent set performs a database query operation to obtain a return information set.
[0009] Step 4: Based on the adjudication strategy, adjudicate the returned information set to obtain the final database query result.
[0010] Furthermore, step 1 also includes preprocessing the query information:
[0011] The text information is subjected to language detection, encoding standard conversion, noise character removal, and punctuation cleaning.
[0012] Endpoint detection, noise suppression, and speech enhancement are performed on the speech information.
[0013] Furthermore, in step 2, the paraphrasing and intent recognition of the question information specifically includes: the MCP client processes the question information by calling a large model based on a preset prompt word template.
[0014] Furthermore, in step 2, the preset MCP connection credential set includes a data source location tuple, model service call parameters, policy descriptor, and response output configuration.
[0015] Furthermore, in step 3, each heterogeneous intelligent agent in the heterogeneous intelligent agent set performs a database query operation, specifically including:
[0016] After receiving the corresponding request information, each heterogeneous agent invokes the large model to perform semantic reasoning and task understanding on the request information to obtain structured instructions. Each heterogeneous agent, in conjunction with the MCP connection credential set, identifies databases semantically matching the request information from multi-source heterogeneous databases and determines the corresponding execution strategy. Based on the structured instructions, it accesses the corresponding databases and generates query results. After receiving the query results, the large model processes them according to the response output template of the heterogeneous agents and outputs return information.
[0017] Furthermore, in step 3, each heterogeneous agent in the heterogeneous agent set performs a database query operation, and the process also includes a retrieval enhancement technique that incorporates a graph structure:
[0018] Each heterogeneous intelligent agent obtains the corresponding database metadata in the form of structured text; the database metadata includes the source type of each data source, the table name and table description of the data table in the data source, and the table creation statement of the data table.
[0019] Generate graph-structured text information based on the database metadata;
[0020] The graph structure text information is vectorized and encoded, and then stored in a vector database in the form of a graph structure.
[0021] Each heterogeneous agent in the heterogeneous agent set constructs a graph retrieval agent, which retrieves and recalls the most similar text fragments from the vector database according to the corresponding request information, thereby obtaining a recall information set.
[0022] Each heterogeneous agent in the heterogeneous agent set accesses the target database by executing the tool invocation strategy based on the corresponding request information and recall information, combined with the MCP connection credential set, and generates a query result set.
[0023] Furthermore, generating graph-structured text information based on the database metadata specifically includes:
[0024] A key-value pair structure is formed based on the database metadata; wherein, the key-value pair structure includes two types of key-value pair structures; the key in the first type of key-value pair structure is the source type of the data source, and the value in the first type of key-value pair structure is the table name and table description of the data table in the data source corresponding to the key in the first type of key-value pair structure; the key in the second type of key-value pair structure is the table name, table description, and business domain of the data table in the data source, and the value in the second type of key-value pair structure is the table creation statement of the data table corresponding to the key in the second type of key-value pair structure;
[0025] The table descriptions in different table creation statements are deduplicated and merged to form a graph structure based on the key-value pairs. The nodes of the graph structure are the source type of the data source, the table name and description of the data table within the data source, and the table creation statement of the data table. The node attribute of the source type of the data source is set as a first-level node; the node attributes of the table name, table description, and business domain of the data table within the data source are set as second-level nodes; and the node attributes of the table creation statement of the data table are set as third-level nodes. The edge connections of the graph structure are the mapping relationships of the key-value pair structure.
[0026] Furthermore, the graph retrieval tool performs retrieval based on a two-layer retrieval method using parent node information and leaf node information.
[0027] Furthermore, in step 4, the adjudication strategy includes format verification and consensus voting, specifically including:
[0028] All returned information in the returned information set is format-validated. If any returned information does not conform to the encoding format and output format, the returned information fails the validation.
[0029] The returned information that passes the format validation is used to perform a consensus vote based on the large number decision strategy. The returned information is split into key-value pairs according to the output format. The value corresponding to the generated result of different heterogeneous agents is determined based on the key. The data with the most data consistency is selected as the final database query result.
[0030] Secondly, the present invention provides an intrinsically secure database query device based on a large model and the MCP protocol, comprising:
[0031] The question information acquisition module is used to acquire question information input by the user; wherein, the question information includes text information and voice information;
[0032] The request processing and sending module is used to receive the question information in the MCP client, perform synonym rewriting and intent recognition on the question information to obtain a set of request information, and send the set of request information, together with a preset set of MCP connection credentials, to the MCP server; wherein, the set of request information contains multiple request information that are synonymous with the question information;
[0033] The parsing and execution module is used to receive the request information set and parse the MCP connection credential set on the MCP server, construct a heterogeneous intelligent agent set based on the request information set, and each heterogeneous intelligent agent in the heterogeneous intelligent agent set performs a database query operation to obtain a return information set; wherein, each heterogeneous intelligent agent in the heterogeneous intelligent agent set has equivalent functions but different structures.
[0034] The strategy adjudication module is used to adjudicate the returned information set based on the adjudication strategy to obtain the final database query result.
[0035] The beneficial effects of this invention are as follows:
[0036] This invention organically combines large-scale models and the MCP protocol. The MCP protocol acts like a USB-C interface, providing a standardized method for large-scale models. By configuring the MCP connection credential set, the large-scale model gains the ability to understand database configuration information in the target cluster and call the target database. This enables efficient adaptation and seamless interaction with different types of data sources and tool systems. This technical solution effectively solves the query problem under multi-source heterogeneous data structures, significantly improves the unified query and collaborative processing capabilities across database environments, and can be applied to data querying in data analysis systems and intelligent customer service systems to achieve intelligent upgrades of business scenarios.
[0037] This invention, based on the theory of intrinsic security, constructs functionally equivalent heterogeneous intelligent agents through a combination of "redundant control and large model services." It then applies policy decisions to the returned information from these heterogeneous agents to obtain the final database query results. Compared to traditional database query models (such as Text2SQL), this invention effectively reduces the risks of field mismatches, out-of-bounds data access, and unauthorized operations caused by semantic understanding biases during the query process, thus enhancing security from the source. Attached Figure Description
[0038] Figure 1 A flowchart illustrating an intrinsically secure database query method based on a large model and the MCP protocol, provided in an embodiment of the present invention;
[0039] Figure 2 This is a flowchart illustrating an intrinsically secure database query device based on a large model and the MCP protocol, provided as an embodiment of the present invention. Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0041] Example 1
[0042] like Figure 1 As shown, an embodiment of the present invention provides an intrinsically secure database query method based on a large model and the MCP protocol, comprising:
[0043] Step 1: Obtain the user's input question information. The question information includes text information and voice information.
[0044] Specifically, the system acquires query information input by users through their terminal devices. The text input consists of queries described in natural language, such as "What percentage of users completed payments through the app in the past month?" or "What was the average daily active users (DAU) over the past seven days?". This text information can be input through a browser's graphical user interface (GUI), a command-line interface, or other interactive input components. The voice input refers to audio data acquired through a microphone or voice capture device. The acquired raw voice signal is processed by an Automatic Speech Recognition (ASR) module and transcribed into a standardized text representation. This text representation and the original text information constitute the query information, which serves as input for subsequent semantic understanding and task parsing.
[0045] Preferably, the process of obtaining the user's input question information may further include:
[0046] ① Perform text preprocessing operations on user-input text information, such as language detection, encoding standard conversion (e.g., UTF-8), noise character removal, and punctuation cleaning;
[0047] ② In the process of acquiring user-input voice information, audio processing operations such as endpoint detection, noise suppression, and voice enhancement are performed on the raw voice signal acquired by the microphone device or voice acquisition device.
[0048] Step 2: The MCP client receives the query information, performs paraphrasing and intent recognition on the query information to obtain the request information set, and sends the request information set, together with the preset MCP connection credential set, to the MCP server.
[0049] The request information set contains multiple request messages that are synonymous with the question message. In this embodiment of the invention, the generated request information set contains three request messages: a first request message, a second request message, and a third request message.
[0050] Specifically, in this embodiment, the MCP client is the user interaction and service request entry point deployed between the user terminal and the backend large model service. It is the interface through which the user directly interacts with the application, which can be a web application, a mobile app, or a programming API (such as a RESTful API). The role of the MCP client is to receive user input to initiate requests and to subsequently receive the response results from heterogeneous intelligent agents.
[0051] Specifically, the MCP client, based on a preset prompt word template, invokes a large model service to implement a semantic processing flow, thereby processing the query information to generate N ≥ 3 request messages. The prompt word template is used to guide the large model in generating N request messages with different expressions but consistent semantics.
[0052] As an feasible approach, the preset prompt word template is as follows:
[0053] [Function]: You are a database intelligent analysis assistant. Please complete the following tasks:
[0054] [Input Issue]:
[0055] {User Question Information}
[0056] [Task Requirements]:
[0057] 1. Please generate a query question with a different phrasing but the same semantic meaning as the original question, and name it: Request Information;
[0058] 2. Extract the query intent for this question (e.g., statistical query / trend analysis / indicator comparison, etc.);
[0059] 3. The output format is JSON.
[0060] Output format:
[0061] {
[0062] "Request Information": " ",
[0063] "Query Intent": " "
[0064] }
[0065] The large models are mainstream pre-trained models with powerful language understanding and generation capabilities, such as DeepSeek, Qwen, or Spark. The deployment of these large models can be based on local computing resources deployed in a private cluster, using a local large model service interface built with the Ollam framework, or it can be an API service interface provided by an external AI capability platform, including but not limited to Alibaba Cloud AI Platform, DeepSeek Open Platform, and Silicon-based Streaming Large Model Development Platform.
[0066] Specifically, for the same user query, the MCP client calls the large language model service three times based on the preset prompt word module, thereby obtaining the first request information, the second request information, and the third request information.
[0067] For example, when a user enters a text message: "How many users have been active on the app in the last 7 days?", the MCP client can obtain the first request information, the second request information, and the third request information by calling the large model service three times for paraphrasing and intent recognition, as follows:
[0068] 1. The first request information is:
[0069] {
[0070] Request Information: "How many active users have there been in the app in the past 7 days?"
[0071] Query Intent: Statistical Query
[0072] }
[0073] 2. The second request information is:
[0074] {
[0075] Request information: "What is the total number of active users on the app in the past week?"
[0076] Query Intent: Statistical Query
[0077] }
[0078] 3. The third request information is:
[0079] {
[0080] Request information: "How many users have been active in the app in the past 7 days?"
[0081] Query Intent: Statistical Query
[0082] }
[0083] Specifically, the MCP connection credential set refers to a set of configuration parameters used to support secure, stable, and standardized connection access between large models and external database resources under the MCP protocol. This credential set not only covers the basic connection parameters of the database ontology but also includes control information for model invocation, authentication, invocation strategies, and result format management, comprising the following components:
[0084] 1. Data source location tuple: target database type, target database instance terminal address, port number, database name, authentication credentials (such as username, password or token);
[0085] 2. Model service call parameters: service URL or API interface of the large model (e.g., / v1 / chat / completions), model API call license (e.g., API-KEY), model name and version (e.g., qwen-plus), request parameters (e.g., temperature, Top-p, etc.);
[0086] 3. Strategy descriptor: query timeout threshold, maximum connection pool size, concurrency limit, maximum number of rows in the query result set;
[0087] 4. Response output configuration: output encoding format (UTF-8), output format (such as streaming output), sensitive data desensitization rules.
[0088] Furthermore, the MCP connection credential set can be installed, obtained, and loaded from the MCP protocol open platform (such as https: / / mcp.so / ), or it can be a custom MCP protocol following standard procedures. For example, the following is an example of connection credentials for locating tuples in a MySQL data source:
[0089] {
[0090] "mcpServers": {
[0091] "mysql": {
[0092] "command": "node",
[0093] "args": [" / path / to / mysql-mcp-server / build / index.js"],
[0094] "env": {
[0095] "MYSQL_HOST": "your-mysql-host",
[0096] "MYSQL_PORT": "3306",
[0097] "MYSQL_USER": "your-mysql-user",
[0098] "MYSQL_PASSWORD": "your-mysql-password",
[0099] "MYSQL_DATABASE": "your-default-database"
[0100] },
[0101] "disabled": false,
[0102] "autoApprove": []
[0103] }
[0104] }
[0105] }
[0106] Preferably, given that the large model itself currently lacks the ability to perceive and understand the current time context, this embodiment obtains an MCP protocol capable of querying and perceiving the current time from the MCP protocol open platform and configures it into the MCP connection credential set. In this way, the large model can access and parse the current time, gaining the ability to process time semantics. During subsequent database calls, when faced with query requests containing a time dimension, the large model can perform more accurate semantic reasoning on time information, thereby significantly improving the accuracy of database queries. The implementation of the MCP protocol can include, but is not limited to, standardized MCP protocols that support time service capabilities, such as the Baidu Maps MCP protocol and the Amap Maps MCP protocol.
[0107] Step 3: The MCP server receives the request information set and parses the MCP connection credential set. Based on the request information set, it constructs a heterogeneous intelligent agent set. Each heterogeneous intelligent agent in the heterogeneous intelligent agent set performs a database query operation to obtain the returned information set.
[0108] Specifically, the MCP server receives a request from the MCP client and parses the MCP connection credential set, then begins processing the request. Based on the concepts of mimicry defense and dynamic heterogeneous redundancy in intrinsic security, it constructs a first heterogeneous intelligent agent, a second heterogeneous intelligent agent, and a third heterogeneous intelligent agent according to the first request information, the second request information, and the third request information, respectively. The first, second, and third heterogeneous intelligent agents respectively perform context information understanding and invoke available tools, and perform database queries on relevant devices to obtain the first, second, and third return information.
[0109] Specifically, the MCP server is an intelligent middleware component that follows the MCP protocol. It is responsible for coordinating semantic instruction conversion, resource scheduling, connection control, and result set processing between large models and external databases. It is a large model carrier and inference engine used for model loading, storage, and computation scheduling.
[0110] Because semantic biases, field matching errors, or unauthorized field calls may occur when the model performs semantic understanding and generates database query statements, such queries pose significant uncertainties and information security risks. Therefore, based on the idea of dynamic heterogeneous redundancy in intrinsic security, functionally equivalent but structurally different heterogeneous intelligent agents are constructed. At the same time, to ensure the dynamic diversity of the executors, the MCP server schedules and allocates computing resources to construct a set of heterogeneous intelligent agents, namely the first heterogeneous intelligent agent, the second heterogeneous intelligent agent, and the third heterogeneous intelligent agent.
[0111] Specifically, heterogeneous intelligent agents refer to multi-instance intelligent processing units built based on different model configurations, context awareness capabilities, and tool invocation methods after multi-path semantic parsing and task decomposition of user natural language request information through a large model. Each heterogeneous intelligent agent has independent inference logic, execution chain, and verification mechanism, and is a data security query intelligent execution body architecture composed of "redundancy control + large model service".
[0112] Specifically, after receiving the request information, the first, second, and third heterogeneous intelligent agents respectively invoke the large model to perform semantic reasoning and task understanding on the request information. Combining the MCP connection credential set, they identify target database instances that semantically match the request from multi-source heterogeneous databases, determine the corresponding execution strategy, access the target database according to the parsed structured execution instructions, and return query results that conform to the permission and format specifications. After receiving the query results, the large model processes the results according to the response output template of the heterogeneous intelligent agents and outputs a response answer that conforms to the natural language expression habits and format requirements.
[0113] Preferably, to further improve the semantic understanding accuracy of heterogeneous agents regarding contextual information and the accuracy of target database location, this embodiment introduces graph-structured retrieval enhancement technology. During the semantic reasoning and execution strategy generation process of the heterogeneous agent set (first heterogeneous agent, second heterogeneous agent, and third heterogeneous agent), dynamic semantic completion support from a vector database is provided. The specific method is as follows:
[0114] ① Different heterogeneous intelligent agents obtain database metadata in the form of structured text. Database metadata is information describing various data structures in the database, used to help understand, locate, manage, and operate the database, including but not limited to:
[0115] The source type of each data source (e.g., MySQL, Hive, etc.);
[0116] The table name and description of each data table in each data source;
[0117] The table creation statement for each table in the database contains information such as the field name, field type, field meaning, enumeration value description, and constraints between fields (such as primary key and foreign key).
[0118] ② Generate graph structure text information based on the database metadata, specifically including:
[0119] Based on database metadata, key-value pair structures with indexing value are formed. In the first type of key-value pair structure, the "key" is the source type of the data source, and the "value" is the table name and table description of the data table in the data source corresponding to the source type of the data source. In the second type of key-value pair structure, the "key" is the table name and table description of the data table in the data source, and the "value" is the table creation statement corresponding to the table name and table description of the data table in the data source.
[0120] Then, the table descriptions in different table creation statements are deduplicated and merged to reduce unnecessary calculations, and a graph structure is formed based on the key-value pairs.
[0121] In this graph structure, the node information includes the source type of the data source, the table name and description of the data table within the data source, and the table creation statement. The node attribute for the source type of the data source is set as a first-level node; the node attributes for the table name, table description, and business domain of the data table within the data source are set as second-level nodes; and the node attributes for the table creation statement are set as third-level nodes. The edge connections in the graph structure are key-value pair mappings.
[0122] ③ After vectorizing the graph structure text information, store it in the vector database in the form of a graph structure.
[0123] Specifically, the embedding model used in the vectorization encoding can be Qwen3-Embedding, bge-m3, or m3e-base.
[0124] ④ Each heterogeneous agent in the heterogeneous agent set constructs a graph retrieval agent, and retrieves and recalls the most similar text fragments from the vector database according to the corresponding request information, thereby obtaining a set of recall information, namely the first recall information, the second recall information, and the third recall information.
[0125] Specifically, the search engine performs a two-level search based on parent node information and leaf node information. Root node information consists of first-level and second-level nodes in the graph structure, while leaf node information consists of third-level nodes. To enhance the accuracy of search results, when searching for leaf node information, the parent node information corresponding to the retrieved leaf node is also returned as a search result.
[0126] Retrieval based on parent node information provides contextual information about the database structure for the large model, while retrieval based on leaf node information provides contextual information about statistical analysis and aggregation fields when parsing data queries for the large model.
[0127] Specifically, the search results from the parent node information retrieval and the leaf node information retrieval are merged to obtain the recall information.
[0128] The search strategy (search_type) used by the search engine can be various, including but not limited to:
[0129] 1. Similarity: An algorithm that directly calculates the cosine distance between the query vector and the document vector and returns the Top K results;
[0130] 2. MMR: Maximum Marginal Relevance Algorithm that avoids returning content with duplicate semantics;
[0131] 3. similarity_score_threshold: A similarity threshold filtering algorithm that only returns similarity scores exceeding a set threshold.
[0132] ⑤ Each heterogeneous agent in the heterogeneous agent set accesses the target database based on the corresponding request information and recall information, combined with the vector MCP connection credential set, and executes the tool invocation strategy to generate a query structure set.
[0133] Specifically, the first heterogeneous intelligent agent, the second heterogeneous intelligent agent, and the third heterogeneous intelligent agent, based on their respective request information and recall information, combined with the MCP connection credential set, execute the tool invocation strategy to access the target database, generate a set of query results, and thus obtain the first query result, the second query result, and the third query result.
[0134] Furthermore, the heterogeneous intelligent agent inputs the corresponding request information and recall information into the large model to construct a natural language prompt template that includes both user query requirements and database context information. This prompt template has database semantic awareness and structured query processing capabilities, thus improving accuracy.
[0135] Furthermore, the intrinsic security property in the intrinsically secure database query method is embodied by functionally equivalent heterogeneous intelligent agents. Its dynamic characteristics are manifested in the existence of three equivalent but heterogeneous intelligent agents, and its heterogeneous redundancy characteristics include:
[0136] 1. Differences in request information: First request information, second request information, and third request information after synonym rewriting and intent recognition;
[0137] 2. Different large models are invoked: Alibaba's Qwen large model, DeepSeek large model, and iFlytek's Xinghuo large model are invoked through the local Ollam framework or external AI capability platforms;
[0138] 3. Different embedding models are used when generating the vector database: Qwen3-Embedding, bge-m3, and m3e-base;
[0139] 4. Different retrieval and recall strategies: direct calculation of cosine distance, maximum marginal relevance algorithm, and similarity threshold filtering algorithm.
[0140] Furthermore, the heterogeneous intelligent agent is implemented based on the Spring AI framework. Spring AI is an artificial intelligence application development framework built on the Spring ecosystem. This framework includes standardized access components for large models, prompt engineering management tools, document processing and vectorization tools, vector databases and retrieval tools, and Model Context Protocol (MCP) support. It seamlessly integrates with the Spring ecosystem, such as Spring Boot, which facilitates the application of the intrinsically secure database query method in this embodiment to data analysis systems and intelligent customer service systems, thereby achieving intelligent upgrades to business scenarios.
[0141] Step 4: Based on the adjudication strategy, adjudicate the returned information set to obtain the final database query results.
[0142] Specifically, the adjudication strategy in this embodiment includes format verification and consistency voting. For the first, second, and third return information returned by the first, second, and third heterogeneous intelligent agents, format verification is first performed. If there is an encoding format (such as UTF-8) or output format (such as JSON) that does not meet the requirements of this embodiment, the return information of that heterogeneous intelligent agent fails the judgment. Then, for the return information that passes the format verification, consistency voting is performed based on the large number adjudication strategy. That is, the return information is split into key-value pairs according to the output format, and the value corresponding to the result generated by different heterogeneous intelligent agents is determined based on the key value. The data with the most data consistency is selected as the final database query result.
[0143] Example 2
[0144] like Figure 2 As shown, this embodiment of the invention also provides an intrinsically secure database query device based on a large model and the MCP protocol, comprising:
[0145] The question information acquisition module is used to acquire the question information input by the user; the question information includes text information and voice information.
[0146] The request processing and sending module is used to receive the query information in the MCP client, perform synonym rewriting and intent recognition on the query information to obtain a set of request information, and send the set of request information, together with a preset set of MCP connection credentials, to the MCP server; wherein, the set of request information contains N request information that are synonymous with the query information;
[0147] The parsing and execution module is used to receive the request information set and parse the MCP connection credential set on the MCP server, construct a heterogeneous intelligent agent set based on the request information set, and each heterogeneous intelligent agent in the heterogeneous intelligent agent set performs a database query operation to obtain the return information set.
[0148] The strategy adjudication module is used to adjudicate the returned information set based on the adjudication strategy to obtain the final database query results.
[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An intrinsically secure database query method based on a large model and the MCP protocol, characterized in that, include: Step 1: Obtain the user's input question information; wherein, the question information includes text information and voice information; Step 2: The MCP client receives the question information, performs synonym rewriting and intent recognition on the question information to obtain a request information set, and sends the request information set together with a preset MCP connection credential set to the MCP server; wherein, the request information set contains N request information messages that are synonymous with the question information; Step 3: The MCP server receives the request information set and parses the MCP connection credential set. Based on the request information set, it constructs a heterogeneous intelligent agent set. Each heterogeneous intelligent agent in the set performs a database query operation to obtain a return information set. Each heterogeneous intelligent agent obtains the corresponding database metadata in structured text format. The database metadata includes the source type of each data source, the table name and description of the data table in the data source, and the table creation statement. Generate graph-structured text information based on the database metadata; The graph structure text information is vectorized and encoded, and then stored in a vector database in the form of a graph structure. Each heterogeneous agent in the heterogeneous agent set constructs a graph retrieval agent, which retrieves and recalls the most similar text fragments from the vector database according to the corresponding request information, thereby obtaining a recall information set. Each heterogeneous agent in the heterogeneous agent set accesses the target database by executing the tool invocation strategy based on the corresponding request information and recall information, combined with the MCP connection credential set, and generates a set of query results. The generation of graph structure text information based on the database metadata specifically includes: A key-value pair structure is formed based on the database metadata; wherein, the key-value pair structure includes two types of key-value pair structures; the key in the first type of key-value pair structure is the source type of the data source, and the value in the first type of key-value pair structure is the table name and table description of the data table in the data source corresponding to the key in the first type of key-value pair structure; the key in the second type of key-value pair structure is the table name and table description of the data table in the data source, and the value in the second type of key-value pair structure is the table creation statement of the data table corresponding to the key in the second type of key-value pair structure; The table descriptions in different table creation statements are deduplicated and merged to form a graph structure based on the key-value pairs. The nodes of the graph structure are the source type of the data source, the table name and description of the data table within the data source, and the table creation statements of the data table. The node attribute of the source type of the data source is set as a first-level node; the node attributes of the table name and description of the data table within the data source are set as second-level nodes; and the node attributes of the table creation statements of the data table are set as third-level nodes. The edge connections of the graph structure are the mapping relationships of the key-value pair structure. Step 4: Based on the adjudication strategy, adjudicate the returned information set to obtain the final database query result.
2. The intrinsically secure database query method based on a large model and the MCP protocol according to claim 1, characterized in that, Step 1 further includes preprocessing the query information: The text information is subjected to language detection, encoding standard conversion, noise character removal, and punctuation cleaning. Endpoint detection, noise suppression, and speech enhancement are performed on the speech information.
3. The intrinsically secure database query method based on a large model and the MCP protocol according to claim 1, characterized in that, Step 2, specifically the paraphrasing and intent recognition of the question information, includes: the MCP client processing the question information by calling a large model based on a preset prompt word template.
4. The intrinsically secure database query method based on a large model and the MCP protocol according to claim 1, characterized in that, In step 2, the preset MCP connection credential set includes the data source location tuple, model service call parameters, policy descriptor, and response output configuration.
5. The intrinsically secure database query method based on a large model and the MCP protocol according to claim 1, characterized in that, In step 3, each heterogeneous agent in the heterogeneous agent set performs a database query operation, specifically including: After receiving the corresponding request information, each heterogeneous agent invokes the large model to perform semantic reasoning and task understanding on the request information to obtain structured instructions. Each heterogeneous agent, in conjunction with the MCP connection credential set, identifies databases semantically matching the request information from multi-source heterogeneous databases and determines the corresponding execution strategy. Based on the structured instructions, it accesses the corresponding databases and generates query results. After receiving the query results, the large model processes them according to the response output template of the heterogeneous agents and outputs return information.
6. The intrinsically secure database query method based on a large model and the MCP protocol according to claim 1, characterized in that, The graph search engine performs searches using a two-tiered retrieval method based on parent node information and leaf node information.
7. The intrinsically secure database query method based on a large model and the MCP protocol according to claim 1, characterized in that, In step 4, the decision-making strategy includes format verification and consensus voting, specifically including: All returned information in the returned information set is format-validated. If any returned information does not conform to the encoding format and output format, the returned information fails the validation. The returned information that passes the format validation is used to perform a consensus vote based on the large number decision strategy. The returned information is split into key-value pairs according to the output format. The value corresponding to the result generated by different heterogeneous agents is determined based on the key. The data with the most data consistency is selected as the final query result.
8. An intrinsically secure database query device based on a large model and the MCP protocol, characterized in that, include: The question information acquisition module is used to acquire question information input by the user; wherein, the question information includes text information and voice information; The request processing and sending module is used to receive the question information in the MCP client, perform synonym rewriting and intent recognition on the question information to obtain a set of request information, and send the set of request information, together with a preset set of MCP connection credentials, to the MCP server; wherein, the set of request information contains N request information that are synonymous with the question information; The parsing and execution module is used to receive the request information set and parse the MCP connection credential set on the MCP server, construct a heterogeneous intelligent agent set based on the request information set, and each heterogeneous intelligent agent in the set performs a database query operation to obtain a return information set; wherein, the database query operation performed by each heterogeneous intelligent agent in the set also includes a retrieval enhancement technique that introduces a graph structure: Each heterogeneous intelligent agent obtains the corresponding database metadata in the form of structured text; the database metadata includes the source type of each data source, the table name and table description of the data table in the data source, and the table creation statement of the data table. Generate graph-structured text information based on the database metadata; The graph structure text information is vectorized and encoded, and then stored in a vector database in the form of a graph structure. Each heterogeneous agent in the heterogeneous agent set constructs a graph retrieval agent, which retrieves and recalls the most similar text fragments from the vector database according to the corresponding request information, thereby obtaining a recall information set. Each heterogeneous agent in the heterogeneous agent set accesses the target database by executing the tool invocation strategy based on the corresponding request information and recall information, combined with the MCP connection credential set, and generates a set of query results. The generation of graph structure text information based on the database metadata specifically includes: A key-value pair structure is formed based on the database metadata; wherein, the key-value pair structure includes two types of key-value pair structures; the key in the first type of key-value pair structure is the source type of the data source, and the value in the first type of key-value pair structure is the table name and table description of the data table in the data source corresponding to the key in the first type of key-value pair structure; the key in the second type of key-value pair structure is the table name and table description of the data table in the data source, and the value in the second type of key-value pair structure is the table creation statement of the data table corresponding to the key in the second type of key-value pair structure; The table descriptions in different table creation statements are deduplicated and merged to form a graph structure based on the key-value pairs. The nodes of the graph structure are the source type of the data source, the table name and description of the data table within the data source, and the table creation statements of the data table. The node attribute of the source type of the data source is set as a first-level node; the node attributes of the table name and description of the data table within the data source are set as second-level nodes; and the node attributes of the table creation statements of the data table are set as third-level nodes. The edge connections of the graph structure are the mapping relationships of the key-value pair structure. The strategy adjudication module is used to adjudicate the returned information set based on the adjudication strategy to obtain the final database query result.
Citation Information
Patent Citations
System intelligent interaction method and device based on large language model
CN117370493A
Data processing method and device and large language model fine tuning method and device
CN119576964A